---
title: "Sitemap Errors That Keep Pages Out of the Index"
description: "Find out which sitemap errors stop Google from indexing your pages, how to read each Search Console warning, and what to fix first."
canonical: https://adsbeast.pro/blog/en/sitemap-errors-that-keep-pages-out-of-the-index
language: en
published: 2026-09-19T02:34:19.548Z
updated: 2026-09-19T02:34:19.548Z
author: "ADS Beast editorial team"
translations:
  - ar: https://adsbeast.pro/blog/ar/sitemap-errors-that-keep-pages-out-of-the-index
  - es: https://adsbeast.pro/blog/es/sitemap-errores-que-impiden-indexar-tus-paginas
  - he: https://adsbeast.pro/blog/he/sitemap-errors-that-keep-pages-out-of-the-index
  - ru: https://adsbeast.pro/blog/ru/karta-sayta-oshibki-iz-za-kotorykh-stranitsy-ne-v-indekse
  - sk: https://adsbeast.pro/blog/sk/sitemap-xml-chyby-pre-ktore-sa-stranky-nedostanu-do-indexu
  - uk: https://adsbeast.pro/blog/uk/karta-saytu-pomilki-cherez-yaki-storinki-ne-v-indeksi
---
# Sitemap Errors That Keep Pages Out of the Index

A sitemap is a list of URLs you want crawled, nothing more. Google reads it, then decides on its own whether each URL deserves indexing. Pages stay out of the index when the sitemap points to URLs that are blocked, redirected, marked noindex, canonicalized elsewhere, or served from a file Google cannot parse.

## In short

- A sitemap is a hint. It never forces indexing, and a submitted URL can sit unindexed indefinitely.
- The most common blockers are robots.txt disallow rules, noindex tags, canonical tags pointing to another URL, and non-200 status codes.
- File-level errors (wrong format, malformed XML, oversized files) stop Google from reading the sitemap at all, which is a different problem from individual URLs being skipped.
- Every error in the Search Console Sitemaps report maps to one specific fix. Fixing the file first, then the URLs, saves the most time.
- A sitemap stuffed with non-indexable URLs wastes crawl budget and slows discovery of the pages you actually care about.

## What a sitemap does and does not control

An XML sitemap is a file that lists URLs on your site along with optional metadata such as last modification date. It tells search engines where your pages are. It does not tell them what to do with those pages.

Google treats a sitemap as a discovery signal. The crawler picks URLs from it, fetches them, and then applies its own rules: status code, robots directives, canonical tags, content quality, duplication. Any of those can keep a page out of the index even though the URL appears in a perfectly valid sitemap. This is why "the URL is in my sitemap" is never an answer to "why is this page not indexed."

The practical consequence: a sitemap error and an indexing problem are not the same thing. You can have a clean sitemap with zero warnings and still see pages excluded. You can also have a sitemap that Google refuses to read at all, which is worse, because then nothing in it gets discovered through that channel.

## Sitemap file errors that stop Google from reading it

File-level errors mean Google could not process the sitemap. Every URL inside it is invisible until the file itself is fixed.

These show up in the Sitemaps report in Search Console and in Bing Webmaster Tools. The three you will meet most often:

**"Couldn't fetch."** The server returned a 4xx or 5xx status when Google requested the file. Common causes: the sitemap was deleted or moved, a redirect chain loops, the server blocks Googlebot at the firewall level, or the file sits behind authentication. Check that the URL returns a 200 directly, without redirects, from a clean request.

**"Sitemap is HTML."** The URL returns an HTML page instead of XML. This usually happens when a CMS or a rewrite rule intercepts the request and serves a 404 page or a homepage with a 200 status. Open the sitemap URL in a browser and view source. If you see `<!DOCTYPE html>`, the file is not being served.

**"Sitemap contains URLs which are blocked by robots.txt."** The file is readable, but some URLs inside it are disallowed. Google reports the count. Each blocked URL is skipped.

Malformed XML is a fourth category. A single unescaped ampersand, a missing closing tag, or an incorrect namespace declaration makes the whole document unparseable. XML sitemaps require the namespace `http://www.sitemaps.org/schemas/sitemap/0.9` on the `urlset` element. Spaces in URLs and unencoded characters cause the same failure.

Validate before submitting. Any XML validator will catch structural problems, and Search Console will tell you the exact line when it can.

## Limits: when the file itself is the problem

A single sitemap file can hold up to 50,000 URLs and must stay under 50 MB uncompressed. Exceed either limit and Google stops reading past the break, or rejects the file.

If you are near those numbers, split. Create multiple sitemap files, then list them in a sitemap index file. The index is a separate XML document that points to each child sitemap. You submit the index once and Google follows the references.

Google recommends keeping files well below the limit. A sitemap that takes too long to fetch can time out, and a timeout looks like a fetch error even though the file is technically valid. Segmenting by content type (products, blog posts, category pages) also makes the Search Console data readable, because you can see indexed counts per segment instead of one blended number.

| Situation | What to do | Why |
|---|---|---|
| Under 50,000 URLs, one content type | Single sitemap at the domain root | Simplest to maintain and submit |
| Over 50,000 URLs or 50 MB | Split into multiple files plus a sitemap index | Google rejects or truncates oversized files |
| Multiple content types on one site | One sitemap per type, all listed in an index | Per-segment indexing data in Search Console |
| Large site with frequent updates | Generate sitemaps programmatically, refresh lastmod | Stale lastmod values get ignored over time |

## URL-level reasons a page stays out of the index

The sitemap is valid, Google reads it, and the page still does not appear. At this point the problem is the URL, not the file.

Check these in order:

1. **Status code.** The URL must return 200. A 301 or 302 means the page moved, and Google will index the destination, not the source. A 404 or 410 means the page is gone. A 5xx means the server failed during the fetch, and Google will retry later rather than index.
2. **robots.txt.** A disallow rule on the URL or its directory blocks crawling entirely. If a page cannot be crawled, it cannot be indexed, regardless of what the sitemap says.
3. **Meta robots and X-Robots-Tag.** A `noindex` directive in the HTML head or in the HTTP header removes the page from the index. The X-Robots-Tag header is easy to miss because it never appears in page source.
4. **Canonical tag.** If the page declares a different URL as canonical, Google consolidates signals to that URL and typically drops the original from the index. This is the most frequent cause of "indexed, though blocked" and "alternate page with proper canonical tag" statuses.
5. **Content quality and duplication.** Thin pages, near-duplicate pages, and pages with no unique value get crawled and then excluded. No sitemap fix changes this outcome.

Each of these produces a distinct status in the Page Indexing report. Read the status before changing anything. Removing a noindex tag from a page that also canonicalizes elsewhere will not get it indexed.

## Sitemap errors in Google Search Console: what each one means

The Sitemaps report lists a status per submitted sitemap. The wording is specific, and each phrase points at a different layer of the problem.

| Report message | Layer | Fix |
|---|---|---|
| Couldn't fetch | Server or file location | Restore the file, remove redirects, allow Googlebot through the firewall |
| Sitemap is HTML | Serving | Make the URL return XML, not a rendered page |
| Contains URLs blocked by robots.txt | URL rules | Remove the disallow rule or remove those URLs from the sitemap |
| Sitemap index file error | Structure | Check child sitemap URLs inside the index |
| Parsing error | XML syntax | Fix malformed tags, escape ampersands, verify the namespace |
| Submitted URL not found (404) | URL status | Remove dead URLs, or restore and return 200 |

Status "Success" in this report means Google read the file. It does not mean the URLs inside were indexed. Those are separate reports: Sitemaps for file health, Page Indexing for URL outcomes. Teams that only watch the Sitemaps report regularly miss the fact that half their submitted URLs are excluded.

## How to audit a sitemap before Google does

Run the audit yourself, in this order, and you will catch most problems before they reach Search Console.

1. Fetch the sitemap URL and confirm it returns 200 with an XML content type, with no redirect in between.
2. Validate the XML structure and namespace.
3. Crawl every URL in the file and record status codes. Any non-200 URL should not be in the sitemap.
4. Compare the sitemap against your robots.txt disallow rules and remove overlaps.
5. Check for noindex directives in both the HTML head and the X-Robots-Tag header.
6. Check canonical tags and remove URLs that canonicalize to a different address.
7. Strip parameter-heavy URLs (`?sort=`, `?session=`, tracking parameters) unless they are genuinely distinct pages.
8. Confirm the sitemap is referenced in robots.txt and submitted in Search Console and Bing Webmaster Tools.

The order matters. Fixing URL-level issues before the file is readable wastes effort, because Google is not reading the list yet.

If your site also runs paid campaigns, the same discipline applies elsewhere: a clean technical foundation makes landing pages easier to work with, whether that is [online advertising platforms compared for agencies](/blog/en/online-advertising-platforms-capabilities-compared-for-agencies) or an account you just set up with [Google Ads](/blog/en/how-to-create-and-set-up-a-google-ads-account-2026). Paid traffic does not fix an indexing problem, but it does amplify whatever the page already does.

## What to leave out of a sitemap

Exclude any URL that returns 404, 301, or 302. Exclude anything with a noindex directive. Exclude URLs blocked in robots.txt. Exclude duplicates and parameter-heavy variations. Exclude pages that canonicalize to a different address.

A sitemap should contain only the canonical, indexable, 200-status URLs you want in search results. Every extra URL costs crawl budget. On a large site, a sitemap bloated with redirects and parameter variants can slow discovery of new content by days or weeks, depending on crawl rate and site authority.

Two things worth keeping in: recently published pages, so they get discovered fast, and pages you have updated substantially, with an accurate `lastmod` value. Inaccurate `lastmod` timestamps get ignored, and once Google stops trusting the field, it stops using it as a recrawl signal.

## Where to place the sitemap file

Put it at the root of your domain: `https://example.com/sitemap.xml`. A sitemap can live in a subdirectory, but the root location is easier to find and simpler to reference in robots.txt.

A typical `sitemap example` entry in robots.txt is a single line pointing to the file. You can also submit the sitemap directly in Google Search Console and Bing Webmaster Tools. Do both. The robots.txt reference helps discovery across all crawlers, and the direct submission gives you reporting per sitemap.

If you use a sitemap index, submit the index URL, not the individual child files. Submitting children separately works but fragments your reporting.

## Next step

Open the Sitemaps report in Search Console and fix file-level errors first: fetch failures, wrong format, parsing problems. Once every submitted sitemap reads as Success, move to the Page Indexing report and work through the URL-level causes one status at a time. When you are ready to automate sitemap generation and keep it in sync with your content, see how [sitemap](/features/en/content-seo) management works alongside the rest of your SEO tooling.

For sites that also publish machine-readable files for AI crawlers, the same logic of accurate, current references applies to [llms.txt and what to put in it](/blog/en/llms-txt-why-your-site-needs-it-and-what-to-put-in). And if you are diagnosing why a competitor outranks you on pages you cannot get indexed, [Facebook Ads Library research](/blog/en/facebook-ads-library-find-competitors-and-analyze-ads) and [TikTok creative preparation](/blog/en/tiktok-ads-how-creative-preparation-differs) cover the paid side of the same discoverability question.

## FAQ

**Why is my sitemap not being indexed?**

A sitemap is a hint, not a guarantee. Google indexes URLs from a sitemap only when they are crawlable, return a 200 status code, and contain content worth indexing. If robots.txt blocks the URLs, if pages carry noindex tags, or if canonical tags point elsewhere, Google skips them even though they appear in the sitemap.

**How many URLs can a single sitemap file contain?**

A sitemap file can hold up to 50,000 URLs and must stay under 50 MB uncompressed. Exceed either limit and you need to split the file into multiple sitemaps listed in a sitemap index. Google recommends keeping files well below the limit to avoid timeouts during fetching.

**What causes sitemap errors in Google Search Console?**

The most common errors are "Couldn't fetch" (the server returned a 4xx or 5xx), "Sitemap contains URLs which are blocked by robots.txt," and "Sitemap is HTML" (the file returns an HTML page instead of XML). Others include malformed XML, incorrect namespace declarations, and URLs with spaces or unescaped characters. Each error class has a specific fix in the Sitemaps report.

**Where should I put my sitemap file?**

Place it at the root of your domain, for example `https://example.com/sitemap.xml`. It can live in a subdirectory, but the root location is easier to find and reference in robots.txt. You can also submit the sitemap directly in Google Search Console and Bing Webmaster Tools.

**What URLs should I exclude from a sitemap?**

Leave out any URL that returns a 404, 301, or 302, any page with a noindex meta tag, and any URL blocked in robots.txt. Also exclude duplicate or parameter-heavy URLs (like `?sort=` or `?session=`) and pages that canonicalize to a different URL. A sitemap full of non-indexable URLs wastes crawl budget and slows discovery of pages you want indexed.

**Does a valid sitemap guarantee my pages will be indexed?**

No. A valid sitemap only guarantees Google can read your list. Indexing depends on each URL's status code, robots directives, canonical tag, and content quality. You can have a sitemap with zero errors and still see most of its URLs excluded from the index.

## Questions and answers

### Why is my sitemap not being indexed?

A sitemap is a hint, not a guarantee. Google indexes URLs from a sitemap only if they are crawlable, return a 200 status code, and contain content worth indexing. If robots.txt blocks the URLs, if the pages have noindex tags, or if canonical tags point elsewhere, Google will skip them even though they appear in the sitemap.

### How many URLs can a single sitemap file contain?

A sitemap file can hold up to 50,000 URLs and must stay under 50 MB uncompressed. If you exceed either limit, split the sitemap into multiple files and list them in a sitemap index file. Google recommends keeping sitemaps well below the limit to avoid timeouts during fetching.

### What causes sitemap errors in Google Search Console?

The most common errors are 'Couldn't fetch' (server returned a 4xx or 5xx), 'Sitemap contains URLs which are blocked by robots.txt', and 'Sitemap is HTML' (the file returns an HTML page instead of XML). Others include malformed XML, incorrect namespace declarations, and URLs with spaces or unescaped characters. Each error class has a specific fix in the Sitemaps report.

### Where should I put my sitemap file?

Place it at the root of your domain, for example https://example.com/sitemap.xml. It can live in a subdirectory, but a sitemap at /sitemap.xml is easier to find and reference in robots.txt. You can also submit the sitemap directly in Google Search Console and Bing Webmaster Tools.

### What URLs should I exclude from a sitemap?

Leave out any URL that returns a 404, 301, or 302, any page with a noindex meta tag, and any URL blocked in robots.txt. Also exclude duplicate or parameter-heavy URLs (like?sort= or?session=) and pages that canonicalize to a different URL. A sitemap full of non-indexable URLs wastes crawl budget and slows down discovery of pages you actually want indexed.

