Knowledge Hub Topic Cluster

XML Sitemaps & Feeds Engineering Hub

How Googlebot discovers and ingests XML sitemaps, evaluates lastmod signals, and processes index feeds at enterprise scale.

1. XML Sitemaps Architecture & Signals

An XML sitemap is an inventory of your canonical URLs designed specifically for search engines. It acts as a primary discovery mechanism for new content, a canonical declaration signal, and a re-crawl scheduling indicator via authentic modification timestamps.

3. Googlebot Rules for XML Sitemaps

Google Ignores <priority> & <changefreq>

Google Search Central explicitly states that Googlebot does not evaluate <priority> or <changefreq>. Focus strictly on accurate canonical <loc> and authentic <lastmod> dates.

50,000 URLs / 50MB Uncompressed Limit

An individual sitemap file cannot exceed 50,000 URLs or 50MB uncompressed. Sites exceeding this limit must use a Sitemap Index file referencing child sitemaps.

4. Common GSC Sitemap Errors

"Couldn't fetch"

Googlebot attempted to fetch the sitemap but encountered a network timeout, 5xx server error, or robots.txt block. Often resolves automatically if transient, but check server logs for 504 gateway timeouts.

"Sitemap is HTML"

The submitted sitemap URL returned an HTML document (such as a 404 page or a login redirect) instead of XML. Ensure the endpoint returns Content-Type: application/xml.

"URL not allowed"

The sitemap contains URLs that do not reside within the path or host level of the sitemap location (e.g. submitting a sitemap at example.com/blog/sitemap.xml that attempts to declare URLs at example.com/store/).