An XML sitemap is a machine-readable list of a site’s preferred URLs that helps search engines discover pages but does not guarantee crawling or indexing.
Meaning and scope
An XML sitemap is a machine-readable file that lists URLs a site wants search engines to know about. It is especially useful when site size, architecture or internal linking makes discovery harder. A sitemap supplements crawlable internal links; it does not make a URL indexable, canonical, useful or guaranteed to be crawled.
How to inspect it
Audit a sitemap as a declaration of the preferred indexable set, not as a file that merely returns 200. Sample entries for final response, self-canonical state, indexability and presence in the intended inventory. Check that last-modified dates represent meaningful page changes rather than a boilerplate timestamp, and compare sitemap counts with live sections instead of repeating an old total.
- Generate entries from the maintained content or product inventory rather than a hand-kept list.
- Include only final canonical URLs that a search engine can fetch and index.
- Split large sets into sitemap files and a sitemap index within protocol limits.
- Declare the sitemap in robots.txt or submit it through search-engine tools where appropriate.
- Recount and revalidate after launches, removals, redirects and CMS migrations.
Practical decisions
I generate sitemap sections from the actual content inventory and recount live output instead of repeating a number from an old document. The same principle applies to any site: wire sitemap generation to the source that knows whether a URL is published and preferred, then review exceptions explicitly. Keep staging and unpublished routes out of the production declaration.
Limits and mistakes
Submitting a sitemap does not force crawling, indexing or a rank. Do not add redirected, noindexed, duplicate, blocked or error URLs merely to increase the count. A sitemap cannot compensate for orphaned pages, broken navigation or a canonical policy that contradicts its entries.