301→302301→302
A temporary redirect where a permanent one belongs.
Found by Redirect checker →
redirect loopredirect loop
The URL redirects back to itself and never settles.
Found by Redirect checker →
Crawling

Sitemap and URL list crawler

Besides following links, SpiderHead can crawl the URLs a sitemap declares or an explicit URL list, and reconcile what the sitemap says with what the crawl found.

Sitemap and URL list crawler

Sitemap, list and URL-file modes

Fetch an explicit urls list instead of following links at all, or walk sitemap indexes recursively, gzip included.

  • Sitemap index and urlset
  • Duplicates and stale lastmod values
Sitemap and URL list crawler

Sitemap against crawl

The audit compares the sitemap URL set with the crawl: non-indexable or redirecting URLs in the sitemap, and indexable pages missing from it.

  • Sitemap orphans
  • Out-of-sync URL sets
How it works

Three steps, one saved scan

  1. Read the sitemap

    Give one sitemap URL. Sitemap indexes, child sitemaps, text lists and gzip files are followed down to the URL entries.

  2. Check each declared URL

    Every URL in the sitemap gets a response code, so redirects and broken pages show up in the list. Fetches run three at a time by default.

  3. Compare with the crawl

    The audit sets the sitemap URLs against the pages it crawled and flags what is out of sync.

What it checks

Checks in this feature

CheckWhat it meansSeverityEvidence
SITEMAP_URL_4XX_5XXA sitemap URL returns a 4xx or 5xx responseCriticalSitemap URL and its status code
SITEMAP_URL_3XXA sitemap URL redirects instead of returning 200ImportantSitemap URL and its redirect status
SITEMAP_URL_NON_INDEXABLEThe sitemap lists a URL that is not indexableImportantSitemap entry
SITEMAP_DESYNCThe sitemap and the crawled URL sets do not matchImportantSitemap set against crawl set
URL_NOT_IN_SITEMAPAn indexable page is missing from the sitemapTipCrawled canonical URL
SITEMAP_STALE_LASTMODThe sitemap has stale or boilerplate lastmod valuesTiplastmod values in the sitemap
How to run it

One command

bash
seohead sitemap-crawl --url https://example.com/sitemap.xml

Install first: installation guide. Every command also runs as an MCP tool for AI agents.

Related checks

From the check registry

SITEMAP_URL_3XXSITEMAP_URL_NON_INDEXABLEURL_NOT_IN_SITEMAPSITEMAP_DESYNCSITEMAP_STALE_LASTMOD
FAQ

Questions

Sitemap and URL list crawler
Sitemap indexes, urlset files, text sitemaps (.txt) and gzip files, both .gz files and gzip-encoded responses. Nested indexes are followed to the end.
Not with the sitemap command, which takes one sitemap URL. To import a local URL list, use the scan import command (scan-import-urls). It writes into a saved scan and creates a verified backup first.
No. It reads the sitemap and requests the URLs listed in it. It does not edit the sitemap or the pages.
Those checks need the crawled URL set, so they run in the audit that compares the sitemap with the crawl. The sitemap command on its own gives you the response codes for each declared URL.

Check your redirects locally