301→302301→302
A temporary redirect where a permanent one belongs.
Found by Redirect checker →
redirect loopredirect loop
The URL redirects back to itself and never settles.
Found by Redirect checker →
Crawling

Polite crawl budget and rate limits

SpiderHead is polite by default: about two requests per second per host, bounded read-only requests, and an explicit URL cap that is a budget rather than a guess about site size.

Polite crawl budget and rate limits

Budgets you set

Request and time budgets, scope and template rules, include and exclude rules by extension or media type.

  • --max-urls is an explicit budget
  • Politeness adapts to the origin
Polite crawl budget and rate limits

Safe by design

Network tools block private targets unless explicitly allowed; the crawler sends only bounded, read-only requests to the site you name.

How it works

Three steps, one saved scan

  1. Set a URL budget

    --max-urls caps the pages a crawl fetches (default 200), and going past a project budget needs --approve-large-crawl.

  2. Pace each host

    The default delay is 0.5 seconds between requests, about two per second per host; --max-urls-per-second sets the rate, and adaptive back-off slows the crawl when the server struggles.

  3. Honor robots.txt and private targets

    robots.txt is obeyed by default, and private or local addresses stay blocked until you allow them by name.

What it checks

Checks in this feature

CheckWhat it meansSeverityEvidence
SLOW_RESPONSEThe server took longer than 1.5 seconds to respondImportantResponse time per page against the 1.5 s threshold
BLOCKED_BY_ROBOTSThe URL is blocked by robots.txtImportantBlocked URL from the crawl or SF export
IMPORTANT_URL_BLOCKED_BY_ROBOTSA page that receives internal links is blocked by robots.txt, so link discovery stopsImportantBlocked URL and its internal inlinks
ROBOTS_BLOCKS_RESOURCESrobots.txt blocks JavaScript or CSS needed to render the pageTipBlocked resource URL
SITEMAP_NOT_IN_ROBOTSrobots.txt has no Sitemap directiveTiprobots.txt content
How to run it

One command

bash
seohead crawl-site --config-help

Install first: installation guide. Every command also runs as an MCP tool for AI agents.

Related checks

From the check registry

SLOW_RESPONSE
FAQ

Questions

Polite crawl budget and rate limits
It waits 0.5 seconds between requests by default and backs off on its own when the server slows or times out. A crawl stops after 5 timeouts in a row.
Yes, by default. --robots report_only fetches robots.txt, reports what would be blocked, and crawls anyway. --robots ignore skips robots.txt entirely, so use it only on sites you own.
Only after you opt in. SEOHEAD_ALLOW_PRIVATE_HOSTS allows named hostnames, and SEOHEAD_ALLOW_PRIVATE_NETWORKS=1 opens every private range.
Yes, with --max-urls-per-second. It is a floor, not a guarantee: adaptive back-off still slows the crawl if the server struggles. Raise it on servers you control.

Check your redirects locally