301→302301→302
A temporary redirect where a permanent one belongs.
Found by Redirect checker →
redirect loopredirect loop
The URL redirects back to itself and never settles.
Found by Redirect checker →
Crawling

Resume an interrupted crawl

Crawls are resumable. A scan is a retained SQLite artifact with provenance, so a stopped crawl continues from where it stopped instead of starting again.

Resume an interrupted crawl

Retained scans

Every scan is kept: inspect, export, snapshot, pin, prune, requeue, diff bodies and reanalyse offline without contacting the site again.

  • scan-reanalyze reparses retained HTML
  • Snapshots and pins for baselines
Resume an interrupted crawl

Recovery

Interrupted runs are recorded and can be resumed; crawl-diagnose explains what stopped a crawl.

How it works

Three steps, one saved scan

  1. The scan file is the checkpoint

    A SQLite scan keeps the queue, the evidence and the full settings in one file. They are written in the same transaction, so a stop does not split them.

  2. Resume from the file alone

    Run seohead crawl-site --resume with the scan path. The start URL and settings are read from the file, so the crawl continues under the same rules.

  3. Diagnose what stopped it

    crawl-diagnose reads the scan offline. It shows the queue, robots, scope and budget state, the content types and the recorded failures, with no network request.

What it checks

Checks in this feature

CheckWhat it meansSeverityEvidence
NO_RESPONSEThe URL timed out, failed DNS or refused the connection. A stopped origin often shows up here.CriticalURL and failure type in the scan
SERVER_ERROR_5XXThe page returns a 5xx status, for example when the origin stops responding under load.CriticalURL and status code
BROKEN_PAGE_4XXThe page returns a 4xx status, so it is broken for users and crawlers.CriticalURL and status code
BLOCKED_BY_ROBOTSrobots.txt blocks the URL. Confirm the block is intended before resuming.ImportantURL and the robots.txt rule
How to run it

One command

bash
seohead crawl-diagnose --scan ./scans/audit.sqlite

Install first: installation guide. Every command also runs as an MCP tool for AI agents.

FAQ

Questions

Resume an interrupted crawl
No. Passing --url, --config or --max-urls together with --resume is refused. The stored start URL and settings are used, so one crawl never runs under another crawl's rules.
Yes. A resume from a different source revision is refused by name. If this checkout cannot be verified as the build that wrote the file, pass --producer-build with its SHA.
No. A page already committed to the scan is not fetched twice. The resumed run continues from the stored queue.
No. The JSON output and the stored audit record partial, finish_reason and resumed, so a partial scan never reads as a whole site.
Run the identical command again with the same --out-dir. The crawl_state.json checkpoint in that folder decides the resume. If the scope or limits changed, the crawl starts over on purpose.

Check your redirects locally