301→302301→302
A temporary redirect where a permanent one belongs.
Found by Redirect checker →
redirect loopredirect loop
The URL redirects back to itself and never settles.
Found by Redirect checker →
Crawling

Custom extraction from crawled pages

Content selectors and declarative extraction rules run over retained page bodies, and saved-scan source search finds which pages carry a marker such as a GTM snippet.

Custom extraction from crawled pages

Extraction rules

Declarative rules over retained bodies, so you can extract again without recrawling.

Custom extraction from crawled pages

Search the saved HTML

Find every page that contains a string or pattern in a saved scan.

  • scan-content-search
  • Bounded, paged results
How it works

Three steps, one saved scan

  1. Crawl once and keep the bodies

    A saved scan keeps complete page bodies within the body-retention limit. Rules and searches then run on that copy, not on the live site.

  2. Run extraction rules offline

    Write rules as JSON with kind text, attribute, structured, presence or count, and operator equals, contains, matches or exists. Run them with seohead scan extract; nothing is fetched and the result is not saved.

  3. Search the saved HTML for a marker

    scan-content-search finds a literal string or a regex in raw HTML, head markup, body text or a CSS selector, and writes a local package. Read it 100 rows at a time with scan-content-search-page.

What it checks

Checks in this feature

CheckWhat it meansSeverityEvidence
TITLE_MISSINGTitle element is missing. A rule with kind text on the title selector finds the same gap across the saved scan.Criticaldocs/CHECKS.md registry row: severity critical, fix: add a unique, descriptive title.
CANONICAL_MISSINGIndexable page has no canonical URL. A presence rule on link[rel=canonical] lists the pages that lack it.Importantdocs/CHECKS.md registry row: severity warning (mapped to Important).
SCHEMA_VALIDATION_ERRORStructured data validation errors. A structured rule with a JSON pointer lets you read the JSON-LD values behind the error.Importantdocs/CHECKS.md registry row: severity warning, source SF Structured Data Validation Errors.
STRUCTURED_DATA_MISSINGStructured data is missing on pages where you expect it. A presence rule on the JSON-LD block shows the pages that have none.Tipdocs/CHECKS.md registry row: severity notice (mapped to Tip).
LOREM_IPSUM_PLACEHOLDERLorem Ipsum placeholder text appears in the page content area. A content search for the placeholder phrase finds it across the whole scan.Importantdocs/CHECKS.md registry row: severity warning, source crawl:lorem_ipsum_count.
How to run it

One command

bash
seohead scan-content-search --scan ./scans/audit.sqlite --query gtm.js

Install first: installation guide. Every command also runs as an MCP tool for AI agents.

FAQ

Questions

Custom extraction from crawled pages
No. scan-extract and scan-content-search read the retained bodies in the saved scan file. They never fetch a URL, so you can change a rule and run it again without a new crawl.
Not necessarily. A search hit shows the tag is in the HTML. It does not prove the tag fired for a visitor. The search reports where the code is present.
The page is reported as unknown, not as missing. A search never says a marker is absent when the body is not in the scan.
scan-extract does not save its result; it prints it for that run. scan-content-search writes a new local package, so you keep its matches.
Yes, with --kind regex for Python regular expressions. If a page takes longer than the regex time budget, it is reported as unavailable, not as a non-match.

Check your redirects locally