301→302301→302
A temporary redirect where a permanent one belongs.
Found by Redirect checker →
redirect loopredirect loop
The URL redirects back to itself and never settles.
Found by Redirect checker →
Audit

Duplicate and thin content checker

Exact duplicates by hash, near-duplicates, thin pages, low text ratio, boilerplate and placeholder text across a crawl.

Duplicate and thin content checker

Content analysis

Near-duplicates, boilerplate, semantic inputs and similarity over retained page bodies.

Duplicate and thin content checker

Readability

Flesch score and long sentences as optional content checks.

How it works

Three steps, one saved scan

  1. Crawl the site into a scan

    Run a crawl. It saves a native SQLite scan that keeps each page's extracted text for offline analysis.

  2. Compare every page body

    Exact copies are found by content hash. Near-duplicates are found with simhash and LSH, so pages are not compared pair by pair.

  3. Read the groups and fix them

    Each finding lists the pages in its group. Consolidate with a canonical tag or rewrite the pages so each one serves its own search intent.

What it checks

Checks in this feature

CheckWhat it meansSeverityEvidence
DUPLICATE_BY_HASHThe page body is identical to another page's bodyImportantHash group with the URLs that share it
NEAR_DUPLICATEThe page text is almost identical to another page's text, above the similarity threshold (default 0.92)ImportantCluster of similar URLs
THIN_CONTENTThe page has too few words to stand on its own in searchImportantWord count for the page
LOW_TEXT_RATIOVisible text is a small share of the HTML, so markup outweighs contentTipText-to-HTML ratio for the page
LOREM_IPSUM_PLACEHOLDERLorem Ipsum placeholder text is in the page's main content areaImportantMatched placeholder passage
READABILITY_DIFFICULTText has a low Flesch reading-ease score. This is an optional content checkTipFlesch score for the page
How to run it

One command

bash
seohead duplicate-check --scan ./scans/audit.sqlite

Install first: installation guide. Every command also runs as an MCP tool for AI agents.

Related checks

From the check registry

DUPLICATE_BY_HASHNEAR_DUPLICATETHIN_CONTENTLOW_TEXT_RATIOLOREM_IPSUM_PLACEHOLDER
FAQ

Questions

Duplicate and thin content checker
No. The crawl and the duplicate check only read pages. The check runs offline on the saved scan file.
An exact duplicate has the same content hash. A near-duplicate has different bytes but text similar enough to pass the threshold. You can change the threshold with --threshold, from 0 to 1.
Not by default. Only indexable pages are compared, because a page that canonicalizes to another URL is not a defect. Add --all-pages to include non-indexable pages.
Yes. The offline check analyzes up to 10,000 documents and 16 MiB of retained extracted text. If a bound is reached, the report states the exact partial coverage. It never reports a clean result for pages it did not analyze.
Header, nav and footer markup is grouped by the separate boilerplate-report command. It is not part of duplicate-check.

Check your redirects locally