301→302301→302
A temporary redirect where a permanent one belongs.
Found by Redirect checker →
redirect loopredirect loop
The URL redirects back to itself and never settles.
Found by Redirect checker →
Crawling

Crawl scope, robots.txt and directives

Decide what a crawl may touch: scope rules, robots policy, and a robots.txt check that explains which rule blocks which URL.

Crawl scope, robots.txt and directives

robots.txt check

robots-check reads robots.txt and reports blocked paths, declared sitemaps and resources that rendering needs.

  • Blocked JS or CSS for rendering
  • Sitemap directive present or missing
Crawl scope, robots.txt and directives

Directives

Page-level noindex, nofollow, nosnippet, unavailable_after and X-Robots-Tag are read and reported.

How it works

Three steps, one saved scan

  1. Set the crawl scope

    Include and exclude patterns, named segments, extension filters, a depth limit and a URL cap decide which pages a crawl fetches.

  2. Pick the robots policy

    Each crawl can obey robots.txt, only report what it blocks, or ignore it. The flag is --robots respect, report_only or ignore.

  3. Test a path before you crawl

    robots-check reads the file for one user-agent and tests the paths you give it, so you know what a crawler is allowed to request.

What it checks

Checks in this feature

CheckWhat it meansSeverityEvidence
IMPORTANT_URL_BLOCKED_BY_ROBOTSA live page that receives internal links is blocked by robots.txt. The links carry no value for discovery, so the page is lost from the crawl path. Robots.txt blocks crawling, not indexing.ImportantBlocked URL, the internal links pointing to it, and the matching Disallow rule
BLOCKED_BY_ROBOTSA URL is blocked by robots.txt. Confirm the block is intentional.ImportantBlocked URL and the matching rule
ROBOTS_BLOCKS_RESOURCESrobots.txt blocks JavaScript or CSS needed for rendering, so search engines may render the page incompletely.ImportantBlocked .js or .css URL and the rule that blocks it
DIRECTIVES_OUTSIDE_HEADA robots meta directive sits outside the head element, so it is not applied. A noindex or nofollow in that position has no effect.CriticalThe element that precedes the directive in the document
UNAVAILABLE_AFTERThe page carries an unavailable_after directive, which removes it from the index on a set date.ImportantDirective value and date
NOINDEXThe page contains a noindex directive. Confirm that leaving the index is intentional.TipDirective source: meta robots or X-Robots-Tag
SITEMAP_NOT_IN_ROBOTSrobots.txt has no Sitemap directive. Adding one with the absolute sitemap URL helps crawlers find it.TipSitemap directives found in robots.txt
How to run it

One command

bash
seohead robots-check --url https://example.com/

Install first: installation guide. Every command also runs as an MCP tool for AI agents.

Related checks

From the check registry

BLOCKED_BY_ROBOTSIMPORTANT_URL_BLOCKED_BY_ROBOTSROBOTS_BLOCKS_RESOURCESSITEMAP_NOT_IN_ROBOTSNOINDEXNOSNIPPET
FAQ

Questions

Crawl scope, robots.txt and directives
No. It stops crawling. A blocked page can still be indexed from links, and a noindex on that page is never read while the page is blocked. To keep a page out of the index, allow crawling and use noindex.
A 404 means no robots.txt exists, so the report says crawling is allowed. A 429 or 5xx response returns an error, and the tool never claims that crawling is allowed.
Yes. Use scope include or exclude patterns, scope segments_only to fetch named segments, and extension filters. Depth and URL caps keep the crawl inside the limits you set.
The flag sets the policy: respect obeys robots.txt, report_only records what it blocks without obeying it, and ignore skips it. Use respect when you want a crawl to behave like a search engine crawler.

Check your redirects locally