301 chain301 chain
Permanent redirects stacked one after another.
Found by Redirect checker →
302302
A temporary redirect on a page that moved for good.
Found by Redirect checker →
IndexingUpdated

robots.txt

In short

robots.txt is a root-level text file that tells compliant crawlers which URL paths they may or may not crawl, using user-agent rules.

Mechanism and decision

robots.txt is a public root-level text file that gives crawl rules to compliant user agents. Rules are path- and agent-specific, so an overbroad prefix or a malformed served file can block more than intended.

Use it for narrow crawl-access decisions, not secrecy or removal from results. Keep rendering resources for public pages accessible and use a suitable noindex or removal policy when a URL must not appear in search.

The diagram shows a crawler reading allow and disallow path rules before requesting site sections.

What you can control

robots.txt neither authenticates visitors nor guarantees deindexation of URLs already known to a crawler.

Practical workflow

  1. Fetch the root file on the canonical host.
  2. Test actual prefixes and user-agent groups.
  3. Allow resources necessary for public rendering.
  4. Re-test after host, framework or CDN changes.

FAQ

It controls crawling, not removal from the index. An externally discovered URL can still be known to a search engine.
No. Protect private material with authentication or other access controls.