In short
robots.txt is a root-level text file that tells compliant crawlers which URL paths they may or may not crawl, using user-agent rules.
Mechanism and decision
robots.txt is a public root-level text file that gives crawl rules to compliant user agents. Rules are path- and agent-specific, so an overbroad prefix or a malformed served file can block more than intended.
Use it for narrow crawl-access decisions, not secrecy or removal from results. Keep rendering resources for public pages accessible and use a suitable noindex or removal policy when a URL must not appear in search.
What you can control
robots.txt neither authenticates visitors nor guarantees deindexation of URLs already known to a crawler.
Practical workflow
- Fetch the root file on the canonical host.
- Test actual prefixes and user-agent groups.
- Allow resources necessary for public rendering.
- Re-test after host, framework or CDN changes.
FAQ
It controls crawling, not removal from the index. An externally discovered URL can still be known to a search engine.
No. Protect private material with authentication or other access controls.