The process in which a search engine discovers and requests URLs to collect content and signals; crawling is distinct from indexing, ranking, and reporting.
The process in which a search engine’s crawler discovers and requests URLs to collect content and signals; crawling is distinct from indexing, ranking, and reporting.
The decision and its boundary
Crawling is the retrieval stage in which a search engine discovers URLs through links, sitemaps, and other signals and requests them. A successful fetch does not mean that the page will be indexed, shown, or ranked. Conversely, an indexed URL may not be fetched immediately after a change. The actionable work is to make important URLs discoverable, accessible, internally linked, and consistent with the canonical and index-control policy.
What to inspect
Define a URL set before crawling: templates, important directories, sitemaps, and known legacy routes. Record response code, final URL, canonical, robots directives, rendered content availability, and link depth. Compare crawl output with sitemap and search-console evidence without treating their populations as interchangeable. Mark unreachable or unmeasured URLs explicitly.
A practical implementation path
Give important pages ordinary HTML links from relevant hubs, use a sitemap as a declaration of final indexable URLs, and keep robots rules aligned with the desired crawl policy. Fix response errors and redirect chains before asking for recrawl. After a change, repeat the same bounded crawl and inspect server or search evidence for the defined sample rather than making a sitewide conclusion from one URL.
Mistakes and measurement limits
The common confusion is saying that crawling caused ranking or that a robots block is an indexation solution. Crawling, indexation, and ranking are separate stages with different evidence. Do not call a 200 response a pass until the page’s final URL, canonical, and visible content have also been checked.
Worked crawl example. When a new category is launched, list its canonical URL, expected response, parent hub link, sitemap inclusion, robots policy, and rendering requirement. Crawl the parent and the category with a defined user-agent and rendering mode, then compare the collected URL with the intended state. If the crawler discovers a redirect, blocked resource, or different canonical, investigate that condition instead of calling the category crawled. Crawling one URL cannot demonstrate that a whole template family is discoverable.
Check criteria. A useful crawl record states the start set, date, response handling, rendering assumptions, limits, failures, and URL-normalization policy. For important URLs, inspect status, redirect chain, final URL, canonical, index directives, link source, and whether essential content rendered. Compare the crawl with sitemap entries and search-console coverage cautiously because they answer different questions and may use different timestamps. Keep unmatched populations visible.
Common errors. Do not use robots.txt to hide a URL while expecting it to communicate a noindex rule. Do not mistake a 200 response for a valid crawl target when it has an incorrect canonical, thin rendered body, or redirecting assets. Avoid saying a crawler ‘found all pages’ without a documented scope. Discovery, fetch, indexation, and ranking must remain separate in the report.