301 chain301 chain
Permanent redirects stacked one after another.
Found by Redirect checker →
302302
A temporary redirect on a page that moved for good.
Found by Redirect checker →
Technical SEOUpdated

Crawling: how search bots request and discover URLs

In short

The process in which a search engine discovers and requests URLs to collect content and signals; crawling is distinct from indexing, ranking, and reporting.

The process in which a search engine’s crawler discovers and requests URLs to collect content and signals; crawling is distinct from indexing, ranking, and reporting.

The decision and its boundary

Crawling is the retrieval stage in which a search engine discovers URLs through links, sitemaps, and other signals and requests them. A successful fetch does not mean that the page will be indexed, shown, or ranked. Conversely, an indexed URL may not be fetched immediately after a change. The actionable work is to make important URLs discoverable, accessible, internally linked, and consistent with the canonical and index-control policy.

The diagram highlights the main decision points behind Crawling: how search bots request and discover URLs.

What to inspect

Define a URL set before crawling: templates, important directories, sitemaps, and known legacy routes. Record response code, final URL, canonical, robots directives, rendered content availability, and link depth. Compare crawl output with sitemap and search-console evidence without treating their populations as interchangeable. Mark unreachable or unmeasured URLs explicitly.

A practical implementation path

Give important pages ordinary HTML links from relevant hubs, use a sitemap as a declaration of final indexable URLs, and keep robots rules aligned with the desired crawl policy. Fix response errors and redirect chains before asking for recrawl. After a change, repeat the same bounded crawl and inspect server or search evidence for the defined sample rather than making a sitewide conclusion from one URL.

Mistakes and measurement limits

The common confusion is saying that crawling caused ranking or that a robots block is an indexation solution. Crawling, indexation, and ranking are separate stages with different evidence. Do not call a 200 response a pass until the page’s final URL, canonical, and visible content have also been checked.

Worked crawl example. When a new category is launched, list its canonical URL, expected response, parent hub link, sitemap inclusion, robots policy, and rendering requirement. Crawl the parent and the category with a defined user-agent and rendering mode, then compare the collected URL with the intended state. If the crawler discovers a redirect, blocked resource, or different canonical, investigate that condition instead of calling the category crawled. Crawling one URL cannot demonstrate that a whole template family is discoverable.

Check criteria. A useful crawl record states the start set, date, response handling, rendering assumptions, limits, failures, and URL-normalization policy. For important URLs, inspect status, redirect chain, final URL, canonical, index directives, link source, and whether essential content rendered. Compare the crawl with sitemap entries and search-console coverage cautiously because they answer different questions and may use different timestamps. Keep unmatched populations visible.

Common errors. Do not use robots.txt to hide a URL while expecting it to communicate a noindex rule. Do not mistake a 200 response for a valid crawl target when it has an incorrect canonical, thin rendered body, or redirecting assets. Avoid saying a crawler ‘found all pages’ without a documented scope. Discovery, fetch, indexation, and ranking must remain separate in the report.

FAQ

No. Crawling is a fetch and discovery process. Indexation and ranking are separate decisions.
Use relevant internal links, include the final canonical URL in a sitemap where appropriate, and make the page accessible to crawlers.
The requested and final URL, response, canonical, robots directives, render state, link discovery context, and scope or gaps.