Indexing is the process by which a search engine analyses an eligible crawled page and may store it for retrieval in search results.
Indexing is the process by which a search engine analyses an eligible crawled page and may store it for retrieval in search results.
What it means
Crawling discovers or fetches a URL; rendering interprets content; indexing is a separate eligibility decision.
A 200 response, sitemap entry or crawl does not guarantee indexing.
How to assess it
Use URL inspection and a crawl to compare rendered response, robots directives, canonical target, internal links and sitemap membership.
Practical workflow
- Define the user task and the set in scope.
- Use URL inspection and a crawl to compare rendered response, robots directives, canonical target, internal links and sitemap membership.
- Choose the smallest reversible change.
- State the intended indexable set, allow discovery, use consistent canonicals, and use noindex deliberately.
- Re-check the rendered outcome and preserve the measurement scope and date.
A concrete decision and verification example
An actual inspection example is a product filter URL that returns 200, appears in a sitemap and has internal links, but declares a canonical to the parent category. The correct conclusion is not “the filter is broken.” First decide whether the filtered view deserves its own indexable document: it needs a distinct search need, enough useful content or inventory, and a stable canonical URL. Otherwise the parent category may be the intended representative.
Work in an explicit indexation matrix. For each URL class, record discovery route, HTTP status after redirects, robots accessibility, noindex state, canonical target, sitemap inclusion and expected index state. The matrix makes conflicting controls visible: a noindex page blocked in robots.txt cannot be reliably processed for noindex because the crawler cannot read the directive. Google’s documentation is explicit about this dependency.
Verify in two passes. Crawl the rendered site to test controls at scale, then inspect representative URLs in the search engine’s own tools to see what it reports. Preserve the inspection date and exact URL. Indexing may take time and search engines choose canonicals using multiple signals, so an implementation can be correct before the result is reflected in every report.
Limits and common mistakes
Blocking robots.txt alone is not an index-control mechanism.