Indexing
Indexing is the process by which search engines store crawled web pages in a searchable database so they can be retrieved as results — a page must be indexed before it can rank or be cited.
Also known as: search indexing, index inclusion, page indexation
Indexing is the process by which search engines store and organize crawled web pages in a searchable database. A page must be indexed before it can appear in search results or be cited in AI answers built on top of search. Indexing is distinct from crawling — crawling fetches the page, while indexing decides whether to add it to the searchable database and how to represent it there.
What Indexing Means
Indexing is the database step search engines run after crawling. It evaluates each crawled page for inclusion, checking quality signals, duplicate detection, and content value, then stores accepted pages in the searchable index that powers query responses. Pages can be excluded at multiple points: blocked by robots, noindexed by directive, deemed duplicates of canonical alternatives, or judged too thin or low-value to merit a slot. Search console reports the status of each tracked URL — indexed, excluded, with reasons — which is the authoritative source for what’s actually in the index versus only crawled.
How Indexing Works
Indexing works through a sequence: a crawler fetches a page, the rendering engine processes any JavaScript needed to see the content, the indexing system evaluates the page for inclusion (checking quality signals, duplicate detection, content value), and an indexed page becomes eligible to rank for relevant queries. Submission through search console’s URL Inspection tool can accelerate indexing for individual pages, and inclusion in the XML sitemap signals priority. Mobile-first indexing means Google primarily uses the mobile version of a site for indexing and ranking, so the mobile version is the source of truth for content, links, and metadata.
Common Pitfalls and Misconceptions
A common mistake is assuming every page on a site is automatically indexed. Modern search engines are selective: a fully crawlable page can still be excluded from the index if quality signals don’t meet a threshold, if the content duplicates other indexed pages, or if the engine simply decides it doesn’t need to keep that page. Another misconception is that indexing and crawling are the same thing. They are sequential — a page must be crawled before it can be indexed, but being crawled doesn’t guarantee indexing. The ‘Crawled but not indexed’ status in search console is one of the most common quality-based exclusions and signals that the page hasn’t earned an index slot.
Indexing in Practice
The mature practice is to monitor indexation as a portfolio metric, not just at the individual-page level. Track the share of important URLs that are indexed, the patterns of exclusions (‘Crawled — currently not indexed’ versus ‘Discovered — currently not indexed’ versus ‘Duplicate without user-selected canonical’), and the trend over time. Sites that watch indexation as a system metric catch quality and architecture problems early — usually months before the symptoms show up in traffic — and they make different decisions about which pages to invest in versus consolidate or remove. AI answer engines built on search indexes only cite pages in those indexes.
Common questions.
What is the difference between crawling and indexing?
How do you check if a page is indexed?
Why do pages get excluded from the index?
How do you get a page indexed faster?
What is mobile-first indexing?
Can a page be partially indexed?
How does indexation relate to AI answer engines?
Related Terms
More from Search & AEO.
Let’s Talk
Let’s talk about what your next quarter could look like.
Tell us what you’re working on. A senior practitioner reads it, not an SDR queue, and replies, usually within one business day.
- Reviewed personally, not routed through a queue.
- A conversation about what you’re actually working on, not a generic pitch.
- No pressure, just a chance to talk it through.