Upgrade to Pro — share decks privately, control downloads, hide ads and more …

How Google Manages its Index (SearchNorwichXL)

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest. →
Avatar for Adam Gent Adam Gent
September 25, 2026
52

How Google Manages its Index (SearchNorwichXL)

Avatar for Adam Gent

Adam Gent

September 25, 2026

Transcript

  1. How Google Manages its Index 1 What is Google’s Index?

    2 How Does Google Index Useful Pages? 3 Why Pages Are Actively Removed by Google? 4 How Quality Thresholds Manage Google’s Index? 5 What is the Connection Between Crawling and Indexing?
  2. Google’s index is just a database… “Google's index technically is

    just a large database sitting on thousands of computers.” Gary Illyes, ‘The Grumpy One’, Analyst on Google Search team How Google Search indexes pages
  3. Eligible to be served in search results. “Once a page

    is indexed, it becomes eligible to show up in Google Search results...” Source: Search Console Help, Google Search Indexing and Ranking FAQ
  4. Google wants to index useful content for its users from

    the web… Source: USA vs. Google antitrust trial
  5. …but there is A LOT of spam. 60% of the

    web is duplicate. 40 billion pages of spam everyday. 96.55% of content gets NO traffic from Google. Source: 60% Of The Internet Is Duplicate, Google Web Spam Report 2020, Ahrefs Study
  6. Googlebot fetches and downloads content… Robots.txt Web Crawler Indexing pipeline

    Source: Olga Zarr SEO, JavaScript SEO in 2024 Q&A with Martin Splitt from Google Scheduler URL Queue
  7. …indexing pipeline processes content… Indexing Pipeline Eligibility Checks HTML Parser

    Web Rendering Service Canonicalization Signal Collection Spam Detection Source: Olga Zarr SEO, JavaScript SEO in 2024 Q&A with Martin Splitt from Google
  8. …and decides if a page should be indexed. Content Analysis

    Canonicalization Rendered DOM Source: Olga Zarr SEO, JavaScript SEO in 2024 Q&A with Martin Splitt from Google Database
  9. “When Google Search notices a pattern of low-quality or thin

    content on pages, they might be removed from the index and might stay in Discovered.” Martin Splitt, ‘The Happy One’, Googlebot Whisperer on the Search Relations team, Google Search Central, Help! Google Search isn’t indexing my pages
  10. Our study of 1.4 million pages found that 86% of

    not indexed pages are caused by “quality” issues on the website. Source: New Study: The Biggest Reason Why Your Pages are Not Indexed in Google, Indexing Insight
  11. Big vs Small Brands: Quality No. 1 Reason Source: New

    Study: The Biggest Reason Why Your Pages are Not Indexed in Google, Indexing Insight
  12. ‘Crawled - previously indexed’ report was created to monitor removals.

    Source: ‘Crawled – previously indexed’ report from Indexing Insight
  13. Processed content has a folder of collected signals and annotations.

    Historic HTML Vector Embeddings Inverted Index (IR Scores) Anchors (Links) Document (DocID) Duplicate Cluster (Canonical URL)
  14. "A document in our database looks more like a folder

    with lots of information in that folder… There's all sorts of scores and identifiers and bits and pieces attached to it. The content is just the smallest part." Martin Splitt, ‘The Happy One’, Googlebot Whisperer on the Search Relations team, Search Engine Journal, Google Passages – What They Are & What They Are Not
  15. A signal is any measurable data point or computed value

    that Google's algorithms use to assess relevance, quality, or user satisfaction of a piece of content. Source: Summarised from How Search Works, Google
  16. Raw Signals Computed Signals Raw data pulled from crawled content.

    Signals calculated from raw data points. • Links • PageRank • Body content • Vector Embeddings • Canonical tag • Canonical URL
  17. Two computed signals Google uses to assess if a page

    should be indexed. Search Interest Does the page appear in search? Page Quality Is the page getting good engagement?
  18. Search Interest Search Interest Does the page appear in search?

    Page Quality Is the page getting good engagement?
  19. Google’s logging and storage system Search Logs Glue Other Services

    Raw query and interaction data is logged by Google’s search logging infrastructure. Glue aggregates and structures the log, interaction and SERP data into interaction tables. Features like Navboost, QBST, RankEmbed BERT and Tetris use the aggregated data.
  20. Google’s logging system helps it keeps a record of search

    interest for each page. Glue Aggregated interaction data (demand). Document (DocID) Navboost 13 months of click data (engagement).
  21. Google uses interaction data to tier its index by search

    demand. RAM 🔥 High search demand SSD ☕ Medium search demand HDD 🥶 Low search demand Source: Search Engine Roundtable, Google Search Indexing Tiers
  22. A page with little search interest: Does not get served

    in search results OR It gets removed from search results
  23. “...PhD publication my PhD thesis is sitting somewhere on a

    web server no one is interested in it why would we index it.”. Gary Illyes, , ‘The Grumpy One’, Grumpy Analyst on Google Search team My site isn't indexed! (Search Off the Record)
  24. Page Quality Search Interest Is there interest in the topic?

    Page Quality Is the page getting good engagement?
  25. 3 Top-Level Signals in Google Search Popularity (P*) Topicality (T*)

    Source: DOJ Trial Document 1436 Quality (Q*)
  26. What is Page Quality (Q*)? 1) Content Google creates vector

    phrase model of each topic and compares your page’s main body content to high-quality pages. Nearest Seed PageRank is key signal in understanding page 2) Links quality. Google measures the distance from “trusted” pages. A key signal of page quality is the measuring the user-satisfaction 3) Clicks over time in Google SERPs. Source: DOJ Trial Document 1436 & Interview with Google Engineer
  27. Our analysis found that if an indexed page has poor

    user engagement over time it gets actively removed by Google Search.
  28. Google has confirmed there is a relationship between quality and

    indexing. "Index selection, while it's largely about (RAM/flash/disk) space, it's tightly tied to quality of content.” Gary Illyes, Grumpy Analyst on Google Search team Late night tweet from Gary
  29. Soft limit is a threshold below the hard limit that

    stops disk space being used up.
  30. Gary confirmed a similar process within Google’s own index. "If

    we have tons of free space available, we're more likely to index crappier content. If we don't, we might deindex stuff to make space for higher quality docs.” Gary Illyes, Grumpy Analyst on Google Search team Late night tweet from Gary
  31. “Quality and popularity signals, for instance, help Google determine how

    frequently to crawl web pages to ensure the index contains the freshest web content.” Source: Document 1436, Pg. 138, Google Antitrust Trial
  32. The 130 Day Indexing Rule Source: New Study: The 130

    Day Indexing Rule, Indexing Insight
  33. The 190 Day Indexing Rule Source: New Study: After 190

    Days Since Last Crawl Googlebot Forgets, Indexing Insight
  34. “Those [URL is Unknown to Google] have no priority; they

    are not known to Google (Search) so inherently they have no priority whatsoever. URLs move between states as we collect signals for them, and in this particular case the signals told a story that made our systems "forget" that URL exists.” Gary Illyes, Grumpy Analyst on Google Search team LinkedIn Comment
  35. How Google Manages its Index 1 Page Indexing report shows

    ALL processed content. 2 Indexed means you are eligible to appear in search. 3 Page quality is a BIG reason why pages are removed. 4 Page quality is used by Google to manage its index. 5 Page quality is used for crawl priority.