Millions of pages removed from Google's index after May 2025 "silent" update

Downward-trending line graph showing a sharp drop in indexed pages starting May 2025, labeled "Deindexed"

For two weeks I watched “Crawled - currently not indexed” pages surge in Google Search Console, especially on my clients’ very large websites. The spike started around May 26 to 29, 2025, and I was not alone: webmasters across unrelated industries reported the same drop in the same window, which points to a change in Google’s indexing behavior rather than isolated site issues.

Google Search Console chart showing affected pages rising to 7.93 million with a sharp spike in late May 2025
Affected pages climbed to 7.93 million, spiking in late May 2025

Site owners first flagged it in the final week of May. By early June, SEO forums and social feeds were full of matching reports. Many GSC charts show a hard inflection on May 27, 2025, where excluded pages, especially in the “Crawled - currently not indexed” bucket, jumped.

What likely caused the deindexing spike

The core question is why so many pages, across so many sites, were suddenly dropped. Google never confirmed a cause, so the explanations below are the strongest theories the SEO community landed on, ordered from most to least supported by evidence.

Quality-based pruning (most supported)

Google appears to have raised its quality bar and culled thin, unoriginal, or duplicate pages at scale. This is the theory with the most direct evidence behind it.

Marie Haynes analyzed affected sites and found a consistent pattern in what got removed: thin, unoriginal, low-value pages. On a travel site, old posts that rehashed information available elsewhere were dropped. On a recipe blog, simple “What is the difference between X and Y?” posts were dropped, the exact queries Google’s own AI can answer. On an attorney site, generic “SEO fluff” tips that added nothing were removed.

This lines up with Google’s direction in its January 2025 Quality Rater Guidelines, which added prominent new guidance on paraphrased content (sections 4.6.6 and 4.6.7) and instructs raters to give the Lowest rating to content that is copied, paraphrased, or AI-generated with little effort or added value.

A silent algorithm update (unconfirmed)

Many SEOs believe Google rolled out an unannounced indexing adjustment. The tight, shared timing across sites argues against random fluctuation.

Barry Schwartz of Search Engine Roundtable reported a wave of complaints about Google indexing fewer pages from late May onward. Some called it an “index pruning update.” Google stayed silent, so it sits in the “silent update” category: real in effect, never officially acknowledged.

Crawl budget or index-capacity limits (plausible)

Google may have tightened how many pages it keeps indexed on very large sites. Google never indexes everything it crawls, and it prioritizes pages it judges important.

A Consainsights user noted that even active sites with good content saw pages dropped on May 27 to 29, so this was not purely a garbage-page cleanup. Real estate portals are the clearest example: huge page counts, many similar or transient (expired listings, near-duplicate profiles). If Google recalibrated how much of a site deserves indexing, low-priority pages would be first out. It is hard to separate this from quality pruning, since low-value pages also carry low crawl priority.

A tie to AI-generated answers (speculative)

Some suspect Google is dropping pages whose answers its AI already covers. Treat this as a hypothesis, not a finding.

If a definition or basic FAQ page is now redundant next to an AI snippet, Google may see less reason to index it. That fits the observation that many dropped pages were basic-info pages. Google has never linked its AI answers to indexing behavior, so this stays speculative.

Google’s response: “normal and expected”

Google issued no formal statement and treated the change as routine. The Search Liaison channels stayed quiet, which itself signals Google saw normal fluctuation.

The closest thing to a response came from John Mueller on Bluesky in early June. Asked about millions of pages being dropped, he first asked for specific examples. After reviewing some, he said he saw no technical issue on Google’s side or on the sites, and that adjustments to what gets crawled and indexed are, in his words, “normal and expected.” In a follow-up he added that Google does not index all content and that what it indexes changes over time.

For context, Google’s Martin Splitt has explained why pages fall out of the index after first getting in: Google indexes a page, sees that searchers do not actually use it, and lets better-performing pages take its place. That is the same logic operating at scale here.

What I saw across my own properties

I analyzed multiple sites I manage, mine and clients’, to find patterns that reproduce rather than one-off anecdotes. Four stood out.

Large sites took the hit. On sites under 10,000 pages I saw little to no change. On portals above 50,000 URLs, the drop in indexed pages was significant.

Google Search Console chart for a small blog under 1000 pages showing a flat, stable line around 697 affected pages
A blog under 1,000 pages showed no notable change in “Crawled - currently not indexed”

Google ramped crawling in late April, then pruned. Almost every affected domain showed a crawl spike at the end of April 2025, peaking around April 21. It read like Google reassessing entire sites before dropping pages weeks later.

Google Search Console crawl stats showing total crawl requests spiking to over 821,000 on April 21, 2025 before dropping
A pronounced crawl spike in late April 2025, peaking near 821,000 requests on April 21

Internal linking stopped saving thin pages. A strong internal linking setup used to lift low-value pages. Not this time. Pages with dozens of internal links still dropped because the content itself did not clear the quality bar. Google looks willing to ignore link equity when the content does not hold up.

“Thin” means depth, not just word count. On one directory site, pages with only 1 to 2 listings got deindexed even though the templates were clean, structured, and internally linked. Pages with 4 or more listings stayed in. Google set a soft usefulness threshold, and sparse pages fell below it.

Taken together, this reads as a deliberate index cleanup: shed bloat, lift SERP quality, and cut redundancy ahead of AI-generated results. What survives is pages with clear value, authority, or at least a floor of usefulness.

How to recover indexing

The through-line for every fix below: give Google a reason to keep the page. Ordered by impact.

Improve content quality and originality

This is the durable fix. If a page was dropped for being thin or redundant, the only lasting fix is making it genuinely better.

Expand dropped pages with real substance: worked examples, current data, first-hand analysis that was not there before. Turn a dropped three-line FAQ into a proper guide with subheads, visuals, and expert detail. Then push past “same as everyone else”: add your own research, small experiments, user-generated input, or a clear point of view. Google’s guidelines reward Experience and Expertise, meaning first-hand knowledge, not rephrased consensus.

Strengthen E-E-A-T signals

Add author bylines with real credentials, cite authoritative sources, and verify your facts. These will not force indexing on their own, but if Google read a page as low-trust or low-expertise, better signals help over time.

Prune or merge dead weight

Sometimes the right move is to let a page go. If a page truly serves no unique purpose, merge it into a stronger page, 301-redirect it, or leave it out. Google not indexing it is a signal worth taking.

Many dropped pages had zero external links. You cannot conjure links, but auditing the backlink profile of deindexed pages is revealing. For pages that matter and that you have genuinely improved, promote them: share in communities, do light outreach, get a handful of legitimate links. Even a few can signal that real people value the page. Focus this effort on strategic pages, not everything.

Use fresh content to trigger re-crawls

Link improved old pages from new posts Google will crawl quickly. A new article referencing and linking a dropped page can send Googlebot back to reassess it. Pair that with the improved content and a link or two, and you tip the odds.

Check sitemaps and inspect URLs

Keep your XML sitemap current with the pages you want indexed. It will not override a “crawled, not indexed” decision, but it reinforces that the pages matter. Then use the URL Inspection tool for specific URLs: check the last crawl date, run “Test Live URL,” and look at the canonical Google chose. If Google picked a different canonical or flagged a duplicate, fix the duplication.

Monitor and be patient

Some pages return on their own as the site earns more trust. Track the GSC Pages report weekly or monthly. Are improved pages getting reindexed? Are new pages hitting the same wall? Reindexing after your edits means the approach worked. If nothing moves, revisit why Google still sees the page as not worth keeping.

Looking ahead

Do not expect Google to clarify this further. Given Mueller’s stance, Google considers it business as usual, and confirmation is unlikely unless enough confusion forces a documentation note.

The takeaway holds regardless: Google does not index the whole web and does not intend to. It aims to index the best of the web. Keep your pages unique, useful, and relevant for their topic, and you protect their place in the index and their staying power in search.

Update: July 2026

The May 2025 pruning was the leading edge of a direction Google kept pushing, not a one-off. Google’s March 2026 Core Update (March 27 to April 8, 2026) rewarded original, first-hand content and demoted pages that mostly rephrase what already ranks. Analysts framed the core signal as Information Gain: how much genuinely new knowledge a page adds compared with what already ranks.

The recovery playbook above is still the play. Depth, originality, and first-hand expertise protect indexing. Thin and paraphrased pages keep losing ground with each cycle.

Susanna Marsiglia
Written by Susanna Marsiglia

SEO, GEO & UX consultant since 2013. I help websites get found by humans and AI alike.