• 4 mins read
  • Published

Why Google Leaves So Many Pages Crawled But Not Indexed

Paul Christiano Journalist FAYFO.com

by Paul Christiano

Why Google Leaves So Many Pages Crawled But Not Indexed FAYFO.com
Why Google Leaves So Many Pages Crawled But Not Indexed

Thousands of site owners are seeing their pages stuck in Google's 'crawled - currently not indexed' status. Learn what causes this, why quality matters more than ever, and which tools can help you recover visibility.

More website owners are reporting that their pages are stuck in Google's 'crawled - currently not indexed' status, according to recent findings from Google Search Console. In nearly every case, the underlying issue is content quality. Most affected pages are considered 'commodity content'-material that repeats what’s already widely available online, without offering new insights or unique value.

At the Google Search Central event in Toronto in April 2026, presenters explained that Google's indexing process has become more selective. After crawling a page, Google only adds it to the index if it deems the content useful. With AI making content creation easier, Google now prioritizes pages that demonstrate personal experience or provide knowledge unavailable elsewhere.

There are two main reasons a page might be crawled but not indexed. The first is technical: for example, a misconfigured robots.txt file can block Google from accessing key resources, as seen in a recent site migration where a 'Disallow: /*?*' rule inadvertently hid essential CSS and JavaScript files. If Google's live test in Search Console shows your content is visible, a technical issue is unlikely. Some pages, such as /feed/ or paginated URLs, are also normally excluded from the index.

The second-and far more common-reason is quality. As one Google representative put it, if thousands of pages cover the same topic, Google may decide your version adds little value. Sometimes, Google will temporarily index a page to test user engagement, but if it doesn't outperform existing results, it may be dropped.

Commodity content is the main culprit. Google emphasized that content anyone could write, especially with AI, is less likely to be indexed. Non-commodity content, by contrast, offers unique perspectives or first-hand experience. For example, this article draws on direct observations from professional SEO work and insights from industry events, rather than simply summarizing public documentation.

When analyzing pages stuck in this status, start by reviewing the 'Crawled - currently not indexed' report in Search Console. Identify URLs you want indexed, then search for their target queries. If Google's AI-generated answers already satisfy user intent, your page may not offer enough additional value to warrant inclusion. Tools like Gemini in Chrome can help compare your content against top-ranking pages, checking for originality, depth, and unique analysis.

It's important to note that authority can sometimes compensate for commodity content, but most sites need to demonstrate clear, original value. Gemini and similar tools can't reveal Google's exact ranking logic, but they can highlight where your content may fall short compared to competitors.

Google's John Mueller and Martin Splitt recently discussed this issue on the 'Search Off the Record Podcast.' They confirmed that a high volume of 'crawled - not indexed' pages often signals a site-wide quality concern. If Google's systems doubt the overall value of a website, they will crawl and index fewer pages. Quality isn't just about text-user experience, page layout, and accessibility all play a role. Pages hidden behind ads, interstitials, or excessive filler are less likely to be indexed, even if the main content is unique.

Fixing this issue is challenging. Technical problems can be resolved by correcting site configurations and requesting reindexing. But when quality is the problem, significant effort is required. Sites that rely on mass-producing content-especially with AI-risk being filtered out. Google's June 2026 spam update reportedly impacted many such sites, though no manual actions appear in Search Console.

To improve, focus on drawing from real experience and adding new knowledge to your topics. Use prompts in Gemini or other LLMs to assess whether your content is likely to be seen as commodity. Ask for ideas to inject first-hand insights and make your articles stand out.

Several tools can help. At tools.mariehaynes.com, you can filter your 'crawled - not currently indexed' URLs to focus on those that matter, and use the GSC Index Checker to monitor indexing status. These resources streamline the process of identifying and addressing problematic pages.

Google's approach to indexing is evolving, with stricter standards for what gets included. For more on how Google tests new content visibility, see this analysis of how Google Discover handles brand-new articles.

Related articles