The Great Indexing Dilemma
No two status messages cause more debate among developers and SEOs than:
- Discovered - currently not indexed
- Crawled - currently not indexed
While they look similar in the interface, their technical root causes exist at opposite ends of Google’s pipeline.
Detailed Comparative Diagnostic
┌───────────────────────────────────────────────┬───────────────────────────────────────────────┐
│ DISCOVERED - CURRENTLY NOT INDEXED │ CRAWLED - CURRENTLY NOT INDEXED │
├───────────────────────────────────────────────┼───────────────────────────────────────────────┤
│ The URL was discovered (via sitemap or link), │ Googlebot requested, downloaded, and parsed │
│ but Googlebot has NOT yet fetched or crawled │ the HTML/DOM, but deliberately chose NOT to │
│ the URL due to crawl budget or server load. │ write the page into the primary search index. │
├───────────────────────────────────────────────┼───────────────────────────────────────────────┤
│ ROOT CAUSE: │ ROOT CAUSE: │
│ • Server latency / host crawl capacity limits │ • Content quality / low information gain │
│ • Excessive crawl bloat (faceted navigation) │ • Near-duplicate content / thin templates │
│ • Domain authority / trust threshold too low │ • Algorithmic quality evaluation threshold │
└───────────────────────────────────────────────┴───────────────────────────────────────────────┘
Deep Dive: “Discovered - Currently Not Indexed”
When Google marks a URL as Discovered - currently not indexed, Google’s crawler made a deliberate decision to postpone crawling.
Primary Causes & Solutions:
- Server Overload Signals: If your origin server experiences high TTFB (>1,000ms) or frequent HTTP
503or429spikes, Googlebot reduces its crawl rate to protect your infrastructure.- Fix: Optimize server caching, deploy edge caching (Cloudflare, Fastly), or upgrade origin compute resources.
- Crawl Waste & URL Bloat: Generating thousands of parameterized URLs (
?color=blue&size=m&sort=price_asc) exhausts Googlebot’s crawl budget before it ever reaches genuine content.- Fix: Apply
robots.txtdisallows or canonical tags on parameter matrices.
- Fix: Apply
- Internal Linking Orphans: URLs submitted in an XML sitemap but receiving zero internal links across the website are deprioritized by Google’s discovery scheduler.
- Fix: Link to newly published articles or products directly from high-authority hub pages or category nodes.
Remediation:
The engineering team added a Disallow: /*?*filter= rule in robots.txt and purged parameter URLs from the XML sitemap. Within 14 days, Googlebot re-allocated its crawl bandwidth back to primary product pages.
Deep Dive: “Crawled - Currently Not Indexed”
When a URL enters Crawled - currently not indexed, Googlebot successfully connected to your server, executed rendering, and read your page. It simply decided that the page was not worth storing in the index.
Primary Causes & Solutions:
- Thin or Auto-Generated Content: Pages with only 1–2 sentences, placeholder boilerplate, or purely scraped vendor descriptions.
- Fix: Merge thin pages into comprehensive pillar guides, or apply
<meta name="robots" content="noindex">to low-value landing pages.
- Fix: Merge thin pages into comprehensive pillar guides, or apply
- Near-Duplicate Content: Multiple landing pages created for slight geographic variations (e.g., “Plumber in Dallas”, “Plumber in Fort Worth”) with 95% identical text.
- Fix: Add unique localized value, customer testimonials, specific local pricing, or canonicalize to a regional hub.
- Low Information Gain: If 10 other websites already provide identical explanations with higher domain authority, Google’s Helpful Content System may discard your article to save index space.