Troubleshooting Guide Crawl Queue & Scheduling

Resolving "Discovered – Currently Not Indexed"

When Googlebot knows your URL exists but has placed it in a pending queue and has not yet fetched it. Here is how to diagnose crawl queue bottlenecks and unstick your URLs.

1. What "Discovered – Currently Not Indexed" Means

Official Google documentation describes this status as follows:

"The page was found by Google, but not crawled yet. Typically, Google wanted to crawl the URL but the site was overloaded; therefore Google rescheduled the crawl. This is why the last crawl date is empty on the report."
— Google Search Console Help

The Key Technical Distinction: Unlike Crawled – currently not indexed, Googlebot has not yet fetched the page. Google discovered the URL from an XML sitemap, an internal link, or an external backlink, added it to the crawl queue, and subsequently deferred the crawl due to server capacity limits or low initial crawl priority.

2. Root Causes: Why Google Defers Crawling

A. Host Latency & Crawl Capacity Limit

If your server response times (TTFB) rise or return sporadic 5xx gateway errors during crawling, Googlebot automatically throttles down crawl rate to protect your infrastructure.

B. Massive Programmatic URL Ingestion

Publishing 100,000+ new programmatic URLs overnight on a domain with modest crawl demand overwhelms Googlebot's scheduling queue.

C. Deep Architecture & Click Depth (>3 Hops)

If URLs are discovered solely via sitemaps and are buried 4 to 6 clicks away from the homepage, Googlebot assigns them lower crawl scheduling priority.

D. Fresh Domain Authority Ramp-up

Brand new websites naturally experience a conservative crawl schedule as Googlebot gradually assesses site stability, reliability, and content freshness.

3. Technical Diagnostic Protocol

  1. Step 1: Check the Crawl Stats Report

    Navigate to Settings → Crawl stats → Open report. Check Host status (Robots.txt fetch, DNS resolution, Server connectivity). If server connectivity shows degradation, your host is throttling Googlebot.

  2. Step 2: Inspect Average Response Time

    In Crawl Stats, inspect the "Average response time (ms)" chart. If latency spikes above 500–1,000ms, Googlebot will decrease its crawl rate to prevent taking down your server.

  3. Step 3: Audit Discovery Source

    In URL Inspection, inspect the Discovery card. If Referring page is "None detected" and Sitemaps is your only listed source, your internal linking architecture is failing to surface this URL.

4. Engineering Remediation Playbook

  • Optimize Server Response Latency: Implement edge caching (Cloudflare, Fastly), optimize database queries, and ensure HTML documents return with sub-200ms TTFB.
  • Promote URLs in Internal Navigation: Place prominent contextual links on high-traffic parent pages (categories, recent articles, hub landing pages) to signal priority to Googlebot's scheduler.
  • Prune Low-Priority Sitemap Feeds: Remove expired, out-of-stock, or low-utility pages from your XML sitemaps to prevent dilute crawl queues.
  • Flatten Site Architecture: Ensure critical pages are reachable within 2 to 3 hops from the domain root.

5. Comparison: Discovered vs. Crawled

Dimension Discovered – currently not indexed Crawled – currently not indexed
Googlebot Action Found URL, but has NOT fetched it yet Fetched & rendered HTML, but decided not to index
Primary Bottleneck Crawl capacity, queue backlog, server latency Content quality, uniqueness, domain quality threshold
Primary Fix Improve TTFB, increase internal links, prune sitemaps Improve content substance, consolidate near-duplicates