Escaping the Web Interface Constraints
While the Google Search Console web UI is convenient for day-to-day inspections, it enforces strict operational constraints:
- UI Table Export Limit: Exactly 1,000 rows of queries or pages.
- Manual Repetition: Exporting data every week requires human intervention.
- Data Expiration: Data drops off after 16 months.
The Google Search Console API (Search Analytics API) allows engineers to extract up to 25,000 rows per request, paginate through millions of query combinations, and schedule automated nightly ETL pipelines directly into company data lakes.
Authentication: Service Account vs. OAuth 2.0
For automated, headless backend cron jobs and server pipelines, a Google Cloud Service Account is the industry standard:
┌────────────────────────────────────────────────────────────────────────┐
│ SERVICE ACCOUNT SETUP PROTOCOL │
│ │
│ 1. Create a Project in Google Cloud Console │
│ 2. Enable "Google Search Console API" │
│ 3. Create a Service Account (e.g. [email protected]) │
│ 4. Generate and download JSON Private Key Credentials │
│ 5. In GSC UI (Settings > Users), add the service account email as │
│ a "Full User" on the verified property! │
└────────────────────────────────────────────────────────────────────────┘
Production Python Script: Search Analytics Extraction
Install the official client libraries:
pip install google-api-python-client google-auth
Here is a complete, production-ready Python script to query GSC with pagination:
import datetime
from google.oauth2 import service_account
from googleapiclient.discovery import build
KEY_FILE_LOCATION = "service-account-credentials.json"
SITE_URL = "https://example.com/" # Or "sc-domain:example.com" for Domain property
def get_search_console_service():
credentials = service_account.Credentials.from_service_account_file(
KEY_FILE_LOCATION,
scopes=["https://www.googleapis.com/auth/webmasters.readonly"]
)
return build("searchconsole", "v1", credentials=credentials)
def fetch_gsc_performance(start_date, end_date, row_limit=25000):
service = get_search_console_service()
request_body = {
"startDate": start_date,
"endDate": end_date,
"dimensions": ["query", "page", "country", "device"],
"rowLimit": row_limit,
"startRow": 0,
"dataState": "final" # Exclude volatile hourly data
}
response = service.searchanalytics().query(
siteUrl=SITE_URL,
body=request_body
).execute()
rows = response.get("rows", [])
print(f"Successfully retrieved {len(rows)} rows of search performance data.")
return rows
if __name__ == "__main__":
today = datetime.date.today()
start = (today - datetime.timedelta(days=7)).strftime("%Y-%m-%d")
end = (today - datetime.timedelta(days=3)).strftime("%Y-%m-%d")
data = fetch_gsc_performance(start, end)
Bypassing the 25,000-Row Single Request Limit
If your website receives queries across millions of pages, a single API call of 25,000 rows will only capture a fraction of your dataset.
The Chunking & Pagination Strategy:
- Iterate by Single Date: Instead of querying a 30-day window at once, query day-by-day (
start_date == end_date). - Loop by
startRow: Paginate requests usingstartRow = 0, 25000, 50000...until the API returns zero rows. - Partition by Device or Country: Filter by
country: USA,country: GBR, etc., to unlock deeper query layers that would otherwise be truncated by the top 25k ceiling.