DEV Community

Cover image for My Search Console script timed out for 5 minutes. The API was fine — it was ignoring my proxy.
孙永瑞
孙永瑞

Posted on

My Search Console script timed out for 5 minutes. The API was fine — it was ignoring my proxy.

I spent an hour last night trying to answer one question: is it worth chasing a keyword where a competitor just dropped out of the top 20?

Answering it properly meant pulling my own Search Console data — impressions, clicks, average position, per query. I had a service account, I had credentials, I had written the request. It timed out. Every single time.

Not a fast failure either. It hung for about five minutes and then raised TimeoutError('timed out').

The trap: googleapiclient doesn't read your proxy env vars

Here's what I had:

from googleapiclient.discovery import build
service = build("searchconsole", "v1", credentials=creds)
service.searchanalytics().query(siteUrl=site, body=payload).execute()
Enter fullscreen mode Exit fullscreen mode

Textbook. And completely dead behind a proxy, because the default Google API client builds its transport on httplib2, and httplib2 does not pick up HTTPS_PROXY or https_proxy from the environment. Setting the variable felt like configuring something. It configured nothing. The library went straight out the front door, and the connection just sat there until it gave up.

What made this genuinely annoying is that the failure mode looks like an outage. Nothing in the error says "your transport is misconfigured." Five minutes of silence reads like "the API is slow today," so my first instinct was to retry, then to suspect the service account, then to suspect the quota.

The fix: AuthorizedSession and a plain REST call

Skip the discovery client and use Google's own auth transport on top of requests, which does honour proxy env vars:

from google.oauth2 import service_account
from google.auth.transport.requests import AuthorizedSession

creds = service_account.Credentials.from_service_account_file(
    SA_PATH, scopes=["https://www.googleapis.com/auth/webmasters.readonly"]
)
session = AuthorizedSession(creds)

r = session.post(
    f"https://searchconsole.googleapis.com/webmasters/v3/sites/{site}/searchAnalytics/query",
    json={"startDate": start, "endDate": end, "dimensions": ["query"], "rowLimit": 250},
    timeout=60,
)
rows = r.json().get("rows", [])
Enter fullscreen mode Exit fullscreen mode

Then run it with the proxy actually set:

HTTPS_PROXY=http://127.0.0.1:7897 python3 gsc_page_perf.py
Enter fullscreen mode Exit fullscreen mode

That worked on the first attempt. Same credentials, same endpoint, same payload — the only thing that changed was the HTTP layer underneath.

The general lesson I keep relearning: when a client library hangs, suspect the transport before the service. The library is a wrapper someone else wrote around HTTP, and wrappers are where your environment quietly disappears.

Trap two: the data is always three days late

My first successful run returned a wall of zeros for the most recent dates. Nothing was broken. Search Console data has a fixed delay of roughly three days, so if you set endDate to today, the tail of your series is empty and your averages get dragged toward zero.

Set endDate = today - 3 days. My real window ended up being 2026-08-28 to 2026-09-24. This is documented, and I still lost ten minutes to it, because a response full of zeros looks exactly like a query that matched nothing.

Trap three: service accounts are scoped per property

I assumed one service account could read all six of my sites. sites.list returned exactly one property. Delegation isn't inherited across properties — each site has to grant access to that service account individually. The other five simply weren't there, and "not authorized" and "no data" are indistinguishable in an empty response.

Trap four: the token exchange is a separate request, and it can be blocked too

Two days after the fix above worked, I ran the same script and got nothing. Not a timeout this time — an error:

TransportError: HTTPSConnectionPool(host='oauth2.googleapis.com', port=443):
Max retries exceeded with url: /token
(Caused by ProxyError('Unable to connect to proxy',
OSError('Tunnel connection failed: 502 Bad Gateway')))
Enter fullscreen mode Exit fullscreen mode

All three queries came back empty. Same machine, same proxy, same credentials.

The detail that mattered: AuthorizedSession makes two network calls, not one. First it exchanges your service-account JWT for an access token at oauth2.googleapis.com, then it calls the API. I'd only been thinking about the second one.

So I tested the first one directly:

curl -s -o /dev/null -w "%{http_code}" --max-time 10 \
  -x http://127.0.0.1:7897 https://oauth2.googleapis.com/
# 404
Enter fullscreen mode Exit fullscreen mode

A 404 means the proxy connected fine and the host answered. requests couldn't reach the same host that curl reached through the same proxy, six seconds apart. That asymmetry is the whole diagnosis: it isn't the network, and it isn't the credentials — it's this HTTP client talking to this endpoint.

The fix is to stop asking Python to do the token exchange. Sign the JWT locally, hand it to curl, and use curl for both hops:

import jwt, json, subprocess, time

sa = json.load(open(SA_PATH))
now = int(time.time())
assertion = jwt.encode({
    "iss": sa["client_email"],
    "scope": "https://www.googleapis.com/auth/webmasters.readonly",
    "aud": "https://oauth2.googleapis.com/token",
    "iat": now, "exp": now + 3600,
}, sa["private_key"], algorithm="RS256")

out = subprocess.run([
    "curl", "-s", "--max-time", "60", "-x", "http://127.0.0.1:7897",
    "-X", "POST", "https://oauth2.googleapis.com/token",
    "-d", "grant_type=urn:ietf:params:oauth:grant-type:jwt-bearer",
    "--data-urlencode", f"assertion={assertion}",
], capture_output=True, text=True)

token = json.loads(out.stdout)["access_token"]
Enter fullscreen mode Exit fullscreen mode

Then POST to searchAnalytics/query with -H "Authorization: Bearer {token}". Zero errors, all queries returned data on the next run.

So, was the keyword worth chasing?

No — and the data said it plainly. I was curious about a head term where a competing domain had just fallen out of the top 20. Tempting. But the page targeting that space had 0 clicks, 10 impressions and an average position of 36.8 over 28 days, and the four queries it actually matched were all "alternatives" phrasing, not the head term itself.

Site-wide: 1 click, 117 impressions, 0.855% CTR, average position 62.73.

The useful finding was the one I wasn't looking for: a different query sitting at position 11 — one place off the first page, with real impressions behind it. That's the one worth pushing. Not the exciting gap in a competitor's rankings, but the boring term that's already close.

Two hours of tooling pain to learn "don't chase the shiny keyword." I'll take it — I'd have spent two weeks writing content for a term I had zero exposure on.

Update, three days later: the boring keyword fell out of the top 20

Remember the query I said was worth pushing, sitting at position 11? Three days after I wrote that, a rank-tracking alert told me it had dropped out of the top 20 entirely. I checked by hand. The page was gone from the first two pages — not demoted a few places, just gone.

Here's the part I got wrong. I hadn't been ignoring it. I'd re-pulled the Search Console data and it still looked healthy: same position, same impressions. I'd even written a note to myself that the drop wasn't worth acting on.

Search Console's position is an average over your date window, not where the page ranks today. That query had six impressions across 28 days, and they clustered in the first few days of the window. Those six pulled the average to a respectable-looking number and held it there. I re-ran it at 14 days and 7 days and got identical figures — which I read as confirmation, when what it actually meant was that there had been no impressions at all recently.

No impressions doesn't mean no change. It means the page stopped appearing often enough to register. The average can't show you that, because a period with no data contributes nothing to it.

Two lessons, and the second one is the expensive one:

  1. An average answers "how did this do over the window?" — never "where is it now?" If you need to know whether you're ranking today, use something that checks today.
  2. The alert that caught this wasn't Search Console. GSC's three-day delay means it literally cannot tell you about yesterday. The thing that caught it was a daily rank check from my own tool — slower to build, and the only reason I found out in hours instead of a week.

The page itself was fine, incidentally: HTTP 200, still in the sitemap, still allowed in robots.txt. Not a technical failure. Just a ranking that moved while my average sat still.


If you want your own Search Console data pulled apart properly — real numbers, not vibes — I do that: https://yongrui-services.pages.dev/

Top comments (0)