You rotate to a fresh residential IP, open Google in a real browser, type a query - and get a captcha. Again. We assumed that was the proxy. We measured it, and it mostly wasn't: what mattered most was whether that exit had done anything before the search.
The numbers
Every attempt below opens google.com in a real, headful browser (Chromium 151
through Patchright), types a query into the search box and checks whether Google
served results or a captcha. Each attempt gets a fresh sticky session on a
residential pool, so each one starts from a new exit.
The cold arm searches straight away. The warm arm first browses six ordinary
pages on the same session, then searches.
| served | |
|---|---|
| cold | 20% (46 of 232) |
| warm, six pages | 87% (409 of 471) |
That is four runs on two machines. Inside every run the cold and warm arms are
interleaved, because the hour of the day is one of the biggest effects we have
measured, and running the two arms an hour apart would compare the hours.
One run tried more depths side by side:
| warm-up depth | 0 | 2 | 4 | 7 |
|---|---|---|---|---|
| served | 11% | 24% | 33% | 82% |
Depth 7 is the six-page warm-up above. A separate test of a single page of
warm-up moved nothing (30% against 32%). The effect builds with depth, and the
big jump is at the deepest rung.
What we don't know
Four of the six warm-up pages are Google's own. So two explanations fit every
row: Google's infrastructure has already seen this exit behave normally, or the
browser has simply lived through several navigations. The run that separates
them - the same depth with third-party pages only - has not been done yet. This
is an effect without a mechanism, and I'd rather say so than guess.
The machine matters too
The same code, the same gateway and the same parameters, in overlapping hours:
a Windows workstation was served 39% (24 of 61), a Linux VPS 0% (0 of 84). We
haven't isolated which property of the Linux host Google reads. If your scraper
works on your laptop and dies on a server, this is a candidate.
What it costs
Warm-up is not free. In a smoke run of six queries, warming one exit took about
three minutes and 34 MB of proxy traffic. After that, each results page took
17-21 seconds and 0.2-1.6 MB. So the strategy pays only if you keep a warm exit
for several queries instead of burning a new one each time.
What to do with this.
Don't send a search from a cold exit. Give it a few ordinary pages first, keep the sticky session, and reuse a warm exit for as many queries as Google keeps serving it. Rotating to a new IP on every request is the most expensive way to get captchas.
Reproduce it
Everything is public: the harness, the run files and the scripts that turn rows
into the tables above.
- Harness and raw rows: https://github.com/nodemaven/proxy-benchmark
- The scraper we built on it: https://github.com/nodemaven/google-browser-scraper
pip install google-browser-scraper
patchright install chromium
google-browser-scraper search "best running shoes" --proxy "http://USER-session-{session}:PASS@your-gateway:port"
{session} is where your provider expects a session id. It works with any
provider; I work at one (NodeMaven), which is why we had the pool to measure
this on. If you re-run it and get different numbers, I'd happy to see them.
Top comments (0)