<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alexandr Kazmin</title>
    <description>The latest articles on DEV Community by Alexandr Kazmin (@alkaznodemaven).</description>
    <link>https://dev.to/alkaznodemaven</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166105%2F83835fb5-dbb7-4de8-928f-57b9d45b4393.jpg</url>
      <title>DEV Community: Alexandr Kazmin</title>
      <link>https://dev.to/alkaznodemaven</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alkaznodemaven"/>
    <language>en</language>
    <item>
      <title>A cold proxy exit gets Google results 20% of the time. A warmed one gets 87%</title>
      <dc:creator>Alexandr Kazmin</dc:creator>
      <pubDate>Tue, 06 Oct 2026 10:54:43 +0000</pubDate>
      <link>https://dev.to/alkaznodemaven/a-cold-proxy-exit-gets-google-results-20-of-the-time-a-warmed-one-gets-87-3el8</link>
      <guid>https://dev.to/alkaznodemaven/a-cold-proxy-exit-gets-google-results-20-of-the-time-a-warmed-one-gets-87-3el8</guid>
      <description>&lt;p&gt;You rotate to a fresh residential IP, open Google in a real browser, type a query - and get a captcha. Again. We assumed that was the proxy. We measured it, and it mostly wasn't: what mattered most was whether that exit had done anything before the search.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;Every attempt below opens google.com in a real, headful browser (Chromium 151&lt;br&gt;
through Patchright), types a query into the search box and checks whether Google&lt;br&gt;
served results or a captcha. Each attempt gets a fresh sticky session on a&lt;br&gt;
residential pool, so each one starts from a new exit.&lt;/p&gt;

&lt;p&gt;The cold arm searches straight away. The warm arm first browses six ordinary&lt;br&gt;
pages on the same session, then searches.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;served&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;cold&lt;/td&gt;
&lt;td&gt;20% (46 of 232)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;warm, six pages&lt;/td&gt;
&lt;td&gt;87% (409 of 471)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is four runs on two machines. Inside every run the cold and warm arms are&lt;br&gt;
interleaved, because the hour of the day is one of the biggest effects we have&lt;br&gt;
measured, and running the two arms an hour apart would compare the hours.&lt;/p&gt;

&lt;p&gt;One run tried more depths side by side:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;warm-up depth&lt;/th&gt;
&lt;th&gt;0&lt;/th&gt;
&lt;th&gt;2&lt;/th&gt;
&lt;th&gt;4&lt;/th&gt;
&lt;th&gt;7&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;served&lt;/td&gt;
&lt;td&gt;11%&lt;/td&gt;
&lt;td&gt;24%&lt;/td&gt;
&lt;td&gt;33%&lt;/td&gt;
&lt;td&gt;82%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Depth 7 is the six-page warm-up above. A separate test of a single page of&lt;br&gt;
warm-up moved nothing (30% against 32%). The effect builds with depth, and the&lt;br&gt;
big jump is at the deepest rung.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we don't know
&lt;/h2&gt;

&lt;p&gt;Four of the six warm-up pages are Google's own. So two explanations fit every&lt;br&gt;
row: Google's infrastructure has already seen this exit behave normally, or the&lt;br&gt;
browser has simply lived through several navigations. The run that separates&lt;br&gt;
them - the same depth with third-party pages only - has not been done yet. This&lt;br&gt;
is an effect without a mechanism, and I'd rather say so than guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  The machine matters too
&lt;/h2&gt;

&lt;p&gt;The same code, the same gateway and the same parameters, in overlapping hours:&lt;br&gt;
a Windows workstation was served 39% (24 of 61), a Linux VPS 0% (0 of 84). We&lt;br&gt;
haven't isolated which property of the Linux host Google reads. If your scraper&lt;br&gt;
works on your laptop and dies on a server, this is a candidate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Warm-up is not free. In a smoke run of six queries, warming one exit took about&lt;br&gt;
three minutes and 34 MB of proxy traffic. After that, each results page took&lt;br&gt;
17-21 seconds and 0.2-1.6 MB. So the strategy pays only if you keep a warm exit&lt;br&gt;
for several queries instead of burning a new one each time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this.
&lt;/h2&gt;

&lt;p&gt;Don't send a search from a cold exit. Give it a few ordinary pages first, keep the sticky session, and reuse a warm exit for as many queries as Google keeps serving it. Rotating to a new IP on every request is the most expensive way to get captchas.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;Everything is public: the harness, the run files and the scripts that turn rows&lt;br&gt;
into the tables above.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Harness and raw rows: &lt;a href="https://github.com/nodemaven/proxy-benchmark" rel="noopener noreferrer"&gt;https://github.com/nodemaven/proxy-benchmark&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The scraper we built on it: &lt;a href="https://github.com/nodemaven/google-browser-scraper" rel="noopener noreferrer"&gt;https://github.com/nodemaven/google-browser-scraper&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;google-browser-scraper
patchright &lt;span class="nb"&gt;install &lt;/span&gt;chromium
google-browser-scraper search &lt;span class="s2"&gt;"best running shoes"&lt;/span&gt; &lt;span class="nt"&gt;--proxy&lt;/span&gt; &lt;span class="s2"&gt;"http://USER-session-{session}:PASS@your-gateway:port"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;{session}&lt;/code&gt; is where your provider expects a session id. It works with any&lt;br&gt;
provider; I work at one (NodeMaven), which is why we had the pool to measure&lt;br&gt;
this on. If you re-run it and get different numbers, I'd happy to see them.&lt;/p&gt;

</description>
      <category>webscraping</category>
      <category>python</category>
      <category>playwright</category>
      <category>proxies</category>
    </item>
  </channel>
</rss>
