Proxy plans are sold per GB, but scrapers think in pages. So when someone asks "how many GB do I need?", the usual answer is "it depends". That's true, and it doesn't help anyone pick a plan.
So I measured it. 34 public pages, fetched three ways, with every byte counted at the tunnel in both directions.
The short answer
| How you fetch | Median per page | Pages per GB |
|---|---|---|
| HTTP client, HTML only | 22.7 KB | 44,083 |
| Playwright, full load | 1.10 MB | 906 |
| Playwright, images/media/fonts blocked | 454.6 KB | 2,199 |
Same pages, same day. At the median, the full browser load used about 49 times the traffic of a plain HTML fetch.
How I measured
- 34 public pages: docs, news articles, blog posts, forum threads, shop category and product pages, and a few JavaScript apps. Only pages robots.txt allows, nothing behind a login.
- Three modes: Requests 2.34.2 (one GET, compression on), Playwright 1.63.0 full load, and Playwright with images, media and fonts blocked.
- Three runs per page, fresh process and empty cache every time. 306 loads, all HTTP 200.
- All traffic went through scrapescope, a small local proxy I built that counts bytes per tunnel. So the numbers include TLS, headers and upload, not just page weight.
- October 5, 2026, from Helsinki, direct connection.
Three things that surprised me
1. Scripts, not images. In full loads, scripts were 49% of the bytes. Images, media and fonts together were 35%. Blocking them saved 28% at the median, less than I expected. On news pages the blocked loads were actually heavier (2.24 MB vs 1.32 MB median), because pages behave differently when things fail to load.
2. Half the bytes go to other companies. 49.9% of full-load bytes went to third-party hosts. Google Tag Manager alone was on 17 of the 34 pages at about 308 KB per load. That's 12% of all full-load bytes, for analytics your scraper will never use.
3. Compression does a lot of quiet work. All 34 HTML responses were compressed (17 Brotli, 14 gzip, 3 zstd). The median HTML document was 83.5 KB decoded but 22.7 KB on the wire. A client that doesn't ask for compression pays for the bigger number.
Upload counts too
Upstream was 6.3% of the traffic for HTML-only fetches and about 3% for browser loads. Small, but if your provider bills both directions, it's on your bill.
Estimate your own
GB per day = pages per day × bytes per page × attempts per page ÷ 1,000,000,000
50,000 pages a day with a full browser (1.10 MB median, 1.2 attempts per page) comes to about 66 GB a day. The same job with an HTTP client is about 1.4 GB.
Your pages won't match mine, so treat these as a starting point. Better still, run your own job through scrapescope for a day and use your numbers.
Limits
One location, one day, 34 pages. Direct connection, so no proxy exit location, blocks or retries. CONNECT bytes are estimated. Pages change.
The full write-up has the breakdown by page type, cost per 1,000 pages at different per-GB prices, the raw CSV of all 306 runs and the scripts: How many GB of proxies do you need?
scrapescope is open source: github.com/ipvolt/scrapescope
Disclosure: I'm building ipvolt, a proxy service that bills per GB, which is why I wanted real numbers. scrapescope works with any provider.
What do your scrapers average per page? I'm curious whether anyone sees numbers very different from these.
Top comments (0)