requests vs httpx for Scraping: When the Faster Client Stops Saving You Money
requests is the default Python HTTP client; httpx is the modern one that added async, HTTP/2, and a compatibility layer. The pitch is "httpx is strictly newer, migrate." The scraping reality is more boring and more useful: the client is rarely your bottleneck — and when it is, the savings have a ceiling worth pricing before you touch code.
What each one actually buys you
requests: the ecosystem default. Every scraping snippet, every library example, every Stack Overflow answer assumes it. Session objects give you keep-alive and cookie persistence, which covers 90% of polite collection. It is synchronous. That's not a bug in the library; it's a design contract: one request, one wait, one result.
httpx: requests-shaped API plus three real features —
- async I/O: hundreds of in-flight requests from one process.
- HTTP/2: multiplexed connections to servers that support it.
- Proxy/transport hooks that are cleaner for pools and rotations.
The migration line is one import. The trap is thinking that line buys you speed.
Where the money actually goes
Model a collection run: N pages, per-page server latency L, network RTT, politeness delay P (what robots.txt and your conscience demand), and rate-limit ceiling R for the target.
-
If P and R dominate — polite collection against one site — concurrency is dead weight. The target says "one request per 3 seconds";
httpxwith 200 async slots will happily queue 199 of them. requests with asleepdoes the same job at the same throughput, with less code. The async upgrade costs you: event-loop debugging,async-infection through your whole pipeline, harder tracing. The savings: zero, because the bottleneck was never your client. - If you hit many independent sites — search engines, APIs, DNS-adjacent lookups, per-site spot checks — async pays. 200 targets × 500 ms: sequential = 100 seconds; async = under a second of wall time per batch. This is where httpx earns its migration.
- HTTP/2 is situational. It helps when a target serves many small resources over one connection and the server supports it. Against a plain nginx with HTTP/1.1 enabled, you get parity with a version negotiation handshake.
The cost of the async switch, honestly
-
Everything up and down the stack becomes async. BeautifulSoup parsing stays sync (fine), but your pipeline, scheduler, CLI — if you use
asyncio, your functions become coroutines. For a script you maintain alone, that's a taste; for a team, it's a training cost. -
Debugging changes shape. Tracebacks through an event loop are a genre of their own.
requestsfailures read like books; async failures read like crime scenes. - Ecosystem drag. Older tutorials and helper libraries assume requests; you'll occasionally be the compatibility layer.
Decision table
| Collection shape | Right client | Why |
|---|---|---|
| One site, polite pacing | requests | bottleneck is P and R, not the client |
| Hundreds of independent targets | httpx async | wall-time wins are real |
| HTTP/2 APIs you poll continuously | httpx | multiplexing cuts connection churn |
| Team codebase, mixed maintainers | requests | debugging and onboarding cost less |
| Drop-in modernization of existing code | httpx (sync mode) | free compatibility layer, no async tax |
The clients are both free. The real ledger is wall-clock time you save against code complexity you add — do the arithmetic on your actual bottleneck first. For most polite, single-site scrapes, the answer is: you're already at the speed limit; the new car changes nothing.
The collector and dedupe pipeline behind my public-data runs is here; the free public-source field guide is here. Related: Scrapy vs BeautifulSoup, where each starts costing you.
Top comments (0)