DEV Community

Yuhe He
Yuhe He

Posted on

requests vs httpx for Scraping: When the Faster Client Stops Saving You Money

requests vs httpx for Scraping: When the Faster Client Stops Saving You Money

requests is the default Python HTTP client; httpx is the modern one that added async, HTTP/2, and a compatibility layer. The pitch is "httpx is strictly newer, migrate." The scraping reality is more boring and more useful: the client is rarely your bottleneck — and when it is, the savings have a ceiling worth pricing before you touch code.

What each one actually buys you

requests: the ecosystem default. Every scraping snippet, every library example, every Stack Overflow answer assumes it. Session objects give you keep-alive and cookie persistence, which covers 90% of polite collection. It is synchronous. That's not a bug in the library; it's a design contract: one request, one wait, one result.

httpx: requests-shaped API plus three real features —

  • async I/O: hundreds of in-flight requests from one process.
  • HTTP/2: multiplexed connections to servers that support it.
  • Proxy/transport hooks that are cleaner for pools and rotations.

The migration line is one import. The trap is thinking that line buys you speed.

Where the money actually goes

Model a collection run: N pages, per-page server latency L, network RTT, politeness delay P (what robots.txt and your conscience demand), and rate-limit ceiling R for the target.

  • If P and R dominate — polite collection against one site — concurrency is dead weight. The target says "one request per 3 seconds"; httpx with 200 async slots will happily queue 199 of them. requests with a sleep does the same job at the same throughput, with less code. The async upgrade costs you: event-loop debugging, async-infection through your whole pipeline, harder tracing. The savings: zero, because the bottleneck was never your client.
  • If you hit many independent sites — search engines, APIs, DNS-adjacent lookups, per-site spot checks — async pays. 200 targets × 500 ms: sequential = 100 seconds; async = under a second of wall time per batch. This is where httpx earns its migration.
  • HTTP/2 is situational. It helps when a target serves many small resources over one connection and the server supports it. Against a plain nginx with HTTP/1.1 enabled, you get parity with a version negotiation handshake.

The cost of the async switch, honestly

  1. Everything up and down the stack becomes async. BeautifulSoup parsing stays sync (fine), but your pipeline, scheduler, CLI — if you use asyncio, your functions become coroutines. For a script you maintain alone, that's a taste; for a team, it's a training cost.
  2. Debugging changes shape. Tracebacks through an event loop are a genre of their own. requests failures read like books; async failures read like crime scenes.
  3. Ecosystem drag. Older tutorials and helper libraries assume requests; you'll occasionally be the compatibility layer.

Decision table

Collection shape Right client Why
One site, polite pacing requests bottleneck is P and R, not the client
Hundreds of independent targets httpx async wall-time wins are real
HTTP/2 APIs you poll continuously httpx multiplexing cuts connection churn
Team codebase, mixed maintainers requests debugging and onboarding cost less
Drop-in modernization of existing code httpx (sync mode) free compatibility layer, no async tax

The clients are both free. The real ledger is wall-clock time you save against code complexity you add — do the arithmetic on your actual bottleneck first. For most polite, single-site scrapes, the answer is: you're already at the speed limit; the new car changes nothing.


The collector and dedupe pipeline behind my public-data runs is here; the free public-source field guide is here. Related: Scrapy vs BeautifulSoup, where each starts costing you.

Top comments (0)