There's a specific kind of frustration that comes from deploying a scraper that worked perfectly on your laptop, only to watch it fail on every request in production. Same code. Same target. Different result.
I've hit this wall more than once. Eventually I stopped treating it as a mystery and started treating it as a testable problem. Now I run a short checklist before any scraper goes to a server.
1. Does your TLS fingerprint match a real browser?
This is the one most developers have never heard of, and it's often why a scraper fails immediately on a server.
Every HTTPS connection starts with a handshake that sends a specific set of ciphers, extensions, and curves. That combination is a fingerprint — it identifies the library making the request. Python's default urllib3 produces a fingerprint that looks nothing like Chrome or Firefox.
If your headers claim you're Chrome but your TLS handshake says Python, that mismatch is an immediate red flag.
How to test it: Go to tls.browserleaks.com in your browser and note the JA3 fingerprint. Then run the same check from your scraper. If they differ significantly, you have a problem.
The fix is usually a client that can impersonate a real browser's TLS signature, like curl_cffi or tls_client. It's a one-line change that eliminates a whole category of blocks.
2. Are your HTTP/2 settings consistent with a browser?
If you're using HTTP/2 — and most modern scrapers are — the settings frame you send during connection setup is another fingerprint.
Real browsers send specific values for window sizes, max concurrent streams, and header table size. HTTP client libraries often send different values or omit fields entirely.
How to test it: Does your scraper get blocked on sites that work fine in a browser with the same headers? If headers match but requests still fail, HTTP/2 settings are a likely cause.
The fix is the same: use a client that handles this for you. curl_cffi and Playwright both match real browsers out of the box.
3. Does your session behavior look human?
Individual requests can look perfect and still get blocked if the pattern across requests is wrong.
The two most common mistakes:
Rotating on every request for a multi-step workflow. If your scraper loads a category page, then a product page, then a review page — and each request comes from a different IP — the session looks like three unrelated strangers. Sticky sessions fix this.
Firing requests too fast or too uniformly. A script that hits a page every 200ms like clockwork looks automated regardless of how clean each request is.
How to test it: Log the IP used for each request in a multi-step workflow. If it changes mid-flow, you have a session consistency problem. Then add jitter to your timing if intervals are perfectly uniform.
4. Have you tested the target's actual threshold?
This is the one people skip most. You don't know where the line is until you find it.
Run a small test against the real target before deploying at full speed. Start with 10 requests from a single IP, then 50, then 100. Watch when the first block appears and what kind it is — 403, CAPTCHA, silent drop, rate-limit header.
How to test it: Run the test from your production environment, not your laptop. The IP type, network path, and library versions should all match what production will use.
5. Do you know why a request failed?
Less a test, more a habit — but the one that saves the most time long-term.
When a request fails, most scrapers immediately rotate the IP and retry. That's often wrong, because it treats every failure as an IP problem. If the real cause was a session inconsistency or a TLS mismatch, rotating just masks the symptom.
A simple pattern: before any retry logic runs, log what actually failed. A 403, a timeout, an empty response, and a CAPTCHA each point to a different root cause — and only one is solved by rotating.
A pre-deployment checklist
- [ ] TLS fingerprint checked against a real browser
- [ ] HTTP/2 settings confirmed to match
- [ ] Session consistency tested on a multi-step workflow
- [ ] Request timing has natural variation
- [ ] Target's threshold tested from production
- [ ] Failure logging distinguishes between block types
It takes maybe 30 minutes. It has saved me days.
Most scraper failures in production aren't caused by bad code. They're caused by the gap between where you developed and where you deployed. Testing that gap directly is the whole point.
Top comments (0)