The 429 Playbook: Handling Rate Limits Before Your IP Gets Banned
Every scraping tutorial shows the happy path. The page that matters is the one you get after an hour of collecting: 429 Too Many Requests.
Here's the playbook I run, ordered by what each response should cost you.
1. Read the headers first, back off second
Most servers tell you what to do in the response:
-
Retry-After: the contract. If it says 30, sleep 30. Not 5, not "I'll retry in a loop for a second". -
X-RateLimit-Remaining/X-RateLimit-Reset: your budget and its refill time. WhenRemaininghits 2, you throttle before the 429, not after. - No headers (common on free-tier APIs): treat every 429 as a warning shot. The next escalation is usually a temporary IP ban, and those are invisible — you get a 403 with a "blocked" page, not a rate-limit message.
2. The three-stage ladder
1st 429 → exponential backoff: 2^n * base (cap at 60s), 2 retries
2nd 429 → drop concurrency to 1, switch to a paced schedule (e.g. 1 req / 5s)
3rd 429 → stop. Persist the cursor, exit non-zero, let the scheduler retry later
The third stage is the one people skip. A tight retry loop that never gives up is how a free-tier API learns your IP address very well, very permanently. My collection jobs are resumable (cursor in a state file), so "stop and come back in an hour" is cheap — the pipeline picks up exactly where it died.
3. The failure mode nobody shows you
A 429 is not the only rate signal. Watch for:
- Silent partial responses: 200 OK with an empty list, where yesterday's identical call returned 50 items. Treat "200 + empty" as a soft 429 and re-verify against a known-good channel/endpoint before trusting the gap.
- Latency creep: responses slowing from 200 ms to 2 s is a rate limiter warming up. Throttle proactively.
-
Quality cliffs: the payload structure changes (
nullfields, truncated text) — some services degrade responses before they refuse them.
4. Budget math before you write code
Free tiers are rate-limited per IP, per key, per day. Do the arithmetic up front: 45 requests/min free tier × 60 min = 2,700 requests. At 200 items per page, a 20,000-message collection is 100 pages — fits easily on one channel, dies instantly at 30 channels. The playbook only works if your request budget covers the job; otherwise you're choosing between key rotation, sampling, or a paid tier today, not at 2 a.m. mid-collection.
The short version: 429s are a contract, not an obstacle. Read it, back off on schedule, persist your cursor, and know when stopping is the cheapest move in the pipeline.
Top comments (4)
The "200 OK with an empty list" failure mode is the one that would have bitten me. I always treated rate limiting as something that announces itself, but latency creep as a warm-up signal is a subtle and useful framing. Do you also persist the backoff state across process restarts, or is it only kept in memory for a single run?
The empty-list failure mode is the nastiest because it looks like success — we ended up asserting a minimum expected row count per window as a guard, which catches silent truncation instantly. Latency creep as the warm-up signal was the funniest lesson: the API warns you in whispers before it shouts in 429s.
Such a clean explanation of a nuanced problem. Bookmarking this for reference, thanks for sharing!
Thanks Jason! The recovery-behavior part was the trickiest to write — treating 429 as a signal to back off, not a failure to hammer. Glad it's bookmark-worthy.