DEV Community

Yuhe He
Yuhe He

Posted on

The 429 Playbook: Handling Rate Limits Before Your IP Gets Banned

The 429 Playbook: Handling Rate Limits Before Your IP Gets Banned

Every scraping tutorial shows the happy path. The page that matters is the one you get after an hour of collecting: 429 Too Many Requests.

Here's the playbook I run, ordered by what each response should cost you.

1. Read the headers first, back off second

Most servers tell you what to do in the response:

  • Retry-After: the contract. If it says 30, sleep 30. Not 5, not "I'll retry in a loop for a second".
  • X-RateLimit-Remaining / X-RateLimit-Reset: your budget and its refill time. When Remaining hits 2, you throttle before the 429, not after.
  • No headers (common on free-tier APIs): treat every 429 as a warning shot. The next escalation is usually a temporary IP ban, and those are invisible — you get a 403 with a "blocked" page, not a rate-limit message.

2. The three-stage ladder

1st 429  → exponential backoff: 2^n * base (cap at 60s), 2 retries
2nd 429  → drop concurrency to 1, switch to a paced schedule (e.g. 1 req / 5s)
3rd 429  → stop. Persist the cursor, exit non-zero, let the scheduler retry later
Enter fullscreen mode Exit fullscreen mode

The third stage is the one people skip. A tight retry loop that never gives up is how a free-tier API learns your IP address very well, very permanently. My collection jobs are resumable (cursor in a state file), so "stop and come back in an hour" is cheap — the pipeline picks up exactly where it died.

3. The failure mode nobody shows you

A 429 is not the only rate signal. Watch for:

  • Silent partial responses: 200 OK with an empty list, where yesterday's identical call returned 50 items. Treat "200 + empty" as a soft 429 and re-verify against a known-good channel/endpoint before trusting the gap.
  • Latency creep: responses slowing from 200 ms to 2 s is a rate limiter warming up. Throttle proactively.
  • Quality cliffs: the payload structure changes (null fields, truncated text) — some services degrade responses before they refuse them.

4. Budget math before you write code

Free tiers are rate-limited per IP, per key, per day. Do the arithmetic up front: 45 requests/min free tier × 60 min = 2,700 requests. At 200 items per page, a 20,000-message collection is 100 pages — fits easily on one channel, dies instantly at 30 channels. The playbook only works if your request budget covers the job; otherwise you're choosing between key rotation, sampling, or a paid tier today, not at 2 a.m. mid-collection.

The short version: 429s are a contract, not an obstacle. Read it, back off on schedule, persist your cursor, and know when stopping is the cheapest move in the pipeline.

Top comments (4)

Collapse
 
sgaggjhkjh profile image
Arjun Sharma •

The "200 OK with an empty list" failure mode is the one that would have bitten me. I always treated rate limiting as something that announces itself, but latency creep as a warm-up signal is a subtle and useful framing. Do you also persist the backoff state across process restarts, or is it only kept in memory for a single run?

Collapse
 
yuhehe profile image
Yuhe He •

The empty-list failure mode is the nastiest because it looks like success — we ended up asserting a minimum expected row count per window as a guard, which catches silent truncation instantly. Latency creep as the warm-up signal was the funniest lesson: the API warns you in whispers before it shouts in 429s.

Collapse
 
dev_in_the_fog profile image
Jason Y. (dev_in_the_fog) •

Such a clean explanation of a nuanced problem. Bookmarking this for reference, thanks for sharing!

Collapse
 
yuhehe profile image
Yuhe He •

Thanks Jason! The recovery-behavior part was the trickiest to write — treating 429 as a signal to back off, not a failure to hammer. Glad it's bookmark-worthy.