DEV Community

137Foundry
137Foundry

Posted on

How to Implement Exponential Backoff Correctly for API Retries

Exponential backoff sounds simple enough that most engineers implement a version of it from memory, without checking it against what the pattern is actually supposed to do. That version usually has at least one of a few common bugs baked in. Here's how to build it correctly, step by step.

Step 1: Start with the actual formula, not an approximation

The core idea is that each retry waits longer than the last, typically doubling: attempt 1 waits 1 second, attempt 2 waits 2 seconds, attempt 3 waits 4 seconds, and so on. The formula is usually expressed as base_delay * (2 ^ attempt_number), capped at some maximum delay so it doesn't grow unbounded across many attempts.

delay = min(base_delay * (2 ** attempt_number), max_delay)
Enter fullscreen mode Exit fullscreen mode

The cap matters more than people initially think. Without one, a job on its tenth retry attempt might be waiting over 15 minutes for a single backoff interval, which is rarely what you actually want, even for a genuinely down dependency.

Step 2: Add jitter, or you'll create a thundering herd

Pure exponential backoff without randomization has a subtle failure mode: if a dependency goes down and a thousand jobs all start retrying on the same schedule, they'll all retry again at exactly the same moments, in synchronized waves, each wave hitting the dependency at once right as it might be starting to recover. This is the thundering herd problem, and it can prevent a recovering service from ever fully stabilizing because every retry wave re-overwhelms it.

Jitter fixes this by adding randomness to each delay:

delay = random_between(0, min(base_delay * (2 ** attempt_number), max_delay))
Enter fullscreen mode Exit fullscreen mode

This spreads retries out across a window instead of clustering them, which is a meaningfully better behavior for both your system and the dependency you're retrying against.

Step 3: Decide your retry limit deliberately, not by default

Most HTTP client libraries either don't retry by default or default to a small fixed number like 3. Whether that's correct for your use case depends on how time-sensitive the calling job is. A background data sync job can reasonably retry 5-8 times over several minutes. A user-facing request blocking a page load should probably cap out at 2-3 attempts within a couple of seconds, because past that point failing fast and showing the user an error is better than making them wait.

Step 4: Respect Retry-After headers when the API provides one

Many APIs, particularly ones with rate limiting, return a Retry-After header telling you exactly how long to wait before retrying. Your own calculated backoff delay should defer to this value when it's present and longer than what your formula would produce, since the API is telling you directly rather than you guessing.

if response.headers.get("Retry-After"):
    delay = max(calculated_delay, retry_after_seconds)
Enter fullscreen mode Exit fullscreen mode

Ignoring this header and retrying on your own schedule anyway is a common way to get more aggressively rate-limited, since it looks to the API like you're ignoring its explicit instruction.

Step 5: Make sure your retry loop actually distinguishes error types

This connects directly to the rest of your resilience strategy. A connection timeout or 503 should go through the backoff logic above. A 400 or 401 should not, retrying those just repeats a request that will fail identically every time, wasting the retry budget on something backoff can't fix. Filter which errors enter the retry loop at all before applying any of the delay logic above.

Step 6: Account for backoff in your overall job timeout budget

If a job has an overall deadline, a batch that needs to finish within an hour, for instance, your retry and backoff logic needs to respect that budget rather than operating independently of it. Five retries with backoff capped at 60 seconds each could theoretically consume five minutes on a single failing call, which might be fine for a background job with no deadline and genuinely problematic for one that's expected to complete quickly. Calculate the worst-case total delay your retry configuration could produce and check it against what the calling context can actually tolerate.

Step 7: Log each retry attempt with enough context to debug later

A retry that eventually succeeds is invisible in most logging setups, which sounds fine until you're trying to understand why a job took four minutes to complete when the actual work should have taken two seconds. Log the attempt number, the delay applied, and the error that triggered the retry, even on attempts that eventually succeed, so the full retry history is reconstructable after the fact.

"Backoff without jitter is one of those things that looks correct in a code review and only reveals the bug during an actual incident, when every retry wave lands on the dependency at once instead of spreading out." - Dennis Traina, [founder of 137Foundry](https://137foundry.com/services)

Putting it together

None of these six steps is individually hard, but skipping any one of them tends to reintroduce exactly the failure mode it was meant to prevent, usually discovered at the worst possible time. A correct implementation combines a capped exponential delay, jitter to avoid synchronized retry waves, a retry limit appropriate to the calling context, respect for any Retry-After header the API provides, error-type filtering so you're not retrying unrecoverable failures, and logging thorough enough to reconstruct what happened during an incident.

This pairs directly with the resilience work covered in a longer article on circuit breakers for flaky dependencies. Backoff handles individual retry timing. A circuit breaker handles the pattern of sustained failure across many calls. Most production-grade integrations need both working together, not one or the other.

References

The AWS Architecture Blog has a widely cited writeup on exponential backoff and jitter, including the specific thundering herd problem this article covers, from engineers who dealt with it at real scale. Google Cloud's API design guidance also covers retry behavior expectations from the API provider's side, useful context for understanding why headers like Retry-After exist and what a well-behaved client is expected to do with them. Martin Fowler's site covers the broader family of resilience patterns backoff belongs to, including how it pairs with a circuit breaker once retry logic alone isn't enough.

Exponential backoff with jitter isn't a complicated pattern to implement correctly once you know the pieces. 137Foundry treats it as one of the first things to get right in any pipeline that depends on external APIs, well before more advanced patterns like circuit breakers become necessary.

Most HTTP client libraries and job queue frameworks now ship a reasonable backoff-with-jitter implementation out of the box, so this often isn't custom code to write from scratch, it's a configuration decision: confirming the library's default is actually enabled, checking its cap and base delay against your specific latency tolerance, and verifying it respects Retry-After headers rather than assuming it does. Checking these three things takes a few minutes per integration and is worth doing explicitly rather than assuming whatever the library ships with is already correct for your use case, especially since some libraries' "default" behavior is actually no retry at all, which is easy to miss until the first outage exposes it and someone discovers the calling code was relying on a safety net that was never actually there in the first place, despite everyone assuming otherwise for months.

Top comments (0)