Picture this. It is 2:14 in the morning and your checkout service starts returning errors. Every request fails for about twenty seconds, then your deploy finishes rolling back and the error rate drops to zero. You exhale.
Then the dashboards spike again. Harder than before. Nothing changed on your end. The service was healthy. So what knocked it over the second time?
Your own services did. Every client that failed during those twenty seconds had been retrying, and the moment the service came back, they all arrived at once. The recovery got flattened by a second wave made entirely of retries. That second wave has a name engineers have used for decades: the thundering herd.
The naive retry
The first retry loop most of us write looks like this: call the service, and if it fails, try again. Immediately. Maybe three times.
For one client, that is harmless. But failures in distributed systems are rarely lonely. A deploy, a network blip, a database failing over: these knock out hundreds or thousands of clients in the same second. Now all of them retry in the same second too. The service comes back to life and instantly receives ten times its normal traffic, all at once, and falls over again. Your retry logic just turned a twenty second outage into a rolling one.
(A quick term check: a distributed system is just a bunch of services talking to each other over a network instead of one big program. The network between them is the part that fails in interesting ways.)
Fixed delays do not break the herd
The obvious fix is to wait between retries. Retry every five seconds, say. That helps the service, but the clients are still marching in lockstep. They all failed together, so they all wait five seconds together, and they all retry together. The herd is still a herd. It is just a slower one.
Exponential backoff spreads them out
The next step is to make each wait longer than the last: one second, then two, then four, then eight. This is exponential backoff, and it does real work. A short blip only costs your clients a second or two, while a long outage backs everyone further and further off, giving the service room to breathe.
But there is still a crack in it. Clients that failed at the same moment are on the same schedule. Their retries still land in the same windows, just further apart. The herd is thinner now, but it is still moving as one.
Jitter scatters the herd
The final ingredient is randomness. Instead of waiting exactly four seconds, wait a random amount of time up to four seconds. One client waits half a second, another waits three, another waits the full four. The synchronized schedule dissolves. The retries that used to arrive as one hammer blow now arrive as a drizzle the service can actually absorb.
This combination has a name too: exponential backoff with jitter. The AWS Architecture Blog wrote the classic post on it ("Exponential Backoff And Jitter"), and the formula most people use is called full jitter:
wait = a random number between 0 and the smaller of (cap, base times 2 to the power of attempt)
The cap keeps the wait from growing forever. The base times 2 to the power of attempt is the exponential backoff. The random part is the jitter. All three matter, and most homegrown retry loops I have seen are missing at least one of them.
The code
Here is a retry helper in Go that does all three. Read it top to bottom. There is nothing clever in it, which is the point.
package main
import (
"context"
"math/rand"
"time"
)
// retryWithBackoff runs fn, retrying failures with exponential
// backoff and full jitter. It gives up after maxAttempts.
func retryWithBackoff(ctx context.Context, maxAttempts int, fn func() error) error {
const base = 100 * time.Millisecond
const maxWait = 5 * time.Second
var err error
for attempt := 0; attempt < maxAttempts; attempt++ {
if err = fn(); err == nil {
return nil
}
// Exponential part: 100ms, 200ms, 400ms, ...
wait := base << attempt
if wait > maxWait {
wait = maxWait
}
// Jitter part: random between 0 and wait.
jitter := time.Duration(rand.Int63n(int64(wait)))
select {
case <-ctx.Done():
return ctx.Err()
case <-time.After(jitter):
}
}
return err
}
A few things worth noticing. The wait doubles every attempt (that << is a bit shift, a cheap way to multiply by powers of two). The jitter picks a random wait anywhere from zero up to that doubled value, so two clients on the same attempt number almost never sleep the same amount. And the whole thing respects the context, so a caller can cancel the retries instead of letting them run to the bitter end. Modern Go seeds its random number generator for you, so there is no seeding ceremony needed.
The same idea in Python is even shorter:
import random
import time
def retry_with_backoff(fn, max_attempts=5, base=0.1, cap=5.0):
for attempt in range(max_attempts):
try:
return fn()
except Exception:
if attempt == max_attempts - 1:
raise
wait = min(cap, base * (2 ** attempt))
time.sleep(random.uniform(0, wait))
Same three ingredients. Exponential growth, a cap, and a random pick inside the window.
Three ways people still get this wrong
First, retrying things that should never be retried. Only retry failures that might succeed next time: timeouts, 503s, 429s. Do not retry a 400; the request itself is broken and trying again changes nothing. And only retry requests that are safe to repeat. Engineers call this idempotent: doing it twice has the same effect as doing it once. A retried "charge this card" that is not idempotent can charge it twice.
Second, backoff without a cap. Left alone, base times 2 to the power of attempt grows fast. By attempt ten you are waiting almost two minutes between tries. Pick a ceiling that matches how long your caller can actually afford to wait.
Third, retries with no overall deadline. A retry loop that never gives up will keep a dying request alive long past the point where anyone still wants the answer. Bound the attempts, or bound the total time, and let the caller decide what failure means.
The takeaway
Retries are load. Every retry policy you write is a decision about how much extra traffic your system will generate at the exact moment it is least able to handle it. Backoff and jitter are how you make that extra traffic survivable. The next time you write a retry loop, check for all three: growing waits, a cap, and randomness. If any one is missing, you are one bad twenty seconds away from finding out why.
Have you ever watched a retry storm flatten a service that was already recovering? What did the graphs look like?
Top comments (0)