Retries + Backoff + Jitter: How a Small Failure Becomes a Retry Storm.
Retries improve availability until they become the outage. The dangerous question isn't How many retries should we configure? It's which layer owns the retry, is the operation idempotent, and what happens when 10,000 clients retry together?
The key mental model is a retry consumes additional capacity from the system that may already be failing. AWS's Builders' Library gives a particularly memorable example: with a five-deep service stack and three retries independently occurring at each layer, load against the database can increase 243×
Backoff alone isn't enough. Correlated clients can wake up together and generate another burst, which is why jitter spreads retry arrival times.
The other architect-level subtlety is timeout ambiguity. A timeout tells the caller that it didn't receive a successful response; it doesn't prove the remote operation didn't happen. Side-effecting operations such as payment creation, order submission or label purchase therefore need idempotency semantics before automatic retries are safe.
I hope below visual model helps to understand it better. Let me what kind of issues you faced with retried in in which layer, and do you track it?

Top comments (0)