The API returns 503 Service Unavailable.
Your client immediately retries three times.
That sounds resilient.
Now imagine ten thousand clients doing the same thing at the same time.
The struggling service does not receive less traffic. It receives several times more traffic precisely when it has the least capacity to handle it.
A retry is not automatically a recovery mechanism.
Without timeouts, backoff, jitter, retry limits, and safe operation semantics, retry code can turn a temporary failure into a longer outage.
The retry loop that makes failures worse
This implementation looks reasonable at first glance:
async function requestWithRetry(url) {
for (let attempt = 0; attempt < 3; attempt++) {
try {
const response = await fetch(url);
if (response.ok) {
return response;
}
} catch {
// Try again immediately
}
}
throw new Error("Request failed");
}
It has several problems:
- every attempt can wait indefinitely,
- all failures are treated as retryable,
- retries happen immediately,
- every client follows the same timing,
- the total operation has no time budget,
- the method does not consider whether repeating the request is safe,
- the final error loses useful context.
The fix is not simply “add more retries.”
The fix is to make every retry earn its place.
Start with a timeout
A request that never completes can occupy connections, memory, and other limited resources. Set a maximum wait for each attempt.
Modern JavaScript provides AbortSignal.timeout():
const response = await fetch(url, {
signal: AbortSignal.timeout(5_000),
});
The signal aborts the request after the configured active time and rejects with a TimeoutError DOM exception.
If the caller also needs to cancel the operation, combine signals:
async function fetchWithTimeout(url, callerSignal) {
const signal = AbortSignal.any([
callerSignal,
AbortSignal.timeout(5_000),
]);
return fetch(url, { signal });
}
A timeout does not tell you whether the server completed an operation before the client stopped waiting. That distinction becomes critical for requests with side effects.
Retry only failures that may be temporary
A 400 Bad Request usually will not improve after a pause. Neither will an authentication failure caused by an invalid token.
Good retry candidates often include:
-
408 Request Timeout, -
429 Too Many Requests, - selected
5xxresponses, - a temporary network failure,
- a per-attempt timeout when repeating the operation is safe.
Treat this list as a policy for the API you are calling, not a universal law.
const RETRYABLE_STATUS_CODES = new Set([
408,
429,
500,
502,
503,
504,
]);
function isRetryableResponse(response) {
return RETRYABLE_STATUS_CODES.has(response.status);
}
Do not retry every 5xx response automatically without considering the service contract. Some APIs document narrower behavior, while others provide a specific Retry-After response.
Add exponential backoff
Immediate retries create more load when the service is already struggling.
Exponential backoff increases the delay after each failed attempt:
function backoffDelay(attempt, baseDelay = 250) {
return baseDelay * 2 ** attempt;
}
That produces delays such as:
250 ms
500 ms
1000 ms
2000 ms
Backoff gives the dependency time to recover and reduces repeated pressure from one client.
It does not solve synchronization across many clients.
Add jitter so clients do not retry together
If thousands of clients fail at the same moment and use the same deterministic backoff, they wake up together after the same delay. The retry spike has merely moved.
Jitter adds randomness:
function fullJitterDelay(attempt, baseDelay = 250, cap = 8_000) {
const maximum = Math.min(cap, baseDelay * 2 ** attempt);
return Math.floor(Math.random() * maximum);
}
Clients now spread attempts across a time window rather than forming another synchronized wave.
A small sleep helper keeps the retry loop readable:
function sleep(milliseconds, signal) {
return new Promise((resolve, reject) => {
const timeout = setTimeout(resolve, milliseconds);
signal?.addEventListener(
"abort",
() => {
clearTimeout(timeout);
reject(signal.reason);
},
{ once: true },
);
});
}
In production code, remove abort listeners when they are no longer needed or use a well-tested utility that already handles cleanup.
Respect Retry-After
A server may tell the client when to try again:
Retry-After: 5
The value can be a delay in seconds or an HTTP date.
function parseRetryAfter(value) {
if (!value) return null;
const seconds = Number(value);
if (Number.isFinite(seconds)) {
return Math.max(0, seconds * 1_000);
}
const date = Date.parse(value);
if (Number.isNaN(date)) {
return null;
}
return Math.max(0, date - Date.now());
}
When the server provides a valid instruction, prefer it over an arbitrary client delay, while still enforcing your own maximum wait and total operation budget.
Do not retry unsafe operations blindly
Retrying a read is usually easier to reason about than retrying a payment, order creation, email send, or database mutation.
Consider this sequence:
Client sends payment request
Server charges the card
Response is lost
Client times out
Client retries
Server charges the card again
A timeout proves that the client did not receive a response. It does not prove that the server did nothing.
Operations with side effects need an idempotency strategy. A common pattern is a unique key representing one logical operation:
const response = await fetch("/payments", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify(payment),
});
The server stores the key and returns the original result when the same logical operation is submitted again.
This works only when the server actually implements idempotency. A client-generated header alone does not make an endpoint safe.
Build a bounded retry helper
Here is a compact implementation that combines the main ideas:
const RETRYABLE_STATUS_CODES = new Set([
408,
429,
500,
502,
503,
504,
]);
function fullJitterDelay(attempt, baseDelay, cap) {
const maximum = Math.min(cap, baseDelay * 2 ** attempt);
return Math.floor(Math.random() * maximum);
}
function parseRetryAfter(value) {
if (!value) return null;
const seconds = Number(value);
if (Number.isFinite(seconds)) {
return Math.max(0, seconds * 1_000);
}
const date = Date.parse(value);
return Number.isNaN(date) ? null : Math.max(0, date - Date.now());
}
function sleep(milliseconds) {
return new Promise((resolve) => setTimeout(resolve, milliseconds));
}
export async function fetchWithRetry(
url,
options = {},
{
attempts = 3,
timeoutMs = 5_000,
baseDelayMs = 250,
maxDelayMs = 8_000,
} = {},
) {
let lastError;
for (let attempt = 0; attempt < attempts; attempt++) {
try {
const response = await fetch(url, {
...options,
signal: AbortSignal.timeout(timeoutMs),
});
if (response.ok || !RETRYABLE_STATUS_CODES.has(response.status)) {
return response;
}
if (attempt === attempts - 1) {
return response;
}
const retryAfter = parseRetryAfter(
response.headers.get("Retry-After"),
);
const delay = Math.min(
maxDelayMs,
retryAfter ?? fullJitterDelay(
attempt,
baseDelayMs,
maxDelayMs,
),
);
await sleep(delay);
} catch (error) {
lastError = error;
if (attempt === attempts - 1) {
throw error;
}
await sleep(
fullJitterDelay(attempt, baseDelayMs, maxDelayMs),
);
}
}
throw lastError ?? new Error("Request failed");
}
This is a starting point, not a universal networking library. A production client may also need:
- caller cancellation,
- a total time budget across all attempts,
- domain-specific retry rules,
- observability,
- circuit breaking,
- connection pooling,
- rate-limit coordination,
- idempotency keys,
- body replay rules.
Measure retries as part of the system
A retry that eventually succeeds can hide a degrading dependency.
Track at least:
- requests by final outcome,
- attempts per logical operation,
- time spent waiting between attempts,
- status codes that triggered retries,
- requests that exhausted the retry budget,
- operations completed after one or more retries,
- duplicate-operation prevention,
- total latency including backoff.
A dashboard showing only successful final responses may report a healthy service while clients quietly perform two or three times the expected traffic.
A practical retry checklist
Retry-After is respected when appropriate.
The takeaway
Retries are useful because distributed systems fail temporarily.
Retries are dangerous for exactly the same reason: a struggling dependency has less capacity available, not more.
A safer sequence is:
Set a timeout
↓
Classify the failure
↓
Check whether repeating is safe
↓
Wait with backoff and jitter
↓
Respect the retry budget
↓
Measure the result
Do not ask only whether the request eventually succeeded.
Ask how many attempts it took, how much extra load it created, and whether repeating the operation could have caused a second side effect.
What is the hardest retry failure you have diagnosed: synchronized clients, duplicate writes, ignored rate limits, or something else?
Read the AWS retry with backoff guidance
Sources and further reading
- AWS Prescriptive Guidance: Retry with backoff pattern
- AWS Builder Center: Timeouts, retries, and backoff with jitter
- MDN: AbortSignal.timeout()
- MDN: AbortSignal
Connect with Me
If you found this article helpful, let's connect!
- 💻 GitHub: johnnylemonny
Top comments (3)
Great point! 👍
Retries can look harmless, but without things like backoff and a clear retry limit, they can easily make the situation worse instead of helping 😸
I also liked the reminder that a timeout doesn’t necessarily mean the server didn’t complete the operation. That’s easy to forget.
Thanks!
That second point is exactly the one that surprises people most. A timeout only tells us what the client experienced, not what actually happened on the server. Once side effects enter the picture, retries become much more than a networking concern.
And yes, retries can be incredibly useful, but without limits, backoff, and a bit of randomness, they can quickly turn into an accidental traffic amplifier. Glad you found the article helpful! 👍
Việc thiếu jitter trong chiến lược retry là lỗi mà mình từng mắc phải và nó thực sự gây ra thảm họa khi hàng loạt client cùng ập vào server ngay khi vừa phục hồi. Nếu không có độ trễ ngẫu nhiên, các request sẽ tạo thành những đợt sóng xung kích (thundering herd problem) khiến hệ thống không kịp thở. Ngoài việc áp dụng exponential backoff, mình thấy việc kết hợp chặt chẽ với Retry-After header từ phía server là cực kỳ quan trọng để điều tiết lưu lượng một cách thông minh thay vì chỉ dựa vào logic client cứng nhắc. Đảm bảo tính idempotency cho các request retry cũng là yếu tố sống còn để tránh tình trạng side effect dữ liệu khi mạng chập chờn.