DEV Community

Cover image for Who gave your frontend permission to retry?
Siddharth Pandey
Siddharth Pandey

Posted on

Who gave your frontend permission to retry?

Almost every frontend codebase has this somewhere, usually in a file called api.ts:

async function fetchWithRetry(url: string, retries = 3) {
  try {
    return await fetch(url);
  } catch (err) {
    if (retries === 0) throw err;
    return fetchWithRetry(url, retries - 1);
  }
}
Enter fullscreen mode Exit fullscreen mode

It reads as resilience. From the server's side it is something else: a rule that says whenever the backend is struggling, every open tab sends it four times the traffic, immediately, all at once.

A retry is a request the server did not ask for. Backend engineers have treated that as a design problem for years, with backoff, jitter, budgets and circuit breakers. Frontend code mostly has not, and the frontend is the layer where it matters most. A backend service has a known number of callers. A web app has as many callers as there are open tabs, none of them coordinate, and all of them noticed the outage in the same second.

This post is about what a respectful retry looks like from the browser: backoff, jitter, and the four things the browser adds that backend retry guides never mention.

Backoff alone only reschedules the stampede

The first fix everyone reaches for is exponential backoff: wait 1s, then 2s, then 4s. It feels like it should spread the load. It does not, because every client computes the same delays from the same starting moment.

To put numbers on it, here is a small simulation. One thousand clients all fail at t=0, the server stays down, and each client retries four times. Delays use a 1 second base and a 30 second cap. The script counts how many retries land in each 100ms window.

const strategies = {
  'immediate (no delay)':       () => 0,
  'fixed 1s':                   () => BASE,
  'exponential':                (a) => Math.min(CAP, BASE * 2 ** a),
  'exponential + equal jitter': (a) => { const c = Math.min(CAP, BASE * 2 ** a); return c / 2 + rnd() * c / 2; },
  'exponential + full jitter':  (a) => rnd() * Math.min(CAP, BASE * 2 ** a),
};
Enter fullscreen mode Exit fullscreen mode

The output:

immediate (no delay)         peak/100ms= 4000  busy buckets=   1  last retry at 0.0s
fixed 1s                     peak/100ms= 1000  busy buckets=   4  last retry at 4.0s
exponential                  peak/100ms= 1000  busy buckets=   4  last retry at 15.0s
exponential + equal jitter   peak/100ms=  208  busy buckets= 114  last retry at 14.7s
exponential + full jitter    peak/100ms=  164  busy buckets= 137  last retry at 14.3s
Enter fullscreen mode Exit fullscreen mode

Read the third row carefully. Plain exponential backoff has exactly the same peak as a fixed one second delay: all 1,000 clients arrive in the same 100ms window, four separate times. Backoff made the waves further apart. It did not make any wave smaller. A server that fell over under 1,000 simultaneous requests gets 1,000 simultaneous requests at 1s, at 3s, at 7s and at 15s, which is a good way to knock it down again each time it gets up.

The total work is identical in every row: 4,000 retries. The only thing that changes is whether they arrive as four walls or as a drizzle.

Jitter is the part that does the work

Jitter means adding randomness to the delay so clients stop moving in lockstep. The reference for this is Marc Brooker's 2015 post on the AWS Architecture Blog, Exponential Backoff And Jitter, which compared the variants and concluded that jittered backoff "should be considered a standard approach for remote clients".

Full jitter is the simplest version. Compute the exponential ceiling, then pick a uniformly random delay between zero and that ceiling:

const delay = Math.random() * Math.min(CAP, BASE * 2 ** attempt);
Enter fullscreen mode Exit fullscreen mode

In the simulation that one Math.random() drops the peak from 1,000 requests per window to 164, and spreads the same 4,000 retries across 137 windows instead of 4. It also finishes slightly sooner, because the average delay is half the ceiling.

Equal jitter keeps half the backoff as a guaranteed floor and randomises the other half. It is the intuitive compromise, and it did worse here (208 vs 164), which matches Brooker's own verdict that equal jitter "is the loser". The floor it protects is exactly the part that keeps clients bunched together.

The uncomfortable implication for frontend code: a delay of nearly zero is a valid outcome of full jitter. That is fine. The goal was never for each client to wait a long time. The goal is for clients to disagree with each other.

What the browser changes

Everything above applies to any client. The next four points are specific to code running in a tab, and they are where frontend retry logic usually goes wrong.

1. The server may have told you when to come back, and you could not read it

A 429 or 503 response can carry a Retry-After header, either as a number of seconds or as an HTTP date. If the server is telling you when it will be ready, guessing with your own backoff curve is rude and less accurate.

The browser trap is CORS. Retry-After is not on the list of response headers that cross-origin JavaScript is allowed to read by default. If your API lives on a different origin from your app, res.headers.get('retry-after') returns null even though the header is plainly visible in the network tab. The server has to opt in:

Access-Control-Expose-Headers: Retry-After
Enter fullscreen mode Exit fullscreen mode

Without that line, the backend team can implement rate limiting perfectly and the frontend will ignore it, silently. Libraries do not save you here. axios-retry, for example, does read retry-after and uses it as a minimum delay, but it reads from the same header object the browser already filtered.

One more detail: when a rate limiter sends Retry-After: 30 to a thousand clients, it has just synchronised them. Honour the header as a floor and add your jitter on top of it, not instead of it.

2. A hidden tab should not be retrying at all

A backend worker that is retrying has a caller waiting on it. A browser tab that is retrying may have been in the background for three hours. Nobody is waiting for that response, and every retry it sends is pure cost to a server that is already in trouble.

The same applies offline. If navigator.onLine is false, the request cannot succeed, so spending retry attempts on it only guarantees that the attempts are used up before the connection returns. (The reverse is not reliable: true means a network interface exists, not that your API is reachable.)

The respectful behaviour is to pause. Hold the retry until visibilitychange says the tab is visible again, then go.

3. The user is also a retry loop

When a page hangs, people click the button again. Then they hit refresh. Each of those starts a fresh request with a fresh retry counter, layered on top of whatever your code was already doing. No backend client behaves like this.

Two things follow. First, cancel the old work: pass an AbortSignal through, and when the component unmounts or the user navigates, abort. A retry loop that keeps running for a screen nobody is looking at is the frontend version of a zombie process. Second, give the user a reason not to mash the button. A visible "Retrying in 4s" with a manual "Try now" converts an anxious user into a patient one. Silent retries behind a spinner invite the refresh.

4. Retries multiply across layers

The retry loop in api.ts is rarely the only one. A typical stack has a fetch wrapper with 3 retries, inside a data fetching library with 3 retries, behind an API gateway or BFF that retries the origin 3 times. Each layer makes up to 4 attempts, so one user action can become 4 x 4 x 4 = 64 requests at the origin.

Google's SRE book states the rule directly in its chapter on handling overload: "requests should only be retried at the layer immediately above the layer that is rejecting them." For a frontend, that means picking one layer to own retries and setting every other layer to zero.

Know what your library already does

Before writing any retry code, check what is already retrying on your behalf. The defaults differ more than most people expect.

Library Default retries Default delay Jitter Pauses when hidden
TanStack Query (queries) 3 in the browser, 0 on the server Math.min(1000 * 2 ** failureCount, 30000) No Yes
TanStack Query (mutations) 0 n/a n/a n/a
SWR on, exponential from a 5s interval randomised Yes Yes, on default settings
axios-retry 3 none (noDelay) Only with exponentialDelay, up to 20% No

A few observations. TanStack Query's default delay has no random term, so a fleet of tabs on defaults retries in step. It does the considerate thing on visibility, though: its retryer pauses "if the document is not visible or when the device is offline". Its choice to not retry mutations by default is also correct, for reasons covered in the next section. axios-retry out of the box is the immediate-retry row of the simulation, softened only by the fact that it limits itself to network errors and idempotent methods.

Adding jitter to TanStack Query is one option:

new QueryClient({
  defaultOptions: {
    queries: {
      retryDelay: (attempt) => Math.random() * Math.min(30_000, 1000 * 2 ** attempt),
    },
  },
});
Enter fullscreen mode Exit fullscreen mode

Not everything deserves a second attempt

Backoff and jitter decide when to retry. The harder question is whether to.

Status codes. Retrying a 400, 401, 403, 404 or 422 asks the same question again and expects a different answer. The useful set is small: 408, 429, 502, 503, 504, plus network failures where fetch rejects. Whether to include a bare 500 is a judgement call; it more often means a bug than a blip, and bugs do not heal in four seconds.

Idempotency. A GET can be repeated freely. A POST /orders cannot, because a timeout does not tell you whether the server processed the request before the connection died. Retrying it is how users get charged twice. The standard fix is an idempotency key: the client generates a unique ID once per user intent, sends it on every attempt, and the server deduplicates.

const key = crypto.randomUUID(); // once per click, not once per attempt
await fetchWithRetry('/orders', {
  method: 'POST',
  headers: { 'Idempotency-Key': key },
  body: JSON.stringify(order),
});
Enter fullscreen mode Exit fullscreen mode

This needs server support. Without it, the honest frontend behaviour for a failed non-idempotent request is to stop and tell the user.

Budgets. A per-request cap of three retries still allows traffic to quadruple during an outage. The SRE book describes a second limit: a per-client retry budget, where "a request will only be retried as long as this ratio is below 10%", which by their numbers cuts worst-case growth from just under 3x to about 1.1x. In a tab, that is a small token bucket: each first attempt earns a tenth of a token, each retry spends one.

Putting it together

Here is the whole thing in one function. Jittered backoff, Retry-After as a floor, no retries for unsafe requests, a budget, cancellation, and a pause for hidden tabs.

const BASE = 1000, CAP = 30_000, MAX_SERVER_WAIT = 60_000;
const RETRYABLE = new Set([408, 429, 502, 503, 504]);
const IDEMPOTENT = new Set(['GET', 'HEAD', 'OPTIONS', 'PUT', 'DELETE']);

// Retry budget: every first attempt earns 0.1 token, every retry costs 1.
// Steady state, retries cannot exceed ~10% of traffic from this tab.
let tokens = 10;
const deposit = () => { tokens = Math.min(10, tokens + 0.1); };
const withdraw = () => (tokens >= 1 ? (tokens -= 1, true) : false);

function retryAfterMs(res: Response): number {
  const h = res.headers.get('retry-after');
  if (!h) return 0;
  const secs = Number(h);
  if (Number.isFinite(secs)) return secs * 1000;
  const date = Date.parse(h);
  return Number.isNaN(date) ? 0 : Math.max(0, date - Date.now());
}

const sleep = (ms: number, signal?: AbortSignal | null) =>
  new Promise<void>((resolve, reject) => {
    const t = setTimeout(resolve, ms);
    signal?.addEventListener('abort', () => { clearTimeout(t); reject(signal.reason); }, { once: true });
  });

function untilVisible(): Promise<void> {
  if (document.visibilityState === 'visible') return Promise.resolve();
  return new Promise((resolve) => {
    document.addEventListener('visibilitychange', function on() {
      if (document.visibilityState !== 'visible') return;
      document.removeEventListener('visibilitychange', on);
      resolve();
    });
  });
}

export async function fetchWithRetry(url: string, init: RequestInit = {}, maxRetries = 3): Promise<Response> {
  const headers = new Headers(init.headers);
  const method = (init.method ?? 'GET').toUpperCase();
  const safeToRepeat = IDEMPOTENT.has(method) || headers.has('Idempotency-Key');
  deposit();

  for (let attempt = 0; ; attempt++) {
    let res: Response | undefined;
    let err: unknown;
    try { res = await fetch(url, { ...init, headers }); } catch (e) { err = e; }

    if (res && !RETRYABLE.has(res.status)) return res;

    const serverWait = res ? retryAfterMs(res) : 0;
    const giveUp = init.signal?.aborted || !safeToRepeat || attempt >= maxRetries
      || serverWait > MAX_SERVER_WAIT || !withdraw();
    if (giveUp) { if (res) return res; throw err; }

    const jitter = Math.random() * Math.min(CAP, BASE * 2 ** attempt);
    await sleep(serverWait + jitter, init.signal);
    await untilVisible();
  }
}
Enter fullscreen mode Exit fullscreen mode

A few choices in there are worth defending. If the server asks for more than a minute, the function gives up and returns the response, because no user is watching a spinner for that long and the UI should say so. A plain POST without an idempotency key gets exactly one attempt. And the budget is shared across every call in the tab, so one broken endpoint cannot spend retries on behalf of the whole app.

The budget does the most work when things are worst. Running this function against a mocked endpoint that always returns 503, 100 consecutive requests produced 14 retries in total. Without the budget, the same 100 requests would have produced 300.

If you already use a data fetching library, do not add this underneath it. Configure the library's own retry options to do the same things, and set the other layers to zero.

Conclusion

The three-line retry loop is written from the point of view of one request that wants to succeed. Every improvement here comes from switching to the point of view of the server receiving all of them. Backoff gives it room. Jitter stops the clients arriving together. Retry-After lets it speak for itself. Budgets, idempotency rules and paused tabs all remove requests that had no business being sent.

None of this makes an individual request more likely to succeed. It makes the outage shorter for everyone, which is the same thing over any timescale that matters.

Key Takeaways

  • Exponential backoff without jitter has the same peak load as a fixed delay. In the simulation, both hit 1,000 requests per window; full jitter hit 164.
  • Use full jitter: Math.random() * Math.min(cap, base * 2 ** attempt). Check whether your library's default includes a random term. TanStack Query's does not.
  • If your API is cross-origin, Retry-After is invisible to JavaScript until the server sends Access-Control-Expose-Headers: Retry-After. Treat the header as a floor and still add jitter.
  • Retry in one layer only. Three layers with three retries each is up to 64 requests per click.
  • Only retry what is safe to repeat: a short list of status codes, idempotent methods, and writes that carry an idempotency key. Cap the rest with a retry budget of around 10%.
  • Pause retries in hidden tabs and abort them when the user leaves.

The part I have not settled is the user. A person hitting refresh during an outage resets every counter and budget in this post, and the only defences I know of are better feedback in the UI or enforcement on the server. Should a frontend persist its retry budget across reloads, or is that the point where client-side politeness stops and the rate limiter has to take over?

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to