Maya taps Pay for a $20 ride. Ten seconds later: request timed out. The app retries, and Maya gets charged $40.
The retry went through. So did the first attempt. The first request never failed: it worked perfectly, and the answer just never made it back.
Stripe's fix for this is one extra header on the request, a random string called an idempotency key. This post walks through how it works and how to build it yourself. The animated version:
Three ways to fail
Stripe's engineers start from a blunt fact: the network is unreliable. A call from a client to a server can fail in three places:
- Before the request reaches the server. Nothing happened, so a retry is safe.
- Halfway through. The server started working and something broke. Maybe the charge ran, maybe it didn't.
- After the work is done. The money moved, then the connection dropped before the response got back.
From the client's side, all three look exactly the same: a timeout and no answer.
Retry, or give up?
Both obvious options are wrong. Never retry, and case 1 gives Maya a free ride. Always retry, and case 3 charges her twice. Waiting longer doesn't help you choose: a slow server and a dead connection look identical from the outside.
What you want is a retry that's safe no matter which case happened.
What "idempotent" means
An operation is idempotent if doing it twice has the same effect as doing it once. A GET only reads. DELETE the same thing twice and it's just as deleted. Stripe's article uses a PUT to a DNS provider: "set this record to this value" can be sent twice and the record ends up the same.
A charge isn't like that. A POST /v1/charges with amount=2000 (that's $20 in cents) says make a new charge, not which charge. The server can't tell a retry from a new purchase. Send it twice and you get two charge objects and $40 gone.
The idempotency key
The server needs a name for the attempt. That's the idempotency key: the client makes up a unique string and sends it in an Idempotency-Key header.
- Stripe suggests a V4 UUID, or any string random enough that two clients never pick the same one. Keys can be up to 255 characters.
- The crucial part: the key is made once, before the first attempt, and every retry of that payment reuses it.
Think of a coat check: show ticket 42 twice and you get the same coat back, not a second coat.
What the server remembers
Each key points to a saved result. Stripe saves the first response for every key: its status code and body.
- First time a key arrives: no entry, so the server runs the charge and stores the response against the key.
- Second time: the entry is there, so it charges nothing and returns the saved response.
It saves the result whether the request succeeded or failed. If the first attempt returned a 500, the retry gets that same 500. A key doesn't mean "try again from scratch". It means "tell me what happened to this request".
The hard case: crashing halfway
Replay the three failures with a key. Case 1: the server never saw the key, so the retry runs normally. Case 3: the retry finds the key and gets the saved reply. Case 2 is the hard one: the server died partway through, so it has to remember how far it got.
Brandur Leach, who wrote Stripe's article, sketches a Postgres design where each key's row stores a recovery point: started, ride_created, charge_created, finished. A retry picks up from there, and each step between calls to other services commits all at once or not at all.
When his server calls Stripe, it passes its own idempotency key along, so even a crash in the middle of the charge can't charge Maya twice.
Edge cases that bite
Two retries at once. Both look up the key and find no finished result. So the key's row gets a lock: the first request stamps a locked_at time and does the work, and the second gets a 409 Conflict. Nothing is saved or charged, and the client retries later.
Same key, different request. If a bug reuses Maya's key for a $50 charge, Stripe compares it with the original and rejects it. Keys are also scoped per user, and should never be built from personal data like an email address.
How long to remember. Stripe lets keys be pruned once they're at least 24 hours old. For your own system Brandur suggests about 72 hours, so a bad Friday deploy still has its keys through the weekend. A background "reaper" deletes old keys, and a "completer" pushes abandoned requests to the finish.
Back off, and add jitter
Keys make a retry safe. They don't make it polite. If the server goes down and thousands of clients retry instantly, they hammer the thing that needs to recover.
Stripe's answer is exponential backoff: wait 2^n, where n is the number of failures so far. But if an outage hits thousands of clients in the same second, they all back off on the same schedule and come back in synchronized waves: the thundering herd.
The fix is jitter: add a random amount to every wait, and the waves smear into a steady trickle. Stripe's own client libraries can do both for you, and attach an idempotency key to every request they retry.
The whole recipe
- Handle failures consistently: clients retry.
- Handle failures safely: servers are idempotent, using keys.
- Handle failures responsibly: back off exponentially and add jitter.
At Stripe, every POST accepts an idempotency key. GET and DELETE don't need one: they're idempotent by definition.
Watch the full animated breakdown: https://youtu.be/OE_RnizyyEM
Sources: Stripe's Designing robust and predictable APIs with idempotency, the Stripe API docs on idempotent requests, and Brandur Leach's Implementing Stripe-like idempotency keys in Postgres. Maya, the charge ids and timings in the images are illustrative.
Practice designing systems like this yourself at systemdesignlab.in.






Top comments (0)