The duplicate charges showed up in support tickets, both the same shape: one order, two rows in payments, same amount, same card, a few hundred milliseconds apart. The gateway logged two POSTs with different request ids and no shared key. Our retry wrapper had done exactly what we told it to.
The retry policy came from a shared HTTP client that had only ever wrapped reads. max_attempts=3, exponential backoff, retry on timeout, retry on connection errors, retry on any 5xx. Someone reused the call site for a write, and nothing in that client objected. The policy was correct for GETs and wrong for POSTs, and the code could not tell the difference.
A timeout is not a failure
A ReadTimeout means the client stopped waiting. It does not mean the server stopped working. The usual sequence is: TCP connect succeeded, request line, headers and body went out, the server started processing, the response came back late or never. The socket gave up on our side. On their side the row was already committed.
On a read path that mistake costs one extra query, so nobody notices. On a write path it costs one extra order. Same exception, same handler, different blast radius.
Before the request left, and after
These are different failures and they need different handling. Name the exception rather than catching Exception.
A ConnectTimeout or ConnectionRefusedError on a fresh connection happens before anything reached the server. Nothing was written. Retrying is reasonable, with one caveat: make sure no bytes of the request body were sent, because a ConnectionError raised after a partial send is not the same event.
A ReadTimeout, or RemoteDisconnected after the body was sent, is the unknown case. The server may have processed the request completely. It may have processed it and then failed while writing the response. There is no local signal that separates "never arrived" from "arrived and committed." The two situations collapse into one exception class, which is exactly why a blanket decorator is unsafe.
Retrying a 500 from a peer
A 500 is not evidence that nothing happened. It is evidence that something broke after the request arrived. The handler ran. The write may have committed and the error may have come from a post-commit step: sending a receipt, updating a search index, serializing a response. Retrying that POST inserts the row a second time.
A 502 or 503 from a proxy is more ambiguous. The upstream may never have been reached, or it may have received the request and answered too slowly for the proxy. Unless you own that layer and know its behavior, treat it as unknown too.
The only condition under which retrying a write is safe is idempotency: the server can distinguish a repeat from a new intent. That requires a key that travels with the original request, not one generated when the retry fires.
Make the key part of the request
# wrong: a fresh key on every attempt, so the retry looks like a new charge
def charge(order_id, amount):
for attempt in range(3):
key = str(uuid4())
try:
return gateway.post("/charges", json={"amount": amount},
headers={"Idempotency-Key": key})
except (ReadTimeout, ConnectionError):
continue
# right: the key comes from the intent and is stable across attempts
def charge(order_id, amount):
key = "charge:%s" % order_id
for attempt in range(3):
try:
return gateway.post("/charges", json={"amount": amount},
headers={"Idempotency-Key": key})
except ConnectTimeout:
continue # nothing was sent
except ReadTimeout:
raise UnknownWrite(order_id) # arrived, outcome unknown
The second version makes the retry a repeat of the same request. The gateway sees charge:4471 again, finds its stored result, and returns that instead of taking the money twice. If the endpoint does not implement the key, no client-side change makes the retry safe.
Treat a timed-out write as unknown
The honest state after a ReadTimeout on a POST is unknown. Not failed. Code that models it as failed retries, and the retry is what creates the duplicate.
What to do instead:
- Retry only writes you have made idempotent, and derive the key from the operation: order id, job id, a value you already persist. Never generate it inside the retry loop.
- Classify before you retry. Retry a connect failure that provably happened before the body was sent. Do not retry a response timeout, a 500 from a peer that received the request, or an ambiguous 502 unless the operation is idempotent.
- On an unknown write, stop and reconcile. Read the resource back by the key you sent, or query the server's idempotency record, before doing anything else.
- Store the key on the write path so a retry after a process restart uses the same value. A key built in memory does not survive a crash.
- If the endpoint has no idempotency key and you cannot add one, the only safe retry is one where no request was sent. Everything else needs a reconciliation job or a human.
To see how often this bites you, measure the timeout and 5xx rates on write endpoints separately from read endpoints, and count duplicate-key violations or duplicate rows per endpoint. A write endpoint with a nonzero timeout rate and no idempotency key is already producing duplicates. They show up as support tickets before they show up in your dashboards.
I write about production failures in Postgres, queues, and distributed systems.
Top comments (0)