Traced a double-charge bug that had nothing to do with retry timing or refunds. It was write order.
Our idempotency middleware called the processor first, then wrote the idempotency key to the database once the charge succeeded. Made sense on paper: don't mark a key as "used" until you know the outcome.
Problem shows up when the process dies between those two steps. Processor accepts the charge, returns a 201, and the pod gets OOM-killed before the key write commits. Client sees a timeout, retries with the same key. Middleware checks the table, finds nothing, treats it as a fresh request, charges again.
Fixed it by flipping the order: write the key with status "pending" before calling the processor, then update to "complete" after. A retry that lands mid-flight now sees "pending" and can poll or reject instead of reprocessing. Slightly more writes per request, but the failure window shrinks from "anytime during the whole request" to a single row update.
Most idempotency writeups show the happy path — client retries, server checks key, returns cached response. Fewer show what your key-write timing actually protects against, which turns out to matter more than the retry logic itself.
Where do you write your idempotency key — before the call to the processor, or after?
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.