DEV Community

challan116-ux
challan116-ux

Posted on

Idempotency Keys: The Retry Tax Every Integration Pays (and How EDI Paid It First)

Idempotency Keys: The Retry Tax Every Integration Pays (and How EDI Paid It First)

Every integration team learns the same lesson from the same incident: a network blip, a timeout that was not really a failure, a retry that did its job exactly as designed — and a customer charged twice, or a purchase order booked twice, or an inventory allocation doubled. The retry was not the bug. The missing idempotency key was.

I build EDI and API integrations for a living, and idempotency is the tax every integration pays whether it budgets for it or not. Modern APIs pay it with idempotency keys and dedupe tables. EDI has been paying it for forty years with control numbers, functional acknowledgments, and duplicate-checking translators. The vocabulary differs. The failure modes are identical.

What an idempotency key actually buys you

An idempotency key is a client-supplied identifier — usually a UUID — attached to a mutating request, with a simple contract: if the server sees the same key twice, it returns the result of the first execution instead of executing again. That is the whole idea. The hard part is everything the one-sentence version leaves out.

Scope. A key only means something inside a scope: this endpoint, this account, this trading partner. A key that dedupes globally will eventually collide with an unrelated request and return somebody else's result. In EDI terms, an interchange control number is only unique per sender — ISA13 from Walmart and ISA13 from Target are different namespaces that happen to use the same field.

Lifetime. Keys expire. Twenty-four hours is a common API default; EDI duplicate windows are often days, because a partner can legitimately resend an interchange after a long outage. Pick the window based on how long a retry can plausibly arrive, not on how long storage is cheap. A key that expires before the last retry lands is a duplicate waiting for a slow network.

Stored result, not just a flag. The naive implementation stores "seen: yes/no" and returns an error on the second attempt. That defeats the purpose — the retrying client does not know whether the first attempt succeeded. Store the response (status, body, or a reference to the created resource) and replay it. The second caller should get exactly what the first caller got.

The race that passes every sequential test

The implementation bug I see most often is check-then-act without atomicity:

  1. Request A arrives with key K. Server checks: K not seen. Proceed.
  2. Request B arrives with key K, milliseconds later. Server checks: K still not seen — A has not finished writing yet. Proceed.
  3. Both execute. Both side effects land. Both return success.

Sequential tests never catch this, because sequential tests never overlap. The fix is to make the claim atomic: insert the key with a unique constraint before doing the work, and let the database tell the loser to wait or replay. Insert-first, process-second. If the process crashes after the claim, you have a claimed key with no result — which is why the claim needs a state (in-flight, completed, failed) and a lease, so a later retry can take over a dead claim instead of replaying an empty one.

EDI translators learned the same lesson with duplicate interchanges: the control number is registered when the envelope is opened, not when translation finishes. Register early, or two copies of the same 850 both become orders.

Same key, different body

The second bug is quieter: a client reuses a key with a different payload. Maybe the caller regenerated the key from an order ID and the order was edited between attempts. Maybe two different operations share a key namespace. If you dedupe on the key alone, the second request silently returns the first request's result for a different operation — worse than a duplicate, because now the data is wrong and nothing errored.

Hash the payload alongside the key. Same key, same hash: replay the stored result. Same key, different hash: that is a client bug — return a clear error (409 or 422) and log it loudly. EDI has the analog: same control number, different document content is a partner-side mapping bug, and the correct response is a rejected acknowledgment with the reason in it, not a silent merge.

EDI paid this tax first

None of this is new if you come from EDI:

  • Control numbers are idempotency keys. ISA/GS/ST control numbers exist so a receiver can say "I have already seen interchange 000000042" and refuse the second copy. The envelope is the dedupe table.
  • Acknowledgments are layered. An AS2 MDN says "received intact." A 997 says "parsed, structure valid." A business-level response (855, 997's transactional cousin, an application advice) says "accepted, and here is what we did with it." Each layer answers a different question, and conflating them — treating "received" as "processed" — is how duplicates slip through during retries.
  • Duplicates are expected traffic. Partners resend when they do not get an ack, and they do not always wait politely. Any EDI flow that assumes exactly-once delivery is a design document, not a system.

The API world keeps rediscovering these three ideas and naming them fresh. That is fine — the ideas are good. It is just worth knowing they have forty years of production scar tissue behind them.

A checklist before you ship

  • Keys scoped per account/partner and per operation, never global
  • Atomic insert-first claim with a unique constraint — no check-then-act gap
  • Claim carries a state and a lease, so a crashed attempt can be taken over
  • Stored result replayed on duplicate, not a bare error
  • Payload hash stored with the key; same-key-different-body is a loud client error
  • Key lifetime matched to real retry horizons (including the partner who resends after a weekend outage)
  • Acknowledgment layers kept separate: received ≠ parsed ≠ accepted
  • One test that fires the same key concurrently — because production will

If you are wiring up order flows, supplier events, or shipment notifications where a duplicate creates a real-world mess — a double order, a double shipment, a double charge — that dedupe-and-acknowledge layer is exactly what we built SignalEDI to own, so an X12 interchange and an API event get the same duplicate protection and the same audit trail.

Retries are not the enemy. Retries without memory are. Give your integration a memory, and the 2 a.m. page becomes a log line instead.

— Chris, founder of SignalEDI

Top comments (0)