Picture an integration. Your app sends a webhook: order 42 is confirmed, please create the shipment. The partner commits the shipment and returns success. The connection breaks before your sender reads the response.
From your side, delivery failed. From theirs, the work is done.
A timeout tells you what you know. It says nothing about what the receiver did. So you retry, and the only thing standing between your customer and a second parcel is whether the receiver can recognize that this is the same thing again.
I ran into this while documenting the webhook module of Semitexa, the PHP framework I'm building. The mechanics were fine. What turned out to be hard was the vocabulary: "the key" meant three different things depending on who was asking.
One key, three questions
When people say "make it idempotent", they usually mean one key. But there are three separate questions:
- Publication identity: did we already queue this outbound intent for this endpoint?
- Event identity: did the receiver already record this event?
- Business identity: is the effect already done, one shipment for this fulfilment intent?
For the example, that could be order:42:confirmed:v1, evt-order-42-v1, and a business rule. They protect different things, and mixing them up fails in two opposite directions:
- A fresh key on every retry and the receiver can never recognize the repeat.
- One key for everything about order 42 and the next legitimate shipment (a replacement, a second confirmation) gets swallowed as a duplicate.
The rule I ended up writing down: the key describes what may happen once, and it has to survive every attempt at that same thing.
Retry the right failures
Not every failure deserves another attempt. The policy we settled on:
| Result | Decision |
|---|---|
| 2xx | delivered |
| 408, 429 | retry if attempts remain |
| other 4xx | permanent rejection |
| 5xx, timeout, no status | retry with exponential backoff and jitter, within a budget |
A 400 is the receiver telling you the request is wrong. Sending it again ten times is just noise. A 503 or a timeout is the opposite: you don't know, so you ask again, and the receiver's memory makes asking again safe.
"Duplicate" is not "completed"
This one is quiet. The receiver gets the event, records it in its inbox, and crashes before creating the shipment. Your retry arrives. The inbox finds the record and says: duplicate, ignore.
The shipment never happens, and every log line looks healthy.
An inbox match tells you the event was seen. It does not tell you the work was done. The recovery decision has to look at business state, not only at the inbox. Where the action is a local database write, the completion record and the business update should commit together. Where it's an external side effect, you need the recipient's own idempotency mechanism or an explicit reconciliation path.
A lease protects the record, not the request
Delivery workers claim a record with a lease. Suppose a worker sends the HTTP request, then loses its lease before writing "delivered". A well-behaved worker won't overwrite a record it no longer owns, and ours doesn't. Good. But the request has already gone out. The receiver may already have acted.
The lease coordinates your bookkeeping. It cannot retract anything on the other side. Which brings it back to the receiver: the repeat has to be recognizable there, by the event identity, and the effect has to be protected where it happens, by the business identity.
What I'd check in any integration
- Can you name all three identities for your most important webhook?
- Does a retry carry the same event identity as the first attempt?
- If the receiver records and then crashes, what finishes the work?
- Do you inspect business state before you replay anything by hand?
If question 3 has no answer, that's the one I'd fix first.
Which of the three identities does your integration actually store today?
The longer version, with the retry ranges, the inbox key format and the recovery cases, is on the Semitexa blog: Your Webhook Was Delivered Twice. Now What?
Top comments (0)