In an ERP, a duplicated webhook is not a technical bug. It's a wrongly issued invoice.
I work on the ecommerce integrations of a SaaS ERP: Shopify, WooCommerce, PrestaShop. Each store sends us events about its orders, and the expensive mistake is always the same: the system does the same thing twice, or does today what was already corrected yesterday. In this post I explain the layers I use to avoid it and why each one exists.
What you can't assume about a webhook
A webhook does not guarantee that an event arrives only once or in order. Shopify, for example, does not guarantee ordering within a topic or across topics for the same resource and recommends ignoring duplicates with the X-Shopify-Webhook-Id header.
With one store it's rare. With thousands of integrations, rare stops being rare: any combination of duplicate, concurrency and disorder happens every day. That's why there is no single solution. There are layers, and each one covers a gap the others don't see.
Before the layers there is a base rule: the endpoint only verifies the signature, stores the event exactly as it arrives and responds 200. Everything else happens in a worker, so the response does not depend on how long processing takes. Storing the raw event also lets you reprocess and audit it when there is a bug.
And a webhook is a notification, not the truth. For a critical decision, like issuing an invoice, it's better to request the order again from the store's API instead of trusting the payload you received.
Layer 1: discard the identical before processing it
The cheapest one. You compute a hash of the payload and store it in Redis with an expiry. If it already exists, the event is a repeat and you discard it without touching the database.
// Illustrative example, not the real code
$key = "webhook:{$shopDomain}:" . hash('sha256', $payload);
// SET NX EX: check and mark in a single atomic operation
if (!$redis->set($key, '1', ['NX', 'EX' => 300])) {
return; // already received, discard
}
try {
$this->inbox->store($rawEvent); // durable, unique per event id
} catch (\Throwable $e) {
$redis->del($key); // if it wasn't stored, the provider's retry must be able to get through
throw $e;
}
What matters is that checking and marking is a single operation, which is exactly what SET with NX and EX does. If you do a GET and then a SET, two copies of the same event that arrive at the same time both get through. The key includes the store because the same id can repeat across different stores.
Even so, this layer is an optimisation: Redis can lose or expire the key, so the deduplication that guarantees anything is the durable one, with a unique constraint on the event id. What Redis does do is take work off your plate: in my case it discards about a quarter of the events before they use up resources.
Layer 2: a lock per order
The hash doesn't help against two different events for the same order that arrive almost at once. If two workers process them in parallel, each one reads the order's state, modifies it and writes it, and the last one to write wins.
The solution is to serialise the work per order with a distributed lock: before processing, the worker takes that order's lock and, if another worker holds it, it retries a bounded number of times, with some randomness in the wait, or puts the event back on the queue. But the lock coordinates, it doesn't prove correctness. What really prevents losing a write is saving with a conditional write by version, which fails if someone got there first.
// Illustrative example, not the real code
$lockKey = "lock:{$tenantId}:order:{$orderId}";
$token = bin2hex(random_bytes(16)); // unique per acquisition
if (!$redis->set($lockKey, $token, ['NX', 'PX' => 30000])) {
throw new RetryLater(); // another worker is on this order
}
try {
$this->process($event); // the final save is a conditional write by version
} finally {
// Lua script: if get(key) == token then del(key); never GET and then DEL
$redis->eval(self::RELEASE_IF_OWNER, [$lockKey, $token], 1);
}
Three details that matter: the lock expires (it's a lease, with a duration longer than the measured time the work takes), the value is a unique token per acquisition and it's released by comparing that token atomically. Even so, if the worker takes longer than the lease, another one gets in while the first is still writing. That's why the lock doesn't replace the conditional write or business idempotency.
Layer 3: ignore what arrives late
Even if they don't overwrite each other, events can arrive out of order. An ordering guard compares the event's date with the date of the last state you have already applied, discards the one that is older and records that it discarded it.
// Illustrative example, not the real code
// The watermark lives in the order itself
if ($event->emittedAt() < $order->lastEmittedAt()) {
$this->logger->info('webhook.stale_discarded', ['orderId' => $orderId]);
return; // older than what has already been applied
}
// when saving, the same check goes inside the conditional write
You have to decide what you compare. The event's date at the provider, not the time it reached you, because delivery can be delayed. If the provider also gives a version or a sortable id, store it: two updates in the same second tie on date. And the check must travel in the conditional write, not only in a previous if, because between the if and the save another worker may have got ahead.
Layer 4: what a hash can't see
The same refund can arrive by two different routes, with different payloads. To the hash they are two unrelated events, so both get through.
Here I don't deduplicate messages, I apply business idempotency: the key comes from the identity of the refund (the store and the refund id), not from the content or the delivery. The question changes from "have I seen this message?" to "have I already applied this refund?". That key has to live in the database with a unique constraint and store the result, so a retry returns the same thing. In Redis, which expires, it's not enough.
Don't overwrite what the merchant edited by hand
A merchant can edit a document by hand after we have generated it. If another sync event arrives, you don't want to overwrite that edit.
For that I compare two hashes: one of the fiscal fields and one of the non-fiscal fields. That way I see what has really changed between what I have and what arrives, and I don't flatten a manual edit. So that the hash doesn't invent differences, you have to normalise the amounts to the same precision and sort the keys before computing it.
The exchange rate, fixed to the original document
A detail that causes false differences: if an order is in another currency and you recalculate with the exchange rate of each event, the volatility between one event and the next makes the amounts not add up even though nothing has changed.
I fix the exchange rate of the original document across the whole chain of events, so the amount only changes when the order changes.
Summary of the layers
| Layer | What it prevents | Where the guarantee is |
|---|---|---|
| 1. Hash in Redis | Processing an identical event twice | In durable deduplication by event id: Redis only saves work |
| 2. Lock per order | Two workers processing the same order at once | In the conditional write by version |
| 3. Ordering guard | An old event overwriting a new one | In the check inside the conditional write |
| 4. Business idempotency | Applying the same refund twice | In a durable key with a unique constraint |
What's left
None of these pieces is sophisticated on its own. What's hard is having all of them, and having every new incident come in with a test that stops it from happening again; how I automate that cycle is in From a Slack thread to a pull request. With the four layers and those two precautions, a repeated or out-of-order event should not change the order twice.
And since webhooks also get lost, the system has to converge even if none arrives: that takes a periodic reconciliation against the store, which Shopify also recommends.
Top comments (0)