Here are three of the most-integrated webhook producers on the internet, and what each one does when your endpoint is down.
GitHub doesn't retry at all. Its docs say it plainly: "GitHub does not automatically redeliver failed webhook deliveries." You get a failed entry in the delivery log, and you can redeliver it by hand or through the API for the past 3 days. Shopify gives you a one-second connection timeout and five seconds for the whole request, then retries 8 times over 4 hours, and if all of those fail, a subscription created through the Admin API is deleted. Stripe retries in live mode for up to three days with exponential backoff.
Three respected platforms, three completely different contracts. And every consumer integrating with them has to discover the contract by reading docs, or more often by losing an event in production. That's the real problem with webhooks. Not the HTTP call. The promise around the HTTP call is almost never written down, and when it is, it's different everywhere.
So if you're the one building the webhook system, this article is about what you can actually promise, and how to make each promise cheap for the consumer to rely on. Spoiler: "we deliver every event exactly once" isn't on the list.
The Only Honest Guarantee Is At-Least-Once
Let's get this out of the way, because it shapes every other decision.
A webhook delivery is an HTTP request from you to a server you don't control. Picture the failure that matters: the consumer receives the request, writes the order to its database, and then the 200 OK never reaches you. Maybe their load balancer timed out. Maybe your worker was killed mid-read. From where you sit, the delivery failed. From where they sit, it succeeded.
You have two choices. Retry, and they see the event twice. Don't retry, and in every other failure mode (the request never arrived at all) they see it zero times. There's no third option where you know which case you're in. That's the whole reason "exactly-once delivery" over an unreliable network doesn't exist, and the big producers say so. Shopify's docs describe minimizing duplicates and tell you to process webhooks with idempotent operations. Stripe's say that endpoints "might occasionally receive the same event more than once."
So the producer's job isn't to prevent duplicates. It's to make duplicates trivially detectable. That means one thing above all else: a stable event ID that stays the same across every retry of the same event.
The Standard Webhooks spec (more on it later) defines exactly this. Its webhook-id header is "a unique identifier associated with a specific event triggered, and it remains the same no matter how many times a webhook that has failed is retried." Stripe has the event id. Shopify has X-Shopify-Webhook-Id.
The gotcha: one event, two IDs
Here's the part that bites people who think an event ID solves deduplication on its own. Stripe's own best-practices section admits: "In some cases, two separate Event objects are generated and sent." Their advice for catching those is to dedupe on the ID of the object in data.object combined with event.type, not only the event ID.
Shopify has a similar split, for a different reason. If a merchant action triggers deliveries to multiple subscriptions on the same topic, "each has a different X-Shopify-Webhook-Id but shares the same X-Shopify-Event-Id." Webhook ID dedupes individual deliveries; event ID correlates deliveries that came from the same action.
If you're the producer, the lesson is: be precise in your docs about which ID is stable across what. "Unique ID" is ambiguous. "Stable across retries of the same delivery to the same endpoint" is a contract.
The Event You Never Sent: The Producer's Dual Write
Before we get to the consumer side, there's a bug that lives entirely in the producer, and it's the one that causes the scariest kind of webhook failure: the event that was never sent at all.
Look at this handler. It's the first thing almost everyone writes.
// Looks fine. Isn't.
await db.orders.update(orderId, { status: "paid" });
await webhooks.send("order.paid", { orderId });
If the process crashes between those two lines, the order is paid and no webhook ever goes out. No retry schedule can save you, because the retry machinery never heard about the event. Flip the order and you get the opposite bug: a webhook for a payment that got rolled back.
That's a dual write: two systems, no shared transaction. The fix is the transactional outbox. You write the event into a table in the same database transaction as the state change, and a separate dispatcher reads that table and does the HTTP work.
BEGIN;
UPDATE orders SET status = 'paid' WHERE id = $1;
INSERT INTO webhook_outbox (event_id, event_type, payload, created_at)
VALUES (gen_random_uuid(), 'order.paid', $2::jsonb, now());
COMMIT;
Now the event exists if and only if the state change committed. The dispatcher picks up rows, fans them out to each subscribed endpoint, and records every attempt. The event_id generated here is the one that becomes webhook-id on every delivery attempt, forever. gen_random_uuid() is built into PostgreSQL 13 and later.
Notice what this also gives you for free: a durable record of every event you ever emitted. Hold that thought, because it's the foundation for replay later.
Sign the ID and the Timestamp, Not Just the Body
Now the part everyone thinks they've got right: signatures.
The simplest scheme, and the one GitHub uses, is an HMAC-SHA256 of the raw request body with a shared secret, sent as X-Hub-Signature-256: sha256=<hex>. It proves the body came from someone who knows the secret and wasn't modified. That's good. It's also incomplete in two ways.
No timestamp means no replay window. If an attacker ever captures one valid GitHub-style delivery (a leaked log, a debugging proxy, a misconfigured tunnel), that body and signature pair verifies forever. There's nothing in the signed content that expires.
An unsigned ID can be swapped. GitHub does send an X-GitHub-Delivery GUID, but it isn't part of the signed content. If a consumer dedupes on that header, an attacker replaying a captured body can put any fresh GUID they like in it, and the dedupe check passes.
Stripe fixes the first problem. The Stripe-Signature header carries t=<timestamp> and one or more v1=<signature> values, and the signed payload is the timestamp, a ., and the raw body. Tamper with the timestamp and the signature breaks, so an old capture can't be dressed up as fresh. Stripe's libraries reject anything more than 5 minutes off by default.
Standard Webhooks closes the second hole too. The signed content is msg_id.timestamp.payload: the event ID, the attempt timestamp, and the body, joined by full stops. Change any one of the three and verification fails. The ID you dedupe on is now covered by the same signature that proves authenticity.
| Scheme | Signed content | Replay window | Dedupe key signed? |
|---|---|---|---|
GitHub X-Hub-Signature-256
|
body | none | no |
Stripe Stripe-Signature
|
t.body |
timestamp, 5 min default | event ID is in the body, so yes |
Standard Webhooks webhook-signature
|
id.timestamp.body |
timestamp | yes, in the header |
If you're designing from scratch, there's no reason to pick anything weaker than the third row.
Verifying it, in three languages
Here's a complete Standard Webhooks v1 verifier using only each language's standard library. It's worth reading even if you'll use the official standardwebhooks package, because the details below are exactly where hand-rolled verifiers go wrong.
:::tabs
import { createHmac, timingSafeEqual } from "node:crypto";
const TOLERANCE_SECONDS = 5 * 60;
export function verifyWebhook(
rawBody: Buffer,
headers: Record<string, string | undefined>,
secret: string,
): void {
const id = headers["webhook-id"];
const ts = headers["webhook-timestamp"];
const signatures = headers["webhook-signature"];
if (!id || !ts || !signatures) throw new Error("missing webhook headers");
const timestamp = Number.parseInt(ts, 10);
const now = Math.floor(Date.now() / 1000);
if (!Number.isFinite(timestamp) || Math.abs(now - timestamp) > TOLERANCE_SECONDS) {
throw new Error("timestamp outside tolerance");
}
const key = Buffer.from(secret.replace(/^whsec_/, ""), "base64");
const expected = createHmac("sha256", key)
.update(`${id}.${ts}.`)
.update(rawBody)
.digest();
// Space-delimited list: more than one signature appears during secret rotation.
for (const entry of signatures.split(" ")) {
const [version, value] = entry.split(",", 2);
if (version !== "v1" || !value) continue; // ignore unknown schemes
const received = Buffer.from(value, "base64");
// timingSafeEqual throws on a length mismatch, so check first.
if (received.length === expected.length && timingSafeEqual(received, expected)) {
return;
}
}
throw new Error("no matching signature");
}
package webhooks
import (
"crypto/hmac"
"crypto/sha256"
"encoding/base64"
"errors"
"net/http"
"strconv"
"strings"
"time"
)
const tolerance = 5 * time.Minute
func Verify(secret string, h http.Header, body []byte) error {
id, ts, sigs := h.Get("webhook-id"), h.Get("webhook-timestamp"), h.Get("webhook-signature")
if id == "" || ts == "" || sigs == "" {
return errors.New("missing webhook headers")
}
unix, err := strconv.ParseInt(ts, 10, 64)
if err != nil {
return errors.New("bad timestamp")
}
if d := time.Since(time.Unix(unix, 0)); d > tolerance || d < -tolerance {
return errors.New("timestamp outside tolerance")
}
key, err := base64.StdEncoding.DecodeString(strings.TrimPrefix(secret, "whsec_"))
if err != nil {
return err
}
mac := hmac.New(sha256.New, key)
mac.Write([]byte(id + "." + ts + "."))
mac.Write(body)
expected := mac.Sum(nil)
for _, entry := range strings.Split(sigs, " ") {
version, value, ok := strings.Cut(entry, ",")
if !ok || version != "v1" {
continue
}
got, err := base64.StdEncoding.DecodeString(value)
if err == nil && hmac.Equal(got, expected) { // constant-time, length-safe
return nil
}
}
return errors.New("no matching signature")
}
import base64
import hashlib
import hmac
import time
TOLERANCE_SECONDS = 5 * 60
def verify(secret: str, headers: dict[str, str], body: bytes) -> None:
msg_id = headers.get("webhook-id")
ts = headers.get("webhook-timestamp")
signatures = headers.get("webhook-signature")
if not (msg_id and ts and signatures):
raise ValueError("missing webhook headers")
if abs(time.time() - int(ts)) > TOLERANCE_SECONDS:
raise ValueError("timestamp outside tolerance")
key = base64.b64decode(secret.removeprefix("whsec_"))
signed = f"{msg_id}.{ts}.".encode() + body
expected = base64.b64encode(hmac.new(key, signed, hashlib.sha256).digest())
for entry in signatures.split(" "):
version, _, value = entry.partition(",")
if version == "v1" and hmac.compare_digest(value.encode(), expected):
return
raise ValueError("no matching signature")
:::
A few things in there are load-bearing:
-
Raw bytes, not parsed JSON. Re-serialize a parsed body and the whitespace or key order changes, and so does the HMAC. Stripe's docs spell it out: "Any manipulation to the raw body of the request causes the verification to fail." In Express that means mounting
express.raw({ type: "application/json" })on the webhook route instead ofexpress.json(). -
Constant-time comparison. GitHub's docs: "Never use a plain
==operator." Usecrypto.timingSafeEqualin Node,hmac.Equalin Go,hmac.compare_digestin Python. -
Node's
timingSafeEqualthrows on unequal lengths. The Node docs say an error is thrown "ifaandbhave different byte lengths." An attacker sending a short signature turns your 401 into an unhandled exception and a 500, which, depending on the producer, may trigger retries. Go'shmac.Equaljust returnsfalse, so the length check isn't needed there. - Loop over every signature. During secret rotation there's more than one. Stripe lets you keep the old secret active for up to 24 hours and "generates one signature per secret until expiration." A verifier that only checks the first value breaks every rotation.
-
Ignore unknown schemes, never fall back. Stripe tells you to ignore every scheme that isn't
v1"to prevent downgrade attacks." Test events even carry an extra signature under a fakev0scheme, so a verifier that only checksv1is exactly what they expect.
The tolerance-versus-retries trap
This one lives on the producer side, and it's subtle. A 5-minute tolerance window only works if every retry is re-signed with a fresh timestamp. Stripe does this: "If Stripe retries an event ... then we generate a new signature and timestamp for the new delivery attempt." Standard Webhooks bakes it into the definition: the header carries "the timestamp of the attempt," which "may be different to the timestamp of the event."
Now imagine a producer that signs once, at event creation, and stores the signed request in the outbox to resend verbatim. Attempts one and two fail. Attempt three, 5 minutes and 5 seconds after the first on a Standard Webhooks-style schedule, arrives with a timestamp that's already outside the window. Every retry from then on is rejected by a correctly written consumer. The retry schedule becomes decoration.
So: store the event, not the signed request. Sign at send time, every time.
And the dedupe-window trap
The spec's own consumer advice contains a small trap worth calling out. It suggests using webhook-id as an idempotency key, "e.g. save the IDs in redis for 5 minutes."
Five minutes is enough to block a malicious replay, since anything older fails the timestamp check anyway. It is not enough for the duplicate you'll actually see in production: the lost 200. If your consumer processed an event, the ack got dropped, and the producer's next retry lands two hours later, the ID is long gone from a 5-minute cache and you process it again. Keep processed IDs at least as long as the producer's full retry horizon, which for a multi-day schedule means days, not minutes.
A Retry Schedule Is a Promise, So Publish It
The comparison from the opener, laid out next to the Standard Webhooks recommendation:
| Response timeout | Automatic retries | After retries run out | |
|---|---|---|---|
| GitHub | 10 seconds | none; manual redelivery for the past 3 days | failure recorded |
| Shopify | 1 s connect, 5 s total | 8 attempts over 4 hours | Admin API subscription deleted |
| Stripe (live mode) | no number on the webhooks page; "quickly return" a 2xx | up to 3 days, exponential backoff | failed; manual resend available |
| Standard Webhooks (recommended) | 15 to 30 seconds | multi-day, exponential, with jitter | notify the consumer; disable the endpoint |
The Standard Webhooks example schedule is worth copying outright: immediately, then after 5 seconds, 5 minutes, 30 minutes, 2 hours, 5 hours, 10 hours, 14 hours, 20 hours, and 24 hours, for a total of a little over 75 hours. The spec also recommends "some level of random jitter to retries to prevent cases where the failures are due to recurring load caused by the webhook attempts themselves."
// Delay before each retry, per the Standard Webhooks example schedule.
const RETRY_DELAYS_SECONDS = [
5, 5 * 60, 30 * 60, 2 * 3600, 5 * 3600, 10 * 3600, 14 * 3600, 20 * 3600, 24 * 3600,
];
export function nextAttemptAt(failedAttempts: number, now = new Date()): Date | null {
const base = RETRY_DELAYS_SECONDS[failedAttempts - 1];
if (base === undefined) return null; // schedule exhausted: notify and disable
const jitter = 1 + (Math.random() * 0.2 - 0.1); // plus or minus 10%
return new Date(now.getTime() + base * jitter * 1000);
}
My take on the table: a 4-hour horizon is too short for anything a business depends on. A consumer's deploy goes wrong on a Friday evening, nobody notices until Saturday morning, and the subscription is gone. Deleting the subscription is the harshest possible outcome, because it silently turns "temporarily lost events" into "permanently lost events plus a broken integration." Disabling is recoverable. Deletion isn't.
GitHub's no-retry model is at least honest: it tells you up front that recovery is your job, and gives you the delivery log and a redelivery API to do it. That's a defensible contract for a developer tool. It's the wrong one for payments.
Status codes are part of the contract too
The spec's status-code guidance is short and worth adopting as written:
-
2xxis success. Anything else, including timeouts and connection resets, is a failure. -
3xxis a failure. Don't follow redirects; tell the consumer to update their URL. Stripe and Shopify both treat redirects as failures too. -
410 Gonemeans the consumer is done with you. Disable the endpoint. -
429,502, and504mean "slow down." Throttle rather than hammer. - Honour a
retry-afterheader when one comes back.
The 410 rule is the one most producers skip, and it's the cheapest one to implement. It gives consumers a clean, machine-readable way to unsubscribe from a URL they've decommissioned, instead of eating days of retries against a dead route.
Don't Promise Ordering. Give Consumers a Way Not to Need It
Stripe states it outright: "Stripe doesn't guarantee the delivery of events in the order that they're generated." Creating a subscription can emit customer.subscription.created, invoice.created, invoice.paid, and charge.created, and they can arrive in any order.
This isn't laziness. Once you retry with backoff, ordering is gone by construction. Event 1 fails and waits 5 minutes; event 2 succeeds immediately. The only way to keep order is to block every later event for an endpoint behind a failing one, and then one poison message stalls an entire customer's integration for three days.
There's a sneaky gotcha on the consumer side here too. Stripe's snapshot events record created in seconds, "so distinct events can share a timestamp. Don't use created to determine event order or whether you've already processed an event."
So what do you give consumers instead? Two good options:
-
Thin payloads plus fetch-back. Send just the IDs and the event type, and let the consumer fetch the current state from your API. Out-of-order doesn't matter when every event means "go look at the latest." Standard Webhooks lists the advantages: better performance, it's future-proof ("you can always make a thin one full, but not the other way around"), and it keeps data access auditable, because the consumer has to authenticate to read the data. Stripe's newer thin events work this way; its SDK exposes a
fetch_related_object()call to pull the current object. - A version or sequence number on the resource. If you send full payloads, include a monotonically increasing version per resource, and consumers can drop anything older than what they've already stored.
Full payloads aren't wrong. They save the consumer an API call, which matters if your API has tight rate limits. But if you go full, the version field isn't optional. And keep them small; the spec suggests usually under 20kb.
What a Trustworthy Consumer Looks Like
Everything above is aimed at producers, but your docs should hand consumers the receiving side of the contract, because the two halves only work together.
Ack fast, work later. Stripe: return a 2xx "before any complex logic that could cause a timeout." With Shopify's 5-second budget, a handler that calls three downstream services will time out on a bad day, and then the retry piles duplicate load onto the same struggling dependencies. (That amplification loop is the same one covered in Designing for Partial Failure.) Verify, record, enqueue, return.
Make the dedupe record and the side effect atomic. A separate "have I seen this ID?" lookup followed by an insert is a race: two concurrent retries both see "no" and both proceed. Let a unique constraint decide:
BEGIN;
INSERT INTO processed_webhooks (webhook_id, received_at)
VALUES ($1, now())
ON CONFLICT (webhook_id) DO NOTHING
RETURNING webhook_id;
-- If RETURNING produced no row, this is a duplicate: COMMIT and return 200.
-- Otherwise apply the side effect in this same transaction:
UPDATE orders SET status = 'paid' WHERE external_id = $2;
COMMIT;
If the transaction rolls back, the ID isn't recorded, and the retry gets another chance. If it commits, the ID and the effect land together. That's the consumer's version of the producer's outbox: the same idea, pointed the other way.
Return 200 for duplicates. A duplicate is a success from the producer's point of view. Return an error and you've just asked for another retry of something you already handled.
Recovery: The Feature That Makes Everything Else Survivable
Every retry schedule eventually runs out. Consumer outages last longer than 75 hours sometimes. A consumer ships a bug that returns 200 and silently drops a whole event type for a week. None of the mechanics above help with that. What helps is being able to ask the producer, "what did I miss?"
Look at what the big producers offer:
-
Stripe: Resend from the Dashboard for up to 15 days after event creation, or
stripe events resend <event_id> --webhook-endpoint=<endpoint_id>from the CLI for up to 30 days. For bulk recovery, the List Events API takesdelivery_success=falseand returns the events that failed to reach at least one of your endpoints, going back 30 days. - GitHub: Redeliver any delivery from the past 3 days, in the UI or through the REST API.
Standard Webhooks calls this out as "immensely important": give consumers "a way to manually replay specific webhooks or failures within a range in order to recover from long outages without missing message delivery."
This is where the outbox pays off twice. You already have a durable log of every event, keyed by the same ID the consumer dedupes on. A replay endpoint is a query over that table plus a re-send through the normal dispatcher. And because consumers dedupe by webhook-id, "replay everything since Tuesday" is safe. Anything they already processed is a no-op.
My opinion, stated plainly: a webhook system with no replay or event-list API is a lossy system with good manners. The retries make loss rarer. Only replay makes it recoverable.
The Producer's Own Attack Surface: SSRF
One last thing, because it's the security hole that's specific to the sending side. A webhook system is a feature where customers type in a URL and your servers make a POST request to it. Standard Webhooks puts it bluntly: implementations "are especially vulnerable to SSRF as they let their consumers (customers) add any URLs they want, which will be called from the internal webhook system."
Point an endpoint at http://169.254.169.254/ or an internal hostname, and a naive dispatcher will happily make that call from inside your network. Validating the URL at registration time isn't enough either, because DNS can resolve to something different at send time.
Stripe's answer is open source. Smokescreen is an HTTP CONNECT proxy that "proxies most traffic from Stripe to the external world (e.g., webhooks)." It resolves each hostname and ensures it's "a publicly routable IP address and not an internal IP address," which, in their words, "prevents a class of attacks where, for instance, our own webhooks infrastructure is used to scan Stripe's internal network." The spec recommends the same two layers: route webhook traffic through a filtering proxy like that, and run the webhook workers in their own subnet that can't reach internal services.
A side benefit: all egress through one proxy gives you a stable set of source IPs, which enterprise consumers will ask for so they can allowlist you.
Just Use Standard Webhooks
If you're building a new webhook system in 2026, don't invent a signature scheme. Standard Webhooks was announced in December 2023, initiated by Svix, with a technical steering committee that includes people from Zapier, Twilio, Lob, Mux, ngrok, Supabase, and Kong. Its site lists OpenAI, Anthropic, Google Gemini, Twilio, PagerDuty, and Etsy among compatible implementations.
The announcement names the exact problem this article opened with: "For consumers, this means handling webhooks differently for every provider, relearning how to verify webhooks, and encountering gotchas with bespoke implementations."
What you get by adopting it:
- The
webhook-id/webhook-timestamp/webhook-signatureheaders, with the ID and timestamp inside the signature. -
whsec_-prefixed symmetric secrets between 24 and 64 bytes, with space-delimited multi-signature rotation built in. - An optional asymmetric scheme,
v1ausing Ed25519, for when you'd rather consumers hold only a public key. - Reference verification libraries in multiple languages, so your consumers don't write the verifier above by hand.
And if you already have a legacy scheme, the migration path is gentle: add the standard headers alongside your existing ones. The spec notes you can even reuse the same signing secrets for both, so existing integrations don't notice.
The symmetric-versus-asymmetric choice is the one real tradeoff. HMAC is simpler and faster, and every language's standard library has it. Ed25519 means a leaked consumer config can't be used to forge events, because the consumer never holds anything that can sign. For most producers, HMAC with per-endpoint secrets and painless rotation is the right default. Reach for v1a when your consumers are numerous, less trusted, or when a forged event would be expensive.
The Contract, in One Paragraph
Trust in a webhook system isn't a feature of the HTTP call. It's a contract the producer writes down and keeps. Every event gets an ID that never changes across retries. That ID, the attempt timestamp, and the raw body are signed together, and re-signed on every attempt. Events are written to an outbox in the same transaction as the change that caused them. Retries follow a published multi-day schedule with jitter, redirects count as failures, and 410 means stop. Ordering isn't promised, so payloads are thin or versioned. Outbound calls go through an SSRF-filtering proxy. And when all of that still isn't enough, the consumer can ask for everything since Tuesday and replay it safely. None of these ideas are new. Writing all of them down in one place for your consumers is the part most producers skip.
Originally published at andriiboyko.com.


Top comments (0)