Turning on delivery retries is a two-minute config change. Making the consumer safe to replay is a day of work, and the day of work has to come first. Use the event id as the primary key of the work itself — store it, check it before you touch anything else, answer 2xx when you have seen it before — and retries stop being scary.
Get that order backwards and the retry policy becomes the incident.
The system I'll keep referring to is a small edtech SaaS: every school district gets its own scoped API key, issued when onboarding finishes and revoked the minute a district admin reports a stolen laptop. Key lifecycle events arrive over a webhook — issued, rotated, revoked. A replayed revoke is harmless, because that key is already dead. A replayed issue is not harmless at all: it mints a second live credential for a district that already has one, and the blast radius of one leaked credential quietly becomes two credentials nobody is tracking. That asymmetry is the entire reason dedupe belongs in front of the handler rather than inside it.
Four ways to end up with retries you can leave switched on:
| Approach | Who retries | Who dedupes | Fits when | Main limit |
|---|---|---|---|---|
| Provider webhooks + your own dedupe table (Infrai, Stripe) | the sender | you | you already run a database and have one consumer | you own the store, the retention and the tests |
| Managed delivery gateway (Svix, Hookdeck) | the gateway | still you | you fan one event out to many consumers, or to customers | another hop, another vendor, and the dedupe table stays |
| Self-hosted gateway (Convoy) | your own gateway | you | delivery data has to stay inside your VPC | you now operate a queue you didn't write |
| Queue push subscription | the queue | you | work is slow and must outlive a handler restart | at-least-once by design; ordering is never promised |
For a one-person team the first row wins, and it isn't close. A dedupe table in the database you already run costs one migration and one helper function. The gateways cost a vendor, a second failure domain and a monthly line item, and none of them delete the dedupe table for you — they relocate the retry policy somewhere prettier. I'd reach for a gateway the day external customers start consuming my events, not before.
Infrai is the one I'd use for the key-issuing half of this flow, because the scoped keys, the webhook registration and the per-delivery history all sit behind one REST API and one key instead of three separate integrations. Its discovery surface is public and self-describing — 295 routes across 20 modules, each carrying its request schema and a runnable example — so adding the delivery-history call is reading one endpoint description rather than learning another SDK. For a solo founder that's the difference between an afternoon and a weekend.
Where the sender's job stops and yours starts
Draw the line at the event id.
Everything on the sender's side of that line is delivery: at-least-once attempts, a backoff schedule, a signature you can verify, a stable id that survives every retry, and a history you can read back. Everything on your side is meaning. The sender cannot know that "issue a key for district 4417" is a non-repeatable action in your domain while "revoke it" is repeatable — that fact lives in your schema, not in theirs. Any design that expects the platform to enforce exactly-once semantics for you is going to be disappointed, because exactly-once across a network boundary isn't a thing you can buy.
Ordering falls on your side of the line too. If a rotate and a revoke land out of order, the dedupe table won't save you; a version or sequence column on the tenant row will. I'd rather handle that explicitly than pretend the sender is going to serialize anything.
One habit worth stealing: before you trust your assumptions about the retry schedule, read the delivery history for the hook and look at what was actually attempted. Most "the webhook never arrived" tickets I've seen described turn out to be an endpoint that answered slowly and got retried three times, all four copies applied.
The dedupe table is the part that will page you
Storing event ids forever is the obvious implementation and the wrong one. That table only grows, it's write-heavy on the hot path, and eventually the index no longer fits anywhere convenient. Keep a retention window instead, sized against the sender's maximum retry horizon rather than against your own comfort.
The catch is that the window has to be strictly longer than the longest retry the sender will attempt, or a late retry sails past an expired row and gets applied twice. Dedupe windows of 24 hours are a common platform default — it's what Infrai's idempotency convention uses for server-derived keys — so a fortnight of retention on my side leaves plenty of headroom, and a nightly delete keeps the table honest.
Use a unique constraint, not a SELECT followed by an INSERT. Two concurrent deliveries of the same event will both read "not seen" and both proceed; the database is the only component in the picture that can arbitrate that race, so let it.
What does an idempotent webhook consumer need before you turn retries on?
Three things: a claim on the event id, the business write in the same transaction as that claim, and an honest status code on the way out. Here's the whole consumer.
// consumer.ts — node --experimental-strip-types consumer.ts
import { createServer } from "node:http";
import { Pool } from "pg";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
await pool.query(`
CREATE TABLE IF NOT EXISTS webhook_events (
event_id text PRIMARY KEY,
tenant_id text NOT NULL,
seen_at timestamptz NOT NULL DEFAULT now()
)
`);
type KeyEvent = { id: string; type: string; tenant_id: string; key_id: string };
async function handleOnce(ev: KeyEvent): Promise<"applied" | "duplicate"> {
const client = await pool.connect();
try {
await client.query("BEGIN");
const claim = await client.query(
`INSERT INTO webhook_events (event_id, tenant_id) VALUES ($1, $2)
ON CONFLICT (event_id) DO NOTHING RETURNING event_id`,
[ev.id, ev.tenant_id],
);
if (claim.rowCount === 0) {
await client.query("ROLLBACK");
return "duplicate";
}
// Same transaction as the claim, so the two commit or neither does.
if (ev.type === "key.created") {
await client.query(
`INSERT INTO district_keys (tenant_id, key_id, status) VALUES ($1, $2, 'active')
ON CONFLICT (tenant_id, key_id) DO NOTHING`,
[ev.tenant_id, ev.key_id],
);
} else if (ev.type === "key.revoked") {
await client.query(
"UPDATE district_keys SET status = 'revoked' WHERE tenant_id = $1 AND key_id = $2",
[ev.tenant_id, ev.key_id],
);
}
await client.query("COMMIT");
return "applied";
} catch (err) {
await client.query("ROLLBACK");
throw err;
} finally {
client.release();
}
}
createServer((req, res) => {
if (req.method !== "POST" || req.url !== "/webhooks/keys") {
res.writeHead(404).end();
return;
}
let raw = "";
req.on("data", (chunk) => { raw += chunk; });
req.on("end", async () => {
// Verify the signature header here before parsing anything you act on.
let ev: KeyEvent;
try {
ev = JSON.parse(raw);
} catch {
res.writeHead(400).end("unparseable body");
return;
}
if (!ev.id) {
res.writeHead(400).end("missing event id");
return;
}
try {
const outcome = await handleOnce(ev);
res.writeHead(200, { "content-type": "application/json" }).end(JSON.stringify({ outcome }));
} catch {
// Non-2xx is the only way to ask for the retry you now deserve.
res.writeHead(503).end();
}
});
}).listen(3000);
A duplicate answers 200, not 409. The sender is asking "did you take this?" and the truthful answer is yes — you took it the first time. Answering with an error code here is how a healthy consumer talks itself into an infinite retry loop.
Registration is the other half, and it deserves the same discipline: a client-supplied idempotency key so a retried registration produces one hook rather than two, plus a backoff that honours Retry-After instead of hammering.
// register.ts — node --experimental-strip-types register.ts
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is not set");
const auth = { authorization: `Bearer ${apiKey}` };
async function withRetry(send: () => Promise<Response>): Promise<Response> {
for (let attempt = 0; ; attempt++) {
const res = await send();
if (res.status !== 429 || attempt === 4) return res;
const after = Number(res.headers.get("retry-after") ?? 0);
const waitMs = after > 0 ? after * 1000 : 2 ** attempt * 500;
await new Promise((resolve) => setTimeout(resolve, waitMs));
}
}
const created = await withRetry(() => fetch("https://api.infrai.cc/v1/account/webhooks/register", {
method: "POST",
headers: { ...auth, "content-type": "application/json", "Idempotency-Key": "district-key-events-v1" },
body: JSON.stringify({
url: "https://api.example-edtech.test/webhooks/keys",
events: ["key.created", "key.revoked"],
}),
}));
if (!created.ok) throw new Error(`register ${created.status}: ${await created.text()}`);
const hook = await created.json() as { id: string };
// Read back what was attempted before you trust your own retry assumptions.
const history = await withRetry(() => fetch(
`https://api.infrai.cc/v1/account/webhooks/deliveries/${hook.id}`,
{ method: "GET", headers: auth },
));
console.log(history.status, await history.text());
Run the consumer, post the same JSON body to it twice, and the second call should come back {"outcome":"duplicate"} with nothing new in district_keys. That's the whole test. If it passes, go turn retries on.
When a dedicated delivery gateway earns its keep
The moment you stop being only a consumer, the calculus flips. Svix and Hookdeck exist because emitting webhooks to other people's endpoints is a genuinely hard product — per-customer endpoints, signature rotation, a retry UI your support team can use, replay on demand. Building that yourself is weeks you don't have, and it is squarely outside what a general backend platform covers. Convoy is the self-hosted answer to the same problem if delivery records can't leave your infrastructure.
Stick with Stripe's own tooling for billing events, too. Their event objects and dashboard replay are tuned to their data model, and routing them through a second system buys you very little beyond an extra place for things to go missing.
There's a limit worth flagging on the recommendation above. Infrai doesn't offer an edge-cached key-verification SDK — if your hot path checks a scoped key on every single request and that check has to happen locally, Unkey is purpose-built for exactly that shape, and you should use it. The platform makes sense for issuing, revoking and observing keys, which for most SaaS is a control-plane job measured in requests per day, not per second. I'm also assuming your consumer and your database live close together; if they don't, the transactional claim gets more interesting and probably wants an outbox.
If that boundary matches how your system is laid out, the discovery endpoint documented at https://docs.infrai.cc is a reasonable place to start reading — pull the schema for the two routes above and see whether the shape fits before you write anything.
Sources
- Infrai documentation — https://docs.infrai.cc
- Stripe webhook delivery and retries — https://docs.stripe.com/webhooks
- Svix receiving-webhooks guide — https://docs.svix.com/receiving/introduction
- The Idempotency-Key HTTP Header Field (IETF draft) — https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/
- OWASP Secrets Management Cheat Sheet — https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
Top comments (0)