Last month we published a minimal webhook receiver in Go. It ended with a question: how do you handle retries and idempotency? This post is the full answer, in Node.js.
By the end you’ll have a receiver that:
Verifies signatures on the raw request bytes, with a constant-time compare and secret rotation
Acknowledges in milliseconds and never does slow work on the request path
Deduplicates retried deliveries, so a goal is never counted twice
Retries failed processing with exponential backoff and jitter, then parks poison messages in a dead-letter state
Reconciles against the REST API, so a missed webhook doesn’t become a permanently wrong score
Ships with tests and a script that simulates signed deliveries locally
Everything runs on Node 20+ with three dependencies.
Why webhooks (and when to use something else)
A webhook means the provider calls you when something happens, so you never poll. That’s ideal for server-side triggers: notifications, score updates in your database, downstream jobs.
If you need a continuous live stream for a UI, a WebSocket is the better transport, and our 50-line WebSocket scoreboard tutorial shows that side. A common pattern is both: WebSocket for the browser, webhooks for backend logic. Live odds can also arrive over webhooks, as described on the Odds API page.
Read this before you copy the code
I want to be upfront about what is verified and what isn’t, because webhook bugs are expensive.
Topic Status in this tutorial
Webhooks exist and push events to your endpoint Documented
Registering an endpoint needs an API key Documented, see the API reference
Event payload field names (type, match_id, minute, team) Assumed, modeled on the Go tutorial and the documented WebSocket event. Check real payloads in the sandbox
Signature header name, algorithm, encoding Not assumed. All three are environment variables below. Confirm them in the webhooks docs
Retry schedule and delivery guarantees Not assumed. The code is built for the safe worst case: at-least-once delivery, duplicates, and out-of-order events
That last row is the design principle. If your receiver is correct under at-least-once delivery, it’s correct under every stricter guarantee too.
The architecture
text
Provider ──POST──▶ /webhooks/orbistats
│
1. verify signature (raw bytes)
2. parse + validate (zod)
3. INSERT into inbox ← idempotency key = PRIMARY KEY
4. respond 200 immediately
│
▼
inbox table (durable)
│
worker: claim → handle → done
│ failure
├──▶ retry with backoff
└──▶ dead after N attempts
Safety net: REST poll ──▶ compare with local state ──▶ log drift
This is the transactional inbox pattern. The request handler does exactly two things: prove the sender is authentic, and durably write the event. Everything else happens later, off the request path. That one decision removes most webhook failure modes: timeouts, lost events on crash, and double processing.
Step 1: Project setup
bash
mkdir webhook-receiver && cd webhook-receiver
npm init -y
npm install express zod better-sqlite3
mkdir src scripts test
Edit package.json:
json
{
"name": "orbistats-webhook-receiver",
"private": true,
"type": "module",
"engines": { "node": ">=20.6" },
"scripts": {
"start": "node --env-file=.env src/index.js",
"send": "node --env-file=.env scripts/send-test-event.js",
"test": "node --test"
},
"dependencies": {
"better-sqlite3": "^11.0.0",
"express": "^4.21.0",
"zod": "^3.23.0"
}
}
(Versions are those current at writing; use the latest compatible releases.) Node 20.6+ supports --env-file, so we don’t need dotenv.
Create .env.example and copy it to .env:
bash
PORT=8080
DB_PATH=webhooks.db
Comma-separated: lets you rotate secrets without downtime
WEBHOOK_SECRET=change-me
CONFIRM THESE THREE in the webhooks docs. Do not assume.
WEBHOOK_SIGNATURE_HEADER=x-signature
WEBHOOK_SIGNATURE_ENCODING=hex
WEBHOOK_SIGNATURE_PREFIX=
Optional: header carrying a unique delivery/event ID, if the docs define one
WEBHOOK_DELIVERY_ID_HEADER=
MAX_ATTEMPTS=6
ORBISTATS_API_KEY=
Step 2: Configuration
src/config.js:
javascript
const required = (name) => {
const v = process.env[name];
if (!v) throw new Error(Missing env var ${name});
return v;
};
export const config = {
port: Number(process.env.PORT ?? 8080),
dbPath: process.env.DB_PATH ?? "webhooks.db",
secrets: required("WEBHOOK_SECRET").split(",").map((s) => s.trim()),
signatureHeader: (process.env.WEBHOOK_SIGNATURE_HEADER ?? "x-signature").toLowerCase(),
signatureEncoding: process.env.WEBHOOK_SIGNATURE_ENCODING ?? "hex", // "hex" | "base64"
signaturePrefix: process.env.WEBHOOK_SIGNATURE_PREFIX ?? "", // e.g. "sha256="
deliveryIdHeader: (process.env.WEBHOOK_DELIVERY_ID_HEADER ?? "").toLowerCase(),
maxAttempts: Number(process.env.MAX_ATTEMPTS ?? 6),
apiKey: process.env.ORBISTATS_API_KEY ?? "",
};
Failing at startup on a missing secret is deliberate. A receiver that silently runs without verification is worse than one that refuses to boot.
Step 3: The inbox table
src/db.js:
javascript
import Database from "better-sqlite3";
import { config } from "./config.js";
export const db = new Database(config.dbPath);
db.pragma("journal_mode = WAL");
db.exec(`
CREATE TABLE IF NOT EXISTS inbox (
idempotency_key TEXT PRIMARY KEY,
received_at INTEGER NOT NULL,
event_type TEXT,
match_id TEXT,
payload TEXT NOT NULL,
status TEXT NOT NULL DEFAULT 'pending', -- pending | done | dead
attempts INTEGER NOT NULL DEFAULT 0,
next_attempt_at INTEGER NOT NULL,
last_error TEXT
);
CREATE INDEX IF NOT EXISTS idx_inbox_due ON inbox (status, next_attempt_at);
-- Domain table: one row per goal event, keyed so duplicates are harmless
CREATE TABLE IF NOT EXISTS goals (
idempotency_key TEXT PRIMARY KEY,
match_id TEXT NOT NULL,
team TEXT,
minute INTEGER
);
`);
Notice the idempotency key is a primary key. The database, not your application code, guarantees uniqueness, so two concurrent retries can’t both win a race.
SQLite keeps this tutorial self-contained. The same schema works in Postgres, and I show the multi-worker version later.
Step 4: Signature verification (on the raw bytes)
This is where most receivers go wrong, so slow down.
An HMAC signature is computed over the exact bytes the sender transmitted. If you let a JSON middleware parse the body first and then re-serialize it, key order and whitespace can change and every signature will fail (or, worse, you’ll “fix” it by skipping verification). We capture the raw Buffer and verify that.
src/signature.js:
javascript
import crypto from "node:crypto";
import { config } from "./config.js";
export function computeSignature(rawBody, secret, encoding = config.signatureEncoding) {
return crypto.createHmac("sha256", secret).update(rawBody).digest(encoding);
}
export function verifySignature(rawBody, headerValue) {
if (!headerValue) return false;
const received =
config.signaturePrefix && headerValue.startsWith(config.signaturePrefix)
? headerValue.slice(config.signaturePrefix.length)
: headerValue;
const receivedBuf = Buffer.from(received.trim());
// Try every configured secret so rotation never drops events
return config.secrets.some((secret) => {
const expected = Buffer.from(computeSignature(rawBody, secret));
return (
expected.length === receivedBuf.length &&
crypto.timingSafeEqual(expected, receivedBuf)
);
});
}
Three details matter here:
timingSafeEqual, not ===. Plain string comparison returns early at the first mismatch, which leaks timing information. timingSafeEqual throws if lengths differ, so we check length first.
Raw bytes in, always. The function takes a Buffer, never a parsed object.
Multiple secrets. During rotation you deploy WEBHOOK_SECRET=new,old, switch the sender, then drop the old one. Zero downtime.
If the docs say the signature covers a timestamp plus the body (many providers do this for replay protection), build the signed string accordingly and reject deliveries older than a few minutes:
javascript
// Only if the docs define a timestamp header. Illustrative.
const age = Math.abs(Date.now() / 1000 - Number(timestampHeader));
if (!Number.isFinite(age) || age > 300) return false;
Step 5: Tolerant payload validation
src/schema.js:
javascript
import { z } from "zod";
export const eventSchema = z
.object({
id: z.union([z.string(), z.number()]).optional(),
type: z.string().optional(),
event: z.string().optional(), // some payloads name it event
match_id: z.union([z.string(), z.number()]).transform(String),
minute: z.number().int().optional(),
team: z.string().optional(),
timestamp: z.string().optional(),
})
.passthrough() // keep unknown fields; providers add them over time
.refine((e) => e.type || e.event, { message: "payload needs type or event" });
export const eventTypeOf = (e) => e.type ?? e.event;
Two choices worth defending:
.passthrough(): APIs add fields without notice. A strict schema turns a harmless addition into an outage.
Accepting type or event: until you’ve inspected real payloads, being liberal in what you accept costs nothing.
Step 6: The HTTP server and the ACK rules
src/server.js:
javascript
import express from "express";
import crypto from "node:crypto";
import { config } from "./config.js";
import { db } from "./db.js";
import { verifySignature } from "./signature.js";
import { eventSchema, eventTypeOf } from "./schema.js";
const insertInbox = db.prepare();
INSERT INTO inbox (idempotency_key, received_at, event_type, match_id, payload, next_attempt_at)
VALUES (@key, @now, @type, @matchId, @payload, @now)
ON CONFLICT(idempotency_key) DO NOTHING
const insertDead = db.prepare();
INSERT INTO inbox (idempotency_key, received_at, payload, status, next_attempt_at, last_error)
VALUES (@key, @now, @payload, 'dead', @now, @err)
ON CONFLICT(idempotency_key) DO NOTHING
export function idempotencyKey(req, rawBody, parsed) {
// 1. A delivery/event ID header, if the docs define one
if (config.deliveryIdHeader) {
const h = req.get(config.deliveryIdHeader);
if (h) return hdr:${h};
}
// 2. An ID inside the payload
if (parsed?.id !== undefined) return id:${parsed.id};
// 3. Fallback: hash of the exact bytes (retries resend identical bodies)
return sha:${crypto.createHash("sha256").update(rawBody).digest("hex")};
}
export function createApp() {
const app = express();
app.disable("x-powered-by");
app.get("/healthz", (_req, res) => res.json({ ok: true }));
app.post(
"/webhooks/orbistats",
express.raw({ type: "/", limit: "256kb" }), // raw Buffer, not parsed JSON
(req, res) => {
const rawBody = req.body;
if (!Buffer.isBuffer(rawBody) || rawBody.length === 0) {
return res.status(400).json({ error: "empty body" });
}
// 1. Authenticity first. Nothing below runs for forged requests.
if (!verifySignature(rawBody, req.get(config.signatureHeader))) {
return res.status(401).json({ error: "invalid signature" });
}
// 2. Parse and validate
const text = rawBody.toString("utf8");
let parsed;
let parseError;
try {
const result = eventSchema.safeParse(JSON.parse(text));
if (result.success) parsed = result.data;
else parseError = result.error.message;
} catch (e) {
parseError = `invalid JSON: ${e.message}`;
}
const key = idempotencyKey(req, rawBody, parsed);
const now = Date.now();
try {
// 3. Authentic but unusable: keep it for inspection, but still ACK
if (!parsed) {
insertDead.run({ key, now, payload: text, err: parseError });
return res.status(200).json({ status: "stored_unparsed" });
}
// 4. Durable write. `changes` is 0 when the key already exists.
const info = insertInbox.run({
key,
now,
type: eventTypeOf(parsed),
matchId: parsed.match_id,
payload: text,
});
return res.status(200).json({ status: info.changes === 1 ? "accepted" : "duplicate" });
} catch (err) {
// We could not persist: tell the sender to retry later
console.error(JSON.stringify({ level: "error", msg: "inbox write failed", err: err.message }));
return res.status(500).json({ error: "temporary failure" });
}
}
);
return app;
}
The status-code rules (this is the heart of reliable webhooks)
Situation Respond Why
Bad or missing signature 401 Not from the provider. Retrying won’t help and shouldn’t
Authentic event, written to inbox 200 You own it now. Process later
Authentic duplicate 200 The sender retried because it missed your last ACK. Acknowledge so retries stop
Authentic event you can’t parse or don’t recognize 200, store it Rejecting authentic traffic invites pointless retries and can get your endpoint disabled
You failed to write to the database 500 Genuinely temporary. You want a retry
The most common mistake is returning an error for a duplicate. The sender is retrying precisely because it didn’t see your 200. Answering the retry with an error just extends the loop.
Also notice what’s absent: no business logic in the handler. Slow handlers cause timeouts, timeouts cause retries, and retries cause duplicates. Keeping the request path to “verify, write, ACK” breaks that chain.
Step 7: Idempotent handlers
src/handlers.js:
javascript
import { db } from "./db.js";
const insertGoal = db.prepare();
INSERT INTO goals (idempotency_key, match_id, team, minute)
VALUES (?, ?, ?, ?)
ON CONFLICT(idempotency_key) DO NOTHING
export const scoreboard = db.prepare();
SELECT team, COUNT(*) AS goals FROM goals WHERE match_id = ? GROUP BY team
// event type -> handler. Add more as you inspect real payloads.
const handlers = {
"match.goal": (event, key) => {
insertGoal.run(key, event.match_id, event.team ?? null, event.minute ?? null);
},
};
export async function handleEvent(event, key) {
const handler = handlers[event.type];
if (!handler) {
console.log(JSON.stringify({ level: "info", msg: "no handler", type: event.type }));
return; // unknown types are not errors
}
await handler(event, key);
}
Two principles at work:
The handler is idempotent too. Even if the inbox dedupe somehow failed, ON CONFLICT DO NOTHING on the goal row makes a second run harmless. Defense in depth: assume each layer can fail.
Store events, derive state. We don’t keep a mutable “score” counter that we += 1. We store goal rows and compute the score with a query. If a correction ever arrives (a goal disallowed after review), you delete or void one row and the score fixes itself. Patching counters is how scores drift permanently.
Step 8: The worker, with retries and a dead-letter state
src/worker.js:
javascript
import { db } from "./db.js";
import { config } from "./config.js";
import { eventSchema, eventTypeOf } from "./schema.js";
import { handleEvent } from "./handlers.js";
const due = db.prepare();
SELECT idempotency_key, payload, attempts FROM inbox
WHERE status = 'pending' AND next_attempt_at <= ?
ORDER BY received_at
LIMIT 20
const markDone = db.prepare(UPDATE inbox SET status = 'done', last_error = NULL WHERE idempotency_key = ?);
const markRetry = db.prepare(UPDATE inbox SET attempts = ?, next_attempt_at = ?, last_error = ? WHERE idempotency_key = ?);
const markDead = db.prepare(UPDATE inbox SET status = 'dead', attempts = ?, last_error = ? WHERE idempotency_key = ?);
// 2s, 4s, 8s ... capped at 5 min, with jitter so retries don't synchronize
export function backoffMs(attempt) {
const base = Math.min(2 ** attempt * 1000, 5 * 60_000);
return base / 2 + Math.random() * (base / 2);
}
export async function drainOnce(now = Date.now()) {
const rows = due.all(now);
for (const row of rows) {
try {
const event = eventSchema.parse(JSON.parse(row.payload));
await handleEvent({ ...event, type: eventTypeOf(event) }, row.idempotency_key);
markDone.run(row.idempotency_key);
} catch (err) {
const attempts = row.attempts + 1;
const msg = String(err.message).slice(0, 500);
if (attempts >= config.maxAttempts) {
markDead.run(attempts, msg, row.idempotency_key);
console.error(JSON.stringify({ level: "error", msg: "event dead-lettered", key: row.idempotency_key, err: msg }));
} else {
markRetry.run(attempts, now + backoffMs(attempts), msg, row.idempotency_key);
}
}
}
return rows.length;
}
let running = false;
export function startWorker(intervalMs = 500) {
const timer = setInterval(async () => {
if (running) return; // never overlap runs
running = true;
try {
while ((await drainOnce()) > 0) { /* keep draining while there's work */ }
} catch (e) {
console.error("worker error", e);
} finally {
running = false;
}
}, intervalMs);
return () => clearInterval(timer);
}
What this gives you:
Exponential backoff with jitter: failures wait 1 to 2s, then 2 to 4s, then longer. Jitter prevents a thundering herd after an outage.
A bounded number of attempts: after MAX_ATTEMPTS, the event becomes dead. A poison message can’t loop forever or block the queue behind it.
Visibility: last_error and attempts tell you exactly what went wrong and how many times.
Dead letters are not a trash can. Alert on them. A growing dead count means your handler or your assumptions about payloads are broken, and you’ll want to fix the bug and replay those rows by setting them back to pending.
Step 9: Reconciliation, the safety net
Here’s an uncomfortable truth about webhooks: any push system can miss an event. Your server was down past the retry window, a network partition ate a delivery, a deploy dropped a connection. If a goal is missed and you never check, your scoreboard stays wrong forever.
The fix is cheap: periodically compare your local view against the REST API. We use the live-matches endpoint from the quickstart pattern (GET /v1/{sport}/matches/live).
src/reconcile.js:
javascript
import { config } from "./config.js";
import { scoreboard } from "./handlers.js";
const BASE = "https://api.orbistats.com/v1";
export async function reconcileLiveMatches(sport = "football") {
const res = await fetch(${BASE}/${sport}/matches/live, {
headers: { Authorization: Bearer ${config.apiKey} },
signal: AbortSignal.timeout(10_000),
});
if (!res.ok) throw new Error(live endpoint returned ${res.status});
const body = await res.json();
const matches = Array.isArray(body) ? body : body.data ?? [];
const drift = [];
for (const m of matches) {
const local = Object.fromEntries(
scoreboard.all(String(m.match_id)).map((r) => [r.team, r.goals])
);
const gotHome = local[m.home?.name] ?? 0;
const gotAway = local[m.away?.name] ?? 0;
if (gotHome !== m.home?.score || gotAway !== m.away?.score) {
drift.push({
match_id: m.match_id,
api: { home: m.home?.score, away: m.away?.score },
webhook_view: { home: gotHome, away: gotAway },
});
}
}
return drift;
}
A few honest caveats:
The response shape here follows the example on the Orbistats homepage (match_id, home.score, away.score). Confirm it in the API reference, because the code accepts either a bare array or { data: [...] }.
If you start your receiver mid-match, drift is expected, since you missed earlier goals. Reconciliation is for detecting trouble.
Own goals can attribute differently (“team that scored” vs “team credited”), so the first time drift appears, check whether it’s real before panicking.
In production, don’t just log drift. Repair it by re-fetching the match’s event list and rebuilding rows.
Reconciliation also covers a failure mode no amount of signature checking can: the provider’s webhook for an event simply never arriving. The Sports Data API and Live Scores API are the pull-based sources you reconcile against.
Step 10: Wire it together
src/index.js:
javascript
import { createApp } from "./server.js";
import { startWorker } from "./worker.js";
import { reconcileLiveMatches } from "./reconcile.js";
import { config } from "./config.js";
const app = createApp();
const server = app.listen(config.port, () =>
console.log(JSON.stringify({ level: "info", msg: "listening", port: config.port }))
);
const stopWorker = startWorker();
// Safety net: only runs if you provided an API key
if (config.apiKey) {
setInterval(async () => {
try {
const drift = await reconcileLiveMatches();
if (drift.length) console.warn(JSON.stringify({ level: "warn", msg: "score drift", drift }));
} catch (e) {
console.error("reconcile failed:", e.message);
}
}, 60_000);
}
for (const sig of ["SIGINT", "SIGTERM"]) {
process.on(sig, () => {
stopWorker();
server.close(() => process.exit(0)); // finish in-flight requests, then exit
});
}
Run it:
bash
npm start
curl localhost:8080/healthz
Step 11: Simulate signed deliveries locally
You shouldn’t need a live match (or your provider) to test this. scripts/send-test-event.js signs a payload exactly the way your receiver expects and sends it twice:
javascript
import crypto from "node:crypto";
const secret = process.env.WEBHOOK_SECRET;
const url = process.env.TARGET_URL ?? "http://localhost:8080/webhooks/orbistats";
const header = (process.env.WEBHOOK_SIGNATURE_HEADER ?? "x-signature").toLowerCase();
const encoding = process.env.WEBHOOK_SIGNATURE_ENCODING ?? "hex";
const prefix = process.env.WEBHOOK_SIGNATURE_PREFIX ?? "";
const event = {
id: process.argv[2] ?? evt_${Date.now()},
type: "match.goal",
match_id: "48213",
minute: 72,
team: "Manchester City",
timestamp: new Date().toISOString(),
};
const body = JSON.stringify(event);
const signature = prefix + crypto.createHmac("sha256", secret).update(body).digest(encoding);
async function send(label) {
const res = await fetch(url, {
method: "POST",
headers: { "content-type": "application/json", [header]: signature },
body,
});
console.log(label, res.status, await res.text());
}
await send("first ");
await send("replay"); // same id, so this must come back as a duplicate
bash
npm run send -- evt_1
first 200 {"status":"accepted"}
replay 200 {"status":"duplicate"}
Then check the database:
bash
sqlite3 webhooks.db "SELECT status, COUNT(*) FROM inbox GROUP BY status;"
sqlite3 webhooks.db "SELECT * FROM goals;"
You should see one inbox row and one goal row, even though you sent two deliveries.
Step 12: Automated tests
test/webhook.test.js (uses Node’s built-in test runner, no extra dependencies):
javascript
import { test, before, after } from "node:test";
import assert from "node:assert/strict";
import crypto from "node:crypto";
// Env must be set BEFORE the modules load
process.env.WEBHOOK_SECRET = "test-secret";
process.env.DB_PATH = ":memory:";
const { createApp } = await import("../src/server.js");
const { drainOnce, backoffMs } = await import("../src/worker.js");
const { db } = await import("../src/db.js");
let server;
let base;
before(async () => {
server = createApp().listen(0);
await new Promise((r) => server.once("listening", r));
base = http://127.0.0.1:${server.address().port};
});
after(() => server.close());
const sign = (body) => crypto.createHmac("sha256", "test-secret").update(body).digest("hex");
const post = (body, sig = sign(body)) =>
fetch(${base}/webhooks/orbistats, {
method: "POST",
headers: { "content-type": "application/json", "x-signature": sig },
body,
});
const evt = (id, type = "match.goal") =>
JSON.stringify({ id, type, match_id: "1", minute: 10, team: "A" });
const count = (sql) => db.prepare(sql).get().n;
test("rejects a bad signature", async () => {
const res = await post(evt("e1"), "deadbeef");
assert.equal(res.status, 401);
});
test("accepts a valid event once and ACKs duplicates", async () => {
const body = evt("e2");
assert.equal((await (await post(body)).json()).status, "accepted");
assert.equal((await (await post(body)).json()).status, "duplicate");
assert.equal(count("SELECT COUNT(*) AS n FROM inbox WHERE idempotency_key = 'id:e2'"), 1);
});
test("processing is idempotent", async () => {
await post(evt("e3"));
await post(evt("e3"));
await drainOnce();
assert.equal(count("SELECT COUNT(*) AS n FROM goals WHERE idempotency_key = 'id:e3'"), 1);
});
test("unknown event types are ACKed and completed, not errored", async () => {
const res = await post(evt("e4", "match.something_new"));
assert.equal(res.status, 200);
await drainOnce();
assert.equal(count("SELECT COUNT(*) AS n FROM inbox WHERE idempotency_key = 'id:e4' AND status = 'done'"), 1);
});
test("authentic but malformed payloads are stored, not rejected", async () => {
const body = "not json at all";
const res = await post(body);
assert.equal(res.status, 200);
assert.equal(count("SELECT COUNT(*) AS n FROM inbox WHERE status = 'dead'"), 1);
});
test("backoff grows with attempts", () => {
assert.ok(backoffMs(1) < backoffMs(5));
});
bash
npm test
These six tests encode the rules from Step 6. If anyone later “simplifies” the handler into returning 409 for duplicates, a test fails.
Step 13: Going live without wasting your access window
When you’re ready to receive real events:
Expose your local server over HTTPS with a tunnel (ngrok or Cloudflare Tunnel), or deploy it.
Get an API key. Request access and follow the activation steps in the quickstart.
Inspect a real payload in the sandbox, then adjust schema.js and your handlers to match.
Confirm the three signature settings (header, encoding, prefix) in the webhooks docs, and put them in .env.
Register your public URL (https://your-domain/webhooks/orbistats) using the registration steps in the API reference.
One practical tip: Orbistats access works in time windows. The Free plan is one hour with no card, and the timer starts when you open your activation link. Starter is 24 hours and Growth is five days, with every sport and endpoint on every plan. Current details are on the pricing page.
So build and test everything above locally with simulated, signed events before you start the clock. Steps 1 to 12 need no live data. Then spend the window verifying real payloads and signatures, and let a live match prove the pipeline.
Also bookmark the status page. When events stop arriving, it’s the fastest way to learn whether the problem is yours or the provider’s.
Step 14: Production hardening checklist
The tutorial version is correct. A production version is also operable:
Move the worker to its own process so slow handlers can’t affect request latency. better-sqlite3 is synchronous, so heavy handlers block the event loop in a single process.
Use Postgres for multiple workers. Replace SQLite’s claim step with a row-locking claim:
sql
UPDATE inbox
SET status = 'processing', locked_at = now()
WHERE idempotency_key IN (
SELECT idempotency_key FROM inbox
WHERE status = 'pending' AND next_attempt_at <= now()
ORDER BY received_at
LIMIT 20
FOR UPDATE SKIP LOCKED
)
RETURNING *;
Add a lease timeout so rows stuck in processing after a crash return to pending.
Monitor three numbers: pending depth, age of the oldest pending row, and dead-letter count.
sql
SELECT status, COUNT(*) FROM inbox GROUP BY status;
SELECT (strftime('%s','now')*1000 - MIN(received_at)) / 1000 AS oldest_pending_seconds
FROM inbox WHERE status = 'pending';
Alert on silence. If a match is live and zero events have arrived for several minutes, something is wrong even if no errors are firing. Reconciliation catches this.
Never log secrets or full signatures. Log the idempotency key and event type instead.
Terminate TLS properly, apply rate limiting at your proxy, and keep the body size limit small.
Retain the inbox for a while (days, not forever). It’s your audit trail and your replay source.
Handle match-day spikes with a queue and caching in front of your readers. Our post on scaling a sports data consumer for match-day spikes covers that side.
Common pitfalls (each one has bitten someone)
Symptom Likely cause Fix
Every signature fails A JSON parser ran before verification Use express.raw on this route and verify the Buffer
Signature fails only sometimes You re-serialized parsed JSON Verify the original bytes, never JSON.stringify(req.body)
Same goal counted twice Dedupe in memory, or only in the handler Unique key in the database, enforced on insert
Provider keeps retrying forever You return non-2xx for duplicates or unknown types ACK anything authentic that you’ve stored
Slow responses, then more duplicates Business logic on the request path Write to the inbox, ACK, process in the worker
Scores drift over a week Mutable counters and no reconciliation Store events, derive state, reconcile periodically
Endpoint suddenly disabled by sender Sustained 4xx/5xx or timeouts Fix the cause, monitor, and see the retry/disable policy in the docs
Duplicates after a deploy In-memory dedupe lost on restart Persist idempotency keys
Wrapping up
Reliable webhook handling comes down to a handful of habits:
Verify authenticity on the raw bytes, with a constant-time compare.
Acknowledge fast and keep the request path to verify, write, ACK.
Make duplicates harmless with a database-enforced idempotency key, in the inbox and in the handler.
Retry with backoff and jitter, cap the attempts, and alert on dead letters.
Store events, derive state, and reconcile against a pull API so a missed push can’t become a permanent error.
If you want the wider picture of how REST, WebSocket and webhooks fit together, our sports odds API guide walks through delivery methods and a production checklist.
Question for the comments: for your idempotency key, do you trust the provider’s event ID, or do you hash the body as a fallback? And have you ever been bitten by a webhook that never arrived? I’d like to hear how you caught it.
Top comments (1)
tr.ee/dev-to