A webhook retry policy is safe only after the consumer can authenticate a delivery, atomically claim its event ID inside the correct account scope, and return the recorded result for every duplicate. Short answer: store one durable receipt per (consumer_scope, event_id) before acknowledging success, put the business work in the same local transaction, and test duplicates plus crashes before enabling retries. A global event-ID set is too broad for a multi-account platform; an in-memory set disappears at exactly the wrong time.
The leaked-key drill makes this concrete. Revoke one credential, rotate it, replay captured deliveries, and verify that the affected integration is bounded to one account while legitimate redeliveries still converge on one business result. Authentication must happen before deduplication, or an unauthenticated request can reserve a real event ID and suppress the valid delivery that follows.
What does the replay ledger actually cost?
The bill is mostly the durable state retained per accepted delivery: the scoped event ID, processing state, timestamps, a small result reference, and index overhead. Payload archives can dominate that term if the consumer keeps entire request bodies in the hot deduplication table, so do the arithmetic before selecting a retention period.
Use a capacity equation instead of a storage slogan:
def retained_bytes(events_per_day: int, retention_days: int, bytes_per_receipt: int) -> int:
return events_per_day * retention_days * bytes_per_receipt
# Planning assumptions, not benchmark results.
example = retained_bytes(10_000_000, 30, 320)
print(example) # 96,000,000,000 bytes before replicas and backups
That 96 GB decimal figure follows only from the stated assumptions. It is not a promise about a database engine, because row headers, indexes, replication, backup policy, compression, and allocator behavior change the physical footprint. Measure the actual table and index sizes under representative event IDs. The dominant lever in this example is retention: halving retained days halves the logical receipt bytes, while shaving a few bytes from a status field barely moves the total.
| Choice | Hot-state effect | Failure-mode consequence |
|---|---|---|
| Keep full payloads with receipts | Largest | Easier forensic reconstruction, but more sensitive material and a wider storage footprint |
| Keep event ID, scope, state, and result reference | Bounded per delivery | Enough for deduplication; investigation depends on separate logs or business records |
| Delete receipts at the retry horizon | Stops unbounded growth | A very late replay can execute again |
I would deliberately stop keeping full webhook bodies in the replay ledger once validation and dispatch finish. I would expire compact receipts only after the sender's maximum retry window, operational repair window, and clock-skew allowance have all passed. The cost of deletion is real: after expiry, the ledger can no longer prove that an old event ran, and a replay may become new work.
Why isn't an event ID alone enough?
Because uniqueness belongs to a namespace. Two independent producers can both emit evt_42; two customer accounts can also be connected to the same producer. A primary key on event_id alone lets one account suppress another account's work, which turns a collision or hostile replay into a cross-tenant denial of service.
Scope first.
Use a stable consumer scope derived from the authenticated account and integration, not from user-controlled body fields. During credential rotation, both the retiring and replacement keys should resolve to that same scope. Otherwise the same delivery accepted under the new key bypasses a ledger entry written under the old key. The secret authenticates a caller; it should not redefine the business namespace whenever it rotates.
This order matters:
- Read the raw body with a strict size limit.
- Locate the credential by a non-secret identifier, then verify the signature and freshness policy.
- Resolve the credential to an account and integration scope.
- Validate the event ID and event type.
- Atomically insert the receipt and durable work item.
- Acknowledge only after that transaction commits.
A duplicate that finds a completed receipt gets the previously recorded outcome. One that finds pending work should receive a response consistent with the retry contract while the original worker continues. Do not hold the incoming connection open while waiting for slow business work merely to make the handler look synchronous.
The smallest durable implementation
The example uses SQLite to expose the transaction boundary without tying the design to a hosted product. In a horizontally scaled service, use a transactional database shared by every consumer instance and preserve the same composite uniqueness constraint. The handler assumes signature verification has already returned a trusted consumer_scope; algorithms and header formats are sender-specific and should not be guessed in generic code.
import json
import sqlite3
from dataclasses import dataclass
from datetime import datetime, timezone
@dataclass(frozen=True)
class AcceptedEvent:
consumer_scope: str
event_id: str
event_type: str
payload: dict
def utc_now() -> str:
return datetime.now(timezone.utc).isoformat()
def initialize(connection: sqlite3.Connection) -> None:
connection.executescript(
"""
CREATE TABLE IF NOT EXISTS webhook_receipts (
consumer_scope TEXT NOT NULL,
event_id TEXT NOT NULL,
status TEXT NOT NULL CHECK (status IN ('pending', 'complete', 'failed')),
accepted_at TEXT NOT NULL,
completed_at TEXT,
result_ref TEXT,
PRIMARY KEY (consumer_scope, event_id)
);
CREATE TABLE IF NOT EXISTS webhook_jobs (
consumer_scope TEXT NOT NULL,
event_id TEXT NOT NULL,
event_type TEXT NOT NULL,
payload_json TEXT NOT NULL,
created_at TEXT NOT NULL,
PRIMARY KEY (consumer_scope, event_id),
FOREIGN KEY (consumer_scope, event_id)
REFERENCES webhook_receipts (consumer_scope, event_id)
);
"""
)
def accept(connection: sqlite3.Connection, event: AcceptedEvent) -> str:
try:
with connection:
connection.execute(
"""
INSERT INTO webhook_receipts
(consumer_scope, event_id, status, accepted_at)
VALUES (?, ?, 'pending', ?)
""",
(event.consumer_scope, event.event_id, utc_now()),
)
connection.execute(
"""
INSERT INTO webhook_jobs
(consumer_scope, event_id, event_type, payload_json, created_at)
VALUES (?, ?, ?, ?, ?)
""",
(
event.consumer_scope,
event.event_id,
event.event_type,
json.dumps(event.payload, separators=(",", ":"), sort_keys=True),
utc_now(),
),
)
return "accepted"
except sqlite3.IntegrityError:
row = connection.execute(
"""
SELECT status FROM webhook_receipts
WHERE consumer_scope = ? AND event_id = ?
""",
(event.consumer_scope, event.event_id),
).fetchone()
if row is None:
raise
return f"duplicate:{row[0]}"
The receipt and job enter durable storage together. If the process dies before commit, neither exists and a retry can claim the event. If it dies after commit, the durable job remains available to a worker. The uniqueness constraint resolves concurrent deliveries; an application-level SELECT followed by INSERT does not, because both requests can observe absence before either inserts.
There is still a hard edge. If a worker calls an external system and crashes before marking the receipt complete, the local database cannot know whether that remote side effect happened. Carry the event ID as the downstream idempotency key when the downstream interface supports it; otherwise reconcile against a stable business key, or accept and document at-least-once effects. A local replay ledger cannot manufacture an atomic transaction across an unrelated service.
This design has limitations. It consumes durable storage, adds a database dependency to acceptance, and guarantees no more than its retention window. It is a poor fit when events have no stable producer-assigned identity, when every effect occurs in an external system that offers neither an idempotency key nor a queryable business key, or when the consumer cannot tolerate the write latency of a durable claim. In those cases, fix the event contract or choose an explicitly reconciled at-least-once workflow; shortening the table name will not repair the missing boundary.
Which failures will retries amplify?
Start with concurrent duplicates, because sequential tests miss the race. Send the same authenticated event through multiple consumer instances and require one receipt, one job, and one business transition. Then terminate a worker before local commit and after local commit. Those two cuts should produce different recovery paths but the same final business state.
Test malformed IDs, oversized bodies, stale signed requests, an unknown credential identifier, and a valid signature attached to the wrong account route. These requests must not reserve ledger keys. Also test a producer reusing an ID with a different body. Retaining a digest or immutable event metadata long enough to flag the conflict is safer; silently treating the second body as an ordinary duplicate hides producer corruption.
Retries need bounds. Use exponential delay with jitter, cap the number or age of attempts, and move persistently failing work to a reviewable terminal state. Those are policy choices, so emit attempt count, event age, scope, state transition, and a privacy-safe event identifier into metrics or structured logs. Never put the signing secret or full sensitive payload into those records.
The sharpest trap is returning success before durable commit. The sender then has no reason to retry, while a crash can erase the only copy of the work. The opposite trap is returning failure after a committed job merely because later asynchronous work is slow; that creates needless duplicate traffic. Define acknowledgment as proof of durable acceptance, not proof that every downstream action has finished.
Retries magnify that mistake.
Run the leaked-key drill before enabling retries
Treat the drill as an end-to-end test of blast radius, not a ceremonial secret rotation. Create two test accounts with separate credentials and integration scopes. Deliver distinct events plus one duplicate to each, then mark only account A's credential revoked. Requests signed by that credential should fail authentication without changing the replay ledger; account B should continue accepting new work.
Rotate account A to a replacement credential mapped to the same consumer scope. Replay an already accepted event with the replacement key and confirm that the original receipt wins. Next, replay an event older than the configured receipt retention boundary in an isolated environment and observe the documented consequence: without another business-level guard, it can run again.
This is the point. The blast radius is one account only if credential lookup, deduplication scope, queues, logs, and revocation controls all preserve that boundary.
OWASP's secrets-management guidance calls for lifecycle handling that includes creation, rotation, revocation, and expiration, with auditing around secret use and administrative action. Apply that lifecycle to signing credentials: keep secret material out of source code and logs, restrict who can retrieve or rotate it, record revocation events, and verify that caches stop accepting a revoked credential within the system's declared propagation bound. Do not claim instant revocation unless the complete lookup path provides it.
Only then enable automatic retries. The decision rule is narrow: duplicates inside the retained window must converge on one scoped receipt and one business outcome under concurrency, crash, rotation, and revocation tests. If any layer keys by event ID globally, acknowledges before commit, or loses the scope during rotation, the retry policy is amplifying an unresolved correctness or isolation defect.
Top comments (0)