DEV Community

SunspireValerius59
SunspireValerius59

Posted on

How to Make a Webhook Consumer Idempotent: 3 Credential Boundaries Before Retries

An access reviewer cannot sign off on webhook retries if one leaked signing credential could authenticate events for every clinic. First constrain that credential's blast radius; then make each authenticated event ID produce at most one database effect. Short answer: scope the signing key to one tenant and endpoint, claim (tenant_id, event_id) in the same transaction as the business update, and acknowledge only after commit. A retry after a lost response then sees the existing claim and returns success without repeating the update.

This matters for an appointment reminder service. An event may authorize a reminder schedule change, while an unrelated OTP delivery workflow shares the same patient account. Event-ID deduplication prevents repeated effects from delivery retries; it does not prove that an event was authorized, or prevent someone with a broad signing key from submitting a new event ID. Keep those two questions separate in the review.

Which credential can submit a new event?

Start the access review with a table of signing credentials, tenant bindings, permitted event types, owners, rotation paths, and verification logs. One credential spanning every clinic creates a different incident scope from one credential restricted to a clinic's reminder endpoint. The latter requires a provisioning and rotation process per clinic; that operational load buys a smaller authentication boundary. A reviewer should be able to trace the binding from the credential record to the consumer's tenant context, rather than trusting a tenant ID supplied in the payload.

Verify the message authentication code over the exact received bytes before parsing or mutating data. Use constant-time comparison and reject unknown key IDs. For an HMAC-based integration, bind the key lookup to the provisioned tenant and endpoint; do not pick a tenant solely from attacker-controlled JSON. Set a maximum body size before buffering. A signed timestamp with a bounded acceptance window can limit delayed replay, but the precise signed fields and rotation overlap must match the sender's documented protocol. Never log the key or full health payload. The OWASP secrets guidance covers storage, access control, rotation, and audit requirements for the signing material.

That's the first boundary.

The credential inventory tells the reviewer how many clinics a compromise can affect. A perfect dedupe table cannot shrink that scope. For example, if a key can authenticate events for two clinics, an attacker holding that key can submit fresh IDs for either clinic; deduplicating a prior event is irrelevant to those new submissions. Restrict the verifier's key lookup, event type, and tenant binding together, then ask the reviewer to approve the resulting scope rather than a promise about duplicate suppression.

How do you claim an event without losing the business update?

Use a uniqueness constraint on the authenticated tenant and event ID, then insert the claim and apply the reminder change in one transaction. The following Python example assumes that a trusted request verifier has already produced tenant_id, event_id, and a validated reminder status. It shows the storage boundary, not signature verification or a complete HTTP handler.

import sqlite3


def apply_reminder_event(connection, tenant_id, event_id, reminder_id, status):
    if status not in {"scheduled", "canceled"}:
        raise ValueError("unsupported reminder status")

    with connection:
        inserted = connection.execute(
            """INSERT INTO processed_events (tenant_id, event_id)
               VALUES (?, ?) ON CONFLICT(tenant_id, event_id) DO NOTHING""",
            (tenant_id, event_id),
        ).rowcount
        if not inserted:
            return "duplicate"

        changed = connection.execute(
            """UPDATE reminders SET status = ?
               WHERE tenant_id = ? AND reminder_id = ?""",
            (status, tenant_id, reminder_id),
        ).rowcount
        if changed != 1:
            raise ValueError("reminder not found or ambiguous")
    return "applied"
Enter fullscreen mode Exit fullscreen mode

Create processed_events with a composite primary key (tenant_id, event_id) and reminders with a unique (tenant_id, reminder_id) key. SQLite's conflict clause and transaction context make the example runnable against that schema. In a deployed service, use the database's equivalent atomic insert and transaction behavior, and test concurrent duplicate deliveries against that database. A preflight SELECT followed by an insert is not enough: two workers can both see an absent ID.

There is a trap in claiming the ID before starting the business transaction. If the worker crashes after that separate claim commits, every retry appears processed while the reminder never changes. Here, a missing reminder raises inside the transaction, rolling back the claim as well. The handler should treat that persistent validation failure differently from a transient database error: acknowledge or quarantine invalid events under a documented policy, and retry transient failures. Do not return success before commit.

No half-committed claim.

What should a retry actually repeat?

The same authenticated event ID is a duplicate even if its HTTP delivery has a different request ID. Return a successful acknowledgment for a committed duplicate; returning an error only invites another delivery of the same event. An event with a new ID but identical payload is not necessarily a duplicate: it may represent a second legitimate transition. If the source can reuse IDs across endpoints, include the source or endpoint binding in the uniqueness scope as well. Choose the key from the sender's documented uniqueness guarantee, not from a guessed global namespace.

The limitation of this transaction design is its database boundary. For outbound SMS or email, a local transaction cannot atomically commit both the database update and an external delivery. Record a delivery intent in the same transaction, then have a separate worker send it using an idempotency mechanism supported by the destination or reconcile uncertain outcomes. Otherwise a crash after sending but before recording completion can send twice. OTP sends deserve their own policy: a duplicate webhook must never mint a fresh OTP merely because a network acknowledgment went missing. If a sender offers no stable event ID, this design is not suitable as written: establish a documented compound key or a reconciliation process before switching on automatic retries.

The comparison belongs here, after the failure modes are clear:

Design Crash after claim Concurrent deliveries Credential compromise
In-memory ID set State disappears on restart Depends on shared state No boundary
Separate durable claim Business effect can be lost Uniqueness can stop duplicates No boundary
Transactional claim and update Both roll back together Unique key admits one writer Still requires scoped keys

These are architectural properties, not a product ranking. The transaction protects one database effect. External effects still need an outbox-style workflow or equivalent reconciliation, and scoped authentication still needs an access review.

Roll out the boundary before enabling retries

First deploy the uniqueness constraint and transactional handler while delivery retries are disabled. Test duplicate IDs arriving concurrently, a process interruption before commit, a committed update followed by a lost acknowledgment, an unknown tenant, and an invalid signature. In each case inspect both the processed-event row and the reminder row; counting HTTP responses alone misses the partial-commit failure. Keep test records free of patient-identifying data.

Next review each credential's tenant and endpoint scope, storage permissions, rotation owner, and audit trail. Measure duplicate acknowledgments, failed verification, transaction rollbacks, and delivery-intent backlog separately, with tenant-scoped identifiers that do not expose the payload. Keep processed IDs at least as long as the sender can redeliver; if that horizon is undocumented, confirm it before setting a retention policy. A bounded retry schedule and a quarantine path for persistent failures keep a poison event from occupying workers indefinitely.

Only then enable retries for a small tenant cohort and verify that duplicate delivery leaves exactly one committed reminder transition and one intended notification. The sign-off artifact is concrete: credential-to-tenant mapping, schema constraint, concurrency and crash test results, retention decision, and a rollback switch for the retry policy. That is the evidence a reviewer can approve without mistaking idempotency for authorization.

References

Sources

Top comments (0)