Short answer: use the FIFO queue's five-minute duplicate filter to reduce immediate retries, but make a persistent idempotency receipt the authority for whether a customer's weekly digest may be sent; queue-only filtering is acceptable only when a later duplicate cannot cause a second external effect.
That distinction matters in fintech. A digest containing account activity may be less urgent than a payment authorization, yet sending it twice still creates support work and makes customers question the underlying records. The primary trade-off is latency versus cost: an in-memory check is quick and cheap, while a durable conditional write adds storage work to every attempt. The durable write wins once duplicate delivery can cross the broker's short window or a worker can lose its acknowledgement.
How do duplicate webhook events escape a five-minute FIFO queue dedupe window?
It shouldn't decide the business outcome.
A FIFO queue can order work and collapse matching submissions within its documented interval, but a five-minute dedupe window is a transport property, not a lifetime guarantee for a weekly digest. A webhook sender may retry later, an operator may replay a dead-lettered message, or the same business event may arrive with a new delivery identifier. None of those cases is unusual enough to justify betting a customer-facing side effect on a volatile timer.
Define two identities instead. The transport identity answers, "Have I recently enqueued this delivery?" The business identity answers, "Have I already performed this digest operation?" For the latter, derive a canonical key from stable fields such as customer ID, digest week, digest type, and content version. Don't derive it from arrival time, worker ID, or a random request ID; those values make every retry look new.
For example, customer_8421|2026-W32|weekly_activity|v3 describes one intended digest, no matter how many webhook envelopes request it. Hashing that canonical string keeps the stored key compact, but hashing isn't encryption. If identifiers are sensitive, apply the organization's key-management policy and use a keyed construction rather than assuming a bare digest conceals predictable inputs. OWASP's key-management guidance is the relevant baseline for key lifecycle, storage, and rotation.
The catch is retention. A five-minute queue filter has a bounded, easy-to-estimate footprint; a durable receipt ledger grows with customers and periods. Keep receipts at least as long as the maximum replay horizon, plus the operational interval in which support or audit tooling can legitimately resubmit work. I'm not sure there is one defensible universal duration for fintech digests, because regulatory, dispute, and internal replay policies differ. The written replay policy resolves that uncertainty, not a convenient database TTL.
Put the idempotency key around the effect
The receipt belongs beside the send decision, in storage that supports an atomic create-if-absent or an equivalent uniqueness constraint. Checking first and inserting later is unsafe: two workers can both observe absence, both send, and only then discover that one receipt conflicts. FIFO ordering narrows concurrency under normal operation, but it doesn't erase redelivery, multiple consumers, manual replay, or a lease expiring while work continues.
A useful ledger records more than seen = true. Give each key a small state machine: claimed, sent, and retryable, with timestamps, attempt count, payload fingerprint, and a provider correlation value when one exists. A worker atomically claims an absent or retryable row, verifies that the fingerprint matches any existing row, builds the digest from a fixed snapshot, sends it, and marks the row sent. A conflicting fingerprint under the same key is a data-integrity failure, not a duplicate to ignore; quarantine it for investigation because the producer has assigned one identity to two meanings.
The hardest failure mode sits between the external send and the sent update. Consider one concrete sequence: the worker claims customer_8421|2026-W32|weekly_activity|v3 at 08:00, the recipient accepts the digest at 08:01, and the worker loses its lease before committing sent; an automated replay at 08:12 is outside the queue filter, yet the durable row still says claimed. The retry cannot infer the external fact from the queue, and exactly-once execution across an independent mail system and a local ledger is not created by a hash key. Prefer an outbound provider that accepts your idempotency key or exposes a stable status lookup, then reconcile the ambiguous row before sending again. If neither capability exists, the residual choice is explicit: risk a duplicate, or hold the digest for operator review. Pretending the interval solves this is worse than either choice because it hides ambiguity precisely where the two systems stopped agreeing.
Here is a compact Python sketch using a generic transactional store. The important operations are the conditional claim and the reconciliation branch; the interface names are illustrative application boundaries, not vendor endpoints.
from dataclasses import dataclass
from datetime import datetime, timezone
import hashlib
import hmac
@dataclass(frozen=True)
class DigestRequest:
customer_id: str
iso_week: str
digest_type: str
content_version: str
payload_fingerprint: str
def receipt_key(request: DigestRequest, secret: bytes) -> str:
canonical = "|".join(
(
request.customer_id,
request.iso_week,
request.digest_type,
request.content_version,
)
).encode("utf-8")
return hmac.new(secret, canonical, hashlib.sha256).hexdigest()
def process_digest(request, ledger, sender, secret):
key = receipt_key(request, secret)
now = datetime.now(timezone.utc)
claim = ledger.claim_if_absent_or_retryable(
key=key,
fingerprint=request.payload_fingerprint,
claimed_at=now,
)
if claim.state == "sent":
return {"status": "duplicate", "receipt": key}
if claim.fingerprint != request.payload_fingerprint:
ledger.quarantine(key, reason="idempotency_key_payload_mismatch")
return {"status": "quarantined", "receipt": key}
if claim.needs_reconciliation:
accepted = sender.lookup_by_idempotency_key(key)
if accepted:
ledger.mark_sent(key, provider_reference=accepted.reference)
return {"status": "reconciled", "receipt": key}
result = sender.send_weekly_digest(request, idempotency_key=key)
ledger.mark_sent(key, provider_reference=result.reference)
return {"status": "sent", "receipt": key}
This sketch leaves transaction syntax inside ledger, where it belongs.
Break the worker on purpose before comparing designs
The ledger contract must be tested against real contention: start two claims for the same absent key, block both before commit if the database permits it, and assert that only one caller receives permission to send. Then terminate a worker after the external acceptance but before mark_sent, advance the claim lease, and verify that the replacement worker performs a status lookup before it considers another send. Also test a duplicate with a changed fingerprint, a retry following a 429 response, and replay after the queue interval. These are storage and integration tests, not mock-only unit tests, because a mock that returns claim=True cannot demonstrate that a uniqueness constraint arbitrates two transactions.
One invariant should survive every test: a business key never silently changes meaning.
Compare receipt claims with queue filtering on latency and cost
The controls are complementary, but only one owns correctness. Queue filtering avoids needless worker starts during a retry burst. The ledger prevents a repeated business effect across a longer horizon and supplies evidence for support and audit work.
| Decision factor | Five-minute FIFO filtering | Persistent receipt ledger |
|---|---|---|
| Immediate retry latency | Lowest path overhead; duplicate work can be suppressed before a worker runs | Adds a storage claim to the worker path |
| Duplicate horizon | Limited to the broker's configured or documented interval | Set by receipt retention and replay policy |
| Concurrent workers | Depends on queue delivery and lease behavior | A uniqueness constraint or conditional write arbitrates claims |
| Changed payload under one key | Usually treated as the same transport identity | Fingerprint comparison can quarantine the conflict |
| Ambiguous external send | Cannot establish whether the effect occurred | Stores state needed for lookup and reconciliation |
| Operating cost | Lower storage and fewer reads | Ongoing writes, retained rows, indexes, cleanup, and reconciliation |
For a weekly digest, the extra claim latency is normally off the customer's interactive path, so trading a storage round trip for durable suppression is a reasonable default. Cost still deserves measurement: record claim latency, ledger write volume, index size, receipts removed, duplicate hits by age, and the number of ambiguous sends awaiting reconciliation. A high duplicate-hit count under five minutes argues for keeping the queue filter. Hits hours or days later prove why it cannot be the authority.
Backpressure belongs in the same design. If the sender replies with HTTP 429 Too Many Requests, the worker should respect Retry-After when it is present, leave the receipt retryable rather than sent, and reschedule without changing the business key. MDN documents that Retry-After may accompany a 429 response, while rate-limit implementation details vary by server. Preserve the key. Otherwise each delayed attempt defeats the ledger it is meant to consult.
Queue-only suppression remains suitable for effects that are genuinely disposable: cache warming, replaceable metrics aggregation, or work whose downstream operation is independently idempotent and whose late duplicate has no material consequence. It is not suitable when a digest can be replayed after the queue interval, when support needs an audit trail, or when a second send changes customer behavior. Conversely, a receipt ledger isn't free: for very high-volume ephemeral work, its write amplification and retention burden can outweigh the harm of repeating the task. Stick with the simpler control when that harm has been explicitly bounded, not merely assumed away.
Roll out the ledger without delaying the weekly job
Begin in observation mode. Compute the prospective business key, retain no sensitive canonical input beyond policy, and record how often it repeats at less than five minutes, later the same day, and after the normal replay horizon. This establishes the shape of duplicates without letting a new state machine suppress production sends.
Next, enable atomic claims for a small customer cohort while leaving queue filtering active. Alert on key-and-fingerprint conflicts, stale claimed rows, 429 retry age, and reconciliation backlog. Expand only after a forced-replay exercise shows that one semantic digest produces one accepted external operation and one durable sent receipt. Rollback should disable ledger enforcement while preserving the rows for diagnosis; deleting evidence during a rollback makes the next decision harder.
Keep the rule plain: the queue optimizes the common retry, and the ledger governs the customer-visible effect.
Top comments (0)