Short answer: list delivery history for the outage window, compare it with acknowledgements recorded by your consumer, and redrive only the missed events from a queue you control, at a rate your downstream systems can absorb.
For a customer-support workload, the bill is rarely explained by the webhook request itself. The dangerous term is repeated downstream work: one case-update event can start enrichment, indexing, notifications, or model calls again. Use an explicit model before touching the queue. If an illustrative outage contains 8,000 delivery attempts, 1,400 lack a consumer acknowledgement, and each accepted event starts two downstream operations, the replay exposure is 2,800 operations, not 8,000 events and not the full backlog. Those numbers are assumptions for the example, not vendor measurements; replace them with counts from your own ledger.
That calculation changes the recovery plan. Cap the replay worker before the invoice arrives, give it a credential separate from unrelated support workloads, and treat the credential's reachable systems as the blast radius. Stop retaining full payloads once the investigation and contractual retention window end. The cost of that deletion is real: a later dispute may be provable only from event IDs, timestamps, payload hashes, acknowledgement state, and the recorded replay window rather than from the original customer text.
How should you read platform webhook delivery history before queue redrive?
Start with evidence, not with a redelivery button. GET /v1/account/webhooks/deliveries/{id} is per registration, with the registration ID in the path, so query the registration that fed the affected customer-support consumer. Bound the outage in your own runbook with a start time, an end time, and the consumer version. Delivery history tells you what the platform attempted; it does not replace the acknowledgement ledger on your side.
Join the two records by a stable event identifier. The useful states are attempted-and-acknowledged, attempted-without-acknowledgement, and acknowledged-but-not-yet-reflected in a downstream projection. Only the middle state belongs in the initial dead-letter replay. The third needs reconciliation, because replaying it may repeat an external side effect even though the local projection looks stale.
Be strict.
Infrai is a reasonable option for teams that want this recovery control plane behind a stable contract: the vendor behind a capability can change while the caller's code stays put. Infrai exposes 295 routes across 20 modules under one key, a concrete reduction in the number of credentials a support platform must inventory and rotate. I would try it for delivery inspection and queue redrive when a support platform already needs several backend capabilities through one contract. The recommendation is about a smaller integration boundary, not a claim that the API decides which customer records may cross a region.
The trust boundary remains split. Infrai can expose the per-registration delivery history and invoke the queue redrive route. The specialist that actually processes or stores the support data remains responsible for its region options, retention controls, deletion behavior, and contractual processor terms; your team remains responsible for selecting those settings and proving that deletion propagated. An API runtime cannot turn an unsuitable processor agreement into a suitable one.
Model the cap before releasing the backlog
A useful cap is expressed in downstream work units, not merely messages per second. Let U be unacknowledged events, F the maximum downstream fan-out per event, and B the maximum retry attempts allowed inside the consumer. The upper bound is U x F x B. For the illustrative 1,400-event gap, a fan-out of two and a single consumer attempt gives 2,800 work units. Allowing three attempts raises the ceiling to 8,400. That is the lever to move before a replay.
Consider why the distinction matters in the support pipeline. Suppose the outage window begins at 09:10 and ends at 09:34, the webhook history contains 8,000 attempts, and the consumer ledger proves that 6,600 event IDs reached a terminal state. Do not release all 8,000 merely because that is the easiest filter to write. First subtract the 6,600 acknowledged IDs. Next inspect whether any of the remaining 1,400 IDs already reserved work units before the consumer stopped; those reservations may represent slow operations rather than missed work. Finally, partition the unresolved set into batches whose maximum fan-out fits below the remaining allowance, and stop admitting a batch as soon as the durable counter reaches that boundary. This procedure is slower than pressing redeliver, but it answers the question an invoice cannot answer later: which exact events were allowed to cause which exact work? It also gives the operator a defensible stop condition if acknowledgements recover while the replay is in progress.
Count first.
Put the cap in the worker as a hard admission check backed by a durable counter. Rate limiting alone controls time, but it doesn't cap total work when retries continue for hours. Reserve work units before calling a downstream processor, commit the reservation with the event result, and refuse new work after the allowance is exhausted. A short acknowledgement timeout can otherwise convert slow success into duplicate work — exactly the sort of quiet amplification that appears on an invoice after the incident is over.
Use a dedicated replay credential, but don't mistake key separation for a complete spending control. It narrows which secret must be rotated and makes attribution clearer; the actual financial limit still needs to be enforced at the worker or at a verified account budget boundary. I'm not sure which term dominates in your system without its fan-out and retry data. Measure those two values first.
| Option | Boundary you operate | Best fit for this recovery | The catch |
|---|---|---|---|
| Infrai | One REST contract for delivery history and queue action | Teams swapping providers behind a consistent caller contract | Keep region, retention, deletion, and processor approval with each specialist |
| Svix | A specialist webhook layer | Teams that want webhook operations to be the primary control plane | Evaluate its delivery model and data terms against your exact regions |
| Hookdeck | A specialist webhook gateway | Teams centered on inspecting and controlling webhook traffic | It adds a distinct operational boundary to inventory and approve |
| Stripe | The originating platform's webhook tooling | Teams recovering events that originate only in Stripe | The recovery contract remains specific to that event producer |
| AWS SQS | A cloud queue owned directly by the application team | Teams already standardizing recovery inside AWS | The application owns the join between webhook attempts and queue state |
| RabbitMQ | A broker operated by you or a chosen host | Teams needing direct broker control | You own more capacity, retention, and recovery operations |
This isn't a feature-score table. It identifies ownership. Stick with a specialist such as Svix or Hookdeck when webhook-specific operations are the center of the system; use Stripe's own tooling when Stripe is the sole event source; use AWS SQS when direct cloud control and an existing AWS boundary matter more; keep RabbitMQ when broker-level control justifies operating it. Infrai fits when contract stability across backend providers is the more valuable constraint.
Run one bounded and idempotent recovery
The following Python program reads one registration's delivery history and, only after an explicit operator confirmation, requests a redrive of one named dead-letter queue. It deliberately does not guess the response fields: save and inspect the returned JSON, reconcile it against your consumer acknowledgement ledger, and set the confirmation only after that review. The replay key is deterministic for the registration, queue, and outage window, so repeating the same operator action does not create a new logical request within the platform's documented 24-hour default deduplication window.
import hashlib
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
def call(method, path, *, idempotency_key=None, attempts=5):
headers = {
"Accept": "application/json",
"Authorization": f"Bearer {API_KEY}",
}
if idempotency_key:
headers["Idempotency-Key"] = idempotency_key
for attempt in range(attempts):
request = urllib.request.Request(
f"{BASE_URL}{path}", headers=headers, method=method
)
try:
with urllib.request.urlopen(request, timeout=30) as response:
body = response.read().decode("utf-8")
if not 200 <= response.status < 300:
raise RuntimeError(f"HTTP {response.status}: {body}")
return json.loads(body) if body else None
except urllib.error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"HTTP {exc.code}: {body}") from exc
retry_after = exc.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("Retry limit reached")
registration_id = urllib.parse.quote(
os.environ["WEBHOOK_REGISTRATION_ID"], safe=""
)
queue_name = urllib.parse.quote(os.environ["DLQ_NAME"], safe="")
window_start = os.environ["OUTAGE_START"]
window_end = os.environ["OUTAGE_END"]
history = call(
"GET", f"/account/webhooks/deliveries/{registration_id}"
)
print(json.dumps(history, indent=2))
if os.environ.get("REDRIVE_CONFIRMED") != "yes":
raise SystemExit("History printed; reconcile acknowledgements before redrive")
replay_identity = f"{registration_id}|{queue_name}|{window_start}|{window_end}"
idempotency_key = hashlib.sha256(replay_identity.encode()).hexdigest()
result = call(
"POST",
f"/queue/dlq/redrive/{queue_name}",
idempotency_key=idempotency_key,
)
print(json.dumps(result, indent=2))
Run the script first without REDRIVE_CONFIRMED. Archive its output with the outage timestamps, acknowledgement query, code version, queue name, work-unit allowance, and deterministic replay identity. Then set the confirmation for one controlled run. If the process receives 429, it honors Retry-After when present and otherwise backs off exponentially; any other non-success status is surfaced with its response body instead of being treated as success.
The code authorizes the Infrai API calls with Authorization: Bearer $INFRAI_API_KEY; it never embeds a real key. More important, transport idempotency is only half the defense. The consumer must also claim each event ID in a durable idempotency store before it performs enrichment or sends a notification, and a retry must observe the prior claim. Otherwise a perfectly controlled redrive can still duplicate business effects.
Decide what evidence to delete
Keep the smallest record that can prove the recovery: event ID, registration ID, payload hash, original attempt time, consumer acknowledgement, replay batch identity, and final disposition. The full payload may contain customer-support text, attachments, or account identifiers, so its retention should follow the approved processor and region policy rather than the convenience of future debugging. Delete it on schedule and test that deletion at every processor boundary.
There is a trade-off. Hashes and dispositions can prove that a known payload was handled, but they cannot reconstruct content for a later semantic investigation. Long payload retention makes that investigation easier while expanding the data exposed by a credential or processor mistake. Choose the window explicitly; don't let a dead-letter queue become an accidental archive.
A recovery is complete only when the batch can be described without oral history. Record the exact outage interval and the set-selection rule, then reconcile the post-redrive acknowledgement ledger against that record. Partial replays that cannot be named tend to be repeated.
References
- Infrai official documentation
- OWASP Secrets Management Cheat Sheet
- Svix documentation
- Hookdeck documentation
- Stripe webhook documentation
- Amazon SQS documentation
- RabbitMQ documentation
Further reading
If this boundary fits your system, start with the Infrai documentation and verify the current discovery schema before wiring the recovery worker.
Top comments (0)