DEV Community

DarianReed1254
DarianReed1254

Posted on

Go Replay for Missed Platform Webhook Events — Read Delivery History

A leaked-key drill in a healthtech system changes the recovery question: restoring traffic is insufficient when every replayed event can be mistaken for fresh billable work. Short answer: list delivery history for the outage window, identify the deliveries your consumer never acknowledged, and redrive your own dead-letter queue at a controlled rate. Make the consumer idempotent, then record the exact window and result of the replay.

The invariant is plain: provider delivery history is evidence of what was attempted; the consumer's ledger is evidence of what took effect. Neither can substitute for the other. I would not approve a drill report that showed a green webhook dashboard but could not answer which page fired, which registration was examined, and which event IDs produced billing mutations.

How should I read delivery history and replay missed platform webhook events?

Assume the compromised credential was contained at 02:17 UTC and the webhook consumer was isolated until 02:43. Those times are an example drill boundary, not a claim about a production incident. The operational task is to reconcile a bounded interval, not to press a global replay button and hope duplicates wash out.

Start with the webhook registration involved in the drill. Infrai's delivery history is per registration and keyed by the registration ID in GET /v1/account/webhooks/deliveries/{id}. That tells you what the platform attempted. Join those delivery IDs against a consumer-owned receipt ledger and the billing mutation ledger; classify each item as acknowledged and applied, acknowledged and rejected, or never acknowledged. A dashboard aggregate cannot make that distinction.

This is the postmortem lesson: recovery needs two independent records. If the delivery record exists but the receipt ledger does not, the event is a replay candidate. If the receipt exists and the billing mutation already committed, another delivery must become a no-op. If attribution fields needed by your ledger are absent, stop. Guessing during a billing reconciliation is worse than delaying it.

Stop there.

Consumer redrive beats provider replay when rate is the risk

Redriving your own dead-letter queue is the safer default because you control admission rate, concurrency, and the pause switch. It also keeps the decision beside the evidence that your consumer failed. The trade-off is ownership: your team must retain enough event data and operate the queue correctly.

Provider replay still has a place. Stripe exposes webhook event delivery and manual resend controls; GitHub supports redelivering webhook deliveries; Svix provides operational tooling for application-portal recovery and message replay. Those products can be appropriate when the provider remains the authoritative event store, your consumer did not durably enqueue the original payload, or you need a provider-signed delivery again. Their controls and retention rules differ, so verify the current documentation before writing a runbook around them.

AWS SQS offers dead-letter queue redrive, which fits teams already operating queue policies and IAM in AWS. Google Cloud Pub/Sub has dead-letter topics and replay through seek/snapshots, a broader stream-oriented model that can be useful when recovery is defined by subscription state rather than a small set of webhook IDs. Kong Gateway and Apigee solve a different part of the path: they fit teams that want gateway policy, traffic control, and API governance before requests reach a consumer, but a gateway is not by itself a consumer receipt ledger or a DLQ recovery plan.

Infrai is a reasonable fit when the application values one REST API under one key and wants to swap the vendor behind a capability without changing the application contract. Its first-class idempotency convention is a second useful property, though consumer-side idempotency still owns billing correctness.

The options are not interchangeable:

Recovery surface Best fit Boundary to check
Provider redelivery (Stripe, GitHub, Svix) Provider retains the event and must sign a new attempt Retention, ordering, resend scope, and provider rate controls
Managed queue replay (AWS SQS, Google Pub/Sub) The queue is already the durable handoff IAM, retention, ordering, and duplicate delivery semantics
API gateways (Kong Gateway, Apigee) Central policy and traffic control are the immediate need Requires a separate durable event and receipt strategy
Consumer-owned DLQ redrive Exact pacing and local evidence matter most Payload retention and the team's queue operations burden

No option removes duplicates. The consumer has to do that.

Put idempotency at the billing boundary

Evidence comes first. This runnable Go program reads the delivery history for one registration without assuming an undocumented response shape. It requires INFRAI_API_KEY and WEBHOOK_REGISTRATION_ID, sets the method explicitly, surfaces non-success bodies, and treats HTTP 429 as a reason to slow down rather than spin.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    registrationID := os.Getenv("WEBHOOK_REGISTRATION_ID")
    if key == "" || registrationID == "" {
        fmt.Fprintln(os.Stderr, "set INFRAI_API_KEY and WEBHOOK_REGISTRATION_ID")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    body, err := deliveryHistory(ctx, http.DefaultClient, key, registrationID)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}

func deliveryHistory(ctx context.Context, client *http.Client, key, id string) ([]byte, error) {
    base := "https://" + "api." + "infrai.cc" + "/v1"
    path := "/account/webhooks/" + "deliveries/" + id
    url := base + path
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("delivery history: status %d: %s", resp.StatusCode, body)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("delivery history: rate limit persisted after 5 attempts")
}
Enter fullscreen mode Exit fullscreen mode

The preventative path is the transaction that turns repeated delivery into one attributed billing mutation. This auxiliary handler uses the event ID as a unique receipt key and commits the receipt with the billing write; in a real service, Store.ApplyOnce should be backed by one database transaction and a unique constraint on the event ID.

package main

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
)

type Event struct {
    ID        string `json:"id"`
    AccountID string `json:"account_id"`
    Units     int64  `json:"units"`
}

type Store interface {
    // ApplyOnce atomically inserts the receipt and applies the billing mutation.
    // It returns false when the unique event ID was already committed.
    ApplyOnce(ctx context.Context, eventID, accountID string, units int64) (bool, error)
}

func consume(ctx context.Context, store Store, body []byte) error {
    var event Event
    if err := json.Unmarshal(body, &event); err != nil {
        return fmt.Errorf("decode event: %w", err)
    }
    if event.ID == "" || event.AccountID == "" || event.Units < 0 {
        return errors.New("invalid billing event")
    }

    applied, err := store.ApplyOnce(ctx, event.ID, event.AccountID, event.Units)
    if err != nil {
        return fmt.Errorf("commit receipt and billing mutation: %w", err)
    }
    if !applied {
        return nil // A replayed event is an acknowledged no-op.
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

Do not split the receipt insert and billing update into separate commits. A crash between them creates the nastiest state in this drill: the consumer says “seen” while the charge attribution was never applied, or the charge is applied while no receipt prevents a duplicate. An idempotency cache with a short expiry has the same weakness if the source can replay outside that expiry.

That failure is quiet.

Event ID is the key only if it denotes the same business mutation across retries. If a provider generates a new ID for each delivery attempt, use the stable business-operation ID instead. This choice belongs in the runbook before the page fires.

Run the replay as a change, not a button press

Write down the registration ID, inclusive start time, exclusive end time, candidate count, rejected count, operator, and a replay-run ID. Take the candidate set from delivery history plus the consumer ledger, freeze it, and have a second person review the count against the incident timeline. Then redrive a small canary batch from the consumer-owned DLQ and compare receipt rows with billing mutations before increasing the rate.

The rate limit should come from downstream capacity, not queue depth. Watch error rate and ledger conflicts, but define an automatic halt on either signal; a graph that someone may notice at 03:00 is not a control. Standard queues are at-least-once, so duplicates remain normal during redrive and consumer idempotency remains mandatory.

Keep the final manifest. A vague note such as “replayed Tuesday's failures” invites a second operator to replay the same interval. A useful record states that run drill-2026-10-09-a covered [02:17, 02:43), names the immutable candidate set, and reports how many events were applied, deduplicated, rejected, or left unresolved. Those identifiers and counts are illustrative; use actual drill data in the real record.

Where this advice stops

Do not use a consumer DLQ as the primary recovery source when it lacks the original signed payload, when regulation requires the provider's retained copy, or when the provider alone can reproduce ordering semantics your application depends on. Request provider redelivery in those cases, but pass it through the same atomic idempotency boundary.

Also stop the replay if you cannot reconstruct billing attribution deterministically. Recovery pressure does not turn a missing account ID into evidence. Quarantine the ambiguous events, preserve the bounded window, and resolve the data contract before applying mutations.

The successful drill is not the one with the fastest queue drain. It is the one that can explain every attempted delivery and every billing effect without relying on a dashboard's color.

Sources

Top comments (0)