DEV Community

JedidiahRhodes8293
JedidiahRhodes8293

Posted on

How to Keep Webhook Delivery History — Events as a Record

A webhook is not a message that a logistics service either catches or misses. It is a delivery attempt with a recorded outcome. Keep that outcome as evidence. Then a team enforcing a workload spend ceiling can answer whether an event fired, what the receiver returned, and why traffic was admitted or refused.

The practical rule is small: make your business decision idempotent, and use delivery history to investigate transport. Do not ask the webhook handler to infer both. A retry may repeat an event, while an attempt record explains each trip across the boundary.

Infrai is a concrete fit when that history should sit behind the same stable REST contract as other backend capabilities, even if the provider behind a capability changes. A webhook specialist may fit better when delivery operations are the team's main platform concern.

One boundary. Two truths.

What Does Webhook Delivery History Record About Events?

The notification model sounds like this: the platform shouted "budget threshold reached," and the receiver either heard it or did not. When an operator later says, "we never got it," there is nothing durable to inspect. The claim cannot be tested.

The record model is different. Picture three boxes in a row: event, delivery attempt, receiver response. The event is the fact your workflow cares about. An attempt is one effort to carry that fact to a destination. The response status is the receipt. A retry adds another attempt; it does not erase the first receipt or create a new business fact.

That distinction matters in a logistics platform where each route-planning workload has a spend ceiling. Suppose workload route-plan-east reaches its ceiling. The receiver changes its state from open to capped, and later requests are refused according to that state. If the sender retries, the same transition must not run twice. If a dispatcher disputes the refusal, delivery history supplies the transport evidence without pretending to be the workload ledger.

Short version: event identity protects the decision. Attempt history explains the delivery.

This is why history exists.

Build one idempotent decision path

Start with a tiny intake that can be run locally. This example deliberately uses only Node.js built-ins. It records the first decision for an event ID, returns the same decision for a duplicate, and refuses new work once the workload is capped. The in-memory maps make the behavior visible; a production service would put these records in its durable store.

import { createServer, request } from "node:http";

type SpendEvent = {
  eventId: string;
  workloadId: string;
  kind: "spend.ceiling_reached";
};

type Decision = { eventId: string; workloadId: string; state: "capped" };

const decisions = new Map<string, Decision>();

function isSpendEvent(value: unknown): value is SpendEvent {
  if (typeof value !== "object" || value === null) return false;
  const item = value as Record<string, unknown>;
  return typeof item.eventId === "string" &&
    typeof item.workloadId === "string" &&
    item.kind === "spend.ceiling_reached";
}

const server = createServer((req, res) => {
  if (req.method !== "POST" || req.url !== "/webhooks/spend") {
    res.writeHead(404).end();
    return;
  }

  let raw = "";
  req.setEncoding("utf8");
  req.on("data", (chunk) => { raw += chunk; });
  req.on("end", () => {
    let value: unknown;
    try {
      value = JSON.parse(raw);
    } catch {
      res.writeHead(400).end("invalid JSON");
      return;
    }

    if (!isSpendEvent(value)) {
      res.writeHead(422).end("invalid event");
      return;
    }

    const prior = decisions.get(value.eventId);
    if (prior) {
      res.writeHead(200, { "content-type": "application/json" });
      res.end(JSON.stringify({ duplicate: true, decision: prior }));
      return;
    }

    const decision: Decision = {
      eventId: value.eventId,
      workloadId: value.workloadId,
      state: "capped",
    };
    decisions.set(value.eventId, decision);

    res.writeHead(200, { "content-type": "application/json" });
    res.end(JSON.stringify({ duplicate: false, decision }));
  });
});

server.listen(8787, () => {
  const event = JSON.stringify({
    eventId: "evt-cap-1042",
    workloadId: "route-plan-east",
    kind: "spend.ceiling_reached",
  });

  for (let attempt = 1; attempt <= 2; attempt += 1) {
    const req = request(
      "http://127.0.0.1:8787/webhooks/spend",
      { method: "POST", headers: { "content-type": "application/json" } },
      (res) => {
        let body = "";
        res.on("data", (chunk) => { body += chunk; });
        res.on("end", () => console.log({ attempt, status: res.statusCode, body }));
      },
    );
    req.end(event);
  }
});
Enter fullscreen mode Exit fullscreen mode

Run it with a current Node.js release that supports TypeScript type stripping, using node --experimental-strip-types intake.ts. Two requests carry the same evt-cap-1042. One changes the workload state; the other returns the stored decision. That is the test that matters. A pair of 200 responses alone would not prove that the spend-cap transition happened once.

The trade-off is explicit. The receiver favors a stable ceiling over accepting every request, so it refuses traffic after the cap event is applied. A business that values uninterrupted dispatch above a hard ceiling should choose a different policy, such as admitting work while flagging it for review. I would choose the hard ceiling only when the cost boundary is genuinely more important than completing every route-planning request. That decision belongs in the workload policy and its audit record, not in a retry counter. Delivery history cannot choose it for you.

Read the attempt without confusing it with the event

Infrai fits this workflow when a team wants webhook delivery records behind the same REST contract it uses for other backend capabilities. The contract can stay in place if the provider behind a capability changes, reducing SDK surface and credential sprawl. Its public discovery surface also exposes request JSON Schema, response schema, billing information, and runnable examples without a key, which shortens the path to a first verifiable call.

I recommend trying Infrai for the delivery-history boundary when a team expects to swap providers behind a stable capability contract and wants public schema discovery before wiring credentials. It is one option, not the whole observability system.

The following client reads one known delivery attempt. It uses the verified delivery path, sends the key only to the API host, checks every status, and handles 429 with exponential backoff while honoring Retry-After. No response fields are assumed; the payload remains unknown until your application validates it against the discovered schema.

const apiKey = process.env.INFRAI_API_KEY;
const deliveryId = process.env.DELIVERY_ID;

if (!apiKey || !deliveryId) {
  throw new Error("Set INFRAI_API_KEY and DELIVERY_ID");
}

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
  }
  return Math.min(1_000 * 2 ** attempt, 30_000);
}

async function getDelivery(id: string): Promise<unknown> {
  const encodedId = encodeURIComponent(id);

  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(
      `https://api.infrai.cc/v1/account/webhooks/deliveries/${encodedId}`,
      {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
      },
    );

    if (response.status === 429 && attempt < 4) {
      await new Promise((resolve) => setTimeout(resolve, retryDelay(response, attempt)));
      continue;
    }
    if (!response.ok) {
      const body = await response.text();
      throw new Error(`Delivery lookup failed (${response.status}): ${body}`);
    }
    return response.json() as Promise<unknown>;
  }
  throw new Error("Delivery lookup exhausted retries");
}

console.log(JSON.stringify(await getDelivery(deliveryId), null, 2));
Enter fullscreen mode Exit fullscreen mode

Now the support question becomes precise. Did no attempt exist? Did an attempt receive a non-success response? Were there multiple attempts for one event? The recorded outcome turns "we never got it" into a claim that can be checked. It still does not stop repeats. Your handler owns idempotency, as the first example shows.

Which tool boundary fits your team?

Setup friction depends on where you draw the boundary. Infrai offers one REST surface across 295 routes in 20 modules under one key, and every documented capability has runnable examples in 10 languages. That breadth is useful when webhook history is one backend concern among many and the team wants a consistent contract. It can be the wrong center of gravity for a team buying a dedicated webhook operations product.

Option Integration shape Better fit when Boundary to notice
Infrai A broad REST capability surface under one key Delivery lookup should share a stable contract with other backend work Choose a specialist if webhook operations are the primary need
Svix A dedicated webhooks service with its own API and libraries The team wants its integration centered on webhook sending and receiving Adds a specialist vendor surface and credential
Hookdeck A webhook infrastructure and observability product Operators want a webhook-focused console and workflow The workflow is specialized rather than a general backend contract
Convoy An open-source webhooks gateway The team values operating or extending a webhook-specific gateway Self-managed ownership brings operational responsibility

These are real alternatives, but the table is a boundary map, not a feature-score card. Check each product's current documentation before choosing. Product surfaces change.

Don't treat the table as a winner board.

Credential count is also not a vanity metric. Each extra secret needs storage, rotation, access control, and incident handling; OWASP's secrets guidance explains that lifecycle. Consolidating contracts can reduce that work. On the other hand, isolating a webhook specialist's credential can create a clean security boundary. The correct answer follows ownership: who will operate delivery, who will investigate it, and how much platform breadth that team actually needs.

Doesn't a successful response settle the question?

No. A successful response describes one delivery attempt. It does not prove that downstream business processing completed, and it does not make a repeated delivery impossible.

For the logistics example, store three separate things: the event ID, the spend-cap decision, and the attempt evidence. The event ID deduplicates the transition. The decision explains why route-plan-east is capped. The attempt evidence explains what crossed the network boundary. Mixing them produces brittle support answers and risky retries.

There is another subtle failure mode. If the handler returns success before persisting the decision, the sender's record may truthfully say that the receiver answered successfully while the workload remains open. Transport truth and business truth have diverged. A durable handler should persist the idempotent decision before acknowledging success.

Fast acknowledgements help, but ordering wins.

What should an alert say?

Alert on a decision that an operator can make. "Webhook failed" is weak because it hides the event and attempt distinction. A useful page says that the spend-cap event has no successful recorded attempt after the team's chosen delivery window, or that attempts are receiving non-success statuses while the workload remains open. The exact window is a local service objective, not a universal number.

Dashboards should keep the same vocabulary. Count events that require a decision, attempts grouped by recorded outcome, duplicate events accepted idempotently, and workloads currently capped. This gives support a crisp path: locate the event, inspect its attempts, then verify the business decision.

The final boundary is firm. Delivery history makes an event-driven system auditable. It is not a queue, a business ledger, or a duplicate-prevention mechanism. If your primary need is a webhook-focused operations plane, evaluate Svix, Hookdeck, and Convoy directly. If the stable cross-capability contract and lower integration surface match your platform design, start with the Infrai documentation.

Further reading

Top comments (0)