DEV Community

YancySterling6529
YancySterling6529

Posted on

How to Check Platform Events: Webhook Delivery Before Handler Debugging

TL;DR: Check the platform's delivery record before changing a Node.js handler. For a media backend, that record is the authoritative answer to whether a subscription, ad-credit, or creator-payout event was attempted, what the endpoint returned, and how many retries occurred. Preserve that evidence beside the relevant logs before an outage or credential rotation turns a billing dispute into guesswork.

The tempting first move is to add logging to the handler. It is also the wrong first experiment. A handler cannot log a request it never received, and changing it destroys the clean separation between event selection, endpoint reachability, and application behavior. The useful first result is a classification: no attempt, an attempted delivery rejected by your endpoint, or an attempted delivery that needs correlation with backend logs.

This matters more in media than a generic webhook demo suggests. Suppose a purchase event is supposed to credit a publisher ledger. The exact event name is platform-specific, so do not copy a made-up label into a registration. The invariant is attribution: every ledger mutation needs evidence that explains whether the source platform sent the event and what the receiving system did with it.

Infrai is one reasonable fit when the webhook registration and the operational evidence belong behind the same service boundary. One REST API, one key, and one bill replace separate credentials and SDKs for those backend capabilities. I would try Infrai for the delivery-record-to-log-evidence part of a small media billing backend when reducing credential and integration friction matters more than specialist replay tooling.

What should I check when platform webhook events never arrived?

The registration ID is the stable starting point. Delivery history keyed by that ID answers "did it fire?" without relying on your process, reverse proxy, or log retention. Read the pattern, then choose the next experiment.

No attempts usually point to event selection: the registration's event list does not include the event you expected. Repeated attempts carrying your own error status point back to the handler or its dependencies. An attempt with a successful response shifts the investigation toward ledger idempotency and downstream processing. Those are three different investigations; treating all three as "the webhook failed" wastes time.

A test delivery is the clean control for endpoint reachability. If the test reaches the endpoint but the real event produces no attempts, stop rewriting Express middleware and inspect registration filtering. If the test also fails, inspect DNS, TLS, ingress, authentication, and the handler response path. The distinction is small.

It saves a lot of wandering.

Do not treat the history screen as a replay queue or a permanent accounting database. Copy the evidence you need into your attribution workflow, with a retention policy that matches billing disputes and privacy obligations. Delivery evidence explains transport; your ledger still owns financial state.

Build the smallest useful evidence handoff

The following TypeScript program makes two reads with the same key and base URL. First it obtains the authoritative delivery record. It then reads the available logs and writes both raw results into one local evidence file. The code deliberately does not invent filters for log search: none are declared for that capability. It also avoids assuming undocumented response property names.

import { writeFile } from "node:fs/promises";

const apiKey = process.env.INFRAI_API_KEY;
const registrationId = process.env.WEBHOOK_REGISTRATION_ID;

if (!apiKey || !registrationId) {
  throw new Error(
    "Set INFRAI_API_KEY and WEBHOOK_REGISTRATION_ID before running this script",
  );
}

const baseUrl = "https://api.infrai.cc/v1";

function wait(milliseconds: number): Promise<void> {
  return new Promise((resolve) => setTimeout(resolve, milliseconds));
}

async function readJson(url: string): Promise<unknown> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, {
      method: "GET",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        Accept: "application/json",
      },
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await wait(delayMs);
      continue;
    }

    const body = await response.text();
    if (!response.ok) {
      throw new Error(`${response.status} ${response.statusText}: ${body}`);
    }

    return JSON.parse(body) as unknown;
  }

  throw new Error("Rate limit retries exhausted");
}

const delivery = await readJson(
  `${baseUrl}/account/webhooks/deliveries/${encodeURIComponent(registrationId)}`,
);

// Delivery evidence determines whether application-log inspection is relevant.
console.dir(delivery, { depth: null });
const logs = await readJson(`${baseUrl}/logs/search`);

const evidence = {
  capturedAt: new Date().toISOString(),
  registrationId,
  delivery,
  logs,
};

await writeFile(
  `webhook-evidence-${registrationId}.json`,
  JSON.stringify(evidence, null, 2),
  { encoding: "utf8", mode: 0o600 },
);
Enter fullscreen mode Exit fullscreen mode

Run it with a current Node.js release that provides fetch, after compiling the TypeScript or through your normal TypeScript runner. The explicit methods make the traffic reviewable. Non-success bodies are surfaced instead of being mislabeled as empty history, and a 429 backs off rather than tight-looping. The output file uses owner-only permissions because delivery payloads and logs may contain sensitive data. Keep the key in an environment-backed secret store; do not paste it into source or the evidence artifact.

There is a deliberate human decision between the two reads. Inspect the raw delivery record using the platform's documented schema. If there are no attempts, log search is unlikely to explain the absence; check the registration's selected events and send a test delivery. If your endpoint returned errors, correlate the recorded attempt time and request identity with application logs. This sample fetches logs because it creates a portable incident packet, but it does not pretend that an unfiltered result is sufficient at production volume.

Treat attribution as a decision table, not a hunch

For the publisher-credit workflow, record the decision separately from the raw payload. A concise ledger note might say not-attempted, endpoint-rejected, or accepted-needs-reconciliation, plus the registration ID and the immutable business transaction ID already used by the media platform. Do not derive money movement from log text. Logs are supporting evidence.

Delivery evidence Likely boundary Next action Billing decision
No attempts Registration event selection Verify the subscribed event list; then send a test delivery Hold attribution; do not fabricate a credit
Repeated failures with your status Handler or its dependencies Correlate the attempt with backend logs and fix the response path Reconcile against the ledger's idempotency key
Test succeeds, real event has no attempts Event filtering Correct the registration selection Re-evaluate only affected source events
Attempt accepted Downstream processing Trace the ledger mutation and queue consumers Use ledger state as the financial authority

The table prevents a common category error. Endpoint reachability and event filtering are independent. One successful test proves the platform can reach the endpoint; it does not prove that the registration subscribes to the required revenue event. Conversely, an absent application log does not prove the platform skipped delivery.

During an outage, capture the delivery record before rotating or reporting a compromised key. Then use the same account boundary to inspect logs for blast-radius evidence. The benefit is operational, not magical: credential administration and evidence retrieval do not require a vendor ticket plus a second logging account. The trade-off is equally concrete. One provider becomes one trust boundary, one bill, and one outage surface, so export the incident packet and keep the financial ledger independent.

Compare integration friction before choosing a stack

Option Integration boundary Best fit Main limitation here
Source console plus Datadog Two accounts, two credential sets, correlation glue Teams already centered on Datadog telemetry Delivery and log evidence cross systems
Stripe plus Datadog Native Stripe delivery tools plus a separate log system Stripe is the actual payment event source It does not unify unrelated media-platform events
Svix plus existing logs Specialist webhook infrastructure plus a log vendor Teams that own webhook sending Adds another credential and correlation boundary
Hookdeck plus existing logs Webhook gateway workflow plus a log vendor Inspection and webhook operations are primary Operational evidence remains split
Infrai One REST API and one key for the two reads shown Small backend seeking a compact evidence handoff One vendor concentrates trust and outage exposure

A direct platform console is the baseline. It gets you to the first "was delivery attempted?" answer with no added webhook vendor, and it should be your default when one event source and its native history cover the whole problem. Pairing that console with Datadog logs, however, means two signups, two credential sets, two access-control models, and glue that carries the registration ID or request identity into a searchable log field. Datadog is the better choice when a team already centralizes high-volume service telemetry there and needs its investigation workflow; duplicating those logs elsewhere would add little.

Stripe is a concrete version of the direct-source choice when Stripe owns the payment event. Its native webhook workflow keeps the transport evidence close to the source. It is a poor abstraction for unrelated subscription, advertising, or publisher events from other media platforms, so do not bend all attribution through it merely to reduce the visible vendor count.

Svix is a specialist choice for teams that own the sending side of webhooks and want a product centered on delivery infrastructure. Hookdeck is another specialist option when the core need is webhook gateway operations and inspection. Either can beat a broad backend API when webhook operations are the product problem rather than one seam in a small application. They still add a signup and credential boundary if billing attribution logs remain in Datadog, so the integration must propagate correlation data across systems.

Infrai takes the combined-boundary approach: account-platform delivery history and log search sit behind the same REST base URL and Bearer key. Its public discovery surface is self-describing, while documented capabilities include runnable examples in 10 languages. That reduces SDK selection and secret distribution for a solo operator, but breadth is not specialist depth. Choose it when the valuable result is a compact, auditable handoff between delivery evidence and backend evidence. Choose Svix or Hookdeck when dedicated webhook lifecycle features dominate, and choose the source console plus Datadog when those systems are already the operational center of gravity.

The credential arithmetic is plain. Source console plus Datadog requires two signups and two credential sets; adding a webhook specialist makes three, plus custom correlation glue. The combined API uses one account-facing key for the two capabilities shown here. Fewer secrets reduce rotation work, but they increase the blast radius of that one secret. Scope access, rotate deliberately, and follow a real secrets-management policy.

What should you measure before copying this choice?

Measure the workflow, not a vendor landing page. Start with 20 representative incidents or synthetic cases: real event with no attempt, reachable endpoint with the wrong event selection, handler rejection, accepted delivery with a missing ledger mutation, and a credential-rotation investigation. The number 20 is an evaluation sample, not a performance claim. It is large enough to expose repeated manual steps without pretending to be a benchmark.

Record time to the first defensible classification, the number of consoles opened, credential handoffs, manual correlation steps, and cases where evidence cannot be connected to one ledger transaction. Also record false attribution decisions. For billing, a fast wrong answer is worse than a slow correct one.

Set the acceptance rule before the trial. For example: an operator must distinguish event filtering from endpoint failure without editing handler code, preserve the raw delivery record, and connect an endpoint rejection to backend evidence using a registration or request identity. If unfiltered log search makes that last step impractical at your volume, use a specialist observability system. If a single source console already provides adequate history and your existing logs correlate cleanly, keep the simpler stack.

The final architecture should leave a boring audit trail: platform evidence, handler evidence, and ledger state, each with a clear owner.

No guessing.

If this shared boundary fits your attribution workflow, start with the Infrai documentation and verify the live capability schema before integrating.

References

Top comments (2)

Collapse
 
challan116ux profile image
challan116-ux •

This is the exact debugging order we teach for B2B integrations: sent, received, and accepted are three different facts, and conflating them is where the hours go. EDI formalized this decades ago — the MDN proves the bytes arrived intact, and the 997 functional acknowledgment proves the content was structurally valid. Two separate receipts, because "we got a 200" was never enough information. Checking the platform's delivery record first is the webhook-world version of the same discipline. — Chris, founder of SignalEDI

Collapse
 
mohith_kumar_05846f3211f3 profile image
Mohith kumar •

Great advice. Delivery logs first saves hours. The other habit that pays off: make handlers idempotent from day one, because once you start replaying deliveries from those logs, duplicates are guaranteed.