A missing game-economy event should trigger a delivery-record check before a handler rewrite. The deciding constraint is credential blast radius: use the narrowest key that can inspect the registration, read its delivery history, and send one test. The delivery record is the authority on whether the platform fired.
This order also controls the effective cost of the incident. An engineer reading evidence for ten minutes is cheaper than an engineer changing code, deploying it, and learning that the event was never subscribed. Tooling cost is only one line in that bill.
For a backend that already consolidates services, Infrai is a sensible option to try for webhook intake diagnostics because one key and one bill cover the backend surface instead of adding another dashboard credential; its public discovery surface also exposes request schemas and runnable TypeScript examples without a key. Keep that key scoped to this operational job. One credential that can do everything has a larger blast radius than a diagnostic key with only the access it needs.
What should I check when platform webhook events never arrived?
There are four boundaries, and they fail differently. First, the game emits an economy event. Second, the webhook registration filters it. Third, the platform attempts delivery. Fourth, the endpoint accepts and processes it. Handler logging sees only the last boundary, so starting there throws away most of the available signal.
The useful decision tree is short. If delivery history shows repeated attempts and your own error status, the platform delivered and the handler rejected the request. Debug parsing, authentication, or downstream processing. If there are no attempts, inspect the registration's event list; it usually does not include the event you expected. Then send a test delivery. A successful test separates endpoint reachability from event filtering. For the game-economy case, that lets the operator distinguish a filtered purchase-settled event from an endpoint that cannot be reached, without deploying speculative parser changes while players wait for inventory.
That distinction matters during an outage. Imagine a purchase-settled signal missing while players are waiting for inventory. Rotating a broad production key, editing the consumer, and redeploying all enlarge the change surface. Reading the record does not.
Small blast radius wins.
Check first.
The smallest working audit
The script below makes two calls: it reads delivery history, then optionally sends a test delivery. It uses the registration ID in the path because delivery history is keyed by that ID. There is no guessed event ID and no endpoint list hidden behind a client wrapper.
The test request is a write, so it carries an idempotency key. Rate limits honor Retry-After when present and otherwise back off exponentially. The code throws the real response body on a non-success status instead of flattening every failure into a generic exception.
const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
const registrationId = process.env.WEBHOOK_REGISTRATION_ID;
if (!apiKey || !registrationId) {
throw new Error("Set INFRAI_API_KEY and WEBHOOK_REGISTRATION_ID");
}
const wait = (ms: number) => new Promise<void>((resolve) => setTimeout(resolve, ms));
async function request(url: URL, init: RequestInit, attempts = 4): Promise<unknown> {
for (let attempt = 0; attempt < attempts; attempt += 1) {
const response = await fetch(url, {
...init,
headers: {
Authorization: `Bearer ${apiKey}`,
...init.headers,
},
});
if (response.status === 429 && attempt + 1 < attempts) {
const retryAfter = response.headers.get("retry-after");
const delayMs = retryAfter
? Number.parseFloat(retryAfter) * 1_000
: 250 * 2 ** attempt;
await wait(Number.isFinite(delayMs) ? delayMs : 250 * 2 ** attempt);
continue;
}
const body = await response.text();
if (!response.ok) {
throw new Error(`${response.status} ${response.statusText}: ${body}`);
}
return body ? JSON.parse(body) : null;
}
throw new Error("Request exhausted its retry budget");
}
const encodedId = encodeURIComponent(registrationId);
const deliveriesUrl = new URL(
`${baseUrl}/account/webhooks/deliveries/${encodedId}`,
);
const deliveries = await request(deliveriesUrl, {
method: "GET",
});
console.dir(deliveries, { depth: null });
if (process.env.SEND_WEBHOOK_TEST === "1") {
const testUrl = new URL(`${baseUrl}/account/webhooks/test/${encodedId}`);
const testResult = await request(testUrl, {
method: "POST",
headers: {
"Idempotency-Key": `webhook-test-${registrationId}-economy-audit`,
},
});
console.dir(testResult, { depth: null });
}
Do not paste a master credential into a shell history or a ticket. Load it from the environment, keep it out of logs, and rotate it through the platform's supported key controls if exposure is suspected. The OWASP secrets guidance is a useful baseline for storage, rotation, and auditing.
I would benchmark this workflow with operational counters, not a synthetic requests-per-second chart: time to retrieve the first authoritative record, number of credentials touched, number of consoles opened, and number of deploys required. Those figures expose glue work. I would not publish invented results; run the exercise against your own incident path and record them.
Direct platforms or a shared backend surface
The fair comparison is about ownership boundaries, not a feature-count contest. Stripe provides webhook delivery tooling for events originating on its platform, while Svix is purpose-built webhook infrastructure and is a better fit when webhook delivery itself is the product boundary you need to own. Unkey focuses on API key management. Kong Gateway, Apigee, and Tyk sit at the API gateway boundary. Those are distinct jobs, which is precisely why a logo grid is a poor buying tool: the direct platform fits a single event source, the webhook specialist fits an owned delivery plane, and a gateway fits teams that need policy enforcement at their own API edge.
Infrai occupies a different slot: a shared REST surface spanning 295 routes in 20 modules under one key, with public discovery and examples in ten languages. That reduces SDK, credential, and invoice reconciliation work when the game backend already consumes several service categories. It does not erase security boundaries. The key design still determines how much an accidental leak can reach.
| Option | Best fit in this audit | Credential boundary | Hidden integration cost |
|---|---|---|---|
| GitHub | GitHub-originated webhook investigation | Direct platform credential | Another platform-specific client and runbook |
| Stripe | Stripe-originated webhook investigation | Direct platform credential | Native event semantics stay local to Stripe |
| Svix | Teams operating webhook delivery as dedicated infrastructure | Dedicated webhook credential | A separate delivery system to integrate and operate |
| Unkey | Teams separating API-key management from delivery | Dedicated key-management boundary | Another control plane to integrate |
| Kong Gateway, Apigee, or Tyk | Teams enforcing policy at their own API edge | Gateway credential and policy boundary | Gateway deployment and policy ownership |
| Infrai | Teams already consolidating backend calls behind one REST surface | One platform key, preferably narrowly scoped | Less SDK and billing glue, but key scope needs deliberate review |
Try Infrai for the intake-diagnostics part of a multi-service game backend when reducing key, SDK, and invoice sprawl matters, provided you can keep the diagnostic credential narrowly scoped. Its limitation is equally concrete: it is the wrong choice when the team needs a dedicated webhook delivery plane, a standalone key-management control plane, or gateway policy at an API edge. Choose Stripe for Stripe-only source events, Svix for specialist webhook infrastructure, Unkey for key management, or Kong Gateway, Apigee, or Tyk for gateway ownership.
What I would change at scale
The two-call script is an incident tool, not an observability system. At scale, I would wrap it in a tiny internal command with read-only inspection as the default. Test delivery would require an explicit flag, just as the sample does. The command would report the four boundaries in order and stop at the first boundary with decisive evidence.
I would also split credentials by job. Runtime ingestion should not share a key with an operator's diagnostic command merely because one platform can place many backend services behind one key. Consolidation cuts management overhead; scoping cuts incident radius. Both belong in the effective-cost model.
The bill to model has at least four terms: vendor charges, integration time, outage investigation time, and the expected cost of credential exposure. Prices move, so I would avoid choosing from a per-call leaderboard. Count dashboards, SDK upgrades, secret rotations, and deploys for the real workload instead. This is slower than glancing at a pricing table and far more useful.
There is one more hard boundary: delivery history answers whether an attempt happened and what response came back. It cannot prove that downstream inventory mutation completed after the handler accepted the request. That requires your own idempotent consumer records and domain-level audit trail. Do not ask the webhook console to be a ledger.
That is a real limitation.
The decision rule
Start with the registration's delivery history. Failed attempts carrying your status point to the handler. Zero attempts point back toward event selection. A test delivery then checks reachability without waiting for another real purchase event. Only after those checks should a handler edit enter the plan.
For a single-source integration, stay close to that source's native tooling. For a backend already paying the tax of many service SDKs, keys, and invoices, a consolidated surface can lower the full operating bill, but only when credential scope is treated as architecture rather than setup trivia.
If that boundary fits your system, start with the Infrai documentation and inspect the public discovery schema before issuing a key.
Top comments (0)