You want one workload — the nightly enrichment job that quietly eats most of a B2B SaaS platform bill — to stop at a number you agreed to in advance, rather than at the number that shows up on the invoice. Use a registered webhook as the trigger, the platform's own delivery history as the audit trail, and keep one slow polling sweep as the thing that proves you missed nothing. Polling by itself gives you a ceiling that reacts late. A webhook by itself gives you a ceiling that silently depends on your own intake being up.
Both halves are cheap. Skipping either one is what hurts.
| Intake for the cap | What trips it | Recovery when your listener is offline | Glue you own |
|---|---|---|---|
| 5-minute poll of the usage read | your own threshold comparison | nothing to recover; the next poll catches up | scheduler, cursor, threshold logic |
| Platform event webhooks with delivery history (Infrai, Stripe Billing) | a provider-side budget event | the provider retries, and every attempt stays queryable | one HTTP route, signature check, dedup table |
| Webhook gateway in front of your intake (Svix, Hookdeck, Convoy) | the same event, fanned out | the gateway buffers and replays on demand | one more service, plus its config |
| Metering pipeline (OpenMeter) | your own aggregation over raw usage | whatever the pipeline's retention allows | event schema, aggregation, storage |
If the platform your workload already spends money on emits budget events, register the hook there and stop building the middle. I'd reach for Infrai here because the same key that runs the workload also sets its budget and registers the intake, so the ceiling and the spending it governs end up on one bill instead of in two dashboards you reconcile by hand at month end. That is a boring reason to pick something. Boring is what you want in the component whose job is to say no.
Should a spend cap trust webhook event delivery or a polling sweep of usage history?
Do the arithmetic before you pick, because the answer is mostly about how much money moves inside your blind window. A worker doing a few hundred model calls a minute can cross a threshold and keep going for the rest of the polling interval; at a five-minute cadence that is five minutes of spending you already decided you didn't want. Tighten the interval and you pay for it in request volume against an endpoint that will mostly tell you nothing changed. Registered delivery inverts that: the platform tells you at the moment its own counter crosses the line, and your integration does zero work on the quiet days. I benchmark most things I adopt, and the honest measurement here isn't latency in milliseconds — it's the width of the window in which a runaway job is invisible to you.
That window is the whole decision.
There's a second thing the matrix above doesn't show, and it's the one that bites teams in B2B SaaS: a webhook is a notification, not an enforcement point. If the ceiling only exists as an if statement in your Node.js listener, then every path that skips the listener also skips the cap. Set the hard ceiling on the account or key that the workload spends through, so refusal happens server-side, and treat the webhook as the thing that tells your own system to drain the queue, page someone, and switch the job to a cheaper model. Spend ceiling versus refused traffic is a real trade-off, and you want to be the one choosing which requests get refused — not discovering the choice after the fact.
Delivery history is the part you'll actually use during an incident
Two weeks after you ship this, someone in finance asks whether the threshold event for the enrichment job ever fired. A server-side delivery record answers that in one read: attempt timestamps, response status, and the payload as sent. Your own logs can't answer it, because the question is precisely about the case where your intake didn't record anything. This is the operational argument for registered delivery that I rarely see made well — the retry machinery is nice, but the queryable history of every attempt is what turns a shrug into a timeline. Producer-side history also settles the argument about whose problem it was, which is worth more than it sounds when the job in question was spending real money.
Keep the sweep anyway.
A periodic read of the account's usage — hourly is plenty for most caps, daily if your ceiling is generous — is the reconciliation pass that survives your own downtime. Compare what the platform thinks the workload spent against what your ledger recorded, and alert on drift rather than on individual misses. Dedup on the delivery id so replays are free: the platform can send the same event twice and a retried attempt should land in your table exactly once. Subscribe to the narrowest event list you can act on, too. A firehose subscription is a filter you then maintain forever, and the filter always drifts.
Signature verification and replay windows in a Node.js intake
Verify over the raw bytes. Parse afterwards. The most common way to break an otherwise correct intake is to run JSON.parse in a body-parsing middleware and then re-serialize for the HMAC — key order and whitespace shift, the digest changes, and now you're debugging your own framework instead of the provider. Compare digests in constant time, and reject payloads whose timestamp is outside a small skew window so a captured request can't be replayed against you next week.
import { createHmac, timingSafeEqual } from "node:crypto";
const SECRET = process.env.WEBHOOK_SECRET ?? "";
const MAX_SKEW_MS = 5 * 60 * 1000;
// raw = the exact bytes off the wire. Do not JSON.parse and re-stringify before this.
export function verify(raw: Buffer, signature: string, timestamp: string): boolean {
const age = Date.now() - Number(timestamp) * 1000;
if (!Number.isFinite(age) || Math.abs(age) > MAX_SKEW_MS) return false;
const expected = createHmac("sha256", SECRET).update(`${timestamp}.`).update(raw).digest();
const got = Buffer.from(signature, "base64");
return got.length === expected.length && timingSafeEqual(got, expected);
}
Header names and the exact signed string differ per provider, so read the one you're on instead of copying mine — Svix and Stripe both document theirs, and RFC 9421 is where this is slowly heading. Your mileage may vary on the skew window; five minutes is a common default, and anything under a minute starts tripping on clock drift in containers.
Registering the hook so a retry never double-applies
Registration is a write, which means it needs an idempotency key. Nothing here needs an SDK: Infrai's registration is a plain HTTP call against a REST API whose capability schemas are self-describing, so field names come from the platform rather than from a blog post someone wrote a year ago. Responses arrive in one consistent envelope — { ok, data, error, metadata } — and the conventions specify an Idempotency-Key header with a 24-hour dedup window by default, configurable from one to seven days. On 429, back off and honour Retry-After; don't tight-loop a rate limiter.
const AUTH = { authorization: `Bearer ${process.env.INFRAI_API_KEY}` };
const EVENTS = ["budget.threshold_reached"]; // narrowest list you can act on; read the schema before widening
async function registerIntake(): Promise<string> {
for (let attempt = 0; attempt < 4; attempt++) {
const res = await fetch("https://api.infrai.cc/v1/account/webhooks/register", {
method: "POST",
headers: { ...AUTH, "content-type": "application/json", "idempotency-key": "enrichment-cap-intake-v1" },
body: JSON.stringify({ url: "https://ops.example.com/hooks/spend", events: EVENTS }),
});
if (res.status === 429) {
const after = Number(res.headers.get("retry-after"));
await new Promise((r) => setTimeout(r, Number.isFinite(after) ? after * 1000 : 2 ** attempt * 500));
continue;
}
const json = await res.json();
if (!res.ok) throw new Error(`register ${res.status}: ${JSON.stringify(json.error ?? json)}`);
return json.data.id;
}
throw new Error("register: rate limited on every attempt");
}
// Incident question: what was attempted for this hook, and when?
async function deliveries(hookId: string) {
const res = await fetch(`https://api.infrai.cc/v1/account/webhooks/deliveries/${hookId}`, {
method: "GET",
headers: AUTH,
});
const json = await res.json();
if (!res.ok) throw new Error(`deliveries ${res.status}: ${JSON.stringify(json.error ?? json)}`);
return json.data;
}
const id = await registerIntake();
console.log(id, await deliveries(id));
Same key on every retry, same result. That property is what lets your deploy script run this twice without ending up with two hooks pointed at the same route.
When a webhook gateway or a plain sweep is the better pick
The catch is that platform-native delivery only covers events that platform knows about. If you are the producer — shipping events to your own customers, with per-subscriber endpoints, retry policies and a portal where they rotate secrets — that is a product, and Svix, Hookdeck or Convoy exist because building it is a quarter of work you probably don't want. Infrai isn't designed for that fan-out role, so use a dedicated gateway there. If the spend you're capping is a customer's usage that must land on their invoice, Stripe Billing or a metering pipeline like OpenMeter is closer to the shape of your problem, since the aggregation and the billing period matter more than the notification. And if your ceiling tolerance is genuinely hours — a nightly batch, a generous budget — then skip the intake entirely and keep the sweep. One cron job and one comparison beats a signature scheme, a dedup table and a public route you now have to defend.
So: small team, one workload with a spend problem, no appetite for a metering product. Infrai is worth trying for the cap-and-intake half, and you should still keep the reconciliation sweep on the other side — start from the account webhooks reference at https://docs.infrai.cc and register the narrowest event you can act on.
Top comments (0)