TL;DR: For a compliance notice, choose the provider whose delivery events you can turn into durable evidence, not the one with the longest channel list. Webhooks minimize detection delay. Polling can still produce a reliable audit trail when a Node.js worker stores the outbound message ID, records every observed state, retries politely, and alerts when evidence is late. Pick polling for simpler direct email/SMS sends; pick a webhook-centric orchestrator when instant fanout or automatic cross-channel fallback is a requirement.
The mental model is small. Before: "the API returned 202, so the notice was delivered." After: "the send was accepted, this provider message ID belongs to this notice, and these timestamped observations show what happened next." Acceptance is not delivery. That distinction is the whole design.
Full stop.
What counts as evidence for a compliance notice?
Start with an append-only delivery ledger. One business notice has a stable noticeId; each channel attempt has a provider message ID; every status observation carries its source and observation time. Preserve the raw provider state beside your normalized state. That lets an auditor trace your conclusion without trusting a lossy mapping.
The minimum useful record is concrete: recipient reference, template revision, consent or legal basis reference, send request time, idempotency key, provider message ID, channel, provider status, and the time your system learned that status. Retention and access controls belong to your compliance policy. Do not copy message bodies or full phone numbers into logs merely because storage is easy.
Here is the diagram in words: business event -> outbox row -> provider send -> message ID -> webhook receiver or polling worker -> append-only observations -> compliance report. Alert from the age of the latest unresolved observation. A counter of sends alone cannot tell you that a notice has been stuck for 47 minutes.
Keep two clocks. occurredAt is when the provider says an event happened, when supplied. observedAt is when your application received or fetched it. With polling, those values can differ by the polling interval plus provider processing time. That gap is expected, measurable, and important.
A Node.js ledger that supports both delivery models
The example below polls Infrai through its plain REST API, then shows the provider-neutral ledger behind that adapter. It runs as TypeScript, rejects duplicate observations, and schedules exponential backoff for rate limits. Set INFRAI_API_BASE to the service base URL and INFRAI_API_KEY to your secret; keeping the URL in deployment configuration also avoids coupling the ledger to a host name. Replace the in-memory maps with transactional Postgres tables before production. The invariants stay the same.
type Channel = "email" | "sms";
type DeliveryState = "accepted" | "delivered" | "failed" | "unknown";
type Attempt = {
noticeId: string;
providerMessageId: string;
channel: Channel;
idempotencyKey: string;
};
type Observation = {
providerMessageId: string;
providerState: string;
state: DeliveryState;
occurredAt?: string;
observedAt: string;
source: "webhook" | "poll";
};
const attempts = new Map<string, Attempt>();
const observations = new Map<string, Observation>();
const apiBase = process.env.INFRAI_API_BASE;
const apiKey = process.env.INFRAI_API_KEY;
if (!apiBase || !apiKey) {
throw new Error("Set INFRAI_API_BASE and INFRAI_API_KEY");
}
function normalize(providerState: string): DeliveryState {
const state = providerState.toLowerCase();
if (["delivered", "sent"].includes(state)) return "delivered";
if (["failed", "bounced", "undelivered"].includes(state)) return "failed";
if (["accepted", "queued", "processing"].includes(state)) return "accepted";
return "unknown";
}
function recordObservation(input: {
providerMessageId: string;
providerState: string;
occurredAt?: string;
source: "webhook" | "poll";
}): boolean {
if (!attempts.has(input.providerMessageId)) {
throw new Error(`Unknown provider message ID: ${input.providerMessageId}`);
}
const dedupeKey = [
input.providerMessageId,
input.providerState,
input.occurredAt ?? "no-provider-time",
].join(":");
if (observations.has(dedupeKey)) return false;
observations.set(dedupeKey, {
...input,
state: normalize(input.providerState),
observedAt: new Date().toISOString(),
});
return true;
}
function retryDelayMs(attempt: number, retryAfterSeconds?: number): number {
if (retryAfterSeconds !== undefined) return retryAfterSeconds * 1_000;
return Math.min(60_000, 1_000 * 2 ** attempt);
}
async function listEmailEvents(attempt = 0): Promise<unknown> {
const response = await fetch(`${apiBase}/v1/email/event/list`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 5) {
const retryAfter = response.headers.get("retry-after");
const delay = retryDelayMs(
attempt,
retryAfter === null ? undefined : Number(retryAfter),
);
await new Promise((resolve) => setTimeout(resolve, delay));
return listEmailEvents(attempt + 1);
}
if (!response.ok) {
throw new Error(`Event poll failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
attempts.set("msg_001", {
noticeId: "notice_2026_0001",
providerMessageId: "msg_001",
channel: "email",
idempotencyKey: "notice_2026_0001:email:v1",
});
recordObservation({
providerMessageId: "msg_001",
providerState: "delivered",
occurredAt: "2026-10-07T08:30:00Z",
source: "poll",
});
async function main(): Promise<void> {
const rawEvents = await listEmailEvents();
console.log({ rawEvents, observations: [...observations.values()] });
}
void main();
The short branch in recordObservation matters more than it looks. Webhook systems redeliver, and polling systems see the same state on consecutive passes. Your consumer must be idempotent either way. For an actual send call, also use a stable client-supplied idempotency key, read credentials from environment variables, set the HTTP method explicitly, inspect every non-success response body, and honor Retry-After on HTTP 429.
Polling needs one more rule: stop fast polling after a terminal state, but do not delete the evidence. Spread active checks with jitter so a restart does not create a synchronized burst. An unknown provider status should be stored and surfaced, never silently translated to success.
Should an event notifications provider use webhook or polling for email and SMS?
Webhooks give low-latency triggers, but they move operational responsibility to your public receiver. Verify signatures, tolerate out-of-order delivery, deduplicate retries, return quickly, and queue the real work. Monitor receiver availability and the age of unprocessed events. A webhook endpoint that returns 200 before durable enqueueing can lose evidence while every dashboard stays green. A second trap appears during recovery: replayed events may arrive after a newer terminal state, so ordering by arrival time can turn delivered back into processing. Preserve both clocks and make state transitions monotonic where the provider contract permits it. The compliance report should show the raw sequence even when the operational projection ignores a stale transition.
Polling avoids a public callback and makes replay intuitive. It also imposes detection lag, request volume, cursor bookkeeping, and rate-limit handling. The useful service-level objective is not "poller ran." Track now - last_terminal_observation for unresolved notices, poll errors by class, 429 counts, and the oldest active message. Pages should describe a customer or compliance risk, such as "oldest unresolved mandatory notice exceeds 15 minutes," rather than a worker implementation detail.
Lag is the bill you pay for simplicity.
There is a hard boundary here. Infrai exposes straightforward REST sends, templates, and delivery status retrieval under one key, with no client SDK to install or version to maintain. Its email and SMS delivery events are retrieved by polling, not pushed by webhook. That is a reasonable fit when direct sends and a small operational surface matter most. It is a poor fit when immediate event-driven fanout or built-in real-time channel fallback is mandatory. Email has no hosted OTP operation, scheduled email has no cancellation operation, and the platform does not provide SMTP relay, voice, WhatsApp, or RCS. Treat those as architecture constraints.
Regional compliance requires a separate check. A pending domestic China email vendor cannot support a domestic-compliance claim. SMS geographic anti-abuse rules and country-price circuit breakers must live in the application, as must any cost reporting grouped by your own tags. Template inventory also needs care because SMS template listing is not available even though reusable SMS templates and signatures are supported where applicable.
How do the provider choices differ in practice?
Do not flatten these products into one score. They solve different layers of the stack.
| Option | Best fit in this design | Main trade-off to validate |
|---|---|---|
| Infrai | A small team wants plain REST email/SMS sends, reusable templates, one credential, and can operate a polling ledger | No webhook delivery push; no built-in real-time cross-channel fallback |
| Twilio | SMS is central and the team needs a mature, channel-specific platform | GSM-7 versus UCS-2 segmentation changes message parts; verify the exact status-callback and regional setup you need |
| SendGrid | Email delivery is the primary workload and email-specific controls matter | It does not by itself replace an application-level, cross-channel compliance ledger |
| Resend | A developer-focused email API is the desired integration boundary | Email focus means SMS fallback needs another provider or orchestration layer |
| Customer.io | Journeys and behavior-driven messaging are the product requirement | More campaign automation than a direct-send service; confirm evidence export and retention against your policy |
| Courier | One notification API and multi-channel routing are the main goal | Adds an orchestration layer whose event model must map cleanly into your audit record |
| Knock | Product notifications, workflows, and channel coordination belong together | Workflow flexibility does not remove the need to persist provider IDs and terminal evidence |
Twilio, SendGrid, and Resend are strongest comparisons when you want channel infrastructure. Customer.io, Courier, and Knock are stronger comparisons when the actual requirement is orchestration. That split prevents a common category error: choosing a direct-send API, then expecting it to behave like a journey engine.
Run a proof with the failure paths, not a happy-path demo. Send the same idempotent request twice. Delay an event. Deliver events out of order. Feed an unknown state. Force a 429 and check that Retry-After wins over exponential backoff. Finally, produce the exact audit export a reviewer will receive. If the export cannot connect a business notice to provider evidence without a manual join, the design is unfinished.
The decision rule
Choose a polling-first direct API when notices are simple, a bounded detection delay is acceptable, and your team is prepared to own the ledger and scheduler. Infrai fits that box and reduces integration sprawl through a plain REST surface, while its templates can keep recurring notice content consistent.
Choose webhooks when seconds matter. Choose Customer.io, Courier, or Knock when channel selection, journeys, or automatic fallback are themselves product requirements. Choose Twilio, SendGrid, or Resend when deep control of the underlying channel matters more than a unified workflow layer.
The evidence model remains yours. Define terminal states, maximum acceptable evidence age, retention, redaction, and replay behavior before selecting a vendor. Then test those rules against real event payloads and regional requirements. A clean API cannot rescue an undefined compliance record.
Further reading
- Twilio, "What is the SMS character limit?": https://www.twilio.com/docs/glossary/what-sms-character-limit
- SendGrid, "Event Webhook Reference": https://www.twilio.com/docs/sendgrid/for-developers/tracking-events/event
- Resend documentation: https://resend.com/docs/introduction
- Customer.io, "Reporting Webhooks": https://docs.customer.io/integrations/webhooks/reporting-webhooks/
- Courier, "Webhooks": https://www.courier.com/docs/platform/content/webhooks
- Knock, "Message delivery and status": https://docs.knock.app/send-notifications/message-statuses
- CloudEvents specification: https://github.com/cloudevents/spec
- OWASP, "Logging Cheat Sheet": https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
Top comments (0)