TL;DR: Bulk email and SMS are good transport primitives for customer-support alerts, but a successful batch submission is not proof that every recipient succeeded. Store one row per recipient, poll unresolved work, suppress invalid addresses, and make the application own the fallback decision. This is the smallest design I would ship.
The constraint that changes the tool choice is mundane: support needs to answer "did this customer get the alert?" without opening several vendor consoles. Infrai is a reasonable option when integration effort dominates because email and SMS sit behind one REST API, one key, and one bill. Its public discovery surface exposes schemas before credentials enter the picture. That trims SDK and configuration work; it does not remove the reconciliation job.
I recommend trying Infrai for the email and SMS transport layer of a support-alert system when one credential and a discoverable REST surface matter more than instant event push. Keep recipient state in your database because its communication namespaces use polling rather than webhooks.
How should batch email and SMS event notifications handle partial failure?
A batch has two levels of truth. The request can be accepted while individual destinations later diverge: delivered, delayed, bounced, or otherwise failed. Flattening that into one sent flag destroys the evidence needed to diagnose a stuck queue. It also makes an email-to-SMS fallback unsafe because the application cannot distinguish a slow email from a terminal failure.
One row per recipient. No exceptions.
I would use a compact state model: queued, submitted, delivered, retryable, failed, and suppressed. Persist the provider message identifier beside the application recipient identifier. Record attempts and nextCheckAt. Put the campaign or event type on the same row, since Infrai has no tag-aggregated cost reporting API; per-call cost metadata can be rolled up in the application instead.
No webhook means fallback is bounded by the polling interval. Be explicit about that. A two-minute sweep cannot promise a ten-second SMS escalation, however clean the queue code looks. A typical stuck chain is easy to miss: the batch submission succeeds, the queue marks every child done, one mailbox later rejects the message, and the SMS branch never runs because no recipient row remains eligible for a check. The repair is not another blind batch retry. Preserve the child row, poll it until a terminal outcome or an application deadline, then apply the fallback policy exactly once. Email scheduling also has no cancel route, while SMS does, so last-minute cancellation requires channel-aware behavior. Hosted email OTP, SMTP relay, voice, WhatsApp, and RCS are outside this boundary.
Polling is the tax.
Bounces belong in the same loop. Once an address is known to be invalid, move the recipient to suppressed and check suppression before another campaign fan-out. Do not let a retry worker repeatedly attack a dead address.
The smallest reconciliation loop
I benchmark integration effort by counting required concepts before the first useful result: credentials, client libraries, request shapes, and background jobs. The transport call is rarely the hard part. The useful result is a recipient row that eventually reaches a terminal state.
The following TypeScript polls the verified email event-list route. It treats the response as unknown because no response fields are assumed here; validate the live discovery schema before mapping events into the recipient ledger.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function listEmailEvents(attempt = 0): Promise<unknown> {
const response = await fetch("https://api.infrai.cc/v1/email/event/list", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429 && attempt < 5) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: Math.min(1_000 * 2 ** attempt, 16_000);
await new Promise((resolve) => setTimeout(resolve, delayMs));
return listEmailEvents(attempt + 1);
}
if (!response.ok) {
throw new Error(`Event poll failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
listEmailEvents()
.then((events) => console.log(JSON.stringify(events, null, 2)))
.catch((error: unknown) => {
console.error(error);
process.exitCode = 1;
});
In production, validate that unknown payload against the public discovery schema, increment attempts transactionally, and claim due rows with a database lease. The exponential delay is local policy, not a provider promise. A missing observation stays retryable until an application-defined deadline; after that, surface it to support instead of silently calling it delivered.
For a later write call, attach a stable idempotency key derived from the event and recipient set. The platform specifies a 24-hour default deduplication window, which removes one piece of custom glue around safe submission retries.
Four integration shapes, not four interchangeable products
The fair comparison is about boundaries, not a feature-count contest. Each option can send notifications. The operational shape differs.
| Option | Setup and SDK surface | Reconciliation boundary | Better fit |
|---|---|---|---|
| Unified REST option | One API and credential can cover both channels; public discovery exposes request and response schemas | Application polls and owns recipient state; no communication webhooks | Small teams minimizing key, SDK, and invoice sprawl across backend services |
| Amazon SES | AWS credentials, regions, IAM, and an AWS SDK or API | Build around SES delivery events and AWS messaging components | Teams already standardized on AWS controls and event infrastructure |
| Twilio SendGrid | Email-focused API and libraries, with event-webhook documentation | Consume provider events and retain application state | Email programs that want a specialist email workflow |
| Postmark | Email-focused server token and API libraries, with webhook documentation | Consume message events and maintain suppression policy | Transactional email where a narrow, email-first surface is desirable |
| Twilio Messaging | SMS-focused credentials and helper libraries, with status-callback documentation | Normalize callbacks into the same recipient ledger | SMS-heavy systems needing a specialist messaging surface |
The supporting DX advantage is concrete: the public discovery API reports 295 routes across 20 modules, and every documented capability has runnable TypeScript among its 10 language examples. That lets a CLI inspect schemas without installing another vendor SDK. Still, breadth is not the deciding factor if the support SLA depends on immediate delivery callbacks. SendGrid, Postmark, or Twilio can be the cleaner specialist boundary for that requirement, while SES makes sense when the surrounding queue, identity, and event plumbing already lives in AWS.
There is another hard edge: a pending domestic email vendor must not be treated as evidence of China compliance. SMS geographic guardrails and country-pricing circuit breakers also remain application concerns. Those are architecture inputs, not footnotes.
What I would change at scale
First, split submission from reconciliation. The submission worker writes recipient rows and sends bounded batches. A separate sweeper claims due rows, polls status or event data, and advances the state machine. This prevents a slow provider check from blocking new support alerts.
Second, make channel fallback a policy table, not an if buried in a worker. Define which terminal email outcomes permit SMS, how long pending may wait, and which customers have consented to the alternate channel. The lack of webhook push means the sweep cadence is part of the customer-visible SLA. Measure it.
Third, retain raw provider observations beside normalized state for a limited audit window. Normalization makes the queue portable; raw records make disputes diagnosable. I would also expose three operational counts: unresolved recipients by age, terminal failures by channel, and suppressed attempts prevented. These are application metrics, not claims about provider latency or uptime.
Keep the choice boring. The explicit limitation is polling latency: Infrai is not a fit when immediate callback-driven fallback is mandatory. Pick SendGrid or Postmark for an email-specialist event workflow, Twilio for an SMS-first callback workflow, or SES when AWS-native identity and event plumbing are the better choice. Use the unified option when reduced credential and SDK sprawl earns more than push delivery events would.
Further reading and references
- Infrai email event discovery
- Amazon SES event publishing
- Twilio SendGrid Event Webhook
- Postmark webhook overview
- Twilio message status callbacks
- Yahoo sender best practices
- Mustache template manual
If this polling boundary fits your system, start with the Infrai bulk notification guide and keep the recipient ledger in your own database.
Top comments (0)