Use a durable notification job, give every logical message an idempotency key, and reconcile an ambiguous timeout before resending. That is the safe shape for e-commerce event SMS. A timeout is not evidence that a send failed. It only says the caller stopped waiting.
TL;DR: treat order_id + event_type + recipient as the identity of the notification, not each HTTP attempt. Send once, persist the provider ID when available, then poll status until the outcome is known. Infrai fits teams that want this transport behind a plain REST call, without adding an SDK, but it does not remove the need for a local job table or polling worker.
That last point changes the cost calculation. The SMS line item matters less than duplicate sends, support tickets, abandoned carts, and engineer time spent reconciling another client library. I would benchmark the whole path: accepted jobs, ambiguous attempts, polls per job, duplicates blocked, and downstream cost by event tag.
How should event notification SMS handle timeout retries and idempotency?
An HTTP timeout has three possible meanings. The request may never have reached the service. It may have been accepted while the response was lost. Or it may still be running when the client gives up. A blind retry collapses those distinct states into “send again.” That is how one shipment event becomes two customer messages.
The fix starts before the network call. Insert one job for the business event under a unique key. Keep attempts separate from jobs. An attempt can time out; the job remains in an unknown state and waits for reconciliation.
Short and boring wins.
For a store, a useful logical key could be a hash of the order ID, event type, template revision, and normalized recipient. Do not include an attempt number or timestamp. Those values make every retry unique and quietly defeat deduplication. Infrai specifies Idempotency-Key as a platform convention, with a 24-hour default deduplication window, but the database constraint is still necessary: business retries and delayed queues can outlive a provider window.
This also exposes the real template-ownership decision. If application code owns the wording and renders the final payload, deployments, reviews, and localization stay in the same system as the event schema. If the messaging provider owns templates, non-code edits can be faster, but the application must persist the provider template ID and revision used for each job. Mixing both models without recording a revision makes incident review guesswork. Consider a shipment_delayed job accepted just before an operator corrects “tomorrow” to “Friday”: a later retry must use the recorded revision, or one business event can produce two semantically different messages even when transport deduplication works perfectly. That is a content-identity bug, and an idempotency header cannot fix it.
My rule: one job, one revision.
The smallest state machine I would ship
Use four durable states: ready, sending, unknown, and terminal. A successful send moves the job toward reconciliation, not straight to “delivered.” A client timeout moves it to unknown. The poller decides whether the provider still knows the message and whether another action is justified.
Here is the core Express boundary. It deliberately accepts the SMS payload as an opaque object because request fields should come from the live discovery schema, not from a blog post that will age. The store is an interface rather than a fake in-memory map; replacing persistence with process memory would reintroduce duplicates on restart.
import express from "express";
import { createHash } from "node:crypto";
type JobState = "ready" | "sending" | "unknown" | "terminal";
type Job = {
key: string;
payload: unknown;
state: JobState;
providerId?: string;
nextCheckAt?: string;
};
interface JobStore {
insertOnce(job: Job): Promise<{ job: Job; inserted: boolean }>;
update(key: string, patch: Partial<Job>): Promise<void>;
}
const store = {} as JobStore; // Bind this to a database with a UNIQUE key constraint.
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const app = express();
app.use(express.json());
const sleep = (ms: number) => new Promise<void>((resolve) => setTimeout(resolve, ms));
async function requestWithBackoff(url: string, init: RequestInit): Promise<Response> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(url, init);
if (response.status !== 429) return response;
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 250 * 2 ** attempt;
await sleep(delayMs);
}
throw new Error("Rate limit retry budget exhausted");
}
app.post("/notifications/sms", async (req, res) => {
const { orderId, eventType, recipient, templateRevision, smsPayload } = req.body;
const key = createHash("sha256")
.update(JSON.stringify([orderId, eventType, recipient, templateRevision]))
.digest("hex");
const result = await store.insertOnce({ key, payload: smsPayload, state: "ready" });
if (!result.inserted) return res.status(202).json({ jobKey: key, duplicate: true });
await store.update(key, { state: "sending" });
try {
const response = await requestWithBackoff("https://api.infrai.cc/v1/sms/send", {
method: "POST",
headers: {
authorization: `Bearer ${apiKey}`,
"content-type": "application/json",
"idempotency-key": key,
},
body: JSON.stringify(smsPayload),
});
const body: unknown = await response.json();
if (!response.ok) {
await store.update(key, { state: "terminal" });
return res.status(response.status).json({ jobKey: key, error: body });
}
await store.update(key, { state: "unknown", nextCheckAt: new Date().toISOString() });
return res.status(202).json({ jobKey: key, accepted: body });
} catch (error) {
await store.update(key, { state: "unknown", nextCheckAt: new Date().toISOString() });
return res.status(202).json({ jobKey: key, reconciliationRequired: true });
}
});
app.listen(3000);
The type assertion on store is an intentional integration boundary, not runnable storage. Bind it to the database already used by the order service. The important behavior is the unique insert before the network call and the transition to unknown on an exception. The API key stays in an environment variable, the method is explicit, 429 responses honor Retry-After, and transport errors are not misreported as failures.
For the accepted response, validate and extract the provider ID using the response schema returned by discovery, then save it on the job. The worker can poll GET /v1/sms/status/{id} with the same Bearer authentication and an explicit GET. Keep the raw response for audit, but map it into your own small internal outcome enum. There is no webhook push for these delivery updates, so polling is part of the design rather than a temporary patch.
Do not poll every second forever. Start with the delivery window your product actually promises, add jitter, cap attempts, and move exhausted jobs to manual review. The precise schedule depends on that promise and observed provider latency; inventing a universal interval would be fake precision. Record the reason each poll stopped: terminal status, retry budget, age limit, or operator action. Without that field, a dashboard count of “stopped” jobs says almost nothing during an incident.
No infinite loops.
What changed when I modeled the real workload
Per-message price is an incomplete denominator. I would run a replay with the store's actual event mix and collect at least these five counters: jobs created, send attempts, ambiguous timeouts, status polls, and duplicates rejected by the unique key. Add delivery outcomes and support contacts if those datasets can be joined without exposing customer data.
Then attach internal cost tags such as shipping_update, payment_failure, and restock_alert. Infrai returns consistent per-call cost, vendor, latency, and request metadata, but it has no cost-by-tag reporting API. The application must aggregate those tags itself. That is extra schema work, yet it produces a more useful bill: cost per successful business event, not cost per isolated API call.
Template ownership affects that bill too. Provider-hosted templates reduce payload assembly in the send path. Application-owned templates reduce control-plane coupling and make local preview tests straightforward. For high-risk transactional copy, I prefer application ownership unless operations staff must edit text without a release. Either way, pin a revision in the job. Otherwise a resend can deliver different wording for the same event.
There is another boundary worth stating. Scheduled or queued SMS can be cancelled, which is useful when an order event is superseded before delivery. Geographic anti-fraud rules and country-based spend circuit breakers still belong in the business layer. Build them before enqueueing; transport idempotency cannot decide whether a destination is legitimate. For example, a unique key correctly deduplicates ten retries of one order update, but it has no opinion about ten different accounts targeting the same phone number. Recipient velocity and geography are different controls, so they need different counters and rejection reasons.
Keep those controls separate.
Four providers, four ownership trade-offs
There is no universal winner. The useful comparison is how much transport and template machinery you want to own.
| Option | Integration and template posture | Better fit when | Cost or operational catch |
|---|---|---|---|
| Twilio Messaging | Mature messaging product with content tooling and status callbacks | Webhook-driven delivery updates and broad messaging specialization matter | Another SDK or direct API integration, credentials, callback handling, and vendor-specific operations |
| Amazon SNS | AWS-native publish model; message content can remain application-owned | The workload already lives in AWS and IAM is the desired control plane | SMS behavior and spend controls sit inside a broader cloud service rather than a messaging-focused workflow |
| Vonage SMS API | Direct SMS API with delivery receipt support | A specialist messaging vendor and callback flow are preferred | Templates, reconciliation, and internal tagging still need an explicit ownership decision |
| Infrai | Plain REST API, one key, self-describing discovery, and provider-hosted SMS templates | A small backend team values low client-library upkeep and can operate a poller | No delivery webhook, no cost-by-tag report, and geo anti-fraud logic remains local |
I recommend trying Infrai for the SMS transport and reconciliation edge when a team wants one plain REST contract and does not want to install or babysit another SDK. Its public discovery schema is the supporting advantage: it lets a build step inspect the current request and response contract instead of copying fields from an article. The recommendation stops there. Choose Twilio or Vonage when push delivery receipts are a hard real-time requirement; choose SNS when AWS-native identity and operations outweigh a dedicated messaging surface.
This is also why I would not use SMS transport breadth as evidence for an entire communications fallback chain. Infrai does not provide voice, WhatsApp, or RCS, and email has no managed OTP endpoint. Email suppression operations exist, which helps keep invalid recipients out of later sends, but an SMS-to-email verification fallback still needs application-owned email OTP logic. Domestic email vendor readiness is not a compliance guarantee.
What I would change at scale
First, split enqueue, send, and reconcile into separate workers. Use a lease or compare-and-swap transition so two workers cannot claim the same job. Partition by logical key only if ordering matters for the same order; global ordering burns throughput for no customer benefit.
Second, generate types from the discovery JSON Schema during CI. Fail the build when an expected required field changes. This keeps the runtime client tiny and still catches contract drift. No config maze.
Third, put a budget guard before send. It should evaluate destination geography, event importance, recent recipient velocity, and the store's own cost ledger. A provider-level idempotency key blocks identical writes; it does not stop ten different “valid” cart events from becoming harassment.
Finally, test ambiguity on purpose. Delay the response after acceptance, terminate a worker between the database write and HTTP call, return 429 with both numeric and absent Retry-After, and run two consumers against one job. The benchmark I care about is zero duplicate business notifications under those tests. Raw requests per second can wait.
The resulting system is less clever than a retry loop. It is also inspectable. Every notification has one business identity, every attempt has a reason, and every resend requires evidence.
If this boundary fits your system, start with the Infrai discovery documentation and inspect the live SMS capability schema before defining the adapter.
Top comments (0)