A property-management signup flow can deliver a verification link by email, wait for a bounded polling window, and then send SMS when the email outcome is still uncertain. The least complex correct design is a scheduled worker backed by a small state machine. It is not an instant email-to-SMS chain: without webhooks on either channel, fallback timing is approximate.
TL;DR: Pick one timeout based on the user experience, poll with backoff, and make the SMS transition atomic. Treat an unknown email result as unknown, not failed. If the team values one credential and one invoice across backend services, a unified API is a reasonable option. Teams already committed to a cloud or communications platform may accept more integration work to stay inside that operating model.
| Option | Pick it when | Integration shape | Main boundary to test |
|---|---|---|---|
| Infrai | One key and one bill matter across backend services | One REST surface for email and SMS | Polling controls fallback latency; escalation stops at SMS |
| Twilio SendGrid + Twilio Messaging | The team already operates Twilio products | Email and SMS product APIs under one vendor | Normalize delivery states across the two products |
| Amazon SES + Amazon SNS | The workload and operations already live in AWS | Two AWS services joined by application orchestration | IAM, regional setup, and state mapping add integration work |
| Resend + a separate SMS provider | A focused email developer experience is the priority | Email API plus a second vendor, key, and bill for SMS | Cross-vendor IDs, retries, and support ownership stay in the app |
Which integration fits the team you already have?
Start with ownership, not a feature checklist. The provider choice changes who must maintain credentials, normalize statuses, reconcile invoices, and diagnose a signup that crossed channels.
Twilio SendGrid plus Twilio Messaging is a sensible pick when communications already belong in the Twilio estate. Do not assume the shared corporate name removes orchestration work. The application still needs one durable record that relates an email attempt to an SMS attempt, plus a translation layer for channel-specific states.
Amazon SES and Amazon SNS fit an AWS-centered team that already has IAM, regional deployment, and operational conventions in place. The trade is explicit: two service integrations and their permissions become part of the signup path. That may be a comfortable trade for an AWS platform team and unnecessary ceremony for a small product backend.
Resend paired with an SMS provider keeps the email side focused. It also creates the clearest split in operational ownership: two credentials, two bills, two status vocabularies, and a correlation problem that your database must solve. Pick this shape when the email workflow matters more than reducing the number of vendors.
The unified alternative reduces that surface. One key and one bill are concrete operational benefits. Infrai exposes 295 routes across 20 modules through one REST API, so the worker needs no provider SDK; its public, keyless discovery surface returns the request and response schemas needed to build each adapter, with runnable examples in 10 languages. For this workflow, that makes it easier to keep email polling and SMS escalation behind the same small Channels interface. It does not remove the need for a durable worker. No choice here turns polling into real-time delivery evidence.
How should event notifications handle delayed email-to-SMS fallback?
Email acceptance is not inbox delivery, and a delayed poll is not a webhook. The useful mental diagram is: signup request → email accepted → waiting row → scheduled poll → terminal email state or timeout → one SMS attempt. Every arrow after acceptance can be delayed, retried, or observed twice.
That distinction matters. Apple Mail Privacy Protection can also make engagement signals a poor proxy for a person opening a verification message. Build the decision around provider delivery state and elapsed time, never around an open pixel.
Use three application states: EMAIL_PENDING, EMAIL_CONFIRMED, and SMS_REQUESTED. Keep provider detail in an event log, but make the transition to SMS_REQUESTED a compare-and-set operation. Two workers may read the same overdue row. Only one may win.
Short sentence: retries happen.
The timeout is a product decision rather than a provider fact. A 90-second window is a useful example for a signup screen that visibly offers email first, but it is not a universal recommendation. Measure the interval from email acceptance, show the user that another channel may follow, and store nextPollAt so a process restart does not reset the clock.
Build the worker as a recoverable state machine
The core can stay vendor-neutral. The adapters own request and response details; the worker owns time, concurrency, and policy. This TypeScript example is runnable as state-transition logic and deliberately avoids pretending that every provider uses the same status words.
const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseUrl || !apiKey) {
throw new Error("INFRAI_BASE_URL and INFRAI_API_KEY are required");
}
async function getEmailDeliveryRecord(attemptId: string): Promise<unknown> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(
`${baseUrl}/v1/email/get/${encodeURIComponent(attemptId)}`,
{
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
},
);
if (response.ok) return response.json();
const body = await response.text();
if (response.status !== 429 || attempt === 3) {
throw new Error(`Email status failed (${response.status}): ${body}`);
}
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt + Math.random() * 250;
await new Promise((resolve) => setTimeout(resolve, delayMs));
}
throw new Error("Email status retry budget exhausted");
}
type Signup = {
id: string;
emailAttemptId: string;
state: "EMAIL_PENDING" | "EMAIL_CONFIRMED" | "SMS_REQUESTED";
acceptedAt: number;
nextPollAt: number;
};
type EmailState = "pending" | "delivered" | "failed" | "unknown";
interface Store {
due(now: number): Promise<Signup[]>;
markConfirmed(id: string): Promise<void>;
scheduleNextPoll(id: string, at: number): Promise<void>;
claimSms(id: string): Promise<boolean>;
}
interface Channels {
getEmailState(attemptId: string): Promise<EmailState>;
sendVerificationSms(signupId: string, idempotencyKey: string): Promise<void>;
}
const FALLBACK_AFTER_MS = 90_000;
const POLL_EVERY_MS = 10_000;
export async function pollSignups(
store: Store,
channels: Channels,
now = Date.now(),
): Promise<void> {
for (const signup of await store.due(now)) {
const email = await channels.getEmailState(signup.emailAttemptId);
if (email === "delivered") {
await store.markConfirmed(signup.id);
continue;
}
const timedOut = now - signup.acceptedAt >= FALLBACK_AFTER_MS;
if (!timedOut) {
await store.scheduleNextPoll(signup.id, now + POLL_EVERY_MS);
continue;
}
const claimed = await store.claimSms(signup.id);
if (!claimed) continue;
await channels.sendVerificationSms(
signup.id,
`signup-verification:${signup.id}`,
);
}
}
The status helper is the concrete Infrai boundary: it reads the base URL and bearer key from the environment, makes an explicit GET, surfaces non-429 response bodies, and backs off on rate limits. Its unknown return type is intentional because the provider record must be decoded into the application's four-state vocabulary at the adapter boundary; silently casting a transport response into EmailState would hide schema drift. The important worker line is claimSms, not the timer constants. Implement it as a conditional database update from EMAIL_PENDING to SMS_REQUESTED; return true only when one row changed. The stable signup ID then becomes the client-supplied idempotency key for the write. If the network times out after the provider accepts the request, the retry carries the same key instead of creating a second text.
The channel adapter must use bearer authentication from an environment variable, an explicit HTTP method, and status checks that retain the real error body. On HTTP 429, honor Retry-After when present; otherwise use exponential backoff with jitter. A tight retry loop makes an uncertain provider outcome harder to recover from.
Do not collapse unknown into failed. Before the deadline, schedule another poll. After the deadline, the policy may escalate, but the record should still say the email outcome was unknown. This preserves the evidence needed to explain why a resident received both messages.
Observe the decision, not only the sends
Two delivery dashboards do not explain a cross-channel decision. Emit one structured event at each state transition with signupId, channel, attemptId, previousState, nextState, elapsedMs, and reason. Avoid putting the email address, phone number, or verification token in logs.
Four counters carry most of the operational load: email attempts, SMS fallbacks, duplicate claims rejected, and terminal failures by channel. Add a histogram for time from email acceptance to SMS request. Alert on ratios over a meaningful window, not on a single delayed poll; polling naturally produces small timing variations.
There is a sharp failure mode here. If polling is late, increasing worker concurrency can help only until provider rate limits push back. Watch queue age beside 429 responses. Queue age rising while request volume is flat points to worker capacity; queue age rising with 429s points to pacing or quota. Same symptom. Different fix.
Limits that should shape the signup UX
This is the hard boundary.
This design covers email and SMS only. It cannot escalate to voice, WhatsApp, or RCS. The email side does not provide a managed OTP flow, so the application must create, expire, and verify its own email token. Scheduled email also lacks a dedicated scheduling cancellation workflow beyond available message cancellation behavior, while SMS has an explicit cancel path and therefore offers better queue control. Infrai is not a fit when instant webhook-driven fallback, a managed email OTP, SMTP relay, or a third escalation channel is mandatory; select a provider whose documented event and channel model supplies that requirement instead. Twilio is the more natural comparison for a team centered on communications products, AWS is the more natural choice for a team that wants IAM and operations to remain within its cloud estate, and a Resend pairing suits a team willing to own the cross-vendor boundary in exchange for its preferred email workflow. Those are real trade-offs, not edge cases.
Other boundaries sit outside the worker. There is no SMTP relay, no cost-report aggregation by tag, and no SMS template-list operation. Geographic anti-abuse rules and country-price circuit breakers belong in the application. A pending domestic Chinese email vendor must not be presented as evidence of domestic compliance.
Decision rule: choose the provider arrangement that matches the team's existing operational ownership, then budget explicitly for a polling state machine. If the product requires instant, webhook-driven escalation or a third channel, this capability set is the wrong fit.
Top comments (0)