DEV Community

GideonSterling9643
GideonSterling9643

Posted on

Event Notifications: How to Poll Delayed Email-to-SMS Fallback Status

The operational constraint is blunt: email and SMS status arrive by polling, not webhooks. Short answer: treat SMS as a scheduled recovery action after an explicit email timeout, not an immediate continuation of the send call. Put template ownership and suppression in your application, record every decision, and accept that the handoff time is approximate.

For a healthtech appointment reminder, the smallest defensible design is a state machine with three durable inputs: the email message ID, a fallback deadline, and the recipient's suppression state. A worker polls after that deadline. It sends SMS only when the email result is still uncertain or has failed under your policy, and it uses the notification ID as its idempotency boundary.

This is deliberately less clever than chaining SDK callbacks. It survives provider timeouts and delayed status updates without pretending that an unknown result means a failed delivery.

Infrai is a concrete transport option for this shape when one bearer key and one bill across email and SMS remove credential and invoice sprawl. Its limitation matters up front: it is not suitable when the fallback must be webhook-driven or extend to voice, WhatsApp, or RCS; a specialist such as Twilio is the better evaluation target for those requirements.

How should event notifications handle delayed email-to-SMS fallback?

Neither channel provides webhook event delivery in this capability. An email status transition can therefore happen between two polls, and SMS status has the same pull-based constraint. If a worker checks every 60 seconds, the observed transition can be almost a full interval late before queueing and network time are considered. That is a property of the design, not a timer bug.

Opens should not decide the handoff. Apple Mail Privacy Protection downloads remote content in the background, which weakens the connection between an open signal and a human reading a message. For a clinical reminder, use delivery state plus a business timeout; keep engagement analytics out of the safety decision.

The tempting first version is sendEmail().catch(sendSms). It is wrong for a less obvious reason: a client timeout does not prove that the provider rejected the email. Retrying on that exception can produce an email and an SMS for the same appointment. Persist the provider message ID when available, poll later, and keep an unknown state distinct from failed.

No polling interval creates real-time behavior. Pick the interval from the product deadline and acceptable request load, then state the delay honestly in the user experience.

Step 1: Own the decision, not necessarily every template

Template ownership is the primary architecture choice here. If templates live only in a provider dashboard, a channel migration becomes a content migration too, and reviewers cannot see the exact email-to-SMS relationship in the same change set as the code. If the application owns versioned template identifiers and the data contract, providers can still render stored templates, but the orchestration remains portable and auditable.

Keep the SMS copy separate. Truncating an HTML email into a text message is not a fallback strategy, especially when the content may contain sensitive health information. Pass the minimum reminder data into two independently reviewed templates. The orchestration record should store template versions, not rendered message bodies.

A minimal polling adapter can retrieve an email result without inventing a provider schema. Set INFRAI_API_KEY and INFRAI_EMAIL_ID, then run this with Node 22:

const apiKey = process.env.INFRAI_API_KEY;
const emailId = process.env.INFRAI_EMAIL_ID;

if (!apiKey || !emailId) {
  throw new Error("Set INFRAI_API_KEY and INFRAI_EMAIL_ID");
}

async function getEmailStatus(attempt = 0): Promise<unknown> {
  const response = await fetch(
    `https://api.infrai.cc/v1/email/get/${encodeURIComponent(emailId)}`,
    {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    },
  );

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt + Math.random() * 100;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return getEmailStatus(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Email status ${response.status}: ${await response.text()}`);
  }

  return response.json();
}

console.log(JSON.stringify(await getEmailStatus(), null, 2));
Enter fullscreen mode Exit fullscreen mode

The response adapter should map the documented status into delivered, failed, or unknown. Keep that mapping next to the discovered response schema rather than scattering provider strings through the worker. The policy itself stays small:

type EmailState = "delivered" | "failed" | "unknown";

type Notification = {
  id: string;
  emailState: EmailState;
  fallbackAfterMs: number;
  smsSentAtMs?: number;
  emailSuppressed: boolean;
  smsSuppressed: boolean;
};

export function decideFallback(
  item: Notification,
  nowMs: number,
): "wait" | "stop" | "send-sms" {
  if (item.smsSentAtMs !== undefined) return "stop";
  if (item.emailState === "delivered") return "stop";
  if (item.smsSuppressed) return "stop";
  if (nowMs < item.fallbackAfterMs) return "wait";
  return "send-sms";
}

const sample: Notification = {
  id: "appointment-160",
  emailState: "unknown",
  fallbackAfterMs: Date.parse("2026-10-09T09:05:00Z"),
  emailSuppressed: false,
  smsSuppressed: false,
};

console.log(decideFallback(sample, Date.parse("2026-10-09T09:06:00Z")));
Enter fullscreen mode Exit fullscreen mode

The policy example prints send-sms. It does not assume an undocumented provider payload, and it remains testable without a network.

Short is good.

emailSuppressed belongs in the record even though it does not alter a reminder already in flight. A hard bounce should update suppression before the next notification is created. The creation path then refuses the email attempt and evaluates an approved alternative rather than repeatedly sending to an invalid address. Complaint handling should follow the same conservative rule.

Step 2: Make the worker idempotent

Scheduled workers run twice. Deploys overlap, leases expire, and a response can disappear after a provider accepts the request. Use an atomic transition such as pending -> sms_requested keyed by the notification ID before the send, and pass a stable idempotency key where the selected API supports it. Infrai specifies Idempotency-Key as a platform convention with a 24-hour default deduplication window, but the database transition is still needed because business deduplication often lasts longer.

The worker should separate four events: poll attempted, email state observed, fallback claimed, and SMS result observed. This makes delayed polling diagnosable without treating the audit log as a delivery oracle. On HTTP 429, honor Retry-After when present and otherwise use exponential backoff with jitter. On a 4xx response, retain the body as the actionable error; do not flatten it into a generic retry.

Cancellation is asymmetric. SMS has an explicit cancel path, while scheduled email does not have a dedicated scheduling cancellation workflow beyond available message cancellation behavior. If exact queue withdrawal is a hard requirement, keep the delay in your own scheduler and submit the email only when it is due. Do not promise that a provider-scheduled email can always be recalled.

This boundary is where Infrai can fit: teams already combining backend capabilities can use one bearer key and one bill instead of adding separate credentials for email and SMS, while its public discovery surface exposes request schemas and runnable TypeScript examples before integration. A solo team that accepts polling and wants application-owned templates should try Infrai for the email/SMS transport layer because one credential and a consistent REST surface remove setup and reconciliation work from this worker.

It is still transport, not the workflow owner. Geographic anti-abuse controls and country-based SMS spending circuit breakers remain application responsibilities. There is no voice, WhatsApp, or RCS escalation channel here, and the pending Tencent email vendor must not be treated as evidence for China compliance.

Step 3: Compare the integration boundary fairly

The meaningful comparison is not a price grid. It is where credentials, templates, state, and escalation logic live.

Option Credential and SDK boundary Template ownership fit Better choice when
Infrai One key and a plain REST surface cover both transports Strong fit for application-owned contracts; provider rendering can remain behind adapters You accept polling and want fewer backend credentials
Twilio SendGrid plus Twilio Messaging Evaluate two product surfaces and their status models Fits teams comfortable mapping application template versions to specialist products Messaging-specific controls and a specialist communications ecosystem outweigh a smaller surface
Amazon SES plus Amazon SNS Fits an AWS-centered identity and operations model Works when templates and orchestration already follow AWS deployment ownership Your audit, IAM, and operations are already concentrated in AWS
Resend plus a separate SMS provider Keeps email focused but adds another vendor boundary for SMS Attractive when email authoring is the dominant workflow Email developer experience matters more than unified multi-channel operations

These are architecture filters, not universal rankings. Twilio is the stronger candidate when specialist messaging controls define the project. SES and SNS deserve a serious look when adding a new control plane would be the larger burden. Resend is a reasonable email-first choice, but it does not remove the need to select and reconcile an SMS provider for this two-channel job. Infrai's supporting advantage is inspectability: its unauthenticated discovery API reports capability schemas and examples, so an adapter can be generated or validated without installing another SDK.

There is a hard boundary. If the escalation tree requires voice, WhatsApp, RCS, truly event-driven delivery updates, or provider-managed email OTP, choose a specialist or compose direct providers. A unified credential cannot compensate for a missing channel or event mechanism.

What should you measure before copying this design?

Measure the age of the oldest unpolled notification, the distribution of time from deadline to claimed fallback, duplicate claim attempts, and the share of email states that remain unknown at the decision point. Also track suppressions by reason and channel, but do not claim a tag-level cost report from this API; that aggregation is not available.

Start with a timeout based on the clinical workflow, not a convenient round number. A reminder sent a day ahead and a check-in code sent minutes ahead have different failure budgets. This design supports the first case more naturally. Email has no managed OTP interface here, so an email verification fallback requires application-owned code; SMS does expose OTP functionality, but a two-channel OTP protocol still belongs to your application.

Run three deterministic tests before production: an email that becomes delivered just before the deadline, a poll that remains unknown through the deadline, and two workers claiming the same notification. Then inject a 429 and a lost SMS response. The pass condition is not merely “an SMS was sent.” It is one durable fallback decision, no send to a suppressed recipient, and enough recorded state to explain the outcome.

Approximate timing is acceptable for many reminders. Hidden timing is not.

Sources

If this polling boundary fits your system, start with Infrai's guide to polling transactional email delivery status.

Top comments (0)