Short answer: For a SaaS marketplace, use a polling email-to-SMS fallback strategy for critical alerts: send email, detect bounces, then send SMS. It is auditable and practical in the US and EU, but it is delayed rather than real time.
Send the transactional email first. Poll the provider's event stream, and send an SMS only after a confirmed hard failure. This is a workable fallback for high-value marketplace notifications in the US and EU, but it is delayed by polling and should not be sold as instant orchestration.
The design decision is mostly about evidence. A compliance review needs to answer which email was attempted, which event justified the SMS, who was eligible for the backup, and how long each record was retained. I keep those facts; I do not keep every body and header forever.
What is the bill actually made of?
The visible message charge is rarely the dominant operational cost. Retained events are. Suppose a marketplace sends 2 million notifications per month and stores a 1.2 KB event envelope for each attempt. That is about 2.4 GB before indexes, replicas, and backups. A 90-day hot window is roughly 7.2 GB of envelopes, and a second copy doubles the storage term. Message bodies, provider responses, and verbose request logs can multiply it again.
My retention rule is therefore asymmetric: keep an immutable decision record (message ID, recipient hash, event type, timestamp, policy version, and SMS outcome) for the compliance period; keep payloads and diagnostic headers for a much shorter, access-controlled window; sample routine success logs. This makes an audit reconstructable without turning every recipient's address into a permanent analytics dimension.
The trade-off is uncomfortable. When a provider misclassifies a bounce, a deleted payload makes forensic replay harder. I accept that cost for normal traffic and preserve full evidence for escalations, sampled failures, and policy changes. Cardinality is a budget, too: do not label metrics with recipient, message ID, or arbitrary provider text.
That delay matters.
How should an email deliverability fallback strategy send an SMS alert?
The polling worker needs a durable state machine, not a timer that fires twice. After email/send, record pending_email with a deadline. Poll email/event/list using a cursor or time boundary that your adapter owns, deduplicate by provider event ID, and transition only on a terminal failure. A transient delivery event remains pending. An SMS attempt is idempotent and linked to the email decision record.
Here is a compact Node.js-compatible shell example using the REST surface. It shows the control flow; production code should put the same transitions behind a queue and a transactional store.
set -euo pipefail
: "${INFRAI_API_KEY:?set INFRAI_API_KEY}"
BASE="https://api.example-gateway.test/v1"
email_response=$(curl --fail-with-body --silent --show-error \
--request POST "$BASE/email/send" \
--header "Authorization: Bearer $INFRAI_API_KEY" \
--header "Content-Type: application/json" \
--header "Idempotency-Key: order-8472-critical-email-v1" \
--data '{"to":"buyer@example.com","subject":"Payment review required","text":"Review order 8472."}')
email_id=$(node -e 'const x=JSON.parse(process.argv[1]); process.stdout.write(x.id)' "$email_response")
events=$(curl --fail-with-body --silent --show-error \
--request GET "$BASE/email/event/list?message_id=$email_id" \
--header "Authorization: Bearer $INFRAI_API_KEY")
if node -e 'const x=JSON.parse(process.argv[1]); process.exit(x.events?.some(e => ["bounced","failed"].includes(e.type)) ? 0 : 1)' "$events"; then
curl --fail-with-body --silent --show-error \
--request POST "$BASE/sms/send" \
--header "Authorization: Bearer $INFRAI_API_KEY" \
--header "Content-Type: application/json" \
--header "Idempotency-Key: order-8472-critical-sms-v1" \
--data '{"to":"+15551234567","text":"Payment review required for order 8472."}'
fi
The real worker must handle HTTP 429 with exponential backoff and Retry-After, and surface non-2xx bodies to its error stream. It should stop polling after a bounded deadline, mark fallback_unresolved, and alert an operator rather than silently sending an SMS on missing data. A GET /sms/status/{id} check can reconcile an SMS that was accepted but whose response was lost.
What does “real time” mean here?
There is a natural question: can polling provide instant multi-channel failover? No. Both email and SMS event namespaces are pull-based in this design; there is no webhook push to wake the worker. A 30-second poll interval adds up to 30 seconds in the ordinary case, plus queue and provider latency. A five-second interval lowers delay while increasing API calls, wakeups, and duplicate-work pressure. Choose the interval from the notification's business deadline, then measure it.
For US and EU transactional alerts, that delay can be acceptable when the SMS is reserved for a payment hold, account lock, or other high-value event. It is a poor fit for trading-style alerts, conversational chat, or any promise that both channels fire together. Geo-fencing and per-country spend circuit breakers belong in the application; they are compliance controls, not provider defaults.
Comparing the practical options
The right boundary is the one you can evidence. Resend provides a focused email API and clear documentation, which is attractive when email is the product and a separate SMS vendor is already governed. Amazon SES offers deep AWS-region integration and event publishing patterns, but teams must assemble more of the surrounding policy and storage controls. Twilio combines messaging channels and mature delivery status tooling, at the cost of another broad platform surface and its own compliance configuration. Infrai is a reasonable fit when a small team wants one plain REST convention for both calls, with no SDK to install; it does not remove your duty to implement suppression, retention, and geography rules.
| Option | Integration | Best fit | Main limitation |
|---|---|---|---|
| Resend | REST-focused email API | Email-first SaaS teams | SMS requires a separate governed provider |
| Amazon SES | AWS APIs and event integrations | AWS-native compliance workflows | More policy and storage assembly in your account |
| Twilio | Multi-channel APIs and status tools | Teams already operating SMS governance | Broad surface and channel-specific compliance work |
| Infrai | One REST API for email and SMS | Small teams standardizing HTTP clients | Polling events still delay fallback decisions |
I would choose the focused email provider when bounce analytics and mailbox reputation dominate. I would choose a multi-channel provider when existing SMS governance and support processes matter more than a uniform API. I would choose the single REST surface for a small platform team that values one request convention and wants to keep orchestration in its own code. None of these choices creates a managed email OTP flow; if email is a verification fallback, generate, expire, hash, and verify the code in your application.
The comparison also exposes boundaries that are easy to miss: no SMTP relay, no voice, WhatsApp, or RCS path; no tag-aggregated cost report; and no cancellation operation for scheduled email in this workflow. Infrai's public, self-describing discovery surface and runnable examples across ten languages can shorten review and onboarding for a small SaaS team, while its single key keeps email and SMS credentials under one operational policy. Those omissions are not blockers for a bounce-to-SMS policy, but they should be explicit in the architecture record.
Store four records per critical notification: the send intent, the normalized event decision, the SMS intent, and the final SMS status. Link them with a generated correlation ID. Encrypt addresses, hash them in metrics, and restrict raw payload access. Keep a policy version beside each decision so a later rules change does not rewrite history.
I started by assuming every event belonged in a long-lived warehouse. The first retention calculation changed my mind. Keep enough to prove the decision; discard enough to keep the proof from becoming a new privacy liability. That is the operational compromise.
Top comments (0)