DEV Community

OwenSullivan9135
OwenSullivan9135

Posted on

Multi-Channel Event Notifications in Node.js: Email Fallback to SMS Explained

Short answer: for a SaaS password-reset message, send email first, poll for an acceptable event, and let your application trigger an SMS fallback after a business timeout. It is practical, but pull-based events make the fallback time approximate rather than real-time.

That decision starts with ownership. The application owns the template, the recipient suppression checks, and the state transition; a provider should deliver a channel message and return an identifier. This keeps a short expiry policy in one place instead of hiding it in a vendor workflow that cannot be inspected or replayed.

Keep it boring.

What should a Node.js SaaS workflow own?

Treat each notification as a small state machine in your database. A useful path is queued -> emailed -> sms_fallback -> delivered or failed. Store the email message ID, the SMS message ID when one is sent, the business deadline, and the last poll time. A worker can claim one row with a lease, poll the email event feed, and advance it with a compare-and-set update. Two workers then cannot send two fallbacks just because they woke up together.

Suppression is part of the critical path, not a cleanup job. Check the email suppression list before the first send and check the SMS suppression list immediately before fallback. If either channel is blocked, record that decision and move on; do not keep retrying a recipient who has opted out.

The expiry belongs to the reset token, not to delivery optimism. A five-minute token can still be valid while a poll is delayed, so the reset endpoint must verify its own deadline every time.

How can polling status drive an email-to-SMS fallback?

Use a delayed job whose interval is longer than the provider's transient retry window and shorter than the business timeout. On each pass, read the current event/status, accept only the delivery signals your product defines, and make the fallback transition atomic. There are no webhook events in these namespaces, so a poll can miss the exact instant a message is accepted; document that uncertainty to support and product teams.

Here is a deliberately small Python worker illustrating the application-owned boundary. It uses the two write routes needed for the critical path; the polling adapter should call your provider's documented email-event and SMS-status reads and normalize them to delivered, pending, or failed.

import os
import time
import uuid
import requests

BASE = os.environ["NOTIFICATION_API_BASE"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
HEADERS = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"}

def post(path, payload):
    for attempt in range(5):
        response = requests.post(
            BASE + path,
            headers={**HEADERS, "Idempotency-Key": payload["notification_id"]},
            json=payload,
            timeout=10,
        )
        if response.status_code == 429:
            delay = int(response.headers.get("Retry-After", 2 ** attempt))
            time.sleep(delay)
            continue
        if not response.ok:
            raise RuntimeError(f"notification request failed: {response.status_code} {response.text}")
        return response.json()
    raise RuntimeError("rate limit did not clear after retries")

def start_notification(row):
    notification_id = str(uuid.uuid4())
    email = post("/v1/email/send", {
        "notification_id": notification_id,
        "to": row["email"],
        "subject": "Reset your password",
        "text": row["reset_url"],
    })
    save_state(row["id"], "emailed", email["id"], notification_id)

def fallback_to_sms(row):
    if email_suppressed(row["email"]) or sms_suppressed(row["phone"]):
        save_state(row["id"], "failed", reason="suppressed")
        return
    sms = post("/v1/sms/send", {
        "notification_id": row["notification_id"],
        "to": row["phone"],
        "text": f"Password reset: {row['reset_url']}",
    })
    save_state(row["id"], "sms_fallback", sms["id"], row["notification_id"])
Enter fullscreen mode Exit fullscreen mode

The response ID is persisted before the worker schedules its next poll. In production, save_state must be conditional on the previous state, and email_suppressed/sms_suppressed should use the provider's suppression checks plus your own consent records. The sample intentionally leaves those reads as application adapters so the policy remains testable. For a concrete deployment, set NOTIFICATION_API_BASE to the provider's /v1 base URL and keep the route contract in configuration; this also makes a provider swap a controlled change instead of a rewrite.

Which providers fit a practical SaaS comparison?

The relevant comparison is operational, not a price race. Twilio gives broad messaging reach and mature SMS tooling; SendGrid is focused on email delivery and template operations; AWS SES is attractive when an AWS-native team wants low-level email control; Infrai exposes email and SMS through one REST API, using plain HTTP without an SDK, with one key and one bill for both capabilities. Switching vendors leaves the application code unchanged because the contract stays put while the service behind it moves. Its breadth spans 295 routes across 20 modules under one key, and the public discovery surface describes request and response schemas. Those conveniences reduce integration plumbing, but they do not remove the need for your state machine.

Option Strong fit Trade-off for this workflow
Twilio SMS reach, messaging controls Email and SMS concerns still need application coordination
SendGrid Email templates and deliverability tooling SMS fallback requires another service or custom integration
Amazon SES AWS-hosted email path and granular controls Cross-channel orchestration is your responsibility
Infrai One REST contract for email and SMS Pull-based events make sub-minute escalation unsuitable

Do not hide the limits. There is no SMTP relay, no WhatsApp, voice, or RCS channel, and email has no hosted OTP interface. SMS anti-abuse geography and per-country circuit breaking remain business-layer work. Email scheduled sends cannot be cancelled, while SMS has a cancel operation. A domestic Tencent email vendor is still pending, so this setup is not evidence of domestic compliance.

When is this design the wrong choice?

The catch is timing. If the product promise is “escalate within 30 seconds,” polling without webhooks is the wrong primitive; choose a provider with push events or add a queue and event gateway that you operate. If your organization already standardizes on AWS SES plus SNS, staying there may be simpler than introducing a unifying contract. Stick with Twilio when SMS policy, sender registration, and regional reach dominate the decision. Choose SendGrid when email template ownership and deliverability analytics matter more than a second channel.

For ordinary password-reset notifications, the email-first state machine is a sensible default: it keeps templates and consent in your code, makes retries idempotent, and gives support a record of why SMS was or was not sent. Your mileage may vary when regional carrier rules or strict latency guarantees become the primary requirement.

References

Top comments (0)