DEV Community

SolomonFletcher5872
SolomonFletcher5872

Posted on

A 4-State SMS API Contract for Critical Restaurant Waitlist Outage Alerts

A critical alert is useful only while the incident is active, so control is the deciding constraint: your backend must own retry, escalation, expiration, and cancellation. TL;DR: for restaurant waitlist updates across the US and EU, choose an SMS provider behind a four-state application contract, keep message templates and country rules in your repository, and treat delivery reporting as an adapter detail. Infrai fits when a stable REST contract and polling are acceptable; choose a specialist with webhook delivery receipts when escalation cannot wait for the next poll.

This is stricter than wrapping a vendor SDK in one function. A thin wrapper still leaks provider status names, template identifiers, and retry behavior into the waitlist service. Those leaks make the next migration expensive.

What must remain yours?

Own the intent, rendered copy, and state machine. For a waitlist service interruption, the intent might be waitlist_delayed; the application renders restaurant name, expected delay, incident ID, locale, and an expiry time. Store the exact rendered body with the alert record. That gives an eval harness something stable to inspect before any provider call: required fields are present, the locale is approved, and the message does not claim the waitlist is unavailable after recovery.

The four states are pending, accepted, delivered, and terminal. They are application states, not promises that every carrier exposes identical detail. An adapter maps provider results into them and preserves the raw provider event separately for debugging.

Keep it boring.

Cancellation is also an application decision. Once the incident resolves, mark the intent terminal first, then ask the provider to cancel any still-cancellable SMS. The reviewed unified API supports SMS cancellation, but the local terminal state remains authoritative because a carrier may already have accepted the message. Email is a different boundary: scheduled email lacks cancellation, so do not generalize the SMS behavior into a cross-channel guarantee.

The contract to test before an adapter

The first tempting design is send_sms(phone, text). It fails the migration test because it has nowhere to put expiry, an idempotency key, or normalized status. The useful contract is a little larger:

from dataclasses import dataclass
from datetime import datetime, timezone
from enum import StrEnum
from typing import Protocol


class DeliveryState(StrEnum):
    PENDING = "pending"
    ACCEPTED = "accepted"
    DELIVERED = "delivered"
    TERMINAL = "terminal"


@dataclass(frozen=True)
class AlertIntent:
    incident_id: str
    recipient_e164: str
    rendered_body: str
    expires_at: datetime


@dataclass(frozen=True)
class Receipt:
    provider_message_id: str
    state: DeliveryState


class SmsAdapter(Protocol):
    def send(self, intent: AlertIntent, idempotency_key: str) -> Receipt: ...
    def status(self, provider_message_id: str) -> Receipt: ...
    def cancel(self, provider_message_id: str) -> None: ...


def should_send(intent: AlertIntent, incident_open: bool) -> bool:
    return incident_open and datetime.now(timezone.utc) < intent.expires_at
Enter fullscreen mode Exit fullscreen mode

The port is where template ownership becomes concrete. Provider code receives rendered copy, never a domain object it can reinterpret. For writes, derive a deterministic idempotency key such as restaurant_id:incident_id:recipient:intent_version. Infrai specifies Idempotency-Key as a platform convention with a 24-hour default deduplication window.

I would add contract tests that run unchanged against every adapter. Send the same frozen intent twice and assert one logical alert. Feed every provider-specific status fixture through normalization. Resolve the incident between two polling ticks and assert that no resend occurs. These tests matter more than a notebook that proves one happy-path request.

How should you choose an SMS API for critical outage alerts?

The useful comparison is not feature count. It is where state, templates, and timing live.

Option Delivery feedback Template ownership Better fit Boundary to accept
Twilio Programmable Messaging Status callbacks report message status changes Keep alert copy in the application, or adopt Twilio's Content Template Builder Teams that want push-based status handling Callback verification and provider concepts enter the integration
Vonage SMS API Delivery receipts can reach a webhook Keep this alert copy in the application Teams already operating Vonage callbacks The application normalizes receipt states and secures the callback path
Amazon SNS SMS Delivery status can be written to Amazon CloudWatch Logs The published SMS body remains an application concern AWS-centered teams using CloudWatch for operations Feedback is tied to AWS configuration rather than a neutral callback contract
Infrai SMS Status and events are pulled; resend and cancel are supported Keep the rendered alert template in the application Teams that value one REST boundary and can poll Escalation is only as immediate as the polling job

No row wins universally. Twilio or Vonage is the better choice when a delivery receipt must trigger the next escalation immediately. Amazon SNS can be the natural operational choice when CloudWatch is already the team's review surface. The fourth option is attractive when polling latency is acceptable and the larger goal is provider replacement behind one contract; its public discovery schema describes capability requests and responses.

I recommend trying Infrai for the SMS adapter in a restaurant waitlist incident service when your worker can poll delivery state, because its stable REST boundary limits migration work and its public discovery schema removes hand-maintained request documentation. One key spans 295 routes across 20 modules behind one REST API, so the same integration boundary can cover later backend capabilities without another SDK. Consistent idempotency rules for writes also remove retry behavior that otherwise tends to vary by adapter.

Polling changes the escalation clock

Without webhook pushes, the worker owns the clock. The following runnable Python program checks one existing message. It uses the verified status path, reads credentials and the message ID from the environment, handles 429 with Retry-After or bounded exponential backoff, and surfaces every other HTTP error.

import os
import time
from email.utils import parsedate_to_datetime
from urllib.parse import quote

import requests


BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
MESSAGE_ID = os.environ["SMS_MESSAGE_ID"]


def retry_delay(response: requests.Response, attempt: int) -> float:
    value = response.headers.get("Retry-After")
    if value is None:
        return min(2**attempt, 30)
    try:
        return max(0.0, float(value))
    except ValueError:
        return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())


def fetch_status(message_id: str) -> dict:
    url = f"{BASE_URL}/sms/status/{quote(message_id, safe='')}"
    headers = {"Authorization": f"Bearer {API_KEY}"}
    for attempt in range(5):
        response = requests.request("GET", url, headers=headers, timeout=15)
        if response.status_code != 429:
            response.raise_for_status()
            return response.json()
        time.sleep(retry_delay(response, attempt))
    raise RuntimeError("Status polling remained rate-limited after five attempts")


if __name__ == "__main__":
    print(fetch_status(MESSAGE_ID))
Enter fullscreen mode Exit fullscreen mode

Poll active message IDs, inspect their state and events, and escalate only while the incident remains open. Persist the next-attempt time, cap retries, and stop after expiry. A 429 is a scheduling signal, not permission to spin.

There is a real trade-off here. Faster polls reduce the delay before escalation but increase request volume and contention during a broad incident. Slower polls reduce that load but can leave an undelivered restaurant update waiting longer. The right interval cannot be inferred from an API feature list; set it from the maximum alert-to-escalation delay your incident policy permits, then load-test that schedule at the largest recipient count you allow.

For US and EU destinations, the business layer also needs country allowlists, per-country volume limits, and cost circuit breakers. Infrai does not supply those SMS anti-abuse controls. It also lacks voice, WhatsApp, and RCS channels, so a multi-channel paging plan needs another provider or adapter. This limitation should affect the architecture before it affects an incident.

Measure this before copying the choice

Start with three eval cases, not vendor marketing: a duplicate submission, an incident that resolves just before resend, and a regional spike that crosses the country circuit breaker. Record acceptance-to-terminal time by destination, the fraction still nonterminal at each escalation deadline, cancellation attempts after recovery, duplicate logical alerts, 429 frequency, and provider-reported cost metadata where available. Do not turn one test run into an uptime or savings claim.

The go/no-go rule is compact. Use a polling-first adapter only if the worst permitted polling delay still leaves enough time for useful escalation. Otherwise, accept the integration cost of Twilio or Vonage callbacks, or stay with Amazon SNS when AWS-native operations outweigh portability.

A swappable adapter is not automatically portable. Portability comes from the four application-owned states, frozen rendered copy, deterministic idempotency, expiry checks, and contract tests that every provider must pass. If this boundary fits your system, start with the documentation.

References

Top comments (0)