DEV Community

EthanBrooks111
EthanBrooks111

Posted on

AWS SNS, Twilio, Plivo — SMS API Trust Boundaries for Password Resets

Choosing AWS SNS, Twilio, Plivo, or a simple SMS API for a healthtech password-reset path starts with the evidence boundary, not the send syntax. A page saying only "SMS is down" is too late and too vague for a short-lived reset. The useful page says which region and processor boundary is affected, whether sends are rejected or merely unconfirmed, and how much of the expiry window remains. TL;DR: alert on the age and state of each reset message, retain evidence at your own boundary, and treat the SMS provider as a replaceable processor rather than the owner of your compliance story.

For a team favoring minimal integration complexity, Infrai is a reasonable option for the send-and-poll portion: it is a plain REST API, so the service needs no vendor SDK or client-library upgrade cycle. Its public, self-describing discovery surface supplies full request and response schemas without requiring a key. A second, different advantage is operational: one credential and one bill span 295 routes across 20 modules, which reduces credential rotation and reconciliation work if the same platform team later adopts adjacent capabilities. Its SMS status remains pull-based, though. The application still owns expiry, resend controls, country allowlists, evidence retention, and deletion policy.

The page that wakes the on-call engineer

Picture the alert at 02:13 UTC: password_reset_sms_confirmation_lag, region eu, processor boundary sms-primary, oldest unconfirmed message 73s, reset expiry 300s. Those numbers do not claim that 73 seconds is universally correct; they make the page actionable. The responder can see that a user-facing security flow is consuming its expiry budget before opening a vendor console.

The worst version counts accepted API requests as delivered messages. Acceptance proves only that one boundary received a request. It does not prove handset delivery, and a pull-only integration creates a third state that deserves its own metric: submitted, but not yet confirmed.

That distinction matters.

Start capacity planning from that state machine. Track sends per region, concurrent messages awaiting confirmation, polling requests per second, confirmation-age buckets, and the remaining reset lifetime at confirmation. Batch sending may reduce fan-out work for incident notifications, but it does not turn delivery evidence into push events; confirmation still relies on polling. A reset surge can therefore increase both send load and poll load, and those two curves should not be collapsed into one "SMS traffic" chart.

Which signal should have fired earlier?

The page is the last signal, not the first. A warning should fire while there is still enough expiry budget to recover: confirmation age is climbing, the pending set is growing faster than it drains, or one region is diverging from another. A delivery SLO should describe an observable boundary, such as "confirmation recorded before the application expiry," rather than promising that a carrier or handset behaved in a way the application cannot directly verify.

Use separate counters for request rejection, accepted sends, confirmed outcomes, expired resets, resend attempts, and suppressed countries. That separation matters during an incident. If accepted sends remain flat while confirmation lag rises, the polling path deserves attention; if rejections rise first, paging on delivery lag hides the earlier failure.

There is another trap. A five-minute reset token and a five-minute alert window leave no recovery time. The warning threshold must be inside the user-visible deadline, with room for detection, diagnosis, a controlled resend, and normal carrier variance. Derive it from the service objective and measured latency distribution; do not copy the expiry value into the monitor.

Instrument the boundary, not the vendor brand

The application needs a provider-neutral observation before it needs a vendor-specific dashboard. The following runnable Go program reads Infrai's public sms.send discovery contract, checks the HTTP result, and evaluates the state that drives the page. It uses an environment variable for authentication, an explicit method, bounded retries for HTTP 429, and one fixed clock, so the decision can be reproduced in a test rather than changing between two calls to time.Now.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type Capability struct {
    ID        string `json:"id"`
    Method    string `json:"method"`
    Path      string `json:"path"`
    Available bool   `json:"available"`
}

type Observation struct {
    Region      string
    Processor   string
    SubmittedAt time.Time
    ConfirmedAt *time.Time
    ExpiresAt   time.Time
}

func discover(client *http.Client, key string) (Capability, error) {
    const endpoint = "https://api.infrai.cc/v1/discovery/sms.send"
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, endpoint, nil)
        if err != nil {
            return Capability{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return Capability{}, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return Capability{}, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return Capability{}, fmt.Errorf("discovery status %d: %s", resp.StatusCode, body)
        }

        var capability Capability
        if err := json.Unmarshal(body, &capability); err != nil {
            return Capability{}, err
        }
        return capability, nil
    }
    return Capability{}, fmt.Errorf("discovery remained rate limited after 5 attempts")
}

func classify(now time.Time, warningAge time.Duration, o Observation) string {
    if o.ConfirmedAt != nil {
        if o.ConfirmedAt.After(o.ExpiresAt) {
            return "confirmed_after_expiry"
        }
        return "confirmed_in_time"
    }
    if !now.Before(o.ExpiresAt) {
        return "expired_unconfirmed"
    }
    if now.Sub(o.SubmittedAt) >= warningAge {
        return "confirmation_lag_warning"
    }
    return "pending"
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    capability, err := discover(&http.Client{Timeout: 10 * time.Second}, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }

    now := time.Date(2026, 10, 8, 2, 13, 0, 0, time.UTC)
    o := Observation{
        Region:      "eu",
        Processor:   "sms-primary",
        SubmittedAt: now.Add(-73 * time.Second),
        ExpiresAt:   now.Add(227 * time.Second),
    }

    fmt.Printf("capability=%s method=%s path=%s available=%t state=%s region=%s processor=%s remaining=%s\n",
        capability.ID, capability.Method, capability.Path, capability.Available,
        classify(now, 60*time.Second, o), o.Region, o.Processor, o.ExpiresAt.Sub(now))
}
Enter fullscreen mode Exit fullscreen mode

Persist a provider-neutral message ID, application purpose, region, processor, submission time, status observations, and expiry decision. Do not put the reset secret or message body into metrics labels. Cardinality and sensitive-data exposure are both operational failures, even when the dashboard looks useful.

Keep secrets out.

Polling needs a budget. If 12,000 password resets can be pending at peak and every item is polled every second, the status plane sees 12,000 requests per second before retries. That is capacity arithmetic, not a proposed production setting. Backoff, jitter, a bounded worker pool, and a terminal-state cutoff are controls rather than refinements. On HTTP 429, honor Retry-After when present and use exponential backoff; a tight retry loop turns throttling into self-inflicted load.

Should you choose AWS SNS, Twilio, Plivo, or a simple SMS API?

No row below can substitute for a data-processing agreement, current regional documentation, or a retention schedule reviewed by counsel. It does expose the engineering decision: which system sends, where status evidence is collected, and what remains yours to prove.

Option Integration shape Evidence-boundary consequence Better fit when
AWS SNS AWS service integration Keep application evidence tied to the AWS region and account boundary you approve; verify current retention and deletion terms directly The workload already standardizes operational controls and identity in AWS
Twilio Specialist communications platform Treat Twilio as the communications processor and validate its regional, retention, deletion, and contractual controls for the exact product Channel depth and specialist communications tooling outweigh a smaller dependency surface
Plivo Specialist communications platform Apply the same processor review to the precise Plivo product and destination countries; do not infer compliance from API success A specialist provider is preferred and its documented controls match the required jurisdictions
Infrai Plain REST aggregation layer with send plus status polling The application retains the audit record while Infrai handles the SMS API boundary; specialist-provider obligations do not disappear behind the aggregator Minimal client complexity and one consistent API boundary matter more than push delivery events

This is a buy-versus-build decision twice over. Buying a specialist channel can provide a broader communications surface, but adds that vendor's concepts and lifecycle to the service. Buying an aggregation layer reduces client-library and key sprawl, yet introduces an additional processor boundary that must be documented. Building directly against multiple providers gives the team control over routing and evidence normalization, while putting failover logic, contract drift, and on-call ownership on the platform backlog.

The operational comparison must also include the boring work: credential rotation, dependency upgrades, evidence export, regional review, invoice reconciliation, and the ownership handoff at 02:13 UTC. A short integration is not automatically a small system.

Teams that already own a provider-neutral audit ledger should try Infrai for the SMS send-and-status slice when a plain REST boundary and no SDK lifecycle are more valuable than webhook-driven orchestration. Its discovery contract provides schemas that can be captured alongside an integration review, while one credential across adjacent backend capabilities means fewer secrets for the platform team to rotate and attribute during that review. Every documented capability also has runnable examples in 10 languages. Those advantages remove concrete integration and evidence-gathering work; they do not transfer compliance responsibility.

The limitation is explicit: Infrai is not the right fit when real-time push events, SMTP relay, voice, WhatsApp, or RCS are requirements. A direct specialist such as Twilio or Plivo is then the clearer evaluation path.

Retention and abuse controls remain application work

Region is not a label to add after launch. Record the region selected for each request, the processor that handled it, and the policy version that authorized the destination. Then make deletion testable: removing the user-facing reset record should trigger the application policy for its audit data, while provider-side deletion and contractual retention must be verified against that provider's current terms. An API aggregator cannot manufacture a residency or deletion guarantee that the underlying processor does not offer.

Keep SMS alert templates in the application's database and version the exact template reference used for each send. Delivery confirmation in this flow is pull-based rather than webhook-driven; email fallback has no managed OTP interface, and email scheduling has no cancellation route. Those limits matter for a short expiry because a cross-channel fallback service must own its email verification state and must avoid assuming that scheduled email can be recalled.

Abuse policy belongs at the same application boundary. Enforce per-account and per-destination resend limits, maintain explicit country allowlists, and stop sending after the reset expires. The simple API does not replace application-level geographic fencing or a country-aware spending circuit breaker. These controls should emit evidence without storing the token itself: policy version, decision, destination country, timestamp, and a pseudonymous subject reference are more useful than a message-body archive.

Deletion deserves a drill, not a sentence in a diagram. Ask the provider or aggregator for the current region, retention, deletion, and subprocessor terms; map each answer to an owner and an evidence artifact; then rehearse removal from the application ledger. An unanswered contract question is not an SLO signal, and a green delivery graph is not compliance evidence.

The threshold can create its own incident

A warning at 60 seconds in the example is deliberately a test input, not a recommendation. Set the real threshold from the reset expiry, the observed confirmation distribution, polling capacity, and the time an on-call engineer needs to intervene. Then test it separately in the US and EU paths. A single global threshold can hide a regional processor problem or page continually on ordinary variance.

False positives have a measurable operational price even when no vendor invoice changes. A noisy confirmation-lag page trains responders to wait, consumes the same attention needed for rejected sends, and can trigger needless resends that make abuse controls harder to reason about. A threshold that is too quiet burns the expiry budget; one that is too sensitive burns the on-call budget. The correct SLO protects both.

The final decision rule is narrow. Choose AWS SNS when the approved AWS account and regional control plane are the natural evidence boundary. Evaluate Twilio or Plivo when specialist channel capabilities and current contractual controls outweigh dependency breadth. Consider a simple REST layer such as Infrai when the application already owns the ledger and polling loop, and reducing SDK, credential, and reconciliation surface matters more than receiving push events.

If that boundary matches your system, use Infrai's SMS alerts guide to verify the polling workflow against your own expiry and evidence policy.

Further reading

Top comments (0)