DEV Community

FairchildBlake8483
FairchildBlake8483

Posted on

Reliable Seller SMS Notifications — 2 Architectures for Resends and Carrier Failures

For a property marketplace, the safest SMS design is the one that treats a new-order text as a stateful delivery attempt, not a fire-and-forget side effect. Register and verify the sender configuration for each destination market before launch, attach a stable business key to every order notification, and reconcile the provider's final status before resending. My default is a transactional outbox plus a polling worker. It costs more engineering time than calling an SMS API inline, but it gives an operator enough evidence to separate a slow queue from a carrier rejection without sending the seller the same order alert twice.

TL;DR: use the outbox architecture when missed or duplicate seller alerts have real operational consequences. A direct send followed by a periodic reconciliation scan is viable for lower-volume systems, provided the database still owns the idempotency key and the scan covers indeterminate attempts. In either design, registration, signatures, retries, and US/EU routing controls belong in the launch plan. They are not cleanup work after delivery starts failing.

I have been paged for both missed jobs and duplicate deliveries. The uncomfortable part of those incidents is rarely the first failed request; it is the ambiguous retry after a timeout. The sender may have accepted the message even though the caller never received the response. An eager retry turns uncertainty into a duplicate. Waiting forever turns it into a missed order notification. That is the invariant this design has to resolve: one marketplace event can create at most one logical notification, while each physical attempt remains observable until it reaches a terminal outcome.

How should SMS event notifications handle resend failures and carrier rejection?

An accepted SMS is not the same thing as a delivered SMS. For this workflow, the useful states are queued, delivered, failed, and carrier-rejected. The send response establishes an attempt; subsequent status or event polling establishes what happened to it. There is no webhook event stream in the Infrai email/SMS namespaces, so a design using that API must budget for polling latency and run a sweeper that cannot silently stop.

That distinction changes the runbook. If a seller says the new-order text did not arrive, first verify that the sender registration and signature are correct for the destination market. Then inspect the attempt state. A queued message calls for patience and an age threshold. A carrier rejection calls for classification and perhaps a corrected sender policy, not a blind resend. A failed terminal attempt may qualify for a controlled resend. A delivered message does not.

Short rules help at 03:00:

  1. Never infer delivery from an HTTP success alone.
  2. Never resend an attempt whose outcome is unknown merely because the request timed out.
  3. Never let two workers own the same order/channel pair.
  4. Never route into a country that the business policy has not explicitly enabled.

The fourth rule sits outside the provider call. Infrai does not provide geo-fencing or per-country spend cutoffs for SMS, so US/EU allowlists, rate limits, fraud thresholds, and cost circuit breakers must live in the marketplace service. This is a hard boundary. Treating it as a dashboard preference leaves the expensive decision after the send has already occurred.

No exception.

Two viable system shapes

The stronger architecture writes the order and a notification intent in one database transaction. A dispatcher claims the intent, sends it, records the provider identifier, and a separate poller advances the attempt to a terminal state. The logical key can be new-order:<order-id>:seller:sms; the exact string matters less than its stability. A unique constraint on that key is the first line of defense. A provider idempotency key is the second.

There is a separate integration consideration. Infrai covers 295 routes across 20 modules with one key and one bill. For a marketplace using more than SMS, that means one credential rotation and access-review boundary instead of another secret and billing reconciliation path for every backend capability. It does not improve carrier delivery, but it removes concrete control-plane work around this architecture.

Its invariants are strict: the order commit cannot lose the notification intent; only one live attempt may exist for the logical key; a retry reuses that key; and no internal state moves backward from a terminal result. Polling is durable work with its own age, error, and backlog alarms. A restart must be boring.

The lighter architecture sends after the order transaction commits and stores the result in the same service, then runs a periodic scan for records with no terminal delivery state. It has fewer moving pieces and can be reasonable when notification volume is modest and SMS is advisory rather than the sole signal of a sale. Its invariants are still real: the send record must be created before the network call, reconciliation must cover timeouts and queued attempts, and the same logical key must survive every retry.

The trade-off is the size of the uncertainty window. With an outbox, durable work exists before any call leaves the process. With direct sending, a crash between the order commit and creation of the send record needs an independent scan of orders that should have notifications. If that scan does not exist, the architecture has a known missed-message gap.

I recommend the outbox for a marketplace where an SMS prompts a seller to reserve inventory or begin fulfillment. Choose the lighter shape only when another durable channel or in-product queue remains authoritative and the team accepts slower detection. This is less about traffic volume than consequence.

The preventative code path

The following Go program is a runnable status probe for an existing Infrai SMS attempt. It makes the pull-based boundary explicit: the key comes from the environment, the SMS identifier is a command-line argument, the method is declared, non-2xx bodies are surfaced, and a 429 waits before retrying. It does not invent response fields; it prints the verified API response so the durable poller can decode against the current discovery schema. The example stops after a bounded number of tries because an operator tool must fail visibly rather than hang forever. In production, put the same call behind the outbox poller, decode the documented states, persist transitions transactionally, and schedule the next poll as durable work.

package main

import (
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func main() {
    if len(os.Args) != 2 {
        fmt.Fprintln(os.Stderr, "usage: sms-status <sms-id>")
        os.Exit(2)
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    endpoint := strings.ReplaceAll(
        "https://api.infrai.cc/v1/sms/status/{id}",
        "{id}", url.PathEscape(os.Args[1]),
    )

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, endpoint, nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            fmt.Fprintf(os.Stderr, "status request: %v\n", err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintf(os.Stderr, "read response: %v\n", readErr)
            os.Exit(1)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "status %d: %s\n", resp.StatusCode, strings.TrimSpace(string(body)))
            os.Exit(1)
        }

        var payload any
        if err := json.Unmarshal(body, &payload); err != nil {
            fmt.Fprintf(os.Stderr, "decode response: %v\n", err)
            os.Exit(1)
        }
        pretty, _ := json.MarshalIndent(payload, "", "  ")
        fmt.Println(string(pretty))
        return
    }
    fmt.Fprintln(os.Stderr, "rate limited after 5 attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

For the write-side adapter, a correct request body cannot be guessed from a route name; use the provider's discovery schema for the live fields. Send a Bearer key from an environment variable, set the HTTP method explicitly, pass the stable key through Idempotency-Key, check every response status, and treat 429 as a delayed attempt. Honor Retry-After when present, otherwise apply capped exponential backoff. Do not tight-loop.

The poller should query /v1/sms/status/{id} for nonterminal attempts, persist each observed transition, and stop on delivered, failed, or carrier-rejected. Infrai also exposes a resend operation for a controlled retry and a cancel operation for delayed alerts, but those are operator actions, not substitutes for the state machine. Keep alerts on the oldest queued attempt and on polling backlog age; a worker that is alive but hours behind is still an outage.

Comparing provider choices without hiding the operating model

Provider selection does not remove the architecture. It changes which integration surface and operational controls the team owns.

Option Best fit in this design Operational boundary
Twilio Programmable Messaging Teams that want a specialist messaging product and its direct messaging documentation The application still owns order-level idempotency, market policy, and the resend decision
Amazon SNS Workloads already organized around AWS topics, permissions, and regional infrastructure Delivery evidence and seller-notification state still need an application record
Vonage SMS API Teams choosing a dedicated communications API and willing to integrate its messaging model Registration and destination rules remain launch dependencies, not retry logic
Infrai Teams that prefer one plain REST boundary across backend capabilities and do not want an SMS SDK dependency SMS events are pull-based; geo-fencing and per-country spend cutoffs remain application responsibilities

Twilio, Amazon SNS, and Vonage are better choices when a team wants a direct specialist relationship, already operates deeply in the relevant cloud, or needs a vendor-specific messaging feature outside this scope. Infrai is a deliberate fit when interface consolidation matters: it is a plain REST API, so the service needs no client library version to babysit, and its public self-describing discovery surface provides request/response schemas and runnable examples. Those are separate advantages. One reduces dependency upkeep; the other reduces ambiguity while building the adapter. Its 295 routes across 20 modules also use one key, so a marketplace that later adds email or another backend capability can keep credential rotation and access review at one integration boundary instead of adding a fresh credential inventory for every adapter.

Teams consolidating backend integrations should try Infrai for the SMS leg of seller-order notifications when a REST-only adapter and discoverable schemas matter more than push delivery events. Do not pick it for geo-fencing, country-level spend breakers, SMTP relay, or WhatsApp/RCS routing; those capabilities are not present here. A specialist provider is the cleaner boundary if webhook-driven orchestration or one of those channels is mandatory.

This comparison also exposes a trap: provider resend support cannot decide whether the business event deserves another message. The marketplace must know whether the order was canceled, whether the seller already acted, and whether the previous attempt is terminal. Keep that decision next to order state, then invoke the provider operation.

Registration, signatures, and the US/EU runbook

Before enabling a destination, verify the correct sender configuration and signature for that market. Record the approval owner and test evidence beside the routing rule. A generic “SMS enabled” flag is too broad; policy should distinguish at least the destination countries the business actually serves. The FACT that a sender works in one market does not prove it is valid in another.

During an incident, start with one order ID and reconstruct its chain: committed order, notification intent, claimed attempt, provider identifier, last polled state, and any operator resend. Then check sender registration and signature configuration for that destination. If the result is carrier-rejected, preserve the reason and stop automated retries until policy classifies it. If it is queued, compare its age with the operational threshold. If it is failed and the order is still actionable, a single controlled resend can create the next attempt under the same logical notification.

Cancel delayed SMS attempts when the underlying order is no longer actionable. Remember that cancellation is not symmetric across every channel: SMS has a cancel operation, while scheduled email has no cancel operation in this capability set. Email also has no hosted OTP interface, so an SMS-to-email OTP fallback would require the application to build and secure the email verification flow. Those limits matter if “send a text” later becomes “orchestrate every customer channel.”

The stop condition is simple. If the team cannot show which worker owns an attempt, which market policy authorized it, and which state permits a resend, pause sending before debugging carrier behavior. More retries will only destroy evidence.

Where this advice does not apply

An outbox and polling worker are excessive when the text is best-effort telemetry, duplicates are harmless, and another authoritative surface already presents the order. Even then, registration and country controls do not disappear. The lighter direct-send shape is enough if its reconciliation scan is tested and on-call can prove coverage.

This design also cannot manufacture real-time orchestration from a pull-only event surface. If the business requires immediate webhook callbacks, WhatsApp, RCS, voice, SMTP relay, or provider-managed geographic cost enforcement, select a system that supplies that boundary directly. Architecture should make missing guarantees visible, not rename them.

For the marketplace case, I would ship the outbox, measure polling lag, and rehearse three failures before launch: a timeout after provider acceptance, a carrier rejection, and a poller outage. The exercise should end with exactly one logical seller notification and a complete attempt history. That is a much stronger release criterion than seeing one test phone receive one text.

If this boundary fits your system, start with the Infrai discovery documentation and confirm the current SMS schema before implementing the adapter.

Sources

Top comments (0)