Compliance evidence changes the SMS API decision: choose the provider and architecture that let you prove what your e-commerce signup system requested, what the carrier-facing network reported, and how your application reacted, without storing the verification secret in logs. Short answer: model delivery as an asynchronous state machine, poll through a narrow internal interface, retain append-only status observations under a documented policy, and set an SLO on terminal evidence rather than treating an accepted API call as delivery.
Acceptance is not delivery.
This is not a claim that SMS proves account ownership. NIST treats use of the public switched telephone network for out-of-band authentication as a restricted authenticator and tells verifiers to consider risks such as SIM change and number porting. A verification link sent by SMS is therefore one signal in a signup control, not a reason to discard session binding, expiration, rate limits, and abuse detection.
The bounded incident worth designing around is ordinary: a shopper requests a verification link, the messaging API accepts the transactional notification, and the SaaS app marks the signup notification as complete. Later, support needs to answer whether the message reached a terminal delivery state. The original response cannot answer that question. The invariant is sharper than “send SMS reliably”: every accepted request needs a correlation ID, a sequence of timestamped observations, and an explicit terminal outcome or timeout. Education attendance alerts expose the same evidence gap, but they have different content, consent, and escalation policies; this implementation stays with e-commerce account signup so those policies are not casually conflated.
What should a Node.js SaaS app require from an SMS alerts API?
Start with the questions an investigator will actually ask. Which signup initiated the send? Which destination region and policy applied? What content version was selected? When did the provider accept the request? Which status transitions were observed, and when did the system stop checking? Evidence must also show that retries did not create duplicate user-visible messages.
Do not log the link.
It is a bearer secret until it expires. Store a template identifier, a content-version hash, the provider's opaque message identifier, an internal correlation ID, coarse destination jurisdiction, timestamps, and normalized delivery states. Keep the phone number in the account data store under its own access and retention controls; the delivery ledger usually needs only a stable, scoped recipient reference. A hash is useful for integrity comparison, but blindly hashing a phone number does not make a low-entropy identifier anonymous. For a Node.js application, keep this ledger contract outside the SDK wrapper; the Go worker below can sit behind a queue while the app remains responsible for signup authorization. That boundary avoids making a runtime library the system of record and keeps provider replacement from rewriting the account flow.
The data model below separates a send attempt from its observations. One attempt can accumulate several observations, while the unique idempotency key prevents a worker retry from becoming a second logical send.
package delivery
import "time"
type State string
const (
StateAccepted State = "accepted"
StateQueued State = "queued"
StateSent State = "sent"
StateDelivered State = "delivered"
StateFailed State = "failed"
StateUnknown State = "unknown"
)
type Attempt struct {
ID string
SignupID string
RecipientRef string
IdempotencyKey string
ProviderMessageID string
TemplateVersion string
RegionPolicy string
AcceptedAt time.Time
}
type Observation struct {
AttemptID string
State State
Reason string
ObservedAt time.Time
RawDigest string
}
Retention is a policy decision, not a constant copied from an SDK example. Define the purpose, access roles, deletion schedule, and legal basis with the people responsible for privacy and compliance in each operating region. The useful engineering property is append-only observation during the retention window, followed by verifiable deletion. Restricting mutation makes the history easier to explain, but it does not replace access logs or a retention job.
Polling is a state machine, not a timer
A polling-only setup can be simple, provided “simple” does not mean an endless fixed-interval loop. The worker should distinguish terminal from nonterminal states, bound the total observation window, add jitter, respect server retry guidance when present, and preserve an unknown result when the available evidence cannot justify a stronger claim. This design deliberately uses delivery status polling with no webhook; it is useful when the team cannot expose and authenticate an inbound callback, though the request volume becomes a capacity concern.
Do not translate “accepted” into “delivered.” Acceptance usually establishes only that the downstream system received the submission. Your normalized vocabulary should be deliberately smaller than any provider vocabulary, and the mapping should be reviewed whenever an upstream contract changes. Preserve a digest of the raw response, or an encrypted raw record if policy requires it, so the normalization can be examined later without letting provider-specific fields leak through the application.
package delivery
import (
"context"
"errors"
"math/rand"
"time"
)
type Receipt struct {
State State
Reason string
RawDigest string
RetryAfter time.Duration
}
type StatusReader interface {
Status(ctx context.Context, messageID string) (Receipt, error)
}
type Recorder interface {
Append(ctx context.Context, observation Observation) error
}
func ObserveUntilTerminal(
ctx context.Context,
reader StatusReader,
recorder Recorder,
attempt Attempt,
deadline time.Time,
) (State, error) {
delay := 2 * time.Second
for time.Now().Before(deadline) {
receipt, err := reader.Status(ctx, attempt.ProviderMessageID)
if err == nil {
err = recorder.Append(ctx, Observation{
AttemptID: attempt.ID,
State: receipt.State,
Reason: receipt.Reason,
ObservedAt: time.Now().UTC(),
RawDigest: receipt.RawDigest,
})
if err != nil {
return StateUnknown, err
}
if receipt.State == StateDelivered || receipt.State == StateFailed {
return receipt.State, nil
}
if receipt.RetryAfter > 0 {
delay = receipt.RetryAfter
}
}
jitter := time.Duration(rand.Int63n(int64(delay / 4)))
timer := time.NewTimer(delay + jitter)
select {
case <-ctx.Done():
timer.Stop()
return StateUnknown, ctx.Err()
case <-timer.C:
}
if delay < 30*time.Second {
delay *= 2
}
}
return StateUnknown, errors.New("delivery evidence window expired")
}
Those durations are example control values, not measured recommendations. Capacity planning comes first: if the peak signup rate is R, each message receives P polls on average, and the observation window is W seconds, the steady polling demand is roughly R × P requests per signup cohort, distributed across W. Test the distribution, not merely the mean. A retry surge after an upstream slowdown can synchronize workers unless backoff includes jitter and the queue applies concurrency limits. Suppose the planned peak is 40 signup sends per second and observation averages six status reads; that assumption implies 240 reads associated with each second's signup cohort before retries, so it belongs in a load test and provider quota review, not in a spreadsheet nobody pages from. Replace both figures with measured workload data before launch.
No receipt, no claim.
The SLO should describe what the system controls. “99.9% of SMS messages are delivered” mixes application performance with handset availability and carrier behavior. A defensible internal SLO is closer to: for accepted verification sends, the evidence pipeline records either a terminal normalized status or an explicit observation timeout within the configured window. Measure user-facing verification completion separately, because delivery receipts and completed signups answer different questions.
Build the preventative path before choosing the API
Make the send path transactional at the application boundary. In one database transaction, create the pending signup and an outbox row with a unique idempotency key. A worker claims the outbox row, submits the message once, records the opaque provider ID, and schedules observation work. If the process stops between remote acceptance and local persistence, reconciliation must use the same idempotency key or a provider-supported lookup rather than guessing and sending again.
package delivery
import (
"context"
"crypto/sha256"
"encoding/hex"
)
type SendRequest struct {
Destination string
Body string
IdempotencyKey string
}
type Sender interface {
Send(ctx context.Context, request SendRequest) (messageID string, err error)
}
func ContentDigest(templateVersion, renderedBody string) string {
sum := sha256.Sum256([]byte(templateVersion + "\x00" + renderedBody))
return hex.EncodeToString(sum[:])
}
Keep the rendered body out of the ledger even though this function needs it transiently to calculate a digest. The verification link should carry a random, single-use, short-lived token; bind it to the intended signup session, invalidate it after successful use, and avoid putting account details in the SMS body. The endpoint that consumes the link must enforce those properties. SMS transport does not. A Node.js web process can create the outbox record and return promptly while the Go worker performs submission and polling; the queue contract, rather than shared client code, is the integration surface.
Test four failure boundaries before production: an ambiguous send timeout, repeated execution of the same outbox item, a status that never becomes terminal, and a provider state the normalizer does not recognize. The correct response to a new state is unknown plus an alert for mapping review. Inventing a favorable mapping destroys the evidence the design was meant to preserve.
Deployment deserves similar restraint. Release a new status mapping behind recorded fixtures, canary the polling worker at limited concurrency, and verify that rollback does not erase scheduled observations. Alert on queue age, observation timeouts, unknown-state rate, duplicate idempotency conflicts, and the age of the oldest accepted attempt without a terminal result. Page only on symptoms tied to an SLO or imminent data loss; a single transient polling error belongs in metrics, not on a person's phone.
The buy-versus-build boundary
The provider choice follows from the evidence contract. Ask candidates to demonstrate their status semantics, data location options, retention controls, subprocessor information, access controls, exportability, idempotency behavior, rate-limit signaling, and the procedure for obtaining records during an investigation. Marketing labels do not answer those questions.
| Boundary | Managed capability | Team-owned capability | Decision pressure |
|---|---|---|---|
| Carrier connectivity | Network submission and downstream status inputs | Normalize status and track gaps | Coverage by destination and documented semantics |
| Evidence storage | Optional provider-side history | Correlated, access-controlled ledger | Retention, export, deletion, and audit needs |
| Retry behavior | API throttling and request handling | Outbox, idempotency, backoff, reconciliation | Duplicate risk and on-call load |
| Verification control | Message transport | Token issuance, session binding, expiry, abuse controls | Account-takeover risk |
| Regional processing | Published processing locations and subprocessors | Routing policy and recipient classification | Contractual and regulatory requirements |
Self-hosting the orchestration layer buys a stable interface and keeps application policy under your control, but it also creates migrations, key management, backups, capacity work, and an on-call obligation. Buying more of the evidence workflow can reduce that operational surface, yet the team still owns the correctness of signup authorization and the ability to export proof. Choose the narrowest boundary your team can operate to its SLO, then test exitability before signing.
Run a proof with synthetic numbers that cannot reach a real shopper. Confirm that the candidate rejects malformed destinations predictably, returns a durable message identifier, exposes documented terminal and nonterminal states, and lets the system reconcile an ambiguous submission without duplication. Repeat the exercise for both US and EU routing policies if those are in scope, because a global product label is not evidence that contracts, processing paths, sender requirements, or status behavior are identical.
When does this design not apply?
This approach has limitations and is not suitable everywhere. Do not use SMS as the sole recovery or high-assurance authentication factor when your risk assessment cannot accept number reassignment, interception, porting, or SIM-change risks. NIST's restricted designation is a reason to assess alternatives and compensating controls, not a decorative footnote. The trade-off is explicit: SMS has broad handset reach, while possession of a phone number is weaker evidence than a phishing-resistant authenticator and delivery status is weaker evidence than completed verification.
Polling is also the wrong default when a trustworthy webhook is available and your evidence window or volume makes polling wasteful. The ledger and state machine still apply; only the ingestion mechanism changes. A webhook reduces repeated status traffic but adds a public callback surface, signature or credential validation, replay defense, and ingress availability to the team's responsibilities. Conversely, a low-volume system may accept bounded polling because it removes that public surface, as long as capacity calculations, timeout semantics, and rate limits still close. For high-volume attendance alerts with a narrow delivery deadline, push status plus reconciliation polling may be a better fit than the polling-first setup shown here.
Email fallback requires its own delivery and security analysis. DMARC, defined by RFC 7489, lets a domain publish policy tied to SPF and DKIM identifier alignment and receive reports; it does not prove that a recipient opened a verification message or that a signup should be authorized. Keep email authentication evidence distinct from application verification evidence.
The conclusion is operational: establish the audit questions and normalized states first, make duplicate prevention and secret handling application responsibilities, then evaluate transport providers against that contract. A short integration is useful. An explainable one is the requirement.
Sources
- NIST SP 800-63B, Digital Identity Guidelines: Authentication and Lifecycle Management: https://pages.nist.gov/800-63-3/sp800-63b.html
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)