The useful outcome from a low-cost SMS alert service is a defensible record that a passwordless account notification was attempted and resolved, without putting the generated report or a reusable secret in the message. The least complex design uses an outbox, one delivery worker, status callbacks, and four signals: queue age, acceptance, final outcome, and user action.
TL;DR: For passwordless backup and account notifications across the US and EU, subject Twilio, Vonage, and Telnyx to the same failure drill and evidence checklist. The lowest advertised SMS rate cannot establish operational fitness. A viable route has to support consent evidence, an immutable intent record, correlated delivery updates, duplicate suppression, and an operator-readable history. For a B2B SaaS report, email carries the attachment; SMS only says that the report is ready or that delivery needs attention.
The page usually arrives too late: report_notification_unresolved is above its threshold, and support already has a customer asking why a monthly compliance report never appeared. The on-call view should contain a tenant-safe correlation ID, notification class, destination region, outbox age, attempt count, last normalized state, and the timestamp of the last transition. It should not expose the phone number, report name, one-time code, or attachment URL.
That distinction matters under pressure.
A provider can accept a request quickly and still leave the team unable to distinguish a delayed carrier handoff from a stale callback, a duplicate worker, or a user who received the message but never completed the action. My default decision rule is therefore conservative: treat acceptance, delivery, and user action as three separate facts, because collapsing them produces a comforting graph and a useless incident record.
How should a low-cost SMS alert service handle passwordless accounts?
Work backward from the page. A final failure counter is useful, but it is a lagging signal. The earlier warning is usually queue age: notification intents remain ready while workers are stalled, rate-limited, or unable to claim work. Alert on the oldest eligible item by notification class and destination region. A single global average hides a small, important queue behind ordinary traffic.
Next, separate provider acceptance from terminal delivery. accepted means the downstream service took responsibility for an attempt; it does not establish handset receipt or user action. Normalize every provider-specific update into a small internal state machine, while retaining the original status as evidence. A practical sequence is pending, submitted, delivered, failed, with unknown for updates that cannot be mapped safely. Never turn an unfamiliar value into success.
The fourth signal belongs to the application. For a passwordless backup flow, the meaningful result is that the short-lived recovery action was completed, expired, or was replaced. For a generated report, it may be that the email attachment was produced and its notification acknowledged. Delivery and action are different clocks. Combining them makes an SMS route look healthy while the account workflow is broken.
One sharp page is enough.
Page when a user-facing objective is in danger and an operator can act; create a ticket for slower evidence gaps. A missing callback for one message may be noise. A growing regional cohort with old outbox rows is a service condition. This is a deliberate trade-off: a per-message page has fast detection but poor signal quality, while a cohort threshold waits longer but gives the responder evidence of shared impact.
Build an evidence trail, not a message log
The outbox row records intent before network I/O. Give it a stable idempotency key derived from the business event and notification class, then enforce uniqueness in storage. The worker may crash after a remote service accepts a request but before the local transaction records that acceptance. Retries are unavoidable. The stable key and a reconciliation path make that ambiguity containable. It does not guarantee exactly-once delivery across two independent systems; no local uniqueness constraint can do that by itself. The limitation is important enough to put in the runbook, because the ambiguous interval is where an operator must reconcile before retrying.
Keep the payload sparse. A report notification can say that a report is available and direct the recipient to the normal authenticated application. The email system sends the attachment under its own policy. An SMS recovery message should use a random, single-use, expiring token or code and apply consistent responses and rate controls. The OWASP Forgot Password Cheat Sheet also recommends protecting the account until a valid token is presented, avoiding account-existence leaks, and invalidating the token after use. Those are application duties; changing SMS services does not remove them.
Consent evidence needs similar discipline. GDPR Article 7 places the burden of demonstrating consent on the controller when processing is based on consent, requires a distinguishable request in clear language, and gives the data subject a right to withdraw. Store the policy or notice version, purpose, capture time, and withdrawal time as first-class records. Do not infer consent from a successful delivery receipt.
For each attempt, retain the internal notification ID, provider-neutral route key, request time, response classification, downstream reference, callback receipt time, normalized transition, and a hash or redacted representation of the destination. Define retention and access rules with legal and security owners. More logs are not automatically better evidence; uncontrolled message bodies create another sensitive-data store.
Instrument the transition boundary
Instrumentation belongs where state changes, not around a convenient HTTP client timer. The following Go sketch shows the narrow contract. Storage must enforce the unique idempotency key, and callback ingestion should apply the same transition rules in a separate transaction.
package notify
import (
"context"
"errors"
"time"
)
type Intent struct {
ID string
IdempotencyKey string
Region string
Class string
Destination string
Body string
CreatedAt time.Time
}
type Receipt struct {
RemoteID string
Status string
}
type Sender interface {
Send(ctx context.Context, destination, body, idempotencyKey string) (Receipt, error)
}
type Store interface {
Claim(ctx context.Context, now time.Time) (Intent, error)
MarkSubmitted(ctx context.Context, id, remoteID string, at time.Time) error
MarkRetryable(ctx context.Context, id, reason string, at time.Time) error
}
type Metrics interface {
ObserveAttempt(class, region, result string, elapsed time.Duration)
}
var ErrNoWork = errors.New("no eligible notification")
func DeliverOne(ctx context.Context, store Store, sender Sender, metrics Metrics, now time.Time) error {
intent, err := store.Claim(ctx, now)
if err != nil {
return err
}
started := time.Now()
receipt, err := sender.Send(ctx, intent.Destination, intent.Body, intent.IdempotencyKey)
if err != nil {
metrics.ObserveAttempt(intent.Class, intent.Region, "retryable", time.Since(started))
return store.MarkRetryable(ctx, intent.ID, "send_error", now)
}
if receipt.RemoteID == "" {
metrics.ObserveAttempt(intent.Class, intent.Region, "invalid_receipt", time.Since(started))
return store.MarkRetryable(ctx, intent.ID, "missing_remote_id", now)
}
metrics.ObserveAttempt(intent.Class, intent.Region, "accepted", time.Since(started))
return store.MarkSubmitted(ctx, intent.ID, receipt.RemoteID, now)
}
This code deliberately does not mark a message delivered. Only an authenticated, validated status update can advance it to that state. Callback handlers must tolerate repeats and out-of-order arrival; a late submitted event must not move a terminal delivered record backward. Record rejected transitions for investigation without mutating the current state.
Avoid high-cardinality metric labels. Notification IDs and phone numbers belong in access-controlled traces or audit records, not metric dimensions. Metrics need bounded labels such as class, region, normalized outcome, and route. Logs carry the correlation ID. The runbook joins them.
Compare routes with one failure drill
Run the same controlled workload against every candidate in the US and the EU regions you actually serve, using test recipients and approved content. Confirm current regional registration, sender, consent, retention, subprocessor, and data-transfer requirements directly with the relevant provider and counsel; these conditions can vary by destination and account configuration. This test is intentionally poor at judging marketing reach or negotiating commercial terms. It is suitable for evidence and failure handling, which is the narrower decision being made here.
Score evidence you can reproduce:
| Decision area | Test | Evidence to retain |
|---|---|---|
| Duplicate control | Retry after an ambiguous client timeout | Number of user-visible sends and correlated attempt records |
| Status integrity | Replay and reorder signed callbacks | Accepted, rejected, and ignored state transitions |
| Regional operation | Exercise each required destination class | Timestamped results by country and sender configuration |
| Compliance support | Add and withdraw a test recipient's consent | Purpose, notice version, capture, suppression, and access history |
| Incident response | Pause a worker and delay callbacks | Queue-age alert, page context, runbook action, and recovery record |
Do the paperwork as part of the test, not after the technical bake-off. The service boundary includes account access controls, callback verification, audit export, retention controls, regional processing terms, and escalation paths. Record unknowns as unknowns. A polished dashboard is not evidence that the route meets your obligations.
Cost still belongs in the decision, but use a workload model rather than a headline unit price. Include destination mix, message segments, sender requirements, retries, support, engineering time, and the cost of retaining or exporting evidence. Rates and rules change.
The comparison should remain useful when they do.
Tune the page against the cost of being wrong
A threshold set to zero tolerance will page on ordinary asynchronous behavior. A loose global threshold will miss a quiet regional failure. Start with separate service-level indicators for oldest ready intent, time from intent to acceptance, time from acceptance to terminal state, and time from delivery to application action. Then choose windows from the user promise and the workflow's risk, not from a tidy round number copied from another queue.
Test both sides before rollout. Pause consumers to verify that queue age fires early. Deliver callbacks late and out of order. Force an ambiguous timeout after remote acceptance. Withdraw a test recipient's consent while an intent is queued and verify suppression at send time. For report delivery, confirm that the SMS contains neither the attachment nor a bearer link and that email failure cannot be mistaken for SMS success.
The on-call action should be explicit: inspect cohort size, freeze unsafe retries, confirm callback ingestion, reconcile ambiguous attempts, and communicate through the incident process. If a page does not change an operator's next action, demote or redesign it.
False positives have a compliance cost too. Repeated pages train responders to acknowledge without investigation, and aggressive automated retries can create duplicate account messages. Tune with reviewed incident data, preserve changes to alert policy as evidence, and require a reason when a threshold moves. The goal is not a silent pager.
It is an early, credible signal tied to a safe action.
Further reading
- OWASP, “Forgot Password Cheat Sheet”: https://cheatsheetseries.owasp.org/cheatsheets/Forgot_Password_Cheat_Sheet.html
- GDPR, “Article 7: Conditions for consent”: https://gdpr-info.eu/art-7-gdpr/
Top comments (0)