The page says verification_completion_rate < 92% for 10m. An SMS OTP delivery can fail after API acceptance because carrier filtering, sender registration, shared routes, and anti-fraud controls sit outside the 2FA login handler. The on-call sees no obvious application errors, yet a growing set of gaming account claims is stuck at otp_sent. This is legal intake verification for account-ownership claims: the generated claim report, which should be emailed as an attachment after verification, never exists because the workflow did not clear its identity gate.
The immediate answer is to treat SMS acceptance, handset delivery, and successful verification as three separate signals. Register the sender and campaign where the destination requires it, avoid mixing unrelated traffic on shared routes, make retries idempotent, and measure completion by destination class. A successful API response proves that the messaging service accepted a request; it does not prove that a carrier delivered the code or that a person entered it.
Short answer: instrument the state transitions before tuning an alert. Otherwise a carrier filter, an anti-fraud control, an expired code, and a broken client all look like the same falling conversion line.
Why can SMS OTP delivery fail under carrier filtering?
Work backward from the page. Completion is a late signal: it includes carrier handling, handset availability, user behavior, expiry, and the login UI. The earlier warning should compare accepted sends with delivery outcomes, split by country, sender type, and route class. Do not put phone numbers or codes in metric labels or logs.
Acceptance is not delivery.
For this claim flow, I would keep a small state machine: requested, accepted, delivered, verified, expired, and failed. “Delivered” must come from an authenticated delivery callback, not from the initial send response. Some networks do not provide equally timely or detailed delivery receipts, so unknown remains a real state rather than being silently counted as delivered.
The first tempting alert is raw failure count. It is a poor page because traffic volume dominates it. Ten failures during eleven attempts are urgent; ten during a large launch may be ordinary noise. Use a minimum sample size, then alert on a ratio and require persistence.
| Signal | What it can establish | What it cannot establish |
|---|---|---|
| Send request accepted | The upstream API accepted the request | Carrier or handset delivery |
| Delivery receipt | A downstream status was reported | That the claimant saw or entered the code |
| Verification success | The submitted code matched an active challenge | Why abandoned challenges failed |
| Report email accepted | The email request entered its delivery path | Inbox placement or attachment use |
That distinction prevents a common postmortem error: assigning every abandoned login to “SMS delivery.” The evidence may only support “verification did not complete.”
Instrument the challenge rather than the message
The challenge is the unit that matters. One challenge can cause more than one send attempt, while a retrying worker can accidentally submit the same attempt twice. Give both objects stable identifiers. Store the challenge state durably, and enforce a unique idempotency key at the send boundary.
This complete Go example records low-cardinality transition counters without exposing the destination. It also rejects duplicate transitions, which keeps a redelivery from inflating the accepted count.
package main
import (
"fmt"
"sync"
)
type Event struct {
ChallengeID string
EventID string
Country string
SenderClass string
State string
}
type Recorder struct {
mu sync.Mutex
seen map[string]struct{}
counts map[string]int
}
func (r *Recorder) Record(e Event) bool {
r.mu.Lock()
defer r.mu.Unlock()
if _, ok := r.seen[e.EventID]; ok {
return false
}
r.seen[e.EventID] = struct{}{}
key := fmt.Sprintf("%s|%s|%s", e.Country, e.SenderClass, e.State)
r.counts[key]++
return true
}
func main() {
r := Recorder{seen: map[string]struct{}{}, counts: map[string]int{}}
e := Event{"claim-7f3", "evt-001", "US", "registered-long-code", "accepted"}
fmt.Println(r.Record(e), r.Record(e))
fmt.Println(r.counts)
}
Run it with go run main.go; it prints true false, then one accepted event. In production, the same uniqueness rule belongs in the database or durable queue, not only in process memory.
Keep the OTP out of telemetry. Hashing a phone number is not automatically safe if the input space can be guessed; an internal opaque challenge ID is a cleaner correlation key. Logs should retain the selected route class, destination country, provider message identifier, normalized status, attempt number, and timestamps needed to reconstruct the trace.
Trace the route and registration boundary
US application-to-person traffic sent over 10-digit long codes is subject to A2P 10DLC registration. Registration is not a magic delivery flag: the declared use case, consent flow, message content, sender, and actual traffic still need to agree. A legal-intake OTP should not share an identity with promotional game announcements merely because both fit through the same API.
Europe is not one sender regime. Rules and supported sender types vary by destination, so resolve policy from the destination country rather than from a broad EU label. The operational requirement is stable even when the local rule changes: maintain a reviewed mapping from destination and message purpose to an approved sender class, reject unknown mappings before enqueueing, and record the selected policy version on the attempt.
Shared routes increase integration ambiguity. If unrelated tenants or use cases share upstream capacity or sender identity, a reputation or throughput event outside this claim flow can affect it. A dedicated route can reduce that coupling, but it adds registration, capacity planning, and failover work. The decision axis here is integration effort: choose the simplest route whose identity, callback detail, regional coverage, and isolation meet the claim workflow's risk, then test the fallback as a separate route rather than assuming it behaves identically.
Registration does not bypass filtering.
Anti-fraud controls add another boundary. Rate limits should apply to the account, challenge, destination, device, and network signals appropriate to the threat model. A global resend button limit alone is easy to abuse and hard to diagnose. Return a neutral response to the claimant, but emit an internal reason code such as challenge_rate_limited or risk_rejected; never collapse those decisions into carrier failure.
Make retries boring in Go
Retries are allowed only while the challenge is active, and each attempt gets a new event ID under the same challenge ID. Backoff reduces pressure, but it does not make a non-idempotent send safe. The worker must claim an outbox row atomically, submit with a stable idempotency key when the upstream supports one, and persist the result before acknowledging the queue item.
This runnable scheduler shows the timing rule without binding the design to a commercial API:
package main
import (
"fmt"
"time"
)
func nextAttempt(created time.Time, attempt int) (time.Time, bool) {
delays := []time.Duration{0, 15 * time.Second, 45 * time.Second}
if attempt < 0 || attempt >= len(delays) {
return time.Time{}, false
}
next := created.Add(delays[attempt])
if next.After(created.Add(2 * time.Minute)) {
return time.Time{}, false
}
return next, true
}
func main() {
created := time.Date(2026, 1, 2, 15, 4, 5, 0, time.UTC)
for attempt := 0; attempt < 4; attempt++ {
next, ok := nextAttempt(created, attempt)
fmt.Printf("attempt=%d scheduled=%t at=%s\n", attempt, ok, next.Format(time.RFC3339))
}
}
The two-minute window is example policy, not a universal standard. Set expiry from the threat model and claimant experience, then keep the server, UI countdown, and resend rules consistent. A late delivery from an older attempt must not validate a replaced challenge.
After verification, generate the report once and store its immutable object reference with the claim. The email worker should attach that exact artifact and use its own idempotency key. SMS retry logic must never regenerate or resend the report. This boundary keeps a duplicate delivery callback from becoming a duplicate legal document email.
Test the states you cannot reproduce on a laptop
A happy-path test with a real handset proves very little. Build a fake messaging adapter that can accept a request and later emit delivered, failed, delayed, duplicated, or out-of-order callbacks. Sign callbacks in the test exactly as production callbacks are authenticated. Then assert that an unknown event ID is rejected, a duplicate is harmless, and an expired challenge cannot move back to verified. Exercise a more awkward sequence too: the first attempt is accepted, the claimant requests a replacement, the replacement is delivered and verified, and only then does a delayed receipt arrive for the first attempt. The old callback may update its own attempt record, but it must not reopen the challenge, send another attachment, or move the claim backward.
Before deployment, run a small destination matrix covering each supported country and sender class. The pass condition is not “a text arrived once.” Verify the recorded route, callback normalization, expiry behavior, resend copy, and the handoff that creates and emails the report attachment. Keep test recipients controlled and consented.
Deployment should canary by route policy version. Watch accepted-to-delivered and delivered-to-verified ratios independently, plus callback age and unknown-status share. If accepted-to-delivered falls only for one destination and sender class, investigate registration, content, route, and carrier handling. If delivery is steady while verification falls, look at expiry, UI submission, risk decisions, and challenge replacement. This method has a clear limitation: sparse traffic and incomplete delivery receipts can make a route-level ratio inconclusive. In that case, hold the page, widen the observation window, and use controlled destination tests plus upstream evidence; do not manufacture certainty from three messages.
No guessing.
The runbook should also say what not to do during an incident: do not spray the same OTP across alternate senders, disable fraud controls, or repeatedly retry an expired challenge. Those actions muddy the trace and can worsen filtering. Offer a separately verified recovery path when policy permits it, and preserve the audit record for the claim.
Tune the page against claimant harm
Start with two alerts. Page on a sustained accepted-to-delivered regression for a supported route after a minimum event count. Create a lower-urgency ticket for a delivered-to-verified regression unless it crosses a severe, persistent threshold or blocks the entire intake flow. The exact numbers must come from the service's baseline and error budget; no public source can supply the right threshold for this workload.
Review alert decisions after incidents. Record which dimensions isolated the fault, how long the page preceded user reports, and whether the responder could act. If the alert fired on a low-volume country with three attempts, raise the sample floor or use a longer window. If aggregation hid a single sender class, split that class out.
Sensitivity has a cost. Thresholds set too low page on normal receipt delays and small denominators, training responders to distrust the signal. Thresholds set too high wait until claimants abandon verification and the report-email stage stops. The useful threshold is the one that catches a sustained, actionable route regression early enough to protect the intake flow without turning ordinary variance into an emergency.
Further reading
- https://www.twilio.com/docs/messaging/compliance/a2p-10dlc
- https://www.twilio.com/docs/messaging/guides/track-outbound-message-status
- https://pages.nist.gov/800-63-4/sp800-63b.html
- https://www.rfc-editor.org/rfc/rfc4226
- https://www.rfc-editor.org/rfc/rfc6238
- https://docs.aws.amazon.com/ses/latest/dg/Welcome.html
Top comments (0)