For a React Native mobile app, an SMS OTP login recovery API belongs on the backend. A server-owned, single-use challenge carries a short expiry and an attempt limit. The resend path must not extend the challenge deadline. The app should request and submit codes; it should never decide whether a code is valid, how many attempts remain, or when another message may be sent.
Short answer: treat SMS OTP as an administrator account-recovery ceremony, not as a messaging feature. Bind each challenge to the recovery transaction, store only a keyed digest of the code, make verification atomic, and rate-limit by account, destination, device, and network. Autofill is a usability hint. It is not proof of possession.
The page says “administrator recovery completion rate below SLO.” On-call sees requests, sends, deliveries where the carrier exposes them, verification attempts, and successful recoveries split by country and carrier. A logistics SaaS cannot let a dispatcher remain locked out during an overnight exception, but lowering every control during an incident turns an availability problem into an account-takeover problem.
What should page before recovery failures reach users?
A page on raw SMS send errors fires too late and often points at the wrong layer. The earlier signal is the recovery funnel: accepted challenge requests are rising, but verified challenges are not. Instrument each state transition with a low-cardinality reason such as expired, attempts_exhausted, superseded, send_rejected, or verified; keep phone numbers, codes, and free-form carrier text out of metric labels and logs. Then inspect the shape before changing policy: a broad increase in send_rejected points toward transport or destination quality, a concentration of expired outcomes may indicate late arrival, and a jump in superseded challenges usually means the resend experience deserves attention. None of those signals proves a root cause by itself. Correlate them with queue age, transport feedback, verification latency, and deployment markers, while keeping the event vocabulary independent of any provider so a routing change does not erase the baseline.
Do not page on sends alone.
Work backward from the page. A fall in completions with stable request volume can come from delivery latency, invalid or suppressed destinations, users requesting repeated codes and entering an older one, or verification storage failing to provide atomic consume semantics. Email bounce processing belongs in the same contact-hygiene model: a permanent bounce should suppress future recovery email to that address, while an SMS destination that is known invalid should not be hammered by retries. Suppression is a state transition with provenance and review rules, not a CSV someone remembers to upload.
The SLO should measure the user outcome, with a separately reported security guardrail. For example, define the service-level indicator as eligible recovery transactions completed within the allowed ceremony window divided by eligible recovery transactions started. “Eligible” must be written down before an incident; silently removing provider failures or invalid-recipient responses makes the graph calm and the service unreliable.
How should a React Native mobile app handle SMS OTP?
The app sends an account-recovery request to the backend, receives an opaque challenge ID plus a server-derived resend time, and displays one code input. React Native can use the platform's one-time-code autofill affordance, but the backend contract must still work when autofill is unavailable, delayed, or selects an old message.
Do not return whether an administrator account exists. Use the same outward response shape and broadly similar work path for known and unknown accounts, then notify through established channels after a recovery attempt where policy permits. A resend should normally supersede the previous code, preserve the original ceremony deadline, and return the same neutral response when throttled. This removes the common race in which three valid codes coexist and on-call cannot explain which message the user entered.
One detail is easy to miss: client countdowns drift. The server owns retry_at; the app renders it and refreshes from the next response. Disabling a button locally is good interface behavior, but it is not abuse prevention.
The server is authoritative.
Make verification one atomic decision
The following Go sketch shows the boundary even if the HTTP application itself runs on Node.js. It omits code generation and transport wiring on purpose; those require a cryptographically secure generator, secret management, and a transactional data store, and pretending otherwise in a short sample would be dangerous.
package recovery
import (
"context"
"errors"
"time"
)
var ErrInvalidChallenge = errors.New("invalid or expired challenge")
type Challenge struct {
ID string
CodeDigest []byte
ExpiresAt time.Time
Attempts int
UsedAt *time.Time
}
type Store interface {
// Consume atomically checks expiry, attempt budget, digest, and unused state.
// A failed comparison increments Attempts in the same transaction.
Consume(ctx context.Context, id string, candidateDigest []byte, now time.Time) (bool, error)
}
type Verifier struct {
Store Store
Now func() time.Time
Digest func(string) []byte // Implement with a server-secret keyed construction.
}
func (v Verifier) Verify(ctx context.Context, challengeID, code string) error {
if challengeID == "" || code == "" {
return ErrInvalidChallenge
}
ok, err := v.Store.Consume(ctx, challengeID, v.Digest(code), v.Now())
if err != nil {
return err
}
if !ok {
return ErrInvalidChallenge
}
return nil
}
Atomicity is the mechanism, not an implementation flourish. Two concurrent submissions must not both consume one challenge, and a failed guess must not escape the attempt counter because another request updated the row first. Use a transaction, compare-and-swap, or a datastore primitive with equivalent guarantees. After success, rotate the administrator's sessions or apply the product's documented recovery policy; recovery is incomplete if an attacker-controlled session remains valid.
Race tests are mandatory.
Capacity planning starts with attempts, not users. Estimate peak recovery starts during a regional logistics incident, multiply by the maximum sends permitted per ceremony, then separately budget verification requests because guesses can exceed sends. Queueing outbound work can absorb a brief transport slowdown, but a queued message that will arrive after the challenge expires has negative value. Give jobs a deadline and discard stale work.
Abuse controls need several weak signals
A single IP limit punishes offices and carrier-grade NAT; a phone-only limit lets an attacker distribute requests across destinations; an account-only limit can be used to lock out a known administrator. Layer modest limits across challenge, account, normalized destination, device signal, and network prefix, then escalate friction rather than exposing which key fired. Keep a hard global circuit breaker for spend and traffic anomalies, with a controlled operational override that does not disable verification limits.
The trade-off is explicit: stricter resend and attempt budgets reduce guessing and SMS pumping, while increasing lockout risk when delivery is slow. Measure both. Alert on sudden changes in requests per completed recovery, destinations per network, resend ratio, invalid-recipient responses, and age-at-verification; investigate distributions rather than choosing a universal threshold from intuition. NIST SP 800-63B also treats the public switched telephone network as a restricted authenticator channel, which is a useful architectural constraint: SMS recovery needs risk review and an alternative recovery path rather than being treated as permanent proof of identity.
| Decision | Managed component | Self-hosted component | SRE question |
|---|---|---|---|
| Message transport | Less carrier integration work | More routing control | Who absorbs provider and carrier changes on-call? |
| Challenge state | Faster operational start | Direct atomicity and retention control | Can one transaction enforce consume-once semantics? |
| Abuse scoring | Broader external signals may exist | Easier policy inspection | Can operators explain and safely override a denial? |
| Recipient suppression | Delivery feedback may be normalized | Data stays near account policy | How quickly does a permanent failure stop retries? |
There is no honest winner without workload data. Buy-versus-build review should include lock-in at the event schema and suppression-list layers, not only the send API, because migrations fail when historical recipient state cannot move cleanly.
Instrument the change, then tune the page
Emit one event for each recovery transition with a correlation ID, channel, coarse region, provider-independent outcome, and latency. Trace the request through enqueue, transport acceptance, delivery feedback when available, verify, and consume. Retain the minimum data required for security review and support; hash or tokenize destinations consistently only when the threat model and retention policy justify correlation.
Roll out the new funnel alert in shadow mode. Compare it with completed recoveries and support contacts across ordinary peaks and planned logistics events, then page only on sustained burn against the recovery SLO. A multi-window burn-rate alert is preferable to “five failures in five minutes” because it ties urgency to error-budget consumption rather than arbitrary traffic volume.
Bad thresholds have a bill. Too loose, and dispatch administrators discover the incident first. Too tight, and every carrier wobble wakes someone who soon learns to distrust the page; aggressive automated suppression can also strand valid recipients after a transient error. The control is only finished when the team has a runbook that distinguishes transport degradation, invalid-recipient growth, verification-store faults, and abuse, and each branch names a reversible action.
References
- Google, “Email sender guidelines”: https://support.google.com/a/answer/81126
- Twilio, “SMS documentation”: https://www.twilio.com/docs/sms
- NIST, “Digital Identity Guidelines: Authentication and Lifecycle Management”: https://pages.nist.gov/800-63-3/sp800-63b.html
- OWASP, “Forgot Password Cheat Sheet”: https://cheatsheetseries.owasp.org/cheatsheets/Forgot_Password_Cheat_Sheet.html
Further reading
- Google, “SMS Retriever API”: https://developers.google.com/identity/sms-retriever/overview
- Apple, “Enabling AutoFill for domain-bound SMS codes”: https://developer.apple.com/documentation/security/enabling-autofill-for-domain-bound-sms-codes
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.