The page says that a marketplace seller cannot open a new property order because the login code never arrived. Support sees an impatient seller. On-call sees an authentication attempt, one accepted SMS request, and no trustworthy evidence of delivery.
TL;DR: model SMS and backup email as one expiring challenge with separate delivery attempts. Poll delivery status only as evidence, never as permission to authenticate. After a bounded SMS wait, offer an explicit email fallback, reuse the challenge, hash every code, and make the fallback transition idempotent. This is the least complex design that keeps login correctness inside the application while still exposing enough state to operate delivery.
Do not automatically send both channels. A surprise email expands the disclosure surface and can train users to accept codes they did not request. For a seller trying to acknowledge a new order, the useful promise is narrower: one visible path forward, even when a carrier receipt is late or absent.
How should two-factor authentication handle SMS failure and backup email?
Work backward from the page. The seller reports “no code,” but that phrase collapses several states: the application may not have submitted the message, the delivery service may have rejected it, a carrier may still be processing it, or the handset may not have displayed it. A provider acceptance response is not proof that a person received anything. This distinction should shape both the UI and the alert.
The signal that should fire earlier is a growing age of unresolved delivery attempts combined with user impact. Count challenges that remain usable but have neither a confirmed delivery outcome nor a successful verification. Split that count by channel and destination region, then correlate it with fallback requests and completed logins. A single delayed receipt is noise; a cohort of sellers repeatedly asking for fallback is a service symptom.
Receipts lag.
For each challenge, persist a stable challenge ID, purpose, account ID, expiry, consumed time, and code hash. For each channel attempt, persist a separate attempt ID, state, creation time, last observation time, and a redacted destination fingerprint. Never put the raw code, phone number, or email address in logs or metric labels.
The state transitions matter more than the transport SDK:
| Event | Challenge state | Delivery attempt state | Operator meaning |
|---|---|---|---|
| SMS accepted for processing | pending | accepted | Transport has custody; delivery is unknown |
| Receipt reports delivery | pending | delivered | Delivery evidence exists; code is still unverified |
| Bounded wait expires | pending | unresolved | UI may offer fallback |
| Seller requests email | pending | accepted | Same challenge, second attempt |
| Correct code is submitted | consumed | any | Authentication succeeds exactly once |
| Challenge expires | expired | any | Neither channel can revive it |
Keep “delivered” and “verified” separate. That one boundary prevents a delayed callback from turning into an authentication decision.
Implement one challenge and idempotent channel attempts
The application owns the challenge. A delivery adapter owns submission and status lookup, but it cannot consume a code. The following Go types are deliberately small enough to sit in front of an SMS gateway, an SMTP relay, or a queue-backed internal service.
package otp
import (
"context"
"errors"
"time"
)
type Channel string
const (
SMS Channel = "sms"
Email Channel = "email"
)
type DeliveryState string
const (
Accepted DeliveryState = "accepted"
Delivered DeliveryState = "delivered"
Failed DeliveryState = "failed"
Unresolved DeliveryState = "unresolved"
)
type Challenge struct {
ID string
AccountID string
Purpose string
CodeHash []byte
ExpiresAt time.Time
ConsumedAt *time.Time
}
type Attempt struct {
ID string
ChallengeID string
Channel Channel
State DeliveryState
CreatedAt time.Time
ObservedAt time.Time
}
type Sender interface {
Send(ctx context.Context, channel Channel, destination, code, idempotencyKey string) (string, error)
Status(ctx context.Context, attemptID string) (DeliveryState, error)
}
var ErrAlreadyRequested = errors.New("channel already requested")
Put a unique constraint on (challenge_id, channel). The fallback handler first creates the email attempt in a transaction; only the winner submits it. A retry returns the existing attempt instead of sending a second code. Use an outbox if the database transaction and the delivery submission cross process boundaries. The outbox worker can retry, keyed by the stable attempt ID, without inventing another logical send.
I treat duplicate delivery as a first-class failure because retries are normal. The trade-off is a little more stored state and cleanup after expiry. It is worth paying: an ambiguous timeout should not become two SMS messages followed by two emails.
Codes need a cryptographically secure random source, a short lifetime chosen for the threat model, rate limits on issue and verification, and a slow keyed comparison against a stored hash. OWASP also recommends single use and invalidation after success. Do not encode the account ID or an order ID into the code itself.
Poll for evidence, not authority
Polling belongs in a bounded background reconciler. The browser should poll your application for a coarse UI state, not contact a delivery provider or receive transport credentials. Likewise, the reconciler should stop after a terminal receipt or after the observation window closes.
package otp
import (
"context"
"math/rand"
"time"
)
type AttemptStore interface {
Due(ctx context.Context, before time.Time, limit int) ([]Attempt, error)
UpdateState(ctx context.Context, id string, state DeliveryState, observedAt time.Time) error
}
func Reconcile(ctx context.Context, store AttemptStore, sender Sender, now time.Time) error {
attempts, err := store.Due(ctx, now, 100)
if err != nil {
return err
}
for _, attempt := range attempts {
state, err := sender.Status(ctx, attempt.ID)
if err != nil {
continue // Preserve the last known state; retry with backoff.
}
if err := store.UpdateState(ctx, attempt.ID, state, now); err != nil {
return err
}
}
return nil
}
func NextPoll(base time.Duration) time.Duration {
// Jitter prevents every pending attempt from landing on the same tick.
return base + time.Duration(rand.Int63n(int64(base/2)))
}
The sample limit of 100 is a batch-control choice, not a universal capacity target. Tune batch size, backoff, and the maximum observation window from queue latency and downstream quotas. Persist the next-poll time so a restart does not reset every attempt to immediate work.
A callback can shorten the wait, but it does not remove reconciliation. Authenticate callbacks according to the transport's documented signing scheme, retain the provider event ID for deduplication, and accept out-of-order events through monotonic state rules. A late “accepted” event must not overwrite “delivered.” If email is the fallback, configure domain authentication; DKIM defines a domain-level signature that receivers can validate, while SPF and DMARC cover related authorization and policy concerns.
Make fallback a deliberate user action
Offer “Send code by email” after the SMS observation window, or immediately after a definitive failure that is safe to expose as a generic delivery problem. Do not reveal whether a phone number or email exists. Render a masked destination already associated with the authenticated pre-login session, and require the same challenge token when requesting fallback.
The email must carry a code for the existing challenge. Generating an independent email challenge creates races: an SMS can arrive after email starts, support cannot tell which challenge is current, and successful verification may leave another credential live. With one challenge, either code representation can be checked against one verifier and the first successful transaction consumes the challenge.
Make consumption atomic:
package otp
import (
"context"
"crypto/subtle"
"errors"
"time"
)
var ErrInvalidCode = errors.New("invalid or expired code")
type ChallengeStore interface {
LoadForUpdate(ctx context.Context, id string) (Challenge, error)
Consume(ctx context.Context, id string, at time.Time) error
}
func Verify(ctx context.Context, store ChallengeStore, id string, candidateHash []byte, now time.Time) error {
challenge, err := store.LoadForUpdate(ctx, id)
if err != nil || challenge.ConsumedAt != nil || !now.Before(challenge.ExpiresAt) {
return ErrInvalidCode
}
if len(candidateHash) != len(challenge.CodeHash) ||
subtle.ConstantTimeCompare(candidateHash, challenge.CodeHash) != 1 {
return ErrInvalidCode
}
return store.Consume(ctx, id, now)
}
In production, keep load-and-consume in one transaction and increment failed verification counters independently. The caller should receive the same error for unknown, expired, consumed, and incorrect challenges. Short code spaces make online rate limiting mandatory even when hashes are protected at rest.
Test the incident before shipping the feature
A happy-path unit test proves very little here. Exercise the timeline: SMS submission times out after the downstream accepted it; the fallback request is retried; email submission is delayed; the SMS receipt arrives after email; then two verification requests race. The expected result is two logical channel attempts at most and exactly one consumed challenge. Also inspect the evidence an operator would have at each step. Before the receipt, the attempt must say “accepted,” not “delivered.” After the fallback retry, there must still be one email attempt ID. After both verification requests, one response may succeed, but the database must show one consumption timestamp and the other response must use the generic invalid-code result. This sequence catches the dangerous gap between a transport test and an authentication test.
Use a fake clock and scripted sender. Assert state, not sleep duration. Integration tests should also cover a missing receipt, a terminal rejection, callback replay, out-of-order callbacks, queue redelivery, process restart, expiry during verification, and destination changes in another session. For property marketplace access, add an authorization assertion after verification: passing the challenge establishes the seller session but does not by itself grant access to an order owned by another seller.
Deployment needs a kill switch for initiating new fallback sends while preserving verification of codes already issued. Roll out by a small cohort, watch unresolved-attempt age and verification completion, then expand. The runbook should let on-call answer four questions quickly: Are submissions succeeding? Are receipts advancing? Are fallback requests rising? Are sellers ultimately completing verification?
Instrument counters for attempts by channel and coarse outcome, histograms for submission latency and time to terminal observation, and gauges for the oldest due reconciliation item. Trace challenge_id and attempt_id across the application, outbox, sender, and callback handler. Keep account identifiers out of metric dimensions.
Alert on sustained user-impact signals, not one provider error. A practical alert combines an aging backlog with a rise in fallback demand or a drop in verification completion, evaluated over more than one scheduler interval. Document the exact query and a dashboard link in the runbook, then test the alert by pausing the fake transport in staging.
Close the loop without paging on noise
The finished trace is straightforward: a seller requests a login code, the SMS attempt becomes visible, the reconciler observes its progress, the UI offers an idempotent email action after a bounded wait, and one atomic verification consumes the shared challenge. On-call can locate the stalled stage without reading message content or guessing from an acceptance response.
Thresholds still carry a cost. Set the unresolved-age alert too aggressively and normal carrier latency pages the team; responders will learn to ignore it. Set it too loosely and sellers waiting to process a new order become the monitor. Start with observed baseline distributions, require enough affected challenges to avoid single-user noise, and pair delivery lag with completion impact. Revisit the threshold after changes to retry timing, queue cadence, or regional traffic.
That is the operational bargain: delayed transport evidence is tolerable; ambiguous authentication state is not.
Further reading
- OWASP Authentication Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Authentication_Cheat_Sheet.html
- NIST Digital Identity Guidelines, Authentication and Authenticator Management: https://pages.nist.gov/800-63-4/sp800-63b.html
- RFC 6376, DomainKeys Identified Mail (DKIM) Signatures: https://datatracker.ietf.org/doc/html/rfc6376
- RFC 7208, Sender Policy Framework (SPF): https://datatracker.ietf.org/doc/html/rfc7208
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)