Short answer: use bounded polling only to improve the login screen, never to decide whether an OTP is valid. The least complex defensible design records one immutable send attempt, polls a provider-neutral status adapter for a short deadline, and writes every observed transition to an audit log. Verification remains a separate, atomic operation. A terminal invalid-recipient result suppresses later sends; an unknown result ends the poll without pretending that delivery succeeded or failed.
The bill is made of message attempts, status lookups, stored evidence, and operational review. For a symbolic workload of N OTP sends, p polls per send, and r retained status rows, the dominant variable term is visible without a price sheet: polling creates N * p lookups, while evidence storage creates roughly N * r rows. A policy of six checks therefore creates up to six status reads per send. Reducing six checks to three halves that lookup count; shortening retention does not. Settle that arithmetic before arguing about databases.
This separation also fits the security boundary described by the OWASP Forgot Password Cheat Sheet: codes should be generated securely, linked to an individual user, invalidated after use, and protected by rate limiting. Delivery state is useful evidence and useful UX input. It is not proof that the person entering a code is entitled to the account.
How should polling SMS OTP delivery status work without webhooks?
It means a bounded sequence of stale snapshots. The first lookup may still report an accepted or queued state after the handset has received the message; a later lookup may expose a terminal failure. No polling interval removes that epistemic limit. The login page should say that a code was sent, offer a controlled new attempt after a server-defined delay, and avoid promises such as "delivered" unless the adapter has mapped a terminal delivery state while preserving its source evidence.
For an illustrative policy, poll after 1, 2, 4, 8, 12, and 18 seconds, then stop at 45 seconds. These are design parameters, not carrier guarantees. Add per-attempt jitter so a traffic spike does not produce synchronized status traffic, and cancel outstanding checks when a terminal state appears. The short deadline contains load and latency; the widening gaps spend fewer reads on states unlikely to change immediately.
The awkward case matters most: the poll deadline expires while the OTP is still valid. Keep those clocks independent. Expiring observation must not expire authentication, and an authentication timeout must not rewrite the delivery record. Delivery can be duplicated or remain uncertain; verification and new-attempt eligibility should therefore use compare-and-swap semantics around one attempt identifier.
Bounded polling has a clear limitation: it is unsuitable when the product requires immediate server-side reaction at high message volume, because every additional check multiplies read traffic while still observing a snapshot. An authenticated event callback or a queue-fed delivery event is the better mechanism in that case, provided signatures, replay protection, deduplication, and evidence persistence are implemented. Polling remains reasonable where callbacks are unavailable and a short-lived UX hint, rather than instantaneous automation, is the actual requirement. This is the central trade-off.
package otp
import (
"context"
"errors"
"time"
)
type State string
const (
Queued State = "queued"
Delivered State = "delivered"
InvalidRecipient State = "invalid_recipient"
Failed State = "failed"
Unknown State = "unknown"
)
type Snapshot struct {
AttemptID string
State State
SourceRef string
Observed time.Time
}
type StatusReader interface {
Read(context.Context, string) (Snapshot, error)
}
type EvidenceWriter interface {
Append(context.Context, Snapshot) error
}
func Observe(ctx context.Context, id string, reader StatusReader, log EvidenceWriter, waits []time.Duration) (State, error) {
for _, wait := range waits {
timer := time.NewTimer(wait)
select {
case <-ctx.Done():
timer.Stop()
return Unknown, ctx.Err()
case <-timer.C:
}
snap, err := reader.Read(ctx, id)
if err != nil {
continue // A transport error is not a delivery verdict.
}
if err := log.Append(ctx, snap); err != nil {
return Unknown, errors.New("status observed but evidence was not persisted")
}
switch snap.State {
case Delivered, InvalidRecipient, Failed:
return snap.State, nil
}
}
return Unknown, nil
}
The adapter owns status normalization. The ledger owns history. Keeping those responsibilities apart prevents a provider vocabulary change from silently changing suppression policy.
Suppression is a state machine, not a Boolean
A single blocked flag loses the reason, scope, and evidence behind a decision. Model the normalized recipient, channel, reason class, source attempt, observation time, policy version, and reversal record. Preserve the provider's opaque reference, but do not make business logic depend on it. The write that creates a terminal invalid-recipient observation and the write that activates suppression should share one transactional boundary or one idempotent event key.
| Observed condition | Login response | Suppression action | Evidence |
|---|---|---|---|
| Queued or accepted | Continue the OTP screen | None | Append the snapshot |
| Delivered | Continue verification | None | Append the terminal snapshot |
| Invalid recipient | Offer another recovery path | Suppress this channel and address | Record reason, source, and policy version |
| Generic failure | Permit policy-controlled retry | Do not infer invalidity | Record the failure class |
| Poll deadline reached | Keep verification independent | No suppression | Record observation timeout |
Do not collapse a transient lookup error into Failed. Do not suppress on Unknown. Those shortcuts make dashboards look decisive while corrupting the recipient registry, and the resulting false suppression is hard to distinguish from a legitimate compliance control months later.
Email bounces may enter the same suppression ledger when email is a recovery channel, but they should retain their own reason taxonomy. The FTC's CAN-SPAM compliance guide applies to commercial email and describes obligations including a functioning opt-out mechanism; it should not be stretched into an SMS OTP rule. Shared evidence storage is reasonable. Shared legal semantics are not.
Build the audit trail around decisions
An audit record should answer four questions without replaying application logs: what was observed, which rule evaluated it, what changed, and which actor or process authorized that change. Use append-only observations plus a separately versioned current projection. Corrections then become new records rather than edits that erase history.
I would make the idempotency key (attempt_id, normalized_state, source_sequence) when the upstream system supplies a stable sequence; otherwise, use a canonical hash of stable source fields and retain the raw source reference for investigation. This is a design choice, not a universal recipe. A lossy status vocabulary can map several upstream states to queued, so the evidence record should carry both the normalized and original values even though policy reads only the normalized one.
Test the failure boundaries, not merely the happy path. Replay the same terminal snapshot twice and assert one suppression decision. Deliver observations out of order and assert that a late nonterminal state cannot reverse a terminal one. Simulate a timeout between evidence append and projection update, then verify that replay repairs the projection without creating a second decision. Finally, race two correct OTP submissions and require one successful consumption.
Five checks expose the important mistakes:
- Duplicate observations produce one effective transition.
- Unknown states never suppress a recipient.
- A repeated send creates a new attempt while preserving the earlier trail.
- Logs and metrics contain attempt identifiers but no OTP values.
- Retention deletion removes payload detail without erasing the minimum decision provenance required by policy.
The fifth check is where compliance evidence and cost collide. Keep compact decision records longer than verbose polling payloads if policy permits, and document the two retention classes.
Retention is an explicit loss budget
Storage rarely dominates while raw payloads are small, yet indefinite retention expands privacy exposure and turns every schema change into a migration obligation. A practical ledger separates short-lived source payloads from longer-lived normalized decisions, with durations set by counsel, regulation, contractual duties, and investigation needs rather than copied from an unrelated system. US and EU deployments may require different policy configurations; the evidence model can remain common while approved durations and access controls differ.
What should be deleted first? Drop repetitive nonterminal polling bodies after their short investigation window, retaining the normalized transition, timestamp, source reference, and policy version for the approved evidence period. Stop keeping OTP secrets in observability data at all; OWASP's guidance that codes are stored securely, single-use, and expire makes broad log retention the wrong place for them.
There is a real cost to this restraint. If a carrier or provider dispute arrives after raw payload deletion, investigators may be able to prove the decision taken but not reconstruct every upstream field that preceded it. I accept that loss only when the retention policy names it, the deletion job is itself auditable, and a legal or incident hold can suspend deletion through an authorized process.
Keep less, deliberately.
Further reading
- OWASP, "Forgot Password Cheat Sheet": https://cheatsheetseries.owasp.org/cheatsheets/Forgot_Password_Cheat_Sheet.html
- US Federal Trade Commission, "CAN-SPAM Act: A Compliance Guide for Business": https://www.ftc.gov/business-guidance/resources/can-spam-act-compliance-guide-business
Top comments (0)