DEV Community

YvesSterling6854
YvesSterling6854

Posted on

Python SMS OTP Delivery Status: Polling Without Webhooks for 2FA Login

Short answer: poll SMS delivery only for a short, bounded diagnostic window, while letting OTP submission drive the login state. A delivery receipt is transport evidence, not proof that the person received the message, and waiting for one can turn an ordinary carrier delay into a frozen 2FA screen. For a developer tool that creates a report and emails it as an attachment after login, the application should own the verification UX and template contract; the messaging layer should report transport state without deciding access.

That split was the deciding constraint in my notebook-to-production evaluation. The tempting design was one loop that waited for delivered, unlocked the code field, and then enabled report generation. It looked wonderfully linear in a notebook. Under delayed, missing, or terminal receipts, though, it coupled three different events: message acceptance, carrier reporting, and proof of possession. The chosen design keeps them separate and gives polling a deadline.

Should SMS OTP delivery status use polling without webhooks?

The user can type a valid code before the latest delivery state arrives. Conversely, a delivered state cannot establish that the intended person has the phone. The verification attempt is therefore the only event that should authenticate the session. Polling can still improve the interface: it can replace “sending” with “sent,” surface a likely failure, or decide when offering another code becomes reasonable. It must not become a second authentication factor by accident.

OWASP's forgot-password guidance supplies the security shape that also matters here: responses should avoid revealing whether an account exists, timing should remain consistent, side-channel messages should be rate limited, tokens should be random, single-use, securely stored, and expire after an appropriate period. Those constraints belong in the OTP service regardless of how transport status is obtained.

Keep the public response boring. “If the account can receive a code, one has been sent” leaks less than exposing unknown_number, blocked, or a carrier-specific reason. Detailed transport evidence belongs in restricted telemetry, joined by an opaque attempt ID rather than displayed as an account oracle.

No receipt is definitive.

The architectural choice still has real trade-offs:

Status path Best fit Main limitation
Bounded polling A client already has a short-lived login session and callback infrastructure is unavailable Adds read traffic and can only approximate real-time UX
Webhook ingestion The backend can expose, authenticate, deduplicate, and monitor callbacks Requires public callback operations and careful handling of retries and reordering
No delivery tracking Verification success is enough and transport detail would not change the UX Gives support and routing diagnostics less evidence

Polling is a poor fit when large login bursts make repeated reads expensive, when the provider's status is too coarse to change an action, or when the client may disappear immediately after requesting a code. Webhooks are a better transport-observation mechanism when the team can operate them and needs asynchronous evidence after the browser closes. Neither choice verifies the user.

A small state machine beats an endless loop

Use two state machines. The authentication state can move from challenge_created to verified, expired, or locked. The transport observation can move among pending, accepted, delivered, failed, and unknown, according to the normalized vocabulary your adapter supports. These names are an application contract, not a claim that every carrier or provider emits each state.

The browser should ask your backend for a normalized snapshot. It should never receive messaging credentials, and it should stop asking after a fixed deadline. Poll slowly enough that a login surge does not multiply into a status-query surge. Add jitter so many clients created on the same second do not synchronize.

Here is the focused Python policy I would put under an eval harness before wiring it to a web route:

from dataclasses import dataclass
from enum import Enum


class TransportState(str, Enum):
    PENDING = "pending"
    ACCEPTED = "accepted"
    DELIVERED = "delivered"
    FAILED = "failed"
    UNKNOWN = "unknown"


@dataclass(frozen=True)
class PollDecision:
    poll_again: bool
    delay_seconds: float | None
    show_code_entry: bool
    offer_another_code: bool


def decide_poll(
    state: TransportState, elapsed_seconds: float, attempt: int
) -> PollDecision:
    terminal = state in {TransportState.DELIVERED, TransportState.FAILED}
    deadline_reached = elapsed_seconds >= 45.0
    poll_again = not terminal and not deadline_reached
    delay = min(2.0 * (1.5**attempt), 8.0) if poll_again else None

    return PollDecision(
        poll_again=poll_again,
        delay_seconds=delay,
        show_code_entry=True,
        offer_another_code=state is TransportState.FAILED or deadline_reached,
    )
Enter fullscreen mode Exit fullscreen mode

Notice the intentionally unconditional show_code_entry=True. The code box does not wait on transport telemetry. The 45-second observation window and 2-to-8-second backoff are example policy values, not universal recommendations; measure them against real receipt latency and query load before adopting them. Short-lived random jitter should be added by the caller, where a seeded generator can make tests deterministic.

Also make a request for another code create a new challenge or follow a clearly tested rule for invalidating the old one. Never let repeated clicks create an unbounded stream of messages. The security limit should be enforced server-side, even if the button is disabled in the browser.

Template ownership is an authentication boundary

Template ownership sounds like a copywriting decision, but it determines who can change security semantics. The application team should own the meaning of the message: what action the code authorizes, how long the challenge remains useful, and which text must stay consistent with the login screen. A transport adapter may render a provider-specific envelope, yet it should receive a versioned template identifier and constrained variables rather than arbitrary prompt output.

Generated prose should never invent an OTP, expiration, destination, or support URL. Keep those values outside the model boundary. In the report workflow, the same rule extends to the post-login email: the report generator owns the attachment bytes and filename, while the email template owns the surrounding transactional copy. Authentication success authorizes report generation; SMS delivery status does not.

This is where prompt-cost awareness helps. There is little value in spending model tokens on a six-digit-code message whose correctness depends on stable wording. Generate the report content where generation is the feature, then use deterministic templates for the OTP and attachment email. It is easier to evaluate, cache, localize, and audit.

US and EU rollout needs separate evidence

Do not turn a country field into a legal conclusion. US and EU destinations can differ in consent records, sender registration, routing, filtering, data retention, and the status detail available from intermediaries. Those differences deserve configuration and review, but a receipt still has the same architectural limit: it reports transport progress, not human possession.

Classify every message before launch. An OTP is security messaging; the emailed report is transactional when the user requested it. Marketing content mixed into either message changes the compliance analysis. The FTC's CAN-SPAM guide says commercial email is subject to requirements including accurate header information, non-deceptive subject lines, identification as an advertisement, a physical postal address, and an opt-out mechanism. It also distinguishes transactional or relationship content, while warning that mixed-content messages require evaluating the primary purpose. That is directly relevant to the report attachment email, not a shortcut for SMS rules or EU requirements.

Get region-specific legal review for the actual sender, recipient relationship, message purpose, and routing setup. Technically, store the policy version used for each challenge and email. Operationally, avoid logging OTP values or full report attachments; log identifiers, normalized outcomes, timings, and template versions.

What should the experiment measure?

Before copying the polling policy, replay delayed, missing, contradictory, and terminal transport observations. The key product measure is time from challenge creation to successful verification, segmented by region and route. Pair it with receipt latency, verification success, repeat-code frequency, poll requests per challenge, rate-limit events, and abandonment. For the report workflow, also measure time from verified session to attachment handoff and final email outcome as separate spans.

The eval should assert behavior, not merely response text. A valid OTP must be accepted while transport remains pending. A delivered receipt must never authenticate. A failed receipt may expose a neutral recovery action without confirming account existence. The loop must stop at its deadline, and duplicate browser polls must not send duplicate messages.

Start with a tiny matrix: early verification, late receipt; terminal failure, then successful verification; no receipt at all; another-code request under the server limit; and an expired code arriving after a newer challenge. Add region and template version as dimensions only when they reveal a real policy branch. Otherwise the matrix grows while confidence does not.

The production choice is bounded observation plus verification-led state. Keep template semantics with the application, keep transport detail behind a normalized adapter, and validate the timing values using your own traffic before presenting polling as real-time UX.

References

Top comments (0)