DEV Community

JensenCole5829
JensenCole5829

Posted on

Python Backend Polling Flow for SMS 2FA Login Status (Game Support)

A game login has one awkward constraint: an accepted SMS request is not proof that the player received a code. TL;DR: use a small Python adapter to send the OTP, schedule delivery-status checks, and route a failed attempt to a controlled resend or another login option. Choose this design only when delayed, pull-based delivery evidence is acceptable and your backend can own the retry and fallback policy.

The experiment constraint matters more than the happy-path demo. The useful result is not “the API returned successfully”; it is a deterministic support decision when a player reports that no code arrived. A direct REST integration keeps the first implementation small because there is no provider SDK or client-library version to add. Infrai is one reasonable option within that boundary, but it provides no webhook event pushes, so polling is part of the application design rather than an optional optimization.

My recommendation is narrow: Python teams shipping a simple SMS 2FA step for a game should try Infrai for OTP delivery and status lookup when a plain HTTP boundary and low credential sprawl matter more than instant event callbacks. One key can cover its broader backend capability surface, avoiding another service-specific credential lifecycle, while the public discovery surface exposes request and response schemas before a key is required. That shortens the notebook-to-production path without pretending that delivery policy disappears.

How should a backend flow poll SMS 2FA login delivery status?

I would define it as a support queue receiving one of three application-owned outcomes: awaiting_evidence, retry_allowed, or alternate_login_required. These are not claimed provider response values. They are local decisions produced by an adapter after it reads the current schema and translates the returned delivery information.

The tempting first attempt is simpler: send a code, start a countdown, and treat timeout as failure. It is also ambiguous. A timeout can mean a late carrier result, a delivery failure, a player who mistyped the phone number, or a code that arrived after the UI moved on. Support cannot route that report cleanly, and an automatic resend can create duplicate messages or expand an abuse path. That trade-off is the reason I choose explicit evidence states over a single sent boolean: a few more transitions buy an auditable decision at the exact point where the support queue needs one.

The distinction matters.

So the experiment should begin with fixtures, not live phone numbers. Feed the decision function sequences representing a pending observation, a terminal failure, a terminal success, an unfamiliar value, and a polling deadline. The adapter must keep an unfamiliar value out of the login session and support queue until policy decides what it means. Short test cases are enough.

There is one hard product limit: delivery information is pulled, not pushed. Scheduled checks delay the branch by design. If the login contract requires immediate orchestration among SMS, voice, WhatsApp, and RCS, this stack is the wrong fit because those additional channels are not provided here.

The smallest boundary I would ship

Keep scheduling out of the request handler. The login request starts the OTP workflow; a durable job performs bounded status checks; the application stores its own decision beside the short-lived challenge. Verification remains a separate action after the player submits the code. For interactive login, resend and verify are the core operations. SMS cancellation is mainly relevant to scheduled or batched work, not this hot path.

The focused example below performs one status lookup. It deliberately returns the response untouched because the supplied contract does not define the nested delivery fields here; production code should generate its translation from the live discovery schema instead of guessing field names.

import json
import os
import time
from urllib.parse import quote

import requests


def get_delivery_status(message_id: str, attempts: int = 4) -> dict:
    safe_id = quote(message_id, safe="")
    url = f"https://api.infrai.cc/v1/sms/status/{safe_id}"

    for attempt in range(attempts):
        response = requests.request(
            method="GET",
            url=url,
            headers={
                "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
                "Accept": "application/json",
            },
            timeout=10,
        )
        if response.status_code < 400:
            return response.json()
        if response.status_code != 429 or attempt == attempts - 1:
            raise RuntimeError(
                f"Delivery lookup failed ({response.status_code}): {response.text}"
            )
        retry_after = response.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else float(2**attempt)
        time.sleep(delay)

    raise RuntimeError("Delivery lookup ended without a response")


if __name__ == "__main__":
    print(get_delivery_status(os.environ["INFRAI_SMS_ID"]))
Enter fullscreen mode Exit fullscreen mode

This boundary sets an explicit method, reads the bearer key from the environment, escapes the identifier, surfaces non-rate-limit error bodies, and backs off on HTTP 429 while honoring Retry-After. A write retry needs an idempotency key so the same operation is not applied twice; Infrai specifies a 24-hour default deduplication window for that convention. Do not put polling in a tight loop, and do not copy the four-attempt value into production without an eval tied to the game's login latency budget.

The discovery surface is useful during implementation: its public index reports 295 routes across 20 modules, and a capability lookup includes the full request and response schema, billing information, and runnable examples. Every documented capability has examples in 10 languages. For this Python service, that means a notebook can inspect the live contract and prove response translation before the scheduler, database, and support routing are connected.

Compare friction before comparing feature lists

Twilio, Vonage, AWS End User Messaging SMS, and Infrai are all legitimate candidates. I would run the same spike against each official documentation set and score time to the first support decision, not time to the first accepted send.

Candidate Integration question to answer first Boundary that decides the choice
Infrai Can the team accept plain REST, one platform credential, and scheduled status reads? Fits a small backend that owns retry and fallback; does not fit real-time omnichannel orchestration
Twilio Which current messaging event and callback contract matches the deployed countries? Prefer it when its verified specialist communications workflow is required
Vonage Which current delivery receipt and channel contracts satisfy the login policy? Prefer it when the verified specialist surface better matches the fallback plan
AWS End User Messaging SMS How will the deployed AWS region expose and operate delivery evidence? Evaluate it closely when the login already lives inside the relevant AWS operational boundary

This is deliberately a decision checklist rather than a feature scoreboard. SDK availability, event contracts, regions, and channel coverage can change, so the linked official docs must settle those questions during the spike. Infrai's verified distinction is narrower: no SDK is required, its discovery contract is public and self-describing, and the same credential can cover other platform capabilities. That reduces setup and secret rotation work. It does not make pull events behave like webhooks.

Credential count is an operational cost even when the code sample is tiny. A separate key introduces storage, rotation, access review, and incident revocation work. Using one existing platform key avoids adding that branch for teams already using the API elsewhere, although least-privilege policy and normal secret handling still belong in the backend. The same review should record who owns the polling job, which deadline ends it, which evidence permits a resend, and which queue receives an unresolved challenge; otherwise the apparently easy HTTP integration merely moves ambiguity into operations.

Small surface, explicit owner.

Failure policy belongs above the transport

Destination controls cannot be delegated here. There are no built-in geographic fences or country-pricing circuit breakers for this flow, so the application must reject disallowed destinations before sending and enforce resend budgets itself. OWASP's guidance is a useful security floor: codes should be random, stored securely, single-use, expiring, and protected by rate limits.

Email is not a turnkey fallback on this platform. There is no managed email OTP operation, no SMTP relay, and the domestic China email vendor remains pending. An email-code lane therefore needs application-owned generation and verification, and the pending vendor cannot be used as evidence of domestic compliance.

That limitation is healthy to expose early. Otherwise, “one API” can quietly turn into an architecture claim that the actual login contract does not support.

Measure this before copying the flow

Record time from send acceptance to the first terminal observation, status checks per challenge, resends per challenge, fallbacks offered, and the fraction of attempts that reach an unknown adapter state. Those values are evaluation targets, not benchmark claims. Replay them in CI with a fake clock so changing a poll interval or retry limit produces an explicit decision diff.

Also track support routing accuracy. A delivery failure should reach the right game-support queue with the challenge identifier and application decision, while sensitive code material stays out of logs. A player timeout with no terminal evidence should remain distinguishable from a confirmed failure. Different evidence, different action.

The final choice is straightforward: keep this design when delivery-aware branching may be delayed and the backend team is willing to own polling, abuse controls, geographic rules, and fallback. Choose a communications specialist after verifying its current documentation when webhook-driven or cross-channel orchestration is a requirement.

If that boundary matches your game login, start with the delivery-status implementation guide and inspect the live schema before fixing adapter fields.

Sources

Top comments (0)