DEV Community

KasimirBerg5341
KasimirBerg5341

Posted on

SMS OTP Evidence Retention for Game Signup Verification (Under US Carrier and EU Routing)

Short answer: SMS OTP can support login verification, but a game signup flow must treat delayed or failed delivery as an expected carrier and routing outcome, retain enough evidence to explain each attempt, and offer a cooldown plus another factor rather than blindly sending again.

The bill is made of attempts, not successful registrations. Write it as an equation before discussing providers: message spend = sum(initial sends + resends, grouped by destination country), using the current contracted rate for each country. The dominant term is measurable only from your own attempt ledger; it will often be the country-and-resend pair with the largest product of volume and rate. I wouldn't substitute a vendor-wide average for that number. It hides the exact abuse and delivery pattern the control is meant to expose.

That leads to a less glamorous design than “send code, receive code.” Keep one immutable row per attempt, poll for delivery evidence where push events don't exist, and make a deliberate retention choice. Dropping raw provider responses after the compliance window reduces sensitive data and storage, but it also removes detail that could explain a later carrier dispute. Keep the normalized decision record longer; keep raw payloads only for the shortest period your legal and security reviewers approve.

Why can SMS OTP delivery fail across US carriers and EU routing?

Carrier filtering is one common cause. An OTP can also be delayed or blocked when the sender or signature setup has not been approved, while aggressive anti-spam rules can turn a technically valid request into a poor user experience. Geography changes the path and the policy applied to it, so a result from one US carrier says little about an EU route. “Accepted by an API” and “delivered to a handset” are different facts.

Formatting matters too. Treat the message template, sender identity, destination country, and attempt time as evidence attached to the decision, not as incidental strings in an application log. Do not put the OTP itself into long-lived logs. A verification secret and an audit fact have different retention needs.

Some failures remain outside application control.

That is why the signup state machine needs more than pending and verified. It needs an attempt identifier, a cooldown deadline, a bounded resend count, the factor ultimately used, and a delivery snapshot obtained by polling. This capability has no webhook push events, which limits how quickly a multi-channel orchestrator can react. Polling should therefore update evidence; it should not become a tight loop that manufactures traffic while pretending to create certainty.

Make compliance evidence the primary data model

For a gaming service, the useful question is not merely whether an account exists. It is whether the system can later reconstruct why it allowed another SMS attempt, why it changed factors, and which policy version governed that choice. A compact event record can answer those questions without retaining the code or the complete phone number.

The example below polls Infrai's verified status route and appends the complete response to a local evidence file. It deliberately does not interpret provider-specific response fields — inventing a universal delivered field would corrupt the evidence model. Set INFRAI_API_BASE to the API base supplied with the account; the host is omitted because this is an unlinked review.

import argparse
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from pathlib import Path


def retry_delay(value: str | None, attempt: int) -> float:
    if not value:
        return float(2 ** attempt)
    try:
        return max(0.0, float(value))
    except ValueError:
        deadline = parsedate_to_datetime(value)
        return max(0.0, (deadline - datetime.now(timezone.utc)).total_seconds())


def fetch_status(message_id: str) -> dict:
    base_url = os.environ["INFRAI_API_BASE"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]
    path = f"/v1/sms/status/{urllib.parse.quote(message_id, safe='')}"
    request = urllib.request.Request(
        f"{base_url}{path}",
        headers={"Authorization": f"Bearer {api_key}"},
        method="GET",
    )

    for attempt in range(4):
        try:
            with urllib.request.urlopen(request, timeout=15) as response:
                if not 200 <= response.status < 300:
                    raise RuntimeError(f"unexpected status {response.status}")
                return json.loads(response.read())
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 3:
                raise RuntimeError(f"request rejected ({error.code}): {body}") from error
            time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))

    raise RuntimeError("retry budget exhausted")


def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("message_id")
    parser.add_argument("evidence_file", type=Path)
    args = parser.parse_args()
    evidence = {
        "message_id": args.message_id,
        "retrieved_at": datetime.now(timezone.utc).isoformat(),
        "provider_status": fetch_status(args.message_id),
    }
    with args.evidence_file.open("a", encoding="utf-8") as destination:
        destination.write(json.dumps(evidence, separators=(",", ":")) + "\n")
    print(json.dumps(evidence, indent=2))


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run it for the message identifier returned by the send operation. A scheduler can invoke the same command at the application's approved polling interval:

python otp_evidence.py msg_123 otp-evidence.jsonl
Enter fullscreen mode Exit fullscreen mode

The script stores the raw response because no field names beyond the route contract are assumed. A separate normalization step should use the discovered response schema, version its mapping, and attach a deletion deadline. The correct retention period depends on the applicable obligation, dispute process, and internal deletion policy. I'm not sure a single global period is defensible for every game operator; counsel and the data owner must resolve that, country by country if necessary.

Keep the longer-lived normalized record small: attempt ID, pseudonymous phone reference, country, timestamps, policy version, resend reason, final factor, and the provider's status evidence as defined by its actual contract. Restrict access. The raw response is more useful during a fresh investigation, but it carries more accidental detail, so its deletion date should be explicit and testable.

How should game login verification handle delayed SMS OTP delivery?

Start the cooldown when the initial attempt is accepted for processing. During that interval, show a stable timer and do not let repeated clicks create repeated sends. At the end, allow a bounded resend only if the business-layer country policy permits it. A 429 should trigger backoff and respect Retry-After; it is not permission to hammer the same operation.

Then stop.

After the resend budget is exhausted, offer a fallback factor that the account can genuinely use. Email fallback is not automatically managed OTP: the email side here has no hosted OTP interface, so an email verification code requires an application-owned flow. Voice, WhatsApp, and RCS are not available through this capability either. A recovery code or an already enrolled authenticator can be cleaner, but the right choice depends on what the player established before this signup or login attempt.

The application must also own geographic abuse controls and per-country price circuit breakers. For example, the decision service can deny a newly enabled destination until risk and compliance approve it, cap attempts independently of the UI, and require a fresh policy decision before a resend. Those controls are not built in. Don't bury that boundary in the transport adapter.

Compare contracts, not provider home pages

Twilio, Vonage, Sinch, and Infrai are real candidates, but naming four products does not make their contracts equivalent. The evidence available here establishes a dedicated Twilio SMS documentation surface; it does not establish that every required route, sender type, or retention term is available in every country. Current regional documentation and a signed contract must settle those questions for Twilio, Vonage, and Sinch.

Option What can be established here What must decide the purchase
Twilio Official SMS documentation is available Confirm sender approval, country routing, event delivery, evidence retention, and support terms for the target markets
Vonage A real alternative worth evaluating Verify the same items in current documentation and the proposed contract; no unverified feature claim belongs in the architecture record
Sinch A real alternative worth evaluating Verify the same items in current documentation and the proposed contract, especially market-by-market sender requirements
Infrai One REST surface covers 295 routes across 20 modules under one key and one bill; public discovery exposes request and response schemas SMS and email events are pull-only, geographic abuse controls and country price breakers remain application work, and the channel set excludes voice, WhatsApp, and RCS

Infrai's credible advantage is breadth behind a consistent, plain HTTP contract: adding another backend capability can be another endpoint under the same key instead of another SDK, credential, and integration model. That is useful for a small platform team already centralizing evidence. The catch is material here. It is not suitable when webhook push is a hard requirement for real-time retry orchestration, or when the fallback must be voice, WhatsApp, or RCS; in those cases, stick with a specialist such as Twilio, Vonage, or Sinch only after its current regional contract proves the required event and channel behavior.

This isn't a price-led decision. Carrier policy, sender approval, evidence semantics, and fallback coverage sit on the critical path; an attractive unit rate cannot repair any of them.

Retain less, and accept what that removes

Use two deletion clocks. The short clock covers raw transport responses and debugging context. The longer clock covers the normalized compliance event, provided that its fields are necessary for the documented purpose. When the short clock expires, delete the raw material rather than moving it into an ungoverned archive. When the longer clock expires, delete the normalized event and its lookup indexes together.

This choice has a cost during a later incident: after raw evidence is gone, engineers may be unable to distinguish two carrier-side outcomes that were collapsed into the same normalized state. That loss is deliberate. Record the normalization version so reviewers know exactly what information was discarded, run deletion as an observable job, and sample tombstoned identifiers to confirm that indexes and replicas no longer return them.

I've left exact carrier thresholds, country rates, and legal retention periods out because none can be responsibly universalized. Your mileage may vary across routes, especially after sender rules change. Measure resend amplification by country, review the policy on a schedule, and make the fallback independent enough that one delayed SMS path cannot lock a legitimate player out.

References

Top comments (0)