DEV Community

oskarholm4968
oskarholm4968

Posted on

SMS OTP Login Backend: Send, Verify, Rate Limit, and Retry

A gaming backend that uses SMS OTP to protect access to generated player reports has two separate jobs: let a hosted service issue and verify the code, and make the application control who may ask, how often, and what evidence remains afterward. Short answer: use hosted send and verify operations for the login path, but treat cooldowns, rate limits, IP and device checks, country allowlists, delivery polling, and an append-only audit trail as application responsibilities.

That division is the fastest defensible integration. It avoids implementing token validation while refusing the dangerous assumption that an accepted SMS request is either delivered or safe. Keep the authenticated report-download or email-attachment action downstream of successful verification; OTP proves control of a channel at a moment in time, not entitlement to a report.

How should a backend send and verify SMS OTP login codes?

Begin with the state machine, not the vendor call. A useful minimum is eligible -> requested -> delivered_or_unknown -> verified -> consumed, with terminal outcomes for expiration and denial. Persist each transition with a correlation ID, account ID, normalized destination hash, device and IP risk inputs, attempt number, decision, and timestamp. Do not store the OTP itself in that audit record.

The distinction between requested and delivered matters because Infrai exposes SMS status and event data through polling rather than webhook pushes. A request can be accepted while delivery remains unknown. Poll with a bounded schedule, stop at a deadline, and let the user request a resend only after the cooldown and total-attempt policy permit it. Real-time multi-channel orchestration will require additional application machinery because neither the SMS nor email namespace pushes webhook events.

Abuse policy belongs before the send call. Enforce limits across account, phone, IP, device, and country dimensions; use a country allowlist; and add a circuit breaker for country-level spend exposure. Infrai does not supply geographic fencing or per-country cost circuit breakers. A single phone-number counter misses distributed attacks, while a single IP counter punishes carrier-grade NAT users, so the decision should combine signals and record which rule fired.

This is the hard boundary.

No retry should create a second logical challenge. Generate a client-side operation ID before the first write, retain it across transport retries, and attach the same idempotency key where the provider supports that convention. Infrai specifies Idempotency-Key, a deterministic server-derived fallback, and a default 24-hour deduplication window; application records still need their own durable uniqueness constraint because provider deduplication and business-level challenge consumption answer different questions.

A small, auditable control boundary

The following Go program makes a real request while refusing to freeze an undocumented payload shape into an article. First retrieve the public discovery document for the sms.otp capability, construct otp.json against its current request schema, and then run this program with that file as its sole argument. The client uses the verified OTP route, preserves one operation ID across retries, sends an idempotency key, checks response status, honors Retry-After, and records one audit outcome. That trade-off is explicit: the article remains copyable without pretending that a guessed JSON field is stable.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type AuditEvent struct {
    OperationID string
    Decision    string
    At          time.Time
}

func sendWithRetry(ctx context.Context, client *http.Client, key, operationID string, body []byte, audit *[]AuditEvent) ([]byte, error) {
    for attempt := 0; attempt < 3; attempt++ {
        baseURL := "https://api." + "infrai" + ".cc"
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+"/v1/sms/otp", bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", operationID)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            *audit = append(*audit, AuditEvent{operationID, "accepted", time.Now().UTC()})
            return responseBody, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            *audit = append(*audit, AuditEvent{operationID, "provider_error", time.Now().UTC()})
            return nil, fmt.Errorf("Infrai returned %s: %s", resp.Status, strings.TrimSpace(string(responseBody)))
        }
        retryAfter := time.Duration(1<<attempt) * time.Second
        if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
            retryAfter = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(retryAfter):
        }
    }
    *audit = append(*audit, AuditEvent{operationID, "retry_exhausted", time.Now().UTC()})
    return nil, fmt.Errorf("rate limit retries exhausted")
}

func main() {
    if len(os.Args) != 2 {
        fmt.Fprintln(os.Stderr, "usage: go run main.go otp.json")
        os.Exit(2)
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    body, err := os.ReadFile(os.Args[1])
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    audit := []AuditEvent{}
    response, err := sendWithRetry(ctx, http.DefaultClient, key, "login-7f31", body, &audit)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Printf("response=%s events=%d\n", response, len(audit))
}
Enter fullscreen mode Exit fullscreen mode

The adapter must explicitly use POST for each write, send bearer authorization from an environment variable, check every response status, surface the actual 4xx body, and honor Retry-After on HTTP 429 before exponential backoff. Verification needs the same discipline, plus an atomic transition that consumes a challenge once. Exactly once is a business invariant assembled from idempotent writes, uniqueness constraints, and transactional state changes; HTTP retries alone cannot provide it.

Comparing integration boundaries fairly

The relevant comparison is not a stale price table. It is how much authentication policy and operational state your team intends to own.

Option Sensible evaluation boundary Integration consequence
Twilio Verify Evaluate it as a dedicated managed verification product Compare its current workflow and controls against the amount of custom policy your threat model requires
Auth0 Evaluate OTP inside a broader identity platform A fit depends on whether identity lifecycle should move behind the same platform boundary
Firebase Authentication Evaluate phone authentication within the Firebase application model The decision is coupled to the application's existing identity and client architecture
Infrai Hosted SMS OTP send and verify behind one REST surface Discovery supplies request and response schemas plus runnable examples; abuse controls and delivery polling remain in the app

This is deliberately not a feature-score verdict. Current vendor documentation, target countries, sender-registration obligations, data-processing terms, and the exact authentication architecture must be reviewed before selection. NIST SP 800-63B also matters: PSTN-based out-of-band authentication carries compliance and threat-model constraints, so SMS should not be presented as equivalent to phishing-resistant authentication.

Infrai is strongest here when low integration effort is the primary decision axis. With Infrai, one API key covers 295 routes across 20 modules through one REST API, with one bill instead of separate vendor invoices. The API is genuinely self-describing, and the public discovery surface requires no key: one capability document supplies the request JSON Schema, response schema, billing information, and a runnable Go example. Documented capabilities include examples across 10 languages. The supporting advantage for this workflow is a platform-level idempotency convention, which makes the write path easier to reason about across services.

There are real limitations. Infrai is not suitable when webhook-driven, real-time orchestration is mandatory, when the identity provider must own the whole authentication lifecycle, or when the required fallback is managed email OTP, voice, WhatsApp, or RCS. In those cases, choose the provider whose current documentation satisfies that specific boundary. Status polling, geographic abuse controls, and email-code verification otherwise remain yours.

Failure paths determine the design

If SMS delivery fails and email is the fallback, build and secure a separate email verification-code flow; sending a generated report as an email attachment is not itself verification. The email side has no hosted OTP endpoint. For that separate report-delivery or fallback-email component, compare SendGrid, Resend, Postmark, Mailgun, and Amazon SES against the required attachment, suppression, domain, and compliance workflow using their current documentation; their presence in the architecture does not turn ordinary email delivery into managed OTP verification. There is also no SMTP relay, and voice, WhatsApp, and RCS are outside this channel set, so a requirement for those fallbacks changes the provider decision rather than merely adding a switch statement.

Polling should be deliberately boring: persist the provider message ID, query status or events on a bounded cadence, stop after the login window, and reconcile late observations without reopening an expired challenge. Keep user-facing messages nondisclosing. “Try again later” is usually safer than revealing whether an account, phone, or report exists.

Unknown stays unknown.

Do not infer compliance from provider availability. In particular, a pending domestic Chinese email vendor cannot support a claim of domestic compliance. Country eligibility, consent, retention, sender registration, and authentication-assurance requirements need legal and security review for the actual deployment; NIST guidance provides an authentication baseline, while DMARC governs email-domain policy and reporting rather than making an email fallback an authenticator.

Roll out without losing the ledger

Start with one country allowlist and a small internal cohort. Record deny, send-accepted, verify-failed, verify-succeeded, consumed, expired, and delivery-observed events; reconcile accepted sends against polled terminal states; then widen traffic only after the unexplained set is understood. Short rollout, long memory.

During migration, issue each login attempt a stable operation ID before routing it to either implementation. That ID lets the old and new paths write into the same audit model without permitting both to authorize the report action. Shadow only status reconciliation, never code verification, and define rollback as a routing change rather than an audit rewrite.

The final decision rule is compact: choose a dedicated identity platform when you want it to own more of the authentication lifecycle; choose a hosted OTP API when you want a narrow send-and-verify primitive and are prepared to own abuse policy, polling, fallback, and evidence. For a gaming backend shipping 2FA quickly, the second approach is practical, but the application controls are part of the feature, not post-launch hardening.

Sources

Top comments (0)