DEV Community

UlricDonovan1564
UlricDonovan1564

Posted on

OTP Login Providers: SMS Polling Status Without Webhooks for Gaming Verification

An OTP login provider without webhooks makes SMS verification status a polling concern, but it should not make carrier state part of your security decision. For a short-expiry gaming password reset, design for delayed and out-of-order evidence first: the authentication service must own expiry, attempt limits, resend cooldowns, and an append-only decision trail.

TL;DR: Polling-only SMS OTP is enough for a straightforward gaming account reset or 2FA flow. Keep the provider behind a small contract so changing the delivery vendor does not change the reset state machine. Poll message status or the verification result, make every local transition idempotent, and treat resend as a controlled state transition rather than a button that sends on demand. Pick a broader verification product when voice, WhatsApp, RCS, or real-time omnichannel failover is a requirement.

That separation gives compliance reviewers evidence about what the application decided even when carrier delivery evidence is delayed. It also prevents a provider callback, or the absence of one, from becoming the authority on whether a code is still valid.

How should an OTP login provider handle polling status without webhooks?

The bounded failure scenario I use in a runbook is mundane. A player requests a password-reset code with a 5-minute application expiry. The SMS remains queued long enough for the player to press resend, then both messages reach the handset out of order. The first code must not become valid again merely because its delivery status changed late.

Late means late.

That is the invariant: delivery state can inform the reset flow, but it cannot extend or recreate authorization state. The auth service records a challenge ID, an expiry instant, a hash of the active code, a resend count, a next-resend instant, and failed verification attempts. A resend rotates or supersedes the active challenge under one transaction. Replaying the same request key returns the same transition instead of producing another message.

I first assumed a sequence diagram would settle this review; it did not. A diagram makes the happy path look tidy while hiding the race that matters: request A can reach a handset after resend B has already superseded it. The useful artifact is an event ledger instead: challenge_created, delivery_requested, resend_accepted, challenge_superseded, verification_failed, and challenge_consumed, each with an actor, request ID, policy version, and timestamp. Those are application facts. Provider status samples belong beside them as observations, never commands, so an auditor can follow the decision even if the final carrier update arrived after expiry.

No callback is required for correctness. A worker can poll while a challenge is live, with jitter and a firm deadline, then stop. The login request should never block until the carrier reports a terminal delivery state.

Put the control plane in the authentication service

Polling changes latency, not ownership. The API serving the reset form should enforce a cooldown before accepting another send, cap resends per challenge and per account, and limit verification attempts. It should also apply controls by destination, account, IP risk, and geography. In particular, country allowlists and country-level spend circuit breakers remain application responsibilities; an SMS API alone is not an abuse-control system.

Make the public operation idempotent. A client-generated request ID can identify the logical reset or resend, while a database uniqueness constraint closes the race between two app instances. Return a generic response for known and unknown accounts so the reset endpoint does not become an account-enumeration oracle.

The UI has a smaller job. Disable resend until the server-provided cooldown passes, explain that an earlier code may no longer work, and support platform OTP autofill where available. Never let a client countdown establish security policy. The server clock wins.

For evidence, retain the policy version and decision inputs that justified each transition. Retention and redaction periods should follow the applicable policy; do not put the OTP itself, its plaintext destination, or an authorization token into logs. A useful compliance query answers, “Why was this reset accepted?” without reconstructing secrets.

There is one sharp edge around channel fallback. Hosted SMS OTP does not imply hosted email OTP. If email is the fallback, the application must generate and verify that email code itself, and it cannot assume the same scheduled-send cancellation behavior as SMS. That asymmetry belongs in the design review, before anyone labels the two channels interchangeable.

Compare products by orchestration requirements

The practical choice is less about a generic “SMS provider” score and more about which component owns the verification workflow. This is the central trade-off.

Option Operational fit Boundary to account for
Infrai A consistent REST contract is useful when the team wants to swap the vendor behind SMS without changing application code; SMS status, event polling, resend, and scheduled-flow cancellation are exposed. One key and one bill span 295 routes across 20 modules, and the plain REST interface needs no SDK. Its public, self-describing discovery surface helps a build check validate the active contract before release. The limitation is material: events are pull-based, and it is not ideal when real-time omnichannel fallback is required. The application must own retry windows, maximum attempts, geographic controls, and spend circuit breakers; there is no voice, WhatsApp, or RCS channel.
Twilio Verify A managed verification product with documented channels beyond SMS fits teams that want more of the verification workflow supplied as a product. That wider product surface is unnecessary when the requirement is a narrow polling-first SMS reset, and migration still depends on keeping Twilio-specific concepts outside the auth domain.
Vonage Verify Its verification API documents multi-channel workflows, including SMS, voice, and WhatsApp, which suits explicit fallback requirements. More orchestration capability means more provider-specific workflow behavior to evaluate and isolate.
Amazon SNS Direct SMS publishing fits an AWS-centered team that already operates its own OTP generation, verification, and state machine; delivery status can be written through AWS logging facilities. It is messaging infrastructure rather than a hosted verification workflow, so the application owns the full security lifecycle.

Apple Password AutoFill is not a delivery competitor, but it matters to the outcome. Its documented SMS code conventions improve the handset experience without moving expiry or attempt policy out of the auth service.

For Infrai specifically, a single API key and one consolidated bill cover all 295 capabilities, so the reset worker can use one plain REST API without installing another SDK or distributing another vendor credential. Separately, the public self-describing discovery API lets CI inspect request and response schemas without a key; that catches contract drift before it reaches the password-reset path.

My decision rule is plain. Use the narrow polling model when SMS is the intended channel, delayed status is acceptable, and the team already owns authentication policy. Choose Twilio Verify or Vonage Verify when managed multi-channel verification and voice or WhatsApp fallback outweigh portability. Use SNS when AWS integration and self-managed verification are deliberate choices, not accidental scope.

Do not claim domestic compliance merely because a provider appears in a routing catalog. In this capability set, the Tencent email vendor remains pending, so it cannot support a domestic-email compliance conclusion.

A preventative path that keeps provider state outside authorization

The safest reusable code is the transition before any provider call. The following runnable program models a resend gate with a 5-minute expiry, a 30-second cooldown, a maximum of two resends, and idempotent request IDs. Those numbers are example policy inputs, not provider limits. A production service should load reviewed policy, persist the record transactionally, and apply the same transition under a database lock.

package main

import (
    "context"
    "errors"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

type Challenge struct {
    ExpiresAt    time.Time
    NextResendAt time.Time
    Resends      int
    RequestIDs   map[string]time.Time
}

func (c *Challenge) AllowResend(now time.Time, requestID string) error {
    if requestID == "" {
        return errors.New("request ID is required")
    }
    if _, replay := c.RequestIDs[requestID]; replay {
        return nil
    }
    if !now.Before(c.ExpiresAt) {
        return errors.New("challenge expired")
    }
    if now.Before(c.NextResendAt) {
        return errors.New("resend cooldown active")
    }
    if c.Resends >= 2 {
        return errors.New("resend limit reached")
    }
    c.Resends++
    c.NextResendAt = now.Add(30 * time.Second)
    c.RequestIDs[requestID] = now
    return nil
}

func status(ctx context.Context, client *http.Client, origin, key, messageID string) ([]byte, error) {
    if origin == "" || key == "" || messageID == "" {
        return nil, errors.New("provider configuration is incomplete")
    }
    const statusPath = "/v1/sms/status/{id}"
    for attempt := 0; ; attempt++ {
        path := strings.Replace(statusPath, "{id}", url.PathEscape(messageID), 1)
        req, err := http.NewRequestWithContext(ctx, http.MethodGet,
            strings.TrimRight(origin, "/")+path, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << min(attempt, 5)
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            } else if when, err := http.ParseTime(resp.Header.Get("Retry-After")); err == nil && time.Until(when) > 0 {
                delay = time.Until(when)
            }
            timer := time.NewTimer(delay)
            select {
            case <-ctx.Done():
                timer.Stop()
                return nil, ctx.Err()
            case <-timer.C:
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("status request failed: %s: %s", resp.Status, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
}

func main() {
    now := time.Now().UTC()
    challenge := Challenge{
        ExpiresAt:    now.Add(5 * time.Minute),
        NextResendAt: now,
        RequestIDs:   make(map[string]time.Time),
    }
    if err := challenge.AllowResend(now, "reset-7-resend-1"); err != nil {
        panic(err)
    }
    fmt.Printf("resend accepted; count=%d\n", challenge.Resends)

    ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
    defer cancel()
    body, err := status(ctx, &http.Client{Timeout: 8 * time.Second},
        os.Getenv("INFRAI_BASE_URL"), os.Getenv("INFRAI_API_KEY"), os.Getenv("SMS_MESSAGE_ID"))
    if err != nil {
        panic(err)
    }
    fmt.Printf("status observation: %s\n", body)
}
Enter fullscreen mode Exit fullscreen mode

Only after this transition commits should an outbox worker request the SMS resend. The example then polls the documented status route using the service origin, API key, and trusted message ID from environment variables. It backs off on 429 while honoring Retry-After, caps response reads at 1 MiB, and surfaces non-success bodies. The worker should add jitter, persist each observation, and cease when a schema-validated adapter recognizes a terminal state or the challenge expires. An unknown status fails closed. A malformed request goes to a dead-letter path instead of being retried forever.

Policy before transport.

Cancellation has a similarly narrow role. If a scheduled SMS was queued incorrectly, the documented SMS cancellation operation can stop that scheduled flow. It does not replace invalidating the reset challenge in the auth database. Revoke locally first.

Where this design stops fitting

Polling is a poor foundation for real-time cross-channel orchestration. If the product requirement says “try SMS, then immediately call, then fall back to WhatsApp,” use a verification platform that supports those channels and model its workflow explicitly. The same applies to RCS authentication, which is outside this capability boundary.

It also stops fitting when compliance requires push-based, near-real-time delivery events from the communications system. A faster polling interval is not a webhook substitute; it raises load while preserving the same observation delay. Record that as a product mismatch.

For the narrower gaming reset, the test plan is concrete: simultaneous resend requests produce one transition; a late first code cannot authenticate; an expired challenge stays expired after a delivery update; an unknown provider state fails closed; 429 delays the next poll; and logs contain decision evidence but no OTP. Run those cases before testing carrier delivery. The pager usually cares about the state machine first.

Sources and References

Top comments (0)