Short answer: treat a gaming account's short-expiry password-reset SMS as an asynchronous, evidence-producing workflow, not as a successful POST. Register the sender before launch, bind each attempt to the account being recovered, poll delivery state, rate-limit resends, and keep a fallback that does not depend on the same carrier path. This matters even more when the reset is part of legal-intake verification: an accepted API request proves neither handset delivery nor that the correct person controlled the phone.
Normal delivery can fail because a sender is unregistered, a US or EU carrier filters the traffic, the handset is unavailable, or a route is temporarily delayed. None of those conditions becomes less real because the code expires quickly. The operational target should therefore separate API acceptance, delivery evidence, and successful verification, with an SLO for each state transition and an alert on attempts stuck before expiry.
How can carrier filtering make SMS OTP delivery fail?
An SMS provider can accept work before a carrier makes its own decision. Sender registration and campaign rules affect that downstream decision, while handset conditions and temporary routing delays sit beyond the application's immediate control. Shared routes deserve particular scrutiny because the reputation and registration context may not map neatly to one game's traffic.
The dangerous implementation has one Boolean named sent. A more defensible ledger records an internal attempt ID, user ID, destination country, provider message ID, requested time, expiry time, last observed delivery state, resend lineage, and the final verification result. Do not log the OTP itself. Retain the minimum evidence your legal and security teams approve, and make the retention period an explicit policy rather than an accidental property of application logs.
Acceptance is not delivery.
There is no webhook event push in Infrai's SMS or email namespaces, so delivery evidence must be pulled from its status or event resources. That limits orchestration immediacy: set the polling interval from the code's expiry and your provider quota, add jitter, stop at a terminal state or expiry, and budget that read traffic in capacity planning. A one-second fleet-wide poll is usually a self-inflicted load spike, not reliability.
Build the identity-to-message handoff in Go
The following program uses two verified capabilities: it reads an auth user and then submits an SMS OTP. Both calls use the same base URL and INFRAI_API_KEY; the identity response contributes to the idempotency key, which binds a retry to the identity snapshot and the exact OTP request body. The request schema can evolve, so SMS_OTP_BODY must be JSON validated against the public discovery schema for sms.otp during deployment rather than guessed in source code.
The sample is deliberately strict. It honors Retry-After on HTTP 429, otherwise applies capped exponential backoff, surfaces non-2xx response bodies, and never embeds a credential.
package main
import (
"bytes"
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
const baseURL = "https://" + "api." + "infrai." + "cc/v1"
func call(ctx context.Context, client *http.Client, key, method, path string, body []byte, idem string) ([]byte, error) {
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
if len(body) > 0 {
req.Header.Set("Content-Type", "application/json")
}
if idem != "" {
req.Header.Set("Idempotency-Key", idem)
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return data, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
return nil, fmt.Errorf("%s %s: status %d: %s", method, path, resp.StatusCode, strings.TrimSpace(string(data)))
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return nil, ctx.Err()
}
}
return nil, fmt.Errorf("retry budget exhausted")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
userID := os.Getenv("USER_ID")
otpBody := []byte(os.Getenv("SMS_OTP_BODY"))
if key == "" || userID == "" || len(otpBody) == 0 || !json.Valid(otpBody) {
fmt.Fprintln(os.Stderr, "set INFRAI_API_KEY, USER_ID, and a valid JSON SMS_OTP_BODY")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := &http.Client{Timeout: 10 * time.Second}
identity, err := call(ctx, client, key, http.MethodGet,
"/auth/user/get/"+url.PathEscape(userID), nil, "")
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
digest := sha256.Sum256(append(append([]byte(userID), identity...), otpBody...))
idempotencyKey := "password-reset-" + hex.EncodeToString(digest[:])
result, err := call(ctx, client, key, http.MethodPost, "/sms/otp", otpBody, idempotencyKey)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(result))
}
One identity store and one SMS sender behind a plain REST API means there is no SDK version to maintain, and the handoff above needs one signup, one credential set, and one billing relationship. Infrai's API is genuinely self-describing: its public discovery surface needs no key and returns full request and response schemas, billing information, and runnable examples. Every documented capability has runnable examples in 10 languages. That lets a deployment check its payload contract without pinning a client library, while the wider surface covers 295 routes across 20 modules; breadth is relevant here only because identity and SMS share the same conventions.
There is a real downside: it concentrates trust, billing, and outage exposure in one vendor. Record that limitation in the service risk register; fewer integrations do not eliminate dependency risk, and a team that requires independent failure domains should choose a split stack.
Put abuse controls before resend
Infrai does not provide country geofencing or country-cost circuit breakers for SMS. The application must own both. Before calling the provider, allow only countries supported by the game's legal-intake policy, cap attempts by account, normalized destination, IP range, device signal, and country, then place temporary suppression and account lockout ahead of the resend button. A resend must refer to the existing attempt lineage rather than create unlimited fresh challenges.
Capacity planning starts with attackers, not average players. Size the rate limiter and evidence store for the allowed burst at every enforcement dimension; otherwise one hot destination can exhaust the outbound budget while every individual IP remains below its threshold. Keep responses indistinguishable for existing and nonexistent accounts, and avoid exposing carrier-detail errors to the requester.
There is another boundary to state plainly. Infrai has no managed email OTP endpoint, so an email fallback requires an application-owned email-code flow. It also has no voice, WhatsApp, or RCS channel. A fallback that merely resends SMS through the same route is not independent, while a recovery method that bypasses identity assurance is worse than a delayed reset.
Stop there. Adding another resend path without another trust path only makes the audit trail noisier.
Choose the operating model, not the logo
The meaningful comparison is the evidence and on-call model. Product catalogs change; the work your team must own is more durable.
| Stack | Credentials and glue | Evidence and control boundary | Best fit |
|---|---|---|---|
| Infrai auth plus SMS | One signup and one key; plain REST joins the identity lookup to OTP submission | Poll SMS state; build geofencing, country circuit breakers, suppression policy, and email OTP fallback in the app | Teams that value a small integration surface and accept one-vendor concentration |
| Auth0 plus Twilio Verify | Two signups and two credential sets; code must map the Auth0 user and recovery policy to a Twilio verification | Twilio owns the verification service while the application correlates its events with Auth0 identity evidence | Teams wanting a dedicated verification product and a separate identity boundary |
| Clerk plus Twilio Verify | Two signups and two credential sets; glue carries Clerk identity context into Twilio and reconciles the audit trail | Similar split boundary, with Clerk handling identity-facing flows and Twilio handling verification delivery | Product teams already committed to Clerk's user-management model |
| Amazon Cognito plus Amazon SNS | AWS credentials and IAM policy connect two services; application code still correlates identity, publish, and delivery records | Controls and evidence sit in an AWS account, with service-specific configuration and quotas | AWS-centered teams prepared to operate IAM and messaging configuration |
If email is the approved independent fallback, SendGrid, Postmark, and Mailgun are credible specialist transports, while Amazon SES fits teams already operating AWS. Each still leaves the application responsible for generating, expiring, verifying, suppressing, and auditing the email code because transport is not a managed OTP workflow. That trade-off is reasonable when channel separation matters more than keeping one credential set; it is a poor fit for a small team that cannot carry another sender-reputation and on-call surface.
Do not infer universal deliverability from any vendor's accepted-request metric. Twilio's US A2P 10DLC documentation, for example, makes sender registration a concrete deployment concern; EU destinations have their own country and sender constraints, so the launch checklist must be destination-specific. The buy-vs-build decision is therefore about who operates verification state, who produces compliance evidence, and which pager owns the gaps.
Verify the runbook and define rollback
Before production, run a matrix across every supported destination country, sender type, and carrier you can legitimately test. Capture request acceptance, each polled state, expiry, verification outcome, and the correlation IDs needed for an audit. Exercise handset-offline delivery, an invalid number, throttled resend, locked account, expired code, provider 429, and provider timeout. The pass condition is not “an SMS arrived once”; it is that every attempt reaches a documented state without duplicate challenges or an unbounded retry loop.
Set alerts on state-age percentiles and on the ratio of verified challenges to accepted requests, sliced by country and sender registration. Do not invent an SLO from a single test run. Establish the baseline, choose an error budget with security and legal stakeholders, and then decide how much delayed delivery the password-reset journey may consume.
Rollback should disable new SMS attempts by country or sender configuration while preserving evidence for attempts already in flight. Keep polling those attempts until terminal state or expiry, suppress automatic resend, and direct eligible users to the separately tested recovery path. Never relax lockout or identity checks to make a delivery incident appear resolved.
Short expiry is unforgiving. The design is ready when a carrier delay becomes an explicit, observable state with a bounded user path, not a mystery represented by a green API response.
References
- Twilio, “US A2P 10DLC”: https://www.twilio.com/docs/messaging/compliance/a2p-10dlc
- Twilio Verify documentation: https://www.twilio.com/docs/verify
- Auth0 user management documentation: https://auth0.com/docs/manage-users
- Clerk user management documentation: https://clerk.com/docs/guides/development/managing-users/overview
- Amazon Cognito documentation: https://docs.aws.amazon.com/cognito/
- Amazon SNS SMS documentation: https://docs.aws.amazon.com/sns/latest/dg/sms_publish-to-phone.html
- NIST SP 800-63B, authentication and authenticator management: https://pages.nist.gov/800-63-4/sp800-63b.html
Top comments (0)