An SMS provider can deliver and verify a code, but it cannot see enough of your application to decide whether a request is abusive. Short answer: design a secure SMS OTP login flow with per-user, per-IP, and per-device limits in front of every send; give each challenge a short expiry and an attempt ceiling; consume it exactly once; temporarily lock repeated failures; and enforce country policy before spending a send. Provider selection comes after those controls, not before them.
For a healthtech marketplace notifying a seller about a new order, keep that notification outside the login challenge. An order event may justify a message, but it must never create, extend, or implicitly verify an authentication challenge. The operational invariant is blunt: no unauthenticated caller gets an unmetered path from an HTTP request to a billable SMS.
How should a secure SMS OTP login flow fail?
Start the incident review with a bounded exercise: one account receives repeated OTP requests from rotating IP addresses, then the correct code is submitted twice. No invented outage or benchmark is needed to expose the design error. If the only alert is a provider dashboard showing message volume, it arrives after the application has already surrendered control.
I distrust that dashboard as the primary signal because it cannot answer the useful questions. Which user, device, IP, and destination country crossed policy? Did suppression prevent a send? Was a challenge accepted after it had already been consumed? The page should fire on an application-owned decision, such as a sharp rise in denied sends or locked challenges, with dimensions that let the responder distinguish a broken client release from a distributed attack. Provider delivery status is supporting evidence.
The postmortem invariant is therefore narrower than “add rate limiting.” Every state transition must be attributable and atomic: requested, denied, issued, failed, locked, expired, or consumed. Count denied requests as well as successful sends. Otherwise an attacker can hammer the gate while the quiet SMS graph looks healthy.
Page on policy failure.
Put policy before delivery
Evaluate controls from cheapest and broadest to most specific. Normalize the phone number and derive its country, apply an explicit allowlist or deny rule, check account state and suppression, then test limits for the user, IP, and device. Only after all four decisions pass should the backend ask a provider to send an OTP. Geographic anti-fraud throttles and country-price kill switches are application responsibilities here; assuming the messaging layer supplies them leaves a costly gap.
The three rate-limit keys catch different attacks. A user limit contains harassment of one account. An IP limit slows a noisy source. A device limit still has value when addresses rotate, although a device identifier is evidence rather than identity and should never become a permanent fingerprint masquerading as authentication. NATs, hospitals, and corporate networks also make a hard IP-only rule prone to collateral lockouts, so the decision should combine dimensions rather than treating any one as infallible.
Keep the send response deliberately vague. “If this account can receive a code, it will arrive shortly” leaks less than separate messages for unknown users, suppressed numbers, or locked accounts. Internally, retain a precise reason code. Operations needs truth even when an unauthenticated caller should not get it.
Suppression deserves an early check because repeatedly attempting blocked or opted-out numbers is neither a useful authentication control nor a reasonable notification strategy. The option evaluated here exposes SMS OTP and suppression capabilities through a plain REST API, so a service that already speaks HTTP can integrate without installing or tracking a client SDK; its country rules and abuse budgets still belong in the service that knows the user and device. That boundary is the reason to consider it, not a claim that one API can infer business risk.
There is a second, less visible integration advantage. Infrai uses one API key and one bill across its capability surface, while its genuinely self-describing public discovery surface exposes request and response schemas without a key; the snapshot covers 295 routes across 20 modules with runnable examples in 10 languages. For this marketplace, the same credential can cover OTP and seller-notification calls, reducing credential rotation and invoice reconciliation across those paths, and schema review can happen before a credential enters the build pipeline. The trade-off remains operational coupling to a broad platform, and breadth does not compensate for a missing channel or security control.
Make the challenge a one-way state machine
Store a keyed hash of the code, never the cleartext code. Bind the record to the intended user and normalized destination, use a short expiration chosen for the product's delivery conditions, and set a finite maximum attempt count. The exact duration and count are policy choices absent from the provider facts, so test them against real delivery latency and support burden instead of copying a magic number from an example.
Verification must be one atomic operation. Compare the submitted value, increment failure state, apply a temporary lockout when the ceiling is reached, and mark a successful challenge consumed in the same transaction. A second verifier racing with the first must lose. So must a replay arriving one millisecond later.
Once means once.
Do not extend expiry on resend. A resend should either refer to the same bounded challenge or replace it while invalidating the previous one; allowing each request to push the deadline forward turns retry traffic into indefinite validity. Likewise, a successful verification should revoke other outstanding challenges for the same login intent when the product permits only one active flow.
The distinction between retry classes matters at 3 a.m. Retrying a local transaction after a serialization conflict is safe when the transaction is atomic. Retrying a provider write requires an idempotency key so a timeout cannot create a second send. Retrying a rejected code is a security event and consumes an attempt. These are three different mechanisms, even though a dashboard may label all of them “retries.”
A small preventative path in Go
The business gate belongs before this program. The runnable adapter below makes the actual OTP POST after that gate passes, but it does not guess at provider fields: export a request object validated against the current public discovery schema as OTP_REQUEST_JSON. It reads the key and base URL from the environment, uses a stable challenge ID as the idempotency key, retries HTTP 429 with bounded exponential delay while honoring Retry-After, and returns the real error body for every other non-2xx response.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func retryDelay(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func sendOTP(ctx context.Context, client *http.Client, body []byte, challengeID string) ([]byte, error) {
base := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
key := os.Getenv("INFRAI_API_KEY")
if base == "" || key == "" || challengeID == "" {
return nil, fmt.Errorf("INFRAI_BASE_URL, INFRAI_API_KEY, and CHALLENGE_ID are required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, base+"/v1/sms/otp", bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", challengeID)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return responseBody, nil
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
return nil, fmt.Errorf("OTP send returned %s: %s", resp.Status, responseBody)
}
timer := time.NewTimer(retryDelay(resp.Header.Get("Retry-After"), attempt))
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
}
return nil, fmt.Errorf("retry budget exhausted")
}
func main() {
body := []byte(os.Getenv("OTP_REQUEST_JSON"))
if !json.Valid(body) {
panic("OTP_REQUEST_JSON must be valid JSON matching the current discovery schema")
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
result, err := sendOTP(ctx, &http.Client{Timeout: 10 * time.Second}, body, os.Getenv("CHALLENGE_ID"))
if err != nil {
panic(err)
}
fmt.Println(string(result))
}
Do not mistake the adapter for the security design. The pre-send counters need explicit windows or token buckets and durable atomic updates; verification needs its own atomic consumed flag, expiry, attempt count, and temporary lockout. The adapter's idempotency key prevents one accepted business decision from becoming two provider writes, but it cannot decide whether that business decision was legitimate.
Compare integration boundaries, not feature checklists
Twilio Verify, Vonage Verify, and Firebase Authentication all offer managed phone-verification workflows, while Amazon SNS is a more general messaging primitive. Those are real alternatives, but they place the seam in different locations. A fair selection starts with how much authentication state the team wants to own.
| Option | Integration boundary | Best fit | Boundary to inspect |
|---|---|---|---|
| Twilio Verify | Managed verification service with its own service and verification flow | Teams wanting an established OTP-specific workflow | Confirm how application-level user, IP, device, geography, and lockout policy wraps the service |
| Vonage Verify | Managed verification workflow | Teams that want provider-managed verification rather than building code delivery | Model the same pre-send abuse gate and check current regional behavior in its documentation |
| Firebase Authentication | Client and backend authentication product with phone sign-in | Applications already willing to adopt Firebase's identity boundary | It is a larger identity integration than adding an SMS call to an existing auth service |
| Amazon SNS | General SMS publishing infrastructure | Teams that already own challenge generation and verification state | More authentication machinery remains in the application |
| A plain REST OTP API | HTTP boundary without a required SDK | Polyglot backends minimizing library and key-management integration | Native geography throttles and country-price circuit breakers still require application policy |
No row removes the need to decide what happens before a send. Managed verification can reduce code-secret handling, while a general SMS primitive provides more control and more responsibility. The fastest proof is not a successful happy-path message; it is a test matrix covering suppressed destinations, disallowed countries, duplicate sends, concurrent verification, replay after success, expiry, maximum attempts, and temporary lockout. Run two verification requests concurrently against the same challenge and demand one winner; advance the clock beyond expiry; exhaust the attempt counter; rotate the source IP while holding the user and device constant; then hold the IP constant while changing users. This sequence catches policy that exists in prose but not at the transaction boundary.
There are also cases where SMS OTP is the wrong center of gravity. NIST treats PSTN out-of-band authentication as restricted, so a system with stronger assurance requirements should evaluate phishing-resistant authenticators rather than polishing SMS indefinitely.
The limitation is concrete. Infrai does not fit a product that requires immediate webhook events across email and SMS, voice fallback, WhatsApp, RCS, SMTP relay, or managed email OTP; its communication events are pulled, and the fallback orchestration remains yours. Choose Twilio Verify or Vonage Verify when their managed verification workflow and supported delivery options better match the requirement, Firebase Authentication when adopting its identity boundary is acceptable, or Amazon SNS when the team explicitly wants to own the full challenge lifecycle over a general SMS primitive.
The decision I would sign off is conditional. Choose the provider whose boundary matches the authentication system you are prepared to operate, but ship the abuse gate and atomic challenge state as application controls. Then page on the gate. A provider success graph cannot tell you that the wrong person was allowed to ask.
Sources
References consulted:
- NIST SP 800-63B, Digital Identity Guidelines: https://pages.nist.gov/800-63-3/sp800-63b.html
- Twilio Verify API overview: https://www.twilio.com/docs/verify/api
- Vonage Verify API overview: https://developer.vonage.com/en/verify/overview
- Firebase phone number sign-in: https://firebase.google.com/docs/auth/web/phone-auth
- Amazon SNS SMS documentation: https://docs.aws.amazon.com/sns/latest/dg/sms_publish-to-phone.html
- OWASP Authentication Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Authentication_Cheat_Sheet.html
Top comments (0)