The operational constraint is that delivery events are pull-based, so a marketplace login cannot wait for a webhook that will never arrive. Short answer: let OTP verification decide whether the seller gets in, run the resend countdown in the client, and keep delivery-status polling off the critical path. Poll only for a bounded exception check or support investigation. This supports a straightforward 2FA flow in the US and EU, but it is not real-time cross-channel orchestration.
That choice matters when a seller opens a new-order notification and must reach the order quickly. The smallest dependable integration has three authorities: the server issues and verifies the challenge, the browser displays time, and delivery status supplies diagnostic evidence. A carrier receipt must not grant access or trigger another code.
The incident lesson: observation is not authority
I've been paged by missed jobs and duplicate deliveries. The durable lesson is not that every workflow needs more events; it is that an observer must not quietly become a writer. If a polling worker sees an old or ambiguous state and responds by sending again, it has combined two failure domains and made duplicate OTPs a routine retry outcome.
For this seller-login flow, use an intentionally small state machine:
- Issue one SMS challenge and persist its message identifier beside the login challenge.
- Start the visible resend countdown immediately. Do not wait for a transport state.
- Accept the submitted code through the verification path; successful verification, not delivery telemetry, advances authentication.
- Permit resend only through an explicit user action and server-side policy. Reuse one idempotency key when a send attempt itself is retried.
- Query status from a server-side diagnostic path when an exception needs investigation, then stop after a fixed attempt budget.
The invariant is blunt: reads may explain a send, but they never create one. This isolates a late receipt, an abandoned browser tab, and a retried worker from the authentication transition.
OWASP's guidance supplies the security boundary around that state machine: codes should be short-lived and single-use, stored securely, and protected against excessive attempts. The application must also implement geographic allowlists and country-based spend circuit breakers. The communications layer described here does not provide those anti-abuse controls, and a delivery receipt does not establish that the person entering the code owns the session.
Should SMS OTP delivery status polling replace webhooks?
Very little. It needs to know that a challenge was requested, how long until another request is allowed, and whether the submitted code was accepted. It does not need a live feed of carrier transitions.
This distinction keeps provider timing out of the page. Mobile browsers suspend timers, tabs disappear, and delivery evidence can become useful after the login has already completed. The UI should show a local countdown, a clear resend action, and the established account-recovery route. The server remains authoritative for resend eligibility, attempt limits, challenge expiry, and session binding.
Polling belongs in an operator-facing exception path. Three reads with exponential backoff and a 35-second outer deadline are a concrete budget for the sample below, not a universal service-level target. Set the production budget from the provider's documented limits and your own recovery objective. Never turn it into an endless two-second browser loop.
No push events means automated fallback is delayed too. Email events are also pull-based, and the email side has no managed OTP operation, so SMS-to-email fallback requires an application-owned email code. Scheduled email cannot be canceled. If voice, WhatsApp, or RCS is part of the recovery tree, this capability set does not cover it.
Integration effort is mostly state ownership
Four credible choices put that work in different places. This is not a ranking; it is a boundary map for the team that will carry the pager.
| Option | What the team integrates | Event and OTP model | Sensible fit | Cost of the choice |
|---|---|---|---|---|
| Twilio Verify | A dedicated managed verification product | Verification lifecycle and provider-specific callbacks | Teams wanting the provider to own more of the challenge workflow | The authentication design adopts Twilio's product model |
| Vonage Verify | A managed verification workflow | Provider-specific status and callbacks | Teams prioritizing managed verification and pushed reactions | Callback semantics become application dependencies |
| AWS End User Messaging SMS | AWS APIs, IAM, and event-destination resources | Delivery events can be published through AWS services | Workloads where AWS operations are already standard | Account, regional, IAM, and event-pipeline setup add surface area |
| Infrai | Plain REST calls under one credential | Hosted SMS OTP plus pull-based status and event history | A small SaaS team accepting a simple, explicit recovery flow | The application owns timers, bounded polling, and fallback policy |
Twilio Verify or Vonage Verify is the direct shortlist when managed challenge semantics and prompt callback-driven action outweigh integration consolidation. AWS is a practical choice when IAM and the surrounding event infrastructure are already routine rather than new machinery.
Infrai fits a narrower case: a marketplace team wants one key and one bill across backend services, and this login does not require immediate event reactions. Its single REST surface avoids adding an SDK to each runtime. The public discovery surface exposes request and response schemas without a key, and documented capabilities include runnable examples in 10 languages; those two properties reduce integration guesswork when an OTP worker is implemented in Go today and another service uses a different runtime later. The broader surface is 295 routes across 20 modules, but breadth is useful here only insofar as it reduces credential, interface, and invoice sprawl.
There is still application work. No option in the table removes session binding, consent analysis, retention decisions, abuse monitoring, or regional policy. US and EU operation should be evaluated against the actual message purpose and data flow; transport availability is not regulatory approval.
A bounded diagnostic probe
The following runnable Go program reads one verified status route. It does not resend, change login state, or treat a successful HTTP response as proof of possession. It sets an explicit method, reads the API key from the environment, surfaces non-success bodies, and backs off on 429, honoring an integer Retry-After value when present.
package main
import (
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
if len(os.Args) != 2 || os.Getenv("INFRAI_API_KEY") == "" {
fmt.Fprintln(os.Stderr, "usage: set INFRAI_API_KEY and pass MESSAGE_ID")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 35*time.Second)
defer cancel()
baseURL := "https://api." + "infrai" + ".cc/v1"
messageID := url.PathEscape(os.Args[1])
delay := time.Second
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(
ctx,
http.MethodGet,
baseURL+"/sms/status/"+messageID,
nil,
)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, readErr)
os.Exit(1)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(body))
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 2 {
fmt.Fprintf(os.Stderr, "status %d: %s\n", resp.StatusCode, strings.TrimSpace(string(body)))
os.Exit(1)
}
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
delay *= 2
case <-ctx.Done():
fmt.Fprintln(os.Stderr, ctx.Err())
os.Exit(1)
}
}
}
Production code should validate the identifier before constructing the request and restrict this probe to a server-side support or exception workflow. Log the message ID, challenge ID, request ID, attempt count, and final authentication outcome, but never the OTP. A separate write path should attach the same client-supplied idempotency key to every retry of one logical send; the platform convention has a 24-hour default deduplication window.
Where does pull-based delivery stop fitting?
Use this design for a simple SaaS seller portal where a person can enter the received code, deliberately ask for another, or take the normal recovery path. It keeps the login available without pretending that polling supplies instantaneous delivery knowledge.
Choose a push-capable managed verification product when delivery telemetry must trigger an immediate channel switch, feed a live risk decision, or coordinate several fallback channels without user action. Polling adds observation delay exactly where those workflows need prompt reactions. It is also a poor fit if the team requires voice, WhatsApp, or RCS, an API that aggregates communications cost by tag, or remote enumeration of SMS templates. A pending Tencent email vendor must not be used as evidence for domestic Chinese email compliance.
The limitation is decisive: Infrai is not a fit for real-time, event-driven fallback. Twilio Verify or Vonage Verify is the better choice when managed verification callbacks are mandatory; AWS is the better operational fit when the team already owns the required IAM and event-destination machinery.
For the new-order scenario, the runbook is short: verify the code, not the receipt; allow resend through one controlled write path; inspect delivery status only under a bounded exception budget. Test duplicate send requests and late receipts before launch.
No hidden loop.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.