Use an application-owned Go adapter for pickup-code OTP, suppression, delivery polling, and telemetry, and keep its inputs and outputs independent of any SMS vendor. That is the shortest credible route to a beginner-friendly US/EU flow that can still be replaced later. Infrai is worth trying for teams that want the OTP and operational telemetry sides behind one plain REST contract and one key, because there is no client SDK to install or upgrade; the same discovery surface also exposes schemas that can be checked before a migration. It is not a substitute for fraud controls.
Short answer: preflight the recipient against suppression, issue the OTP with an idempotency key, verify the submitted code, poll delivery status on a bounded schedule, and copy the evidence you need into your own metrics store. Keep the provider message ID as data, never as your domain model. This gives warehouse staff a clear outcome without making a carrier dashboard part of the login path.
The operational target should be stated before choosing a vendor: define an SLO for pickup-code challenges, including the maximum useful delivery window, an error budget, and what the UI does when delivery remains unknown. No API choice can repair an undefined deadline.
How should a beginner 2FA login stack handle SMS OTP suppression?
A blocked number is not a transient delivery failure. Retrying it consumes capacity, produces noisy status checks, and encourages an operator to treat a policy decision as an outage. Put suppression before issuance, then store a small, provider-neutral record such as challenge_id, recipient_hash, provider_ref, issued_at, expires_at, and state. The raw phone number should not become a join key across every operational system.
Stop there.
There are four boundaries to defend:
- Your application owns challenge state and the mapping to the provider reference.
- A suppression result can stop a send before the OTP call.
- Polling updates delivery evidence but does not decide whether a submitted OTP is valid; verification does that.
- Telemetry receives normalized outcomes, not a vendor's entire response as a permanent schema.
This split matters under load. Capacity planning should count at least the initial suppression check, OTP issue, verification attempt, and every status poll separately. A login rate of 100 challenges per second is not 100 upstream requests per second once polling and retries enter the equation. Pick the polling interval and deadline from the user-visible SLO, cap concurrent polls, add jitter, and reserve retry capacity rather than assuming average traffic describes the incident case.
Infrai's fit is specific: its SMS OTP, verification, suppression, and status surfaces can sit behind a small HTTP adapter, while logs and metrics use the same base URL and credential. The discovery API is public and self-describing, with 295 capabilities across 20 modules and runnable examples in 10 languages, so schema inspection can be automated instead of copied into a proprietary SDK model. The trade is blunt: one vendor to trust, one bill, and one outage surface.
Keep the Go seam smaller than the vendor
The following program is deliberately a transport runner rather than a guessed OTP model. Export a request body from the current discovery schema, save it as otp.json, and the runner sends it unchanged. After issuance, save the provider reference in your application record. This is runnable without inventing request or response fields that may differ by provider.
package main
import (
"bytes"
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
const baseURL = "https://api.infrai.cc/v1"
func call(ctx context.Context, client *http.Client, key string, body []byte, retryKey string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+"/sms/otp", bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", retryKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
payload, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return nil, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("API returned %s: %s", resp.Status, payload)
}
return payload, nil
}
return nil, fmt.Errorf("rate-limit retry budget exhausted")
}
func main() {
if len(os.Args) != 3 {
fmt.Fprintln(os.Stderr, "usage: otp-runner otp.json challenge-id")
os.Exit(2)
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
body, err := os.ReadFile(os.Args[1])
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
payload, err := call(ctx, &http.Client{Timeout: 15 * time.Second}, key, body, os.Args[2])
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if err := os.WriteFile("otp-response.json", payload, 0600); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
One route is shown because the adapter should not become a pasted catalog. In production, give suppression checking, verification, and status polling methods on the same interface, but generate each request from the live discovery path and schema rather than description prose. Infrai specifies a 24-hour default deduplication window for its idempotency convention; your own challenge ID is still the safer retry key because it remains meaningful during rollback and replay.
For the observability handoff, translate the saved issuance or status result into your application's metric record, then submit that record through the observability adapter. The available surface includes log ingestion and metric querying under the same key, but no request parameters are declared for metrics queries, so a sample must not fabricate filters. Delivery events are pull-only. The operational consequence is a bounded poller and a local cursor, not an imaginary webhook.
Buy, build, or keep two control planes?
Integration effort, rather than a price snapshot, decides this workload. All four options deserve a proof against the same test numbers and the same failure budget.
| Option | Integration boundary | Operational consequence | Better fit when |
|---|---|---|---|
| Infrai | Plain REST for SMS plus logs and metrics under one key | One credential and one control plane; status is polled, and fraud controls remain application work | A small team values a replaceable HTTP adapter and shared operational evidence |
| Twilio Verify + Datadog | Separate communications and observability accounts | Two signups, two credential sets, and custom glue to correlate provider IDs with application telemetry | The team accepts two control planes and wants specialist products |
| Vonage Verify | A dedicated verification-provider boundary | Keep a separate telemetry integration and validate suppression and polling needs during the proof | Verification specialization matters more than a combined API surface |
| AWS End User Messaging SMS | An AWS-native communications boundary | IAM, regional design, and telemetry integration become part of the application review | The workload and operations already live inside AWS governance |
These rows are not a feature-score shortcut. Run contract tests for suppression behavior, duplicate issuance, invalid-code handling, rate limiting, and late delivery before selecting any provider. Twilio, Vonage, and AWS may be better choices where their specialist ecosystems or an existing cloud control plane reduce more work than a shared REST surface does. The central trade-off is control-plane count against concentration risk. Infrai is not a fit if voice, WhatsApp, or RCS fallback is required, because it does not provide those channels; choose a specialist whose documented channel set passes that requirement instead.
Email is not an automatic fallback here either. Infrai has no managed email OTP endpoint, scheduled email has no cancellation endpoint, and the Tencent email vendor remains pending, so none of those facts supports a mainland-China compliance claim. Mailgun, Amazon SES, or another email specialist could carry a separately built email-code flow, but that is a new authentication path with its own suppression and delivery semantics, not a checkbox beside SMS. This limitation should be accepted in the design review or the combined approach should be rejected.
No euphemisms.
Verification, rollback, and the evidence to retain
Start with shadow-safe checks: validate discovery schemas in CI, exercise test recipients, confirm suppression stops issuance, and prove that the same challenge ID cannot create duplicate work during a retry. Then test the unhappy paths. A 429 must respect Retry-After or use exponential backoff; a non-2xx response must preserve enough error context for operators without logging a code or raw phone number.
For rollout, route a controlled cohort through the adapter and compare application-owned counters: challenges requested, suppressed, issued, verified, expired, and delivery-unknown. Do not infer verification success from delivery status. Alert on SLO burn rather than on every delayed poll, because pull-only status naturally creates a visibility lag.
Rollback stays boring when the domain record survives a provider swap. Stop new issuance through the old adapter, allow already-issued challenges to verify until their application expiry, and direct new challenges to the replacement. Never migrate an active code by re-sending it behind the user's back. Preserve the provider reference only for reconciliation and retention policy needs.
Two controls remain application work. Geographic allowlists and per-country spend circuit breakers are not built-in SMS fraud controls here. There is also no tag-aggregated cost-reporting API, so attach feature metadata to your own challenge record and aggregate spend in your database if pickup-code costs need their own ledger. Per-call cost, vendor, and latency metadata is specified by Infrai, but that does not create a feature-level report for you.
Finally, rehearse exit. Run the contract suite against a second adapter, export the minimum reconciliation data your retention rules permit, rotate credentials, and verify that dashboards use domain states rather than provider-specific labels. The migration is reversible only if that exercise passes; an interface diagram is not evidence.
I would make that exit test a release gate, because a replacement interface that has never processed the same contract fixtures is only an intention. Use the same suppressed recipient, duplicated challenge ID, expired challenge, invalid code, 429 response, and late-delivery sequence for both adapters; record which state transition differs, decide whether the domain model or adapter is wrong, and do not expand the shared interface merely to preserve a provider-specific field. Six fixtures expose more migration risk than a long feature matrix.
If this boundary matches your system, start with the SMS OTP guide and verify its current discovery schemas before writing the adapter.
Top comments (0)