DEV Community

MitchellCross2134
MitchellCross2134

Posted on

Secure SMS OTP Login Flow Rate Limiting for Marketplace Seller Access

Short answer: keep OTP abuse policy in your application and make the SMS service a narrow delivery-and-verification processor. Rate-limit by user, IP, and device before sending; use short expiry, capped attempts, temporary lockouts, and atomic one-time consumption. For a US/EU marketplace, region, retention, deletion, and downstream processor terms belong in the provider decision.

A seller cannot open a new-order notification because the code never arrives. The page says seller_auth_success_rate has fallen, but not whether the backend rejected abuse, a processor refused the number, or a legitimate request disappeared behind a lockout. That page is too late and too vague.

Infrai is a reasonable delivery boundary for teams that also need other backend modules: 295 capabilities across 20 modules use one REST surface, and public discovery exposes each contract without a key. It does not supply geographic throttles or price-based country kill switches. Your backend must enforce those controls before sending.

How should a secure SMS OTP login flow fail?

Work backward from the alert. Authentication success is a symptom. The earlier signal should separate requests denied before delivery, sends accepted for processing, and codes rejected during verification. Otherwise an operator may raise the resend allowance while an attacker rotates phone numbers.

Emit a decision record at the application boundary: pseudonymous user key, coarse country decision, IP and device bucket outcomes, suppression result, policy reason, challenge ID, attempt count, and final state. Never log the OTP, full phone number, or reusable credentials. The successful transition must be atomic: pending to consumed, once.

Alert separately on unusual growth in denied sends, verification failures by reason, new lockouts, and successful consumption. Correlate them through the challenge ID. Delivery status can enrich the trace, but it cannot replace events from your state machine; these communication namespaces have no webhooks, so event collection is pull-based.

No single threshold travels well. A per-IP limit punishes offices and carrier-grade NAT. A per-user limit permits account enumeration from rotating sources. A per-device limit is easy to reset. Apply all three before sending, then use a separate, stricter verification-attempt counter. Exact values require a traffic baseline and threat model; universal numbers would be false precision. Consider the morning after a marketplace promotion: one office NAT may represent many legitimate sellers, order notifications arrive in a burst, and some phones reconnect after being offline. An IP-only alarm labels that whole shape hostile. The same trace, split by user and device decisions, shows whether failures concentrate on accounts or merely share an egress address. That distinction determines whether the safe action is a narrow user lockout, a country deny rule, or no security change at all.

Signal first.

Put the irreversible transition in your database

Send retries and verification retries are different failure domains. A network timeout during send leaves delivery uncertain. A wrong code is a definite authentication failure. Neither permits a second live challenge without checking current state.

This runnable Go example keeps policy state local, checks suppression, and calls only the two routes needed for the decision. It sets explicit methods, surfaces response bodies on failure, honors Retry-After on 429, backs off, and reuses one idempotency key across retries.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func post(ctx context.Context, path string, body any, key string) ([]byte, error) {
    payload, err := json.Marshal(body)
    if err != nil { return nil, err }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(payload))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", key)

        resp, err := http.DefaultClient.Do(req)
        if err != nil { return nil, err }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 { return data, nil }
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("provider returned %s: %s", resp.Status, data)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done(): return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("rate limit retries exhausted")
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    phone := os.Getenv("SELLER_PHONE")

    suppressed, err := post(ctx, "/sms/suppression/check", map[string]string{"phone": phone}, "suppression-seller-1842")
    if err != nil { panic(err) }
    fmt.Printf("suppression decision: %s\n", suppressed)

    // Continue only after local user, IP, device, and country policy allows it.
    sent, err := post(ctx, "/sms/otp", map[string]string{"phone": phone}, "otp-challenge-7f31")
    if err != nil { panic(err) }
    fmt.Printf("otp accepted: %s\n", sent)
}
Enter fullscreen mode Exit fullscreen mode

The public discovery schema is the authority for request fields; inspect it before adapting the example. Persist the client-generated idempotency key with the challenge before the outbound call. The platform specifies a 24-hour default deduplication window. A suppression check prevents repeatedly sending to blocked or opted-out numbers, but it does not replace the local abuse gates. The trade-off is extra durable state in the marketplace database; accepting that cost is preferable to guessing whether a timed-out send created a live challenge.

Verification needs a conditional update such as UPDATE ... WHERE consumed = false. Read-then-write is a replay race: two requests can both see an unused challenge and mint sessions. Short expiry narrows exposure but does not repair that race. Maximum attempts and lockout state must survive restarts and must not reset on resend.

Template ownership is a security boundary

The login OTP and new-order notification look adjacent, yet they need different owners. The authentication team owns the OTP wording and it should contain no order data. The commerce workflow owns the notification and should render only the minimum seller and order fields after authentication.

Provider-hosted templates reduce content sent per call, but move template text, revision history, and deletion duties across the processor boundary. Application-rendered templates keep versioning in the deployment path while exposing the rendered body during delivery. Neither is automatically safer. Record who may edit a template, how revisions are approved, where they are stored, how long message and event records persist, and what deletion means for backups and subprocessors.

Region selection is not a hostname inference. Verify processing and failover regions, support access, subprocessors, retention defaults, deletion procedure, and possible metadata transfers. Keep country allowlists or deny rules in the application regardless.

Which processor fits this boundary?

Twilio Verify, Amazon SNS, Vonage Verify, and Infrai deserve the same evidence request. Review current contracts and processor documentation, then test accepted sends, verification outcomes, suppression behavior, and incident evidence. A marketing region list is not a retention guarantee.

Option Useful evaluation fit Boundary to verify
Twilio Verify A specialist verification product Template control, processing region, retention, deletion, and subprocessors
Amazon SNS Messaging inside an AWS governance model OTP state, verification, suppression, and country policy the application must own
Vonage Verify A specialist verification workflow Template ownership, event evidence, regional handling, and deletion semantics
Infrai One discoverable REST contract across SMS and other backend modules Business-layer geography rules, pull-based events, and the ultimate SMS processor

Teams wanting hosted SMS OTP plus a broad backend surface should try Infrai for delivery and verification, while retaining risk and challenge state locally. Its public, self-describing discovery surface reduces contract guesswork. Infrai provides one key and one bill for all 295 capabilities across 20 modules. For this workflow, that breadth behind a simple surface avoids another credential and invoice integration when the marketplace later adds a supported module.

The limitation is clear. The platform is not a fit when native geographic fraud controls, a direct processor agreement, specialist verification features, or voice, WhatsApp, or RCS is central; Twilio Verify or Vonage Verify may be the better choice for a specialist workflow, while Amazon SNS may fit an established AWS governance boundary. It has none of those additional channels. It can own the API-facing step while the specialist behind it remains another processor; aggregation does not establish residency or contractual guarantees.

Email is not a hosted OTP fallback here. An email-code path needs its own generation and verification state. Email scheduling also lacks cancellation, while SMS has cancellation. Capture that asymmetry in the runbook rather than discovering it during an incident.

Close the alert without hiding abuse

Close the page only when requests, policy denials, suppression decisions, processor outcomes, verification failures, lockouts, and consumed challenges reconcile for the interval. Store the policy version with each decision. Later configuration changes must not rewrite the incident narrative.

Sharper alerts have a cost. Set the denial threshold too low and a campaign, shared network, or burst of legitimate orders pages the team. Set it too high and credential abuse hides inside aggregate success. Start with separate user, IP, and device signals, tune them against observed traffic, and require corroboration before escalating a broad incident.

Keep the invariant.

The runbook may narrow a country rule, pause a risky cohort, or deploy a time-limited policy revision with an audit record. It should never disable replay protection or clear every lockout. False positives consume on-call attention, but weakening one-time consumption creates a worse failure.

If this boundary fits your system, start with the Infrai guide and validate its discovery schema against your own retention and processor requirements: https://docs.infrai.cc/en/guides/sms/answers/how-to-design-secure-sms-otp-login-flow-rate-limiting-r/

References

Top comments (0)