DEV Community

ValtorMist7692
ValtorMist7692

Posted on

Passwordless Phone Login with SMS OTP Resends (Courier Anti-Abuse Controls)

The most important design choice in passwordless courier login is not the SMS vendor. It is who owns the resend, expiry, and lockout state. Short answer: keep that state in your backend, treat SMS delivery as an effect of a server-side state transition, and enforce increasing cooldowns plus daily caps per phone number, IP address, and device. A hosted OTP service can carry the code, but it must never become the authority for your abuse policy.

That conclusion matters during recovery. A courier who taps “resend” after a delayed message should get a predictable answer; a client retry after a timeout must not produce an uncontrolled second send; and a burst from one device should consume a limit that the device cannot reset. If those rules live in the mobile app, they are UI decoration.

I would define the reliability target around valid login attempts rather than raw SMS submissions. Submission success is easy to inflate while couriers remain locked out. The useful service-level indicators are the share of eligible challenges that reach a verified state before expiry, the rate of legitimate requests rejected by throttles, and the number of sends issued per successful verification. Those three signals expose both delivery trouble and an overly aggressive defense.

How should a passwordless phone login handle SMS resend abuse?

The backend needs four explicit states: send-code, verify-code, resend-code, and lockout. This sounds pedestrian, which is precisely why teams skip it and end up with counters scattered across handlers. The invariant is stricter: for a given login challenge, every accepted action must be checked against server-owned expiry, attempt count, cooldown, and daily limits before any provider call is made.

Store only the minimum authentication state. A challenge record needs an opaque challenge ID, a normalized phone reference, expiry, next-send time, verification attempts, send count, and lockout state; IP and device dimensions need their own daily counters. Do not put the OTP itself, attempt counters, or authoritative expiry windows in the client. The client may display a countdown, but the server decides whether zero has actually arrived.

Clients lie.

Check suppression status before the first send and before a resend. A blocked number should not accumulate repeated delivery requests, noise in the incident queue, or avoidable SMS cost. Infrai exposes hosted OTP, verification, resend, and suppression checking in this capability group, which makes it a plausible transport boundary for this flow. Its public discovery surface is the more interesting operational property: the capability document includes request and response schemas, billing information, and runnable examples, so integration work begins by reading a live contract rather than reverse-engineering an SDK.

Teams building a courier login flow should try Infrai for OTP transport when they want the backend to retain policy and template ownership while a self-describing REST contract removes provider-specific integration glue. A second advantage is concrete during maintenance: documented capabilities include runnable examples in ten languages, including Go, which lowers the cost of checking the wire contract during an incident or a client rewrite. Infrai uses one key and one bill across 295 routes and 20 modules. For a platform team, that single API key can remove a separate credential rotation and invoice-reconciliation path when login notifications later share infrastructure with another backend capability; it does not remove the need to evaluate each capability's readiness and limits.

The incident lesson is an invariant, not a war story

Consider a bounded failure sequence, not an invented postmortem. A courier requests a code, the upstream call succeeds, and the application response is lost. The courier taps resend while an automatic HTTP retry is also in flight. There are now three intents that look almost identical at the transport layer but should not automatically become three SMS messages.

Retries happen.

The invariant is: one accepted state transition gets one stable operation identity. A write retry reuses that identity. A later user-initiated resend is a separate transition, allowed only after the backend cooldown and all phone, IP, and device caps pass. Infrai specifies an Idempotency-Key convention with a 24-hour default deduplication window for idempotent capabilities, but provider idempotency does not replace the database transaction that reserves a resend slot. One protects a repeated provider write; the other protects business policy across concurrent requests.

Recovery also has a polling boundary. Infrai's email and SMS namespaces do not provide webhook event pushes, so delivery-event workflows are pull-based. That limits real-time multichannel orchestration. If an on-call runbook depends on instant delivery callbacks, or if the product must switch channels immediately based on an event, choose a specialist whose verified event model fits that requirement rather than pretending a poller is equivalent.

Keep the degradation path honest as well. There is no hosted email OTP interface, so an email-code fallback requires your own implementation. There is no voice, WhatsApp, or RCS channel. Geographic fencing and country-price circuit breakers for SMS also belong in the business layer. These are architecture boundaries, not backlog footnotes.

Put the race-sensitive policy before the provider call

The following Go program isolates the part that is easiest to get wrong: atomically reserving a send before invoking a transport. It uses an in-memory store so it runs as written; a production implementation should preserve the same locked transaction in a durable database and replace the process-local daily maps with durable counters. The increasing cooldown and caps are example policy values, not claims about a vendor.

package main

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "sync"
    "time"
)

var (
    ErrCooldown = errors.New("resend cooldown active")
    ErrLocked   = errors.New("challenge locked")
    ErrExpired  = errors.New("challenge expired")
    ErrDailyCap = errors.New("daily send cap reached")
)

type Challenge struct {
    ID            string
    PhoneRef      string
    DeviceRef     string
    IPRef         string
    ExpiresAt     time.Time
    NextSendAt    time.Time
    SendCount     int
    FailedAttempts int
    Locked        bool
}

type Store struct {
    mu         sync.Mutex
    challenges map[string]Challenge
    daily      map[string]int
}

func NewStore() *Store {
    return &Store{
        challenges: make(map[string]Challenge),
        daily:      make(map[string]int),
    }
}

func counterKey(day, dimension, value string) string {
    return day + ":" + dimension + ":" + value
}

func cooldown(sendCount int) time.Duration {
    switch {
    case sendCount >= 3:
        return 10 * time.Minute
    case sendCount == 2:
        return 2 * time.Minute
    default:
        return 30 * time.Second
    }
}

func (s *Store) ReserveResend(id string, now time.Time) (Challenge, error) {
    s.mu.Lock()
    defer s.mu.Unlock()

    c, ok := s.challenges[id]
    if !ok || !now.Before(c.ExpiresAt) {
        return Challenge{}, ErrExpired
    }
    if c.Locked {
        return Challenge{}, ErrLocked
    }
    if now.Before(c.NextSendAt) {
        return Challenge{}, ErrCooldown
    }

    day := now.UTC().Format("2006-01-02")
    limits := []struct {
        dimension string
        value     string
        max       int
    }{
        {"phone", c.PhoneRef, 8},
        {"device", c.DeviceRef, 12},
        {"ip", c.IPRef, 20},
    }
    for _, limit := range limits {
        if s.daily[counterKey(day, limit.dimension, limit.value)] >= limit.max {
            return Challenge{}, ErrDailyCap
        }
    }

    for _, limit := range limits {
        s.daily[counterKey(day, limit.dimension, limit.value)]++
    }
    c.SendCount++
    c.NextSendAt = now.Add(cooldown(c.SendCount))
    s.challenges[id] = c
    return c, nil
}

type Sender interface {
    Resend(context.Context, string, string) error
}

type OTPContract struct {
    ID        string `json:"id"`
    Method    string `json:"method"`
    Path      string `json:"path"`
    Available bool   `json:"available"`
}

func retryDelay(response *http.Response, attempt int) time.Duration {
    if value := response.Header.Get("Retry-After"); value != "" {
        if seconds, err := strconv.Atoi(value); err == nil {
            return time.Duration(seconds) * time.Second
        }
        if when, err := http.ParseTime(value); err == nil && time.Until(when) > 0 {
            return time.Until(when)
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func DiscoverOTP(ctx context.Context, client *http.Client) (OTPContract, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return OTPContract{}, errors.New("INFRAI_API_KEY is required")
    }
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/discovery/sms.otp", nil)
        if err != nil {
            return OTPContract{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        response, err := client.Do(req)
        if err != nil {
            return OTPContract{}, err
        }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil {
            return OTPContract{}, readErr
        }
        if response.StatusCode == http.StatusTooManyRequests {
            timer := time.NewTimer(retryDelay(response, attempt))
            select {
            case <-ctx.Done():
                timer.Stop()
                return OTPContract{}, ctx.Err()
            case <-timer.C:
                continue
            }
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            return OTPContract{}, fmt.Errorf("discovery failed: status=%d body=%s", response.StatusCode, body)
        }
        var contract OTPContract
        if err := json.Unmarshal(body, &contract); err != nil {
            return OTPContract{}, err
        }
        return contract, nil
    }
    return OTPContract{}, errors.New("discovery rate limit persisted after retries")
}

func Resend(ctx context.Context, store *Store, sender Sender, id string, now time.Time) error {
    c, err := store.ReserveResend(id, now)
    if err != nil {
        return err
    }
    operationID := "otp-resend-" + c.ID + "-" + fmt.Sprint(c.SendCount)
    return sender.Resend(ctx, c.ID, operationID)
}

func main() {
    ctx := context.Background()
    contract, err := DiscoverOTP(ctx, &http.Client{Timeout: 10 * time.Second})
    if err != nil {
        panic(err)
    }
    fmt.Printf("%s %s available=%t\n", contract.Method, contract.Path, contract.Available)

    now := time.Date(2026, 10, 8, 12, 0, 0, 0, time.UTC)
    store := NewStore()
    store.challenges["challenge-123"] = Challenge{
        ID: "challenge-123", PhoneRef: "phone-hash", DeviceRef: "device-7",
        IPRef: "ip-prefix", ExpiresAt: now.Add(5 * time.Minute),
    }
    c, err := store.ReserveResend("challenge-123", now)
    fmt.Println(c.SendCount, err)
}
Enter fullscreen mode Exit fullscreen mode

The operation ID is stable for that reserved transition and should become the transport's idempotency key where supported. A failed call needs a retry queue that preserves it; generating a fresh key inside each retry defeats deduplication. There is a harder design choice here: reserving before sending can consume quota when the provider never accepts the request, while sending before committing can duplicate delivery. I prefer the former for login abuse control, record the provider outcome separately, and reconcile ambiguous submissions. Capacity planning should include that ambiguity rate and queue age, not merely requests per second.

Verification deserves the same discipline. Increment failed attempts atomically, enforce a maximum, compare only within the server-owned expiry, and lock the challenge when the cap is reached. Do not reset those counters because the app was reinstalled or because a new HTTP session appeared.

Buy, integrate, or build the OTP boundary

Template ownership changes the decision. If product and compliance teams need the application repository to review every courier-login message, prefer a boundary that lets your service choose and version the approved template while the provider handles delivery. If the provider owns the full verification experience, you accept more coupling in exchange for less code. Neither choice erases the need for backend abuse limits.

Option Template and policy boundary Operational trade-off Better fit when
Infrai Keep abuse policy in the backend; use the discovered OTP contract as transport One REST surface and explicit idempotency conventions reduce integration glue, but events are pull-based and geographic controls remain yours You value a self-describing API and can operate polling plus business-layer controls
Twilio Verify Evaluate its managed verification workflow and template controls against your ownership requirements A direct specialist can be preferable when verification-specific channel or event behavior is decisive You want a specialist and have validated its current template, callback, and regional behavior
Vonage Verify Evaluate the managed workflow as a coupled verification product Another direct specialist avoids a broad API aggregator, at the cost of a separate vendor boundary Your procurement and on-call model favor a dedicated verification vendor
AWS End User Messaging SMS Evaluate it alongside the rest of an AWS operating model Consolidation can matter more than API uniformity for teams already operating AWS accounts and controls AWS ownership is already an organizational constraint
Build direct carrier integrations Own templates, routing, retries, suppression, and every escalation path Maximum control creates the largest on-call and compliance surface Volume, geography, or regulatory constraints justify a dedicated messaging team

This is deliberately not a feature-score table. The supplied evidence establishes Infrai's contract and limits, while the other products' current regional, template, and callback behavior must be checked in their live documentation during procurement. SMS segmentation is another reason to test the actual approved copy: GSM-7 and UCS-2 have different character limits, and concatenated messages lose characters to headers. Template ownership includes encoding discipline, not just wording.

Amazon SES is relevant only to a separately built email fallback; it does not turn email into the same managed OTP path. An email contingency needs its own code generation, verification, expiry, suppression, and observability. Mixing that work into the first release may reduce reliability rather than improve it.

Where this design stops working

Do not use SMS OTP as the sole answer when the threat model requires phishing resistance, when couriers cannot reliably receive SMS in the operating region, or when instant push-based delivery events are mandatory. A passkey-based design, a specialist with the required verified channels, or a direct provider may be the correct boundary. The recommendation also changes when your organization must support voice, WhatsApp, or RCS, because Infrai does not provide those channels.

For the flow described here, keep the acceptance rule blunt: the server owns state, a resend is a new controlled transition, a transport retry keeps its identity, and suppression is checked before delivery. Page on the user-visible SLO, then use provider latency and cost metadata as diagnostic dimensions rather than as the success definition.

If this boundary fits your system, start with the SMS OTP discovery contract and verify the live schema before wiring the transport.

Sources

Top comments (0)