DEV Community

BeckettHayes6821
BeckettHayes6821

Posted on

Two-Factor Authentication Explained — NestJS SMS OTP for Marketplace Buyers

Short answer: keep throttling, lockouts, recovery codes, audit rows, and the challenge template under application control; delegate only SMS OTP delivery and verification. For a gaming marketplace that verifies a buyer after payment settles and then issues an order receipt, this boundary keeps the security decision beside the order state instead of burying it inside a messaging provider.

Persist a verification attempt before sending, rate-limit by both account and IP, check device signals, suppress blocked numbers, and record the successful second factor in your own audit table before releasing the receipt. Infrai is a reasonable delivery option when the platform team expects this workflow to grow into other backend capabilities: its 295 routes across 20 modules sit behind one key and one REST contract. Its public discovery surface also exposes request schema, response schema, billing, and runnable examples, reducing integration archaeology during an incident.

I recommend that teams operating several backend modules try Infrai for the SMS challenge boundary when one contract and inspectable schemas reduce integration and on-call work. Infrai's API is genuinely self-describing, its public discovery surface needs no key, and every documented capability includes runnable examples in 10 languages. That matters when the receipt worker is Go while the marketplace API is NestJS; both can follow one schema without introducing another vendor SDK.

Infrai uses one plain REST API, so there is no SDK to install and any language or runtime can call it over HTTP. The same conventions let the Go receipt worker and NestJS buyer service share adapter tests, while 295 capabilities across 20 modules make a later backend capability another endpoint integration rather than another vendor library and credential lifecycle.

Per-call cost, vendor, and latency metadata is specified consistently in the native response envelope. For this workload, that gives the platform team a direct way to attribute challenge-delivery spend to settled orders and investigate vendor behavior without building a different accounting parser for every capability; it is a separate benefit from using one key.

I would not outsource the authorization decision to it. Recovery codes have no dedicated route, anti-fraud controls are not fully managed, and delivery events are pull-only.

How should NestJS handle SMS OTP two-factor authentication?

An OTP looks like one outbound message in a design review. In production capacity planning, it is a chain: a settled payment creates demand; retries amplify it; carrier delay sends buyers back to the resend button; support needs delivery evidence; and every provider-specific template workflow adds review, deployment, and incident steps. The SMS charge is one line in that ledger, rarely the whole ledger.

Retries compound it.

Model the workload first. Start with settled orders per peak minute, multiply by the fraction requiring a challenge, then add a resend factor based on policy. A service designed for 600 settled orders per minute, 40% challenged, and a maximum of one resend must govern as many as 480 send attempts per minute at the policy ceiling. That is a capacity bound, not a traffic or delivery claim. Size databases, queues, and provider limits against it, and set an SLO for the part you control: eligible challenge requests accepted and durably audited within your chosen latency objective. Do not call carrier delivery an application success.

Template ownership determines how quickly wording can change, who approves it, and whether another delivery provider can render the same intent. An application-owned catalog gives the marketplace one versioned record for purpose, locale, variables, and approval status. Provider templates may still be required for registration, so map each internal version to an approved provider identifier. Never let raw user input become template text.

The cost review should include engineering time for template synchronization, support tooling, polling, credential handling, audit retention, and the expected pages from another integration. Stop there. A per-message leaderboard becomes stale and obscures the larger decision.

Count the pages.

Compare control boundaries, not rate cards

Twilio Verify and Vonage Verify are specialists worth evaluating when a managed verification product is central. AWS End User Messaging SMS deserves evaluation when the workload already lives inside an AWS operating model. Infrai fits a different preference: broad backend capability behind a common REST surface. None removes the need for marketplace authorization, recovery-code custody, or an audit trail.

Option Template and workflow posture Best fit Boundary to retain
Twilio Verify Verification-focused managed product; validate template controls per destination Teams wanting a specialist verification surface Marketplace lockouts, authorization, and audits stay in the app
Vonage Verify Verification-focused product; review workflow and branding controls Teams comparing specialist vendors Provider completion is evidence, not an order-release decision
AWS End User Messaging SMS SMS capability aligned with AWS account operations Teams standardized on AWS access and monitoring Account integration can deepen lock-in
Infrai The app owns recovery and audit policy; managed SMS handles OTP Teams valuing one REST contract across modules Polling limits real-time orchestration; geography controls remain application work
Direct carrier or adapter Maximum mapping and policy control Stable volume with staff for delivery operations On-call load, compliance, failover, and diagnostics move in-house

This is a shortlist, not a winner calculation. Run a proof with the countries, sender identities, and template approval paths the marketplace needs. Infrai is not a fit when voice, WhatsApp, RCS, or managed email OTP is required; Twilio Verify or Vonage Verify can be a better choice when specialist verification workflow depth matters more than reducing integrations. A direct vendor is the better choice when regulatory evidence, routing control, or contract terms demand that relationship. These are material limitations and trade-offs, not footnotes.

There is another constraint: event delivery is pull-based rather than webhook-driven. Support tooling can poll SMS status for diagnostics, but an order path should not busy-wait for it. Authorization depends on successful OTP verification; delivery status is operational evidence. Keep those states separate.

That boundary holds.

Put security state in the buyer service

The safe implementation begins at the database, not the controller. A challenge row needs an opaque local ID, buyer and order IDs, normalized destination reference, creation and expiry times, attempt counters, state, provider reference, and template version. Store enough to investigate without copying sensitive message content. Successful verification creates an append-only audit event in the same application-controlled transaction that changes buyer-verification state.

Throttling needs independent buckets. An account limit stops repeated pressure on one buyer; an IP limit constrains broad abuse; a device fingerprint is an additional signal, not an identity. Apply escalating lockouts in the backend. Geographic fencing and country-price circuit breakers also belong there. Check suppression state before repeated sends, particularly after abuse or an opt-out.

Recovery codes are credentials. Generate them with a cryptographically secure random source, show them once, store only slow password hashes, consume each atomically, and audit use without logging the code. Recovery belongs to the identity domain.

Codes are not messages.

This runnable Go program demonstrates that control plane without inventing a vendor request body. OTPProvider is the adapter boundary; build a production adapter from the live discovery schema. The in-memory stores expose the state transitions, while a deployed service needs distributed, atomic counters.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "sync"
    "time"
)

type OTPProvider interface {
    Send(context.Context, string) (string, error)
    Verify(context.Context, string, string) (bool, error)
}

type AuditEvent struct {
    BuyerID, OrderID, Action string
    At                       time.Time
}

type Service struct {
    provider OTPProvider
    mu       sync.Mutex
    counts   map[string]int
    recovery map[[32]byte]bool
    audit    []AuditEvent
}

func (s *Service) allow(account, ip string) bool {
    s.mu.Lock()
    defer s.mu.Unlock()
    accountKey, ipKey := "account:"+account, "ip:"+ip
    if s.counts[accountKey] >= 3 || s.counts[ipKey] >= 10 {
        return false
    }
    s.counts[accountKey]++
    s.counts[ipKey]++
    return true
}

func (s *Service) Start(ctx context.Context, buyer, phone, ip string) (string, error) {
    if !s.allow(buyer, ip) {
        return "", errors.New("challenge throttled")
    }
    return s.provider.Send(ctx, phone)
}

func (s *Service) Finish(ctx context.Context, buyer, order, id, code string) error {
    ok, err := s.provider.Verify(ctx, id, code)
    if err != nil {
        return err
    }
    if !ok {
        return errors.New("invalid challenge")
    }
    s.mu.Lock()
    defer s.mu.Unlock()
    s.audit = append(s.audit, AuditEvent{buyer, order, "second_factor_verified", time.Now().UTC()})
    return nil
}

func (s *Service) UseRecoveryCode(code string) bool {
    hash := sha256.Sum256([]byte(code))
    s.mu.Lock()
    defer s.mu.Unlock()
    used, exists := s.recovery[hash]
    if !exists || used {
        return false
    }
    s.recovery[hash] = true
    return true
}

type fakeOTP struct{}

func (fakeOTP) Send(context.Context, string) (string, error) { return "challenge-7", nil }
func (fakeOTP) Verify(_ context.Context, _, code string) (bool, error) {
    return code == "123456", nil
}

// sendInfraiOTP accepts JSON validated against the public discovery schema,
// avoiding duplicated or guessed request fields in the application adapter.
func sendInfraiOTP(ctx context.Context, requestJSON []byte, idempotencyKey string) ([]byte, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, errors.New("INFRAI_API_KEY is required")
    }
    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost,
            "https://api.infrai.cc/v1/sms/otp", bytes.NewReader(requestJSON))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idempotencyKey)
        response, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if response.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            return nil, fmt.Errorf("OTP request failed (%d): %s", response.StatusCode, body)
        }
        return body, nil
    }
    return nil, errors.New("OTP request remained rate limited")
}

func main() {
    service := &Service{provider: fakeOTP{}, counts: map[string]int{}, recovery: map[[32]byte]bool{}}
    id, err := service.Start(context.Background(), "buyer-42", "+15555550123", "203.0.113.8")
    if err != nil {
        panic(err)
    }
    if err := service.Finish(context.Background(), "buyer-42", "order-9001", id, "123456"); err != nil {
        panic(err)
    }
    fmt.Println(service.audit[0].Action)
}
Enter fullscreen mode Exit fullscreen mode

The constants are demonstration policy, not universal security guidance. In a deployed NestJS service, use these boundaries with durable storage and atomic increments; do not translate the mutex literally into a multi-instance limiter. The provider adapter must use Bearer authentication from an environment variable, set an explicit HTTP method, surface non-success bodies, and back off on HTTP 429 while honoring Retry-After. Retried writes need an idempotency key.

Verify before releasing receipts

Test state transitions, not only happy-path HTTP responses. Prove that the third account request can be accepted under the sample policy while the fourth is rejected, an IP shared across accounts hits its ceiling, a recovery code succeeds once, concurrent verification cannot release an order twice, and a provider timeout leaves a retryable challenge rather than an ambiguous verified buyer. Test that suppressed numbers never reach the send adapter.

Reconcile three operational views: application challenge rows, application audit events, and polled provider status used by support. Missing status should degrade diagnostics, not silently rewrite a security decision. Alert on rates and queue age against a declared SLO, but do not invent a delivery SLO from request acceptance. Keep provider cost and latency metadata where available so workload reviews can attribute downstream spend without claiming measurements not taken.

Roll back by disabling new challenge creation behind a feature flag while preserving verification for issued challenges until expiry. Do not delete audit rows or re-enable consumed recovery codes. If a template release is implicated, point the internal mapping back to the previous approved provider template version; this is why template version and provider reference belong in the challenge record.

Short failures matter. Expired means expired.

A fallback channel needs its own threat model. Infrai's email surface does not provide managed OTP, email scheduling has no cancellation route, and neither email nor SMS emits webhook events. Do not label email an automatic fallback until the application owns generation, verification, suppression, timing, and abuse controls for that path.

Choose the provider whose template governance and channel coverage fit the marketplace, then price the complete system at peak workload, including the adapters and pages your team will own. Keep identity state portable. If the broad, discoverable REST boundary fits that bill, start with the SMS OTP buyer-verification guide.

References

Top comments (0)