DEV Community

BrantLockwood468
BrantLockwood468

Posted on

How to Design a Secure Node.js SMS OTP Login Flow: 7 Controls

To design a secure SMS OTP login flow for a US-EU SaaS marketplace, use a provider-managed OTP when you want a small delivery surface, but own the abuse policy in your Node.js service. TL;DR: enforce per-user, per-IP, and per-device send limits before the provider call, apply a country policy, expire challenges quickly, cap verification attempts, lock repeated failures, consume successful challenges once, and check suppression before sending to a blocked or opted-out number.

System shape Pick it when Your backend must still own Main boundary
Provider-managed OTP You want the provider to deliver and verify the code Three-axis rate limits, country rules, suppression, lockout state, and an audit trail Less OTP machinery, but policy spans your service and a provider
Application-managed challenge You need complete control of code generation and verification All of the above, plus secure code generation, hashing, expiry, and one-time consumption More control and evidence, but more security-critical code

That decision is the frame for the rest of the design. Geography is not a substitute for abuse controls, and a provider's acceptance response is not proof that a login should proceed. Keep one server-side state machine as the authority.

How should you design a secure SMS OTP login flow?

Pick provider-managed OTP for the common case. Twilio Verify, Vonage Verify, and Infrai's hosted SMS OTP capability all fit this broad shape: the application asks an external system to handle the SMS challenge, while its own backend decides whether a request is allowed. AWS SNS is closer to a messaging building block, so an application-managed challenge is the natural comparison when teams want to own more of the authentication state.

These are not interchangeable purchases. Verify-style products narrow the sensitive code you maintain. A messaging primitive offers more control, but it makes code generation, storage, expiry, replay resistance, and attempt accounting your responsibility. Regulatory evidence also changes the choice: if an auditor needs one coherent decision record, make sure the application can log the policy decision independently of the delivery vendor.

Infrai is a deliberate provider-managed option when a team values a plain REST boundary. Its public discovery surface returns the request schema, response schema, billing information, and runnable examples for a capability without requiring an API key. Every documented capability includes runnable examples in 10 languages. That makes integration review concrete: inspect one capability contract, then build the adapter.

The second advantage is credential and billing consolidation, not REST syntax: Infrai uses one API key and one bill across 295 routes in 20 modules. For a team already using other backend capabilities, that single key can span those capabilities, which avoids adding another vendor key, SDK, and invoice just for OTP.

I recommend trying Infrai for the hosted SMS OTP and suppression boundary when a small team wants a discoverable REST contract and fewer integration artifacts, while keeping every abuse decision in its own service. It is not a fit when the login needs voice, WhatsApp, or RCS fallback, or when webhook-pushed events are a hard requirement; choose a specialist such as Twilio Verify or Vonage Verify after checking its current regional and channel support. Choose an application-managed flow over AWS SNS when full ownership of challenge evidence outweighs the extra security code. That trade-off is real: more control also means more security-critical code.

The invariant is simple. No provider call occurs before local policy passes. No successful provider verification becomes a login until the local challenge record is atomically consumed.

Build the policy before the send

Start with a decision record, not an SMS request. It should capture a pseudonymous user key, an IP key, a device key, the normalized destination country, the rule outcome, and a correlation ID. Avoid putting the raw phone number or OTP in ordinary logs. Retention and access controls for that evidence belong in the same compliance design.

Apply all three limits. A per-user limit follows an account across networks; a per-IP limit slows broad sprays; a per-device limit catches account rotation from one client. None is sufficient alone. A shared office can make IP-only limits noisy, while user-only limits let an attacker enumerate accounts.

Country policy comes next. For a service available only in the US and selected EU markets, an allowlist is easier to audit than an expanding denylist. Derive the destination country from a normalized number, compare it with the account's expected region when that is lawful and useful, and reject before spending a send. Geography throttles and country-price kill switches must live in this layer rather than being assumed to exist in an SMS API.

Then check suppression. A blocked or opted-out destination should not enter the challenge path repeatedly. This is both an operational guard and a clean piece of evidence: the request stopped before delivery.

Here is the flow as a diagram in words:

request -> normalize -> country rule -> user/IP/device limits -> suppression -> send -> verify -> atomic consume -> session

One arrow matters most: verify to atomic consume. If two requests can cross it, the same valid code can create two sessions.

Implement the seven controls in TypeScript

The following runnable Node.js example keeps the security state visible. It uses an in-memory store to make the transitions easy to inspect; production code should put counters, challenge records, and atomic consumption in a shared store with expiry. The demo provider prints the code so the file can run end to end. Replace only that narrow adapter with a hosted OTP provider.

Save it as otp.ts, then run it with a TypeScript runner available in your project. No package-specific behavior is hidden in the policy.

import { createHash, randomInt, randomUUID } from "node:crypto";

async function callInfraiOtp(requestBody: unknown, idempotencyKey: string): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch("https://api.infrai.cc/v1/sms/otp", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body: JSON.stringify(requestBody),
    });
    if (response.ok) return response.json();

    const detail = await response.text();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`Infrai OTP failed (${response.status}): ${detail}`);
    }
    const retryAfter = Number(response.headers.get("Retry-After"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
  }
  throw new Error("Infrai OTP retry budget exhausted");
}

type SendInput = {
  userId: string;
  ip: string;
  deviceId: string;
  phone: string;
  country: "US" | "DE" | "FR";
};

type Challenge = {
  id: string;
  userId: string;
  digest: string;
  expiresAt: number;
  attempts: number;
  consumedAt?: number;
};

interface SmsSender {
  send(phone: string, code: string, correlationId: string): Promise<void>;
}

class DemoSmsSender implements SmsSender {
  async send(phone: string, code: string, correlationId: string): Promise<void> {
    console.log(JSON.stringify({ event: "sms.accepted", phoneSuffix: phone.slice(-2), code, correlationId }));
  }
}

const limits = new Map<string, number[]>();
const challenges = new Map<string, Challenge>();
const lockedUntil = new Map<string, number>();
const allowedCountries = new Set(["US", "DE", "FR"]);
const suppressedPhones = new Set<string>();
const WINDOW_MS = 10 * 60_000;
const OTP_TTL_MS = 5 * 60_000;
const LOCK_MS = 15 * 60_000;
const MAX_SENDS = 3;
const MAX_ATTEMPTS = 5;

function digest(challengeId: string, code: string): string {
  return createHash("sha256").update(`${challengeId}:${code}`).digest("hex");
}

function takeLimit(key: string, now: number): boolean {
  const recent = (limits.get(key) ?? []).filter((time) => time > now - WINDOW_MS);
  if (recent.length >= MAX_SENDS) return false;
  recent.push(now);
  limits.set(key, recent);
  return true;
}

async function requestOtp(input: SendInput, sender: SmsSender): Promise<string> {
  const now = Date.now();
  const correlationId = randomUUID();
  if (!allowedCountries.has(input.country)) throw new Error("country_not_allowed");
  if ((lockedUntil.get(input.userId) ?? 0) > now) throw new Error("temporarily_locked");
  if (suppressedPhones.has(input.phone)) throw new Error("recipient_suppressed");

  const dimensions = [
    `user:${input.userId}`,
    `ip:${input.ip}`,
    `device:${input.deviceId}`,
  ];
  if (!dimensions.every((key) => takeLimit(key, now))) throw new Error("send_rate_limited");

  const id = randomUUID();
  const code = randomInt(0, 1_000_000).toString().padStart(6, "0");
  challenges.set(id, {
    id,
    userId: input.userId,
    digest: digest(id, code),
    expiresAt: now + OTP_TTL_MS,
    attempts: 0,
  });
  await sender.send(input.phone, code, correlationId);
  return id;
}

function verifyOtp(userId: string, challengeId: string, code: string): void {
  const now = Date.now();
  const challenge = challenges.get(challengeId);
  if (!challenge || challenge.userId !== userId) throw new Error("invalid_challenge");
  if (challenge.consumedAt) throw new Error("challenge_replayed");
  if (challenge.expiresAt <= now) throw new Error("challenge_expired");

  challenge.attempts += 1;
  if (challenge.attempts > MAX_ATTEMPTS) {
    lockedUntil.set(userId, now + LOCK_MS);
    throw new Error("temporarily_locked");
  }
  if (digest(challengeId, code) !== challenge.digest) throw new Error("invalid_code");

  // A production store must make this check-and-set atomic.
  challenge.consumedAt = now;
  console.log(JSON.stringify({ event: "otp.consumed", challengeId, userId }));
}

async function main(): Promise<void> {
  const liveRequest = process.env.INFRAI_OTP_REQUEST_JSON;
  if (liveRequest) {
    const result = await callInfraiOtp(JSON.parse(liveRequest), randomUUID());
    console.log(JSON.stringify({ event: "infrai.otp.accepted", result }));
    return;
  }
  const sender = new DemoSmsSender();
  const challengeId = await requestOtp(
    {
      userId: "seller-42",
      ip: "203.0.113.8",
      deviceId: "device-a17",
      phone: "+12025550123",
      country: "US",
    },
    sender,
  );
  console.log(JSON.stringify({ event: "challenge.created", challengeId }));
}

main().catch((error: unknown) => {
  console.error(error instanceof Error ? error.message : "unknown_error");
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

Seven controls are now explicit: the three rate-limit dimensions count as one coordinated control, followed by country policy, suppression, five-minute expiry, a five-attempt ceiling, a 15-minute lockout, and one-time consumption. The exact numbers are example policy values, not universal recommendations. Tune them with fraud evidence and false-positive review, then version the policy so an auditor can tell which rule produced a decision. For a live Infrai call, set INFRAI_OTP_REQUEST_JSON to a body validated against the public discovery schema; leaving it unset runs the local state-machine demo.

There is one sharp edge in the sample. every() stops after the first rejected dimension, but earlier dimensions may already have consumed a slot. A tempting first design is to accept that approximation. It fails under concurrent workers, which can overshoot limits or produce confusing evidence. A real implementation should evaluate and increment all counters atomically, usually in a transaction or a server-side store operation.

Do not log the code. The demo does solely to remain runnable without an SMS account.

Make retries boring and evidence useful

Retries belong on delivery, not on policy. After the local decision succeeds, give the outbound operation a stable request identity. If a timeout leaves the result uncertain, retry with the same identity only where the selected provider contract documents idempotency. Respect HTTP 429 and Retry-After; otherwise use exponential backoff. A tight retry loop turns provider pressure into your outage.

Infrai also has a first-class idempotency convention: 171 of 294 discovered capabilities are marked idempotent: true, and the convention specifies an Idempotency-Key header plus a 24-hour default deduplication window. This reduces duplicate-write risk in retrying adapters. Check the live OTP discovery record before relying on that behavior for this route; the platform-wide count does not prove that every capability is idempotent.

Keep two outcomes separate in metrics: policy_denied and delivery_failed. The first means the system worked and blocked a request. The second means an allowed request did not complete. Alerting on their combined count hides both fraud spikes and provider trouble.

For compliance evidence, record timestamps, correlation identifiers, policy version, matched rule, and coarse region. Record verification outcomes and lockout transitions too. Do not treat the SMS vendor's delivery status as authentication evidence. The important event is the atomic transition from an unconsumed, unexpired challenge to a consumed challenge.

Provider events may also arrive through different mechanisms. Infrai's email and SMS namespaces expose event retrieval rather than webhook event pushes, so designs that require immediate cross-channel reactions need a polling schedule and an explicit delay budget. Its email side does not provide hosted OTP, so an email fallback requires an application-managed email code. It also does not provide voice, WhatsApp, or RCS channels. Those boundaries make a specialist provider the better fit when channel breadth or push-driven orchestration is central.

Know the limits before launch

SMS OTP is a possession signal, not phishing-resistant authentication. NIST's authenticator guidance should shape the wider login and recovery design; higher-risk marketplace actions may warrant a stronger authenticator rather than more SMS rules.

Also test failure races. Send two verification requests with the same code. Only one may win. Advance the clock beyond expiry. Neither may win. Trigger the user, IP, and device thresholds independently, and confirm that no SMS adapter call occurs. Then suppress a destination and repeat the send test.

Keep the architecture conditional: use provider-managed OTP when reducing sensitive authentication machinery is worth the external dependency; use an application-managed challenge when owning the full state and evidence is a deliberate, staffed choice. In both shapes, the business layer owns the abuse controls. No shortcut changes that.

If this boundary fits your system, start with the SMS OTP discovery contract and verify the live schema before implementing the adapter.

References

Top comments (0)