The hard part of SMS OTP login is not generating six digits. It is deciding what happens when a customer-support user taps “send again” three times, a carrier delivers the first message late, or an attacker targets one phone number all afternoon. For a Node.js backend, the practical choice is a short-lived, single-use challenge with separate send and verify budgets, an explicit retry cooldown, and logs that never contain the code.
Short answer: create one server-side challenge per login attempt, hash the OTP, expire it quickly, allow a small number of verification attempts, and make resend idempotent during a cooldown. Treat SMS as a possession signal, not a standalone identity proof; NIST describes SMS and other out-of-band methods with important limits, so offer a stronger authenticator for sensitive support actions.
The decision matrix for a support desk
| Concern | Conservative default | Why it matters |
|---|---|---|
| Code lifetime | 5 minutes | Late delivery should not create an endless valid window. |
| Code format | Six random decimal digits | Familiar on phones; rate limits carry the security load. |
| Verification attempts | 5 per challenge | Stops guessing without punishing a mistyped code forever. |
| Resend cooldown | 30 seconds | Limits spend and message floods while allowing a genuine retry. |
| Active challenges | One per user and purpose | Prevents an old message from racing a newer state. |
That is a baseline, not a law. A help-desk workflow may need a longer lifetime for rural delivery, while a high-risk password reset should use a stronger factor. Measure delivery latency and failed verification by carrier and region before changing these values.
Five tries. Then stop.
The integration axis is simple: keep the challenge state behind a small interface. Your HTTP handler should not know whether the state lives in Redis, Postgres, or a self-hosted service. It should know the user, purpose, expiry, attempt count, and a digest comparison result. Less glue means fewer places to forget a limit.
How should Node.js send, verify, rate-limit, and cool down SMS codes?
The send endpoint first normalizes the phone number to E.164, checks an account-level and IP-level budget, then looks for an active challenge. During the cooldown it returns the existing challenge metadata without sending another message. That response is deliberately boring. A client can display “try again in 18 seconds” without learning whether the account exists.
The code below is the core state transition. The storage and SMS adapter are intentionally generic; the important contract is that put is atomic with the expiry and that consume cannot succeed twice.
import { createHash, randomInt } from "node:crypto";
type Challenge = {
userId: string;
purpose: "support-login";
digest: string;
expiresAt: number;
nextSendAt: number;
attempts: number;
};
const CODE_TTL_MS = 5 * 60_000;
const RESEND_COOLDOWN_MS = 30_000;
const MAX_ATTEMPTS = 5;
const digest = (code: string) =>
createHash("sha256").update(code).digest("hex");
export async function issueCode(
userId: string,
now = Date.now(),
): Promise<{ sent: boolean; retryAfter?: number }> {
const active = await store.get<Challenge>(userId, "support-login");
if (active && active.nextSendAt > now) {
return { sent: false, retryAfter: active.nextSendAt - now };
}
const code = randomInt(100000, 1000000).toString();
const challenge: Challenge = {
userId,
purpose: "support-login",
digest: digest(code),
expiresAt: now + CODE_TTL_MS,
nextSendAt: now + RESEND_COOLDOWN_MS,
attempts: 0,
};
await store.put(challenge, CODE_TTL_MS);
await sms.send({ template: "support-login", code });
return { sent: true };
}
export async function verifyCode(userId: string, code: string) {
const challenge = await store.get<Challenge>(userId, "support-login");
if (!challenge || challenge.expiresAt <= Date.now()) return { ok: false };
if (challenge.attempts >= MAX_ATTEMPTS) return { ok: false };
challenge.attempts += 1;
await store.put(challenge, challenge.expiresAt - Date.now());
if (digest(code) !== challenge.digest) return { ok: false };
await store.consume(userId, "support-login");
return { ok: true };
}
There are two details here that are easy to miss. Hashing limits the damage of a database read, and consuming after a successful comparison makes the code single-use. In production, compare digests with a constant-time primitive and make the increment plus consume operation transactional; the sample keeps those seams visible instead of pretending an in-memory object is enough.
A retry is not the same as a resend. The client may repeat a timed-out request, so give the send operation an idempotency key and persist the result. Otherwise a mobile reconnect can create two messages while the user sees one spinner.
What fails in real OTP flows, and how do you observe it?
Most incidents are state-machine bugs. A late first message arrives after a second code, a verification route accepts a code after expiry, or a queue retries delivery but the API reports failure. Draw the states: created, cooldown, expired, locked, and consumed. Every transition needs an owner and an event name. When a support agent reports that the “new” code failed, trace the request ID across those transitions: check which challenge digest was active, whether the resend was inside the cooldown, which carrier timestamp arrived first, and whether a worker replayed a delivery event. That investigation is longer than the six-digit code, but it tells you whether to fix client retry behavior, storage atomicity, or carrier routing. Without the timeline, teams tend to raise limits and make the attack surface wider.
Log identifiers, not secrets. A useful event has a request ID, user pseudonym, purpose, region, provider response class, latency, and decision (sent, cooldown, rejected, verified). Never log the raw phone number or OTP. Alert on spikes in send volume, repeated failures for one destination, and a widening gap between accepted sends and delivered messages.
The customer-support angle changes the runbook. Agents need a safe explanation for “I never got it” that does not reveal account existence. They also need a recovery path that is stronger than disabling the limit. A manual override should require an audited, higher-assurance check, with a reason code and a second reviewer for privileged accounts.
Where is this design a poor fit?
SMS is unsuitable when the action protects high-value data, when users commonly lack cellular service, or when regulatory policy requires phishing-resistant authentication. Use a passkey or a hardware-backed authenticator for those cases. Stick with email links or an authenticator app when the user already has a verified channel and delivery speed matters more than phone reach.
It is also a poor fit for bulk customer notifications. OTP traffic needs per-user state, strict expiry, and abuse controls; a campaign system has different queues and opt-out rules. Mixing the two makes both systems harder to reason about.
Your mileage may vary on the exact numbers. I would tune them from delivery percentiles and support tickets, then re-run abuse tests after every change. The invariant is stable: one purpose, one active challenge, bounded attempts, and an auditable decision at each step.
Top comments (0)