DEV Community

AndersonBlake6857
AndersonBlake6857

Posted on

Node.js 2FA Login Fallback Favors One Template Over SMS and Email Copies

TL;DR: Keep OTP meaning, variables, and versioning under one authentication-domain template, then let SMS and email adapters render channel-specific layouts. For a Node.js gaming SaaS that cannot consume webhooks, poll SMS delivery for a short, bounded window and switch to email only by advancing the same login challenge. Do not create a second independent code.

That ownership choice matters beyond login. The same gaming platform may email an order receipt after payment settles. A receipt belongs to the order domain; a login code belongs to the authentication domain. Owning both inside an email service feels tidy at first, but it moves business wording, required fields, and release timing away from the systems that understand them. The decision is firm: domain-owned templates beat channel-owned copies when the same event can travel over more than one channel. The boundary is equally firm. Channel adapters still own encoding, subject lines, SMS length constraints, and transport metadata. They do not own what the message means.

How should 2FA login switch from SMS to email OTP fallback?

Picture the before state. The SMS service owns login_code_sms_v4. The email service owns login_code_email_v7. Each has its own variables and deployment schedule. When SMS polling reaches a fallback condition, the application asks email to create another message. Now the team has two templates, two audit trails, and a tempting path to two independently issued codes.

The after state is shorter: login challenge -> central message definition -> SMS renderer -> bounded status polling -> email renderer, if policy permits. One challenge remains authoritative throughout. The channel changes; the security decision does not.

One challenge. One expiry.

This is also the clean boundary for the payment flow: settled order -> receipt definition -> email renderer. The receipt can contain order identifiers and line items because its owner is the order domain. The OTP definition cannot quietly accumulate purchase data just because both messages use email.

OWASP's forgot-password guidance provides a useful floor for code handling: tokens or codes should be cryptographically random, sufficiently long, stored securely, single-use, and expired after an appropriate period. It also recommends rate limiting and consistent responses so an account cannot be enumerated. Those properties belong to the challenge state machine, not to either delivery adapter.

Decision One domain-owned template Separate channel-owned copies
Semantic fields Defined once by authentication Can drift between SMS and email
Channel layout Rendered by each adapter Bundled with business meaning
Version activation One challenge version Two release schedules
Best boundary Multi-channel OTP fallback Unrelated, channel-specific messages

The polling state machine in Node.js

The copyable unit below is intentionally vendor-neutral. A transport returns an opaque message ID and exposes status lookup. The authentication store performs a compare-and-set transition, so two app instances cannot both activate fallback. Production storage must make that transition atomic.

import { createHash, randomInt } from "node:crypto";

type DeliveryState = "queued" | "sent" | "delivered" | "failed" | "unknown";
type ChallengeState = "sms_pending" | "email_pending" | "verified" | "expired";

type LoginChallenge = {
  id: string;
  userId: string;
  codeHash: string;
  expiresAt: number;
  state: ChallengeState;
  smsMessageId: string;
  templateVersion: "login-otp-v1";
};

interface SmsTransport {
  getStatus(messageId: string): Promise<DeliveryState>;
}

interface EmailTransport {
  send(input: {
    to: string;
    templateVersion: LoginChallenge["templateVersion"];
    variables: { code: string; expiresInMinutes: number };
  }): Promise<{ messageId: string }>;
}

interface ChallengeStore {
  get(id: string): Promise<LoginChallenge>;
  transition(
    id: string,
    from: ChallengeState,
    to: ChallengeState,
  ): Promise<boolean>;
}

const hashCode = (challengeId: string, code: string): string =>
  createHash("sha256").update(`${challengeId}:${code}`).digest("hex");

export async function pollThenFallback(input: {
  challengeId: string;
  email: string;
  code: string;
  sms: SmsTransport;
  mail: EmailTransport;
  store: ChallengeStore;
  now?: () => number;
}): Promise<"wait" | "emailed" | "stop"> {
  const now = input.now ?? Date.now;
  const challenge = await input.store.get(input.challengeId);

  if (challenge.state !== "sms_pending" || challenge.expiresAt <= now()) {
    return "stop";
  }
  if (hashCode(challenge.id, input.code) !== challenge.codeHash) {
    return "stop";
  }

  const delivery = await input.sms.getStatus(challenge.smsMessageId);
  if (delivery === "queued" || delivery === "sent" || delivery === "unknown") {
    return "wait";
  }
  if (delivery === "delivered") {
    return "stop";
  }

  const won = await input.store.transition(
    challenge.id,
    "sms_pending",
    "email_pending",
  );
  if (!won) return "stop";

  const expiresInMinutes = Math.max(
    1,
    Math.ceil((challenge.expiresAt - now()) / 60_000),
  );
  await input.mail.send({
    to: input.email,
    templateVersion: challenge.templateVersion,
    variables: { code: input.code, expiresInMinutes },
  });
  return "emailed";
}

export const createNumericCode = (): string =>
  randomInt(0, 1_000_000).toString().padStart(6, "0");
Enter fullscreen mode Exit fullscreen mode

The example reuses one six-digit code only while the challenge remains live. That is a deliberate trade-off: it avoids concurrent credentials during a channel change, but it means both channels carry the same secret. A stricter policy can rotate the code during the atomic transition, invalidate the earlier hash, and update the central template variables before email is sent. That choice should be made in the authentication domain and covered by a race test. Do not poll forever. Use a capped schedule with jitter, stop at the challenge expiry, and treat an unavailable status endpoint as unknown rather than as proof of failure. A transport-level sent state is not evidence that the player received or read a message. Fast fallback may improve access, yet it also gives an attacker another route to the same challenge. The policy has to balance those outcomes explicitly. This is the concrete tension: a shorter polling window opens the email path sooner, while a longer one gives SMS more time to reach a terminal state. Neither interval can be selected from transport status alone; authentication risk, challenge lifetime, measured regional delivery behavior, and support policy all matter.

Make the template contract observable

Logs should describe transitions without recording the code, its hash, an email address, or a phone number. Useful fields are challenge_id, template_version, channel, attempt, delivery_state, transition, and a coarse region such as us or eu. Keep user_id out unless the logging policy permits a protected pseudonymous identifier. Metrics answer a different question. Count challenges started, terminal SMS states, fallback transitions, successful verifications by final channel, expirations, polling errors, and compare-and-set losses. Measure latency from challenge creation to verification, not merely provider acceptance. The crisp before/after is operational too: before, dashboards group two unrelated template names; after, both channels share login-otp-v1 and one challenge ID.

Alert on symptoms a user feels: a sustained rise in expirations, a missing polling worker heartbeat, or a sharp shift in fallback ratio for one region. A single failed send is an event, not a page.

Template deployment needs the same discipline. Validate the central variable schema in CI. Render SMS and email snapshots from the same fixture. Then canary a version and retain the prior renderer until active challenges using it have expired. For the gaming example, keep fixtures separate: login-otp-v1 accepts a code and expiry; order-receipt-v3 accepts a settled order. Accidental field sharing should fail the build.

What about compliance and a second delivery path?

Email fallback is not a reason to add promotional copy to a security message or receipt. In the United States, the FTC explains that CAN-SPAM distinguishes transactional or relationship content from commercial content based on the message's primary purpose. Mixing a coupon into an OTP or settled-order receipt muddies that purpose. Keep operational templates focused, and have counsel review the actual content and sending practice. US and EU deployment also creates a data-handling question that transport code cannot answer alone: where are destination addresses, delivery events, and logs processed and retained? Record that decision in the adapter configuration and data inventory. Do not infer regional compliance from an endpoint label. The evidence must come from contracts, system configuration, and legal review. A second channel expands the attack surface. Email account compromise, SIM-related attacks, mailbox forwarding, and recovery flows have different risks. Fallback should therefore obey a server-side policy based on the account and risk context; it should not appear as an unlimited "send another way" button. Rate limits must cover the challenge across both channels, not reset when the renderer changes.

Two objections worth resolving before launch

The first objection is that channel teams need freedom to tune copy. They do. Give each adapter ownership of presentation rules and a review path for channel constraints, while the authentication domain owns semantic fields, expiry language, and version activation. This split prevents a central template from becoming a lowest-common-denominator blob.

The second objection is that no webhooks means delivery cannot be reliable. Polling can support a bounded decision, provided the provider exposes a trustworthy status resource and the application models uncertainty. It cannot prove human receipt. Test queued, delivered, terminal failure, unknown, timeout, expiry, duplicate worker, and late status change cases. The race test is the one teams skip: run two pollers against the same failed SMS and assert that exactly one email transition wins.

For a Node.js gaming SaaS, the final rule is compact. Put authentication meaning and versioning in the authentication domain. Put settled-order receipt meaning in the order domain. Let shared channel adapters render and transport both. When SMS status is polled, advance one atomic challenge before email fallback, preserve one expiry and one rate-limit budget, and observe the state transition rather than the secret.

References

Top comments (0)