Treat passwordless phone login as a small SMS OTP state machine, not as two Express or Node.js handlers named "send" and "check." For an edtech contact form, the useful decision rule is to authenticate control of the phone number first and route the submitted issue second; keep cooldowns, send quotas, max attempts, and one-time consumption in one durable record.
Short answer: issue a random six-digit code, store only a keyed digest, expire it quickly, allow one verification, and update every counter atomically. Return the same public response for known and unknown phone numbers. A request for another code must invalidate the previous code, but it must not reset the failure budget. That last detail closes an easy loop in which an attacker alternates new codes and guesses forever.
One code. One use.
The data flow is compact. A learner or instructor submits a phone number plus a support category. The service normalizes the number, evaluates anti-abuse controls, creates an OTP challenge, and asks an asynchronous delivery adapter to send it. Successful verification consumes that challenge and releases the already-stored form to a deterministic queue such as account access, billing, or course support. Routing does not belong in the SMS callback, and message delivery does not belong in the routing rule. A food delivery courier login has the same security states, although its post-verification action opens a courier session rather than releasing an edtech support form. Express/Node.js changes the HTTP layer, not the datastore conditions shown below.
How Should Passwordless Phone Login Handle SMS OTP Cooldowns?
The repeat-send flow is safe if it rotates a challenge under a bounded policy rather than creating fresh permission to guess. There are four separate clocks and counters to model: OTP expiry, minimum cooldown, sends within a rolling or fixed window, and failed verification attempts. Combining them into one counter loses information and makes production behavior hard to explain.
A practical starting policy for this example is a 10-minute OTP lifetime, a 60-second cooldown, no more than five sends per phone number in an hour, and five failed checks for the active challenge. Those are example configuration values, not universal security constants. Tune them with an evaluation set that covers delayed messages, impatient double-clicks, shared devices, recycled phone numbers, and deliberate high-rate traffic. The important invariant is stable: issuing another code rotates the secret while retaining the window-level pressure.
The trade-off is explicit. A longer cooldown suppresses bursts but makes a delayed SMS more frustrating; a shorter one improves recovery while increasing delivery volume and attack surface. There is no honest universal value, so keep these four settings in configuration and evaluate them against a written abuse suite before changing them.
SMS can also split into multiple segments depending on encoding and length. Keep the verification message short and avoid user-controlled text in it; besides making the prompt easier to recognize, this makes segment behavior more predictable. The message should identify the purpose, include the code and expiry, and never ask the recipient to reply with a password. GSM-7 messages have different segment limits from messages containing UCS-2 characters, so a harmless-looking change in copy or localization can alter delivery shape; the referenced segmentation guide gives the exact limits.
A runnable challenge store
This example uses Python 3.11 and SQLite so the transaction boundary is visible. In a deployed service, the same compare-and-update operations can live in a relational database or another store that provides atomic conditional writes. The delivery function is deliberately an interface: replacing an SMS gateway should not rewrite authentication policy.
import hashlib
import hmac
import secrets
import sqlite3
import time
from dataclasses import dataclass
OTP_TTL_SECONDS = 10 * 60
RESEND_COOLDOWN_SECONDS = 60
SEND_WINDOW_SECONDS = 60 * 60
MAX_SENDS_PER_WINDOW = 5
MAX_VERIFY_ATTEMPTS = 5
@dataclass(frozen=True)
class IssueResult:
accepted: bool
retry_after_seconds: int = 0
def normalize_phone(raw: str) -> str:
phone = raw.strip().replace(" ", "").replace("-", "")
if not phone.startswith("+") or not phone[1:].isdigit():
raise ValueError("Use an E.164-style number such as +15551234567")
return phone
def code_digest(secret: bytes, challenge_id: str, code: str) -> str:
payload = f"{challenge_id}:{code}".encode()
return hmac.new(secret, payload, hashlib.sha256).hexdigest()
def init_db(db: sqlite3.Connection) -> None:
db.execute(
"""
CREATE TABLE IF NOT EXISTS otp_challenges (
phone TEXT PRIMARY KEY,
challenge_id TEXT NOT NULL,
digest TEXT NOT NULL,
expires_at INTEGER NOT NULL,
last_sent_at INTEGER NOT NULL,
window_started_at INTEGER NOT NULL,
sends_in_window INTEGER NOT NULL,
failed_attempts INTEGER NOT NULL,
consumed_at INTEGER
)
"""
)
def issue_code(
db: sqlite3.Connection, secret: bytes, raw_phone: str, now: int | None = None
) -> tuple[IssueResult, str | None]:
phone = normalize_phone(raw_phone)
now = int(time.time()) if now is None else now
with db:
row = db.execute(
"SELECT * FROM otp_challenges WHERE phone = ?", (phone,)
).fetchone()
if row is not None:
columns = [item[0] for item in db.execute(
"SELECT * FROM otp_challenges LIMIT 0"
).description]
current = dict(zip(columns, row))
elapsed = now - current["last_sent_at"]
if elapsed < RESEND_COOLDOWN_SECONDS:
return IssueResult(False, RESEND_COOLDOWN_SECONDS - elapsed), None
if now - current["window_started_at"] >= SEND_WINDOW_SECONDS:
window_started_at, sends = now, 0
else:
window_started_at = current["window_started_at"]
sends = current["sends_in_window"]
if sends >= MAX_SENDS_PER_WINDOW:
retry = SEND_WINDOW_SECONDS - (now - window_started_at)
return IssueResult(False, max(1, retry)), None
failed_attempts = current["failed_attempts"]
else:
window_started_at, sends, failed_attempts = now, 0, 0
challenge_id = secrets.token_hex(16)
code = f"{secrets.randbelow(1_000_000):06d}"
digest = code_digest(secret, challenge_id, code)
db.execute(
"""
INSERT INTO otp_challenges VALUES (?, ?, ?, ?, ?, ?, ?, ?, NULL)
ON CONFLICT(phone) DO UPDATE SET
challenge_id = excluded.challenge_id,
digest = excluded.digest,
expires_at = excluded.expires_at,
last_sent_at = excluded.last_sent_at,
window_started_at = excluded.window_started_at,
sends_in_window = excluded.sends_in_window,
failed_attempts = excluded.failed_attempts,
consumed_at = NULL
""",
(
phone,
challenge_id,
digest,
now + OTP_TTL_SECONDS,
now,
window_started_at,
sends + 1,
failed_attempts,
),
)
return IssueResult(True), code
def verify_code(
db: sqlite3.Connection,
secret: bytes,
raw_phone: str,
submitted_code: str,
now: int | None = None,
) -> bool:
phone = normalize_phone(raw_phone)
now = int(time.time()) if now is None else now
with db:
row = db.execute(
"""
SELECT challenge_id, digest, expires_at, failed_attempts, consumed_at
FROM otp_challenges WHERE phone = ?
""",
(phone,),
).fetchone()
if row is None:
return False
challenge_id, expected, expires_at, failures, consumed_at = row
if consumed_at is not None or now >= expires_at or failures >= MAX_VERIFY_ATTEMPTS:
return False
candidate = code_digest(secret, challenge_id, submitted_code)
if not hmac.compare_digest(candidate, expected):
db.execute(
"UPDATE otp_challenges SET failed_attempts = failed_attempts + 1 WHERE phone = ?",
(phone,),
)
return False
updated = db.execute(
"""
UPDATE otp_challenges SET consumed_at = ?
WHERE phone = ? AND consumed_at IS NULL AND expires_at > ?
""",
(now, phone, now),
)
return updated.rowcount == 1
The caller sends the returned code only after the transaction commits. In a fuller design, use an outbox row committed with the challenge and let a worker deliver it. That avoids the ambiguous case where a provider accepts a message but the application transaction later fails. Never log the plaintext code; logging the phone number also deserves redaction or tokenization because operational logs tend to have a wider audience than the authentication store.
There is one intentional trade-off here: failed attempts survive another-code requests during the one-hour window. That is stricter for a legitimate user who mistypes, but it prevents the send flow from becoming an attempt-reset button. A production policy may restore eligibility after a longer lock period or a different recovery check. Make that transition explicit and test it rather than quietly deleting the row.
That distinction matters.
Keep abuse controls outside the happy path
Per-number limits are necessary but incomplete. An attacker can rotate numbers while sending from one network, or distribute traffic across networks toward one number. Evaluate several signals at issuance time: normalized phone number, source network, device or session identifier, account identifier when available, and aggregate delivery volume. Do not collapse them into a permanent device fingerprint; retain only what the abuse decision and audit policy require.
Responses need discipline too. The public API can say, "If this number can receive verification messages, a code will arrive," regardless of account state. Internally, record a reason code such as cooldown, phone_window, attempt_limit, or delivery_rejected. Uniform public wording reduces account discovery, while specific internal labels keep dashboards useful.
Race it.
Two repeat-send requests arriving together must not both pass a read-then-write cooldown check. The sample serializes a database write transaction, but a production test harness should launch concurrent requests against the actual datastore and assert that only one challenge becomes deliverable. Also test an old code after a new issue, the same valid code twice, the fifth and sixth bad guesses, and verification exactly at expiry. Then run the identical cases through the Express/Node.js boundary if that is where the HTTP layer lives: the language changes, but the max-attempt and cooldown invariants do not. Boundary tests catch more than another polished prompt ever will.
Do not let an AI classifier decide authentication. It can help classify the verified contact text into likely support categories, provided the system evaluates accuracy and has a low-confidence fallback. The security state machine, quotas, and queue permissions should remain deterministic. Run classification once on the submitted issue, after verification, rather than on every send or code check.
Route only after verification
Authentication answers one narrow question: did this session demonstrate access to the phone number? It does not prove identity, entitlement, or the truth of the contact-form claims. After a successful check, attach the challenge ID and verification timestamp to the pending form, then apply a small routing table to trusted form fields. Free text can inform triage, but it should not be allowed to name an internal queue directly.
For example, account_access can go to an identity-support queue, payment_question to billing operations, and an unknown category to general support. Preserve the original category and the chosen route so an evaluator can compare expected and actual decisions later. If classification is added, build a labeled set from reviewed tickets, track abstentions as well as correct routes, and pin the prompt and model configuration used for each evaluation run. A notebook can discover a useful rule; the deployable artifact needs versioned inputs, deterministic fallbacks, and replayable results.
Channel fallback also needs a product decision. Email may be appropriate for a case receipt or a recovery path, but it should have its own verification and abuse policy rather than inheriting SMS state. Email and SMS have different delivery mechanisms, message constraints, and failure signals. Keep both behind narrow adapters and make the support workflow depend on outcomes such as accepted, temporarily_blocked, and verified, not provider-specific response bodies.
Operational readiness is part of the feature
Before release, walk the state transitions in prose with support and security teams. A user requests a code, waits through the cooldown, asks for another, enters an old code, enters a wrong current code, then enters the right one. State exactly which counters change at each step and what the user sees. Repeat the exercise for an expired challenge and for a delivery rejection. If the team cannot narrate the state, the implementation will be difficult to operate. For the contact form, add three route assertions: a verified account-access issue reaches identity support, a verified payment question reaches billing operations, and an unverified issue reaches neither. For the food delivery variant, assert that a courier session is minted only after consumption succeeds. Those checks belong beside the OTP tests because authentication and the protected action form one user-visible login flow.
Keep the trace boring.
Instrument issuance decisions, delivery outcomes, verification success, expiry, lockout, and time-to-verify without recording secrets. Alert on changes in rates rather than on raw traffic alone, because enrollment periods and assignment deadlines can legitimately shift volume. Keep a kill switch for message issuance that does not disable already-verified support routing. Review retention for challenge rows and abuse signals, and ensure an on-call engineer can distinguish delivery trouble from an application-side throttle.
Finally, evaluate integration effort at the boundary that is hardest to replace. A thin delivery adapter, an explicit challenge repository, and a pure routing function are cheap to test independently. Provider SDK calls scattered through request handlers are not. The winning design is the one the team can replay, inspect, and change without weakening the authentication invariants.
References
- NIST, Digital Identity Guidelines: Authentication and Authenticator Management: https://pages.nist.gov/800-63-4/sp800-63b.html
- OWASP, Authentication Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Authentication_Cheat_Sheet.html
- OWASP, Forgot Password Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Forgot_Password_Cheat_Sheet.html
- Twilio, SMS character limits and segmentation: https://www.twilio.com/docs/glossary/what-sms-character-limit
- Amazon Web Services, Amazon SES documentation: https://docs.aws.amazon.com/ses/latest/dg/Welcome.html
Top comments (0)