Short answer: for an automotive SaaS that sends service updates, issue a short-lived, one-time SMS OTP from a server-side state machine; bind it to one login attempt, rate-limit every relevant identity and network dimension, and make lockout and replay checks explicit. Treat US and EU delivery as untrusted transport, not as proof that a code was used safely.
The bill is not the hard part. Message retention, fraud review, and support evidence are. Keeping every OTP and every message body forever increases breach impact, while keeping nothing makes a dispute impossible to investigate. I retain a salted digest of the code, policy version, timestamps, destination fingerprint, and outcome; I discard the clear code after verification and apply a short, documented retention window to delivery metadata.
Which records should an automotive OTP flow retain?
An order for a service update might trigger a sign-in challenge for a fleet manager in Texas or a driver in Germany. The application creates attempt_id, normalizes the phone number to an E.164 form, and records a purpose such as service_update_login. The SMS provider receives the code, but the database owns the state: issued, verified, expired, or locked.
Store a keyed digest rather than reversible encryption. A random 128-bit attempt identifier, a server-side pepper, and a slow password-hash function make a database snapshot less useful to an attacker. Keep the attempt bound to the account or pending account, destination fingerprint, and session nonce. A code copied from one browser should not authenticate another browser that merely knows the phone number.
The retention decision has a real cost. If you delete delivery receipts immediately, a support engineer cannot distinguish a carrier delay from a rejected code. If you retain message text, a log export becomes a credential archive. My default is to keep event IDs and coarse timestamps for the audit period approved by counsel, then delete payloads and rotate the pepper. Your mileage may vary; the right duration depends on statutory retention and the evidence your incident process requires.
Retention is a security control.
How should secure SMS OTP login use rate limiting, retry, lockout, and replay protection?
Use several buckets, each answering a different abuse question. A per-attempt counter limits guesses against one code. A per-account and per-phone bucket limits repeated sends. An IP or device-risk bucket catches distributed automation. A carrier or country bucket can detect a delivery campaign without assuming that geography is identity. Return the same public response for an existing and nonexistent account so an attacker cannot enumerate customers.
The retry policy needs two clocks. The resend clock controls how often a new message can be requested; the verify clock controls how long the current code remains valid. Issuing a new code should invalidate the previous digest, increment a challenge version, and preserve the attempt's audit trail. Verification must be an atomic compare-and-consume operation: check status, expiry, challenge version, and attempt count, then mark the record used in the same transaction.
Replay should fail even after a successful request is retried by a queue worker. A request token or idempotency key belongs to the send operation, while the OTP belongs to the login attempt. They solve different duplication problems. Never accept a previously verified code because a delivery callback arrived late.
Here is the failure sequence I design against: the worker submits an SMS, times out before seeing the provider response, and retries with the same send key; the carrier delivers both copies; the user enters the first code; a delayed callback arrives after the database has marked the attempt used; then a support replay or an attacker tries the second copy. The send ledger should converge those submissions to one logical message, and the verifier should reject both stale and already-consumed challenges. Keep the callback event ID, attempt ID, and challenge version together so an audit query can explain the order without exposing the message body. If those records are split across systems, a “successful delivery” dashboard can hide a replay path for days.
No shortcuts.
Lockout is a brake, not a punishment. A staged policy can add delay after the fifth failed guess, require a fresh challenge after ten, and place the account in review after a risk threshold. The exact numbers must come from threat modelling and observed abuse; NIST SP 800-63B requires rate limiting for activation secrets and discusses retry limits, but it does not choose your product's threshold. Avoid permanent lockout for a lost phone: offer a separately authenticated recovery path and log that transition.
| Control | Protects against | What to measure | Cost or limitation |
|---|---|---|---|
| Per-attempt guess limit | Online code guessing | Failed verifies by attempt ID | Too low a limit punishes mistyped codes |
| Per-account and phone send limit | SMS pumping and resend abuse | Sends per hour and per destination | Shared family or fleet numbers create false positives |
| IP/device bucket | Distributed automation | Distinct accounts per source | NAT and mobile networks blur identity |
| Challenge version + consume | Replay after success or resend | Reuse attempts and stale-version rejects | Requires a transactional write path |
| Progressive lockout | Sustained attack | Delay, lock, and recovery rates | A hard lock can become denial of service |
Keep the response boring. A public 202-style acknowledgement can say that a challenge was requested without confirming the account, while internal telemetry records the precise reason for suppression. Error text should not reveal whether a phone exists, whether a code was close, or which limit fired; don't let a helpful-looking message become an account-enumeration oracle.
What changes for US and EU SaaS delivery and service updates?
Regional policy belongs in configuration with an owner, version, and effective date. US and EU traffic may use different sender registration, consent, quiet-hour, and data-residency rules; the application should not infer legal permission from a country code alone. Record the consent or transactional basis attached to the service-update notification, and keep marketing opt-out separate from a security login challenge.
The SMS channel itself is not a cryptographic authenticator under your control. SIM swaps, number recycling, malware, and compromised notification previews remain possible. For higher-risk actions, step up to a stronger authenticator or require an already authenticated support workflow. NIST's digital identity guidance treats out-of-band secrets as bounded by the channel and emphasizes verifier-side replay resistance; that is a reason to narrow what an OTP can authorize, not a reason to pretend SMS proves a person is driving the vehicle.
Assume the phone is lost.
Email can be a service-update fallback, but it needs its own sender controls. DKIM, specified in RFC 6376, authenticates a domain signature; it does not fix OTP replay, phone ownership, or an unsafe recovery policy. Keep the same attempt ledger and risk gates across channels, and make a channel switch invalidate the old challenge.
A minimal transactional verifier in Python
The following sketch shows the critical ordering. The database transaction and row lock are the important parts; the hash parameters and limits are policy values that should be versioned and tested.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
import hashlib
import hmac
import secrets
@dataclass
class Challenge:
attempt_id: str
account_id: str
digest: bytes
expires_at: datetime
version: int
failed: int = 0
used: bool = False
def issue(account_id: str, phone_fingerprint: str, pepper: bytes, now: datetime) -> tuple[str, str, bytes]:
code = f"{secrets.randbelow(1_000_000):06d}"
attempt_id = secrets.token_urlsafe(16)
digest = hashlib.sha256(pepper + attempt_id.encode() + code.encode()).digest()
expires = now + timedelta(minutes=5)
# Persist attempt_id, digest, expires, version=1, and phone_fingerprint atomically.
return attempt_id, code, digest
def verify(row: Challenge, attempt_id: str, supplied: str, pepper: bytes, now: datetime) -> bool:
if row.attempt_id != attempt_id or row.used or now >= row.expires_at:
return False
if row.failed >= 5:
return False
expected = hashlib.sha256(pepper + attempt_id.encode() + supplied.encode()).digest()
if not hmac.compare_digest(expected, row.digest):
row.failed += 1
return False
row.used = True # In production, this is a conditional UPDATE in one transaction.
return True
The example intentionally returns no reason to the caller. Production code should enforce the conditional update in the database, attach a policy version, and emit a correlation ID without logging the code. A queue retry may repeat the send request, so its idempotency record must be separate from Challenge.used.
Decide by failure shape, not delivery promises
Choose the smallest design that can explain a failed login six weeks later. A hosted SMS API may reduce integration work, while a direct carrier relationship may provide more control over routing and evidence; neither removes the need for your own ledger, rate limits, and recovery policy. The appropriate choice depends on the assurance level of the action and the jurisdictions in which the SaaS operates.
The catch is that SMS OTP is not suitable for approving irreversible vehicle ownership changes, high-value refunds, or administrator role transfers on its own. Use a phishing-resistant authenticator or a human-reviewed step for those actions. Stick with SMS for bounded login and low-risk service-update access when the fallback and lockout paths are tested.
I would load-test the limiter with bursty fleet traffic, inject delayed and duplicated delivery callbacks, and run a replay corpus against every endpoint that accepts an attempt ID. A green delivery dashboard is not evidence that a code was consumed once. That proof lives in the state transition and its audit record.
Top comments (0)