Short answer: use a single-use email link as the default password-reset path, offer SMS OTP only after the email path has had a fair delivery window, and keep both channels behind one server-side recovery state machine. For a property-management system that emails generated owner reports as attachments, recovery must restore account access without turning the report message, phone number, or delivery telemetry into a second identity database.
Start with the bill because it exposes the architecture. The variable communication cost is email attempts × email rate + SMS attempts × SMS rate; storage is audit events × bytes × retention time; support adds failed recoveries × handling time. Insert your contracted rates instead of borrowing a public list price. In most real designs, the useful change isn't shaving a few bytes from an audit row. It is preventing retries, duplicate sends, and an SMS fallback that fires before the mailbox has had time to accept the reset message.
The least complex reliable option is therefore one recovery record, one hashed secret at a time, and a channel policy that can be measured. Keep the generated PDF in the authenticated report workflow. Don't attach it to a password-reset email.
How do password reset email links and fallback SMS OTP change cost?
Compare failure boundaries, not the shape of the API call. An email link can be opened on a different device, can be delayed by mailbox filtering, and can be consumed by an automated security scanner before a person clicks. An SMS OTP depends on a previously verified phone number, expires into a short input flow, and introduces carrier delivery plus number-reassignment and SIM-swap risk. Neither channel proves identity by itself; each proves control of a destination already bound to the account.
For US and EU users, geography should select messaging policy, sender configuration, consent evidence, and retention policy. It should not silently weaken the reset rules. GDPR Article 7 requires a controller to demonstrate consent where consent is the legal basis and makes withdrawal as easy as giving consent. Transactional recovery messages still need counsel to identify the correct legal basis; don't reuse a recovery phone number for marketing because the database happens to contain it. I'm not sure which basis applies to your exact tenant and lease workflow. That answer depends on purpose, jurisdiction, and the records your organization actually keeps.
Mail privacy is another reason to avoid treating an open event as delivery or user intent. Apple's Mail Privacy Protection can download remote content in the background and prevent senders from learning whether a recipient opened a message. A reset-link request, token redemption, password change, and authenticated report download are useful security events. A tracking pixel isn't.
No click tracking.
Use a decision table before writing channel code:
| Condition | Action | Why |
|---|---|---|
| Verified email exists | Send one single-use link | Lowest-friction primary path |
| Email is merely slow | Wait for the configured delivery window | Early fallback creates duplicate live challenges |
| Window elapsed and a verified mobile number exists | Offer SMS OTP | A distinct route can recover from mailbox-specific delay |
| No verified second destination | Move to manual recovery | Sending to unverified data expands the attack surface |
| Attempt, destination, or account limit is reached | Return the same public response and suppress the send | Prevent enumeration and retry amplification |
That last row matters. OWASP recommends a consistent message and response time for existing and nonexistent accounts, side-channel delivery, random single-use expiring tokens, and protections against excessive submissions. A public 429 can be appropriate at an edge that limits abusive clients, but the account lookup response must not reveal whether the address exists. Internally, record the reason as a bounded code such as destination_rate_limited; don't leak it into the browser copy.
Turn the channel budget into one recovery transaction
The state machine owns policy; email and SMS are adapters. A minimal transaction can move through requested, email_sent, sms_offered, verified, completed, and expired. Store a hash of the active secret, never the raw link token or OTP. Bind it to a transaction ID, account ID, purpose, expiry, attempt count, and channel. Issuing an SMS challenge invalidates the email secret unless your risk team explicitly accepts two live credentials. I wouldn't accept that complexity for ordinary report access.
The following Python sketch shows the contract. A Node.js service can implement the same transitions with its database transaction primitive and a cryptographically secure random API; the state rules are the portable part.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
import hashlib
import hmac
import secrets
@dataclass
class Challenge:
transaction_id: str
account_id: str
channel: str
secret_hash: str
expires_at: datetime
attempts_left: int
consumed_at: datetime | None = None
def digest(secret: str) -> str:
return hashlib.sha256(secret.encode("utf-8")).hexdigest()
def issue_email_challenge(account_id: str) -> tuple[Challenge, str]:
raw_token = secrets.token_urlsafe(32)
challenge = Challenge(
transaction_id=secrets.token_hex(16),
account_id=account_id,
channel="email_link",
secret_hash=digest(raw_token),
expires_at=datetime.now(timezone.utc) + timedelta(minutes=15),
attempts_left=1,
)
return challenge, raw_token
def verify(challenge: Challenge, candidate: str, now: datetime) -> bool:
if challenge.consumed_at is not None or now >= challenge.expires_at:
return False
if challenge.attempts_left <= 0:
return False
challenge.attempts_left -= 1
valid = hmac.compare_digest(challenge.secret_hash, digest(candidate))
if valid:
challenge.consumed_at = now
return valid
Fifteen minutes and one email-token attempt are example policy values, not universal standards. Put them in reviewed configuration, test their boundary conditions, and make the database update atomic. Two concurrent redemptions must not both observe consumed_at = NULL and proceed. The password update, session revocation decision, and challenge consumption belong in one transaction or in a design with equivalent idempotency guarantees.
The request endpoint should always produce bland public copy such as “If the account can be recovered, instructions will arrive shortly.” Behind it, normalize the identifier, apply limits by account plus destination plus network signal, create an opaque transaction, enqueue the channel message, and return. Keep SMTP or carrier latency out of the request. Fast responses are nice; indistinguishable responses are the security property.
For an OTP, generate it with a cryptographically secure source, store only its hash, cap guesses, and invalidate it after successful use. A six-digit code has only one million possible values, so rate limits and a short expiry are part of the credential rather than optional middleware. Don't log the code. Don't put the email token in analytics URLs either.
def issue_sms_challenge(account_id: str, transaction_id: str) -> tuple[Challenge, str]:
otp = f"{secrets.randbelow(1_000_000):06d}"
challenge = Challenge(
transaction_id=transaction_id,
account_id=account_id,
channel="sms_otp",
secret_hash=digest(otp),
expires_at=datetime.now(timezone.utc) + timedelta(minutes=5),
attempts_left=5,
)
return challenge, otp
Short code. Sharp edges.
Keep report attachment retention out of recovery
After a successful reset, the user signs in and requests the generated property report. That authenticated action creates a report job with an idempotency key. The worker renders the PDF, virus-scans or otherwise validates it according to the organization's file policy, and submits one attachment message. Recovery establishes account control; authorization still decides which building, owner, or tenant records the account may receive.
The attachment path needs its own limits. MIME encoding increases the transmitted size beyond the raw file, while provider and receiving-mailbox limits vary. Reject or switch to an authenticated download before enqueueing a file that exceeds your configured envelope budget. Never discover the limit by retrying the same oversized message. Keep report classification, recipient, content hash, generation version, and delivery state in the job record, but don't retain the PDF forever just because the message log remains useful.
from dataclasses import dataclass
@dataclass(frozen=True)
class ReportMessage:
job_id: str
account_id: str
recipient: str
subject: str
pdf_bytes: bytes
def enqueue_report(message: ReportMessage, max_raw_bytes: int, queue) -> None:
if len(message.pdf_bytes) > max_raw_bytes:
raise ValueError("report exceeds the configured attachment budget")
queue.publish(
topic="report-email",
key=message.job_id,
payload={
"job_id": message.job_id,
"account_id": message.account_id,
"recipient": message.recipient,
"subject": message.subject,
"attachment_sha256": hashlib.sha256(message.pdf_bytes).hexdigest(),
},
)
The example deliberately leaves binary storage behind an internal reference rather than putting PDF bytes on the queue. The worker can claim the job once, send it with a stable message ID, and classify the result. Retry temporary SMTP 4xx responses with bounded exponential backoff and jitter; treat permanent 5xx SMTP replies as terminal for that recipient. Those are protocol reply classes, not web-service outage claims. An SMS fallback belongs to account recovery, not report attachment delivery, because SMS cannot carry the report and an unexpected text can disclose that an account or property workflow exists.
Provider topology changes the adapter boundary. Amazon Web Services documents email through Amazon SES and SMS through AWS End User Messaging SMS; Twilio documents SendGrid for email and Verify for verification channels; Microsoft documents both Email and SMS resources under Azure Communication Services. These are examples of different credential, event, and configuration surfaces, not a ranking. A team already operating one cloud may value consolidated identity controls; a team that needs provider independence may prefer narrow adapters and its own event schema. The catch is that abstraction hides provider-specific delivery signals, so preserve the raw provider event in short-lived restricted storage while mapping it to a small internal status vocabulary.
Put deletion deadlines in the delivery budget
Observability should answer four questions without exposing credentials: Was a challenge created? Was a send accepted by the channel adapter? Was the challenge verified? Was the password actually changed? Track rates and latency between those transitions, segmented by channel and coarse region. Alert on a jump in fallback offers, repeated OTP guesses, queue age, and report jobs that remain nonterminal beyond the service objective. Delivery acceptance is not inbox placement, and an email open isn't verification.
Use structured reason codes. Keep token values, OTPs, full report contents, and reset URLs out of logs. Hash or otherwise pseudonymize destination fields where operational correlation is required, and restrict access to the mapping system. Retention should be purpose-specific: short-lived raw delivery events for diagnosis, longer-lived minimal security events where policy and law require them, and generated reports only as long as the business record schedule demands. The exact periods need legal, security, and operational agreement; your mileage may vary across US states and EU member states.
This is where the earlier cost equation pays off. Deduplicate jobs by transaction ID, suppress sends after completion, and measure the fallback window before changing it. Then deliberately stop keeping raw secrets immediately, provider payloads after their diagnostic window, and report binaries after their records window. The cost is reduced forensic detail when a complaint arrives late. Keep bounded status codes and timestamps long enough to explain the decision path, but accept that data minimization means some old incidents cannot be reconstructed byte for byte.
This design is not suitable when the phone number is unverified, shared by a household, or used as a high-assurance recovery factor. Stick with manual, risk-reviewed recovery for those cases. Email-only recovery is also the better choice when SMS consent, sender registration, or regional delivery coverage isn't established. Conversely, a verified SMS fallback earns its place when measured mailbox delays are causing recoveries to expire and the added attack surface has an explicit owner.
Test the awkward transitions before deployment: duplicate reset requests, a scanner visiting the link, email arriving after SMS issuance, five wrong OTPs, redemption at the exact expiry boundary, two simultaneous valid submissions, a password change followed by a queued retry, and a report job requested twice. Run synthetic delivery probes only to destinations you control. They reveal channel drift without turning customer addresses into test fixtures.
Further reading
- OWASP Forgot Password Cheat Sheet
- NIST SP 800-63B: Authentication and Lifecycle Management
- Apple: Use Mail Privacy Protection on iPhone
- GDPR Article 7: Conditions for consent
- RFC 5321: Simple Mail Transfer Protocol
- Amazon SES documentation
- AWS End User Messaging SMS documentation
- Twilio SendGrid documentation
- Twilio Verify documentation
- Azure Communication Services documentation
Top comments (0)