Short answer: treat an SMS OTP as an authentication attempt with an evidence trail, not as a message that becomes trustworthy when an API accepts it. For a B2B SaaS workflow that releases a generated report by email, record sender registration context, routing class, provider acceptance, delivery evidence, verification outcome, and report-release authorization as separate events.
That decision rule matters because “the API returned success” answers only one narrow question. It doesn't prove that a US or EU carrier accepted the message, that the subscriber received it, or that the person who entered a code should receive a report attachment. Carrier filtering, sender registration, shared routes, and anti-fraud controls act at different points. A single sent=True field erases those boundaries and makes a compliance review mostly guesswork.
Keep the flow plain: create a short-lived login challenge, ask a messaging adapter to send it through an approved route, consume asynchronous delivery evidence, verify one submitted code under a strict attempt budget, and only then authorize a separate email service to send the generated attachment. The attachment service should receive an authorization record, not the OTP itself. This separation is small enough for a notebook and explicit enough for production.
How should an audit record connect SMS OTP delivery to a 2FA login?
Carrier filtering is downstream of application acceptance. Sender registration is upstream configuration. Shared routes are transport choices. Anti-fraud checks may reject an attempt before any message is submitted, while 2FA login verification happens after a user supplies a code. Those aren't interchangeable failure labels — and treating them as one retryable “SMS error” can amplify traffic at exactly the wrong moment.
For US application-to-person traffic, registration requirements can depend on sender type and messaging program. Twilio's A2P 10DLC documentation is one concrete example of the registration concepts an implementation has to preserve in its configuration and change records. It is evidence for a US path, not a universal rule for every sender or country. The EU should likewise remain a region-specific policy input; the supplied sources don't establish one EU-wide sender-registration recipe, so a compliance owner needs current country and carrier requirements from the chosen providers before launch. I'm not sure a static country matrix can stay authoritative for long. Versioned policy decisions are safer than claims baked into code comments.
The useful internal status model is deliberately modest:
| Evidence state | What it establishes | What it does not establish |
|---|---|---|
policy_allowed |
The application allowed this destination and use case | Carrier acceptance or user possession |
provider_accepted |
The messaging provider accepted the submission | Handset receipt |
delivery_observed |
A delivery event was recorded for the attempt | That the intended person entered the code |
challenge_verified |
The submitted code passed the application's checks | Permission to release every report |
report_authorized |
Policy allowed this report release | Successful email delivery or attachment access |
Names beat optimism.
The evidence record should also preserve timestamps, an opaque attempt ID, destination region, sender-policy version, route class, provider message ID, callback event ID, verification result, and report authorization ID. Don't store the plaintext OTP in logs. Hashing an OTP without a server-side secret is also a poor audit shortcut because the code space is intentionally small; keep verification material in a bounded challenge store and keep audit events free of reusable secrets.
Implement the evidence contract in Python
The following example uses only the Python standard library. It doesn't call a commercial API; MessagingPort and ReportMailerPort are the boundaries where production adapters belong. The sample records acceptance separately from delivery and requires both a verified challenge and an explicit report grant before the attachment stage. Its event codes are application-defined labels, not carrier error codes.
from __future__ import annotations
from dataclasses import asdict, dataclass
from datetime import datetime, timedelta, timezone
from hashlib import sha256
import hmac
import json
import secrets
from typing import Protocol
def now_utc() -> datetime:
return datetime.now(timezone.utc)
@dataclass(frozen=True)
class AuditEvent:
attempt_id: str
event_type: str
occurred_at: str
details: dict[str, str]
class Ledger:
def __init__(self) -> None:
self.events: list[AuditEvent] = []
def append(self, attempt_id: str, event_type: str, **details: str) -> None:
self.events.append(
AuditEvent(
attempt_id=attempt_id,
event_type=event_type,
occurred_at=now_utc().isoformat(),
details=details,
)
)
class MessagingPort(Protocol):
def submit(self, *, destination: str, body: str, route_class: str) -> str: ...
class ReportMailerPort(Protocol):
def send_report(self, *, recipient: str, report_name: str, grant_id: str) -> str: ...
class DemoMessagingAdapter:
def submit(self, *, destination: str, body: str, route_class: str) -> str:
del destination, body, route_class
return f"msg_{secrets.token_hex(6)}"
class DemoReportMailer:
def send_report(self, *, recipient: str, report_name: str, grant_id: str) -> str:
del recipient, report_name, grant_id
return f"mail_{secrets.token_hex(6)}"
@dataclass
class Challenge:
attempt_id: str
digest: str
expires_at: datetime
attempts_left: int
verified: bool = False
class LoginFlow:
def __init__(
self,
messaging: MessagingPort,
mailer: ReportMailerPort,
ledger: Ledger,
secret: bytes,
) -> None:
self.messaging = messaging
self.mailer = mailer
self.ledger = ledger
self.secret = secret
self.challenges: dict[str, Challenge] = {}
def _digest(self, attempt_id: str, code: str) -> str:
payload = f"{attempt_id}:{code}".encode()
return hmac.new(self.secret, payload, sha256).hexdigest()
def start(self, destination: str, region: str, policy_version: str) -> tuple[str, str]:
attempt_id = f"otp_{secrets.token_hex(8)}"
code = f"{secrets.randbelow(1_000_000):06d}"
route_class = "registered_transactional"
self.ledger.append(
attempt_id,
"policy_allowed",
region=region,
policy_version=policy_version,
route_class=route_class,
)
message_id = self.messaging.submit(
destination=destination,
body=f"Your login code is {code}",
route_class=route_class,
)
self.challenges[attempt_id] = Challenge(
attempt_id=attempt_id,
digest=self._digest(attempt_id, code),
expires_at=now_utc() + timedelta(minutes=5),
attempts_left=5,
)
self.ledger.append(
attempt_id,
"provider_accepted",
provider_message_id=message_id,
)
return attempt_id, code
def observe_delivery(self, attempt_id: str, callback_event_id: str) -> None:
self.ledger.append(
attempt_id,
"delivery_observed",
callback_event_id=callback_event_id,
)
def verify(self, attempt_id: str, submitted_code: str) -> bool:
challenge = self.challenges[attempt_id]
if now_utc() >= challenge.expires_at or challenge.attempts_left == 0:
self.ledger.append(attempt_id, "verification_denied", reason="OTP_EXPIRED_OR_LOCKED")
return False
challenge.attempts_left -= 1
challenge.verified = hmac.compare_digest(
challenge.digest,
self._digest(attempt_id, submitted_code),
)
result = "challenge_verified" if challenge.verified else "verification_denied"
self.ledger.append(attempt_id, result, attempts_left=str(challenge.attempts_left))
return challenge.verified
def release_report(self, attempt_id: str, recipient: str, report_name: str) -> str:
challenge = self.challenges[attempt_id]
if not challenge.verified:
raise PermissionError("REPORT_AUTH_REQUIRED")
grant_id = f"grant_{secrets.token_hex(8)}"
self.ledger.append(attempt_id, "report_authorized", grant_id=grant_id)
mail_id = self.mailer.send_report(
recipient=recipient,
report_name=report_name,
grant_id=grant_id,
)
self.ledger.append(attempt_id, "report_email_accepted", mail_id=mail_id)
return mail_id
ledger = Ledger()
flow = LoginFlow(
messaging=DemoMessagingAdapter(),
mailer=DemoReportMailer(),
ledger=ledger,
secret=secrets.token_bytes(32),
)
attempt_id, demo_code = flow.start(
destination="+12025550123",
region="US",
policy_version="sms-policy-2026-08",
)
flow.observe_delivery(attempt_id, callback_event_id="evt_demo_001")
assert flow.verify(attempt_id, demo_code)
flow.release_report(attempt_id, "analyst@example.com", "quarterly-risk.pdf")
print(json.dumps([asdict(event) for event in ledger.events], indent=2))
The demo returns the code only so the file can run end to end. A deployed adapter sends the code and never returns it to a browser, report worker, analytics pipeline, or log sink. The same rule applies to notebook experiments: use a fake destination and a deterministic adapter in tests, then exercise the real callback parser against signed fixtures. A notebook that can only prove the happy path isn't a deployment test.
Evaluate carrier and anti-fraud outcomes as separate cohorts
Classify failure by the component that can act on it. A policy denial belongs before submission and should not trigger a provider retry. Provider rejection belongs to sender configuration or request validation. A later delivery event belongs to the asynchronous message timeline. An incorrect or expired code belongs to authentication. A report authorization denial belongs to the SaaS permission model. This taxonomy gives an eval harness stable labels even if an underlying messaging supplier changes.
Retries need the same care. A timeout in an application process doesn't prove that a submission failed, so every send attempt needs an idempotency strategy at the adapter boundary and a stable internal attempt ID. Delivery callbacks need deduplication by callback event ID. Resends should create a new message record linked to the same login session, invalidate superseded codes according to policy, and consume a fraud budget. Otherwise, a user can receive several valid codes in a confusing order while the audit log reports one vague attempt.
For observability, measure the funnel rather than a single delivery percentage: policy-allowed to provider-accepted, accepted to delivery-observed, delivery-observed to challenge-verified, and verified to report-authorized. Slice by country, sender-policy version, route class, and provider without putting phone numbers in metric labels. Watch latency distributions as well as counts. A fast acceptance event followed by no delivery evidence points to a different investigation than a verified login followed by no report grant.
Anti-fraud is the uncomfortable trade-off. Tight destination, velocity, and attempt limits reduce abuse but can block legitimate users who travel, share corporate egress, or request a second code after a delay. Loose controls improve apparent completion while increasing pumping and account-takeover exposure. There is no honest universal threshold. Tune limits with labeled outcomes, review false denials, and keep the model or ruleset version in the decision event so an evaluation can reproduce why attempt otp_… was allowed.
This architecture is not suitable when SMS possession is too weak for the account's risk level or when users cannot reliably receive mobile messages. Those cases require a separately approved authentication or recovery design instead of aggressive resends. SMS can remain one channel in a layered login design, but a report containing regulated or highly sensitive data may also need document-level access controls instead of a conventional email attachment. The catch is operational weight: alternate authenticators and protected download links require enrollment, recovery, and support workflows that a plain attachment does not.
Set the report release policy and retention boundary
Start eval-driven. Unit tests should cover code expiry, constant-time comparison, attempt exhaustion, superseded challenges, authorization denial, and callback deduplication. Contract tests should feed recorded, redacted callback shapes into each adapter. Staging tests should verify that region and sender-policy selection are configuration decisions, not guesses derived from a UI locale. None of those tests should send a real report to a real customer.
Then run controlled probes through approved test destinations and compare events across the funnel. The evidence question for each probe is specific: did policy allow it, did the provider accept it, was delivery observed, was the expected challenge verified, and did the exact report grant reach the mail boundary? An alert such as OTP_FUNNEL_DROP_US_POLICY_17 is far more useful than “SMS failed.” It points to the cohort and policy version without pretending the alert already knows the carrier's reason.
Keep generated report content out of the authentication trace. Store a digest or immutable object reference for the report artifact, bind it to the grant, and let the email adapter record its own acceptance identifier. Amazon SES documents email sending through its API and SMTP interface; whichever mail transport is selected, its acceptance evidence remains separate from OTP delivery evidence. This clean boundary also limits prompt and token-cost spillover: the report-generation job can be evaluated for factual quality and cost without coupling those measurements to carrier retries.
Before production, read the ledger as if the support engineer has no dashboard and the auditor has no source code. They should be able to reconstruct one report release from opaque identifiers, timestamps, policy versions, and causal links, while learning neither the OTP nor unnecessary report content. Confirm that callback authentication and deduplication are active, retention matches policy, sensitive fields are redacted, region rules have named owners, resends consume explicit budgets, and a verified challenge cannot authorize a different tenant's report. Finally, rehearse a provider change with the same eval corpus. If the adapter swap changes the meaning of provider_accepted, the abstraction is leaking.
Ship only when those statements are testable.
Top comments (0)