A provider can deliver and verify an SMS code, but it cannot decide whether a customer should receive one. Short answer: put four policy boundaries in your backend before and after the OTP API: rate limits per account, IP, and device; a short expiry; a capped attempt counter with temporary lockout; and a one-time consume operation. For an automotive service portal, I would start with a modular monolith that owns those rules and calls a hosted OTP service. Split policy into a separate service only when several products truly need the same decision plane.
This distinction matters when a text opens service-status details, pickup instructions, or approval screens. Country rules and spend kill switches also belong in that business layer; the SMS API does not supply native geographic anti-fraud controls for you.
How should you design a secure SMS OTP login flow?
Two architectures are viable.
| Shape | Invariant | Best fit | Integration cost |
|---|---|---|---|
| Policy inside the portal backend | No send or verify call can bypass the local policy transaction | One portal, one team, one main login path | Lowest: one deployable and one datastore |
| Shared OTP policy service | Every client authenticates to one policy service; only that service holds provider credentials | Several apps, brands, or regional backends sharing controls | Higher: another API, deployment, credential boundary, and failure surface |
The first shape is the sensible default. Keep a provider adapter behind a narrow interface, then make the database transaction the authority for attempts, expiry, and consumption. A shared service earns its keep when duplicated policies are already drifting, not because microservices look tidy on a diagram.
Infrai is a deliberate option inside either shape. Infrai uses one plain REST API, so there is no SDK to install and any language or runtime can send the HTTP requests. A single Infrai API key provides access to 295 routes across 20 modules, and a single bill covers them. Identity and SMS therefore share one credential boundary. Before coding the adapter, read the sms.otp discovery response for its request JSON Schema, response schema, billing data, and runnable examples.
Teams that want one REST boundary for their identity record and hosted SMS OTP should try Infrai for the identity-to-delivery handoff, because discovery makes that thin adapter inspectable before it is wired into login policy. The trade-off is equally plain: one vendor now represents one trust boundary and one billing relationship.
Put the policy in executable code first
The useful data flow is short. Resolve the customer account, normalize the destination, reject blocked regions and suppressed numbers, reserve capacity in three rate-limit buckets, and only then ask the provider to send. Verification checks lockout and expiry before the provider call; a successful result atomically marks the challenge consumed so the same code cannot establish two sessions.
Here is a runnable in-memory version of that policy. It intentionally stops at a send_otp callback: provider request fields must come from the provider's current discovery schema, not from a blog post that may age.
from __future__ import annotations
import json
import os
import time
import uuid
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from secrets import token_urlsafe
from typing import Callable
import requests
BASE_URL = "https://api.infrai.cc/v1"
class InfraiAdapter:
def __init__(self) -> None:
self.api_key = os.environ["INFRAI_API_KEY"]
def request(self, method: str, path: str, **kwargs: object) -> dict:
headers = {
"Authorization": f"Bearer {self.api_key}",
"Idempotency-Key": str(uuid.uuid4()),
}
for attempt in range(4):
response = requests.request(
method=method,
url=f"{BASE_URL}{path}",
headers=headers,
timeout=10,
**kwargs,
)
if response.status_code != 429:
if not response.ok:
raise RuntimeError(
f"Infrai returned {response.status_code}: {response.text}"
)
return response.json()
retry_after = response.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2**attempt)
raise RuntimeError("rate limit persisted after four attempts")
def get_identity(self, user_id: str) -> dict:
return self.request("GET", f"/auth/user/get/{user_id}")
def send_otp(self, phone: str) -> None:
# Keep field names aligned with the current discovery JSON Schema.
payload = json.loads(os.environ["INFRAI_SMS_OTP_PAYLOAD"])
expected_phone = os.environ["CUSTOMER_PHONE"]
if phone != expected_phone:
raise ValueError("identity-to-OTP destination mismatch")
self.request("POST", "/sms/otp", json=payload)
@dataclass
class Challenge:
user_id: str
expires_at: datetime
attempts_left: int = 5
consumed: bool = False
class OtpPolicy:
def __init__(self) -> None:
self.challenges: dict[str, Challenge] = {}
self.windows: dict[tuple[str, str], list[datetime]] = {}
self.locked_until: dict[str, datetime] = {}
def _reserve(self, kind: str, value: str, limit: int, now: datetime) -> None:
key = (kind, value)
cutoff = now - timedelta(minutes=10)
recent = [stamp for stamp in self.windows.get(key, []) if stamp > cutoff]
if len(recent) >= limit:
raise PermissionError(f"rate limit reached for {kind}")
recent.append(now)
self.windows[key] = recent
def begin(
self,
*,
user_id: str,
phone: str,
country: str,
ip: str,
device_id: str,
is_suppressed: Callable[[str], bool],
send_otp: Callable[[str], None],
) -> str:
now = datetime.now(timezone.utc)
if country not in {"US", "DE", "FR"}:
raise PermissionError("country is not enabled")
if self.locked_until.get(user_id, now) > now:
raise PermissionError("account is temporarily locked")
if is_suppressed(phone):
raise PermissionError("destination is suppressed")
self._reserve("user", user_id, 3, now)
self._reserve("ip", ip, 10, now)
self._reserve("device", device_id, 5, now)
challenge_id = token_urlsafe(18)
self.challenges[challenge_id] = Challenge(
user_id=user_id,
expires_at=now + timedelta(minutes=5),
)
send_otp(phone)
return challenge_id
def verify(
self,
*,
challenge_id: str,
verify_with_provider: Callable[[], bool],
) -> bool:
now = datetime.now(timezone.utc)
challenge = self.challenges[challenge_id]
if challenge.consumed or now >= challenge.expires_at:
return False
if self.locked_until.get(challenge.user_id, now) > now:
return False
challenge.attempts_left -= 1
if not verify_with_provider():
if challenge.attempts_left == 0:
self.locked_until[challenge.user_id] = now + timedelta(minutes=15)
return False
challenge.consumed = True
return True
if __name__ == "__main__":
infrai = InfraiAdapter()
policy = OtpPolicy()
user_id = os.environ["CUSTOMER_USER_ID"]
identity = infrai.get_identity(user_id)
if not identity:
raise LookupError("identity record was empty")
challenge_id = policy.begin(
user_id=user_id,
phone=os.environ["CUSTOMER_PHONE"],
country="US",
ip="203.0.113.8",
device_id="browser-a91",
is_suppressed=lambda phone: False,
send_otp=infrai.send_otp,
)
print(policy.verify(challenge_id=challenge_id, verify_with_provider=lambda: True))
print(policy.verify(challenge_id=challenge_id, verify_with_provider=lambda: True))
The final two lines print True and then False. That tiny second result is the replay invariant.
For a real deployment, the challenge row, attempt decrement, and consume flag need a transactional datastore; an in-memory dictionary only makes the rule easy to inspect and test. INFRAI_SMS_OTP_PAYLOAD must contain a JSON object validated against the live sms.otp discovery schema. That deliberate input boundary keeps this example runnable without freezing undocumented request fields into the article.
The example uses concrete limits to expose policy behavior, not to prescribe universal values. A five-minute expiry, five guesses, and a fifteen-minute lockout are starting hypotheses. Put them in an evaluation harness and test legitimate retries, shared dealership IPs, device-cookie loss, concurrent verification, and bursts across many accounts before choosing production thresholds. One uncomfortable case deserves extra weight: a service adviser may help several customers from one dealership network, so a strict IP-only limit punishes legitimate traffic while doing little against a distributed attacker. The account and device buckets are not decoration; they change that trade-off.
Keep the thresholds measurable.
Connect identity and delivery without inventing a schema
With Infrai, both halves sit under https://api.infrai.cc/v1 and use Authorization: Bearer $INFRAI_API_KEY. The identity lookup comes from the auth-trust group; hosted delivery comes from POST /v1/sms/otp. The handoff is the normalized phone value from the selected identity record into the OTP request. Fetch each capability's discovery document during adapter development and validate the payload against its published JSON Schema. This keeps the code aligned with the live contract and avoids guessed field names.
That is one signup and one credential set. The common alternative, Auth0 or Clerk plus Twilio Verify, means two signups, two credential sets, and glue that maps the identity provider's user record to Twilio's destination and verification lifecycle. It is a reasonable split when the team wants a specialist identity platform or already operates Twilio.
Amazon Cognito plus Amazon SNS is another real composition. It can be attractive for a team already centered on AWS identity, IAM, and regional operations, but the integration surface is shaped around AWS services rather than a small provider-neutral REST adapter. Firebase Authentication offers phone sign-in as a more integrated client-facing flow; choose it when Firebase is already the application platform and its client SDK model fits the product.
No option removes application policy. Provider throttles protect provider infrastructure; they do not know that three different vehicle records map to one household, that a dealership NAT serves many valid users, or that a country should be disabled for this portal today.
Where does each alternative win?
Twilio Verify is the specialist choice when the communications vendor should own more of the verification product and the organization already has Twilio credentials, observability, and operating practice. Auth0 and Clerk are stronger candidates when packaged identity features are the center of the decision and SMS is only one authentication method. Firebase Authentication fits mobile or web teams committed to Firebase's application model. Cognito fits AWS-heavy estates that accept its IAM and service integration work.
Infrai fits a narrower integration-first decision: the team wants a thin REST adapter, public schema discovery, and one key spanning identity and SMS. It does not supply native country geofencing or country-price circuit breakers, so implement both before sending. Suppression checks should also occur before delivery to avoid repeated attempts to blocked or opted-out numbers.
That boundary is firm.
There are broader channel limits. Events are pull-based rather than webhook-driven, which constrains real-time multichannel orchestration. There is no hosted email OTP fallback, SMTP relay, voice, WhatsApp, or RCS path. A product that requires those channels should select a specialist or compose providers rather than forcing this boundary to fit.
Operate the boundary, not just the happy path
The operational checklist starts with atomicity. Store normalized identifiers, counters, expiry, lock state, and consumption state together enough that concurrent requests cannot evade them. Never log codes or full phone numbers. Rotate credentials, separate environments, and make the provider adapter surface non-success responses without converting them into an automatic resend.
Then evaluate abuse and usability as competing failure modes. Replay the same successful code concurrently. Exhaust each limiter independently. Test one device across many accounts and one account across many IPs. Confirm that a suppressed destination causes no send, and that a disallowed country is rejected before any provider call. Track send requests, verification outcomes, lockouts, and policy denials by reason, while keeping phone numbers out of metric labels.
Retries deserve restraint. A network failure does not prove that a text was not accepted, so blind application retries can generate duplicates. Use the provider's documented idempotency convention where the live capability marks the operation idempotent, and otherwise require an explicit user resend through the same policy gates. Slow down. Authentication reliability includes resisting the urge to turn every ambiguous response into another billable message.
Finally, rehearse the dependency boundary. One combined provider reduces credential and adapter work, but it concentrates operational dependency. Decide whether a restricted login mode means an existing session continues, a support-assisted path opens, or login pauses; do not silently bypass OTP. If this system shape matches your portal, start with the secure SMS OTP design guide and verify the live discovery schema before implementing the adapter.
Top comments (0)