TL;DR: For a logistics password-reset flow, keep the message template, expiry policy, and abuse controls in your application; let the SMS API transport the final text. Compare providers only after the same Python test harness can exercise each adapter. The lowest quoted message rate is not the cheapest system when retries, blocked destinations, support work, and account-takeover exposure enter the equation.
My evaluation constraint is strict: a provider swap must not change the words a dispatcher receives, extend the reset window, or bypass a per-account and per-destination limit. A simple first design hands a template identifier and variables to the transport provider. That looks tidy in a notebook. It also makes copy review, localization, expiry changes, and failover depend on provider-specific template state.
The chosen boundary is less magical and more useful. Python renders one approved message, records a template revision, and passes opaque text to a narrow transport adapter. This is the part I would settle before comparing Plivo, Telnyx, Vonage, Twilio, or another API. Those names identify candidates, not a ranking; commercial terms and destination support must be checked for the actual US and EU traffic profile.
Who should own the reset template?
The application team should own it when one security policy must survive routing changes. In this logistics scenario, a warehouse supervisor may request a reset while a driver is on the road. The text needs a recognizable service name, one action, and a short expiry. It should not contain a password, a reusable secret, or operational shipment data.
Template ownership is also change ownership. Keep the source beside the code that creates the reset challenge, review it with the same pull request, and give every approved version an immutable identifier. The transport layer then has a deliberately boring contract: destination, rendered body, idempotency key, and metadata needed for observation. Provider delivery identifiers belong in adapter output, not in business logic.
There is a trade-off. Provider-managed templates can centralize features offered by that provider, while application-managed rendering creates work for localization and compliance review. For a reset path that may fail over, I accept that work because identical wording and expiry semantics are part of the security boundary. Marketing messages can make a different choice.
A focused Python boundary
The example below does not send a real reset or prescribe an HTTP route. It shows the seam I put under tests. The reset token is generated and stored elsewhere; the sender receives only the already-created one-time code and its expiry.
from dataclasses import dataclass
from typing import Protocol
@dataclass(frozen=True)
class ResetMessage:
destination: str
body: str
template_revision: str
idempotency_key: str
@dataclass(frozen=True)
class DeliveryReceipt:
transport_id: str
accepted: bool
class SmsTransport(Protocol):
def send(self, message: ResetMessage) -> DeliveryReceipt: ...
def build_reset_message(
destination: str,
code: str,
expires_in_minutes: int,
challenge_id: str,
) -> ResetMessage:
if not 1 <= expires_in_minutes <= 15:
raise ValueError("expiry must be between 1 and 15 minutes")
body = (
f"Fleet Portal reset code: {code}. "
f"Expires in {expires_in_minutes} minutes. "
"If you did not request this, ignore this message."
)
return ResetMessage(
destination=destination,
body=body,
template_revision="password-reset-v3",
idempotency_key=challenge_id,
)
The 15 is an example policy ceiling, not a universal standard. Put the actual value in configuration, but test the permitted range and never let an adapter silently replace it. I would snapshot the exact body for every locale and fail CI when punctuation, service naming, or expiry language changes without an intentional fixture update.
This keeps prompt cost at zero for the critical path. A language model has no job generating security copy at request time; nondeterministic wording makes evaluation harder and adds latency. If AI helps draft translations, the output still goes through human review and lands as a versioned static template before production.
How should SMS API alternatives handle OTP abuse?
Rate limiting cannot be delegated entirely to an SMS account. A provider sees sends; the application sees failed logins, reset challenges, accounts, sessions, and risk signals. Enforce limits before creating another challenge, and combine several keys: normalized destination, account, IP or device signal, and a broader tenant or region budget. One global counter is too blunt. One phone-number counter is too easy to distribute around. Return the same public response for an existing and nonexistent account so the endpoint does not become an account-discovery tool. Store only the minimum operational data, apply retention rules, and keep the code or token out of logs. The OWASP Forgot Password guidance recommends consistent responses, random single-use tokens, secure storage, expiry, and rate limiting against excessive automated submissions. Retries need their own rule because a network timeout does not prove that no message was accepted. Retrying with a fresh reset challenge can produce two valid codes and confuse the user, so reuse the challenge and idempotency key for transport retries; create a new challenge only through the reset policy. Short-circuit blocked or malformed destinations before consuming a send budget.
Retries are ambiguous.
Fast feedback matters. Yet an acceptance response from any transport is not evidence that a handset received the text. Model accepted, delivered, failed, and unknown as four separate states, then let signed delivery events update the receipt asynchronously. Verify event authenticity according to the selected provider's documented mechanism, reject stale replays, and make the handler idempotent.
Unknown is not delivered.
Compare the candidates with one experiment
Do not begin with a price table. Begin with a replayable data set that reflects the route: US and EU destinations, the supported character sets, realistic message lengths, invalid numbers, opt-out cases, and simulated timeouts. Use synthetic or explicitly consented test destinations. A logistics company should also split results by country and carrier class; an aggregate pass rate can hide a weak route.
I would run every adapter through the same checks:
| Dimension | Evidence to collect | Failure that matters |
|---|---|---|
| Template control | Rendered body and revision | Copy changes between routes |
| Abuse boundary | Counter decisions before send
|
A retry consumes a new challenge |
| Delivery state | Accepted-to-final state transitions | Unknown is counted as delivered |
| Failover | Same idempotency key and body | Two active reset codes reach one user |
| Regional fit | Results split by country | US success masks an EU gap |
| Operations | Redacted logs and event verification | Codes or full numbers enter logs |
The commercial comparison comes after those tests. Normalize quotes against the same destination mix and include required registration, number type, inbound handling, support tier, and retry behavior where applicable. Do not publish “cheapest” from a single list price: the answer changes with country, carrier, traffic pattern, taxes, and contract terms. Record the date and assumptions for every quote.
Plivo, Telnyx, Vonage, and Twilio should therefore face the identical adapter contract and evidence sheet. Their current documentation and account-specific terms are the sources for supported regions, sender requirements, authentication, event signatures, and commercial conditions. A decision is defensible when the evidence is reproducible, not when one dashboard has the friendliest first impression.
What should you measure before copying this design?
Measure completion, not sends. The useful numerator is successfully completed resets; supporting signals include challenge creation, limit denials, transport acceptance, final delivery state, time to delivery, duplicate delivery, expiry before use, and successful reset. Segment by template revision, route, country, and adapter without placing the token or full destination in telemetry.
Also measure suppression. A rising number of attempts stopped per account, destination, or network range may be an attack, a broken client retry loop, or a user who cannot receive the first message. Those cases require different responses, so preserve reason codes internally while keeping the public endpoint response uniform.
Start with a shadow adapter test and synthetic destinations. Then use a small, explicitly controlled slice of eligible traffic, with a rollback rule based on reset completion and duplicate-code rate. Keep one evaluation fixture for rendering and another for transport state transitions. Notebook exploration is useful for inspecting the distributions; promotion to production should require those fixtures to pass in CI.
The final choice is about ownership: application-owned policy and templates, provider-owned transport, and measured evidence at the boundary. That structure keeps a logistics reset consistent across US and EU routes and makes future comparisons ordinary engineering work rather than a rewrite.
Further reading
- OWASP Forgot Password Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Forgot_Password_Cheat_Sheet.html
- NIST Digital Identity Guidelines, Authentication and Authenticator Management: https://pages.nist.gov/800-63-4/sp800-63b.html
- NIST Secure Software Development Framework: https://csrc.nist.gov/pubs/sp/800/218/final
- Resend documentation: https://resend.com/docs/introduction
- Yahoo Sender Best Practices: https://senders.yahooinc.com/best-practices/
Top comments (0)