DEV Community

LukasSchmidt295
LukasSchmidt295

Posted on

SMS OTP API Ownership: 6 Rate Limit Rules for SaaS Login Recovery

If you are evaluating an SMS OTP API for SaaS login recovery, give the application ownership of the password-reset template, code or link policy, rate limit, and message intent; let delivery adapters own only transport details. This is the simplest design that keeps a short-lived B2B SaaS reset flow consistent across email and SMS without welding security behavior to a provider.

Short answer: generate a single-use, random reset secret on the server, store only a protected representation with an explicit expiry, and render channel-specific copy from one versioned template contract. A short expiry belongs in application policy. Retries belong in a bounded delivery worker. Verification must be an atomic consume operation, never a comparison performed by a message provider.

This split matters when a user is locked out five minutes before an event-ticket check-in window. The message may arrive late, twice, or through the fallback channel. None of those outcomes should extend the credential, create a second valid reset path, or change what “expired” means.

What should a SaaS login SMS OTP API own?

The application should own six things: the reset intent, secret generation, expiry, one-time consumption, template variables, and the user-visible meaning of each message. The transport layer should own addressing syntax, provider authentication, delivery submission, normalized status events, and provider-specific throttling.

Keep that boundary boring.

That boundary is deliberate. If an email template says “expires in 10 minutes” while an SMS template says “expires soon,” but the database accepts the secret for 30 minutes, the system has three policies. A dashboard edit can silently create that mismatch when copy and authentication rules live in different control planes.

Keep templates in the same review path as the code that defines their variables. A template change can then be tested against the reset state machine before deployment. Operations staff may still edit non-security copy through a controlled workflow, but expiry language, links, and required variables need code review and versioning.

The plain-language data flow is small: an API accepts a reset request, creates an opaque secret and a database record, then places a message intent on a queue. A worker renders the requested template and calls a generic email or SMS adapter. The reset endpoint hashes the presented secret, atomically marks the matching unexpired record as used, and only then permits a password change. This distinction also resolves a common naming problem in API selection: a password-reset link is not an OTP merely because it arrives by SMS. If the product requirement is account recovery, model recovery explicitly instead of squeezing it into a login-code abstraction whose verification and expiry rules may be owned elsewhere.

A runnable boundary before provider details

The following example uses the Python standard library and SQLite so the ownership boundary is visible. It does not send a real message. Instead, the outbox record is the stable handoff a delivery worker would claim after the transaction commits.

import hashlib
import json
import secrets
import sqlite3
import time
from dataclasses import dataclass


RESET_TTL_SECONDS = 600


@dataclass(frozen=True)
class ResetMessage:
    user_id: str
    channel: str
    destination: str
    reset_url: str
    expires_in_minutes: int


def digest(secret: str) -> str:
    return hashlib.sha256(secret.encode("utf-8")).hexdigest()


def request_reset(
    db: sqlite3.Connection,
    *,
    user_id: str,
    channel: str,
    destination: str,
    now: int | None = None,
) -> str:
    issued_at = int(time.time()) if now is None else now
    secret = secrets.token_urlsafe(32)
    expires_at = issued_at + RESET_TTL_SECONDS
    message = ResetMessage(
        user_id=user_id,
        channel=channel,
        destination=destination,
        reset_url=f"https://accounts.example/reset?token={secret}",
        expires_in_minutes=RESET_TTL_SECONDS // 60,
    )

    with db:
        db.execute(
            """INSERT INTO password_resets
               (token_digest, user_id, expires_at, used_at)
               VALUES (?, ?, ?, NULL)""",
            (digest(secret), user_id, expires_at),
        )
        db.execute(
            """INSERT INTO message_outbox
               (kind, template_version, payload, attempts, available_at)
               VALUES (?, ?, ?, 0, ?)""",
            ("password_reset", 3, json.dumps(message.__dict__), issued_at),
        )
    return secret


def consume_reset(
    db: sqlite3.Connection, secret: str, *, now: int | None = None
) -> str | None:
    checked_at = int(time.time()) if now is None else now
    with db:
        row = db.execute(
            """UPDATE password_resets
               SET used_at = ?
               WHERE token_digest = ?
                 AND used_at IS NULL
                 AND expires_at >= ?
               RETURNING user_id""",
            (checked_at, digest(secret), checked_at),
        ).fetchone()
    return None if row is None else str(row[0])
Enter fullscreen mode Exit fullscreen mode

The UPDATE ... RETURNING statement makes competing submissions converge on one winner. Production code also needs a supported database transaction level, retention cleanup, key-management decisions, and a response that does not reveal whether an account exists. Those details belong beside the authentication model, not inside a transport callback.

Notice what is absent: a provider template identifier, a provider-generated code, and delivery status as a condition for validity. Delivery is evidence about transport. It is not authentication state.

Why not let each channel define the reset?

Provider-hosted templates can be useful when non-engineers need localized copy or a regulated approval flow. They also create a second deployment surface. The cost is not merely migration work; it is the chance that a live template expects reset_link while the application emits reset_url, or that an emergency edit promises an expiry the verifier does not enforce.

Application-rendered templates have the inverse trade-off. They make review, fixtures, and local tests straightforward, but the team must build localization, escaping, preview, and approval controls. A hybrid can work: keep a versioned semantic contract in the application, allow reviewed prose variants elsewhere, and reject publication when required variables or protected sentences change.

Here is a compact decision table:

Ownership model Strong fit Main risk Required control
Application-rendered Security copy changes with code Engineers become a content bottleneck Snapshot and localization tests
Provider-rendered Operations owns frequent copy changes Policy drifts outside code review Version pinning and contract validation
Hybrid contract Several channels and locales Two systems can disagree Publish-time schema checks

For a password reset with a short expiry, I would choose application ownership unless an approval requirement rules it out. The reason is narrow: the sentence describing expiry is part of the security contract. Brand punctuation is not.

How should retries and rate limits behave?

Treat “create reset intent” and “attempt delivery” as separate operations. The request endpoint should return the same generic response for registered and unregistered identifiers, then enforce limits around the account, destination, client, and recent reset activity. A single IP limit is weak protection for a multi-tenant SaaS product, while destination-only limits can let an attacker lock a victim out of receiving useful messages.

Do not create a fresh secret on every transport retry. Retry the outbox item with the same template version and reset record, add jitter to capped backoff, and stop when the remaining validity window is too short to be useful. Exact retry counts and delays are operational choices that should come from measured delivery latency and provider responses, not copied folklore.

Fast failure matters here. A worker should classify failures into retryable transport conditions, permanent addressing failures, and ambiguous outcomes. If submission times out after the remote system accepted it, an idempotency key or an internal deduplication key limits duplicate sends. A duplicate message is annoying; two independently valid secrets are a larger state-space problem.

NIST SP 800-63B treats use of the public switched telephone network for out-of-band authentication as restricted and asks verifiers to consider risks such as SIM changes and number porting. That makes SMS a recovery delivery option with explicit risk analysis, not proof that a phone number is a durable identity. Email has a different threat model, and DMARC addresses domain-level message authentication policy and reporting rather than the validity of a reset secret.

Can the template be tested like an AI feature?

Yes, and the eval should be small enough to run on every change. I use the same notebook-to-production instinct here that serves prompt work: begin with a table of adversarial cases, convert it into executable fixtures, and gate the deployment on stable assertions. No model is needed for the core checks.

Test template version 3 with a destination in each supported locale, an expiry at the boundary, escaped user-controlled display text, a missing required variable, and the longest permitted organization name. Assert that the rendered message contains one reset URL, never includes the stored digest, names the correct duration, and produces an SMS segment count within the team’s chosen limit. Then test the state machine separately: before expiry succeeds once, the same secret fails twice, and the expiry boundary follows one documented clock rule.

For AI-assisted copy changes, pin the security sentences and evaluate only the prose that is allowed to vary. Record prompt and model costs in the experiment, but do not put a generative call on the reset delivery path. A password reset is latency-sensitive, security-sensitive, and repetitive; runtime generation adds variability without solving the template-ownership problem.

One trap deserves its own line.

Do not log the reset URL.

Log an internal intent ID, template version, channel, normalized outcome, attempt number, timestamps, and a trace correlation value. Keep destinations redacted according to the organization’s privacy policy. Useful service-level views include request-to-queue time, queue age, submission latency, delivery outcome where available, expiry-before-delivery counts, and successful consumption time. These measurements separate authentication defects from transport delays.

Operating the flow without policy drift

Before release, walk one template version from request through render, submission, expiry, and consumption in a staging environment whose messages cannot reach real users. Confirm that an unknown account receives the same API response shape and timing treatment expected by the threat model. Review link origins, redirects, cookie behavior, and the password-change session; the message is only the entry point.

During deployment, publish code and templates as a compatible unit, keep the previous version renderable while queued messages drain, and watch queue age rather than send volume alone. Rollback must preserve the ability to consume secrets already issued. When a template contract changes, migrate by version instead of reinterpreting old outbox payloads.

Afterward, sample redacted render artifacts, reconcile submission and status events, expire old reset records, and alert on sudden changes in requests, permanent failures, or consumption ratios. Tune limits using those signals. The final decision rule stays simple: security semantics live with the application; delivery mechanics live behind adapters; editable prose may move only behind a validated contract.

Further reading

Top comments (0)