DEV Community

ColeMitchell4991
ColeMitchell4991

Posted on

Password Reset Email API 429 Rate Limits in 4 Delivery States

TL;DR: Put signup verification or password reset delivery behind one small interface and model each attempt as one of four states: ready, sending, cooling_down, or accepted. When a transactional email API returns 429, honor a valid Retry-After, persist the next eligible send time, and return a neutral response to the browser. Don't sleep inside the request handler, and don't create a new secret for every click. This keeps a marketplace signup flow straightforward to integrate without turning retries into duplicate-message storms.

The least complex useful design is a database row, a background worker, and a replaceable delivery adapter. The browser asks for a link; the application creates or reuses an unexpired challenge; the worker claims the row and calls the adapter. An accepted request ends the delivery loop. A throttled request moves it into a durable cooldown. The signup endpoint remains quick even when delivery is not.

How should a password reset email API handle 429 rate limits?

The button that requests another message is a user-interface control, while 429 Too Many Requests is feedback about traffic admitted by another system. Treating them as the same mechanism creates a nasty loop: a person sees no message, clicks again, and each click adds another immediate API call. More traffic arrives precisely when the downstream service has asked for less. The same failure occurs with a password reset link, though the concrete example here is a marketplace account verification link.

There are also two clocks. The product clock answers, "When may this account request another message?" The transport clock answers, "When may this delivery operation be attempted again?" Keep both. A 60-second product cooldown can reduce accidental clicks, but it cannot substitute for a server-provided Retry-After; conversely, a transport retry time should not reveal whether an address belongs to an account.

For a marketplace signup, the public response should remain neutral and stable. The service can say that a verification message will be sent if the request is eligible, then expose only a generic cooldown to the page. Internally, it can record exact states and reasons. That split limits account-discovery clues and keeps transport details out of the front end.

Short handlers win. Every time.

Build the four-state loop

The example below uses only the Python standard library. The storage methods are deliberately an interface: in production they need an atomic claim, durable timestamps, and a uniqueness rule for the active signup challenge. The delivery adapter is similarly narrow, so changing an HTTP client or email service does not leak into signup logic.

from __future__ import annotations

from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from email.utils import parsedate_to_datetime
from enum import Enum
from typing import Mapping, Protocol


class DeliveryState(str, Enum):
    READY = "ready"
    SENDING = "sending"
    COOLING_DOWN = "cooling_down"
    ACCEPTED = "accepted"


@dataclass(frozen=True)
class DeliveryResult:
    status_code: int
    headers: Mapping[str, str]
    receipt_id: str | None = None


@dataclass
class VerificationJob:
    job_id: str
    recipient: str
    verification_url: str
    state: DeliveryState
    attempt: int
    next_attempt_at: datetime


class MailAdapter(Protocol):
    def send_signup_verification(
        self, recipient: str, verification_url: str, idempotency_key: str
    ) -> DeliveryResult: ...


class JobStore(Protocol):
    def claim_ready(self, now: datetime) -> VerificationJob | None: ...
    def mark_accepted(self, job_id: str, receipt_id: str | None) -> None: ...
    def schedule(self, job_id: str, next_attempt_at: datetime) -> None: ...
    def release(self, job_id: str, next_attempt_at: datetime) -> None: ...


def retry_after_at(value: str | None, now: datetime) -> datetime | None:
    if value is None:
        return None

    value = value.strip()
    if value.isdigit():
        return now + timedelta(seconds=int(value))

    try:
        parsed = parsedate_to_datetime(value)
    except (TypeError, ValueError, OverflowError):
        return None

    if parsed.tzinfo is None:
        parsed = parsed.replace(tzinfo=timezone.utc)
    return max(now, parsed.astimezone(timezone.utc))


def capped_backoff(attempt: int) -> timedelta:
    # Local policy for missing or invalid Retry-After values.
    seconds = min(15 * (2 ** min(attempt, 6)), 15 * 60)
    return timedelta(seconds=seconds)


def deliver_one(store: JobStore, mail: MailAdapter, now: datetime) -> bool:
    job = store.claim_ready(now)
    if job is None:
        return False

    try:
        result = mail.send_signup_verification(
            recipient=job.recipient,
            verification_url=job.verification_url,
            idempotency_key=job.job_id,
        )
    except TimeoutError:
        store.release(job.job_id, now + capped_backoff(job.attempt))
        return True

    if 200 <= result.status_code < 300:
        store.mark_accepted(job.job_id, result.receipt_id)
    elif result.status_code == 429:
        retry_at = retry_after_at(result.headers.get("Retry-After"), now)
        store.schedule(job.job_id, retry_at or now + capped_backoff(job.attempt))
    else:
        store.release(job.job_id, now + capped_backoff(job.attempt))

    return True
Enter fullscreen mode Exit fullscreen mode

claim_ready is the critical boundary. It should change ready or an expired cooling_down row to sending in the same transaction that selects it. Without that claim, two workers can both observe an eligible row and send the same link. The state names are intentionally about our operation, not about delivery to an inbox: a 2xx API response means the request was accepted, not that a person received or opened the message.

The parser accepts the two common shapes of Retry-After: a non-negative delay in seconds or an HTTP date. It also normalizes the date to UTC and refuses malformed input. A local capped backoff is a fallback, not permission to override a valid later time.

Notice what is absent. There is no sleep, recursive retry, or provider-specific response object in the signup path. Those omissions make the integration easier to test in a notebook and much safer to move into a worker. A Node.js service can use exactly the same states and transitions; Python is used here because a small protocol and fake adapter make the timing policy quick to exercise before production wiring.

Keep challenge lifetime separate from delivery attempts

Generating a fresh link for every retry looks harmless. It is not. Several messages can arrive out of order, leaving the person to guess which link is current. Reuse one active challenge until it expires or is consumed, and rotate it only under an explicit security policy. Store a digest of the secret rather than the secret itself, bind it to the intended signup, and make consumption atomic so two clicks cannot complete two transitions. This distinction also makes resend behavior easier to reason about. A resend request usually means "try delivery again," not "change the credential." The delivery job may have many attempts, while the challenge has one lifecycle. The same discipline applies to logs: record the job ID, attempt number, resulting state, status class, and scheduled retry time, but don't log the raw verification URL or secret. Email addresses should be minimized or transformed according to the system's privacy requirements. One dense record of the transition is more useful during an incident than five disconnected log lines, especially when workers overlap and the question is which one owned the job at a specific instant.

For browser-visible cooldowns, use server time as the authority. A countdown rendered by the client is convenient feedback, but refreshing the page or changing a device clock must not bypass the stored eligibility timestamp. Return an absolute eligibility time or remaining duration from the application, then recheck it on the next request.

Test the timing rules before wiring an API

I favor an eval-like table of cases here because delivery timing has the same failure mode as a prompt pipeline: the happy path is obvious, while edge cases quietly change behavior. Freeze now, feed synthetic adapter results into the worker, and assert the persisted transition. This can run before any network credential exists.

Input Expected state Expected next action
202 Accepted accepted No automatic resend
429 with Retry-After: 120 cooling_down Eligible 120 seconds after now
429 with a valid future HTTP date cooling_down Eligible at that UTC instant
429 with malformed Retry-After cooling_down Use capped local backoff
Timeout with an unknown outcome cooling_down Retry with the same job ID
Two workers claim one row one sending claim Only one adapter call

The timeout case deserves extra attention. The remote system may have accepted the request before the connection failed. Reusing a stable idempotency key gives an adapter the information it needs to suppress a duplicate when that facility exists. The local database must still prevent concurrent claims, because an idempotency key is not a replacement for queue ownership.

Unknown means unknown.

Test the public endpoint separately. Repeated requests for the same signup should produce the same neutral response shape, respect the product cooldown, and avoid disclosing internal states. Then test the worker with a fake clock and a scripted adapter. No real inbox is needed for most of this suite, which keeps feedback fast and avoids spending delivery quota on deterministic cases.

Measure integration effort at the boundaries

The useful comparison is not how many lines make the first message appear. Count the boundaries the team must own after launch: atomic job claiming, challenge storage, retry scheduling, adapter normalization, logs, metrics, and the user-facing cooldown. A library that performs an HTTP request may reduce initial typing while leaving all seven concerns in the application. This is the concrete integration-effort trade-off: a four-state worker adds a table and scheduler, but it removes transport waiting from web processes and makes rate-limit behavior testable.

Watch three families of signals. Queue age shows whether eligible jobs are waiting too long. State-transition counts show whether 429 or timeout paths are growing. Challenge outcomes connect delivery work to the actual goal: completed marketplace signup. Keep those metrics aggregated; observability should not become a second store of verification secrets.

Cost belongs in the review, but it is not the primary decision. Retries consume requests and operational attention, so cap attempts and define a terminal review path. More important is predictable ownership: somebody must know which component decides eligibility, which one interprets throttling, and which one can invalidate a challenge.

The limitation is real: this pattern is too much machinery for a prototype that can tolerate manual recovery, and it is incomplete for very high throughput without partitioning, dead-letter handling, and capacity planning. Direct synchronous sending has less integration effort at tiny scale. The worker becomes the better practice when signup must stay responsive during an email API rate limit and the team needs deterministic retry-after behavior. That choice spends storage and operational complexity to gain control over timing; it isn't free infrastructure.

Before release, walk through the system in prose. Confirm that signup creates at most one active challenge and one eligible job, the worker claims atomically, and every 429 schedules work instead of blocking a process. Confirm that accepted jobs stop, ambiguous outcomes reuse the job ID, expired challenges cannot be consumed, and public responses stay neutral. Finally, rehearse a sustained throttle with the fake adapter and verify that queue age, transition metrics, and alerts tell the same story.

Further reading

Top comments (0)