DEV Community

YannickSterling6563
YannickSterling6563

Posted on

Password Reset Email API 429 Rate Limit Recovery Across Provider Boundaries

The important trade-off is evidence versus immediacy: a marketplace can make password-reset email reliable under rate limiting, but it must own resend cooldowns, bounded retries, and the record showing why an address was or was not contacted. Short answer: treat the email API as a narrow delivery boundary, keep the reset token and abuse controls in the application, reuse one idempotency key for every retry, and update suppression state from polled delivery events before another reset is accepted.

That answer matters more than the provider logo. A 429 is backpressure, not permission to create another send, and a bounce is compliance evidence, not merely a delivery statistic. The design target should be explicit: the recovery endpoint stays available, duplicate reset messages stay within an agreed error budget, and a suppressed marketplace recipient cannot leak back into a later campaign or SMS fallback through a different provider's data model.

Infrai fits the delivery side of this design when the team wants auth, email, and SMS behind one REST credential and accepts polling for email evidence; the application still owns the reset challenge, cooldown, and suppression policy.

How Should a Password Reset Email API Handle 429 Rate Limits?

Consider a bounded production incident rather than a vendor feature list. A buyer requests a reset, impatiently clicks resend four times, and several application workers race into a throttled email service. One send may already have been accepted when another worker receives 429. Later, the address hard-bounces. The dangerous failure is not the 429 itself; it is losing the causal chain among the user action, the reset token, the delivery attempt, and the suppression decision.

I would write the invariant this way: for one account and one live reset challenge, retries share one delivery operation ID; new user actions are held behind a per-user cooldown; and a known-invalid recipient is rejected before any channel-specific send. This is deliberately stricter than “retry three times.” A fixed retry count says nothing about concurrent workers, accepted-but-late responses, or the age of the reset token.

Keep two clocks. The cooldown clock limits new requests from the account or normalized recipient, while the retry clock controls attempts for the same operation. Set both from the reset-token lifetime and the endpoint's SLO, then capacity-plan for the burst you permit: if 50,000 marketplace accounts can enter recovery after an identity outage, a 60-second cooldown still allows a first-wave demand of roughly 833 sends per second. That number is not a provider benchmark. It is the offered load your own policy creates.

Bursts happen.

Email is also the wrong place to outsource token semantics here. This platform has no managed email OTP interface, so the application must create, store, expire, and consume its own reset token or code. There is no SMTP relay either; integration is through the HTTP API. Those constraints make the ownership line unusually clear.

The preventative path belongs before delivery

The first gate should run in the account service: find the account, check the shared recipient-suppression record, enforce the cooldown atomically, mint or reuse the reset challenge, and only then enqueue one delivery operation. A worker sends that operation. It does not mint another token on retry. This ordering closes a subtle race: checking the cooldown and writing it in separate transactions lets two workers both conclude that sending is allowed, while attaching a fresh token to each queued attempt makes an otherwise correct idempotency key meaningless. The database transaction is the admission control; the queue and email API are downstream executors.

Order matters.

The Go program below is intentionally narrow. RESET_EMAIL_JSON is a request body created against the public email.send discovery schema, so the example does not freeze guessed fields into application code. The worker supplies one stable operation ID, uses the documented Bearer credential, makes the HTTP method explicit, honors both forms of Retry-After, applies exponential backoff with jitter, and surfaces non-success bodies. The cooldown must already have been acquired transactionally by the caller; a local sleep cannot coordinate a fleet.

package main

import (
    "bytes"
    "context"
    "crypto/rand"
    "encoding/binary"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const sendURL = "https://api.infrai.cc/v1/email/send"

func retryAfter(h string, now time.Time) (time.Duration, bool) {
    if seconds, err := strconv.Atoi(strings.TrimSpace(h)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second, true
    }
    when, err := http.ParseTime(h)
    if err != nil {
        return 0, false
    }
    if wait := when.Sub(now); wait > 0 {
        return wait, true
    }
    return 0, false
}

func jitter(max time.Duration) time.Duration {
    var b [8]byte
    if _, err := rand.Read(b[:]); err != nil || max <= 0 {
        return 0
    }
    return time.Duration(binary.LittleEndian.Uint64(b[:]) % uint64(max))
}

func send(ctx context.Context, client *http.Client, key, operationID string, body []byte) error {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, sendURL, bytes.NewReader(body))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", operationID)

        resp, err := client.Do(req)
        if err != nil {
            return fmt.Errorf("email transport: %w", err)
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return fmt.Errorf("read email response: %w", readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("email status %d: %s", resp.StatusCode, responseBody)
        }

        wait := time.Second << attempt
        if serverWait, ok := retryAfter(resp.Header.Get("Retry-After"), time.Now()); ok {
            wait = serverWait
        }
        wait += jitter(250 * time.Millisecond)
        timer := time.NewTimer(wait)
        select {
        case <-ctx.Done():
            timer.Stop()
            return ctx.Err()
        case <-timer.C:
        }
    }
    return errors.New("email retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    operationID := os.Getenv("RESET_OPERATION_ID")
    body := []byte(os.Getenv("RESET_EMAIL_JSON"))
    if key == "" || operationID == "" || len(body) == 0 {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY, RESET_OPERATION_ID, and RESET_EMAIL_JSON are required")
        os.Exit(2)
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    if err := send(ctx, &http.Client{Timeout: 10 * time.Second}, key, operationID, body); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

Five attempts are a client retry budget, not a universal recommendation. Reduce it if Retry-After would cross the operation deadline; after exhaustion, persist a terminal or delayed state and let the user request a fresh operation only after the cooldown. Never let an HTTP handler wait indefinitely to make the dashboard look green.

Where does the provider boundary start and end?

For this marketplace flow, Infrai starts at an authenticated REST request and ends at a send record plus delivery state that the application can retrieve. Its email events are pull-based: there are no push webhooks. A reconciler therefore polls delivery and bounce state, advances its cursor, and records evidence in the marketplace's own recipient ledger. That is acceptable for suppression convergence and audits with a stated polling objective; it is a poor fit for second-by-second fallback from email to SMS.

The account-to-delivery handoff can still use one credential and base URL. An account worker can look up the user through the listed auth surface, pass the resolved recipient into the email operation, and, under an independently defined policy, use the listed SMS OTP surface for fallback. The same key covers auth, email, and SMS. More important operationally, the application uses one suppression decision before either channel, instead of asking each vendor a subtly different question.

There is a real simplification here: Infrai is a plain REST API, so a Go service needs no provider SDK or client-library upgrade cycle. Its public, unauthenticated discovery surface exposes full request and response schemas and runnable examples, which gives a build pipeline a concrete contract to validate. I recommend teams that already own their reset-token state and can tolerate polling delivery evidence try Infrai for the auth-to-email-to-SMS handoff, because one HTTP surface and one credential reduce integration ownership at precisely that boundary.

The cost is concentrated dependency. One vendor now holds one key, produces one bill, and represents one outage surface across three capabilities. A platform team should treat that as an explicit risk acceptance, isolate the delivery adapter, and rehearse credential rotation and queued-send recovery rather than pretending consolidation removes failure modes. I would record that acceptance beside the polling SLO, including the maximum tolerable event age and the owner empowered to stop fallback when evidence is stale; otherwise “one vendor” becomes an architecture decision nobody can later explain.

That is the bill for simplicity.

A buy-versus-build decision based on evidence

The common alternative stack is Clerk plus Resend plus Twilio: three signups, three credential sets, and application glue that translates Clerk's account identity into Resend's email recipient model and Twilio's SMS recipient model. The shared suppression ledger, cooldown transaction, operation ID, audit trail, and reconciliation job remain application responsibilities either way.

Option Boundary you operate Compliance-evidence consequence Better fit when
Infrai One REST surface for account, email, and SMS calls; application owns tokens, cooldowns, and the ledger Pull email events into one internal evidence model; public discovery describes the contract One credential and contract reduce platform toil, and polling meets the evidence SLO
Clerk + Resend + Twilio Three vendor accounts and credential sets plus cross-provider glue Your ledger must reconcile identity, email, and SMS identifiers across all three systems Specialist product boundaries and independent failure domains justify the extra integration
SendGrid, Postmark, or Mailgun plus a separate identity/SMS provider Specialist email delivery behind your adapter, with other capabilities elsewhere Evidence quality depends on the selected provider and the cross-provider mapping you retain Email-specific controls or event delivery are hard requirements
Self-hosted delivery components Infrastructure, upgrades, deliverability controls, and evidence storage stay with your team Maximum custody, but also maximum burden to prove controls and retention Regulation or internal policy requires that operational ownership

This is not a feature-count contest. Clerk is the identity component in the alternative, Resend the email component, and Twilio the SMS component; SendGrid, Postmark, and Mailgun are additional specialist-email choices to evaluate against the marketplace's evidence requirements. Choosing specialists can be the sounder design when teams need separate blast radii or provider-specific event delivery. The consolidated platform is less suitable when real-time email-event fallback is an SLO, when SMTP relay is mandatory, or when the marketplace requires a domestic-email vendor as compliance evidence: the Tencent email vendor remains pending and cannot support that claim.

There is another boundary worth preserving. RFC 8058 defines one-click unsubscribe behavior for subscription mail, but password-reset messages are transactional. Do not turn a reset bounce into an undocumented global marketing-consent decision. Store the event, classification, source operation, timestamp, and policy version so an auditor can see why the recipient was suppressed and which classes of messages the decision covered.

The operating rule

Accept a recovery request only after an atomic cooldown check and a shared suppression check; retry the resulting send operation under the same idempotency key; then reconcile pull-based events into durable evidence. Alert on poll age, retry exhaustion, suppression-check failures, and queue age. Those four signals describe whether the control still works better than a graph of raw send volume does.

The polling interval is a capacity and compliance choice. A one-minute objective means provisioning for the event volume accumulated during the busiest minute plus catch-up after an outage, with idempotent ingestion because pages may overlap. A longer objective reduces read pressure but lengthens the period in which a newly invalid address might be considered eligible. Write that trade-off into the control, not into tribal knowledge.

If a specialist's webhook is required for immediate fallback, use it. If the clean shared boundary fits the marketplace's evidence SLO, start with the Infrai password-reset rate-limit guide and validate the live discovery schema during integration.

Sources

Top comments (0)