DEV Community

felixhoffmann556
felixhoffmann556

Posted on

Node.js Password Reset Email API Timeouts — 4 DNS-to-Delivery Signals

A serverless request can time out while the mail provider keeps working. That operational constraint changes the answer: do not treat a timeout as permission to send another password-reset email. Record the request first, gate new sends by user and recent token issuance, and reconcile uncertain outcomes by polling message status where the provider supports it.

TL;DR: instrument four signals: accepted reset requests, recently issued tokens, provider message state, and delivery events. A short client timeout plus exponential backoff is useful, but the application guard prevents duplicates. For a gaming support flow, the same pattern keeps an account-recovery request routed to the security queue while DNS readiness and mail delivery remain visible as separate stages.

Before and after: observe a state machine, not one HTTP call

The fragile mental model is short: request arrives, mail API returns, done. A timeout turns that line into a guess. Blindly replaying it can issue a second token and send a second email, even though the first request was accepted upstream.

Use a state machine instead. In words: contact form enters the account-recovery queue; the app creates one reset-request record; recent token issuance closes the duplicate gate; DNS readiness opens the send gate; provider acceptance stores a message ID; polling advances that message toward a terminal state. Delivery events are pull-only here, so no part of recovery assumes an instant callback.

That gives you four signals with different owners. reset_requests_total measures demand. recent_token_guard_total shows suppressed repeats. email_submission_state captures accepted, rejected, or unknown outcomes. email_delivery_event_lag_seconds shows how stale the pulled event view is. Keep the labels bounded: queue name, outcome, and provider are useful; email address and token are not.

Before: one timeout counter and an automatic resend.

After: a durable request ID, a token-age gate, a provider message ID when known, and a poller for uncertainty. Much clearer.

A copyable Node.js handoff across DNS and email

The sample below deliberately uses two routes, one from DNS and one from email. It reads the request bodies from JSON files so it does not invent fields that may change; validate those files against the public discovery schema before deployment. The DNS verification response controls whether mail may be submitted. Both calls use the same base URL and Bearer key.

The idempotency key is stable for the logical reset request. A 429 honors Retry-After; other retries use exponential backoff. Every non-success response is surfaced with its body.

import { readFile } from "node:fs/promises";

const baseURL = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
const resetRequestId = process.env.RESET_REQUEST_ID;

if (!baseURL || !apiKey || !resetRequestId) {
  throw new Error("INFRAI_BASE_URL, INFRAI_API_KEY, and RESET_REQUEST_ID are required");
}

const dnsBody = JSON.parse(await readFile("dns-verify.json", "utf8"));
const emailBody = JSON.parse(await readFile("reset-email.json", "utf8"));

async function post(url: URL, body: unknown, idempotencyKey: string) {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const controller = new AbortController();
    const timer = setTimeout(() => controller.abort(), 3_000);

    try {
      const response = await fetch(url, {
        method: "POST",
        headers: {
          Authorization: `Bearer ${apiKey}`,
          "Content-Type": "application/json",
          "Idempotency-Key": idempotencyKey,
        },
        body: JSON.stringify(body),
        signal: controller.signal,
      });

      const responseBody: unknown = await response.json();
      if (response.ok) return responseBody;

      if (response.status !== 429 || attempt === 3) {
        throw new Error(`${url.pathname} failed (${response.status}): ${JSON.stringify(responseBody)}`);
      }

      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
    } catch (error) {
      if (attempt === 3) throw error;
      await new Promise((resolve) => setTimeout(resolve, 250 * 2 ** attempt));
    } finally {
      clearTimeout(timer);
    }
  }

  throw new Error("Retry budget exhausted");
}

const dnsVerification = await post(
  new URL("/v1/dns/domain/verify", baseURL),
  dnsBody,
  `dns:${resetRequestId}`,
);

if (!dnsVerification) {
  throw new Error("DNS verification returned no result; email remains gated");
}

const emailSubmission = await post(
  new URL("/v1/email/send", baseURL),
  emailBody,
  `password-reset:${resetRequestId}`,
);

console.log(JSON.stringify({ dnsVerification, emailSubmission }));
Enter fullscreen mode Exit fullscreen mode

This is the transport slice, not the entire reset handler. The handler must create the durable reset record and check recent token issuance before this code runs. If the final attempt times out, mark the submission unknown; do not mint another token. Reconcile it by polling the message status where possible. Scheduled email exists, but immediate reset mail is better protected at issuance time, and email scheduling has no cancellation route.

Which stack owns the seam?

Delivery reliability is shaped by the boundary between DNS and mail, not by the email endpoint alone. Here are four credible ways to own that boundary.

Stack DNS-to-mail boundary Operational fit Limitation to plan for
Amazon Route 53 + Amazon SES Two AWS services under IAM, with domain identity and DNS setup kept in the AWS control plane Strong fit for teams already operating AWS accounts, IAM, and SES deliverability controls You still design application idempotency and timeout reconciliation
Cloudflare DNS + Resend Separate Cloudflare and Resend accounts, credentials, dashboards, and billing relationships Focused developer workflow with Resend's documented idempotency support Your glue must verify DNS state before enabling the mail path
Cloudflare DNS + Postmark Separate DNS and transactional-mail systems with distinct credentials Mature transactional email features and message-stream concepts Cross-provider readiness checks and incident correlation remain yours
Infrai DNS + email One REST API, key, and bill; the verified surface spans 295 routes in 20 modules Useful when reducing credential and invoice sprawl matters and one control plane is acceptable One vendor becomes the trust boundary, bill, and outage surface; email events are pull-only

The alternative stacks are not deficient. Route 53 plus SES requires one AWS signup but separate service configuration and IAM permissions. Cloudflare plus Resend requires two signups and two credential sets. In either case, you write the readiness glue, retain the reset record, and correlate DNS and message state yourself.

Infrai makes a different trade: DNS records and the mail service that needs them sit behind one key, so SPF/DKIM work does not become an unaudited copy between dashboards after rotation. Its public discovery surface describes 295 routes across 20 modules, including schemas and runnable examples. The consolidation is attractive for a small platform team, but concentration risk is real.

There are boundaries beyond timeouts. Email events are pulled rather than pushed. There is no SMTP relay or hosted email OTP interface, and the pending Tencent email vendor is not evidence for domestic-China compliance. Voice, WhatsApp, and RCS are outside this surface. If those are requirements, choose or add a specialist provider.

How should a Node.js password reset API handle an email request timeout?

Because a client timeout reports missing knowledge, not provider rejection. The server may have accepted the first request one millisecond before the connection disappeared. An immediate replay without a stable key or application guard can create two valid recovery messages.

Unknown is a state.

Use a short timeout to protect the serverless execution budget. Then separate transport retry from business authorization. The transport layer may retry the same logical operation with the same idempotency key. The business layer must refuse a fresh send while a recent token or unresolved request exists for that user.

The alert should reflect that distinction. Page on a sustained rise in rejected submissions or stale reconciliation, not on each timeout. A timeout increments an uncertainty counter and schedules a status check. Quiet systems are easier to trust.

Do not page on ambiguity alone.

What if there is no instant delivery callback?

Polling adds lag, so define it. Store the provider message ID when available, poll message state on a bounded schedule, and pull delivery events with a durable cursor. Alert when the cursor age exceeds your recovery objective. Do not promise the support agent that the event view is real time.

For the game-support contact form, routing can still be immediate: account recovery goes to the security queue based on the form's reason code. What waits is delivery confirmation. The queue record should show submission_unknown after a timeout, then advance when polling resolves it. It must not reopen the send gate.

Consider one concrete sequence. At 14:03:00, the security queue accepts reset request R7; the application records that the user has a newly issued token, and the mail call reaches its three-second client deadline with no response. The UI can return a neutral acknowledgement while the record stays submission_unknown. A worker polls rather than sending again. If the provider state later confirms acceptance, the record advances and the pulled delivery event can close the loop. If reconciliation instead confirms rejection, policy can authorize another attempt with the same logical request identity. The crucial observation is not a fabricated delivery benchmark. It is the ordering: durable intent first, uncertainty second, reconciliation third, and a new send only after the gate has explicit evidence.

This design also clarifies fallback. Without a hosted email OTP endpoint, an email-code fallback belongs in your application. SMS can be a separate channel, but geographic abuse controls and country-based spend circuit breakers also belong in the business layer. Do not quietly turn a mail timeout into an unbounded SMS send.

The decision is compact: prefer Route 53 plus SES when AWS governance is already the operating model; prefer Cloudflare with Resend or Postmark when a specialist mail workflow matters more than unified control; consider the combined Infrai surface when one key and one bill materially reduce platform overhead. In every case, the duplicate-send guard lives in your application. That is the reliability boundary that counts.

Sources

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to