DEV Community

BartholomewVance6831
BartholomewVance6831

Posted on

Node.js Custom Password Reset Email API vs Outbox Polling (Choose Outbox)

TL;DR: Send directly from the auth callback while a marketplace is still an MVP and delivery state does not drive support or risk decisions. Choose an outbox relay once signup verification and password reset share a production path, especially when the email API has no webhooks. The relay costs one table and one worker, but it keeps request latency, retries, and polling outside the token-issuing transaction.

Choice Integration work Failure boundary Best fit
Direct send One HTTP call in the auth path Signup and email fail together Prototype or internal marketplace
Outbox relay Table, worker, status poller Auth commits before transport work Public marketplace with support controls

My choice is the outbox relay after MVP. A direct call wins until the team needs independent retries, auditable state, or one transport path for both verification and recovery. Then its tiny initial diff becomes glue spread across controllers, scheduled jobs, and support queries.

Should auth own the custom password email provider API?

The auth system should create and validate the verification secret. The mail layer should receive a short-lived link or opaque send payload; it should not mint its own reset token or decide whether an account is verified. Keeping that line sharp prevents a retry from creating a second security state.

The products named in the usual version of this question expose different boundaries. Supabase Auth documents custom SMTP for auth messages. Clerk documents email templates and custom flows. Auth.js, formerly NextAuth.js, documents email providers for email sign-in, while a credentials-based password-recovery flow remains application work. Those surfaces can all lead to an email, but they do not assign token creation, templates, and delivery tracking to the same component. Compare ownership before API syntax.

For a marketplace signup, persist an intent such as marketplace.verify_signup, the account ID, a recipient reference, an expiry, and a deduplication key. Avoid storing the raw token in logs or general event metadata. Recovery can use the same transport contract with a different intent and policy.

Do less here. The auth transaction inserts one outbox row and returns. It does not wait for an external email service, and it does not poll.

One write. Then stop.

Why is polling harder than sending?

A successful Fetch promise is not proof of delivery. The Fetch API resolves when an HTTP response is available, including responses with error status codes, so code must inspect response.ok or the status explicitly. An accepted send response describes provider intake, not necessarily mailbox delivery.

Without webhooks, the application has to reconcile later. Poll by the provider message identifier, map remote states into a small local vocabulary, and stop when the state is terminal or the local observation window expires. Keep accepted, delivered, temporarily_failed, permanently_failed, and unknown distinct. In particular, unknown must not become delivered because a polling request timed out.

The nasty edge is ambiguity. If the POST times out after the remote service accepted it, a blind retry can send a second verification message. Give each logical email a stable idempotency key when the API supports one. If it does not, record the attempt before the call, retain the remote ID as soon as it arrives, and make duplicate behavior explicit. One extra message may be tolerable; two independently valid tokens may not be. Token policy decides that outcome, not the mail API.

This is the trade-off I care about: the relay adds moving parts before it removes uncertainty. A worker can crash after a remote acceptance but before saving the remote ID; a poll can outlive the provider's status-retention window; and a delayed verification message can reach someone after the account requested another link. The design needs a lease on each job, a stable logical ID across retries, and a token policy that makes older links invalid when appropriate. None of that comes free with a queue. It is still less invasive than teaching every auth handler its own retry rules once verification and password recovery share the transport.

Use backoff with jitter, honor Retry-After when supplied, and cap concurrent status requests. HTTP defines 429 Too Many Requests for rate limiting and permits a server to indicate how long to wait. A tight one-second loop is configuration disguised as code. It fails precisely when the provider is under pressure.

A small TypeScript boundary

The useful abstraction is not a universal provider SDK. It is a narrow adapter around the two operations this system owns: submit and observe.

type DeliveryState =
  | "accepted"
  | "delivered"
  | "temporarily_failed"
  | "permanently_failed"
  | "unknown";

type VerificationEmail = {
  jobId: string;
  recipient: string;
  verificationUrl: string;
  expiresAt: string;
};

interface MailTransport {
  submit(message: VerificationEmail): Promise<{ remoteId: string }>;
  observe(remoteId: string): Promise<DeliveryState>;
}

class HttpMailTransport implements MailTransport {
  constructor(
    private readonly baseUrl: string,
    private readonly token: string,
  ) {}

  async submit(message: VerificationEmail): Promise<{ remoteId: string }> {
    const response = await fetch(`${this.baseUrl}/messages`, {
      method: "POST",
      headers: {
        authorization: `Bearer ${this.token}`,
        "content-type": "application/json",
        "idempotency-key": message.jobId,
      },
      body: JSON.stringify(message),
    });

    if (!response.ok) {
      throw new Error(`mail submission failed with ${response.status}`);
    }
    return (await response.json()) as { remoteId: string };
  }

  async observe(remoteId: string): Promise<DeliveryState> {
    const response = await fetch(`${this.baseUrl}/messages/${remoteId}`);
    if (!response.ok) return "unknown";
    const body = (await response.json()) as { state: DeliveryState };
    return body.state;
  }
}
Enter fullscreen mode Exit fullscreen mode

Those endpoint names and that response schema form an adapter contract, not a claim about a commercial API. Each implementation must validate the real schema. Keep credentials in the worker environment, set request timeouts with AbortSignal, and redact the verification URL from structured logs.

Test with a fake transport that returns a scripted sequence: accepted, then temporarily_failed, then delivered. Also test a submit timeout followed by a retry with the same job ID. These tests catch more integration bugs than snapshotting an HTML template.

In production, expose counts by intent and normalized state, retry age, and the oldest unprocessed outbox row. Alert on backlog age rather than raw failure count; traffic changes make counts noisy. Support needs lookup by account and job ID, but the view should reveal neither token nor full verification URL.

The two criteria I would benchmark

First, measure time-to-first-safe-call. That is not how quickly a demo email arrives. Count the code and configuration required for an authenticated request, timeout, idempotency behavior, redacted logging, and one deterministic test. An API that needs fewer lines for the POST but offers no stable way to reconcile an ambiguous response can be the larger integration.

Second, trace state ownership. Make a one-page sequence diagram and mark who owns token issuance, outbox commit, message acceptance, delivery observation, and user-visible resend. Any state with two owners will produce glue. Any state with no owner will become a support ticket.

Run the same contract test against each candidate adapter and record behavior: accepted status, error shape, retry guidance, idempotency support, status retention, and rate-limit signals. Do not invent a composite score. A hard requirement should fail a candidate plainly; the remaining differences belong in code review.

Config bloat is a warning sign. If switching transport requires auth-controller changes, template-policy changes, and new database states, the boundary is leaking. If it requires one adapter plus credentials, the design is doing its job.

Five states are enough for this adapter. More states need evidence.

When direct send is the better choice

Use direct send when verification is non-critical, traffic is modest, a failed request can be retried by the user, and nobody needs delivery evidence. It is also reasonable for an internal marketplace pilot where the team accepts coupled availability and can replace the path before launch.

The relay is not a fit for a throwaway prototype, a system with no worker runtime, or a flow where the auth platform fully owns delivery and exposes the evidence the team needs. In those cases, use the platform's supported SMTP or email integration instead of duplicating its state. Direct send also has an honest advantage: fewer credentials, deployments, dashboards, and failure modes on day one.

That limit matters.

Keep even that version behind the same MailTransport interface. Set a timeout. Check HTTP status. Return a generic UI response so account existence is not disclosed by recovery behavior. The OWASP Forgot Password guidance recommends consistent messages and response timing, side-channel delivery, single-use expiring tokens, and rate limiting. Those controls belong in the application regardless of transport shape.

Do not build a queue because distributed systems look serious. Build it when one auth request can no longer carry the operational contract for retries and observation. That is the boundary.

References

Top comments (0)