DEV Community

kevindev
kevindev

Posted on Edited on

Email Verification Events Need an Outbox

Email verification looks like one step in a signup flow: create a user, send a message, and wait for a click. In a real REST API, it is a small distributed workflow. PostgreSQL commits locally, an email provider accepts work remotely, and the browser may retry the request before either side has finished.

The useful design question is not “how do I send an email?” It is “how do I make the state change and the request to send an email agree when something fails?” A transactional outbox gives that question a bounded answer.

Why email verification is a distributed workflow

Suppose POST /signups creates an account and then calls an email provider. If the database commit succeeds but the process crashes before the provider call, the user never receives the message. If the provider accepts the message and the API times out, a retry can send two messages.

Some teams first test the happy path with a temporary email inbox or a temp email generator. That is useful for checking the UI, but it does not prove delivery behavior under a crash. For a practical phishing-safety checklist around temporary inboxes, see How Temporary Email Helps Avoid Phishing. A test named “tem email” can still miss the most important failure window.

The outbox pattern separates these concerns:

  1. The API transaction stores the user change and an email intent together.
  2. A worker reads committed intents and talks to the email provider.
  3. Retries and status fields make the result inspectable.

This make the signup endpoint predictable even when email delivery is slow.

The transaction boundary

The signup handler should write an account, a verification token, and an outbox row in one PostgreSQL transaction. The outbox row contains the event type, a stable aggregate ID, a deduplication key, and a JSON payload. It should not contain a password, a raw token, or more personal data than the worker needs.

The API can return 201 Created after the transaction commits. It should not wait for the email provider. A response such as { "status": "verification_pending" } tells the client what happened without claiming that the message has arrived.

For an implementation that uses an ORM, keep the outbox insert inside the same callback or unit-of-work object as the user insert. For a lower-level driver, use one client connection and one explicit BEGIN/COMMIT. Mixing a pool checkout for the account with a second connection for the outbox defeats the boundary.

A small PostgreSQL outbox schema

The schema can stay deliberately boring:

CREATE TABLE email_outbox (
  id              uuid PRIMARY KEY,
  event_type      text NOT NULL,
  aggregate_id    uuid NOT NULL,
  dedupe_key      text NOT NULL UNIQUE,
  payload         jsonb NOT NULL,
  status          text NOT NULL DEFAULT 'pending',
  attempts        integer NOT NULL DEFAULT 0,
  available_at    timestamptz NOT NULL DEFAULT now(),
  locked_at       timestamptz,
  sent_at         timestamptz,
  last_error      text,
  created_at      timestamptz NOT NULL DEFAULT now()
);

CREATE INDEX email_outbox_pending_idx
  ON email_outbox (available_at, created_at)
  WHERE status = 'pending';
Enter fullscreen mode Exit fullscreen mode

The unique dedupe_key is normally derived from the account ID and verification generation, not from the request ID. That distinction matters when a client retries the same signup request. A partial index keeps the polling query small as completed events accumulate.

Publishing events from Node.js

Workers should claim a bounded batch with row locks, then commit the claim before performing network I/O. FOR UPDATE SKIP LOCKED allows several workers to make progress without waiting on each other:

WITH next_events AS (
  SELECT id
  FROM email_outbox
  WHERE status = 'pending'
    AND available_at <= now()
  ORDER BY created_at
  FOR UPDATE SKIP LOCKED
  LIMIT 50
)
UPDATE email_outbox AS e
SET status = 'processing', locked_at = now(), attempts = attempts + 1
FROM next_events
WHERE e.id = next_events.id
RETURNING e.*;
Enter fullscreen mode Exit fullscreen mode

The Node.js worker then sends each message with a provider idempotency key, when the provider supports one. After success it sets status = 'sent'; after a temporary failure it returns the row to pending and calculates a future available_at. A worker crash leaves processing rows behind, so a lease timeout should move stale rows back to pending.

A queue library can still be useful, but the database remains the source of truth for the signup event. That makes local recovery and replay much more clear during an incident. For a broader view of preserving restore context around email operations, see restore context for email-driven operations. For browser-side cancellation, abortable email checks in a form flow covers a related boundary.

Retries, idempotency, and expiry

Not every error deserves a retry. Timeouts, connection resets, and provider 5xx responses are usually temporary. Invalid recipient data and a rejected domain are usually permanent. Store a short, sanitized error classification rather than dumping provider responses into logs.

The verification token should have an expiry in the database, and the confirmation endpoint should check it in the same transaction that marks the token as consumed. Replaying an old email must not restore access. A consumed_at column and a conditional update are enough for most systems:

UPDATE verification_tokens
SET consumed_at = now()
WHERE token_hash = $1
  AND consumed_at IS NULL
  AND expires_at > now()
RETURNING user_id;
Enter fullscreen mode Exit fullscreen mode

The worker may deliver a message more than once if the provider result is lost. Idempotent token consumption limits the damage. It also reduce pressure to build perfect exactly-once delivery, which is not a realistic promise across independent systems.

What to observe in production

Track outbox age, pending count, processing leases, attempt count, permanent failures, and time from commit to provider acceptance. Alert on the oldest pending event, not only on a worker process being alive. A worker that is running but cannot authenticate to the provider is still an outage.

Keep logs structured around outbox_id, aggregate_id, and dedupe_key hashes. Do not log verification tokens or complete email addresses. The logs is most useful when an engineer can follow one event without exposing the secret that it carries.

A practical checklist

  • Write the account, token, and outbox row in one transaction.
  • Give every event a stable dedupe key.
  • Claim work with a lease and SKIP LOCKED.
  • Separate temporary from permanent provider errors.
  • Use exponential backoff with a maximum attempt policy.
  • Expire and consume tokens atomically.
  • Measure oldest-event age and delivery latency.
  • Test crash windows, duplicate requests, and stale leases.

If a team calls a disposable test inbox “temp org mail,” that is fine as fixture vocabulary, but keep the production contract explicit: the API records intent, and a worker delivers it.

Questions engineers usually ask

Should the API wait for the provider?

No. Waiting couples request latency to provider latency and still cannot remove the ambiguous timeout case. Return after the local transaction commits and expose verification-pending state.

Is an outbox enough without a queue?

Often yes at modest volume. PostgreSQL polling is simple to operate. Add a dedicated queue when throughput, scheduling, or cross-service fan-out requires it, while retaining an auditable event record.

Can I use a temporary email address for production users?

That is a product and abuse-control decision, not an outbox decision. The delivery workflow should remain correct for normal addresses, temporary email testing, and provider failures alike.

The pattern is small: one transaction, one durable intent, and a worker with explicit retry state. Those boundaries make a REST API easier to explain, test, and operate.

Top comments (0)