An email verification endpoint often starts as a small feature: create a token, send a message, and return 202 Accepted. It becomes more interesting when the provider times out after accepting the request, a user clicks “send again” three times, or a worker crashes between delivery and acknowledgement.
The useful unit of design is not the send call. It is the verification attempt: a durable record of why a message was requested, what happened to it, and whether another attempt is allowed. This gives an authentication service a clear state model and makes failures inspectable.
Why a send button is not a delivery model
There are at least three different events in a verification flow:
- The user requests verification.
- The application asks a delivery provider to send a message.
- The provider, webhook, or a periodic poll reports an outcome.
Treating these as one operation creates ambiguous states. If the HTTP request times out, the server does not know whether the provider rejected the message or accepted it just before the connection closed. Retrying blindly can create duplicate emails and confusing token races.
The first version usually look simple, but the ambiguity appears as soon as there is real traffic. A verification attempt should therefore have its own identity, timestamps, and state transitions.
Model the verification attempt explicitly
A small PostgreSQL table is enough for the durable part of the workflow:
create table email_verification_attempts (
id uuid primary key,
user_id bigint not null references users(id),
token_digest bytea not null,
status text not null check (status in ('queued', 'sending', 'sent', 'confirmed', 'failed', 'expired')),
idempotency_key text not null,
provider_message_id text,
attempt_number integer not null default 1,
expires_at timestamptz not null,
created_at timestamptz not null default now(),
updated_at timestamptz not null default now(),
unique (user_id, idempotency_key)
);
Store a digest of the token instead of the raw token. The raw value is only returned in the email link, while the database needs to verify it. A unique idempotency key lets a client retry the same logical request without creating another attempt. A separate attempt_number supports an explicit “send again” action with a new key.
That distinction are important: a network retry is not the same as a user request for a new message. The API can safely replay the first result, while a deliberate resend can invalidate or supersede earlier attempts according to a documented policy.
For a useful application-level comparison, see this versioned invite email flow. The same separation between a user action and a message attempt is especially useful when more than one Node.js instance can receive the same request.
Use PostgreSQL to make retries safe
Create the attempt and its outbox event in one transaction. The application should not call an external provider while holding the database transaction open.
begin;
insert into email_verification_attempts (
id, user_id, token_digest, status, idempotency_key, expires_at
)
values ($1, $2, $3, 'queued', $4, now() + interval '30 minutes')
on conflict (user_id, idempotency_key) do nothing;
insert into email_outbox (event_type, aggregate_id, payload)
select 'email_verification_requested', id, jsonb_build_object('attempt_id', id)
from email_verification_attempts
where user_id = $2 and idempotency_key = $4
on conflict (event_type, aggregate_id) do nothing;
commit;
The worker claims queued rows with a short lease. A lease prevents two workers from processing the same row at once, but it does not pretend that an external provider is transactional with PostgreSQL. If the worker loses its connection after the provider accepted the message, the lease eventually expires and the row may be retried. The provider request should include the attempt ID as its own idempotency key when that feature exists.
Separate the API transaction from delivery
The API should commit a queued attempt quickly and return a stable response. A worker then changes it to sending, calls the provider, and records sent or failed. Webhook processing can move it to confirmed only after validating the token and its expiry. These details often gets skipped in the first implementation.
This boundary makes operational behaviour easier to reason about:
- A database failure returns an error and creates no queued attempt.
- A provider timeout leaves an attempt that can be reconciled.
- A duplicate HTTP request returns the existing attempt rather than sending again.
- A user-initiated resend creates a new attempt under a rate limit.
If the service emits events to another system, explicit typed verification states keep the database change and event publication aligned across the UI and worker. It also gives you a place to attach a correlation ID, which is handy when debugging a tem email report from production. The retry path also need a clear owner.
Design useful status responses
Do not expose internal provider details or reveal whether an account exists. A public response can be deliberately boring:
{
"status": "accepted",
"request_id": "req_7f1...",
"next_check_after_seconds": 5
}
For a repeated idempotent request, return the same logical status and request ID when possible. For a resend, return a new request ID. A 429 response should include a retry hint, but the limit must apply to both successful and failed requests or attackers can use errors to bypass the budget.
The confirmation endpoint should distinguish an expired token, an already-used token, and a malformed token internally, while presenting a safe public message. Log the reason with a correlation ID, never the token itself. It is easy to accidentally paste a tamp mail com address or token into a debug line during an incident, so redaction deserves a test.
For test environments, a disposable email address generator can be useful for isolated fixtures, but production verification still needs provider controls, retention rules, and abuse monitoring. If a test workflow needs to create temporary mail, keep that dependency in the fixture layer and never use it as a substitute for production identity policy.
A practical review checklist
Before shipping an email verification API, check the following:
- Is every logical request identified by an idempotency key?
- Can a worker retry after a timeout without creating an unbounded duplicate?
- Are
queued,sending,sent,confirmed,failed, andexpiredtransitions explicit? - Is token material hashed at rest and excluded from logs?
- Are resend and confirmation endpoints rate-limited separately?
- Can an operator find the attempt from a request or correlation ID?
- Are old attempts expired and deleted according to a retention policy?
An attempt ledger does not make email delivery perfect. It makes uncertainty visible, bounded, and recoverable. That is the real improvement: authentication code can fail without losing the story of what the system tried to do.
Top comments (0)