Email verification is often treated as a small side effect of a signup request. In a real backend, it is a distributed workflow: an HTTP request writes state, a worker sends a message, and a provider may respond slowly or not at all. A worker crash between those steps can leave a user waiting or cause a duplicate email.
This post describes a PostgreSQL lease pattern I use for verification jobs. It applies whether the address is a normal customer mailbox, a throwaway email address used in a test, or data produced by a disposable email generator. The goal is not perfect delivery. The goal is to make ownership, retry, and recovery explicit.
Why the worker needs a lease
A queue consumer normally claims a job, performs work, and acknowledges it. The risky part is the gap between claiming and acknowledging. If the process dies in that gap, the job can be invisible forever. If the process retries without an ownership check, two workers may send the same verification message.
A lease gives each claim a short ownership window. The worker receives a random lease token and an expiry time. It may complete the job only while both values still match. If it crashes, another worker can reclaim the row after the lease expires.
This is simpler to reason about than a boolean such as is_processing. A boolean tells us that somebody touched the job; a lease tells us who owns it and until when.
For related delivery sequencing, see this practical guide to a PostgreSQL outbox for signup emails. The outbox solves the database-to-queue boundary, while the lease below solves worker ownership.
The PostgreSQL lease model
The table needs state that describes both the business result and the temporary worker claim:
CREATE TABLE verification_jobs (
id BIGSERIAL PRIMARY KEY,
user_id BIGINT NOT NULL,
email TEXT NOT NULL,
token_hash BYTEA NOT NULL,
status TEXT NOT NULL CHECK (status IN ('pending', 'processing', 'sent', 'failed')),
attempts INTEGER NOT NULL DEFAULT 0,
available_at TIMESTAMPTZ NOT NULL DEFAULT now(),
lease_token UUID,
lease_expires_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
sent_at TIMESTAMPTZ
);
CREATE INDEX verification_jobs_claim_idx
ON verification_jobs (available_at, lease_expires_at, id)
WHERE status IN ('pending', 'processing');
token_hash is stored instead of the raw token. The API can give the raw token to the user once, but a database read should not be enough to reconstruct a valid link. For a useful companion, this explanation of a privacy budget for email verification covers what else should not leak through logs and metrics.
Claiming work without duplicate sends
FOR UPDATE SKIP LOCKED lets many workers scan the same table without waiting on rows already claimed by another transaction. The update and select happen as one statement, so the lease token is created at the point of ownership:
WITH candidate AS (
SELECT id
FROM verification_jobs
WHERE available_at <= now()
AND (
status = 'pending'
OR (status = 'processing' AND lease_expires_at < now())
)
ORDER BY id
FOR UPDATE SKIP LOCKED
LIMIT 1
)
UPDATE verification_jobs AS job
SET status = 'processing',
lease_token = gen_random_uuid(),
lease_expires_at = now() + interval '2 minutes',
attempts = attempts + 1
FROM candidate
WHERE job.id = candidate.id
RETURNING job.id, job.email, job.token_hash, job.lease_token, job.lease_expires_at;
The interval should be longer than the normal provider call, but not so long that a dead worker blocks recovery. In practice, the worker should renew the lease for slow calls or split the operation into smaller steps. Two minutes is an example, not a universal setting.
The send call itself must also be designed for retries. Pass a stable message or operation identifier to a provider that supports idempotency. If the provider does not support it, the database can prevent two active workers, but it cannot undo a successful send followed by a lost network response. That limit needs to be documented for the REST API.
Making retries and expiry explicit
When the provider accepts the message, complete the job only with the original lease:
UPDATE verification_jobs
SET status = 'sent',
sent_at = now(),
lease_token = NULL,
lease_expires_at = NULL
WHERE id = $1
AND status = 'processing'
AND lease_token = $2;
If this updates zero rows, the worker no longer owns the job. It should not report success locally. The message may already have been sent by a worker whose lease expired, so the safest action is to record the ambiguous result and let the reconciliation policy decide.
For a temporary provider failure, set status = 'pending' and calculate available_at with bounded exponential backoff. Permanent failures, such as an invalid destination, should become failed with a reason that is safe to expose. Do not keep retrying an address typed as dummy e mail forever just because the first network request timed out.
The API should also expire the verification token independently of the job lease. A two-minute worker lease is about processing ownership; a verification link may be valid for 15 minutes or one hour. Mixing those clocks is a common bug.
Operational checks that matter
The useful metrics are about age and ownership, not only total jobs:
- oldest pending job age;
- number of processing jobs with an expired lease;
- attempts per job, including the retry distribution;
- provider latency and ambiguous timeout count;
- sent, failed, and expired-token rates.
Add a bounded cleanup query for old terminal rows. Keep enough history for support and incident review, but do not retain email addresses and tokens forever. The retention decision should match the privacy requirements of the product, even when the test data comes from a tool called tempail.
An alert on expired leases is especially valuable. It tells the team that workers are crashing, running too slowly, or using a lease that is too short. A dashboard showing only queue depth can miss this failure because the rows are still present.
Questions engineers usually ask
Should this be a separate queue service?
Not necessarily. PostgreSQL is a reasonable first queue when the workload is moderate and the job already belongs to a transaction. A dedicated broker becomes attractive when throughput, fan-out, or delivery isolation demands it. The lease contract remains useful either way.
How many attempts should be allowed?
Choose a limit from the provider’s failure behavior and the user experience. Persist the count, stop retrying after the limit, and surface a support-safe reason. Unlimited attempts turn a provider outage into a database workload problem.
Does a lease guarantee exactly-once email delivery?
No. It gives at-most-one active owner in the database, not exactly-once delivery across an external provider. Stable idempotency keys, reconciliation, and clear ambiguous-outcome handling are still required.
Final checklist
Before shipping this pattern, verify that the service has:
- a conditional claim with
SKIP LOCKED; - a random lease token and an expiry timestamp;
- conditional completion using the same token;
- bounded backoff and a maximum attempt count;
- separate token expiry and job lease clocks;
- metrics for stale leases and ambiguous provider results.
That contract makes the email path a maintainable backend component. When a worker fails, the system has a defined next action instead of a hidden row and a guess.
Top comments (0)