This is part of my build-in-public series: I am shipping a small store of n8n workflow templates (AI Automation Lab) and a free open-source lab on GitHub. Previous parts: The Crash Test Nobody Runs and No Idempotency Key? How I Make Gmail Sends Safe to Retry Anyway.
Part 1 of this series was about where crashes bite. Part 2 was about giving yourself a reconciliation handle when the provider offers no idempotency key. This part is about the component that does the actual reconciling: the sweeper.
A sweeper is a scheduled job that scans for work left in an ambiguous state and decides its fate. It exists because crashes don't clean up after themselves. Somewhere in your queue there is a row that says claimed, and nobody knows whether the email went out, the charge happened, or the webhook fired. The sweeper's entire job is to turn "unknown" into a decision — without guessing.
The lease is the whole trick
The sweeper only works if "claimed" has an expiry. Every job row carries a lease:
CREATE TABLE jobs (
id TEXT PRIMARY KEY,
state TEXT NOT NULL,
payload JSONB NOT NULL,
provider TEXT NOT NULL,
provider_ref TEXT,
attempts INT DEFAULT 0,
claimed_at TIMESTAMPTZ,
lease_seconds INT DEFAULT 600
);
Workers claim atomically — only the worker that gets the row proceeds:
UPDATE jobs
SET state = 'claimed', claimed_at = now(), attempts = attempts + 1
WHERE id = $1 AND state = 'queued'
RETURNING *;
And the sweeper touches only expired leases:
SELECT * FROM jobs
WHERE state = 'claimed'
AND claimed_at < now() - (lease_seconds || ' seconds')::interval;
The lease is the line between "the worker is slow" and "the worker is dead." Ten minutes is my default.
The sweeper's decision table
| Provider class | Handle available | Action |
|---|---|---|
| Forced-unique (Stripe Idempotency-Key, natural-key upsert) | idempotency key or provider reference | Re-drive with the same key, or query the provider by reference. Safe by construction. Ends done. |
Suggestion-key (SMTP, plain webhooks, Gmail raw messages.send) |
reconciliation handle (deterministic Message-ID, searchable log) | Look it up. Found → done. Still absent after the settle window → review. Never auto-resend.
|
| Suggestion-key | no handle at all |
review. Always. |
Note the asymmetry: for forced-unique providers the sweeper recovers; for suggestion-key providers the sweeper classifies. Recovery and classification are different verbs.
Five ways to screw up a sweeper
1. Re-running the workflow as the recovery action. The sweeper must reconcile state, not re-run logic.
2. Letting the sweeper crash unsafely. Every recovery action must be idempotent, and the sweeper's own claim on a row must be a lease too.
3. Sweeping without single-flight. Two sweeper instances racing over the same expired leases will double-recover. One row in a sweep_locks table, one atomic claim, one holder.
4. Treating "not found" as "not sent." Provider search has indexing delay. The sweeper needs a settle window: mark the row sweeping, wait, re-scan, and only then decide.
5. Counting in n8n. The sweeper's verdict must come from provider-side evidence. n8n execution history tells you which workflows ran, not which side effects landed.
The kill matrix
-
Kill after the effect, before mark-done. Provider must show exactly one effect; row ends
done. No duplicate. - Kill mid-sweep. Sweeper dies halfway through recovery. Re-run it. Still exactly one effect, no double-recovery.
- Kill before the effect. After the settle window, the sweeper re-drives or reconciles. Exactly one effect.
-
Relay forward-then-drop. The ambiguous class: provider accepted, response never observed. Sweeper must land the row in
reviewand never auto-resend.
Assert all four against provider-side counts, and report mock results as mock results.
What the sweeper doesn't give you
It does not give you exactly-once. What the sweeper gives you is smaller and more honest: a bounded recovery path, a review state that means "a human looked at this" instead of "this silently duplicated," and a kill matrix you can re-run after every change.
I'm building AI Automation Lab — n8n workflow templates for AI automation — and keeping a free lab on GitHub. I also send occasional notes on n8n reliability patterns: get them here.
Top comments (0)