DEV Community

HaoPeng Zhang
HaoPeng Zhang

Posted on

The Sweeper Pattern: How I Reconcile Crashed n8n Jobs Without Guessing

This is part of my build-in-public series: I am shipping a small store of n8n workflow templates (AI Automation Lab) and a free open-source lab on GitHub. Previous parts: The Crash Test Nobody Runs and No Idempotency Key? How I Make Gmail Sends Safe to Retry Anyway.

Part 1 of this series was about where crashes bite. Part 2 was about giving yourself a reconciliation handle when the provider offers no idempotency key. This part is about the component that does the actual reconciling: the sweeper.

A sweeper is a scheduled job that scans for work left in an ambiguous state and decides its fate. It exists because crashes don't clean up after themselves. Somewhere in your queue there is a row that says claimed, and nobody knows whether the email went out, the charge happened, or the webhook fired. The sweeper's entire job is to turn "unknown" into a decision — without guessing.

The lease is the whole trick

The sweeper only works if "claimed" has an expiry. Every job row carries a lease:

CREATE TABLE jobs (
  id              TEXT PRIMARY KEY,
  state           TEXT NOT NULL,
  payload         JSONB NOT NULL,
  provider        TEXT NOT NULL,
  provider_ref    TEXT,
  attempts        INT DEFAULT 0,
  claimed_at      TIMESTAMPTZ,
  lease_seconds   INT DEFAULT 600
);
Enter fullscreen mode Exit fullscreen mode

Workers claim atomically — only the worker that gets the row proceeds:

UPDATE jobs
SET state = 'claimed', claimed_at = now(), attempts = attempts + 1
WHERE id = $1 AND state = 'queued'
RETURNING *;
Enter fullscreen mode Exit fullscreen mode

And the sweeper touches only expired leases:

SELECT * FROM jobs
WHERE state = 'claimed'
  AND claimed_at < now() - (lease_seconds || ' seconds')::interval;
Enter fullscreen mode Exit fullscreen mode

The lease is the line between "the worker is slow" and "the worker is dead." Ten minutes is my default.

The sweeper's decision table

Provider class Handle available Action
Forced-unique (Stripe Idempotency-Key, natural-key upsert) idempotency key or provider reference Re-drive with the same key, or query the provider by reference. Safe by construction. Ends done.
Suggestion-key (SMTP, plain webhooks, Gmail raw messages.send) reconciliation handle (deterministic Message-ID, searchable log) Look it up. Found → done. Still absent after the settle window → review. Never auto-resend.
Suggestion-key no handle at all review. Always.

Note the asymmetry: for forced-unique providers the sweeper recovers; for suggestion-key providers the sweeper classifies. Recovery and classification are different verbs.

Five ways to screw up a sweeper

1. Re-running the workflow as the recovery action. The sweeper must reconcile state, not re-run logic.

2. Letting the sweeper crash unsafely. Every recovery action must be idempotent, and the sweeper's own claim on a row must be a lease too.

3. Sweeping without single-flight. Two sweeper instances racing over the same expired leases will double-recover. One row in a sweep_locks table, one atomic claim, one holder.

4. Treating "not found" as "not sent." Provider search has indexing delay. The sweeper needs a settle window: mark the row sweeping, wait, re-scan, and only then decide.

5. Counting in n8n. The sweeper's verdict must come from provider-side evidence. n8n execution history tells you which workflows ran, not which side effects landed.

The kill matrix

  1. Kill after the effect, before mark-done. Provider must show exactly one effect; row ends done. No duplicate.
  2. Kill mid-sweep. Sweeper dies halfway through recovery. Re-run it. Still exactly one effect, no double-recovery.
  3. Kill before the effect. After the settle window, the sweeper re-drives or reconciles. Exactly one effect.
  4. Relay forward-then-drop. The ambiguous class: provider accepted, response never observed. Sweeper must land the row in review and never auto-resend.

Assert all four against provider-side counts, and report mock results as mock results.

What the sweeper doesn't give you

It does not give you exactly-once. What the sweeper gives you is smaller and more honest: a bounded recovery path, a review state that means "a human looked at this" instead of "this silently duplicated," and a kill matrix you can re-run after every change.


I'm building AI Automation Lab — n8n workflow templates for AI automation — and keeping a free lab on GitHub. I also send occasional notes on n8n reliability patterns: get them here.

Top comments (0)