DEV Community

HaoPeng Zhang
HaoPeng Zhang

Posted on

The Crash Test Nobody Runs: What I Learned About "Exactly-Once" Building AI Workflows in n8n

This is part of my build-in-public series: I am shipping a small store of n8n workflow templates (AI Automation Lab) and a free open-source lab on GitHub.

I used to believe my n8n workflows were reliable because the execution list was green. Then a payment-notification workflow sent the same email twice, and the green checkmarks were still green.

n8n can tell you a workflow ran. It cannot tell you the side effect happened exactly once.

Green means the execution finished. It says nothing about whether the external action - the email, the charge, the row - happened zero times, once, or twice. Once you see that gap, "exactly-once" stops being a checkbox you flip and becomes a boundary you have to test.

Here is what I actually learned, mostly by testing that boundary the hard way.

Two crash shapes, and only one is reachable from inside n8n

When people say "we test failure handling", they usually mean the easy shape. There are two:

  1. Accepted and known, receipt lost. The HTTP node returns, you hold the provider response, and then the process dies before you persist "done". A Code node that throws after the HTTP call reproduces this exactly. It is deterministic, and it is the easy class - because you already observed acceptance, so you can reconcile from the provider id if you logged it anywhere durable.

  2. Submitted, response never observed. The request reached the provider and the side effect committed, but your process never saw the response - connection reset after commit, timeout after the write landed, process killed between send and receive. This is the case that makes exactly-once impossible and forces an Unknown/review state. And you cannot reproduce it reliably from inside n8n, because whether the request already left the socket when you kill the process is a race, not a control.

That second point is the whole article in one sentence: the test everyone runs only exercises class 1, so the bug that actually doubles your emails - class 2 - never gets tested.

The crash test nobody runs: put a relay between n8n and the provider

Instead of killing n8n and hoping the kill landed in the right window, make the failure a switch. Put a tiny local HTTP endpoint in front of the provider (or a mock that still records acceptance), and select the behavior with a header or env var:

  • fail-before-forward - the provider never sees it. This must be safe to retry.
  • forward-then-drop-response - the provider commits, then the relay hangs or closes without returning. This is the ambiguous case. Run your sweeper here.
  • forward-and-respond - the happy path. It must end in done, and the sweeper must not touch it.

forward-then-drop-response is cleaner than "kill n8n after send" because the relay decides the interleaving, so you hit the same boundary on every run instead of hoping the kill landed in the right millisecond.

Then, and only then, run the workflow-injected throw as a cheaper second test for class 1.

And get the assertions right, because this is where most people fool themselves:

  • Count side effects at the provider (or the relay's acceptance log), never at n8n. The n8n execution count tells you nothing about whether the send happened.
  • After the sweeper runs, class 2 must land in Unknown/review, never auto-retry. Only class 1 - acceptance observed plus a persisted provider id - may be closed from the receipt.
  • Run the proxy test twice with the same idempotency key and assert the provider ends with one effect, not two. If the key is only advisory (SMTP-style), expect two and require the review path - that is the correct outcome, not a bug.

n8n is the worker, never the queue

Once you accept that a crashed execution can't be "resumed" safely, the design changes: stop trying to resume the execution; recover the job.

n8n execution state is not a durable queue. After a restart, the original execution is gone - and anything a Wait node held is only safe where the wait is short (waits over 65 seconds offload to the DB, which tells you the waiting state survives a restart, but says nothing about whether a side effect already committed before the crash). Resuming a Wait is not evidence of external state.

So keep the deferred work in a store you control (Postgres is the default; Sheets/Airtable only if it has to be codeless) and treat n8n purely as the worker:

  1. Persist a self-contained job row - a stable dedup key derived from the source event (never $execution.id, $now, or a fresh uuid), plus state, payload, attempts, claimed_at, updated_at, last_error, provider_ref.
  2. Claim atomically, not check-then-insert. The classic bug is a lookup followed by an insert: two simultaneous webhook deliveries both read "not seen yet" and both proceed. Make it one statement:
INSERT INTO idem (idem_key, status, claimed_at)
VALUES ($1, 'pending', now())
ON CONFLICT (idem_key) DO NOTHING
RETURNING idem_key;
Enter fullscreen mode Exit fullscreen mode

The run that inserts gets a row; the loser gets zero rows - and n8n does not execute downstream nodes on an empty branch, so only the winner reaches the side-effect node. No IF node needed; the SQL is the gate.

  1. Make completion monotonic. Keep the guard so a late duplicate can't flip state backward:
UPDATE idem SET status='done', done_at=now()
WHERE idem_key=$1 AND status='pending';
Enter fullscreen mode Exit fullscreen mode
  1. Recover on a schedule, not on restart. A sweeper over state='claimed' AND claimed_at < now() - interval '10 minutes' is what actually cleans up after a crash.
  2. Keep three states, not two: pending / done / review. pending should be short-lived; a long-lived pending is exactly the thing that should page a human.

Idempotency is a property of the receiving system, not your payload

A key carried in metadata is a reconciliation handle, not an idempotency guarantee. Idempotency is a property of the receiving system - your payload only tells the receiver which logical operation it is. That splits providers into two classes, and you should design differently for each:

  1. Providers that enforce uniqueness server-side (Stripe Idempotency-Key, a DB unique constraint / ON CONFLICT, order APIs with a client-supplied reference). Here a blind retry is safe and you genuinely get exactly-once for the effect.
  2. Providers where the key is advisory (SMTP, generic webhooks, most append-style APIs). Nothing stops a second execution, so exactly-once is not achievable from the caller side. The honest target is at-least-once plus visibility.

The uncomfortable truth about class 2: the workflow can never distinguish "committed but unconfirmed" from "not committed". It can only refuse to auto-repeat. I would much rather page a human on an uncertain outbound email than retry it silently - and that is the correct trade-off, not a shortcoming.

Sweepers must classify before acting

The sweeper is where a nice design quietly turns back into the double email. It must classify before it acts, per provider class:

  • Key-enforced provider: re-issue with the same key, or query by reference. Found -> mark done. Not found -> safe to retry. This is the only real exactly-once recovery path.
  • Key-advisory provider: you cannot tell submitted-but-unconfirmed from not-submitted. Do not blind-resend. Move the row to review and alert. A human, or a second workflow, decides.

For the review record, write the marker before the side effect and carry the reconciliation handle plus the execution id, so a person can look it up provider-side instead of guessing. A silent resend here is precisely how the double email gets created in the first place.

The one test that actually finds the crash-window bug

Two tests are worth more than all the green checkmarks:

  1. Fire two identical deliveries concurrently and assert exactly one row is inserted and exactly one side effect happens. That tests the atomic claim.
  2. Inject a failure between the side effect and the mark-done (a Stop or throw node), then run the sweeper and assert it lands in review rather than a resend.

The second one is the test that actually finds the crash-window bug - and it is the one almost nobody runs.

What I'd decide before building the next workflow

Before you write a single node, answer one question for the send cycle: is the provider one that enforces uniqueness server-side, or is the key only a hint? That single fact determines whether your recovery is "resend automatically" or "page a human." Everything else - the relay, the job table, the sweeper - is implementation detail around that fork.

n8n is excellent at running the workflow. It is not, and was never meant to be, your source of truth about what happened in the outside world.


If this was useful, the free lab (workflows, MIT) is here: github.com/zhp910318/n8n-ai-automation-lab - and I keep a small set of production-tested n8n templates at AI Automation Lab.


📬 Get the free n8n workflow templates — plus new ones by email. Join the free list and I will send the free templates plus occasional automation tips. No spam, unsubscribe anytime.

Top comments (0)