This is part of my build-in-public series: I'm shipping a small store of n8n workflow templates (AI Automation Lab) and a free open-source lab on GitHub. This post grew out of a real thread on the n8n forum — a reader asked what happens when every AI provider fails mid-send, and my answer kept getting longer. So here's the full version.
Here's a sentence that should scare anyone building AI workflows that send email:
Gmail's messages.send has no idempotency-key parameter.
That means the most common "customer notification" path in n8n — generate a draft with an LLM, send it via Gmail — has no built-in protection against the exact failure that matters: you send, the response never comes back, and now you don't know if the customer got the email.
When I first hit this, my instincts were all wrong. Local dedup table: "we already tried job 123, skip it." Gmail thread ID: "search the thread, if it's there we sent it." Neither of those is a server-side guarantee. They were warm blankets. This post is what I do instead — in real workflows, not theory.
1. Two buckets, opposite recovery actions
Every failed send falls into one of two buckets, and the recovery action is opposite, so the classification is the whole game:
Bucket 1 — "definitely not accepted." 4xx validation errors, 429 rate-limit rejections. The provider told you no. The message did not go out. Safe to back off and retry. These are the easy ones.
Bucket 2 — "maybe accepted." Timeout. Connection reset. Socket closed. Your process died after writing to the socket but before reading the response. You genuinely cannot know whether Gmail committed the send.
Bucket 2 is the one that ends careers in billing systems and ends marriages in customer-support inboxes. The rule I follow now: Bucket 2 never auto-retries. It lands in a review queue with everything a human needs to decide.
2. The Message-ID trick: reconcile against the mailbox
Since Gmail gives me no idempotency key, I manufacture the next best thing: a deterministic Message-ID that I control.
When I build the raw MIME for the send, I set the header myself:
Message-ID: <job-9f3a2c1b@yourdomain>
where job-9f3a2c1b is derived from the job's dedup key — the same stable key from the source event, never $execution.id, never a fresh uuid. Same job, same Message-ID, every retry. Forever.
Now the sweeper doesn't blindly resend. It asks the mailbox:
GET gmail/v1/users/me/messages/list?q=rfc822msgid:<job-9f3a2c1b@yourdomain>
Gmail stores the Message-ID you gave it, so the mailbox becomes your acceptance log.
3. The three caveats I'd be lying to omit
Caveat 1 — the search index lies for a minute. Gmail's messages.list search has indexing delay. So the sweeper's rule is: not-found → wait 1–2 minutes → sweep again.
Caveat 2 — "found in mailbox" is not a uniqueness proof. If you want to assert "no duplicate was created," you have to look for two hits with the same Message-ID. Count, don't just check existence.
Caveat 3 — the effect log lives outside n8n. The sweeper asserts: (a) retrying the same key produces exactly 1 effect, (b) drop-response produces no duplicate, (c) an ambiguous job lands in review and is never auto-retried. The counting happens in the mock, never in n8n.
4. Choose your asymmetry per channel
- Customer email: a duplicate is mildly annoying; a missed email is worse. I accept at-least-once with a dedup hint.
- Money, inventory, irreversible actions: Bucket 2 → forced review, always. No exceptions.
5. The shape of the job row (Postgres)
CREATE TABLE ai_jobs (
id TEXT PRIMARY KEY,
message_id TEXT NOT NULL,
state TEXT NOT NULL,
payload JSONB,
attempts INT DEFAULT 0,
claimed_at TIMESTAMPTZ,
provider_ref TEXT
);
Claiming is atomic, the Message-ID is derived from id, and the sweeper reconciles against the mailbox before anything is re-sent.
That's the whole pattern. If you're building n8n workflows that touch the outside world, the free lab on GitHub has the workflow templates this series is built on, and the store has the paid ones.
Top comments (0)