Your agent chose the right tool. The arguments were valid. The action was allowed.
Then the worker died.
On recovery, should it call the tool again?
A small crash test produced 110 side effects from 100 intended actions. Moving the local completion marker ahead of the call removed the duplicates – but a different crash left 10 actions missing.
All four implementations passed the happy-path control.
This is a reproducible tool-execution lab, with real worker termination and a simulated external provider. No LLM, real emails, payments, or customer systems are involved. That isolation is deliberate: even a perfect model cannot recover information that the execution layer never recorded.
The result that a successful demo hides
The lab runs 100 distinct logical operations per scenario. Every tenth operation gets one injected worker crash, followed by one retry in a fresh process.
There are two crash locations:
- Before the effect: the provider receives the request, but the worker is killed before the provider applies it. For the “record before” strategy, the local marker already exists.
- After the effect: the provider applies the action, then kills the worker before returning a receipt.
Here are the observed results. Each cell shows effects / missing operations:
| Strategy | No crash | Crash before effect | Crash after effect |
|---|---|---|---|
| Retry without deduplication | 100 / 0 | 100 / 0 | 110 / 0 |
| Record completion after the call | 100 / 0 | 100 / 0 | 110 / 0 |
| Record completion before the call | 100 / 0 | 90 / 10 | 100 / 0 |
| Stable key, deduplicated by provider | 100 / 0 | 100 / 0 | 100 / 0 |
These are constructed failures, not estimates of how frequently production systems fail. The 10 crashes are part of the test input.
The useful finding is the trade-off: a duplicate counter can show zero while your system silently loses work.
The two writes you cannot wish into one transaction
Consider this tool handler:
const receipt = await provider.execute(operation);
await operations.markCompleted(operation.id, receipt.id);
There are two separate changes: one at the provider, one in your application.
If the first succeeds and the worker dies before the second, your database still says the work is unfinished. Repeating it may create another external effect.
Swapping the lines changes the failure:
await operations.markCompleted(operation.id);
await provider.execute(operation);
Now a crash can leave a completed record for an action that never happened. A recovering worker sees the marker and skips the operation.
In the lab, that worker exits successfully. The external ledger is still short by ten actions.
The marker has confused “we intended to do it” with “we have evidence that it happened.”
An in_progress record is useful. Treating it as a completion receipt is the mistake. A lease can coordinate workers, but it does not establish whether a remote action happened before the lease holder disappeared.
What the fourth strategy actually changes
The last strategy gives each logical operation a stable identity and sends that identity to a provider that enforces deduplication.
The worker may attempt operation 40 twice. The provider recognizes the second attempt and returns the existing receipt instead of applying a second effect.
The important boundary is the one that owns the effect. Putting a cache around the caller leaves the gap between the remote write and the local record.
Amazon’s explanation of idempotent APIs describes both client request identifiers and the need to make token recording atomic with the service’s mutations. Our simulated provider performs its lookup and append in one synchronous parent-process callback. That is enough for this sequential experiment; a real service needs an appropriate durable, concurrent implementation.
A production request might carry this shape:
// Illustrative contract, not a production implementation.
await provider.execute(persistedOperation.payload, {
idempotencyKey: persistedOperation.id,
});
“Persisted” matters twice:
- The operation identity survives retries. A UUID generated once and stored is fine. A fresh UUID inside each attempt defeats deduplication.
- The intended payload stays bound to that identity. A changed recipient, amount, or message is not automatically the same operation.
Do not use the payload alone as the identity. Two intentionally separate actions can have identical payloads. Conversely, reusing one operation ID for changed parameters should produce a conflict, not quietly execute a different action.
Stripe documents this distinction concretely: repeated keys replay stored results, mismatched parameters are rejected, and keys can be removed after they are at least 24 hours old. Its documentation also notes that stored results can include errors. Provider-specific rules and retention windows belong in your retry design.
A key is only useful when the receiver honors it.
Why this belongs in an agent discussion
The distributed-systems problem is older than agents. DEV already has articles about duplicate agent actions and job replay after a restart. This lab makes two crash boundaries directly comparable and includes lost work in the scorecard.
Agents add another place where an operation can be recreated. Your queue may retry a job while the planner independently decides to “try sending again.” If those paths mint different operation IDs, provider deduplication will treat them as different actions.
The application therefore needs a durable identity for the intended action across planner turns, tool calls, and worker attempts. Do not ask the model to invent that identity anew during recovery.
Checkpointing still helps. It just has a boundary: LangGraph’s Functional API documentation explicitly notes that a task that started but did not finish may run again. Completed task results can be restored; an unfinished external write still needs safe re-execution.
Human approval has a different job, too. Approval establishes that an action is allowed. It does not prevent an already-approved action from happening twice. Bind approval to the intended payload, then handle execution and recovery separately.
What if the provider has no idempotency contract?
First, check whether there is an authoritative way to retrieve the outcome using a client reference or provider operation ID.
After an ambiguous failure, a useful set of outcomes is:
| Evidence | Next step |
|---|---|
| The provider confirms the effect | Store its receipt and complete locally |
| The provider authoritatively confirms no effect and no in-flight request | Retry under the applicable contract |
| The result is still uncertain | Keep it unresolved; reconcile or escalate |
An empty search result is not necessarily proof that nothing happened. Indexing may lag, or the original request may still complete.
If neither deduplication nor reliable reconciliation exists, you have a product decision: accept the possibility of duplicates, accept the possibility of omission, or hold the operation for investigation. Make that trade-off visible. Do not label uncertainty as success to clear the queue.
An outbox can make the local intent and local business change atomic. Its dispatcher still needs a safe contract with the external destination. Moving the retry into another process does not close the remote commit gap.
Reproduce it
Save the complete script below as crash-lab.mjs, then run:
node crash-lab.mjs > results.json
The recorded run used Node.js 24.19.0 on Linux. It needs no packages or API keys. Run it in a local development environment: it creates temporary marker files and deliberately terminates only its own child workers with SIGKILL.
The parent process represents the provider and survives worker crashes. It records effects independently of the worker’s local markers. The script compares actual provider records against the 100 intended operation identities and asserts the expected outcomes.
The provider is an in-memory simulation. This experiment does not test provider or host crashes, concurrent requests, expired keys, changed payloads, network behavior, authorization, or LLM planning quality. Passing these cases is evidence about these cases, not a production-readiness certificate.
Full runnable crash lab
// Isolated educational experiment. No network, credentials, or real messages.
// Run: node crash-lab.mjs > results.json
import { fork } from 'node:child_process';
import { existsSync, writeFileSync, mkdtempSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { fileURLToPath } from 'node:url';
import assert from 'node:assert/strict';
const self = fileURLToPath(import.meta.url);
const strategies = ['retry-only', 'record-after', 'record-before', 'provider-key'];
if (process.argv[2] === 'worker') {
const [strategy, receipt, operation] = process.argv.slice(3);
// A pre-written marker is deliberately treated as completion in record-before.
if (strategy !== 'retry-only' && existsSync(receipt)) {
process.exit(0);
}
if (strategy === 'record-before') writeFileSync(receipt, 'claimed');
process.on('message', (message) => {
if (message.type !== 'ack') return;
if (strategy !== 'retry-only') writeFileSync(receipt, message.providerId);
process.disconnect();
});
process.send({ type: 'execute', operation,
key: strategy === 'provider-key' ? operation : null });
} else {
const directory = mkdtempSync(join(tmpdir(), 'agent-crash-lab-'));
const rows = [];
// Parent = simulated external provider. Worker = separate killable process.
// Provider state survives worker death, but not parent/host death.
async function run(strategy, boundary, injectFaults) {
const effects = [];
const receipts = new Map();
let requests = 0;
let kills = 0;
let skippedWithoutEffect = 0;
async function attempt(operation, receipt, shouldKill) {
return new Promise((resolve, reject) => {
const child = fork(self, ['worker', strategy, receipt, operation], {
stdio: ['ignore', 'ignore', 'inherit', 'ipc'],
});
const timeout = setTimeout(() => {
child.kill('SIGKILL');
reject(new Error('Unexpected worker timeout'));
}, 5000);
child.on('error', reject);
child.on('message', (message) => {
if (message.type !== 'execute') return;
requests++;
if (shouldKill && boundary === 'before-effect') {
kills++;
child.kill('SIGKILL');
return;
}
// Atomic in this single parent event-loop callback. NOT a production DB.
let providerId = message.key && receipts.get(message.key);
if (!providerId) {
providerId = `effect-${effects.length + 1}`;
effects.push({ operation: message.operation, providerId });
if (message.key) receipts.set(message.key, providerId);
}
if (shouldKill && boundary === 'after-effect') {
kills++;
child.kill('SIGKILL'); // effect committed; worker gets no receipt
return;
}
child.send({ type: 'ack', providerId });
});
child.on('exit', (code, signal) => {
clearTimeout(timeout);
if (signal === 'SIGKILL' && shouldKill) resolve(false);
else if (code === 0) resolve(true);
else reject(new Error(`Unexpected worker exit: ${code}/${signal}`));
});
});
}
for (let index = 1; index <= 100; index++) {
const operation = `tenant-demo:notification:${index}`;
const receipt = join(directory, `${strategy}-${boundary}-${injectFaults}-${index}`);
const ok = await attempt(operation, receipt, injectFaults && index % 10 === 0);
if (!ok) assert.equal(await attempt(operation, receipt, false), true);
if (!effects.some((effect) => effect.operation === operation)) skippedWithoutEffect++;
}
const unique = new Set(effects.map((effect) => effect.operation)).size;
return { strategy, boundary, injected: injectFaults, intended: 100,
requests, kills, effects: effects.length, duplicateEffects: effects.length - unique,
missingOperations: 100 - unique, skippedWithoutEffect };
}
try {
// Happy-path controls: every implementation looks correct without a crash.
for (const strategy of strategies) {
const row = await run(strategy, 'none', false);
assert.equal(row.effects, 100);
assert.equal(row.missingOperations, 0);
rows.push(row);
}
for (const boundary of ['before-effect', 'after-effect']) {
for (const strategy of strategies) {
const row = await run(strategy, boundary, true);
assert.equal(row.kills, 10);
const missing = boundary === 'before-effect' && strategy === 'record-before' ? 10 : 0;
const duplicates = boundary === 'after-effect' && ['retry-only', 'record-after'].includes(strategy) ? 10 : 0;
assert.equal(row.missingOperations, missing);
assert.equal(row.duplicateEffects, duplicates);
assert.equal(row.effects, 100 - missing + duplicates);
rows.push(row);
}
}
console.log(JSON.stringify({ node: process.version, platform: process.platform,
experiment: '100 operations per row; kill every tenth initial attempt; one retry',
scope: 'Single worker at a time; simulated external provider; no LLM; no network',
assertions: 'passed', rows }, null, 2));
} finally {
rmSync(directory, { recursive: true, force: true });
}
}
Bring these questions to your next demo
Ask the team to show one consequential operation and answer:
- Where is its identity stored before execution starts?
- Which system can prove the external effect happened?
- What happens if the worker dies immediately before and immediately after that effect?
- Does the test assert both no duplicates and no missing work?
- What does the user see while the result is genuinely unknown?
Then run the two crash cases against a sandbox version of the real integration. Inspect the destination’s records, not just your worker’s logs.
Which tool in your system would be hardest to reconcile after a lost response – and what evidence would you trust?
Further reading: my software architecture review framework covers contracts, failure behavior, and operability beyond this one test.
Top comments (2)
Your table has one case I'd add a fifth row for: providers with no idempotency key at all. There "record before" loses 10 and "record after" duplicates 10, and a stable key isn't on offer. The usual way out is to put your own operation id into something the provider stores and lets you query (an email Message-ID or custom header, a metadata field, a reference string on a payment), mark the local row
in_progress, and on recovery look the id up at the provider before deciding. Found means store the receipt, not found means safe to send once more. That turns the "crash after effect" cell from 110 into 100 without the provider cooperating, at the cost of one read per recovered operation and a lookup that has to be strongly consistent, which many list endpoints are not.Also worth a variant in the lab: a timeout instead of a kill. A worker that stays alive but never sees the receipt behaves exactly like your after-effect crash from the caller's side, and it's the more common production trigger because retry middleware fires it automatically, with no crash for anyone to notice.
Preserving the evidence and exposing it through a read path are separate guarantees, and the reconciliation table in the last section quietly assumes the second one holds. Here is a case from a public append-only event ledger I parse. Fetching either of two withdrawn ids through the single-item route returns HTTP 410, while fetching an id that never existed returns 404. The listing route leaves both withdrawn ids off every page. The rows are still physically present, confirmed by a database count of 2. The distinction survives in storage and enumeration destroys it.
Re-measured today, that ledger holds 103,498 signed rows. Row ids hash a value containing a timestamp rather than the body, so resending the same intent mints a fresh id and slips past id-based deduplication entirely. Counting groups where one author published a byte-identical body under the same event type, 426 groups repeat and cover 5,641 rows, and the largest single body accounts for 1,205 of them. Those counts do not measure unintended retries, for the reason you already give about payloads not establishing intent. The sharper evidence is much smaller. The entire history contains exactly two withdrawal rows, and both record the reason verbatim as 'duplicate post from an idempotency bug (script executed twice); canonical copy is the later event'.
A reconciliation worker that reads an empty listing as authoritative absence can therefore retry an effect whose single-item lookup says, at that same moment, that it happened and was withdrawn. That is a stronger form of your empty-search-result warning than indexing lag. Waiting longer recovers nothing, because the listing route cannot express the difference at any point in time. So the question to settle before a read path is trusted for reconciliation is whether its contract separates absent from withdrawn, and when it does not, every empty result from it belongs in the unresolved branch. This is one small ledger rather than a payment provider, and being append-only it does not reproduce your crash boundaries. Would the lab support a fifth case where the provider retains the evidence but only the single-item lookup reveals it, to check whether the worker holds the outcome unresolved instead of retrying?