An agent calls send_email. The mail API is slow, the HTTP client gives up at 30 seconds, and the model reads "timeout" as "not sent." It rewrites the subject line and sends again. The customer gets two emails.
The obvious fix is to dedupe on the tool's arguments. It fails here, because the model changed the arguments.
The fix that works: assume every tool call will run twice, and decide in advance what the second run does. This post covers where agent retries come from, the key that makes the second run harmless, and three tests that prove it. It does not cover exactly-once delivery in message queues.
I first hit this in GroundedDocs, my RAG system, and not in the agent. It was the indexer. A job that failed halfway and ran again stored every chunk twice, and the same passage came back twice in search results, pushing other evidence out.
What does a duplicated tool call look like?
This is an illustrative trace I wrote to show the pattern. It is not output from a real system.
step 4 send_email(subject="Your refund is approved") -> timeout after 30s
step 5 send_email(subject="Refund approved, order 8812") -> ok, msg_1042
mail provider, sent log:
msg_1041 "Your refund is approved" sent at step 4
msg_1042 "Refund approved, order 8812" sent at step 5
Step 4 timed out, but the email went out. The model saw a failure, reworded the subject, and tried again. Nothing in the agent's own trace says the customer got two emails.
The root cause is always the same. The action happened, but the system did not record that it happened. So some layer does the safe-looking thing and tries again. Temporal's docs describe the cleanest version: a worker finishes its task, then crashes just before it reports completion. The task is retried, and without protection you get "duplicate charges in a payment processing scenario" (Temporal).
The property that stops this has a name. An operation is idempotent when doing it twice has the same effect as doing it once. Backend engineers have relied on it for years in payments and queues. Agents bring the problem back in a harder form.
Where do agent retries come from?
The wrong model is "one tool call in my trace means one side effect." In an agent stack, four layers can retry the same action, and most of them do it without telling you.
1. The HTTP client. The library retries timeouts and server errors quietly. One real default: the Anthropic Python SDK retries 2 times on connection errors, 408, 409, 429, and 5xx, and retries timeouts too. Your tool's HTTP client, gateway, or service mesh may do the same.
What you see: one tool call in your trace and two requests in the downstream logs.
2. The orchestrator. Your agent framework retries a failed node or tool call.
What you see: the same step logged twice with the same arguments.
3. The model. The tool returns an error, and the model calls it again. This is by design. The Model Context Protocol returns tool errors to the model so it can "self-correct and retry with adjusted parameters."
What you see: two tool calls with slightly different arguments.
4. A resumed run. The process crashes or pauses for approval, and the run restarts from the last checkpoint. LangGraph does not re-run nodes that finished, but nodes after the checkpoint "re-execute," including any LLM calls, API requests, or interrupts. " A node that did its work and crashed before its checkpoint was saved runs again.
What you see: a duplicate that shows up only after a deploy, a crash, or an approval pause.
Agents make this harder than a normal backend in three ways:
- A normal client retries the same request. A model retries with reworded arguments, so the second call does not look like a duplicate.
- A normal client stops on an error. A model reads the error and tries another route.
- Agent runs are long and get resumed, and a resume re-runs steps.
And the layers multiply. A client that tries 3 times, inside an orchestrator that tries 3 times, driven by a model that tries twice, can run one action up to 18 times for a single user request. A resume can repeat all of that.
What makes the second run safe?
Seven rules. They are in the order I apply them.
1. Sort every tool into three groups. Read-only tools are safe to repeat. Some writes are naturally repeatable: "set status to closed" gives the same result twice. The dangerous group is tools that send, create, charge, or append. Only that group needs the rules below.
2. Let one layer own retries. For dangerous writes, turn retries off in the other layers. Otherwise the counts multiply.
3. Give every dangerous write an idempotency key. The orchestrator makes the key. The model must not. Build it from what the action is about: run ID, tool name, and a business ID such as an order number. Do not build it from a hash of the model's arguments, because the model rewords them. Keep personal data like the customer's email out of the key; Stripe advises the same.
4. Record the key and the result. Store the key, a status, and the result in a table with a unique constraint, or in Redis with set-if-absent. Where you can, write that record in the same transaction as the action. Stripe does this: it saves the status code and body of the first request for a key, and a repeat with that key gets the same response.
5. A repeat returns the first result, as a success. AWS calls this a "semantically equivalent response" (AWS Builders' Library). Never send "already exists" back to the model as an error. It will read that as a failure and try another route. If the same key arrives with different parameters, reject it. Stripe and AWS both do.
6. After a timeout, look before you retry. Ask the downstream service whether that key completed. A timeout tells you nothing about whether the action ran.
7. Put a gate on what cannot be undone. Payments, sends, and deletes get a human approval or a hard limit per run. Keep keys longer than your longest possible resume. Stripe keeps them for at least 24 hours.
What does this look like in code?
A sketch of the idea in Python, not production code:
def run_write_tool(run_id, tool, args):
key = f"{run_id}:{tool.name}:{tool.business_id(args)}"
row = store.get(key)
if row and row.status == "done":
return row.result # repeat: same answer as the first time
store.insert_if_absent(key, status="started")
result = tool.call(args, idempotency_key=key) # pass the key downstream too
store.update(key, status="done", result=result)
return result
The line that matters is the first one. The key comes from business_id(args), the order number, and not from the subject line the model wrote. That is why the reworded retry in the trace above would stop at the check.
In GroundedDocs this became three decisions, from the cheapest to the most work.
Ingestion is repeatable. Each document is identified by a hash of its content, and each chunk has a stable ID. Running the indexer again on documents it has already seen writes zero new chunks.
The agent's tool is read-only on purpose. The agent path has one tool, a document lookup. It changes nothing, so all four kinds of retry are harmless for it. The cheapest idempotency fix is to not have a dangerous write at all.
The writes around the agent get a key. Every request still writes two things: an audit event and a usage count against the rate limit. A retried step should not be logged twice or counted twice, so each carries a key made from the request ID and the step.
When does an idempotency key not save you?
The crash lands between "started" and "done." The wrapper records "started," the email goes out, and the process dies before "done" is written. On resume, the store cannot tell you whether the email was sent. Only two things close that gap: the downstream service honors the key you passed it, or you ask it (rule 6). If it supports neither, treat the tool as irreversible and gate it (rule 7).
The key expires before the resume. A run waits three days for approval, but keys live for 24 hours. The resumed call looks brand new. Stripe says it plainly: a key reused after it was pruned becomes a new request.
The key is too coarse. run_id:send_email:order_8812 blocks a second email about that order, including the legitimate "refund paid" email that should follow "refund approved." Put the intent in the key, for example the email template name, so two real actions get two keys.
How do you test your agent for duplicates?
Do this on one dangerous write in the next 20 minutes:
- List every tool and mark it read-only, repeatable write, or dangerous write.
- For one dangerous write, write down every layer that can retry it and how many times. Multiply.
- Crash after success: kill the process right after the tool returns, then resume the run.
- Lost reply and fake error: make the tool succeed but return a timeout, then make it succeed but return an error to the model.
- Count the side effects after each test. The target is exactly one.
The rule I use now: assume every tool call will run twice, and decide in advance what the second run does.
Which layer in your stack retried a write you did not know it could retry, and what did it duplicate?


Top comments (2)
Rule 6 gets harder when the agent drives an app's UI instead of an API. (I'm on the Auten team; we build an MCP server that lets agents click and type on a real screen.) When the agent clicks Send in a mail client, there's no key to pass downstream and no endpoint to ask "did this complete". The only check left is the screen itself: open Sent and look before you retry.
Your reworded-subject trace breaks that check too, though. If the second attempt searches Sent for the new subject, it finds nothing and sends again. I'd match on things the model doesn't rewrite, like recipient plus a time window, and pause for a human when the answer is unclear rather than guess.
Where would you put a UI-only send? Straight into rule 7 every time, or is a read-back of the sent log enough when it gives a clean yes or no?
The point about building the idempotency key from a business ID like an order number instead of a hash of the model's arguments makes the rewording problem disappear. If the orchestrator owns retries and the downstream service is the thing that actually honors the key, the model layer does not need to know idempotency is in play at all. Does passing the same key through an approval-pause resume risk the downstream treating it as a fresh request if the pause is long enough to cross a key TTL?