Your agent's worker sends a payout request, and then the process gets killed before the response comes back. The supervisor restarts it. The agent reloads its plan, sees that the "pay invoice 4417" step never finished, and runs it again. The payment API supports idempotency keys, and the agent sent one both times. The vendor still got paid twice, because the agent generated a fresh UUID on each run.
Yesterday I wrote about gating a risky tool call with a permit and a receipt and ended by asking where an agent's key comes from. This post is my answer, and it applies whatever sits behind the tool.
Two terms
An idempotency key is a value the client sends with a request so the server can recognize a repeat and return the first result instead of acting again. It only works if the repeat carries the same key.
A logical action is the thing you want to happen once in the business sense: pay invoice 4417, delete the staging database created for run 812, send the renewal email for account 77. Attempts are how you try to make it happen. One logical action can take many attempts.
The rule that follows: the key identifies the logical action, not the attempt.
The per request UUID trap
Most retry code mints the key where the request is built. That is the bug.
This is a simplified sketch written for this post, not code from any product:
# Wrong: every attempt looks like a brand new action.
for attempt in range(3):
try:
return pay(invoice, idempotency_key=str(uuid.uuid4()))
except TimeoutError:
continue
The timeout is exactly the case where the first attempt may have succeeded, and exactly the case where this loop sends a key the server has never seen. Moving uuid4() above the loop fixes the in process retry, but not the restart in the opening story, because the key dies with the process.
Where the key should come from
Here is the flow I use, in order:
- Mint or derive the key when the action is planned, not when it is executed. The planner decides "pay invoice 4417 for $250". That decision is the moment the action gets its identity.
- Persist the key with the task state before the first attempt. If the step is in your database, queue message, or workflow history, the key goes there too, written before anything is sent.
- Every attempt reads the key, never creates it. In process retries, a restarted worker, and a resumed workflow all send the same value.
- A new key means a new intent. If the amount or recipient changes, that is a different action and deserves a different key. A transient error is not a reason for a new key.
- Treat a key conflict as a signal, not an obstacle. If the server says "this key was used with a different request", stop and look. Do not mint a fresh key to get past it.
The fixed loop, again a sketch:
# Right: the key was created with the task and stored next to it.
key = task.idempotency_key
for attempt in range(3):
try:
return pay(invoice, idempotency_key=key)
except TimeoutError:
continue
Stored or derived?
There are two reasonable ways to get task.idempotency_key.
Stored: generate a random value when the task row or message is created, and keep it there. Simple, and safe as long as that storage is durable and is the same record every retry reads.
Derived: compute the key from identifiers that are already stable for this action. Also a sketch:
import hashlib
def action_key(run_id: str, step_id: str, business_ref: str) -> str:
raw = f"payout:v1:{run_id}:{step_id}:{business_ref}"
return hashlib.sha256(raw.encode()).hexdigest()
Deriving survives even if you lose the stored key, as long as the inputs are truly stable. Choose them carefully:
- Include a workflow or run ID, the step ID, and the business reference (invoice number, account ID).
- Exclude timestamps, attempt counters, hostnames, and any text the model wrote. Each of those changes between attempts and quietly turns your retry into a new action.
- Do not go too coarse. A key of just the invoice number blocks a legitimate second, partial payment on the same invoice. Add whatever makes two real actions different.
-
Version the format (the
v1above), so changing the recipe later does not collide with old keys.
The agent specific twist
With a normal service, a retry resends the same bytes. With an LLM agent, a retry often means the model generates the tool call again, and it may not produce the same arguments. The memo field gets reworded, or the amount is parsed slightly differently from the invoice text.
If your key is derived from the task, a properly built server will see the same key with a different request and refuse it. That is the system working: it failed closed instead of paying twice. Two things make this smooth in practice:
- Fill consequential arguments from task state, not from the model. Let the model decide to pay; let code supply the amount, currency and recipient from the stored task.
- Route conflicts to a person or a planner step. The fix is to figure out which version of the action is right, not to retry harder.
How AMW handles keys
I build Agent Middleware (AMW), a gateway that puts a permit and receipt boundary in front of MCP tool calls, so here is how it treats keys today. Agents act under a permit; a same-key retry returns the original receipt, no second call or charge. Governed calls must carry a key, scoped per wallet. The same key with a different request is refused with idempotency_key_reused, and a retry that arrives while the first call is still running gets idempotency_in_progress.
What it does not do matters just as much for this topic:
- A new key is a new call. If the permit sets a per tool call limit and it is used up, the retry is refused. Do not count on anything else to catch it. Where your key comes from is your job, which is why this post exists.
- Once only at the remote tool depends on the tool. AMW controls its own side: at most one dispatch, one charge and one receipt per accepted key. Whether the side effect itself happens once at the remote tool depends on that tool de-duplicating on an idempotency key of its own.
-
A crash in the gateway can make a same-key retry wait. After a crash mid call, a same-key retry can see
idempotency_in_progressfor about three hours until the stuck call is reconciled. Your agent should treat that as "pending", not as "failed, try a new key".
Disclosure
I'm Chris, the founder and sole owner of AMW, which is pre-revenue. This post was drafted with help from an AI assistant; the Python is illustrative and is not AMW's source, and the AMW claims were checked against its code. More at thisisatest.tech.
My question for you: when an agent's plan changes partway through, say the amount gets corrected after a failed attempt, do you treat it as the same action and resolve the conflict, or as a new action with a new key? I go back and forth on where that line belongs.
Top comments (0)