The agent paid the invoice once. Then the model regenerated the tool call for a retry, kept the same idempotency key, and changed the memo field by a few words. The server saw the familiar key, returned the original receipt, and the caller moved on. In the ledger it looks like the corrected call went through. It never did.
Yesterday I wrote about where an AI agent's idempotency key should come from. That post was about minting the key for the logical action and keeping it across retries. This one is about the other half: what the server must do when that same key shows up with a different request.
The key-only matching bug
An idempotency key is not a free pass to replay whatever you stored last time. It is a pointer into a record that also captured what was asked for.
Key-only matching is the bug. The server stores (key → response), looks up by key alone on the next request, and returns the stored response. That is fine when every retry truly repeats the first call. It is a silent lie when the arguments changed.
With ordinary HTTP clients, the body usually stays the same on retry. With LLM agents, a "retry" often means the model writes the tool call again. Amounts get re-parsed. Memo text gets reworded. A recipient id shifts by one digit after a copy from a PDF. The key can still be the stable task key you carefully persisted, and the payload is no longer the same action.
If the server only checks the key, the agent gets the old receipt for a call it never made. That is worse than a clean double charge in one important way: the books say the new intent succeeded.
Builders keep rediscovering this. The usual warning is: if you only check the key, a retry with the same key but different arguments returns the old receipt and the caller thinks the new call went through. Hash the args into the key, or refuse the mismatch loudly.
Refuse the mismatch (hash the request)
The correct server behavior is simple:
- When you accept a request under a key, store a hash of the effectful request alongside the response.
- On a later request with the same key, compare hashes.
- Same key and same hash: return the original result. Skip a second side effect and a second charge.
- Same key and different hash: refuse. Prefer not to replay or overwrite. Surface a conflict the client has to handle.
Here is a sketch written for this post, not code from any product:
import hashlib
import json
def request_hash(tool: str, arguments: dict) -> str:
# Canonicalize so key order does not create false mismatches.
canonical = json.dumps(
{"tool": tool, "arguments": arguments},
sort_keys=True,
separators=(",", ":"),
)
return hashlib.sha256(canonical.encode()).hexdigest()
def begin(store, key: str, tool: str, arguments: dict):
incoming = request_hash(tool, arguments)
existing = store.get(key)
if existing is None:
# First time: reserve the key, run the call, store hash + result.
return "new", incoming
if existing["request_hash"] == incoming:
return "replay", existing["result"]
# Same key, different request: fail closed.
raise Conflict("idempotency_key_reused")
Two design notes matter in practice:
- Hash the fields that change the effect. Tool name, amount, currency, recipient, and other business fields belong in the hash. Purely cosmetic headers usually do not. If you hash nothing consequential, you are back to key-only matching in disguise.
-
Make the conflict loud. A soft 200 with a warning header will be ignored by most agent loops. Prefer a distinct error the client must handle, such as
idempotency_key_reusedor HTTP 409 with a clear body.
There is a second school of thought: bind the arguments into the key itself, so a changed payload naturally becomes a new key. That works too, as long as every caller uses the same recipe. The hash-and-refuse approach is usually safer for a shared API, because clients that already send a stable task key still get protection even when they forget to include the args in the key string.
What the client should do on conflict
When the server says the key was reused with a different request, minting a fresh key and firing again is usually the wrong move. That turns a detected conflict into a second real call.
Treat it as a signal:
- Stop the automatic retry. A conflict is not a timeout.
- Compare the two intents. Was the change accidental (model reworded a memo) or intentional (amount corrected after review)?
- If accidental, resend the original arguments under the same key, so you get a true replay.
- If intentional, that is a new logical action. Give it a new key, after you have decided the first action is cancelled, superseded, or left as-is.
- Route unclear cases to a person or a planner step. Agent frameworks that retry every non-200 will happily dig a deeper hole unless you catch this error class explicitly.
Filling consequential arguments from task state, not from free-form model text, cuts how often this fires. Let the model decide to pay; let code supply amount, currency, and recipient from the stored task.
How I am building this in AMW
I build Agent Middleware (AMW), a gateway that puts a permit and receipt boundary in front of MCP tool calls. Agents act under a permit; a same-key retry returns the original receipt, no second call or charge. The request is hashed with the call. Same key with a different payload (or a different permit) is refused as idempotency_key_reused, not replayed. A retry that arrives while the first call is still running gets idempotency_in_progress.
What it does not do:
- A new key is a new call. Nothing silently de-duplicates across keys. If you change the key to "get past" a conflict, you asked for a second dispatch.
- Remote side effects are a separate problem. AMW bounds its own side: at most one dispatch, one charge, and one receipt per accepted key. Whether the remote tool itself runs once still depends on that tool honoring an idempotency key of its own.
I'm biased: this is the product I am building, and it is pre-revenue. The pattern above is older than AMW. Stripe-style APIs have taught request binding for years. Agents just make the mismatch case common enough that leaving it soft is no longer optional.
Disclosure
I'm Chris, the founder and sole owner of AMW, which is pre-revenue. This post was drafted with help from an AI assistant; the Python is illustrative and is not AMW's source, and the AMW claims were checked against its code. More at thisisatest.tech.
My question for you: when your agent hits a same-key conflict because the model rewrote the arguments, do you force a resend of the original payload, or do you treat it as a new action with a new key? I want to hear how people decide that line in production.
Top comments (0)