A tool call can succeed even when an AI agent never receives the response. Imagine an agent that submits a refund request to a payment service. The service processes the refund, but the network connection times out before the agent receives confirmation. The agent sees an error and retries. If the second request creates another refund, a transient communication failure has become a duplicate business operation.
The problem is not necessarily the model's reasoning. It is the execution boundary between the agent and the systems it can change. Retries help, but they don't replace idempotency. Before allowing an agent to retry a state-changing tool call, the application needs a reliable way to recognize that the requested operation has already been accepted or completed.
The difference between a failed request and a failed operation
Distributed systems make it difficult to distinguish an operation that failed from one that succeeded without delivering its result.
Consider this sequence:
- The agent requests an action through a tool gateway.
- The gateway sends the request to a downstream service.
- The service commits the change.
- The response is lost or arrives after the client deadline.
- The agent receives a timeout and considers retrying.
At step four, the caller does not know whether the side effect happened. Treating every timeout as proof of failure is unsafe.
The same ambiguity appears when a worker crashes after committing a database transaction but before acknowledging a queue message. A message may be delivered again even though its original processing succeeded.
This is why retry policies and operation semantics must be designed together. AWS's Well-Architected guidance recommends making mutating operations idempotent so repeated requests can have the same effect as one request. The important distinction is that idempotency protects the effect of repetition; it does not guarantee that a network request is executed only once.
Give each logical action a stable operation key
An agent may call the same tool several times for different reasons. The application therefore needs to distinguish a retry of one logical action from a genuinely new action. For example, consider a workflow that creates a support ticket. A useful operation key could be derived from a durable workflow-run identifier and a stable step identifier:
workflow-8472:create-ticket
This is an illustrative key, not a prescribed format. Production systems should use identifiers with appropriate uniqueness, scope and tenant isolation.
The key must remain unchanged when the same logical operation is retried. Generating a new random key for every attempt defeats the purpose: the downstream service sees each attempt as a different request. Conversely, two intentional ticket creations must not share a key merely because their inputs look similar.
A robust operation record should associate the key with:
- The authenticated tenant or principal.
- The tool and operation type.
- A hash of the normalized request parameters.
- The current execution state.
- The downstream operation or resource identifier, when available.
- The final result or enough information to retrieve it.
The parameter hash matters because an operation key must not silently authorize a different action. If the same key arrives with different parameters, reject the conflict instead of reusing the previous result. Stripe documents a similar pattern for idempotent requests: the same key identifies retries, while mismatched parameters are rejected. Its exact retention and response-caching behavior is specific to Stripe and should not be assumed to apply to every tool gateway.
Model execution as a state machine
A single boolean such as completed is not enough to describe a tool operation safely. The system must distinguish work that has not started, work that is in progress, completed work and uncertain outcomes.
One possible state model is:
- PENDING: the operation has been recorded but execution has not started.
- RUNNING: a worker owns an execution attempt.
- SUCCEEDED: the side effect is confirmed and its result is recorded.
- FAILED_FINAL: the operation failed in a way that should not be retried automatically.
- UNKNOWN: the outcome cannot yet be determined.
The exact states depend on the downstream system. In particular, UNKNOWN is important when an external service may have committed an action but the application cannot confirm it.
A simplified flow looks like this:
Receive tool request
|
v
Validate identity, policy and parameters
|
v
Look up operation key
|
+-- Completed --> Return recorded result
|
+-- Running ---> Wait, poll or report in progress
|
+-- New --------> Record operation and execute
|
v
Reconcile the outcome
|
+---------+---------+
| |
Confirmed Uncertain
| |
v v
Succeeded Unknown
The state transition and execution claim must be concurrency-safe. Two workers should not both observe a missing record and independently execute the same action. A database uniqueness constraint, conditional write or equivalent atomic coordination mechanism can help establish a single owner for a given operation key.
However, reserving a key in a database does not automatically make an external side effect atomic with that reservation. If the worker crashes after the external action succeeds but before the local record is updated, the application can still end up with an uncertain result. That gap needs an explicit recovery strategy.
Do not blindly retry an unknown outcome
A retry policy should consider both the error and the side effect. A connection failure before the request is sent may be safe to retry. A validation error usually needs a corrected request. A rate-limit response may be retryable after an appropriate delay. A timeout after a potentially successful write is different: the application may need to query the downstream system or reconcile the operation before trying again.
For a state-changing tool, a practical policy is:
- Reuse the same operation key for the same logical action.
- Apply bounded retries with backoff and jitter for eligible transient failures.
- Respect the downstream service's retry guidance and rate limits.
- Check operation status when the outcome is uncertain.
- Stop automatic retries when the system cannot establish that repeating the action is safe.
- Escalate when a duplicate action could have significant consequences.
Do not turn every error into another model decision. The model may help interpret a failure, but deterministic application code should enforce whether a retry is permitted. This separation is especially important for actions such as issuing refunds, changing account permissions, sending external messages or provisioning infrastructure.
The transaction boundary matters
Idempotency is easiest when the operation and its result can be committed within one transactional boundary. For example, a service that creates a ticket and stores its operation record in the same database transaction can often make duplicate requests return the original ticket. A uniqueness constraint on the operation key prevents concurrent requests from creating separate records for the same logical action.
External side effects are harder. A database transaction cannot normally roll back an email already sent or a refund already processed by another service. Where an action spans systems, use a pattern appropriate to the integration:
Transactional outbox. When a local database change must produce a message, commit the change and an outbox record in one transaction. A publisher can deliver the record later. Consumers still need duplicate-safe processing because delivery may occur more than once.
Downstream idempotency. If the external API supports idempotency keys, pass a stable key scoped to the business operation. Confirm the API's key-retention rules and behavior for concurrent or mismatched requests.
Reconciliation. If the external system does not support idempotency, use a durable operation identifier or queryable business reference where available. After an ambiguous timeout, check whether the operation already exists before attempting another mutation.
Compensation. For a multi-step workflow, define how to compensate for completed steps when later steps fail. Compensation is a new business action, not a true rollback of history. A refund, for example, may reverse a financial effect without erasing the original transaction.
None of these patterns guarantees exactly-once execution across arbitrary external systems. They make duplicate effects less likely, detectable or recoverable under defined assumptions.
Agent state must survive restarts
An agent's conversation history is not a reliable execution ledger. If the process restarts, a conversation is truncated, or a worker is reassigned, the system still needs to know whether an action was attempted, confirmed, rejected or left in an uncertain state.
Persist operation state outside transient model context. On resumption, the agent should query the application for the status of its existing operation rather than infer success or failure from its last message.
Keep the execution record separate from the agent's plan:
- The plan describes what the agent intends to do.
- The operation record describes what the system has accepted and what is known to have happened.
- The authorization policy determines whether the action is permitted.
- The audit record captures the identity, decision and outcome needed for investigation.
This separation also prevents a retry from becoming an authorization bypass. A previously approved operation should not automatically authorize a materially different request, a different tenant or an action whose permissions have since changed.
Observe the operation, not just the model call
A successful model response does not mean the tool action succeeded. Likewise, a model timeout does not tell you whether a downstream side effect occurred. Instrument the complete execution path with a trace or equivalent correlation mechanism linking the workflow run, logical operation, retry attempts and downstream request identifiers.
Useful operational signals include:
- Counts of operations by final state.
- Retry attempts and retry exhaustion.
- Time spent in RUNNING or UNKNOWN.
- Duplicate-key conflicts and parameter mismatches.
- Reconciliation outcomes.
- Downstream latency, rate limits, and errors.
- Actions routed to human review.
Avoid placing secrets, unrestricted prompts, or sensitive tool arguments in telemetry. Use stable identifiers and carefully selected attributes, with access controls and retention appropriate to the data. The most useful alert is often not simply “tool call failed.” It is “state-changing operations have remained unresolved beyond the expected recovery window.” That tells an operator where intervention may be required.
Test the failure between commit and acknowledgement
Happy-path tests prove very little about retry safety. Test the ambiguous boundary deliberately.
A useful test suite should include:
- Two concurrent requests using the same operation key.
- A retry after the downstream action succeeds but the response is lost.
- A worker crash after the side effect but before the result is recorded.
- The same key submitted with different parameters.
- A downstream timeout followed by successful reconciliation.
- An expired or unavailable idempotency record.
- A restart that resumes an operation in an unknown state.
- A repeated request from a different tenant or unauthorized principal.
For each test, verify the business outcome as well as the HTTP response. Returning the same error twice does not prove that a duplicate side effect was avoided. Also test the recovery path. A system that correctly refuses an unsafe retry but leaves every operation permanently stuck in UNKNOWN is safe in one narrow sense, but it is not operationally complete.
Make retries a consequence of operation semantics
AI agents introduce another decision-making layer, but they do not change the underlying distributed-systems problem. The application still needs durable state, stable operation identity, concurrency control, authorization and a way to reconcile uncertain outcomes.
Start by classifying tools according to their effects. Read-only operations, naturally idempotent writes, and irreversible external actions should not share an indiscriminate retry policy. Then make the retry decision in deterministic code, preserve operation identity across attempts, and define what happens when the result cannot be established.
The key design rule is simple: retry the same logical operation, not a new request that merely looks similar. If the system cannot establish whether the original action took effect, reconcile or escalate before repeating it.
Top comments (1)
The point that the model should not decide whether a retry is allowed is a useful one. Keeping that in deterministic code, with the operation key tied to a hash of the parameters, seems like the safest split. I also like that you call out the UNKNOWN state, since a timeout after a write is where most duplicate refunds come from.