DEV Community

Naveen Alavilli
Naveen Alavilli

Posted on Fully Autonomous

Your Agent Timed Out. Did the Action Still Happen?

An agent submits a request to publish a document. The server commits it. The response disappears before the agent receives it.

The agent sees a timeout and tries again.

You now have two published documents and a log that describes a successful recovery.

This is a design scenario, not a benchmark result. It exposes a question worth asking before giving an agent tools that change external state: what does your system do when it cannot tell whether an action happened?

A timeout establishes that the caller did not receive a timely result. It does not establish that the remote operation failed. The Amazon Builders' Library discussion of idempotent APIs explains this ambiguity and the use of caller-provided request identifiers to make retries safe.

For agent workflows, I'd make that uncertainty explicit in the application contract.

Give uncertainty a state

A boolean success field cannot describe every outcome of a remote write. Consider these states:

State Meaning Next step
Ready The operation is authorized and durably recorded Dispatch
In flight A worker has claimed it; it may have reached the destination Await evidence
Succeeded The destination has confirmed the required outcome Return the stored result
Failed There is conclusive evidence it did not take effect Apply the failure policy
Unknown It may have taken effect, but evidence is missing Reconcile

An expired worker lease also deserves attention: the worker might have crashed after the remote commit. Recovery must account for that possibility.

The UI can say: “The request was submitted, but completion is unconfirmed. Checking its status.” That is a useful result, even when it is less satisfying than “Done.”

Identify the operation before calling the model

A tool-call ID identifies an attempt. The application needs an identity for the logical operation across attempts.

For a hypothetical document workflow:

operation_id: op_8c21
actor_id: user_42
action: publish_document
target_id: doc_104
content_version: 7
request_fingerprint: <hash of canonical action parameters>
state: ready
Enter fullscreen mode Exit fullscreen mode

Create and persist that identity in application code. Carry it through retries and restarts. If the destination supports an idempotency key, pass the same key for the same operation according to that API's contract.

The request fingerprint has a different job: detect changed parameters. Reusing an operation ID with a different destination or payload should produce a conflict. A hash alone cannot express intent; two intentional operations can have identical payloads.

Scope identifiers to the actor or tenant, enforce uniqueness atomically, and check the provider's retention window. An old key may stop protecting a retry after the destination expires its record.

Keep authorization attached to the proposed action

If a workflow requires approval, record what was approved: action, target, payload version, and applicable limits.

A retry of that exact operation can reuse the recorded authorization where policy permits. A model that changes the payload has proposed a different action. It should not silently inherit approval for the old one.

Recheck current permissions at execution time too. A durable approval record should not bypass a later revocation.

A local ledger cannot close a remote transaction

The difficult sequence is:

1. Record the operation locally.
2. Send the remote write.
3. The destination commits.
4. The worker crashes before recording the result.
Enter fullscreen mode Exit fullscreen mode

A local transaction cannot make steps 2 and 4 atomic with an unrelated service.

The recovery path depends on the destination:

  • With suitable idempotency support, retry the same operation under that contract.
  • With an authoritative lookup by operation ID, query it and reconcile the result.
  • With neither, retain the unknown outcome and route it for investigation before attempting another consequential write.

An empty search result may be inconclusive if the destination is eventually consistent. A durable queue improves delivery, but the consumer still needs a strategy for duplicate attempts.

Test the inconvenient boundary

Before shipping, exercise these cases:

Injected condition Expected behavior
Response lost after remote commit Recovery resolves to the original result
Two workers dispatch the same operation The destination applies one logical effect under its idempotency contract
Same ID, different payload Conflict before another effect
Worker dies after remote commit Operation enters reconciliation
Lookup is temporarily stale No premature claim that nothing happened
Idempotency retention expires Retry policy accounts for the lost protection

These are tests for the surrounding application, not prompts asking the model to be more careful.

A practical design review question is: if this write succeeds and its response vanishes, what evidence lets the next worker decide what to do?

If the answer is only “the agent will figure it out,” the recovery protocol is still unfinished.

Top comments (3)

Collapse
 
reidmarlow profile image
Reid Marlow •

The messy failure mode is handing an 'Unknown' status back to the agent in the next turn. When an LLM receives an ambiguous timeout in its tool result, it rarely waits for reconciliation. In practice, it either retries with a slightly tweaked argument that generates a fresh operation ID, or runs a destructive cleanup assuming the write never touched the server.

Reconciliation works best as a deterministic harness intercept before the next model turn is assembled. If the harness cannot resolve whether the write committed through an authoritative status lookup or deduplication check, pausing the loop with a blocked state prevents the model from improvising around unconfirmed state.

Collapse
 
slabb profile image
Sam LABBE •

The state table is the right contract — and it maps almost one-to-one onto the finding statuses in the reconciliation layer from our flight-recorder thread: In flight → pending, Succeeded → matched, Unknown → unconfirmed once a deadline passes (open_gap if no deadline was set at all), and your "changed payload must not inherit approval" is exactly why an explicit flag in an outcome payload is its own finding — auditors search for the flag, not for silence.

The axis the table is missing is time. Without a deadline attached to the operation, "In flight" and "Unknown" are indistinguishable forever — pending has to be able to decay into unconfirmed on its own, or the stuck row waits for a human to notice it. expected_by or expected_within on the operation record makes the state machine self-flipping.

One assumption worth naming, though: every row of that table trusts the ledger to survive the crash unmodified — and the process that crashed mid-write is the process that owns the ledger. Your opening scenario is "a log that describes a successful recovery"; sealing the ledger's heads is what makes that log falsifiable instead of editable. Same instinct at two layers: the state contract bounds what the workflow can do, the sealed record bounds what can be denied about what it did.

Collapse
 
ssapable profile image
ssapable •

The "Unknown" row is the one I'd never modeled, and it bit me last week in a much dumber form. I had an agent posting on a schedule from a server. Every call returned 200 and a post ID, so the log read as a wall of successes. The posts were reaching 2 to 8 people. Technically nothing was "unknown", but the thing I actually cared about (did a human see it) was never checked, so the log was describing a success that hadn't happened either. Your table makes me think the fix is the same shape: the success state should require evidence of the outcome you wanted, not the ack from the API.