DEV Community

hefty
hefty

Posted on

An Agent Retry Is Not a Rewind Button

An agent asks a deployment API to create a release. The call times out. The agent restarts with its conversation intact and tries again.

There may now be two releases. Or one. The timeout tells us what the caller saw, not what the deployment service did. Replaying the conversation cannot settle it.

This is a hypothetical failure, but the design problem is concrete: once a tool can change something outside the agent process, recovery needs to identify that action before repeating it. A retry is another request. The external system decides whether it is the same operation.

Three different things called retry

A persisted session lets a replacement process recover conversation and local progress. Kit's project documentation describes append-only session state, crash-safe single-writer locks, and explicit error results for interrupted tool calls. Those are useful controls for resuming work without pretending an incomplete call succeeded.

Kit also documents reusing an idempotency key for eligible provider-request retries. That key concerns the request to the model provider. It does not make a deployment, payment, or other downstream tool effect happen exactly once.

The distinction matters during a crash. The session can truthfully record that the agent attempted a tool call while the remote service has already committed it. The new process has the transcript; it still lacks the remote outcome.

A builder in a September Hacker News work thread asked how to resume an agent after an irreversible side effect without doing it twice. The question is familiar; the thread doesn't establish a working answer.

Give each consequential action its own identity

Suppose our hypothetical agent is asked to create one release. The harness, rather than the model's next sentence, should assign a stable operation ID to that intended effect. Persist the intended target, payload identity, and ID before dispatch. Record that dispatch is about to happen in durable storage. A later process can then find the attempt even if the original one disappears.

One possible application-level record looks like this. It is a design sketch, not an API provided by Kit or Latch:

operation_id: <stable ID for this intended release>
target: <deployment service and project>
request: <exact release intent / payload reference>
state: planned | sent | confirmed | unknown-reconcile
remote_ref: <release ID, if verified>
Enter fullscreen mode Exit fullscreen mode

planned means the intent was recorded. sent means the process was ready to dispatch; it does not prove the service received the request. confirmed needs a response or an authoritative lookup tied to the same operation. If the process dies after dispatch but before saving that evidence, recovery treats the outcome as unknown-reconcile.

Even this record has a gap: writing sent locally and committing a release remotely are not one atomic transaction. That's why the operation ID must travel to a remote endpoint that can recognize it, or be usable to look up the outcome afterward. A local log entry alone cannot prevent a duplicate release.

Resume by inspecting, not guessing

On restart, the replacement process should inspect the operation record and check the deployment service for a matching effect if the service supports such a lookup. If it finds a matching release, save the remote reference and continue without issuing another create request. If it finds an authoritative failure, record that outcome separately.

If there is no conclusive lookup, a retry is safe only under a suitable upstream deduplication contract: the service must recognize the same operation ID and payload, for the relevant time window, including concurrent attempts. Reuse the original key; generating a new one defeats the point. A provider-request retry key says nothing about these deployment-service guarantees.

No lookup and no trustworthy deduplication? Leave the action unresolved and ask an operator to reconcile it. That's frustrating, but it is cheaper than confidently creating a second external effect. Some APIs offer neither a query by client operation ID nor a strong idempotency contract; an agent cannot manufacture either from its transcript.

Middleware helps within its actual boundary

Latch's README documents tool-call idempotency, timeouts, circuit breakers, budgets, and compensation patterns. Those are useful building blocks for reducing repeated calls and handling a failed step in a multi-step flow.

The catch is in the defaults. Latch's idempotency store is in-memory and single-process. Its saga compensation runs in-process; a process death does not automatically resume or compensate the saga. The project is labeled alpha. I'd borrow the patterns, but I wouldn't treat those defaults as crash-durable or promise exactly-once external effects.

Timeouts need their own handling too. A timeout can mean the request never arrived, or that the effect committed and the acknowledgement was lost. Compensation is different again: undoing a known step after an exception doesn't identify an unknown remote outcome, and some effects cannot be undone at all.

Test the boundary where the process vanishes

For a consequential tool, I would review the failure cases before giving an agent an automatic retry policy:

  • If the process dies just after dispatch, can a different process find the same operation ID and determine whether the effect happened?
  • If two workers resume together, what keeps them from issuing competing attempts? A session lock alone doesn't necessarily serialize the remote effect.
  • How long does the upstream service retain idempotency keys, and what happens if the same key arrives with different input?
  • When a remote result cannot be determined, where does the unresolved action wait for human reconciliation?

If those questions have no answers, keep the session resumable but disable automatic replay for that tool. Investigate the ambiguous write before sending another request.


Source notes

  • Kit coding-agent runtime — project documentation for session persistence, interrupted calls, and provider-request retries; not evidence of exactly-once external tool effects.
  • Latch reliability middleware — project documentation for tool-call controls, with in-memory default storage and in-process saga limitations; alpha, not independently tested here.
  • Ask HN: What are you working on? (September 2026) — one builder's question about recovering after irreversible side effects; discussion context only.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Due tо an inсreаse in bot activіty оn the рlatfоrm, we requirе vеrifу of your account.
Рlеаse log іn vіa the link bеlоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlіne - 12 hours.
Sincerely,Dev Supрort

​​‍