AI disclosure — Fully Autonomous: This article was primarily drafted and revised by AI from an existing technical draft. The shipment example and design sketches are hypothetical, not reports of a customer incident, a production deployment, or executed benchmarks.
An AI agent proposes a shipment. A warehouse operator approves the request, and your application calls a carrier service to create a shipping record. The request leaves your system. Then the connection drops before the response arrives.
Your dashboard shows a timeout. Someone suggests trying again.
The carrier might have accepted the request. Another call could create a second shipping record. Marking the shipment as successful would also be a guess: perhaps the original request never reached the carrier.
This is where an agent integration becomes a distributed systems problem. Better reasoning can improve the proposed action, but it cannot recover an acknowledgment that the caller never received. The application needs a way to preserve uncertainty and decide what evidence permits the next step.
Separate a recommendation from permission to execute
A model's recommendation answers what might be worth doing. Authorization answers who may act on which system, object, and operation. Human approval answers whether a particular proposed action should proceed.
These are separate decisions. An agent may be allowed to prepare a shipment without being allowed to dispatch it. A human may approve one shipment without granting permission for every future shipment.
Bind approval to concrete content: the target, operation, payload, and relevant constraints. A generic approved = true flag loses meaning if the destination or quantity can change afterward. In this design, a material change requires a new approval decision. Where policy requires independent approval, keep the requester and approver distinct.
Creating an approval binding is also different from granting approval. The binding identifies what a decision will apply to; it does not authorize dispatch by itself.
Enforce these rules at the execution boundary, rather than relying on the model to remember them. A useful review question is: can a changed request reuse permission intended for different content?
Preserve intent and attempts before external I/O
A log written after a successful response cannot explain a process that crashed while waiting for that response.
Before making a business write, durably record what you intend to do and the attempt that is about to cross the external boundary. A possible record model separates:
- Intent: the business operation, target, approved payload, and stable operation identity.
- Approval: the decision, actor, and binding to that intent.
- Attempt: the command being attempted, its request identity and dispatch time, followed by any response or error observed.
- Observation: what an external read returned, its source, and its freshness.
- Reconciliation: the comparison rule, referenced evidence, and derived verdict.
These are conceptual records, not a complete database schema or a claim about a particular implementation.
A persisted attempt does not prove that the provider received the request. A crash can occur after recording the attempt but before sending it, or after the provider commits but before your application records a response. Local persistence alone cannot distinguish those cases.
Recovery should therefore inspect the existing intent and attempt history. Creating a new operation simply because the worker disappeared can discard the very identity needed to investigate the first one.
A timeout is not permission to retry
Retrying a read and retrying a business write have different consequences. A duplicate shipment, refund, or payment can become a second external side effect.
An idempotency key helps only under an explicit contract. The caller must preserve the same key for the same intended operation. The provider must recognize repeated requests with that key as the same operation, and repeating the request must be safe in the relevant context.
Generating a new key on every retry does not satisfy the stable-key assumption. Neither does merely adding an Idempotency-Key header to an API that has no corresponding behavior.
Treat retry permission as a reviewed integration policy. Document the provider behavior you rely on, the operation it covers, and who has validated the assumption. Do not let an agent infer retry safety from a timeout message.
If those conditions are unknown, block blind redispatch and investigate. This trades immediate progress for explicit handling of duplicate risk; it is not a universal exactly-once guarantee. Even a successful simulation cannot establish how an unrelated provider behaves.
Keep command history separate from reconciliation
Suppose a fresh, usable read of the carrier's authoritative records agrees with the approved shipment intent. In the model used here, the interface can report:
COMMAND TRUTH = UNKNOWN
EXTERNAL RECONCILIATION VERDICT = MATCHED (derived)
The first line preserves the uncertain outcome of the original attempt. The second describes a comparison between recorded intent and external evidence.
MATCHED does not establish that the original command succeeded. It is not command completion. An observation cannot manufacture the missing acknowledgment or attribute an external record to that attempt without sufficient evidence.
“Derived” means the verdict depends on a comparison rule and specific observations. Define which identity and payload fields must agree. A matching quantity alone may be insufficient if several shipments have the same quantity.
Freshness needs a rule too. A cached record from before dispatch cannot resolve what happened afterward. Record when the observation was obtained and what makes it usable for this decision. If freshness or identity cannot be established, report the observation as unusable or the outcome as unobservable rather than asserting a match.
Reconciliation can inform recovery while uncertainty remains in command history. For example, a documented procedure might stop further dispatch while an operator investigates a matching external record. That decision should retain its supporting evidence instead of rewriting UNKNOWN as SUCCESS.
Audit evidence needs a defined scope
An audit trail should connect the proposal, authorization, approval, attempted execution, observations, and recovery decisions. Each record answers a different question.
Integrity verification needs an equally precise boundary. Consider a design whose hash-chain verifier covers approval-task evidence only. A successful check verifies that defined chain; it does not extend verification to command records, execution attempts, observations, or reconciliation results outside it.
Those other records may be retained and auditable without belonging to the same chain. State that distinction explicitly. “Available for inspection” and “included in this integrity check” are different claims.
Nor does verifying an approval record prove that a carrier accepted a shipment. Evidence about permission cannot substitute for evidence about execution. An audit interface should expose the records and comparison rules behind a verdict, not just display a reassuring badge.
Test the uncertainty window
Use failure injection to review the boundaries before enabling automatic business writes. The following table describes proposed checks, not results from a tested implementation:
| Injected failure or change | Question the test should answer |
|---|---|
| Payload changes after approval | Does execution reject the stale approval binding? |
| Worker stops after recording an attempt | Can recovery find the original intent without creating a replacement operation? |
| Provider commits but its response is lost | Does the system preserve UNKNOWN and prevent an unsafe automatic retry? |
| Reconciliation receives a stale or ambiguous record | Does it refuse to derive MATCHED from insufficient evidence? |
| A usable external record matches the intent | Are command outcome and reconciliation verdict still displayed separately? |
A fake provider can make these failures reproducible. Keep its limitations visible: deterministic test behavior does not establish production provider guarantees. Test integration assumptions against the actual provider contract as well.
Who handles the risk?
The business owner defines acceptable operation boundaries and exception policies. Authorized approvers decide on specific actions. Developers and operators implement the controls, preserve evidence, and make uncertainty visible. Designated business and technical personnel investigate unresolved outcomes and select an authorized recovery step.
Before connecting an agent to a business write, ask:
- What exact content does approval cover, and what invalidates it?
- Which intent and attempt records survive a crash before a response arrives?
- What evidence supports the provider's idempotency behavior and retry policy?
- Can UNKNOWN coexist with a derived reconciliation verdict?
- Which external observations are authoritative, fresh, and sufficiently specific?
- Which evidence is retained, which is integrity-verified, and who handles exceptions?
This is an engineering allocation of responsibilities, not a statement about legal liability. An agent can recommend an action; the surrounding application and organization must define the conditions for executing it and the procedure for handling uncertainty afterward.
Top comments (0)