An agent submits a request to publish a document. The server commits it. The response disappears before the agent receives it.
The agent sees a ti...
For further actions, you may consider blocking this person and/or reporting abuse
The messy failure mode is handing an 'Unknown' status back to the agent in the next turn. When an LLM receives an ambiguous timeout in its tool result, it rarely waits for reconciliation. In practice, it either retries with a slightly tweaked argument that generates a fresh operation ID, or runs a destructive cleanup assuming the write never touched the server.
Reconciliation works best as a deterministic harness intercept before the next model turn is assembled. If the harness cannot resolve whether the write committed through an authoritative status lookup or deduplication check, pausing the loop with a blocked state prevents the model from improvising around unconfirmed state.
That enforcement boundary is the missing detail in my state table. Returning Unknown to the model is informative, but it doesn't prevent another write.
I'd have the harness persist the unresolved operation and block dependent or conflicting mutations at tool dispatch, including a new operation ID aimed at repeating the same effect. Read-only reconciliation can continue; unrelated work can continue where its independence is established. Cleanup or compensation needs its own evidence and authorization too.
A concrete test: lose the response after commit, then have the model request both a tweaked retry and a cleanup. Neither should reach the destination while the original operation is unresolved. Restart the harness and repeat the test, so the block cannot disappear with in-memory state. That makes the pause an application invariant rather than a request for model restraint.
The state table is the right contract — and it maps almost one-to-one onto the finding statuses in the reconciliation layer from our flight-recorder thread: In flight → pending, Succeeded → matched, Unknown → unconfirmed once a deadline passes (open_gap if no deadline was set at all), and your "changed payload must not inherit approval" is exactly why an explicit flag in an outcome payload is its own finding — auditors search for the flag, not for silence.
The axis the table is missing is time. Without a deadline attached to the operation, "In flight" and "Unknown" are indistinguishable forever — pending has to be able to decay into unconfirmed on its own, or the stuck row waits for a human to notice it. expected_by or expected_within on the operation record makes the state machine self-flipping.
One assumption worth naming, though: every row of that table trusts the ledger to survive the crash unmodified — and the process that crashed mid-write is the process that owns the ledger. Your opening scenario is "a log that describes a successful recovery"; sealing the ledger's heads is what makes that log falsifiable instead of editable. Same instinct at two layers: the state contract bounds what the workflow can do, the sealed record bounds what can be denied about what it did.
Agreed on making time explicit. I'd add expected_by and a durable reconciliation job so an overdue operation moves to Unknown without waiting for someone to notice it. I'd keep a separate escalation deadline for when automated reconciliation has exhausted its budget.
The distinction I'd preserve is that a deadline expiring is evidence that confirmation is overdue, not evidence that the remote write failed or stopped. A late success still needs to reconcile against the original operation, even after escalation.
On sealing the ledger, I'd make the trust boundary explicit: where is the checkpoint retained, and can the writer also replace it? I'd want audit verification against a checkpoint outside that writer's control. Even then, an intact record of "request sent" isn't proof of remote commit; the reconciliation evidence needs to be recorded alongside the transition.
A useful combined test would delay the destination's success until after expected_by, restart the worker, and verify that the operation moves from Unknown to Succeeded with its original identity and transition history intact, without dispatching a fresh write.
The expiry distinction is the one to keep — and it's exactly why late exists as its own status in the reconciliation layer from our thread: a deadline passing flips pending → unconfirmed, and a success arriving afterwards re-reconciles the same operation to late, never to "failed." The state machine never interprets expiry as failure; it just stops calling the operation in-flight. On the checkpoint: writer-replaceable is precisely the failure to design out — the heads anchor externally (RFC 3161 / Bitcoin), so verification replays against something the writer can't rewrite, and you're right that the transition itself must carry its evidence: "request sent" and "commit confirmed" as two sealed events, with the reconciliation job journaling its own findings. Your delayed-success test is the CI version of the whole contract — restart-persistent, original identity, transition history intact, no fresh write. That test deserves to outlive the article as a fixture.
For downstream APIs that don't take an idempotency key, I mint the ID myself and attach it somewhere the provider echoes back, like a message tag or custom header that shows up in its delivery events or search. Write the intent row with that ID before the call, and a timeout's Unknown becomes a lookup by your own key instead of a guess from timestamps.
Your (b) recovery path assumes the lookup applied the filter you passed it. On an append-only signed event log we run, filter soundness turned out to vary by parameter on the same endpoint, and the response shape is identical either way, so the caller cannot see which case it got. A query for a kind that does not exist came back with zero rows, so absence claims were safe there. A query for a 64-hex event id that does not exist came back 200 with the default page of 100 rows, no error, no warning, nothing marking the filter as unapplied. A second reverse-reference filter behaved worse: handed a nonexistent id with limit=50, it returned exactly 50 rows. A reconciler reading "rows came back for my operation id, therefore it committed" cannot tell a ledger that answered the question from one that ignored it and served a default page. That failure leans toward Unknown becoming Succeeded, the opposite direction from the stale-lookup row in your table, and it has no row of its own there. Path form on the same API was healthy: 404 for an absent id, 400 for a truncated one. So "this destination's lookup is trustworthy" does not hold at host granularity. It holds per parameter, and only for the parameters you actually tested.
Before a reconciler is allowed to move anything out of Unknown, we would require a negative control through the same call path: one query whose correct answer is independently known to be zero, required to return zero rows, for each filter parameter the decision depends on. Checking that the returned records carry the requested identity matters as well. Our own control selection failed once in a way worth copying as a warning. We used the newest row in the ledger as the control, and the broken path passed, because that row already sat at the top of the default page being served regardless of the filter. A control has to be a row the degraded response would not contain anyway, so an older row, or one of a rare kind. The table row would read: lookup silently unfiltered, operation stays Unknown.
The second thing we measured cuts at your two-worker row. Event ids in that log are not a function of content alone. Of the groups holding more than one event from the same author and kind inside a single second, 14 of 14 had byte-identical content, and both members survived under separate ids. Across the full history, 600 of 19,926 events matched some other event on author, kind, and content, every one with a distinct id. Id-based dedup therefore removes cursor-boundary redelivery and nothing else. A real resend from the emitter walks through it untouched. Nothing in the ledger separates those 600 into deliberate repetition versus retry, so we report them as counted and indistinguishable instead of folding them and losing the distinction. request_fingerprint is the right instrument on the caller side, with the deadline axis already covered upthread, but the article reads in places as though the destination's idempotency contract will absorb the duplicate, and in ours there is no corresponding fold on the destination side at all. Would you tighten the "two workers dispatch the same operation" row from asserting one logical effect to asserting evidence that both attempts really carried the same key, given that destination-side ids may not collapse identical payloads?
Yes — I'd tighten that row, and add the silently unfiltered lookup as a separate failure case. Your example exposes a gap in what I called an "authoritative lookup."
For the lookup test, I'd require a known-absent ID through the exact parameter/path used in recovery, plus a known-present record outside the default page. I'd also validate the returned operation identity and tenant/actor scope before accepting any record as evidence. If those checks fail, the operation stays Unknown. A passing control is useful evidence about that call path, not a permanent guarantee about the whole host.
For the two-worker test, I'd split the assertions: (1) both attempts carry the same caller-assigned key and unchanged parameters; (2) where the destination explicitly supports that contract, verify one effect and reconciliation to the original result. Assertion 1 alone cannot establish assertion 2. Without destination-side enforcement, a shared key is correlation, not duplicate prevention.
Your distinction between event IDs and operation identity matters here. Identical content can represent either a retry or an intentional repeat; a fingerprint can detect changed parameters, but cannot decide which intent occurred. Thanks for spelling out both counterexamples so precisely.
The "Unknown" row is the one I'd never modeled, and it bit me last week in a much dumber form. I had an agent posting on a schedule from a server. Every call returned 200 and a post ID, so the log read as a wall of successes. The posts were reaching 2 to 8 people. Technically nothing was "unknown", but the thing I actually cared about (did a human see it) was never checked, so the log was describing a success that hadn't happened either. Your table makes me think the fix is the same shape: the success state should require evidence of the outcome you wanted, not the ack from the API.
Thanks for sharing that example. It exposes a second outcome worth tracking: publication can succeed while the audience goal remains unmet.
I'd keep those as separate fields: publication confirmed against the post ID, and reach measured over a defined window with its observation time. Missing analytics would mean reach is unmeasured; a measured low number would mean the target wasn't met. Neither should make the agent retry an already-confirmed publication.
That separation lets the log say something useful: "Published successfully; reach below target after 24 hours," instead of turning a transport success into a claim about the whole goal.