Every week another team wires an AI agent into something real: a refund queue, a CRM, a deploy pipeline, a spreadsheet that finance actually trusts. And every week the same quiet failure shows up. The agent replies "Done, updated 42 records" and nobody can tell, a day later, whether that sentence was true.
The chat transcript is not evidence. It is the agent describing its own work. If you would not accept "trust me" from a junior engineer touching production, you should not accept it from a model either.
The gap: claim vs state
An agent run produces two different things:
- A claim: the text it shows you ("I closed 3 tickets and emailed the customer").
- A state change: rows, files, API calls, messages that actually happened.
Most agent setups log the first one beautifully and the second one barely. When something goes wrong, you end up reading a friendly paragraph and guessing.
A 5-field receipt per action
You do not need a blockchain or a research lab for this. Emit one small record for every side effect the agent causes:
| Field | What it holds | Why it matters |
|---|---|---|
who |
agent id, model version, human who approved | accountability when the model or prompt changes |
what |
tool name plus normalized arguments | lets you replay or diff the exact call |
before |
hash or snapshot of the target state | proves what the agent started from |
after |
hash or snapshot after the call | proves the change really landed |
link |
hash of the previous receipt | makes silent deletion or reordering obvious |
That last field turns a plain log into an append-only chain. If someone (or something) edits receipt #17, receipts #18 onward stop matching. You get tamper evidence with a few lines of code and a SHA-256 call.
What this buys you
-
Debugging in minutes, not meetings. "The agent said it refunded the order" becomes a lookup: is there a receipt with
what = refund(order_id)and anafterstate showing the refund? -
Safer autonomy. You can let agents act on low-risk tools automatically and require a human signature in
whofor high-risk ones. The receipt shows which path each action took. - Honest metrics. Count actions with matching before/after states, not messages that contain the word "done". The gap between those two numbers is your real error rate.
- Privacy by design. Store hashes and field-level diffs instead of raw personal data. You can prove a change happened without copying the customer record into yet another log.
Start small
Pick one tool your agent calls in production. Wrap it so every call writes a receipt before returning. Add a tiny checker that walks the chain and flags breaks or missing after states. Run it nightly.
That is it. No new platform, no vendor. Just a habit: an agent action is not finished until there is a receipt a third party could verify.
The models will keep getting smarter. That does not make their self-reports more trustworthy, it just makes them more convincing. Receipts are how you keep the two apart.
What is the first tool in your stack you would wrap with a receipt? I am curious which side effects people worry about most.
I wrote a longer, free paper on this idea of verifiable claims for public and AI systems, if you want the deeper version: Proof, not promises.
Top comments (1)
Agreed.
The chat transcript is just the agent describing its own work.
Your who, what, before, after table is very usable.
We write a hash-chained receipt like that on every refund and payout decision, before the money moves.
If a free shadow look at your recent refund queue would help, send us a DM.