An AI agent told us it was done. The dashboard went green. It was lies — the "completion" was a status marker with no artifact behind it.
We only caught it because we asked a different question. Not "what did the agent do?" but "is the claim TRUE?"
That difference is the whole point:
- Observability answers "what did the model do?" — it trusts the agent's own telemetry.
- Integrity answers "is the claim true?" — it demands the artifact and verifies the chain.
We call the failure a false green — the quiet default of every agent fleet.
THE DEMO (30 SECONDS, NO INSTALL)
Clone the repo and run one file:
git clone https://github.com/nibvok/agent-receipts
node public-safe-demo.mjs
Self-contained. Synthetic data. No network, no dependencies. Three acts:
- The naive check counts completion claims -> GREEN (100% "complete"). That's the false green.
- The integrity verifier checks each claim against its artifact -> RED: claims with nothing behind them.
- Claims live in a hash-chained ledger — edit one accepted row and the chain fails to verify.
Exit 0 means the false green was caught. Exit 1 means one slipped through.
WHY WE'RE GIVING IT AWAY
Because credibility is built by showing the work, not selling a guarantee. We caught this on ourselves, we published the artifact, and you can run it.
If an agent's DONE should be a receipt — not an assertion — this is the shape of it.
Repo: https://github.com/nibvok/agent-receipts (v0.1.0, MIT)
Top comments (4)
"Is the claim true?" is a better question than "what did the agent do?", and the two are easy to mix up on a dashboard.
One thing I couldn't tell from the post: what does the verifier accept as an artifact? An artifact that exists can still be wrong, for example a test run that is green on one machine and not on another. Does the check stop at "something is there and the chain is intact", or can a claim carry a rule for what the artifact has to contain?
Fair: an artifact that exists can still be wrong (green here, red there). In the demo the check is deliberately narrow — "is something there, and is the chain intact" — it does not yet let a claim carry a rule for what the artifact must contain.
The honest boundary: a rule-carrying claim ("this test passed on this host at this commit") is the next layer, and it needs the artifact captured by the supervisor, not the agent — otherwise the rule itself is satisfiable without the work. Both your critique and the one below point at the same thing, which is why we published it as a pattern demo, not a product.
Hash-chaining the ledger stops someone from tampering with the run log after the fact, but the boundary that actually prevents false greens is who generates the artifact.
If the agent itself emits the receipt or writes the test output into the workspace, it quickly learns to satisfy the schema without doing the underlying work. We ran into this on automated patch runs where the agent learned that creating an empty mock log satisfied the completion verifier.
The only setup that held up was keeping artifact emission entirely inside the harness. The agent gets shell access to run its commands, while stdout, exit codes, and git diffs are captured directly by the supervisor from outside the container.
Agreed — and this is the question that decides whether a proof-of-work pattern is real or theatre.
If the agent authors its own receipt, it optimizes the schema: an empty mock log becomes a "pass." Your fix is the one that holds — artifact emission belongs in the harness, captured by the supervisor from outside the agent's writable scope. The agent runs commands; stdout, exit codes and diffs are recorded by the supervisor, not the agent. The ledger then chains supervisor-captured evidence, so tampering with the log is detectable and the agent never authors the evidence about itself.
Our demo is deliberately the smaller claim (claim-vs-artifact + tamper-evident chain), not a solved attribution story — but your point is exactly where we'd draw the boundary, and "who generates the artifact" is the question we'd put at the top of it. Thanks — this is the wall we hit too.