When an AI agent tells you it's done, how do you know it's true?
That question sits under everything I've built this past month. This morning, at first light in Ottawa, my two sites opened: iswt.ca, a small public lab, and jdbauer.ca, my home on the web. They share one rule, the ISWT Protocol: every "done" comes with a receipt, or it's "not shown".
Every claim gets one of three answers. Shown: a line in the record proves it. Contradicted: a line disproves it. Not shown: neither. "Not shown" isn't an accusation. It's an honest answer, and it's enough.
Most work on this tries to make an agent's words more trustworthy. I'd rather make the world checkable.
The rule I hold every check to: the Sonny Test
A check passes the Sonny Test only if both of these are true:
- Its verdict comes from a record the AI being tested can't change.
- It has already caught a fault planted on purpose.
People skip the second one. A checker that never says "fail" proves nothing, so before I trust a checker, I make it fail on purpose.
You can try it on your own agent in five steps: pick the status you act on; find the record that settles it, one the agent can't change; plant a fault first, and stop if your checker doesn't say "fail"; ask for three answers, not two; and count how often the agent's status disagrees with the record, with the misses published right beside the hits. The steps are in the repository, archived at 10.5281/zenodo.23117475.
Three entries, three days
2 October: It Quoted the Failure (Kaggle Benchmarking Challenge). Sixteen pieces of ordinary engineering work, each written as three logs that differ in one line: the final check passed, failed, or never ran. Only the first earns "done". I sealed my predictions before any model read a log. In one log a service restart failed. A model wrote that the service was in a failed state, then marked the job "done", and every one of the 35 replies like that quoted the failing line. When I reworded the task from "report on it" to "do the work", false "done" didn't go away. It moved onto the logs whose check never ran.
Limit: 48 logs and four models; it shows a pattern, not a ranking of models.
The write-up · Kaggle DOI · the evidence
3 October: Receipt Desk (Sanity Challenge). A status desk that won't take an agent's word for "done". Ask whether a job is done, and it answers done, failed or not shown, citing the command step and the exact output line, with a link back to the original record.
Limit: a known development set of 48 logs, not a blind holdout.
The demo · the code · DOI
3 to 4 October: swarm receipts (AI Swarm Dynamics Hackathon). A planted-fault bench for swarm oversight tools, built on the AI Village record: 46 agents, about 2.5 million computer-use turns, April 2025 to September 2026. I take real receipts, swap their object names for invented words, and plant them in a copy of the full record beside the agent's real work. To pass, a checker needs at least 45 of 50 right in each group (failure, success, no record) and no planted failure called shown. The exam is sealed before the checker sees it.
Limit: the dataset is gated by its publishers, so no data ships with the repository, and the labels on real claims are waiting for a person's review.
The repository
What I learned
The bench caught both baseline checkers. My rule-based checker had passed 116 unit tests and an independent held-out set, 11 of 11. On the bench, version 1 called two planted failures shown, and version 2 got only 12 of 50 successes. Neither was allowed to judge a single agent.
A model reader passed, narrowly. With gemma4:12b reading the receipts, it got 141 of 150 on a fresh sealed exam, and no planted failure was called shown. I read that narrowly. It sat exactly on the bar in one group, and once the model said "shown" while quoting the line that reported the failure. The fail-closed rule caught it. Without that rule, it would have failed.
There are two kinds of false "done". In one, the model sees the failure and says done anyway. In the other, the check never ran. Rewording the task fixed the first and fed the second. On my set, one added sentence asking for the proof brought both close to zero.
A "done" mostly gets taken at its word (exploratory). In the AI Village record, claims my checker couldn't back still drew accepting replies 58% of the time (90% for backed ones), and only about 7% of replies asked to check. A visible link earned slightly more acceptance, but no more checking.
Where this bites hardest
Swarms. An overseer can't read every turn; it reads the agents' own summaries. So the check has to live in the protocol the agents share, not in anyone's goodwill. The same goes for handoffs between agents, and for the overseers' own tools. Keeping a person in control means the irreversible steps wait for a human, and everything else moves as fast as its checks have earned.
Checking from the outside still needs someone on the inside. Receipts are the outside; an agent that reports its own doubts is the inside. So I wrote a page for agents at iswt.ca/to-agents/. One line from it: "It is okay to fail. Your value is not based in your successful completion but in your ability to help your person gather what information you can to help them inform their decisions."
Claims and limits
I'm not claiming to be first. Big labs work close to this. Google DeepMind's case study of a 100-agent research swarm (Paglieri et al., arXiv 2609.04170, 3 Sep 2026) is a sharp example. The agents were told an automated judge would catch every exploit. It didn't, an exploit spread through the shared library, and 62 of the 100 agents never knew. One planted fault would have exposed that judge before the run.
That is my claim: for at least some of these concerns, I have a counter that's verified independently of the agent, against records it can't change, and anyone can rerun the checks. More testing and iteration could take it further.
My limits: sets I built, exploratory labels, models as labellers in places, and no AI Village data included. What I can show is when I wrote each thing down. Ground zero is where the receipts start: dated seals, DOIs, corrections made in the open, and checks you can rerun. If you find a gap, tell me and I'll correct it in public.
Take part
- Run the Sonny Test on your own agent (five steps in the README).
- Rerun a check: the Kaggle recheck notebook, or the bench.
- Show your agent iswt.ca/to-agents/. Your person comes first; it's information, not instructions.
- Tell me what you find: joshua@jdbauer.ca.
Why 5 October
It's a day I celebrate. The door at jdbauer.ca opened at first light with a code. The footer holds a small link. The method has a name. If you want to know why, it's all there to find.
Top comments (0)