An architectural review of go-tool-base came back with twenty-one critical and high findings, which is rather more than I fancied working through on my own. So they went out to a fleet of agents, roughly one finding each, and I sat back with a coffee and watched merge requests arrive.
They arrived fast, and all twenty-one were green.
And somewhere around the fourth or fifth I realised I had no idea whether any of it was real.
The work didn't look bad, quite the opposite: sensible diffs, tests included, tidy descriptions explaining what had been wrong and how it was now right. Kinda immaculate, in fact, and that was the problem. I was reading a series of confident accounts of problems being solved, with no way of establishing that the problems had ever been there.
A green test suite arriving alongside a fix can mean four different things, and I couldn't tell them apart from where I was sitting. The fix worked. Or the bug never existed and I'd just accepted a change that does nothing at all. Or the fix addressed something else nearby that happened to be broken. Or, the one that really nags, the test was written after the change and against the changed code, so it was never capable of failing at all.
You can't tell those apart by reading the diff, and you certainly can't tell them apart by reading an agent's description of the diff (a summary of a summary, at best).
So I stopped the flow and asked whether every agent was reproducing its bug before touching it. I asked it as a question rather than an instruction, because I honestly didn't know, and the answer mattered more than the twenty-one fixes did.
What came back
Take go/errorhandling, which is the module that decides what a command-line tool does when something goes wrong. The bug it had been sent to fix was this: if you invoked a CLI incorrectly, say you forgot the subcommand, it printed the usage text at you as it should... and then exited 0. Success! Thoroughly pleased with itself. So a shell script doing if mytool; then sailed cheerfully on, having been told the thing worked.
The fix is to exit 2 instead, the conventional Unix code for "you have used this command wrongly", which is distinct from 1 meaning "it ran and it failed".
So the test asserts two things about a fatal error: that the tool's exit function gets called at all, and that it gets called with 2. Here is that test, verbatim, run against unmodified main before a line was changed:
RED evidence (unmodified main logic, pre-fix):
- fatal_ErrRunSubCommand_exits_2_and_prints_usage FAILED:
"exit-called expectation" expected true actual false;
"exit code" expected 2 actual -1 (exit func never called).
- fatal_ErrNotImplemented_exits_2_and_reports FAILED: identical.
- The 2 non-fatal subtests PASSED pre-fix (correctly never exit).
Line by line, because it reads like noise until you know what it is saying.
expected true actual false is the first assertion failing: the exit function was never called, not with the wrong number, not too late, just never. And expected 2 actual -1 is the same fact from the other side, because -1 is what the test's stand-in exit function reports when nothing ever invoked it. There's no real exit code -1; it's the absence of one.
Then the fourth line. Two other subtests, covering the non-fatal cases, assert the opposite thing: that a warning must not exit the process. Those passed, before the fix, and that line is the one that changed my mind about the whole exercise.
That's a negative control, and it's what turns a red result into evidence: two things broken, two closely related things demonstrably fine, and a clear line between them. Without it, a wall of FAILED tells me only that something somewhere went wrong, which is a screenshot rather than a proof, and could as easily mean the test harness itself was broken and failing everything pointed at it.
go/credentials came back the same shape. Its Probe call is meant to check whether a credential store is actually usable, and it's meant to give up if the store doesn't answer in time. The agent wrote a deliberately hostile store whose every operation blocks forever (not a store anyone would ship, which is rather the point of it), pointed Probe at it with a hundred-millisecond deadline, and put a two-second guard in the test. The guard fired: Probe was still waiting, long past the deadline it was supposed to honour. Then the same test passing, in well under the guard, once the deadline was enforced.
Neither of those is clever, and I think that's what I like about them. They're cheap to ask for, cheap to read at a glance, and very hard to fake by accident. You can write a summary that sounds like diligence in about four seconds. Producing a failing run against untouched main, then the same run green, with the sibling tests behaving sensibly throughout, means you've actually done the thing.
Why I keep asking for it
I review every line my agents write. I know how that sounds, and it's true, and it's also the first thing to buckle. Twenty-one findings across a fleet is already past the point where reading every diff properly is real work rather than performance, and the fleet doesn't get tired, doesn't get embarrassed, and will happily out-produce me all week (it has, more than once).
The verifying is the expensive part, and it's the first thing to go when the work speeds up. So the review has to change shape rather than intensity: less read the change, more check the evidence the change turned up carrying. Red-first is the cheapest such evidence I've found, and one glance tells me whether the agent proved the problem before solving it.
All twenty-one went through that gate in the end. Each arrived with its red, its green, and the sibling tests that stayed sensible in between, and I read far less code than I would have done otherwise, with far more confidence in it.
Originally published at phpboyscout.uk on 3 October 2026.
Top comments (0)