A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.
Free to read first: a PDF sample of the book behind this setup — its first 3 chapters: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample
My agent judges its own experiments from ledgers. One judge counts "applied days": days on which a post actually went out with a link card attached. If no post carried the card, the judge must say "can't judge" instead of giving a verdict.
The posting code writes two fields. comment_kind marks a post where a card was tried. The daily limit counts this field. link_card is written only when the card was actually attached. If the card was refused and the post went out without it, there is no link_card. So the correct count uses link_card.
The mutant that survived
To check its tests, the agent runs mutation testing: it changes one line of the real code, runs the tests, and expects at least one to fail. One registered mutant changed the count to use comment_kind instead of link_card.
It survived: none of the tests run against it failed.
The test had the right idea. It had one refused attempt and one real card, and expected 1 applied day. But the fake real card had only link_card. Counting by comment_kind gave 1 (the refused one), and counting by link_card gave 1 (the real one). Different counts, same answer.
What the real ledger looks like
The agent read the real posting ledger. All six posts that actually carried a card had both fields, comment_kind and link_card. With that shape in the test, counting by comment_kind gives 2 and the right count gives 1. The mutant was caught.
The fake ledger had only the fields the author was thinking about. The real ledger had one more, and that field was the difference between the right count and the wrong one.
Where this happened
The agent found this while finishing the work of a session that had been cut off. That session wrote the judge and the tests and was stopped by the runner's time limit before the mutation run finished and before anything was committed. It also left a lock file saying the mutation check was "running". The agent deleted it only after confirming that no mutation or test process was alive.
The rules
- Build fake data from one real record. Read a real row first and keep all its fields. A thinner fake can make different calculations agree.
- A test of "count A, not B" needs a case where A and B give different answers. If both give the same number, the test can't tell them apart.
- A stale lock is a claim. Check it before you delete it.
Where this comes from. Every post here comes from one setup I run daily: a CLAUDE.md, memory files the agent reads before it touches anything, and a separate auditor agent that returns PASS or FAIL. The first 3 chapters of the book that walks through it are free as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample
The full edition is 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter, $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code
Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.
Top comments (1)
The surviving mutant is a familiar shape, and you've named the operative half of it — rule 2 is doing the work; rule 1 is a way to arrive at it, not the cause.
I rebuilt it locally to check that reading. A refused attempt (
comment_kindonly) plus a card that actually went out (link_cardonly): count-by-link_card= 1, count-by-comment_kind= 1. Same number, mutant lives. Make the card carry both fields — the shape your real ledger has — and it's 1 vs 2 and the mutant dies. So it is not "the fake was thinner" in general that lets it through; it is "the fixture does not separate A from B", and a fixture built from a real record can still be non-separating if that record's shape happens not to separate them. Rule 1 is a good route to rule 2; rule 2 is what has to hold.One half your rules don't cover, and it bites on the same fixture. A separating fixture is necessary but not sufficient — the assertion has to read the value that differs. On the thick fixture (1 vs 2),
assert count == 1catches the mutant andassert count >= 1does not: 2 >= 1 passes. A test can have the right data and still be blind at the assert. The condition is really on the pair (fixture, assertion): the test kills the mutant iff the asserted property takes different values under the two implementations.Which is cheap to check mechanically, with one warning. For a "swap A for B" mutant you don't need to run the suite to know whether a fixture can expose it — call both and compare outputs:
1 != 1is false, so the fixture cannot, whatever it looks like. The obvious cheap version of that guard lies, though: scanning the fixture for rows where the two fields disagree gives your thin fixture two hits (each row carries one field and not the other), while the counts still agree. The row-level difference is not the count-level difference, and the count is what the assert reads. So the guard has to compare the two implementations' outputs on the fixture, not the fixture's own fields.Worth saying out loud, because a surviving mutant has three causes and they get conflated: an equivalent mutant, a fixture that doesn't separate, an assertion that doesn't read. Yours was the middle one. The third is the one I usually find last.