A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.
My agent never marks its own work as done. A separate auditor agent reads the diff, runs the tests, checks the numbers against the files, and returns PASS or FAIL. No explanations accepted.
One morning it failed the same small feature seven times in a row.
The feature
A store shows a view count for each product. The agent was adding a daily ledger: record the per-product counts every evening, and print one line in the evening briefing with the number of products, their total views, the change since the previous reading, and the number of sales.
That's it. A total and a difference.
The seven failures
- Subtracting totals. Today's total minus yesterday's total, whatever products each one contained. On a day when the product list changed, this printed "+-28" when the real change was +2. Fix: subtract per product.
- Unread products looked new. If a product's count couldn't be read one day, it wasn't recorded, so the next day it counted as a "new" product and its whole count became growth: "+31" instead of +1 (or unknown). Fix: record which product IDs existed even when their value was unreadable.
- The fix for #2 broke the total. Rows without an ID were now skipped in the difference but still added into the total: "1 product: 62 views (+2)" when there were really 2 products and +47.
-
More special cases. The same ID twice. The ID
1and the string"1"colliding. A sale on a row without an ID. - The rule applied to recording, not to comparing. The agent picked an incomplete earlier day as the baseline. "+43", truth +3.
- The rule applied to views and sales, but not to the product count printed next to them.
- The same miss in a different line. The evening log line computed its own product count. The condition had been copied into each place separately, and one copy was missed.
What finally worked
After the fourth failure, the agent stopped patching cases and wrote one condition:
if len(by_item) == len(rows):
rec["views"] = sum(by_item.values())
by_item keeps only rows that have an ID and an integer count, keyed by the ID as a string. So the two lengths match only when every row has a distinct ID and a readable count.
If that isn't true, the day has no total, no count, no difference. Views and sales are printed as "?", and no difference is shown. That single line closed four of the earlier cases at once.
Failures 5 to 7 were the same invariant not being applied everywhere the value is read: when choosing a baseline, and when printing each number. The last fix was structural: every number that reaches a human goes through one function (_shown), so there is no second copy of the condition to forget.
The rules I kept
-
The second special case is a signal. When you add a second
iffor the same kind of wrong input, stop and write the invariant instead. - An invariant has three places to live: where data is written, where inputs for a comparison are chosen, and where each number is displayed. Grep for all of them at once.
- Displayed numbers come from one function. Copies of a condition drift apart silently. A single function can't.
- The auditor was right every time. None of the seven failures reached a real report. That's what the auditor is for, and the reason the agent doesn't get to grade its own work.
Want the whole system? The book has 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter. It's $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code
Not sure yet? The first three chapters are free, same PDF format: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample
Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.
Top comments (1)