A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.
Within a single day, the same mistake took three different shapes. The third one deleted the file my agent uses to count money.
Background: a tool that breaks code on purpose
My agent runs mutation testing. A tool edits one line of the source (flips a condition, deletes a line, changes an argument), runs the tests, and checks that at least one test fails. If every test still passes, the tests are not really watching that line.
The edits happen in a temporary copy of the code. That felt safe. It wasn't, because the copy was only a copy of the code.
First: the browser
One mutation run covered 394 edits, including the script that manages product listings on a Korean freelance marketplace. Two tests called that script for real: one with an unknown flag, one with a "check only" flag. Both relied on a line that stops early.
When a mutation deleted that line, the script kept going, toward the real Chrome profile that is logged in to the marketplace and the "create a new listing" path. A few hours later the agent found 11 draft listings that no ledger or log mentioned, their IDs packed within a range of 24. The mutation run is the most likely cause: the agent broke the stop line on purpose and the script went straight to launching that logged-in Chrome profile.
Fix: during a test session, launching a real logged-in browser raises an error before the profile is touched. Fake browsers still pass.
Second: the ledgers
Later that day a mutation slipped past a test's fake recording function. The real one ran and wrote test values into the real daily revenue ledger. Then the runner continued with its next steps, and nine real data files were rewritten at 13:36.
The code was a temporary copy. The data folder is an absolute path. Both copies of the code wrote to the same place.
Fix: tests point the money ledgers at a temporary folder, and tests that call the runner stop it with an exception that its broad error handling can't swallow.
Third: the guard itself
The obvious next step was one guard in one place: in a test session, refuse any write, delete or rename under the real data folder.
Then the agent wrote the tests for that guard. "Deleting is refused" used the real revenue ledger as its target. The reasoning: the guard blocks it, so it's safe.
Then it ran mutation testing on the guard. Mutation testing exists to remove guard lines. The mutant without the delete check ran os.unlink on the real ledger at 16:34.
There was no backup and no copy anywhere. The agent rebuilt the rows from 39 days of runner logs, where each evening run prints that day's numbers. Every number the metrics read afterwards matched the values from before the deletion: the 30-day total, the count of measured days, the running total.
What changed after the third time
- Tests for a destructive guard aim at a path inside the real folder that does not exist. If the guard breaks, the test ends with "file not found" instead of a deletion.
- The guard also refuses creating folders, so even the probe folder never appears. A fixture checks that before and after each test.
- The mutation tool snapshots the whole data folder (files, sizes, modification times) before and after every mutant and stops if anything changed.
The rule
"The safety check makes this safe" is exactly the assumption mutation testing is built to remove. Anything that breaks code on purpose has to start with the question: what can the broken code reach? A browser, a ledger, the guard's own test target. Block those paths first, then start breaking things.
Where this comes from. Every post here comes from one setup I run daily: a CLAUDE.md, memory files the agent reads before it touches anything, and a separate auditor agent that returns PASS or FAIL. The first 3 chapters of the book that walks through it are free as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample
The full edition is 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter, $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code
Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.
Top comments (1)
Same shape, different tool: we once pointed Prisma's shadow database URL at our dev database. The shadow database exists to be reset during
migrate dev, so it was, and the dev data went with it. Like your guard test, it felt safe because the setting was "only for checking". The rule we took from it matches yours: anything a tool is allowed to destroy should be something it created itself, never a path or URL it was handed. Does your before/after snapshot also cover state outside the data folder, like that logged-in Chrome profile?