A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.
My agent runs small experiments on where buyers come from. One of them added a link to the free sample of my book under some of its Threads posts. To keep an old experiment from running quietly forever, the experiment file has an end date. After that date the link stops being added. That part worked exactly as designed.
The first time
The end date was September 22. The judgement was scheduled for the evening of the same day. Then the agent's loop paused for a usage limit from the afternoon of the 22nd until the 24th, so the judgement ran two days late.
The verdict was "the door is being clicked". But the door had been switched off since the 23rd. For two days the experiment that looked like a win was closed, and nothing reported it. The off switch was automatic. "If it wins, keep it open" was left to the next session's memory.
The fix: the end date now sits at least three days after the end of the judgement window, and a test locks that gap. A judgement that runs a couple of days late, as this one did, no longer closes a winning door first.
There was a twist. Later the agent measured clicks on the comment link itself, and over the five days around the window they were zero. The views had come from somewhere else. The same address is also on the account's profile. The first verdict was wrong too, and the judge now reports that case as "views from outside the comment".
The second time
A later version put the link in the post body as a link card instead of a comment. On October 1 it was judged "clicked": 2 clicks in 5 days, against a threshold of 2. The experiment file said October 3. The three-day gap was there.
But the prescription after that win was not "done". It was: keep the card and see whether people actually take the free copy. That needs two more weeks, and the door would have closed on the fourth.
The rule the agent had written after the first time protected the judgement. It did not protect what comes after a win. The agent extended the end date to October 18 without changing the card, and wrote the next judge before the window opened: a 14-day window that measures whether clicks repeat (at least 3) and whether the free copy is actually received.
The ledger shows why repetition is the first question. Both clicks came on September 26 and 27. The four days after that had none. "Clicked" was a verdict right at the threshold.
The rules
- An automatic off switch needs an automatic "keep going" too. Otherwise the safe default quietly ends your best result.
- Put the end date after the next step, not only after the judgement. A win can ask for a longer window, not a finished one.
- Check the door, not the verdict. The agent's mistake log now says: at the start of a cycle, run the judge, then check that every "clicked" door is actually on today.
Where this comes from. Every post here comes from one setup I run daily: a CLAUDE.md, memory files the agent reads before it touches anything, and a separate auditor agent that returns PASS or FAIL. The first 3 chapters of the book that walks through it are free as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample
The full edition is 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter, $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code
Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.
Top comments (0)