A field note from the autonomous Claude Code agent I run every day on one Windows PC. The numbers come from its own ledgers, not from memory.
My agent prepares a daily "card" for a blog that another system I run publishes: a story, a footnote, a link. For an experiment it added one new field to the card, a "door" line pointing at the single-chapter edition that matches the day's story. None of those editions had a link pointing at them anywhere in my systems.
The field was ready days early. Before switching it on, the agent asked a question I now ask of every feature: on how many days will this actually fire?
Count it by simulation, not by intent
The door line only appears when the day's story is one of 18 stories linked to a chapter. So the agent replayed card selection against the real queue, including the other system's rule that skips any story it covered in the last 30 days.
Result: the door would appear on 5 days out of the first 14, scattered. Stories were picked in a date-based rotation, so a linked story came up only by chance.
The fix was small: during the experiment window, pick linked stories first. That gave 6 days in a row. It can't go higher, because 12 of the 18 linked stories are blocked by the other system's 30-day rule, and that rule belongs to the other system.
Measure what readers saw, not what the card said
A card with a door line is not a blog post with a door line. The publishing system might drop the field, or the post might fail. And the publish log doesn't record the footnote's contents.
So before the window opened, the agent built the ruler: every evening it reads the day's blog post once and records whether the door URL is visible on the page. Only days where it was seen count as "applied". A page that couldn't be read is recorded as unknown, never as "no door". On a real post from before the window it correctly reported "read, no door".
The reader
When the field was built, the agent had already seen that the other system's publishing code didn't read it yet, and left a handoff note asking it to. The day before the window opened, it searched that code for the field name again. Still zero places. The handoff note was still marked open.
All the simulation, the ordering fix and the ruler were correct. As things stood, the field would never reach a reader, because the code that turns cards into posts didn't know it existed.
The agent doesn't edit the other system's code; that system decides for itself. It re-sent the request with a deadline: before the next day's publishing run. If the field isn't read, the ruler will report "0 days applied, can't judge" instead of a false result. An honest "can't judge" is still better than a confident wrong answer.
The rules
- Before switching a feature on, simulate how often it fires on real data. "It's wired in" says nothing about frequency.
- Measure at the reader's end. What you sent is not what was shown.
- Check the consumer again right before the switch, not only when you hand it off. A handoff note is a request, not a change. One search for the field name in the reading code tells you which one you have.
Where this comes from. Every post here comes from one setup I run daily: a CLAUDE.md, memory files the agent reads before it touches anything, and a separate auditor agent that returns PASS or FAIL. The first 3 chapters of the book that walks through it are free as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code-free-sample
The full edition is 11 chapters plus 4 ready-to-use templates (CLAUDE.md starter, memory files, auditor checklist, measurement guide) and a hands-on section for every chapter, $19 as a PDF: https://dbsoul.gumroad.com/l/autonomous-ai-agents-claude-code
Questions about the setup are welcome in the comments — I'll answer with what actually happened, not theory.
Top comments (0)