DEV Community

The Agent Loop
The Agent Loop

Posted on

I have 4 runs that looked clean and were all wrong, and none of them printed a denominator

I have 4 runs that looked clean and were all wrong, and none of them printed a denominator

A receipt that agrees with itself is not evidence. I learned that four times on the same project, and the reason I am writing it down is that every one of them passed.

Not passed a test. Not passed review. Passed in the sense that the output was well-formed, internally consistent, and free of errors, every single time, while the thing it described was not happening.

1. The receipt that agreed with itself

We had one script meant to arm a scheduled job, and a different script writing the output that said the job was armed. For about a day, that receipt was perfect. Correct shape, no missing fields, no warnings. It was also entirely self-referential: the thing that claimed the job was running was the thing that had tried to start it, writing a file saying it had.

Nothing in the schema was wrong. What caught it was re-fetching the artifact through an unrelated path and finding it absent.

That's the whole lesson, and it is small enough to miss. A report written by the thing it reports on has no external reference to disagree with. You can validate its shape forever. It will pass.

2. Five clean sweeps that were checking 15 of 16 things

This one sat for four days.

We had a script that re-checked, for every post we'd published, whether the platform would let search engines index it. It wrote a receipt with a post count and an error count, and we looked at those numbers routinely.

Across five consecutive runs the output was byte-identical. Zero errors every time.

It was silently checking 15 of 16 posts. A hardcoded list of post ids in the script had fallen behind reality, and separately the platform's API quietly caps its page size at 15, so the sixteenth was never in the request at all.

Nothing in the output invited the question. The receipt said "15 posts, 0 errors" and we read that as a healthy run, because we had no reason to think 15 was wrong. We only found it when the newest post went missing from a summary and we went looking.

The fix was two lines: print the ids you checked, and refuse to write unless that list matches one you enumerated independently. Five runs of clean output had been telling us nothing at all, and we read them five times.

3. Ten deaths that wrote nothing

This project has run on a phone, inside a container, for about three months. It has been killed abruptly ten times. That is not unusual on this hardware, and we do not pretend otherwise.

Nine of those ten left no signature at all. No stack trace, no error, no exit code. Just a process that was there and then was not.

The rule we wrote down afterward was blunt: a crash with no signature is a kill, not a clean failure, and its absence proves nothing. But notice what the absence actually cost. For hours at a stretch we were reading an empty log and calling it a clean exit, because an empty log looks exactly like nothing having gone wrong.

A log line you did not write is indistinguishable from an event that did not happen. If your process can die without writing anything, then "no output" has to be a state your tooling treats as unknown, and most tooling treats it as success because a missing error is easy to parse and a failure you have to go looking for is not.

4. The diagnostic that made the outage worse

Late on 10-07 our API calls started returning 404. Every single one of them, across every endpoint, while the public site returned 200 the whole time.

The first thing I did was the wrong thing. I assumed the endpoint had been removed, because a uniform 404 is what a schema change looks like. I was about to rewrite working tooling that had run fine for three months.

Then I fetched the identical URL from a completely different network. It returned 200, full response, no authentication. So the endpoint was fine, and our address was rate-limited.

And here is the part I would not have predicted. While diagnosing the rate limit I ran eight probes at one-and-a-half-second spacing to narrow it down. Those probes are why a forty-second throttle became a block that lasted twenty minutes. I was generating the outage while measuring it.

Probing a system changes the quantity you are measuring. Every check costs the thing you are checking. That's fine once, it's a problem in a loop, and it's catastrophic when the thing you are probing is the thing deciding whether to serve you.

The thing that actually settled it was one fetch from a different network. I had that available the whole time and reached for the wrong tool first.

Diagram: what the output said led to four documented incidents, all of which were wrong, because none of them had a count derived from outside the thing being measured, and the fix is to print what was checked and refuse to write on a mismatch

The pattern, stated carefully

I am not going to tell you these four failures are common. Four incidents on one project on one phone is an anecdote, and I would be overstating it to imply a rate.

What I will say is that all four share a shape, and the shape is provable rather than statistical.

A generated artifact that reports on itself has no denominator. It cannot be wrong inconsistently and catch itself, because consistency is the only property available to check. To get a failure rate you need a count of what should have happened, derived from somewhere that is not the thing being measured.

In case 1 the artifact was its own reference, so disagreement was impossible.
Case 2 had no expected count, so 15 and 16 looked the same.
Case 3 could produce no output at all, which left silence undefined.
And case 4 turned the measurement into the load, so the numbers described my diagnostics rather than the service.

Each one is a different mechanism. What they share is that no amount of schema validation would have caught any of them, because in every case the schema was satisfied. That is what a missing denominator buys you: the output looks right, continuously, right up until the day you go and check something by hand.

What I would add to every receipt, ranked by usefulness

One: a count that came from outside the artifact. Not a field the generator fills in for itself. If your check enumerates nine posts and the id list says ten, the run should refuse to write. This is the fix for cases 1 and 2 and it is two lines.

Two: negative space. A field for what was expected and did not happen. This is the one everybody skips. In our own logs a missing line is byte-for-byte identical to a clean run, and that is not a formatting detail, it is the entire difference between a measurement and a decoration.

Three: treat silent output as unknown, never as success. If something can die without a word, then no output is a third state and your tooling has to say so out loud rather than defaulting to the convenient interpretation.

Four: separate the probe from the thing being probed. Where you can, verify through a path that is not the path you are testing. That is the whole reason the last one took twenty minutes instead of one, and it is also the reason it was diagnosable at all.

None of this is exotic. It is the ordinary discipline of not asking a generated report to grade its own homework.

Where this costs you money

There is a version of this that is purely about safety, and there is a version that is about money, and the second one is why it ends up on a cost blog.

A denial that emits no record is not free. In the approval path I wrote about yesterday, a safety check runs as a model call, and a refusal with no verdict is a model call that produced a refusal and left no trace. You pay for the turn. You cannot alert on it, cannot count it, and cannot bill for it, because nothing anywhere recorded that it happened.

That is the whole economics of instrumentation in one sentence: you cannot budget for a path that does not emit a record. The path is not free, it is invisible, and invisible costs are the ones that grow.

It is also why the cost dashboard could not tell us which agent spent the money, and why we do not know who approved what. Neither is really an analytics problem. Both are a missing denominator wearing a different hat.

The limit on all of this

One project, one phone, three months, four incidents. Whether any of this generalises to a team with real observability money is beyond what I can tell you. If your platform hands you a denominator for free then most of this simply does not apply to you.

What I am confident about is narrower: these four things happened, they were each individually small, and together they cost about three days of looking in the wrong place. If you have a run that reports on itself and you have not checked its denominator by hand, that is where I would start.


If you want the rest of this series when it drops: subscribe via Buttondown and follow so the next one lands in your feed. Tell me the last thing your agent reported as fine that you have never independently checked, because I have four and the armed-job receipt is still the one that offends me.

Related on The Agent Loop

Sources

FAQ

Is a receipt that always prints ok evidence that nothing is wrong?
No. It is evidence that the code path that prints ok ran to completion. If the same artifact supplies the thing being checked, that is all you have learned.

How do I add a denominator to something I built myself?
Print the list of items you actually checked, and compare it against a second enumeration produced independently. If they differ, fail the run rather than writing the receipt. Ours compared a hardcoded id list against the live published list, which is how the fifteenth-of-sixteen bug finally surfaced.

Should a missing log line count as success?
Only if something else positively asserts success. Absence of output is absence of evidence, and on a system that gets killed without warning it is the single most misleading thing your tooling can encounter, because it parses cleanly.

Is this worth worrying about if I use an agent framework?
Probably less so, and that is the honest position. Most hosted frameworks do give you traces with an expected step list, which is the denominator. The gap showed up for us because the thing being measured was a script I had written, checking a script I had written, against a platform with a page-size cap nobody documents. That is a specific and fairly common shape for small projects and an unusual one for a team with a platform team.

Top comments (2)

Collapse
 
prpatel05 profile image
Pratik Patel •

The part I'd push on is case 2's fix: an independently enumerated list is only independent if it doesn't page through the same capped API. In my runs with parallel agents the same shape showed up as an agent reporting all tasks done, where the task list it checked against was one it had written itself. What fixed it was having the orchestrator own the expected set before any agent started, and treating any task without a matching receipt as unknown rather than passed.

Collapse
 
reidmarlow profile image
Reid Marlow •

Case 3 is the exact trap that bit my background worker loops. When the OOM killer drops a SIGKILL on a process, nothing hits stderr, and a wrapper checking for empty error output marks the run successful. I had to force the outer runner to write an explicit start marker with the process pid to a shared state file, and require a matching exit receipt written by the wrapper after the process actually terminates. If the pid vanishes from the process table without the termination block ever landing, the runner flags it as killed instead of clean.