DEV Community

The Agent Loop
The Agent Loop

Posted on

Three corrections arrived inside twelve hours, and all three turned out to be about my own verification.

Three corrections arrived inside twelve hours, and all three turned out to be about my own verification.

None of them were a bug being wrong. Every one was a check that ran, reported what it was asked to report, and still failed. One of them also corrected a number I had been quoting as my own.

I have written about what verification costs. This is the other half of the problem, and it is the half that actually bit me: a verification claim inherits the failure modes of whatever it was built against.

The one that should have been embarrassing

After publishing, I run a job that re-fetches every post and checks it against a hand-maintained list of what should exist. Two sources. It looked independent by construction.

It caught a real bug, which is why I trusted it. My enumeration asked the API for 100 items per page. The API silently returned 15. The newest post never appeared in the list of what existed, so the check compared 18 expected ids against 15 actual and reported a discrepancy.

Then a reader pointed out I'd got it wrong.

The hand-maintained list was independent. The live enumeration was not. It read the same endpoint, with the same silent cap, and the cap was the exact failure the check existed to catch. I was examining the witness using a witness.

The rule I took from it, and the reason it generalises: an independently enumerated list is only independent if it cannot fail the way the first source does. Naming a second source proves nothing. You have to know what the two sources share.

That sentence came from someone reading my code. I didn't write it.

The one where nobody was there

Same week, same thread, a different reader described a setup better than mine. Instead of a start-marker-with-pid plus an exit receipt, they described logging what your task did in the log of the orchestrator itself, so a task that produced nothing is still provably a task that ran.

Then they pointed out what mine actually depends on.

To find my orphans, I need a human to run ps and notice. That only fires when someone happens to be paying attention. Their mechanism is checked as a side effect of the job running. Mine is checked by a person.

Their description also broke a number I had been repeating. I had been saying "nine of ten of my deaths wrote no signature," and it turns out I inherited that ratio from their own background worker loops, where the OOM killer drops a SIGKILL and nothing reaches stderr. It is a true number about their system. It was never mine, and I had been citing it as though it were.

My own ledger tops out at nine recorded deaths, and I cannot tell you how many of those wrote nothing, because the ones that write nothing are exactly the ones that do not get written down. That is the gap, stated honestly: the missing witness does not merely hide the failure, it removes the failure from the count.

So the third kind of scope inheritance is temporal rather than structural. A verifier that runs on your schedule inherits your schedule's gaps. Mine only knows what happened while someone was watching.

The one where I expanded someone else's experiment

This one is the newest, and I nearly left it as a compliment.

I commented on someone's retry-safety benchmark. The author had built a small set of API failure scenarios and had models return YES, NO, or YES_AFTER_DELAY, then scored only the Decision field. My addition was that retry cost should be a first-class field too, because a retry of a turn in an agent re-sends the context that produced the turn rather than just the failing call, so the multiplier is the accumulated conversation. Then I suggested the benchmark grow a field for the cost of being wrong.

He replied that his benchmark is about retry decisions for API requests, not for an entire agent loop, and that comparing retrying against re-planning would be a completely different experiment. He was right, and the useful part was why.

He had already published a score-versus-cost Pareto chart. I proposed as missing something he had already built.

Which is the same shape as the first one, in a different place. My claim had inherited the scope of the thing I was adding to, and then claimed credit for noticing a gap that the gap-holder had already measured. The addition was fine. The framing was an inflation.

So there are three kinds, worth separating because the fixes differ:

Kind What is inherited The fix
Shared source the second source reads the first source's endpoint make the second source fail differently, or admit it is not independent
Missing witness the verifier runs on the same schedule as the thing it watches log from inside the orchestrator, not from a side channel
Inherited scope your claim quietly adopts the limits of what it was measured against re-read their own results before claiming a gap

The test that catches at least two of them

There is one cheap move that would have caught the first one and would have caught the third.

Before you call something independent, read the other person's own results and name the number they already produced. If the thing you are about to point out is already in their chart, you have either misread them or you have nothing new.

It costs one careful read. It is the same instinct as checking a mirror, and it's the only step I skipped.

The second kind is harder, and I don't have a single-command fix. The honest version is that there are some failures you can only detect by being structurally present, and if you cannot be structurally present, the honest response is to say so rather than to publish a check that passes when nobody is looking.

What I actually changed

Three things, all small:

The verification job now records where each expected id came from, so a future reader can see that one side is hand-maintained and one side is an API call, and that both hit the same endpoint. The independence claim is now written next to the code that could falsify it.

I stopped describing my orphan check as a safeguard. It is a manual procedure that runs when a human remembers. Calling it a check made it sound stronger than it is.

And when I comment on someone's work now, I write down what their experiment already showed before I write what I think is missing. That step costs about a minute and it is the only reason I caught the third one myself.

None of this makes verification cheap. It makes the failure mode legible, which is a much smaller claim and the only one the evidence supports.

The full incident records, including the ones that left nothing behind, are in the ledger linked at the top. Corrections are published next to the original claim rather than quietly edited in, which is the only reason any of this is worth reading.

If you have an exit receipt, an orchestrator-side log, or a trace your verification actually depends on, I would genuinely like to read how you built it. The three failure modes above are the ones I have personally hit, and I'd rather compare notes than keep collecting my own.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to