DEV Community

Cover image for Done Is Testimony. Terminal State Is the Grade.
Igor Eduardo
Igor Eduardo

Posted on

Done Is Testimony. Terminal State Is the Grade.

Two things landed in the same week and they are the same argument.

Microsoft and Hugging Face published ThinkingBox, a benchmark that grades agents on the backend state and side effects they leave behind instead of the sentences they generate, and then asks whether they can do it twenty times in a row. Their opening example is an agent that makes nine clean tool calls, closes a ticket as resolved, and is wrong: the carrier exception was still open, so the required end state was "on hold." A grader reading tool calls sees nine well-formed calls. The database disagrees.

On Dev.to, The Witness Was the Suspect made the same point from the incident side: when the agent writes its own success log, the record of the incident was authored by the cause of the incident.

Here is what I would require before I trusted an agent's "done," and what I would not claim yet.

The stake

Teams will ship the wrong default.

Most agent dashboards still answer "did it say it succeeded?" Status strings, tool-call traces, and a tidy final message are all produced by the same run that took the action. They are useful for debugging. They are not a grade. When the eval and the agent share a witness, a confident wrong action scores the same as a correct one, and the failure shows up later as a customer, a ledger, or an auditor.

AI failures rarely fail loud. They fail plausible. That is exactly why the grade has to come from somewhere the agent cannot write.

Opinion (one sentence)

An agent's "done" is testimony, not evidence. The eval contract should grade the terminal state through a read path the agent does not control, and it should grade it more than once.

That is a preference with a limiter, not a benchmark result. I have not run ThinkingBox, and its numbers (507 stateful workflows, 20 runs each) are theirs. What I am formalizing is the shape: separate the actor from the grader, and treat "said done while the state is red" as a named failure, not as noise.

What separating the witness looks like (menu, not recipe)

Field Question it answers Failure if skipped
Required end state What should the record look like when the task is truly finished? "Resolved" accepted when "on hold" was correct
Independent read path Who reads the final state, and can the agent influence that read? The agent's own log grades the agent
Side-effect ledger What else changed (tickets, refunds, emails, rows)? Correct primary record, wrong collateral writes
Done-while-red Did the run claim success while the state check failed? Confident wrong actions blend into the pass rate
Repeat consistency Does the same task reach the same end state across N runs? One lucky green run treated as a capability

None of that needs a private harness. It needs someone to write down the required end state before the run, and a grader that reads the system of record after the run.

Why repeats belong in the contract

ThinkingBox's "twenty times in a row" framing is the part I would steal first.

A single green run tells you the agent can reach the right state. Production needs to know whether it will. If the same ticket ends "resolved" on some runs and "on hold" on others, you do not have a capability with a small error rate. You have a coin with a nice transcript. I would report consistency next to the pass rate, not bury it in an appendix.

Where this connects to retrieval

My public work is mostly retrieve-first systems, and the same split shows up there. A RAG answer can cite cleanly and still be wrong about the source, the same way an agent can log cleanly and still leave the wrong record. In both cases the fix is the same move: grade against something the generator did not produce. For retrieval that is a gold source set; for agents it is the terminal state.

The hybrid retrieval pipeline I keep public (hybrid-rag-pipeline) is built around that habit: measure the retrieval half on its own before letting a fluent answer vouch for it.

Steal this check

Pick one agent workflow you already run and add three things:

  1. Write the required end state first. Not "the agent responds," but "the ticket is on hold, no refund issued, customer told why."
  2. Grade from the system of record. Read the final state with a query the agent never touches, and ignore the agent's own status message for scoring.
  3. Count done-while-red separately. Every run where the agent claimed success but the state check failed goes in its own column. That column is the one that becomes an incident.

Then run it more than once and report how often it lands in the same place.

What I am not claiming

  • I am not reporting ThinkingBox results or reproducing their benchmark.
  • I am not saying traces and logs are useless. They are how you debug. They are just the wrong thing to score.
  • I am not publishing thresholds. What counts as acceptable consistency depends on what the action costs when it is wrong.

The point

The process that acted should not be the only witness that it worked. Grade the record, not the report.


Igor Eduardo builds retrieve-first and evaluation-as-contract systems. More at igoreduardo.com.

Top comments (0)