Originally published on AI Tech Connect.
What you need to know One number is not an evaluation. Two agents that both score 40 per cent on a long-horizon task can be failing for entirely different reasons, and those reasons need entirely different fixes. Variance is a first-class metric. Running an eval once and reporting the result is the single most common mistake in agent evaluation as of September 2026. Behaviour is measurable. Solution framing, execution and feedback control can all be instrumented with rule-based signals you already have in your logs. Curves beat endpoints. Recording a progress metric per step turns each run into a shape, and the shape tells you which layer to fix. Harness is a confound. Reports that "the model got worse" frequently turn out to be harness regressions that were never attributed correctly. In…
Top comments (0)