On 6 October, MLflow merged a fix of mine for a model-validation bug in MetricThreshold (mlflow/mlflow#26252). It is a small bug, and a good example of a broader failure mode in evaluation software: a tool can return a perfectly well-formed number or verdict that is not supported by the evidence it claims to measure.
One example
MetricThreshold can gate model promotion on relative improvement over a baseline. The implementation divided the change by the signed baseline value. When that baseline was negative, as R² can be, the sign of the relative-change calculation reversed. A worse candidate could pass while a better one failed.
Nothing crashed. No malformed value appeared. The gate produced a clean answer to the wrong measurement.
The merged fix measures change against the magnitude of the baseline and adds 26 regression cases across positive and negative baselines and both metric directions.
Where it sits
Between 14 and 26 September I read the scoring and gating code of seven open-source evaluation tools: the agent-skills eval harness and a reference gate script, NVIDIA SkillEvaluator, DSPy, LangSmith, DeepEval, Harbor and MLflow. I reproduced 13 defects, 12 of them offline, and filed each upstream with a public reproduction. Twelve carried a proposed fix. At the paper's 30 September cutoff, five fixes had merged and another had been approved. Since then, the MLflow fix has merged as well.
In every case the pattern was the same: the tool returned a result its evidence did not support. The 13 fall into six kinds.
- Absent evidence scored as a measurement
- A result bound to the wrong unit
- A proxy credited as the act
- Stale state read as current
- A sign error in a threshold
- A check aimed at a target the artefact never claimed
The write-up is on Zenodo: Reported, Not Measured: An Empirical Study of Measurement Defects in LLM and Agent Evaluation Tools (CC BY 4.0). It lists each case, the reproduction, the upstream link and its status at the cutoff.
The same standard applied to my own tool
I maintain an evaluation tool myself, so the paper classifies defects from Driftproof's public history under the same six kinds: a parser that converted a non-numeric judge response into a score, a badge that verified generation text but not judge text, and six false-pass cases fixed in 0.12.0. They sit in the catalogue with the same weight as the others.
Why this is not a statistics problem
A noisy measurement can be qualified: more samples, confidence intervals, an effect floor. A well-formed number that was never a valid measurement cannot be repaired that way. More runs of the MLflow gate with a negative baseline give you the same wrong answer with more confidence.
This matters because these outputs sit behind decisions: whether a model is promoted, whether a regression is accepted, whether an agent passes CI, or whether one system is reported as better than another.
The paper proposes nine fields an evaluation record should carry, each paired with the check that would read it, and maps each of the 13 defects to the field and check that could have exposed it. That mapping is a design argument. It has not been tested, and the paper says so.
Try the idea on one of your own skills
Driftproof 0.14.0 also shipped today. If you already have a Claude Code skill, /driftproof:start drafts test cases from the skill with you, writes them only after you approve, and runs a short smoke test. The smoke run is deliberately non-conclusive: its receipt says on its face that it cannot produce a verdict. A full run is required for a measured result.
- Claude Code:
/driftproof:start - CLI:
npx driftproof@0.14.0 - Release notes: v0.14.0
What this is not
The 13 cases are a catalogue of mechanisms found where I looked. Seven tools, chosen by my judgement, over twelve days. It is a lower bound on what exists, not a prevalence estimate, and it says nothing about tools I did not read.
The point is not that evaluation tools are uniquely unreliable. It is that evaluation software deserves the same suspicion we apply to every other measuring instrument: before asking how precise a number is, first ask whether the code measured what the number says it measured.
Disclosure: I build and maintain Driftproof (open source, Apache 2.0).
Top comments (0)