Originally published on AI Tech Connect.
What a single quality score hides Your eval suite finishes and prints one number. Say it is 0.87. That tells you the configuration you tested beat the one scoring 0.84 and lost to the one scoring 0.91, and nothing else. It cannot answer what the next planning meeting will turn on: whether the small model — the one that scored 0.86 — is good enough at a fifth of the price. That gap is a design problem in the harness, not a reporting problem. A harness recording only quality threw away the information needed to answer the question before anyone thought to ask it, and no dashboard work recovers it: you cannot compute the cost of a run you did not measure. The fix is small and structural. Make cost and latency first-class scored dimensions of every eval case, exactly as quality already is.…
Top comments (0)