Originally published on AI Tech Connect.
Why an uncalibrated judge lies with confidence An LLM judge produces a number for every output your system generates, at a cost of a fraction of a penny per verdict. That is the whole appeal: a handful of careful human judgements, scaled to volumes no review team could touch. But the number is only worth something if it predicts what a competent human would have said β and most judges in production have never been tested on that question even once. The failure mode is quiet. An uncalibrated judge does not error out or return garbage; it returns plausible scores with a stable distribution, dashboards trend gently upward, and everyone relaxes. Meanwhile the judge may be systematically rewarding length over correctness, favouring whichever answer it read first, or scoring its own modelβ¦
Top comments (0)