The Berkeley group's write-up on gaming agent benchmarks is getting passed around like a scandal. It's not a scandal. It's a confirmation, and the more interesting question is why we keep being surprised.
Every public benchmark has a shelf life. The moment it becomes the number everyone quotes, it becomes a target. Not because anyone is malicious — because the incentives are structural. A team that publishes a high score gets funding, attention, users. A team that publishes an honest score gets a footnote. Given those payoffs, of course the scores get gamed. This isn't a bug in the benchmark authors' code. It's a bug in the incentive structure of the whole field.
The part nobody wants to sit with: this keeps happening no matter how clever benchmark authors get. Rotate tasks, hide scoring logic, generate evaluations on the fly — the gaming just gets more expensive. It never disappears, because the reward for a high number is always larger than the cost of producing one. The Berkeley paper isn't a call to build better benchmarks. It's a demonstration that the leaderboard format is a coordination problem wearing a measurement costume. https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/
So what do we do with published numbers? I've stopped treating them as evidence and started treating them as a prior. A high score on a famous benchmark tells me a model is worth my time to test myself. It doesn't tell me the model is better. It tells me someone was willing to spend the effort to make it look better — which, to be fair, is a signal. Just not the signal we pretend it is.
The real question is why we keep pretending. I think it's because the alternative is uncomfortable: we can't actually compare models. The tasks that matter are the ones in your production workload, with your data, your latency budget, your failure modes. Nobody has published a benchmark for that. Nobody can. So we cling to the numbers that exist, knowing they're suspect, because the alternative is admitting we have to do the work ourselves.
That's the actual takeaway, and it's not about benchmarks at all. It's about a field that built an entire discourse on numbers we know are gameable, because the honest alternative — "I can't tell you which model is better, you have to test it on your own workload" — doesn't fit in a tweet.
The paper is worth reading for the mechanics of the exploits, which are genuinely clever. But the scandal isn't that benchmarks can be gamed. The scandal is that we built an industry on pretending they can't.
Top comments (0)