Every few weeks new open-weight model is released with a table of benchmark results, and every few weeks we asked the same practical question: is it better than the one we already run? A single ranked list should answer that. Building one turned out to be harder than we expected, and our first attempt failed in an instructive way. This is the first of three posts about that work.
The problem space
There is no shortage of benchmarks. Epoch AI's capabilities index alone draws on more than 50 distinct benchmarks. The trouble is that no model has been measured on all of them, and numbers that do exist are not always comparable.
Vendors report results with their own prompts, harnesses and reasoning-effort settings. Hugging Face found that different implementations of MMLU "give widely different numbers and even change the ranking order of the models". Independent boards run models themselves, which helps, but every board makes its own choices, and a model appears only on the boards whose maintainers got to it.
Benchmarks also age. The authors of MMLU-Pro note that performance on earlier benchmarks had begun to plateau, making differences between models hard to discern. GSM8K is a familiar example: a freshly written look-alike set, GSM1k, showed accuracy drops of up to 8% for some model families. New, harder benchmarks appear to separate the frontier models, and the older ones stay on the books.
Some boards rotate their questions on purpose. LiveBench replaces about one sixth of its questions in each update, so snapshot from one month is not strictly comparable to one from another. We treat each LiveBench snapshot as a unit and never mix scores across snapshots.
The result is that every model is measured on a different, sparse subset of the available boards. Any composite has to decide what to do about the holes.
What we built first
Our first version, which launched in June 2026, was deliberately plain. We normalized each result with 0 to 1 range between the board's random-guess baseline and its ceiling. Within each category (reasoning, coding, math and so on) we averaged whatever normalized scores a model had. The overall number was then an equal-weight mean of the category sub-scores. Missing entries were never filled with zeros; category a model had no data for was simply absent.
To keep thin evidence out of the ranking we added a coverage gate: a model needed results in at least three categories, and at least one of them had to be reasoning or coding. There was no minimum number of benchmarks.
This looked reasonable. It resembles the way several public leaderboards combine results, it is easy to explain, and no benchmark is privileged. We shipped it.
How it broke
On 2026-09-20 we looked at the open-model ranking after adding several new boards and found something we could not defend. Phi 3.5 MoE Instruct was measured on three benchmarks: PIQA (0.886), GSM8K (0.887) and BoolQ (0.846). Its composite was 83.5. Kimi K3 had results on eighteen benchmarks, including some of the hardest current ones, and its composite was 66.7. The ranking put the three-benchmark model well above the eighteen-benchmark one. Overall, Kimi K3 sat at number 8, behind Llama 2 70B and Phi 3.5 MoE. The top three open entries each had exactly three boards.
The mechanism is straightforward once it is visible. Hard modern boards produce low scores, even for excellent models. A model that has been run on many of them gets lower mean for it. A model that was only ever submitted to saturated classics, where nearly everyone scores near the ceiling, is never exposed to that penalty. The average rewards avoiding difficult tests.
We checked that this was not an artifact of the new boards. Recomputing without the six boards added that day still left the older models on top, with Kimi K3 at number 13.
Two fixes we measured and rejected
The first obvious fix is a minimum benchmark count. We tried it. Models with four to eleven easy boards still beat flagships, because the problem is which boards a model has, not how many.
The second is to standardize each benchmark with a z-score across all models, so an easy board stops being easy. This also failed. A legacy board has a legacy population: a 2023 model compared against other 2023-era models still earns a high z-score. With z-scores over the full field, Kimi K3 only reached number 2, and Falcon 180B and StableBeluga2 were still in the top six.
We kept a floor of four benchmarks anyway, partly because it matched how we already labeled confidence: three-benchmark composites swing wildly.
The fix that worked for v2: an active panel
What the failures had in common was that old boards were still voting. A benchmark carries information about today's models only while today's models are still being submitted to it. So we changed the rule: a benchmark counts toward the composite only if the newest result on it comes from a model released within the last 365 days.
The rule is derived from the data rather than curated by hand. Nobody on our side decides that a board is obsolete; if a board goes quiet upstream, it leaves the panel by itself. With this rule, 27 boards were active, and ten saturated classics dropped out, among them GSM8K, HellaSwag, PIQA, BoolQ and MMLU. The four-benchmark floor stayed.
The chart shows the panel as it stands on 2026-10-05: with more boards imported since, 40 are active, and the same ten classics remain frozen.
There was a cost, and we accepted it deliberately. 60 of 126 previously ranked models lost their overall rank, because they had been scored mostly on the frozen boards and fell below the floor of four active ones. Qwen2.5-7B-Instruct, with 9.7 million downloads, was among them. We chose not to hide this: a model's benchmarks page now says plainly that it no longer qualifies, instead of silently dropping the section. An absent score means there is not enough current evidence, not that the model is weak.
After the change Kimi K3 ranked number 1 of 67.
Why v2 was still not enough
The active panel fixed the visible absurdity, but it is a patch on the same estimator, and the weakness remained. Averaging still penalizes models that were submitted to hard boards. We saw this directly later on: when we considered folding in two new Epoch boards, our check showed that doing so would move 107 standings, because a drop-and-average penalizes the frontier models that a hard, narrowly covered board was run on. We held those boards out at weight 0, which treats the symptom.
It also cannot compare models whose benchmark sets do not overlap. If one model was tested on boards A and B, and another on boards C and D, an average of raw scores says nothing about which is stronger. The two sets of boards may differ in difficulty, and nothing in the average tells us by how much.
That is not a new problem. It is the situation psychometrics has dealt with for decades: people sit different exams, and the exams have to be placed on one scale, with the difficulty of each test estimated from the responses rather than assumed. The tool for that is item response theory. In the second part we will describe how we applied it to benchmark results, and what it changed in the rankings.
The live leaderboard
The current leaderboard, and the written methodology behind it, are here: llmrun Score leaderboard. The full method is documented on the methodology page.


Top comments (1)
The failure mode you caught with Kimi K3 illustrates why arithmetic means break down under selective submission. In standard benchmark aggregation, missing evaluations are never missing completely at random. Vendors choose which test suites to publish, creating heavy survivorship bias where legacy models sit safely on saturated classics while frontier architectures absorb low raw scores on uncalibrated frontier tests.
Moving to item response theory addresses the non-overlapping panel problem by estimating test difficulty simultaneously with latent capability rather than assuming an arbitrary scale. One structural hazard to watch in the estimation step is item discrimination versus training set contamination. A benchmark can register as having erratic difficulty purely because a specific training run memorized that distribution, which distorts the discrimination parameter across the rest of the model pool. Looking forward to reading your parameter fitting setup in part two.