DEV Community

Cover image for Fitting one score across 40 sparse benchmarks with item response theory
Alex Fank
Alex Fank

Posted on

Fitting one score across 40 sparse benchmarks with item response theory

Part 1 described how our averaged composite failed: a model measured on three saturated legacy boards outranked flagships measured across eighteen, because hard boards lower a mean and easy boards raise it. Minimum counts and per-board z-scores did not fix it, so we changed the model. This part covers the replacement and the decisions behind it.

The idea from psychometrics

Psychometrics has a standard way to compare people who sat different exams. Each person has one latent ability, written θ. Each test item has a difficulty b and a discrimination a, how sharply it separates weaker from stronger test takers. In the two-parameter logistic (2PL) model, the expected result is a sigmoid of the gap between ability and difficulty.

We treat each benchmark as one item. A result is first rescaled to 0–1 between the board's random-guess baseline and its ceiling, and the model is then:

y(i, j) = sigmoid( a_j * (θ_i − b_j) )
Enter fullscreen mode Exit fullscreen mode

Here y is the normalized result of model i on board j, θ_i is the model's ability, b_j is the board's difficulty and a_j its discrimination (we parameterize a_j = exp(α_j) so it stays positive).

With this form a hard board stops penalizing whoever was run on it: it sits higher on the difficulty axis, and a low raw score there is what the model expects. Models that share no board also become comparable. If A and B have no board in common but each shares boards with a third model C, the fitted a and b values tie all three to one axis. Most models in our data appear on a fraction of the boards, so the whole scale rests on this linking; we check below whether it holds.

Epoch AI's Capabilities Index uses the same family of model, performance = σ(α_b [C_m − D_b]), fitted at the benchmark level; its methodology is at https://epoch.ai/data/eci-documentation/methodology. IRT has also been applied to LLM evaluation at the item level, to estimate ability from far fewer questions: tinyBenchmarks (arXiv 2402.14992), metabench (arXiv 2407.12844) and Fluid Benchmarking (arXiv 2509.11106). Like Epoch, we work with per-board aggregates, since leaderboards do not publish per-question responses.

Item characteristic curves for four real benchmarks

The figure shows four boards with their fitted reference parameters. MMLU-Pro is easy (b = −1.05, a = 1.44) and separates models only at the low end. Humanity's Last Exam is hard (b = 1.87, a = 1.03), and only high-ability models leave the floor. ARC-AGI-2 is the most discriminating board in the reference (a = 3.64, b = 1.05). Its curve is nearly a step. ForecastBench sits at the other extreme (a = 0.31, b = −0.82); its curve is almost flat, and θ depends on it only weakly. A plain average would give all four the same weight.

Decisions

Fitting, and what one board can do

We fit by maximum a posteriori estimation with weakly informative priors: θ ~ N(0, 1), b ~ N(0, 2²) and log a ~ N(0, 0.5²). The last keeps discrimination near 1 unless a board's data say otherwise. The optimizer is full-batch Adam with no random initialization, so the same data always produce the same numbers.

The prototype showed us how little a standard error from one board is worth. Hy4 Preview ranked first among open models on the strength of a single math board, and its standard error came out small. With one observation and fixed item parameters, the Laplace approximation returns a number that is too tight. The prototype report proposed gating on standard error below a threshold; we did not adopt it, since that gate relies on the very quantity that had misled us. We gate on coverage instead: a model needs at least 4 boards, spanning at least 2 categories, including reasoning or coding.

Every published score carries a 90% range from the same approximation, θ ± 1.645 standard errors mapped to the display scale. On the methodology page we give a hypothetical pair, both at 130.0: the well-covered model has a range of 126–134 and the thinly covered one 112–149.

The scale: changed once, then frozen

We first shipped the score as a bounded 0–100 number, the expected result on a typical benchmark. It saturated near 100, and the top models were hard to tell apart. On 2026-10-04, a day after launch, we moved to an open linear scale, 100 + 25θ. A score of 100 is the average model of the 2026-Q4 reference field (θ = 0), 25 points is one standard deviation of ability, and there is no ceiling. Every number changed and the order of models did not.

A free refit re-standardizes θ each time, so one new hard board would shift every model's score. We froze the a and b of 40 boards as a reference named 2026-Q4. Live fits hold those values as anchors and estimate θ, plus parameters for any board not yet in the snapshot, and a model's number moves only when new data about that model arrive. The price is that 100 means the average of a fixed field, so scores compare within one version; a rebase would ship as a new version with an old-to-new table.

Which boards count

Elo-style boards such as LMArena are relative ratings with no baseline or ceiling to normalize against, so they cannot be placed on the 0–1 result axis. They stay on separate leaderboards; Epoch's ECI excludes them on the same grounds. We also drop a board if more than 25% of its normalized results sit exactly at 0 or 1, since that suggests the baseline or ceiling we assumed is wrong.

Some boards that our earlier composite held out at weight 0 are admitted again. The weights existed because a plain average let a hard, narrowly covered board penalize the few frontier models run on it, which the IRT model does not do.

Domain scores for coding, agents, math, science and reasoning use the same global board parameters. Each is the ability estimated from that domain's boards alone, needs at least 2 of them, and is directly comparable with the overall score.

Validation

We hold out results and predict them. In 5-fold cross-validation, the IRT model reaches a held-out RMSE of 0.093 on the 0–1 scale. A baseline with a per-model offset plus a per-board mean reaches 0.177, and the plain board average 0.246.

Held-out error

Coverage improved too: in the prototype, 239 open models were estimable, against 78 that passed the earlier composite's gate. In a leave-one-board-out check, dropping any single large board moved the top 20 open models by at most about 5 places. The largest moves came from one board (5 places) and from ARC-AGI-2 (4).

Some rankings changed a great deal, and we went through the large moves one by one. DeepSeek R1 0528 fell from 5th under the earlier composite to 36th in the prototype. Under the average its rank depended on which boards it happened to be measured on; per-board difficulty places the same results below newer models that did well on those boards. Qwen3 235B A22B moved from 14th to 46th, and Seed OSS 36B, which had only 4 boards, from 7th to 37th. In the other direction MiniMax M2.7 rose from 76th to 29th and Inkling Small from 51st to 9th.

What readers of Part 1 asked

Two readers asked in the comments on Part 1 about things we had assumed and not measured. One asked whether models measured on disjoint benchmark sets end up on one scale at all, whether neighbouring ranks can be told apart, and which date the 365-day rule for active boards uses. The other raised contamination and missing data, which we come back to at the end. We ran the checks on 2026-10-09 with the production fitting code and data; our refit reproduced the live open ranking exactly.

Is it one scale?

We built the bipartite graph of models and boards and counted connected components. There is one, holding all 1,076 entities with data and all 40 boards. No entity sits outside it, gate-passing or otherwise.

Had there been a second component, a free fit would have located it only through the priors. θ ~ N(0, 1) would pull its models toward θ = 0, which displays as 100, the reference average, and that 100 would look like a measurement without being one. In the frozen fit this cannot happen for a board in the 2026-Q4 reference, because its a and b are fixed and a model is placed on the scale through its own boards whatever the graph looks like. The risk returns only for a new non-reference board observed solely by models with no anchored boards; today all 40 boards are anchored.

Can neighbours be told apart?

The published ranges come from a Laplace approximation and ignore which boards a model happened to be run on. To capture that, we resampled boards: 500 replicates, each drawing the 40 boards with replacement, with item parameters frozen at their reference values and each model's θ refit on the drawn boards (a result counts as many times as its board was drawn). The eligible set was held fixed at the 99 open models that pass the gate on full data, and ranks are among those 99.

90% rank intervals from resampling benchmarks

Kimi K3 is first in 93.8% of replicates and never below second, for a 90% rank interval of 1–2. All 14 adjacent pairs in the open top 15 have overlapping 90% rank intervals. If we ask instead how often the higher-placed model stays ahead, only one pair clears 95%: DeepSeek V4 Pro 0813 over DeepSeek V4 Flash 0731 (#3 over #4), at 97.4%. Kimi K3 over GLM 5.3 comes close at 94.4%, and from #9 down each adjacent order holds in 48–59% of replicates, about a coin flip.

The top three stay within ranks 1–4, #4 to #7 form a second block, and from #8 down the intervals spread over roughly 6 to 16. Width tracks coverage more than score: DeepSeek V4.1 Flash, scored on 10 boards, has an interval of #2–11, and MiMo V2.5 Pro, on 6 boards, #9–23, while GLM 5.2 with 24 boards sits at #5–8. We now read the open leaderboard in these tiers and would not defend a one-place gap below the top.

What date the 365-day rule uses

A board counts toward the score while it is still receiving current results, meaning its newest date is at most 365 days old. That date is the latest of three per-board values: the newest as_of on an open-model result, the newest as_of on an external result, and the newest release date of an open model scored on the board. Our code comments call the release date a fallback, but the code takes the maximum.

That makes the rule looser than it reads. We counted, for each of the 40 active boards, the entities with a result dated on or after 2025-10-10. Four boards have fewer than five: Fiction.LiveBench and SWE-bench Lite one each, Aider Polyglot and MATH Level 5 two each. SWE-bench Lite is the weakest case. Its one recent entry is DeepSeek V3.2, whose result is dated 2025-01-11; the board stays active because that model's catalogue release date is 2025-12-01, and under an as_of-only rule it would already be out. MATH Level 5 ages out on 2026-10-15 unless new results arrive. We treat this as a known weakness and intend to tighten the rule, for example by requiring a minimum number of recent results per board.

Contamination and missing data

The second reader pointed out that which boards a model is run on is not random, and that uneven contamination could distort how discriminating a board appears. Non-random coverage is what broke the average in Part 1; in the IRT model a missing board carries no penalty and only widens the range, and the resampling above shows how far ranks move with the choice of boards. We list contamination as a limitation and have not measured its effect on a. The a values are fitted once on the 2026-Q4 field and then frozen, so they reflect that field's published results as they stand.

Where it stands, 2026-10-09

Top open models with 90% ranges

On the live fit, 40 reference boards and 4,344 observations produce a published overall score for 384 models, 99 of them open. The open top three are Kimi K3 at 133.6 (90% range 131.9–135.4), GLM 5.3 at 129.9 (127.9–131.8) and DeepSeek V4 Pro 0813 at 129.2 (127.4–131.0); given the rank intervals above, the order of the second and third is not settled. The best model overall, GPT 6 Astra at 157.1, is 23.5 points ahead of the best open one, about 0.94 standard deviations of ability. It has results on only 6 boards, and its range of 150.6–163.7 is more than three times as wide as Kimi K3's.

Limitations

Contamination inflates scores whenever test questions reach training data; rotating or private sets reduce it and do not remove it. Some leaderboards accept developer-submitted runs with their own scaffolds, and we use each source's published figure as is. Where a proprietary model has several effort variants, we use the best published one. The inputs are English-centric, and knowledge, long context, instruction following, multilingual ability and vision are not covered yet. The scale is tied to a frozen snapshot of the field. The score also says nothing about whether a model fits your hardware; we estimate that separately.

Part 3 looks at how other composite indices approach the same problem (Epoch's ECI, Artificial Analysis, Arena and others) and what each choice trades away.

The full method is on the methodology page, and current scores are on the leaderboard.

Top comments (0)