TL;DR
- A leaderboard rank is a point estimate, not a fact. Every accuracy score carries a confidence interval, and on most public evals that interval is wide enough to swallow the gap between the top few models.
- You can estimate the error bar yourself from two numbers you already have: the accuracy and the number of eval items. For a 300-item test at 70 percent accuracy, the 95 percent confidence half-width is about 5 points. A 0.4-point lead inside that band is noise.
- For paired comparisons on the same questions, use McNemar's test, not two separate intervals. It is strictly more sensitive and it is what the gap actually deserves.
- Before trusting any ranking, ask whether the benchmark can even discriminate. If every top model scores above 95 percent, the test is saturated and the order is close to random.
- Popularity is not quality. The most-liked open models on the Hub (cover chart, pulled 2026-10-06) are a social signal, not an accuracy ranking. Treat the two as separate axes.
This is the second post in Leaderboard Forensics. The first covered how the Hub's benchmark plumbing actually works. This one is about the single most common misreading of those numbers: believing a rank order that the sample size cannot support.
Why do the top open models cluster so tightly?
Open-model evaluation has matured to the point where the frontier is crowded. On a typical reasoning or knowledge benchmark, the top ten entries often sit inside a three-point band. The cover chart above shows the eight most-liked open text-generation repositories on the Hugging Face Hub as of 2026-10-06, and likes are a popularity proxy, but the same crowding shows up in accuracy tables.
Here is the uncomfortable part. When scores cluster, the rank order is mostly determined by measurement noise rather than capability. The model in position one and the model in position four may be statistically indistinguishable. The leaderboard still has to print them in some order, so it prints the noise.
The curve above is illustrative. It plots the 95 percent confidence half-width for a model scoring 70 percent, as a function of how many questions the benchmark contains. At 100 items the half-width is about 9 points. At 1,000 items it drops to roughly 2.8 points. Only around 10,000 items does it fall under one point. Most popular public benchmarks live at the left end of that curve.
How do I compute the error bar on a single accuracy score?
Accuracy on a fixed test set is a binomial proportion. If a model answers k of n questions correctly, the estimated accuracy is p = k / n, and the standard error is sqrt(p * (1 - p) / n). A rough 95 percent interval is p plus or minus 1.96 * SE.
The Wald interval above is fine for a quick gut check, but it misbehaves near 0 and 100 percent and for small n. The Wilson interval is better behaved and still closed-form. Here is a compact, dependency-light helper.
import math
def wilson_interval(k, n, z=1.96):
"""95% Wilson score interval for a binomial proportion.
k = correct answers, n = total eval items."""
if n == 0:
return (0.0, 0.0, 0.0)
p = k / n
denom = 1 + z * z / n
center = (p + z * z / (2 * n)) / denom
half = (z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))) / denom
return (p, center - half, center + half)
for n in (100, 300, 1000):
p, lo, hi = wilson_interval(int(0.70 * n), n)
print(f"n={n:5d} acc={p:.3f} 95% CI=[{lo:.3f}, {hi:.3f}] width={hi-lo:.3f}")
Output:
n= 100 acc=0.700 95% CI=[0.604, 0.782] width=0.178
n= 300 acc=0.700 95% CI=[0.646, 0.748] width=0.102
n= 1000 acc=0.700 95% CI=[0.671, 0.727] width=0.056
Read that bottom line carefully. Even with 1,000 questions, the interval around a 70 percent score spans almost six points. If two models are reported as 70.3 and 69.9 on that benchmark, the difference is well inside the noise floor. The rank order is not evidence of anything.
When is a leaderboard gap actually significant?
The single-score interval is the wrong tool for comparing two models, because it ignores that they were graded on the same questions. Scores on a shared test set are correlated. The correct instrument is McNemar's test, which looks only at the items where the two models disagree.
Build a 2x2 table of agreement. Let b be the number of questions where model A is right and model B is wrong, and c the reverse. The questions both got right or both got wrong carry no information about which is better. McNemar's statistic is (b - c)^2 / (b + c), compared against a chi-square distribution with one degree of freedom.
from scipy.stats import chi2
def mcnemar(b, c):
"""b: A right & B wrong; c: A wrong & B right."""
n = b + c
if n == 0:
return 1.0
stat = (abs(b - c) - 1) ** 2 / n # continuity-corrected
return 1 - chi2.cdf(stat, df=1)
# Two models 0.4 points apart on a 300-item eval:
# say 18 items where A wins, 17 where B wins, rest tied.
print(f"p-value = {mcnemar(18, 17):.3f}") # p-value = 1.000
A one-item edge out of 300 produces a p-value of essentially 1.0. There is no detectable difference. You would need the disagreement counts to be lopsided, not the headline accuracy to be marginally higher, before the gap earns a rank.
This is why serious evaluation reports publish disagreement matrices, bootstrap confidence intervals, or at least the per-benchmark item counts. If a leaderboard gives you only a sorted column of three-decimal numbers, it is giving you false precision.
How do I tell whether a benchmark can discriminate at all?
Before arguing about margins, check that the test has headroom. A benchmark that is saturated, where the top cluster all score above 95 percent, cannot separate frontier models no matter how many items it has. The remaining questions are either ambiguous, mislabeled, or trivially easy, and the order among the leaders is driven by which model happened to win the coin flips on the hard residue.
Three quick discrimination checks:
- Ceiling check. If the top five are all within one point of 100 percent, the benchmark is done. Move on.
- Baseline check. Compute what a trivial strategy scores. If majority-class guessing or a tiny n-gram model gets 90 percent, the benchmark measures format compliance, not capability.
- Spread check. Look at the standard deviation across all submitted models. A benchmark where everyone scores 68 to 72 has almost no discriminative range left.
A benchmark earns your trust when the baseline is low, the ceiling is far away, and the spread across models is several times the single-score confidence width.
What should I actually do with a leaderboard, then?
Treat the leaderboard as a filter, not a verdict. Use it to shortlist the handful of models in the top band, then evaluate those candidates on your own held-out data that resembles your real traffic. The rank inside that top band is the least reliable information on the page. The membership of the band is the useful part.
And keep popularity and accuracy on separate axes. The cover chart is real Hub data, and it tells you what the community has adopted, which matters for tooling, documentation, and community support. It says nothing about which model is most accurate on your task. Both signals are valuable. Conflating them is how teams end up defending a 0.4-point lead that never existed.
FAQ
Is the Wald interval ever good enough?
For a sanity check on a single score with a few hundred or more items, yes. For scores near 0 or 100 percent, or for small n, switch to Wilson, which stays inside the valid range and does not collapse to zero width at the extremes.
Why McNemar instead of a two-sample proportion test?
Because both models answer the same questions, their scores are paired and correlated. A two-sample test assumes independence and will be too conservative. McNemar conditions on the disagreements, which is exactly the information that distinguishes the two models.
How many eval items do I need to resolve a one-point gap?
Roughly, to get a 95 percent half-width near one point at mid-range accuracy you need on the order of 10,000 independent items. Most public benchmarks have far fewer, which is why one-point gaps are usually unresolvable.
Do bootstrap confidence intervals help?
Yes, especially when scoring is not a clean right-or-wrong binary, such as graded or rubric-scored tasks. Resample the items with replacement a few thousand times, recompute the metric, and read the 2.5 and 97.5 percentiles. It makes no distributional assumption and handles weird metrics gracefully.
Does a larger model automatically win on a saturated benchmark?
No. On a saturated test the residual questions are dominated by noise and label quality, so scale buys you almost nothing. The honest move is to find a harder benchmark with headroom, or build one from your own data.
Can I trust the three decimals a leaderboard prints?
Only if the item count justifies them. Three decimals on a 300-item eval is false precision. The real resolution of that test is closer to one or two points, so mentally round the column before you rank it.
Further reading
Methodology note: Hub likes in the cover chart were pulled live from the public Hugging Face models API (sort by likes, text-generation filter) on 2026-10-06. The confidence-width curve and the McNemar example use illustrative inputs and are labeled as such.
Top comments (0)