TL;DR
On 2026-10-05 we pulled every dataset tagged benchmark:official from the public Hugging Face Hub API, and every leaderboard attached to those datasets, twice: once in the default view and once with base_model=false. Every number below comes from that measurement, and about 40 lines of Python will reproduce it.
- 48 official benchmark datasets, 44 with at least one leaderboard entry.
- 1,469 entries from 158 organizations or user namespaces.
- Gini 0.71 for entries per organization, HHI 0.0373 (low concentration in market terms, high inequality in the tail).
- The top 5 organizations hold 33.8% of all entries, the top 10 hold 49.5%.
- 84 organizations (53%) appear on exactly one board.
- Median relative gap between #1 and #2: 3.71%. On 9 of 43 comparable boards the gap is under 1%.
-
446 entries (30.4%) do not appear in the default leaderboard view, because the default view filters out models whose model card declares a
base_model. 34 of 44 boards are affected.
In short, many organizations take part, but entry counts are uneven. On about one board in five the top two are close enough that the order could flip on a rerun, and nearly a third of all submissions only show up if you change a query parameter.
Why measure benchmark leaderboards at all?
QuantID works on measurement and diagnostics. A leaderboard is a measurement instrument: a fixed test, a scoring rule and a published ordering. Like any instrument, it has a resolution, a bias, and a display layer that can hide part of the signal. Most talk about leaderboards is about who is first. We care more about the instrument itself:
- How many independent parties actually participate?
- How unequal is that participation?
- How far apart are the leaders, compared with plausible run-to-run noise?
- How much of the data is filtered away before a human sees it?
Hugging Face now lets model owners attach evaluation results to a model repository through .eval_results/*.yaml files. Datasets tagged as official benchmarks then collect those files into a leaderboard. The whole system can be queried through a public API, so it can be audited like any other measurement pipeline.
What data did we use?
Three public calls, no authentication:
GET https://huggingface.co/api/datasets?filter=benchmark:official&limit=200
GET https://huggingface.co/api/datasets/{id}/leaderboard
GET https://huggingface.co/api/datasets/{id}/leaderboard?base_model=false
Each leaderboard row contains rank, value, modelId, lower_is_better, a source link and author metadata. We define the organization of an entry as the namespace before the slash in modelId (so Qwen/Qwen3.8-Flash-Next belongs to Qwen). User namespaces count as organizations. We do not merge related accounts: the API does not say which accounts belong together, and guessing would add our own bias to the measurement.
For each board, the larger of the two responses is the full population (in practice, the base_model=false response). The default response is the visible population.
Snapshot details:
| Quantity | Value |
|---|---|
| Measurement date | 2026-10-05 |
| Official benchmark datasets | 48 |
| Boards with at least one entry | 44 |
| Total entries (full population) | 1,469 |
| Entries in default view | 1,023 |
| Distinct organizations | 158 |
New YAML files are merged every day, so the leaderboards keep changing. Treat every number here as a reading taken on one date, not as a constant.
How concentrated is AI benchmark leadership?
"Concentration" covers two separate questions: who submits, and who wins.
Concentration of participation (entries)
Let $x_i$ be the number of entries from organization $i$, $N = \sum_i x_i$, and $n$ the number of organizations.
Gini coefficient (sorted ascending, 1-indexed):
$$G = \frac{\sum_{i=1}^{n} (2i - n - 1)\, x_{(i)}}{n \sum_{i=1}^{n} x_{(i)}}$$
$G = 0$ means every organization has the same number of entries. As $G$ approaches 1, one organization holds everything.
Herfindahl-Hirschman Index:
$$HHI = \sum_{i=1}^{n} s_i^2, \quad s_i = x_i / N$$
The HHI is driven mostly by the largest shares. Its reciprocal $1/HHI$ is the "effective number" of equally sized participants.
Top-k share:
$$S_k = \frac{\sum_{i=1}^{k} x_{[i]}}{N}$$
with $x_{[i]}$ sorted descending.
Measured:
| Metric | Value |
|---|---|
| Gini (entries by org) | 0.71 |
| HHI | 0.0373 |
| Effective number of orgs (1/HHI) | about 27 |
| Top-5 share | 33.8% |
| Top-10 share | 49.5% |
| Orgs on exactly one board | 84 of 158 |
These two indicators seem to disagree, and the disagreement tells us something. In antitrust terms, an HHI of 0.0373 is unconcentrated (below 0.15), because no single organization holds a large slice. The largest, Qwen, has 190 of 1,469 entries (12.9%). The Gini of 0.71, though, says the distribution is very unequal. The long tail explains both readings. More than half of all organizations appear only once, on a single board, while a few large labs submit many model sizes to many boards.
How to read the Lorenz curve: the bottom 50% of organizations account for only 7.0% of entries, and the bottom 90% account for 38.4%. The top 10% of organizations account for the remaining 61.6%.
The ten organizations with the most entries are Qwen (190), zai-org (85), deepseek-ai (82), nvidia (72) and moonshotai (67), then a four-way tie at 48 (ornith-ai, OrionLLM, google, openai), then ibm-granite (39). Entry volume mostly follows how many model sizes a lab releases. It says little about quality.
Concentration of wins (#1 positions)
We find the #1 entry of each board by sorting on value: descending by default, ascending when lower_is_better is true. Across 44 boards, 24 distinct organizations hold at least one #1 position. Among those winners, the Gini of #1 counts is 0.35 and the HHI is 0.072 (about 14 effective winners). If we add the 134 organizations with no #1 at all, the Gini rises to 0.90. That is expected when participants far outnumber boards.
The most frequent #1 holders in this snapshot were FINAL-Bench (9 boards), moonshotai (4), zai-org (4), deepseek-ai (3) and XiaomiMiMo (3). ornith-ai, tsinghua-sigs-robot-lab, tencent and MiniMaxAI held 2 each, and seventeen organizations held exactly one.
Leadership is therefore spread across about two dozen teams, and the biggest submitters are mostly not the most frequent winners. Qwen, for example, has the most entries overall but does not hold #1 on more than one board in this reading.
How do you measure benchmark competitiveness?
Concentration tells you who shows up. Competitiveness tells you whether the order at the top means anything.
For each board with at least two entries, we compute the relative gap between #1 and #2:
$$g = \frac{|v_1 - v_2|}{|v_1|} \times 100\%$$
Here $v_1$ and $v_2$ are the scores of the first and second entries under the board's own sort direction. A relative measure lets us compare boards scored in percent, in error rate or in arbitrary units.
Measured over 43 comparable boards:
| Statistic | Value |
|---|---|
| Median relative gap | 3.71% |
| Boards with gap < 1% | 9 |
| Boards with gap 1 to 2% | 5 |
| Boards with gap 2 to 5% | 10 |
| Boards with gap 5 to 10% | 9 |
| Boards with gap 10 to 20% | 6 |
| Boards with gap 20 to 50% | 3 |
| Boards with gap > 50% | 1 |
Why does a gap under 1% matter? Most of these benchmarks have a few hundred to a few thousand items. For a binomial accuracy $p$ on $n$ items, the standard error is $\sqrt{p(1-p)/n}$. At $p = 0.8$ and $n = 1{,}000$ that is about 1.3 percentage points, or about 1.6% in relative terms. Self-reported runs also differ in sampling temperature, prompt template and harness. With all that, a lead under 1% usually cannot be told apart from a tie. On about one board in five, the current #1 could plausibly change on a rerun.
At the other end, a handful of boards show leads above 20%. That usually means one of three things: a real new capability, a board with very few entries, or entries scored under different conditions. A large gap is a reason to open the source link, not a conclusion in itself.
We use a simple rule for any ranking. Report the gap together with an uncertainty estimate (2 SE or a bootstrap interval), and call it a tie when the interval contains zero.
What share of Hugging Face leaderboard entries are hidden by default?
The default /leaderboard response (and the default web view) leaves out entries from models whose model card declares a base_model. In practice that removes fine-tunes, merges, quantized variants and many other derivative releases. Passing base_model=false brings them back.
$$\text{hidden share} = 1 - \frac{\sum_b |L_b^{default}|}{\sum_b |L_b^{full}|}$$
Measured: 446 of 1,469 entries (30.4%) are hidden in the default view, and 34 of 44 boards hide at least one entry.
The filter is a design choice, not a bug: it keeps the default view focused on original models. It still has measurable consequences:
- A fine-tuned model can outscore its base model and still stay invisible to anyone who does not change the query.
- Rankings quoted from the default page and rankings computed from the API can disagree.
- Concentration statistics computed only on the default view understate the long tail, because small teams publish a large share of the derivative work.
If you report a position on one of these boards, say which view it comes from.
How can I reproduce this measurement?
The script below needs only requests and runs in under a minute.
import requests, collections, statistics as st
import concurrent.futures as cf
H = "https://huggingface.co/api"
ds = requests.get(f"{H}/datasets",
params={"filter": "benchmark:official", "limit": 200},
timeout=60).json()
ids = [d["id"] for d in ds]
def fetch(i):
def get(p):
try:
r = requests.get(f"{H}/datasets/{i}/leaderboard", params=p, timeout=60).json()
return r if isinstance(r, list) else []
except Exception:
return []
return i, get({}), get({"base_model": "false"})
rows = list(cf.ThreadPoolExecutor(8).map(fetch, ids))
def org(e): return e["modelId"].split("/")[0]
def gini(x):
x = sorted(x); n = len(x); s = sum(x)
return sum((2*i - n - 1) * v for i, v in enumerate(x, 1)) / (n * s)
entries, firsts = collections.Counter(), collections.Counter()
boards_of = collections.defaultdict(set)
gaps, n_default, n_full = [], 0, 0
for bid, default, full in rows:
full = full if len(full) >= len(default) else default
n_default += len(default); n_full += len(full)
for e in full:
entries[org(e)] += 1; boards_of[org(e)].add(bid)
if not full:
continue
lib = full[0].get("lower_is_better", False)
s = sorted(full, key=lambda e: e["value"], reverse=not lib)
firsts[org(s[0])] += 1
if len(s) > 1 and s[0]["value"]:
gaps.append(abs(s[0]["value"] - s[1]["value"]) / abs(s[0]["value"]) * 100)
v = sorted(entries.values(), reverse=True); N = sum(v)
print("boards", len(ids), "orgs", len(entries), "entries", N)
print("gini", round(gini(v), 3), "hhi", round(sum((x/N)**2 for x in v), 4))
print("top5", round(100*sum(v[:5])/N, 1), "top10", round(100*sum(v[:10])/N, 1))
print("single-board orgs", sum(len(b) == 1 for b in boards_of.values()))
print("median gap %", round(st.median(gaps), 2), "gap<1%", sum(g < 1 for g in gaps))
print("hidden %", round(100*(1 - n_default/n_full), 1))
print("top #1 holders", firsts.most_common(10))
Run it on another day and the numbers will differ. That is expected: the method stays fixed, and each run records the state of the boards on its own date.
What are the limitations of this measurement?
- An organization here is a namespace, not a legal entity. One company may publish under several namespaces, and one namespace may host models from several teams. We did not merge or split any.
-
Scores are mostly self-reported. Most rows carry
verified: falseand cite a model card as the source. We measure the published leaderboard, not the truth behind each number. -
Ties and sort direction. We trust the
lower_is_betterflag. Datasets with several metrics are treated exactly as the API returns them. - Snapshot effect. Entries arrive through pull requests to model repositories. One large release can move the top-k share by several points in a day.
- The gap metric ignores noise. A relative gap has no units, but it does not account for noise. The binomial SE estimate above is a rough guide, not a per-board significance test.
Disclosure
QuantID is a technology alliance partner of VIDRAFT, which appears on several boards (including under the FINAL-Bench namespace); data shown is computed identically for all organizations.
FAQ
How many official benchmark leaderboards are on Hugging Face?
On 2026-10-05 the API returned 48 datasets tagged benchmark:official. Of those, 44 had at least one leaderboard entry.
Is AI benchmark leadership dominated by a few labs?
Participation is unequal: the Gini is 0.71 and the top 10 organizations hold 49.5% of entries. Still, #1 positions are spread across 24 organizations, and no organization holds more than 13% of entries.
Why don't I see some models on a Hugging Face leaderboard?
The default view hides entries from models that declare a base_model in their card. Add ?base_model=false to the API call to see them. In our snapshot this hid 30.4% of all entries.
How close are the top two models on a typical board?
The median relative gap between #1 and #2 was 3.71%. On 9 of 43 boards it was under 1%, which is usually within run-to-run noise.
What is the difference between Gini and HHI here?
HHI is driven by the biggest shares, so it stays low when no single organization is large. Gini measures inequality across the whole distribution, including the long tail of single-entry participants. Reporting both keeps "no monopoly" from being read as "equal participation".
Can I get these numbers live?
Yes. The script above reproduces them, and we maintain a live map of all official boards as a public Space.
Links
- Live map of all official benchmark leaderboards: huggingface.co/spaces/quantid/huggingface-official-benchmark-leaderboards
- QuantID on Hugging Face: huggingface.co/quantid
- QuantID: quantid.co.kr
Top comments (0)