TL;DR: Hugging Face currently tags 48 datasets as official benchmarks (benchmark:official). Their leaderboards are not uploaded by the benchmark owners: they are assembled automatically from small .eval_results/*.yaml files that model authors commit to their own model repos (or propose through pull requests). On October 5, 2026 I queried every board twice through the public API. The default view returned 1,023 entries. With ?base_model=false, which keeps derivative models, it returned 1,469. So 446 entries (30.4%) never show up in the default view, and the deciding field is the base_model key in each model card. Another finding: 772 of the 1,023 visible entries (75.5%) come from open pull requests, and none were flagged verified.
What are Hugging Face official benchmark leaderboards?
An official benchmark is just a dataset repo that carries the benchmark:official tag. The Hub renders a leaderboard widget on its dataset page. Each row on that widget is one model's score on that benchmark. A dataset owner does not have to run any evaluation service for this. The Hub collects scores that model repos self-report and matches them to the dataset.
You can pull the list of boards with one request:
curl -s "https://huggingface.co/api/datasets?filter=benchmark:official&limit=200" \
| python -c "import sys,json; d=json.load(sys.stdin); print(len(d)); print('\n'.join(x['id'] for x in d))"
When I ran it on October 5, 2026, it returned 48 datasets. Some of the better known ones: TIGER-Lab/MMLU-Pro, Idavidrein/gpqa, cais/hle, SWE-bench/SWE-bench_Verified, hf-audio/open-asr-leaderboard, openai/gsm8k, MathArena/aime_2026, MMMU/MMMU_Pro and harborframework/terminal-bench-2.0. The list also includes newer, narrower boards such as llamaindex/ParseBench, llamaindex/ExtractBench, LEXam-Benchmark/LEXam, FutureMa/EvasionBench and crosbylegal/RedlineBench.
Four of the 48 (tiiuae/PBench, mercor/ACE, ChrisHayduk/nanofold-public, harborframework/terminal-bench-science) had zero entries in both views when I checked. Tagged boards can exist before anyone has reported a score to them.
How are Hugging Face leaderboard entries built from .eval_results YAML?
Every row traces back to one YAML file in a model repo, under a directory named .eval_results/. The leaderboard API includes the path in each row, so you can see it yourself:
curl -s "https://huggingface.co/api/datasets/Idavidrein/gpqa/leaderboard" \
| python -c "import sys,json; e=json.load(sys.stdin)[1]; print({k:e[k] for k in ('rank','modelId','value','filename','verified','pullRequest')})"
Each row has these fields: rank, filename, value, verified, source, pullRequest, modelId, author, lower_is_better and num_parameters. filename is the YAML path inside the model repo, and pullRequest holds a PR number when the file lives on a PR ref instead of main.
Here is a real file. moonshotai/Kimi-K3 keeps five of them on main:
curl -s "https://huggingface.co/api/models/moonshotai/Kimi-K3/tree/main/.eval_results" \
| python -c "import sys,json; [print(f['path'], f['size']) for f in json.load(sys.stdin)]"
.eval_results/apex-agents.yaml 157
.eval_results/deep-swe.yaml 156
.eval_results/gpqa.yaml 152
.eval_results/hle.yaml 139
.eval_results/moonshotai__Kimi-K3.yaml 211
The GPQA one (/raw/main/.eval_results/gpqa.yaml) is just 152 bytes:
- dataset:
id: Idavidrein/gpqa
task_id: diamond
value: 93.5
source:
url: https://huggingface.co/moonshotai/Kimi-K3
name: Model Card
That is the whole contract:
-
dataset.idis the benchmark dataset repo. It has to match one of the official board IDs exactly, or the score goes nowhere. -
dataset.task_idpicks a sub-task or split when a board has several (for GPQA,diamond). -
valueis the number that gets ranked. The board'slower_is_betterflag decides the sort order (WER on the ASR board, for example). -
sourceis an attribution link. In practice it usually points back at the model card. - Optional keys show up in the wild too, like
dateand a free-textnotesfield for the protocol (sample count, temperature, token budget, precision). Use them. They are the only place a reader can see how the number was produced.
The file is a YAML list, so a single file can hold several benchmark results. Many repos still use one file per benchmark, which keeps diffs and PRs small.
Why does my model not appear on the Hugging Face leaderboard? (the base_model filter)
This is where most authors get stuck. The leaderboard endpoint has a parameter called base_model. Without it, the API (and the default widget) drops models whose card declares a base_model, meaning fine-tunes, merges, quantizations and adapters. With ?base_model=false, those derivatives come back:
B=Idavidrein/gpqa
curl -s "https://huggingface.co/api/datasets/$B/leaderboard" | python -c "import sys,json;print(len(json.load(sys.stdin)))"
curl -s "https://huggingface.co/api/datasets/$B/leaderboard?base_model=false" | python -c "import sys,json;print(len(json.load(sys.stdin)))"
For GPQA that is 111 vs 175. I ran the same comparison for all 48 boards with this script:
import requests
from concurrent.futures import ThreadPoolExecutor
API = "https://huggingface.co/api/datasets"
boards = [d["id"] for d in requests.get(f"{API}?filter=benchmark:official&limit=200").json()]
def count(bid):
d = requests.get(f"{API}/{bid}/leaderboard", timeout=60).json()
a = requests.get(f"{API}/{bid}/leaderboard?base_model=false", timeout=60).json()
return bid, len(d), len(a)
rows = list(ThreadPoolExecutor(8).map(count, boards))
dflt = sum(r[1] for r in rows); allr = sum(r[2] for r in rows)
print(len(rows), dflt, allr, f"{(allr - dflt) / allr:.1%} hidden")
Output on 2026-10-05: 48 1023 1469 30.4% hidden.
How much is hidden varies a lot from board to board:
| Board | Default | With derivatives | Hidden share |
|---|---|---|---|
| TIGER-Lab/MMLU-Pro | 141 | 216 | 34.7% |
| Idavidrein/gpqa | 111 | 175 | 36.6% |
| cais/hle | 89 | 108 | 17.6% |
| SWE-bench/SWE-bench_Verified | 67 | 89 | 24.7% |
| hf-audio/open-asr-leaderboard | 62 | 86 | 27.9% |
| likaixin/ScreenSpot-Pro | 23 | 48 | 52.1% |
| openai/gsm8k | 18 | 46 | 60.9% |
| MMMU/MMMU_Pro | 10 | 24 | 58.3% |
| llamaindex/ExtractBench | 11 | 25 | 56.0% |
| LiquidAI/ifstruct-v1.0 | 6 | 18 | 66.7% |
| mteb/arguana | 3 | 16 | 81.3% |
| crosbylegal/RedlineBench | 13 | 13 | 0.0% |
The pattern makes sense. Older or easier benchmarks like GSM8K attract many fine-tunes, so derivatives are the majority there and the default view shows a minority of the reported scores. On agentic and legal boards where most submissions are base models from labs (RedlineBench, skillsbench, Long-Horizon-Terminal-Bench), the two views are identical.
The filter reads the model card metadata, not the weights. A model trained from scratch that lists a base_model anyway (some teams do this to credit an architecture or tokenizer source) gets hidden just like a LoRA. A heavily modified derivative that leaves the field out shows up as if it were a base model. So the default view is a filter on declared lineage, not real lineage.
Who verifies scores on Hugging Face official leaderboards?
According to the API, nobody does yet. Of the 1,023 default-view rows, 0 had verified: true. Every number on these boards is self-reported by whoever controls the model repo or opens a PR to it.
The PR path is bigger than most people expect. 772 of 1,023 visible rows (75.5%) carried a pullRequest number. So the YAML sits on an open (unmerged) PR ref such as refs/pr/1, not on main. Reading those files directly needs the PR ref:
# file on an open PR, not on main
curl -s "https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro/raw/refs%2Fpr%2F1/.eval_results/gsm8k.yaml"
If you try tree/main/.eval_results on such a model, you get .eval_results does not exist on "main" even though the row shows on the leaderboard. That is expected behavior, not a bug.
Practical implications for anyone citing these boards:
- A row is a claim, not a measurement. Follow the
source.urland read thenotes, if there are any. - Protocols differ across rows on the same board. Majority voting over many samples, different thinking budgets, different task subsets: none of that is normalized. The
task_idfield is the only structural guard, and it only separates named splits. - Rows can come from someone other than the model owner, through a PR. Check the PR author before treating a number as the vendor's.
How do I submit a model score to a Hugging Face benchmark leaderboard?
There is no submission form. You commit a file:
from huggingface_hub import HfApi
import yaml, io
result = [{
"dataset": {"id": "Idavidrein/gpqa", "task_id": "diamond"},
"value": 71.2,
"date": "2026-10-05",
"source": {"url": "https://huggingface.co/your-org/your-model", "name": "Model Card"},
"notes": "GPQA Diamond, 198 items, pass@1, temperature 0, bf16",
}]
buf = io.BytesIO(yaml.safe_dump(result, sort_keys=False).encode("utf-8"))
HfApi().upload_file(
path_or_fileobj=buf,
path_in_repo=".eval_results/gpqa.yaml",
repo_id="your-org/your-model",
commit_message="Add GPQA Diamond eval result",
# create_pr=True, # use this when you are not the repo owner
)
The numbers above are placeholders for illustration. Then check whether the row appears, and which view it appears in:
import requests
def where_am_i(board, model):
base = f"https://huggingface.co/api/datasets/{board}/leaderboard"
d = [e for e in requests.get(base).json() if e["modelId"] == model]
a = [e for e in requests.get(base + "?base_model=false").json() if e["modelId"] == model]
return {"default_view": bool(d), "with_derivatives": bool(a),
"rank_default": d[0]["rank"] if d else None,
"rank_all": a[0]["rank"] if a else None}
print(where_am_i("Idavidrein/gpqa", "your-org/your-model"))
If with_derivatives is true and default_view is false, the base_model field in your card is what hides you.
Checklist for model authors before reporting a leaderboard rank
-
Match
dataset.idexactly to one of the 48 IDs fromfilter=benchmark:official. Case and owner prefix matter. -
Set
task_idwhen the board has subsets. A GPQA "main" score sitting next to "diamond" scores is misleading. -
Decide on
base_modelon purpose. If your model really is a fine-tune, keeping the field is honest, but you will only appear in the derivative-inclusive view. If the field is left over from a template or only credits a tokenizer, fix the card. Don't remove it from a real derivative just to get a better spot. -
Verify with the parameterless API, not the
?base_model=falseURL and not a cached screenshot. The parameterless call is what most visitors see by default. - Quote the right rank. "Rank 3 including derivatives" and "rank 3" are different claims. With 30.4% of entries hidden, the gap can be large.
-
Write the protocol in
notes: sample count, decoding settings, token budget, precision, subset size. Nobody verifies rows centrally, so this is your credibility. -
Merge your own PRs. If a contributor opened the eval PR, the row already counts. Merging it puts the file on
mainand makes the history auditable.
FAQ
How many official benchmark leaderboards are on Hugging Face?
48 datasets carried the benchmark:official tag on October 5, 2026. Four of them had no entries yet. Query https://huggingface.co/api/datasets?filter=benchmark:official&limit=200 for the current list.
What does base_model=false do in the Hugging Face leaderboard API?
It adds back models whose card declares a base_model (fine-tunes, merges, quantizations, adapters). The default request leaves them out. Across all boards that is 1,469 entries vs 1,023, so 30.4% are hidden by default.
Where does a Hugging Face leaderboard score come from?
From a YAML file under .eval_results/ in the model repo, on main or on an open PR ref. The leaderboard row's filename field gives the path, and pullRequest gives the PR number if there is one.
Are Hugging Face official leaderboard scores verified?
Not at the time of measurement. All 1,023 default-view rows had verified: false. Treat each row as a self-reported claim and check its source and notes.
Why does my model show on the board in one browser link but not another?
You are probably comparing the default view with the derivative-inclusive view. Run both API calls for your modelId. If only the ?base_model=false call finds you, your model card's base_model field explains it.
Can someone else add a score for my model?
Yes. Anyone can open a pull request that adds .eval_results/*.yaml to your repo, and 75.5% of visible rows came from PR refs when I measured. Review those PRs like any other change to your model card.
Methodology: all counts come from unauthenticated requests to the public Hugging Face API on 2026-10-05. Leaderboards change daily, so re-run the scripts above for current numbers.
Top comments (1)
The notes field is where I'd push hardest, and I'd add two things to your list: how many runs were made, and whether value is their mean or the best of them. A self-reported row that's the best of five runs sits on average about 1.16 run-to-run standard deviations above that model's typical run, and nothing in the YAML says which kind of number you're looking at. That matters because the gaps are small next to the sampling noise: Kimi-K3's 93.5 on GPQA Diamond's 198 questions carries a standard error of about 1.8 points from the question sample alone. The same selection works across boards too: a model that did badly on HLE can simply not commit hle.yaml, so missing rows can be missing for a reason.