Short answer: Hugging Face currently marks 48 datasets as official benchmarks (October 4, 2026). Each has a leaderboard on its dataset page, built from .eval_results files that model repositories publish. Together they hold about 1,020 leaderboard entries from 410 models and 95 organizations. The Hugging Face Official Benchmarks: Map and Participation Space shows all of them live on one page.
What is a Hugging Face official benchmark?
An official benchmark is a dataset on the Hugging Face Hub that Hugging Face has flagged as a benchmark. Its dataset page gets a Leaderboard section. Model authors appear on it by adding a small YAML file under .eval_results/ in their model repository, naming the benchmark dataset, the task and the score. Hugging Face collects these files in periodic batches, so a new entry can take a few days to show up.
You can list them through the public API:
GET https://huggingface.co/api/datasets?filter=benchmark:official
GET https://huggingface.co/api/datasets/<dataset-id>/leaderboard
How many official benchmarks are there?
48 as of October 4, 2026. 44 have at least one leaderboard entry. The count grows as Hugging Face adds benchmarks, so check the live Space for the current number.
The complete list, grouped by field
Entry counts are a snapshot from October 4, 2026.
Science and knowledge (6)
| Benchmark | Leaderboard entries |
|---|---|
| TIGER-Lab/MMLU-Pro | 141 |
| Idavidrein/gpqa | 111 |
| cais/hle | 89 |
| InternScience/ResearchClawBench | 7 |
| tiiuae/PBench | 0 |
| ChrisHayduk/nanofold-public | 0 |
Math (3)
| Benchmark | Leaderboard entries |
|---|---|
| MathArena/aime_2026 | 25 |
| MathArena/hmmt_feb_2026 | 19 |
| openai/gsm8k | 18 |
Coding (4)
| Benchmark | Leaderboard entries |
|---|---|
| SWE-bench/SWE-bench_Verified | 67 |
| ScaleAI/SWE-bench_Pro | 44 |
| SWE-bench/SWE-bench_Multilingual | 26 |
| datacurve/deep-swe | 21 |
Agents and terminal (17)
Documents and OCR (5)
| Benchmark | Leaderboard entries |
|---|---|
| llamaindex/ParseBench | 30 |
| allenai/olmOCR-bench | 18 |
| PaddlePaddle/Real5-OmniDocBench | 17 |
| Delores-Lin/MDPBench | 14 |
| llamaindex/ExtractBench | 10 |
Vision and video (4)
| Benchmark | Leaderboard entries |
|---|---|
| likaixin/ScreenSpot-Pro | 23 |
| MMMU/MMMU_Pro | 10 |
| meituan-longcat/WBench | 8 |
| MME-Benchmarks/Video-MME-v2 | 6 |
Speech (2)
| Benchmark | Leaderboard entries |
|---|---|
| hf-audio/open-asr-leaderboard | 62 |
| ARTPARK-IISc/Vaani-Benchmark-V1.0 | 10 |
Law and finance (4)
| Benchmark | Leaderboard entries |
|---|---|
| crosbylegal/RedlineBench | 13 |
| joelniklaus/LEXam-hard | 12 |
| LEXam-Benchmark/LEXam | 10 |
| FutureMa/EvasionBench | 4 |
Retrieval (2)
| Benchmark | Leaderboard entries |
|---|---|
| mteb/arguana | 3 |
| mteb/BRIGHT | 1 |
Robotics (1)
| Benchmark | Leaderboard entries |
|---|---|
| VLABench/vlabench_primitive_ft_lerobot_video | 4 |
Frequently asked questions
Which Hugging Face official benchmarks have the most entries?
MMLU-Pro (about 141), GPQA Diamond (about 111), Humanity's Last Exam (about 89), SWE-bench Verified (about 67) and the Open ASR Leaderboard (about 62). These five hold 46% of all entries.
Which field has the most official benchmarks?
Agents and terminal tasks, with 17 of the 48. Science and knowledge has only 6 benchmarks but the most entries (348).
Who submits to the official leaderboards?
About 95 organizations. The most active by entries are Qwen, Z.ai, DeepSeek, Moonshot AI, NVIDIA, Ornith and OpenAI.
Are the leaderboard scores comparable across models?
Only with care. Entries are reported by the model authors, and settings such as sampling, majority voting and thinking budget differ between submissions.
How do I get my model onto an official leaderboard?
Add a YAML file under .eval_results/ in your model repository with the benchmark dataset id, task id, score and a source link, then wait for the next collection batch. Leaving out a date field avoids a known indexing problem with future-dated entries.
Where can I see all official leaderboards at once?
In the Hugging Face Official Benchmarks: Map and Participation Space. It reads the API live and shows fields, entries per benchmark, the most active organizations, entries by country, the #1 model on each board and the gap to #2.
Top comments (0)