TL;DR — Hugging Face now marks selected benchmark datasets as official and builds their leaderboards from .eval_results/*.yaml files in model repositories. Scores are self-reported, so the evaluation protocol is what makes a number comparable. This post walks through the protocol behind Darwin-180B-RSI, which ranks #1 on five of these boards.
How official leaderboards work
- A benchmark dataset carries the
benchmark:officialtag and aneval.yamldefining its tasks. - A model publisher adds a file such as
.eval_results/gpqa_diamond.yamlto their model repository with the dataset id, task id, value and date. - The dataset page aggregates those files into a ranked leaderboard.
You can query any board directly:
curl https://huggingface.co/api/datasets/Idavidrein/gpqa/leaderboard
Because entries are self-reported, the leaderboard shows the value but not always how it was obtained. The notes field and the model card are where the protocol lives.
The protocol behind five #1 scores
Common settings for every run of Darwin-180B-RSI:
| Setting | Value |
|---|---|
| Thinking budget | 131,072 tokens per sample |
| Sampling | temperature 1.0 · top_p 0.95 · top_k 20 |
| Precision | bf16 |
| Engine | vLLM, tensor parallel, expert parallel |
Per-benchmark settings and results:
| Benchmark | Samples | Reported score | Value |
|---|---|---|---|
| AIME 2026 | 16 | majority vote (mean 98.75) | 100.0 |
| HMMT Feb 2026 | 16 | majority vote (mean 96.59) | 100.0 |
| GPQA Diamond | up to 16 | majority vote | 94.44 |
| MMLU-Pro | 1 | single sample | 88.12 |
| MMMU-Pro (vision) | 3 | majority vote | 79.48 |
Three things to check before comparing numbers
- Samples and voting. A majority-of-16 score is a system score and is not directly comparable with a single-sample score. Look for both.
- Thinking budget. Reasoning models lose points to truncation when the budget is small; a cut-off answer is not a wrong answer.
- Contamination. Training data should be filtered against every reported test set. VIDRAFT reports an 8-gram overlap filter against all benchmarks it lists.
Reproducing
All five benchmarks are public datasets and the weights are open at huggingface.co/FINAL-Bench/Darwin-180B-RSI. Serving requires 8 × B200-class GPUs (4 at minimum) in bf16:
vllm serve FINAL-Bench/Darwin-180B-RSI \
--tensor-parallel-size 8 --enable-expert-parallel \
--max-model-len 135168 --trust-remote-code
Top comments (1)
Dear User,
Duе to an incrеase in bоt аctivity оn thе plаtfоrm, wе rеquire verіfy of уоur account.
Pleasе log іn viа the link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеаdlіne - 12 hours.
Sincerely,Dev Suррort