DEV Community

AI OpenFree
AI OpenFree

Posted on

Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility

TL;DR — Hugging Face now marks selected benchmark datasets as official and builds their leaderboards from .eval_results/*.yaml files in model repositories. Scores are self-reported, so the evaluation protocol is what makes a number comparable. This post walks through the protocol behind Darwin-180B-RSI, which ranks #1 on five of these boards.

How official leaderboards work

  1. A benchmark dataset carries the benchmark:official tag and an eval.yaml defining its tasks.
  2. A model publisher adds a file such as .eval_results/gpqa_diamond.yaml to their model repository with the dataset id, task id, value and date.
  3. The dataset page aggregates those files into a ranked leaderboard.

You can query any board directly:

curl https://huggingface.co/api/datasets/Idavidrein/gpqa/leaderboard
Enter fullscreen mode Exit fullscreen mode

Because entries are self-reported, the leaderboard shows the value but not always how it was obtained. The notes field and the model card are where the protocol lives.

The protocol behind five #1 scores

Common settings for every run of Darwin-180B-RSI:

Setting Value
Thinking budget 131,072 tokens per sample
Sampling temperature 1.0 · top_p 0.95 · top_k 20
Precision bf16
Engine vLLM, tensor parallel, expert parallel

Per-benchmark settings and results:

Benchmark Samples Reported score Value
AIME 2026 16 majority vote (mean 98.75) 100.0
HMMT Feb 2026 16 majority vote (mean 96.59) 100.0
GPQA Diamond up to 16 majority vote 94.44
MMLU-Pro 1 single sample 88.12
MMMU-Pro (vision) 3 majority vote 79.48

Three things to check before comparing numbers

  • Samples and voting. A majority-of-16 score is a system score and is not directly comparable with a single-sample score. Look for both.
  • Thinking budget. Reasoning models lose points to truncation when the budget is small; a cut-off answer is not a wrong answer.
  • Contamination. Training data should be filtered against every reported test set. VIDRAFT reports an 8-gram overlap filter against all benchmarks it lists.

Reproducing

All five benchmarks are public datasets and the weights are open at huggingface.co/FINAL-Bench/Darwin-180B-RSI. Serving requires 8 × B200-class GPUs (4 at minimum) in bf16:

vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code
Enter fullscreen mode Exit fullscreen mode

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Duе to an incrеase in bоt аctivity оn thе plаtfоrm, wе rеquire verіfy of уоur account.
Pleasе log іn viа the link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеаdlіne - 12 hours.
Sincerely,Dev Suррort

‌‍