TL;DR — Seoul-based AI startup VIDRAFT has released Darwin-180B-RSI, an open-weight 180B mixture-of-experts model that now sits at #1 on five Hugging Face official benchmark leaderboards: AIME 2026 (100%), HMMT February 2026 (100%), GPQA Diamond (94.44%), MMLU-Pro (88.12%) and MMMU-Pro (79.48%). The full weights are downloadable at huggingface.co/FINAL-Bench/Darwin-180B-RSI.
What happened
On September 28, 2026, VIDRAFT published Darwin-180B-RSI and submitted its self-measured scores to five benchmark leaderboards that Hugging Face lists as official benchmarks (datasets carrying the benchmark:official tag). Leaderboard entries on these datasets are populated from .eval_results files in each model repository.
| Benchmark | Darwin-180B-RSI | Previous #1 |
|---|---|---|
| AIME 2026 | 100.0 | 97.1 (Inkling, A.X-K2) |
| HMMT Feb 2026 | 100.0 | 92.7 (Kimi-K2.6) |
| GPQA Diamond | 94.44 | 93.5 (Kimi-K3) |
| MMLU-Pro | 88.12 | 88.0 (MiniMax-M2.1, Intern-S2-Preview) |
| MMMU-Pro (vision) | 79.48 | 79.4 (Kimi-K2.6) |
According to the public leaderboards, no model had previously reported 100% on the AIME 2026 or HMMT February 2026 boards.
Why these five benchmarks matter
- GPQA Diamond — 198 graduate-level physics, chemistry and biology questions written by domain PhDs; experts in the field reach roughly 65%.
- MMLU-Pro — 12,032 ten-option questions across 14 disciplines, a harder successor to MMLU.
- MMMU-Pro (vision) — 1,730 college-level questions where the question itself is embedded in an image (charts, scores, medical images) across 30 subjects.
- AIME 2026 / HMMT Feb 2026 — the 2026 American Invitational Mathematics Examination and the Harvard-MIT Mathematics Tournament, both olympiad-track competitions with exact answers.
How the scores were measured
The model card publishes the protocol. Every run used a 131,072-token thinking budget, sampling at temperature 1.0 / top_p 0.95 / top_k 20, in bf16.
| Benchmark | Samples per question | Reported |
|---|---|---|
| AIME 2026 | 16 | majority vote (mean over 16 = 98.75) |
| HMMT Feb 2026 | 16 | majority vote (mean over 16 = 96.59) |
| GPQA Diamond | up to 16 | majority vote |
| MMLU-Pro | 1 | single sample |
| MMMU-Pro | 3 | majority vote |
Readers comparing numbers should note that leaderboard values are self-reported by each publisher and that settings (sample count, voting, thinking budget) differ between models.
What is inside
The model's parent is Qwen3.8-Flash-Next (180B MoE, Qwen Community License). VIDRAFT says it modified only about 0.02% of the parameters — attention paths and shared experts — while leaving all 512 routed experts, the router and the vision encoder unchanged. The company attributes the result to three in-house techniques:
- Darwin — a model "diagnose-then-evolve" framework (arXiv:2605.14386)
- RSI (recursive self-improvement) — the model solves practice problems, its answers are checked against verifiable references, and it is retrained on its correct reasoning. VIDRAFT reports the same accuracy with 11% shorter reasoning than the parent.
- ZTC (Zero-Token Confidence) — a readout of the model's internal state that estimates, before generation, the probability that an answer will be correct.
Deep dives
- Perfect scores on AIME 2026 and HMMT 2026 — why thinking budget mattered
- Darwin: evolving a 180B parent model by changing 0.02% of it
- Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility
FAQ
Is the model open? Yes. Weights, config and tokenizer are public on Hugging Face under the Qwen Community License 1.0.
Who is VIDRAFT? A Korean AI deep-tech startup (CEO Minsik Kim) that describes itself as an "AI foundry". Its Darwin model family counts 50+ official models and 400+ community derivatives on Hugging Face.
Can I reproduce the scores? The model card lists the full evaluation protocol; all five benchmarks are public datasets.
Top comments (0)