DEV Community

AI OpenFree
AI OpenFree

Posted on

Korean startup's open 180B model tops five Hugging Face official leaderboards, including perfect AIME 2026 and HMMT 2026 scores

TL;DR — Seoul-based AI startup VIDRAFT has released Darwin-180B-RSI, an open-weight 180B mixture-of-experts model that now sits at #1 on five Hugging Face official benchmark leaderboards: AIME 2026 (100%), HMMT February 2026 (100%), GPQA Diamond (94.44%), MMLU-Pro (88.12%) and MMMU-Pro (79.48%). The full weights are downloadable at huggingface.co/FINAL-Bench/Darwin-180B-RSI.

What happened

On September 28, 2026, VIDRAFT published Darwin-180B-RSI and submitted its self-measured scores to five benchmark leaderboards that Hugging Face lists as official benchmarks (datasets carrying the benchmark:official tag). Leaderboard entries on these datasets are populated from .eval_results files in each model repository.

Benchmark Darwin-180B-RSI Previous #1
AIME 2026 100.0 97.1 (Inkling, A.X-K2)
HMMT Feb 2026 100.0 92.7 (Kimi-K2.6)
GPQA Diamond 94.44 93.5 (Kimi-K3)
MMLU-Pro 88.12 88.0 (MiniMax-M2.1, Intern-S2-Preview)
MMMU-Pro (vision) 79.48 79.4 (Kimi-K2.6)

According to the public leaderboards, no model had previously reported 100% on the AIME 2026 or HMMT February 2026 boards.

Why these five benchmarks matter

  • GPQA Diamond — 198 graduate-level physics, chemistry and biology questions written by domain PhDs; experts in the field reach roughly 65%.
  • MMLU-Pro — 12,032 ten-option questions across 14 disciplines, a harder successor to MMLU.
  • MMMU-Pro (vision) — 1,730 college-level questions where the question itself is embedded in an image (charts, scores, medical images) across 30 subjects.
  • AIME 2026 / HMMT Feb 2026 — the 2026 American Invitational Mathematics Examination and the Harvard-MIT Mathematics Tournament, both olympiad-track competitions with exact answers.

How the scores were measured

The model card publishes the protocol. Every run used a 131,072-token thinking budget, sampling at temperature 1.0 / top_p 0.95 / top_k 20, in bf16.

Benchmark Samples per question Reported
AIME 2026 16 majority vote (mean over 16 = 98.75)
HMMT Feb 2026 16 majority vote (mean over 16 = 96.59)
GPQA Diamond up to 16 majority vote
MMLU-Pro 1 single sample
MMMU-Pro 3 majority vote

Readers comparing numbers should note that leaderboard values are self-reported by each publisher and that settings (sample count, voting, thinking budget) differ between models.

What is inside

The model's parent is Qwen3.8-Flash-Next (180B MoE, Qwen Community License). VIDRAFT says it modified only about 0.02% of the parameters — attention paths and shared experts — while leaving all 512 routed experts, the router and the vision encoder unchanged. The company attributes the result to three in-house techniques:

  • Darwin — a model "diagnose-then-evolve" framework (arXiv:2605.14386)
  • RSI (recursive self-improvement) — the model solves practice problems, its answers are checked against verifiable references, and it is retrained on its correct reasoning. VIDRAFT reports the same accuracy with 11% shorter reasoning than the parent.
  • ZTC (Zero-Token Confidence) — a readout of the model's internal state that estimates, before generation, the probability that an answer will be correct.

Deep dives

  1. Perfect scores on AIME 2026 and HMMT 2026 — why thinking budget mattered
  2. Darwin: evolving a 180B parent model by changing 0.02% of it
  3. Reading an official Hugging Face leaderboard: protocols, majority vote and reproducibility

FAQ

Is the model open? Yes. Weights, config and tokenizer are public on Hugging Face under the Qwen Community License 1.0.

Who is VIDRAFT? A Korean AI deep-tech startup (CEO Minsik Kim) that describes itself as an "AI foundry". Its Darwin model family counts 50+ official models and 400+ community derivatives on Hugging Face.

Can I reproduce the scores? The model card lists the full evaluation protocol; all five benchmarks are public datasets.

Top comments (0)