VIDRAFT Darwin-180B-RSI Tops 5 Hugging Face Leaderboards: AIME & HMMT Perfect Scores, 94.44% on GPQA Diamond
TL;DR: VIDRAFT's Darwin-180B-RSI is a 180B-parameter reasoning model that achieves #1 rankings across five Hugging Face-certified leaderboards — including the first-ever perfect scores on AIME 2026 and HMMT 2026 benchmarks. It combines three proprietary techniques (Darwin selective merging, Recursive Self-Improvement, and Zero-Token Confidence) to push frontier-level math, science, and multimodal reasoning. If you're evaluating reasoning models for hard STEM workloads, this is a result worth tracking.
What it is
Darwin-180B-RSI is VIDRAFT's latest reasoning-focused large language model. The "180B" refers to its parameter count, and "RSI" stands for Recursive Self-Improvement — a core training methodology baked into the model's name.
Key characteristics from the source:
- Scale: 180 billion parameters (VIDRAFT has separately extended the same Darwin merging technology to a ~397-billion-parameter model class)
- Focus: High-difficulty reasoning across mathematics, science, and multimodal domains
- Certification: Rankings verified on Hugging Face's officially certified leaderboards — not self-reported numbers
- Efficiency signal: Achieves the same accuracy as prior iterations while reducing inference reasoning length by 11%
How it works
Darwin-180B-RSI is built on three proprietary technologies that VIDRAFT describes at a conceptual level:
1. Darwin — Selective Model Merging
Darwin analyzes performance at the layer and expert (MoE-style) unit level across multiple source models, then selectively identifies and combines the strongest components from each. Rather than averaging model weights indiscriminately, it surgically extracts high-performing substructures to produce a merged model that inherits complementary strengths. VIDRAFT has scaled this technique up to the ~397B parameter range.
2. RSI — Recursive Self-Improvement
RSI is a training loop in which the model:
- Attempts to solve problems autonomously
- Verifies whether its own answers are correct
- Trains on the verified, high-quality reasoning traces
The improved model then becomes the solver for the next iteration of this loop — a self-bootstrapping cycle that incrementally raises reasoning quality without requiring human-labeled solutions at each step.
3. ZTC — Zero-Token Confidence
Before generating a response, ZTC analyzes the model's internal state to compute a confidence estimate for the candidate answer. If confidence falls below a threshold, the model is designed to withhold its response rather than produce a low-confidence output. This is a calibration mechanism aimed at reducing hallucination by giving the model a principled "abstain" option.
Benchmarks & results
All figures below are from Hugging Face-certified leaderboards as reported by VIDRAFT on 2026-09-28:
| Benchmark | Darwin-180B-RSI | Previous SOTA (on that leaderboard) | Notes |
|---|---|---|---|
| AIME 2026 | 100% | 97.1% | Math Olympiad-level problems; first perfect score on this leaderboard |
| HMMT 2026 | 100% | 92.7% | Math Olympiad-level problems; first perfect score on this leaderboard |
| GPQA Diamond | 94.44% | 93.5% (Kimi-K3) | PhD-level expert-authored science questions |
| MMLU-Pro | 88.12% | — | 14 professional domains incl. law, medicine, economics |
| MMMU-Pro | 79.48% | — | Multimodal reasoning: charts, sheet music, medical imaging |
Darwin-180B-RSI ranked #1 on all five leaderboards simultaneously. The AIME and HMMT perfect scores are noted as the first time any model has achieved 100% on those Hugging Face-certified boards.
How to try it
The source article does not include explicit Hugging Face model IDs, GitHub repository links, or API endpoint details for Darwin-180B-RSI at the time of publication. Access details have not been publicly confirmed in this report.
What is publicly known from related VIDRAFT coverage:
- VIDRAFT models have been distributed via Hugging Face (the company has surpassed 1.6 million model downloads on the platform as of September 2026)
- VIDRAFT has previously open-released evaluation datasets and diagnostic tooling (AX-RAY) on Hugging Face
To follow release announcements, monitor VIDRAFT's Hugging Face organization page and their official channels directly. As soon as a model card or repository goes public, standard access patterns would apply.
FAQ
Q: Are the benchmark results independently verified or self-reported?
A: VIDRAFT explicitly states these rankings come from Hugging Face's certified leaderboards — meaning the evaluation pipeline is managed by Hugging Face, not run in-house by VIDRAFT. That's a meaningful distinction from self-reported numbers.
Q: What makes the 11% reasoning-length reduction significant?
A: Shorter reasoning chains at equivalent accuracy directly reduce inference compute costs and latency. For production deployments where you're paying per token or running latency-sensitive workloads, a model that reasons more concisely without sacrificing correctness is practically valuable — not just a benchmark curiosity.
Q: How does Darwin differ from standard model merging techniques like SLERP or TIES-merging?
A: Standard merging methods operate on full weight tensors with relatively coarse granularity. Darwin's distinguishing claim is that it evaluates and selects at the layer and expert unit level, enabling more surgical recombination of specialized capabilities. The specific algorithmic details beyond this conceptual description have not been publicly disclosed.
Q: Is ZTC related to existing confidence calibration or selective prediction research?
A: Conceptually it occupies similar territory — the model abstains when uncertain rather than guessing. VIDRAFT's specific implementation approach has not been detailed in public materials beyond the mechanism described above.
Originally reported by 지디넷코리아 (2026-09-28) — source article.
Top comments (0)