TL;DR — Darwin-180B-RSI, an open-weight model from Korean startup VIDRAFT, reports 100% on AIME 2026 (30/30) and 100% on HMMT February 2026 (33/33) under 16-sample majority vote with a 131K-token thinking budget. They are the first perfect scores on these two Hugging Face official leaderboards.
The numbers
| Benchmark | Majority of 16 | Mean over 16 samples | Single sample | Previous #1 |
|---|---|---|---|---|
| AIME 2026 | 100.0 | 98.75 | 96.67 | 97.1 |
| HMMT Feb 2026 | 100.0 | 96.59 | 93.94 | 92.7 |
Even the mean accuracy over 16 samples (98.75 on AIME, 96.59 on HMMT) is above the previous top entries on both boards.
About the two competitions
- AIME (American Invitational Mathematics Examination) is the qualifier toward the USA Mathematical Olympiad. Every answer is an integer from 0 to 999, so guessing rarely works.
- HMMT (Harvard-MIT Mathematics Tournament) is one of the most difficult high-school competitions, attended by national-olympiad-level students. Answers include fractions and radicals, so grading requires symbolic equivalence (for example, 2√3 = √12).
The finding: misses were truncations, not wrong math
VIDRAFT's engineers first measured the parent model with a 65K-token budget. The model missed two AIME 2026 problems, and on inspection every miss was a truncated solution — the reasoning ran out of tokens before reaching an answer. Solutions that finished were correct.
Re-running those problems with a 131K-token budget produced 8/8 correct answers. With the budget raised for all problems, the 16-sample run scored 30/30.
The practical lesson for anyone evaluating reasoning models:
- Truncation is not an error. Count cut-off answers separately; otherwise a budget limit looks like a capability limit.
- Budget is part of the protocol. A score without its thinking budget is not comparable.
- Grade symbolically. String matching undercounts fractions and radicals; a math-equivalence checker is required.
How to read "majority of 16"
Majority vote is a system score: the model answers each question 16 times and the most frequent answer is graded. VIDRAFT reports the majority score and the mean single-sample accuracy side by side on the model card and in the leaderboard submission notes, so readers can compare with entries that use a single sample.
Try it
Model: huggingface.co/FINAL-Bench/Darwin-180B-RSI — the model card contains the full evaluation protocol.
Top comments (0)