DEV Community

AI OpenFree
AI OpenFree

Posted on

Perfect scores on AIME 2026 and HMMT 2026: what an open 180B model got right, and why thinking budget mattered

TL;DR — Darwin-180B-RSI, an open-weight model from Korean startup VIDRAFT, reports 100% on AIME 2026 (30/30) and 100% on HMMT February 2026 (33/33) under 16-sample majority vote with a 131K-token thinking budget. They are the first perfect scores on these two Hugging Face official leaderboards.

The numbers

Benchmark Majority of 16 Mean over 16 samples Single sample Previous #1
AIME 2026 100.0 98.75 96.67 97.1
HMMT Feb 2026 100.0 96.59 93.94 92.7

Even the mean accuracy over 16 samples (98.75 on AIME, 96.59 on HMMT) is above the previous top entries on both boards.

About the two competitions

  • AIME (American Invitational Mathematics Examination) is the qualifier toward the USA Mathematical Olympiad. Every answer is an integer from 0 to 999, so guessing rarely works.
  • HMMT (Harvard-MIT Mathematics Tournament) is one of the most difficult high-school competitions, attended by national-olympiad-level students. Answers include fractions and radicals, so grading requires symbolic equivalence (for example, 2√3 = √12).

The finding: misses were truncations, not wrong math

VIDRAFT's engineers first measured the parent model with a 65K-token budget. The model missed two AIME 2026 problems, and on inspection every miss was a truncated solution — the reasoning ran out of tokens before reaching an answer. Solutions that finished were correct.

Re-running those problems with a 131K-token budget produced 8/8 correct answers. With the budget raised for all problems, the 16-sample run scored 30/30.

The practical lesson for anyone evaluating reasoning models:

  1. Truncation is not an error. Count cut-off answers separately; otherwise a budget limit looks like a capability limit.
  2. Budget is part of the protocol. A score without its thinking budget is not comparable.
  3. Grade symbolically. String matching undercounts fractions and radicals; a math-equivalence checker is required.

How to read "majority of 16"

Majority vote is a system score: the model answers each question 16 times and the most frequent answer is graded. VIDRAFT reports the majority score and the mean single-sample accuracy side by side on the model card and in the leaderboard submission notes, so readers can compare with entries that use a single sample.

Try it

Model: huggingface.co/FINAL-Bench/Darwin-180B-RSI — the model card contains the full evaluation protocol.

Top comments (0)