DEV Community

QuantID
QuantID

Posted on Originally published at quantid-measurement-lab.static.hf.space

Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions

TL;DR

A single self-improving model family holds eleven public number-one benchmark records at once, across mathematics, science, law, structured output, and decisions. This is a measurement note: the eleven, their scores, and why the spread matters more than any single win.

The eleven records

# Benchmark Field Result
1 AIME 2026 Math 100% (perfect)
2 HMMT 2026 Math 100% (perfect)
3 GPQA Diamond Science 94.44%
4 MMLU-Pro Knowledge 88.12%
5 MMMU-Pro Multimodal 79.48%
6 LEXam Law 68.94%
7 LEXam-hard Law 45.72%
8 ExtractBench Extraction 90.29%
9 IFStruct Structured output 98.95%
10 MDPBench Decision process 83.65%
11 S1MB Decision engine Borda 89.58, Task Avg 66.46

Why the spread is the real signal

A single top score can come from tuning to one test. Eleven number-one records across unrelated fields cannot. Math, science, law, and typed decisions stress different capabilities, so holding the top of all of them at once is evidence of a general method rather than a lucky fit.

Two of the records are perfect scores, on AIME 2026 and HMMT 2026. When a benchmark saturates, the honest reading is that it no longer discriminates at the top, which is itself a measurement result and a reason to move evaluation toward harder, decision-oriented tests such as S1MB.

How these were produced

The family is self-improving: it runs a recursive loop tied to external verification rather than to its own opinion. For decisions, where the output is typed, scoring uses a zero-token method, which keeps the measurement deterministic and reproducible.

Check it yourself

FAQ

How many number-one records does the family hold?
Eleven public number-one benchmark records at once, across math, science, law, structured output, and decisions.

Which are perfect scores?
AIME 2026 and HMMT 2026, both at 100%.

Why does holding many unrelated benchmarks matter?
A single win can reflect tuning to one test. Leading many unrelated fields at once is evidence of a general method.

Are the results reproducible?
Decision results use a deterministic zero-token method, and the leaderboards and models are public.

Top comments (0)