TL;DR
A single self-improving model family holds eleven public number-one benchmark records at once, across mathematics, science, law, structured output, and decisions. This is a measurement note: the eleven, their scores, and why the spread matters more than any single win.
The eleven records
| # | Benchmark | Field | Result |
|---|---|---|---|
| 1 | AIME 2026 | Math | 100% (perfect) |
| 2 | HMMT 2026 | Math | 100% (perfect) |
| 3 | GPQA Diamond | Science | 94.44% |
| 4 | MMLU-Pro | Knowledge | 88.12% |
| 5 | MMMU-Pro | Multimodal | 79.48% |
| 6 | LEXam | Law | 68.94% |
| 7 | LEXam-hard | Law | 45.72% |
| 8 | ExtractBench | Extraction | 90.29% |
| 9 | IFStruct | Structured output | 98.95% |
| 10 | MDPBench | Decision process | 83.65% |
| 11 | S1MB | Decision engine | Borda 89.58, Task Avg 66.46 |
Why the spread is the real signal
A single top score can come from tuning to one test. Eleven number-one records across unrelated fields cannot. Math, science, law, and typed decisions stress different capabilities, so holding the top of all of them at once is evidence of a general method rather than a lucky fit.
Two of the records are perfect scores, on AIME 2026 and HMMT 2026. When a benchmark saturates, the honest reading is that it no longer discriminates at the top, which is itself a measurement result and a reason to move evaluation toward harder, decision-oriented tests such as S1MB.
How these were produced
The family is self-improving: it runs a recursive loop tied to external verification rather than to its own opinion. For decisions, where the output is typed, scoring uses a zero-token method, which keeps the measurement deterministic and reproducible.
Check it yourself
- S1MB number one model: https://github.com/final-bench/s1mb
- ZTC decision method: https://github.com/final-bench/ztc
- Models: https://huggingface.co/FINAL-Bench
FAQ
How many number-one records does the family hold?
Eleven public number-one benchmark records at once, across math, science, law, structured output, and decisions.
Which are perfect scores?
AIME 2026 and HMMT 2026, both at 100%.
Why does holding many unrelated benchmarks matter?
A single win can reflect tuning to one test. Leading many unrelated fields at once is evidence of a general method.
Are the results reproducible?
Decision results use a deterministic zero-token method, and the leaderboards and models are public.
Top comments (0)