DEV Community

GINIGEN AI
GINIGEN AI

Posted on Originally published at ginigen-ai-edge-ai-news.static.hf.space

One Self-Improving Model, Eleven Number-One Titles: What That Takes

TL;DR

One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions. The interesting part is not any single score, it is that a recursive self-improvement loop, bound to external verification, can push a whole spread of benchmarks to the top at once.

The eleven, at a glance

  • Math: AIME 2026 100%, HMMT 2026 100%
  • Science: GPQA Diamond 94.44%
  • Knowledge and multimodal: MMLU-Pro 88.12%, MMMU-Pro 79.48%
  • Law: LEXam 68.94%, LEXam-hard 45.72%
  • Extraction and structure: ExtractBench 90.29%, IFStruct 98.95%
  • Decisions: MDPBench 83.65%, S1MB Borda 89.58

Why a self-improving loop gets here

A model that only trains once is stuck with the data it saw. A self-improving (RSI) model runs a loop: it attempts problems, keeps the attempts that pass an external check, and trains on those. The reward is tied to verification outside the model, such as code execution, external evidence, or an answer key, not to the model grading itself. That is what keeps the loop from drifting into its own confident mistakes.

When the loop is honest, improvement on one reasoning skill tends to carry to others, which is why the records span unrelated fields instead of one narrow test.

The decision layer

For typed decisions, the family uses a zero-token method: it reads the problem in one forward pass and applies a calibrated probe to produce the decision, generating no text. That is deterministic and cheap, which also makes it a good fit for running decisions close to where the data is.

Open to use

FAQ

How many number-one records, and across what?
Eleven at once, across math, science, law, structured output, and decisions.

What is a self-improving (RSI) model?
A model that runs a loop: it solves problems, keeps the solutions that pass an external check, and trains on them, so it improves without new human labels.

Why tie the loop to external verification?
If a model grades its own work, the loop can reinforce confident errors. External checks keep improvement real.

Is it open?
Yes, the decision model and method are public under Apache-2.0.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.