Darwin-180B-RSI, an open-weight 180B model from the Korean startup VIDRAFT, is now #1 on LEXam and LEXam-hard, the two legal-reasoning leaderboards that Hugging Face lists as official benchmarks. The model was never trained on legal data.
With these two, the model holds first place on seven Hugging Face official leaderboards, the most of any organization on the Hub right now (Zhipu AI has 4; DeepSeek, Xiaomi and Moonshot AI have 3 each).
The benchmark: law-school exams, not trivia
LEXam was built by researchers at ETH Zurich, the University of Zurich and the Max Planck Institute, from 340 real law-school exams at Swiss universities, in German and English, with answers verified by legal experts. It doesn't test whether a model has memorized statutes. It tests whether the model can apply law to a set of facts and reason to a conclusion.
It is hard. The multiple-choice set has four options, so guessing gives 25 points, yet the strongest models on the leaderboard sit around 50. LEXam-hard is a subset of the open-ended questions that the strongest open models score lowest on.
The results
| Leaderboard | Darwin-180B-RSI | Previous #1 |
|---|---|---|
| LEXam (MCQ, 1,655 questions) | 68.94 | 52.41 (DeepSeek-R1) |
| LEXam-hard (open-ended, 518 questions) | 45.72 | 40.82 (Inkling) |
Against the results the LEXam authors published themselves: GPT-5 62.65, Claude-4.5-Sonnet 58.01, Gemini-2.5-Pro 55.72. Darwin-180B-RSI scores 68.94 with a majority vote over 4 samples, and 60.54 with a single sample, which is still above Claude-4.5-Sonnet and Gemini-2.5-Pro.
How we measured: LEXam uses 4 samples per question and a majority vote, with a 32K-token thinking budget. LEXam-hard uses a single sample; the 60 answers that hit the 32K limit were regenerated with 120K. LEXam-hard answers were graded by DeepSeek-R1-0528, the judge named in the benchmark's official eval.yaml. Every protocol is on the model card.
How the model learned: model-level RSI
The interesting part is not the rank but how the model got there.
Darwin-180B-RSI uses model-level recursive self-improvement:
- The model solves practice problems.
- Each answer is checked automatically.
- The model trains only on its own solutions that turned out correct.
- The improved model becomes the solver for the next round.
The practice problems were math and science only, chosen because their answers can be checked automatically. No human-written solutions or reasoning traces were used. No legal text was in the training set.
So why does it do well on law? Our reading is that self-improvement on verifiable problems trains a habit of reasoning rather than knowledge: break the problem down, check each step, commit to a conclusion. That habit transfers to a domain the model never studied. We can't prove the transfer is caused by RSI alone, since we haven't measured the parent model on LEXam under the same protocol. One effect we did measure: the model reaches the same accuracy as its parent with about 11% shorter reasoning.
Model-level vs. harness-level RSI
Google recently released RRSI, which keeps the model frozen and lets an AI improve the prompts, tools and workflow around it. That is harness-level RSI. Darwin-180B-RSI is model-level: the weights change, so the gain ships inside the model file and works for anyone who downloads it, with no harness required. The two are complementary.
What's under the hood
- Darwin evolves models instead of pretraining them from scratch. It diagnoses a parent model layer by layer and expert by expert, like an MRI, transplants the strongest parts from several models, and fuses them weighted by diagnostic trust. The Darwin family has 50+ official models and 400+ community derivatives (arXiv 2605.14386).
- Darwin-180B-RSI changes only about 0.02% of the parent's parameters (attention paths and shared experts). All 512 routed experts, the router and the vision encoder are unchanged.
- ZTC (Zero-Token Confidence) reads the model's internal state once, before generation, to estimate the probability that an answer will be correct.
All seven #1 results
- Math: AIME 2026 (100), HMMT Feb 2026 (100)
- Science: GPQA Diamond (94.44)
- General knowledge: MMLU-Pro (88.12)
- Visual reasoning: MMMU-Pro (79.48)
- Law: LEXam (68.94), LEXam-hard (45.72)
Weights and evaluation settings: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
Top comments (0)