Darwin-180B-RSI Tops Global Legal Reasoning Benchmarks — Without Ever Training on Legal Data
TL;DR: VIDRAFT's Darwin-180B-RSI reasoning model has claimed #1 on both the LEXam and LEXam-hard Hugging Face-certified legal reasoning benchmarks, despite never being trained on legal data. It achieves this through Recursive Self-Improvement (RSI) — a self-supervised loop driven entirely by math and science problems. Developers working in legal tech, public sector AI, or high-stakes reasoning applications should take note.
What it is
Darwin-180B-RSI is a 180-billion-parameter reasoning model developed by VIDRAFT (비드래프트), a Korean AI startup. Its headline result: #1 on both LEXam and LEXam-hard, two Hugging Face-certified legal reasoning benchmarks developed by ETH Zurich and collaborators, without the model ever having been fine-tuned on legal corpora.
The Darwin family of models is built on a foundational approach that merges and combines the strengths of existing pre-trained models rather than training from scratch — a design philosophy aimed at dramatically reducing the compute cost typically associated with large-scale pre-training.
The -RSI suffix designates models that have additionally undergone VIDRAFT's Recursive Self-Improvement training procedure (described below).
How it works
Darwin-180B-RSI's performance on legal reasoning is the product of two conceptually distinct techniques applied on top of the base Darwin architecture:
1. Recursive Self-Improvement (RSI)
Instead of relying on human-authored chain-of-thought annotations, the model enters a self-supervised loop:
- The model attempts to solve problems (sourced from math and science domains only).
- An automated verifier checks whether the model's solution is correct.
- Solutions that pass verification are fed back as training signal.
- This cycle repeats, letting the model iteratively refine its own reasoning capability without human labeling.
The key insight is that general reasoning ability — not domain-specific knowledge — is what transfers to legal exam performance. Legal reasoning requires applying rules to facts and drawing conclusions, a structure that maps closely to rigorous mathematical reasoning.
2. Zero-Token Confidence (ZTC)
ZTC is VIDRAFT's technique for estimating the model's internal confidence in a given answer without requiring extra inference tokens. This is used during the self-improvement loop to help the model gauge answer reliability, and it contributes to a reported 11% reduction in average reasoning chain length relative to the parent model — meaning fewer tokens consumed at inference time with no accuracy regression.
The base Darwin architecture itself achieves cost efficiency by merging and composing existing model checkpoints rather than running a full pre-training run from scratch. Only 0.02% of total parameters are modified during the self-improvement phase.
Benchmarks & results
All scores below are sourced directly from the AI타임스 report and reflect results on Hugging Face-certified leaderboards.
LEXam (multiple-choice legal reasoning):
| Model | Score |
|---|---|
| Darwin-180B-RSI | 68.94 |
| GPT-5 | 62.65 |
| Claude 4.5 Sonnet | 58.01 |
| DeepSeek-R1 | 52.41 |
| Qwen-235B-Thinking | 48.19 |
| gpt-oss-120b | 47.71 |
LEXam-hard (open-ended / essay-style legal reasoning):
| Model | Score |
|---|---|
| Darwin-180B-RSI | 45.72 |
| Inkling (Thinking Machines) — previous #1 | 40.82 |
Darwin-180B-RSI's 180B parameter count is notably smaller than DeepSeek-R1's 685B, yet it substantially outperforms it on both benchmarks.
Broader leaderboard standing: With these results, VIDRAFT now holds #1 positions across 7 of 47 categories on Hugging Face-certified leaderboards, spanning math, science, general knowledge, visual reasoning, and legal reasoning — the highest count among Korean AI companies, and ahead of Zhipu AI (4), DeepSeek, Xiaomi, and Moonshot AI (3 each).
LEXam benchmark context: LEXam is based on real university law school exams and is designed to test the ability to apply statutes to factual scenarios and derive legal conclusions — not rote memorization of law. This makes it a meaningful proxy for practical legal reasoning ability.
How to try it
The article states that the model and evaluation conditions are publicly available on Hugging Face for independent verification. You can explore VIDRAFT's public presence at:
- 🤗 Hugging Face: search for
VIDRAFTorDarwin-180B-RSIon huggingface.co
At the time of writing, no public GitHub repository URL, API endpoint, or install command has been announced in this report. Check VIDRAFT's official channels for the latest on API access or model weights availability.
FAQ
Q: How can a model with zero legal training data outperform models that presumably have seen legal text?
A: The results suggest that LEXam is testing structured reasoning — applying rules to facts to reach conclusions — rather than legal memorization. RSI trained on math and science problems appears to build exactly that kind of generalizable deductive capability, which transfers surprisingly well to legal exam formats.
Q: Is Darwin-180B-RSI smaller or larger than the models it beat?
A: Smaller in some key comparisons. At 180B parameters it outperforms DeepSeek-R1 (685B parameters) on both LEXam benchmarks. Parameter count is clearly not the dominant factor here.
Q: What domains is VIDRAFT targeting with this technology?
A: According to CEO Minsik Kim, the target verticals are legal services, public administration, and financial services — domains where accurate, auditable judgment under uncertainty is critical.
Q: Does ZTC mean faster inference?
A: ZTC contributed to an 11% reduction in reasoning chain length compared to the parent model, which translates directly to lower token consumption at inference time while maintaining parent-model-level accuracy.
Originally reported by AI타임스 (2026-10-01) — source article.
Top comments (0)