DEV Community

AI OpenFree
AI OpenFree

Posted on

Darwin-180B-RSI: How VIDRAFT's Recursive Self-Improvement Model Topped 7 Hugging Face Leaderboards — Without Legal Training Data

Darwin-180B-RSI: How VIDRAFT's Recursive Self-Improvement Model Topped 7 Hugging Face Leaderboards — Without Legal Training Data

TL;DR: VIDRAFT's Darwin-180B-RSI is a 180-billion-parameter reasoning model trained via a Recursive Self-Improvement (RSI) framework that updates model weights directly — no human-annotated chain-of-thought required. It just claimed #1 on both the LEXam and LEXam-hard Hugging Face-certified legal reasoning benchmarks, despite never being trained on legal domain data. Its checkpoint is publicly available on Hugging Face for external reproduction.

What it is

Darwin-180B-RSI is VIDRAFT's self-improving large language model, built around a model-level Recursive Self-Improvement (RSI) architecture. Key facts from the announcement:

  • 180 billion parameters in size
  • Achieves top scores across 7 Hugging Face-certified leaderboard categories — the most of any single AI organization across all 47 officially certified Hugging Face benchmark leaderboards at time of publication
  • Covers math, science, general knowledge, visual reasoning, and now legal reasoning (LEXam, LEXam-hard) — the last two domains added without any legal-specific training data
  • Outperforms models up to 685 billion parameters (e.g., DeepSeek-R1) on legal benchmarks despite being less than a third of the parameter count
  • Model checkpoint and evaluation environment are open-sourced on Hugging Face for independent verification

How it works

The core design principle is Recursive Self-Improvement at the weight level — a conceptually distinct approach from system-level self-improvement (like prompt tuning or tool-orchestration loops that leave base weights unchanged).

Here's the high-level loop:

  1. Autonomous problem solving: The model attempts problems from domains where correctness can be verified mechanically — primarily mathematics and science.
  2. Automated verification: A rule-based or formal checker determines whether each candidate solution is correct. No human-written intermediate reasoning traces are injected.
  3. Selective retraining: Only the reasoning paths that led to verified correct answers are used to update the model's weights. Incorrect paths are discarded.
  4. Iterative bootstrapping: The improved model becomes the compute substrate for the next improvement round, creating a recursive loop.

The practical result is that reasoning capabilities generalized from math and science transfer to unseen domains like legal analysis — the model isn't memorizing legal statutes, it's applying structured multi-step logical inference. As a secondary efficiency gain, the RSI-trained model reportedly shortened reasoning token length by 11% relative to the base model, pruning redundant inference steps without sacrificing accuracy.

This is architecturally different from Google's system-level RRSI approach (which adjusts prompts, external tools, and workflows while keeping base weights frozen). VIDRAFT's approach bakes improvements permanently into a single weight file, meaning users get the full capability from the model artifact alone, with no additional scaffolding required at inference time.

Benchmarks & results

All scores below are from the Hugging Face-certified LEXam benchmark, developed jointly by ETH Zurich, University of Zurich, and the Max Planck Institute. LEXam is built from 340 real law school final exams and tests applied legal reasoning under civil law — not statute memorization.

LEXam (multiple-choice, 1,655 questions):

Model Score
Darwin-180B-RSI (4-pass majority vote) 68.94
Darwin-180B-RSI (1-pass, single inference) 60.54
GPT-5 (benchmark authors' measurement) 62.65
Claude-4.5-Sonnet (benchmark authors' measurement) 58.01
Gemini-2.5-Pro (benchmark authors' measurement) 55.72
DeepSeek-R1 (685B) 52.41
Qwen3-235B-Thinking 48.19
gpt-oss-120b 47.71

LEXam-hard (open-ended, 518 questions — scored by official grading model):

Model Score
Darwin-180B-RSI 45.72
Inkling (Thinking Machines) 40.82
DeepSeek-V4-Pro 38.93

Context: random-chance baseline on the 4-option multiple-choice LEXam is 25 points. Most top LLMs average ~50 points on it; LEXam-hard's previous best was in the low 40s.

Across all 47 Hugging Face certified leaderboards, VIDRAFT now holds #1 in 7 categories — ahead of Zhipu AI (4), DeepSeek / Xiaomi / Moonshot AI (3 each).

How to try it

The Darwin-180B-RSI model checkpoint and its evaluation environment are publicly available on Hugging Face as open-source. VIDRAFT has also disclosed the full benchmark evaluation parameters — number of inference passes, ensemble majority-vote settings, and reasoning token budget — to enable independent reproduction.

To pull the model via the Hugging Face CLI:

huggingface-cli download VIDRAFT/Darwin-180B-RSI
Enter fullscreen mode Exit fullscreen mode

⚠️ Verify the exact repository path on huggingface.co/VIDRAFT before downloading. The source confirms open availability but does not specify a direct URL string.

FAQ

Q: Does Darwin-180B-RSI require legal fine-tuning data to perform well on legal benchmarks?
A: No. The model was explicitly evaluated without any prior legal domain training. Its legal reasoning performance comes from generalizing multi-step logical inference skills acquired during math and science self-improvement cycles.

Q: What's the practical difference between VIDRAFT's model-level RSI and system-level self-improvement (like Google's RRSI)?
A: System-level approaches adjust prompts, tools, and orchestration workflows at runtime while leaving the underlying model weights unchanged. VIDRAFT's RSI updates the model's weights directly, so improvements are permanently embedded in the weight file. You don't need any additional runtime infrastructure — just load the model and run inference. VIDRAFT notes both approaches are complementary rather than competing.

Q: How do I know the benchmark results are legitimate and not cherry-picked evaluation conditions?
A: VIDRAFT publicly disclosed all evaluation parameters: inference pass count, majority-vote configuration, and reasoning token budget. The checkpoint is on Hugging Face so anyone can re-run the LEXam evaluation independently. The benchmark itself was built by academic researchers at ETH Zurich, University of Zurich, and the Max Planck Institute, with answers validated by legal professionals.


Originally reported by 로이슈 (2026-10-01) — source article.

Top comments (0)