Darwin-180B-RSI: VIDRAFT's Recursive Self-Improvement MoE Model Claims Five Leaderboard Tops
TL;DR: VIDRAFT, a Korean Pre-AGI AI startup, has released Darwin-180B-RSI — a 180-billion-parameter mixture-of-experts reasoning model built on Qwen3.8-Flash-Next that combines selective model merging with recursive self-improvement. The model reportedly tops five Hugging Face leaderboards across math, science, general, and multimodal reasoning benchmarks. If the reported gains hold on task-specific workloads, it's worth a closer look for reasoning-heavy applications.
What it is
Darwin-180B-RSI is VIDRAFT's open reasoning model, released in late September 2026. Here are the key technical facts as reported:
- Base model: Built on top of Qwen3.8-Flash-Next via selective model merging — making it a high-performance model adaptation rather than a ground-up pretrain
- Architecture: Mixture-of-Experts (MoE) with 180 billion total parameters and 512 experts, of which 10 experts are activated per inference request
- Context window: Approximately 260,000 tokens
- Availability: Described as an open model, with reported leaderboard results hosted on Hugging Face
The sparse expert activation design (10 of 512 experts per token routing) means that despite the 180B parameter count, the effective compute per forward pass is substantially lower than a dense 180B model — a meaningful practical consideration for deployment cost and latency.
How it works
VIDRAFT describes Darwin-180B-RSI as combining two high-level techniques: selective model merging and recursive self-improvement (RSI).
At a conceptual level, the pipeline works roughly like this:
- Model merging as initialization: Rather than training from scratch, VIDRAFT merges capabilities from existing checkpoints selectively — preserving strengths from the base model while injecting targeted improvements.
- Recursive self-improvement loop: The model generates candidate solutions to reasoning problems. Those solutions are automatically evaluated against verifiable ground-truth answers. Only solutions with confirmed correct answers are retained and used to retrain the model in the next iteration.
- Iteration: This generate → filter → retrain cycle repeats, theoretically allowing the model to bootstrap higher reasoning quality from its own verified outputs over multiple passes.
This approach is conceptually related to ideas like rejection sampling fine-tuning and self-play reinforcement learning, but VIDRAFT's specific implementation details — including iteration count, reward modeling, and filtering criteria — are not disclosed in the source reporting.
The MoE routing means that at inference time, each token activates only a small subset of the expert network. This is a well-established technique (see Mixtral, DeepSeek-MoE) for scaling parameter count while keeping per-token FLOPs tractable.
Benchmarks & results
VIDRAFT reports the following scores on publicly recognized benchmarks, with Darwin-180B-RSI claiming first place on five Hugging Face leaderboards:
| Benchmark | Reported Score | Domain |
|---|---|---|
| AIME 2026 | 100% (perfect) | Mathematical reasoning |
| HMMT 2026 | 100% (perfect) | Mathematical competition |
| GPQA Diamond | 94.44% | Graduate-level science Q&A |
| MMLU-Pro | 88.12% | Multi-domain professional knowledge |
| MMMU-Pro | 79.48% | Multimodal understanding & reasoning |
⚠️ Important caveat: These numbers are self-reported by VIDRAFT. The source article explicitly notes that no independent third-party validation of these results has been performed. Leaderboard rankings reflect performance on fixed benchmark datasets and may not generalize to production workloads, novel problem distributions, or domain-specific tasks. Engineers evaluating this model should run it against their own held-out test sets before drawing conclusions about real-world reasoning quality.
How to try it
The source reports Darwin-180B-RSI is an open model available on Hugging Face. If VIDRAFT follows its prior release pattern (the company had a prior Hugging Face feature covered in September), the model weights should be downloadable via the standard Hugging Face toolchain.
To check availability and download weights once the repository is public:
# Search for the model on Hugging Face
huggingface-cli search vidraft/darwin-180b-rsi
# Download when available
huggingface-cli download vidraft/darwin-180b-rsi
Note: The exact repository path is not confirmed in the source. Search for
VIDRAFTorDarwin-180B-RSIdirectly on huggingface.co to find the correct model card and any associated inference instructions.
No GitHub repository, API endpoint, or OpenAI-compatible serving URL is mentioned in the source reporting. Check VIDRAFT's Hugging Face profile and official channels for inference API access details.
FAQ
Q: Is Darwin-180B-RSI practical to self-host given its 180B parameter count?
A: The MoE architecture activates only 10 of 512 experts per request, so active-parameter FLOPs are much lower than a dense 180B model. That said, you still need sufficient memory to load all expert weights. Quantized or offloaded serving strategies (e.g., llama.cpp, vLLM with expert offloading) may be necessary depending on your hardware. Check the model card for official serving recommendations once it is published.
Q: How does the recursive self-improvement approach differ from standard RLHF or SFT?
A: Standard supervised fine-tuning (SFT) trains on human-labeled data; RLHF uses a human preference reward model. VIDRAFT's RSI loop instead uses the model's own outputs filtered by verifiable correctness — closer in spirit to rejection sampling fine-tuning or outcome-reward RL (like GRPO or STaR), but without requiring a separate reward model or human annotators for each iteration. The practical implication is that it scales more easily to domains with automated answer verification, like math and formal science.
Q: Should I trust the perfect AIME and HMMT scores at face value?
A: Treat them as a signal worth investigating, not a guarantee. Perfect scores on competition math benchmarks from self-reported results warrant independent replication. Run the model on problems from your domain and measure it yourself before committing to it for production use cases.
Originally reported by AI Market Watch (미국) (2026-09-28) — source article.
Top comments (0)