TL;DR — Darwin-180B-RSI keeps its parent's 512 routed experts, router and vision encoder intact and modifies about 0.02% of the parameters. VIDRAFT's Darwin framework treats an open model as a "parent", diagnoses its weak paths, and strengthens only those — here combined with recursive self-improvement (RSI) on verified answers.
The parent
The father model is Qwen3.8-Flash-Next, a 180B mixture-of-experts vision-language model:
| Layers / hidden size | 48 / 2,560 |
| Attention | 36 linear-attention + 12 full-attention layers |
| Experts | 512 routed (10 active per token) + shared expert |
| Context | 262,144 tokens |
What Darwin changed
According to the model card, Darwin updated three wiring paths and preserved the rest:
- 🔹 Full-attention layers (12) — the long-range path across the whole context
- 🔹 Linear-attention layers (36) — the fast inference path
- 🔹 Shared expert (48 layers) — the common path every token passes through
- 🔒 Preserved: all 512 routed experts, the router and the vision encoder
That is roughly 30 million trainable parameters out of 177 billion. Because the experts — where most factual knowledge is stored in an MoE model — are untouched, the parent's knowledge is retained.
Recursive self-improvement on verified answers
The RSI loop is simple to state:
- The model solves practice problems that do not overlap with any reported benchmark (8-gram overlap filter).
- Its answers are checked against verifiable references.
- Only the correct reasoning is kept, and the model is retrained on it.
- The improved model becomes the next solver.
Two design choices matter. Unverified solutions are never learned, so errors are not reinforced. And among correct solutions, shorter ones are preferred — which is why VIDRAFT reports the same accuracy with 11% shorter reasoning than the parent (MMLU-Pro: 4,320 → 3,833 tokens per answer).
How Darwin compares with evolutionary model merging
The best-known prior work in this area is Sakana AI's evolutionary model merge (Nature Machine Intelligence, 2025), which searches merge recipes over hundreds of generations and demonstrated 7–10B models. Darwin takes a diagnosis-first approach — measuring layer- and expert-level strengths before transplanting or tuning — which VIDRAFT says reduces search cost and has scaled from 4B to 397B models.
On Hugging Face, the Darwin family lists 50+ official models and 400+ community derivatives. The framework is described in Darwin Family (arXiv:2605.14386).
Zero-Token Confidence
VIDRAFT pairs RSI with ZTC, a lightweight readout of the model's final-layer hidden state that estimates, before any token is generated, the probability that the upcoming answer is correct. In an RSI loop, low-confidence areas point to what the model should learn next; in deployment, the same signal can gate tool calls or escalate to a human. The company's earlier Darwin-397B-ZTC ships such a readout.
Links
- Model: FINAL-Bench/Darwin-180B-RSI
- Darwin collection: huggingface.co/collections/FINAL-Bench/darwin-family
Top comments (0)