DEV Community

AI OpenFree
AI OpenFree

Posted on

Darwin: evolving a 180B parent model by changing 0.02% of it

TL;DR — Darwin-180B-RSI keeps its parent's 512 routed experts, router and vision encoder intact and modifies about 0.02% of the parameters. VIDRAFT's Darwin framework treats an open model as a "parent", diagnoses its weak paths, and strengthens only those — here combined with recursive self-improvement (RSI) on verified answers.

The parent

The father model is Qwen3.8-Flash-Next, a 180B mixture-of-experts vision-language model:

Layers / hidden size 48 / 2,560
Attention 36 linear-attention + 12 full-attention layers
Experts 512 routed (10 active per token) + shared expert
Context 262,144 tokens

What Darwin changed

According to the model card, Darwin updated three wiring paths and preserved the rest:

  • 🔹 Full-attention layers (12) — the long-range path across the whole context
  • 🔹 Linear-attention layers (36) — the fast inference path
  • 🔹 Shared expert (48 layers) — the common path every token passes through
  • 🔒 Preserved: all 512 routed experts, the router and the vision encoder

That is roughly 30 million trainable parameters out of 177 billion. Because the experts — where most factual knowledge is stored in an MoE model — are untouched, the parent's knowledge is retained.

Recursive self-improvement on verified answers

The RSI loop is simple to state:

  1. The model solves practice problems that do not overlap with any reported benchmark (8-gram overlap filter).
  2. Its answers are checked against verifiable references.
  3. Only the correct reasoning is kept, and the model is retrained on it.
  4. The improved model becomes the next solver.

Two design choices matter. Unverified solutions are never learned, so errors are not reinforced. And among correct solutions, shorter ones are preferred — which is why VIDRAFT reports the same accuracy with 11% shorter reasoning than the parent (MMLU-Pro: 4,320 → 3,833 tokens per answer).

How Darwin compares with evolutionary model merging

The best-known prior work in this area is Sakana AI's evolutionary model merge (Nature Machine Intelligence, 2025), which searches merge recipes over hundreds of generations and demonstrated 7–10B models. Darwin takes a diagnosis-first approach — measuring layer- and expert-level strengths before transplanting or tuning — which VIDRAFT says reduces search cost and has scaled from 4B to 397B models.

On Hugging Face, the Darwin family lists 50+ official models and 400+ community derivatives. The framework is described in Darwin Family (arXiv:2605.14386).

Zero-Token Confidence

VIDRAFT pairs RSI with ZTC, a lightweight readout of the model's final-layer hidden state that estimates, before any token is generated, the probability that the upcoming answer is correct. In an RSI loop, low-confidence areas point to what the model should learn next; in deployment, the same signal can gate tool calls or escalate to a human. The company's earlier Darwin-397B-ZTC ships such a readout.

Links

Top comments (0)