Running a frontier-class model used to mean a rack of data-center GPUs. Not anymore.
VIDRAFT has released POCKET-Darwin-180B, a compressed build of Darwin-180B-RSI, the model that holds first place on seven Hugging Face official leaderboards (self-reported), that runs without a GPU. It is being released on both Hugging Face and ModelScope, Alibaba's AI model platform and China's largest; the download pages below open with the release.
- Size: 111 GB (4-bit GGUF), down from 360 GB (BF16)
- CPU only: a single server CPU (16 threads) generates 18.4–21.0 tokens/s, peak memory 78.8 GB
- Laptop: RTX 5060 Laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s
- Mini PC: a 128 GB-RAM mini PC holds the whole model in memory, no GPU required
- Same accuracy: MMLU-Pro, 2,000 questions, paired comparison: 87.65% vs 87.65%
From an 8-GPU server to a laptop
| Darwin-180B-RSI (BF16) | POCKET-Darwin-180B | |
|---|---|---|
| Model size | 360 GB (131 files) | 111 GB (4 files) |
| Hardware | 4–8 NVIDIA B200, or 8× H100 (80 GB) | Laptop with 8 GB GPU + 32 GB RAM, 128 GB mini PC, CPU-only server, one DGX Spark |
| MMLU-Pro | 87.65% | 87.65% |
An 8× H100 server is commonly estimated at around US$350K. A gaming laptop that runs POCKET costs roughly 1/250 of that.
Why a 180B model fits on a laptop
Darwin-180B is a Mixture-of-Experts model: of 512 experts, only 10 are used per token, so only about 3 billion of its 180 billion parameters are active at any moment. llama.cpp memory-maps the weights and reads the experts it needs directly from the SSD, so the full 111 GB never has to sit in RAM.
The compression itself is a graft: we took the widely used Unsloth UD-Q4_K_XL layout of the base model as a template and replaced only the 300 tensors that our self-improvement training changed (in Q8_0). Every other byte is identical to the base build, and all 300 replaced tensors passed read-back verification.
The model underneath: Model-level RSI
Darwin-180B-RSI is built on Qwen3.8-Flash-Next. It improves through Model-level Recursive Self-Improvement: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used. Only the attention paths and shared experts are trained; the 512 routed experts and the router are left untouched.
The R3 round compressed here improved over the previous round on 1,000 held-out SuperGPQA questions never used for training or selection (+1.03 points, 95% CI [+0.05, +2.00]).
Leaderboard results of the original model (self-reported, majority voting):
| Benchmark | Score |
|---|---|
| AIME 2026 | 100 |
| HMMT Feb 2026 | 100 |
| GPQA Diamond | 94.44 |
| MMLU-Pro | 88.12 |
| MMMU-Pro | 79.48 |
| LEXam (law) | 68.94 |
| LEXam-hard (law) | 45.72 |
Verified as a research institution on ModelScope
VIDRAFT's organization on ModelScope (FINAL-Bench) carries an official Research Institution verification. Among the organizations we checked, verification is held by only a few, including Alibaba's Qwen, Shanghai AI Laboratory, Zhipu AI and OpenBMB, and VIDRAFT is the only Korean AI organization we found with it.
Who this is for
Organizations that cannot send data to an external cloud (defense, finance, public sector) can now run a top-tier model on an air-gapped server or a single mini PC.
Run it
Requires llama.cpp b11048 or newer.
# Laptop / desktop (8 GB+ VRAM, 32 GB+ RAM; experts streamed from SSD)
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 --cpu-moe -fa on -c 8192 --jinja
# CPU only
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 0 -t 16 -c 8192 --jinja --load-mode none
Recommended sampling: temperature 1.0, top-p 0.95, top-k 20. It is a reasoning model, so leave room for at least 2,048 output tokens.
- Hugging Face: https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF
- ModelScope: https://modelscope.cn/models/FINAL-Bench/POCKET-Darwin-180B-GGUF
- Original model: https://huggingface.co/FINAL-Bench/Darwin-180B-RSI
Top comments (0)