DEV Community

Cover image for Running a 180B Model on a Laptop With No GPU: How 4-bit GGUF Keeps Full Accuracy
GINIGEN AI
GINIGEN AI

Posted on

Running a 180B Model on a Laptop With No GPU: How 4-bit GGUF Keeps Full Accuracy

A frontier-class model used to mean a rack of data-center GPUs. That assumption is what this post takes apart.

POCKET-Darwin-180B-GGUF is a 4-bit build of Darwin-180B-RSI, a 180-billion-parameter model, packaged so it runs without a GPU. It ships as GGUF on Hugging Face and on ModelScope. The headline numbers, all measured in October 2026:

  • Size: 111 GB (4-bit GGUF, 4 files), down from 360 GB (BF16, 131 files)
  • CPU only: a single server CPU (16 threads) generates 18.4 to 21.0 tokens/s, peak memory 78.8 GB
  • Laptop: an RTX 5060 Laptop (8 GB VRAM) with 32 GB RAM runs it at 4.17 tokens/s
  • Mini PC: a 128 GB-RAM mini PC holds the whole model in memory, no GPU
  • Accuracy: MMLU-Pro, 2,000 questions, paired comparison, 87.65% original vs 87.65% quantized

TL;DR

A 180B mixture-of-experts model activates only about 3B parameters per token. Combine that sparsity with a 4-bit quantization that keeps the small fraction of the weights that actually changed during training at high precision, and you get a model that fits in laptop-class memory and still answers like the full model. You run it with llama.cpp in one command. This post explains why it works and how to reproduce it.

How does a 180B model fit on a laptop?

Two things make this possible, and neither is magic.

First, the architecture is sparse. Darwin-180B is a mixture-of-experts (MoE) model: 512 expert sub-networks, of which only 10 are selected per token. Of the 180B total parameters, roughly 3B are active for any given token. The dense parts (attention, routing, embeddings) run every step; the expert weights are read on demand. llama.cpp can memory-map the file and let the operating system page expert tensors in and out, so you do not need all 111 GB resident at once. On a 128 GB mini PC the whole thing sits in RAM; on a 32 GB laptop the OS pages against the GGUF file on disk, which is why the laptop is slower but still runs.

Second, the quantization is selective. The build starts from a community-verified 4-bit quantization of the base weights. On top of that, only the portion that the self-improvement training actually modified (about 3% of total volume) is kept at higher precision. You spend your bit budget where the signal is, not uniformly. That is the difference between "4-bit and lossy" and "4-bit and lossless on the benchmark."

Does 4-bit quantization hurt accuracy here?

On the measured benchmark, no. The honest way to check quantization damage is a paired comparison: same questions, same decoding, original versus quantized, item by item. On MMLU-Pro (2,000 questions) both the original and the 4-bit build scored 87.65%. Identical.

That result is specific and worth reading carefully. It does not mean 4-bit is free in general. It means that for this model, on this benchmark, with a selective scheme that protects the trained delta, the drop is below the measurement's resolution. Quantization error concentrates in the weights that carry the most task-relevant signal, so protecting that 3% is what preserves the score.

There is a second, subtler result. On SuperGPQA (1,000 graduate-level science questions, never used in training), the quantized POCKET build was compared against the same 4-bit quantization of its base model. POCKET scored 61.55% vs 59.10%, a 2.45-point gain that held as statistically significant, while using about 13% fewer tokens per question. The self-improvement training taught the model to commit to an answer instead of re-deriving it, and quantization preserved that behavior. A third-party developer reproduced this at the file level in a public thread.

How do you run POCKET-Darwin-180B-GGUF?

You need a recent llama.cpp build (b11048 or later), the four GGUF files (111 GB total), and enough memory to either hold or page the model. No cloud account, no GPU driver stack required for the CPU path.

# 1. Build llama.cpp (b11048 or later)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j

# 2. Download the GGUF (4 files, ~111 GB) from Hugging Face
#    FINAL-Bench/POCKET-Darwin-180B-GGUF
#    (any standard huggingface download works; place all parts
#     in one directory so the loader finds the shards)

# 3. Run, CPU only, 16 threads
./build/bin/llama-cli \
  -m ./POCKET-Darwin-180B-Q4.gguf \
  -t 16 \
  -c 8192 \
  -p "Explain mixture-of-experts routing in two sentences."
Enter fullscreen mode Exit fullscreen mode

A few practical notes from the measurements:

  • Threads matter. The server CPU result used 16 threads. Match -t to your physical cores. Oversubscribing threads past your core count usually slows generation, not speeds it.
  • Memory decides your path. With 128 GB of RAM the model stays resident and you get the top of the range. With 32 GB the OS pages expert tensors from disk, so throughput drops to a few tokens per second. A fast NVMe SSD helps the paging path noticeably.
  • Context length costs memory too. The 78.8 GB figure is peak at a working context. Raising -c raises the KV-cache footprint on top of the weights.
  • GPU offload is optional. On the 8 GB laptop GPU, only a fraction of layers fit in VRAM, so most of the work still lands on CPU and RAM. The 4.17 tok/s number reflects that mixed path.

When is a CPU-only 180B model actually the right call?

Throughput is the honest tradeoff. 21 tok/s on a server CPU is fine for batch jobs, drafting, extraction, and agent steps that tolerate latency. 4 tok/s on a laptop is a developer convenience, not a chat product. If you need interactive speed at scale, a GPU still wins.

Where the CPU path wins is where the data cannot leave the building. An air-gapped server, a mini PC on a factory floor, a defense or finance or public-sector box with no outbound connection: there a model that runs entirely on local CPU and RAM, with weights you downloaded once, is not a compromise. It is the only option that clears the policy. Edge AI is often less about milliseconds and more about where the bytes are allowed to sit.

FAQ

Can I run this in LM Studio or Ollama?
This post only verifies the llama.cpp path (b11048 or later), which is the engine GGUF targets. Treat other runners as untested here until you confirm they load this specific build.

Do I need the full 111 GB in RAM?
No. With 128 GB you can keep it resident for best speed. With 32 GB the OS memory-maps the GGUF and pages expert tensors on demand, trading throughput for a smaller footprint. You do need the full 111 GB on disk.

Why is the laptop so much slower than the server CPU?
Memory. The 32 GB laptop cannot hold the model, so it pages from disk, and only a few layers fit on the 8 GB GPU. The 16-thread server CPU with ample RAM keeps far more of the working set hot.

Is 4-bit always this lossless?
No, and do not generalize it. The 87.65% match is a measured paired result for this model and benchmark, enabled by keeping the trained delta (about 3% of the weights) at higher precision. Always verify your own quantization with a paired comparison rather than assuming parity.

What makes the active parameter count so low?
The MoE router picks 10 of 512 experts per token, so roughly 3B of the 180B parameters do work on any single token. That sparsity is what makes CPU inference and on-demand expert loading practical.

Where do I get it?
huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF and modelscope.cn/models/FINAL-Bench/POCKET-Darwin-180B-GGUF. The original is huggingface.co/FINAL-Bench/Darwin-180B-RSI.

Further reading

  • Earlier in this series: Offline AI on a Phone: Thread Pinning, Lazy Loading, and a Safety Gate That Knows When to Stay Quiet

Measurements in this post are from October 2026 device testing. Quantization parity was checked with paired, same-question comparisons; treat quantization as model-specific and verify your own.

Top comments (0)