DEV Community

Donald Lee
Donald Lee

Posted on

How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

Mixture-of-experts models like Qwen3.6 35B-A3B, gpt-oss or GLM Flash are too big for most consumer GPUs, but they only read a few experts per token. llama.cpp's --n-cpu-moe N flag exploits that: it keeps the expert FFN tensors of the first N layers in system RAM and runs them on the CPU, while attention, shared weights and the KV cache stay on the GPU.

The hard part is choosing N. Too low and you run out of VRAM; too high and you give away speed for nothing. Guessing and restarting the server a few times works, but you can compute it from the GGUF file itself.

Step 1: read the tensor sizes from the GGUF header

A GGUF file starts with a header that lists every tensor with its type and shape, so you can get exact byte counts without downloading the weights (an HTTP range request on the first few MB is enough). For the unsloth Qwen3.6-35B-A3B-UD-Q4_K_M.gguf (22,123,538,944 bytes of tensors):

Part Bytes
Expert tensors (ffn_*_exps), per layer 486,539,264 (about 0.45 GiB), on most of the 40 layers
Everything else (attention, shared, norms, output) 2,555,013,632
of which token_embd 540,344,320

Two details matter:

  • token_embd never goes to the GPU. llama.cpp keeps the input embedding on the CPU whatever -ngl says, so subtract it from the VRAM side.
  • Expert size differs slightly per layer in some quants, so sum the actual layers you move rather than multiplying one number.

Step 2: add the KV cache for your context

Qwen3.6 35B-A3B is a hybrid model: only 10 of its 40 layers use full attention (2 KV heads × 256 dims), the rest are linear-attention layers with a small fixed state. The FP16 KV cache is therefore

2 (K and V) × 10 layers × 2 heads × 256 × 2 bytes = 20,480 bytes per token

which is only 0.625 GiB at 32K context and 2.5 GiB at 128K. On dense models of this size it would be several times more, so always check the config's layer_types / full-attention interval before assuming.

Step 3: find the smallest N that fits

VRAM needed = non-expert weights − token_embd + expert bytes of the layers kept on the GPU + KV cache + about 1 GiB for CUDA context and compute buffers. Increase N (moving whole layers' experts to RAM) until that is below your card's memory:

GPU Context --n-cpu-moe VRAM used In system RAM Rough decode speed
12 GB (RTX 3060, 360 GB/s) 32K 22 11.8 GiB 10.5 GiB 24–41 tok/s
12 GB 128K 26 11.8 GiB 12.3 GiB 16–27 tok/s
16 GB (RTX 5060 Ti, 448 GB/s) 32K 13 15.8 GiB 6.4 GiB 31–53 tok/s
16 GB 128K 17 15.9 GiB 8.2 GiB 21–35 tok/s
24 GB (RTX 4090) 32K 0 21.7 GiB 0.5 GiB (only token_embd) 81–142 tok/s
24 GB 128K 0 23.6 GiB 0.5 GiB 52–91 tok/s

Speeds assume dual-channel DDR5-5600 (about 90 GB/s) and are ranges from memory bandwidth, not benchmarks: each token reads the active experts (8 of 256 per layer here), so the layers on the CPU are limited by RAM bandwidth and the rest by VRAM bandwidth.

So on a 16 GB card:

llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 99 --n-cpu-moe 13 -c 32768
Enter fullscreen mode Exit fullscreen mode

Things that trip people up

  • Leave headroom. The 1 GiB buffer figure is an estimate; if you see OOM at the first long prompt, add one more layer to N or lower -ub.
  • Quantizing the KV cache (-fa on -ctk q8_0 -ctv q8_0) roughly halves the cache, which on hybrid models like this buys little, but on MoE models where every layer keeps a full KV cache it can save you a layer or two of offload at long context.
  • Split GGUFs: pass the first shard (-00001-of-0000N.gguf); llama.cpp loads the rest.
  • Models with a huge per-layer embedding table (e.g. Qwen3.8 Flash Next's n-gram table) can leave it on disk with --lazy-mode, so it doesn't count against RAM or VRAM.

A calculator that does this for you

Disclosure: I made a free planner that does exactly these steps from the GGUF headers for about 24 popular MoE files, including the per-layer sizes, the smallest N for your card and context, and the command line: modelvram.com/moe-offload-calculator. The underlying numbers are also available as an open dataset at modelvram.com/data if you want to check them.

Top comments (0)