Mixture-of-experts models like Qwen3.6 35B-A3B, gpt-oss or GLM Flash are too big for most consumer GPUs, but they only read a few experts per token. llama.cpp's --n-cpu-moe N flag exploits that: it keeps the expert FFN tensors of the first N layers in system RAM and runs them on the CPU, while attention, shared weights and the KV cache stay on the GPU.
The hard part is choosing N. Too low and you run out of VRAM; too high and you give away speed for nothing. Guessing and restarting the server a few times works, but you can compute it from the GGUF file itself.
Step 1: read the tensor sizes from the GGUF header
A GGUF file starts with a header that lists every tensor with its type and shape, so you can get exact byte counts without downloading the weights (an HTTP range request on the first few MB is enough). For the unsloth Qwen3.6-35B-A3B-UD-Q4_K_M.gguf (22,123,538,944 bytes of tensors):
| Part | Bytes |
|---|---|
Expert tensors (ffn_*_exps), per layer |
486,539,264 (about 0.45 GiB), on most of the 40 layers |
| Everything else (attention, shared, norms, output) | 2,555,013,632 |
of which token_embd
|
540,344,320 |
Two details matter:
-
token_embdnever goes to the GPU. llama.cpp keeps the input embedding on the CPU whatever-nglsays, so subtract it from the VRAM side. - Expert size differs slightly per layer in some quants, so sum the actual layers you move rather than multiplying one number.
Step 2: add the KV cache for your context
Qwen3.6 35B-A3B is a hybrid model: only 10 of its 40 layers use full attention (2 KV heads × 256 dims), the rest are linear-attention layers with a small fixed state. The FP16 KV cache is therefore
2 (K and V) × 10 layers × 2 heads × 256 × 2 bytes = 20,480 bytes per token
which is only 0.625 GiB at 32K context and 2.5 GiB at 128K. On dense models of this size it would be several times more, so always check the config's layer_types / full-attention interval before assuming.
Step 3: find the smallest N that fits
VRAM needed = non-expert weights − token_embd + expert bytes of the layers kept on the GPU + KV cache + about 1 GiB for CUDA context and compute buffers. Increase N (moving whole layers' experts to RAM) until that is below your card's memory:
| GPU | Context | --n-cpu-moe |
VRAM used | In system RAM | Rough decode speed |
|---|---|---|---|---|---|
| 12 GB (RTX 3060, 360 GB/s) | 32K | 22 | 11.8 GiB | 10.5 GiB | 24–41 tok/s |
| 12 GB | 128K | 26 | 11.8 GiB | 12.3 GiB | 16–27 tok/s |
| 16 GB (RTX 5060 Ti, 448 GB/s) | 32K | 13 | 15.8 GiB | 6.4 GiB | 31–53 tok/s |
| 16 GB | 128K | 17 | 15.9 GiB | 8.2 GiB | 21–35 tok/s |
| 24 GB (RTX 4090) | 32K | 0 | 21.7 GiB | 0.5 GiB (only token_embd) | 81–142 tok/s |
| 24 GB | 128K | 0 | 23.6 GiB | 0.5 GiB | 52–91 tok/s |
Speeds assume dual-channel DDR5-5600 (about 90 GB/s) and are ranges from memory bandwidth, not benchmarks: each token reads the active experts (8 of 256 per layer here), so the layers on the CPU are limited by RAM bandwidth and the rest by VRAM bandwidth.
So on a 16 GB card:
llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 99 --n-cpu-moe 13 -c 32768
Things that trip people up
-
Leave headroom. The 1 GiB buffer figure is an estimate; if you see OOM at the first long prompt, add one more layer to N or lower
-ub. -
Quantizing the KV cache (
-fa on -ctk q8_0 -ctv q8_0) roughly halves the cache, which on hybrid models like this buys little, but on MoE models where every layer keeps a full KV cache it can save you a layer or two of offload at long context. -
Split GGUFs: pass the first shard (
-00001-of-0000N.gguf); llama.cpp loads the rest. -
Models with a huge per-layer embedding table (e.g. Qwen3.8 Flash Next's n-gram table) can leave it on disk with
--lazy-mode, so it doesn't count against RAM or VRAM.
A calculator that does this for you
Disclosure: I made a free planner that does exactly these steps from the GGUF headers for about 24 popular MoE files, including the per-layer sizes, the smallest N for your card and context, and the command line: modelvram.com/moe-offload-calculator. The underlying numbers are also available as an open dataset at modelvram.com/data if you want to check them.
Top comments (0)