This article was originally published on runaihome.com
The number that makes people assume they've misread something: an RTX 4090 running Qwen3-30B-A3B at up to 196 tokens per second. A 7B dense model on the same GPU benchmarks at around 135 tok/s. A dense 30B model wouldn't even fit in 24GB VRAM at Q4 without extreme quantization. How does a 30B model go faster than a 7B? That's the MoE architecture at work.
What's happening is a Mixture-of-Experts (MoE) architecture that fundamentally changes how local inference scales — and unless you understand it, you'll either write off this model as "too big for consumer hardware" or reach for it when the dense 32B is actually the smarter choice. This is the practical guide: how the architecture works, what VRAM you need, real tok/s numbers across GPU tiers, and exactly how to run it.
What "30B-A3B" Means, and Why Your Bandwidth Math Is Wrong
For dense LLMs, the bandwidth rule is simple: each token generated requires reading every parameter in the model from VRAM. A 30B dense model at Q4 needs to move roughly 15 GB of data through the GPU's memory bus per token. That's why a 30B model runs at roughly one-fifth the speed of a 7B on the same GPU.
Qwen3-30B-A3B breaks this rule with a sparse architecture. Inside the model are 128 separate "expert" networks. For each token, a lightweight routing layer examines the current hidden state and selects only 8 of those 128 experts to activate. Every other expert sits idle — its weights sit in VRAM but don't touch the compute units.
The practical breakdown:
- Total parameters: 30.5 billion (all experts combined)
- Active parameters per token: 3.3 billion (8 of 128 experts)
- Effective inference cost per token: similar to a 4B–8B dense model
- VRAM footprint: still sized for the full 30.5B (everything must be loaded)
This is the MoE bargain: you pay the VRAM cost of a 30B model, but you get the generation speed of a model three to five times smaller. Quality sits somewhere between the two, shaped by the fact that during training, those 128 experts developed genuine specialization — routing pushes different types of reasoning to different expert subsets.
The full architecture: 48 transformer layers, 32 query attention heads with 4 KV heads (grouped query attention), 128 total experts with 8 activated per forward pass. Native context is 32,768 tokens, extended to 131,072 tokens with YaRN RoPE scaling. License is Apache 2.0, which means commercial use without royalties.
For comparison, Llama 3.3 70B — the other strong 24GB-runnable option — is a standard dense model where all 70B parameters load into VRAM and participate in every token. Fitting it in 24GB requires heavy quantization (Q3 or aggressive Q2), which costs noticeably more quality than Q4_K_M on a smaller model.
VRAM Requirements by Quantization Level
The "30B" label on the tin causes unnecessary hardware anxiety. At Q4_K_M, the weights sit at roughly 19 GB — comfortably inside a 24GB GPU with headroom for a reasonable KV cache.
| Quantization | File Size | Minimum VRAM | Fits On |
|---|---|---|---|
| Q4_K_M (default) | ~19 GB | 24 GB | RTX 4090, RTX 3090, RTX 4080 |
| Q5_K_M | ~22 GB | 24 GB | RTX 4090/3090; tight — reduces KV cache room |
| Q8_0 | ~31 GB | 40+ GB or dual 24GB | Used A6000 (48GB), Mac Studio 64GB+ |
| BF16 (full precision) | ~61 GB | 80 GB+ | H100, multi-GPU with NVLink |
A few cards that won't work cleanly:
RTX 4060 Ti 16GB: Q4_K_M doesn't fit (19 GB > 16 GB). You can use --n-gpu-layers in llama.cpp to keep the first N layers on GPU and offload the rest to system RAM — but the PCIe bottleneck between GPU and system RAM guts your tok/s, and you've lost the main reason to run this model over Qwen3-14B.
RTX 3060 12GB: Not viable at any useful quantization. The model needs 19 GB just for weights; the card has 12 GB. Full CPU offload would result in sub-5 tok/s performance, slower than Qwen3-8B running entirely on GPU.
Mac Studio M2/M3 with 64GB unified memory: Works cleanly at Q4_K_M (19 GB of 64 GB used) via MLX. Mac Studio 96GB has comfortable headroom for Q8 inference.
If you're on 16GB, the right model is Qwen3-14B or Qwen3-8B, not this one.
Tokens Per Second Across GPU Tiers
Community benchmarks from April–May 2026 using Q4_K_M quantization in llama.cpp:
| GPU | VRAM | Memory Bandwidth | Qwen3-30B-A3B tok/s |
|---|---|---|---|
| RTX 4090 | 24 GB | 1,008 GB/s | 120–196 tok/s |
| RTX 3090 | 24 GB | 936 GB/s | ~73 tok/s |
| RTX 4060 Ti 16GB | 16 GB | 288 GB/s | Not recommended (CPU offload required; PCIe bottleneck kills throughput) |
| RTX 3060 12GB | 12 GB | 360 GB/s | Not viable (weights exceed VRAM by 7+ GB) |
The RTX 4090 range (120–196 tok/s) reflects variation across test conditions: different quant variants (Q4_K_M vs Unsloth UD-Q4_K_XL), context window sizes, and whether llama.cpp or Ollama is the inference backend. Ollama adds a Go server layer that typically costs 3–10% throughput compared to raw llama.cpp; the upper bound (196 tok/s) comes from optimized llama.cpp setups with modest context windows.
The RTX 3090 figure (73 tok/s) is lower than many expect given it's only 7% slower than the RTX 4090 on memory bandwidth. MoE inference involves more irregular memory access patterns than dense models — the routing mechanism causes non-contiguous expert weight reads — which appears to amplify the sensitivity to GPU architecture differences beyond raw bandwidth numbers.
To put the speed in perspective: a dense Qwen3-32B model on an RTX 4090 runs substantially slower, because all 32B parameters load from VRAM for every token — at Q4_K_M the weights alone occupy ~19 GB, leaving very little KV cache headroom at 24 GB. Community benchmarks consistently report the MoE 30B-A3B running 3–5× faster than the dense 32B on the same GPU, at the cost of around 2–3 points on standard benchmarks.
Quality Benchmarks: The Honest Numbers
Qwen3-30B-A3B vs Qwen3-32B dense — the direct matchup that matters for 24GB GPU owners:
| Benchmark | Qwen3-30B-A3B | Qwen3-32B (dense) |
|---|---|---|
| MMLU | 81.38 | ~83 |
| Arena Hard | 91.0% | 93.8% |
| AIME 2024 | 80.4% | 81.4% |
| AIME 2025 | 70.9% | ~72% |
The gap is real but narrow: the dense 32B holds a 1–3 point advantage across benchmarks. Both models blow past Llama 3.3 70B on math — Llama 3.3 70B scores MATH 77.0% (MATH benchmark) and has a strong 88.4% on HumanEval, but its reasoning under AIME-style competition math is significantly weaker than either Qwen3 variant. Qwen3's training methodology produces genuinely better mathematical reasoning at comparable or smaller VRAM footprints.
In everyday chat and coding tasks you won't notice the 2–3 point quality gap between the two Qwen3 models. In structured math problems or complex multi-step reasoning, the dense 32B will occasionally produce a more complete chain of reasoning. Whether that marginal accuracy gain is worth 3–5× slower generation is a use-case question, answered below.
Thinking Mode: /think and /no_think
Qwen3-30B-A3B includes a built-in thinking mode switch — you don't need a separate model file. Add /think anywhere in a prompt to activate extended chain-of-thought reasoning. The model generates its reasoning inside <think>...</think> tags before producing the final response. Add /no_think to turn it off within the same session.
When thinking mode helps:
- Multi-step math problems where you'd spot-check the intermediate steps
- Code debugging with non-obvious logic errors
- Planning task
Top comments (0)