Model selection usually starts with a number: 8B, 27B, 30B. It is the first thing on the model card and the first thing in every comparison thread. It is also close to useless for the question most teams are actually asking, which is how much traffic will this serve on the card I have.
Below is that question worked out for four current open models on a single 96 GB GPU. Everything here is arithmetic on published configs — no benchmarks were run. The script is at the end and takes about a second.
Units: GB throughout means GiB (2³⁰ bytes), which is what nvidia-smi reports.
The four models
| Model | Total params | Architecture | Native context | License |
|---|---|---|---|---|
| Qwen3-8B | 8.2B | Dense, 36 layers, GQA 32Q/8KV, head dim 128 | 32,768 (131K with YaRN) | Apache-2.0 |
| Qwen3.5-9B | ~9B | Hybrid: 3 linear-attention layers per 1 full, 32 total, 4 KV heads, head dim 256 | 262,144 (1.01M with YaRN) | Apache-2.0 |
| Gemma-4-26B-A4B-it | 25.2B (3.8B active) | MoE, 128 experts, top-8 + 1 shared, 30 layers, 5 full-attention + 25 sliding (window 1024), 8 KV heads, head dim 256 | 262,144 | Apache-2.0 |
| Qwen3.8-27B | ~27B | Hybrid: same 3:1 pattern, 64 layers, 4 KV heads, head dim 256 | 262,144 (1M with YaRN) | Apache-2.0 |
Every figure in that table comes from the model card or config.json of the repo it links to.
Weights: where MoE stops helping
Mixture-of-Experts models are usually introduced with their active parameter count, because that is what determines compute per token. Gemma-4-26B-A4B activates 3.8B parameters out of 25.2B.
Memory does not work that way. Routing is per-token and unpredictable, so every expert has to be resident. At BF16:
- Qwen3-8B — 15.3 GB
- Qwen3.5-9B — 16.8 GB
- Gemma-4-26B-A4B — 46.9 GB
- Qwen3.8-27B — 50.3 GB
The "4B active" model costs nearly three times the VRAM of the 9B. What you buy for that is compute: it runs at roughly 4B-model speed per token while holding 26B-model knowledge. That is a real and valuable trade. It is just not a memory saving, and model cards that lead with the active count invite exactly that misreading.
KV cache: the term nobody budgets
Weights are static. KV cache grows with every token in flight, and it is where single-GPU deployments actually fall over. For a standard attention layer:
bytes per token = 2 (K and V) × layers × kv_heads × head_dim × dtype_bytes
The variable that matters most is layers — specifically, how many layers keep a cache that grows with sequence length. Three of these four models do something to reduce that count:
-
Qwen3-8B is conventional. All 36 layers cache.
2 × 36 × 8 × 128 × 2 = 144 KB per token. -
Qwen3.5-9B interleaves 3 linear-attention layers for every 1 full-attention layer. Linear attention carries a fixed-size recurrent state instead of a growing cache, so only 8 of 32 layers contribute:
2 × 8 × 4 × 256 × 2 = 32 KB per token. - Qwen3.8-27B uses the same 3:1 pattern over 64 layers, so 16 layers cache: 64 KB per token.
- Gemma-4-26B-A4B splits differently: 5 full-attention layers grow with the sequence (40 KB per token), while the other 25 use a 1024-token sliding window and stay flat at about 200 MB per sequence no matter how long the context gets.
So an 8B dense model spends 4.5× more KV memory per token than a 9B hybrid. Parameter count predicted the opposite.
What that buys in concurrency
Take 96 GB, subtract weights, reserve 4 GB for activations and framework overhead, and divide by the KV cost of one 32,768-token sequence:
| Model | KV per 32K sequence | Concurrent 32K sequences |
|---|---|---|
| Qwen3-8B | 4.50 GB | 17 |
| Qwen3.5-9B | 1.00 GB | 75 |
| Gemma-4-26B-A4B | 1.45 GB | 31 |
| Qwen3.8-27B | 2.00 GB | 20 |
The ranking has almost nothing to do with size. Qwen3.5-9B — nominally the second-smallest model here — serves four times the concurrent long-context traffic of Qwen3-8B, and nearly four times that of the 27B. The 26B MoE, despite spending 47 GB on weights, still fits more concurrent sequences than the 8B dense model does.
If your workload is RAG, document processing, or anything else that puts long prompts in flight, this table is a better predictor of cost per request than any leaderboard.
Why there is no quality column here
There is an obvious missing dimension, and it is missing on purpose.
The four model cards report largely disjoint benchmark sets. Gemma-4-26B-A4B-it reports MMLU-Pro 82.6%, GPQA Diamond 82.3%, LiveCodeBench v6 77.1%. Qwen3.5-9B reports MMLU-Pro 82.5 and C-Eval 88.2. Qwen3.8-27B reports GPQA Diamond 89.2 and LiveCodeBench v6 90.3. Qwen3-8B's card reports neither.
Two of those MMLU-Pro numbers are within 0.1 of each other, which is exactly the kind of coincidence that produces a confident and wrong blog post. Vendor-run evaluations differ in prompt format, few-shot count, decoding parameters, and whether reasoning mode was enabled — none of which are consistently disclosed. Putting them in one table implies a comparison that the underlying numbers do not support.
The honest version: use these figures to shortlist, then run your own eval on your own prompts. Memory math transfers between deployments. Benchmark scores often do not.
What this analysis does not tell you
- Nothing about throughput. Memory determines what fits. Tokens per second depends on memory bandwidth, kernel quality, batching policy, and the serving stack. A model that fits 75 sequences will not necessarily serve them fast.
- Quantization changes every number. FP8 weights roughly halve the weight column; FP8 or INT8 KV cache halves the cache column. The relative ordering mostly survives, the absolute headroom does not.
- Linear attention is a trade, not a free win. The architectures that cut KV cost do so by compressing history into a fixed-size state. That is very good for throughput and is not free on tasks that need precise long-range recall. Test it rather than assuming the memory win is costless.
-
Real allocators are messier. vLLM and SGLang preallocate a KV pool via
gpu_memory_utilization; fragmentation, CUDA graphs, and multimodal encoders all take a cut. Treat the concurrency numbers as ceilings. - Parameter counts are approximate. Qwen publishes "9B" and "27B" rather than exact totals, so those weight figures carry a percent or two of slack.
Reproduce it
GiB = 1024**3
BYTES = 2 # BF16 KV cache
# name, params(B), full_attn_layers, sliding_layers, window, kv_heads, head_dim
MODELS = [
("Qwen3-8B", 8.2, 36, 0, 0, 8, 128),
("Qwen3.5-9B", 9.0, 8, 0, 0, 4, 256),
("Gemma-4-26B-A4B", 25.2, 5, 25, 1024, 8, 256),
("Qwen3.8-27B", 27.0, 16, 0, 0, 4, 256),
]
def kv_for_seq(full_l, sl_l, win, kvh, hd, seq):
growing = 2 * full_l * kvh * hd * BYTES * seq
capped = 2 * sl_l * kvh * hd * BYTES * min(seq, win)
return growing + capped
VRAM, RESERVE, SEQ = 96, 4, 32_768
for name, p, full_l, sl_l, win, kvh, hd in MODELS:
weights = p * 1e9 * BYTES / GiB
per_seq = kv_for_seq(full_l, sl_l, win, kvh, hd, SEQ) / GiB
free = VRAM - weights - RESERVE
print(f"{name:<18} weights {weights:5.1f} GB | "
f"{per_seq:4.2f} GB/seq | {int(free/per_seq):3d} concurrent {SEQ//1024}K seqs")
Change SEQ to your real context length and BYTES to 1 if you serve an FP8 KV cache. The layer counts come straight from each repo's config.json — if a model ships a new revision, re-read it rather than trusting this table.


Top comments (0)