DEV Community

yyyysa4
yyyysa4

Posted on

Parameter count is a bad way to pick an open model for one GPU

Model selection usually starts with a number: 8B, 27B, 30B. It is the first thing on the model card and the first thing in every comparison thread. It is also close to useless for the question most teams are actually asking, which is how much traffic will this serve on the card I have.

Below is that question worked out for four current open models on a single 96 GB GPU. Everything here is arithmetic on published configs — no benchmarks were run. The script is at the end and takes about a second.

Units: GB throughout means GiB (2³⁰ bytes), which is what nvidia-smi reports.

The four models

Model Total params Architecture Native context License
Qwen3-8B 8.2B Dense, 36 layers, GQA 32Q/8KV, head dim 128 32,768 (131K with YaRN) Apache-2.0
Qwen3.5-9B ~9B Hybrid: 3 linear-attention layers per 1 full, 32 total, 4 KV heads, head dim 256 262,144 (1.01M with YaRN) Apache-2.0
Gemma-4-26B-A4B-it 25.2B (3.8B active) MoE, 128 experts, top-8 + 1 shared, 30 layers, 5 full-attention + 25 sliding (window 1024), 8 KV heads, head dim 256 262,144 Apache-2.0
Qwen3.8-27B ~27B Hybrid: same 3:1 pattern, 64 layers, 4 KV heads, head dim 256 262,144 (1M with YaRN) Apache-2.0

Every figure in that table comes from the model card or config.json of the repo it links to.

Weights: where MoE stops helping

Mixture-of-Experts models are usually introduced with their active parameter count, because that is what determines compute per token. Gemma-4-26B-A4B activates 3.8B parameters out of 25.2B.

Memory does not work that way. Routing is per-token and unpredictable, so every expert has to be resident. At BF16:

  • Qwen3-8B — 15.3 GB
  • Qwen3.5-9B — 16.8 GB
  • Gemma-4-26B-A4B — 46.9 GB
  • Qwen3.8-27B — 50.3 GB

The "4B active" model costs nearly three times the VRAM of the 9B. What you buy for that is compute: it runs at roughly 4B-model speed per token while holding 26B-model knowledge. That is a real and valuable trade. It is just not a memory saving, and model cards that lead with the active count invite exactly that misreading.

KV cache: the term nobody budgets

Weights are static. KV cache grows with every token in flight, and it is where single-GPU deployments actually fall over. For a standard attention layer:

bytes per token = 2 (K and V) × layers × kv_heads × head_dim × dtype_bytes
Enter fullscreen mode Exit fullscreen mode

The variable that matters most is layers — specifically, how many layers keep a cache that grows with sequence length. Three of these four models do something to reduce that count:

  • Qwen3-8B is conventional. All 36 layers cache. 2 × 36 × 8 × 128 × 2 = 144 KB per token.
  • Qwen3.5-9B interleaves 3 linear-attention layers for every 1 full-attention layer. Linear attention carries a fixed-size recurrent state instead of a growing cache, so only 8 of 32 layers contribute: 2 × 8 × 4 × 256 × 2 = 32 KB per token.
  • Qwen3.8-27B uses the same 3:1 pattern over 64 layers, so 16 layers cache: 64 KB per token.
  • Gemma-4-26B-A4B splits differently: 5 full-attention layers grow with the sequence (40 KB per token), while the other 25 use a 1024-token sliding window and stay flat at about 200 MB per sequence no matter how long the context gets.

So an 8B dense model spends 4.5× more KV memory per token than a 9B hybrid. Parameter count predicted the opposite.

What that buys in concurrency

Take 96 GB, subtract weights, reserve 4 GB for activations and framework overhead, and divide by the KV cost of one 32,768-token sequence:

Model KV per 32K sequence Concurrent 32K sequences
Qwen3-8B 4.50 GB 17
Qwen3.5-9B 1.00 GB 75
Gemma-4-26B-A4B 1.45 GB 31
Qwen3.8-27B 2.00 GB 20

The ranking has almost nothing to do with size. Qwen3.5-9B — nominally the second-smallest model here — serves four times the concurrent long-context traffic of Qwen3-8B, and nearly four times that of the 27B. The 26B MoE, despite spending 47 GB on weights, still fits more concurrent sequences than the 8B dense model does.

If your workload is RAG, document processing, or anything else that puts long prompts in flight, this table is a better predictor of cost per request than any leaderboard.

Why there is no quality column here

There is an obvious missing dimension, and it is missing on purpose.

The four model cards report largely disjoint benchmark sets. Gemma-4-26B-A4B-it reports MMLU-Pro 82.6%, GPQA Diamond 82.3%, LiveCodeBench v6 77.1%. Qwen3.5-9B reports MMLU-Pro 82.5 and C-Eval 88.2. Qwen3.8-27B reports GPQA Diamond 89.2 and LiveCodeBench v6 90.3. Qwen3-8B's card reports neither.

Two of those MMLU-Pro numbers are within 0.1 of each other, which is exactly the kind of coincidence that produces a confident and wrong blog post. Vendor-run evaluations differ in prompt format, few-shot count, decoding parameters, and whether reasoning mode was enabled — none of which are consistently disclosed. Putting them in one table implies a comparison that the underlying numbers do not support.

The honest version: use these figures to shortlist, then run your own eval on your own prompts. Memory math transfers between deployments. Benchmark scores often do not.

What this analysis does not tell you

  • Nothing about throughput. Memory determines what fits. Tokens per second depends on memory bandwidth, kernel quality, batching policy, and the serving stack. A model that fits 75 sequences will not necessarily serve them fast.
  • Quantization changes every number. FP8 weights roughly halve the weight column; FP8 or INT8 KV cache halves the cache column. The relative ordering mostly survives, the absolute headroom does not.
  • Linear attention is a trade, not a free win. The architectures that cut KV cost do so by compressing history into a fixed-size state. That is very good for throughput and is not free on tasks that need precise long-range recall. Test it rather than assuming the memory win is costless.
  • Real allocators are messier. vLLM and SGLang preallocate a KV pool via gpu_memory_utilization; fragmentation, CUDA graphs, and multimodal encoders all take a cut. Treat the concurrency numbers as ceilings.
  • Parameter counts are approximate. Qwen publishes "9B" and "27B" rather than exact totals, so those weight figures carry a percent or two of slack.

Reproduce it

GiB = 1024**3
BYTES = 2  # BF16 KV cache

# name, params(B), full_attn_layers, sliding_layers, window, kv_heads, head_dim
MODELS = [
    ("Qwen3-8B",         8.2, 36,  0,    0, 8, 128),
    ("Qwen3.5-9B",       9.0,  8,  0,    0, 4, 256),
    ("Gemma-4-26B-A4B", 25.2,  5, 25, 1024, 8, 256),
    ("Qwen3.8-27B",     27.0, 16,  0,    0, 4, 256),
]

def kv_for_seq(full_l, sl_l, win, kvh, hd, seq):
    growing = 2 * full_l * kvh * hd * BYTES * seq
    capped  = 2 * sl_l   * kvh * hd * BYTES * min(seq, win)
    return growing + capped

VRAM, RESERVE, SEQ = 96, 4, 32_768
for name, p, full_l, sl_l, win, kvh, hd in MODELS:
    weights = p * 1e9 * BYTES / GiB
    per_seq = kv_for_seq(full_l, sl_l, win, kvh, hd, SEQ) / GiB
    free    = VRAM - weights - RESERVE
    print(f"{name:<18} weights {weights:5.1f} GB | "
          f"{per_seq:4.2f} GB/seq | {int(free/per_seq):3d} concurrent {SEQ//1024}K seqs")
Enter fullscreen mode Exit fullscreen mode

Change SEQ to your real context length and BYTES to 1 if you serve an FP8 KV cache. The layer counts come straight from each repo's config.json — if a model ships a new revision, re-read it rather than trusting this table.


Top comments (0)