Researched October 2026. Architecture specs from public model cards; all numbers below are arithmetic, not my own benchmarks.
In my last post I showed that a 7B model's weights need ~16 GB in FP16 — a comfortable fit on a 24 GB RTX 4090. Several readers asked the obvious follow-up: then why does my 7B inference server OOM the moment I enable long context?
Because weights are only half the story. The other half is the KV cache, and it scales in a way that surprises almost everyone.
The One Formula
For each token in context, the model must store a key and a value vector per layer:
KV cache per token = 2 × layers × kv_heads × head_dim × bytes_per_value
Take Llama 3.1 8B (32 layers, 8 KV heads, 128 head dim, FP16):
2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes ≈ 128 KiB per token
Now multiply by context length:
| Context | KV cache (Llama 3.1 8B, FP16) | + 16 GB weights → total |
|---|---|---|
| 8K | ~1 GB | ~17 GB fits 4090 |
| 32K | ~4 GB | ~21 GB fits 4090 |
| 128K | ~16 GB | ~33 GB needs 48 GB+ |
Same model, same weights — the context window alone decides whether you need a $0.35/hr card or a $1.50+/hr one. (Prices from my provider comparison.)
Not All 7B Models Are Equal
KV head count varies wildly between architectures:
| Model | Layers | KV heads | KiB/token (FP16) | 128K context cache |
|---|---|---|---|---|
| Llama 3.1 8B | 32 | 8 | 128 | ~16 GB |
| Mistral 7B v0.3 | 32 | 8 | 128 | ~16 GB |
| Qwen2.5 7B | 28 | 4 | 56 | ~7 GB |
Qwen2.5 7B at 128K needs roughly half the cache of Llama 3.1 8B. Two "7B models," very different GPU bills.
The 70B Reality Check
Llama 3.1 70B (80 layers, 8 KV heads): 320 KiB/token → 128K context = ~41 GB of KV cache. Add 140 GB of FP16 weights and you're at ~181 GB — that's three 80 GB cards. Even in FP8 (70 GB weights), you're at ~111 GB: still two cards minimum. Long context is where 70B deployments quietly become multi-GPU projects.
Three Things Most Guides Skip
1. Batch size is a silent multiplier. KV cache scales with batch × sequence length. Serving 8 concurrent users at 8K context costs roughly the same cache as 1 user at 64K. Your "fits on a 4090" math breaks the moment traffic grows.
2. PagedAttention doesn't shrink the cache. vLLM's PagedAttention eliminates fragmentation waste, but the bytes are the bytes — it can't make 33 GB fit in 24 GB.
3. Quantize the cache too. KV cache in FP8 halves these numbers with minimal quality loss on most workloads. It's the same lever as weight quantization, applied to the forgotten half of VRAM.
A 30-Second Sizing Rule
context_budget = (VRAM − weights − 2 GB overhead) ÷ KiB_per_token ÷ batch_size
If your planned context exceeds the budget, you have three moves: quantize the cache, pick a GQA-heavy architecture, or rent the bigger card — in that order of cost-effectiveness.
I built a free GPU Advisor that walks through these tradeoffs with four questions and recommends the cheapest fitting setup. No signup.
What's the longest context you're actually serving in production — and did the KV cache math surprise you the first time? Curious what caught people off guard.
Top comments (0)