Everyone guesses VRAM. Here is how to actually measure it.
Ask five people "how much VRAM for a 70B model" and you get five answers: 40GB, 48GB, "about 35GB at Q4", "you need an A100", "just use the API". They are all approximating the same broken formula.
The formula is params x bits / 8. It is wrong in four ways:
- It ignores the embedding table, which stays in higher precision even in aggressive quants.
- It ignores the KV cache, which grows with context length and attention architecture (GQA vs MLA vs MHA differ by 5-10x).
- It ignores runtime overhead (CUDA context, fragmentation, the compute graph) - realistically 8-20%.
- For MoE models it does not matter how many experts exist; what matters is how many activate per token.
The honest answer
VRAM needed = measured GGUF file size + KV cache (from the real model config) + overhead.
The GGUF file size is not a formula - it is the actual number of bytes you download from Hugging Face. The KV cache is computed from the model's config.json (layers, kv-heads, head-dim for GQA; kv_lora_rank for MLA). No guessing.
I built a free tool that does exactly this for 110+ models, refreshed daily from the Hugging Face API:
https://modelfit-eight.vercel.app/guides/
It includes a full guide on the VRAM math, plus:
- why Q4_K_M is the sweet spot (and when to drop to IQ3 or go up to Q5/Q6)
- Ollama vs llama.cpp vs LM Studio - which to pick
- RTX 3090 vs 4090 for local LLMs (spoiler: the used 3090 wins on value)
- how people run 70B models on a single 24GB card (MoE offload + KV quantization)
- which GPU to buy for a given model (reverse picker: pick the model, see the cheapest card)
The one thing I want you to take away
Stop trusting params x bits / 8. Open the model's Hugging Face repo, look at the actual .gguf file sizes, add KV cache for your context length, add 10% overhead. That number is real. Every estimate is a lie with extra steps.
The site is static, no signup, no email wall, data refreshes daily. Criticism welcome - especially if you think my KV-cache formula or overhead percentages are off.
Build: measured GGUF sizes via Hugging Face API, KV from config.json (GQA/MLA), daily cron rebuild.
Top comments (0)