DEV Community

Kata Omel
Kata Omel

Posted on Originally published at modelfit-eight.vercel.app

How much VRAM does a 70B LLM actually need? I wrote the honest math down (free, no signup)

Everyone guesses VRAM. Here is how to actually measure it.

Ask five people "how much VRAM for a 70B model" and you get five answers: 40GB, 48GB, "about 35GB at Q4", "you need an A100", "just use the API". They are all approximating the same broken formula.

The formula is params x bits / 8. It is wrong in four ways:

  1. It ignores the embedding table, which stays in higher precision even in aggressive quants.
  2. It ignores the KV cache, which grows with context length and attention architecture (GQA vs MLA vs MHA differ by 5-10x).
  3. It ignores runtime overhead (CUDA context, fragmentation, the compute graph) - realistically 8-20%.
  4. For MoE models it does not matter how many experts exist; what matters is how many activate per token.

The honest answer

VRAM needed = measured GGUF file size + KV cache (from the real model config) + overhead.

The GGUF file size is not a formula - it is the actual number of bytes you download from Hugging Face. The KV cache is computed from the model's config.json (layers, kv-heads, head-dim for GQA; kv_lora_rank for MLA). No guessing.

I built a free tool that does exactly this for 110+ models, refreshed daily from the Hugging Face API:

https://modelfit-eight.vercel.app/guides/

It includes a full guide on the VRAM math, plus:

  • why Q4_K_M is the sweet spot (and when to drop to IQ3 or go up to Q5/Q6)
  • Ollama vs llama.cpp vs LM Studio - which to pick
  • RTX 3090 vs 4090 for local LLMs (spoiler: the used 3090 wins on value)
  • how people run 70B models on a single 24GB card (MoE offload + KV quantization)
  • which GPU to buy for a given model (reverse picker: pick the model, see the cheapest card)

The one thing I want you to take away

Stop trusting params x bits / 8. Open the model's Hugging Face repo, look at the actual .gguf file sizes, add KV cache for your context length, add 10% overhead. That number is real. Every estimate is a lie with extra steps.

The site is static, no signup, no email wall, data refreshes daily. Criticism welcome - especially if you think my KV-cache formula or overhead percentages are off.

Build: measured GGUF sizes via Hugging Face API, KV from config.json (GQA/MLA), daily cron rebuild.

Top comments (0)