DEV Community

Cover image for GGUF VRAM Calculator: Check Before You Download
Mr Say Nothing
Mr Say Nothing

Posted on Originally published at mrsaynothing.dev

GGUF VRAM Calculator: Check Before You Download

Originally published on mrsaynothing.dev. The GGUF VRAM calculator itself is open source — one HTML file, no build.

ggml_backend_cuda_buffer_type_alloc_buffer: failed to allocate 3412 MiB. The file downloaded fine. The card was never going to take it — and the arithmetic that predicted this would fit in the margin of a receipt. The usual workflow is backwards: hit download, watch the progress bar, let the out-of-memory error do the math. That ordering is what today's tool refuses. The GGUF VRAM calculator is live: model size, quantization and context in — weights, KV cache and a verdict for every common card out. Or run it in reverse from your VRAM budget to the largest quant that fits. It runs entirely in your browser, nothing is uploaded, and it joins the run-LLMs-locally cluster.

It has two modes, because the question comes in two shapes:

  • "How much VRAM does this model need?" — parameters, quant, context, architecture in; a breakdown and a card-by-card verdict table out (6 through 96 GB).
  • "What fits in my VRAM?" — budget, model size, context in; every quant from Q2_K to F16 with its own verdict out.

How much VRAM does a GGUF model need?

Three parts, always the same three:

VRAM ≈ weights + KV cache + overhead
weights  = params × bits-per-weight / 8
KV cache = 2 × layers × kv_dim × context × 2 bytes   (f16)
overhead ≈ 0.7 GB CUDA/runtime + ~5% compute buffers
Enter fullscreen mode Exit fullscreen mode

Worked example, the most common one: an 8B model at Q4_K_M with 32k context. Weights: 8 × 4.85 / 8 = 4.85 GB. KV cache: 2 × 32 layers × 1024 kv-dim × 2 bytes = 128 KB per token × 32,768 = 4.0 GB. With overhead: ~9.8 GB total — "Q4 8B", the model everyone calls an 8-GB-card model, misses an 8 GB card by 23% the moment you give it a long context. The quant was never the villain; the context was. That failure mode has its own field note, because it also eats system RAM when you offload.

Which GGUF quant fits my card?

Mode 2 exists because the reverse question is what people actually have: a fixed card, a model in mind. A 7B at 4096 context against an 8 GB budget returns exactly this:

Quant Weights Total (est.) Verdict
Q4_K_M 4.24 GB ~5.7 GB fits with headroom
Q6_K 5.77 GB ~7.3 GB tight
Q8_0 7.44 GB ~9.0 GB partial CPU offload, much slower

One screen, and the age-old forum debate "Q4 or Q8" collapses into your hardware's answer instead of somebody else's.

Where the numbers come from

The bits-per-weight table comes from llama.cpp's ggml quant formats — the same numbers the files on Hugging Face are built with:

Quant bpw Quant bpw
Q2_K 3.35 Q5_K_M 5.69
Q3_K_M 3.91 Q6_K 6.59
IQ4_XS 4.25 Q8_0 8.50
Q4_K_M 4.85 F16 16.0

The KV cache half uses each family's public geometry. Grouped-query attention is the whole story there: a Llama-3-class 8B keeps 8 KV heads × 128 dims × 32 layers, which is 128 KB per token — while a Llama 2 13B, with no GQA, burns 640 KB per token, five times more, from an older and nominally smaller-era model. The architecture dropdown carries the four common shapes plus a custom mode for anything you can read off a config.json.

The file size told you what you would download. It never told you what you could run.

What the estimates deliberately miss

Three things, on purpose. Mixture-of-experts routing: only the active experts get touched at inference, but this tool prices the whole weight set, so MoE totals read high. Mixed quantization: Q4_K_M is itself an average across tensors, so a real file lands a few percent off the table value. And CUDA-side padding varies by backend version — the flat 0.7 GB is a middle-of-the-road figure, not a constant of nature. When a decision is worth more than a few hundred megabytes, skip the estimates and read the file's header with llama.cpp's gguf_dump.py — header-reading calculators do exactly that. This one stays arithmetic on purpose: three inputs you can re-derive by hand, and a verdict you can argue with.

The tool is at github.com/mrsaynothing/gguf-vram-calculator — one HTML file, no dependencies, no build step. It sits alongside the local GGUF guide, llama.cpp vs Ollama and the best-local-LLM shortlist, with the serving-side counterpart in can vLLM run GGUF.

If a verdict disagrees with your load times, open a repository issue with the model, quant and context — the table should survive contact with real hardware, and anywhere it doesn't is a row to fix.

If this saved you one failed load this year, the weekly letter carries more of the same — one email, real numbers.

Top comments (1)

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The napkin that started this tool: an 8B at Q4_K_M with 32k context lands around 9.8 GB — 4.85 of weights and almost 4 more of pure KV cache. The Q4-fits-an-8GB-card rule dies the moment the context grows, and the file size never warns you. What is the biggest model you have watched OOM at load time — and was it the file size that misled you, or the context?