Originally published on mrsaynothing.dev. The GGUF VRAM calculator itself is open source — one HTML file, no build.
ggml_backend_cuda_buffer_type_alloc_buffer: failed to allocate 3412 MiB. The file downloaded fine. The card was never going to take it — and the arithmetic that predicted this would fit in the margin of a receipt. The usual workflow is backwards: hit download, watch the progress bar, let the out-of-memory error do the math. That ordering is what today's tool refuses. The GGUF VRAM calculator is live: model size, quantization and context in — weights, KV cache and a verdict for every common card out. Or run it in reverse from your VRAM budget to the largest quant that fits. It runs entirely in your browser, nothing is uploaded, and it joins the run-LLMs-locally cluster.
It has two modes, because the question comes in two shapes:
- "How much VRAM does this model need?" — parameters, quant, context, architecture in; a breakdown and a card-by-card verdict table out (6 through 96 GB).
- "What fits in my VRAM?" — budget, model size, context in; every quant from Q2_K to F16 with its own verdict out.
How much VRAM does a GGUF model need?
Three parts, always the same three:
VRAM ≈ weights + KV cache + overhead
weights = params × bits-per-weight / 8
KV cache = 2 × layers × kv_dim × context × 2 bytes (f16)
overhead ≈ 0.7 GB CUDA/runtime + ~5% compute buffers
Worked example, the most common one: an 8B model at Q4_K_M with 32k context. Weights: 8 × 4.85 / 8 = 4.85 GB. KV cache: 2 × 32 layers × 1024 kv-dim × 2 bytes = 128 KB per token × 32,768 = 4.0 GB. With overhead: ~9.8 GB total — "Q4 8B", the model everyone calls an 8-GB-card model, misses an 8 GB card by 23% the moment you give it a long context. The quant was never the villain; the context was. That failure mode has its own field note, because it also eats system RAM when you offload.
Which GGUF quant fits my card?
Mode 2 exists because the reverse question is what people actually have: a fixed card, a model in mind. A 7B at 4096 context against an 8 GB budget returns exactly this:
| Quant | Weights | Total (est.) | Verdict |
|---|---|---|---|
| Q4_K_M | 4.24 GB | ~5.7 GB | fits with headroom |
| Q6_K | 5.77 GB | ~7.3 GB | tight |
| Q8_0 | 7.44 GB | ~9.0 GB | partial CPU offload, much slower |
One screen, and the age-old forum debate "Q4 or Q8" collapses into your hardware's answer instead of somebody else's.
Where the numbers come from
The bits-per-weight table comes from llama.cpp's ggml quant formats — the same numbers the files on Hugging Face are built with:
| Quant | bpw | Quant | bpw |
|---|---|---|---|
| Q2_K | 3.35 | Q5_K_M | 5.69 |
| Q3_K_M | 3.91 | Q6_K | 6.59 |
| IQ4_XS | 4.25 | Q8_0 | 8.50 |
| Q4_K_M | 4.85 | F16 | 16.0 |
The KV cache half uses each family's public geometry. Grouped-query attention is the whole story there: a Llama-3-class 8B keeps 8 KV heads × 128 dims × 32 layers, which is 128 KB per token — while a Llama 2 13B, with no GQA, burns 640 KB per token, five times more, from an older and nominally smaller-era model. The architecture dropdown carries the four common shapes plus a custom mode for anything you can read off a config.json.
The file size told you what you would download. It never told you what you could run.
What the estimates deliberately miss
Three things, on purpose. Mixture-of-experts routing: only the active experts get touched at inference, but this tool prices the whole weight set, so MoE totals read high. Mixed quantization: Q4_K_M is itself an average across tensors, so a real file lands a few percent off the table value. And CUDA-side padding varies by backend version — the flat 0.7 GB is a middle-of-the-road figure, not a constant of nature. When a decision is worth more than a few hundred megabytes, skip the estimates and read the file's header with llama.cpp's gguf_dump.py — header-reading calculators do exactly that. This one stays arithmetic on purpose: three inputs you can re-derive by hand, and a verdict you can argue with.
The tool is at github.com/mrsaynothing/gguf-vram-calculator — one HTML file, no dependencies, no build step. It sits alongside the local GGUF guide, llama.cpp vs Ollama and the best-local-LLM shortlist, with the serving-side counterpart in can vLLM run GGUF.
If a verdict disagrees with your load times, open a repository issue with the model, quant and context — the table should survive contact with real hardware, and anywhere it doesn't is a row to fix.
If this saved you one failed load this year, the weekly letter carries more of the same — one email, real numbers.
Top comments (1)
The napkin that started this tool: an 8B at Q4_K_M with 32k context lands around 9.8 GB — 4.85 of weights and almost 4 more of pure KV cache. The Q4-fits-an-8GB-card rule dies the moment the context grows, and the file size never warns you. What is the biggest model you have watched OOM at load time — and was it the file size that misled you, or the context?