How much VRAM for an LLM is the wrong first question. The better one is how fast that VRAM is, because a local model generating text reads its entire set of weights from memory for every single token. That makes memory bandwidth, in gigabytes per second, the number that decides whether your coding agent types or crawls. This is the bandwidth-first companion to my Local AI Hardware Guide (2026) from March: fewer shopping lists, more of the arithmetic behind them.
TL;DR
- Token generation at batch size 1 is memory-bandwidth bound. A rough ceiling is
bandwidth ÷ model size in memory: a 10 GB model on a 936 GB/s RTX 3090 tops out around 90 tokens a second. - When the model or its context does not fit, layers spill to DDR5 (about 50 GB/s) over PCIe 4.0 (31.5 GB/s): a 20x bandwidth drop, and 42.5 tok/s becomes 3.8 in the benchmark below.
- At 4-bit, weights cost about 5 GB for 8B, 10 GB for 14B, 20 GB for 32B and 40 GB for 70B, before the KV cache. A 32k-token agent context adds several more gigabytes.
- The best VRAM per dollar is still a used RTX 3090 24GB, around $650–750, because it pairs 24 GB with a 384-bit bus.
Why memory bandwidth, not VRAM size alone, sets LLM speed
A game is compute-bound: the CPU prepares draw calls, the shaders fill pixels, a faster core means more frames. Autoregressive decoding, the way an LLM writes one token at a time, is the reverse. To compute the next token the GPU streams every weight of the model from VRAM through its compute units, does a little maths on each, and starts again. The cores mostly wait on the memory bus. The back-of-envelope formula:
max tokens/sec ≈ memory bandwidth (GB/s) ÷ model size in memory (GB)
RTX 3090, 14B model at 4-bit (~10 GB): 936 / 10 ≈ 94 tok/s ceiling
same model spilled to DDR5 (~50 GB/s): 50 / 10 ≈ 5 tok/s ceiling
Real runs land below the ceiling (roughly 60–75 tok/s on the 3090, 1.5–2 on DDR5), but the ratio holds: same model, same machine, twenty times slower because the weights moved. That is why a 5.8 GHz Core i9 and a liquid loop do nothing here. The CPU only feeds the GPU; a six-core Ryzen 5 is plenty.
GPU memory bandwidth by tier
Bandwidth is set by the memory type and the bus width; the marketing tier says nothing about it. The same 16 GB can be fast or slow:
| Card | VRAM | Bus | Bandwidth | In practice |
|---|---|---|---|---|
| RTX 4060 Ti 16GB | 16 GB GDDR6 | 128-bit | 288 GB/s | starter; 14B fits |
| RTX 4070 Ti Super | 16 GB GDDR6X | 256-bit | 672 GB/s | fast, but no 32B |
| RTX 3090 (used) | 24 GB GDDR6X | 384-bit | 936 GB/s | 32B plus context |
| RTX 4090 | 24 GB GDDR6X | 384-bit | 1,008 GB/s | 3090 capacity, ~8 % faster |
| DDR5-6000, dual channel | system RAM | 128-bit | ~48–55 GB/s | where spilled layers go |
| PCIe 4.0 x16 | link | — | 31.5 GB/s | the pipe in between |
Figures are from Nvidia's spec pages for the RTX 3090 and RTX 4090, plus the PCIe 4.0 and JEDEC DDR5 standards.
The 4060 Ti row explains itself: 288 GB/s over a 5 GB model is a ~58 tok/s ceiling, and it measures 42.5 below. Fine for a starter card, out of its depth at 32B.
The 20x cliff: what happens when a model does not fit in VRAM
Runtimes like llama.cpp, Ollama and vLLM don't refuse a model that is too big. They split it: some layers in VRAM, the rest in system RAM across the PCIe bus. The GPU finishes its layers in microseconds, then stalls.
The benchmark from the episode, Llama 8B on an RTX 4060 Ti 16GB:
| Where the weights live | Speed | Change |
|---|---|---|
| 100 % in VRAM | 42.5 tok/s | baseline |
| 80 % VRAM, 20 % in DDR5 | 3.8 tok/s | −91 % |
| 100 % on CPU / DDR5 | 1.6 tok/s | −96 % |
Remember the middle row. Moving a fifth of the model off the card cost 91 % of the speed, because every token waits for the slowest pipe. "It almost fits" is not a thing: either the model and its context fit in VRAM, or you run at DDR5 speed with extra steps.
How much VRAM do you need for a 32B model? Weights plus KV cache
Weights are the entry fee. At 4-bit quantization (GGUF Q4_K_M or AWQ):
| Model class | Weights at 4-bit | VRAM with working context | Fits on |
|---|---|---|---|
| 7B / 8B (Llama 3.1 8B, DeepSeek-R1 Distill 7B) | ~5 GB | ~8–10 GB | RTX 4060 Ti 16GB, Mac 16GB |
| 14B (Qwen 2.5 14B) | ~10 GB | ~14–16 GB | 16 GB cards |
| 32B (Qwen 2.5 32B, DeepSeek Distill 32B) | ~20 GB | ~24–26 GB | RTX 3090 / 4090 24GB, Mac with 64 GB |
| 70B (Llama 3.3 70B) | ~40–43 GB | ~48–52 GB | dual RTX 3090 (48 GB), Mac Studio 96–128 GB |
The third column is where agent users get surprised. Coding agents like Cline, Roo Code or Cursor keep the system prompt, tool definitions and every file chunk in context, and the model stores keys and values for each token in the KV cache, which grows linearly:
# simplified sketch: KV cache size for one sequence
def kv_cache_gb(n_layers, n_kv_heads, head_dim, context_tokens, bytes_per_value=2): # 2 = fp16
return 2 * bytes_per_value * n_layers * n_kv_heads * head_dim * context_tokens / 1e9
# Llama 3.1 8B config: 32 layers, 8 KV heads, head_dim 128
kv_cache_gb(32, 8, 128, 8_192) # ≈ 1.1 GB
kv_cache_gb(32, 8, 128, 32_768) # ≈ 4.3 GB
kv_cache_gb(32, 8, 128, 131_072) # ≈ 17 GB, more than a 16 GB card holds in total
The layer and head counts are from the model's config on Hugging Face. Qwen 2.5 32B at 32k context needs about 19.5 GB of weights plus 5.5 GB of cache, roughly 25 GB. So a 16 GB RTX 4070 Ti Super cannot run it properly despite 672 GB/s, and the used 3090's extra 8 GB matters more than the 4090's extra 72 GB/s.
Best GPU for local LLMs: three budget tiers
Tier 1, starter, $1,200–1,500. An RTX 4060 Ti 16GB (about $450), a Ryzen 5 7600, 64 GB of DDR5, a 2 TB NVMe drive. Never the 8 GB version: it saves about $70, halves the VRAM, and anything above 8B with real context runs out of memory. Apple equivalent: a Mac mini with 16 GB. Expect 40–50 tok/s on 8B, a tight fit for 14B.
Tier 2, the sweet spot, about $1,650. Built for 32B coding models like Qwen 2.5 32B. Path A, a new RTX 4070 Ti Super 16GB for $800, forces 32B down to aggressive 3-bit quantization. Path B is my pick: a used RTX 3090 24GB for $650–750, with a Ryzen 7 7700X, 64 GB DDR5 and an 850 W PSU. It runs 32B at 35–40 tok/s with 16–32k of context. The Mac counterpart, a Mac mini M4 Pro with 64 GB at about $2,200, does 11–12 tok/s on 32B: slower, but silent at about 30 W and able to hold 48 GB-plus models.
Tier 3, enthusiast, $3,500–5,000+. One RTX 4090 or two used 3090s (48 GB, about $1,400 in GPUs), a Ryzen 9 7950X, 128 GB of RAM, a 1,200 W supply: 32B at 65+ tok/s, 70B across two cards at 20+. On Apple, a Mac Studio Ultra with 96–128 GB, whose strength is residency: an embedding model, an 8B router and a 32B or 70B reasoning model stay loaded together, with no swapping between agent steps.
Ollama vs LM Studio, and GGUF vs AWQ
Ollama is a background daemon with a one-line CLI (ollama run qwen2.5:32b) and an API that Cursor, Cline and VS Code extensions connect to directly. LM Studio is the graphical option: model downloads, a GPU-layers slider, a local server. Whichever you use, the goal of that slider is every layer on the card.
Match the format to the silicon: GGUF on Apple Silicon (the llama.cpp format, built for Metal and unified memory), AWQ or EXL2 on Nvidia, which use the Tensor Cores; the episode's figure was 20–30 % more throughput than GGUF on CUDA.
The Raspberry Pi 5 trap
Every few weeks a post claims you can run coding agents on an $80 Raspberry Pi 5. I tried: a 1.5B model gave barely three tokens a second while the board throttled at 85 °C. A fun weekend project, punishment as a daily driver, for the same reason as everything above: little bandwidth, few tokens.
Local vs cloud: the home gym rule
Local hardware is a home gym: you train every day with no commute, no subscription, full privacy. That is the 80 %: completion, local refactors, boilerplate, unit tests, private document search, at zero dollars per token with the code never leaving your machine. The other 20 %, a repo-wide refactor across fifty files or multi-step frontier reasoning, goes to a cloud API for fifteen minutes.
The Tier 2 build with a used 3090 is about $1,600 all in. At $50–100 a month of API spend the episode's estimate was payback in about 18 months; at the $50 end it is closer to three years, so divide by your own bill.
Verdict: SHIP IT
I stamped this field test SHIP IT. For local inference, VRAM capacity and bandwidth beat clock speed every time, and the cheapest way to buy both is a used RTX 3090. Size the card for the model plus its context, keep every layer on the GPU, send the heavy 20 % to the cloud.
What do you run local models on, and what tokens per second do you actually see? Real numbers in the comments beat any spec sheet, mine included.
FAQ
How much VRAM do I need to run an LLM locally?
At 4-bit, with context: 8–10 GB for 8B, 14–16 GB for 14B, 24–26 GB for 32B, 48–52 GB for 70B. Long agent contexts add more.
Is a used RTX 3090 still good for LLMs in 2026?
Yes. It has 24 GB on a 384-bit bus at 936 GB/s, enough for a 32B model at 4-bit with 16–32k of context, for about $650–750 used.
Why is my local LLM so slow?
Usually part of the model or its KV cache spilled into system RAM; offloading 20 % of the layers cost 91 % of the speed above. Shrink the context or the model until everything fits on the GPU.
Sources
- My earlier guide on dev.to, The Local AI Hardware Guide (2026): https://dev.to/axrisi/the-local-ai-hardware-guide-2026-4mk
- NVIDIA GeForce RTX 3090 specifications: https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090-3090ti/
- NVIDIA GeForce RTX 4090 specifications: https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/
- llama.cpp (layer offloading, GGUF): https://github.com/ggml-org/llama.cpp
- Ollama: https://ollama.com
- LM Studio: https://lmstudio.ai
- vLLM (PagedAttention, KV cache management): https://github.com/vllm-project/vllm
- Llama 3.1 8B model card and config: https://huggingface.co/meta-llama/Llama-3.1-8B
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode · Subscribe on YouTube · the written diff lands in your inbox every morning at thedailydiff.dev.



Top comments (1)
Your 20%-spill row is the most damning number in the piece: on the same Llama 8B, moving a fifth of the model off the card drops you from 42.5 to 3.8 tok/s (-91%) — every spilled layer has to cross the 31.5 GB/s PCIe link on every token. One caveat that changes the buying advice: this cliff only exists on discrete GPUs. On unified-memory machines the offloaded layers sit in the same LPDDR pool, so a spill degrades gently instead of falling off the PCIe wall — which is why your own table lists a 64GB Mac alongside the 3090/4090 as fitting a 32B model. The 'minimum 24GB VRAM' takeaway is really 'minimum 24GB of fast-attached memory,' a very different shopping list for Mac buyers.