DEV Community

Omniverse Compute (OMC)
Omniverse Compute (OMC)

Posted on

Your local LLM is a VRAM problem, not a compute problem

Your local LLM is a VRAM problem, not a compute problem

Every week someone asks whether their GPU is "fast enough" to run a local model. The question is almost always wrong. For local inference, throughput is rarely the constraint that stops you — capacity is. Weights have to fit in VRAM. Once they fit, a five-year-old card serves tokens at reading speed and nobody notices.

So the numbers that matter are about memory, and they're easy to compute once you separate the two things that consume it.

The two consumers of VRAM

  1. Weights — the model itself, in whatever quantization you chose.
  2. KV cache — the attention keys and values for the tokens in your context window. This scales with context length, batch size and model architecture, and it is the number people forget.

A rough planning formula for KV cache (fp16, no fancy tricks):

KV bytes ≈ 2 × layers × kv_heads × head_dim × 2 bytes × seq_len × batch
Enter fullscreen mode Exit fullscreen mode

The important term in there is seq_len. A model that fits comfortably at 4K context can fall off a cliff at 32K, and the GGUF file size on your disk tells you nothing about that — it's roughly two-thirds of what actually runs.

Model sizing, in 4-bit

7B   → ~4–5 GB weights   → fits almost anything with real context headroom
14B  → ~8–9 GB weights   → comfortable on a 12 GB card at moderate context
32B  → ~18–20 GB weights → the 24 GB sweet spot, with room for a real context window
70B  → ~38–42 GB weights → two 24 GB cards, or one very large one
Enter fullscreen mode Exit fullscreen mode

Those are weights only. Add KV cache, CUDA context and framework overhead, then add headroom — because the first time you actually need the model is not the time to discover you're 800 MB over.

Checking what you actually have

Before shopping for hardware, measure:

nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
nvidia-smi --query-compute-apps=pid,used_memory --format=csv
Enter fullscreen mode Exit fullscreen mode

On a machine that "should" fit the model but doesn't, those two commands usually explain it: a browser with hardware acceleration, another inference server, or a leftover process is holding gigabytes you forgot about.

Quantization is a memory decision, not a disk decision

The standard 4-bit formats (GGUF/AWQ-style) cut weight memory roughly four-fold with a modest quality cost that most people can't detect for drafting, summarizing and retrieval over personal documents. What they don't shrink proportionally is the KV cache — so on long-context workloads, quantization buys you less headroom than the file size suggests.

If you're choosing a quantization level, decide it against your target context length, not against disk space: "fits on disk" and "fits at the context I need" are different questions with different answers.

What this means for hardware choices

· One 24 GB card (a used RTX 3090 is still the default answer in 2026) serves 7B–14B fast, 32B comfortably in 4-bit, and does image, speech and vision work on the side.
· Two of them get you into 70B-class territory — and realistically into the "TCO beats renting" zone only if you were already going to buy the hardware.
· MoE models change the arithmetic again: total parameters no longer equal active parameters, so memory pressure and compute pressure decouple. Worth checking before you conclude a model is out of reach.
· CPU offload works, badly. It's a great way to run a model once to see if you like it, and a terrible way to run it daily.

The distributed angle, honestly

There is a real question underneath this: if VRAM is the constraint, why not pool VRAM across machines? That's the thesis behind decentralized GPU networks like Omniverse Compute (which, full disclosure, I work on). Pooling helps for the cases where a job can be sharded — batch rendering, certain training patterns, embarrassingly parallel inference. It helps much less for a single latency-sensitive chat session, where inter-node round trips eat the advantage. Anyone who tells you a network "adds VRAM to your laptop" is skipping that distinction.

The short version

Size the model against VRAM at the context length you need, not against your disk or your benchmark envy. Measure what's actually resident before you buy anything. And treat "faster" as a tuning problem — quantization, batching, flash attention, context management — long before you treat it as a purchasing problem.


Originally published at the Decentralized Compute Forum (https://forum.omc.network) — an editorial site on GPU markets and the economics of AI compute, operated by the OMC team. Source piece: "Quantization is a VRAM problem, not a disk problem" → https://forum.omc.network/posts/quantizing-local-llm-vram-math-2026

Top comments (0)