DEV Community

Helgard
Helgard

Posted on

Running the 510 GB DeepSeek-V4.1-Flash on an 8 GB GPU — and three bugs that never raise an error

My home machine for local models is modest: an RTX 5060 with 8 GB, a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. DeepSeek-V4.1-Flash is 510 GB on disk. It now runs on that box as a normal chat model in Open WebUI: ~1.6 tokens/s reading experts from disk only, ~2.4 tokens/s with a 16 GB RAM cache. That is slow, and this post is honest about why — but it works, and every speed-up keeps the output token-for-token identical to a plain version that does all the math with DeepSeek's own code.

Code: https://github.com/helgard-orlm/deepseek-v41-flash-8gb

Who did what, up front: I set the goal, chose the model, proposed some of the ideas (predicting the next layer's experts among them) and made the calls. The engine, the server and the bug hunts were done by Claude (Anthropic) in my sessions; the overlap scheme and the one-token short path came from Codex (OpenAI) reviewing the engine. "We" below means the three of us.

This is the second model we run this way. The first, a 133 GB Qwen at 11 tokens/s, is described in the previous post. The comparison between the two is the most useful part of this one.

Bytes per token decide everything

For a mixture-of-experts model, the size on disk barely matters. What matters is how many expert bytes one token pulls in.

DeepSeek-V4.1-Flash: 40 layers, 384 experts per layer, 6 chosen per token, each expert 17.93 MiB in FP4 with scales.

6 experts × 40 layers × 17.93 MiB ≈ 4.2 GiB per token
Enter fullscreen mode Exit fullscreen mode

The Qwen model needs 1.27 GiB per token. Same machine, same disk, same method — 3.3× more bytes, and that alone explains most of the speed gap.

Where the 475 GiB go:

part size where it lives
attention (FP8), shared experts, norms, router ~6.7 GiB GPU
routed experts (15,360 of them) 269 GiB NVMe → RAM cache → GPU
Engram tables (a per-token memory, FP8) 189 GiB NVMe, ~48 rows per token, 5 ms
embeddings + output head RAM, head computed on the CPU in FP32 like the reference

The approach: don't rewrite the model, only move its storage

DeepSeek publishes its reference inference code (model.py + kernel.py, MIT). We use it as is for all the arithmetic and replace only where the weights come from. That has a nice consequence: there is no conversion step. In the original safetensors files each expert's three matrices already sit next to each other (17.7 MB) and its scales are one more block (1.1 MB), so an expert is two pread calls with O_DIRECT straight from the Hugging Face files. A read benchmark showed this already uses 94–97% of what the drive can do, so re-laying-out the data on disk had nothing to give.

We first measured the whole thing on a rented RTX 5090 with VRAM artificially capped at 7.3 GiB, before buying the Gen5 drive. Once the numbers looked worth it, the model moved home.

Three bugs on RTX 50xx that never raise an error

This is the part I'd want to read before trying anything similar on a Blackwell consumer card.

1. TileLang 0.1.8 computes garbage on sm_120, silently. The reference pins this version. On our card the FP4 matrix multiply had cosine 0.0006 to a plain torch computation (i.e. unrelated numbers), the FP8 one returned NaN, and the model happily wrote "athaatha Stone". The reference self-test passed — it checks shapes, not values. TileLang 0.1.9: cosine 0.9996.

2. A race in the reference act_quant kernel. With DeepSeek's scale format (ue8m0, always on in this model) the kernel is built with num_stages = 0. On the RTX 5060, for more than 64 rows, part of the FP8 output comes out NaN — and differently every time: the same input five times gave 367, then up to 2065 NaN values. Generating one token at a time is clean, so only prompts were hit. It wasn't the compiler (CUDA 12.8 did the same). num_stages = 2 gives zero NaN and is bit-exact with a torch implementation at 1, 17, 300 and 1024 rows. That one line is the only change the setup script makes to DeepSeek's code.

3. The sparse attention kernel wants 141 KB of shared memory; consumer Blackwell has ~99 KB. Heads are independent, so we call the same kernel in groups of 16 heads. Small trap inside: .contiguous() on a head slice of a 1-token tensor doesn't change its strides, and the kernel checks strides — clone(memory_format=torch.contiguous_format) does.

The common lesson: on new hardware, "it ran without errors" says nothing. Compare every kernel with a dumb torch version on real weights before trusting a single generated word.

Making it faster: v1 to v4

The rule for every step: the 136 generated tokens of a fixed 4-question check must stay identical to the plain version (v1). Anything that changed a token was rejected or fixed. To be precise about what that proves: the speed-ups changed nothing. It does not prove v1 equals DeepSeek's untouched program run side by side — that program can't load the model on 8 GB, so we never ran it. What backs v1 instead: it calls DeepSeek's own model and kernel code for every computation, the individual kernels were compared with plain torch on real weights (above), the expanded wo_a matches DeepSeek's convert.py bit for bit, and the answers are right (Canberra, a correct TCP/UDP explanation, working merge code).

v1 — plain streaming: 0.89 tokens/s from disk, 1.17 with a 14 GB RAM cache. Read the layer's 6 experts, copy them to the GPU, wait, compute. Disk, bus and compute simply added up.

v2 — overlap (Codex's scheme): 1.28. A separate copy thread with CUDA events instead of a global synchronize. The surprise: issuing all 6 expert reads at once was slower. They share the drive, so they all finish together at the end of the layer and computing can't start any earlier. What worked was a deep queue but experts one after another: at most 2 in flight, each read in 4 parallel pieces. Waiting for the disk went from 0.65 to 0.31 s per token. (The first version of this had a race of its own — a ring buffer overwrote an expert that hadn't reached the GPU yet, and token 29 changed. The token check caught it.)

v3 — a short path for one token: +5.7%. The reference finds each expert's tokens with torch.where(indices == e) — a CPU↔GPU sync, 240 times per token, and each one delays the next copy. For a single token the indices are now moved to the CPU once per layer. Same arithmetic.

v4 — prefetch two experts of the next layer into the gap: 1.64 from disk (+21%), 2.3 with RAM. Apply layer i+1's router to layer i's input (rescaled by the two layers' norm weights) and you get the right top-6 71% of the time. Once layer i's real reads are in flight, read the top 2 guesses that aren't already in RAM. 94% of them get used. Three or four guesses are worse: they start competing with the reads that are actually needed.

One thing we got wrong along the way: I tried to reduce "bytes read in vain" by counting guesses that were already in RAM against the budget. Fewer wasted bytes, lower speed. Reads that land in the gap, while the disk would otherwise idle, are free; the only thing that matters is how many experts arrive on time.

version disk only with RAM cache
v1 plain 0.89 1.17
v2 overlap 1.28 1.42–1.67
v3 short path 1.35 1.95
v4 prefetch 1.64 2.28–2.34 (14 GB)
live chat, RAM 16 GB 2.43

A 311-token prompt takes 18 s (27 s in v1). Peak VRAM 7.06 GiB. Follow-up messages don't re-read the conversation: the server snapshots the full model state (30 MiB — KV, the sliding window, compression tails, indexer cache, Engram history) after each answer and processes only the new tail.

What we measured and dropped

Some of these were my ideas, some Claude's; all were measured on the real weights rather than argued about.

  • Store deltas between experts or layers, they'll compress better. The FP4 codes have 3.895 bits of entropy; the difference to the same expert in the next layer, or to a neighbour, has 3.997 — worse than the original. Even after the best neuron permutation and per-neuron scaling, the remaining difference is 0.999–1.000 of the original, exactly what two random matrices give.
  • Skip "silent" neurons and don't read their part of the down projection. To keep each expert within 5% error you still need 61–75% of the neurons: ~12% fewer bytes, and the read would have to wait for the first half of the computation. This model's SwiGLU neurons simply don't go quiet the way ReLU² ones do.
  • Smarter cache policies. A simulation over recorded routes: LRU beat LFU with ageing, per-layer LRU and SLRU. The theoretical optimum is ~15 points higher, but nothing simple gets near it.
  • More reader threads, splitting reads into smaller pieces — the drive was already near its limit.

Why it stops here

From disk, a token takes about 0.6 s, and just reading its 4.2 GiB at the 8.2 GiB/s the drive delivers on these reads is ~0.5 s of that. There's very little software left between us and that number. The only ways to go faster are fewer bytes (more RAM for the cache — the whole expert set is 269 GiB, so a cache is always partial) or a second drive. With Qwen the same method gives 11 tokens/s because each token needs a third of the bytes. If you're choosing a MoE model to stream on a home PC, look at bytes per token before parameter count.

Try it

You need a fast NVMe drive with ~520 GB free and a Blackwell GPU (the fixes above are for sm_120; other cards may need fewer of them). setup.sh downloads the exact model revision we used if it isn't there, checks that kernel.py is the expected file, and builds a copy of DeepSeek's code with the one-line fix. Before publishing, we cloned the repository fresh on the same machine, ran the setup against the downloaded model and started the server from the clone: it loaded with every parameter accounted for and answered correctly. The engine and server files are byte-identical to what runs at home. The full 510 GB download from scratch was not repeated for this test.

Engine, server, bug hunts and this write-up: Claude (Anthropic). Overlap scheme and the one-token short path: Codex (OpenAI). Goal, model choice, the expert-prediction idea and decisions: me. All numbers are from logs on the machines described above.

Top comments (0)