<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Helgard</title>
    <description>The latest articles on DEV Community by Helgard (@helgard_orlm).</description>
    <link>https://dev.to/helgard_orlm</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3980095%2F54f659c3-a39f-45b2-84f8-3d10deac3dd1.png</url>
      <title>DEV Community: Helgard</title>
      <link>https://dev.to/helgard_orlm</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/helgard_orlm"/>
    <language>en</language>
    <item>
      <title>Running the 510 GB DeepSeek-V4.1-Flash on an 8 GB GPU — and three bugs that never raise an error</title>
      <dc:creator>Helgard</dc:creator>
      <pubDate>Sat, 03 Oct 2026 12:25:53 +0000</pubDate>
      <link>https://dev.to/helgard_orlm/running-the-510-gb-deepseek-v41-flash-on-an-8-gb-gpu-and-three-bugs-that-never-raise-an-error-5647</link>
      <guid>https://dev.to/helgard_orlm/running-the-510-gb-deepseek-v41-flash-on-an-8-gb-gpu-and-three-bugs-that-never-raise-an-error-5647</guid>
      <description>&lt;p&gt;My home machine for local models is modest: an &lt;strong&gt;RTX 5060 with 8 GB&lt;/strong&gt;, a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. DeepSeek-V4.1-Flash is 510 GB on disk. It now runs on that box as a normal chat model in Open WebUI: &lt;strong&gt;~1.6 tokens/s reading experts from disk only, ~2.4 tokens/s with a 16 GB RAM cache.&lt;/strong&gt; That is slow, and this post is honest about why — but it works, and every speed-up keeps the output token-for-token identical to a plain version that does all the math with DeepSeek's own code.&lt;/p&gt;

&lt;p&gt;Code: &lt;strong&gt;&lt;a href="https://github.com/helgard-orlm/deepseek-v41-flash-8gb" rel="noopener noreferrer"&gt;https://github.com/helgard-orlm/deepseek-v41-flash-8gb&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who did what, up front:&lt;/strong&gt; I set the goal, chose the model, proposed some of the ideas (predicting the next layer's experts among them) and made the calls. The engine, the server and the bug hunts were done by &lt;strong&gt;Claude&lt;/strong&gt; (Anthropic) in my sessions; the overlap scheme and the one-token short path came from &lt;strong&gt;Codex&lt;/strong&gt; (OpenAI) reviewing the engine. "We" below means the three of us.&lt;/p&gt;

&lt;p&gt;This is the second model we run this way. The first, a 133 GB Qwen at 11 tokens/s, is described in &lt;a href="https://dev.to/helgard_orlm/running-a-133-gb-moe-model-on-an-8-gb-gpu-at-11-tokenss-by-streaming-experts-from-nvme-207b"&gt;the previous post&lt;/a&gt;. The comparison between the two is the most useful part of this one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bytes per token decide everything
&lt;/h2&gt;

&lt;p&gt;For a mixture-of-experts model, the size on disk barely matters. What matters is how many expert bytes one token pulls in.&lt;/p&gt;

&lt;p&gt;DeepSeek-V4.1-Flash: 40 layers, 384 experts per layer, 6 chosen per token, each expert 17.93 MiB in FP4 with scales.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;6 experts × 40 layers × 17.93 MiB ≈ 4.2 GiB per token
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Qwen model needs 1.27 GiB per token. Same machine, same disk, same method — 3.3× more bytes, and that alone explains most of the speed gap.&lt;/p&gt;

&lt;p&gt;Where the 475 GiB go:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;part&lt;/th&gt;
&lt;th&gt;size&lt;/th&gt;
&lt;th&gt;where it lives&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;attention (FP8), shared experts, norms, router&lt;/td&gt;
&lt;td&gt;~6.7 GiB&lt;/td&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;routed experts (15,360 of them)&lt;/td&gt;
&lt;td&gt;269 GiB&lt;/td&gt;
&lt;td&gt;NVMe → RAM cache → GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engram tables (a per-token memory, FP8)&lt;/td&gt;
&lt;td&gt;189 GiB&lt;/td&gt;
&lt;td&gt;NVMe, ~48 rows per token, 5 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;embeddings + output head&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;RAM, head computed on the CPU in FP32 like the reference&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The approach: don't rewrite the model, only move its storage
&lt;/h2&gt;

&lt;p&gt;DeepSeek publishes its reference inference code (&lt;code&gt;model.py&lt;/code&gt; + &lt;code&gt;kernel.py&lt;/code&gt;, MIT). We use it &lt;strong&gt;as is&lt;/strong&gt; for all the arithmetic and replace only where the weights come from. That has a nice consequence: there is no conversion step. In the original safetensors files each expert's three matrices already sit next to each other (17.7 MB) and its scales are one more block (1.1 MB), so an expert is two &lt;code&gt;pread&lt;/code&gt; calls with &lt;code&gt;O_DIRECT&lt;/code&gt; straight from the Hugging Face files. A read benchmark showed this already uses 94–97% of what the drive can do, so re-laying-out the data on disk had nothing to give.&lt;/p&gt;

&lt;p&gt;We first measured the whole thing on a rented RTX 5090 with VRAM artificially capped at 7.3 GiB, before buying the Gen5 drive. Once the numbers looked worth it, the model moved home.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three bugs on RTX 50xx that never raise an error
&lt;/h2&gt;

&lt;p&gt;This is the part I'd want to read before trying anything similar on a Blackwell consumer card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. TileLang 0.1.8 computes garbage on sm_120, silently.&lt;/strong&gt; The reference pins this version. On our card the FP4 matrix multiply had cosine 0.0006 to a plain torch computation (i.e. unrelated numbers), the FP8 one returned NaN, and the model happily wrote "athaatha Stone". The reference self-test passed — it checks shapes, not values. TileLang 0.1.9: cosine 0.9996.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A race in the reference &lt;code&gt;act_quant&lt;/code&gt; kernel.&lt;/strong&gt; With DeepSeek's scale format (&lt;code&gt;ue8m0&lt;/code&gt;, always on in this model) the kernel is built with &lt;code&gt;num_stages = 0&lt;/code&gt;. On the RTX 5060, for more than 64 rows, part of the FP8 output comes out NaN — and differently every time: the same input five times gave 367, then up to 2065 NaN values. Generating one token at a time is clean, so only prompts were hit. It wasn't the compiler (CUDA 12.8 did the same). &lt;code&gt;num_stages = 2&lt;/code&gt; gives zero NaN and is bit-exact with a torch implementation at 1, 17, 300 and 1024 rows. That one line is the only change the setup script makes to DeepSeek's code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The sparse attention kernel wants 141 KB of shared memory&lt;/strong&gt;; consumer Blackwell has ~99 KB. Heads are independent, so we call the same kernel in groups of 16 heads. Small trap inside: &lt;code&gt;.contiguous()&lt;/code&gt; on a head slice of a 1-token tensor doesn't change its strides, and the kernel checks strides — &lt;code&gt;clone(memory_format=torch.contiguous_format)&lt;/code&gt; does.&lt;/p&gt;

&lt;p&gt;The common lesson: on new hardware, "it ran without errors" says nothing. Compare every kernel with a dumb torch version on real weights before trusting a single generated word.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making it faster: v1 to v4
&lt;/h2&gt;

&lt;p&gt;The rule for every step: the 136 generated tokens of a fixed 4-question check must stay &lt;strong&gt;identical&lt;/strong&gt; to the plain version (v1). Anything that changed a token was rejected or fixed. To be precise about what that proves: the speed-ups changed nothing. It does not prove v1 equals DeepSeek's untouched program run side by side — that program can't load the model on 8 GB, so we never ran it. What backs v1 instead: it calls DeepSeek's own model and kernel code for every computation, the individual kernels were compared with plain torch on real weights (above), the expanded &lt;code&gt;wo_a&lt;/code&gt; matches DeepSeek's &lt;code&gt;convert.py&lt;/code&gt; bit for bit, and the answers are right (Canberra, a correct TCP/UDP explanation, working merge code).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v1 — plain streaming: 0.89 tokens/s&lt;/strong&gt; from disk, 1.17 with a 14 GB RAM cache. Read the layer's 6 experts, copy them to the GPU, wait, compute. Disk, bus and compute simply added up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v2 — overlap (Codex's scheme): 1.28.&lt;/strong&gt; A separate copy thread with CUDA events instead of a global synchronize. The surprise: issuing all 6 expert reads at once was &lt;em&gt;slower&lt;/em&gt;. They share the drive, so they all finish together at the end of the layer and computing can't start any earlier. What worked was a deep queue but experts &lt;strong&gt;one after another&lt;/strong&gt;: at most 2 in flight, each read in 4 parallel pieces. Waiting for the disk went from 0.65 to 0.31 s per token. (The first version of this had a race of its own — a ring buffer overwrote an expert that hadn't reached the GPU yet, and token 29 changed. The token check caught it.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v3 — a short path for one token: +5.7%.&lt;/strong&gt; The reference finds each expert's tokens with &lt;code&gt;torch.where(indices == e)&lt;/code&gt; — a CPU↔GPU sync, 240 times per token, and each one delays the next copy. For a single token the indices are now moved to the CPU once per layer. Same arithmetic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v4 — prefetch two experts of the next layer into the gap: 1.64 from disk (+21%), 2.3 with RAM.&lt;/strong&gt; Apply layer i+1's router to layer i's input (rescaled by the two layers' norm weights) and you get the right top-6 71% of the time. Once layer i's real reads are in flight, read the top 2 guesses that aren't already in RAM. 94% of them get used. Three or four guesses are worse: they start competing with the reads that are actually needed.&lt;/p&gt;

&lt;p&gt;One thing we got wrong along the way: I tried to reduce "bytes read in vain" by counting guesses that were already in RAM against the budget. Fewer wasted bytes, lower speed. Reads that land in the gap, while the disk would otherwise idle, are free; the only thing that matters is how many experts arrive on time.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;version&lt;/th&gt;
&lt;th&gt;disk only&lt;/th&gt;
&lt;th&gt;with RAM cache&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;v1 plain&lt;/td&gt;
&lt;td&gt;0.89&lt;/td&gt;
&lt;td&gt;1.17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v2 overlap&lt;/td&gt;
&lt;td&gt;1.28&lt;/td&gt;
&lt;td&gt;1.42–1.67&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v3 short path&lt;/td&gt;
&lt;td&gt;1.35&lt;/td&gt;
&lt;td&gt;1.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;v4 prefetch&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.64&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2.28–2.34&lt;/strong&gt; (14 GB)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;live chat, RAM 16 GB&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.43&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A 311-token prompt takes 18 s (27 s in v1). Peak VRAM 7.06 GiB. Follow-up messages don't re-read the conversation: the server snapshots the full model state (30 MiB — KV, the sliding window, compression tails, indexer cache, Engram history) after each answer and processes only the new tail.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we measured and dropped
&lt;/h2&gt;

&lt;p&gt;Some of these were my ideas, some Claude's; all were measured on the real weights rather than argued about.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Store deltas between experts or layers, they'll compress better.&lt;/strong&gt; The FP4 codes have 3.895 bits of entropy; the difference to the same expert in the next layer, or to a neighbour, has 3.997 — worse than the original. Even after the best neuron permutation and per-neuron scaling, the remaining difference is 0.999–1.000 of the original, exactly what two random matrices give.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skip "silent" neurons and don't read their part of the down projection.&lt;/strong&gt; To keep each expert within 5% error you still need 61–75% of the neurons: ~12% fewer bytes, and the read would have to wait for the first half of the computation. This model's SwiGLU neurons simply don't go quiet the way ReLU² ones do.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Smarter cache policies.&lt;/strong&gt; A simulation over recorded routes: LRU beat LFU with ageing, per-layer LRU and SLRU. The theoretical optimum is ~15 points higher, but nothing simple gets near it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;More reader threads, splitting reads into smaller pieces&lt;/strong&gt; — the drive was already near its limit.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why it stops here
&lt;/h2&gt;

&lt;p&gt;From disk, a token takes about 0.6 s, and just reading its 4.2 GiB at the 8.2 GiB/s the drive delivers on these reads is ~0.5 s of that. There's very little software left between us and that number. The only ways to go faster are fewer bytes (more RAM for the cache — the whole expert set is 269 GiB, so a cache is always partial) or a second drive. With Qwen the same method gives 11 tokens/s because each token needs a third of the bytes. If you're choosing a MoE model to stream on a home PC, look at &lt;strong&gt;bytes per token&lt;/strong&gt; before parameter count.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;You need a fast NVMe drive with ~520 GB free and a Blackwell GPU (the fixes above are for sm_120; other cards may need fewer of them). &lt;code&gt;setup.sh&lt;/code&gt; downloads the exact model revision we used if it isn't there, checks that &lt;code&gt;kernel.py&lt;/code&gt; is the expected file, and builds a copy of DeepSeek's code with the one-line fix. Before publishing, we cloned the repository fresh on the same machine, ran the setup against the downloaded model and started the server from the clone: it loaded with every parameter accounted for and answered correctly. The engine and server files are byte-identical to what runs at home. The full 510 GB download from scratch was not repeated for this test.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Engine, server, bug hunts and this write-up: Claude (Anthropic). Overlap scheme and the one-token short path: Codex (OpenAI). Goal, model choice, the expert-prediction idea and decisions: me. All numbers are from logs on the machines described above.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>cuda</category>
      <category>python</category>
    </item>
    <item>
      <title>Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe</title>
      <dc:creator>Helgard</dc:creator>
      <pubDate>Sat, 03 Oct 2026 11:46:52 +0000</pubDate>
      <link>https://dev.to/helgard_orlm/running-a-133-gb-moe-model-on-an-8-gb-gpu-at-11-tokenss-by-streaming-experts-from-nvme-207b</link>
      <guid>https://dev.to/helgard_orlm/running-a-133-gb-moe-model-on-an-8-gb-gpu-at-11-tokenss-by-streaming-experts-from-nvme-207b</guid>
      <description>&lt;p&gt;I run local models on one home machine: an &lt;strong&gt;RTX 5060 with 8 GB&lt;/strong&gt;, a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. When NVIDIA published &lt;code&gt;Qwen3.8-Flash-Next&lt;/code&gt; in NVFP4 (133 GB on disk), the obvious answer was "it doesn't fit". It does now: it runs as a normal chat model in Open WebUI at &lt;strong&gt;9–12 tokens/s&lt;/strong&gt;, and its output matches the Hugging Face reference implementation.&lt;/p&gt;

&lt;p&gt;Code: &lt;strong&gt;&lt;a href="https://github.com/helgard-orlm/qwen-flash-next-8gb" rel="noopener noreferrer"&gt;https://github.com/helgard-orlm/qwen-flash-next-8gb&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who did what, up front:&lt;/strong&gt; I set the goal, chose the model and made the calls along the way. The engine, kernels and server were written by &lt;strong&gt;Claude&lt;/strong&gt; (Anthropic) in my sessions; the CUDA-graph speed-up was written by &lt;strong&gt;Codex&lt;/strong&gt; (OpenAI) in a parallel session, and Claude then found and fixed a crash in it. "We" below means the three of us.&lt;/p&gt;

&lt;p&gt;This post is about the method, the measurements, and what didn't work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this model, of all models
&lt;/h2&gt;

&lt;p&gt;Most of the weights of a mixture-of-experts model are experts, and each token uses only a few of them. So the question is not "how big is the model" but &lt;strong&gt;how many expert bytes one token needs&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Qwen3.8-Flash-Next: 48 layers, 512 experts per layer, top-10 routing, and each expert is tiny — three 2560×640 matrices, 2.70 MiB in NVFP4 with scales.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10 experts × 48 layers × 2.70 MiB ≈ 1.27 GiB per token
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For comparison, DeepSeek-V4.1-Flash, which we run on the same box with the same trick, needs about 4.2 GiB per token. Same disk, a third of the bytes. (How DeepSeek went — 2.4 tokens/s and three silent bugs on RTX 50xx — is in &lt;a href="https://dev.to/helgard_orlm/running-the-510-gb-deepseek-v41-flash-on-an-8-gb-gpu-and-three-bugs-that-never-raise-an-error-5647"&gt;a separate post&lt;/a&gt;, code at &lt;a href="https://github.com/helgard-orlm/deepseek-v41-flash-8gb" rel="noopener noreferrer"&gt;https://github.com/helgard-orlm/deepseek-v41-flash-8gb&lt;/a&gt;.)&lt;/p&gt;

&lt;p&gt;Everything else (embeddings, attention, the 36 Gated DeltaNet layers, router, shared expert, lm_head) is about 7.2 GB in BF16. Storing the big DeltaNet matrices in FP8 with one scale per row brings the resident part to &lt;strong&gt;6.13 GiB of VRAM&lt;/strong&gt;. The 63 GiB of experts live on the NVMe drive.&lt;/p&gt;

&lt;h2&gt;
  
  
  The engine
&lt;/h2&gt;

&lt;p&gt;There was no runtime for this architecture that could stream experts, so it's a from-scratch PyTorch + Triton engine. The math was ported from &lt;code&gt;transformers&lt;/code&gt; 5.18 (&lt;code&gt;qwen4_exp&lt;/code&gt;): DeltaNet, full attention with a block indexer that kicks in beyond ~2048 tokens, four hyper-connection streams, and an n-gram embedding table (PLE, used in layer 1) with 51 billion parameters.&lt;/p&gt;

&lt;p&gt;The pieces, in the order they mattered:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Repack experts for the disk.&lt;/strong&gt; In the original files an expert's six tensors are scattered. We rewrote them into one 2,764,800-byte block per expert (675 pages of 4 KiB): &lt;code&gt;gate | up | down | gate_scale | up_scale | down_scale&lt;/code&gt;. One expert = one &lt;code&gt;pread&lt;/code&gt; with &lt;code&gt;O_DIRECT&lt;/code&gt; into page-aligned pinned memory = one copy to the GPU.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A RAM cache in front of the disk.&lt;/strong&gt; LRU over pinned memory (16 GB holds about a quarter of all experts), a pool of reader threads, and two banks of GPU slots so the next layer can load while this one computes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Prefetch into the gap.&lt;/strong&gt; After a layer's required reads are issued, guess the next layer's experts and read two of them while the GPU is busy (91% of guesses are used). The first version started the prefetch &lt;em&gt;at the same time&lt;/em&gt; as the required reads — they shared the disk and everything got slower. It has to go into the gap, not next to the real work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Own kernels.&lt;/strong&gt; An NVFP4 SwiGLU expert kernel for a single token using Blackwell's hardware &lt;code&gt;cvt.rn.f16x2.e2m1x2&lt;/code&gt; instruction (115 µs per layer, 3.6e-7 from a torch reference), and an FP8 row-scaled matrix-vector kernel. The generic unpack path for the FP8 DeltaNet matrices had been costing ~106 ms per token on its own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Don't read 66 GB per prompt chunk.&lt;/strong&gt; The first prompt path processed 512 tokens at a time, and each chunk needed nearly every expert, so the whole expert set was read again for every chunk. Running the whole prompt through each layer at once, with experts read in batches of 32, took a 2451-token prompt from &lt;strong&gt;71.6 s to 23.3 s&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Read only the PLE rows you need.&lt;/strong&gt; The n-gram table is a 53.7 GB file; each token needs a handful of rows. Random reads from NVMe: 0.3–0.5 ms per token. (On the HDD it would be a disk seek per row — a second AI acting as reviewer caught that the table was still on the HDD in the first run.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the time actually goes
&lt;/h2&gt;

&lt;p&gt;With every expert already in RAM, a token took 80–88 ms. Splitting it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;39 ms&lt;/strong&gt; — Python issuing GPU work (the GPU was waiting on Python, not the other way round);&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;47 ms&lt;/strong&gt; — moving 1.33 GB of experts over PCIe (measured 28.6 GB/s, PCIe 5.0 ×8);&lt;/li&gt;
&lt;li&gt;and the two ran &lt;strong&gt;one after the other&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are 52 CPU↔GPU syncs per token (mostly reading the router's chosen experts). We expected them to be the problem; measured, they are cheap. They only hurt because the expert copy for layer L can't start until layer L's router has answered. That pointed to three changes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Speculative copies.&lt;/strong&gt; Run layer L+1's router on layer L's MoE input. It picks the right top-10 65% of the time — enough to start most copies one layer early. (Prefetching top-16 to raise the hit rate moved more bytes over the bus than it saved.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CUDA graphs for the 36 one-token DeltaNet layers&lt;/strong&gt; (Codex). One shared memory pool: 36 private pools ran out of memory on an 8 GB card. Graphs are dropped before any prompt ≥1024 tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin the decode thread to a P-core.&lt;/strong&gt; On a hybrid CPU the unpinned thread wandered between P and E cores: 93–108 ms per token, jumping around. Pinned to core 0: a flat 81.6 ms.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;configuration&lt;/th&gt;
&lt;th&gt;tokens/s&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;first version, experts from disk, no cache&lt;/td&gt;
&lt;td&gt;2.68&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;disk only + prefetch&lt;/td&gt;
&lt;td&gt;4.72–4.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RAM cache 16 GB + prefetch&lt;/td&gt;
&lt;td&gt;8.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;live chat, 1000-token answer, before the last three changes&lt;/td&gt;
&lt;td&gt;9.71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;live chat, 1000-token answer, after&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;11.65&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Prompt processing: 24 tokens in 4.5 s, 2451 in 23.3 s, 7182 in 98 s. Follow-up messages are fast because the server snapshots the whole recurrent state (DeltaNet matrices + KV + indexer keys) in RAM and only processes the new tail.&lt;/p&gt;

&lt;h2&gt;
  
  
  How we know it's right
&lt;/h2&gt;

&lt;p&gt;Speed without correctness is a random number generator, so every change was checked against the &lt;code&gt;transformers&lt;/code&gt; reference fed with the &lt;strong&gt;same&lt;/strong&gt; quantized weights:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;next-token argmax: &lt;strong&gt;32/32&lt;/strong&gt;, mean |Δlogit| 0.077;&lt;/li&gt;
&lt;li&gt;a 2451-token prompt, long enough for the attention indexer to start selecting blocks: &lt;strong&gt;11/11&lt;/strong&gt;, and the needle-in-a-haystack fact comes back right;&lt;/li&gt;
&lt;li&gt;perplexity EN 2.21 / RU 2.22 — plus control runs that &lt;em&gt;must&lt;/em&gt; break: with nibbles swapped it's 2016 / 33,269, with experts removed 588 / 5,351. That checks the NVFP4 unpacking independently of the reference.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What didn't work
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A fused DeltaNet kernel&lt;/strong&gt; (one kernel for the recurrence, 4.3 → 0.43 ms per token across 36 layers). 5 of 256 tokens came out different from the reference. Rejected — 4 ms isn't worth a model that says different things.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 8192-token context didn't actually work&lt;/strong&gt;, in the old version either: a 7182-token prompt ran out of VRAM in attention. Fixed by sizing prompt pieces so that &lt;code&gt;piece × (position + piece) ≤ 3072²&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A crash only after long prompts:&lt;/strong&gt; CUDA graph capture in the default &lt;code&gt;global&lt;/code&gt; mode was invalidated by expert-reader threads calling &lt;code&gt;event.synchronize()&lt;/code&gt;. &lt;code&gt;capture_error_mode="thread_local"&lt;/code&gt; fixed it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The repository has the engine, the server (OpenAI-compatible, streaming, a thinking-mode model id), the repack and reference-check tools, and a &lt;code&gt;setup.sh&lt;/code&gt; that prepares a working directory from a Hugging Face snapshot. Before publishing, we ran it the way a stranger would: a fresh &lt;code&gt;git clone&lt;/code&gt; of the repo on the same box, &lt;code&gt;setup.sh&lt;/code&gt; into an empty directory (15 min 48 s, mostly reading ~120 GB from an HDD; spot check of 204 repacked experts against the originals: 0 differences), then the server from the clone. The repacked experts and the extracted non-expert weights came out &lt;strong&gt;byte-identical&lt;/strong&gt; to the production directory, and three greedy test prompts gave identical answers from both.&lt;/p&gt;

&lt;p&gt;You need a Blackwell GPU (the NVFP4 kernel uses the sm_120a &lt;code&gt;e2m1&lt;/code&gt; instruction), a fast NVMe drive with ~130 GB free, and patience for a first-time setup that reads most of the 133 GB.&lt;/p&gt;

&lt;p&gt;The bigger point: on a mixture-of-experts model, "does it fit in VRAM" is the wrong question. The right one is &lt;strong&gt;bytes per token&lt;/strong&gt;, and a home PC with a fast drive can serve a lot more of them than it looks.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Engine, kernels, server and this write-up: Claude (Anthropic). CUDA graphs and the speculation/pinning A/B: Codex (OpenAI). Goal, model choice and decisions: me. All numbers are from logs on the machine described above.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>cuda</category>
      <category>python</category>
    </item>
    <item>
      <title>An AI hardware advisor that walks every path — and gets its math checked by the rules in Sanity</title>
      <dc:creator>Helgard</dc:creator>
      <pubDate>Sat, 03 Oct 2026 05:49:32 +0000</pubDate>
      <link>https://dev.to/helgard_orlm/an-ai-hardware-advisor-that-walks-every-path-and-gets-its-math-checked-by-the-rules-in-sanity-21db</link>
      <guid>https://dev.to/helgard_orlm/an-ai-hardware-advisor-that-walks-every-path-and-gets-its-math-checked-by-the-rules-in-sanity-21db</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path One: Ship an Agent That Queries Real Content&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;An agent that answers "what computer do I need to run &lt;em&gt;this&lt;/em&gt; AI model?" — including "don't buy, rent" and&lt;br&gt;
"you already have enough". It reads a &lt;strong&gt;Sanity dataset of criteria&lt;/strong&gt;, not a catalog: laws (formulas for weights&lt;br&gt;
memory, KV cache, speed ceiling, own-vs-rent breakeven), machine-checkable rules, 11 &lt;strong&gt;solution paths&lt;/strong&gt;&lt;br&gt;
(keep what you have · upgrade · used · one GPU · multi-GPU · unified memory · MoE experts in system RAM ·&lt;br&gt;
CPU-only small model · smaller model · cloud API · rented GPU), run reports and reference hardware.&lt;br&gt;
New models appear every week, so anything not in the base is looked up on the web &lt;strong&gt;by the base's own checklist&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sanity project: &lt;code&gt;onwa0wvs&lt;/code&gt;, public dataset &lt;code&gt;v2&lt;/code&gt; (200 documents), Studio: &lt;a href="https://hw-for-ai-lab-v2.sanity.studio" rel="noopener noreferrer"&gt;https://hw-for-ai-lab-v2.sanity.studio&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Agent connection: &lt;strong&gt;Sanity Context MCP&lt;/strong&gt; on the full dataset with &lt;strong&gt;embeddings enabled&lt;/strong&gt;
(&lt;code&gt;/context/mcp/onwa0wvs/v2&lt;/code&gt;); all numbers below are measured through this connection&lt;/li&gt;
&lt;li&gt;Second entry: a &lt;strong&gt;Knowledge Base&lt;/strong&gt; (&lt;code&gt;kbW7wsbtJQkl&lt;/code&gt;, built from the core 138 documents of the same dataset, refreshed weekly)
behind its own Context MCP endpoint in KB mode — created and connected, &lt;strong&gt;not measured separately&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demo — try it with your own agent:&lt;/strong&gt; &lt;a href="https://hw-advisor.helgardorlm.tech" rel="noopener noreferrer"&gt;https://hw-advisor.helgardorlm.tech&lt;/a&gt; (MCP &lt;code&gt;https://hw-advisor.helgardorlm.tech/mcp&lt;/code&gt;, no login) · &lt;strong&gt;Replay of all 27 measured runs:&lt;/strong&gt; &lt;a href="https://helgard-orlm.github.io/sanity-ai-hardware-advisor/" rel="noopener noreferrer"&gt;https://helgard-orlm.github.io/sanity-ai-hardware-advisor/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Repo: &lt;a href="https://github.com/helgard-orlm/sanity-ai-hardware-advisor" rel="noopener noreferrer"&gt;https://github.com/helgard-orlm/sanity-ai-hardware-advisor&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live — bring your own agent (no login, no keys):&lt;/strong&gt; &lt;a href="https://hw-advisor.helgardorlm.tech" rel="noopener noreferrer"&gt;https://hw-advisor.helgardorlm.tech&lt;/a&gt; · MCP endpoint &lt;code&gt;https://hw-advisor.helgardorlm.tech/mcp&lt;/code&gt;&lt;br&gt;
Replay of all 27 measured runs: &lt;a href="https://helgard-orlm.github.io/sanity-ai-hardware-advisor/" rel="noopener noreferrer"&gt;https://helgard-orlm.github.io/sanity-ai-hardware-advisor/&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/helgard-orlm/sanity-ai-hardware-advisor" rel="noopener noreferrer"&gt;https://github.com/helgard-orlm/sanity-ai-hardware-advisor&lt;/a&gt; (the public proxy with &lt;code&gt;check_answer&lt;/code&gt; is in &lt;a href="https://github.com/helgard-orlm/sanity-ai-hardware-advisor/tree/main/public_mcp" rel="noopener noreferrer"&gt;&lt;code&gt;public_mcp/&lt;/code&gt;&lt;/a&gt;)&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Project ID &lt;strong&gt;&lt;code&gt;onwa0wvs&lt;/code&gt;&lt;/strong&gt;, public dataset &lt;strong&gt;&lt;code&gt;v2&lt;/code&gt;&lt;/strong&gt; — 200 documents, 15 types (law, rule, solutionPath, aiModel, gpu, cpu, runReport, offer, cloudOffer, situationTemplate, …)&lt;/li&gt;
&lt;li&gt;Query it directly: &lt;a href="https://onwa0wvs.api.sanity.io/v2025-01-01/data/query/v2?query=*%5B_type==%22law%22%5D" rel="noopener noreferrer"&gt;https://onwa0wvs.api.sanity.io/v2025-01-01/data/query/v2?query=*[_type=="law"]&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Studio: &lt;a href="https://hw-for-ai-lab-v2.sanity.studio" rel="noopener noreferrer"&gt;https://hw-for-ai-lab-v2.sanity.studio&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it with your own agent (no keys, no login)
&lt;/h2&gt;

&lt;p&gt;The demo is a public MCP endpoint, not a hosted chatbot: &lt;strong&gt;you bring the agent, the base brings the knowledge and the checking.&lt;/strong&gt;&lt;br&gt;
Add &lt;code&gt;https://hw-advisor.helgardorlm.tech/mcp&lt;/code&gt; to ChatGPT (developer mode), Claude, Claude Code, Codex, Cursor or VS Code.&lt;br&gt;
There are one-click buttons for Cursor and VS Code on the page. Or paste one line into any agent that can run commands:&lt;br&gt;
&lt;em&gt;"Read &lt;a href="https://hw-advisor.helgardorlm.tech/setup.md" rel="noopener noreferrer"&gt;https://hw-advisor.helgardorlm.tech/setup.md&lt;/a&gt; and follow it."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Behind it is a ~230-line stdlib proxy. It passes the four &lt;strong&gt;Sanity Context MCP&lt;/strong&gt; tools through unchanged and keeps the token&lt;br&gt;
server-side. It puts today's date and the advisor's work order in front of &lt;code&gt;initial_context&lt;/code&gt;. It adds one tool, &lt;strong&gt;&lt;code&gt;check_answer&lt;/code&gt;&lt;/strong&gt;,&lt;br&gt;
which is the same validator the app uses: it recomputes every calculation with the &lt;code&gt;law.formula&lt;/code&gt; stored in Sanity.&lt;br&gt;
In the first outside test (Codex, its own subscription), &lt;code&gt;check_answer&lt;/code&gt; rejected the first draft with 4 errors. The second draft passed:&lt;br&gt;
5 calculations recomputed, 11/11 solution paths walked. A test from ChatGPT (developer mode) also went through to a passing check. Agents that can only browse get the same tools as plain HTTPS links&lt;br&gt;
(&lt;code&gt;/api/context&lt;/code&gt;, &lt;code&gt;/api/query?q=…&lt;/code&gt;, &lt;code&gt;/api/check&lt;/code&gt;), listed in &lt;code&gt;/llms.txt&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Known limit:&lt;/strong&gt; the check covers numbers, completeness, sources and identity. It does not cover whether a verdict makes sense. In one test an agent said &lt;em&gt;yes&lt;/em&gt; to "MoE experts in RAM" for a dense model, and the check passed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;Sanity Context MCP tools the agent calls: &lt;code&gt;initial_context&lt;/code&gt;, &lt;code&gt;groq_query&lt;/code&gt;, &lt;code&gt;schema_explorer&lt;/code&gt;, &lt;code&gt;array_field_reader&lt;/code&gt; (plus our &lt;code&gt;check_answer&lt;/code&gt;). What it does with the content it retrieves — the structure is used, not just searched:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Paths are data.&lt;/strong&gt; The agent must give a yes/no/maybe with a number for &lt;em&gt;every&lt;/em&gt; &lt;code&gt;solutionPath&lt;/code&gt; document.
Before this, the same model answered a 321B-MoE question with "4×H200" and never mentioned the
one-GPU + large-RAM path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rules are the search checklist.&lt;/strong&gt; For any model/card/CPU that is not in the base, the app requires the
fields that the base's &lt;code&gt;rule&lt;/code&gt; documents test for that type (e.g. &lt;code&gt;cpu.instructionSets&lt;/code&gt;, &lt;code&gt;aiModel.variants.fileGb&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Laws are recomputed by code.&lt;/strong&gt; The agent writes every calculation with its inputs into a hidden check block;
the app evaluates the &lt;code&gt;law.formula&lt;/code&gt; from Sanity and sends the answer back if the number is wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The person's situation is stored&lt;/strong&gt;, not remembered: country, budget, existing PC survive a change of mind.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The same code checks the &lt;strong&gt;base itself&lt;/strong&gt;: every &lt;code&gt;law&lt;/code&gt; variable and &lt;code&gt;rule&lt;/code&gt; path must be a real schema field&lt;br&gt;
(coverage check), and laws get property tests. That caught the filler model writing an own-vs-rent formula that&lt;br&gt;
&lt;em&gt;added&lt;/em&gt; electricity instead of subtracting it and billed rent for 24 h/day — &lt;strong&gt;12.3 vs 110 months&lt;/strong&gt; on the same inputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  A real trace (GLM-5.3-Flash, not in the base) — &lt;a href="https://helgard-orlm.github.io/sanity-ai-hardware-advisor/?sc=A2_not_in_base&amp;amp;run=1" rel="noopener noreferrer"&gt;open it in the replay&lt;/a&gt;
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;GROQ over the base → no &lt;code&gt;aiModel&lt;/code&gt; document → 8 web searches &lt;strong&gt;by the base's own checklist&lt;/strong&gt;: the model card and
&lt;code&gt;config.json&lt;/code&gt; (layers, KV heads, head dim), file sizes, and the GPU fields the base's rules test (slots, bandwidth, compute capability).&lt;/li&gt;
&lt;li&gt;Laws from Sanity: weights &lt;code&gt;320 × 4 / 8 = 160 GB&lt;/code&gt; (NVFP4), KV for 8k context 22.5 GiB (flagged as rough for hybrid attention).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validator round 1 sent the answer back:&lt;/strong&gt; the speed ceiling was written as &lt;strong&gt;19.91&lt;/strong&gt; tok/s — code evaluated the
&lt;code&gt;law.formula&lt;/code&gt; from Sanity with the agent's own inputs and got &lt;strong&gt;199.1&lt;/strong&gt; (a 10× slip); plus three entities missing
fields that the base's rules check (e.g. &lt;code&gt;variants.fileGb&lt;/code&gt;, &lt;code&gt;slots&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Round 2: no errors. All 11 solution paths get a verdict — multi-GPU = yes, MoE-in-RAM / unified memory = maybe, cloud = no
(the person ruled it out) — and the build is 3 × RTX PRO 6000 (288 GB) with the numbers above.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Other catches by the same recomputation in the measured run: KV cache 0.949 → 6.75 GiB (A4), monthly electricity cost 0.675 → 2.025 (A5).&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured, honestly
&lt;/h2&gt;

&lt;p&gt;Blind grader (a different model, sees only the dialogue and a frozen checklist), 9 scenarios × 3 runs,&lt;br&gt;
78 checklist items, same base state for all arms:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;version&lt;/th&gt;
&lt;th&gt;score&lt;/th&gt;
&lt;th&gt;median time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;v2: base + instructions&lt;/td&gt;
&lt;td&gt;47/78&lt;/td&gt;
&lt;td&gt;62 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v3: + solution paths, validator&lt;/td&gt;
&lt;td&gt;53/78&lt;/td&gt;
&lt;td&gt;104 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;v3.1 via &lt;strong&gt;Sanity Context MCP&lt;/strong&gt;: + rules-as-checklist, recomputed laws, dated prices&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;59/78&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;181 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What this does &lt;strong&gt;not&lt;/strong&gt; show: that structured retrieval beats plain text for reading. The same 200 documents&lt;br&gt;
pasted as text notes scored the same as v3 (53/78) — at this size the model reads either. Where structure&lt;br&gt;
mattered in our runs is &lt;strong&gt;checking&lt;/strong&gt;: fewer first answers rejected by the validator with structured queries&lt;br&gt;
(8/27 vs 14/27), and the gains of v3.1 are exactly in the scenarios where code recomputes Sanity's laws&lt;br&gt;
(KV for 8 users 7/9 → 9/9, own vs rent 5/12 → 7/12, "don't buy" 4/6 → 6/6, v2 → v3.1). The price: answers take ~3× longer.&lt;br&gt;
An earlier full run of v3.1 scored 62/78 with two bugs in my own checker (nested fields, European number format) that sent&lt;br&gt;
correct answers back for rework; after fixing them the clean run scored 59/78 — with N=3 per scenario, 59 vs 62 is noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it was built
&lt;/h2&gt;

&lt;p&gt;Architect method: a &lt;strong&gt;risk catalog&lt;/strong&gt; (16 classes of agent failure seen in earlier sessions) → acceptance scenarios&lt;br&gt;
written &lt;em&gt;before&lt;/em&gt; changes → build → blind grading. The base was filled by a model (GPT Luna) from an empty schema and&lt;br&gt;
two prompts; I never wrote data by hand. An independent judge model watched the session and corrected the method&lt;br&gt;
several times (frozen checklist, snapshot baseline, text control).&lt;/p&gt;

&lt;h2&gt;
  
  
  Known limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Answers are slow (minutes) because the validator can send them back up to twice.&lt;/li&gt;
&lt;li&gt;Ratings N=3 per scenario; differences of 1–2 points are noise.&lt;/li&gt;
&lt;li&gt;"unknown after search" is allowed and can be abused for vague entities.&lt;/li&gt;
&lt;li&gt;In the measured run 12 of 27 answers came out in Russian to English questions (an account language setting leaked in);
the app now checks the answer's language in code and sends it back (3/3 English on a re-test).&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>We create a way to unload Qwen2.5 KV cache to RAM.</title>
      <dc:creator>Helgard</dc:creator>
      <pubDate>Thu, 11 Jun 2026 20:14:41 +0000</pubDate>
      <link>https://dev.to/helgard_orlm/we-create-a-way-to-unload-qwen25-kv-cache-to-ram-57fo</link>
      <guid>https://dev.to/helgard_orlm/we-create-a-way-to-unload-qwen25-kv-cache-to-ram-57fo</guid>
      <description>&lt;p&gt;Two separate problems with local LLMs, one mechanism fixes both:&lt;/p&gt;

&lt;p&gt;Long context doesn't fit. The KV-cache grows linearly with every token; on 8 GB it dies around 110k.&lt;br&gt;
Memory doesn't survive a restart. Kill the process and the cache is gone — next session re-reads everything from scratch (that's all RAG is: re-reading text, every query).&lt;br&gt;
What the KV-cache actually is: as the model reads, each layer stores a key/value vector per token — this is the model's working state, the thing it attends over to pick the next token. It's not text, it's computed internal state. Normally it lives entirely in VRAM and only exists while the process is alive.&lt;/p&gt;

&lt;p&gt;The mechanism:&lt;/p&gt;

&lt;p&gt;Keep only the last N tokens' KV in VRAM (the hot window).&lt;br&gt;
Stream everything older to CPU RAM (8-bit), then disk. It's not deleted, just moved somewhere 10× cheaper.&lt;br&gt;
Keep a tiny index in VRAM: one vector per sentence. On a query, that index finds the few relevant chunks and copies only those KV vectors back into VRAM for the attention step.&lt;br&gt;
So VRAM stays flat — it holds the hot window plus a small retrieval budget, regardless of whether total context is 50k or 800k. The cap moves from VRAM to your RAM stick.&lt;/p&gt;

&lt;p&gt;The part that makes it memory and not just offload: you write that cold KV to disk and reload it in a different process days later. The model resumes from its own stored state — a few ms of memcpy, no re-reading the source. Two things make the positions line up across the gap: keys are stored already rotary-rotated at their absolute position, and the query is injected at its true position, so relative distances stay exact even between tokens 700k apart.&lt;/p&gt;

&lt;p&gt;Result (RTX 5060 8 GB, 4-bit): a planted fact is the top prediction from 120k–500k tokens, top-2 at 800k, VRAM flat ~6 GB — where a normal full cache OOMs at 65k. The live store behind my assistant is 2.5M tokens of past sessions in 467 MB, answered in ~1s on CPU. Miss → LOW CONFIDENCE, not a hallucination.&lt;/p&gt;

&lt;p&gt;Why Qwen2.5-7B-1M and not anything else — two independent requirements, most models fail one:&lt;/p&gt;

&lt;p&gt;Architecture must keep a growing KV. Mamba/SSM and linear-attention models (Granite-4, RWKV, Qwen3.5's hybrid layers) compress history into one fixed-size state — there's no per-token KV to stream out or pull back. The whole method is inapplicable to them.&lt;br&gt;
The model must actually read at that range. Effective context is ~50–70% of the advertised number (RULER). MiniCPM-1B (128k nominal) goes to garbage by ~120k even with a perfect full cache in VRAM — that's the model's limit, not the cache's. No infrastructure fixes it.&lt;br&gt;
Qwen2.5-7B-1M passes both and fits 8 GB in 4-bit. Substrate alone OOMs at 65k; the model alone can't persist across restarts. The combination is the whole point.&lt;/p&gt;

&lt;p&gt;Limits: confidence drops with distance (p 0.98 → 0.21 at 800k); many near-identical facts confuse the cheap index (need full attention there); on messy logs generation can override the retrieved KV with its own prior, so I quote stored text verbatim instead. Research PoC — the core (spilling KV to CPU) has precedent (ArkVale, RetrievalAttention), credited in the README.&lt;/p&gt;

&lt;p&gt;Code (MIT): &lt;a href="https://github.com/helgard-orlm/living-kv-cache" rel="noopener noreferrer"&gt;https://github.com/helgard-orlm/living-kv-cache&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>performance</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
