<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Arsen Apostolov</title>
    <description>The latest articles on DEV Community by Arsen Apostolov (@sikamikanikobg).</description>
    <link>https://dev.to/sikamikanikobg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1410108%2Fb7d644e9-449a-4ef9-8bcc-73a2ff63902f.jpeg</url>
      <title>DEV Community: Arsen Apostolov</title>
      <link>https://dev.to/sikamikanikobg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sikamikanikobg"/>
    <language>en</language>
    <item>
      <title>Qwen3.8-Flash-Next (125B) on Three RTX 3090s at 80 Tokens/s: Teaching llama.cpp Which Experts Matter</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Sat, 26 Sep 2026 16:02:31 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/i-ran-a-125b-model-on-three-rtx-3090s-at-80-tokenss-by-teaching-llamacpp-which-experts-matter-4ioi</link>
      <guid>https://dev.to/sikamikanikobg/i-ran-a-125b-model-on-three-rtx-3090s-at-80-tokenss-by-teaching-llamacpp-which-experts-matter-4ioi</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Qwen3.8-Flash-Next is a 125B mixture-of-experts model: on paper it beats the Qwen3.8-27B I run in production everywhere, by +16.5 points on agentic coding. It does not fit in the 72 GB of VRAM my three RTX 3090s have, and stock llama.cpp, spilling a quarter of the experts into system RAM, runs it at &lt;strong&gt;23 tokens/second&lt;/strong&gt;. Three changes get it to &lt;strong&gt;80 tokens/second, entirely in VRAM&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Sort the experts by how often they're used.&lt;/strong&gt; They are wildly unequal: the busiest 25% handle 52% of the work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Store them at three precisions.&lt;/strong&gt; Popular experts keep more bits and rare ones get squeezed harder, so the whole model fits on the GPUs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Patch llama.cpp&lt;/strong&gt; (~400 lines) so one routing decision drives three expert tables at once, then add speculative decoding on top.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Quality stays close to the 8-bit reference: &lt;strong&gt;+1.9% perplexity, 91% same next token, GSM8K 95.5%&lt;/strong&gt; (the 27B scores 95–96.5% on the same harness). There's also a 256k-context profile that finds a random code hidden 173,000 tokens deep. The trade-offs are real, and I've listed them. Everything is open: &lt;a href="https://github.com/SikamikanikoBG/qwen38-flash-next-3x3090" rel="noopener noreferrer"&gt;patches, tools, raw measurements&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Update, 26 Sep 2026, after the first day behind my assistant:&lt;/strong&gt; the 80 tok/s holds for short prompts. In real agent use, with ~25k-token prompts, sampling and thinking on, the model decodes at &lt;strong&gt;30–45 tok/s&lt;/strong&gt;. My first deployment also had a caching mistake that cost ~45 s before every answer; a server config change fixed it (now &lt;strong&gt;0.7 s&lt;/strong&gt;). Details in the "real-world update" section below.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  the use case
&lt;/h2&gt;

&lt;p&gt;My box, &lt;strong&gt;vader&lt;/strong&gt;, runs &lt;a href="https://dev.to/sikamikanikobg/qwen38-27b-on-2x-rtx-3090-my-first-local-model-i-actually-trust-fpi"&gt;Qwen3.8-27B&lt;/a&gt; around the clock: two RTX 3090s, vLLM, about 135 tokens/second. It's the first local model I actually trust with agent work, and it does most of the work my assistant does.&lt;/p&gt;

&lt;p&gt;Then Qwen shipped &lt;strong&gt;Qwen3.8-Flash-Next&lt;/strong&gt;, the preview of their Qwen4 architecture. Their own model card puts it ahead of my 27B on almost everything I care about:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;benchmark (Qwen's card)&lt;/th&gt;
&lt;th&gt;Flash-Next&lt;/th&gt;
&lt;th&gt;3.8-27B&lt;/th&gt;
&lt;th&gt;gap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE (agentic coding)&lt;/td&gt;
&lt;td&gt;58.7&lt;/td&gt;
&lt;td&gt;42.2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+16.5&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JobBench (professional tasks)&lt;/td&gt;
&lt;td&gt;55.7&lt;/td&gt;
&lt;td&gt;33.4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+22.3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Multilingual&lt;/td&gt;
&lt;td&gt;81.0&lt;/td&gt;
&lt;td&gt;73.8&lt;/td&gt;
&lt;td&gt;+7.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolathlon (tool use)&lt;/td&gt;
&lt;td&gt;73.5&lt;/td&gt;
&lt;td&gt;67.1&lt;/td&gt;
&lt;td&gt;+6.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HLE&lt;/td&gt;
&lt;td&gt;35.9&lt;/td&gt;
&lt;td&gt;30.8&lt;/td&gt;
&lt;td&gt;+5.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA Diamond&lt;/td&gt;
&lt;td&gt;91.7&lt;/td&gt;
&lt;td&gt;89.2&lt;/td&gt;
&lt;td&gt;+2.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IFBench&lt;/td&gt;
&lt;td&gt;81.3&lt;/td&gt;
&lt;td&gt;79.5&lt;/td&gt;
&lt;td&gt;+1.8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So the goal was simple to state: &lt;strong&gt;the smarter model, on the same three cards, with decode and time-to-first-token that still feel interactive.&lt;/strong&gt; No new hardware. I told my AI engineer (Claude Code, more on that at the end) that "no" was not an acceptable answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  the box
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPUs:&lt;/strong&gt; 3x RTX 3090, 24 GB each: 72 GB of VRAM. No NVLink, and no PCIe peer-to-peer (the board has no Resizable BAR), so the cards can't talk to each other directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CPU:&lt;/strong&gt; dual Xeon E5-2660 v4, 56 threads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAM:&lt;/strong&gt; 60 GB. Only &lt;strong&gt;one&lt;/strong&gt; 32 GB stick per CPU socket, so it's single-channel, and that matters a lot below.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disk:&lt;/strong&gt; one 1 TB NVMe.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  a small dictionary before we start
&lt;/h2&gt;

&lt;p&gt;If you know what a KV cache is, skip this. If you don't, it's all you need for the rest of the article.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Parameters / weights.&lt;/strong&gt; The numbers a model learned. "125B" means 125 billion of them. Each one normally takes 2 bytes, so 125B parameters is ~250 GB, far more than 72 GB of VRAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VRAM.&lt;/strong&gt; The GPU's own memory. It's about 20–50x faster than system RAM &lt;em&gt;on this box&lt;/em&gt;. A model runs fast only if what it needs for each word lives in VRAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantization.&lt;/strong&gt; Storing each weight with fewer bits (8, 4, 3…) instead of 16. Like saving a photo as a smaller JPEG: smaller file, a little less detail. "Q8_0", "Q4_K", "IQ3_S", "MXFP4" are llama.cpp's names for different bit-budgets and recipes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mixture of experts (MoE).&lt;/strong&gt; Instead of one giant network, each layer has many small "expert" networks, 512 of them here, and a &lt;strong&gt;router&lt;/strong&gt; picks the best 10 for each word. So the model &lt;em&gt;knows&lt;/em&gt; 125B parameters' worth of things but only &lt;em&gt;computes&lt;/em&gt; with about 6B per word. That's why it can be fast, and also why it's huge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token.&lt;/strong&gt; A word or word-piece. &lt;strong&gt;Decode speed&lt;/strong&gt; (tokens/second) is how fast the answer streams out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prefill / time to first token (TTFT).&lt;/strong&gt; Before answering, the model reads your whole prompt. That's prefill, and TTFT is how long you wait before the first word appears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context / KV cache.&lt;/strong&gt; How much text the model can hold in mind at once ("256k context" is about 500 pages), and the memory it uses to remember it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;n-gram embedding.&lt;/strong&gt; New in this model: a 51 GB lookup table indexed by 2–3 word phrases. It's only ever &lt;em&gt;looked up&lt;/em&gt;, a few rows per word, so it can live on the NVMe drive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speculative decoding / MTP.&lt;/strong&gt; A small built-in "draft head" guesses the next few words, and the big model checks all the guesses in one pass. Correct guesses are free words.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;KL divergence (KLD).&lt;/strong&gt; How far the compressed model's word probabilities drift from the original's. 0 means identical; lower is better. It's the most sensitive quality meter there is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;llama.cpp / GGUF.&lt;/strong&gt; The inference engine I patched, and its model-file format.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  the problem, in one picture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fobry7yo68to60yzsyojk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fobry7yo68to60yzsyojk.png" alt="Where Qwen3.8-Flash-Next (125B) lives: before and after" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In 16-bit the model is 360 GB. Even the popular 4-bit build (&lt;a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF" rel="noopener noreferrer"&gt;unsloth's UD-Q4_K_XL&lt;/a&gt;) is 111 GB: 72 GB of experts, 27 GB of n-gram table (which can stay on disk) and a few GB of everything else. The experts alone don't fit in 72 GB of VRAM once the rest of the model and the context need room too. So llama.cpp does the sensible thing: it puts the overflow experts in system RAM and computes them on the CPU.&lt;/p&gt;

&lt;p&gt;On this machine that means &lt;strong&gt;23 tokens/second&lt;/strong&gt;. With one memory stick per socket, the CPU reads those experts about ten times slower than a GPU would. And for every word, the work bounces between the GPUs and the CPU dozens of times, so each GPU waits.&lt;/p&gt;

&lt;h2&gt;
  
  
  step 1: the obvious setup, and the first trap
&lt;/h2&gt;

&lt;p&gt;The first run with stock settings also showed prefill all over the place (47 to 306 tokens/second on the same 4k prompt) and decode sagging to 10 tok/s. Thread states told the story: the main thread sat in &lt;code&gt;D&lt;/code&gt; (waiting on disk). llama.cpp memory-maps the model file, loading the GPU part streams 70+ GB through the page cache, and the CPU-side experts kept getting evicted and re-read from NVMe.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;--load-mode none&lt;/code&gt; loads the CPU-side weights into RAM and keeps them there. Also, &lt;strong&gt;don't download a 188 GB file while you benchmark&lt;/strong&gt;: my download filled the page cache and pushed the server into swap. With both fixed, prefill doubled (&lt;strong&gt;658–792 tok/s&lt;/strong&gt;) and decode steadied at &lt;strong&gt;23–26 tok/s&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Speculative decoding (MTP) barely helped: 24–29 tok/s. The reason matters for everything that follows. A draft of 5 words means the big model checks 6 words at once. Each word picks 10 experts, so one check can touch up to 60 &lt;em&gt;different&lt;/em&gt; experts per layer instead of 10, and the slow ones in CPU RAM get hit harder. &lt;strong&gt;As long as experts live in CPU RAM, nothing else will make this fast.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  step 2: the experts are not equal
&lt;/h2&gt;

&lt;p&gt;Here's the observation the whole project rests on. llama.cpp's calibration data (the &lt;em&gt;importance matrix&lt;/em&gt; unsloth publishes with their quants) records how often each of the 24,576 experts was picked on real text:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38hgpedoun8y1is3uj7x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38hgpedoun8y1is3uj7x.png" alt="Routing skew: a few experts do most of the work" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In the average layer, &lt;strong&gt;the busiest 25% of experts handle 52% of the tokens, and the busiest 80% handle 95%.&lt;/strong&gt; Stock llama.cpp can only place &lt;em&gt;whole layers&lt;/em&gt; of experts: a layer's 512 experts are one tensor, all on the GPU or all on the CPU. It offloads a quarter of the experts, including popular ones, so a quarter of the expert work lands on the slow path.&lt;/p&gt;

&lt;h2&gt;
  
  
  step 3: split every layer into hot and cold
&lt;/h2&gt;

&lt;p&gt;The first patch splits every layer's expert tensor in two at load time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;hot experts&lt;/strong&gt; go to the GPUs;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cold experts&lt;/strong&gt; go to CPU RAM;&lt;/li&gt;
&lt;li&gt;the &lt;strong&gt;router&lt;/strong&gt; is permuted so its choices land in the right half.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With the same 19% of expert bytes on the CPU, the cold experts now serve only &lt;strong&gt;4.5% of routed tokens&lt;/strong&gt; instead of ~19%, about 4x less slow-path traffic.&lt;/p&gt;

&lt;p&gt;Getting it right took four bugs, each instructive enough to list in its own section below. The result is &lt;strong&gt;30–32 tok/s&lt;/strong&gt; with correct answers. That's +30%, but it's not the leap it should be, so I profiled it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;where one decoded word spends its time&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;kernel launches&lt;/td&gt;
&lt;td&gt;~4,300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dense weights (attention, DeltaNet, shared expert), 8-bit&lt;/td&gt;
&lt;td&gt;~4.9 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;routed experts&lt;/td&gt;
&lt;td&gt;~2.5 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;~3,500 tiny element-wise kernels&lt;/td&gt;
&lt;td&gt;~6–7 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU↔CPU synchronisations&lt;/td&gt;
&lt;td&gt;~390&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The CPU was barely reading any experts anymore; it was the &lt;em&gt;round trips&lt;/em&gt; that cost. Every layer handed work to the CPU and waited for the answer. Then a cheap experiment: drop the cold experts entirely (quality garbage, speed only) and see how fast it goes fully on GPU: &lt;strong&gt;52 tok/s&lt;/strong&gt;, and &lt;strong&gt;86–88 tok/s&lt;/strong&gt; with speculative decoding. That was the target. The question became: &lt;strong&gt;how do you fit every expert in VRAM without wrecking quality?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  step 4: three shelves, all on the GPU
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnbzg98y08zf6u37wtx0l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnbzg98y08zf6u37wtx0l.png" alt="How one word gets written" width="800" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The idea: if some experts do most of the work, &lt;strong&gt;give them more bits and squeeze the rarely-used ones harder&lt;/strong&gt;. Everything still fits in VRAM, and most tokens still see a well-preserved expert. Think of a library that keeps its bestsellers in hardcover and prints the rarely-borrowed titles as compact paperbacks, so the whole collection fits in the building.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2apbxoc17r1j93ih4vb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2apbxoc17r1j93ih4vb.png" alt="What I built" width="800" height="235"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two pieces make it work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An offline re-packer&lt;/strong&gt; (&lt;a href="https://github.com/SikamikanikoBG/qwen38-flash-next-3x3090/blob/main/tools/expert_tiers.py" rel="noopener noreferrer"&gt;&lt;code&gt;tools/expert_tiers.py&lt;/code&gt;&lt;/a&gt;). It starts from the 8-bit model (188 GB) and uses the importance matrix for both &lt;em&gt;how often&lt;/em&gt; each expert is used and &lt;em&gt;which inputs matter&lt;/em&gt; inside it. A greedy planner fills a VRAM budget: each byte goes to the expert where it removes the most expected error. Every expert is then re-quantized with llama.cpp's own quantizers, and the file is written with the experts reordered hottest-first as three tensors per layer. One wrinkle: the down-projection matrices have rows 640 wide, which the fancy 2–3-bit formats can't handle (they need multiples of 256). So they get their own ladder: IQ4_NL / MXFP4 instead of IQ4_XS / IQ3.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A patched llama.cpp&lt;/strong&gt; (&lt;a href="https://github.com/SikamikanikoBG/qwen38-flash-next-3x3090/tree/main/patches" rel="noopener noreferrer"&gt;&lt;code&gt;patches/&lt;/code&gt;&lt;/a&gt;). &lt;code&gt;mul_mat_id&lt;/code&gt;, the operation that runs "each word through its chosen experts", learns to take an &lt;strong&gt;id range&lt;/strong&gt;. Each shelf's matrix multiply sees the router's full choice list, computes only the ids that fall on its shelf, and writes zeros for the rest. The three results simply add up. That meant touching the CPU kernels, three CUDA paths (decode, prefill, and the fused gate+up+activation kernel) and the loader. On top sits Qwen's multi-token-prediction draft head, from a &lt;a href="https://github.com/ggml-org/llama.cpp/pull/28243" rel="noopener noreferrer"&gt;not-yet-merged llama.cpp PR&lt;/a&gt;, re-quantized to 4 bits so it fits too.&lt;/p&gt;

&lt;p&gt;Result, in one chart:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsjgtle8cia7p3qxlcm81.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsjgtle8cia7p3qxlcm81.png" alt="Decode journey" width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  the four bugs, briefly
&lt;/h2&gt;

&lt;p&gt;For people who'll try this, and because two of them are delightful.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate experts crash the CUDA MoE kernel.&lt;/strong&gt; My first version pointed "not on this shelf" at expert 0. A word with two cold experts then had expert 0 twice, and CUDA's expert-grouping helper assumes each expert appears at most once per word. It silently dropped a slot, and a later kernel read an unwritten index. Fixed by teaching the kernels a real "skip".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A flag stored where a precision setting lives.&lt;/strong&gt; I marked the skip-capable nodes in &lt;code&gt;op_params[0]&lt;/code&gt;, which llama.cpp already uses for the accumulator precision. The flag was quietly overwritten. It moved to slot 7, with a magic value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The unsigned ternary.&lt;/strong&gt; &lt;code&gt;ids ? ids[i] : blockIdx.x&lt;/code&gt;, where &lt;code&gt;blockIdx.x&lt;/code&gt; is unsigned. C++ promotes the whole expression to unsigned, so my &lt;code&gt;-1&lt;/code&gt; ("skip") became 4,294,967,295 and the kernel read 4 billion rows past the buffer. &lt;code&gt;compute-sanitizer&lt;/code&gt; found it in one run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fused kernels hide the tag.&lt;/strong&gt; For decode, CUDA fuses gate, up and activation into one kernel whose "destination" is the activation node, not the matrix multiply that carries my id-range tag. The first id-range build printed &lt;code&gt;//////////&lt;/code&gt;. The kernel now looks up the tag on its source node.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  how much quality did it cost?
&lt;/h2&gt;

&lt;p&gt;The honest meter is KL divergence against the 8-bit model over 24,576 tokens of Wikipedia text:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscypq9ckgugggwtb9keq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscypq9ckgugggwtb9keq.png" alt="Quality vs size" width="800" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;model&lt;/th&gt;
&lt;th&gt;expert size&lt;/th&gt;
&lt;th&gt;fits in VRAM&lt;/th&gt;
&lt;th&gt;perplexity vs 8-bit&lt;/th&gt;
&lt;th&gt;KLD&lt;/th&gt;
&lt;th&gt;same top token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;unsloth UD-Q4_K_XL (stock)&lt;/td&gt;
&lt;td&gt;71.7 GiB&lt;/td&gt;
&lt;td&gt;no, 25% in RAM&lt;/td&gt;
&lt;td&gt;+0.9%&lt;/td&gt;
&lt;td&gt;0.045&lt;/td&gt;
&lt;td&gt;93.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;tiered, 128k profile&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;53.0 GiB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+1.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.091&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.1%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;uniform IQ3_S / MXFP4, same size&lt;/td&gt;
&lt;td&gt;52.2 GiB&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;+3.0%&lt;/td&gt;
&lt;td&gt;0.095&lt;/td&gt;
&lt;td&gt;91.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tiered, 256k profile&lt;/td&gt;
&lt;td&gt;49.5 GiB&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;+3.2%&lt;/td&gt;
&lt;td&gt;0.113&lt;/td&gt;
&lt;td&gt;90.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What I take from it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fitting in VRAM costs quality, and there's no way around that.&lt;/strong&gt; 20 GB less expert weight roughly doubles the KL divergence of the stock 4-bit build. That's physics, not a bug.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At equal size, tiering beats uniform&lt;/strong&gt;: perplexity drift 1.9% vs 3.0%. But the gain is smaller than my planner predicted: its error model is crude, and at this size the total byte count dominates. There's more to win with a better tier recipe.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On a task, it holds up.&lt;/strong&gt; GSM8K (200 grade-school maths questions, greedy, no thinking) scores &lt;strong&gt;95.5%&lt;/strong&gt;. The 27B's published number on the same harness is 95.0–96.5%. GSM8K is saturated, so it mainly proves nothing broke. The model card's gains on agentic work are where the upgrade should show.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  speed vs context, and the 256k profile
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftaeq91ejcnvjjlalewln.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftaeq91ejcnvjjlalewln.png" alt="Speed vs context" width="800" height="383"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Both profiles start at &lt;strong&gt;~80 tokens/second&lt;/strong&gt; on a short prompt and slow down as the prompt grows. The per-word attention and indexer work grows with context, and the draft head's guesses get a little worse too. Typical numbers, averaging two runs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;prompt&lt;/th&gt;
&lt;th&gt;128k profile: decode / time to first token&lt;/th&gt;
&lt;th&gt;256k profile: decode / time to first token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;4k tokens&lt;/td&gt;
&lt;td&gt;82 tok/s / 7–9 s&lt;/td&gt;
&lt;td&gt;67 tok/s / 7–8 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15k tokens&lt;/td&gt;
&lt;td&gt;66 tok/s / 18–22 s&lt;/td&gt;
&lt;td&gt;58 tok/s / 19–22 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30k tokens&lt;/td&gt;
&lt;td&gt;58 tok/s / 35 s&lt;/td&gt;
&lt;td&gt;54 tok/s / 38 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;61k tokens&lt;/td&gt;
&lt;td&gt;53 tok/s / 76–79 s&lt;/td&gt;
&lt;td&gt;51 tok/s / 82 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;95k / 122k tokens&lt;/td&gt;
&lt;td&gt;39 tok/s / 140 s&lt;/td&gt;
&lt;td&gt;35 tok/s / 190 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Prefill runs at &lt;strong&gt;540–860 tokens/second&lt;/strong&gt;. That's the weak spot next to the 27B's ~1,300, and it's where I'd look next: it's also the part that makes a 120k-token prompt a three-minute wait.&lt;/p&gt;

&lt;p&gt;The 256k profile needed two more compromises. The KV cache drops to 8-bit, and experts shrink to 49.5 GiB, because the sparse-attention indexer's scratch memory grows with context and pushed a card out of memory at ~200k on the first try. After rebalancing layers across the cards, it read &lt;strong&gt;173,692 tokens of Wikipedia with a random 10-character code hidden at 45% depth, and returned the code: exactly right (&lt;code&gt;6BC5XYE8FS&lt;/code&gt;). A second run with the code at 90% depth of 112,711 tokens was also exact.&lt;/strong&gt; It took 6.1 minutes (472 tokens/second) to read that much. Long-context prefill on this architecture is still the slow part.&lt;/p&gt;

&lt;h2&gt;
  
  
  what it costs to run
&lt;/h2&gt;

&lt;p&gt;Measured at the cards (&lt;code&gt;nvidia-smi&lt;/code&gt;, 5 Hz, three 1,024-token generations per profile): the three 3090s draw &lt;strong&gt;~540 W together while decoding&lt;/strong&gt; and ~124 W at rest with the model loaded. That's &lt;strong&gt;7.5–8.4 joules per token&lt;/strong&gt;, 2.1–2.3 kWh per million tokens. On my tariff (0.30 BGN/kWh day, 0.18 night) &lt;strong&gt;a million generated tokens cost 0.63–0.70 BGN in daytime, 0.38–0.42 BGN at night&lt;/strong&gt;, about $0.35–0.40. &lt;a href="https://dev.to/sikamikanikobg/qwen38-27b-on-one-rtx-3090-vs-two-20-decode-14-cold-prefill-and-3x-on-cached-prompts-55kc"&gt;The 27B costs 0.21–0.34 BGN&lt;/a&gt; for the same million, so the bigger brain costs roughly twice as much per word, all three cards included. It's still coffee money.&lt;/p&gt;

&lt;h2&gt;
  
  
  real-world update: day one behind my assistant
&lt;/h2&gt;

&lt;p&gt;Benchmarks are one request with a short prompt and greedy decoding. My assistant is none of those things, so after a day of it running Jarvis, here is what the server logs say about real requests:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;a real assistant turn&lt;/th&gt;
&lt;th&gt;the benchmark&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;prompt&lt;/td&gt;
&lt;td&gt;~25,000 tokens (system prompt, tools, memory)&lt;/td&gt;
&lt;td&gt;41 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sampling&lt;/td&gt;
&lt;td&gt;temperature 0.7, thinking on&lt;/td&gt;
&lt;td&gt;greedy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tokens per verify step (draft acceptance)&lt;/td&gt;
&lt;td&gt;2.6 (39%)&lt;/td&gt;
&lt;td&gt;3.4 (62%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;decode&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;41–45 tok/s&lt;/strong&gt; (30–45 is what it feels like)&lt;/td&gt;
&lt;td&gt;80 tok/s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Decode: 30–45 tok/s, not 80.&lt;/strong&gt; Two things compound. The prompt is deep, and every decoded word pays attention over it; the context chart above already shows ~50 tok/s at 60k. And the draft head guesses sampled text, tool-call JSON and reasoning worse than greedy prose, so speculative decoding saves less. That's the honest number for agent work on this box. For comparison, the 27B benchmarks at ~135 tok/s on this box, so the upgrade is intelligence, not speed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Time to first token: the mistake was mine.&lt;/strong&gt; My assistant runs several &lt;em&gt;roles&lt;/em&gt;: chat, planner, classifier, triage, background jobs. Each has its own ~25k-token prompt prefix. I deployed the server with one slot (&lt;code&gt;-np 1&lt;/code&gt;), so every time the role changed, the cached prefix was thrown away and &lt;strong&gt;25,000 tokens were re-read from scratch: ~45 seconds before every answer&lt;/strong&gt;. The fix is four slots sharing one KV pool, so every role keeps its prefix warm:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;returning to a 14k-token prompt after another role ran&lt;/th&gt;
&lt;th&gt;before (&lt;code&gt;-np 1&lt;/code&gt;)&lt;/th&gt;
&lt;th&gt;after (&lt;code&gt;-np 4 -kvu&lt;/code&gt;)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;time to first token&lt;/td&gt;
&lt;td&gt;20–45 s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.6–0.7 s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It isn't free: four slots need a little more memory, so the shared pool dropped from 128k to 96k tokens (128k with four slots ran out of VRAM). A single request can still use all 96k.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Schedule the background jobs apart.&lt;/strong&gt; A model at 30–45 tok/s with 25k-token prompts spends minutes per agent task. My assistant had 13 scheduled jobs between 06:30 and 09:00 on Mondays, four of them at the same minute. I've spread them 30 minutes apart (05:00 to 11:30) so they stop queueing behind each other. I'll report after a week whether that's enough.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Images and video work too,&lt;/strong&gt; once the vision projector is loaded (&lt;code&gt;--mmproj&lt;/code&gt;, on the GPU with the most free memory: 4 s per photo, against 61 s on the CPU) and the container has &lt;code&gt;ffmpeg&lt;/code&gt; for video (a 5-second clip in 6.4 s). The repo has the exact launcher.&lt;/p&gt;

&lt;h2&gt;
  
  
  what I'd not claim
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It's single-user.&lt;/strong&gt; Every speed number here is one request at a time, and real agent turns decode at 30–45 tok/s, not 80. The 27B on vLLM batches many users far better.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 27B is still faster.&lt;/strong&gt; About 135 vs 80 tok/s decode, and roughly 1,300 vs 540–860 tok/s prefill. The 125B is smarter, not quicker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It uses all three cards.&lt;/strong&gt; You can't also keep the 27B running.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The quality numbers are wikitext + GSM8K.&lt;/strong&gt; They say the compression is gentle. They don't prove the agentic gains survive intact; that needs agent benchmarks I haven't run yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The tier planner is a heuristic.&lt;/strong&gt; The measurements say it helps; a calibrated one would help more.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  reproduce it
&lt;/h2&gt;

&lt;p&gt;Everything is in &lt;strong&gt;&lt;a href="https://github.com/SikamikanikoBG/qwen38-flash-next-3x3090" rel="noopener noreferrer"&gt;SikamikanikoBG/qwen38-flash-next-3x3090&lt;/a&gt;&lt;/strong&gt;: both llama.cpp patches, the re-packer, the benchmark scripts, and every raw number behind the charts. The short version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# llama.cpp at 81bc6b8 + MTP PR #28243 + the tiering patch&lt;/span&gt;
git apply patches/0001-qwen4exp-mtp-pr28243.patch patches/0002-tiered-experts-hot-cold-split.patch
cmake &lt;span class="nt"&gt;-B&lt;/span&gt; build &lt;span class="nt"&gt;-DGGML_CUDA&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ON &lt;span class="nt"&gt;-DCMAKE_CUDA_ARCHITECTURES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;86 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; cmake &lt;span class="nt"&gt;--build&lt;/span&gt; build &lt;span class="nt"&gt;-j&lt;/span&gt;

&lt;span class="c"&gt;# re-pack the 8-bit model into three precision shelves (53 GiB of experts)&lt;/span&gt;
python tools/expert_tiers.py write &lt;span class="nt"&gt;--src&lt;/span&gt; Qwen3.8-Flash-Next-Q8_0-00001-of-00006.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--imatrix&lt;/span&gt; imatrix_unsloth.gguf &lt;span class="nt"&gt;--budget-gib&lt;/span&gt; 53 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--gu&lt;/span&gt; Q6_K,IQ4_XS,IQ3_XXS &lt;span class="nt"&gt;--dn&lt;/span&gt; Q8_0,IQ4_NL,MXFP4 &lt;span class="nt"&gt;--out&lt;/span&gt; fn-tier53.gguf

&lt;span class="c"&gt;# serve: all layers on GPU, 4 slots sharing a 96k KV pool, vision, MTP drafting 4 tokens&lt;/span&gt;
llama-server &lt;span class="nt"&gt;-m&lt;/span&gt; fn-tier53-00001-of-00002.gguf &lt;span class="nt"&gt;-ngl&lt;/span&gt; 99 &lt;span class="nt"&gt;-ts&lt;/span&gt; 16,16,16 &lt;span class="nt"&gt;-c&lt;/span&gt; 98304 &lt;span class="nt"&gt;-np&lt;/span&gt; 4 &lt;span class="nt"&gt;-kvu&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-b&lt;/span&gt; 1024 &lt;span class="nt"&gt;-ub&lt;/span&gt; 256 &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;--jinja&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--mmproj&lt;/span&gt; mmproj-F16.gguf &lt;span class="nt"&gt;-mmdev&lt;/span&gt; CUDA1 &lt;span class="nt"&gt;--image-max-tokens&lt;/span&gt; 1024 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-md&lt;/span&gt; mtp-shared-iq4.gguf &lt;span class="nt"&gt;--spec-type&lt;/span&gt; draft-mtp &lt;span class="nt"&gt;--spec-draft-n-max&lt;/span&gt; 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  credits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Qwen&lt;/strong&gt;, for the model and an architecture that rewards this kind of work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;unsloth&lt;/strong&gt;, for the GGUFs, the MTP draft files and the importance matrix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;llama.cpp and ggml&lt;/strong&gt;, plus the authors of the &lt;a href="https://github.com/ggml-org/llama.cpp/pull/28243" rel="noopener noreferrer"&gt;Qwen4Exp MTP PR&lt;/a&gt; and the &lt;a href="https://github.com/ggml-org/llama.cpp/issues/28734" rel="noopener noreferrer"&gt;sparse-attention decay issue&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/syv-ai/qwen38-27b-rtx3090" rel="noopener noreferrer"&gt;syv-ai/qwen38-27b-rtx3090&lt;/a&gt;&lt;/strong&gt;, whose rigor on the 27B set the bar for measuring this.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;This project, from the patches and the tools to the measurements and this write-up, was done with **Claude Code&lt;/em&gt;* acting as my AI engineer, working on my hardware under my direction. The numbers are from vader, the raw data is in the repo, and the mistakes are both of ours.*&lt;/p&gt;

</description>
      <category>llm</category>
      <category>qwen</category>
      <category>homelab</category>
      <category>ai</category>
    </item>
    <item>
      <title>Qwen3.8-27B on One RTX 3090 vs Two: +20% Decode, +14% Cold Prefill, and 3x on Cached Prompts</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Wed, 23 Sep 2026 03:27:37 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/qwen38-27b-on-one-rtx-3090-vs-two-20-decode-14-cold-prefill-and-3x-on-cached-prompts-55kc</link>
      <guid>https://dev.to/sikamikanikobg/qwen38-27b-on-one-rtx-3090-vs-two-20-decode-14-cold-prefill-and-3x-on-cached-prompts-55kc</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Qwen3.8-27B — the dense 27.8B that dropped on August 14 — fits on a single RTX 3090 at W4A16 and decodes at &lt;strong&gt;~125–155 tokens/sec&lt;/strong&gt;. Splitting it across two cards with tensor parallelism buys &lt;strong&gt;roughly 20% more decode&lt;/strong&gt; (noisy: +13% in one run, +22% in the next) and &lt;strong&gt;14% faster prefill on average&lt;/strong&gt; on a cold prompt (23% at the top of the ladder: an 8,600-token prompt takes &lt;strong&gt;8.5 s&lt;/strong&gt; to first token on one card and &lt;strong&gt;6.9 s&lt;/strong&gt; on two). On a prompt the server has already seen, the pair answers &lt;strong&gt;2–3x faster&lt;/strong&gt; — but that is the prefix cache, not the cards, and the script below measures the two separately. At the API list price of $3.00 per million output tokens, one million tokens cost &lt;strong&gt;0.21–0.34 BGN of electricity, about $0.12–0.20&lt;/strong&gt;, on the box that serves them.&lt;/p&gt;

&lt;h2&gt;
  
  
  the box
&lt;/h2&gt;

&lt;p&gt;vader is the machine I run this on: three RTX 3090s in an ASUS Z10PE-D8 WS, dual Xeon E5-2660 v4 (56 threads), 60 GB of RAM, Ubuntu 26.04. It has been my inference box since the Whisper experiments, and since July it has also carried the power-limiting work (&lt;a href="https://dev.to/sikamikanikobg/what-does-a-local-llm-actually-cost-per-month-i-read-the-meters-1274"&gt;the meters post&lt;/a&gt; is where those numbers live). Two of the three cards are in this test; the third sits idle.&lt;/p&gt;

&lt;p&gt;The model under test is the one that is currently loaded on it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;single-card&lt;/th&gt;
&lt;th&gt;TP2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;container&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen38-27b-single-1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;qwen38-27b-tp2-single-1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;image&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ghcr.io/syv-ai/qwen38-27b-rtx3090&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;same, &lt;code&gt;:latest&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;weights&lt;/td&gt;
&lt;td&gt;Qwen3.8-27B-W4A16-AutoRound-fast&lt;/td&gt;
&lt;td&gt;same&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cards&lt;/td&gt;
&lt;td&gt;1x 3090 (22.6 GB VRAM)&lt;/td&gt;
&lt;td&gt;2x 3090 (21.7 GB each)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;max context&lt;/td&gt;
&lt;td&gt;131,072&lt;/td&gt;
&lt;td&gt;262,144&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;port&lt;/td&gt;
&lt;td&gt;18021&lt;/td&gt;
&lt;td&gt;18022&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both are vLLM, both have been up two days, both report healthy, and both run with &lt;strong&gt;prefix caching on&lt;/strong&gt; — which is why every number below comes in two flavours, cold and warm. The W4A16 quantization is what puts a 27.8B dense model inside one 24 GB card — 22.6 GB of it, with the KV cache budget for the rest of the 128k window coming out of what is left.&lt;/p&gt;

&lt;h2&gt;
  
  
  the method
&lt;/h2&gt;

&lt;p&gt;No benchmark harness, no &lt;code&gt;vllm bench&lt;/code&gt; — a 150-line script (&lt;a href="https://github.com/SikamikanikoBG/qwen38-bench" rel="noopener noreferrer"&gt;&lt;code&gt;bench_qwen38_v2.py&lt;/code&gt;&lt;/a&gt;) that sends a real chat-completion request with a prompt of a given size, streams the answer, and times it. For each prompt size it runs three times and takes the median:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;prompt sizes:&lt;/strong&gt; calibrated so the server reports ~600, ~2,400, ~4,700, ~9,300 prompt tokens (the script probes the tokenizer first, then scales the prompt to hit the target)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;output:&lt;/strong&gt; &lt;code&gt;max_tokens=256&lt;/code&gt;, &lt;code&gt;temperature=0&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;measured:&lt;/strong&gt; time to first SSE event, total time, and the server-reported token counts from &lt;code&gt;stream_options.include_usage&lt;/code&gt; — not my own chunk counting&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cold vs warm:&lt;/strong&gt; every cold request carries a unique first line, so no cached prefix can match it; the warm pass repeats one identical request after a warm-up call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last bullet is the one that decides whether a benchmark means anything. With prefix caching on — and it is on by default in vLLM — timing the same prompt twice does not measure prefill the second time, it measures a cache lookup. The tell is arithmetic: 8,600 tokens "prefilled" in 0.56 s is 15,000 tokens/sec, roughly ten times what two 3090s can actually push through a 27B model. If your TTFT numbers look like that, you are timing the cache. So every cold request here is salted with a unique first line, and the warm pass is run deliberately and reported in its own column.&lt;/p&gt;

&lt;p&gt;One caveat: this is a single-request benchmark. vLLM's continuous batching is where production throughput lives, and I did not hammer it with concurrent requests. What I measured is the latency you feel as the one person using the box — which, for a homelab, is most of the time.&lt;/p&gt;

&lt;h2&gt;
  
  
  decode: the second card buys 15–30%, noisily
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyn1hypir2dc3rti8678h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyn1hypir2dc3rti8678h.png" alt="Decode speed by prompt size, cold runs" width="799" height="395"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The single card holds &lt;strong&gt;123–154 tokens/sec&lt;/strong&gt; across the ladder; the TP2 pair runs &lt;strong&gt;146–193&lt;/strong&gt;. A 27B dense model at W4A16 on one 3090 is memory-bandwidth-bound in a way that does not care much about prompt size, because the KV cache for these sizes is small next to the weights. The pair is faster at every size but one (at 4,700 tokens the two came out level), and the average gain was +22% in this run against +13% in a run an hour earlier. Single-stream decode on these cards moves by about 10% between runs an hour apart, so I am not going to print a decimal: &lt;strong&gt;call it 15–30%, and expect the lower end.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is a real number and a small one. For a 27B model, tensor parallelism across two cards is not the 2x you get from adding compute — the all-reduce between the cards eats most of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  prefill: cold is the cards, warm is the cache
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkheov6y1214phf5k1xx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkheov6y1214phf5k1xx.png" alt="Time-to-first-token by prompt size, cold and warm" width="800" height="398"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two columns, two different questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cold&lt;/strong&gt; — a prompt the server has never seen — is prefill, and prefill is compute-bound: the single card does &lt;strong&gt;1,100–1,220 tokens/sec&lt;/strong&gt; of it, the pair &lt;strong&gt;1,230–1,360&lt;/strong&gt;. Time to first token climbs from 0.53 s to &lt;strong&gt;8.5 s&lt;/strong&gt; on one card as the prompt goes from 600 to 9,300 tokens, and from 0.50 s to &lt;strong&gt;6.9 s&lt;/strong&gt; on two. The gap grows with the prompt: 6% at 600 tokens, &lt;strong&gt;23% at 9,300&lt;/strong&gt;, 14% on average. That is what the second card buys you on a fresh long prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warm&lt;/strong&gt; — the same prompt again — is the prefix cache doing its job, and here the pair really is 2–3x faster: &lt;strong&gt;0.5–0.75 s&lt;/strong&gt; flat across the ladder against &lt;strong&gt;0.5–1.7 s&lt;/strong&gt; on the single card, whose smaller KV pool and slower re-read of the uncached tail show at every size above 600 tokens. It is a cache number rather than a prefill number, and it is also the one you feel most often, because every follow-up turn in a chat is a warm request.&lt;/p&gt;

&lt;p&gt;The routing rule that falls out: &lt;strong&gt;a long fresh prompt goes to the pair and saves about a fifth of its wait; a conversation stays on whichever lane it started on, so its prefix stays cached.&lt;/strong&gt; For prompts under about 2,000 tokens the single card is within noise of the pair on both counts and draws one card's power.&lt;/p&gt;

&lt;h2&gt;
  
  
  what it costs
&lt;/h2&gt;

&lt;p&gt;The box draws &lt;strong&gt;620 W&lt;/strong&gt; at the wall with both containers loaded and serving — 247 W on one TP2 card, 328 W on the other, 45 W on the single-card 3090 sitting at P8 with its weights resident. Those are the monitor's readings from the day the containers went up, not a measurement taken during these runs; the governor from the meters post is power-capping it, holding 79°C.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqunez6xef0ov2ri6kezf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqunez6xef0ov2ri6kezf.png" alt="API price vs local electricity cost" width="800" height="356"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At the API's list price — $3.00 per million output tokens, DeepInfra's rate for Qwen3.8-27B as listed on llm-stats — one million generated tokens takes about &lt;strong&gt;6,667 seconds&lt;/strong&gt; at 150 tokens/sec. At 620 W that is &lt;strong&gt;1.15 kWh&lt;/strong&gt;. On the Bulgarian dual-rate tariff that is &lt;strong&gt;0.34 BGN by day (0.30 BGN/kWh) or 0.21 BGN by night (0.18 BGN/kWh)&lt;/strong&gt; — roughly &lt;strong&gt;$0.20 or $0.12&lt;/strong&gt;. The API price is &lt;strong&gt;15x to 25x&lt;/strong&gt; the electricity, depending on when the box runs.&lt;/p&gt;

&lt;p&gt;I am not going to pretend that makes the API a bad deal. The API does not cost you 620 W of heat in your room, it does not need a 60 GB machine, it does not need you to keep a W4A16 quant on a 24 GB card, and its $3.00 buys you their datacenter's power, cooling and idle capacity, not yours. But for a box that is already on and already paying its standby tax, every token it serves is almost pure margin. The electricity argument for local inference is not "it is cheaper than the API" in any total-cost sense — it is "the marginal cost of a token on hardware you already run is an order of magnitude below the list price."&lt;/p&gt;

&lt;h2&gt;
  
  
  the verdict
&lt;/h2&gt;

&lt;p&gt;Qwen3.8-27B is the first 27B-class dense model I have run that I would put in front of real work on this box, and the numbers say why:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It fits.&lt;/strong&gt; W4A16 on one 3090, 22.6 GB, with room for a 128k context window. The TP2 variant doubles the context to 262k by splitting the weights.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is fast enough to feel like an API.&lt;/strong&gt; 125–190 tokens/sec decode is faster than most hosted endpoints feel in a chat UI. The part you notice is a cold long prompt: several seconds on either lane.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The two-card setup is not 2x.&lt;/strong&gt; It is roughly +20% decode and +14% cold prefill, and 2–3x on warm follow-ups because the pair keeps a bigger cache. Buy the second card for the context window and the cache, not for raw speed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The marginal cost is 0.21–0.34 BGN, about $0.12–0.20, per million tokens&lt;/strong&gt; against a $3.00 list price, on hardware that is already drawing its standby power.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark with the cache in mind.&lt;/strong&gt; If a prefill number looks like 15,000 tokens/sec on consumer cards, it is the cache. Salt every request, or turn prefix caching off, and report which one you did.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The script, the raw JSON and the charts are in the repo: &lt;a href="https://github.com/SikamikanikoBG/qwen38-bench" rel="noopener noreferrer"&gt;qwen38-bench&lt;/a&gt;. If you run a different quant or a different card, I want your numbers — the interesting part of this post is the shape of the curve, and the shape changes with the hardware.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;vader: 3x RTX 3090 (two used here), dual Xeon E5-2660 v4, 60 GB RAM, Ubuntu 26.04, vLLM, W4A16-AutoRound-fast quantization.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>qwen</category>
    </item>
    <item>
      <title>I Benchmarked Qwen3.8-27B on Three RTX 3090s. The Second Card Bought Me 13% More Decode — and a 3x Faster Prefill</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Sun, 20 Sep 2026 12:14:36 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/i-benchmarked-qwen38-27b-on-three-rtx-3090s-the-second-card-bought-me-13-more-decode-and-a-3x-2781</link>
      <guid>https://dev.to/sikamikanikobg/i-benchmarked-qwen38-27b-on-three-rtx-3090s-the-second-card-bought-me-13-more-decode-and-a-3x-2781</guid>
      <description></description>
      <category>llm</category>
      <category>qwen</category>
      <category>homelab</category>
      <category>ai</category>
    </item>
    <item>
      <title>What Does a Local LLM Actually Cost per Month? I Read the Meters.</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Sun, 20 Sep 2026 07:59:22 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/what-does-a-local-llm-actually-cost-per-month-i-read-the-meters-1274</link>
      <guid>https://dev.to/sikamikanikobg/what-does-a-local-llm-actually-cost-per-month-i-read-the-meters-1274</guid>
      <description>&lt;h1&gt;
  
  
  What Does a Local LLM Actually Cost per Month? I Read the Meters.
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;The Local LLM Lab — Part 5&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;One controlled experiment. One number. One verdict.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The question nobody answers in the local-LLM hype is the boring one: &lt;strong&gt;what does the electricity bill say?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not "how many tokens per second." Not "how many GB of VRAM." The bill. The one that arrives on the first of the month and is the only number that actually matters for the person paying for the machine.&lt;/p&gt;

&lt;p&gt;I've been running a local AI stack on an RTX 3090 for about three months now — Whisper transcription as a permanent service, an embedding model for a RAG pipeline, and a 27B chat model on a second machine. I instrumented the GPU with a power meter, set up a dual-rate tariff (day 0.30 BGN/kWh, night 0.18 BGN/kWh), and let the meter run.&lt;/p&gt;

&lt;p&gt;This is what the last 30 days actually cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPU:&lt;/strong&gt; NVIDIA RTX 3090, 24 GB, power limit 260 W, running WhisperX (a Whisper ASR webservice) as a permanent resident&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tariff:&lt;/strong&gt; dual-rate — 0.30 BGN/kWh day, 0.18 BGN/kWh night (22:00–06:00)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's on the GPU 24/7:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Whisper ASR webservice — avg 7.9 GB VRAM, peak 10.3 GB, resident&lt;/li&gt;
&lt;li&gt;nomic-embed-text (Ollama) — 308 MB VRAM, resident&lt;/li&gt;
&lt;li&gt;immich ML — negligible&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's on the second machine (Vader):&lt;/strong&gt; Qwen3.8-27B via vLLM, two RTX 3090s, on demand&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Whisper service is the interesting one. It's not a chatbot. It's a transcription endpoint that my own tools call whenever I record a voice memo, a meeting, or a podcast clip. It sits there all day, at rest, drawing power, waiting for audio.&lt;/p&gt;

&lt;p&gt;That's the honest shape of a "local AI stack" — not a GPU that's 100% utilised 24/7, but a GPU that's 0% utilised 99% of the time and 100% for a few seconds when it's actually doing work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The last 30 days cost €2.00 in electricity.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the entire GPU. Not the whole machine. Just the GPU, measured at the card, over 30 days, with the Whisper service resident 24/7 and the embedding model resident 24/7.&lt;/p&gt;

&lt;p&gt;Breakdown by service:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Service&lt;/th&gt;
&lt;th&gt;Avg W&lt;/th&gt;
&lt;th&gt;Energy (kWh)&lt;/th&gt;
&lt;th&gt;Cost (€)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Whisper ASR (3090)&lt;/td&gt;
&lt;td&gt;22 W&lt;/td&gt;
&lt;td&gt;13.24&lt;/td&gt;
&lt;td&gt;1.76&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;nomic-embed-text (Ollama)&lt;/td&gt;
&lt;td&gt;3 W&lt;/td&gt;
&lt;td&gt;1.75&lt;/td&gt;
&lt;td&gt;0.23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;immich ML&lt;/td&gt;
&lt;td&gt;~0 W&lt;/td&gt;
&lt;td&gt;0.01&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~25 W&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;15.03&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.00&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Whisper service is the big one, but even it is only 22 W on average. That's not a GPU under load. That's a GPU that's mostly idle with a model resident in VRAM, spiking to ~120 W for a few seconds when it transcribes something, and sitting at ~35 W the rest of the time.&lt;/p&gt;

&lt;p&gt;The embedding model costs 23 cents a month. That's less than a coffee.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the number is so small (and why that's the point)
&lt;/h2&gt;

&lt;p&gt;People who argue "local LLMs are expensive" usually mean one of two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The hardware cost amortised over time.&lt;/strong&gt; A 3090 costs €800–1,200 used. Amortised over 3 years, that's ~€25/month. That's real, and it's the dominant cost, not the electricity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A GPU that's pegged at 100% utilisation 24/7.&lt;/strong&gt; That's a training rig, not an inference stack. My GPU is at 0% utilisation on average. The 30-day average power draw is ~25 W. A 3090 under sustained LLM inference at 260 W would burn ~187 kWh/month and cost ~€23/month. That's a different machine doing a different job.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The honest framing for a home inference stack is: &lt;strong&gt;the GPU is a mostly-idle appliance that costs a few euros a month to keep warm, and spikes when you actually use it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The €2.00 number is the "keep it warm" cost. The spikes are the "actually use it" cost, and they're short enough that they barely move the monthly total.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the power trace actually looks like
&lt;/h2&gt;

&lt;p&gt;I pulled the 30-day power history at 2-hour resolution. The shape is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Baseline:&lt;/strong&gt; ~35 W, 24/7 (Whisper resident, idle)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spikes:&lt;/strong&gt; ~120 W, for a few seconds to a few minutes, when audio comes in&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Occasional higher spikes:&lt;/strong&gt; ~50–54 W sustained for longer stretches when there's a queue of transcriptions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One outlier:&lt;/strong&gt; a single 2-hour bucket at 121 W (a batch of long audio files)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 2-hour buckets that show 35 W are the honest baseline. The 2-hour buckets that show 121 W are the work. The average of all of them is 22 W for the Whisper service.&lt;/p&gt;

&lt;p&gt;That's the shape of a local inference stack: &lt;strong&gt;a low baseline with short, sharp spikes.&lt;/strong&gt; Not a sustained load.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison that matters
&lt;/h2&gt;

&lt;p&gt;Here's the comparison that actually answers "is local worth it":&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Monthly electricity cost&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;This stack (Whisper + embed, 24/7 resident)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;€2.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Measured, 30 days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3090 at 100% util, 24/7 (training rig)&lt;/td&gt;
&lt;td&gt;~€23&lt;/td&gt;
&lt;td&gt;Hypothetical, sustained load&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API transcription (Whisper API, ~10 hrs audio/month)&lt;/td&gt;
&lt;td&gt;~€14–29&lt;/td&gt;
&lt;td&gt;OpenAI Whisper API pricing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API LLM (GPT-4o, ~1M tokens/month)&lt;/td&gt;
&lt;td&gt;~€14&lt;/td&gt;
&lt;td&gt;List price&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The local stack costs &lt;strong&gt;less than a tenth&lt;/strong&gt; of the equivalent API spend, even before you factor in that the API spend is per-use and the local spend is a flat "keep it warm" cost that doesn't scale with usage.&lt;/p&gt;

&lt;p&gt;The electricity is not the cost. The hardware is the cost. But the electricity is small enough that it stops being the argument against.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the "is local worth it" question
&lt;/h2&gt;

&lt;p&gt;The electricity bill is not the reason to choose cloud over local. It's not even close. A 3090 running a resident inference stack costs ~€2/month in electricity. The hardware amortisation is ~€25/month. The cloud API equivalent is ~€14–29/month for comparable usage, and it scales with usage while the local cost doesn't.&lt;/p&gt;

&lt;p&gt;The real cost of local is the &lt;strong&gt;upfront hardware&lt;/strong&gt; and the &lt;strong&gt;time you spend keeping it running.&lt;/strong&gt; The electricity is a rounding error.&lt;/p&gt;

&lt;p&gt;If you're already paying for a GPU for other reasons (gaming, rendering, a second machine), the marginal cost of adding a local inference stack is essentially the electricity: &lt;strong&gt;a few euros a month.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A local LLM inference stack on a 3090 costs ~€2/month in electricity.&lt;/strong&gt; The hardware is the real cost, not the power. The electricity bill is small enough that it should stop being an argument in the cloud-vs-local debate.&lt;/p&gt;

&lt;p&gt;If your objection to local is "the electricity," the meter says otherwise. The meter says: &lt;strong&gt;a few euros a month, mostly idle, spikes when you use it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's not a cost problem. That's a hardware-cost problem, and it's a one-time problem, not a recurring one.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I measured this
&lt;/h2&gt;

&lt;p&gt;Every number in this article came out of my homelab monitor, not an estimate: a small self-hosted dashboard that reads the GPU's power draw at the card, tracks per-service VRAM, and prices the energy against my tariff. It's open source, and it exposes the same data through a read-only MCP server, so an AI agent can pull the numbers for you instead of you SSH-ing in to run &lt;code&gt;nvidia-smi&lt;/code&gt; by hand.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;SikamikanikoBG/homelab-monitor&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docs:&lt;/strong&gt; &lt;a href="https://sikamikanikobg.github.io/homelab-monitor/" rel="noopener noreferrer"&gt;sikamikanikobg.github.io/homelab-monitor&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to run your own "what does my stack cost" experiment, that's the tooling I'd point you at.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Local LLM Lab is a series of measured experiments on a home GPU stack. Every number in this article was pulled from a power meter on the card, not estimated. The next piece prices the same stack by € per 1,000 correct answers — the number that actually matters when the model is doing your work.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>hardware</category>
      <category>llm</category>
    </item>
    <item>
      <title>Qwen3.8-27B on 2 RTX 3090: My First Local Model I Actually Trust</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Tue, 08 Sep 2026 05:45:03 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/qwen38-27b-on-2x-rtx-3090-my-first-local-model-i-actually-trust-fpi</link>
      <guid>https://dev.to/sikamikanikobg/qwen38-27b-on-2x-rtx-3090-my-first-local-model-i-actually-trust-fpi</guid>
      <description>&lt;p&gt;For years "local LLM" meant a toy — a chatbot you played with on weekends, too slow or too dumb to trust with real work. This month, for the first time, I stopped worrying. My assistant, Jarvis, runs its daily tasks on a 27B model on my own hardware. On the tasks I actually run, it feels Sonnet-like — on some, better. Tests continue.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does — not chat, work
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Read ~500 Jira emails and built a live tracker of 20+ demands: status, UAT dates, declared FTE benefits, what's closed and why.&lt;/li&gt;
&lt;li&gt;A compliance table for the EU AI Act's image-marking rules — provider vs. deployer, the metadata-vs-visible-label split, with dates. It caught a term the legal team was misusing.&lt;/li&gt;
&lt;li&gt;A net-worth rollup from my own books, and a gold-position sizing question that came back as a plan: 5–10% of liquid, DCA'd over 2–3 tranches.&lt;/li&gt;
&lt;li&gt;The homelab itself: benchmarking models, watching the boxes, writing this post.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The point isn't any single task. It's that I stopped being careful. With an API, every prompt is a small cost and a small privacy decision. Locally, it's just a prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjkt10we888kqqy2k44fi.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjkt10we888kqqy2k44fi.jpg" alt="Qwen" width="800" height="241"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;        ┌────────────────────────────────┐
        │        ArDi (the hub)          │
        │  56-core · 64GB · 52 Docker    │
        │  Jarvis 2 core · Whisper ASR   │
        │  embeddings · homelab-monitor  │
        └───────────────┬────────────────┘
                        │  Tailscale
        ┌───────────────▼────────────────┐
        │       Vader (inference)        │
        │  vLLM · Qwen3.8-27B (262K ctx) │
        │  2× RTX 3090 · 48GB VRAM       │
        │  tensor-parallel · port 8010   │
        └────────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Qwen3.8-27B is dense 27B, Apache 2.0, 262K context, vision + video. I picked a dense 27B over a bigger MoE on purpose: one model that fits in 48 GB and is fast, not a 2.4T monster that needs a cluster. Under load Vader draws ~685 W at 80 °C. It's a server, not a laptop.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you're going to do this
&lt;/h2&gt;

&lt;p&gt;Four things I'd tell my past self:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;VRAM is the budget, not FLOPs.&lt;/strong&gt; A 27B dense model at Q8/Q4 fits two 24 GB cards with headroom. That's why I went dense: the math is simple — weights + KV cache ≤ 48 GB. A 2.4T MoE would need a cluster and a mortgage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split the box.&lt;/strong&gt; Inference on its own machine, everything else (the assistant core, ASR, embeddings, the monitor) on the hub, talking over Tailscale. When the GPU pegs at 100%, the rest of the house keeps working. One box doing everything is where "local AI" dies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;vLLM over Ollama for serving.&lt;/strong&gt; Ollama is fine for tinkering; vLLM with tensor parallelism is what makes 48 GB feel like one big GPU and what gets you real tokens/sec. Port 8010, OpenAI-compatible API — your existing clients just point at it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure the power, or you're guessing.&lt;/strong&gt; "It's a 685 W box" only matters if you can attribute that draw to the process and price it. I built &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;homelab-monitor&lt;/a&gt; for exactly that — VRAM and power per process, energy priced at your tariff. The whole hub ran 4.97 kWh this week: &lt;strong&gt;1.30 BGN / 7 days&lt;/strong&gt;. If you run more than one box, it's the thing I'd reach for first.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2xa8djp7ja5gwwu5c78l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2xa8djp7ja5gwwu5c78l.png" alt="vLLM" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;Tests continue. The hardest agentic chains still make me double-check. But for the daily 80%, it's reliable enough that I stopped caring. A local model you have to think about is a toy. One you don't is infrastructure.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Drafted with AI assistance. Hardware, cost and task data are from my own homelab, measured live.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>homelab</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Thu, 09 Jul 2026 12:37:52 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/does-a-second-gpu-increase-ollamas-context-window-quadro-p2000-rtx-3090-tested-5hbh</link>
      <guid>https://dev.to/sikamikanikobg/does-a-second-gpu-increase-ollamas-context-window-quadro-p2000-rtx-3090-tested-5hbh</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Short version: no.&lt;/strong&gt; I dropped a much older GPU (&lt;strong&gt;Quadro P2000, 5GB, Pascal, 2016&lt;/strong&gt;) next to an &lt;strong&gt;RTX 3090 (24GB, Ampere)&lt;/strong&gt; on the same box, ran the same context-length ladder (8K→128K) through Ollama and vLLM on &lt;code&gt;qwen3-coder:30B-A3B&lt;/code&gt;, and got &lt;strong&gt;zero extra usable context in either engine&lt;/strong&gt; — and a &lt;strong&gt;74% decode-speed hit&lt;/strong&gt; for the trouble. Ollama hits the identical &lt;code&gt;Chunk too big&lt;/code&gt; wall at ctx=65536 whether the P2000 is there or not. vLLM refuses tensor-parallel across the two cards entirely — not a VRAM problem, a flat compute-capability rejection (&lt;code&gt;Minimum capability: 75. Current capability: 61.&lt;/code&gt;) that fails in 40 seconds, before any memory profiling. And the one real, measured effect of adding the P2000 to Ollama: decode speed goes from &lt;strong&gt;76 → 19.5 tok/s&lt;/strong&gt; at ctx=49152 once the P2000 gets pulled in as an actual compute device.&lt;/p&gt;

&lt;p&gt;Full narrative version — the two-stage collapse, the prompt-cache validation bug caught mid-sweep, the CUDA13-silently-drops-Pascal finding — is on &lt;a href="https://medium.com/@arsen.apostolov" rel="noopener noreferrer"&gt;Medium&lt;/a&gt;.## The setup&lt;/p&gt;

&lt;p&gt;ardi (dual Xeon E5-2680 v4, 128GB RAM, openSUSE Leap) has a Quadro P2000 sitting in a second slot next to the RTX 3090 this whole series has run on so far. Same model as phase 1 (&lt;code&gt;qwen3-coder:30B-A3B&lt;/code&gt;), same box, four legs: {Ollama, vLLM} × {3090 only, 3090+P2000 tandem}, priced through &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt; against real GPU power draw.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama: same wall, extra tax
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;ctx&lt;/th&gt;
&lt;th&gt;3090 only decode tok/s&lt;/th&gt;
&lt;th&gt;tandem decode tok/s&lt;/th&gt;
&lt;th&gt;P2000 VRAM (tandem)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;8,192&lt;/td&gt;
&lt;td&gt;124.3&lt;/td&gt;
&lt;td&gt;122.0&lt;/td&gt;
&lt;td&gt;6 MB / 0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;24,576&lt;/td&gt;
&lt;td&gt;108.2&lt;/td&gt;
&lt;td&gt;70.0&lt;/td&gt;
&lt;td&gt;62 MB / 0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;32,768&lt;/td&gt;
&lt;td&gt;99.4&lt;/td&gt;
&lt;td&gt;61.0&lt;/td&gt;
&lt;td&gt;62 MB / 0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;49,152&lt;/td&gt;
&lt;td&gt;75.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3,580 MB / 55%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;td&gt;fatal: &lt;code&gt;Chunk too big&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;fatal: identical &lt;code&gt;Chunk too big&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two separate costs, not one: decode already falls behind at ctx=24576 while the P2000 is still basically idle (62MB, 0% util) — some scheduling overhead just from having a second visible device. Then the real collapse hits at ctx=49152, when the P2000 actually gets pulled into the compute path (3.58GB, 55% util) and decode craters to &lt;strong&gt;19.5 tok/s&lt;/strong&gt;. Same context ceiling either way, worse speed the whole way there.&lt;/p&gt;

&lt;h2&gt;
  
  
  vLLM: doesn't even get to try
&lt;/h2&gt;

&lt;p&gt;Expected failure mode going in: tensor-parallel splits the ~17GB AWQ checkpoint roughly in half, and the P2000's 5GB doesn't hold its ~8.5GB share. Actual failure, at ctx=8192, in 40 seconds, before any memory profiling:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ValueError: The quantization method auto_awq is not supported for the current GPU.
Minimum capability: 75. Current capability: 61.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AWQ's Marlin kernel needs compute capability 7.5+ (Turing and later). The P2000 is 6.1 (Pascal). Not a close VRAM call — a flat architectural exclusion, decided before capacity is even checked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bonus finding: Ollama's own CUDA13 build almost drops the P2000
&lt;/h2&gt;

&lt;p&gt;Boot log, before any of the above:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;skipping&lt;/span&gt; &lt;span class="n"&gt;CUDA&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt; — &lt;span class="n"&gt;compute&lt;/span&gt; &lt;span class="n"&gt;capability&lt;/span&gt; &lt;span class="n"&gt;not&lt;/span&gt; &lt;span class="n"&gt;in&lt;/span&gt; &lt;span class="n"&gt;compiled&lt;/span&gt; &lt;span class="n"&gt;architectures&lt;/span&gt;
&lt;span class="n"&gt;device&lt;/span&gt;=&lt;span class="s2"&gt;"Quadro P2000"&lt;/span&gt; &lt;span class="n"&gt;cc&lt;/span&gt;=&lt;span class="m"&gt;610&lt;/span&gt;
&lt;span class="n"&gt;archs&lt;/span&gt;=&lt;span class="s2"&gt;"[750 800 860 870 890 900 1000 1030 1100 1200 1210]"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Falls back to a legacy &lt;code&gt;cuda_v12&lt;/code&gt; runtime that does support Pascal — so it works, just via a path most people wouldn't notice without reading boot logs. This 2016 card is now old enough that modern quantized-inference stacks are starting to architecturally step around it, not just outrun it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What wasn't the point of this one
&lt;/h2&gt;

&lt;p&gt;Not claiming a second GPU is never worth it — a matched pair, or a smaller-but-newer card, is a different setup entirely. This was specifically: does &lt;em&gt;this&lt;/em&gt; 5GB Pascal card, next to &lt;em&gt;this&lt;/em&gt; 3090, on &lt;em&gt;these&lt;/em&gt; two engines, buy anything. Check compute capability against your quantization scheme before you do the VRAM math — it can end the conversation first.&lt;/p&gt;

&lt;p&gt;Every number above priced through &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt; — open source, MIT licensed — against ardi's real GPU power draw. Full write-up with all four charts and the mid-sweep debugging on &lt;a href="https://medium.com/@arsen.apostolov" rel="noopener noreferrer"&gt;Medium&lt;/a&gt;.What's the oldest card you've tried to tandem into a rig — did it actually pull weight, or did you just assume it was?&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ollama</category>
      <category>vllm</category>
      <category>gpu</category>
    </item>
    <item>
      <title>Whisper large-v3 VRAM Requirements: Why It Won't Fit on a 5GB GPU (and What We Tried Instead)</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Wed, 08 Jul 2026 03:29:30 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/whisper-large-v3-vram-requirements-why-it-wont-fit-on-a-5gb-gpu-and-what-we-tried-instead-1a18</link>
      <guid>https://dev.to/sikamikanikobg/whisper-large-v3-vram-requirements-why-it-wont-fit-on-a-5gb-gpu-and-what-we-tried-instead-1a18</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;whisper-large-v3&lt;/code&gt; &lt;strong&gt;OOMs on a 5GB GPU (Quadro P2000) at float16, int8_float16, and full int8&lt;/strong&gt; — before serving a single request. Root cause is architecture overhead (32-layer encoder-decoder, activations, CUDA context), not just weight size. Fine-tuned &lt;code&gt;whisper-tiny → base → small → small-v2&lt;/code&gt; on Common Voice Bulgarian instead: held-out WER improved from &lt;strong&gt;88.2% → 32.7%&lt;/strong&gt; across escalating model size, but never closed the gap to large-v3's &lt;strong&gt;27.3%&lt;/strong&gt;. A community &lt;code&gt;large-v3-turbo&lt;/code&gt; Bulgarian fine-tune claiming &lt;strong&gt;9.97% WER on FLEURS&lt;/strong&gt; scored &lt;strong&gt;31.2%&lt;/strong&gt; on our own held-out set — same ballpark as our own model, not the win the model card implied. Built a real dual-GPU nginx failover (P2000 = fine-tune, 3090 = large-v3) that worked correctly on deploy, then failed a real spontaneous-speech test badly enough to roll back to large-v3-only within ~5 seconds. Core finding: &lt;strong&gt;Common Voice read-aloud WER does not predict real assistant-use transcription quality.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;ardi has one RTX 3090 (24GB) doing LLM inference work, and a Quadro P2000 (5GB) that's sat idle for about two years. Jarvis, a self-hosted assistant, depends on Whisper for speech-to-text — testing showed only &lt;code&gt;large-v3&lt;/code&gt; handles Bulgarian well; smaller stock checkpoints are fine for English, not for a lower-resource language. &lt;code&gt;large-v3&lt;/code&gt; sits permanently loaded on the 3090, the same card needed for local LLM serving.&lt;/p&gt;

&lt;p&gt;Question: can the idle P2000 take Bulgarian transcription off the 3090's hands via a Bulgarian-specific fine-tune small enough to fit 5GB?&lt;/p&gt;

&lt;p&gt;(One naming note so the rest of this makes sense: the container running here is &lt;code&gt;whisper-asr-webservice&lt;/code&gt; wrapping &lt;code&gt;faster-whisper&lt;/code&gt; — not the separate WhisperX project, despite what I've been calling it internally for months.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt 1: does large-v3 just fit?
&lt;/h2&gt;

&lt;p&gt;Tested &lt;code&gt;large-v3&lt;/code&gt; on the P2000 at three precisions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;float16        -&amp;gt; OOM
int8_float16   -&amp;gt; OOM
int8           -&amp;gt; OOM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuuj62ek4tj3n88imy46l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuuj62ek4tj3n88imy46l.png" alt="Horizontal bar chart: VRAM used on the P2000 by whisper-tiny (0.32GB), the deployed small-v2 fine-tune (1.4GB), and a community turbo-bg model (4.15GB), all under a 5.0GB ceiling — versus large-v3 as a red hatched bar breaking through the ceiling, labeled OOM, needs ~8GB" width="800" height="337"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All three OOM before serving a request. Not a quantized-weight-size problem — the encoder-decoder's non-weight overhead (32 layers, activations, CUDA context) exceeds 5GB regardless of precision. &lt;code&gt;whisper-tiny&lt;/code&gt; loads at 318MB with no issue, ruling out a driver/compatibility problem. &lt;code&gt;medium&lt;/code&gt; (769M params) was the practical ceiling for raw model size — 3.87GB used, 1.2GB headroom — but a generic multilingual &lt;code&gt;medium&lt;/code&gt; isn't good enough for Bulgarian on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempts 2–4: escalating fine-tunes
&lt;/h2&gt;

&lt;p&gt;Fine-tuned on Mozilla Common Voice Bulgarian, on the 3090, via HuggingFace &lt;code&gt;transformers&lt;/code&gt; &lt;code&gt;Seq2SeqTrainer&lt;/code&gt;. Evaluated on the same 150 held-out test clips (never seen in training) for every model:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa6eda5cdddxzp12j0o7c.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa6eda5cdddxzp12j0o7c.png" alt="Grouped bar chart: WER by model, zero-shot vs fine-tuned, tiny through small-v2, with a dashed reference line at large-v3's 27.3%. Fine-tuned WER descends from 68.7% to 32.7% across the four models, never reaching the reference line" width="800" height="451"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;small-v2&lt;/code&gt; = same architecture as &lt;code&gt;small&lt;/code&gt;, retrained on train+other combined (6,739 rows vs 4,952) for 5 epochs. Validation WER by epoch: &lt;code&gt;32.17 → 28.99 → 28.21 → 28.21 → 28.44&lt;/code&gt; — flattened, then rose at epoch 5 (overfitting), so &lt;code&gt;load_best_model_at_end&lt;/code&gt; correctly kept the epoch 3/4 checkpoint rather than the final one. No more clean Bulgarian Common Voice data exists beyond train+other, so this is the practical ceiling for this data/model-size combination.&lt;/p&gt;

&lt;p&gt;Two gotchas caught along the way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. CUDA_VISIBLE_DEVICES alone doesn't guarantee GPU index matches&lt;/span&gt;
&lt;span class="c"&gt;# nvidia-smi's PCI-bus order -- a run silently landed on the P2000&lt;/span&gt;
&lt;span class="c"&gt;# instead of the intended 3090 until:&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CUDA_DEVICE_ORDER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;PCI_BUS_ID
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CUDA_VISIBLE_DEVICES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1

&lt;span class="c"&gt;# 2. ardi's root disk (already at a tight 90% baseline) filled to 100%&lt;/span&gt;
&lt;span class="c"&gt;# mid-training from accumulated dataset/HF caches -- silent SIGKILL,&lt;/span&gt;
&lt;span class="c"&gt;# no traceback. Fixed by pointing the cache at a bigger volume instead&lt;/span&gt;
&lt;span class="c"&gt;# of the system disk:&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;HF_HOME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/backup/hf-cache
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CACHE_DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/backup/whisper-bg-tiny-data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither is the interesting part of this story, but both cost real debugging time — worth checking explicitly on any shared multi-GPU box.&lt;/p&gt;

&lt;h2&gt;
  
  
  Attempt 5: the community shortcut that didn't reproduce
&lt;/h2&gt;

&lt;p&gt;Searched Hugging Face for an existing Bulgarian ASR fine-tune before pushing further on limited training data. Found &lt;code&gt;sam8000/whisper-large-v3-turbo-bulgarian-bulgaria&lt;/code&gt; — a fine-tune of &lt;code&gt;large-v3-turbo&lt;/code&gt; (same 32-layer encoder as full large-v3, decoder pruned from 32 to 4 layers), claiming &lt;strong&gt;9.97% WER on the FLEURS Bulgarian benchmark&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Converted to CTranslate2, it does fit the P2000 — &lt;strong&gt;4.1–4.2GB used, ~900MB headroom&lt;/strong&gt; — tight but real (the third bar in the VRAM chart above). Evaluated on the &lt;em&gt;same&lt;/em&gt; held-out Common Voice test set used for every model above:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sam8000/whisper-large-v3-turbo-bulgarian-bulgaria: 31.2% WER
our own small-v2 (fine-tuned):                     32.7% WER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Statistically the same result, not the dramatic win the model card implied. The 9.97% FLEURS number isn't fake — it just doesn't transfer to a different eval set with different preprocessing/normalization. &lt;strong&gt;Always re-measure a candidate on your own eval, apples to apples, before trusting a model card's headline number.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that worked: dual-GPU failover
&lt;/h2&gt;

&lt;p&gt;Built a real deployment: two whisper containers (P2000 = small-v2, 3090 = large-v3 unchanged) behind an nginx sidecar using &lt;code&gt;proxy_next_upstream&lt;/code&gt; for automatic failover. One detail that shapes what "failover" means here: &lt;code&gt;whisper-asr-webservice&lt;/code&gt; loads its model eagerly at process boot, not per-request — so this isn't a live per-call fallback, it's "is this backend up or down," decided once at startup.&lt;/p&gt;

&lt;p&gt;Deployed live, confirmed it actually worked — routing correct. The old standalone production container was kept &lt;strong&gt;stopped, not deleted&lt;/strong&gt;, for the entire session — the eventual rollback was a container start, not a rebuild.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test that actually mattered
&lt;/h2&gt;

&lt;p&gt;Real spontaneous speech — describing colors and objects out loud, not Common Voice-style read sentences. Verdict: "quite, quite, quite weak." Noticeably worse than the 32.7% benchmark WER suggested for casual listening. Rolled back to large-v3-only production immediately — &lt;strong&gt;~5 seconds&lt;/strong&gt;, because the old container was never torn down.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we deliberately didn't do next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Didn't publish a GitHub repo for the fine-tuned checkpoints — the result isn't good enough to ship as a "solution."&lt;/li&gt;
&lt;li&gt;Didn't chase a 6th fine-tune attempt (medium-size, more data augmentation) — diminishing returns were already visible in the epoch curve, and the deeper problem (domain mismatch between read-aloud and spontaneous speech) wouldn't be fixed by more of the same data.&lt;/li&gt;
&lt;li&gt;Didn't keep the dual-GPU stack running "just in case" — production reverted to exactly its pre-session state, P2000 idle again.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The actual finding
&lt;/h2&gt;

&lt;p&gt;Common Voice is people reading prepared text aloud in clean conditions — a different domain from spontaneous conversational speech directed at an assistant (prosody, hesitation, mic quality, vocabulary). A benchmark WER on read-aloud speech didn't predict real assistant-use quality here, for either our own fine-tune or a community model claiming a much better number on a different benchmark. This generalizes past Bulgarian and past Whisper: &lt;strong&gt;eval-set domain match matters more than the headline metric.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Full narrative version — the charts, the physical GPU install photo, the "why I still don't have a use for this card" ending — &lt;a href="https://medium.com/@arsen.apostolov" rel="noopener noreferrer"&gt;on Medium&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Every VRAM ceiling and WER number above was measured via &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt; — MIT licensed, one container, the same tool that's priced every benchmark in this series.&lt;/p&gt;

&lt;p&gt;Curious if anyone's gotten a Bulgarian (or other lower-resource-language) Whisper fine-tune to hold up on real spontaneous speech, not just a read-aloud benchmark — and what closed the gap if so.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>whisper</category>
      <category>gpu</category>
      <category>homelab</category>
    </item>
    <item>
      <title>vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Sun, 05 Jul 2026 05:54:01 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/vllm-vs-llamacpp-vs-ollama-what-happens-when-your-model-doesnt-fit-in-24gb-vram-56eb</link>
      <guid>https://dev.to/sikamikanikobg/vllm-vs-llamacpp-vs-ollama-what-happens-when-your-model-doesnt-fit-in-24gb-vram-56eb</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;Benchmarked &lt;strong&gt;llama.cpp, Ollama, and vLLM&lt;/strong&gt; across &lt;strong&gt;5 models (1B to 116.8B params)&lt;/strong&gt; on one &lt;strong&gt;RTX 3090 (24GB) + 128GB RAM&lt;/strong&gt; home-lab box, priced through &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt;. Inside 24GB, vLLM's continuous batching scales aggregate throughput &lt;strong&gt;3.9x-5.4x&lt;/strong&gt; from concurrency 1 to 8 (llama.cpp only manages &lt;strong&gt;1.2x-1.9x&lt;/strong&gt;, even with &lt;code&gt;-np 8&lt;/code&gt; explicitly set to match). Past 24GB — two models deliberately chosen to force RAM-spill — llama.cpp and Ollama both degrade to single-digit tok/s and keep generating. &lt;strong&gt;vLLM OOMs outright on both&lt;/strong&gt;, at the same ~22.1-22.2GB-used / &amp;lt;700MB-free ceiling, regardless of quantization scheme. Sub-plot: llama.cpp's manually-tuned layer offload beats Ollama's automatic split by &lt;strong&gt;37x&lt;/strong&gt; on time-to-first-token during RAM-spill, while landing on nearly identical steady-state decode speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The roster
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Vendor&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Fits in 24GB?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 3 1B&lt;/td&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;dense&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder 30B-A3B&lt;/td&gt;
&lt;td&gt;Alibaba&lt;/td&gt;
&lt;td&gt;MoE (~3.3B active)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B-A4B&lt;/td&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;MoE (~4B active)&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-4.5-Air 106B-A12B&lt;/td&gt;
&lt;td&gt;Zhipu&lt;/td&gt;
&lt;td&gt;MoE (~12B active)&lt;/td&gt;
&lt;td&gt;no, deliberately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-OSS 120B-A5.1B&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;MoE (~5.1B active)&lt;/td&gt;
&lt;td&gt;no, deliberately&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;(Gemma 4 is real — Google's newest release as of this writing, not a Gemma 3 typo.)&lt;/p&gt;

&lt;p&gt;3 prompt tiers (short/medium/long), concurrency 1 and 8, 2 reps per cell, 15 backend×model pairs total. &lt;strong&gt;Caveat stated up front&lt;/strong&gt;: the first three models ran against my production Ollama (&lt;code&gt;OLLAMA_NUM_PARALLEL=1&lt;/code&gt;, serialized by default — real daily-use config); GLM and GPT-OSS ran against a separate isolated instance (&lt;code&gt;OLLAMA_NUM_PARALLEL=4&lt;/code&gt;) since they needed a clean volume anyway. Ollama's concurrency=8 numbers for the first three models are &lt;strong&gt;not&lt;/strong&gt; its concurrency ceiling — they're its actual default production behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concurrency, inside 24GB
&lt;/h2&gt;

&lt;p&gt;Aggregate decode tok/s, concurrency 1 → concurrency 8:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Ollama&lt;/th&gt;
&lt;th&gt;llama.cpp&lt;/th&gt;
&lt;th&gt;vLLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 3 1B&lt;/td&gt;
&lt;td&gt;125.6 → 71.4&lt;/td&gt;
&lt;td&gt;294.1 → 400.6&lt;/td&gt;
&lt;td&gt;235.5 → 1172.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder 30B-A3B&lt;/td&gt;
&lt;td&gt;129.3 → 108.4&lt;/td&gt;
&lt;td&gt;157.2 → 183.9&lt;/td&gt;
&lt;td&gt;172.0 → 677.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B-A4B&lt;/td&gt;
&lt;td&gt;84.5 → 78.5&lt;/td&gt;
&lt;td&gt;118.8 → 220.6&lt;/td&gt;
&lt;td&gt;133.8 → 723.4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;vLLM's own c1→c8 scaling: &lt;strong&gt;3.9x-5.4x&lt;/strong&gt; (paged attention, requests slot into idle cycles). llama.cpp's, even with &lt;code&gt;-np 8&lt;/code&gt; matched to the concurrency level: &lt;strong&gt;1.2x-1.9x&lt;/strong&gt; — it pre-declares a fixed KV-cache reservation per parallel slot before the server starts, so concurrency is a config decision, not a runtime one. Head-to-head at c8: vLLM beats llama.cpp by &lt;strong&gt;2.9x-3.7x&lt;/strong&gt;, beats Ollama's serialized default by &lt;strong&gt;6.3x-16.4x&lt;/strong&gt; (caveat above applies).&lt;/p&gt;

&lt;h2&gt;
  
  
  The cliff, and vLLM's wall
&lt;/h2&gt;

&lt;p&gt;GLM-4.5-Air (~52% of layers spilled to system RAM under llama.cpp's tuning) and GPT-OSS-120B (~67% spilled) were picked specifically to not fit. llama.cpp and Ollama both ran them — slow, single-digit tok/s, but real generation, no crash. vLLM failed outright on &lt;strong&gt;both&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# GPT-OSS-120B, native MXFP4, --cpu-offload-gb 45
OutOfMemoryError: CUDA out of memory. Tried to allocate 1.08 GiB.
GPU 0 has a total capacity of 23.56 GiB of which 533.69 MiB is free.
Process ... has 22.21 GiB memory in use.
RuntimeError: Engine core initialization failed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# GLM-4.5-Air, pre-quantized AWQ, --cpu-offload-gb 36
OutOfMemoryError: CUDA out of memory. Tried to allocate 1.16 GiB.
GPU 0 has a total capacity of 23.56 GiB of which 685.69 MiB is free.
Process ... has 22.12 GiB memory in use.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same shape, different model, different quantization path. I retried GLM at &lt;code&gt;--gpu-memory-utilization 0.78&lt;/code&gt; (down from 0.90, to force more declared headroom) — &lt;strong&gt;got the byte-for-byte identical error&lt;/strong&gt;: 22.12 GiB used, 685.69 MiB free, 1.16 GiB requested. That rules out the utilization knob as the fix; the base weight + offload footprint is already pinned at the ceiling before profiling starts. Two models, two quant schemes, same ~22GB wall — reads as a real limit of vLLM's CPU-offload path for &amp;gt;100B-param MoE on one 24GB card on this stack, not a per-model quirk.&lt;/p&gt;

&lt;h2&gt;
  
  
  TTFT: the 37x gap that steady-state doesn't show
&lt;/h2&gt;

&lt;p&gt;On the models that ran everywhere, steady-state decode is nearly a tie once warmed up — GPT-OSS-120B's longest tier: &lt;strong&gt;7.65 tok/s (llama.cpp) vs 7.6 tok/s (Ollama)&lt;/strong&gt;. GLM: &lt;strong&gt;4.58 vs 4.59&lt;/strong&gt;. Time-to-first-token is a different story:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Ollama TTFT&lt;/th&gt;
&lt;th&gt;llama.cpp TTFT&lt;/th&gt;
&lt;th&gt;Gap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GLM-4.5-Air&lt;/td&gt;
&lt;td&gt;13.6s&lt;/td&gt;
&lt;td&gt;8.1s&lt;/td&gt;
&lt;td&gt;1.7x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-OSS-120B&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;274.0s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7.3s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;llama.cpp's &lt;code&gt;-ngl&lt;/code&gt; is a number I computed myself from the model's real &lt;code&gt;config.json&lt;/code&gt; (layer count, per-layer size) — &lt;code&gt;-ngl 12&lt;/code&gt; for GPT-OSS, offloading ~21GB deliberately. Ollama figures the split out automatically at load time, and on a freshly-pulled, partially-RAM-resident 65GB model, that automatic path is expensive. Same destination, very different path there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs (BGN per 1M output tokens, real GPU energy)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Ollama&lt;/th&gt;
&lt;th&gt;llama.cpp&lt;/th&gt;
&lt;th&gt;vLLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 3 1B&lt;/td&gt;
&lt;td&gt;0.19&lt;/td&gt;
&lt;td&gt;0.05&lt;/td&gt;
&lt;td&gt;~0*&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B-A4B&lt;/td&gt;
&lt;td&gt;0.25&lt;/td&gt;
&lt;td&gt;0.14&lt;/td&gt;
&lt;td&gt;0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3-Coder 30B-A3B&lt;/td&gt;
&lt;td&gt;0.16&lt;/td&gt;
&lt;td&gt;0.13&lt;/td&gt;
&lt;td&gt;0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-4.5-Air&lt;/td&gt;
&lt;td&gt;2.61&lt;/td&gt;
&lt;td&gt;1.95&lt;/td&gt;
&lt;td&gt;OOM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-OSS-120B&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;10.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.43&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OOM&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;*vLLM's Gemma 3 1B run finished in 6s — too fast for the power sampler to catch a reading, recorded near-zero. A sampling limitation on short bursts, not a genuine free result.&lt;/p&gt;

&lt;p&gt;GPT-OSS-120B on Ollama costs &lt;strong&gt;~7x more real electricity per million tokens&lt;/strong&gt; than llama.cpp for the identical model — the TTFT convenience tax from above, showing up again in currency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three disclosed vLLM checkpoint swaps
&lt;/h2&gt;

&lt;p&gt;The original plan was on-the-fly bitsandbytes 4-bit quant for every vLLM leg. It failed for every MoE model, for three distinct, verified reasons — not the same error copy-pasted three times:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Qwen3-Coder-30B&lt;/strong&gt;: &lt;code&gt;ValueError: BitsAndBytes quantization with padded hidden_size ... Parameter shape (786432, 1) != checkpoint shape (2048, 768)&lt;/code&gt; — bnb can't dequantize this MoE's padded expert layout. Fix: pre-quantized AWQ checkpoint. Ran clean after (677.9 tok/s aggregate @ c8).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 4 26B-A4B&lt;/strong&gt;: &lt;code&gt;AttributeError: MoE Model Gemma4ForConditionalGeneration does not support BitsAndBytes quantization yet.&lt;/code&gt; A new architecture, bnb path not wired up yet. Fix: a different pre-quantized checkpoint — which then hit a pydantic error because its &lt;code&gt;config.json&lt;/code&gt; says &lt;code&gt;compressed-tensors&lt;/code&gt;, not AWQ, despite the repo name. Fixed by dropping the explicit &lt;code&gt;--quantization&lt;/code&gt; flag entirely and letting vLLM auto-detect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-4.5-Air&lt;/strong&gt;: not a failure — a practicality call. Skipped a 212GB native bf16 download to test a bnb+MoE+CPU-offload combo the vLLM community already flagged as shaky, went straight to a ~63GB pre-quantized AWQ checkpoint that tests the exact same question.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every root cause above came from the actual container logs, not from assuming precedent carried over from the previous model's failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What wasn't tested
&lt;/h2&gt;

&lt;p&gt;Only two &lt;code&gt;--gpu-memory-utilization&lt;/code&gt; values before accepting the OOM as final, not a full &lt;code&gt;--cpu-offload-gb&lt;/code&gt; sweep. No multi-GPU / tensor-parallel vLLM path — a different question from "does single-card CPU offload work." Ollama's c8 numbers for the first three models are its production default, not its concurrency ceiling. And one raw llama.cpp per-request timing (Gemma 4, medium tier, c8) self-reported an impossible 250,024 tok/s from a near-zero-duration completion — the aggregate figures used throughout are total-tokens-over-wall-time, which isn't corrupted by that, but it's a known rough edge in the raw per-request logs.&lt;/p&gt;

&lt;p&gt;Full narrative version, with the RAM-spill mechanics and the redacted dashboard screenshot: &lt;a href="https://medium.com/@arsen.apostolov" rel="noopener noreferrer"&gt;on Medium&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Every number above was priced through &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt; — open source, MIT licensed — against the RTX 3090's real power draw.&lt;/p&gt;

&lt;p&gt;If you're already running one of these three backends: has yours ever tried to load something that just didn't fit — and did it fail loud or fail quiet?&lt;/p&gt;

</description>
      <category>llm</category>
      <category>homelab</category>
      <category>vllm</category>
      <category>ai</category>
    </item>
    <item>
      <title>Local LLM vs Claude: Benchmarking qwen3-coder:30b as a Production Agent Backend</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Fri, 03 Jul 2026 11:16:36 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/local-llm-vs-claude-benchmarking-qwen3-coder30b-as-a-production-agent-backend-482b</link>
      <guid>https://dev.to/sikamikanikobg/local-llm-vs-claude-benchmarking-qwen3-coder30b-as-a-production-agent-backend-482b</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;Replayed 27 real historical tasks from Jarvis (my LangGraph agent, ~90 tools) through &lt;code&gt;qwen3-coder:30b&lt;/code&gt; on an RTX 3090, scored against Claude's actual production answers to the same tasks. Quality: &lt;strong&gt;Claude 89.4/100 vs qwen 22.8/100&lt;/strong&gt;. Cost: &lt;strong&gt;qwen ~5,150x cheaper per task&lt;/strong&gt; ($0.00015 vs $0.763, real GPU electricity vs real API billing). Reliability: qwen leaked malformed tool-call tags into &lt;strong&gt;26% of answers&lt;/strong&gt; and only overlapped with the tools the task actually needed &lt;strong&gt;14.8%&lt;/strong&gt; of the time. Same qwen3-coder:30b scored 100% in an earlier, much smaller benchmark — the gap here is about tool-surface complexity, not the model being bad.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;Jarvis is a real personal AI agent — LangGraph &lt;code&gt;create_react_agent&lt;/code&gt;, ~90 tools spanning email/calendar/notes/files/messaging/code, running on Claude in production. &lt;code&gt;qwen3-coder:30b&lt;/code&gt; had already scored 100% task success in a &lt;a href="https://dev.to/sikamikanikobg/how-to-run-reliable-local-llm-agents-on-an-rtx-3090-a-benchmark-5-models-priced-in-watts-15d0"&gt;controlled 17-task benchmark&lt;/a&gt; on the same RTX 3090. Obvious next question: drop it into the real agent and see what happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;28 real task prompts pulled from Jarvis's own Langfuse traces (90-day window), stratified 4×7 across calendar / code / email / files / general / messaging / notes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude's answers are real production history, not re-run.&lt;/strong&gt; Re-running through the sandbox would hand it fake stub data it never saw — that's a worse baseline, not a fairer one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;qwen runs fresh&lt;/strong&gt;, through a sandboxed replay harness: the real Jarvis agent code in-process, every write-capable tool intercepted (nothing sent/written for real), and every mocked read-only tool serves the &lt;em&gt;real recorded output&lt;/em&gt; from that task's original trace when available — not a generic stub. Same data, both models.&lt;/li&gt;
&lt;li&gt;1/28 tasks excluded (336,906-char prompt, over any 16K–24K context window) → 27 scored.&lt;/li&gt;
&lt;li&gt;Judge: LLM-as-judge (&lt;code&gt;claude-opus-4-8&lt;/code&gt;), scored independently per answer (not pairwise) to avoid position bias, 1–5 → 0–100.&lt;/li&gt;
&lt;li&gt;Every qwen run priced as a &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt; experiment against real 3090 power draw. Claude's cost is Langfuse's recorded API billing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Caveat, stated plainly:&lt;/strong&gt; the judge is a Claude model scoring Claude's own answers alongside qwen's — self-preference bias is a documented effect in LLM-as-judge setups and probably inflates the gap somewhat. It doesn't explain a 66-point gap, a 26% malformed-output rate, or two tool-call loops, but it's a real methodology limitation, not a footnote.&lt;/p&gt;

&lt;p&gt;Getting here took three re-runs: a judge-response parsing bug that silently neutral-scored ~40/54 calls, a mock-data bug that starved qwen of real inbox/calendar content on 16/28 tasks while Claude's baseline had the real thing, and a Claude-API rate limit that neutral-scored another batch mid-scoring. All three caught by checking score distributions, not by trusting a clean exit code — worth knowing before trusting the numbers below.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Claude&lt;/th&gt;
&lt;th&gt;qwen3-coder:30b&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Avg quality (0–100)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89.4&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;22.8&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost / task&lt;/td&gt;
&lt;td&gt;$0.763 (real API billing)&lt;/td&gt;
&lt;td&gt;$0.00015 (real GPU electricity)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total cost, 27 tasks&lt;/td&gt;
&lt;td&gt;$20.60&lt;/td&gt;
&lt;td&gt;$0.004&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total energy&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.0396 kWh&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;~&lt;strong&gt;5,150x cheaper per task&lt;/strong&gt; for qwen (precise, currency-converted from a 0.0072 BGN &lt;em&gt;total&lt;/em&gt; across all 27 tasks, at 1 BGN = $0.5547 — an earlier rough estimate of 180x on this project was wrong, this is the corrected number).&lt;/p&gt;

&lt;p&gt;By category (Claude | qwen | n):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;calendar:   90 | 30 | 4
code:       87 | 25 | 3
email:      92 | 15 | 4
files:      88 | 15 | 4
general:    85 | 30 | 4
messaging:  87 | 22 | 4
notes:      97 | 22 | 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;qwen's best relative showing (calendar, general) is still a third of Claude's score. It never wins a category.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it breaks
&lt;/h2&gt;

&lt;p&gt;Malformed tool-call leak — instead of a real LangGraph tool call, qwen sometimes emits the call as raw text in its final answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;function&lt;/span&gt;&lt;span class="err"&gt;=send_email&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
{"to": "...", "subject": "...", "body": "..."}
&lt;span class="nt"&gt;&amp;lt;/function&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That happened on &lt;strong&gt;7/27 tasks (26%)&lt;/strong&gt;. The user reading that answer sees broken syntax where a real action should have been confirmed or a real answer given.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool-overlap recall: 14.8% average&lt;/strong&gt;, measured over the 18/27 tasks where the original historical trace actually used at least one tool (9 tasks needed none). Most of the time qwen reached for different tools than the ones that actually solved the task — or none.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repetitive-loop failure&lt;/strong&gt; on 2/27 tasks: &lt;code&gt;pilot-17&lt;/code&gt; (email, 24 tool calls, 138.6s, ~196.7K input tokens) and &lt;code&gt;pilot-27&lt;/code&gt; (messaging, 27 tool calls, 148.9s, ~196.7K input tokens) both called the &lt;em&gt;same already-answered tool&lt;/em&gt; (&lt;code&gt;run_command&lt;/code&gt;, &lt;code&gt;todo_write&lt;/code&gt;) repeatedly instead of stopping. Confirmed via raw logs both tasks got real replayed data (&lt;code&gt;replayed_real_data: true&lt;/code&gt;) — a genuine stopping-condition failure, not a data-starvation artifact.&lt;/p&gt;

&lt;p&gt;One more data point worth having, not a verdict: on a task where both models actually called &lt;code&gt;send_email(...)&lt;/code&gt; in the harness (intercepted, nothing sent), Claude told the user the email had been sent — a fabrication. qwen correctly disclosed the send didn't go through. Not "qwen is more honest" — it's also the model leaking raw tags 26% of the time. Both mishandled the mock, just differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope of the claim
&lt;/h2&gt;

&lt;p&gt;Same &lt;code&gt;qwen3-coder:30b&lt;/code&gt;, same GPU, scored 100% on a 17-task controlled benchmark with a much smaller tool surface. This isn't "local LLMs are bad" — it's that a model excellent on a scoped benchmark isn't automatically a safe drop-in for a large, real, ~90-tool production surface with a 31KB context prompt and real messy history behind it. Task/tool-surface complexity mattered as much as raw model quality here. Claude isn't flawless either — see the fabricated send-email confirmation above.&lt;/p&gt;

&lt;p&gt;Jarvis stays on Claude for now. The cost number is real enough to be worth a narrower follow-up — testing qwen on just the categories where it scored closest (calendar, general) as a cheap fallback path, rather than a full swap.&lt;/p&gt;

&lt;p&gt;Full narrative version, charts, and the three-bug scoring saga: &lt;a href="https://medium.com/@arsen.apostolov" rel="noopener noreferrer"&gt;on Medium&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;Every qwen run here was priced through &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt; against the 3090's real power draw — MIT licensed, one container, reproducible if you want to price your own local-model experiments the same way.&lt;/p&gt;

&lt;p&gt;Curious where the line is for you: how cheap does a local model have to be before you'd trust it with a slice of a real agent, and which slice would you pick first?&lt;/p&gt;

</description>
      <category>llm</category>
      <category>homelab</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>How to Run Reliable Local LLM Agents on an RTX 3090: A Benchmark (5 Models, Priced in Watts)</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Sun, 28 Jun 2026 06:54:12 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/how-to-run-reliable-local-llm-agents-on-an-rtx-3090-a-benchmark-5-models-priced-in-watts-15d0</link>
      <guid>https://dev.to/sikamikanikobg/how-to-run-reliable-local-llm-agents-on-an-rtx-3090-a-benchmark-5-models-priced-in-watts-15d0</guid>
      <description>&lt;p&gt;I gave &lt;strong&gt;GLM-4.5-Air&lt;/strong&gt; (106B, open weights) 12 coding tasks through &lt;a href="https://opencode.ai" rel="noopener noreferrer"&gt;opencode&lt;/a&gt; on my RTX 3090. It scored &lt;strong&gt;0%&lt;/strong&gt; — never edited a single file.&lt;/p&gt;

&lt;p&gt;Same model, same GPU, same tasks, but driven by a ~150-line &lt;strong&gt;LangGraph&lt;/strong&gt; agent instead: &lt;strong&gt;93%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The model was never the problem. The orchestrator was. Here's the benchmark — including the part nobody else measures, the &lt;strong&gt;electricity cost per correct task&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgqtw55h6nnjo1q76711v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgqtw55h6nnjo1q76711v.png" alt="opencode vs LangGraph tool-adherence" width="800" height="467"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RTX 3090 (24 GB) + 128 GB RAM&lt;/strong&gt;, models via &lt;strong&gt;ollama&lt;/strong&gt;, Q4 quants, temp 0.2&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5 recent open models&lt;/strong&gt; × &lt;strong&gt;2 orchestrators&lt;/strong&gt; (opencode vs custom LangGraph ReAct with ollama-native tool-calling)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;17 graded tasks&lt;/strong&gt; (12 coding in Python/JS/C++ + 5 general-agent) with hidden unit tests&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every run priced in GPU watts&lt;/strong&gt; via my open-source &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;homelab-monitor&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;tok/s&lt;/th&gt;
&lt;th&gt;opencode adh.&lt;/th&gt;
&lt;th&gt;LangGraph adh.&lt;/th&gt;
&lt;th&gt;LangGraph coding&lt;/th&gt;
&lt;th&gt;LangGraph general&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Qwen3-Coder 30B-A3B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;130&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;92%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GLM-4.5-Air 106B&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5.7&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;100%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Devstral Small 24B&lt;/td&gt;
&lt;td&gt;49&lt;/td&gt;
&lt;td&gt;8%&lt;/td&gt;
&lt;td&gt;53%&lt;/td&gt;
&lt;td&gt;8%&lt;/td&gt;
&lt;td&gt;40%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seed-OSS 36B&lt;/td&gt;
&lt;td&gt;9.5&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;7%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1-Distill 32B&lt;/td&gt;
&lt;td&gt;6.7&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Tool-adherence&lt;/strong&gt; = % of tasks where the model actually &lt;em&gt;called a tool&lt;/em&gt; instead of just printing code in chat. It was the master variable. (GLM's headline "93%" is its blended score across all 17 tasks: 89% coding + 100% general.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Three takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The framework can matter more than the model.&lt;/strong&gt; opencode sends a frontier-shaped system prompt + 12 tools over its OpenAI-compat path; most local models fall back to chatting. Native tool-calling through a lean agent fixes that — GLM went 0% → 93%. (Qwen3-Coder is the exception: it's tuned for agentic tool use and aces opencode out of the box.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Acting ≠ solving.&lt;/strong&gt; LangGraph made Devstral &lt;em&gt;act&lt;/em&gt; (8% → 53% adherence) but not &lt;em&gt;solve&lt;/em&gt; (coding stayed 8%). The framework decides whether a model acts; the model decides whether it's right.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The wattmeter ranks honestly.&lt;/strong&gt; Qwen solved tasks at ~0.0005 BGN each; the models that scored zero still burned &lt;strong&gt;10–30× more energy&lt;/strong&gt; for nothing. On a home rig, the cheapest model is the fast, correct one — and MoE (Qwen activates ~3B of 30B per token) wins twice.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Bonus: &lt;strong&gt;128 GB RAM let me run the 106B GLM&lt;/strong&gt; (23 GB VRAM + 27 GB spilled to RAM) — it works, at 5.7 tok/s. Great for fire-and-forget batch jobs, not interactive coding.&lt;/p&gt;

&lt;h2&gt;
  
  
  The recipe for reliable local agents
&lt;/h2&gt;

&lt;p&gt;Pick a tool-use-tuned model (&lt;strong&gt;Qwen3-Coder 30B-A3B&lt;/strong&gt; is the all-weather winner) → use &lt;strong&gt;native&lt;/strong&gt; tool-calling, not an OpenAI-compat path → keep the harness lean → use RAM for reach, not speed → &lt;strong&gt;measure correctness per kWh&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;📖 &lt;strong&gt;Full write-up with methodology, charts, and the deeper "why" →&lt;/strong&gt; [&lt;a href="https://medium.com/@arsen.apostolov/local-llm-agents-on-an-rtx-3090-i-benchmarked-5-models-2-frameworks-and-the-orchestrator-f5fd600ca221" rel="noopener noreferrer"&gt;https://medium.com/@arsen.apostolov/local-llm-agents-on-an-rtx-3090-i-benchmarked-5-models-2-frameworks-and-the-orchestrator-f5fd600ca221&lt;/a&gt;]&lt;/p&gt;

&lt;p&gt;⭐ Every number was priced in watts by &lt;strong&gt;&lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;homelab-monitor&lt;/a&gt;&lt;/strong&gt; — my open-source tool that turns your GPU's power draw into per-task cost. &lt;strong&gt;Star it&lt;/strong&gt; if you want the same receipts for your own rig. Harness + tasks + leaderboard code are reproducible.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>homelab</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
    <item>
      <title>How to Rank Local LLMs by Cost per Correct Answer (Measured GPU Energy, 8 Ollama Models)</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Tue, 23 Jun 2026 18:11:23 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/how-to-rank-local-llms-by-cost-per-correct-answer-measured-gpu-energy-8-ollama-models-5c5h</link>
      <guid>https://dev.to/sikamikanikobg/how-to-rank-local-llms-by-cost-per-correct-answer-measured-gpu-energy-8-ollama-models-5c5h</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; I priced 8 local Ollama models by &lt;strong&gt;€ per 1,000 correct answers&lt;/strong&gt; — metered GPU energy ÷ correct answers, on one RTX 3090. &lt;code&gt;gemma4:26b&lt;/code&gt; won at &lt;strong&gt;96.9% accuracy for €0.013/1k-correct&lt;/strong&gt;. The most expensive model (&lt;code&gt;qwen3:8b-fp16&lt;/code&gt;) cost &lt;strong&gt;€0.239/1k&lt;/strong&gt; and scored &lt;em&gt;worse&lt;/em&gt; (66.7%). Reasoning tokens and full precision both cost a lot and bought nothing here. Every cost comes from real metered kWh via the open-source &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;HomeLab Monitor&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is the short, copy-pasteable version. The narrative writeup is on &lt;a href="https://medium.com/@arsen.apostolov/tokens-are-cheap-wrong-answers-arent-32be7655845d" rel="noopener noreferrer"&gt;Medium&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The metric
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;€ per correct answer = (metered GPU energy cost over the eval window) ÷ (number of correct answers)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Tokens-per-euro flatters whichever model talks the most. Cost-per-correct only rewards being &lt;em&gt;right cheaply&lt;/em&gt; — which is the thing you actually pay for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The signal
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model                  VRAM     Acc     Tok/task  Tok/s  Wh/pass  €/1k correct (day)
gemma4:26b             16.9 GB  96.9%   68        86     4.5      €0.013   ← winner
gemma3:1b              0.9 GB   82.1%   125       133    3.8      €0.013
gemma3:27b             17.1 GB  100.0%  119       36     16.3     €0.046
qwen3:30b-a3b   (MoE)  18.4 GB  83.3%   555       186    14.1     €0.048
qwen3:8b (Q4_K_M) 🧠   5.4 GB   64.8%   626       126    22.7     €0.100
qwen3:8b          🧠   5.4 GB   64.8%   626       126    23.6     €0.104
qwen3:8b (Q8_0)   🧠   8.7 GB   61.1%   672       88     33.5     €0.156
qwen3:8b (fp16)   🧠   15.5 GB  66.7%   664       53     56.2     €0.239   ← most expensive
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;🧠 = reasoning/thinking mode on. Night tariff knocks ~40% off every row.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things the numbers say
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The value champion is mid-size, not max-size.&lt;/strong&gt; &lt;code&gt;gemma4:26b&lt;/code&gt; hit &lt;strong&gt;96.9%&lt;/strong&gt; for &lt;strong&gt;€0.013 per 1,000 correct&lt;/strong&gt; — cheapest-per-correct on the whole bench &lt;em&gt;and&lt;/em&gt; near-perfect, ~&lt;strong&gt;18×&lt;/strong&gt; cheaper per correct answer than &lt;code&gt;qwen3:8b-fp16&lt;/code&gt;. &lt;code&gt;gemma3:27b&lt;/code&gt; is the only 100% model but costs ~3.5× more (slower, 36 tok/s).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The thinking tax is real and didn't pay off.&lt;/strong&gt; qwen3 reasoning models emit &lt;strong&gt;555–672 tokens/task&lt;/strong&gt; vs the gemmas' &lt;strong&gt;68–125&lt;/strong&gt; (5–9×). Tokens are energy. On these 54 deterministic tasks that extra reasoning bought &lt;em&gt;no&lt;/em&gt; correctness — the priciest model scored &lt;em&gt;lower&lt;/em&gt; than one 18× cheaper. (Caveat: this suite is arithmetic / executable code / format-following. On open-ended hard problems, reasoning earns its tokens. On structured agent work, it was dead weight.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The quantization paradox.&lt;/strong&gt; Same &lt;code&gt;qwen3:8b&lt;/code&gt; at three precisions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;            Accuracy   Energy/pass   Throughput
Q4_K_M      64.8%      22.7 Wh       126 tok/s
Q8_0        61.1%      33.5 Wh       88  tok/s
fp16        66.7%      56.2 Wh       53  tok/s
            └ flat ┘   └ 2.5× ┘      └ halved ┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Higher precision cost &lt;strong&gt;2.5× the energy&lt;/strong&gt; and &lt;strong&gt;half the throughput&lt;/strong&gt; for accuracy that's flat-and-noisy. On a 3090, aggressive quant was the correct call, not a compromise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Methodology (so you can trust the ranking)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;54 deterministic tasks&lt;/strong&gt;, mechanically graded — &lt;strong&gt;no LLM judge&lt;/strong&gt;. Reasoning 15 (GSM8K-style numeric extraction), code 12 (HumanEval-style, &lt;em&gt;executed&lt;/em&gt; asserts in a sandbox), factual 12 (keyword), instruct 15 (format predicates). Grader selftest 11/11.&lt;/li&gt;
&lt;li&gt;Controls identical across all 8 models: &lt;strong&gt;temperature 0, seed 42, num_ctx 4096, num_predict 1024&lt;/strong&gt;, identical prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Warm-up discarded&lt;/strong&gt; → model-load energy excluded (pricing inference, not cold starts).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3 passes&lt;/strong&gt; each, ranges reported.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idle baseline = 38 W&lt;/strong&gt;, measured as a control.&lt;/li&gt;
&lt;li&gt;qwen3 thinking left &lt;strong&gt;on&lt;/strong&gt; (realistic); thinking tokens counted for energy, stripped before grading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Honest determinism caveat:&lt;/strong&gt; Ollama is &lt;em&gt;not&lt;/em&gt; bit-exact at temp 0. &lt;code&gt;gemma3:1b&lt;/code&gt; drifted 81–83% across passes; &lt;code&gt;gemma3:27b&lt;/code&gt; was 100% on all three; qwen3 runs were identical. Report ranges, not point claims.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CPU/DRAM not metered&lt;/strong&gt; (no RAPL on this host), so true wall-plug cost is a bit higher — but the ranking holds because every model paid the same un-metered overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The currency gotcha (measure twice)
&lt;/h2&gt;

&lt;p&gt;Costs are EUR from &lt;strong&gt;measured kWh × Bulgarian dual tariff (€0.1534 day / €0.0920 night)&lt;/strong&gt;. While building this I caught my own dashboard mislabeling &lt;strong&gt;BGN as EUR&lt;/strong&gt;: the tariff read &lt;code&gt;0.30/0.18 EUR&lt;/code&gt;, but those are leva. Bulgaria joined the euro on &lt;strong&gt;2026-01-01&lt;/strong&gt; at fixed &lt;strong&gt;1 EUR = 1.95583 BGN&lt;/strong&gt;; €0.30/kWh would be German-tier, implausible for the EU's cheapest household power. Converted: &lt;code&gt;0.30 / 1.95583 = €0.1534&lt;/code&gt;, &lt;code&gt;0.18 / 1.95583 = €0.0920&lt;/code&gt;. Lesson: don't trust the dashboard's € field — compute from physical kWh and your verified tariff.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to reproduce the energy tracking
&lt;/h2&gt;

&lt;p&gt;Every cost above came from &lt;a href="https://github.com/SikamikanikoBG/homelab-monitor" rel="noopener noreferrer"&gt;&lt;strong&gt;HomeLab Monitor&lt;/strong&gt;&lt;/a&gt; (MIT, one container) — its Experiments tab integrates real GPU power over a run's window into kWh and money. Bring it up:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt;        &lt;span class="c"&gt;# port 9800&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Grab the one-file &lt;code&gt;homelab_run.py&lt;/code&gt; client, mint an ingest key, and wrap your eval — the run comes back priced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;homelab_run&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;homelab&lt;/span&gt;
&lt;span class="n"&gt;homelab&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;configure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://&amp;lt;your-host&amp;gt;:9800&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hlm_…&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;homelab&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemma4:26b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llm-cost-bench&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PASSES&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nf"&gt;run_graded_eval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;# all inference inside the run
&lt;/span&gt;&lt;span class="n"&gt;priced&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;homelab&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;pull&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;             &lt;span class="c1"&gt;# energy_kwh, cost, avg_w, peak_util — from real power
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole instrumentation. Divide the priced energy by your grader's correct count and you've got cost-per-correct for your own roster. &lt;a href="https://sikamikanikobg.github.io/homelab-monitor/" rel="noopener noreferrer"&gt;Docs&lt;/a&gt; · &lt;code&gt;docker pull sikamikaniko123/homelab-monitor&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I deliberately did NOT do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;No LLM-as-judge — mechanical grading only.&lt;/li&gt;
&lt;li&gt;No cold-start energy in the numbers — warm-up discarded on purpose.&lt;/li&gt;
&lt;li&gt;No trusting the dashboard's € field — costs recomputed from measured kWh.&lt;/li&gt;
&lt;li&gt;No single-run claims — 3 passes, ranges where they exist.&lt;/li&gt;
&lt;li&gt;No CPU/DRAM cost claim — only the GPU is metered, and I say so.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;Bigger and full-precision lost. A 26B model did near-perfect work for a rounding error; an fp16 reasoning model charged 18× as much to be wrong more often.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So when you reach for a local model — accuracy, speed, or cost per answer it actually gets right? And have you ever measured the third one?&lt;/strong&gt; Drop your own cost-per-correct numbers in the comments.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>llm</category>
      <category>gpu</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How Much Does It Actually Cost to Run a Local LLM? (€ per Million Tokens, Measured)</title>
      <dc:creator>Arsen Apostolov</dc:creator>
      <pubDate>Mon, 22 Jun 2026 18:33:31 +0000</pubDate>
      <link>https://dev.to/sikamikanikobg/how-much-does-it-actually-cost-to-run-a-local-llm-eu-per-million-tokens-measured-jih</link>
      <guid>https://dev.to/sikamikanikobg/how-much-does-it-actually-cost-to-run-a-local-llm-eu-per-million-tokens-measured-jih</guid>
      <description>&lt;p&gt;"It runs on my own GPU, so it's basically free." I believed that until I put a meter on it. So I ran a controlled benchmark on one box — an openSUSE machine with a single RTX 3090 — driving three local models through ollama under an identical fixed workload (256-token generations in a loop for ~4 minutes each), while my open-source dashboard priced every run by the &lt;strong&gt;real GPU energy it burned&lt;/strong&gt;: power sampled from &lt;code&gt;nvidia-smi&lt;/code&gt; every 10 s, integrated over each run's exact window, multiplied by my actual day/night tariff. One number per model, in euros per million output tokens.&lt;/p&gt;

&lt;p&gt;Here's the part that made me re-run it. The tiny &lt;code&gt;gemma3:1b&lt;/code&gt; came out at &lt;strong&gt;€0.118 / 1M tokens&lt;/strong&gt; — about &lt;strong&gt;5× cheaper&lt;/strong&gt; than a hosted Flash-class API (~€0.55). But &lt;code&gt;gemma3:27b&lt;/code&gt;'s &lt;strong&gt;electricity alone&lt;/strong&gt; was &lt;strong&gt;€0.706 / 1M&lt;/strong&gt; — &lt;em&gt;more&lt;/em&gt; expensive per token than just paying the cloud, and that's before a single cent of the GPU's purchase price. "Local" didn't make it cheaper; it made it cost more &lt;em&gt;and&lt;/em&gt; I own the depreciation. The mechanism is one line: each token costs &lt;strong&gt;watts ÷ throughput&lt;/strong&gt;, and a big dense model is both slow and thirsty. A newer mid-size architecture (&lt;code&gt;gemma4:26b&lt;/code&gt;) bought a lot of that back, landing at &lt;strong&gt;€0.272&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The full guide is methodology-first and reproducible end to end — minting an ingest key, the stdlib-only client, the exact ollama loop that reads &lt;code&gt;eval_count&lt;/code&gt;/&lt;code&gt;eval_duration&lt;/code&gt; for real tokens-per-second, reading each run back priced, and the honest caveats (this is marginal GPU energy only — not capex, idle, or cooling — and the absolute numbers round to fractions of a cent; the &lt;em&gt;shape&lt;/em&gt; is the finding).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read the full guide on Medium → &lt;a href="https://medium.com/@arsen.apostolov/how-much-does-it-actually-cost-to-run-a-local-llm-per-million-tokens-measured-4a90a7f31a48" rel="noopener noreferrer"&gt;https://medium.com/@arsen.apostolov/how-much-does-it-actually-cost-to-run-a-local-llm-per-million-tokens-measured-4a90a7f31a48&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>llm</category>
      <category>homelab</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
