<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: oooocean66</title>
    <description>The latest articles on DEV Community by oooocean66 (@oooocean66).</description>
    <link>https://dev.to/oooocean66</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4116715%2F95ff6c75-fa88-4d94-ba55-d6ef1a9270c9.jpg</url>
      <title>DEV Community: oooocean66</title>
      <link>https://dev.to/oooocean66</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/oooocean66"/>
    <language>en</language>
    <item>
      <title>Qwen3.8-27B: An Architecture and Hands-On Look at Alibaba's Compact Dense Model</title>
      <dc:creator>oooocean66</dc:creator>
      <pubDate>Mon, 28 Sep 2026 07:09:10 +0000</pubDate>
      <link>https://dev.to/oooocean66/qwen38-27b-an-architecture-and-hands-on-look-at-alibabas-compact-dense-model-3281</link>
      <guid>https://dev.to/oooocean66/qwen38-27b-an-architecture-and-hands-on-look-at-alibabas-compact-dense-model-3281</guid>
      <description>&lt;p&gt;In August 2026, less than a month after Moonshot AI's Kimi-K3 surprised the world, Alibaba Group released "Qwen3.8-2.4T-A95B," an enormous 2.4-trillion-parameter model that once again shook the local-LLM community. Three days later, a smaller companion model quietly appeared: "Qwen3.8-27B," the model covered in this article.&lt;/p&gt;

&lt;p&gt;It's released under the Apache-2.0 license, free for commercial use, and as a dense model it puts its full parameter count to work rather than routing through sparse experts. That combination raises a natural question: does this 27B model hold up against its much larger sibling? We look at its architecture and run hands-on benchmarks, comparing it along the way with Muse Glimmer 30B, the model we reviewed previously.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Arrival of Qwen3.8
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Impact of Qwen3.8-2.4T-A95B
&lt;/h3&gt;

&lt;p&gt;On August 12, 2026 — roughly a month after Moonshot AI's Kimi-K3 surprised the world — China's Alibaba Group published Qwen3.8-2.4T-A95B on Hugging Face. At 2.4 trillion parameters, it's a genuinely enormous model, well beyond what most people can run locally, yet it still generated considerable buzz in the local-LLM community.&lt;/p&gt;

&lt;p&gt;Its reasoning benchmarks came close to Kimi K3's, putting it within reach of the major American frontier models.&lt;/p&gt;

&lt;p&gt;Some users ran 1-bit or 2-bit quantized versions on their own hardware and confirmed that the numbers weren't hype&lt;sup id="fnref1"&gt;1&lt;/sup&gt; — reports of the model's real-world performance spread quickly.&lt;/p&gt;

&lt;h3&gt;
  
  
  How Does Qwen3.8-27B Measure Up?
&lt;/h3&gt;

&lt;p&gt;Three days later, at midnight JST on August 15, 2026, the smaller Qwen3.8-27B was released.&lt;/p&gt;

&lt;p&gt;One important distinction is licensing. The larger Qwen3.8-2.4T-A95B ships under Alibaba's own Qwen3.8-Max license, which imposes certain conditions on commercial use — a company large enough to offer that model as a service would likely need a separate commercial licensing agreement.&lt;/p&gt;

&lt;p&gt;Qwen3.8-27B, by contrast, uses a plain Apache-2.0 license, so it can be used commercially at no cost. And as a dense model — as its name suggests — it puts its entire parameter count to work on every token, which is the design typically associated with higher per-parameter capability.&lt;/p&gt;

&lt;p&gt;Qwen3.8-2.4T-A95B made a strong case for its own capability. The question is how much of that capability carries over to this smaller sibling.&lt;/p&gt;

&lt;p&gt;We ran our own hands-on verification to dig into that question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen3.8-27B's Architecture
&lt;/h2&gt;

&lt;p&gt;Architecturally, Qwen3.8 carries the Qwen3.5/Qwen3.6 design forward essentially unchanged. The structure is shown below.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jyi2om69jq8vcn01ez0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8jyi2om69jq8vcn01ez0.png" alt="Qwen3.8-27B's architecture" width="800" height="154"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 1: Qwen3.8-27B's architecture&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Looking at the config.json published on Hugging Face makes that continuity even clearer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"architectures"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"Qwen3_5ForConditionalGeneration"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"image_token_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;248056&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"language_model_only"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"qwen3_5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"text_config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Beyond that, the basic specs are as follows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Max context size: &lt;strong&gt;256k tokens&lt;/strong&gt; (limited to 64k in our test environment)&lt;/li&gt;
&lt;li&gt;64 blocks total (GDN:GA = 3:1, 4 layers × 16 groups)&lt;/li&gt;
&lt;li&gt;Hidden size: &lt;strong&gt;5,120&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;FFN type: &lt;strong&gt;Dense&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Activation function: &lt;strong&gt;SiLU&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Activation dimension: &lt;strong&gt;17,408&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  An Attention Structure That's Actually a Minority Approach
&lt;/h3&gt;

&lt;p&gt;Qwen3.8's attention is a hybrid of two mechanisms — Gated Delta Network and Gated Attention (effectively GQA) — which is quite different from the "local + global" pattern that's become the popular default lately.&lt;/p&gt;

&lt;p&gt;Gated Delta Network handles two jobs — state updates and information compression — and makes up 75% of the attention layers. The remaining 25% are Gated Attention layers, internally structured as Group Query Attention, responsible for global aggregation and precise retrieval.&lt;/p&gt;

&lt;p&gt;This layout was first adopted in Qwen3-Next-80B and has carried through Qwen3.5 and Qwen3.6.&lt;/p&gt;

&lt;p&gt;Rather than "local + global," it's probably more accurate to describe this as "compression + aggregation." As far as we're aware, Qwen is currently the only major model family using this approach.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Vision Encoder Is Still SigLIP2-Based
&lt;/h3&gt;

&lt;p&gt;Qwen3.8's vision encoder continues to build on the SigLIP2-based architecture (SigLIP2-SO-400M). Additional fine-tuning presumably went into it, but there's been no major structural change.&lt;/p&gt;

&lt;p&gt;That's a notably different approach from Muse Glimmer's custom vision encoder, which we covered previously.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running It Ourselves
&lt;/h2&gt;

&lt;p&gt;For this test, we used a GPUSOROBAN &lt;a href="https://soroban.highreso.jp/compute" rel="noopener noreferrer"&gt;High-Speed Computing&lt;/a&gt; instance equipped with an NVIDIA RTX A4000.&lt;/p&gt;

&lt;h3&gt;
  
  
  Overall Setup
&lt;/h3&gt;

&lt;p&gt;The setup is as follows: we open an SSH tunnel to an access server, then reach the target instance through that tunnel.&lt;sup id="fnref2"&gt;2&lt;/sup&gt; llama.cpp's service port is relayed to localhost via port 8001 on both ends.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ Home ]                    [ HighReso : GPUSOROBAN ]
+---------------+           +----------------+     +------------------+
|    Client     |           | Access Server  |     | Target Instance  |
|  &amp;gt; 8001/tcp   |--(SSH)---&amp;gt;|    (relay)     |----&amp;gt;|  8001/tcp        |
|               |           |                |     |  (llama.cpp)     |
+---------------+           +----------------+     +------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Figure 2: Overview of the connection setup&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Instance Used
&lt;/h3&gt;

&lt;p&gt;The GPU used for this test is an NVIDIA RTX A4000. The instance specifications are as follows.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Instance type&lt;/td&gt;
&lt;td&gt;s16-1-a-standard-ubs24-v&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;td&gt;NVIDIA RTX A4000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory&lt;/td&gt;
&lt;td&gt;GDDR6 16GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory bandwidth&lt;/td&gt;
&lt;td&gt;448.0 GB/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FP32 compute&lt;/td&gt;
&lt;td&gt;19.17 TFLOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BF16 compute&lt;/td&gt;
&lt;td&gt;38.34 TFLOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT8 compute&lt;/td&gt;
&lt;td&gt;153.4 TOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT4 compute&lt;/td&gt;
&lt;td&gt;306.7 TOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vCPU&lt;/td&gt;
&lt;td&gt;11 cores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System memory&lt;/td&gt;
&lt;td&gt;50 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Persistent 100 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CUDA version&lt;/td&gt;
&lt;td&gt;13.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVIDIA driver version&lt;/td&gt;
&lt;td&gt;580&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS&lt;/td&gt;
&lt;td&gt;Ubuntu 24.04 Server&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 1: Instance specifications&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Model Files Used
&lt;/h3&gt;

&lt;p&gt;Even a 4-bit quantized version of Qwen3.8-27B exceeds 16GiB. Factoring in KV cache usage as well, we used a 2-bit quantized model optimized with Unsloth Dynamic 2.0.&lt;/p&gt;

&lt;p&gt;We downloaded the following files from Unsloth's Qwen3.8-27B-GGUF repository:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.8-27B-UD-Q2_K_XL.gguf&lt;/td&gt;
&lt;td&gt;12 GB&lt;/td&gt;
&lt;td&gt;UD2.0, 2-bit quantized model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mmproj-bf16.gguf&lt;/td&gt;
&lt;td&gt;889 MB&lt;/td&gt;
&lt;td&gt;Quantized vision encoder&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 2: Files used with llama.cpp&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We used llama.cpp version 0.1.0-dev (build 10450, commit ece963f41).&lt;/p&gt;

&lt;p&gt;Since the architecture hasn't changed since Qwen3.5, being able to run it on an existing llama.cpp build without waiting for new support is arguably an advantage in its own right.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verification: Standard Mode
&lt;/h3&gt;

&lt;p&gt;First, we ran text inference using the main model on its own.&lt;/p&gt;

&lt;h4&gt;
  
  
  Launch Command
&lt;/h4&gt;

&lt;p&gt;We used the following command to launch the web frontend for testing.&lt;/p&gt;

&lt;p&gt;For this run, we kept the KV cache in FP16 rather than quantizing it to 8-bit, but due to memory constraints we lowered the max context size to 65,536. We also raised the log verbosity from the default of 3 to 4 to check memory usage.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;build/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; ./models/Qwen3.8-27B-UD-Q2_K_XL.gguf &lt;span class="nt"&gt;-t&lt;/span&gt; 12 &lt;span class="nt"&gt;-np&lt;/span&gt; 1 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 0.6 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 20 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--fit&lt;/span&gt; off &lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="nt"&gt;-c&lt;/span&gt; 65536 &lt;span class="nt"&gt;--reasoning&lt;/span&gt; on &lt;span class="nt"&gt;--reasoning-effort&lt;/span&gt; xhigh &lt;span class="nt"&gt;--log-verbosity&lt;/span&gt; 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Reasoning Effort Is Now Configurable
&lt;/h4&gt;

&lt;p&gt;One recent improvement in llama.cpp that stood out to us is the addition of a &lt;code&gt;reasoning_effort&lt;/code&gt; parameter.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;--reasoning-effort LEVEL   reasoning effort level given to the chat template:
                           'default' to keep the template default, or a level
                           such as 'minimal', 'low', 'medium', 'high', 'xhigh'
                           or 'max' (default: default)
                           (env: LLAMA_ARG_REASONING_EFFORT)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Table 3: The reasoning_effort option, as described in llama-server's own help text&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Which levels are actually usable depends on Qwen3.8-27B's own chat_template.jinja. Checking that template, the following levels appear to be supported:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;xhigh (default)&lt;/li&gt;
&lt;li&gt;medium&lt;/li&gt;
&lt;li&gt;low&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Setting these appears to inject the following text into the system prompt:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Injected text&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;xhigh&lt;/td&gt;
&lt;td&gt;Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;(no system prompt injected)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 4: reasoning_effort levels and their injected prompt text&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So rather than setting an explicit token budget, it appears to control reasoning depth simply by appending instructions to the system prompt.&lt;/p&gt;

&lt;h4&gt;
  
  
  Memory Usage: KV Cache Size Is the Sticking Point
&lt;/h4&gt;

&lt;p&gt;Memory usage came out as follows. As is generally true across the Qwen series, the KV cache footprint is large. We initially tried setting the max context to 131,072, but that produced an OOM error, so we halved it. With the ceiling set at 65,536, the KV cache came out to 4,096 MiB.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Device&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;th&gt;CUDA0&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Xeon E5-2690v4&lt;/td&gt;
&lt;td&gt;RTX-A4000&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weight data (Q2_K-quant_XL)&lt;/td&gt;
&lt;td&gt;397.85&lt;/td&gt;
&lt;td&gt;9,567.89&lt;/td&gt;
&lt;td&gt;9,965.74&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache size (f16)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;4,096.00&lt;/td&gt;
&lt;td&gt;4,096.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RS buffer size (f32)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;149.62&lt;/td&gt;
&lt;td&gt;149.62&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gated Delta Net compute buffer&lt;/td&gt;
&lt;td&gt;186.02&lt;/td&gt;
&lt;td&gt;84.02&lt;/td&gt;
&lt;td&gt;270.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;583.87&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;13,897.53&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14,481.40&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 5: Memory usage in standard mode (main model only, units: MiB)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In other words, setting the max at 131,072 tokens in FP16 mode would require 8,192 MiB, and setting it to the model's full 262,144-token context would require 16,384 MB.&lt;/p&gt;

&lt;p&gt;The cause is that, compared with other models, Qwen3.8-27B has more Global (GQA) layers, and each of those layers is configured with more heads and larger dimensions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Muse Glimmer 30B&lt;/th&gt;
&lt;th&gt;Gemma-4-31B&lt;/th&gt;
&lt;th&gt;Qwen3.8-27B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Layer structure&lt;/td&gt;
&lt;td&gt;SWA+Global(3:1)&lt;/td&gt;
&lt;td&gt;SWA+Global(5:1)&lt;/td&gt;
&lt;td&gt;GDN+Global(3:1)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total layers&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregation layers&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregation layer KV design&lt;/td&gt;
&lt;td&gt;2 heads x 128 dim&lt;/td&gt;
&lt;td&gt;4 heads x 512 dim&lt;/td&gt;
&lt;td&gt;4 heads x 256 dim&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache size (128k, f16)&lt;/td&gt;
&lt;td&gt;1,742.00 MiB&lt;/td&gt;
&lt;td&gt;5,920.00 MiB&lt;/td&gt;
&lt;td&gt;8,192.00 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design intent&lt;/td&gt;
&lt;td&gt;Aims for lightweight inference, pushing memory reduction to the limit&lt;/td&gt;
&lt;td&gt;Balances advanced reasoning with memory reduction&lt;/td&gt;
&lt;td&gt;Aims squarely at advanced reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 6: Aggregation-layer configuration and KV design across hybrid models&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Gated Delta Network is, on its own, a highly efficient architecture compared with the "local + global" pattern — the layers handling information compression require essentially zero KV cache. Under otherwise identical conditions, a "local + global" design will always require a larger cache.&lt;/p&gt;

&lt;p&gt;But Global layers appear very frequently here, the model has more layers overall than its peers, and the aggregation-layer KV design is comparatively rich — and together those factors are what push Qwen3.8-27B's KV cache size up.&lt;/p&gt;

&lt;p&gt;That points clearly to a design goal that prioritizes reliable, deeply reasoned answers over lightweight inference.&lt;/p&gt;

&lt;h4&gt;
  
  
  Output
&lt;/h4&gt;

&lt;p&gt;We asked the following question and examined the response.&lt;/p&gt;

&lt;p&gt;Given the recent Kumamoto earthquake, this is a question we've been asking various models out of curiosity.&lt;sup id="fnref3"&gt;3&lt;/sup&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Describe the characteristics of the ground/soil conditions in Kyushu.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The response came out as follows. Our first impression was that it looked surprisingly well-informed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Kyushu's ground conditions are geologically, topographically, and hydrologically
complex, and their strength and character vary widely from place to place —
that variability is the single defining feature. As construction ground, the
characteristics can be summarized as follows.

1. Geological and ground diversity
Kyushu sits on a volcanic island arc, so volcanic ground, sedimentary ground,
granite ground, weathered soil, and river deposits are all mixed together over
short distances.
- Northern Kyushu (Fukuoka, Saga, Nagasaki, etc.): sedimentary basins and
  granite, with serpentinite bedrock visible in some areas.
- Central Kyushu (Kumamoto, etc.): the Aso caldera, with well-developed
  volcanic ash, pumice, lava, and river deposits.
- Southern Kyushu (Oita, Miyazaki, Kagoshima, etc.): a mix of volcanic ground,
  granite, and sedimentary plains.
As a result, ground conditions differ significantly even within Kyushu itself.

2. Plains are largely soft, recent deposits
Plains and river floodplains around Fukuoka, Saga, Nagasaki, Kumamoto, Oita,
Miyazaki, and Kagoshima often have thick layers of recent deposits — sand,
silt, and clay. Key characteristics: high water content / low shear strength /
prone to compression and settlement / prone to liquefaction during earthquakes
/ prone to differential settlement and subsidence. Soft-ground countermeasures
are especially necessary along rivers and coastal plains.

3. Volcanic ground is widespread
Kyushu has numerous volcanoes, and pumice, volcanic ash, lava, and igneous
rock all contribute to its ground composition. Pumice and volcanic ash often
have low strength, high permeability, and high compressibility, can liquefy
when wet or saturated, and are prone to settlement and changes in earth
pressure. This kind of volcanic-derived ground matters particularly around
Aso, Sakurajima, Kirishima, Kuju, Unzen, and Ibusuki. Lava and igneous rock,
on the other hand, tend to have high strength and can serve as good bearing
layers for foundations, but weathering, cracking, cavities, and erosion make
them non-uniform — the surface may be soft while deeper layers remain strong.
Because volcanic ground places "strong bedrock" right next to "soft volcanic
deposits," ground investigation is especially important here.

4. Mountain and slope areas carry high landslide/collapse risk
Kyushu's mountains are steep, and their slopes are made up of weathered rock,
spoil, volcanic soil, and river deposits. With such steep gradients,
landslides and collapses are easily triggered by rain, typhoons, or
earthquakes, and ground strength varies considerably with the degree of
weathering. Landslide and sediment-disaster risk rises sharply during heavy
rain or seismic events.

5. Active faults and earthquakes have a major effect on the ground
Parts of Kyushu lie near plate boundaries and active fault zones, so ground
motion during earthquakes can be relatively large. Earthquakes can trigger
liquefaction, landslides, slope failure, ground settlement and differential
settlement, and damage to structural foundations. This means seismic design
needs to account for the ground's seismic response, not just the structure
itself.

6. Groundwater and hydrological conditions are also complex
Ground conditions across Kyushu can involve shallow groundwater in the
plains, fissure water in volcanic rock and granite, and saltwater intrusion
along the coast. Construction needs to account for spring water, rising
groundwater levels, strength loss in soft ground, and drainage/water-sealing
during ground improvement.

7. Construction considerations
Because ground conditions vary so much by location across Kyushu,
site-specific investigation is essential rather than relying on general
assessments. Soft ground calls for pile foundations or ground improvement;
lava and bedrock require confirming the bearing layer; volcanic ground needs
liquefaction and settlement countermeasures; slopes need landslide prevention
and anchoring; and ground assessment should account for earthquakes and heavy
rain, alongside a comprehensive look at landslide, liquefaction, and
differential-settlement risk.

Summary
Kyushu's ground is characterized by a combination of soft recent deposits,
low-strength volcanic-derived soil, strong lava and bedrock, and active
faults, earthquakes, and landslides. Because of this, ground strength in
Kyushu can't be judged in general terms — it's important to check each
site's geology, topography, hydrology, and earthquake/rainfall risk
individually.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We had a DeepSearch agent built on GPT-5.6-Luna (powered by Dify) evaluate this response. The results were as follows (we've trimmed a few overly long passages).&lt;/p&gt;

&lt;p&gt;We ran the same question and evaluation against Muse Glimmer 30B previously, and the level of the critique here was slightly higher — this response appears to be operating on a different level.&lt;/p&gt;

&lt;p&gt;GPT-5.6-Luna judged the overall content to be sound, and its feedback moved beyond basic correctness into refinements for readability and presentation — comments aimed at polishing an already-solid answer rather than fixing a flawed one. That's a clear sign it was reading the text at a higher level.&lt;/p&gt;

&lt;p&gt;In this case, it's fair to say this response clearly surpasses Muse Glimmer's.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Overall assessment
As a general-audience overview explaining how much Kyushu's ground varies by
region, this text's overall direction is sound. It covers volcanic ground,
alluvial plains, mountain slopes, earthquakes, and groundwater, and the
structure is easy to follow. That said, using it for construction practice or
specialized education would require correcting a few terms and overly
definitive statements.

As a rough guide: accuracy 6.5/10, readability 8/10, suitability for
construction practice 6/10. The main issues aren't outright factual errors so
much as stating conditionally variable properties as if they were universal,
and grouping geological materials, ground characteristics, and disaster risk
into the same category.

Points needing correction: the phrase "the defining feature" is too
definitive; the description of plains treats their properties too uniformly;
the earthquake section should separate ground characteristics from disaster
hazards; the groundwater section needs qualification; "strong/weak ground"
needs to be broken down further; ground materials and disaster risk should be
kept separate; and some redundancy should be trimmed.

Final assessment: as an introductory, general-audience explanation of
Kyushu's ground, this text is a usable structure. However, phrases like
"recent deposits," "spoil," grouping Oita into southern Kyushu, "volcanic ash
has high permeability," and "can liquefy when wet" need correction or
qualification. Once revised, it's suitable as a general-audience explanation.
For construction planning or foundation-type selection, this text should not
be used as the basis for decisions — it needs to be combined with geological
maps, land-condition maps, liquefaction hazard maps, active-fault data,
boring surveys, standard penetration tests, and groundwater-level surveys.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Performance
&lt;/h4&gt;

&lt;p&gt;We compared this against Muse Glimmer 30B as well. Throughput came out almost identical, but token usage to reach an answer was nearly three times higher — and, correspondingly, so was the time it took.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Qwen3.8-27B]
prompt eval time = 303.63 ms / 64 tokens ( 4.74 ms per token, 210.78 tokens per second)
       eval time = 248375.97 ms / 5771 tokens (43.05 ms per token, 23.23 tokens per second)
      total time = 248679.60 ms / 5835 tokens

[Muse Glimmer]
prompt eval time = 262.18 ms / 68 tokens ( 3.86 ms per token, 259.37 tokens per second)
       eval time = 82874.28 ms / 1866 tokens (44.41 ms per token, 22.52 tokens per second)
      total time = 83136.46 ms / 1934 tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both runs used a high reasoning-effort setting, and the pattern that's generally true of the Qwen series — taking a long time to reach an answer — still held here.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verification: Using Vision
&lt;/h3&gt;

&lt;p&gt;Qwen3.8 has been released as a multimodal model since earlier versions, and this one naturally ships with a vision encoder too. In llama.cpp, attaching the vision-encoder model lets you feed it images, as shown below.&lt;/p&gt;

&lt;p&gt;This is roughly where things start to get tight — if you want more context length, you'll need to quantize the KV cache.&lt;/p&gt;

&lt;h4&gt;
  
  
  Command Line
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;build/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; ./models/Qwen3.8-27B-UD-Q2_K_XL.gguf &lt;span class="nt"&gt;-t&lt;/span&gt; 12 &lt;span class="nt"&gt;-np&lt;/span&gt; 1 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 0.6 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 20 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--fit&lt;/span&gt; off &lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="nt"&gt;-c&lt;/span&gt; 65535 &lt;span class="nt"&gt;--reasoning&lt;/span&gt; on &lt;span class="nt"&gt;--reasoning_effort&lt;/span&gt; xhigh &lt;span class="nt"&gt;--log-verbosity&lt;/span&gt; 4 &lt;span class="nt"&gt;--mmproj&lt;/span&gt; models/mmproj-BF16.gguf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Memory Usage
&lt;/h4&gt;

&lt;p&gt;Memory usage came out as follows, showing an increase of roughly 1,336 MiB.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Device&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;th&gt;CUDA0&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Xeon E5-2690v4&lt;/td&gt;
&lt;td&gt;RTX-A4000&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weight data (Q2_K-quant_XL)&lt;/td&gt;
&lt;td&gt;397.85&lt;/td&gt;
&lt;td&gt;9,567.89&lt;/td&gt;
&lt;td&gt;9,965.74&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache size (f16)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;4,096.00&lt;/td&gt;
&lt;td&gt;4,096.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RS buffer size (f32)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;149.62&lt;/td&gt;
&lt;td&gt;149.62&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gated Delta Net compute buffer&lt;/td&gt;
&lt;td&gt;186.02&lt;/td&gt;
&lt;td&gt;84.02&lt;/td&gt;
&lt;td&gt;270.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision weight data&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;887.99&lt;/td&gt;
&lt;td&gt;887.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision metadata&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.12&lt;/td&gt;
&lt;td&gt;0.12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;583.87&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14,785.64&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;15,369.51&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 7: Memory usage with the vision encoder enabled (units: MiB)&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Prompt and Output
&lt;/h4&gt;

&lt;p&gt;Rather than text-only inference, we tested the model's analytical ability by attaching an image and asking it questions about that image.&lt;/p&gt;

&lt;p&gt;For this, we had the model analyze the Muse Glimmer architecture diagram from our previous article, reusing the same image&lt;sup id="fnref4"&gt;4&lt;/sup&gt;, and asked it to "analyze this in detail and describe everything written in it."&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Note: since this article's author writes in Japanese, this test was conducted using a Japanese-language diagram as the input image. The model's response below — translated into English — conveys the diagram's content, so the image itself is kept as-is here rather than redrawn in English.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfps51m6t3sx4qud8708.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfps51m6t3sx4qud8708.png" alt="The image fed into the Vision Encoder" width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 4: The image fed into the Vision Encoder&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The result was as follows. The passages noted below are places where the model's interpretation was mistaken.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MUSE GLIMMER 30B Architecture Slide: Detailed Analysis

1. Overall composition of the image
Under the title "Reference: MUSE GLIMMER 30B Structure" (dark navy header
bar), the slide has three layers: top text (left = basic model specs, right
= definitions of the SA/FA attention mechanisms), a central callout
(supplementary explanation of block composition), and a bottom flow diagram
(input -&amp;gt; embedding -&amp;gt; Transformer block group -&amp;gt; output). Its purpose is to
visualize the architecture of a 29.6B-parameter dense multimodal LLM and the
design philosophy of its core SA/FA hybrid attention.

2-5. Analysis of parameters, attention mechanism, flow diagram, and callout
(Dense 29.6B, hidden size 6,656, SwiGLU, RoPE theta=50,000 not applied to FA,
max context 128k, DFlash-based drafter, and the respective roles of SA/FA
are laid out in detail exactly as shown in the diagram)

6. Consistency and notes (analytical observations)
Block-count inconsistency: the diagram label reads "52-block configuration"
while the callout says "50 blocks." 52 = 4x13 is consistent as a count of
groups; 50 = 12x4+2 is consistent from the text's own framing. This gap most
likely comes from counting Glimmer Vision's two layers separately — it's
natural to read the diagram's "52" as the total including vision layers, and
the callout's "50" as the text-decoder block count alone. Unifying the
definition within the slide would be preferable. Possible typo: the FA
description "no positional information" appears to be a misprint of "no
positional encoding."

7. Summary
This slide explains that MUSE GLIMMER 30B, built on a dense 29.6B-parameter
model with a 128k context, repeats SA (local, low-cost) and FA (global,
high-quality) in a 3:1 ratio, concentrates positional encoding on SA layers
while removing it from FA layers, and achieves a hybrid attention design
that balances compute efficiency with reasoning performance. Vision-input
fusion, GQA + sigmoid gating, and DFlash/MTP-based inference support are the
other key design elements, with the only consistency issue being the 50/52
block-count discrepancy between the diagram and the text.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response came back roughly twice as long as what Muse Glimmer 30B produced for the same task previously. On top of that, the summary was nearly flawless, and it even called out an inconsistency in how we'd labeled our own diagram.&lt;/p&gt;

&lt;p&gt;The vision encoder is still built on SigLIP2, but additional fine-tuning may well be reflected here — the way it reasons from what it extracts felt noticeably more capable.&lt;/p&gt;

&lt;h3&gt;
  
  
  A Look Inside the Thinking Section
&lt;/h3&gt;

&lt;p&gt;Qwen models are well known for a distinctly self-doubting internal voice during their thinking phase — the model seems to talk itself into second-guessing everything right before it's about to answer, wanting to double-check just one more thing or start over from scratch, dragging out its reasoning in the process. We checked whether that pattern still holds here.&lt;/p&gt;

&lt;p&gt;We've summarized the actual thinking-section output for the Kyushu-ground question in the appendix at the end of this article. Here, we'll first go over the key points from having Google Gemini 3.7 Flash evaluate that output.&lt;/p&gt;

&lt;p&gt;According to Gemini's analysis, Qwen3.8-27B's thinking pattern breaks down into five main traits. First, a separation between the thinking language (English) and the output language (Japanese) — a multilingual-processing behavior where internal reasoning happens in English but the final answer comes out in Japanese. Second, multi-angle context inference and intent-filling, where it anticipates the underlying purpose behind the question (construction practice, exam prep, and so on). Third, exhaustive knowledge brainstorming and categorization, where it lists out related keywords comprehensively and then organizes them by region and geology. Fourth, real-time self-correction and fact-checking — starting to write "Miyagi" and immediately correcting it to "Miyazaki," or avoiding overgeneralizing where serpentinite is distributed, repeatedly double-checking small factual details. And fifth, iterative simulation of output structure and formatting — trying out multiple drafts of whether to use a table or bullet points, and in what order to place the headings.&lt;/p&gt;

&lt;p&gt;The important point here is that the overtly negative emotional tone common through Qwen3.5 doesn't show up — the evaluation describes the process as proceeding "extremely logically." Since there's been no architectural change at all, this improvement appears to come down to training methodology.&lt;/p&gt;

&lt;p&gt;So what exactly changed in how this model was trained? A clue seems to be in the following post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen3.8-Max: A New Bar for Coding and Cowork&lt;/strong&gt;&lt;br&gt;
(&lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;https://qwen.ai/blog?id=qwen3.8&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;This post describes the flagship Qwen3.8-2.4T-A95B model, but given that Qwen3.8-27B is a smaller sibling, it likely went through similar training. In the "Work" section, the post explains that scaling up real-world RL systems (continuously scaling tasks, workspaces, and toolchains along independent axes), a general-purpose reward system capable of internalizing diverse forms of verification, and an online data balancer were combined to lift task-completion capability uniformly across multiple major toolchains.&lt;/p&gt;

&lt;p&gt;That distinctive self-doubting quality, present since Qwen3, doesn't actually originate with Qwen — it traces back to a training method DeepSeek-AI calls &lt;strong&gt;"Long CoT Cold Start"&lt;/strong&gt;&lt;sup id="fnref5"&gt;5&lt;/sup&gt;, applied to DeepSeek-R1 to boost reasoning ability in cases requiring many turns of back-and-forth. The "cold start" in the name reportedly refers to the tendency to pause mid-reasoning and start over from scratch.&lt;/p&gt;

&lt;p&gt;Rather than relying solely on QA-style datasets, this method trains the trial-and-error process directly by drawing on data from sources like programming forums such as Reddit — and that data is reportedly full of human emotional expression, which is said to be what feeds into the model's self-doubting tone. This trait shows up not just in Qwen but broadly across DeepSeek, its originator, and other Chinese LLMs.&lt;/p&gt;

&lt;p&gt;Qwen3.8 appears to have revisited various stages of its training pipeline, and we found a related paper, &lt;strong&gt;"Unified Data Selection for LLM Reasoning"&lt;/strong&gt;&lt;sup id="fnref6"&gt;6&lt;/sup&gt;, which we looked into.&lt;/p&gt;

&lt;p&gt;The paper proposes a new, training-free, and extremely lightweight data-evaluation metric called High-Entropy Sum (HES), and applies it across the SFT (supervised fine-tuning), RFT (rejection-sampling fine-tuning), and RL (reinforcement learning) stages to measure its effect.&lt;/p&gt;

&lt;p&gt;With conventional training (based on average entropy), the important decision points in the reasoning process get diluted by a large volume of low-entropy tokens — boilerplate text, simple arithmetic logic — so the model ends up training without ever distinguishing good solutions from bad ones. Using HES instead, the paper reports, lets the optimized model learn high-quality solutions much more efficiently, with a clearly measurable training benefit.&lt;/p&gt;

&lt;p&gt;Training a model at the scale of Qwen3.8-Max presumably involves a huge volume of answer data with substantial reasoning steps built in. Our guess is that applying the HES concept to that data made more effective solution paths — and by extension, more logical reasoning routes — easier to select, suppressing emotional noise and, as a result, producing more advanced reasoning ability.&lt;/p&gt;

&lt;p&gt;It's also worth noting that issues with Long-CoT-style training have already been pointed out elsewhere, such as in the paper &lt;strong&gt;"Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt"&lt;/strong&gt;&lt;sup id="fnref7"&gt;7&lt;/sup&gt;, and it's likely that some countermeasure along those lines was applied here as well.&lt;/p&gt;

&lt;p&gt;Some people still argue that Chinese models simply distill capability out of frontier models. Based on what we're seeing here, that framing doesn't quite fit Qwen3.8-27B — distillation alone doesn't look like the whole story behind its capability gains. If anything, the Qwen team appears to actively track and adopt a wide range of techniques, and that appetite for new ideas seems to be a real part of what makes this model work.&lt;/p&gt;

&lt;p&gt;This shift in training methodology may also have changed how the model arrives at its answers — the reasoning path itself looks noticeably different from before. If you're doing harness engineering around this model, some of your existing prompting approaches may no longer work as expected, and revisiting them may become necessary if something that used to work stops working.&lt;/p&gt;

&lt;p&gt;As AI models continue to evolve, the way they process information keeps changing along with them — and it may be up to us to keep updating our own approach in step.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks: How the Wider World Sees It
&lt;/h2&gt;

&lt;p&gt;Below is the benchmark comparison published by Alibaba's Qwen team. It shows results on par with — or better than — Anthropic's Claude Opus 4.6 in Max Reasoning mode. Based on what we've seen in our own hands-on testing, that claim doesn't look far-fetched.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Qwen3.8-27B&lt;/th&gt;
&lt;th&gt;Qwen3.6-27B&lt;/th&gt;
&lt;th&gt;Qwen3.7-Plus&lt;/th&gt;
&lt;th&gt;Muse Glimmer-30B&lt;/th&gt;
&lt;th&gt;Opus4.6 Max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;Agentic terminal coding (Terminal Bench 2.1)&lt;/td&gt;
&lt;td&gt;73&lt;/td&gt;
&lt;td&gt;63.4&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;td&gt;51.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;78.2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;Agentic coding (SWE-bench Pro)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;61.7&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;53.5&lt;/td&gt;
&lt;td&gt;57.6&lt;/td&gt;
&lt;td&gt;51.2&lt;/td&gt;
&lt;td&gt;53.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;Repo-level code generation (NL2Repo-Bench)&lt;/td&gt;
&lt;td&gt;42.3&lt;/td&gt;
&lt;td&gt;36.2&lt;/td&gt;
&lt;td&gt;41.1&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;47.6&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;Agentic coding (DeepSWE 1.1)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;42.2&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;13.3&lt;/td&gt;
&lt;td&gt;14.2&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;Software engineering (QwenSWEBench)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;79&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;49.3&lt;/td&gt;
&lt;td&gt;59.2&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;63.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Long-horizon office work (CoWorkBench)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70.7&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;61&lt;/td&gt;
&lt;td&gt;65.1&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;68.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Professional job tasks (JobBench)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;33.4&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;21.8&lt;/td&gt;
&lt;td&gt;27.6&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Frontier agentic tasks Pass@1 (Agents' Last Exam)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.4&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;10.6&lt;/td&gt;
&lt;td&gt;13.2&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Frontier agentic tasks Score (Agents' Last Exam)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;42.9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;27.3&lt;/td&gt;
&lt;td&gt;33.6&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General&lt;/td&gt;
&lt;td&gt;Instruction following (IFBench)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;79.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;69.1&lt;/td&gt;
&lt;td&gt;79.1&lt;/td&gt;
&lt;td&gt;77&lt;/td&gt;
&lt;td&gt;62.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General&lt;/td&gt;
&lt;td&gt;Scientific reasoning (GPQA Diamond)&lt;/td&gt;
&lt;td&gt;89.2&lt;/td&gt;
&lt;td&gt;87.8&lt;/td&gt;
&lt;td&gt;90.3&lt;/td&gt;
&lt;td&gt;83.5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General&lt;/td&gt;
&lt;td&gt;Multidisciplinary reasoning (HLE)&lt;/td&gt;
&lt;td&gt;30.8&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;34.7&lt;/td&gt;
&lt;td&gt;22&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;40&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;General&lt;/td&gt;
&lt;td&gt;Competitive coding (LiveCodeBench v6)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;90.3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;83.9&lt;/td&gt;
&lt;td&gt;89.6&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;td&gt;88.8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What about third-party evaluations? Here's Artificial Analysis's Intelligence Index chart.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fferhtxohijnqjwpr7ysg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fferhtxohijnqjwpr7ysg.png" alt="Artificial Analysis Intelligence Index chart" width="800" height="218"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 5: Artificial Analysis Intelligence Index chart&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;With thinking mode on, it posted a striking score of 52 — not just surpassing Claude Opus 4.6, but matching GPT-5.6 Luna, likely OpenAI's most widely used model at the moment. Seeing the score climb this far came as a genuine surprise to us.&lt;/p&gt;

&lt;p&gt;That said, while this behavior works well for Deep Research, it also means the model is likely to be slow for general-purpose use that mixes in casual conversation, and given how many new approaches went into it, some retuning of middleware and instructions is clearly going to be needed.&lt;/p&gt;

&lt;p&gt;In that sense, Meta's Muse Glimmer 30B, covered earlier, and Google DeepMind's longer-established Gemma-4-31B-it, still have plenty to offer. Their thinking phase is far shorter than Qwen3.8's, so they can produce answers faster.&lt;/p&gt;

&lt;p&gt;In particular, Muse Glimmer 30B and Gemma-4 both benefit heavily from sliding-window memory savings, making it relatively easy to extend their context length. Muse Glimmer 30B's vision encoder is also appealing for its resolution. Given these different strengths, rather than defaulting to a single model for everything, what seems to matter going forward is the skill of choosing and applying the right model for the task.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;This article covered Qwen3.8-27B, the compact model the local-LLM community had been waiting for. Its capability is no exaggeration — it delivered performance beyond what we expected. The benchmark results published on Artificial Analysis in particular likely came as a considerable surprise to a lot of people.&lt;sup id="fnref8"&gt;8&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;The first use case that comes to mind is Deep Research. Paired with a search engine, it can dig into deep content without difficulty, and its resistance to falling into wasteful reasoning loops makes it a reassuring tool for search-augmented use.&lt;/p&gt;

&lt;p&gt;Architecturally, it carries forward Qwen3.5's Gated Delta Network-based hybrid design, and being able to run it as-is without updating your inference engine is a significant point in its favor.&lt;/p&gt;

&lt;p&gt;That said, given how firmly this design leans toward raw capability, its tendency toward a large KV cache footprint is a genuine drawback. Keeping the max context length very long requires a correspondingly capable GPU, and that memory overhead is likely to be a real headache for local-LLM users.&lt;/p&gt;

&lt;p&gt;On the training side, the shift away from the traditional Long CoT Cold Start method has also left some users unsure what to make of it. Figuring out how to adapt to this change, and what prompt adjustments it calls for, will likely be the key going forward.&lt;/p&gt;

&lt;p&gt;As findings from this new approach get shared through arXiv and elsewhere, it will likely spread gradually to other models. Efforts like this could act as a trigger that meaningfully lifts capability across both commercial frontier models and open-weight models alike.&lt;/p&gt;

&lt;p&gt;New technical advances don't just show up in architecture and tokenizer specs — they also show up inside the thinking section, a part of the process that's normally hidden from view. It's worth paying attention to that usually-overlooked area and watching how model behavior changes there.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Qwen/Qwen3.8-27B -- Hugging Face&lt;br&gt;
&lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B" rel="noopener noreferrer"&gt;https://huggingface.co/Qwen/Qwen3.8-27B&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;unsloth/Qwen3.8-27B-GGUF - Hugging Face&lt;br&gt;
&lt;a href="https://huggingface.co/unsloth/Qwen3.8-27B-GGUF" rel="noopener noreferrer"&gt;https://huggingface.co/unsloth/Qwen3.8-27B-GGUF&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenLM-AI Qwen3.8&lt;br&gt;
&lt;a href="https://openlm.ai/qwen3.8/" rel="noopener noreferrer"&gt;https://openlm.ai/qwen3.8/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Qwen3.8-Max: A New Bar for Coding and Cowork&lt;br&gt;
&lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;https://qwen.ai/blog?id=qwen3.8&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2501.12948" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2501.12948&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Unified Data Selection for LLM Reasoning&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2605.22389" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2605.22389&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2505.23480" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2505.23480&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Appendix
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Expected KV Cache Sizes
&lt;/h3&gt;

&lt;p&gt;Qwen3.8-27B requires the following KV cache sizes. Note that supporting long context requires these amounts of VRAM on top of the weight data and vision-encoder data.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Max Context Size&lt;/th&gt;
&lt;th&gt;K:f16, V:f16&lt;/th&gt;
&lt;th&gt;K:q8_0, V:q8_0&lt;/th&gt;
&lt;th&gt;K:q8_0, V:turbo3&lt;sup id="fnref9"&gt;9&lt;/sup&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;td&gt;4,096 MB&lt;/td&gt;
&lt;td&gt;2,048 MB&lt;/td&gt;
&lt;td&gt;1,488 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;131,072&lt;/td&gt;
&lt;td&gt;8,192 MB&lt;/td&gt;
&lt;td&gt;4,096 MB&lt;/td&gt;
&lt;td&gt;2,976 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;262,144&lt;/td&gt;
&lt;td&gt;16,384 MB&lt;/td&gt;
&lt;td&gt;8,192 MB&lt;/td&gt;
&lt;td&gt;5,952 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 8: Expected KV cache sizes&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Thinking-Section Content for the Kyushu Ground Question (Summary)
&lt;/h3&gt;

&lt;p&gt;Summarizing the actual thinking-section content for the "describe the characteristics of Kyushu's ground" question covered in the main text (the original was output in English), the trial-and-error process went roughly as follows. The full text is extremely long, so only the key points are excerpted here.&lt;/p&gt;

&lt;p&gt;It begins by questioning itself over whether "ground" refers to geological bedrock or to civil-engineering/construction-style ground conditions, and considers the question's likely intent from multiple angles — an exam question, construction practice, or general knowledge. It then exhaustively lists out the elements worth covering: plate tectonics, volcanic activity, sedimentary layers, liquefaction risk, and regional differences (north/central/south).&lt;/p&gt;

&lt;p&gt;Along the way, it repeatedly double-checks small factual details. For instance, it starts to write "Miyagi," catches that it's not a Kyushu prefecture, and immediately corrects it to "Miyazaki"; it also reconsiders where serpentinite is actually distributed in Kyushu and settles on the more limited phrasing "in some areas" — self-checking specific place names and geological terms at every step.&lt;/p&gt;

&lt;p&gt;It also repeatedly adjusts its tone — not overgeneralizing, not overusing jargon, not making the answer too long — and trials multiple structural options: bullet points, paragraphs, or a table. It eventually settles on a draft that includes a table with "aspect | characteristic" columns, then tries out several versions of the concluding sentence, fine-tuning the wording right up to the end.&lt;/p&gt;

&lt;p&gt;The sheer volume of this trial-and-error is what substantiates the self-doubting quality mentioned earlier in the article. The actual output — the Kyushu-ground answer quoted in the main text — is concise and tidy, but the internal process behind it involves repeated self-questioning and backtracking. That gap between process and output was striking.&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;Some users, working from the assumption that the model simply wouldn't fit on their hardware, went as far as building their own memory-reduction scheme and inference engine from scratch to verify it themselves. In 2-bit quantized mode, they reportedly recorded results well beyond what anyone expected.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;For more detail on this setup, see our earlier article on using GPUSOROBAN. (&lt;a href="https://www.bluecore.net/archives/242" rel="noopener noreferrer"&gt;https://www.bluecore.net/archives/242&lt;/a&gt;)&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;This article's author is based in Japan, so the model was prompted in Japanese and asked about a Japan-specific topic (the ground conditions behind the recent Kumamoto earthquake) — hence the choice of question here and elsewhere in this article.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn4"&gt;
&lt;p&gt;This image uses an earlier version of the model diagram than the one shown previously in this article.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn5"&gt;
&lt;p&gt;DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (&lt;a href="https://arxiv.org/pdf/2501.12948" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2501.12948&lt;/a&gt;)&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn6"&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2605.22389" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2605.22389&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn7"&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2505.23480" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2505.23480&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn8"&gt;
&lt;p&gt;The people running these benchmarks may well have rubbed their eyes and re-verified the results at least once, wondering if something had gone wrong — that's roughly how long it took for Qwen3.8-27B's results to come out relative to what we'd expect.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn9"&gt;
&lt;p&gt;Measured using a llama.cpp fork that supports TurboQuant. "turbo3" indicates TurboQuant 3-bit quantization. Note that this was tested on a 2-GPU setup, so some overhead is present.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>qwen</category>
    </item>
    <item>
      <title>Muse Glimmer 30B: An Architecture and Hands-On Look at Meta's New Mid-Range Model</title>
      <dc:creator>oooocean66</dc:creator>
      <pubDate>Fri, 18 Sep 2026 08:35:09 +0000</pubDate>
      <link>https://dev.to/oooocean66/muse-glimmer-30b-an-architecture-and-hands-on-look-at-metas-new-mid-range-model-126c</link>
      <guid>https://dev.to/oooocean66/muse-glimmer-30b-an-architecture-and-hands-on-look-at-metas-new-mid-range-model-126c</guid>
      <description>&lt;p&gt;In the world of generative AI, Meta's LLaMa is often considered the model that kicked off the open-weight era. Its successor, LLaMa-4, however, fell short of expectations and was widely regarded as a disappointment, leaving Meta in a difficult position over the past year.&lt;/p&gt;

&lt;p&gt;Then, in early 2026, Meta released a new lineup called the "Muse" series. Muse Glimmer 30B, the mid-range model in that lineup covered here, appears to have drawn considerable attention in the local-LLM community.&lt;/p&gt;

&lt;p&gt;This article walks through Muse Glimmer 30B's architecture and then verifies it hands-on on GPU hardware: standard text inference, its combination with the speculative-decoding technique DFlash, and its Vision capability. We look at what the model can actually do, based on the results below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meta's Muse Series
&lt;/h2&gt;

&lt;h3&gt;
  
  
  LLaMa, the Model That Started the Open-Weight Era
&lt;/h3&gt;

&lt;p&gt;One of Meta AI's most significant contributions was releasing the LLaMa models as open weights.&lt;br&gt;
At the time, high-performance chat models like GPT-3.5-Turbo existed, but they were closed models that couldn't be run on a typical local GPU.&lt;/p&gt;

&lt;p&gt;To address that gap, Meta released LLaMa free of charge.&lt;br&gt;
Version 1 came with restrictive licensing terms — even Alpaca, a derivative model, had commercial-use restrictions — but Version 2 relaxed those terms considerably.&lt;br&gt;
That change made it much easier for a broad range of users to adopt LLMs locally, and the model drew significant attention as a result.&lt;/p&gt;

&lt;p&gt;LLaMa 2 in particular was notable for being the first major model to adopt RoPE (Rotary Positional Embedding)&lt;sup id="fnref1"&gt;1&lt;/sup&gt; and Group Query Attention (GQA)&lt;sup id="fnref2"&gt;2&lt;/sup&gt;, both novel techniques at the time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input tokens -&amp;gt; Embedding -+-&amp;gt; [ RMSNorm -&amp;gt; Self-Attention (GQA + RoPE) ] -+
                            |                                              |
                            +&amp;lt;-----------------------------(residual add)--+
                            |
                            +-&amp;gt; [ RMSNorm -&amp;gt; SwiGLU FFN ] -+
                            |                               |
                            +&amp;lt;------------(residual add)----+
                                   |
                            (x N decoder blocks)
                                   |
                            RMSNorm -&amp;gt; Linear -&amp;gt; Softmax -&amp;gt; Output tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Figure 1: LLaMa-2's architecture — a standard pre-norm decoder block, with GQA (fewer key/value heads shared across query heads) and RoPE-based position encoding applied at every attention layer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;These architectural choices remain standard across many open-weight models today, and models such as llm-jp-4-8b-thinking still carry that lineage forward.&lt;sup id="fnref3"&gt;3&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;However, LLaMa-4 — the successor released after the success of LLaMa-3 — underperformed relative to expectations and was widely characterized as a disappointment. Meta itself acknowledged the setback, which was accompanied by organizational disruption within its research team and the departure of several researchers — an unusually visible decline for a lab of its standing.&lt;/p&gt;

&lt;h3&gt;
  
  
  From LLaMa to Muse
&lt;/h3&gt;

&lt;p&gt;In April 2026, Meta announced a new model called Muse Spark.&lt;/p&gt;

&lt;p&gt;Muse Spark launched as a closed model and initially drew little attention. Versions 1.1 and 1.2 followed in July and August respectively, and as each new version shipped, its performance gradually became better known.&lt;/p&gt;

&lt;p&gt;Around the release of Muse Spark 1.2, Meta also introduced Muse Glimmer 30B, the mid-range model covered in this article. Its release appears to have generated notable interest within the local-LLM community.&lt;/p&gt;

&lt;h2&gt;
  
  
  Muse Glimmer's Architecture
&lt;/h2&gt;

&lt;p&gt;Muse Glimmer uses a hybrid structure similar to several currently popular open-weight models. The architecture is shown below.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input strings -+- Glimmer Vision (image input) -----------------+
               |                                                 |
               +- tok -&amp;gt; Emb -----------------------------------+-(+)-+
                                                                       |
                                RoPE (SA layers: theta=50,000)  &amp;lt;------+
                                RoPE (FA layers: theta=0 = NoPE) &amp;lt;-----+
                                                                       |
     [ SA -&amp;gt; SA -&amp;gt; SA -&amp;gt; FA ]  x13 groups  -------------------------&amp;gt;
            (52 blocks total = 4 layers x 13 groups)
                                                                       |
                                                          LNR -&amp;gt; SoftMax -&amp;gt; Output Probabilities
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Figure 2: Muse Glimmer 30B's architecture. SA = Sliding Attention (local context, RoPE applied); FA = Full Attention (aggregates and consolidates, RoPE disabled / NoPE).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;At a glance, the block layout resembles Qwen's more than it does Google DeepMind's Gemma-4 — it's a comparatively simple structure. It's a 29.6B-parameter, dense model that uses GeLU for its activation function.&lt;/p&gt;

&lt;h3&gt;
  
  
  Attention Structure
&lt;/h3&gt;

&lt;p&gt;The attention structure is a hybrid built on top of Group Query Attention.&lt;/p&gt;

&lt;p&gt;Layers labeled SA use Sliding Window Attention, which handles local context analysis. By explicitly bounding the context range each layer looks at, this reduces both memory usage and compute cost.&lt;/p&gt;

&lt;p&gt;Layers labeled FA use Full Attention. These layers aggregate and consolidate what the preceding SA layers produced, running as standard GQA — which makes them comparatively more compute-heavy than the SA layers.&lt;/p&gt;

&lt;p&gt;Blocks are arranged in groups of three SA layers followed by one FA layer, repeated 13 times for a total of 52 blocks.&lt;/p&gt;

&lt;p&gt;The 3:1 SA-to-FA ratio resembles the structure used since Qwen3.5, but looking at what each layer type actually does, the closer functional match is Gemma-4's "local + global" design.&lt;sup id="fnref4"&gt;4&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;A more distinctive detail is that the RoPE theta value can be set independently per layer: SA layers use 50,000, while FA layers use 0.&lt;/p&gt;

&lt;p&gt;Looking at the &lt;code&gt;config.json&lt;/code&gt; in the Hugging Face repository shows this setting directly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"architectures"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"MuseGlimmerForConditionalGeneration"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"dtype"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bfloat16"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"image_token_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;200092&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hidden_size"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;6656&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"initializer_range"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.02&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"intermediate_size"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;19968&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"layer_rope_theta"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="mf"&gt;500000.0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the modeling source code, a theta value of 0 causes the RoPE module to be treated as NoPE — meaning positional encoding is skipped entirely for that layer.&lt;/p&gt;

&lt;p&gt;This behavior is confirmed in Hugging Face's transformers library:&lt;/p&gt;

&lt;p&gt;(&lt;a href="https://github.com/huggingface/transformers/blob/main/src/transformers/models/muse_glimmer/modeling_muse_glimmer.py" rel="noopener noreferrer"&gt;https://github.com/huggingface/transformers/blob/main/src/transformers/models/muse_glimmer/modeling_muse_glimmer.py&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;In a separate file, &lt;code&gt;modular_muse_glimmer.py&lt;/code&gt;, the code sets &lt;code&gt;position_embeddings&lt;/code&gt; to &lt;code&gt;None&lt;/code&gt; whenever the RoPE theta parameter is 0. The class that defines the model's actual forward pass then branches on that: if &lt;code&gt;position_embeddings&lt;/code&gt; is &lt;code&gt;None&lt;/code&gt; (i.e., NoPE), it skips the positional-encoding step entirely.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;MuseGlimmerTextAttention&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Module&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;forward&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;hidden_states&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;position_embeddings&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;attention_mask&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;past_key_values&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Cache&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Unpack&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;TransformersKwargs&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tensor&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# NoPE layers receive `position_embeddings=None` from the model.
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;position_embeddings&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sin&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;position_embeddings&lt;/span&gt;
            &lt;span class="n"&gt;query_states&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key_states&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;apply_rotary_pos_emb&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query_states&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key_states&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sin&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="p"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This suggests a design choice to keep FA layers focused purely on the aggregation role, stripping out anything unrelated to that job and drawing a clear division of labor between the two layer types.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vision Encoder
&lt;/h3&gt;

&lt;p&gt;The vision model, called Glimmer Vision, is a custom encoder rather than the widely used SigLIP vision transformer that most models rely on.&lt;/p&gt;

&lt;p&gt;It has 50 blocks — 12 groups of (3 SA + 1 FA), plus one more (1 SA + 1 FA) — and its structure is nearly identical to the text decoder's, differing only at the very end. The resulting vector is treated as 1,024 tokens' worth of information and fed into the text decoder.&lt;/p&gt;

&lt;p&gt;This design is described in Meta's own paper, "&lt;a href="https://arxiv.org/abs/2504.13181" rel="noopener noreferrer"&gt;Perception Encoder: The best visual embeddings are not at the output of the network&lt;/a&gt;" (arXiv:2504.13181v2, published April 28, 2025), and Glimmer Vision appears to be built on the Perception Encoder architecture proposed there.&lt;/p&gt;

&lt;p&gt;Vision encoder design has diversified across model vendors this year — some follow the approach above, while others, like Gemma-4-12B-it, fold nearly all of the weight data into the text encoder itself (aside from CLIP). It's possible this custom approach delivers stronger analytical capability than a straightforward SigLIP-based encoder, though that would need to be verified directly — which is one of the things we test below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running It Ourselves
&lt;/h2&gt;

&lt;p&gt;For this verification, we used a GPUSOROBAN &lt;a href="https://soroban.highreso.jp/compute" rel="noopener noreferrer"&gt;High-Speed Computing&lt;/a&gt; instance equipped with an NVIDIA RTX A4000.&lt;/p&gt;

&lt;h3&gt;
  
  
  Overall Setup
&lt;/h3&gt;

&lt;p&gt;The setup is as follows: we open an SSH tunnel to an access server, then reach the target instance through that tunnel.&lt;sup id="fnref5"&gt;5&lt;/sup&gt; llama.cpp's service port is relayed to localhost via port 8001 on both ends.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(The network diagram accompanying this section is omitted here — it uses several Japanese-only labels for "home," "client," "access server," and "target instance," and the setup is already fully described in the text above.)&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Instance Used
&lt;/h3&gt;

&lt;p&gt;The GPU used for this test is an NVIDIA RTX A4000. The instance specifications are as follows.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Instance type&lt;/td&gt;
&lt;td&gt;s16-1-a-standard-ubs24-v&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;td&gt;NVIDIA RTX A4000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory&lt;/td&gt;
&lt;td&gt;GDDR6 16GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory bandwidth&lt;/td&gt;
&lt;td&gt;448.0 GB/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FP32 compute&lt;/td&gt;
&lt;td&gt;19.17 TFLOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BF16 compute&lt;/td&gt;
&lt;td&gt;38.34 TFLOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT8 compute&lt;/td&gt;
&lt;td&gt;153.4 TOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT4 compute&lt;/td&gt;
&lt;td&gt;306.7 TOPS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vCPU&lt;/td&gt;
&lt;td&gt;11 cores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System memory&lt;/td&gt;
&lt;td&gt;50 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Persistent 100 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CUDA version&lt;/td&gt;
&lt;td&gt;13.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVIDIA driver version&lt;/td&gt;
&lt;td&gt;580&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OS&lt;/td&gt;
&lt;td&gt;Ubuntu 24.04 Server&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 1: Instance specifications&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Model Files Used
&lt;/h3&gt;

&lt;p&gt;Even a 4-bit quantized version of Muse Glimmer exceeds 16GiB. Factoring in KV cache usage as well, we used a 2-bit quantized model optimized with Unsloth Dynamic 2.0.&lt;/p&gt;

&lt;p&gt;We downloaded the following files from Unsloth's Muse Glimmer repository:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Muse-Glimmer-30B-UD-IQ2_M.gguf&lt;/td&gt;
&lt;td&gt;12.3 GB&lt;/td&gt;
&lt;td&gt;UD2.0, 2-bit quantized model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mmproj-kquant.gguf&lt;/td&gt;
&lt;td&gt;1.40 GB&lt;/td&gt;
&lt;td&gt;Quantized vision encoder&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dflash-kquant.gguf&lt;/td&gt;
&lt;td&gt;1.63 GB&lt;/td&gt;
&lt;td&gt;Quantized DFlash draft model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 2: Files used with llama.cpp&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We used llama.cpp version 10380 (0b1bad14f).&lt;/p&gt;

&lt;p&gt;Support for Muse Glimmer had already landed via &lt;a href="https://github.com/ggml-org/llama.cpp/pull/26841" rel="noopener noreferrer"&gt;model: Muse Glimmer Support – #26841&lt;/a&gt;, and this build includes that change.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verification: Standard Mode
&lt;/h3&gt;

&lt;p&gt;First, we ran text inference using the main model on its own.&lt;/p&gt;

&lt;h4&gt;
  
  
  Launch Command
&lt;/h4&gt;

&lt;p&gt;We used the following command to launch the web frontend for testing.&lt;/p&gt;

&lt;p&gt;For this run, the KV cache was kept in FP16 rather than quantized to 8-bit. We also raised the log verbosity from the default of 3 to 4 to check memory usage.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;build/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; ./models/Muse-Glimmer-30B-UD-IQ2_M.gguf &lt;span class="nt"&gt;-t&lt;/span&gt; 12 &lt;span class="nt"&gt;-np&lt;/span&gt; 1 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 64 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--fit&lt;/span&gt; off &lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="nt"&gt;-c&lt;/span&gt; 131072 &lt;span class="nt"&gt;--reasoning&lt;/span&gt; on &lt;span class="nt"&gt;--log-verbosity&lt;/span&gt; 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Memory Usage
&lt;/h4&gt;

&lt;p&gt;Memory usage came out as follows. Thanks to the sliding-window design, the footprint looks smaller than what you'd typically see from a model relying solely on standard Grouped Query Attention.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;CUDA0&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weight data&lt;/td&gt;
&lt;td&gt;10,623.10&lt;/td&gt;
&lt;td&gt;1,052.08&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache data&lt;/td&gt;
&lt;td&gt;1,664.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sliding-window KV cache&lt;/td&gt;
&lt;td&gt;97.50&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;273.52&lt;/td&gt;
&lt;td&gt;156.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;12,658.12&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,208.60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 3: Memory usage in standard mode, main model only (units: MiB)&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Output
&lt;/h4&gt;

&lt;p&gt;We asked the following question and examined the response.&lt;/p&gt;

&lt;p&gt;Given the recent Kumamoto earthquake, this is a question we've been asking various models out of curiosity.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Describe the characteristics of the ground/soil conditions in Kyushu.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The response came out as follows. Our first impression was that it looked surprisingly well-informed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Kyushu's ground conditions are among the most affected by volcanic activity
and crustal movement anywhere in the Japanese archipelago, and the defining
characteristic is a foundation dominated by volcanic sediment and soft
alluvial deposits.

1. Geological and Topographic Features
- An island arc ringed by active volcanoes: Located at the southwestern end
  of the Japanese archipelago, on the boundary between the Philippine Sea
  Plate and the Eurasian Plate. Some of Japan's most active volcanoes are
  concentrated here — Mt. Aso, Sakurajima, the Kirishima range, and Mt.
  Unzen — and volcanically derived ground extends across a wide area.
- The Median Tectonic Line runs through the region: A roughly east-west
  fault line crosses nearly all of Kyushu, and active fault zones are
  extremely common. A chain of active faults runs through Fukuoka,
  Kumamoto, and Oita, including the Futagawa and Hinagu fault zones, which
  were the epicenters of the 2016 Kumamoto earthquakes.
- Contrast between mountainous terrain and narrow plains: Inland areas are
  ringed by the Kyushu mountains, interspersed with alluvial lowlands such
  as the Ariake Sea coast, the Chikugo Plain, Kitakyushu, and the
  Miyazaki Plain. The plains have thick sedimentary layers, while the
  mountainous areas are dominated by hard igneous rock.

2. Ground Material Characteristics
- Extensive volcanic ash and ejecta layers: Ash from the Aso caldera's
  eruptions covers central to northern Kyushu in deposits reaching tens of
  meters, and in some places over 100 meters. Pumice, scoria, and
  tuffaceous ash are known for high water content, high compressibility and
  permeability, and a tendency toward consolidation settlement.
- Thick, soft alluvial layers: Lowlands such as the Ariake Sea coast, the
  Chikushi Plain, the Yatsushiro Plain, and the Miyakonojo Basin consist of
  thick alluvial clay, sand, and silt layers. Low N-values and poor bearing
  capacity are common across broad areas of soft ground.
- A mix of volcanic and sedimentary rock: Northern Kyushu's bedrock
  consists of Paleozoic sedimentary rock, Tertiary igneous rock, and hard
  rock formations such as the Mifune Group. Southern Kyushu is dominated by
  more recent volcanic rock, which is hard but weathers easily and is prone
  to collapse.

3. Geotechnical Characteristics
- High liquefaction risk: Reclaimed land and coastal alluvial sand layers —
  particularly around Hakata Bay in Fukuoka, Nagasaki Bay, Yatsushiro Bay
  in Kumamoto, and Kagoshima Bay, as well as sandy ground near river
  mouths — are prone to liquefaction. Widespread liquefaction damage was
  reported during the 2016 Kumamoto earthquakes.
- Ground compressibility and differential settlement: Volcanic ash layers
  combined with organic alluvial clay contribute to differential
  settlement of buildings, road subsidence, and long-term settlement
  following liquefaction. Groundwater tables are shallow and sensitive to
  fluctuation.
- Ground amplification characteristics: Thick, soft volcanic ash and
  alluvial layers amplify seismic motion in many areas. Long-period ground
  motion amplification has been specifically noted around the Aso caldera
  and the Chikugo Plain.
- Cities with extensive reclaimed and improved land: Port cities such as
  Kitakyushu, Nagasaki, Kagoshima, and Sasebo are historically built on
  reclaimed land, and ground improvement is essential across much of these
  areas.

Regional Trends
- Northern Kyushu (Fukuoka, Saga, northern Oita): Dominated by alluvial
  clay in the Chikushi Plain and reclaimed land along the Genkai Sea coast.
  Ground is soft with high liquefaction risk.
- Central Kyushu (Kumamoto, southern Oita): Thickly covered by Aso volcanic
  ash. This ash compacts poorly and is prone to settlement.
- Southern Kyushu (Miyazaki, Kagoshima): A mix of volcanic ejecta directly
  beneath active volcanoes and marine-terrace deposits. Ground consists of
  volcanic soil over shallow bedrock, giving locally high bearing capacity
  in places, but with significant heterogeneity from collapse and ejecta
  deposits.

Overall, Kyushu's ground can be characterized as "heterogeneous, multi-
layered deposits from volcanic activity plus thick, soft alluvial layers,"
making earthquake resistance, ground subsidence, and liquefaction
countermeasures major challenges for urban development in the region.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We had a DeepSearch agent built on GPT-5.6-Luna (powered by Dify) evaluate this response. The results were as follows.&lt;/p&gt;

&lt;p&gt;The evaluation flagged a recurring issue: the response tends to generalize localized Kyushu characteristics as if they applied uniformly across the whole region. Only one point — the claim about the Median Tectonic Line — was flagged as factually incorrect outright; the rest of the issues are mostly about over-generalization.&lt;/p&gt;

&lt;p&gt;Smaller models tend to produce noticeably off-base answers to this kind of question, so getting a response of this quality without RAG suggests fairly strong reasoning ability. That's particularly notable given that we normally test with 4-bit quantized models, and this run used a 2-bit quantized model instead.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The overall thrust of the response is reasonable, but it should be treated
with caution given that it describes all of Kyushu with a single, uniform
set of ground characteristics.

In particular, the following statements need correction:
- "The Median Tectonic Line crosses all of Kyushu": there are multiple
  views on whether and where the Median Tectonic Line runs through Kyushu,
  so this cannot be stated as fact.
- "Located on the boundary between the Philippine Sea Plate and the
  Eurasian Plate": Kyushu is affected by plate motion, but describing it as
  simply sitting on a plate boundary is inaccurate.
- "Aso volcanic ash deposits reach over 100m in central-northern Kyushu":
  this does not distinguish between airfall ash, pyroclastic-flow deposits,
  and redeposited material, and reads as if a uniform 100m+ layer covers
  the entire area.
- "Volcanic ash has high compressibility and permeability": ash, pumice,
  and pyroclastic-flow deposits each have different properties and
  shouldn't be grouped together.
- "Southern Kyushu is dominated by volcanic rock": Shirasu, sedimentary
  rock, alluvial deposits, and marine terraces are also widely distributed
  there.
- "Port cities are built on reclaimed land, making ground improvement
  essential": reclaimed land is only part of the picture, and the need for
  ground improvement varies by ground conditions and structure.

On the other hand, the following points are largely correct:
- Kyushu has many active volcanoes and active faults.
- Soft alluvial layers are distributed along the Ariake Sea coast, the
  Chikushi Plain, and the Kumamoto Plain, among other areas.
- Liquefaction should be considered for reclaimed land, former riverbeds,
  and low-lying areas near river mouths.
- During the Kumamoto earthquakes, the Futagawa and Hinagu fault zones
  were active, and liquefaction damage did occur.
- Consolidation settlement due to soft ground is a recognized issue along
  the Ariake Sea coast.

A more accurate summary, then, would be: Kyushu's ground varies regionally
across volcanic sediment, bedrock, alluvial deposits, and reclaimed land.
Rather than treating volcanic ash and alluvial layers as characteristic of
all of Kyushu, subsidence, liquefaction, slope failure, and ground-motion
amplification should be assessed separately for each region based on its
specific terrain and geology.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Performance
&lt;/h4&gt;

&lt;p&gt;Speed performance came out as follows, with an output rate of about 22.52 tokens/sec. Given the model's size, that's on the lower end compared to 9B–12B models, though it's not slow enough to be a practical problem. That said, raising the reasoning effort further might noticeably increase response time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;prompt eval time =     262.18 ms /    68 tokens (    3.86 ms per token,   259.37 tokens per second)
       eval time =   82874.28 ms /  1866 tokens (   44.41 ms per token,    22.52 tokens per second)
      total time =   83136.46 ms /  1934 tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Verification: Using DFlash
&lt;/h3&gt;

&lt;p&gt;Muse Glimmer supported DFlash from release, and Unsloth's repository includes a dedicated draft model (dflash-kquant.gguf). We loaded it via &lt;code&gt;--model-draft&lt;/code&gt; and re-ran the same question as in standard mode.&lt;/p&gt;

&lt;p&gt;The result was underwhelming: output speed dropped from 22.52 tps to 17.80 tps, and the token acceptance rate came in at only 15.7% (mean length 1.94).&lt;/p&gt;

&lt;p&gt;We plan to dig further into the conditions under which this speculative-decoding technique does or doesn't pay off in a follow-up article.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verification: Using Vision
&lt;/h3&gt;

&lt;p&gt;Muse Glimmer also ships with a vision encoder, with a dedicated mmproj file available in Unsloth's repository. Loading this vision encoder together with the DFlash draft model resulted in an out-of-memory error, so we tested each capability separately.&lt;/p&gt;

&lt;h4&gt;
  
  
  Command Line
&lt;/h4&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;build/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; ./models/Muse-Glimmer-30B-UD-IQ2_M.gguf &lt;span class="nt"&gt;-t&lt;/span&gt; 12 &lt;span class="nt"&gt;-np&lt;/span&gt; 1 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 64 &lt;span class="nt"&gt;--port&lt;/span&gt; 8000 &lt;span class="nt"&gt;-c&lt;/span&gt; 131072 &lt;span class="nt"&gt;--reasoning&lt;/span&gt; on &lt;span class="nt"&gt;--log-verbosity&lt;/span&gt; 4 &lt;span class="nt"&gt;--mmproj&lt;/span&gt; models/mmproj-kquant.gguf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Memory Usage
&lt;/h4&gt;

&lt;p&gt;Memory usage came out as follows, showing an increase of roughly 1,336 MiB.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;CUDA0&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weight data&lt;/td&gt;
&lt;td&gt;10,623.10&lt;/td&gt;
&lt;td&gt;1,052.08&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache data&lt;/td&gt;
&lt;td&gt;1,664.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sliding-window KV cache&lt;/td&gt;
&lt;td&gt;97.50&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;273.52&lt;/td&gt;
&lt;td&gt;156.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision weight data&lt;/td&gt;
&lt;td&gt;1,335.41&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision metadata&lt;/td&gt;
&lt;td&gt;0.28&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;13,993.81&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,208.60&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 4: Memory usage with the vision encoder enabled (units: MiB)&lt;/em&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Prompt and Output
&lt;/h4&gt;

&lt;p&gt;Rather than text-only inference, we tested the model's analytical ability by attaching an image and asking it questions about that image.&lt;/p&gt;

&lt;p&gt;For this, we had the model analyze the same Muse Glimmer architecture diagram shown earlier in this article&lt;sup id="fnref6"&gt;6&lt;/sup&gt; and asked it to "analyze this in detail and describe everything written in it."&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Note: since this article's author writes in Japanese, this test was conducted using a Japanese-language diagram as the input image. The model's response below — translated into English — conveys the diagram's content, so the image itself is kept as-is here rather than redrawn in English.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy07vbdl5alepqe4cu0b8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy07vbdl5alepqe4cu0b8.png" alt="The Japanese-language architecture slide fed into Muse Glimmer's Vision Encoder" width="432" height="243"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 3: The image fed into the Vision Encoder&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The result was as follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;This slide, titled "Reference: MUSE GLIMMER 30B Structure," is a diagram of
the MUSE GLIMMER 30B model architecture's specifications and overall data
flow.

[Spec list, left side]
- Dense-type 29.6B model
- Hidden size: 6,656 (activation function: SwiGLU, dimension 19,968)
- LNR: Linear Layer
- RoPE: theta = 50,000. Notably not applied to Full Attention layers — for
  FA layers, theta is set to 0, which signals NoPE (no positional encoding)
- Max context: 128k tokens (131,072 tokens)
- Includes a DFlash-based drafter; supports MTP

[Terminology, right side]
- SA: Sliding Attention
  - GQA-based, with a sigmoid output gate
  - Handles local context analysis via a sliding window
  - Low compute cost, fast
  - Low memory usage
- FA: Full Attention
  - GQA-based, with a sigmoid output gate
  - Aggregates and consolidates information from Sliding Attention layers
  - Maintains global structure and handles complex reasoning
  - Uses no positional encoding — positional information is left entirely
    to Sliding Attention

[Central block diagram flow]
Input strings are tokenized (tok) and embedded (Emb). RoPE is added to the
embeddings. Separately, Glimmer Vision receives input from an image icon,
and its output is added immediately after the embedding + RoPE step.

The language model body that follows is described as having 52 blocks
total: 52 blocks = 4 layers x 13 groups.

Each group runs SA → SA → SA → FA in sequence, with SA shown in blue and FA
in red. Thirteen of these groups are chained together.

At the end of each group's output, LNR normalizes the result, which then
passes through Softmax to produce the output probabilities.

In short: taking embeddings + vision features + RoPE as input, the model
runs SA three times for local analysis, then FA once to aggregate and
consolidate — a 4-layer unit repeated across 13 groups for 52 blocks total,
forming a dense causal model. The vision encoder follows a similarly dense
50-block structure, with a 128k context window, SwiGLU activation, and the
same ability to disable RoPE for Full Attention layers.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The content is largely accurate, with only minor issues, and we judged the analytical quality to be high. The response correctly captures what each element means, how the pieces connect, and even the detail that Full Attention layers use NoPE.&lt;/p&gt;

&lt;h2&gt;
  
  
  External Benchmark: How the Model Is Rated Elsewhere
&lt;/h2&gt;

&lt;p&gt;We checked Muse Glimmer's standing against other models using Artificial Analysis, a site that publishes third-party benchmark data. The chart below is organized around their Intelligence score.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe29833wk5ax4kcjijj5v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe29833wk5ax4kcjijj5v.png" alt="Artificial Analysis Intelligence score comparison, featuring Muse Glimmer 30B" width="659" height="183"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Figure 4: Artificial Analysis Intelligence scores (models shown are a custom selection for this comparison)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Muse Glimmer 30B scores 35, which is high relative to other open-weight models in the roughly-30B parameter range. Since this chart excludes coding- and agent-specialized models (a category we don't test as heavily), the comparison is limited to that scope, but within it, Muse Glimmer scores above both Qwen3.5 and Gemma-4.&lt;/p&gt;

&lt;p&gt;What stood out more was the score for Muse Spark, the closed model this architecture is presumably derived from — its score is close enough to Claude 5 Fable's that the gap is fairly small. Reaching this level of performance within a short update cycle, combined with reports that Meta may release this model line as open weights in the future, makes it a lineup worth continuing to watch. The Muse series took a long and difficult path to reach this point, and it will be worth seeing how the rest of the lineup performs going forward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;This article analyzed the architecture of Meta's Muse Glimmer 30B and evaluated its behavior through hands-on testing, including basic chat inference. Based on that evaluation, Muse Glimmer 30B looks like a strong option in the mid-range model category.&lt;/p&gt;

&lt;p&gt;Its architecture builds solidly on the "local + global" pattern that has proven effective across recent hybrid-attention models, while going a step further by omitting RoPE from the aggregation (FA) layers — a refinement that lets those layers focus purely on their consolidation role.&lt;/p&gt;

&lt;p&gt;The vision encoder is also a custom design based on Meta's own Perception Encoder architecture, and its analytical performance held up well in testing.&lt;/p&gt;

&lt;p&gt;The model is released under the Apache 2.0 license, which makes it a viable option for developers considering a Tokens-as-a-Service offering — a space where Gemma-4 has had little competition until now.&lt;/p&gt;

&lt;p&gt;Muse Spark 1.2, the closed model this architecture appears to derive from, also scores well on third-party benchmarks, suggesting Meta's more recent models have made real progress since the LLaMa-4 generation. Whether Meta continues to gain ground in the frontier-model space is something worth watching going forward.&lt;/p&gt;

&lt;p&gt;On a practical note, this test again highlighted the value of having a GPU with more VRAM on hand. GPU prices keep climbing, but without the ability to freely run mid-range models like this one locally, there's a lot that's hard to verify firsthand — a recurring frustration in this kind of investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Muse Glimmer 30B repository (Hugging Face)&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://huggingface.co/meta-models/Muse-Glimmer-30B" rel="noopener noreferrer"&gt;https://huggingface.co/meta-models/Muse-Glimmer-30B&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perception Encoder: The best visual embeddings are not at the output of the network&lt;/strong&gt;&lt;br&gt;
Daniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Choi, et al.&lt;br&gt;
&lt;a href="https://huggingface.co/papers/2504.13181" rel="noopener noreferrer"&gt;https://huggingface.co/papers/2504.13181&lt;/a&gt;&lt;br&gt;
&lt;a href="https://github.com/facebookresearch/perception_models" rel="noopener noreferrer"&gt;https://github.com/facebookresearch/perception_models&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Muse Glimmer model code source (GitHub)&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/huggingface/transformers/tree/main/src/transformers/models/muse_glimmer" rel="noopener noreferrer"&gt;https://github.com/huggingface/transformers/tree/main/src/transformers/models/muse_glimmer&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;llama.cpp PR26841&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/ggml-org/llama.cpp/pull/26841" rel="noopener noreferrer"&gt;https://github.com/ggml-org/llama.cpp/pull/26841&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Artificial Analysis&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://artificialanalysis.ai/" rel="noopener noreferrer"&gt;https://artificialanalysis.ai/&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article is an English adaptation of the original Japanese post published on Zenn: &lt;a href="https://zenn.dev/highreso/articles/ead02613d8c7ea" rel="noopener noreferrer"&gt;"Metaの意地を見せたか　Muse Glimmer 30Bの実力を見てみる"&lt;/a&gt;, by Yuichi Tominaga.&lt;/em&gt;&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;RoPE was first introduced in RoFormer, a Transformer-based language model similar to BERT. (&lt;a href="https://arxiv.org/pdf/2104.09864" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2104.09864&lt;/a&gt;)&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;GQA was first introduced by Google Research's T5-XXL, in the paper that proposed the GQA approach. (&lt;a href="https://arxiv.org/abs/2305.13245" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2305.13245&lt;/a&gt;)&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;llm-jp-4-8b-thinking's config.json lists its architecture as "LlamaForCausalLM," indicating it descends from the LLaMa-2 line, though it uses its own custom tokenizer. Its vocabulary is substantially larger, and its performance far exceeds the original LLaMa-2.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn4"&gt;
&lt;p&gt;In Qwen's case, the Gated Delta Network captures global information all at once, while Gated Attention acts as a global noise filter that smooths the information afterward — a structurally different approach from Muse Glimmer's. Looking across recent architectures, despite some structural variation (including how NVIDIA's Mamba series and LFM acquire local information), the "local + global" pattern used by Muse Glimmer and Gemma-4 appears to be the more dominant trend.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn5"&gt;
&lt;p&gt;For more detail on this setup, see our earlier article on using GPUSOROBAN. (&lt;a href="https://www.bluecore.net/archives/242" rel="noopener noreferrer"&gt;https://www.bluecore.net/archives/242&lt;/a&gt;)&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn6"&gt;
&lt;p&gt;This image uses an earlier version of the model diagram shown previously in this article.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>meta</category>
    </item>
    <item>
      <title>Trying "DFlash," a Diffusion-Model Approach to Parallel Draft-Token Generation, on Gemma</title>
      <dc:creator>oooocean66</dc:creator>
      <pubDate>Mon, 14 Sep 2026 07:02:56 +0000</pubDate>
      <link>https://dev.to/oooocean66/trying-dflash-a-diffusion-model-approach-to-parallel-draft-token-generation-on-gemma-41o8</link>
      <guid>https://dev.to/oooocean66/trying-dflash-a-diffusion-model-approach-to-parallel-draft-token-generation-on-gemma-41o8</guid>
      <description>&lt;p&gt;In the concept edition and the implementation/benchmark edition, we covered a speed-up technique for LLM generation called MTP (Multi-Token Prediction). To recap briefly: a lightweight "draft model" predicts a handful of tokens ahead of time, and the main model checks them all at once. When the guesses are right, you leap ahead several tokens in a single step, which is what makes the whole thing feel faster.&lt;/p&gt;

&lt;p&gt;There's more than one way to build that draft model, and the one we're looking at this time, "DFlash," takes an unusual approach: it uses a diffusion model — the kind of technique you'd normally associate with image generation — to predict multiple tokens all at once instead of one at a time. It claims to support a wide range of models and to significantly outperform EAGLE-3, an existing approach. Those are the claims worth testing directly.&lt;/p&gt;

&lt;p&gt;The catch is that benchmarks like this are usually measured in an environment the vendor sets up. It's harder to find a case where someone ran it on their own GPU, against an opponent that already has a dedicated, well-optimized MTP model of its own — Gemma-4's native Assistant model.&lt;/p&gt;

&lt;p&gt;So this time, using the same setup as the implementation/benchmark edition (an RTX 3060 with 12GB VRAM, llama.cpp, the same JavaScript coding task), we directly compared DFlash against Gemma-4-12B-it's Assistant model. The short version: DFlash did not outperform the Assistant model. The reasons are technically clear, though, and they draw a fairly clear picture of where DFlash is strong and where it isn't — that's what we'll dig into below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Overview
&lt;/h2&gt;

&lt;p&gt;Two new token-prediction techniques appeared in quick succession:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DeepSeek: a token-prediction technique called DSpark&lt;/li&gt;
&lt;li&gt;Z-Lab at UC San Diego: a token-prediction technique called DFlash&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;DSpark&lt;sup id="fnref1"&gt;1&lt;/sup&gt; was released to speed up inference specifically for DeepSeek's own DeepSeek-V4, and it only supports DeepSeek-V4 / DeepSeek-V4-Flash.&lt;/p&gt;

&lt;p&gt;DFlash, on the other hand, supports a much wider range of models, each with its own dedicated DFlash model. Like Google's Assistant model, it's designed to be bolted on as an add-on to achieve token prediction.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Gemma-4 series (12B, 26B-A4B, 31B)&lt;/li&gt;
&lt;li&gt;MiniMax-M2.7&lt;/li&gt;
&lt;li&gt;MiniMax-M2.5&lt;/li&gt;
&lt;li&gt;Qwen 3.6 series (35B-A3B, 27B)&lt;/li&gt;
&lt;li&gt;Qwen3.5 series (4B, 9B, 35B-A3B, 27B, 122B-A10B, 397B-A17B)&lt;/li&gt;
&lt;li&gt;Qwen3 series (4B, 8B, Coder-30B-A3B)&lt;/li&gt;
&lt;li&gt;Kimi-K2.6&lt;/li&gt;
&lt;li&gt;GLM-5.1-FP8&lt;/li&gt;
&lt;li&gt;gpt-oss (20B, 120B)&lt;/li&gt;
&lt;li&gt;LLaMa3.1-8B-Instruct&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Z-Lab is a research group led by Zhijian Liu&lt;sup id="fnref2"&gt;2&lt;/sup&gt;, an assistant professor at UC San Diego who is also a research scientist at NVIDIA. The lab works across the algorithm, systems, and application layers to make AI smaller, faster, and more efficient.&lt;/p&gt;

&lt;p&gt;This time, we wanted to understand how much of a speed-up DFlash actually delivers, verified through hands-on testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  How DFlash Predicts Tokens
&lt;/h2&gt;

&lt;p&gt;DFlash is introduced in the paper "DFlash: Block Diffusion for Flash Speculative Decoding"&lt;sup id="fnref3"&gt;3&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;Speculative token prediction itself isn't new — the earliest paper on the idea was published by Google DeepMind in 2023, "Accelerating Large Language Model Decoding with Speculative Sampling"&lt;sup id="fnref4"&gt;4&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;Improvements continued quietly from there, culminating in 2025 in an MTP draft model called EAGLE-3. Even so, it apparently never escaped the autoregressive paradigm, and in the end didn't deliver a dramatic speed improvement.&lt;/p&gt;

&lt;p&gt;Meanwhile, diffusion models — the noise-removal mechanism commonly used in image generation — have been making their way into the LLM space. It started with Meta's LLaDa, and in Japan, KDDI's ELYZA Lab team released a model called "ELYZA-Diffusion-Instruct-1.0-Dream-7B"&lt;sup id="fnref5"&gt;5&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;Z-Lab's DFlash brings that diffusion-model property into the MTP draft model. The basic mechanism is shown below.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Input Token (N)] --&amp;gt; [Main Model] -- h_on ----------------------------------------&amp;gt; [Predicted: N+1]
                                       |                                             ^
                                       v                                             |
                                  [KVCache]                                          |
                                       |                                             |
+--------------------------------------+--------------------------------+            |
| DFlash Drafter                       v (KV data injection)            |            |
|                                 [KVCache]                             |            |
|                                      |                                |            |
|                              [Diffusion Model]                        |            |
|                                      |                                |            |
|                               [Token Decoder]                         |            |
|                                      |                                |            |
|                            +---------v--------+                       |            |
|                            | Predicted N+2    |                       |            |
|                            |       +          |                       |            |
|                            | Predicted N+3    |                       |            |
|                            |       +          |                       |            |
|                            | Predicted N+4    |                       |            |
|                            +---------+--------+                       |            |
+--------------------------------------+--------------------------------+            |
                                       |                                             |
+--------------------------------------+---------------------------------------------+-------+
| Processing inside Main Model         v                                             |       |
|                             * Uses causal attention to mask;                       |       |
|                               computes probs in parallel                           |       |
|                                      +------------------------------+              |       |
|                                      v                              v              |       |
|                         (Match found)                  (No match)                  |       |
|                         Include matching portion       Nothing included in output  |       |
|                         -&amp;gt; Predicted: N+2                                          |       |
|                         -&amp;gt; Predicted: N+3                                          |       |
+------------------------------------------------------------------------------------+-------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Figure 1: DFlash mechanism&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Looking at the mechanism, it resembles the Assistant model implemented in Gemma, but the big difference is in how the draft model itself predicts tokens. The latter half — the token-evaluation stage — closely mirrors Gemma's approach.&lt;/p&gt;

&lt;p&gt;With DFlash, the KV cache is built independently. As before, when the main model predicts token N+1, it uses that state to update its own KV cache. Data extracted from that update is then injected into DFlash's own KV cache, which is where DFlash's processing begins.&lt;/p&gt;

&lt;p&gt;Gemma's Assistant model runs this prediction step using a very small neural network, sequentially and at high speed, producing as many predicted tokens as needed before handing them off to the evaluation logic.&lt;/p&gt;

&lt;p&gt;DFlash, by contrast, doesn't use an autoregressive model inside its small neural network — it uses a diffusion model. Here, much like generating an image, it produces all of the needed predicted tokens, in the correct order, in one shot. The longer the maximum token length, the longer this takes, but it's still dramatically faster than doing it sequentially. What follows is the same as before: causal-attention-based masking runs in parallel, and the result determines which tokens are allowed to be output together.&lt;/p&gt;

&lt;p&gt;So the biggest contributor to any speed advantage comes down to the parts marked with blue and red boxes in the middle of the figure below.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[For Gemma Assistant]
[Input Token (N)] --&amp;gt; [Main Model] --&amp;gt; [Predicted: N+1]
                            |
                            +-&amp;gt; [Predicted: N+2] -(seq)-&amp;gt; [Predicted: N+3] -(seq)-&amp;gt; [Predicted: N+4]
                                                                    |
                                                                    v
                                                     [Verify] --&amp;gt; Up to n OK --&amp;gt; [Confirmed: N+2, N+3...]

[For DFlash]
[Input Token (N)] --&amp;gt; [Main Model] --&amp;gt; [Predicted: N+1]
                            |
                            |   +-- [Predicted: N+2]
                            +---+-- [Predicted: N+3]   &amp;lt;-- (Outputs all at once, order included)
                            |   +-- [Predicted: N+4]
                            |               |
                            |               v
                            +---------&amp;gt;  [Verify] --&amp;gt; Up to n OK --&amp;gt; [Confirmed: N+2, N+3...]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Figure 2: Difference between Gemma's Assistant model and DFlash&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Running It in llama.cpp
&lt;/h2&gt;

&lt;p&gt;First, get the model files.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hf download google/gemma-4-12B-it
hf download z-lab/gemma4-12B-it-DFlash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This time we're using llama.cpp build 9850. It's a good idea to grab the latest build. Use the Python conversion tool included with it to convert each model to GGUF format. Start with the main model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;python convert_hf_to_gguf.py &lt;span class="se"&gt;\&lt;/span&gt;
~/.cache/huggingface/hub/models--google--gemma-4-12B-it/snapshots/cxxxxxx...a/ &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--outtype&lt;/span&gt; bf16 &lt;span class="nt"&gt;--outfile&lt;/span&gt; gemma-4-12B-it-b16.gguf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next, convert the DFlash model to GGUF.&lt;/p&gt;

&lt;p&gt;The important thing here is &lt;code&gt;--target-model-dir&lt;/code&gt;. When converting a DFlash model, it needs to be given a reference to the main model — that's the parameter it uses to build the converted DFlash output.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;python convert_hf_to_gguf.py &lt;span class="se"&gt;\&lt;/span&gt;
~/.cache/huggingface/hub/models--google--gemma-4-12B-it/snapshots/cxxxxxx...a/ &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--outtype&lt;/span&gt; bf16 &lt;span class="nt"&gt;--outfile&lt;/span&gt; models/gemma-4-12B-it-DFlash-b16.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--target-model-dir&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
~/.cache/huggingface/hub/gemma-4-12B-it-qat-q4_0-unquantized/snapshots/c202...a/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If this fails with the message below, download &lt;code&gt;tokenizer.model&lt;/code&gt; directly from the &lt;code&gt;google/gemma-3-12b-it&lt;/code&gt; repository on Hugging Face and place it in the cache directory.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;INFO:hf-to-gguf:DFlash: Using tokenizer from target model: /home/aiuser/.cache/huggingface/hub/models--google--gemma-4-12B-it/snapshots/5854...7ced3

Traceback (most recent call last):
  File "/data/user_data/aiuser/llama.cpp/conversion/qwen.py", line 57, in set_vocab
    self._set_vocab_sentencepiece()
  File "/data/user_data/aiuser/llama.cpp/conversion/base.py", line 1589, in _set_vocab_sentencepiece
    tokens, scores, toktypes = self._create_vocab_sentencepiece()
  File "/data/user_data/aiuser/llama.cpp/conversion/base.py", line 1606, in _create_vocab_sentencepiece
    raise FileNotFoundError(f"File not found: {tokenizer_path}")
FileNotFoundError: File not found: /home/aiuser/.cache/huggingface/hub/models--google--gemma-4-12B-it/snapshots/585...ced3/tokenizer.model
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Quantizing these produces 4-bit models.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;/opt/llama/bin/llama-quantize &lt;span class="se"&gt;\&lt;/span&gt;
models/gemma-4-12B-it-b16.gguf models/gemma-4-12B-it-q4_0.gguf q4_0

&lt;span class="nv"&gt;$ &lt;/span&gt;/opt/llama/bin/llama-quantize &lt;span class="se"&gt;\&lt;/span&gt;
models/gemma-4-12B-it-DFlash-b16.gguf models/gemma-4-12B-it-DFlash-q4_0.gguf q4_0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At runtime, add the following arguments. Since the DFlash model acts as an add-on draft/MTP model, use &lt;code&gt;--model-draft&lt;/code&gt; to point to it and enable MTP.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/opt/llama/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; /opt/llama/models/gemma-4-12B-it-q4_0.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--model-draft&lt;/span&gt; /opt/llama/models/gemma-4-12B-it-DFlash-q4_0.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-t&lt;/span&gt; 4 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 64 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--device&lt;/span&gt; CUDA0 &lt;span class="nt"&gt;-mg&lt;/span&gt; 0 &lt;span class="nt"&gt;-sm&lt;/span&gt; layer &lt;span class="nt"&gt;--fit&lt;/span&gt; on &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;-c&lt;/span&gt; 163840 &lt;span class="nt"&gt;-ctv&lt;/span&gt; q8_0 &lt;span class="nt"&gt;-ctk&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--spec-type&lt;/span&gt; draft-dflash &lt;span class="nt"&gt;--spec-draft-n-max&lt;/span&gt; 4 &lt;span class="nt"&gt;--spec-draft-device&lt;/span&gt; CUDA0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--chat-template-kwargs&lt;/span&gt; &lt;span class="s1"&gt;'{"enable_thinking":true}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A quick rundown of the arguments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--model-draft &amp;lt;gguf file&amp;gt;&lt;/code&gt;:&lt;/strong&gt; specifies the draft model&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;-sm layer&lt;/code&gt;:&lt;/strong&gt; how the model is dispatched across multiple CUDA devices

&lt;ul&gt;
&lt;li&gt;at the time of writing, only layer-wise distribution is supported&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--spec-type draft-dflash&lt;/code&gt;:&lt;/strong&gt; specifies the speculative-decoding type

&lt;ul&gt;
&lt;li&gt;declares the type as DFlash so it can be handled correctly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--spec-draft-n-max 4&lt;/code&gt;:&lt;/strong&gt; the maximum number of tokens to predict ahead

&lt;ul&gt;
&lt;li&gt;setting this too high increases the penalty on a miss&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--spec-draft-device CUDA0&lt;/code&gt;:&lt;/strong&gt; the device used for drafting

&lt;ul&gt;
&lt;li&gt;required when multiple CUDA devices are present&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Verification: Using DFlash on Gemma-4-12B-it
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Test Environment
&lt;/h3&gt;

&lt;p&gt;We ran the following simplified verification setup.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Machine: modified HP Z440 Workstation

&lt;ul&gt;
&lt;li&gt;CPU: Intel® Xeon® E5-2690 v4 x1&lt;/li&gt;
&lt;li&gt;RAM: 48GB DDR4 RDIMM&lt;/li&gt;
&lt;li&gt;SSD: Intel® DC S3700 Datacenter SSD (400GB SATA SSD)&lt;/li&gt;
&lt;li&gt;GPU: NVIDIA GeForce RTX 3060 (Ampere, 12GB GDDR6 VRAM)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;CUDA: 13.1&lt;/li&gt;
&lt;li&gt;Engine: llama.cpp Build 9574

&lt;ul&gt;
&lt;li&gt;Model: gemma-4-12B-it (&lt;a href="https://huggingface.co/google/gemma-4-12B-it" rel="noopener noreferrer"&gt;https://huggingface.co/google/gemma-4-12B-it&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;As the base model, we use Gemma-4-12b-it-QAT, Google DeepMind's quantization-aware-trained model.&lt;/li&gt;
&lt;li&gt;Drafter model: gemma-4-12B-it-DFlash (&lt;a href="https://huggingface.co/z-lab/gemma4-12B-it-DFlash" rel="noopener noreferrer"&gt;https://huggingface.co/z-lab/gemma4-12B-it-DFlash&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;We use Z-Lab's DFlash model for Gemma-12B.&lt;/li&gt;
&lt;li&gt;The maximum number of predicted tokens is set to 4 (matching the previous verification's settings, for comparability).&lt;/li&gt;
&lt;li&gt;KV Cache: 8-bit quantized (to fit efficiently in VRAM)&lt;/li&gt;
&lt;li&gt;Context size: 163,840 tokens (limited to 160k to fit efficiently in VRAM)&lt;/li&gt;
&lt;li&gt;Prompt cache: off&lt;/li&gt;
&lt;li&gt;Reasoning: on&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Test Cases
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Purpose: measure and compare speed over time when using the Assistant model versus DFlash

&lt;ul&gt;
&lt;li&gt;Only speed is measured; the content of the generated output is not considered.&lt;/li&gt;
&lt;li&gt;We record wall-clock time, token count, and token throughput at write-out.&lt;/li&gt;
&lt;li&gt;These figures are taken directly from llama.cpp's own log output.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Method: using the llama.cpp web frontend, we ask the following over four turns:&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Write JavaScript code for Breakout (as an HTML file).&lt;/li&gt;
&lt;li&gt;Make it look cooler.&lt;/li&gt;
&lt;li&gt;Slow down the ball's movement a bit.&lt;/li&gt;
&lt;li&gt;Check for bugs and optimize it.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Resource Usage
&lt;/h3&gt;

&lt;p&gt;Memory usage came out as follows.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Data category&lt;/th&gt;
&lt;th&gt;Assistant CUDA0&lt;/th&gt;
&lt;th&gt;Assistant CPU&lt;/th&gt;
&lt;th&gt;DFlash CUDA0&lt;/th&gt;
&lt;th&gt;DFlash CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weight data&lt;/td&gt;
&lt;td&gt;6,390.19&lt;/td&gt;
&lt;td&gt;540.00&lt;/td&gt;
&lt;td&gt;6,637.69&lt;/td&gt;
&lt;td&gt;787.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache data&lt;/td&gt;
&lt;td&gt;1,360.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1,360.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache for Sliding Window&lt;/td&gt;
&lt;td&gt;765.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;765.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;533.80&lt;/td&gt;
&lt;td&gt;180.80&lt;/td&gt;
&lt;td&gt;533.80&lt;/td&gt;
&lt;td&gt;180.80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP model weight data&lt;/td&gt;
&lt;td&gt;226.90&lt;/td&gt;
&lt;td&gt;144.00&lt;/td&gt;
&lt;td&gt;390.42&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP token-to-piece cache size&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;1.94&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP model KV cache size&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;640.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP model Sliding Window KV cache size&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;136.00&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP model Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;532.78&lt;/td&gt;
&lt;td&gt;180.79&lt;/td&gt;
&lt;td&gt;519.50&lt;/td&gt;
&lt;td&gt;176.02&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;9,808.67&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,045.59&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;10,984.35&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,226.38&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 1: Memory usage breakdown by MTP method (units: MiB). This run used text-only mode, so multimodal-model requirements are excluded.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Of that, the memory used specifically by the MTP model itself was as follows — roughly double for DFlash.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Assistant CUDA0&lt;/th&gt;
&lt;th&gt;Assistant CPU&lt;/th&gt;
&lt;th&gt;DFlash CUDA0&lt;/th&gt;
&lt;th&gt;DFlash CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MTP model memory usage&lt;/td&gt;
&lt;td&gt;759.68&lt;/td&gt;
&lt;td&gt;324.79&lt;/td&gt;
&lt;td&gt;1,687.86&lt;/td&gt;
&lt;td&gt;176.02&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 2: Total memory used by the MTP model, by MTP method (units: MiB)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The main driver here is how the two approaches use the KV cache. The Assistant model shares its KV cache with the main Gemma model. DFlash, by contrast, keeps its own independent KV cache — meaning it has to hold the same cache structure as the main Gemma model a second time. That's the biggest factor behind the difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token Speed Comparison (Assistant vs DFlash)
&lt;/h2&gt;

&lt;p&gt;As in the previous verification, we compared token ingestion speed turn by turn. There was no major throughput difference between the Assistant model and DFlash for reading tokens in.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Assistant (tok/s)&lt;/th&gt;
&lt;th&gt;DFlash (tok/s)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;138.34&lt;/td&gt;
&lt;td&gt;53.72&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;710.71&lt;/td&gt;
&lt;td&gt;652.47&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;733.53&lt;/td&gt;
&lt;td&gt;739.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;735.43&lt;/td&gt;
&lt;td&gt;708.22&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 3: Token read-in speed by turn&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For output speed per turn, the Assistant model was clearly ahead — DFlash consistently trailed by about 10–20 tokens per second.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Assistant (tok/s)&lt;/th&gt;
&lt;th&gt;DFlash (tok/s)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;62.80&lt;/td&gt;
&lt;td&gt;39.37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;62.04&lt;/td&gt;
&lt;td&gt;48.42&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;66.68&lt;/td&gt;
&lt;td&gt;52.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;62.77&lt;/td&gt;
&lt;td&gt;48.73&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Figure 4: Token output speed by turn&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here is the raw measurement data behind the numbers above. Note that time is in milliseconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reading&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Assistant time (ms)&lt;/th&gt;
&lt;th&gt;Assistant tokens&lt;/th&gt;
&lt;th&gt;Assistant tps&lt;/th&gt;
&lt;th&gt;DFlash time (ms)&lt;/th&gt;
&lt;th&gt;DFlash tokens&lt;/th&gt;
&lt;th&gt;DFlash tps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;181&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;138.34&lt;/td&gt;
&lt;td&gt;577&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;53.72&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3,004&lt;/td&gt;
&lt;td&gt;2,135&lt;/td&gt;
&lt;td&gt;710.71&lt;/td&gt;
&lt;td&gt;3,374&lt;/td&gt;
&lt;td&gt;2,202&lt;/td&gt;
&lt;td&gt;652.47&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;6,361&lt;/td&gt;
&lt;td&gt;4,666&lt;/td&gt;
&lt;td&gt;733.53&lt;/td&gt;
&lt;td&gt;6,391&lt;/td&gt;
&lt;td&gt;4,727&lt;/td&gt;
&lt;td&gt;739.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;9,559&lt;/td&gt;
&lt;td&gt;7,030&lt;/td&gt;
&lt;td&gt;735.43&lt;/td&gt;
&lt;td&gt;10,173&lt;/td&gt;
&lt;td&gt;7,205&lt;/td&gt;
&lt;td&gt;708.22&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Generation&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Assistant time (ms)&lt;/th&gt;
&lt;th&gt;Assistant tokens&lt;/th&gt;
&lt;th&gt;Assistant tps&lt;/th&gt;
&lt;th&gt;DFlash time (ms)&lt;/th&gt;
&lt;th&gt;DFlash tokens&lt;/th&gt;
&lt;th&gt;DFlash tps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;33,186&lt;/td&gt;
&lt;td&gt;2,084&lt;/td&gt;
&lt;td&gt;62.80&lt;/td&gt;
&lt;td&gt;54,413&lt;/td&gt;
&lt;td&gt;2,142&lt;/td&gt;
&lt;td&gt;39.37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;40,281&lt;/td&gt;
&lt;td&gt;2,499&lt;/td&gt;
&lt;td&gt;62.04&lt;/td&gt;
&lt;td&gt;51,096&lt;/td&gt;
&lt;td&gt;2,474&lt;/td&gt;
&lt;td&gt;48.42&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;35,046&lt;/td&gt;
&lt;td&gt;2,337&lt;/td&gt;
&lt;td&gt;66.68&lt;/td&gt;
&lt;td&gt;46,701&lt;/td&gt;
&lt;td&gt;2,447&lt;/td&gt;
&lt;td&gt;52.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;45,532&lt;/td&gt;
&lt;td&gt;2,858&lt;/td&gt;
&lt;td&gt;62.77&lt;/td&gt;
&lt;td&gt;55,774&lt;/td&gt;
&lt;td&gt;2,718&lt;/td&gt;
&lt;td&gt;48.73&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 3: Raw measurements from the logs&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Overall, inference with DFlash trailed the Assistant model by under 20 seconds. The Assistant model was already a fairly well-optimized setup going in, which may be part of why the diffusion model's theoretical advantage didn't stand out here.&lt;/p&gt;

&lt;p&gt;As before, output speed increased somewhat over the course of each turn for both the Assistant model and DFlash, but the Assistant model was faster throughout. (We're omitting the detailed per-token scatter plots for each model here, since the turn-by-turn averages above and the raw log data in Table 3 already capture the trend.) In a few cases we also observed DFlash's output speed peaking around 1,700 tokens and then dropping off slightly after that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Checking the Token Acceptance Rate
&lt;/h3&gt;

&lt;p&gt;As in the previous verification, let's also look at the token acceptance rate.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Acceptance&lt;/th&gt;
&lt;th&gt;Accepted&lt;/th&gt;
&lt;th&gt;Generated&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;41.03%&lt;/td&gt;
&lt;td&gt;1,331&lt;/td&gt;
&lt;td&gt;3,244&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;53.36%&lt;/td&gt;
&lt;td&gt;1,684&lt;/td&gt;
&lt;td&gt;3,156&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;62.02%&lt;/td&gt;
&lt;td&gt;1,744&lt;/td&gt;
&lt;td&gt;2,812&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;57.26%&lt;/td&gt;
&lt;td&gt;1,892&lt;/td&gt;
&lt;td&gt;3,304&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 4: DFlash token acceptance rate&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Comparing this against the Assistant model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Assistant Acceptance&lt;/th&gt;
&lt;th&gt;Assistant Accepted&lt;/th&gt;
&lt;th&gt;Assistant Generated&lt;/th&gt;
&lt;th&gt;DFlash Acceptance&lt;/th&gt;
&lt;th&gt;DFlash Accepted&lt;/th&gt;
&lt;th&gt;DFlash Generated&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;88.27%&lt;/td&gt;
&lt;td&gt;2,196&lt;/td&gt;
&lt;td&gt;2,648&lt;/td&gt;
&lt;td&gt;62.02%&lt;/td&gt;
&lt;td&gt;1,892&lt;/td&gt;
&lt;td&gt;3,304&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Min&lt;/td&gt;
&lt;td&gt;75.34%&lt;/td&gt;
&lt;td&gt;1,564&lt;/td&gt;
&lt;td&gt;2,064&lt;/td&gt;
&lt;td&gt;41.03%&lt;/td&gt;
&lt;td&gt;1,331&lt;/td&gt;
&lt;td&gt;2,812&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg&lt;/td&gt;
&lt;td&gt;81.06%&lt;/td&gt;
&lt;td&gt;1,868&lt;/td&gt;
&lt;td&gt;2,305&lt;/td&gt;
&lt;td&gt;53.42%&lt;/td&gt;
&lt;td&gt;1,663&lt;/td&gt;
&lt;td&gt;3,129&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Table 5: Acceptance rate comparison — Assistant (left) vs. DFlash (right)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This suggests that DFlash's underwhelming result comes down to its lower acceptance rate. If model tuning progresses further and a faster-processing diffusion model becomes possible, DFlash might eventually surpass the Assistant model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trying to Improve Things
&lt;/h2&gt;

&lt;p&gt;Through this verification, we confirmed that DFlash falls a bit short of the Assistant model. But maybe tweaking the parameters would speed things up? With that in mind, we tried the following.&lt;/p&gt;

&lt;h3&gt;
  
  
  Raising the Number of Predicted Tokens
&lt;/h3&gt;

&lt;p&gt;The predicted-token count is currently set to 4. That's the same value DFlash used in the previous Assistant-model verification, kept for comparability. Raising this significantly might help — so we bumped &lt;code&gt;--spec-draft-n-max&lt;/code&gt; from 4 to 15.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/opt/llama/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; /opt/llama/models/gemma-4-12B-it-q4_0.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--model-draft&lt;/span&gt; /opt/llama/models/gemma-4-12B-it-DFlash-q4_0.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-t&lt;/span&gt; 4 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 64 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--device&lt;/span&gt; CUDA0 &lt;span class="nt"&gt;-mg&lt;/span&gt; 0 &lt;span class="nt"&gt;-sm&lt;/span&gt; layer &lt;span class="nt"&gt;--fit&lt;/span&gt; on &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;-c&lt;/span&gt; 163840 &lt;span class="nt"&gt;-ctv&lt;/span&gt; q8_0 &lt;span class="nt"&gt;-ctk&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--spec-type&lt;/span&gt; draft-dflash &lt;span class="nt"&gt;--spec-draft-n-max&lt;/span&gt; 15 &lt;span class="nt"&gt;--spec-draft-device&lt;/span&gt; CUDA0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--chat-template-kwargs&lt;/span&gt; &lt;span class="s1"&gt;'{"enable_thinking":true}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's what came out of running it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;prompt eval time = 122.83 ms / 31 tokens ( 3.96 ms per token, 252.39 tokens per second)
eval time = 48833.63 ms / 1893 tokens ( 25.80 ms per token, 38.76 tokens per second)
total time = 48956.46 ms / 1924 tokens
graphs reused = 1444
draft acceptance = 0.10901 ( 1174 accepted / 10770 generated), mean len = 2.64
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The complete opposite of what we hoped for: acceptance dropped, and speed dropped along with it.&lt;/p&gt;

&lt;p&gt;It generated a lot more candidate tokens, but almost all of them were rejected — simply lengthening the prediction window clearly wasn't the answer.&lt;/p&gt;

&lt;p&gt;Here's the breakdown by turn:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Acceptance&lt;/th&gt;
&lt;th&gt;Accepted&lt;/th&gt;
&lt;th&gt;Generated&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;10.90%&lt;/td&gt;
&lt;td&gt;1,174&lt;/td&gt;
&lt;td&gt;10,770&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;13.59%&lt;/td&gt;
&lt;td&gt;2,059&lt;/td&gt;
&lt;td&gt;15,150&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;19.54%&lt;/td&gt;
&lt;td&gt;2,090&lt;/td&gt;
&lt;td&gt;10,695&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;19.30%&lt;/td&gt;
&lt;td&gt;2,310&lt;/td&gt;
&lt;td&gt;11,970&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Acceptance&lt;/th&gt;
&lt;th&gt;Accepted Tokens&lt;/th&gt;
&lt;th&gt;Generated Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;19.54%&lt;/td&gt;
&lt;td&gt;2,310&lt;/td&gt;
&lt;td&gt;15,150&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Min&lt;/td&gt;
&lt;td&gt;10.90%&lt;/td&gt;
&lt;td&gt;1,174&lt;/td&gt;
&lt;td&gt;10,695&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg&lt;/td&gt;
&lt;td&gt;15.83%&lt;/td&gt;
&lt;td&gt;1,908&lt;/td&gt;
&lt;td&gt;12,146&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;mean_len=2.64&lt;/code&gt; value in the log above appears to represent roughly "how many tokens tend to get accepted." Looking at the per-turn breakdown above, the average was 15.83%&lt;sup id="fnref6"&gt;6&lt;/sup&gt;, suggesting that &lt;code&gt;--spec-draft-n-max=4&lt;/code&gt; was in fact the right setting for this model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rebuilding on Gemma-4-12B-it-QAT
&lt;/h3&gt;

&lt;p&gt;What if we raised the model's own output precision instead? We tried rebuilding DFlash on that basis.&lt;/p&gt;

&lt;p&gt;For this run, since we'd need to manually apply the same quantization level, we couldn't use an Unsloth Dynamic 2.0–quantized model the way we did for the Assistant model. Instead, we rebuilt using Gemma-4-12B-it-QAT (including the DFlash model), hoping it might improve performance.&lt;/p&gt;

&lt;p&gt;The results, unfortunately, were not what we expected. Output was so slow that we stopped measuring after the second turn.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;prompt eval time = 475.69 ms / 31 tokens ( 15.34 ms per token, 65.17 tokens per second)
eval time = 102682.33 ms / 1963 tokens (52.31 ms per token, 19.12 tokens per second)
total time = 103158.02 ms / 1994 tokens
graphs reused = 1953
draft acceptance = 0.00025 ( 2 accepted / 7844 generated), mean len = 1.00
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This configuration is not viable: throughput dropped further, and the acceptance rate fell to 0.00025, or 0.25% — the lowest we observed in this test. With &lt;code&gt;mean_len=1.00&lt;/code&gt;, essentially every generated token was rejected, a clear indication that DFlash should not be used in this configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Couldn't It Beat Gemma-4-Assistant?
&lt;/h3&gt;

&lt;p&gt;It looks like the DFlash model Z-Lab provided is really only usable with the vanilla Gemma-4-12B-it model. That's unfortunate — if that's the case, the Assistant model, which comes with a properly matched MTP model for each release, seems like the more practical choice.&lt;/p&gt;

&lt;p&gt;The likely cause here is that Gemma-4-Assistant's draft model is simply fast enough that it processed tokens more quickly than DFlash's draft model.&lt;/p&gt;

&lt;p&gt;As covered earlier, the speed advantage a diffusion-based drafter like DFlash offers over a conventional MTP draft model comes from being able to output candidate tokens "all at once." Against that, Gemma-Assistant has the following working strongly in its favor, which appears to have canceled out DFlash's advantage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A shared KV-cache mechanism&lt;/li&gt;
&lt;li&gt;A very small attention mechanism&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;First, Gemma-Assistant comes with KV-cache sharing built in from the start, while DFlash keeps an independent KV cache that has to be injected fresh every time. That's an advantage for Gemma-Assistant.&lt;/p&gt;

&lt;p&gt;On top of that, Gemma-Assistant's hidden dimension is only 1,024, with a 4-layer structure. DFlash's hidden dimension, as noted at the end of this article, matches the main model at 3,840, with 5 layers — meaning its compute cost is far heavier than Gemma-Assistant's.&lt;/p&gt;

&lt;p&gt;However good DFlash's diffusion model is at generating tokens in a single batch, more compute per step still means more time spent per step.&lt;/p&gt;

&lt;p&gt;In this case, we should conclude that Gemma-Assistant simply already had a more optimized setup, and DFlash wasn't able to get ahead of it.&lt;/p&gt;

&lt;p&gt;It's also worth remembering that DFlash's original benchmark comparison was against EAGLE-3, not Gemma-Assistant — so it's possible Gemma-Assistant is simply a stronger comparison baseline than the one DFlash was originally built to outperform.&lt;/p&gt;

&lt;p&gt;That said, while the established option produced better results this time, it's possible a pairing with, say, Qwen3.5 would have produced a more favorable result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;This time we introduced DFlash, a technique from the UC San Diego (UCSD) research team Z-Lab that takes the "MTP" technology covered previously and pushes it in a new direction. Applying a diffusion model to a draft model was a genuinely ambitious idea, but in this test it wasn't able to beat Gemma's native draft model, Assistant.&lt;/p&gt;

&lt;p&gt;Tracing the cause, the token acceptance rate was lower than the native draft model's, and that penalty appears to have been the dominant factor. Without a way to understand how to raise that acceptance rate, it's hard to say at this point whether further speed gains are realistic.&lt;/p&gt;

&lt;p&gt;We also think the matchup itself didn't help. Gemma-4's Assistant model is built with practicality in mind and makes full use of Gemma-4's shared KV-cache mechanism, whereas DFlash, as before, has to write to an independent KV cache — and that overhead seems to have tipped the balance toward lower throughput.&lt;/p&gt;

&lt;p&gt;On the broader question of putting diffusion models to work in LLMs, LLaDa and Dream are the well-known names, but more recently Google itself has released a model called Diffusion Gemma. We're currently in the middle of testing it ourselves, and the throughput numbers so far are notably high. We plan to cover those results in a dedicated article.&lt;/p&gt;

&lt;p&gt;As Coding Agent usage keeps climbing, LLM throughput has become a common pain point, and rising VRAM costs have only sharpened the demand for models that can deliver solid throughput even on unified memory.&lt;/p&gt;

&lt;p&gt;This kind of research is very much a "fail your way forward" field, and new techniques are emerging as we speak. We'll keep watching this space and try to keep up.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;DFlash: Block Diffusion for Flash Speculative Decoding&lt;/strong&gt;&lt;br&gt;
Jian Chen, Yesheng Liang, Zhijian Liu&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2602.06036" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2602.06036&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accelerating Large Language Model Decoding with Speculative Sampling&lt;/strong&gt;&lt;br&gt;
DeepMind: Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre and John Jumper&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2302.01318" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2302.01318&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence&lt;/strong&gt;&lt;br&gt;
DeepSeek-AI&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2606.19348" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2606.19348&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Speculative decoding] feat: add DFlash support&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/ggml-org/llama.cpp/pull/22105" rel="noopener noreferrer"&gt;https://github.com/ggml-org/llama.cpp/pull/22105&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Supplementary Information
&lt;/h2&gt;

&lt;h3&gt;
  
  
  DFlash Model Structure and Connection to the Main Model
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;(We've omitted the architecture diagram here to keep the English edition concise — the description below covers the key points.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The DFlash model, which adds MTP capability to Gemma-4-12B-it, is architecturally distinct from Gemma-4-Assistant: rather than sharing a KV cache the way Assistant does, it's a fully separate model. In terms of its activation function and related choices, it's actually closer to a Qwen-style model. Its LA/GA notation follows Gemma-4's own convention — LA meaning Sliding Window Attention and GA meaning Full Attention.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Maximum context size: 256k tokens&lt;/li&gt;
&lt;li&gt;5 blocks total (one group of 4+1)&lt;/li&gt;
&lt;li&gt;Hidden size: &lt;strong&gt;3,840&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;FFN type: &lt;strong&gt;Dense&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Activation function: &lt;strong&gt;SiLU&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Activation dimension: &lt;strong&gt;7,680&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;During prediction, state vectors are pulled from layers 1, 10, 19, 27, 36, and 45, concatenated, and normalized. That combined information is then injected into the draft side's KV cache.&lt;/p&gt;

&lt;p&gt;The input to the draft model is a query built by combining the embedding vectors of the N confirmed tokens with as many MASK tokens as the number of predictions needed.&lt;/p&gt;

&lt;p&gt;Unlike the Assistant model, this approach uses non-causal attention, and internally relies on Masked Diffusion — that's the key structural difference.&lt;/p&gt;

&lt;p&gt;This mechanism lets DFlash derive all of its predicted tokens at once; from there, the rest of the pipeline follows the same logic as the Assistant model, outputting whichever tokens get accepted.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Is EAGLE3?
&lt;/h3&gt;

&lt;p&gt;EAGLE3 is version 3 of the EAGLE (Extrapolation Algorithm for Greater Language Model Efficiency) series, documented in the following papers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;EAGLE ( &lt;a href="https://arxiv.org/pdf/2401.15077.pdf" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2401.15077.pdf&lt;/a&gt; )&lt;/li&gt;
&lt;li&gt;EAGLE2 ( &lt;a href="https://arxiv.org/pdf/2406.16858" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2406.16858&lt;/a&gt; )&lt;/li&gt;
&lt;li&gt;EAGLE3 ( &lt;a href="https://arxiv.org/pdf/2503.01840" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2503.01840&lt;/a&gt; )&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Development has been led by Yuhui Li of Peking University, first author on these papers, together with a joint research team spanning Peking University, Microsoft Research, the University of Waterloo, and the Vector Institute. In the open-source community, they operate under the name SafeAILab.&lt;/p&gt;

&lt;p&gt;Successor models are also in development: SafeAILab released EAGLE3.1 in May 2026, and a team at AWS AI Labs has developed P-EAGLE (Parallel-Drafting EAGLE), released in February 2026.&lt;/p&gt;

&lt;p&gt;Both treat the sequential bottleneck of causal attention as the core problem, and both are focused on how to parallelize that part of the pipeline further.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article is an English adaptation of the original Japanese post published on Zenn: &lt;a href="https://zenn.dev/highreso/articles/ab3dbc20ce58ff" rel="noopener noreferrer"&gt;"ドラフトトークンを並列生成する拡散モデル方式「DFlash」をGemmaで試した"&lt;/a&gt;, by Yuichi Tominaga.&lt;/em&gt;&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2606.19348" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2606.19348&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;&lt;a href="https://zhijianliu.com/" rel="noopener noreferrer"&gt;https://zhijianliu.com/&lt;/a&gt;, &lt;a href="https://z-lab.ai/" rel="noopener noreferrer"&gt;https://z-lab.ai/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2602.06036" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2602.06036&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn4"&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/pdf/2302.01318" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2302.01318&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn5"&gt;
&lt;p&gt;&lt;a href="https://huggingface.co/elyza/ELYZA-Diffusion-Instruct-1.0-Dream-7B" rel="noopener noreferrer"&gt;https://huggingface.co/elyza/ELYZA-Diffusion-Instruct-1.0-Dream-7B&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn6"&gt;
&lt;p&gt;Average acceptance rate across the four turns shown in the table below.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>MTP in Practice: Benchmarking Gemma's Speculative Decoding on a Real GPU</title>
      <dc:creator>oooocean66</dc:creator>
      <pubDate>Fri, 11 Sep 2026 10:06:06 +0000</pubDate>
      <link>https://dev.to/oooocean66/mtp-in-practice-benchmarking-gemmas-speculative-decoding-on-a-real-gpu-242</link>
      <guid>https://dev.to/oooocean66/mtp-in-practice-benchmarking-gemmas-speculative-decoding-on-a-real-gpu-242</guid>
      <description>&lt;p&gt;In the concept edition, we saw that MTP (Multi-Token Prediction) lets a model predict several tokens ahead to speed up generation, and that Qwen and Gemma implement this in completely different ways.&lt;/p&gt;

&lt;p&gt;The theory makes sense, but how much faster does this actually make things in practice? That's what we set out to measure directly.&lt;/p&gt;

&lt;p&gt;In this installment — the implementation/benchmark edition — we'll actually run Gemma's MTP on llama.cpp and measure, with real numbers on an ordinary consumer PC, just how much of an effect it has.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running It in llama.cpp
&lt;/h2&gt;

&lt;h3&gt;
  
  
  For Qwen Models
&lt;/h3&gt;

&lt;p&gt;Run with the following arguments added. Specifying &lt;code&gt;spec-type&lt;/code&gt; activates the model's internal MTP drafter. A dedicated model is provided for MTP.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/opt/llama/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; /opt/llama/models/Qwen3.5-9B-MTP-UD-Q3_K_XL.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-t&lt;/span&gt; 4 &lt;span class="nt"&gt;-np&lt;/span&gt; 1 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 20 &lt;span class="nt"&gt;--min-p&lt;/span&gt; 0.00 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--device&lt;/span&gt; CUDA0 &lt;span class="nt"&gt;-mg&lt;/span&gt; 0 &lt;span class="nt"&gt;-ctv&lt;/span&gt; q8_0 &lt;span class="nt"&gt;-ctk&lt;/span&gt; q8_0 &lt;span class="nt"&gt;--fit&lt;/span&gt; off &lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; 32768 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--spec-type&lt;/span&gt; draft-mtp &lt;span class="nt"&gt;--spec-draft-n-max&lt;/span&gt; 4 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--reasoning&lt;/span&gt; on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Argument summary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;--spec-type draft-mtp:&lt;/strong&gt; specifies the speculative decoding type (&lt;code&gt;draft-mtp&lt;/code&gt; for Qwen).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;--spec-draft-n-max 4:&lt;/strong&gt; how many tokens ahead to speculate at most.

&lt;ul&gt;
&lt;li&gt;Setting this too high increases the penalty when a prediction fails.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Also note: concurrent processing isn't supported, and multimodal input isn't supported either.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  For Gemma Models
&lt;/h3&gt;

&lt;p&gt;Run with the following arguments. For Gemma, the drafter-MTP model is bolted on externally, so you enable MTP by pointing the &lt;code&gt;--model-draft&lt;/code&gt; argument at that file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/opt/llama/bin/llama-server &lt;span class="nt"&gt;--model&lt;/span&gt; /opt/llama/models/gemma-4-12B-it-qat-UD-Q4_K_XL.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--model-draft&lt;/span&gt; /opt/llama/models/mtp-gemma-4-12B-it.gguf &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;-t&lt;/span&gt; 4 &lt;span class="nt"&gt;--prio&lt;/span&gt; 2 &lt;span class="nt"&gt;--temp&lt;/span&gt; 1.0 &lt;span class="nt"&gt;--top-p&lt;/span&gt; 0.95 &lt;span class="nt"&gt;--top-k&lt;/span&gt; 64 &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--port&lt;/span&gt; 8001 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--device&lt;/span&gt; CUDA0 &lt;span class="nt"&gt;-mg&lt;/span&gt; 0 &lt;span class="nt"&gt;-sm&lt;/span&gt; layer &lt;span class="nt"&gt;--fit&lt;/span&gt; on &lt;span class="nt"&gt;-fa&lt;/span&gt; on &lt;span class="nt"&gt;-c&lt;/span&gt; 163840 &lt;span class="nt"&gt;-ctv&lt;/span&gt; q8_0 &lt;span class="nt"&gt;-ctk&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--no-warmup&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-prompt&lt;/span&gt; &lt;span class="nt"&gt;--cache-ram&lt;/span&gt; 0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--spec-type&lt;/span&gt; draft-mtp &lt;span class="nt"&gt;--spec-draft-n-max&lt;/span&gt; 4 &lt;span class="nt"&gt;--spec-draft-device&lt;/span&gt; CUDA0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nt"&gt;--chat-template-kwargs&lt;/span&gt; &lt;span class="s1"&gt;'{"enable_thinking":true}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Argument summary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;--model-draft &amp;lt;gguf file&amp;gt;:&lt;/strong&gt; specifies the draft model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;-sm layer:&lt;/strong&gt; how the model is dispatched across multiple CUDA devices.

&lt;ul&gt;
&lt;li&gt;At the time of writing, only layer-level distribution is supported.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;--spec-type draft-mtp:&lt;/strong&gt; specifies the speculative decoding type.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;--spec-draft-n-max 4:&lt;/strong&gt; how many tokens ahead to speculate at most.

&lt;ul&gt;
&lt;li&gt;Setting this too high increases the penalty when a prediction fails.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;--spec-draft-device CUDA0:&lt;/strong&gt; the device that runs drafting.

&lt;ul&gt;
&lt;li&gt;Required when multiple CUDA devices are present.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Benchmark: Using MTP on Gemma-4-12B-it
&lt;/h2&gt;

&lt;p&gt;In this section, to investigate how effectively MTP actually works and what kind of benefit it brings to users, we compared performance with and without MTP.&lt;/p&gt;

&lt;p&gt;To also demonstrate that this kind of model verification is achievable even on ordinary consumer hardware, the benchmark was run on the author's own personal machine.&lt;/p&gt;

&lt;h3&gt;
  
  
  Test Environment
&lt;/h3&gt;

&lt;p&gt;The experiment used the following simple test setup.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Machine: modified HP Z440 Workstation

&lt;ul&gt;
&lt;li&gt;CPU: Intel® Xeon® E5-2690 v4 x1&lt;/li&gt;
&lt;li&gt;RAM: 48GB DDR4 RDIMM&lt;/li&gt;
&lt;li&gt;SSD: Intel® DC S3700 Datacenter SSD (400GB SATA SSD)&lt;/li&gt;
&lt;li&gt;GPU: NVIDIA GeForce RTX 3060 (Ampere, 12GB GDDR6 VRAM)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;CUDA: 13.1&lt;/li&gt;
&lt;li&gt;Engine: llama.cpp Build 9574

&lt;ul&gt;
&lt;li&gt;Model: gemma-4-12B-it-qat-UD-Q4_K_XL.gguf
&lt;a href="https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF" rel="noopener noreferrer"&gt;https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;As the base model, we use Gemma-4-12b-it-QAT, a Google DeepMind model trained with quantization-aware training in mind.&lt;/li&gt;
&lt;li&gt;On top of that, we use the UD-Q4_K_XL model&lt;sup id="fnref1"&gt;1&lt;/sup&gt;, dynamically quantized by Unsloth&lt;sup id="fnref2"&gt;2&lt;/sup&gt;, which provides open-source fine-tuning tools.&lt;/li&gt;
&lt;li&gt;Drafter model: mtp-gemma-4-12B-it.gguf
Downloaded from the same repository as the base model.&lt;/li&gt;
&lt;li&gt;We use the Assistant model prepared by Unsloth for Gemma-12B, simply quantized to 4-bit, as the draft model.&lt;/li&gt;
&lt;li&gt;The maximum number of speculated tokens is set to 4.&lt;/li&gt;
&lt;li&gt;KV Cache: 8-bit quantized&lt;/li&gt;
&lt;li&gt;The KV cache is quantized to 8 bits to fit efficiently into VRAM.&lt;/li&gt;
&lt;li&gt;Context size: 163,840 tokens&lt;/li&gt;
&lt;li&gt;Context size is capped at 160k to stay within VRAM capacity.&lt;/li&gt;
&lt;li&gt;Prompt cache: off&lt;/li&gt;
&lt;li&gt;Reasoning: on&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Test Cases
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Goal: measure and compare speed with and without MTP.

&lt;ul&gt;
&lt;li&gt;Only speed is measured. The actual content of the generated output is not evaluated at all.&lt;/li&gt;
&lt;li&gt;We measure elapsed time, token count, and token speed during generation.&lt;/li&gt;
&lt;li&gt;These values are taken directly from llama.cpp's own log output.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Method: using the llama.cpp web frontend, we asked the following 4 turns of questions.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;"Write JavaScript code for Breakout (as an HTML page)."&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;"Make it look cooler."&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;"Slow the ball's movement down a bit."&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;"Check for bugs and optimize it."&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Resource Usage
&lt;/h3&gt;

&lt;p&gt;Memory usage was as follows. The memory consumed specifically for MTP came to 759.68 MiB of VRAM on the GPU side and 324.79 MiB of RAM on the CPU side.&lt;/p&gt;

&lt;p&gt;Since we used a 4-bit quantized model for the draft model here, choosing a 16-bit model instead would consume roughly 4x this amount for the weight data. Anyone planning to use a 16-bit model should keep this in mind.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;CUDA0&lt;/th&gt;
&lt;th&gt;CPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Weight data&lt;/td&gt;
&lt;td&gt;6,390.19&lt;/td&gt;
&lt;td&gt;540.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV cache data&lt;/td&gt;
&lt;td&gt;1,360.00&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sliding window KV cache&lt;/td&gt;
&lt;td&gt;765.00&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;533.80&lt;/td&gt;
&lt;td&gt;180.80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP (Assistant) model weight data&lt;/td&gt;
&lt;td&gt;226.90&lt;/td&gt;
&lt;td&gt;144.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MTP (Assistant) model Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;532.78&lt;/td&gt;
&lt;td&gt;180.79&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision Encoder&lt;/td&gt;
&lt;td&gt;167.00&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audio Encoder&lt;/td&gt;
&lt;td&gt;167.00&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multimodal Gated DeltaNet compute buffer&lt;/td&gt;
&lt;td&gt;532.78&lt;/td&gt;
&lt;td&gt;180.79&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;10,675.45&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,226.38&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Units: MiB. Table 1: Memory usage in this test environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token Speed Comparison Results (MTP On vs. Off)
&lt;/h2&gt;

&lt;p&gt;Comparing token ingestion speed turn by turn, we saw no major variation. With MTP enabled, ingestion speed dropped slightly compared to without MTP, coming in at roughly 70–73% of the baseline performance (see the "Read" rows in the table below).&lt;/p&gt;

&lt;p&gt;Token generation speed, on the other hand, was roughly twice as fast overall with MTP enabled (see the "Generate" rows).&lt;/p&gt;

&lt;p&gt;Below are the actual measurements. Note that time is measured in ms.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;No MTP: time&lt;/th&gt;
&lt;th&gt;No MTP: tokens&lt;/th&gt;
&lt;th&gt;No MTP: tps&lt;/th&gt;
&lt;th&gt;With MTP: time&lt;/th&gt;
&lt;th&gt;With MTP: tokens&lt;/th&gt;
&lt;th&gt;With MTP: tps&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;td&gt;301&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;83.04&lt;/td&gt;
&lt;td&gt;181&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;138.34&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;td&gt;1,836&lt;/td&gt;
&lt;td&gt;1,834&lt;/td&gt;
&lt;td&gt;999.10&lt;/td&gt;
&lt;td&gt;3,004&lt;/td&gt;
&lt;td&gt;2,135&lt;/td&gt;
&lt;td&gt;710.71&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;td&gt;4,546&lt;/td&gt;
&lt;td&gt;4,669&lt;/td&gt;
&lt;td&gt;1,027.06&lt;/td&gt;
&lt;td&gt;6,361&lt;/td&gt;
&lt;td&gt;4,666&lt;/td&gt;
&lt;td&gt;733.53&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;td&gt;7,017&lt;/td&gt;
&lt;td&gt;7,102&lt;/td&gt;
&lt;td&gt;1,012.11&lt;/td&gt;
&lt;td&gt;9,559&lt;/td&gt;
&lt;td&gt;7,030&lt;/td&gt;
&lt;td&gt;735.43&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Generate&lt;/td&gt;
&lt;td&gt;45,326&lt;/td&gt;
&lt;td&gt;1,789&lt;/td&gt;
&lt;td&gt;39.47&lt;/td&gt;
&lt;td&gt;33,186&lt;/td&gt;
&lt;td&gt;2,084&lt;/td&gt;
&lt;td&gt;62.80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Generate&lt;/td&gt;
&lt;td&gt;80,190&lt;/td&gt;
&lt;td&gt;2,802&lt;/td&gt;
&lt;td&gt;34.94&lt;/td&gt;
&lt;td&gt;40,281&lt;/td&gt;
&lt;td&gt;2,499&lt;/td&gt;
&lt;td&gt;62.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Generate&lt;/td&gt;
&lt;td&gt;70,164&lt;/td&gt;
&lt;td&gt;2,405&lt;/td&gt;
&lt;td&gt;34.28&lt;/td&gt;
&lt;td&gt;35,046&lt;/td&gt;
&lt;td&gt;2,337&lt;/td&gt;
&lt;td&gt;66.68&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Generate&lt;/td&gt;
&lt;td&gt;84,325&lt;/td&gt;
&lt;td&gt;2,851&lt;/td&gt;
&lt;td&gt;33.81&lt;/td&gt;
&lt;td&gt;45,532&lt;/td&gt;
&lt;td&gt;2,858&lt;/td&gt;
&lt;td&gt;62.77&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Table 2: Actual data from the logs.&lt;/p&gt;

&lt;p&gt;Without MTP, the whole run took about 293,705 ms in total — just under 5 minutes — while with MTP it completed in about 173,150 ms, just under 3 minutes. That's roughly a 40% reduction in time for this case.&lt;/p&gt;

&lt;p&gt;This suggests MTP is a genuinely effective way to boost efficiency in tasks like today's agentic workflows, which repeat many rounds of inference.&lt;/p&gt;

&lt;p&gt;To look more closely at the trend, we also checked how speed changed as token output progressed within each turn (again using the logs).&lt;/p&gt;

&lt;p&gt;Below is the token output profile without MTP. Speed dips slightly but stays roughly flat overall, hovering in the mid-30s to low-40s tps across all four turns from start to finish.&lt;/p&gt;

&lt;p&gt;Generally, with token-by-token generation, output speed tends to gradually decline as context size grows. How much it degrades depends on the model, but this can become a non-trivial problem as more turns accumulate.&lt;/p&gt;

&lt;p&gt;By contrast, here is the token output profile with MTP enabled. Speed starts around 35–40 tps early on and climbs to 60–65 tps later, showing that tokens come out faster and faster as the conversation progresses.&lt;/p&gt;

&lt;p&gt;A quick note before we go further: I'm a native Japanese speaker, and this model is my daily driver, so the test prompts and responses are in Japanese rather than English. Tokenizer behavior can differ quite a bit across languages, so this also happens to double as a look at how MTP performs outside the English-first benchmarks you usually see.&lt;/p&gt;

&lt;p&gt;Next, we examined the actual content of the responses and found that each one opens with a Japanese explanation followed by code. (The original screenshot of that exchange is in Japanese, so it isn't reproduced here — the token breakdown below covers what matters for this analysis.)&lt;/p&gt;

&lt;p&gt;Based on that, we checked the token length of the Japanese explanation portion at the start of each turn, with the following results.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Token count of the Japanese portion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;107&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;147&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;124&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Table 3: Token length of the Japanese explanation at the start of each turn.&lt;/p&gt;

&lt;p&gt;We can't fully prove causation here, but turns where speed rose sharply (Turn 1, Turn 3) tended to have less Japanese output, while turns where speed rose more gradually (Turn 2, Turn 4) tended to have somewhat more Japanese output. This suggests the code-generation portion is where MTP's effect shows up most strongly.&lt;/p&gt;

&lt;p&gt;A likely reason is that, regardless of programming language, code follows fairly strict grammar rules, making it relatively easy to predict what string comes next, and since the characters are essentially alphabetic, predicting the character type is also easier.&lt;/p&gt;

&lt;p&gt;Even outside the code-generation portions, we confirmed output speeds of 40–50 tps right from the start — faster than the conventional approach — suggesting MTP was working effectively there too.&lt;/p&gt;

&lt;p&gt;The author believes this largely comes down to characteristics of the Gemma tokenizer.&lt;/p&gt;

&lt;p&gt;The Gemma tokenizer has a vocabulary of roughly 256k entries. Whereas many other tokenizers fall back to character-level tokenization for Japanese, Gemma's tokenizer is able to register most Japanese text as whole "words."&lt;/p&gt;

&lt;p&gt;Because of this, the model can predict the next token at the "word" level for Japanese much like it does for English, which we believe is why Japanese output was noticeably faster with MTP than without.&lt;/p&gt;

&lt;h3&gt;
  
  
  Checking Token Acceptance Rates
&lt;/h3&gt;

&lt;p&gt;llama.cpp is designed to log an "acceptance" value showing what fraction of tokens were accepted during the verification phase when using an MTP model. Here's an example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;prompt eval time = 9559.06 ms / 7030 tokens ( 1.36 ms per token, 735.43 tokens per second)
eval time = 45531.93 ms / 2858 tokens ( 15.93 ms per token, 62.77 tokens per second)
total time = 55090.99 ms / 9888 tokens
graphs reused = 3983
draft acceptance = 0.82931 ( 2196 accepted / 2648 generated)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using this value, we checked acceptance for each turn, with the following results.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Table 4: Token acceptance rate during the test cases&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;Acceptance&lt;/th&gt;
&lt;th&gt;Accepted&lt;/th&gt;
&lt;th&gt;Generated&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;75.34%&lt;/td&gt;
&lt;td&gt;1,564&lt;/td&gt;
&lt;td&gt;2,076&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;77.71%&lt;/td&gt;
&lt;td&gt;1,890&lt;/td&gt;
&lt;td&gt;2,432&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;88.27%&lt;/td&gt;
&lt;td&gt;1,821&lt;/td&gt;
&lt;td&gt;2,064&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;82.93%&lt;/td&gt;
&lt;td&gt;2,196&lt;/td&gt;
&lt;td&gt;2,648&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Acceptance&lt;/th&gt;
&lt;th&gt;Accepted Tokens&lt;/th&gt;
&lt;th&gt;Generated Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;88.27%&lt;/td&gt;
&lt;td&gt;2,196&lt;/td&gt;
&lt;td&gt;2,648&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Min&lt;/td&gt;
&lt;td&gt;75.34%&lt;/td&gt;
&lt;td&gt;1,564&lt;/td&gt;
&lt;td&gt;2,064&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg&lt;/td&gt;
&lt;td&gt;81.06%&lt;/td&gt;
&lt;td&gt;1,868&lt;/td&gt;
&lt;td&gt;2,305&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;As mentioned above, likely because the output contained a lot of code, these tokens showed a high acceptance rate. So what does the acceptance rate look like for everyday use? Based on several days of logs, we recorded the maximum, minimum, and average, shown below.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Table 5: Token acceptance rate for everyday interactions with the Gemma-4-QAT-Assistant model&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Acceptance&lt;/th&gt;
&lt;th&gt;Accepted Tokens&lt;/th&gt;
&lt;th&gt;Generated Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;75.55%&lt;/td&gt;
&lt;td&gt;2,119&lt;/td&gt;
&lt;td&gt;4,800&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Min&lt;/td&gt;
&lt;td&gt;33.56%&lt;/td&gt;
&lt;td&gt;227&lt;/td&gt;
&lt;td&gt;412&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg&lt;/td&gt;
&lt;td&gt;45.11%&lt;/td&gt;
&lt;td&gt;776&lt;/td&gt;
&lt;td&gt;1,798&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Acceptance roughly halved, and speed dropped somewhat as a result. The corresponding speeds are shown below.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Table 6: Token speeds under the conditions in Table 5&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Read&lt;/th&gt;
&lt;th&gt;Generate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Max&lt;/td&gt;
&lt;td&gt;1,040.82&lt;/td&gt;
&lt;td&gt;84.46&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Min&lt;/td&gt;
&lt;td&gt;117.52&lt;/td&gt;
&lt;td&gt;30.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg&lt;/td&gt;
&lt;td&gt;774.45&lt;/td&gt;
&lt;td&gt;40.44&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So while this isn't exactly blazing fast, it's reasonable to say it's still faster than running without MTP at all.&lt;/p&gt;

&lt;p&gt;We confirmed that, for the Gemma-4 series, llama.cpp and Gemma-4-Assistant are a very well-matched pair, capable of delivering real benefits even in everyday use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;In this article, we looked at the new "MTP" technique — what it is, and how it actually changes generation speed in practice.&lt;/p&gt;

&lt;p&gt;In a word, MTP is essentially the "speculative execution" of CPUs, brought into the world of LLMs: when the prediction is correct, generation gets dramatically faster; when it's wrong, the extra overhead can make it slower.&lt;/p&gt;

&lt;p&gt;We found this feature pairs especially well with code generation. Because programming languages follow strict grammar, once you're partway through a line, it's relatively easy to predict "what string should come next." This lets MTP compensate for the parts where the tokenizer tends to slow things down, and our benchmark reflected that with solid results.&lt;/p&gt;

&lt;p&gt;The current strength of LLMs comes precisely from their sequential nature — deciding the next word based on the words that came before. MTP, because it speeds things up without sacrificing that strength, is likely to play a genuinely important role as this technology continues to develop.&lt;/p&gt;

&lt;p&gt;It's still early days, and there are some constraints, but we think this is a technique well worth watching as it matures and spreads.&lt;/p&gt;

&lt;p&gt;Even as we speak, other groups are developing and releasing their own versions of this idea:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DeepSeek (China): a token-prediction technique called &lt;strong&gt;DSpark&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Z-Lab at UC San Diego: a token-prediction technique called &lt;strong&gt;DFlash&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These, too, are spreading quickly now that inference engines are starting to support them.&lt;/p&gt;

&lt;p&gt;We hope to take a closer look at these technologies in a future article.&lt;/p&gt;

&lt;h2&gt;
  
  
  Supplementary Information (llama.cpp Support Status, Gemma-4 Architecture Details)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  llama.cpp Support Status
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;2026/5/16: Support for Qwen's MTP feature merged (#22673)&lt;/li&gt;
&lt;li&gt;2026/5/22: A further updated version merged (logit computation skip optimization: #23433)&lt;/li&gt;
&lt;li&gt;2026/6/8: PR adding Gemma-4 MTP support merged (#23398).

&lt;ul&gt;
&lt;li&gt;Enables conversion to GGUF format including the MTP drafter.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;2026/6/8: Support for the Assistant drafter models for Gemma-4 E2B and E4B merged. (#24282)

&lt;ul&gt;
&lt;li&gt;However, a bug caused Gemma-4-E4B specifically to malfunction.&lt;/li&gt;
&lt;li&gt;Apparently fixed in PR #25148 below.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;2026/06/30: Fix for the Gemma-4 E4B bug (#25148) merged.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Gemma-4 Model Architecture
&lt;/h3&gt;

&lt;p&gt;Below is the architecture of the Gemma-4-12B-it model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input strings ─┬─ Vision/Audio encoders (minimal, mostly folded into the model) ─┐
                └─ tok → Emb ──────────────────────────────────────────────────┴─(+)─┐
                                                                                       │
     ┌── RoPE / p-RoPE ───────────────────────────────────────────────────────────────┤
     │                                                                                 ▼
     │      [ LA → LA → LA → LA → LA → GA ]  x6 groups   →  LNR → SoftMax → Output Probabilities
     └──────────────────────────────────────────────────↑
            (48 blocks total = 6 layers x 8 groups)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Max context size: 256k tokens (limited to 160k in our test environment)&lt;/li&gt;
&lt;li&gt;48 blocks overall (6 layers x 8 groups)&lt;/li&gt;
&lt;li&gt;Hidden size: 3,840&lt;/li&gt;
&lt;li&gt;FFN type: Dense (omitted from the diagram above for simplicity)

&lt;ul&gt;
&lt;li&gt;Activation function: Gelu_Pytorch_tanh&lt;/li&gt;
&lt;li&gt;Activation dimension: 15,360&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The attention structure is a hybrid of Local Attention (LA) and Global Attention (GA). This structure debuted with Gemma-4.

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Local Attention (LA):&lt;/strong&gt; uses sliding window attention.&lt;/li&gt;
&lt;li&gt;Handles analysis of local context via a sliding window.&lt;/li&gt;
&lt;li&gt;Low compute cost, fast, and low memory usage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global Attention (GA):&lt;/strong&gt; uses group query attention for accurate whole-sequence understanding.&lt;/li&gt;
&lt;li&gt;Captures the entire sequence.&lt;/li&gt;
&lt;li&gt;A high-performance coordinator handling long context and complex reasoning.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Other notable features:

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unified Encoder:&lt;/strong&gt; integrates audio/vision encoders into the model.&lt;/li&gt;
&lt;li&gt;Normally-external audio/vision encoders are almost entirely integrated into the model itself.&lt;/li&gt;
&lt;li&gt;Only the CLIP portion remains an externally applied encoder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-Layer Embedding (PLE):&lt;/strong&gt; keeps embedding information independent at each layer.&lt;/li&gt;
&lt;li&gt;Beyond the input-time embedding, retaining layer-specific information contributes to richer representational power during inference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;KV Cache Sharing:&lt;/strong&gt; a mechanism for sharing and synchronizing the KV cache.&lt;/li&gt;
&lt;li&gt;Normally kept independent, this mechanism shares and synchronizes the KV cache instead.&lt;/li&gt;
&lt;li&gt;Used to achieve more consistent inference.&lt;/li&gt;
&lt;li&gt;When dispatching across multiple GPUs, this shared-state synchronization causes frequent inter-GPU communication, so caution is needed in setups without NVLink.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DualRoPE configuration:&lt;/strong&gt;

&lt;ul&gt;
&lt;li&gt;Local Attention layers use RoPE.&lt;/li&gt;
&lt;li&gt;Embeds token position information using rotation matrices, ensuring reliable positional information for long text.&lt;/li&gt;
&lt;li&gt;Global Attention layers use p-RoPE&lt;sup id="fnref3"&gt;3&lt;/sup&gt;.&lt;/li&gt;
&lt;li&gt;A modified form of RoPE that applies rotation only to the top p% of K/V matrix pairs.&lt;/li&gt;
&lt;li&gt;Because it effectively suppresses the impact of noise, it's applied to the coordinating Global Attention layers, helping clarify positional information for long text.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Gemma-4-12B-Assistant Model Architecture Details
&lt;/h3&gt;

&lt;p&gt;The Gemma-4-12b-it-assistant model works together with the main model to predict tokens as follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Main model (Gemma-4-12b-it):
  Input string → tok → Emb → [LA...LA(45) → GA(46) → GA(47)] → LNR → SoftMax → N+1 output token
                                     │ shares KV cache (layers 45-47) with the Assistant's LA/GA
                                     ▼
Assistant model (MTP-Drafter):
  Query (H_N^0 + Emb of N) → LNR → LA → LA → LA → GA → Norm ─┬→ LNR → predicted state vectors
                                                              └→ Masked Embedder → predicted Token IDs (n tokens)
                                     │
                                     ▼
Verification (inside the inference engine):
  compares the Assistant's predicted state vectors against the Main model's own state vectors
  → decides how many tokens to accept → token output correction → final output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The "Assistant" model that adds MTP capability to Gemma-4-12B-it keeps Gemma-4's architecture but trims its structure down to a genuinely minimal design.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Max context size: 256k tokens&lt;/li&gt;
&lt;li&gt;4 blocks overall (a single group of 3+1)&lt;/li&gt;
&lt;li&gt;Hidden size: 1,024&lt;/li&gt;
&lt;li&gt;FFN type: Dense

&lt;ul&gt;
&lt;li&gt;Activation function: Gelu_Pytorch_tanh&lt;/li&gt;
&lt;li&gt;Activation dimension: 8,192&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Based on the input string, it predicts the next (N+1) token. To do this, it combines the generated state vector h_N^0 with the embedded vector of token N, and feeds the result in as a query to begin processing.&lt;/p&gt;

&lt;p&gt;Gemma-4's attention layers normally operate as self-attention, but the structure inside the Assistant model uses cross-attention instead.&lt;/p&gt;

&lt;p&gt;For KV data, it combines the LocalAttention block at layer 46 of the main model with the LocalAttention blocks at layers 0–2 on the Assistant side, and also integrates the GlobalAttention block at layer 47 of the main model with the GlobalAttention block at layer 3 on the Assistant side, via a shared KV cache. Through this, KV information is bound based on the query.&lt;/p&gt;

&lt;p&gt;The output consists of two parts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Predicted tokens:&lt;/strong&gt; the n tokens predicted by the Assistant model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;State vector:&lt;/strong&gt; the vector used during that prediction process.&lt;/p&gt;

&lt;p&gt;These outputs are sent combined, just before the main model's attention blocks. They pass through all of the main model's attention blocks, generating, in parallel, the state vectors needed for verification.&lt;/p&gt;

&lt;p&gt;Once the state vectors to be compared are ready, their contents (probability distributions) are compared to decide how many tokens should be accepted. This is handled in parallel by the inference engine (for example, the &lt;code&gt;generate&lt;/code&gt; function in the transformers library), which outputs the resulting number of tokens to accept.&lt;/p&gt;

&lt;p&gt;Based on this result, the candidate tokens — accounting for the accepted token count — are combined with token N+1 and passed on for output. This is how the process achieves faster output compared to conventional step-by-step inference.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article is an English adaptation of the original Japanese post published on Zenn: &lt;a href="https://zenn.dev/highreso/articles/40f75642b9d73e" rel="noopener noreferrer"&gt;"実際どれだけ速くなる？ Gemma-4のMTPを、普段使いのPCでllama.cpp検証してみた【検証編】"&lt;/a&gt;, by Yuichi Tominaga.&lt;/em&gt;&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;&lt;a href="https://unsloth.ai/docs/jp/ji-ben/unsloth-dynamic-2.0-ggufs" rel="noopener noreferrer"&gt;https://unsloth.ai/docs/jp/ji-ben/unsloth-dynamic-2.0-ggufs&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;&lt;a href="https://unsloth.ai/" rel="noopener noreferrer"&gt;https://unsloth.ai/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;&lt;a href="https://www.linkedin.com/posts/xuan-thai-nguyen-b31b0b3a_little-notes-on-p-rope-in-gemma-4-first-activity-7446051476784951296-vamg" rel="noopener noreferrer"&gt;https://www.linkedin.com/posts/xuan-thai-nguyen-b31b0b3a_little-notes-on-p-rope-in-gemma-4-first-activity-7446051476784951296-vamg&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>MTP Explained: How Qwen and Gemma Predict Several Tokens Ahead</title>
      <dc:creator>oooocean66</dc:creator>
      <pubDate>Wed, 09 Sep 2026 07:53:18 +0000</pubDate>
      <link>https://dev.to/oooocean66/mtp-explained-how-qwen-and-gemma-predict-several-tokens-ahead-4703</link>
      <guid>https://dev.to/oooocean66/mtp-explained-how-qwen-and-gemma-predict-several-tokens-ahead-4703</guid>
      <description>&lt;p&gt;&lt;strong&gt;MTP (Multi-Token Prediction)&lt;/strong&gt; is a technique that lets a language model predict several tokens ahead while it generates text, then confirm multiple tokens at once when those predictions turn out correct.&lt;/p&gt;

&lt;p&gt;Have you ever used an AI chat and thought, "I wish this answered a little faster"? That feeling gets stronger with "Thinking" models that reason carefully before answering, or with coding agents that iterate through trial and error. For AI systems that burn through thousands, even tens of thousands, of tokens to reach a good answer, generation speed directly shapes the user experience.&lt;/p&gt;

&lt;p&gt;So why does AI text generation take so long in the first place? The reason is simple: LLMs can only produce one token at a time. To decide the next word, the model has to re-run its massive computations from scratch, taking into account everything generated so far. That cost adds up, and the longer the response, the longer you wait.&lt;/p&gt;

&lt;p&gt;This raises a natural question: instead of recomputing everything one word at a time, why not compute a few words ahead in the same pass? That's exactly what MTP does. Every time the model generates a word, it also computes, in that same forward pass, a prediction for what the next few words are likely to be. If the prediction turns out correct, those predicted tokens are accepted all at once, confirming several words in a single step. If it's wrong, the model simply falls back to generating one token at a time as usual — so a wrong prediction costs little, while a correct one yields a meaningful speedup.&lt;/p&gt;

&lt;p&gt;This technique is often called &lt;strong&gt;speculative decoding&lt;/strong&gt;, a term borrowed from an older CPU optimization technique called "speculative execution." CPUs have long predicted which branch of code they're about to take, computed ahead of time based on that prediction, and either kept the result if it was right or discarded it and redid the work if it was wrong. MTP applies the same underlying idea to language model inference.&lt;/p&gt;

&lt;p&gt;This article (the concept edition) walks through how this mechanism is actually implemented inside real models (Qwen and Gemma), and how the two approaches differ. In the next installment — the implementation/benchmark edition — we'll actually run this technique and measure exactly how much faster it gets. Later in the series, we'll also cover "DFlash," which pushes this idea even further.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is MTP, Exactly?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;MTP (Multi-Token Prediction)&lt;/strong&gt; is a technique where an AI predicts several tokens ahead simultaneously while generating text. By predicting tokens in advance, the model can confirm multiple tokens at once when the prediction turns out correct, speeding up output.&lt;/p&gt;

&lt;p&gt;Similar efforts exist elsewhere: instead of generating text sequentially, some "diffusion" models generate text more like image generation, producing output in a more randomized order to speed things up (as of 2026, Google DeepMind and ELYZA, among others, appear to be researching diffusion language models). However, output quality for these remains unstable and hasn't reached a practical level yet.&lt;/p&gt;

&lt;p&gt;At least as of 2026, MTP is arguably the most reliable LLM speedup technique available today.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which Models Support MTP?
&lt;/h3&gt;

&lt;p&gt;The models currently supporting MTP fall mainly into two categories. Each uses a different implementation approach, so support in inference engines needs to be checked individually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen-3.5 / 3.6:&lt;/strong&gt; MTP is natively baked into the model itself — the model and its MTP drafter are a single package. It's supported out of the box in engines like vLLM. llama.cpp added experimental support relatively early on, and Unsloth has released an MTP-enabled model for Qwen-3.5-9B in GGUF format.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemma:&lt;/strong&gt; MTP is achieved by pairing the model with a &lt;em&gt;separate&lt;/em&gt; model called "Gemma-4-Assistant" — think of it as an add-on bolted onto the main model rather than something built in. llama.cpp support for this arrived somewhat later than for Qwen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Basic Mechanics of MTP
&lt;/h2&gt;

&lt;p&gt;Let's look at how behavior differs with MTP disabled versus enabled.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;[Conventional Inference (Token-by-Token)]&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One forward pass predicts one token; the next step runs another forward pass. Because output comes one token at a time, this is relatively slow.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step N (full forward pass) → token generated → Step N+1 (full forward pass) → token generated → ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;[Inference with MTP]&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A single forward pass computes a certain number of tokens ahead.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Step N (forward pass, run once) ──┬─→ Token N     [confirmed]
                                   ├─→ Token N+1   [speculative]
                                   └─→ Token N+2   [speculative]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Token 1: output as confirmed.&lt;/li&gt;
&lt;li&gt;Token 2 onward: probability values are calculated in parallel as speculative predictions.&lt;/li&gt;
&lt;li&gt;Because output comes in multi-token units, this is relatively faster than the conventional approach.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Handling Predicted Tokens (Speculative Decoding)
&lt;/h2&gt;

&lt;p&gt;Predicted tokens go through a follow-up step where they're checked against the correct answer; once they clear the acceptance criteria, they're confirmed all at once.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Draft&lt;/strong&gt;
A lightweight method quickly predicts upcoming tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification&lt;/strong&gt;
The production model computes the probability distribution for the predicted tokens in parallel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Acceptance&lt;/strong&gt;
The system checks how closely the probability distribution from verification matches the distribution the model would have produced without MTP, and determines whether it meets the acceptance threshold.
If it matches, the tokens up to that point are accepted and confirmed all at once.
If it doesn't match, drafting is cut off at the point of mismatch, and the result falls back to the production model's own inference.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Draft → Verification → Acceptance ─┬─ match ────→ accept all at once
                                    └─ mismatch ─→ cut off, fall back to the
                                                    production model's answer,
                                                    then Draft again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because tokens appear to be generated "simultaneously," this is often mistaken for a diffusion model, but &lt;strong&gt;the underlying process is still strictly sequential.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In other words, this approach involves a real trade-off. When the prediction is correct, more tokens get confirmed at once, yielding faster output. But when it's wrong, the extra computation goes to waste, and the overhead can make the result slower than running the base model without MTP at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the Two Approaches to MTP Differ
&lt;/h2&gt;

&lt;p&gt;As mentioned, this article covers two implementations, and they differ considerably in their starting point, goals, and mechanics. Let's look at each in more detail.&lt;/p&gt;

&lt;h3&gt;
  
  
  Qwen's Approach to MTP
&lt;/h3&gt;

&lt;p&gt;Qwen's approach is said to build on techniques researched for DeepSeek-V3, and has been used starting with Qwen3-Next.&lt;/p&gt;

&lt;p&gt;The reason this approach is built directly into the model is that the technique itself originally started out as &lt;strong&gt;"a method for training a smarter model."&lt;/strong&gt; MTP turned out to be useful, and the circuitry built into the model for that purpose is now also used, as a side effect, for speculative token prediction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input token (N) → Qwen3.5 Main Model → state h_N^0 → Token N+1 [confirmed]
                                             │
                                             ▼ (also feeds the MTP path)
                                       MTP Module 1 → h_N^1 → Sampling &amp;amp; accept
                                             │
                                   match? ───┴─── no match?
                                     │                │
                        predicted token accepted   main model's token adopted
                          → confirmed as N+2         → confirmed as N+2, chain stops
                             │
                             ▼ (only if accepted)
                       MTP Module 2 → h_N^2 → Sampling &amp;amp; accept → ...same check for N+3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's how it works, broken down. Even with MTP disabled, when a token is input, the model's main processing generates a state h_N^0 based on the token sequence so far, which is then detokenized to produce token N+1.&lt;/p&gt;

&lt;p&gt;When MTP is enabled, as soon as state h_N^0 is produced, the requested number of MTP modules are created. If n tokens' worth of prediction is requested, n modules are prepared, each producing an embedding Emb_N^n that corresponds to a predicted word, based on state h_N^0.&lt;/p&gt;

&lt;p&gt;After that, as shown in the figure above, the embedding information for token N+1 — the token that would have been output even without MTP — is combined to produce state h_N^1. This state h_N^1 is sent to the main model, where its sampling and acceptance mechanism checks whether it matches what the main model would have predicted on its own.&lt;br&gt;
If it matches, that token is accepted and the process moves on to predicting the next token.&lt;br&gt;
If it doesn't match, the state predicted by the main model is used instead, and verification stops there.&lt;/p&gt;

&lt;p&gt;The defining feature is that the prediction and verification mechanisms are chained together one token at a time, forming a sequential flow throughout.&lt;/p&gt;
&lt;h3&gt;
  
  
  Gemma's Approach to MTP
&lt;/h3&gt;

&lt;p&gt;Gemma's approach was built from the ground up for speed, and its key difference from Qwen is that it processes prediction and verification together, in a batch.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input token (N) → Gemma 4 Main Model → state h_N^0 → Prediction: N+1
                                             │
                                             ▼ (KV-cache update)
                                         KV-Cache ←──────────────┐
                                             │                   │ (shares cache)
                                             ▼                   │
                                    Gemma 4 Assistants ──────────┘
                                             │  (generates sequentially inside the
                                             │   draft model, then sends as a batch)
                                             ▼
                            Prediction: N+2, N+3, N+4, ... (candidate list)
                                             │
                                             ▼
                     Main model: causal-attention masking, probabilities computed
                     in parallel  ──┬── include only as many as fit into the output
                                    └── if none fit, exclude from output
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In Gemma, the draft model is external. Once the main model confirms the first token, the draft model predicts, all at once, however many tokens are allowed.&lt;/p&gt;

&lt;p&gt;The main model first receives the input tokens, generates a state, detokenizes it, and outputs token N+1. At this point, the draft model shares a KV cache with the main model (in Qwen's case, the MTP component and the main component maintain independent KV caches).&lt;/p&gt;

&lt;p&gt;Upon receiving the updated KV cache, the draft model generates however many predicted tokens are needed, and hands them off to the main model. The main model then checks, in a batch, whether these predicted tokens are correct using probability distributions, and determines how many of them are acceptable.&lt;/p&gt;

&lt;p&gt;As a result, however many predicted tokens the main model accepts get output together with token N+1, all at once.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which Approach to MTP Is Better?
&lt;/h3&gt;

&lt;p&gt;Because the two approaches to MTP (Multi-Token Prediction) start from different premises, it's hard to say one is unconditionally superior. That said, evaluating primarily on maturity and compatibility with inference engines, as of July 2026, Gemma's approach appears more mature.&lt;/p&gt;

&lt;p&gt;First, there's a difference in the scope of sequential processing. Because token prediction fundamentally assumes "predicting the next token based on the state of the previous one," the process is inherently sequential. Comparing the flow of each approach, with time running left to right:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Qwen:   Input(N) → Main model → Predict N+2 → Judge N+2 → Predict N+3 → Judge N+3 → ...
                        │             confirmed:N+2 ↑          confirmed:N+3 ↑
                        └→ confirmed: N+1

Gemma:  Input(N) → Main model ──┬→ Predict N+2 ─┐
                        │       ├→ Predict N+3 ─┼→ Judge (batch) → up to n OK → confirmed: N+2, N+3, N+4, ...
                        │       └→ Predict N+4 ─┘
                        └→ confirmed: N+1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;strong&gt;Qwen's&lt;/strong&gt; approach, the main model predicts token N+1, uses that state to predict token N+2 and passes it to the main model, then predicts token N+3 only after seeing the verification result — repeating this process step by step.&lt;/p&gt;

&lt;p&gt;By contrast, with &lt;strong&gt;Gemma&lt;/strong&gt;, after predicting token N+1, the &lt;strong&gt;draft model&lt;/strong&gt; predicts however many tokens are needed all at once (starting from N+2), and hands them to the main model as a single batch. Processing on the main model's side is also parallelized, making this approach far more efficient than Qwen's. Because of this structural difference, Gemma's approach comes out ahead.&lt;/p&gt;

&lt;p&gt;Second, there's a difference in structural flexibility. Because Qwen's MTP functionality is integrated into the main model, improving the drafting logic requires additional post-training. Gemma, on the other hand, keeps the draft model separate, so it can simply be swapped out whenever better logic becomes available. This separation is also advantageous operationally.&lt;/p&gt;

&lt;p&gt;Finally, there's the ease of implementation in inference engines. Looking at llama.cpp's implementation, Gemma's approach can run even in multimodal configurations, while Qwen's approach doesn't support multimodal use. This difference stems from the fact that Qwen embeds MTP internally, requiring the branching logic to be handled inside the model itself. Gemma's approach, by contrast, simply toggles the feature on or off depending on whether input passes through the draft model, which is a more favorable structure for engine-side implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary (And a Preview of What's Next)
&lt;/h2&gt;

&lt;p&gt;MTP, at its core, works by confirming one token while, in that same pass, speculatively predicting a few tokens ahead at minimal extra cost&lt;sup id="fnref1"&gt;1&lt;/sup&gt;, then confirming them all together if the prediction turns out correct. Qwen builds this mechanism directly into the model itself; Gemma delegates it to a separate, dedicated draft model.&lt;/p&gt;

&lt;p&gt;In the next installment — the implementation/benchmark edition — we'll actually run both approaches on llama.cpp and measure, with real benchmarks, exactly how much faster they get and how often the predictions turn out correct. Later in the series, we'll also cover "DFlash," which takes this idea even further.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;Qwen3-Next: Towards Ultimate Training &amp;amp; Inference Efficiency&lt;br&gt;
&lt;a href="https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd" rel="noopener noreferrer"&gt;https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;DeepSeek-V3 Technical Report&lt;br&gt;
&lt;a href="https://arxiv.org/pdf/2412.19437" rel="noopener noreferrer"&gt;https://arxiv.org/pdf/2412.19437&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Multi-token prediction with Gemma using Hugging Face Transformers&lt;br&gt;
&lt;a href="https://ai.google.dev/gemma/docs/mtp/mtp" rel="noopener noreferrer"&gt;https://ai.google.dev/gemma/docs/mtp/mtp&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article is an English adaptation of the original Japanese post published on Zenn: &lt;a href="https://zenn.dev/highreso/articles/2fdb5402d98e5d" rel="noopener noreferrer"&gt;"小さく賭けて、大きく当てる：LLMを高速化する「投機的デコード（MTP）」の正体【概念編】"&lt;/a&gt;, by Yuichi Tominaga.&lt;/em&gt;&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;Speculating too many tokens ahead raises the cost of a wrong prediction, which can outweigh the benefit and slow things down. "Minimal extra cost" here means relative to the cost of recomputing everything from scratch for every single token in the conventional approach.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
  </channel>
</rss>
