DEV Community

oooocean66
oooocean66

Posted on

DFlash-2: Benchmarking Z-Lab's Successor to DFlash for Accuracy and Throughput Gains

A while back, we covered DFlash, a draft-token prediction technique that uses a diffusion model. At the time, we tested it on Gemma-4-12b-it-QAT, and the native Assistant model came out ahead — DFlash wasn't able to show a clear advantage.

Recently, though, a successor called "DFlash-2" surfaced, with a number of enhancements on top of the original design.

As of August 2026, only a handful of models support it. Using the base model alone, DFlash, and DFlash-2 in turn, we looked at where each one actually pays off, and what parameter settings get the most out of it.

What Is DFlash-2?

DFlash is a technique that predicts upcoming output tokens ahead of time to speed up generation, built by Z-Lab. Several DFlash models have been released for different LLM families.

[Input Token (N)] --> [Main Model] -- h_on ----------------------------------------> [Predicted: N+1]
                                       |                                             ^
                                       v                                             |
                                  [KVCache]                                          |
                                       |                                             |
+--------------------------------------+--------------------------------+            |
| DFlash Drafter                       v (KV data injection)            |            |
|                                 [KVCache]                             |            |
|                                      |                                |            |
|                              [Diffusion Model]                        |            |
|                                      |                                |            |
|                               [Token Decoder]                         |            |
|                                      |                                |            |
|                            +---------v--------+                       |            |
|                            | Predicted N+2    |                       |            |
|                            |       +          |                       |            |
|                            | Predicted N+3    |                       |            |
|                            |       +          |                       |            |
|                            | Predicted N+4    |                       |            |
|                            +---------+--------+                       |            |
+--------------------------------------+--------------------------------+            |
                                       |                                             |
+--------------------------------------+---------------------------------------------+-------+
| Processing inside Main Model         v                                             |       |
|                             * Uses causal attention to mask;                       |       |
|                               computes probs in parallel                           |       |
|                                      +------------------------------+              |       |
|                                      v                              v              |       |
|                         (Match found)                  (No match)                  |       |
|                         Include matching portion       Nothing included in output  |       |
|                         -> Predicted: N+2                                          |       |
|                         -> Predicted: N+3                                          |       |
+------------------------------------------------------------------------------------+-------+
Enter fullscreen mode Exit fullscreen mode

Figure 1: The DFlash mechanism (this structure is unchanged in DFlash-2; reused from our earlier DFlash piece)

DFlash speeds up draft generation by using a non-causal model (the diffusion model shown above), aiming for better throughput than conventional autoregressive drafting.

The tradeoff is that the diffusion model also introduces some accuracy loss — in our earlier piece, it fell behind Gemma-4's Assistant model in a head-to-head comparison.

inco.ai has been working on fixing that. Details on the company are scarce, but the maintainer of inco.ai's page on Hugging Face is Zhijian Liu, who also heads Z-Lab — so it's likely a spinoff team or startup connected to the lab.

DFlash-2 adds two things on top of the original mechanism:

  • Lightweight Path Selector — picking the correct token ordering The original DFlash's strength was that its internal diffusion model could output a whole block of tokens, in order, in one shot. The catch is that each token's ordering is generated independently, so the resulting sequence could end up internally inconsistent, and it looks like that was a common reason predictions got rejected.1

To address this, DFlash-2 adds a very lightweight internal model that checks whether a candidate token sequence is "the most natural ordering" and filters out anything that isn't. That's the Lightweight Path Selector. It sits right after the Target LM Head layer inside the drafter model.

  • Local Convolution — fixing the drop-off in accuracy toward the end of a block This layer limits information exchange to each token's immediate neighbors when linking tokens together. Working alongside the Path Selector, it counters "suffix decay" — the sharp drop in acceptance rate for predicted tokens toward the end of a block. The Local Convolution layer is inserted in multiple places between the attention and MLP layers.

Put together, DFlash-2 is best described as an enhancement of DFlash that specifically improves the accuracy of its draft-token ordering.

Two models are currently available:

  • Qwen3.8-27B-DFlash2
  • Muse-Glimmer-30B-DFlash2

Both are published on the Z-Lab and inco.ai repositories, with GGUF builds available for llama.cpp.

We'd originally planned to run this on Gemma-4-12b-it-qat, but since that's the only pairing currently offered, and given the lighter inference footprint, we decided to apply it to Muse Glimmer 30B, the model Meta released, and evaluate it there instead.

Figure 2: How Muse Glimmer 30B connects to DFlash-2
Figure 2: How Muse Glimmer 30B connects to DFlash-2 (diagram labels are in Japanese; it illustrates the Path Selector / Local Convolution placement already described above, so it isn't required reading)

The reason for picking Muse Glimmer 30B is simple: it's recently become our own day-to-day model of choice, replacing Gemma-4-12b-qat-Assistant.

Running It in llama.cpp

Where does llama.cpp stand as of August 2026?

As of August 23, 2026, DFlash-2 support hasn't landed on llama.cpp's master branch yet.2 That said, Z-Lab has already submitted a pull request, and building from it gets you an inference engine with DFlash-2 support.

spec: add DFlash2 support (local convolution + candidate selector) — #27342

https://github.com/ggml-org/llama.cpp/pull/27342

Building llama.cpp from this PR gives you an inference engine that supports DFlash-2:

# Build llama.cpp with PR #27342 applied
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/27342/head:pr-27342
git switch pr-27342
# Build command for NVIDIA CUDA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j
Enter fullscreen mode Exit fullscreen mode

From there, running it as follows gets you a working inference engine:

../dflash/llama.cpp/build/bin/llama-server \
--model ./models/Muse-Glimmer-30B-UD-Q2_K_XL.gguf \
-t 12 -np 1 --prio 2 --temp 1.0 --top-p 0.95 --top-k 64 \
--port 8001 --host 0.0.0.0 --fit off --no-warmup --no-cache-prompt \
-fa on --cache-ram 0 -c 131072 --reasoning on --log-verbosity 4 \
--model-draft ./models/Muse-Glimmer-30B-DFlash2-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 4
Enter fullscreen mode Exit fullscreen mode

A quick rundown of the arguments:

Argument Recommended value What it does
--model-draft DFlash-2 GGUF filename Specifies the DFlash-2 model file
-sm layer Used when multiple GPUs are present; layer-wise distribution is generally recommended
--spec-type draft-dflash DFlash-2 models fall under draft-dflash
--spec-draft-n-max 4 or 6 Caps how many tokens ahead to predict; 4 or 6 is a reasonable starting point

Verification: Using DFlash-2 on Muse Glimmer

Test Environment

We ran the following simplified verification setup.

The setup connects through an access server over an SSH tunnel to reach the target instance.3 llama.cpp's service port is relayed between localhost and the remote instance through port 8001 on both ends.

 Home                                 Highreso: GPUSOROBAN
+----------------+                  +------------------+      +--------------------+
| [>_]  8001/tcp |==(SSH tunnel)==> | [>_]  8001/tcp   |----->| [llama.cpp]        |
|    Client      |                  |  Access Server   |      |  8001/tcp          |
+----------------+                  +------------------+      |  Target Instance   |
                                                                +--------------------+
Enter fullscreen mode Exit fullscreen mode

Figure 3: Connection setup via GPUSOROBAN

Instance Used

We used an NVIDIA RTX A4000 GPU for this test. Full instance specs below.

Item Spec
Instance type s16-1-a-standard-ubs24-v
GPU NVIDIA RTX A4000
GPU memory GDDR6 16GiB
GPU memory bandwidth 448.0 GB/s
FP32 compute 19.17 TFLOPS
BF16 compute 38.34 TFLOPS
INT8 compute 153.4 TOPS
INT4 compute 306.7 TOPS
vCPU 11 cores
System memory 50 GiB
Storage 100 GiB (persistent)
CUDA version 13.2
NVIDIA driver 580
OS Ubuntu 24.04 Server

Table 1: GPUSOROBAN instance specs

Model Files Used

Even at 4-bit quantization, Muse Glimmer's weights alone exceed 16GiB. Accounting for KV-cache overhead, we used the 2-bit quantized build optimized with Unsloth Dynamic 2.0.

We downloaded the following files from Unsloth's Muse Glimmer repository:

File Size Description
Muse-Glimmer-30B-UD-Q2_K_XL.gguf 12.4 GB UD 2.0 / 2-bit quantized main model
dflash-kquant.gguf 1.63 GB Quantized DFlash draft model
Muse-Glimmer-30B-DFlash2-Q4_K_M.gguf 1.65 GB Quantized DFlash-2 draft model

Table 2: Files applied to llama.cpp

Test Cases

  • Goal: measure and compare throughput across three configurations — main model only, DFlash, and DFlash-2.
    • Only speed is measured; output content is not evaluated.
    • We record wall-clock time, token count, and token throughput at generation.
    • These figures come directly from llama.cpp's own log output.
  • Method: using the llama.cpp web frontend, we ran two patterns, clearing the session between each:
    • Single-turn, Japanese-heavy pattern "Tell me about the geological characteristics of Kyushu."
    • Multi-turn, code-heavy pattern
    • Write JavaScript code for Breakout (as an HTML file).
    • Make it look cooler.
    • Slow down the ball's movement a bit.
    • Check for bugs and optimize it.

Memory Usage Comparison

Here's how memory usage came out. (Main-model-only is excluded from this comparison.)

Memory category DFlash CUDA DFlash CPU DFlash-2 CUDA DFlash-2 CPU
Weight data (UD-Q2_K_XL) 10,803.14 1,052.08 10,803.14 1,052.08
KV-cache data (f16) 1,664.00 — 1,664.00 —
Sliding-window KV cache 97.50 — 97.50 —
Compute buffer 273.52 156.52 273.52 156.52
MTP model weight data (Q4_K_M) 1,543.17 0.77 1,556.95 0.77
MTP model KV-cache size (f16) 50.00 — 50.00 —
MTP model compute buffer 407.62 15.51 462.67 15.51
Total 14,838.95 1,224.88 14,907.78 1,224.88

Table 3: Memory usage comparison (MiB)

Aside from a modest increase in the MTP model's weight data and compute buffer, memory usage is essentially unchanged. The added layers (the selector and the convolution layers) account for the increase, but combined they add up to under 100MiB.

Token Throughput Comparison (Main / DFlash / DFlash-2)

Single-Turn, Japanese-Heavy Results

Turn Read tps: main Read tps: dflash Read tps: dflash-2 Generate tps: main Generate tps: dflash Generate tps: dflash-2
Turn 1 259.37 169.68 159.33 22.52 17.8 20.42

Table 4: Read and generation throughput for each configuration

Here's how generation throughput moved over the course of the response:

Figure 4: Generation throughput over time by configuration (single-turn, Japanese-heavy)
Figure 4: Generation throughput over time by configuration

Throughput started out relatively fast during the initial English portion, then dropped noticeably once the Japanese section kicked in. Even so, DFlash-2 stayed ahead of DFlash throughout.

For this single-turn, Japanese-heavy case at least, neither draft model contributed much.

Multi-Turn, Code-Heavy Results

This is where things get interesting — here's the multi-turn, code-heavy comparison.

Turn Read tps: main Read tps: dflash Read tps: dflash-2 Generate tps: main Generate tps: dflash Generate tps: dflash-2
Turn 1 273.02 167.11 267.91 18.97 29.96 32.9
Turn 2 658.39 545.22 605.06 18.72 31.2 30.57
Turn 3 636.62 616.53 600.89 17.97 34.11 37.07

Table 5: Read and generation throughput for each configuration

Both DFlash and DFlash-2 posted strong throughput here. The upward trend across turns is characteristic of MTP in general. In a coding-agent-style workload, DFlash-2 delivers a substantial throughput gain.

Figure 5: Generation throughput over time by configuration (multi-turn, code-heavy)
Figure 5: Generation throughput over time by configuration

This confirms something our earlier DFlash testing didn't cover: in the "multi-turn, code-heavy" scenario, DFlash is genuinely useful for models like Muse Glimmer that, unlike Gemma-4, don't ship with a purpose-built Assistant-style MTP model of their own.

On top of that, DFlash-2's new additions push acceptance even higher, which translates directly into faster output.

"Multi-turn, code-heavy" is also a reasonably good proxy for typical coding-agent usage, suggesting this setup is especially effective in workloads that generate large volumes of tokens.

Token Acceptance Rate Comparison

Let's look at the predicted-token acceptance rate.

Single-Turn, Japanese-Heavy

Acceptance is fairly low here. DFlash in particular dips under 2, which would normally be read as "not worth using."

Turn Acceptance: main Acceptance: dflash Acceptance: dflash-2 mean len: main mean len: dflash mean len: dflash-2
Turn 1 — 15.73% 31.36% — 1.94 2.25

Table 6: Acceptance and mean-length values for each configuration

Multi-Turn, Code-Heavy

Acceptance is markedly higher here. DFlash-2 in particular pushes mean length past the configured maximum of 4 by turn 3, a notably strong result.

Turn Acceptance: main Acceptance: dflash Acceptance: dflash-2 mean len: main mean len: dflash mean len: dflash-2
Turn 1 — 54.49% 60.97% — 3.18 3.44
Turn 2 — 56.78% 57.34% — 3.27 3.3
Turn 3 — 68.06% 78.06% — 3.72 4.12

Table 7: Acceptance and mean-length values for each configuration

Higher acceptance means more predicted tokens get used, which drives throughput up — the two appear to move together.

In other words: the higher the acceptance rate, the more useful the predicted tokens become, and the faster generation gets.

Getting More Out of DFlash

Finding the Right Draft Length

We reran the multi-turn, code-heavy pattern with the predicted-token length set to 6. Results below.

Turn Acceptance: main Acceptance: 4 tokens Acceptance: 6 tokens mean len: main mean len: 4 tokens mean len: 6 tokens
Turn 1 — 60.97% 46.38% — 3.44 3.78
Turn 2 — 57.34% 47.29% — 3.3 3.84
Turn 3 — 78.06% 58.13% — 4.12 4.49

Table 8: Acceptance / mean-length comparison by draft length

Acceptance drops at 6 tokens, and mean length does increase, but only by less than one token — not much of a net gain.

Let's look at how this affects throughput over time.

Figure 6: Generation throughput over time by draft length
Figure 6: Generation throughput over time by draft length

Overall, the 6-token setting came out lower. There were a few stretches in turn 2 where it briefly pulled ahead, but turns 1 and 3 clearly favored the shorter setting.

Looking at this, mean length and the probability that a given draft length "lands" appear to track each other closely — rather than pushing for a longer draft length, it seems best to pick the smallest integer that exceeds the observed mean length.

That tracks: when you set a maximum draft length, DFlash's internal model tries to generate predicted tokens all the way out to that length. If the realistic prediction horizon is 4 tokens, everything predicted beyond that is wasted work.

Unsurprisingly, not generating wasted tokens is what maximizes throughput. With that in mind, a reasonable approach for tuning DFlash looks like this:

  1. Apply the DFlash model to your LLM.
  2. Run a variety of inference workloads (start with --spec-draft-n-max at 4 or 6).
  3. Check the result:
    1. If mean length comes out under 2, skip DFlash for this model — done.
    2. If mean length is above 2, move to the next step.
  4. Check again:
    1. If mean length consistently exceeds the current --spec-draft-n-max by 1 or more, raise --spec-draft-n-max to the smallest integer above the observed mean length. Example: --spec-draft-n-max is 6, mean length is 7.5 → raise it to 8.
    2. Otherwise, keep the current --spec-draft-n-max setting as final. Example: --spec-draft-n-max is 6, mean length is 6.5 → leave it at 6.

Checking the Impact of KV-Cache Quantization

Next: how much does quantizing (or not quantizing) the KV cache — a meaningful factor in the speed/memory tradeoff — affect DFlash-2's performance?

4-bit quantization usually calls for caution even in ordinary inference, while 8-bit is generally considered to have "little practical impact." That said, it's reasonable to expect some impact when precision in token prediction specifically matters.

Here's what we found, looking at acceptance and mean length:

Turn Acceptance: kv-cache f16 Acceptance: kv-cache q8 mean len: kv-cache f16 mean len: kv-cache q8
Turn 1 60.97% 55.96% 3.44 3.24
Turn 2 57.34% 56.16% 3.3 3.25
Turn 3 78.06% 74.88% 4.12 4

Table 9: Acceptance / mean-length comparison by KV-cache precision

The difference isn't large. Let's turn to raw output speed next.

Figure 7: Generation throughput over time by KV-cache precision
Figure 7: Generation throughput over time by KV-cache precision

A slight offset in where each turn starts — especially pronounced around turn 3 — makes the data around the 5,000-token mark a bit noisy. Early on the impact looks minor, but it appears to compound somewhat as the run goes on.

To check this more carefully, let's look at each turn individually.

Figure 8: Throughput distribution broken down by turn (Turn 1)
Figure 8: Throughput distribution broken down by turn (Turn 2)
Figure 8: Throughput distribution broken down by turn (Turn 3)
Figure 8: Throughput distribution broken down by turn

There's a roughly 3-4 tps gap across the board.

Acceptance and mean length were both close between the two settings, but token and character counts can't be subdivided below whole numbers, which is likely what's driving that gap.

The takeaway: when using token prediction, it's best to keep the KV cache at f16 (the default) wherever possible — though since the gap doesn't appear to widen much turn over turn, it's worth weighing against the extra cache memory it costs.

Does It Hold Up for Everyday Use?

We've made the case that DFlash-2 is genuinely useful for coding-agent workloads — but is it actually "not worth enabling" for more ordinary, Deep-Research-style everyday use? We tried it out.

Here's a sample of the questions I actually asked over the course of one day, along with the relationship between token count and throughput.

I use Muse Glimmer to speed up research when I'm writing articles — I like how diligently it digs into things. Some representative examples:

Table 10: A sample of everyday questions asked in one day
Table 10: A sample of everyday questions asked in one day

These are fairly ordinary day-to-day questions. Some are research for an article; others are just things that came up in day-to-day life around that time.

Here's the relationship between token count and output speed for that session:

Figure 9: Distribution of token count vs. output speed (everyday use)
Figure 9: Distribution of token count vs. output speed (everyday use)

Short exchanges are noticeably fast, and speed tapers off as length increases — but it never dropped below the no-MTP baseline (roughly 12 tps, the red dashed line).

Since turn count isn't factored in here, the longer exchanges are more likely to be single-turn, so it's hard to draw a firm conclusion. Still, for everyday use it's clearly a net positive, and as of August 2026 it's become my main LLM day-to-day.4

Conclusion

This time we covered DFlash-2, from the little-known startup inco.ai.

Through testing, we came away with the following:

  • DFlash-2 adds a consistency-checking mechanism on top of the existing DFlash design, aimed at more reliable token prediction.
  • Under normal conditions (single-turn, Japanese-heavy), the impact is modest.
  • In multi-turn, code-heavy settings — i.e., coding-agent use — it becomes a genuinely powerful tool.
    • DFlash-2 shows a clear improvement over the original DFlash here.
  • A couple of things to keep in mind when using token prediction, which further improve results:
    • Measure the realistic draft length rather than assuming one.
    • Don't quantize the KV cache.

Causal models still have real advantages — the act of linking tokens sequentially reinforces the direction of generation and helps keep output coherent and high-quality — so we expect them to remain the mainstream approach for the foreseeable future.

That makes techniques like this one, which speed up output while preserving that property, worth continuing to watch.

As AI practitioners, we think the job going forward is to keep tracking techniques like this with a full picture in view, not just the surface-level claims.

References

DFlash 2: Keep Drafting Parallel
inco.ai
https://inco.ai/blog/dflash2/

DFlash: Block Diffusion for Flash Speculative Decoding
Jian Chen, Yesheng Liang, Zhijian Liu
https://arxiv.org/pdf/2602.06036

Accelerating Large Language Model Decoding with Speculative Sampling
DeepMind:
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre and John Jumper
https://arxiv.org/pdf/2302.01318

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI
https://arxiv.org/pdf/2606.19348

[Speculative decoding] feat: add DFlash support5
https://github.com/ggml-org/llama.cpp/pull/22105


  1. Indeed, compared against Gemma-4-Assistant, we saw a large gap in acceptance values, which we confirmed was reflected directly in the difference in output speed. ↩

  2. It did in fact land shortly after, on August 28, 2026 — officially supported from version 0.3.0-dev, build 10658, commit b10f9ca onward. To stay consistent with the figures measured for this piece, we've kept the write-up based on the pre-release build used at test time. ↩

  3. For more detail on this setup, see our earlier piece, "Using GPUSOROBAN" (https://www.bluecore.net/archives/242). ↩

  4. We also tried Qwen3.8-27B, but it tends to spend too long thinking, which held us back from using it as much day-to-day — even though its raw reasoning quality is excellent. ↩

  5. That PR thread also documents the steps for a DFlash implementation on the Qwen3.5 series, worth a look if you're interested. ↩

Top comments (2)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The six-token run is a useful counterexample to optimizing acceptance length alone: the mean length increases while overall generation throughput falls. I would expose the convention behind mean len, especially because the four-token configuration reports 4.12 in turn three.

If that metric includes the target's own next token, it is not the number of accepted draft tokens and should not be used directly to choose the next draft maximum. Recording committed tokens, drafting time and verification time per round would make the proposed tuning rule easier to validate. The best draft length should maximize committed tokens per wall-clock time for that workload.

Collapse
 
oooocean66 profile image
oooocean66 •

Hi Ahmet, thanks for the close read — this is a fair catch.

You’re right about the convention: mean length here is defined as 1 (the target model’s own guaranteed token) plus the number of accepted draft tokens per verification round. So for a draft-n-max of 4, the practical ceiling is 5, not 4 — the “past the configured maximum of 4” phrasing in the post was imprecise and we’re correcting it to just report the 4.12 figure against that actual ceiling. We’ll update the article text itself to reflect this.

That also means you’re right that it shouldn’t be used on its own to pick the next draft-max — it conflates the bonus token with genuinely accepted draft tokens, so a config can look like it’s “improving” while real draft-token yield is flat or falling, which is exactly what the six-token case shows.

Good suggestion on logging committed tokens, drafting time, and verification time per round separately — that’s the right level of granularity to actually validate a tuning rule, and committed-tokens-per-wall-clock is a better objective than mean length alone. We’ll fold that into how we evaluate draft-length choices going forward. Appreciate you flagging it.