DEV Community

oooocean66
oooocean66

Posted on

Qwen3.8-27B: An Architecture and Hands-On Look at Alibaba's Compact Dense Model

In August 2026, less than a month after Moonshot AI's Kimi-K3 surprised the world, Alibaba Group released "Qwen3.8-2.4T-A95B," an enormous 2.4-trillion-parameter model that once again shook the local-LLM community. Three days later, a smaller companion model quietly appeared: "Qwen3.8-27B," the model covered in this article.

It's released under the Apache-2.0 license, free for commercial use, and as a dense model it puts its full parameter count to work rather than routing through sparse experts. That combination raises a natural question: does this 27B model hold up against its much larger sibling? We look at its architecture and run hands-on benchmarks, comparing it along the way with Muse Glimmer 30B, the model we reviewed previously.

The Arrival of Qwen3.8

The Impact of Qwen3.8-2.4T-A95B

On August 12, 2026 — roughly a month after Moonshot AI's Kimi-K3 surprised the world — China's Alibaba Group published Qwen3.8-2.4T-A95B on Hugging Face. At 2.4 trillion parameters, it's a genuinely enormous model, well beyond what most people can run locally, yet it still generated considerable buzz in the local-LLM community.

Its reasoning benchmarks came close to Kimi K3's, putting it within reach of the major American frontier models.

Some users ran 1-bit or 2-bit quantized versions on their own hardware and confirmed that the numbers weren't hype1 — reports of the model's real-world performance spread quickly.

How Does Qwen3.8-27B Measure Up?

Three days later, at midnight JST on August 15, 2026, the smaller Qwen3.8-27B was released.

One important distinction is licensing. The larger Qwen3.8-2.4T-A95B ships under Alibaba's own Qwen3.8-Max license, which imposes certain conditions on commercial use — a company large enough to offer that model as a service would likely need a separate commercial licensing agreement.

Qwen3.8-27B, by contrast, uses a plain Apache-2.0 license, so it can be used commercially at no cost. And as a dense model — as its name suggests — it puts its entire parameter count to work on every token, which is the design typically associated with higher per-parameter capability.

Qwen3.8-2.4T-A95B made a strong case for its own capability. The question is how much of that capability carries over to this smaller sibling.

We ran our own hands-on verification to dig into that question.

Qwen3.8-27B's Architecture

Architecturally, Qwen3.8 carries the Qwen3.5/Qwen3.6 design forward essentially unchanged. The structure is shown below.

Qwen3.8-27B's architecture
Figure 1: Qwen3.8-27B's architecture

Looking at the config.json published on Hugging Face makes that continuity even clearer:

{
  "architectures": [
    "Qwen3_5ForConditionalGeneration"
  ],
  "image_token_id": 248056,
  "language_model_only": false,
  "model_type": "qwen3_5",
  "text_config": {
Enter fullscreen mode Exit fullscreen mode

Beyond that, the basic specs are as follows:

  • Max context size: 256k tokens (limited to 64k in our test environment)
  • 64 blocks total (GDN:GA = 3:1, 4 layers × 16 groups)
  • Hidden size: 5,120
  • FFN type: Dense
    • Activation function: SiLU
    • Activation dimension: 17,408

An Attention Structure That's Actually a Minority Approach

Qwen3.8's attention is a hybrid of two mechanisms — Gated Delta Network and Gated Attention (effectively GQA) — which is quite different from the "local + global" pattern that's become the popular default lately.

Gated Delta Network handles two jobs — state updates and information compression — and makes up 75% of the attention layers. The remaining 25% are Gated Attention layers, internally structured as Group Query Attention, responsible for global aggregation and precise retrieval.

This layout was first adopted in Qwen3-Next-80B and has carried through Qwen3.5 and Qwen3.6.

Rather than "local + global," it's probably more accurate to describe this as "compression + aggregation." As far as we're aware, Qwen is currently the only major model family using this approach.

The Vision Encoder Is Still SigLIP2-Based

Qwen3.8's vision encoder continues to build on the SigLIP2-based architecture (SigLIP2-SO-400M). Additional fine-tuning presumably went into it, but there's been no major structural change.

That's a notably different approach from Muse Glimmer's custom vision encoder, which we covered previously.

Running It Ourselves

For this test, we used a GPUSOROBAN High-Speed Computing instance equipped with an NVIDIA RTX A4000.

Overall Setup

The setup is as follows: we open an SSH tunnel to an access server, then reach the target instance through that tunnel.2 llama.cpp's service port is relayed to localhost via port 8001 on both ends.

[ Home ]                    [ HighReso : GPUSOROBAN ]
+---------------+           +----------------+     +------------------+
|    Client     |           | Access Server  |     | Target Instance  |
|  > 8001/tcp   |--(SSH)--->|    (relay)     |---->|  8001/tcp        |
|               |           |                |     |  (llama.cpp)     |
+---------------+           +----------------+     +------------------+
Enter fullscreen mode Exit fullscreen mode

Figure 2: Overview of the connection setup

Instance Used

The GPU used for this test is an NVIDIA RTX A4000. The instance specifications are as follows.

Item Spec
Instance type s16-1-a-standard-ubs24-v
GPU NVIDIA RTX A4000
GPU memory GDDR6 16GiB
GPU memory bandwidth 448.0 GB/s
FP32 compute 19.17 TFLOPS
BF16 compute 38.34 TFLOPS
INT8 compute 153.4 TOPS
INT4 compute 306.7 TOPS
vCPU 11 cores
System memory 50 GiB
Storage Persistent 100 GiB
CUDA version 13.2
NVIDIA driver version 580
OS Ubuntu 24.04 Server

Table 1: Instance specifications

Model Files Used

Even a 4-bit quantized version of Qwen3.8-27B exceeds 16GiB. Factoring in KV cache usage as well, we used a 2-bit quantized model optimized with Unsloth Dynamic 2.0.

We downloaded the following files from Unsloth's Qwen3.8-27B-GGUF repository:

File Size Description
Qwen3.8-27B-UD-Q2_K_XL.gguf 12 GB UD2.0, 2-bit quantized model
mmproj-bf16.gguf 889 MB Quantized vision encoder

Table 2: Files used with llama.cpp

We used llama.cpp version 0.1.0-dev (build 10450, commit ece963f41).

Since the architecture hasn't changed since Qwen3.5, being able to run it on an existing llama.cpp build without waiting for new support is arguably an advantage in its own right.

Verification: Standard Mode

First, we ran text inference using the main model on its own.

Launch Command

We used the following command to launch the web frontend for testing.

For this run, we kept the KV cache in FP16 rather than quantizing it to 8-bit, but due to memory constraints we lowered the max context size to 65,536. We also raised the log verbosity from the default of 3 to 4 to check memory usage.

build/bin/llama-server --model ./models/Qwen3.8-27B-UD-Q2_K_XL.gguf -t 12 -np 1 --prio 2 --temp 0.6 --top-p 0.95 --top-k 20 --port 8001 --host 0.0.0.0 --fit off --no-warmup --no-cache-prompt -fa on --cache-ram 0 -c 65536 --reasoning on --reasoning-effort xhigh --log-verbosity 4
Enter fullscreen mode Exit fullscreen mode

Reasoning Effort Is Now Configurable

One recent improvement in llama.cpp that stood out to us is the addition of a reasoning_effort parameter.

--reasoning-effort LEVEL   reasoning effort level given to the chat template:
                           'default' to keep the template default, or a level
                           such as 'minimal', 'low', 'medium', 'high', 'xhigh'
                           or 'max' (default: default)
                           (env: LLAMA_ARG_REASONING_EFFORT)
Enter fullscreen mode Exit fullscreen mode

Table 3: The reasoning_effort option, as described in llama-server's own help text

Which levels are actually usable depends on Qwen3.8-27B's own chat_template.jinja. Checking that template, the following levels appear to be supported:

  • xhigh (default)
  • medium
  • low

Setting these appears to inject the following text into the system prompt:

Level Injected text
xhigh Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.
medium (no system prompt injected)
low Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.

Table 4: reasoning_effort levels and their injected prompt text

So rather than setting an explicit token budget, it appears to control reasoning depth simply by appending instructions to the system prompt.

Memory Usage: KV Cache Size Is the Sticking Point

Memory usage came out as follows. As is generally true across the Qwen series, the KV cache footprint is large. We initially tried setting the max context to 131,072, but that produced an OOM error, so we halved it. With the ceiling set at 65,536, the KV cache came out to 4,096 MiB.

Device CPU CUDA0 Total
Model Xeon E5-2690v4 RTX-A4000 —
Weight data (Q2_K-quant_XL) 397.85 9,567.89 9,965.74
KV cache size (f16) — 4,096.00 4,096.00
RS buffer size (f32) — 149.62 149.62
Gated Delta Net compute buffer 186.02 84.02 270.04
Total 583.87 13,897.53 14,481.40

Table 5: Memory usage in standard mode (main model only, units: MiB)

In other words, setting the max at 131,072 tokens in FP16 mode would require 8,192 MiB, and setting it to the model's full 262,144-token context would require 16,384 MB.

The cause is that, compared with other models, Qwen3.8-27B has more Global (GQA) layers, and each of those layers is configured with more heads and larger dimensions.

Model Muse Glimmer 30B Gemma-4-31B Qwen3.8-27B
Layer structure SWA+Global(3:1) SWA+Global(5:1) GDN+Global(3:1)
Total layers 52 60 64
Aggregation layers 13 10 16
Aggregation layer KV design 2 heads x 128 dim 4 heads x 512 dim 4 heads x 256 dim
KV cache size (128k, f16) 1,742.00 MiB 5,920.00 MiB 8,192.00 MiB
Design intent Aims for lightweight inference, pushing memory reduction to the limit Balances advanced reasoning with memory reduction Aims squarely at advanced reasoning

Table 6: Aggregation-layer configuration and KV design across hybrid models

Gated Delta Network is, on its own, a highly efficient architecture compared with the "local + global" pattern — the layers handling information compression require essentially zero KV cache. Under otherwise identical conditions, a "local + global" design will always require a larger cache.

But Global layers appear very frequently here, the model has more layers overall than its peers, and the aggregation-layer KV design is comparatively rich — and together those factors are what push Qwen3.8-27B's KV cache size up.

That points clearly to a design goal that prioritizes reliable, deeply reasoned answers over lightweight inference.

Output

We asked the following question and examined the response.

Given the recent Kumamoto earthquake, this is a question we've been asking various models out of curiosity.3

Describe the characteristics of the ground/soil conditions in Kyushu.

The response came out as follows. Our first impression was that it looked surprisingly well-informed.

Kyushu's ground conditions are geologically, topographically, and hydrologically
complex, and their strength and character vary widely from place to place —
that variability is the single defining feature. As construction ground, the
characteristics can be summarized as follows.

1. Geological and ground diversity
Kyushu sits on a volcanic island arc, so volcanic ground, sedimentary ground,
granite ground, weathered soil, and river deposits are all mixed together over
short distances.
- Northern Kyushu (Fukuoka, Saga, Nagasaki, etc.): sedimentary basins and
  granite, with serpentinite bedrock visible in some areas.
- Central Kyushu (Kumamoto, etc.): the Aso caldera, with well-developed
  volcanic ash, pumice, lava, and river deposits.
- Southern Kyushu (Oita, Miyazaki, Kagoshima, etc.): a mix of volcanic ground,
  granite, and sedimentary plains.
As a result, ground conditions differ significantly even within Kyushu itself.

2. Plains are largely soft, recent deposits
Plains and river floodplains around Fukuoka, Saga, Nagasaki, Kumamoto, Oita,
Miyazaki, and Kagoshima often have thick layers of recent deposits — sand,
silt, and clay. Key characteristics: high water content / low shear strength /
prone to compression and settlement / prone to liquefaction during earthquakes
/ prone to differential settlement and subsidence. Soft-ground countermeasures
are especially necessary along rivers and coastal plains.

3. Volcanic ground is widespread
Kyushu has numerous volcanoes, and pumice, volcanic ash, lava, and igneous
rock all contribute to its ground composition. Pumice and volcanic ash often
have low strength, high permeability, and high compressibility, can liquefy
when wet or saturated, and are prone to settlement and changes in earth
pressure. This kind of volcanic-derived ground matters particularly around
Aso, Sakurajima, Kirishima, Kuju, Unzen, and Ibusuki. Lava and igneous rock,
on the other hand, tend to have high strength and can serve as good bearing
layers for foundations, but weathering, cracking, cavities, and erosion make
them non-uniform — the surface may be soft while deeper layers remain strong.
Because volcanic ground places "strong bedrock" right next to "soft volcanic
deposits," ground investigation is especially important here.

4. Mountain and slope areas carry high landslide/collapse risk
Kyushu's mountains are steep, and their slopes are made up of weathered rock,
spoil, volcanic soil, and river deposits. With such steep gradients,
landslides and collapses are easily triggered by rain, typhoons, or
earthquakes, and ground strength varies considerably with the degree of
weathering. Landslide and sediment-disaster risk rises sharply during heavy
rain or seismic events.

5. Active faults and earthquakes have a major effect on the ground
Parts of Kyushu lie near plate boundaries and active fault zones, so ground
motion during earthquakes can be relatively large. Earthquakes can trigger
liquefaction, landslides, slope failure, ground settlement and differential
settlement, and damage to structural foundations. This means seismic design
needs to account for the ground's seismic response, not just the structure
itself.

6. Groundwater and hydrological conditions are also complex
Ground conditions across Kyushu can involve shallow groundwater in the
plains, fissure water in volcanic rock and granite, and saltwater intrusion
along the coast. Construction needs to account for spring water, rising
groundwater levels, strength loss in soft ground, and drainage/water-sealing
during ground improvement.

7. Construction considerations
Because ground conditions vary so much by location across Kyushu,
site-specific investigation is essential rather than relying on general
assessments. Soft ground calls for pile foundations or ground improvement;
lava and bedrock require confirming the bearing layer; volcanic ground needs
liquefaction and settlement countermeasures; slopes need landslide prevention
and anchoring; and ground assessment should account for earthquakes and heavy
rain, alongside a comprehensive look at landslide, liquefaction, and
differential-settlement risk.

Summary
Kyushu's ground is characterized by a combination of soft recent deposits,
low-strength volcanic-derived soil, strong lava and bedrock, and active
faults, earthquakes, and landslides. Because of this, ground strength in
Kyushu can't be judged in general terms — it's important to check each
site's geology, topography, hydrology, and earthquake/rainfall risk
individually.
Enter fullscreen mode Exit fullscreen mode

We had a DeepSearch agent built on GPT-5.6-Luna (powered by Dify) evaluate this response. The results were as follows (we've trimmed a few overly long passages).

We ran the same question and evaluation against Muse Glimmer 30B previously, and the level of the critique here was noticeably higher — this response appears to be operating on a different level.

GPT-5.6-Luna judged the overall content to be sound, and its feedback moved beyond basic correctness into refinements for readability and presentation — comments aimed at polishing an already-solid answer rather than fixing a flawed one. That's a clear sign it was reading the text at a higher level.

In this case, it's fair to say this response clearly surpasses Muse Glimmer's.

Overall assessment
As a general-audience overview explaining how much Kyushu's ground varies by
region, this text's overall direction is sound. It covers volcanic ground,
alluvial plains, mountain slopes, earthquakes, and groundwater, and the
structure is easy to follow. That said, using it for construction practice or
specialist education would require correcting a few terms and overly
definitive statements.

As a rough guide: accuracy 6.5/10, readability 8/10, suitability for
construction practice 6/10. The main issues aren't outright factual errors so
much as stating conditionally variable properties as if they were universal,
and grouping geological materials, ground characteristics, and disaster risk
into the same category.

Points needing correction: the phrase "the defining feature" is too
definitive; the description of plains treats their properties too uniformly;
the earthquake section should separate ground characteristics from disaster
hazards; the groundwater section needs qualification; "strong/weak ground"
needs to be broken down further; ground materials and disaster risk should be
kept separate; and some redundancy should be trimmed.

Final assessment: as an introductory, general-audience explanation of
Kyushu's ground, this text is a usable structure. However, phrases like
"recent deposits," "spoil," grouping Oita into southern Kyushu, "volcanic ash
has high permeability," and "can liquefy when wet" need correction or
qualification. Once revised, it's suitable as a general-audience explanation.
For construction planning or foundation-type selection, this text should not
be used as the basis for decisions — it needs to be combined with geological
maps, land-condition maps, liquefaction hazard maps, active-fault data,
boring surveys, standard penetration tests, and groundwater-level surveys.
Enter fullscreen mode Exit fullscreen mode

Performance

We compared this against Muse Glimmer 30B as well. Throughput came out almost identical, but token usage to reach an answer was nearly three times higher — and, correspondingly, so was the time it took.

[Qwen3.8-27B]
prompt eval time = 303.63 ms / 64 tokens ( 4.74 ms per token, 210.78 tokens per second)
       eval time = 248375.97 ms / 5771 tokens (43.05 ms per token, 23.23 tokens per second)
      total time = 248679.60 ms / 5835 tokens

[Muse Glimmer]
prompt eval time = 262.18 ms / 68 tokens ( 3.86 ms per token, 259.37 tokens per second)
       eval time = 82874.28 ms / 1866 tokens (44.41 ms per token, 22.52 tokens per second)
      total time = 83136.46 ms / 1934 tokens
Enter fullscreen mode Exit fullscreen mode

Both runs used a high reasoning-effort setting, and the pattern that's generally true of the Qwen series — taking a long time to reach an answer — still held here.

Verification: Using Vision

Qwen3.8 has been released as a multimodal model since earlier versions, and this one naturally ships with a vision encoder too. In llama.cpp, attaching the vision-encoder model lets you feed it images, as shown below.

This is roughly where things start to get tight — if you want more context length, you'll need to quantize the KV cache.

Command Line

build/bin/llama-server --model ./models/Qwen3.8-27B-UD-Q2_K_XL.gguf -t 12 -np 1 --prio 2 --temp 0.6 --top-p 0.95 --top-k 20 --port 8001 --host 0.0.0.0 --fit off --no-warmup --no-cache-prompt -fa on --cache-ram 0 -c 65535 --reasoning on --reasoning_effort xhigh --log-verbosity 4 --mmproj models/mmproj-BF16.gguf
Enter fullscreen mode Exit fullscreen mode

Memory Usage

Memory usage came out as follows, showing an increase of roughly 1,336 MiB.

Device CPU CUDA0 Total
Model Xeon E5-2690v4 RTX-A4000 —
Weight data (Q2_K-quant_XL) 397.85 9,567.89 9,965.74
KV cache size (f16) — 4,096.00 4,096.00
RS buffer size (f32) — 149.62 149.62
Gated Delta Net compute buffer 186.02 84.02 270.04
Vision weight data — 887.99 887.99
Vision metadata — 0.12 0.12
Total 583.87 14,785.64 15,369.51

Table 7: Memory usage with the vision encoder enabled (units: MiB)

Prompt and Output

Rather than text-only inference, we tested the model's analytical ability by attaching an image and asking it questions about that image.

For this, we had the model analyze the Muse Glimmer architecture diagram from our previous article, reusing the same image4, and asked it to "analyze this in detail and describe everything written in it."

Note: since this article's author writes in Japanese, this test was conducted using a Japanese-language diagram as the input image. The model's response below — translated into English — conveys the diagram's content, so the image itself is kept as-is here rather than redrawn in English.

The image fed into the Vision Encoder
Figure 4: The image fed into the Vision Encoder

The result was as follows. The passages noted below are places where the model's interpretation was mistaken.

MUSE GLIMMER 30B Architecture Slide: Detailed Analysis

1. Overall composition of the image
Under the title "Reference: MUSE GLIMMER 30B Structure" (dark navy header
bar), the slide has three layers: top text (left = basic model specs, right
= definitions of the SA/FA attention mechanisms), a central callout
(supplementary explanation of block composition), and a bottom flow diagram
(input -> embedding -> Transformer block group -> output). Its purpose is to
visualize the architecture of a 29.6B-parameter dense multimodal LLM and the
design philosophy of its core SA/FA hybrid attention.

2-5. Analysis of parameters, attention mechanism, flow diagram, and callout
(Dense 29.6B, hidden size 6,656, SwiGLU, RoPE theta=50,000 not applied to FA,
max context 128k, DFlash-based drafter, and the respective roles of SA/FA
are laid out in detail exactly as shown in the diagram)

6. Consistency and notes (analytical observations)
Block-count inconsistency: the diagram label reads "52-block configuration"
while the callout says "50 blocks." 52 = 4x13 is consistent as a count of
groups; 50 = 12x4+2 is consistent from the text's own framing. This gap most
likely comes from counting Glimmer Vision's two layers separately — it's
natural to read the diagram's "52" as the total including vision layers, and
the callout's "50" as the text-decoder block count alone. Unifying the
definition within the slide would be preferable. Possible typo: the FA
description "no positional information" appears to be a misprint of "no
positional encoding."

7. Summary
This slide explains that MUSE GLIMMER 30B, built on a dense 29.6B-parameter
model with a 128k context, repeats SA (local, low-cost) and FA (global,
high-quality) in a 3:1 ratio, concentrates positional encoding on SA layers
while removing it from FA layers, and achieves a hybrid attention design
that balances compute efficiency with reasoning performance. Vision-input
fusion, GQA + sigmoid gating, and DFlash/MTP-based inference support are the
other key design elements, with the only consistency issue being the 50/52
block-count discrepancy between the diagram and the text.
Enter fullscreen mode Exit fullscreen mode

The response came back roughly twice as long as what Muse Glimmer 30B produced for the same task previously. On top of that, the summary was nearly flawless, and it even called out an inconsistency in how we'd labeled our own diagram.

The vision encoder is still built on SigLIP2, but additional fine-tuning may well be reflected here — the way it reasons from what it extracts felt noticeably more capable.

A Look Inside the Thinking Section

Qwen models are well known for a distinctly self-doubting internal voice during their thinking phase — the model seems to talk itself into second-guessing everything right before it's about to answer, wanting to double-check just one more thing or start over from scratch, dragging out its reasoning in the process. We checked whether that pattern still holds here.

We've summarized the actual thinking-section output for the Kyushu-ground question in the appendix at the end of this article. Here, we'll first go over the key points from having Google Gemini 3.7 Flash evaluate that output.

According to Gemini's analysis, Qwen3.8-27B's thinking pattern breaks down into five main traits. First, a separation between the thinking language (English) and the output language (Japanese) — a multilingual-processing behavior where internal reasoning happens in English but the final answer comes out in Japanese. Second, multi-angle context inference and intent-filling, where it anticipates the underlying purpose behind the question (construction practice, exam prep, and so on). Third, exhaustive knowledge brainstorming and categorization, where it lists out related keywords comprehensively and then organizes them by region and geology. Fourth, real-time self-correction and fact-checking — starting to write "Miyagi" and immediately correcting it to "Miyazaki," or avoiding overgeneralizing where serpentinite is distributed, repeatedly double-checking small factual details. And fifth, iterative simulation of output structure and formatting — trying out multiple drafts of whether to use a table or bullet points, and in what order to place the headings.

The important point here is that the overtly negative emotional tone common through Qwen3.5 doesn't show up — the evaluation describes the process as proceeding "extremely logically." Since there's been no architectural change at all, this improvement appears to come down to training methodology.

So what exactly changed in how this model was trained? A clue seems to be in the following post.

Qwen3.8-Max: A New Bar for Coding and Cowork
(https://qwen.ai/blog?id=qwen3.8)

This post describes the flagship Qwen3.8-2.4T-A95B model, but given that Qwen3.8-27B is a smaller sibling, it likely went through similar training. In the "Work" section, the post explains that scaling up real-world RL systems (continuously scaling tasks, workspaces, and toolchains along independent axes), a general-purpose reward system capable of internalizing diverse forms of verification, and an online data balancer were combined to lift task-completion capability uniformly across multiple major toolchains.

That distinctive self-doubting quality, present since Qwen3, doesn't actually originate with Qwen — it traces back to a training method DeepSeek-AI calls "Long CoT Cold Start"5, applied to DeepSeek-R1 to boost reasoning ability in cases requiring many turns of back-and-forth. The "cold start" in the name reportedly refers to the tendency to pause mid-reasoning and start over from scratch.

Rather than relying solely on QA-style datasets, this method trains the trial-and-error process directly by drawing on data from sources like programming forums such as Reddit — and that data is reportedly full of human emotional expression, which is said to be what feeds into the model's self-doubting tone. This trait shows up not just in Qwen but broadly across DeepSeek, its originator, and other Chinese LLMs.

Qwen3.8 appears to have revisited various stages of its training pipeline, and we found a related paper, "Unified Data Selection for LLM Reasoning"6, which we looked into.

The paper proposes a new, training-free, and extremely lightweight data-evaluation metric called High-Entropy Sum (HES), and applies it across the SFT (supervised fine-tuning), RFT (rejection-sampling fine-tuning), and RL (reinforcement learning) stages to measure its effect.

With conventional training (based on average entropy), the important decision points in the reasoning process get diluted by a large volume of low-entropy tokens — boilerplate text, simple arithmetic logic — so the model ends up training without ever distinguishing good solutions from bad ones. Using HES instead, the paper reports, lets the optimized model learn high-quality solutions much more efficiently, with a clearly measurable training benefit.

Training a model at the scale of Qwen3.8-Max presumably involves a huge volume of answer data with substantial reasoning steps built in. Our guess is that applying the HES concept to that data made more effective solution paths — and by extension, more logical reasoning routes — easier to select, suppressing emotional noise and, as a result, producing more advanced reasoning ability.

It's also worth noting that issues with Long-CoT-style training have already been pointed out elsewhere, such as in the paper "Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt"7, and it's likely that some countermeasure along those lines was applied here as well.

Some people still argue that Chinese models simply distill capability out of frontier models. Based on what we're seeing here, that framing doesn't quite fit Qwen3.8-27B — distillation alone doesn't look like the whole story behind its capability gains. If anything, the Qwen team appears to actively track and adopt a wide range of techniques, and that appetite for new ideas seems to be a real part of what makes this model work.

This shift in training methodology may also have changed how the model arrives at its answers — the reasoning path itself looks noticeably different from before. If you're doing harness engineering around this model, some of your existing prompting approaches may no longer work as expected, and revisiting them may become necessary if something that used to work stops working.

As AI models continue to evolve, the way they process information keeps changing along with them — and it may be up to us to keep updating our own approach in step.

Benchmarks: How the Wider World Sees It

Below is the benchmark comparison published by Alibaba's Qwen team. It shows results on par with — or better than — Anthropic's Claude Opus 4.6 in Max Reasoning mode. Based on what we've seen in our own hands-on testing, that claim doesn't look far-fetched.

Category Benchmark Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Muse Glimmer-30B Opus4.6 Max
Coding Agentic terminal coding (Terminal Bench 2.1) 73 63.4 64 51.7 78.2
Coding Agentic coding (SWE-bench Pro) 61.7 53.5 57.6 51.2 53.4
Coding Repo-level code generation (NL2Repo-Bench) 42.3 36.2 41.1 -- 47.6
Coding Agentic coding (DeepSWE 1.1) 42.2 13.3 14.2 -- --
Coding Software engineering (QwenSWEBench) 79 49.3 59.2 -- 63.8
Agent Long-horizon office work (CoWorkBench) 70.7 61 65.1 -- 68.2
Agent Professional job tasks (JobBench) 33.4 21.8 27.6 -- --
Agent Frontier agentic tasks Pass@1 (Agents' Last Exam) 20.4 10.6 13.2 -- --
Agent Frontier agentic tasks Score (Agents' Last Exam) 42.9 27.3 33.6 -- --
General Instruction following (IFBench) 79.5 69.1 79.1 77 62.5
General Scientific reasoning (GPQA Diamond) 89.2 87.8 90.3 83.5 91.3
General Multidisciplinary reasoning (HLE) 30.8 24 34.7 22 40
General Competitive coding (LiveCodeBench v6) 90.3 83.9 89.6 -- 88.8

What about third-party evaluations? Here's Artificial Analysis's Intelligence Index chart.

Artificial Analysis Intelligence Index chart
Figure 5: Artificial Analysis Intelligence Index chart

With thinking mode on, it posted a striking score of 52 — not just surpassing Claude Opus 4.6, but matching GPT-5.6 Luna, likely OpenAI's most widely used model at the moment. Seeing the score climb this far came as a genuine surprise to us.

That said, while this behavior works well for Deep Research, it also means the model is likely to be slow for general-purpose use that mixes in casual conversation, and given how many new approaches went into it, some retuning of middleware and instructions is clearly going to be needed.

In that sense, Meta's Muse Glimmer 30B, covered earlier, and Google DeepMind's longer-established Gemma-4-31B-it, still have plenty to offer. Their thinking phase is far shorter than Qwen3.8's, so they can produce answers faster.

In particular, Muse Glimmer 30B and Gemma-4 both benefit heavily from sliding-window memory savings, making it relatively easy to extend their context length. Muse Glimmer 30B's vision encoder is also appealing for its resolution. Given these different strengths, rather than defaulting to a single model for everything, what seems to matter going forward is the skill of choosing and applying the right model for the task.

Conclusion

This article covered Qwen3.8-27B, the compact model the local-LLM community had been waiting for. Its capability is no exaggeration — it delivered performance beyond what we expected. The benchmark results published on Artificial Analysis in particular likely came as a considerable surprise to a lot of people.8

The first use case that comes to mind is Deep Research. Paired with a search engine, it can dig into deep content without difficulty, and its resistance to falling into wasteful reasoning loops makes it a reassuring tool for search-augmented use.

Architecturally, it carries forward Qwen3.5's Gated Delta Network-based hybrid design, and being able to run it as-is without updating your inference engine is a significant point in its favor.

That said, given how firmly this design leans toward raw capability, its tendency toward a large KV cache footprint is a genuine drawback. Keeping the max context length very long requires a correspondingly capable GPU, and that memory overhead is likely to be a real headache for local-LLM users.

On the training side, the shift away from the traditional Long CoT Cold Start method has also left some users unsure what to make of it. Figuring out how to adapt to this change, and what prompt adjustments it calls for, will likely be the key going forward.

As findings from this new approach get shared through arXiv and elsewhere, it will likely spread gradually to other models. Efforts like this could act as a trigger that meaningfully lifts capability across both commercial frontier models and open-weight models alike.

New technical advances don't just show up in architecture and tokenizer specs — they also show up inside the thinking section, a part of the process that's normally hidden from view. It's worth paying attention to that usually-overlooked area and watching how model behavior changes there.

References

Qwen/Qwen3.8-27B -- Hugging Face
https://huggingface.co/Qwen/Qwen3.8-27B

unsloth/Qwen3.8-27B-GGUF - Hugging Face
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

OpenLM-AI Qwen3.8
https://openlm.ai/qwen3.8/

Qwen3.8-Max: A New Bar for Coding and Cowork
https://qwen.ai/blog?id=qwen3.8

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
https://arxiv.org/pdf/2501.12948

Unified Data Selection for LLM Reasoning
https://arxiv.org/pdf/2605.22389

Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt
https://arxiv.org/pdf/2505.23480

Appendix

Expected KV Cache Sizes

Qwen3.8-27B requires the following KV cache sizes. Note that supporting long context requires these amounts of VRAM on top of the weight data and vision-encoder data.

Max Context Size K:f16, V:f16 K:q8_0, V:q8_0 K:q8_0, V:turbo39
65,536 4,096 MB 2,048 MB 1,488 MB
131,072 8,192 MB 4,096 MB 2,976 MB
262,144 16,384 MB 8,192 MB 5,952 MB

Table 8: Expected KV cache sizes

Thinking-Section Content for the Kyushu Ground Question (Summary)

Summarizing the actual thinking-section content for the "describe the characteristics of Kyushu's ground" question covered in the main text (the original was output in English), the trial-and-error process went roughly as follows. The full text is extremely long, so only the key points are excerpted here.

It begins by questioning itself over whether "ground" refers to geological bedrock or to civil-engineering/construction-style ground conditions, and considers the question's likely intent from multiple angles — an exam question, construction practice, or general knowledge. It then exhaustively lists out the elements worth covering: plate tectonics, volcanic activity, sedimentary layers, liquefaction risk, and regional differences (north/central/south).

Along the way, it repeatedly double-checks small factual details. For instance, it starts to write "Miyagi," catches that it's not a Kyushu prefecture, and immediately corrects it to "Miyazaki"; it also reconsiders where serpentinite is actually distributed in Kyushu and settles on the more limited phrasing "in some areas" — self-checking specific place names and geological terms at every step.

It also repeatedly adjusts its tone — not overgeneralizing, not overusing jargon, not making the answer too long — and trials multiple structural options: bullet points, paragraphs, or a table. It eventually settles on a draft that includes a table with "aspect | characteristic" columns, then tries out several versions of the concluding sentence, fine-tuning the wording right up to the end.

The sheer volume of this trial-and-error is what substantiates the self-doubting quality mentioned earlier in the article. The actual output — the Kyushu-ground answer quoted in the main text — is concise and tidy, but the internal process behind it involves repeated self-questioning and backtracking. That gap between process and output was striking.


  1. Some users, working from the assumption that the model simply wouldn't fit on their hardware, went as far as building their own memory-reduction scheme and inference engine from scratch to verify it themselves. In 2-bit quantized mode, they reportedly recorded results well beyond what anyone expected. ↩

  2. For more detail on this setup, see our earlier article on using GPUSOROBAN. (https://www.bluecore.net/archives/242) ↩

  3. This article's author is based in Japan, so the model was prompted in Japanese and asked about a Japan-specific topic (the ground conditions behind the recent Kumamoto earthquake) — hence the choice of question here and elsewhere in this article. ↩

  4. This image uses an earlier version of the model diagram than the one shown previously in this article. ↩

  5. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (https://arxiv.org/pdf/2501.12948) ↩

  6. https://arxiv.org/pdf/2605.22389 ↩

  7. https://arxiv.org/pdf/2505.23480 ↩

  8. The people running these benchmarks may well have rubbed their eyes and re-verified the results at least once, wondering if something had gone wrong — that's roughly how long it took for Qwen3.8-27B's results to come out relative to what we'd expect. ↩

  9. Measured using a llama.cpp fork that supports TurboQuant. "turbo3" indicates TurboQuant 3-bit quantization. Note that this was tested on a 2-GPU setup, so some overhead is present. ↩

Top comments (0)