DEV Community

Cover image for Qwen-Image-2.1-Turbo Isn't 7B: The VRAM You Actually Need to Run It Locally
Alan West
Alan West

Posted on Originally published at blog.authon.dev

Qwen-Image-2.1-Turbo Isn't 7B: The VRAM You Actually Need to Run It Locally

Qwen-Image-2.1-Turbo gets called a "7B" model everywhere. Before you try to run it locally, know that the Hugging Face repo is 32.4 GB, the bundled text encoder is bigger than the image model, and the license rules out anything you get paid for.

Qwen released Qwen-Image-2.1-Turbo on 2026-10-09, 19 days after the base Qwen-Image-2.1. It's an accelerated checkpoint that generates and edits images in 8 denoising steps. The base model's default is 40. I'll cover what changed, how much VRAM it really needs, how to run it with diffusers and ComfyUI, and two config gotchas, one of which has already started a "poor results in comfy" thread.

What changed from Qwen-Image-2.1

Not much in the architecture. The model card says Turbo "uses the same 7B visual generation architecture" and loads with the same QwenImage21Pipeline. The differences:

  • 8 steps instead of the base default of 40 (GitHub README).
  • The schedule is baked in: model_index.json carries a fixed sample_sigmas list, 1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568.
  • CFG=1 by default, so there's one forward pass per step.
  • Features are unchanged: text-to-image, editing with reference images, RGBA transparency output, and the same 2K-level resolution presets (2048×2048 at 1:1, 2752×1536 at 16:9).

So you're doing a fifth of the denoising work. Qwen publishes no Turbo speed numbers or quality benchmarks, either on the card or in the README, so everything you'll read about Turbo being "just as good" comes from showcase images Qwen picked themselves.

For scale, here are base-model numbers from LightX2V's RTX 5090 table: 5.586 s at 1024×1024 and 30.333 s at 2048×2048, both at 40 steps with CFG disabled. Turbo should come in well under that, but I couldn't find a Turbo figure on NVIDIA hardware. The one real-hardware Turbo number I found comes from the author of a native Mac port (HF discussion #5): about 120 s at 1024² on an M1 Pro with 16 GB.

The "7B" number hides most of the download

My research machine has no GPU, so I couldn't time generations. I did read the safetensors headers of every component straight from the Hub (range requests, no full download) and added up the parameters:

Component Params BF16 file size
Transformer (DiT, 32 layers) 7.115B 14.23 GB
Text encoder (Qwen3-VL 8B) 8.767B 17.53 GB
VAE 0.338B 0.68 GB
Total 16.22B 32.44 GB

The "7B" is the denoiser alone. The full pipeline has 16.2B parameters, and over half of them are a vision-language model that encodes your prompt and reference images. That's also why VRAM estimates for this model are all over the place.

vLLM-Omni's recipe measured the base model, which has the same architecture, at 34.0 GB peak on a GB300 at 1024×1024 in BF16. That's more than a 32 GB RTX 5090 holds, unless you offload something.

How much VRAM you need at each precision

The text encoder runs once per prompt, while the DiT runs every step. So you can offload the encoder or quantize it hard and lose very little.

Here are weight totals for the practical combinations, using file sizes from Comfy-Org/Qwen-Image-2.1 and unsloth/Qwen-Image-2.1-Turbo-GGUF:

Setup DiT Text encoder VAE Weights total
Full BF16 (diffusers default) 14.23 GB 17.53 GB 0.68 GB 32.44 GB
ComfyUI INT8 DiT + INT8 encoder 7.26 GB 9.35 GB 0.68 GB 17.29 GB
ComfyUI INT8 DiT + W4A8 encoder 7.26 GB 6.31 GB 0.68 GB 14.25 GB
GGUF Q8_0 + W4A8 encoder 7.64 GB 6.31 GB 0.68 GB 14.63 GB
GGUF Q4_K_M + W4A8 encoder 4.20 GB 6.31 GB 0.68 GB 11.19 GB

These are weight sizes, not measured peaks. Activations come on top, and at native 2K they get big. ComfyUI and diffusers offload both swap models in and out, so the number that matters on a small card is roughly DiT plus activations.

My rough guide:

  • 40 GB+ cards: run everything in BF16 and keep it resident.
  • 24–32 GB: BF16 DiT with the text encoder offloaded, or the INT8 Comfy files.
  • 12–16 GB: GGUF Q4_K_M to Q8_0 for the DiT, a 4-bit encoder, and offloading.
  • Apple silicon, 16 GB: the Siliconed app (Swift/Metal) supports the full weights, Comfy's int8, or Q4_K_M.

On quant quality: Unsloth published LPIPS against the BF16 Turbo transformer (lower is better). Q8_0 scores 0.036, Q4_K_M 0.153, and Q2_K 0.374. They recommend Q4_K_M for most GPUs and warn that the 2- and 3-bit files "drift visibly from bf16 at 8 steps." Few-step models have less room to absorb quantization error than a 40-step model does, so I'd treat Q4_K_M as the floor.

How to run it with diffusers

Turbo needs a diffusers build with pipeline-configured sigmas (PR #14950) and a recent transformers. The card tells you to install diffusers from source:

pip install git+https://github.com/huggingface/diffusers.git
pip install "transformers>=5.17.0" accelerate pillow
Enter fullscreen mode Exit fullscreen mode

That PR also shipped in the diffusers 0.41.0 release on PyPI (2026-10-06), so pip install "diffusers>=0.41.0" works too.

This is the card's example with a shorter prompt, 1024×1024 instead of the card's 1680×2512, and CPU offload added, which is what most people on a single consumer GPU will want:

import torch
from diffusers import QwenImage21Pipeline

pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1-Turbo",
    dtype=torch.bfloat16,
)
pipe.enable_model_cpu_offload()  # drop this and use .to("cuda") if you have ~40 GB

image = pipe(
    prompt="A hand-lettered chemistry study poster titled \"Chemical Reactions and Equations\"",
    width=1024,
    height=1024,
    use_kv_cache=True,
    generator=torch.Generator("cpu").manual_seed(42),
).images[0]
image.save("t2i.png")
Enter fullscreen mode Exit fullscreen mode

Gotcha number one: num_inference_steps does nothing on its own here. The card is explicit: "Setting num_inference_steps alone does not override" the saved schedule. To try other step counts you have to pass sigmas= explicitly, and Qwen says other schedules "have not been evaluated." If you tried 4 steps and the output didn't change, that's why.

For editing, pass image=load_image(...).convert("RGBA") to the same pipeline. The prefix KV cache (use_kv_cache=True) encodes the prompt and reference images once and reuses them across all 8 steps.

How to run it in ComfyUI without getting mush

ComfyUI has supported Qwen-Image-2.1 natively since day 0. Put these files from Comfy-Org/Qwen-Image-2.1 in your models/ folders:

  • diffusion_models/qwen_image_2.1_turbo_bf16.safetensors (or _int8_convrot)
  • text_encoders/qwen3vl_8b_bf16.safetensors (or _int8_convrot / _w4a8)
  • vae/qwen_image_2.1_vae_bf16.safetensors

Start from the official text-to-image workflow.

Gotcha number two: a plain KSampler at 8 steps doesn't reproduce Turbo's schedule. The fix in discussion #4 ("For anyone getting poor results in comfy") is to switch to SamplerCustomAdvanced with a ManualSigmas node set to:

1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568, 0.0
Enter fullscreen mode Exit fullscreen mode

Use euler and CFG 1. Save that node as a subgraph so you don't have to retype it every time.

If you already have the base model, Comfy-Org also ships qwen_image_2.1_turbo_lora_avg_rank_178_bf16.safetensors (0.91 GB), a LoRA extracted from Turbo. You can apply it to the base DiT instead of downloading a second 14 GB checkpoint.

The license is the actual headline

The original Qwen-Image from August 2025 was Apache-2.0. Qwen-Image-2.1 and Turbo use the Qwen Research License Agreement, which grants rights "FOR NON-COMMERCIAL PURPOSES ONLY" and defines non-commercial as "for research or evaluation purposes only." For commercial use you email model-business@notice.qwencloud.com and ask. Disputes go to courts in Hangzhou under Chinese law.

On the same day the Turbo weights went up, Qwen also launched paid Pro and Turbo APIs on Alibaba Cloud Model Studio. Planned or not, that's how it works out: you can run it at home for free, and if you want to put it in a product you pay Alibaba.

So, can it replace Midjourney? For personal work and experiments, an 8-step model with good text rendering, native transparency and multi-image references is a strong offer. For client work, read the license before you cancel anything.

Is Qwen-Image-2.1-Turbo worth running locally?

If you have 16 GB or more of VRAM and you want to try text-heavy posters, UI mockups or transparent assets, yes. Get Q8_0 or INT8 and set the sigmas correctly. If you need clean commercial rights, the Apache-licensed Qwen-Image from 2025 is still the one you're allowed to ship with.

Has Turbo at 8 steps actually replaced a paid image service for you, or does the non-commercial license rule it out before image quality even matters? Post your GPU, quant and seconds per image in the comments. I haven't found proper Turbo benchmarks anywhere yet, so yours would be among the first.

Sources

Top comments (0)