On September 20, 2026, the Qwen team open-sourced Qwen-Image-2.1, a model that unifies text-to-image generation and image editing. Its visual generation component is a 7B-parameter, 32-layer Single-Stream DiT, while text and conditioning-image encoding are handled by Qwen3-VL-8B. The team describes it as a compact image model designed to balance efficiency and general-purpose capability.
Here are the key specifications reported by the project and community testing:
| Attribute | Specification |
|---|---|
| Intended use | Unified text-to-image and image-editing model |
| Visual generation component | 7B parameters; 32-layer Single-Stream DiT (Block-Causal attention) |
| Text/conditioning encoder | Qwen3-VL-8B (text instructions and reference images use the same encoding path) |
| Core pipeline parameter count | Approximately 16B (the ModelScope repository lists 16.22B) |
| VAE | 64-channel RGBA autoencoder with 16× spatial compression |
| Resolution | Native 2K; 2048 × 2048 by default; seven aspect ratios, up to 2752 × 1536 |
| Inference | Flow Matching with the Euler scheduler; 40 steps recommended. Service defaults may differ. |
| Reference images | The model supports up to 10; the current vLLM-Omni recipe allows up to 4 per request. |
| Transparency | Native RGBA output, with the alpha channel generated directly in latent space |
| Checkpoint size | Approximately 33 GB (ModelScope repository) |
| License | Qwen Research License (limited to non-commercial research and evaluation) |
The four upgrades highlighted by the Qwen team are a more compact and efficient architecture (the DiT shrinks from 20.4B parameters in the previous generation to 7B), native transparent-image generation (RGBA layers can be generated and edited directly), more flexible multi-image and local editing (the model supports up to 10 reference images, subject to backend limits), and improvements to text rendering and portrait lighting details. On transparency, the ComfyUI blog puts it this way: “No other major open model does this.”
1. Deploy Qwen-Image-2.1 on GPUStack
Qwen-Image-2.1 is a diffusion (DiT) model, not an autoregressive LLM, so it cannot be run with conventional vLLM alone. It uses vLLM-Omni, the vLLM project's runtime for diffusion and multimodal generation models. The Qwen team announced support on the model's release day (Day 0).
The workflow in GPUStack is similar to deploying an LLM: first add a vLLM version that points to a vLLM-Omni image under Inference Backends. The rest of the deployment can be completed in the web interface.
Test environment: NVIDIA A800-SXM4-80GB; driver 570.172.08; CUDA 12.8.
1) Add a vLLM-Omni version under Inference Backends
Important: Add the image version under Inference Backends first. It will then appear in the Backend Version dropdown on the deployment page.
In the left-hand menu, open Inference Backends, find the vLLM card, and click Edit. Under Version Configuration, click Add Version.
-
Version:
qwen-image21 -
Image:
vllm/vllm-omni:qwen-image21(the image tag used in this guide). Verify that this image was built with the Qwen-Image-2.1 integration. At the time this guide was checked, the vLLM recipe said support had not yet been included in a versioned release package and that the corresponding branch should be used. An image tag alone does not prove that it is an official release. - Framework: Select CUDA.
- Keep the default image entrypoint (
ENTRYPOINT),vllm serve. Add--omnito the default command template, for example:{{model_path}} --omni --host {{worker_ip}} --port {{port}} --served-model-name {{model_name}}.
Save the configuration. You can then select qwen-image21 on the deployment page.
⚠️ Image and startup command: Qwen-Image-2.1 requires the diffusion implementation in vLLM-Omni and must be launched with
vllm serve <model-name> --omni. Make sure the image contains the required model code and that GPUStack actually passes--omnito the process. A standard vLLM image or a configuration that omits this argument will not run the example as shown.
2) Create a deployment and select the model and backend
Return to Deployments and click Deploy Model in the upper-right corner. Enter the following settings:
-
Source: Select ModelScope (this guide uses ModelScope; choose a model source based on your network environment and organizational policy) and enter the repository ID
Qwen/Qwen-Image-2.1. - Backend: Select vLLM.
-
Backend Version: Select the
qwen-image21version you just added.
Submit the deployment and wait for the instance to reach Running. If the model fails to load, check the image version, the --omni startup argument, and access permissions for the model repository.
2. Verify and call the service
1) Generate an image through an OpenAI-compatible API
In --omni mode, vLLM-Omni registers two OpenAI-compatible endpoints:
-
/v1/images/generations: standard image-generation endpoint -
/v1/chat/completions: chat endpoint that supports multimodal input
The /v1/chat/completions endpoint uses a multimodal message structure: messages[].content is an array, and prompt text goes in a text item. To generate an image, also set modalities: ["image"] in the request body:
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "qwen-image-2.1",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "A golden retriever running through a grassy field, backlit, cinematic lighting."}
]
}
],
"modalities": ["image"],
"extra_body": {
"size": "1024x1024",
"num_inference_steps": 40,
"true_cfg_scale": 1.0,
"seed": 42
}
}'
To edit an image using a reference, add an image_url item to content, while keeping modalities: ["image"] and the image-generation parameters:
"content": [
{"type": "text", "text": "Extract the main subject from the reference image as a layer with a transparent background."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<REFERENCE_IMAGE_BASE64>"}}
]
Request parameter notes:
-
model: Use the same model name that you specified when creating the deployment. -
messages[].content: Use the array structure; put the prompt in atextitem and the reference image in animage_urlitem. -
image_url.url: A data URI such asdata:image/png;base64,<content>can be used. The current vLLM-Omni recipe limits each data-URI part to 1 MiB. For larger reference images, use the/v1/images/editsfile-upload endpoint.
The generated image in the response is at choices[0].message.content[0].image_url.url. For /v1/images/generations, the image is at data[0].b64_json.
For /v1/images/generations, pass the resolution, number of steps, and guidance scale directly in the request body:
curl -s http://localhost:8000/v1/images/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "qwen-image-2.1",
"prompt": "A corgi wearing a knitted beanie, sitting at a wooden table by the window, with a steaming cup of coffee nearby.",
"size": "2048x2048",
"num_inference_steps": 40,
"true_cfg_scale": 1.0,
"seed": 42
}'
Parameters to keep in mind:
-
num_inference_steps: Setting this to 40 is recommended. More steps usually increase generation time; results also depend on the prompt and sampling settings. -
seed: Fix the seed when you need reproducible editing results. -
true_cfg_scale: This guide sets it to1.0. Do not mix it up withguidance_scale. If you also pass a negative prompt, consider how the guidance settings may affect latency and output quality.
2) Image editing and multi-image composition
One of the important capabilities of 2.1 is that it can take reference images and reorganize a scene while attempting to preserve subjects' identities and product appearances. Typical scenarios showcased by the Qwen team include:
- Combining six portraits into a group photo: Provide individual portraits and combine them into one group image.
- Combining five fashion items into an outfit: Provide a model, clothing, shoes, a bag, and a hat to generate a virtual try-on image.
- Arranging a room from 10 furniture images: Provide a furniture set and compose it into a complete interior.
Reference-image limits depend on the inference backend. The model specification supports up to 10 images, while the current vLLM-Omni recipe allows up to 4 per request. Check whether your API and deployment version support local-editing controls such as selection, brush painting, and separate masks before relying on them.
Here is an edit of the image generated above. Pass the reference image as an image_url item to /v1/chat/completions and ask the model to change only the specified content while preserving the rest of the composition:
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "qwen-image-2.1",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Replace the cup of coffee on the table in the reference image with a stack of books."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<IMG_B64>"}}
]
}
],
"modalities": ["image"],
"extra_body": {
"size": "1024x1024",
"num_inference_steps": 40,
"true_cfg_scale": 1.0,
"seed": 42
}
}'
The returned image is at choices[0].message.content[0].image_url.url. If the field contains a data URI, extract its Base64 content and decode it as edited.png. For larger reference images, use the /v1/images/edits file-upload endpoint to avoid the data-URI size limit.
3) Native transparent images
The alpha channel is generated directly in latent space rather than added through post-processing. As a result, a PNG can be used in infographics, storyboards, stickers, and logos without first running a background-removal or segmentation model. The capability also works in reverse: you can extract a subject from a regular photo as a transparent layer, or edit text on a transparent layer while preserving the background.
Pass an existing image as an image_url item to /v1/chat/completions and ask the model to extract the subject as a transparent layer. First confirm that the current service configuration outputs RGBA. A model's ability to generate transparency does not mean every service configuration preserves the alpha channel by default.
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "qwen-image-2.1",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Extract the main subject from the reference image as a layer with a transparent background."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<IMG_B64>"}}
]
}
],
"modalities": ["image"],
"extra_body": {
"size": "1024x1024",
"num_inference_steps": 40,
"true_cfg_scale": 1.0,
"seed": 42
}
}'
The returned image is at choices[0].message.content[0].image_url.url. Decoding it should give you a PNG with an alpha channel; verify that the output file actually retains transparency.
3. Important considerations
The license is not Apache-2.0. Qwen-Image-2.1 uses the Qwen Research License, which permits non-commercial research or evaluation. Commercial use requires a separate license. Consult the license text for the exact terms.
Benchmark scores depend on the evaluation method. When citing Qwen-Image-Bench, include the benchmark version, data date, and source. Note that the benchmark is published by the Qwen team; its results should not be presented as an independent third-party evaluation.
Community reports are not universal conclusions. Editing, multi-image composition, and transparent-image generation are capabilities highlighted by the project. Photorealism, artifacts, and consistency across multiple people can vary with the prompt, resolution, and inference configuration. When citing community feedback, include the source and test conditions.
Notes for international readers
- Model hosting is region- and policy-dependent. This guide uses a ModelScope repository ID. The Qwen project also publishes model resources through Hugging Face; choose an accessible source that complies with your organization's network and data policies. A China-based hosting platform should not be described as universally accessible or as the only option.
- The hardware and system details are specific to the author's test environment. The NVIDIA A800-SXM4-80GB, driver 570.172.08, and CUDA 12.8 are reproducibility details, not a universal hardware recommendation. Readers should check availability, export controls, and local deployment requirements in their region.
- The license applies regardless of the deployment region. Do not imply that using a hosted model or deploying it outside China changes the Qwen Research License's restrictions. Commercial use requires separate authorization.
References
- Qwen-Image-2.1 official repository and model documentation
- Qwen Research License
- vLLM Recipes: Qwen-Image-2.1
- vLLM-Omni Qwen-Image-2.1 integration PR
Summary
Qwen-Image-2.1 has a clear design direction: reduce the previous generation's 20.4B-parameter DiT to 7B, bring text-to-image generation, image editing, and transparent-image generation into one checkpoint, and use Qwen3-VL to unify text and conditioning-image encoding. Block-Causal attention enables prefix KV reuse, reducing inference overhead in multi-image editing scenarios.
With GPUStack's pluggable vLLM backend, you can configure a vLLM-Omni image that contains the Qwen-Image-2.1 integration, select the Qwen/Qwen-Image-2.1 repository, and verify that the startup command includes --omni. The image must match the vLLM-Omni code version in use. Check the official recipe for the supported parameters, input formats, and reference-image limits for your version.
Test environment: NVIDIA A800-SXM4-80GB | Driver: 570.172.08 | CUDA: 12.8 | GPUStack v2.2.2







Top comments (0)