Originally published on Peak Evergreen.
If you only do one thing: Test one complete asset workflow. Generate it, make one specific revision, place it in a real layout, and count every attempt. An attractive first image is useful; an asset you can revise and deliver consistently is much more valuable.
A good image-model demo answers one question: can this model make something impressive? A useful evaluation asks several more. Can I change the background without changing the product? Does the transparent asset actually have usable edges? Can I get a second version without starting over? How much time passes before the result is ready to use?
Those are the questions that interest me about Qwen-Image 2.1. I have the weights downloaded and running in ComfyUI Desktop on a MacBook Pro with 48 GB of unified memory. That gives this article a practical direction: understanding where the model could fit, then testing those assumptions on a machine someone might already own.
What this session found: I completed 20 local generations: four baseline images, four transparent assets, four edits, three posters, four higher-resolution images, and one synthetic reference. Warm median execution was 143.5 seconds at 1024 × 1024 and 343.5 seconds at 1536 × 1536 for the still life. The images often looked convincing even when they missed placement, exact-text, or unchanged-color requirements. The examples below include those misses.
What changed, and what deserves a closer look
Qwen released 2.1 on September 20, 2026. Its announcement describes a unified model for generation and editing, native transparent output, up to ten reference images, and improvements to text and visual fidelity. The visual generator uses 7 billion parameters across 32 single-stream diffusion-transformer layers. Qwen's announcement
The most interesting part for me is the combination. A reusable illustration often needs generation, extraction, revision, and placement. Bringing more of those operations into one model could reduce the handoffs between tools. Whether it actually does is a workflow question, not something a showcase image settles.
My evaluation separates three kinds of evidence: capabilities documented by Qwen, behavior demonstrated by a particular saved workflow, and results measured on my Mac. A claim about the first does not automatically establish the other two.
Read the license before planning client work
There is a consequential detail behind the downloadable weights. The released model uses the Qwen Research License Agreement. Section 2 limits the grant to non-commercial purposes and requires a separate license for commercial use. The agreement defines non-commercial use as research or evaluation. This is not a blanket permission to deploy the weights in paid production. Released model license
That affects the examples in this article. A product concept or a home-service campaign is a useful evaluation brief; it is not a claim that these weights are automatically cleared for client deliverables. Before using a local installation commercially, settle the applicable permission with the model provider. A hosted service can have different terms, which need their own review.
I would resolve that before spending days building a production asset pipeline. Download access, local execution, and commercial permission are separate decisions.
The 7B headline is not the whole memory budget
A local image workflow has several working parts. The text encoder prepares the conditioning. The diffusion model generates an image representation through successive steps. The VAE converts between image pixels and that representation. Reference images and intermediate data also occupy memory.
Comfy's packaging includes separate image-model, Qwen3-VL 8B text-encoder, and Qwen Image 2.1 VAE files. It offers BF16 and quantized variants rather than one universal download. Comfy's model package
My installed configuration uses qwen_image_2.1_int8_convrot.safetensors, qwen3vl_8b_int8_convrot.safetensors, and qwen_image_2.1_vae_bf16.safetensors. These names matter: a timing for this combination should not be presented as a timing for an unquantized pipeline.
As a rough weight-storage calculation, seven billion values at two bytes each amount to 14 GB before other components and runtime allocations. That is arithmetic, not a measured memory requirement. Quantization can reduce weight storage, but file size alone does not tell us how much memory inference needs or how quickly a particular backend executes it.
On this Mac, the CPU, GPU, operating system, and other applications share the physical memory budget. A workflow fitting once is different from staying responsive through repeated edits. PyTorch's Apple GPU backend is MPS; CUDA-specific instructions and performance claims cannot simply be transferred to it. PyTorch MPS documentation
A repeatable test on my 48 GB MacBook Pro
The machine checked for this article is an Apple M5 Pro MacBook Pro with 18 CPU cores, 20 GPU cores, and 48 GB unified memory, running macOS 27.0. The local startup log reports ComfyUI 0.37.0, PyTorch 2.12.1, and MPS. The image model and text encoder are the INT8 ConvRot files named earlier; the VAE is BF16.
These are configuration observations, not a claim that every 48 GB Mac will perform similarly. Exporting the exact workflow remains necessary because samplers, resolution, attention choices, custom nodes, and prompt processing can change the workload.
The measured configuration was 25 steps, Euler, simple scheduler, CFG 1, batch one, empty negative prompt, with no LoRA, upscaler, prompt enhancement, or manual retouching. Generation, transparent assets, editing, and higher resolution were separate conditions. The official Comfy template uses these sampling settings; Qwen's reference Python example uses 40 steps, so the results should not be mixed. Comfy text-to-image template · Qwen reference implementation
For each timed condition, I restarted the backend and ran seed 100, followed by seeds 101, 102, and 103 without another deliberate restart. “First” therefore means first after backend restart, not disk-cold. macOS's file cache was not cleared. “Warm” describes that sequence, not a promise that every model stayed resident. The execution histories confirm that warm runs reused the unchanged text-encoding node, and the edit runs also reused the loaded reference. Every changed seed ran the sampler again. These medians measure repeated variations of the same brief, not requests with a new prompt every time. First-run differences include fresh conditioning as well as loading.
Local execution / September 22, 2026 : Measured on one M5 Pro, 48 GB.
25 steps · Euler/simple · CFG 1 · batch one · INT8 generator and encoder
| Condition | First after restart | Warm median | Warm range | Warm n |
|---|---|---|---|---|
| Still life · 1024² | 156.0 s | 143.5 s | 142.5 to 145.0 s | 3 |
| Transparent branch · 1024² | 156.0 s | 140.8 s | 140.7 to 142.1 s | 3 |
| Reference edit · 1024² | 221.2 s | 172.5 s | 172.4 to 172.7 s | 3 |
| Still life · 1536² | 358.3 s | 343.5 s | 312.2 to 346.1 s | 3 |
Server execution time includes graph work and image saving. First: seed 100 after a backend restart. Warm: seeds 101, 102, 103 without another deliberate restart, reusing unchanged prompt encoding. Full individual times and API wall times are in the CSV. This is a small local experiment, not a cross-hardware benchmark.
The table uses ComfyUI's server execution-start and execution-success timestamps. I also recorded local API submission-to-completion time, polling every two seconds, in the downloadable CSV. That wall measurement includes polling delay and excludes the later copy of the saved PNG. Neither number is just denoising seconds per step.
The Mac stayed on AC power with its existing energy setting. Other desktop applications remained open, so this is a working-machine observation rather than an isolated hardware benchmark. System swap was already about 3,788 MB before the first run; it would be misleading to attribute all observed swap to this model. Before/after system swap and numeric memory-pressure observations are retained in the records. They are not peak process or GPU-memory measurements. Apple's memory-monitoring guide
Three warm samples per condition establish a local median and range. They do not establish a meaningful p95, general Mac performance, or a ranking against unrelated CUDA benchmarks. Changing the checkpoint, backend, resolution, prompt, reference image, or graph cache state can change the workload.
A good-looking image can still miss the brief
The baseline prompt requested a cream mug beside a closed forest-green notebook, with an evergreen branch and left-side window light. I ran it four times at 1024 × 1024, changing only the seed. All four images looked coherent. Only seed 100 kept the mug beside the notebook; seeds 101, 102, and 103 put it on top.
That distinction matters when evaluating assets against a layout or a client's instructions. These are appealing images, but three need another attempt if “beside” is part of the specification. This is a small, unblinded review of one prompt, not an overall instruction-following score.
One prompt / all four outputs. Objects present; spatial instruction missed. The prompt asked for the mug beside the notebook.
All four baseline outputs, without retouching. Display copies are WebP conversions. Seeds 100 to 103; 1024 × 1024; 25 steps; Euler/simple; CFG 1. One of four follows the requested placement. This is a single-brief observation, not a model-wide accuracy score.
Native transparency is an asset test, not just a visual effect
Qwen documents both transparent generation and extraction of a subject from an ordinary photograph. That makes cutouts, illustrations, interface decoration, and compositing natural things to evaluate. Qwen model card
The important distinction is between an image that looks transparent in a preview and an exported file with a useful alpha channel. A checkerboard can be painted into the RGB pixels. An RGBA file can also have an alpha channel that is completely opaque. Neither is a usable cutout.
For the transparency experiment, I used an illustrated evergreen branch with fine needles, separated stems, and open space between them. This is more informative than a smooth solid icon because the difficult edges are visible.
Create an RGBA PNG asset: one illustrated evergreen branch with two small pine
cones, fine separated needles, and a few open gaps between stems. Warm botanical
editorial illustration with restrained forest-green and brown colors. Center the
complete branch with generous empty space around it. Use a transparent background
and a real alpha channel. No lettering, frame, ground plane, or cast shadow.
The checks are straightforward: confirm non-opaque alpha values, place the same PNG over cream and dark green, and inspect the edges at full size. Look for a pale fringe, clipped needles, missing interior gaps, and shadows that belong to an imaginary background. Then reduce it to the size it would occupy on a webpage. Some defects disappear; others become more obvious.
The selected output, seed 100, has alpha values from 0 to 255. About 34.6% of its pixels are fully transparent. The two placements below use the same raw PNG; only the solid background changes. No edge cleanup or background removal was applied.
Native RGBA / seed 100. One cutout on two real backgrounds. This display copy bakes in the cream and forest page colors side by side; the downloadable file remains a single transparent PNG.
Same RGBA PNG placed over two solid backgrounds for display. No edge cleanup was applied. Download the original transparent PNG.
Needle-edge crops (same 256 × 256 source coordinates):
1024 × 1024; seed 100; 25 steps; Euler/simple; CFG 1. The PNG retains its generated alpha, with metadata removed for the web copy. Small nonzero alpha values remain in nominally empty corners. Download the RGBA PNG.
Much of the needle detail survives both placements, though the dark needles lose contrast on forest green. That is a placement issue to review separately from the matte. “Native alpha” also does not mean a mathematically empty background. The 100-pixel corner samples contain alpha values between 0 and 3, rather than only zero. That faint residue is not conspicuous in these placements, yet it matters if a downstream pipeline expects a perfectly clean matte. Inspect the exported pixels, not just the file extension.
Three of the four branch outputs met my visual brief at webpage size. Seed 103 added a third cone where the prompt asked for two. Inspect the rejected seed 103. The background in that display copy is a normal cream composite, not a second generation.
Editing should be judged by what stays unchanged
Suppose I photograph a plain mug and ask the model to put it into a different setting. The new background may look excellent while the handle becomes thicker, the rim changes shape, or a small surface mark disappears. That is a successful new image and a failed product-preservation edit.
For this session, I generated a synthetic mug reference because a real product photograph was not available. Before editing, I recorded three checks: the tall right-side handle opening, the narrow brown glaze streak running down the front, and the dark elliptical rim. The complete reference was used at 1024 × 1024 without stretching or cropping.
That is a controlled consistency test on a generated image. It does not establish fidelity on photographs of real inventory. A business evaluation should repeat it with its own permitted product photographs.
In <image1>, replace only the background with a warm cream studio backdrop.
Preserve the mug's silhouette, handle opening, rim, surface markings, position,
and camera angle. Keep the product color unchanged. Add no text, props, or
decoration.
Controlled edit / synthetic source. Recognizable does not mean unchanged.
Both images are unretouched WebP display copies at the same scale. Edit: 1024 × 1024; seed 100; 25 steps; Euler/simple; CFG 1. The entire synthetic reference was supplied through image_1. No mask or corrective composite was used.
The first edit preserves the recognizable handle, glaze streak, and rim, and produces the requested warm backdrop. It also warms the body color and moves the rim slightly lower in the frame. That misses a strict background-only brief even though the image remains plausible.
Matching source coordinates. Inspect the rim, mark, and handle.
Same 560 × 420 source-pixel crop from both images, with no alignment correction. The rim sits lower in the edited frame and the body has a warmer tone. The comparison does not claim pixel-exact preservation.
Review the subject at matching scale, not only the complete composition. The detail crops use identical source coordinates without aligning the mug afterward, so they do not conceal the shift.
The other three seeds keep the mug’s placement closer, but all four visibly warm the subject along with the background. None passes the strict unchanged-product-color requirement. All retain the three recognizable identity features. Compare all four edits. For a concept illustration, that could be acceptable; for accurate catalog color, it needs a different workflow or explicit correction.
For a local edit, keep an untouched original and a separate marked or masked input. A mask supplied as a visual instruction does not by itself prove that every pixel outside it is protected. If exact preservation matters, the surrounding workflow may need an explicit composite that restores the original outside the approved region.
I would also compare two revision paths: original to edit A to edit B, and original directly to edit B. That exposes cumulative drift. If the second revision keeps moving the product away from its source, returning to the original reference may be more dependable than repeatedly editing the latest output.
Multiple references need assigned roles
More reference images are not automatically more helpful. A subject photograph, a palette board, and a composition example are three different kinds of instruction. Without clear roles, it becomes hard to tell whether a change is deliberate or accidental.
Comfy's editing template treats image_1 as the target and exposes additional reference slots. Its notes describe referring to them as <image1>, <image2>, and so on. Official editing workflow
An evaluation brief could say: keep the mug from image one, use only the color palette from image two, and ignore the objects in image two. That tests whether the reference guides the intended attribute without importing unrelated content.
Start with one image, then add a second with one defined purpose. Keep the prompt, output dimensions, and sampling settings otherwise stable. Record whether the additional reference improves the required attribute, whether it damages something else, and what happens to completion time. Testing all ten slots immediately makes diagnosis much harder.
For a business, the reusable part would be an approved reference pack with written roles. An assortment of attractive images is a weaker specification than a known product view, a defined palette, and explicit layout constraints.
Typography is a production decision
A model can render attractive lettering and still produce a deliverable with the wrong phone number. Text quality should therefore be tested at the character level as well as the design level.
My typography test is a fictional poster with only three required lines:
Design a square editorial poster using cream, deep forest green, and restrained
orange. Include exactly these three lines of text, with no additional words:
COOL AIR
CLEAR PLAN
SEPTEMBER 2026
Use a strong headline hierarchy, comfortable margins, and a simple abstract
airflow illustration. This is a fictional design study, not an advertisement for
an actual company.
I checked spelling, omitted or extra words, line breaks, and readability at phone size. The figure keeps all three seeds rather than presenting a single selected output as typical behavior.
Exact text / all three seeds. Readable words are only half the test.
Seed 103 (pass): exactly COOL AIR / CLEAR PLAN / SEPTEMBER 2026, once each, correctly spelled and in requested order.
Seeds 101 and 102 (fail): spelling is fine, but both duplicate COOL AIR and CLEAR PLAN.
Raw generations converted to WebP; no text corrected. 1024 × 1024; seeds 101 to 103; 25 steps; Euler/simple; CFG 1. Brief: exactly COOL AIR / CLEAR PLAN / SEPTEMBER 2026, with no extra words.
Seed 103 contains the three requested lines once each. Seeds 101 and 102 spell the words correctly but duplicate “COOL AIR” and “CLEAR PLAN.” Only one of the three passes the exact-text brief. None of the displayed lettering has been corrected.
For an actual campaign, I would compare generated typography with a second route: generate the illustration without words, then add the approved copy in a layout tool. The latter preserves editable text, makes revisions easier, and avoids putting essential contact information inside model-generated pixels.
The right question is whether generating the lettering improves the final workflow. Beautiful lettering does not necessarily justify slower review or less editable source files.
More resolution does not settle the composition
The resolution test repeats the baseline prompt at 1536 × 1536, with the same sampling settings and seed schedule. That is 2.25 times as many output pixels. It is a new generation, not an upscale or a controlled reconstruction of the smaller image.
Resolution / matching display width. More pixels also mean a new composition.
Both are unretouched WebP display copies. Same prompt, seed 100, 25 steps, Euler/simple and CFG 1; different canvas dimensions. The larger image is not an upscale of the smaller one.
At seed 100, the 1024-pixel output puts the mug beside the notebook. The 1536-pixel output puts it on top. The larger version has convincing ceramic and paper detail, yet misses the placement instruction that the smaller version follows.
Seeds 102 and 103 at 1536 pixels do put the mug beside the notebook. That makes two placement passes in four larger outputs, compared with one in four at 1024. This tiny set does not establish a general quality advantage from resolution. Inspect all four larger outputs.
The larger set’s warm median was 343.5 seconds, with a 312.2 to 346.1 second range, about 2.4 times the baseline median. The warm-edit median was 172.5 seconds, and the transparent branch median was 140.8 seconds. All 20 executions completed without a recorded failure or cancellation; completed images can still fail a visual brief.
Choose resolution for the delivery requirement, then review the result against the same brief. A larger file does not automatically earn acceptance, and a matching seed does not lock composition across different canvas sizes.
From one image to a useful asset set
These capabilities suggest several practical evaluation briefs, once the intended use is licensed appropriately:
- A small product catalog: preserve the actual product while testing a consistent background treatment. Check product identity before aesthetic consistency.
- A local service campaign: create a clearly illustrative seasonal visual, then place verified offers and contact information in a separate layout. Do not present generated scenes as photographs of completed customer work.
- A website asset family: produce a hero illustration and smaller transparent elements that share a palette. Evaluate their appearance at real mobile sizes.
- An internal concept review: explore a room, package, or event visual before committing to production. Label concepts and avoid presenting imagined construction details as a build specification.
Each brief needs a definition of done. For the website example: the cutout has clean alpha, the asset is legible at its intended size, the palette works in both themes, the exported file is reasonably sized, and someone can reproduce the selected version.
That links image generation to the broader marketing workflow: an approved brief, reviewed content, reusable source assets, and a deliberate publishing step.
From a prompt to a deliverable / Four checks before an image becomes an asset.
Evaluation frameworkKeep generation speed, first-pass quality, and total production effort as separate measurements.
Prompt enhancement is a separate experimental variable
Qwen publishes optional prompt-enhancement models for generation and editing. Their job is to expand the input description before image generation. They are additional models, not a synonym for the main image generator. Official prompt-enhancement implementation
For these local tests, I kept that stage off. Otherwise a result mixes the image model's behavior with another model's interpretation of the brief, and timing includes an additional stage that may be loaded or offloaded separately.
After establishing a baseline, compare an original prompt with one saved enhanced prompt. Keep both texts and reuse the same seed list. Judge whether the enhancement improves the required details, not simply whether it makes the image more elaborate. It may also introduce an unrequested object or change the intent.
If a rewriter is used, report its elapsed time separately and include it in the complete workflow total. Reusing a prewritten enhanced prompt and rewriting every request are different workloads.
What caching means inside this model
Qwen describes reuse of text and reference-image conditioning across denoising steps. Static conditioning can be computed and reused while the generated image continues to evolve. Qwen's architecture and cache description
That is related to the general idea in prompt caching, but it is not evidence that two independent image requests receive the same hosted cache discount or share a cross-request cache. Those behaviors depend on the implementation.
There is another trap in local benchmarking: the application may reuse an already computed node output when the graph inputs have not changed. A nearly instant repeat can measure skipped work instead of faster inference. For warm timing runs, change the seed, verify the sampler actually executes, and keep the model loaded if the application can. If it offloads and reloads, record that behavior instead of assuming residency.
A useful cache experiment must name what was reused: a model already in memory, prompt encoding from an unchanged node, conditioning within a sampling run, or a completed result. Those are different explanations for a faster number.
Measure useful output, not just seconds per image
Speed and usefulness need separate measurements. A fast output with a changed product or misspelled headline still requires another attempt.
Before generating, define pass/fail checks for the brief: exact text, preserved subject details, usable alpha, requested composition, and no unexpected objects. Review at full resolution and at delivery size. Keep a reason for every rejection. Where possible, assess the images without seeing which setting produced them first.
Then report both the generation timings and the production effort:
First-pass yield = accepted raw outputs / all completed raw outputs
Workflow minutes per final asset =
(generation wait + review + revisions + manual cleanup + export time)
/ number of final accepted assets
Failed or cancelled attempts still belong in the effort total. With no accepted asset, report that outcome rather than a cost-per-asset number. A manually repaired result can count as a final asset, but it should not retroactively become a first-pass success.
In this session, the baseline produced one placement-compliant image in four attempts, and the poster produced one exact-text pass in three. Those counts explain something that a seconds-per-image figure cannot. They apply to these prompts and checks, not every use of the model. I did not continuously time review and export work, so I am not reporting a measured end-to-end cost per usable asset.
This is also where local and hosted workflows can be compared fairly: same briefs, input references, output requirements, and review criteria. Compare complete time and applicable costs, while reporting the different runtimes and permissions. Local generation has no per-request hosted fee, but the machine and the operator's time are not free.
Download the prompts and measured records
The evaluation bundle contains the prompts, machine settings, per-run ComfyUI API graphs, synthetic edit reference, timing CSV, alpha measurements, and descriptive review notes. The run CSV is also available separately. These are API prompt graphs, not visual-editor workflow exports; the bundle explains how they are used.
The raw PNGs, full execution histories, and local logs are retained separately. Published images are unretouched display conversions unless their captions identify ordinary crops or background composites. Review was unblinded, and the acceptance checks were specific to these briefs. A seed helps repeat an experiment within a fixed setup; it does not promise identical pixels across devices or software versions.
This session does not test multi-reference role adherence, prompt enhancement, 2048-pixel output, or cumulative edit drift. Those remain separate experiments. The controlled edit uses a generated reference, so it also cannot establish preservation of a real photographed product.
The useful result is a more concrete buying and workflow question: can this setup produce the asset you need, with an acceptable amount of review and revision? The examples show why that answer needs both timing records and visible failures.
If you are evaluating an image workflow or fitting creative tools into a larger process, the first decision is still the use case and its constraints; the model comes after that. More notes and the full evaluation records live on Peak Evergreen.








Top comments (0)