I used to compare image models in the laziest possible way: give them the same prompt, put the outputs next to each other, and pick the one I liked most.
It felt reasonable. It also gave me bad conclusions.
The problem showed up when I tried to use those “winning” models for actual creative work. A model that looked fantastic on a cinematic prompt could be irritatingly unreliable on an ad. Another would produce a less impressive first image but follow instructions more closely and save me two or three retries.
That made me stop asking which model makes the prettiest image.
Now I care more about a narrower question:
How quickly can this model get a specific asset close to something I would actually use?
This post covers the small, practical comparison I use to answer that question.
My first benchmark was basically useless
The first version was just a folder full of generations.
Same prompt. Same aspect ratio. Three models.
Then I would stare at the results and write things like:
- “Model A looks more polished.”
- “Model B feels more realistic.”
- “Model C has better detail.”
None of that was wrong, exactly. It just wasn’t very useful.
If I’m making a product ad, “better detail” doesn’t matter much if the model mangles the headline. If I’m editing an existing product photo, visual quality means very little if the bottle changes shape.
So I started looking at things that map more directly to the job:
| Metric | What I actually care about |
|---|---|
| Prompt adherence | Did it follow the brief? |
| Text accuracy | Is the text actually correct and complete? |
| Composition | Is the layout usable? |
| Reference preservation | Did it keep the original subject intact? |
| Iterations needed | How many tries before I got something usable? |
| Latency | Does the speed get annoying during iteration? |
I usually score the first four from 1–5. The numbers are not scientific. They are just enough to stop me from changing the criteria halfway through because one image happens to look cool.
For this comparison, I used:
- GPT Image 2
- Nano Banana 2
- Qwen Image 3.0
Every comparison image below follows the same order: GPT Image 2 on the left, Nano Banana 2 in the center, and Qwen Image 3.0 on the right.
I ran them through Ezier’s AI image workspace so I could switch models without changing the rest of my workflow. The point here is not to produce a universal ranking. It is to see what each model feels like during ordinary creative work.
Test 1: Can it make a usable product ad?
I don’t start with portraits anymore. Most strong image models can make a decent-looking portrait, so it doesn’t tell me much.
A product ad is more revealing.
I used this prompt:
Create a premium square advertising image for black wireless earbuds in a compact charging case. Place the product slightly below center on a warm light-beige studio background. Use soft directional lighting and a subtle natural shadow. Add the headline “LESS NOISE. MORE MUSIC.” above the product in clean bold sans-serif typography. Minimal commercial photography, no additional objects, no logos, no extra text.
The first thing I checked was the headline.
Not the lighting. Not the reflections. Not whether the earbuds looked expensive.
Did it write LESS NOISE. MORE MUSIC. correctly?
Then I looked at the boring details that become very un-boring when you need to ship an asset:
- Did it invent an extra earbud?
- Did the charging case turn into an impossible object?
- Did it ignore the requested placement?
- Did it add random copy?
- Would I keep working from this image, or immediately regenerate?
From left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0. Each panel shows the first result from the same product-ad prompt.
| Model | Prompt adherence | Text | Composition | First output usable? |
|---|---|---|---|---|
| GPT Image 2 | 5/5 | 5/5 | 5/5 | Yes |
| Nano Banana 2 | 5/5 | 5/5 | 4/5 | Yes |
| Qwen Image 3.0 | 5/5 | 5/5 | 5/5 | Yes |
All three models rendered the headline correctly, so this test was closer than I expected.
GPT Image 2 produced the most polished product rendering, with strong lighting and a convincing commercial finish. Qwen Image 3.0 also gave me a balanced, usable composition with generous negative space. Nano Banana 2 followed the minimal brief well, although the product felt slightly undersized compared with the other two results.
I would keep all three first outputs. The difference here was less about correctness and more about which composition I preferred.
That is worth saying because comparison posts often force a dramatic winner even when the honest result is a near tie.
Test 2: Make the text harder
The second test is less visually interesting, which is exactly why I like it.
Prompt:
Design a square social media graphic announcing a fictional developer conference called “BUILD//26”. Large headline: “BUILD BETTER. SHIP FASTER.” Smaller text underneath: “September 18–19 · San Francisco”. Add a small button-style label saying “EARLY ACCESS”. Use a black, white, and electric-blue visual system, strong Swiss-inspired grid layout, modern developer conference branding, crisp typography, minimal abstract geometric elements. No additional text.
There were four strings that needed to survive intact:
BUILD//26BUILD BETTER. SHIP FASTER.September 18–19 · San FranciscoEARLY ACCESS
I used to be too generous with text generation. If a word looked roughly right at thumbnail size, I would count it as a win.
Now I zoom in.
A poster with one broken or missing line is still a broken poster.
From left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0. The same typography-heavy prompt revealed a much clearer difference between the models.
| Model | Exact wording | Spelling | Hierarchy | Layout consistency | Visual quality |
|---|---|---|---|---|---|
| GPT Image 2 | 4/5 | 5/5 | 4/5 | 5/5 | 5/5 |
| Nano Banana 2 | 4/5 | 5/5 | 4/5 | 4/5 | 4/5 |
| Qwen Image 3.0 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
This test produced a much clearer difference. GPT Image 2 and Nano Banana 2 both omitted BUILD//26. Everything they did render was spelled correctly, and both designs looked polished, but the missing event name would still require a manual fix.
Qwen Image 3.0 was the only model that preserved all four required strings. It also kept a clear hierarchy between the conference name, headline, date, and button. For this particular typography-heavy task, Qwen gave me the most usable first output.
The most frustrating failure was not misspelling. It was omission. GPT Image 2 and Nano Banana 2 produced attractive graphics but quietly dropped the one line that identified the event.
This is why I distrust broad “best image model” rankings. Text-heavy graphics and pure image generation are not the same job. A model can be excellent at one and mediocre at the other.
Test 3: Can it keep the product while changing the scene?
Text-to-image gets most of the attention, but a lot of practical creative work continues from an image you already have.
The request is often deceptively simple:
Keep the product. Change the background.
For this test, I first asked each model to create the same kind of neutral product shot:
Photorealistic studio product photograph of a matte white insulated water bottle with a simple cylindrical shape and small stainless-steel cap, standing upright on a pale gray seamless background. Front three-quarter view, soft diffused studio lighting, subtle shadow beneath the bottle, no branding, no text, no props, centered composition, premium e-commerce photography, square image.
Then I gave each model its own image and asked for one edit:
Keep the bottle completely unchanged. Replace only the background with a sunlit Mediterranean stone terrace overlooking a calm blue sea. Add realistic warm late-afternoon sunlight and a natural contact shadow. Do not change the bottle’s shape, proportions, cap, material, color, or camera angle. Do not add text or branding.
So this was not three models editing one shared source file. It was three separate generate-then-edit workflows. That is also closer to how I normally work: create an asset, then keep refining it with the same model.
Columns from left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0. The top row shows each model’s original bottle; the bottom row shows its edited result.
| Model | Product preservation | Background replacement | Overall usability |
|---|---|---|---|
| GPT Image 2 | 5/5 | 5/5 | Ready to use |
| Nano Banana 2 | 5/5 | 5/5 | Ready to use |
| Qwen Image 3.0 | 5/5 | 5/5 | Ready to use |
This was the surprise of the comparison: all three models did a good job.
GPT Image 2 kept the original scale and framing especially close. Nano Banana 2 preserved the distinctive handle on its cap and blended the bottle naturally into a warmer coastal scene. Qwen Image 3.0 kept the rounded cap, dark ring, bottle shape, and viewing angle while adding convincing light and shadow.
Most importantly, none of the three changed the product in a way that would make me reject the edit. The second images all looked polished and usable. At that point, choosing between them became a matter of taste: cleaner product presentation, warmer atmosphere, or a slightly different background composition.
For this task, I would call it a three-way tie.
A note on speed and retries
GPT Image 2 and Nano Banana 2 each took around 30 seconds to return an image. Qwen Image 3.0 was closer to 40 seconds.
I was not running a timer or trying to measure model latency precisely. These were simply the wait times I noticed during normal use. For a single image, the ten-second difference was not enough to affect my choice. It would matter more if I were generating a large batch.
Retries matter more to me than a small speed difference anyway. A beautiful result is less useful if it takes several attempts to reach it.
In this comparison, all three models produced a usable product ad, Qwen was the only one to keep every required line in the typography test, and all three completed the editing task successfully. Those practical differences are more useful to me than a single overall score.
What I would choose after this round
For a polished product visual, I would probably start with GPT Image 2. Qwen Image 3.0 was close, and all three results from the first test were usable.
For a graphic where every line of text has to survive, I would choose Qwen Image 3.0 based on this round. It was the only one that included the conference name, headline, date, and button text without dropping anything.
For background replacement, I would be comfortable using any of the three. The products stayed intact, and the final scenes all looked convincing. I would choose based on which visual mood fit the project.
That is the main thing I took away from the comparison: the useful question is not “Which model is best?” It is “Which model is best for the asset I need right now?”
Test your own boring work
Generic leaderboards can show what a model is capable of, but they do not know what you make every week.
If you work in e-commerce, test packaging text, product preservation, background replacement, and ad variations. If you make social graphics, test typography and layout. If you work with recurring characters, test identity, pose, and expression consistency.
The most useful prompts are usually the boring ones—the tasks you actually repeat.
I no longer care much about finding one image model that wins every comparison. I care about getting the asset I need, with as little correction and as few retries as possible.



Top comments (0)