<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Eric Cheung</title>
    <description>The latest articles on DEV Community by Eric Cheung (@eric_cheung_1030).</description>
    <link>https://dev.to/eric_cheung_1030</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3903644%2F1541f806-6631-484c-b89e-e4d08b9e77c3.png</url>
      <title>DEV Community: Eric Cheung</title>
      <link>https://dev.to/eric_cheung_1030</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eric_cheung_1030"/>
    <language>en</language>
    <item>
      <title>Same Prompt, Different Model: How I Benchmark AI Image Generators Without Fooling Myself</title>
      <dc:creator>Eric Cheung</dc:creator>
      <pubDate>Fri, 04 Sep 2026 13:17:11 +0000</pubDate>
      <link>https://dev.to/eric_cheung_1030/same-prompt-different-model-how-i-benchmark-ai-image-generators-without-fooling-myself-1lbn</link>
      <guid>https://dev.to/eric_cheung_1030/same-prompt-different-model-how-i-benchmark-ai-image-generators-without-fooling-myself-1lbn</guid>
      <description>&lt;p&gt;I used to compare image models in the laziest possible way: give them the same prompt, put the outputs next to each other, and pick the one I liked most.&lt;/p&gt;

&lt;p&gt;It felt reasonable. It also gave me bad conclusions.&lt;/p&gt;

&lt;p&gt;The problem showed up when I tried to use those “winning” models for actual creative work. A model that looked fantastic on a cinematic prompt could be irritatingly unreliable on an ad. Another would produce a less impressive first image but follow instructions more closely and save me two or three retries.&lt;/p&gt;

&lt;p&gt;That made me stop asking which model makes the prettiest image.&lt;/p&gt;

&lt;p&gt;Now I care more about a narrower question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How quickly can this model get a specific asset close to something I would actually use?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This post covers the small, practical comparison I use to answer that question.&lt;/p&gt;




&lt;h2&gt;
  
  
  My first benchmark was basically useless
&lt;/h2&gt;

&lt;p&gt;The first version was just a folder full of generations.&lt;/p&gt;

&lt;p&gt;Same prompt. Same aspect ratio. Three models.&lt;/p&gt;

&lt;p&gt;Then I would stare at the results and write things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Model A looks more polished.”&lt;/li&gt;
&lt;li&gt;“Model B feels more realistic.”&lt;/li&gt;
&lt;li&gt;“Model C has better detail.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that was wrong, exactly. It just wasn’t very useful.&lt;/p&gt;

&lt;p&gt;If I’m making a product ad, “better detail” doesn’t matter much if the model mangles the headline. If I’m editing an existing product photo, visual quality means very little if the bottle changes shape.&lt;/p&gt;

&lt;p&gt;So I started looking at things that map more directly to the job:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What I actually care about&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt adherence&lt;/td&gt;
&lt;td&gt;Did it follow the brief?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text accuracy&lt;/td&gt;
&lt;td&gt;Is the text actually correct and complete?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Composition&lt;/td&gt;
&lt;td&gt;Is the layout usable?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference preservation&lt;/td&gt;
&lt;td&gt;Did it keep the original subject intact?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Iterations needed&lt;/td&gt;
&lt;td&gt;How many tries before I got something usable?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Does the speed get annoying during iteration?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I usually score the first four from 1–5. The numbers are not scientific. They are just enough to stop me from changing the criteria halfway through because one image happens to look cool.&lt;/p&gt;

&lt;p&gt;For this comparison, I used:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT Image 2&lt;/li&gt;
&lt;li&gt;Nano Banana 2&lt;/li&gt;
&lt;li&gt;Qwen Image 3.0&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every comparison image below follows the same order: &lt;strong&gt;GPT Image 2 on the left, Nano Banana 2 in the center, and Qwen Image 3.0 on the right.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I ran them through &lt;a href="https://ezier.ai/ai-image/" rel="noopener noreferrer"&gt;Ezier’s AI image workspace&lt;/a&gt; so I could switch models without changing the rest of my workflow. The point here is not to produce a universal ranking. It is to see what each model feels like during ordinary creative work.&lt;/p&gt;




&lt;h2&gt;
  
  
  Test 1: Can it make a usable product ad?
&lt;/h2&gt;

&lt;p&gt;I don’t start with portraits anymore. Most strong image models can make a decent-looking portrait, so it doesn’t tell me much.&lt;/p&gt;

&lt;p&gt;A product ad is more revealing.&lt;/p&gt;

&lt;p&gt;I used this prompt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Create a premium square advertising image for black wireless earbuds in a compact charging case. Place the product slightly below center on a warm light-beige studio background. Use soft directional lighting and a subtle natural shadow. Add the headline “LESS NOISE. MORE MUSIC.” above the product in clean bold sans-serif typography. Minimal commercial photography, no additional objects, no logos, no extra text.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The first thing I checked was the headline.&lt;/p&gt;

&lt;p&gt;Not the lighting. Not the reflections. Not whether the earbuds looked expensive.&lt;/p&gt;

&lt;p&gt;Did it write &lt;strong&gt;LESS NOISE. MORE MUSIC.&lt;/strong&gt; correctly?&lt;/p&gt;

&lt;p&gt;Then I looked at the boring details that become very un-boring when you need to ship an asset:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Did it invent an extra earbud?&lt;/li&gt;
&lt;li&gt;Did the charging case turn into an impossible object?&lt;/li&gt;
&lt;li&gt;Did it ignore the requested placement?&lt;/li&gt;
&lt;li&gt;Did it add random copy?&lt;/li&gt;
&lt;li&gt;Would I keep working from this image, or immediately regenerate?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2a75dal9e32cmi04igrx.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2a75dal9e32cmi04igrx.webp" alt="Three product-ad results shown from left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0" width="800" height="297"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;From left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0. Each panel shows the first result from the same product-ad prompt.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Prompt adherence&lt;/th&gt;
&lt;th&gt;Text&lt;/th&gt;
&lt;th&gt;Composition&lt;/th&gt;
&lt;th&gt;First output usable?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT Image 2&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Image 3.0&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All three models rendered the headline correctly, so this test was closer than I expected.&lt;/p&gt;

&lt;p&gt;GPT Image 2 produced the most polished product rendering, with strong lighting and a convincing commercial finish. Qwen Image 3.0 also gave me a balanced, usable composition with generous negative space. Nano Banana 2 followed the minimal brief well, although the product felt slightly undersized compared with the other two results.&lt;/p&gt;

&lt;p&gt;I would keep all three first outputs. The difference here was less about correctness and more about which composition I preferred.&lt;/p&gt;

&lt;p&gt;That is worth saying because comparison posts often force a dramatic winner even when the honest result is a near tie.&lt;/p&gt;




&lt;h2&gt;
  
  
  Test 2: Make the text harder
&lt;/h2&gt;

&lt;p&gt;The second test is less visually interesting, which is exactly why I like it.&lt;/p&gt;

&lt;p&gt;Prompt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Design a square social media graphic announcing a fictional developer conference called “BUILD//26”. Large headline: “BUILD BETTER. SHIP FASTER.” Smaller text underneath: “September 18–19 · San Francisco”. Add a small button-style label saying “EARLY ACCESS”. Use a black, white, and electric-blue visual system, strong Swiss-inspired grid layout, modern developer conference branding, crisp typography, minimal abstract geometric elements. No additional text.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There were four strings that needed to survive intact:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;BUILD//26&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;BUILD BETTER. SHIP FASTER.&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;September 18–19 · San Francisco&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;EARLY ACCESS&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I used to be too generous with text generation. If a word looked roughly right at thumbnail size, I would count it as a win.&lt;/p&gt;

&lt;p&gt;Now I zoom in.&lt;/p&gt;

&lt;p&gt;A poster with one broken or missing line is still a broken poster.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faom8bpl5dc78f9ccgrql.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faom8bpl5dc78f9ccgrql.webp" alt="Three typography-test results shown from left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0" width="800" height="297"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;From left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0. The same typography-heavy prompt revealed a much clearer difference between the models.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Exact wording&lt;/th&gt;
&lt;th&gt;Spelling&lt;/th&gt;
&lt;th&gt;Hierarchy&lt;/th&gt;
&lt;th&gt;Layout consistency&lt;/th&gt;
&lt;th&gt;Visual quality&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT Image 2&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Image 3.0&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This test produced a much clearer difference. GPT Image 2 and Nano Banana 2 both omitted &lt;code&gt;BUILD//26&lt;/code&gt;. Everything they did render was spelled correctly, and both designs looked polished, but the missing event name would still require a manual fix.&lt;/p&gt;

&lt;p&gt;Qwen Image 3.0 was the only model that preserved all four required strings. It also kept a clear hierarchy between the conference name, headline, date, and button. For this particular typography-heavy task, Qwen gave me the most usable first output.&lt;/p&gt;

&lt;p&gt;The most frustrating failure was not misspelling. It was omission. GPT Image 2 and Nano Banana 2 produced attractive graphics but quietly dropped the one line that identified the event.&lt;/p&gt;

&lt;p&gt;This is why I distrust broad “best image model” rankings. Text-heavy graphics and pure image generation are not the same job. A model can be excellent at one and mediocre at the other.&lt;/p&gt;




&lt;h2&gt;
  
  
  Test 3: Can it keep the product while changing the scene?
&lt;/h2&gt;

&lt;p&gt;Text-to-image gets most of the attention, but a lot of practical creative work continues from an image you already have.&lt;/p&gt;

&lt;p&gt;The request is often deceptively simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Keep the product. Change the background.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For this test, I first asked each model to create the same kind of neutral product shot:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Photorealistic studio product photograph of a matte white insulated water bottle with a simple cylindrical shape and small stainless-steel cap, standing upright on a pale gray seamless background. Front three-quarter view, soft diffused studio lighting, subtle shadow beneath the bottle, no branding, no text, no props, centered composition, premium e-commerce photography, square image.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then I gave each model its own image and asked for one edit:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Keep the bottle completely unchanged. Replace only the background with a sunlit Mediterranean stone terrace overlooking a calm blue sea. Add realistic warm late-afternoon sunlight and a natural contact shadow. Do not change the bottle’s shape, proportions, cap, material, color, or camera angle. Do not add text or branding.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So this was not three models editing one shared source file. It was three separate generate-then-edit workflows. That is also closer to how I normally work: create an asset, then keep refining it with the same model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzugvqdgasnwtc6cazrf.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzugvqdgasnwtc6cazrf.webp" alt="Before-and-after background-editing comparison. The columns show GPT Image 2, Nano Banana 2, and Qwen Image 3.0 from left to right; original images are on top and edited images are below" width="800" height="597"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Columns from left to right: GPT Image 2, Nano Banana 2, and Qwen Image 3.0. The top row shows each model’s original bottle; the bottom row shows its edited result.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Product preservation&lt;/th&gt;
&lt;th&gt;Background replacement&lt;/th&gt;
&lt;th&gt;Overall usability&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT Image 2&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Ready to use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Ready to use&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen Image 3.0&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Ready to use&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This was the surprise of the comparison: all three models did a good job.&lt;/p&gt;

&lt;p&gt;GPT Image 2 kept the original scale and framing especially close. Nano Banana 2 preserved the distinctive handle on its cap and blended the bottle naturally into a warmer coastal scene. Qwen Image 3.0 kept the rounded cap, dark ring, bottle shape, and viewing angle while adding convincing light and shadow.&lt;/p&gt;

&lt;p&gt;Most importantly, none of the three changed the product in a way that would make me reject the edit. The second images all looked polished and usable. At that point, choosing between them became a matter of taste: cleaner product presentation, warmer atmosphere, or a slightly different background composition.&lt;/p&gt;

&lt;p&gt;For this task, I would call it a three-way tie.&lt;/p&gt;




&lt;h2&gt;
  
  
  A note on speed and retries
&lt;/h2&gt;

&lt;p&gt;GPT Image 2 and Nano Banana 2 each took around 30 seconds to return an image. Qwen Image 3.0 was closer to 40 seconds.&lt;/p&gt;

&lt;p&gt;I was not running a timer or trying to measure model latency precisely. These were simply the wait times I noticed during normal use. For a single image, the ten-second difference was not enough to affect my choice. It would matter more if I were generating a large batch.&lt;/p&gt;

&lt;p&gt;Retries matter more to me than a small speed difference anyway. A beautiful result is less useful if it takes several attempts to reach it.&lt;/p&gt;

&lt;p&gt;In this comparison, all three models produced a usable product ad, Qwen was the only one to keep every required line in the typography test, and all three completed the editing task successfully. Those practical differences are more useful to me than a single overall score.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I would choose after this round
&lt;/h2&gt;

&lt;p&gt;For a polished product visual, I would probably start with GPT Image 2. Qwen Image 3.0 was close, and all three results from the first test were usable.&lt;/p&gt;

&lt;p&gt;For a graphic where every line of text has to survive, I would choose Qwen Image 3.0 based on this round. It was the only one that included the conference name, headline, date, and button text without dropping anything.&lt;/p&gt;

&lt;p&gt;For background replacement, I would be comfortable using any of the three. The products stayed intact, and the final scenes all looked convincing. I would choose based on which visual mood fit the project.&lt;/p&gt;

&lt;p&gt;That is the main thing I took away from the comparison: the useful question is not “Which model is best?” It is “Which model is best for the asset I need right now?”&lt;/p&gt;




&lt;h2&gt;
  
  
  Test your own boring work
&lt;/h2&gt;

&lt;p&gt;Generic leaderboards can show what a model is capable of, but they do not know what you make every week.&lt;/p&gt;

&lt;p&gt;If you work in e-commerce, test packaging text, product preservation, background replacement, and ad variations. If you make social graphics, test typography and layout. If you work with recurring characters, test identity, pose, and expression consistency.&lt;/p&gt;

&lt;p&gt;The most useful prompts are usually the boring ones—the tasks you actually repeat.&lt;/p&gt;

&lt;p&gt;I no longer care much about finding one image model that wins every comparison. I care about getting the asset I need, with as little correction and as few retries as possible.&lt;/p&gt;




&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ezier.ai/ai-image/" rel="noopener noreferrer"&gt;Ezier AI Image&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.openai.com/api/docs/guides/image-generation" rel="noopener noreferrer"&gt;OpenAI image generation documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/image-generation" rel="noopener noreferrer"&gt;Google Gemini image generation documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.alibabacloud.com/help/en/model-studio/qwen-image-generation-and-editing-api-reference" rel="noopener noreferrer"&gt;Qwen Image generation and editing documentation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>nanobanana</category>
      <category>qwen</category>
    </item>
    <item>
      <title>I Built a Watermark Remover — Here’s What I Actually Learned</title>
      <dc:creator>Eric Cheung</dc:creator>
      <pubDate>Wed, 29 Apr 2026 06:52:49 +0000</pubDate>
      <link>https://dev.to/eric_cheung_1030/i-built-a-watermark-remover-heres-what-i-actually-learned-4o86</link>
      <guid>https://dev.to/eric_cheung_1030/i-built-a-watermark-remover-heres-what-i-actually-learned-4o86</guid>
      <description>&lt;p&gt;I'd generate an image with Gemini, like it, want to drop it into a draft or mockup — and there was the visible watermark sitting right on top of the export. Not a huge deal, but annoying enough that I'd break flow every time. Opening Photoshop or GIMP for one overlay felt absurd. Cropping usually ruined the composition.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3xo54rdjgjgdidpqoo91.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3xo54rdjgjgdidpqoo91.png" alt=" " width="800" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So I spent a weekend building something for exactly that: &lt;strong&gt;&lt;a href="https://geminiwatermarkremover.ai/" rel="noopener noreferrer"&gt;Gemini Watermark Remover&lt;/a&gt;&lt;/strong&gt; — upload an image, remove the visible mark in-browser, download a clean PNG.&lt;/p&gt;

&lt;p&gt;This is the story of how I built it and what I got wrong before I got it right.&lt;/p&gt;




&lt;h2&gt;
  
  
  The first decision: do one thing
&lt;/h2&gt;

&lt;p&gt;I started with a constraint: no editor, no layers, no timeline, no format conversion, no "enhance" button. Just this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Upload → Remove the mark → Download a clean PNG.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's it. Every time I felt the urge to add something — batch mode, adjustment sliders, export options — I came back to that constraint and cut it.&lt;/p&gt;

&lt;p&gt;The constraint wasn't laziness. It was a product decision. Tools that do everything require users to think. Tools that do one thing let users just get on with their work.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the browser, and why it actually mattered
&lt;/h2&gt;

&lt;p&gt;The core processing runs in the browser. No server upload, no queue, no storage policy. The image never leaves the tab.&lt;/p&gt;

&lt;p&gt;For a photo editor this might be a trade-off. For a small utility like this, it's the right default. People use it for drafts, client concepts, internal assets — images that probably shouldn't hit a random server in the first place.&lt;/p&gt;

&lt;p&gt;The pipeline is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;File input → Image decode → Canvas render → Mark detection / region processing → Preview → PNG export
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The implementation is mostly standard Canvas API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;loadImageFromFile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;File&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;ImageBitmap&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;image/&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Please upload a valid image file.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;createImageBitmap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;drawToCanvas&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;bitmap&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ImageBitmap&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;HTMLCanvasElement&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;canvas&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;document&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createElement&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;canvas&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;2d&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;width&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;bitmap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;width&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;height&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;bitmap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;height&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;drawImage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;bitmap&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;exportAsPng&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;HTMLCanvasElement&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Blob&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;canvas&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toBlob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;blob&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;blob&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Export failed.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))),&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;image/png&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing glamorous. But getting upload → preview → export to feel seamless is the actual product. A clever removal algorithm doesn't help much if the UI is janky.&lt;/p&gt;




&lt;h2&gt;
  
  
  The hard part isn't removing the mark
&lt;/h2&gt;

&lt;p&gt;Removing a visible overlay sounds simple. Detect the region, patch it. Done.&lt;/p&gt;

&lt;p&gt;The problem is what the mark is sitting on top of.&lt;/p&gt;

&lt;p&gt;Watermarks land on gradients, skin tones, compressed JPEG noise, AI-generated texture, dark backgrounds with subtle detail. If the patch looks blurry or slightly wrong, users notice immediately — even if they can't articulate why.&lt;/p&gt;

&lt;p&gt;The real goal isn't "remove the mark." It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Make the processed area look like nothing happened.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sounds obvious, but it pushes against a common temptation: over-processing. A lot of image tools try to "fix" things they weren't asked to fix — smooth skin, sharpen edges, boost contrast. I specifically didn't want that. The best result for this workflow is boring. The image should look untouched except for the mark being gone.&lt;/p&gt;




&lt;h2&gt;
  
  
  Scope as a feature
&lt;/h2&gt;

&lt;p&gt;The first version only targets the visible Gemini overlay. Not every watermark on the internet, not arbitrary logos, not text burns.&lt;/p&gt;

&lt;p&gt;That focus does three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The UI doesn't need a "configure the region" step — common case just works&lt;/li&gt;
&lt;li&gt;The processing logic can be tuned around a known pattern&lt;/li&gt;
&lt;li&gt;Users with the specific problem immediately understand what the tool is&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One of the more useful mental shifts I've had with small tools: &lt;strong&gt;narrow products are easier to trust.&lt;/strong&gt; If a tool claims to do everything, I'm skeptical. If it claims to do one thing and does it well, I'll actually use it.&lt;/p&gt;




&lt;h2&gt;
  
  
  A note on what this tool doesn't do
&lt;/h2&gt;

&lt;p&gt;Google's Gemini images also carry &lt;strong&gt;SynthID&lt;/strong&gt; — an invisible, embedded watermark for AI provenance tracking. This tool doesn't touch that. It's about the visible overlay in your export, not invisible cryptographic signatures baked into the pixel data.&lt;/p&gt;

&lt;p&gt;Worth being explicit: this is for images you own, generated, or have permission to edit. Not for stripping attribution or bypassing content transparency systems.&lt;/p&gt;




&lt;h2&gt;
  
  
  The browser-first model and what it means for the business
&lt;/h2&gt;

&lt;p&gt;Running in the browser keeps infrastructure costs low, which matters a lot for an indie project. No per-image compute, no storage costs, no deletion policy to maintain.&lt;/p&gt;

&lt;p&gt;It also clarifies where paid features make sense: batch processing, higher-volume workflows, and any future server-side features can live in a paid tier. The free tool can stay fast and simple without subsidizing heavy usage.&lt;/p&gt;

&lt;p&gt;That's a cleaner model than gating the core utility behind an account from day one.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;The things I want to improve are mostly at the edges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Better handling of images with complex backgrounds where the mark overlaps important detail&lt;/li&gt;
&lt;li&gt;Mobile UX — Canvas processing on mobile can be slow and I haven't optimized it properly yet&lt;/li&gt;
&lt;li&gt;A before/after slider that's actually good (the current one is functional, not great)&lt;/li&gt;
&lt;li&gt;Some kind of quality indicator so the user knows when a result is uncertain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The core promise stays the same: upload, clean, download, move on.&lt;/p&gt;




&lt;h2&gt;
  
  
  The actual takeaways
&lt;/h2&gt;

&lt;p&gt;Three things I'd say to anyone building something like this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Constraints are productive.&lt;/strong&gt; Deciding not to build something is a real engineering decision. It's usually the right one on v1.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Runs in your browser" is a feature, not a footnote.&lt;/strong&gt; Privacy-by-default is something users care about, and it's worth building around intentionally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The pipeline matters as much as the algorithm.&lt;/strong&gt; A tool can have a solid core and still feel terrible if the upload, preview, and export experience is rough. Get those right first.&lt;/p&gt;




&lt;p&gt;You can try it at &lt;strong&gt;&lt;a href="https://geminiwatermarkremover.ai/" rel="noopener noreferrer"&gt;geminiwatermarkremover.ai&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Happy to hear from other developers building small AI workflow tools — what problems are you solving, and what trade-offs did you end up making?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
