<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Saihhold Zhao</title>
    <description>The latest articles on DEV Community by Saihhold Zhao (@saihhold_zhao_064664f95f9).</description>
    <link>https://dev.to/saihhold_zhao_064664f95f9</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1833208%2F5916e4a5-1831-4c51-ac39-c6d329430228.jpg</url>
      <title>DEV Community: Saihhold Zhao</title>
      <link>https://dev.to/saihhold_zhao_064664f95f9</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saihhold_zhao_064664f95f9"/>
    <language>en</language>
    <item>
      <title>Annotation &amp; Comment: A Simpler Approach to AI Image Editing</title>
      <dc:creator>Saihhold Zhao</dc:creator>
      <pubDate>Wed, 23 Sep 2026 05:28:13 +0000</pubDate>
      <link>https://dev.to/saihhold_zhao_064664f95f9/annotation-comment-a-simpler-approach-to-ai-image-editing-1058</link>
      <guid>https://dev.to/saihhold_zhao_064664f95f9/annotation-comment-a-simpler-approach-to-ai-image-editing-1058</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Most AI image editing tools today rely on &lt;strong&gt;text-only prompts&lt;/strong&gt; to describe what needs to be changed and how it should be modified.&lt;/p&gt;

&lt;p&gt;However, when an editing task becomes more complex, writing prompts can quickly become cumbersome. More importantly, text alone is sometimes not precise enough to indicate the exact area of an image that needs to be edited.&lt;/p&gt;

&lt;p&gt;For example, suppose I want to make several changes to the image below:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Make the model hold a specific object&lt;/li&gt;
&lt;li&gt;Change the color of her shoes to red&lt;/li&gt;
&lt;li&gt;Replace her outfit with a specific piece of clothing&lt;/li&gt;
&lt;li&gt;Change the background environment to a café&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With the traditional text-only approach, I might need to write a prompt like this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Make the model raise her hand and hold the coffee cup from Image 1, change her shoes to red, replace her outfit with the clothing from Image 2, and change the background to a café.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwajg8lnatxtl97jpbwa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwwajg8lnatxtl97jpbwa.png" alt=" " width="799" height="502"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Annotation &amp;amp; Comment Editing
&lt;/h2&gt;

&lt;p&gt;I simplified this process with a new interaction method.&lt;/p&gt;

&lt;p&gt;Instead of writing a long prompt, users only need to &lt;strong&gt;place markers directly on the image and briefly describe the desired change in the corresponding input fields&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This makes it possible to tell the model exactly &lt;strong&gt;where&lt;/strong&gt; a change should happen through visual annotations, while using short text instructions to describe &lt;strong&gt;what&lt;/strong&gt; should be changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Implementation and Real-World Testing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Choose an Image Editing Model
&lt;/h3&gt;

&lt;p&gt;So far, this method has worked well with the following image editing models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT Images 2.5 Flare / Sunburst&lt;/li&gt;
&lt;li&gt;Nano Banana 2&lt;/li&gt;
&lt;li&gt;Seedream 5.0 Pro&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Based on these tests, I believe the same approach is also worth testing with newer generations of image editing models released after them.&lt;/p&gt;

&lt;p&gt;The following example demonstrates an actual test using &lt;strong&gt;GPT Images 2.5 Sunburst&lt;/strong&gt; in &lt;strong&gt;PoloX AI&lt;/strong&gt; &lt;a href="https://polox.ai" rel="noopener noreferrer"&gt;PoloX AI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The implementation is open source and available on GitHub:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/saihhold-zhao/polox_ai" rel="noopener noreferrer"&gt;https://github.com/saihhold-zhao/polox_ai&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Use the Image Annotation Edit Skill
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Trigger the Skill&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Use:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/annotated-image-edit&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;to trigger the &lt;strong&gt;Image Annotation Edit Skill&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7j5g3z3rb1lgszbppr0l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7j5g3z3rb1lgszbppr0l.png" alt=" " width="799" height="346"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Enter Annotation Editing Mode&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Select &lt;strong&gt;Annotation Editing Mode&lt;/strong&gt; to open the image annotation tool.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Add Markers and Editing Instructions&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Place a marker directly on the part of the image you want to modify, then enter a short instruction in the corresponding input field.&lt;/p&gt;

&lt;p&gt;For example, if you want to change the shoes to red:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Place a marker on the shoes → Enter “red”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's it.&lt;/p&gt;

&lt;p&gt;If you want to use a reference image, you can upload it directly in the corresponding input field.&lt;/p&gt;

&lt;p&gt;For example, if you want the model to hold a specific coffee cup:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Place a marker on her hand → Enter “raise her hand and hold this” → Upload the coffee cup reference image&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This creates a clear relationship between the target location, the editing instruction, and the reference image.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4a93vmqfk8i94emxbe3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4a93vmqfk8i94emxbe3.png" alt=" " width="800" height="726"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Let the Agent Analyze the Annotations&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once all annotations are complete, click Continue.&lt;/p&gt;

&lt;p&gt;A vision-language model analyzes all of the annotations, and the Agent automatically prepares the parameters and prompt required by the image editing model.&lt;/p&gt;

&lt;p&gt;There is one important implementation detail:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Agent sends two images to the image editing model: the original image and a second version containing the visual annotations.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This means there is no need to worry about annotation markers covering parts of the original image.&lt;/p&gt;

&lt;p&gt;The model can use the annotated image to understand &lt;strong&gt;where&lt;/strong&gt; the edits should be applied, while still having access to the clean original image for complete, unobstructed visual information.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6dz6mpmu75ifbxhybh3t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6dz6mpmu75ifbxhybh3t.png" alt=" " width="800" height="461"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Get the Result&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;After these steps, the image editing model generates the final result based on the editing instructions prepared by the Agent.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm9hwa67my3bjpxnb56u3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm9hwa67my3bjpxnb56u3.png" alt=" " width="800" height="1299"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;This &lt;strong&gt;Annotation + Comment&lt;/strong&gt; approach significantly simplifies the AI image editing workflow.&lt;/p&gt;

&lt;p&gt;Instead of writing complicated prompts, users can directly indicate &lt;strong&gt;where they want to make a change&lt;/strong&gt; and briefly describe &lt;strong&gt;what they want to change&lt;/strong&gt;. The Agent then handles prompt construction and parameter configuration automatically.&lt;/p&gt;

&lt;p&gt;You can deploy PoloX AI locally and test the implementation here:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/saihhold-zhao/polox_ai" rel="noopener noreferrer"&gt;https://github.com/saihhold-zhao/polox_ai&lt;br&gt;
&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>tutorial</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
