<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Merl Merl</title>
    <description>The latest articles on DEV Community by Merl Merl (@merl985).</description>
    <link>https://dev.to/merl985</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4028925%2F309a350a-b9f4-41fb-b074-88d8a67eb019.png</url>
      <title>DEV Community: Merl Merl</title>
      <link>https://dev.to/merl985</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/merl985"/>
    <language>en</language>
    <item>
      <title>One Reference Image, Eight Character Sheet Runs, and What the Model Made Up</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Tue, 01 Sep 2026 17:48:55 +0000</pubDate>
      <link>https://dev.to/merl985/one-reference-image-eight-character-sheet-runs-and-what-the-model-made-up-1h7p</link>
      <guid>https://dev.to/merl985/one-reference-image-eight-character-sheet-runs-and-what-the-model-made-up-1h7p</guid>
      <description>&lt;p&gt;Character sheets are a reference format. Front, side, and back views, a range of expressions, a set of poses, enough detail that another person could work from the sheet without asking you questions.&lt;/p&gt;

&lt;p&gt;The interesting engineering question is what happens when the input is underspecified. One image of a character does not contain the back of the coat. It does not contain the object hanging off a strap that runs behind the hip. Ask a model to expand that image into a full sheet, and it has to produce values for fields that were never set.&lt;/p&gt;

&lt;p&gt;So I ran it as a controlled test on &lt;a href="https://eap.pixai.art/go/balazs2" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, the model PixAI currently has in early access, and measured what came back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test setup
&lt;/h2&gt;

&lt;p&gt;One original character, generated once from text, then used as the only reference for every run after it. Eight generations: a turnaround, the same turnaround again on a different seed, a third with one sentence added, a six-panel expression sheet, a nine-panel expression sheet, a four-pose sheet, a single pose, and an outfit sheet.&lt;/p&gt;

&lt;p&gt;Every setting held constant across runs: style preset off, Pro mode, Prompt Helper off, default negative prompt untouched, square format, one reference image attached.&lt;/p&gt;

&lt;p&gt;The character design is the instrument, and it splits into three groups.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NAMED IN EVERY PROMPT      hair length and color, eye color,
                           red scarf, olive coat over grey sweater

REFERENCE ONLY             which side the hair tucks behind the ear,
                           which side the scarf tail hangs on,
                           strap across the chest with a metal slider,
                           notebook in the left chest pocket

ABSENT FROM THE REFERENCE  the back of the coat
                           the object on the end of the strap
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first group is the control. If those drift, nothing else is worth reading. They did not drift in any of the eight runs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2o668mfszzyz3msyodw.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2o668mfszzyz3msyodw.jpg" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 1: silence is deterministic, promises are not
&lt;/h2&gt;

&lt;p&gt;The turnaround came back with three genuine views at a matched scale: 983, 978, and 978 pixels tall in a 1024 pixel frame. The back view had to be invented, since the reference never showed it. It came out as a plain quilted panel with a center seam.&lt;/p&gt;

&lt;p&gt;Then I ran the identical prompt, byte for byte, on a different seed. Same invented back. Same center seam.&lt;/p&gt;

&lt;p&gt;That is the useful finding. An unset field gets a default, and the default is stable enough to rely on.&lt;/p&gt;

&lt;p&gt;The strap behaves differently. The reference shows it crossing the chest and running to something behind the hip. Both turnaround runs dropped it entirely from the back view rather than resolving it, on two different seeds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffq4ivwehvhpoyb8ppdi6.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffq4ivwehvhpoyb8ppdi6.jpg" alt=" " width="800" height="779"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For the third run, I held the first run's seed and appended one sentence naming the bag. The bag appeared, correct material, correct hip. Two things came with it: a second strap crossing the first on the back, which no single shoulder bag produces, and the disappearance of the strap from the front and side views of that same sheet.&lt;/p&gt;

&lt;p&gt;One added constraint, satisfied at the point of the constraint, paid for elsewhere in the same output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 2: panel count is a real budget
&lt;/h2&gt;

&lt;p&gt;Six expressions in a two-by-three grid was the strongest output in the set. All six visually distinct, the face measuring between 243 and 260 pixels wide across the panels, and a metal slider a few pixels across surviving at roughly 341 by 512 pixels per portrait.&lt;/p&gt;

&lt;p&gt;Nine expressions in a three by three grid did something I did not predict. The faces stayed on model. The framing simplified: every panel squared up to the camera, the hair fell symmetrically, and a reference-only detail that held through all six panels of the smaller sheet vanished from all nine of the larger one. Two emotion pairs also converged, so nine slots produced seven clearly separate faces.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result 3: the sheet format itself costs resolution
&lt;/h2&gt;

&lt;p&gt;The pose sheet held proportions and broke layout. Four requested poses produced three distinct ones, and the crouching figure landed in the gap between two others at a different camera height instead of on the row.&lt;/p&gt;

&lt;p&gt;More useful was running the same crouch as a single figure in the full frame. The notebook in the chest pocket renders at 71 pixels wide there against 32 on the sheet, with the cover, page block and spine thickness all readable. The reaching hand resolved into five separated fingers.&lt;/p&gt;

&lt;p&gt;A quarter of a 1024 pixel frame has no room for that. The sheet was not failing at drawing; it was failing at budget.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiynma9kgn4w2aivnywk1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiynma9kgn4w2aivnywk1.jpg" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What survived, ranked
&lt;/h2&gt;

&lt;p&gt;Across every figure, the reference-only details came out in a clear order.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;notebook in the chest pocket   25 of 26 figures wearing the coat
strap across the chest         23 of 26 front and side views
scarf tail on one side         10 of 17 full body figures
hair tucked behind one ear     14 of 26 views that face the viewer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ordering is not random. The notebook is a high-contrast object in a fixed place on the garment, and every format keeps that place in frame. The strap depends on which side of the body is visible. The hair parting has no object holding it, so it goes first.&lt;/p&gt;

&lt;p&gt;One caveat that cuts across all of it. The notebook and the strap are missing from the outfit sheet completely, because the coat they live on was replaced. Identity carried by an object lasts exactly as long as the object does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways if you try this
&lt;/h2&gt;

&lt;p&gt;Design for the format. Put the identifying details on objects, keep those objects on a garment you plan to keep, and treat anything that lives purely in the silhouette as the first thing you will lose.&lt;/p&gt;

&lt;p&gt;Run the turnaround twice before you trust the back. The second run tells you whether the invented answer is a stable default or a coin flip.&lt;/p&gt;

&lt;p&gt;If a detail matters, generate it as a single image rather than a panel. The overview and the detail are different jobs.&lt;/p&gt;

&lt;p&gt;One last thing. Treat the output as a working reference rather than a production model sheet. Everything on it that your original image never showed is a plausible guess, which is a different thing from correct.&lt;/p&gt;

&lt;p&gt;If you have one good picture of your OC sitting in a folder, the whole experiment costs you a few generations. &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;Try it on PixAI&lt;/a&gt; and see which of your character's details are carried by an object and which are carried by luck.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>Adding an Instruction Moved the Bug Instead of Fixing It</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Mon, 31 Aug 2026 19:35:52 +0000</pubDate>
      <link>https://dev.to/merl985/adding-an-instruction-moved-the-bug-instead-of-fixing-it-1m90</link>
      <guid>https://dev.to/merl985/adding-an-instruction-moved-the-bug-instead-of-fixing-it-1m90</guid>
      <description>&lt;p&gt;Image models fail in a way that is hard to debug: the output looks fine. There is no stack trace, no assertion error, just a picture that is plausible and quietly different from what you asked for.&lt;/p&gt;

&lt;p&gt;So I built a test fixture and ran nine controlled generations against it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fixture
&lt;/h2&gt;

&lt;p&gt;One scene, described in exactly the same words in every prompt, with four objects placed so that I could check each one afterwards:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- red toolbox      -&amp;gt; viewer's left of the character
- metal ladder     -&amp;gt; viewer's right of the character
- ticket booth     -&amp;gt; further back, behind the wheel
- the character    -&amp;gt; standing at the foot of the Ferris wheel
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model under test was &lt;a href="https://eap.pixai.art/go/balazs2" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt; on PixAI. Two rules kept the runs comparable. Only the camera sentence changed between prompts, and it came first, since these models read a prompt roughly in order of importance. Every test ran as a single generation instead of a batch, because selecting the best of four measures what the model can produce on the fourth attempt, which is a different question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The baseline
&lt;/h2&gt;

&lt;p&gt;The first prompt named no camera at all, and I ran it four times. All four came back the same way: eye level, character centered, full body, wheel behind him at roughly equal width on both sides.&lt;/p&gt;

&lt;p&gt;The left and right assignments held in all four images. The booth, described as standing behind the wheel, ended up beside it every time. That split showed up in every later run too, so it is worth stating early: &lt;strong&gt;left and right are properties of the frame, behind is a property of the camera&lt;/strong&gt;, and only one of those two moves when the lens does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 1: camera height
&lt;/h2&gt;

&lt;p&gt;Two prompts, identical except for one sentence. One asked for an extreme low angle with the camera almost on the concrete, the other for a bird's-eye view from the top of the wheel.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjxorx4e84wzd0vzmg4za.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjxorx4e84wzd0vzmg4za.jpg" alt=" " width="800" height="552"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Both moved the lens instead of bending the pose. Boots large and close, horizon low, base legs splaying outward in the first. Hard hat crown dominant, legs foreshortened, structural shadows consistent with the viewpoint in the second.&lt;/p&gt;

&lt;p&gt;The failure here is partial. In the low-angle run, the steel base directly behind the character reacted to the new viewpoint, while the wheel itself stayed a clean frontal circle with evenly spaced spokes, exactly as it looked at eye level. In the high-angle run, the same wheel turned correctly, its rim running out of both sides of the frame as two curves. Near geometry followed the camera, while distant geometry sometimes kept the view it already had.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 2: the regression
&lt;/h2&gt;

&lt;p&gt;The depth prompt asked for three layers: toolbox very close to the lens, character in the middle distance beside the ladder, booth far behind him and much smaller than he is.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzkcbvkpzw9ewift6rp5.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzkcbvkpzw9ewift6rp5.jpg" alt=" " width="800" height="615"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Result of run one:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assertion&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;toolbox in the immediate foreground, large&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;character readable in the middle distance&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;booth far back, much smaller than him&lt;/td&gt;
&lt;td&gt;pass, around 40 percent of his height&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;character stands beside the ladder&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;fail&lt;/strong&gt;, ladder moved to the booth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One failing assertion, so I patched the prompt. I added two phrases and changed nothing else: the booth got a frame position, near the right edge of the frame, and the character got one too, in the center of the frame.&lt;/p&gt;

&lt;p&gt;Result of run two:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assertion&lt;/th&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;character stands beside the ladder&lt;/td&gt;
&lt;td&gt;pass, fixed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;booth near the right edge of the frame&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;booth far back, half his height&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;fail&lt;/strong&gt;, around 110 percent of his height&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;character in the center of the frame&lt;/td&gt;
&lt;td&gt;partial, sits right of the center line&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The patch fixed the failing assertion and broke a passing one. The booth moved to the right edge as instructed and came forward while doing it, its awning now above the character's hard hat. Two runs of the same scene, two different failure profiles, and the second was not obviously better than the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 3: the lens is a separate control
&lt;/h2&gt;

&lt;p&gt;Fisheye has two parts. The lens has to bend the image, and the camera has to be somewhere specific for the bend to make sense. I asked for both: a fisheye shot from directly beneath the wheel looking straight up.&lt;/p&gt;

&lt;p&gt;The distortion arrived, and it is convincing, with the concrete apron bowing into a wide arc. The camera position was ignored, and what came back has two viewpoints in one frame. The character is drawn from above, crown of the hard hat toward the lens, legs foreshortened, shadow pooled beneath him. The wheel behind him is drawn from below, base legs splaying downward, top edge tilting away.&lt;/p&gt;

&lt;p&gt;The lens effect applied to the whole frame, while the camera position resolved separately for the figure and for the structure behind him.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 4: five constraints in one prompt
&lt;/h2&gt;

&lt;p&gt;The last prompt stacked five requirements: low camera, toolbox in the immediate foreground on the viewer's left, character in the middle distance climbing the ladder with his back to the camera, an older man in a green coverall further back on the viewer's right holding a clipboard and looking up, wheel filling the background.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7uron0bbwr5p4xqbrdmm.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7uron0bbwr5p4xqbrdmm.jpg" alt=" " width="768" height="1280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All five landed, including the back-turned pose, which I expected to be the weak point since turning a character away removes the face.&lt;/p&gt;

&lt;p&gt;The busiest prompt in the test produced the cleanest result. The wording difference is that every element here carried a frame side alongside its distance, while the depth prompt named only layers. Whether the frame side is what carried the result, two runs cannot tell me.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the nine runs support
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Side of the frame:&lt;/strong&gt; honored in every run where it appeared, including ground level, overhead, and fisheye.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Camera height:&lt;/strong&gt; three clear results out of four, with the overhead run understating the distance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lens distortion:&lt;/strong&gt; present in both runs that asked for it, independent of camera position.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Depth, near layer:&lt;/strong&gt; landed in all three prompts that asked for one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Depth, far layer:&lt;/strong&gt; correct in two runs, slid forward in the third.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Size ratio between two objects:&lt;/strong&gt; returned as written in neither of the two runs that asked for one. Much smaller than he is produced 40 percent; half his height produced 110 percent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The camera is controllable on this model. The spatial relationships between objects are not deterministic across runs of the same prompt, and adding instructions to fix one relocated the variance rather than eliminating it. If you are building a pipeline on top of this, treat camera position as a parameter and object-to-object geometry as something you verify per output.&lt;/p&gt;

&lt;p&gt;If you want to run the same experiment, the method is cheap: fix three or four nameable objects in your scene, write down where each one belongs, change one sentence per run, and check the output against your list instead of against your impression of it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;Try your own composition challenge on PixAI&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>Moving a Value Is Two Operations, and One of Them Is a Delete</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Wed, 26 Aug 2026 18:31:24 +0000</pubDate>
      <link>https://dev.to/merl985/moving-a-value-is-two-operations-and-one-of-them-is-a-delete-a8m</link>
      <guid>https://dev.to/merl985/moving-a-value-is-two-operations-and-one-of-them-is-a-delete-a8m</guid>
      <description>&lt;p&gt;An image editor that takes natural language looks like a function call with named arguments. You describe a target, an operation, and a set of things that should stay put, and something comes back. The interesting failures are the ones where every argument you passed is present in the output, and the result is still wrong.&lt;/p&gt;

&lt;p&gt;Here is the smallest example I have. The instruction was to move an object from A to B. The output contained the object at B. It also still contained the object at A. Every requirement satisfied, one implicit requirement missed: a move is a copy plus a delete, and the delete is the half nobody writes down.&lt;/p&gt;

&lt;p&gt;I spent a session testing which implicit requirements PixAI's Tsubaki.3 model infers on its own. Eight edits, four categories, one fixture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test fixture
&lt;/h2&gt;

&lt;p&gt;A single generated scene, built so every element is countable and nameable afterwards:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;subject:   greenhouse gardener
worn:      one yellow rubber glove, left hand only
held_by:   crow -&amp;gt; the matching right glove, in its beak
markers:   4 wooden plant labels, red shears in one pocket,
           straw hat on a cord down her back
bench:     tipped terracotta pot (centre), watering can (right edge)
control:   second terracotta pot, same clay, never named in any instruction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the important one. Two objects of the same material, only one of them ever mentioned, gives you a free assertion on every material edit.&lt;/p&gt;

&lt;p&gt;Fixed parameters across all runs: Pro mode, no style preset, default negative prompt untouched, prompt rewriting disabled, one edit per run from the same source. Free seeds, so single runs are observations rather than proofs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft9gp70lc7uffveifnm2c.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft9gp70lc7uffveifnm2c.jpg" alt=" " width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Category 1: reference reassignment
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Move the yellow rubber glove from the crow's beak onto her bare right hand,
so that she is now wearing a glove on both hands.
The crow's beak is empty and open.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Passed on all three assertions: object present at destination, absent at origin, and bound by the right relation (worn, rather than held). The duplicate-instance failure never appeared in this run.&lt;/p&gt;

&lt;p&gt;The second spatial test asked for containment with specified occlusion: the crow inside the tipped pot, only its head visible, legs and body and twine hidden.&lt;/p&gt;

&lt;p&gt;The bird went in tail-first, which inverts the occlusion spec exactly. Two unrequested mutations came with it: the pot rotated to face the opposite direction, and the twine unbound from the crow's leg and rebound around the pot. Both are consistent with making a tail-first insertion physically renderable at that camera angle. Read charitably, the model mutated whatever was not explicitly frozen until the requested relation became satisfiable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Category 2: change the type, keep the instance
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Change the terracotta pot she is holding to clear transparent glass,
so that the soil and tangled roots inside become visible.
The pot keeps exactly the same shape, size, and position in her hand.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assertion&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Material changed to glass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interior contents invented plausibly&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Silhouette preserved&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;fail&lt;/strong&gt;, returned a straight-sided jar&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No parts added&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;fail&lt;/strong&gt;, gained a metal screw lid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Control object unchanged&lt;/td&gt;
&lt;td&gt;pass, second pot still terracotta&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The control assertion is the useful one. Material change scoped to the named instance rather than to the type, which is the behaviour you want in a cluttered frame.&lt;/p&gt;

&lt;p&gt;The silhouette failure repeated on the second material edit in a different direction. Asked to turn a straw hat on the character's back into hammered copper while keeping its shape and cord, the model delivered the copper and the brim, then relocated the hat onto her head and swapped the thin cord for a thick knotted rope. A single argument about material mutated a spatial relation that was never in the call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Category 3: one instruction, derived consequences
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Show this same greenhouse one second after the whole potting bench
tipped over onto the floor.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No consequences listed on purpose. The dependency chain resolved without being told: the bench is overturned and everything that had been resting on it is on the floor with the soil spilled.&lt;/p&gt;

&lt;p&gt;Garbage collection was less tidy. The crow is gone from the scene with no trace. A second pair of shears appeared on the floor while the original instance stayed in the character's pocket. The watering can landed upright and undamaged. And with nothing airborne, the render reads as an aftermath state rather than the requested t+1s frame.&lt;/p&gt;

&lt;h2&gt;
  
  
  Category 4: text and layout
&lt;/h2&gt;

&lt;p&gt;Three string operations plus one physical dependency, since the crow was perched on the title being replaced.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assertion&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Title replaced, exact string&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Date line added below, exact string&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Badge added with exact string&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Size hierarchy preserved&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Crow re-perched on the new title&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Badge placed in the named corner&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;fail&lt;/strong&gt;, landed below the title on the right&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Surrounding layout preserved&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Worth noting from the generation side: producing the source poster took three attempts. The first two used a compound word in the display title, and across eight images none came back usable, with the same stray consonant inserted between the two halves every time all the letters were legible. The smaller caption line rendered correctly on every attempt. Text accuracy in generation and text accuracy in editing behaved like separate subsystems.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvuox5f47hx074p6b7056.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvuox5f47hx074p6b7056.jpg" alt=" " width="800" height="548"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Does adding constraints help?
&lt;/h2&gt;

&lt;p&gt;Two of the failures got a second pass with the missing constraint written out.&lt;/p&gt;

&lt;p&gt;Pinning the hat's position worked completely. One clause stating that the hat stays on her back was enough to keep it there, keep the cord thin, and leave both control pots alone.&lt;/p&gt;

&lt;p&gt;Pinning the pot's shape produced a mixed result. Naming the lid removed the lid. Naming the root direction fixed the root direction. Naming the category ("do not turn it into a jar") and describing the overall form did nothing, and the vessel came back as a jar with a threaded neck. The blast radius also widened rather than narrowing: on that run the control pot turned to glass as well.&lt;/p&gt;

&lt;p&gt;The pattern across those two, at one run per variant: constraints naming a relation or a specific removable part held, constraints naming a class or a whole shape did not, and a longer defensive instruction bought no extra protection for anything else in the frame.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would take into a workflow
&lt;/h2&gt;

&lt;p&gt;Relations between separate objects are safe to attempt in one pass, including the delete half of a move. Strings inside an existing design are safe. Type changes on an object need verification every time, and if two objects of the same material sit in one frame, check both.&lt;/p&gt;

&lt;p&gt;The one-line version: name the relation you care about, including the one you want left alone, and put your assertions on the objects rather than on the scene.&lt;/p&gt;

&lt;p&gt;If you want to run your own version, the cheapest possible test is two similar objects in one image and an instruction that names exactly one of them. Try your own advanced editing scenario with Tsubaki.3 on PixAI: &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;https://eap.pixai.art/go/balazs&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>Measuring the Blast Radius of an AI Image Edit</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Mon, 24 Aug 2026 19:31:25 +0000</pubDate>
      <link>https://dev.to/merl985/measuring-the-blast-radius-of-an-ai-image-edit-cam</link>
      <guid>https://dev.to/merl985/measuring-the-blast-radius-of-an-ai-image-edit-cam</guid>
      <description>&lt;p&gt;Instruction-based image editing has a failure mode that is easy to describe and hard to pin down: the requested change lands, and something you never mentioned comes back different. The face is a little off, an accessory is gone, the background warmed up by a few degrees.&lt;/p&gt;

&lt;p&gt;The usual way to evaluate this is to look at the output and form an impression. I wanted numbers instead, so I set up something closer to a controlled experiment against PixAI's Tsubaki.3 and ran nine edits, the four main tests all starting from the same source image.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;p&gt;One source image, held constant. Same model, same mode, no style preset, no LoRA, prompt rewriting disabled, default negative prompt untouched, and a pinned seed. Once a base image is loaded the panel exposes no strength control, so there was nothing else to hold fixed.&lt;/p&gt;

&lt;p&gt;The source was built as a test fixture rather than as a picture. Every identifying detail sits on one side only, which turns drift into something you can count:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;left ear      silver hoop          right ear     nothing
right wrist   watch                left wrist    bare
right side    stethoscope head     left arm      two bandage squares
table         ginger tabby cat, white chest patch, one white front paw, red collar
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cat is the control. It sits in the middle of the frame, adjacent to everything I was about to modify, and exactly one of the nine instructions mentions it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measurement
&lt;/h2&gt;

&lt;p&gt;For each result, I compared regions against the same regions of the source and took the mean per-pixel difference. Calibration matters more than the absolute values here:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Region type&lt;/th&gt;
&lt;th&gt;Value when untouched&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Flat wall&lt;/td&gt;
&lt;td&gt;1 to 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Detailed area, table edge&lt;/td&gt;
&lt;td&gt;3 to 6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A region the edit changed&lt;/td&gt;
&lt;td&gt;tens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0u4koev1yskyt74ok55w.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0u4koev1yskyt74ok55w.jpg" alt=" " width="800" height="1056"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One instruction.&lt;/strong&gt; Recolor a garment. Face 3.2, cat fur 5.0, cat eyes 3.8, watch 5.6, bandage 2.4, background 1.6. The change stayed inside the garment it named.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One structural instruction.&lt;/strong&gt; Replace the outfit entirely, which forces the model to rebuild shape rather than repaint color. Stethoscope survived on top of the rebuilt torso. Both bandages survived. Cat at 4.5. Face at 10.1, against 3.2 in the recolor, with nothing in the prompt referring to the face.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three instructions in one prompt.&lt;/strong&gt; Cardigan color, held object swap, facial expression. All three landed. I then ran the same prompt twice more, once with a list of elements to preserve and once with that list plus the cat named by its markings. Whole-frame difference between the three variants: 1.4 to 2.6, which is background noise. The preservation list changed nothing measurable, because nothing in that run was under threat.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ab0uie9zpsimgqyf9kf.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ab0uie9zpsimgqyf9kf.jpg" alt=" " width="799" height="297"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Four instructions plus an environment change.&lt;/strong&gt; New location, new light source, new jacket, new held object, and a rendered sign reading ROOM 2. All four landed and the lettering came out clean. Color did not survive. The cat's fur went from an average RGB of (233, 181, 122) to (142, 149, 148), collapsing the channel spread from 111 to 7. Skin turned bluish grey and hair went near black, on a run whose preservation list named the face and hair explicitly, as most of the other runs did too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The finding
&lt;/h2&gt;

&lt;p&gt;Instruction count was the wrong variable. One instruction and four were carried with the same accuracy, and nothing was ignored across nine edits.&lt;/p&gt;

&lt;p&gt;What scaled was reach. The disturbed area tracked how much of the picture the instruction obliged the model to reconstruct:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;recolor a garment      -&amp;gt; stops at the garment edge
rebuild an outfit      -&amp;gt; reaches the face
change the light       -&amp;gt; reaches every surface whose color depends on light
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The preservation list follows the same curve. At one instruction it had nothing to save. At three it made no measurable difference. At four plus an environment change it named the face and hair and lost both. It costs nothing to include and it stops being sufficient at the point where you start needing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chaining is worse than batching
&lt;/h2&gt;

&lt;p&gt;The obvious instinct is to split a complex edit into steps. I ran the three-change instruction as three sequential edits, each starting from the previous result, and measured how broken up the flat wall behind the subject became:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Route&lt;/th&gt;
&lt;th&gt;Wall blockiness&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Source&lt;/td&gt;
&lt;td&gt;0.36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One prompt, three changes&lt;/td&gt;
&lt;td&gt;0.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chained, step 1&lt;/td&gt;
&lt;td&gt;0.43&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chained, step 2&lt;/td&gt;
&lt;td&gt;1.21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chained, step 3&lt;/td&gt;
&lt;td&gt;7.35&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;By step three the breakup had spread onto the subject and the cat. Every pass re-encodes the whole image, and the artifacts compound in flat regions first. Instruction following survived the whole chain, and image quality did not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr55sfc4ojgoh8g24i4ya.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr55sfc4ojgoh8g24i4ya.jpg" alt=" " width="800" height="633"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I ran step two twice and got byte-identical files, so the result is reproducible on these settings. Whether a different seed avoids it is untested. Worth noting for anyone reaching for a retry: with the seed pinned, rerunning changes nothing at all. Freeing the seed or rewriting the instruction are the only two levers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Batch related changes into one instruction. Four coordinated changes came through intact.&lt;/li&gt;
&lt;li&gt;Expect a structural rebuild to move the face slightly. Check it against the original.&lt;/li&gt;
&lt;li&gt;Expect a lighting or environment change to rewrite color on skin, hair and fur, and expect a preservation list to fail there.&lt;/li&gt;
&lt;li&gt;Avoid chained passes on the same file. The quality cost is real and compounds.&lt;/li&gt;
&lt;li&gt;Pin your settings before a comparison run. Half of this analysis only works because every edit ran on identical parameters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The full write-up with all the before and after images is on Medium. If you want to run the same experiment on your own file, &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;Tsubaki.3 is on PixAI&lt;/a&gt;. Take something you already like, change one small thing, and go audit the corners of the frame.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>I ran 17 controlled generations to find out where AI manga panels break</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Fri, 21 Aug 2026 11:08:29 +0000</pubDate>
      <link>https://dev.to/merl985/i-ran-17-controlled-generations-to-find-out-where-ai-manga-panels-break-32n4</link>
      <guid>https://dev.to/merl985/i-ran-17-controlled-generations-to-find-out-where-ai-manga-panels-break-32n4</guid>
      <description>&lt;p&gt;Garbled text inside generated images has been a standing complaint for as long as I have been using these tools. Recent releases claim to have fixed it, and I wanted to find out where the line actually sits.&lt;/p&gt;

&lt;p&gt;So I set up a small experiment: one character, one model, seventeen generations, and a rule that every comparison changes exactly one variable. The subject is manga panels, because a panel is the one image format that has to carry artwork and text in the same frame and make both readable. The model is Tsubaki.3 on PixAI. I cannot draw, which for this purpose is a feature: everything below comes out of the prompt, with nothing rescued by hand afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Treat the prompt as a function call with five arguments:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;panel(
  scene,       # what is happening
  character,   # who, in full, every run
  framing,     # camera position and how tight
  dialogue,    # the line, and which corner it goes in
  treatment    # panel border, screentone, speed lines
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The character argument stays identical across every run. That is what makes the rest of it measurable. Settings held constant too: no style preset, no LoRA, quality booster off, portrait 3:4, and the negative prompt left prefilled as it ships, with one deliberate exception below. Every comparison ran as a single generation rather than a batch of four, because picking the best of four is selection bias with extra steps.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg130v2s4huia37bhvf8.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg130v2s4huia37bhvf8.jpg" alt=" " width="799" height="525"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Same character description, same action, two different treatment arguments. The left is a color illustration, centered, no border, no text. The right is the manga instruction: panel border, screentone, a lower camera, and the bubble. The room is identical. The treatment is the whole delta.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding 1: the negative prompt does not need touching
&lt;/h2&gt;

&lt;p&gt;The default negative prompt on this model ships with &lt;code&gt;text&lt;/code&gt; in it. That reads like a direct conflict with a speech bubble, so I ran the same panel twice, once as-is and once with &lt;code&gt;text&lt;/code&gt; deleted.&lt;/p&gt;

&lt;p&gt;There was no difference. Both runs produced the requested line, correctly spelled, in a clean hand-lettered balloon. Requested dialogue behaves like subject matter rather than like stray artifact text, so the negative prompt is not the thing standing between you and a bubble.&lt;/p&gt;

&lt;p&gt;Then I raised the input length from two words to eight, with punctuation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"PICK ANOTHER"                              -&amp;gt; correct
"THIS ONE IS BLURRY. THE CAPTION IS WRONG." -&amp;gt; correct, both periods
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bubble occupied about five percent of the frame in both cases. The model did not resize the balloon to fit more text; it shrank the glyphs and stacked them into three rows. Longer strings cost legibility, not area.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding 2: naming the bubble's position rewrites the composition
&lt;/h2&gt;

&lt;p&gt;This is the result I did not expect, and it is the reason the whole experiment was worth running.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1rjpq17gfvcwls782iqa.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1rjpq17gfvcwls782iqa.jpg" alt=" " width="800" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three runs, one scene, one variable:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Dialogue argument&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;line only, no position&lt;/td&gt;
&lt;td&gt;bubble self-placed on empty window area, cleared the face, no tail drawn&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;line plus "upper left corner"&lt;/td&gt;
&lt;td&gt;bubble in that corner, &lt;strong&gt;and the character moved to the right side of the frame&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;B plus an explicit empty-third layout instruction&lt;/td&gt;
&lt;td&gt;requested layout delivered, character shrank, tail came out as an open line&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run B is the interesting one. I changed four words in the dialogue argument and the model re-solved the whole composition around them: figure to the right, table rotated, upper left cleared. Nothing in the prompt said where she should stand.&lt;/p&gt;

&lt;p&gt;The model treats the bubble as a layout element that has to fit, and it resolves the constraint by moving the artwork. That gives you a clean rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Name the corner. Do not also specify the layout.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Run C is what over-constraining looks like. Two instructions competing for the same rectangle, and the subject is what yields.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding 3: placement is reliable, the tail is not
&lt;/h2&gt;

&lt;p&gt;Across the full set, balloon placement never failed. The tail failed in four different ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;missing entirely&lt;/li&gt;
&lt;li&gt;doubled, two tails aimed at the same speaker&lt;/li&gt;
&lt;li&gt;drawn as a thin open line stopping in mid air, with the balloon outline broken where they should join&lt;/li&gt;
&lt;li&gt;hanging into the gap between two characters, pointing at neither&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are building anything on top of this, the tail is the element to flag for review. The balloon and the lettering held up without exception in my set.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding 4: prompt order predicts which character degrades
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz34guythco2aqw9yy861.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz34guythco2aqw9yy861.jpg" alt=" " width="800" height="1066"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two characters in one frame, described in sequence: the editor first and at length, the photographer second and more briefly. The line belongs to the photographer.&lt;/p&gt;

&lt;p&gt;Visual separation held up: pale cropped hair against a black bob, a height difference, a camera on a strap. The bubble sat on his side of the frame and the tail angled toward him, so the association worked, if faintly.&lt;/p&gt;

&lt;p&gt;The degradation was asymmetric and it repeated across two runs. Of the second character's five specified details, two failed the same way both times: glasses that were supposed to sit pushed up on his forehead came back over his eyes, and a bandage across his nose never rendered. The first-described character lost nothing.&lt;/p&gt;

&lt;p&gt;I cannot see inside the model, so I will state it as a behavior rather than a cause: the subject written later and more briefly is the one that lost details, twice, in the same two places. Write the second subject at the same specificity as the first rather than summarizing it after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding 5: a detailed description left the reference image little to do
&lt;/h2&gt;

&lt;p&gt;Character drift is the standard complaint, so I measured it. Accessories and outfit survived well across the set. The face did not: it rounded out and aged up run to run, worst in an extreme close-up where the crop leaves the model the most to invent.&lt;/p&gt;

&lt;p&gt;The expected fix is a reference image, so I generated a new scene twice, once from the description alone and once with my character sheet loaded as a reference. It changed the framing and gave her more hair. The face came out about the same either way. With a description carrying seven specified identity details, the reference had little left to correct.&lt;/p&gt;

&lt;p&gt;Hairstyle was the outlier that nothing fixed. It ranged from close to the sheet in one panel to a tight pinned-up version in another, and the reference run produced the most extreme version of all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would keep from this
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Lettering held for every line I tried, up to eight words. Budget your cleanup time elsewhere.&lt;/li&gt;
&lt;li&gt;Position the bubble by naming a corner, and let the model handle the rest of the arrangement.&lt;/li&gt;
&lt;li&gt;Review the tail on every panel. It is the least reliable element on the page.&lt;/li&gt;
&lt;li&gt;Write every subject in a multi-character prompt at full specificity, the second one included.&lt;/li&gt;
&lt;li&gt;Describe held objects with the grip stated separately. Every run that asked for a photo held between two fingers returned a raised finger and a floating photo.&lt;/li&gt;
&lt;li&gt;Compare with single generations. Batches let you pick winners and learn nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A four-panel strip also held together in one generation, with the beats in written order and set dressing holding its position across panels, which is the part that makes a sequence read as one scene.&lt;/p&gt;

&lt;p&gt;If you want to run your own version of this, the model is Tsubaki.3 and the useful discipline is changing one argument at a time: &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;try it on PixAI&lt;/a&gt; and see which of these behaviors reproduce for you.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>The Reference Carried the Count, the Prompt Carried the Layout</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Tue, 18 Aug 2026 19:52:36 +0000</pubDate>
      <link>https://dev.to/merl985/the-reference-carried-the-count-the-prompt-carried-the-layout-ep9</link>
      <guid>https://dev.to/merl985/the-reference-carried-the-count-the-prompt-carried-the-layout-ep9</guid>
      <description>&lt;p&gt;There is a piece of advice that circulates around reference-based image generation: once you attach a reference image, keep your prompt short. Describe the new scene, leave the character alone, because re-describing what the reference already shows will fight it.&lt;/p&gt;

&lt;p&gt;I wanted a number on that, so I designed a character to fail in measurable ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing a character as a test fixture
&lt;/h2&gt;

&lt;p&gt;The character is an original design called Asagi, built on PixAI with the Tsubaki.3 model. The visible design has seven components, but two of them exist purely as assertions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A count.&lt;/strong&gt; Three small brass bells on a braided green cord at his sash. Countable at a glance, and wrong answers are unambiguous.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A side.&lt;/strong&gt; A black leather arm guard on one forearm, a white cloth wrap on the opposite wrist. An asymmetric pair, so a failure shows up as either symmetry or a swap.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else in the design - the indigo haori with a mustard band, the jade green eyes, the pale ochre lock of hair, the thin scar across one eyebrow - serves as background signal. The bells and the arm pair are the assertions that either pass or fail.&lt;/p&gt;

&lt;p&gt;This matters because "does it look like the same character" is not a testable claim. "Are there three bells?" is.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test matrix
&lt;/h2&gt;

&lt;p&gt;Three conditions, same scene, same model, same settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;condition A:  prompt only          (full character description, no reference)
condition B:  prompt + reference   (full character description, reference attached)
condition C:  reference only       (scene described, character not mentioned)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scene was identical in all three: the character seated at a low table in a sunlit tea house. Condition C's prompt reads in full:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The same character sitting at a low table inside a sunlit tea house
in the daytime, a small cup in front of him, relaxed expression,
paper screens and warm wooden beams behind him, upper body and
both hands visible, anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9kizt3acvt46bb8owj3s.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9kizt3acvt46bb8owj3s.jpg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Bell count&lt;/th&gt;
&lt;th&gt;Asymmetric pair&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A, prompt only&lt;/td&gt;
&lt;td&gt;5 (and 2 on a repeat run)&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B, prompt + reference&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C, reference only&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;collapsed to matching cuffs on both arms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Condition A never got the count right across two runs. Two bells the first time, five the second. A number in a prompt behaves like a density hint rather than an integer.&lt;/p&gt;

&lt;p&gt;Condition C got the count right and lost the layout. The reference carries what a pixel-level encoder can carry: the shape and color of a bell cluster, the geometry of a face. Which forearm wears what is a relational fact, and the model resolved it toward the symmetric default.&lt;/p&gt;

&lt;p&gt;Condition B was the only pass on both assertions. The advice I started with is too broad: the reference and the description populate different fields, and supplying both costs one extra sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does stacking references help?
&lt;/h2&gt;

&lt;p&gt;Tsubaki.3 accepts up to three reference images, so the obvious follow-up is whether more references buy anything. I ran the hardest prompt in the set, a full-body low-angle shot in a storm, with one, two, and three references attached. Then I repeated the whole ladder with a second fixed seed, so the reference count was the only variable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;seed 1:  1 ref -&amp;gt; 3 bells   2 refs -&amp;gt; 3 bells   3 refs -&amp;gt; 5 bells
seed 2:  1 ref -&amp;gt; 3 bells   2 refs -&amp;gt; 3 bells   3 refs -&amp;gt; 3 bells
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Five out of six correct, and the single failure did not reproduce. Run-to-run variance is larger than any effect from the extra reference slots. One good reference did the job, and I would treat the second and third as optional rather than an upgrade path.&lt;/p&gt;

&lt;p&gt;More inputs usually means more constraint. Here it did not, and the feature invites the opposite assumption.&lt;/p&gt;

&lt;h2&gt;
  
  
  What survives a wardrobe change
&lt;/h2&gt;

&lt;p&gt;The destructive test: I moved the character out of a period haori into a bomber jacket and jeans on a modern crosswalk, which deletes every clothing-bound component of the design by definition.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl8qiv36nctxoi5168w8g.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl8qiv36nctxoi5168w8g.jpg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Surviving: hair, the ochre lock, eye color, face.&lt;br&gt;
Gone with the outfit: haori, bells, arm guard, as expected.&lt;br&gt;
Gone unexpectedly: the eyebrow scar and the cord tying his hair at the nape.&lt;/p&gt;

&lt;p&gt;The useful abstraction here is that identity splits into body-bound and clothing-bound components, and a wardrobe change is a scoped delete on the second group. If a character's recognizability is entirely stitched onto one jacket, the character has a single point of failure.&lt;/p&gt;

&lt;p&gt;The scar went missing again in a six-panel expression sheet generated from the same reference. The pattern across both: the finest facial detail drops when the output format diverges sharply from the reference, and a grid of six small heads is a long way from one figure on a street.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpd5k4qtcb1awhgdqnvh3.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpd5k4qtcb1awhgdqnvh3.jpg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes from the panel
&lt;/h2&gt;

&lt;p&gt;Three implementation details that cost me runs before I understood them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attachment order matters.&lt;/strong&gt; Attach the reference, then write the prompt. Adding it afterward reset my settings more than once.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The output inherits the reference's aspect ratio.&lt;/strong&gt; My reference was 16:9, and every reference-based generation came back wide regardless of what I set. My first prompt-only run, with nothing attached, defaulted to portrait and had to be set by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The prompt you send is not the prompt that runs.&lt;/strong&gt; PixAI displays a rewritten version beside your input, and that version had pulled a physical description of the character out of the attached image. Reading it told me which traits the system extracted and which it invented. If you are debugging an unexpected result, this field is the first place to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Character consistency is not one property. It decomposes, and the components have different failure modes and different owners:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The reference owns counts, faces, palette, silhouette.&lt;/li&gt;
&lt;li&gt;The prompt owns arrangement, sides, and everything relational.&lt;/li&gt;
&lt;li&gt;Neither owns the finest details once the output format moves far enough away.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you work with recurring characters, build one or two countable, one-sided details into the design on purpose. They cost nothing, and they turn consistency into something you can check rather than something you squint at.&lt;/p&gt;

&lt;p&gt;Try it with &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;Character Reference on PixAI&lt;/a&gt; if you have a character sitting in a folder that deserves more than one image.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>An AI Image Edit Is a Full Re-Render, and Nine Runs Showed Me What That Costs</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Sat, 15 Aug 2026 09:12:53 +0000</pubDate>
      <link>https://dev.to/merl985/an-ai-image-edit-is-a-full-re-render-and-nine-runs-showed-me-what-that-costs-p6k</link>
      <guid>https://dev.to/merl985/an-ai-image-edit-is-a-full-re-render-and-nine-runs-showed-me-what-that-costs-p6k</guid>
      <description>&lt;p&gt;Instruction-based image editing looks like a targeted operation. You hand a model an image, type "replace the cup with a bouquet", and get the image back with a bouquet in it. The mental model that suggests is a patch applied to a region.&lt;/p&gt;

&lt;p&gt;That mental model is wrong, and the cost of it being wrong shows up in details you were not looking at.&lt;/p&gt;

&lt;p&gt;I spent a week running edits against a single source image to map the behavior. The subject was an anime bike courier I generated for the purpose: mint green undercut, slate blue jacket with one orange chest stripe, a fingerless glove on one hand, three enamel pins, a red bandana on the other wrist, a coffee cup, and a chalkboard sign next to him reading OPEN. Deliberately full of small, countable details, because those are the things you can measure.&lt;/p&gt;

&lt;p&gt;Nine runs. Eight instruction edits on the same model, one masked run for comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;The edits ran on PixAI with Tsubaki.3, its newest model. Worth noting for anyone reproducing this: the source image goes into a Base Image slot that accepts exactly one file, and once it is loaded, the aspect ratio control disappears from the panel. The edit inherits the frame of the input. All eight outputs came back at the source image's exact dimensions, which makes before and after comparison trivial.&lt;/p&gt;

&lt;p&gt;One more setting matters. There is a prompt helper that rewrites or translates your input before it reaches the model. Leave it off if you want to attribute results to your own wording.&lt;/p&gt;

&lt;h2&gt;
  
  
  Discrete swaps land, structure does not
&lt;/h2&gt;

&lt;p&gt;The clean successes were attribute and object level.&lt;/p&gt;

&lt;p&gt;A hair color change from mint green to navy came back correct, with the haircut, the shaved side, the jacket, and the face intact. A hand-held object swap, coffee cup to a paper-wrapped bouquet, came back with the grip and the wrist position preserved, which is the interesting part, since hand and object are drawn together.&lt;/p&gt;

&lt;p&gt;Structural requests failed. I asked three separate times for a limb to move: twice for the gloved hand to come off the handlebar and point at the sign, once for the arm holding the cup to lift toward his chin. Nothing moved in any of the three. My first hypothesis was frame cropping, since the first arm sat half outside the image. The third attempt used an arm fully inside the frame and produced the same non-result.&lt;/p&gt;

&lt;p&gt;Background replacement landed halfway. Asking for a rainy night street produced rain streaks and warmly lit windows while the daylight on the character and the pavement stayed exactly as bright as before. The additive parts of the request arrived; the parts requiring a global relight did not.&lt;/p&gt;

&lt;p&gt;So the boundary is between local and global, and it is sharper than I expected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foxomilo4gah7tobuyd51.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foxomilo4gah7tobuyd51.jpg" alt=" " width="800" height="811"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The preservation clause is a real control, with limits
&lt;/h2&gt;

&lt;p&gt;Half of my runs carried a long clause listing what had to stay unchanged: face, hair, glove, strap, pins, bandana, sign text, background, lighting, framing, art style.&lt;/p&gt;

&lt;p&gt;The two variants of the gesture instruction looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# variant A, bare
Lift his gloved hand off the handlebar so he is pointing at the chalkboard sign.

# variant B, same instruction plus the keep list
Lift his gloved hand off the handlebar so he is pointing at the chalkboard sign.
Keep everything else unchanged: his face, his mint green undercut and the shaved
side, the silver hoop in his ear, the slate blue windbreaker and its single orange
chest stripe, the fingerless glove on that same hand, his other hand bare, the
crossbody strap, the three enamel pins on his chest, the red bandana on his wrist,
the paper coffee cup he is holding, the bike, the text on the chalkboard sign, the
shopfront and awning behind him, the lighting, the framing, and the art style.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither run performed the requested change, so everything else in the diff is drift. Without the clause, the glove changed from black to brown and his mouth opened. With it, both held.&lt;/p&gt;

&lt;p&gt;The clause narrows the drift. It does not close it. Two items proved immune in every run that carried them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The three enamel pins.&lt;/strong&gt; Count preserved in all eight runs. Colors and designs different in all eight.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The text on the chalkboard.&lt;/strong&gt; Seven runs had a goal other than the sign. All seven damaged it. Six of those seven named the sign text explicitly in the keep list.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The failure modes on the sign are worth listing, because they are not subtle: OPEN became handwritten scribbles, then handwriting with a small drawn heart, then something close to Pittal, then ANLEKN, then COFFFE, then DIUDE, then Nami with an accent over it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx4xzhcc3fbz9b5zc3sk7.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx4xzhcc3fbz9b5zc3sk7.jpg" alt=" " width="800" height="867"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The same model renders text well when text is the target. Asking for the board to read CLOSED produced correct spelling, correct placement, and a convincing chalk texture. This yields the one operational rule I would put in a runbook: &lt;strong&gt;text is the last edit in the chain&lt;/strong&gt;, because any later edit will roll it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the mask wins, quantified
&lt;/h2&gt;

&lt;p&gt;Inpainting means painting a mask over the region you want changed, so the tool knows the edit is confined to that area, then describing what belongs there. It costs you brush work, and it buys you a scope guarantee.&lt;/p&gt;

&lt;p&gt;I ran the same cup-to-bouquet change through a mask and diffed both results against the original.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Share of the frame that changed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Instruction edit&lt;/td&gt;
&lt;td&gt;~17%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Masked inpaint&lt;/td&gt;
&lt;td&gt;~2%, nearly all inside the mask&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The masked run left the chalkboard reading OPEN and the pins in their original colors. It also produced a visibly different bouquet, tied with a ribbon rather than wrapped in paper, which follows from the Edit panel running its own model rather than Tsubaki.3.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3i4wzptxvibhj8vboaq8.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3i4wzptxvibhj8vboaq8.jpg" alt=" " width="800" height="557"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell my past self
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;One change per run. Every run is a full re-render, so each additional request is another chance for something unrelated to move.&lt;/li&gt;
&lt;li&gt;Write the keep list. It measurably narrows drift, and it will not eliminate it.&lt;/li&gt;
&lt;li&gt;Local and discrete for instructions, global for regeneration. Pose and setting belong to a new generation, not an edit.&lt;/li&gt;
&lt;li&gt;Text last.&lt;/li&gt;
&lt;li&gt;Branch from the original for each new edit rather than chaining. Two runs redraw the small stuff twice.&lt;/li&gt;
&lt;li&gt;Diff against the original, not against memory. The pins are the thing you will never notice by eye.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The broader point is that "edit" in these tools is a compatibility label for something that behaves like conditional regeneration. Treat the output as a new image that happens to look like the old one, and your review habits fall into place.&lt;/p&gt;

&lt;p&gt;If you want to try the same tests on your own image, the model I used is here: &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>The Prompt Was Fine. The Input Had Already Consumed the Instruction.</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Thu, 13 Aug 2026 16:44:19 +0000</pubDate>
      <link>https://dev.to/merl985/the-prompt-was-fine-the-input-had-already-consumed-the-instruction-3bk</link>
      <guid>https://dev.to/merl985/the-prompt-was-fine-the-input-had-already-consumed-the-instruction-3bk</guid>
      <description>&lt;p&gt;I spent an afternoon on what looked like a prompt engineering problem and turned out to be an input validation problem.&lt;/p&gt;

&lt;p&gt;The setup: PixAI shipped Studio at the start of August, a node canvas where images, video, text and audio are all assets on the same graph, connected by edges. Image-to-video, so the picture is the input and a text instruction describes the motion. I had a finished illustration of an original character sitting in my gallery and wanted a five second clip out of it. Image in, motion instruction in, video out.&lt;/p&gt;

&lt;p&gt;Three runs, one source image, three instructions. The results ordered themselves in a way I did not predict.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run 1: the greedy instruction
&lt;/h2&gt;

&lt;p&gt;I wrote the version I expected to fail, on purpose. Six requests in four sentences:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;She turns all the way around to face the camera, pushes off the hull
with her right arm, and reaches toward the viewer while debris rushes
past her. The camera orbits around her and pushes in at the same time.
Bright flashes of light, sparks, fast dramatic movement.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Character rotation, a physical push, an arm extension, two simultaneous camera moves, environmental particles, lighting changes. Everything you would expect to blow up.&lt;/p&gt;

&lt;p&gt;It executed all of it. The face held. The signature details held.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pixai.art/artwork/2042584634831661526" rel="noopener noreferrer"&gt;https://pixai.art/artwork/2042584634831661526&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Run 2: the disciplined instruction
&lt;/h2&gt;

&lt;p&gt;So I did what you do after an overloaded call succeeds by accident: I reduced the surface area. One primary action, named invariants, one camera move.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The character drifts slowly in place and turns her head toward the
camera, coming to rest facing the viewer. Her floating hair and the
orange cable sway gently with the motion. The camera pushes in slowly
and steadily. Her face, the lens over her right eye, the collar ring
and the armored glove on her right hand stay unchanged throughout.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tighter, more explicit, better written by every measure I would apply to a prompt.&lt;/p&gt;

&lt;p&gt;The output was a few degrees of head rotation. The supporting elements moved, the camera push landed, and the primary action, the thing the whole clip was built around, effectively no-opped.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pixai.art/artwork/2042584865220486332" rel="noopener noreferrer"&gt;https://pixai.art/artwork/2042584865220486332&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual failure
&lt;/h2&gt;

&lt;p&gt;The instruction was not the problem. The input was.&lt;/p&gt;

&lt;p&gt;She was drawn facing the viewer, gaze locked on the lens. &lt;code&gt;turn her head toward the camera&lt;/code&gt; describes a state transition from A to B where the source image was already at B. The instruction was well formed and semantically empty against that particular input.&lt;/p&gt;

&lt;p&gt;This is the same class of bug as passing a correctly typed argument that happens to be a no-op for the current state. Nothing errors. Nothing warns. You get a successful run and an output that does nothing, and you go looking in the wrong layer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuh4kbvawdi0j7e9t6aum.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuh4kbvawdi0j7e9t6aum.png" alt=" " width="800" height="1473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Run 3: same input, different axis
&lt;/h2&gt;

&lt;p&gt;The fix was not less instruction. It was moving the request onto an axis the input had left unconsumed: depth, camera travel, trailing elements, environmental motion, lighting. Everything except a rotation that had already been applied at draw time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The character pushes off hard and accelerates forward through the
wreckage, her long hair and the orange cable whipping out behind her.
Broken metal fragments and dust streak past the lens in the foreground.
The camera pulls back and tracks with her. A hard white flare of
sunlight sweeps across the frame from the left as she passes. She keeps
looking straight ahead at the camera the whole time.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one shipped.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://pixai.art/artwork/2042585046943800636" rel="noopener noreferrer"&gt;https://pixai.art/artwork/2042585046943800636&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What generalizes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Inspect the input before writing the instruction.&lt;/strong&gt; For image-to-video specifically: what state is the subject already in, what is occluded, how much empty space exists in the direction of intended motion, and how well does the subject separate from the background. Any request that targets a state the input already satisfies produces a silent no-op.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sort the frame into four buckets first.&lt;/strong&gt; Primary movement, supporting movement, invariants, camera. The last one is the cheap lever. A camera move carries a shot while the subject stays nearly static, which is also the safest option when the artwork itself is the payload.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Name your invariants explicitly and pick asymmetric ones.&lt;/strong&gt; My character has a lens over one eye and a glove on one hand only. Asymmetric markers are far easier to verify across frames than "the face looks right," the same way a distinctive sentinel value beats eyeballing a diff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Declare the style contract twice.&lt;/strong&gt; I ran a black and white manga panel through the same pipeline and stated black and white at the start and the end of the instruction. Ink, hatching and screentone, the dot shading that stands in for gray in printed manga, all survived, because the movement went to the camera rather than the drawing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failed runs are free, and version history makes comparison cheap.&lt;/strong&gt; Every attempt stays as a node on the canvas, so three runs sit side by side without any bookkeeping on your side. Rewriting the instruction two or three times is the normal path, and treating run 1 as a probe rather than an attempt saves a lot of time.&lt;/p&gt;

&lt;p&gt;The thing I keep coming back to: run 1 was the sloppy instruction and it worked, run 2 was the careful one and it did nothing. Prompt quality was not the variable. The input was.&lt;/p&gt;

&lt;p&gt;Three runs, five seconds each, and the only variable that mattered was one I had not looked at.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://eap.pixai.art/go/balazs1" rel="noopener noreferrer"&gt;The workspace is here if you want to try it.&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>I Built an Anime Art Pipeline on a Node Canvas and Watched Every Handoff</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Tue, 11 Aug 2026 19:20:25 +0000</pubDate>
      <link>https://dev.to/merl985/i-built-an-anime-art-pipeline-on-a-node-canvas-and-watched-every-handoff-52c6</link>
      <guid>https://dev.to/merl985/i-built-an-anime-art-pipeline-on-a-node-canvas-and-watched-every-handoff-52c6</guid>
      <description>&lt;p&gt;If you have ever wired up a data pipeline, you already know the shape of this problem. Each stage transforms something and hands it to the next one. The interesting failures rarely happen inside a stage. They happen at the boundary, where one thing gets passed forward, and three other things quietly get dropped or silently inherited.&lt;/p&gt;

&lt;p&gt;I ran the same experiment on a creative task. I cannot draw, so I make anime characters with AI tools, and until recently my process was the manual version: generate, download, open the next tool, upload, retype the character description because the next tool has no idea who she is. The file crosses the boundary. State does not.&lt;/p&gt;

&lt;p&gt;So I rebuilt the whole thing on a node canvas, in PixAI Studio, and paid attention to exactly what crossed each edge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The graph
&lt;/h2&gt;

&lt;p&gt;The finished workspace looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;character_sheet (imported)
    └── night_scene (text to image)
            ├── dawn_variant  (edit)
            └── clip          (image to video)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four nodes, three edges. A node is one step holding one asset. An edge means the output of the upstream node is the input of the downstream node. Nothing exotic, which is the point: the mental model is function composition, and the value is that intermediate results stay addressable instead of ending up in a downloads folder as &lt;code&gt;image (7).png&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The task
&lt;/h2&gt;

&lt;p&gt;Deliberately small, because a scoped task makes the measurement clean. Target output: a five-second looping wallpaper, widescreen, of an original fox spirit character sitting on shrine steps at night.&lt;/p&gt;

&lt;p&gt;That one sentence acted like a type signature. Widescreen ruled out a portrait frame. Looping ruled out one-directional motion, at least in theory. A seated pose ruled out anything busy. Writing it first is the equivalent of defining the contract before implementing against it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqotujy72ub38v3mgugps.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqotujy72ub38v3mgugps.jpg" alt=" " width="800" height="488"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The graph as it rendered on the canvas. Every later node draws its input from an earlier one.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What crossed each boundary
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Import to generation.&lt;/strong&gt; The character came from my own library rather than an upload, and it brought its full prompt, its model, and its frame shape along. Three of those four were useful.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generation to selection.&lt;/strong&gt; Four candidate images arrived inside the same node. Picking one is an explicit action, and it is the only step in the whole run where the tool cannot help you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Selection to edit.&lt;/strong&gt; The picked image appeared in the edit step's reference slot with no attachment step on my side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Selection to video.&lt;/strong&gt; Same thing. Dragging an edge out of the image node and choosing video created the downstream node with the source picture already bound.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fragvb7iibt08wpphr4ne.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fragvb7iibt08wpphr4ne.jpg" alt=" " width="800" height="1133"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Left column, what arrived by itself. Right column, what still needed a decision.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The three failures, which are the useful part
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Silent inheritance.&lt;/strong&gt; My character sheet was portrait. The scene node inherited that frame shape, and my widescreen wallpaper came out tall. Defaults propagating downstream is helpful right up to the moment a stage needs something different. Check the format at every node that produces something new.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An overly broad transform.&lt;/strong&gt; I asked the edit step for one change, night to dawn. I got the light I asked for and a gate in the background that had not been there. Continuity across stages is real, and it has edges. Narrow instructions, verified one at a time, preserve more of what you already accepted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A spec that failed to constrain.&lt;/strong&gt; My clip does not loop cleanly, because falling leaves move in one direction and the last frame never meets the first. The bug is in the requirement, not in the video step. "Looping" was a word I wrote down without deciding what it excluded.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the graph is worth building
&lt;/h2&gt;

&lt;p&gt;The overhead pays off when one asset feeds several outputs, when you move from stills into motion, or when you want a process you can rerun next month without reconstructing it from memory. For a single standalone image, a plain generator is faster, and there is nothing to organize. Reach for the canvas at the point where your project has stages.&lt;/p&gt;

&lt;p&gt;Video generation also costs meaningfully more than image generation, which is a decent argument for putting the review gate before the expensive node rather than after it.&lt;/p&gt;

&lt;p&gt;The whole build took an afternoon and produced one five-second clip. What I came away with was a clearer map of which stages carry work forward and which ones quietly hand it back.&lt;/p&gt;

&lt;p&gt;If you want to try the same experiment, &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;start a workspace on PixAI&lt;/a&gt; and take one character all the way to a finished output. Watch the boundaries rather than the nodes.&lt;/p&gt;




</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>The Version Matrix Problem in AI Character Art</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Wed, 05 Aug 2026 20:06:00 +0000</pubDate>
      <link>https://dev.to/merl985/the-version-matrix-problem-in-ai-character-art-2cp5</link>
      <guid>https://dev.to/merl985/the-version-matrix-problem-in-ai-character-art-2cp5</guid>
      <description>&lt;p&gt;Anyone who has managed dependencies will recognize the shape of this problem, even without touching an image generator.&lt;/p&gt;

&lt;p&gt;I spent an afternoon running the same original character through Midjourney and PixAI, and the interesting result had nothing to do with image quality. It was a compatibility matrix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test subject
&lt;/h2&gt;

&lt;p&gt;One character, described once in plain language: an adult woman in her mid-thirties, cropped ash grey hair with a thin braid, a scar through the left eyebrow, steel blue eyes, a brass monocle on a chain, a heavy navy coat with a high collar, a leather wrap on her right forearm, and a small brass clockwork falcon on her left shoulder.&lt;/p&gt;

&lt;p&gt;Eight checkable attributes. Same description on both platforms, no cherry-picking between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The matrix
&lt;/h2&gt;

&lt;p&gt;Midjourney's version dropdown, as of early August 2026, runs from 1 through 8.2, with 8.2 as the default since July 24. Below a divider sit the anime models: niji 7, niji 6, niji 5, niji 4. The compatibility that matters for character work looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;feature                     available on
--------------------------  --------------------------------
current default model       8.2
anime models (niji)         stops at niji 7
Omni Reference (character)  version 7
Character Reference (old)   version 6
editor, pan, zoom           documented as running on 6.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that as a dependency graph, and the problem is visible before you generate anything. The newest model, the anime models, and the character lock resolve to three different versions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk6onzrcy0177mjqqj9t1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk6onzrcy0177mjqqj9t1.png" alt=" " width="799" height="495"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the runtime actually does
&lt;/h2&gt;

&lt;p&gt;Three behaviors showed up in testing that the matrix alone does not tell you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Loading a reference silently changes your model.&lt;/strong&gt; Midjourney's documentation states that adding a reference image runs the prompt in version 7. My reference runs came back in that model's rendering style rather than the painterly output the default model had produced minutes earlier. Nothing in the flow asks whether that tradeoff is acceptable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The anime models reject the parameter outright.&lt;/strong&gt; Loading a reference on niji works, and then a line comes back saying niji does not support OW, the weight setting that controls reference strength. Style reference is offered instead, and a style reference transfers a look rather than an identity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reference state does not persist.&lt;/strong&gt; The reference has to be attached again for every run. I changed the scene, generated, and got a different woman: brown wavy hair instead of ash grey, different face, different age. Zero warnings, zero errors, just a batch produced without the input I assumed was still bound.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqi3auztpnf5a433l6z8v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqi3auztpnf5a433l6z8v.png" alt=" " width="800" height="443"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Set up correctly, version 7 handled the character well across two scene changes. The hair, the build, and the coat carried over. The brass falcon arrived as an ordinary bird every time, and the clothing color drifted from run to run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other path
&lt;/h2&gt;

&lt;p&gt;PixAI approaches this as one runtime rather than a matrix. You pick an anime model, Tsubaki.2 in my case, and the reference tools, the LoRA layer, and the editing tools sit in the same generation panel. A LoRA is a small add-on file trained over a base model that teaches it a style, a character, or a detail treatment, and the platform supports both community LoRAs and &lt;a href="https://blog.pixai.art/en/train-lora-on-pixai/" rel="noopener noreferrer"&gt;training your own&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The result I did not expect: running the same written description across three scenes with no reference image loaded produced the same person each time. Market at noon, workshop at night, rain on a pier. Different outfits, different expressions, and the identifying details held, brass fittings on the falcon included.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fai20s1y9g6gid1rrcof4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fai20s1y9g6gid1rrcof4.png" alt=" " width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is imperfect. The scar migrated to her cheek in one image and coat details vary. What it removes is state management: nothing to re-attach, no version to switch, no feature that exists on one branch and not another.&lt;/p&gt;

&lt;h2&gt;
  
  
  When each one is the right call
&lt;/h2&gt;

&lt;p&gt;Midjourney remains excellent, and for single images or style exploration the version matrix never comes up, because you stay on one model and generate. Its aesthetic tooling is deeper than most, and it needs a paid subscription before the first image, starting at ten dollars a month.&lt;/p&gt;

&lt;p&gt;The matrix starts costing you when a character has to survive twenty images across six months. At that point, the question stops being which model renders better and becomes how much setup stands between you and image number twenty.&lt;/p&gt;

&lt;p&gt;If that is the work you are doing, &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;see how far one description gets you in PixAI&lt;/a&gt;. The &lt;a href="https://blog.pixai.art/en/ai-art-generator-quick-start/" rel="noopener noreferrer"&gt;quick start guide&lt;/a&gt; covers the interface basics.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>beginners</category>
      <category>productivity</category>
      <category>design</category>
    </item>
    <item>
      <title>The Most Downloaded Model Was the Wrong Dependency</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Tue, 04 Aug 2026 14:44:09 +0000</pubDate>
      <link>https://dev.to/merl985/the-most-downloaded-model-was-the-wrong-dependency-3hch</link>
      <guid>https://dev.to/merl985/the-most-downloaded-model-was-the-wrong-dependency-3hch</guid>
      <description>&lt;p&gt;Model libraries work like package registries. You search, you sort by popularity, you read the stars, you take the top result. That heuristic is good enough almost everywhere, so I used it, and it produced a worse result than doing nothing at all.&lt;/p&gt;

&lt;p&gt;Here is the run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The spec
&lt;/h2&gt;

&lt;p&gt;The artifact under test is a character. It is written down, it does not exist as an image, and the description is the only source of truth. Seven assertions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;character&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bram&lt;/span&gt;
&lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;lockmaker&lt;/span&gt;
&lt;span class="na"&gt;age&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;40s&lt;/span&gt;
&lt;span class="na"&gt;build&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;short, broad, barrel chest&lt;/span&gt;      &lt;span class="c1"&gt;# the load bearing one&lt;/span&gt;
&lt;span class="na"&gt;hair&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bald, geometric tattoo on left scalp&lt;/span&gt;
&lt;span class="na"&gt;beard&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;copper red, two braids, brass ring on each&lt;/span&gt;
&lt;span class="na"&gt;eyewear&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;brass jeweler's loupe over right eye&lt;/span&gt;
&lt;span class="na"&gt;torso&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;leather apron with tool loops&lt;/span&gt;
&lt;span class="na"&gt;under&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;moss green tunic, teal embroidery&lt;/span&gt;
&lt;span class="na"&gt;belt&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ring of brass keys&lt;/span&gt;
&lt;span class="na"&gt;arms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;studded leather bracers, old burn scars&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two environments, same spec, same target output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Environment A: resolve the dependency yourself
&lt;/h2&gt;

&lt;p&gt;The generator opens with a model already selected. Search the model library for "anime" and you get zero results, because the search is scoped to whichever ecosystem is currently active and the default ecosystem has no such model. There is a dropdown that sets that scope. Until you find it, every query runs against the wrong index.&lt;/p&gt;

&lt;p&gt;With the scope corrected, the top anime checkpoint reports over 450,000 downloads across its versions and a review score of Overwhelmingly Positive from almost 1,500 reviewers. By every signal a registry can give you, that is the correct resolution.&lt;/p&gt;

&lt;p&gt;Switching to it changed the runtime config underneath. A negative prompt field appeared, and the sampler, step count, and CFG scale all moved to new defaults.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6ppujlwo8moep83ycbiz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6ppujlwo8moep83ycbiz.png" alt=" " width="800" height="771"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Decisions required before the second image:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;Discoverable from the UI?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Find the ecosystem dropdown&lt;/td&gt;
&lt;td&gt;Only after a failed search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Re-run the failed search&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Choose the model family&lt;/td&gt;
&lt;td&gt;Requires prior knowledge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Search the checkpoints&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Choose a checkpoint&lt;/td&gt;
&lt;td&gt;Yes, by popularity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Choose its version&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Set CFG&lt;/td&gt;
&lt;td&gt;Yes, with a labelled preset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Set the sampler&lt;/td&gt;
&lt;td&gt;Yes, with a labelled preset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Set the steps&lt;/td&gt;
&lt;td&gt;Yes, with a labelled preset&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Write a negative prompt&lt;/td&gt;
&lt;td&gt;Field appears, content is on you&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Item 3 is the interesting one. It is the only step where the interface cannot help, and it is the step that determines everything downstream.&lt;/p&gt;

&lt;p&gt;There is also a prompt-language change that costs zero clicks and matters more than any of the ten. Community anime checkpoints in this family are SDXL-derived and expect comma-separated tags. The spec had to be rewritten:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1boy, solo, male dwarf, adult man, 40 years old, stocky build,
broad shoulders, barrel chest, short stature, bald head,
dark geometric head tattoo on left side, long copper red beard,
braided beard, brass beard rings, brass jeweler's loupe over right eye,
brown leather apron, tool loops on apron, moss green tunic,
teal embroidered cuffs, ring of brass keys on belt, leather bracers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Negative prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;child, kid, young boy, teenager, youthful face, beardless,
tall, slender, thin body, feminine, 1girl, bad hands, bad anatomy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assertion&lt;/th&gt;
&lt;th&gt;Environment A, default model&lt;/th&gt;
&lt;th&gt;Environment A, top checkpoint&lt;/th&gt;
&lt;th&gt;Environment B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;build: short, broad&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;fail&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;bald + tattoo&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;beard braids + rings&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;loupe over eye&lt;/td&gt;
&lt;td&gt;partial, held in hand&lt;/td&gt;
&lt;td&gt;partial&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;apron with tool loops&lt;/td&gt;
&lt;td&gt;fail&lt;/td&gt;
&lt;td&gt;fail&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;moss green tunic&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;fail, torso bare&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;brass keys on belt&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two runs at two aspect ratios on the top checkpoint. The build assertion failed both times, with three positive tags asserting it and two negative tags excluding the opposite.&lt;/p&gt;

&lt;p&gt;The default model, resolved by nobody, satisfied it on the first run.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37s6nl1ikatldu9vkji9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37s6nl1ikatldu9vkji9.png" alt=" " width="800" height="617"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the heuristic broke
&lt;/h2&gt;

&lt;p&gt;Popularity ranking optimizes for the median request. A checkpoint with 450,000 downloads is fine-tuned toward what most users generate, and a stocky forty-year-old tradesman sits outside that distribution. The registry sorted correctly. My input was out of sample, and nothing in the metadata exposes that, because the metadata describes adoption rather than coverage.&lt;/p&gt;

&lt;p&gt;This is the same failure mode as picking a library by GitHub stars for a use case its maintainers never targeted. The signal is real. It is measuring somebody else's requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Environment B: keep the default
&lt;/h2&gt;

&lt;p&gt;PixAI loads an anime model when the generator opens. I kept it, wrote the spec as prose instead of tags, set the aspect ratio, generated. Three decisions, one of which was the picture size.&lt;/p&gt;

&lt;p&gt;Worth noting for anyone moving between the two: this model is a DiT architecture rather than SDXL, so there is no negative prompt field and no bracket weight syntax like &lt;code&gt;(tag:1.3)&lt;/code&gt;. Exclusions go in the positive prompt as plain statements. If you carry an SDXL prompt across unchanged, the weights are ignored silently.&lt;/p&gt;

&lt;p&gt;All seven assertions passed. From there, three scene variants ran on random seeds with only the scene clause swapped, and the character held across all of them without a reference image and without a locked seed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;  A bald dwarf man in his forties, short and broad, copper red beard
  in two braids each bound with a brass ring, dark geometric tattoo...
&lt;span class="gd"&gt;- Plain warm grey studio backdrop, even neutral lighting.
&lt;/span&gt;&lt;span class="gi"&gt;+ Walking across a stone bridge market at night, lit lanterns overhead.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One extra observation on LoRAs, since it is the same class of bug. My first pick drifted the entire palette, and its page explained why: it was trained on a different base model from the one I was running. Base model is the compatibility field. Read it first. My second pick matched and still carried its own aesthetic, which is what the strength value is there to attenuate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;Registry rank answers "what do most people use." It does not answer "what satisfies my spec." When your input is unusual, the default that ships with the tool is a legitimate baseline, and the cheapest experiment you can run is to try it before you spend ten decisions replacing it.&lt;/p&gt;

&lt;p&gt;The finished artifact was a character sheet: three views, three scenes, three detail crops, all resolved from one paragraph.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8zofept67iwvno0ba5mv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8zofept67iwvno0ba5mv.png" alt=" " width="800" height="1018"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you want to try this with a character of your own, &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;you can start here&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>beginners</category>
      <category>showdev</category>
    </item>
    <item>
      <title>Symmetry Is the Default Your Character Spec Has to Survive</title>
      <dc:creator>Merl Merl</dc:creator>
      <pubDate>Sat, 01 Aug 2026 08:39:30 +0000</pubDate>
      <link>https://dev.to/merl985/symmetry-is-the-default-your-character-spec-has-to-survive-14ic</link>
      <guid>https://dev.to/merl985/symmetry-is-the-default-your-character-spec-has-to-survive-14ic</guid>
      <description>&lt;p&gt;Character consistency in image generation is usually discussed as an aesthetic problem. I find it more useful to treat it as a contract problem: I hand a model a specification, and I want to know whether the specification survives a change of inputs. So I built a spec with a deliberate failure mode and ran it through two anime image models on the same day.&lt;/p&gt;

&lt;p&gt;The failure mode is symmetry.&lt;/p&gt;

&lt;h2&gt;
  
  
  The spec
&lt;/h2&gt;

&lt;p&gt;Seven attributes, written once, held constant across every run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;character: Shiori
hair:        jet-black, low knot at the nape
hair_pin:    single thin gold pin, through the knot
eyes:        pale grey
face_mark:   small beauty mark, below RIGHT eye
outfit:      black high-collared tailcoat, gold piping, burgundy lining
shirt:       white wing-collar, burgundy ribbon tie
glove:       LEFT hand only
ring:        gold, RIGHT index finger, on bare skin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2zayim50vnzkw4uqaavd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2zayim50vnzkw4uqaavd.png" alt=" " width="800" height="458"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Five of these are cheap. Hair color, eye color, and a palette are the kind of thing a diffusion model reproduces almost by accident, because they are global properties of the image.&lt;/p&gt;

&lt;p&gt;Two of them are expensive, and they are the assertions that matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;glove_hand == LEFT and ring_hand == RIGHT&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;face_mark_side == RIGHT&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both are asymmetric. Asymmetry is the interesting test because symmetry is the prior. A model that has stopped tracking your description will not produce noise; it will produce the average, and the average of "one glove" is either two gloves or none, or both attributes collapsed onto whichever hand it happened to render.&lt;/p&gt;

&lt;p&gt;There is a third assertion that turned out to matter more than I expected, and I did not write it down at first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;both_hands_in_frame == true&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An attribute you cannot see is an attribute you cannot verify, and a model that crops the frame has silently deleted your test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 1: baseline, prompt only
&lt;/h2&gt;

&lt;p&gt;Model A is the Niji Journey model, currently Niji 7. Model B is Tsubaki.2 on PixAI, a diffusion transformer, which matters later because bracket weight syntax like &lt;code&gt;(tag:1.3)&lt;/code&gt; works on SDXL-family models and does nothing on a DiT.&lt;/p&gt;

&lt;p&gt;Both baselines passed the expensive assertion. Glove and ring landed on separate hands. Model A mirrored the sides, which is a known and boring class of failure. Model B put them where the spec said.&lt;/p&gt;

&lt;p&gt;Baseline passing is worth stating plainly, because a lot of comparisons stop here and declare a winner on the strength of a first render. First renders are the easy case. The prompt is doing all the work and nothing has been asked to persist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 2: change the inputs, keep the spec
&lt;/h2&gt;

&lt;p&gt;New outfit, new environment, same seven attributes carried into the prompt verbatim: heavy overcoat, knit scarf, rain-slicked street at dusk, full body.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;glove/ring separated&lt;/th&gt;
&lt;th&gt;both hands in frame&lt;/th&gt;
&lt;th&gt;unrequested additions&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2a&lt;/td&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;no, merged onto one hand&lt;/td&gt;
&lt;td&gt;no, waist-up crop&lt;/td&gt;
&lt;td&gt;gold hoop earring&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2b&lt;/td&gt;
&lt;td&gt;A, identical prompt&lt;/td&gt;
&lt;td&gt;no, merged onto one hand&lt;/td&gt;
&lt;td&gt;no, waist-up crop&lt;/td&gt;
&lt;td&gt;gold hoop earring&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run 2b is the useful one. I repeated 2a byte for byte specifically to find out whether the failure was stochastic or deterministic in effect, and it reproduced. That distinction is the whole engineering point:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A stochastic failure is a retry problem. Loop until green.&lt;/li&gt;
&lt;li&gt;A reproducible failure is an input problem. The same prompt will produce the same misreading indefinitely.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;both_hands_in_frame&lt;/code&gt; assertion failing is what caused the primary assertion to fail. With one hand cropped out, the model had one surface for two mutually exclusive attributes and merged them. My spec had a hole in it, and the tool found the hole.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 3: reference as input
&lt;/h2&gt;

&lt;p&gt;At this point the correct move is to stop describing the character and start passing it. Model A exposes a &lt;code&gt;Character Reference&lt;/code&gt; control. On Niji 7 it renders disabled, and clicking it emits nothing, no error, no tooltip, no explanation. The references panel offers an image prompt and a style reference.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdmteyqu73i7bqwb1l9g3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdmteyqu73i7bqwb1l9g3.png" alt=" " width="624" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So I used the image prompt, feeding the passing baseline back in at default weight.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Run&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3a&lt;/td&gt;
&lt;td&gt;baseline image as image prompt&lt;/td&gt;
&lt;td&gt;face fidelity best of all runs, palette bled (jacket lining reappeared as a turtleneck), glove/ring still merged, still cropped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the finding I care about. The reference improved everything that is a global property of the image and improved nothing that is a discrete fact about the subject. Treating "reference image" as a single capability is a mistake. There are at least two different contracts hiding under that label, and only one of them carries a specification.&lt;/p&gt;

&lt;p&gt;Model B routes this through a separate model rather than a flag on the generation model, which is an implementation detail with a real consequence: the character is created on an anime model and then carried by a reference model, so a style shift between the two is expected rather than a bug. It accepts up to ten reference inputs, which fits how character sheets work, since a character is a set of angles rather than one canonical frame.&lt;/p&gt;

&lt;p&gt;Across three scene changes on Model B, the face, the knot, and the gold pin held.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzql50aetxya7wydwi2zq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzql50aetxya7wydwi2zq.png" alt=" " width="800" height="1357"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 4: adding a LoRA, and its cost
&lt;/h2&gt;

&lt;p&gt;A LoRA is a low-rank adapter, a small set of weights trained on top of a base model. I added a community adapter built for eye rendering, typed its trigger words into an otherwise unchanged prompt, and ran it at &lt;code&gt;0.7&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgx3cbdnbdu5xl409r5nj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgx3cbdnbdu5xl409r5nj.png" alt=" " width="799" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It delivered the eye treatment. It also softened the face geometry and added bangs that appear nowhere in the spec. At 0.7, the adapter reached past style and into identity.&lt;/p&gt;

&lt;p&gt;The generalizable rule: an adapter is a weighted intervention on the same latent space your character occupies, so strength is a tradeoff between the effect you want and the subject you are trying to preserve. If identity is the priority, the weight is the first knob to turn down.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I take from this
&lt;/h2&gt;

&lt;p&gt;Three things transfer to anything you build on top of image models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Put an asymmetric assertion in your spec.&lt;/strong&gt; It is a one-token change, and it converts a subjective "does this look like her" into a boolean. Symmetric attributes will pass even when the model has stopped listening.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assert your frame.&lt;/strong&gt; An attribute outside the crop is untested, and an untested attribute fails silently. Naming the framing explicitly fixed more for me than any prompt-weight trick.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distinguish description from identity.&lt;/strong&gt; A prompt is a description and descriptions are satisfiable by more than one subject. Persistence needs the character passed as data, either as reference input or as trained weights. Anything short of that is a retry loop with a fixed point in the wrong place.&lt;/p&gt;

&lt;p&gt;The tooling question resolves along the same line. If the deliverable is one image, the model with the better first render wins, and in my runs that was Model A. If the deliverable is the same subject across many images, what matters is whether the platform gives you a path to pass identity as data, and that is why the second half of my day happened on &lt;a href="https://eap.pixai.art/go/balazs" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>testing</category>
      <category>beginners</category>
    </item>
  </channel>
</rss>
