<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Naveed W</title>
    <description>The latest articles on DEV Community by Naveed W (@naveedoss).</description>
    <link>https://dev.to/naveedoss</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4020780%2F61a227ff-6e13-4398-bd41-5dfcaaa41147.png</url>
      <title>DEV Community: Naveed W</title>
      <link>https://dev.to/naveedoss</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/naveedoss"/>
    <language>en</language>
    <item>
      <title>Your AI line art probably isn't line art. Here's the one-generation test that tells you.</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Mon, 14 Sep 2026 11:58:46 +0000</pubDate>
      <link>https://dev.to/naveedoss/your-ai-line-art-probably-isnt-line-art-heres-the-one-generation-test-that-tells-you-1jn9</link>
      <guid>https://dev.to/naveedoss/your-ai-line-art-probably-isnt-line-art-heres-the-one-generation-test-that-tells-you-1jn9</guid>
      <description>&lt;p&gt;There is a validation problem at the centre of this whole topic, and almost nobody writes about it.&lt;/p&gt;

&lt;p&gt;A drawing that looks like line art and a drawing that works as line art are two different artifacts, and you cannot tell them apart by looking at either one. Open contours are invisible in black and white. They only announce themselves later, when somebody fills the drawing and the colour runs straight past the edge.&lt;/p&gt;

&lt;p&gt;So this piece is organised backwards from the usual shape. The gate comes first, then the routes that feed it.&lt;/p&gt;

&lt;p&gt;Eight tests, all on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt; inside &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foawctgyt5ydd6eqnrzem.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foawctgyt5ydd6eqnrzem.gif" alt="pixai dashboard" width="80" height="38"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The spec
&lt;/h2&gt;

&lt;p&gt;Four properties decide whether a drawing is usable downstream. Judge an AI line art generator on these rather than on how the output looks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Closed contours.&lt;/strong&gt; Every shape sealed. This is the one you cannot verify visually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No grey.&lt;/strong&gt; Grey tones, washes and screentone mean the model rendered instead of drawing, and they degrade at every later stage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A true white background.&lt;/strong&gt; Off-white, noise and faint washes all cause trouble downstream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Line weight carrying information.&lt;/strong&gt; Thick outside, fine inside. Optional on simple subjects, decisive on dense ones.&lt;/p&gt;




&lt;h2&gt;
  
  
  The gate: fill it
&lt;/h2&gt;

&lt;p&gt;Run this on any drawing before you commit to it. One generation, and it settles the only question the image itself won't answer.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: line art set as a base image, then a flat colour prompt. No edit tool.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add flat colours inside this line art. Keep every line exactly where it is and
do not redraw anything. Fill each enclosed area with one solid colour: pale
yellow hat, white mesh veil, cream canvas jacket, brown leather gloves, grey
metal smoker, tan wooden hives, flat pale blue sky. No gradients, no shading,
no lighting, no texture. The black outlines stay visible on top.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Every colour stayed inside its lines and nothing bled between regions.&lt;/strong&gt; The black outlines all survived in their original positions, and each enclosed area took one flat colour with no gradients or shading.&lt;/p&gt;

&lt;p&gt;Colour floods correctly only when the contours are sealed, so a clean fill is structural proof. Weak or open lines would have leaked and marked their own position on the drawing for you.&lt;/p&gt;

&lt;p&gt;Most anime lineart AI output passes inspection and fails this. Run the gate first; everything below is about feeding it something that survives.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh64u8bb7fdr7zo57wru8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh64u8bb7fdr7zo57wru8.png" alt=" " width="799" height="510"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn23g636chhhpfs2bgbw3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn23g636chhhpfs2bgbw3.png" alt="Left the line art. Right the same drawing coloured" width="800" height="661"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the line art. Right the same drawing coloured&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Route 1: prompt for it directly
&lt;/h2&gt;

&lt;p&gt;The cleanest of the three routes, and the one to default to.&lt;/p&gt;

&lt;p&gt;I picked a subject engineered to cause problems. A mesh veil is transparency, wooden hives are repeating structure, a metal smoker is a hard-edged object, and bees in flight are small repeated shapes.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: new generation, no reference, no LoRA.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Black and white anime line art, drawn with a technical pen on white paper. A
teenage girl beekeeper stands among tall wooden hives, wearing a wide-brimmed
hat with a fine mesh veil hanging over her face, a heavy canvas jacket with
the cuffs buckled, thick gloves, and a metal smoker in her right hand with
smoke curling from the spout. Her face is visible through the mesh. Bees in
flight around her. Even line weight, every contour closed, pure white
background, no grey tones, no screentone, no hatching, no solid black fills,
no shading, no colour.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Passed all 4 spec properties.&lt;/strong&gt; No colour, no solid black fills, a pure white background with no wash or noise, and her face legible under the mesh rather than buried by it. The smoker rests in the right hand with smoke curling from the spout, the cuffs are buckled, and line weight holds even throughout.&lt;/p&gt;

&lt;p&gt;Naming a physical tool and real paper carries weight in that prompt. It gives the model something to imitate that isn't a rendered illustration.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgteian98azf87vykn0gj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgteian98azf87vykn0gj.png" alt="Test 1: the direct ask." width="800" height="1333"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 1: the direct ask.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Edge case: subjects with no edges
&lt;/h3&gt;

&lt;p&gt;Line art describes boundaries, and rain, steam and breath don't have any. This tests whether the model understands the format or only the subject matter.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Black and white anime line art, drawn with a technical pen on white paper. A
girl standing in heavy rain under a broken streetlight, water streaming off
her umbrella, steam rising from a grate beside her, her breath visible in the
cold. Draw the rain, the steam and the breath entirely with line, no grey
tones, no screentone, no black fills, no shading, pure white background, every
contour closed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;All 3 came back as closed outline shapes instead of grey.&lt;/strong&gt; Breath and ground steam as self-contained cloud loops, rain as elongated dashes, water off the umbrella ribs as drawn sheets. Every one of those would take a flat fill.&lt;/p&gt;

&lt;p&gt;Most models render vapour with soft gradients, so this one understood the format rather than the subject.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0m8qxxtn93zb4o6kblgn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0m8qxxtn93zb4o6kblgn.png" alt="test 8" width="800" height="682"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 8 rain, steam and breath, all drawn as line. I selected batch x4, selected these two, let’s go with the right one.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Route 2: strip a render afterwards
&lt;/h2&gt;

&lt;p&gt;Generate the picture you want, then remove everything that isn't line.&lt;/p&gt;

&lt;p&gt;I generated a fully rendered beekeeper first, then opened the edit box and typed "Convert to line art", with Tsubaki.3 selected as the edit model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structure retention is the strength.&lt;/strong&gt; Pose, hive arrangement, buckle placement and smoker geometry all survived without shifting. Every colour and value came out, leaving black lines on white, and the veil mesh translated without collapsing into a black blob.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Line quality is the cost.&lt;/strong&gt; The conversion behaves like a threshold filter, so thin hair strands and the mesh picked up pixelated edges the prompted version never had.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Route selection rule:&lt;/strong&gt; prompt directly when drawing quality is the priority, convert a render when you already have an image whose structure you want preserved.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk9gtnwvr15zv8tpqv8sm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk9gtnwvr15zv8tpqv8sm.png" alt="pixai dashboard" width="800" height="504"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ksh72pxdr46szyyggy6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ksh72pxdr46szyyggy6.png" alt="Left the rendered base. Right after conversion." width="800" height="667"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Route 3: convert a photo
&lt;/h2&gt;

&lt;p&gt;The weakest of the three, and the failure is systematic rather than random.&lt;/p&gt;

&lt;p&gt;I attached a photo of a man in a camel coat standing by a vintage car and asked for the same clean treatment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Convert the photo in the reference image into clean black and white anime line
art. Keep the layout, the proportions and the position of everything in the
frame. Draw it with a technical pen, even line weight, every contour closed,
pure white background, no grey tones, no screentone, no black fills, no
shading, no colour.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Composition transfer worked well. Pose, phone grip, coat drape, belt buckle, wire-frame glasses, rings and the car behind him all landed, and the face stylised into anime while remaining the same person.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then it ignored every negative instruction in the prompt.&lt;/strong&gt; Cross-hatching spread across the coat, the lapels, the trousers, the belt and under the car body. Solid black fills appeared in the belt area. Line weight varied throughout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mechanism: it translated the photograph's light and shadow into hatch marks&lt;/strong&gt; rather than discarding the tonal information. A photo carries shadow values, and the model converts them instead of dropping them. No amount of negative prompting removed that behaviour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Route selection rule:&lt;/strong&gt; treat this as a stylised sketch generator, never as a source of production line art.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyxeyv9u8ura7rdvxs5f8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyxeyv9u8ura7rdvxs5f8.png" alt="Left the source photo, by GlassesShop from Pixabay. Right the line art conversion." width="800" height="498"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the source photo, by GlassesShop from Pixabay. Right the line art conversion.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Dial 1: model choice
&lt;/h2&gt;

&lt;p&gt;This moved results more than any wording change I made.&lt;/p&gt;

&lt;p&gt;I ran the beekeeper prompt word for word on Haruka v2, an SDXL model, expecting it to win on the strength of how much line work stands behind that family.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It failed comprehensively.&lt;/strong&gt; Heavy black fills on the hair, the collar, the inner neck and the straps. A background of screentone foliage, shading and a drawn panel border where white space belonged. A face mask over her lower face rather than a veil hanging from the hat brim.&lt;/p&gt;

&lt;p&gt;The strangest miss: &lt;strong&gt;it parsed "technical pen" out of the medium description and drew a pen in her hand in place of the bee smoker&lt;/strong&gt;, with smoke coming out of her mouth like an exhale.&lt;/p&gt;

&lt;p&gt;Aesthetic bias overrode the negative constraints entirely. The model knows what a manga panel looks like, so it produced one, screentone included. The newest model won a comparison I would not have predicted. The &lt;a href="https://blog.pixai.art/en/sdxl-anime-models-pixai-guide/" rel="noopener noreferrer"&gt;anime models guide&lt;/a&gt; covers where these families diverge.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwg099gf7i3p8h8q9whbf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwg099gf7i3p8h8q9whbf.png" alt="Left Tsubaki.3. Right Haruka v2, same prompt." width="800" height="675"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left Tsubaki.3. Right Haruka v2, same prompt.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Dial 2: line weight
&lt;/h2&gt;

&lt;p&gt;The model has line hierarchy available and will not apply it unless the prompt asks. I ran a dense mechanical subject twice, changing one instruction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Black and white anime line art with perfectly even line weight throughout,
drawn with a 0.3mm technical pen. Every line the same thickness, outer
contours and interior detail identical in weight. A mechanic kneeling beside a
stripped motorcycle engine on a workbench, tools and loose parts scattered
around her. Pure white background, no fills, no shading.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Black and white anime line art with weighted lines, drawn with a brush pen.
Thick bold contours on the outer silhouette, medium lines on major forms, very
fine thin lines for interior detail like bolt heads, cable runs and fabric
folds. A mechanic kneeling beside a stripped motorcycle engine on a workbench,
tools and loose parts scattered around her. Pure white background, no fills,
no shading.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The even-weight run complied fully and produced the worse drawing.&lt;/strong&gt; The engine's outer contour carries the same thin stroke as the bolt heads, the gear teeth and her hair strands, so the mechanical area flattens into a mesh you have to decode rather than read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The weighted run resolved foreground from background immediately.&lt;/strong&gt; Heavy brush contours on her silhouette, the bench and the major blocks, fine strokes inside the engine and the fabric folds. The scattered parts became legible because line weight was building depth.&lt;/p&gt;

&lt;p&gt;Even weight is fine on a simple character. On anything dense, specify the brush pen version.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo1eiygoegd7yqpohr46y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo1eiygoegd7yqpohr46y.png" alt="Left even line weight. Right weighted lines." width="800" height="636"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left even line weight. Right weighted lines.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  Output targets: three densities
&lt;/h2&gt;

&lt;p&gt;Different jobs want different amounts of line, and the vocabulary matters more than people expect. I ran three, each on a subject that suits it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1st: Rough anime construction sketch in pencil, loose searching lines,
overlapping strokes, unfinished. A boy leaping off a skate ramp mid-air, board
sideways under his feet, arms out for balance, hoodie flapping. Gesture and
proportion only, no clean outlines, no fills, white background.

2nd: Clean anime production line art, drawn with a technical pen on white
paper. An elderly barber standing beside his chair in an empty shop, apron on,
comb in his breast pocket, mirrors and bottles behind him. Single confident
outline on every form, even line weight, every contour closed, no hatching, no
screentone, no black fills, no shading, pure white background.

3rd: Inked black and white manga panel art. A woman in a long coat standing
alone on a station platform at night as a train rushes past behind her. Heavy
spot blacks in the shadows, dense cross-hatching on the coat folds, fine
parallel speed lines on the train. Dramatic high contrast, white background,
no grey tones, no screentone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F89m3gnvdyg6gu2hjyvpc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F89m3gnvdyg6gu2hjyvpc.png" alt="test 3" width="799" height="325"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;All 3 passed, and they are 3 different products.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The skate sketch carries visible graphite grain and searching strokes, though the face resolved more finished than a true construction draft would.&lt;/p&gt;

&lt;p&gt;The barber is the deliverable of the three. Every form closes into a sealed cell, the comb teeth render as open outlines rather than filled black, and no part of it needed hatching to read.&lt;/p&gt;

&lt;p&gt;The station platform bends the spec on purpose: spot blacks in the hair and ceiling, hatching on the coat folds, speed lines on the train. &lt;strong&gt;That output is inked art rather than line art&lt;/strong&gt;, and the two get conflated constantly. Request the wrong one and you receive a drawing that cannot be coloured.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fff53gefim6fca1j56b9j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fff53gefim6fca1j56b9j.png" alt="Left the source photo, by GlassesShop from Pixabay. Right the line art conversion. (2)" width="799" height="440"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the source photo, by GlassesShop from Pixabay. Right the line art conversion. (2)&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Routing table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you're making&lt;/th&gt;
&lt;th&gt;What to ask for&lt;/th&gt;
&lt;th&gt;What to check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Colouring base&lt;/td&gt;
&lt;td&gt;Clean production line art, even weight, closed contours&lt;/td&gt;
&lt;td&gt;Fill it and look for leaks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Manga page&lt;/td&gt;
&lt;td&gt;Inked panel art, spot blacks, hatching&lt;/td&gt;
&lt;td&gt;Contrast without grey tones&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Character design pass&lt;/td&gt;
&lt;td&gt;Clean line art, plain background&lt;/td&gt;
&lt;td&gt;Whether the shapes read without colour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dense scene or machinery&lt;/td&gt;
&lt;td&gt;Weighted lines, brush pen&lt;/td&gt;
&lt;td&gt;Foreground separating from background&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concept from a photo&lt;/td&gt;
&lt;td&gt;Reference conversion&lt;/td&gt;
&lt;td&gt;Expect hatching where the shadows were&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Early gesture work&lt;/td&gt;
&lt;td&gt;Rough construction sketch&lt;/td&gt;
&lt;td&gt;Motion and proportion, not finish&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What the 8 runs establish
&lt;/h2&gt;

&lt;p&gt;Route 1 produced the cleanest output, and it worked because Tsubaki.3 honours negative instructions that the SDXL model discarded.&lt;/p&gt;

&lt;p&gt;Route 2 preserves structure at the cost of line quality, which makes it the right call only when you already have an image you want to keep.&lt;/p&gt;

&lt;p&gt;Route 3 has a systematic failure that prompting cannot address, because shadow values in a photo become hatch marks by design.&lt;/p&gt;

&lt;p&gt;The gate is what separated a usable drawing from a picture, and it's the step that gets skipped. Setting line art as a base image and requesting flat colour is the whole verification, and it costs one generation. The &lt;a href="https://blog.pixai.art/en/how-to-use-pixai-guide/" rel="noopener noreferrer"&gt;how to use PixAI guide&lt;/a&gt; covers where the model selector, base image and edit options live if you want to reproduce this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next move
&lt;/h2&gt;

&lt;p&gt;An anime line art generator hands you a drawing. Whether you're holding AI anime line art depends on a property the drawing does not display.&lt;/p&gt;

&lt;p&gt;Generate one, then run the gate. Set it as a base image, request flat colours with no shading, and watch the edges. Clean fills mean sealed contours and something you can build on. Colour creeping past the lines means you have a black and white picture, and the leak tells you which part to redraw.&lt;/p&gt;

&lt;p&gt;Start with the barber rather than the engine. A simple subject tells you whether the prompt is right; only a dense one tells you whether the line weight is.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>I graded an AI character generator against 4 requirements. Here's where it passed and where it didn't.</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Fri, 11 Sep 2026 12:27:43 +0000</pubDate>
      <link>https://dev.to/naveedoss/i-graded-an-ai-character-generator-against-4-requirements-heres-where-it-passed-and-where-it-2edb</link>
      <guid>https://dev.to/naveedoss/i-graded-an-ai-character-generator-against-4-requirements-heres-where-it-passed-and-where-it-2edb</guid>
      <description>&lt;p&gt;Most tool reviews score output quality. That's the wrong axis for a design tool.&lt;/p&gt;

&lt;p&gt;When you're illustrating something you already decided, output quality is the whole job. When you're designing, the tool has a different set of responsibilities, and a model that draws beautifully can still be useless for design if it refuses to give you anything you didn't ask for.&lt;/p&gt;

&lt;p&gt;So I wrote 4 requirements first and graded against them afterwards. Eight tests, one character, all on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt; inside &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The character brief was 1 sentence with nothing visual in it at all:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A former child chess prodigy in his early twenties, working the night shift at a public aquarium, one chess piece in his pocket.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No hair, no build, no clothes, no face. That withholding is the test. An ai anime character generator that can only execute will stall on a brief like this. One that can propose will hand you a menu.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rubric
&lt;/h2&gt;

&lt;p&gt;Four responsibilities, written before any generation ran.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;R1. Propose rather than execute.&lt;/strong&gt; Under-specify the brief and see whether the tool supplies design decisions you hadn't made.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;R2. Hold what you keep.&lt;/strong&gt; Once a detail is named, it has to survive into every later image. Otherwise you spend the session re-winning arguments you already won.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;R3. Support sideways movement.&lt;/strong&gt; Moving a character into another genre, role or rendering style is how you find out which version you want. This is the requirement an anime OC generator fails most often.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;R4. Hand over control in stages.&lt;/strong&gt; Text runs out at some point. References and LoRAs are the next rungs, and how cleanly they work sets the ceiling on the design.&lt;/p&gt;

&lt;p&gt;Grades below, then the evidence for each.&lt;/p&gt;




&lt;h2&gt;
  
  
  R1: propose rather than execute — pass
&lt;/h2&gt;

&lt;p&gt;The first run was a batch of 4 with nothing specified about appearance.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: new generation, batch of 4, no reference, no LoRA.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Design an original anime character. A young man in his early twenties who was
a child chess prodigy and now works the night shift at a public aquarium. He
keeps one chess piece in his pocket. Show him standing alone in front of a
huge lit tank at 3am, hands in his pockets, water throwing moving blue light
across him, the empty walkway stretching away behind. Full body, low camera.
Modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The environment came back on the first attempt: low camera, full body, empty tunnel walkway, blue caustics moving across his shirt, his face and the ceiling. None of that atmosphere was described in the prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The chess piece failed in an instructive way.&lt;/strong&gt; The brief put it in his pocket. The render put a wooden piece straight through his right hip, with his pocketed hand clipping through the fabric around it. &lt;strong&gt;Score: 7.5.&lt;/strong&gt; Composition passed, object placement failed.&lt;/p&gt;

&lt;p&gt;The requirement is about what arrives unrequested, though, and on that axis it delivered hair, build, clothing and face as 4 independent proposals across the batch. I kept 3 of them and discarded everything else.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5a1lp4606yeyiq3qrdaj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5a1lp4606yeyiq3qrdaj.png" alt="Test 1: the blind proposal. Four generations, no design input." width="800" height="265"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 1: the blind proposal. Four generations, no design input.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  R2: hold what you keep — pass
&lt;/h2&gt;

&lt;p&gt;Two tests cover this: one with the kept details written into text, one with them carried by an attached reference.&lt;/p&gt;
&lt;h3&gt;
  
  
  Via text
&lt;/h3&gt;

&lt;p&gt;I wrote the 3 kept details into a new prompt and added 2 the model had never offered: a permanent squint in his left eye from reading chess boards under bad light, and his staff ID lanyard wound twice around his left wrist rather than hanging from his neck.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: new generation, no reference, no LoRA.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;An original anime character, full body. Overgrown ash-grey hair pushed back
off his forehead. A navy aquarium staff windbreaker two sizes too big, sleeves
rolled to the elbow. Heavy black rubber boots. A permanent squint in his left
eye from years of reading chess boards under bad light. His staff ID lanyard
wound twice around his left wrist instead of hanging from his neck. He is
crouched at the base of a tank, scraping algae off the glass with a
long-handled tool, a black king chess piece balanced on the ledge beside him.
Wet floor, blue tank glow, one fluorescent strip flickering overhead. Modern
anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Every anchor arrived. Score: 9.0.&lt;/strong&gt; Ash-grey swept-back hair, oversized navy jacket with rolled sleeves, heavy boots, crouching posture, wet floor reflection, overhead fluorescent. The left-eye squint rendered cleanly. Both the scraper tool and the black king on the ledge came through.&lt;/p&gt;

&lt;p&gt;The single defect is physical logic. The lanyard wraps the wrist correctly, then the badge hangs vertically from the wrap rather than resting against the arm the way gravity would place it. Glass reflections also drift slightly out of register with his face.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4sdqc5dtu5l5sv4270dr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4sdqc5dtu5l5sv4270dr.png" alt="Test 2 (Left) the locked design, we will take it as reference coming up. Right just for you to check" width="800" height="520"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 2 (Left) the locked design, we will take it as reference coming up. Right just for you to check&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Via reference
&lt;/h3&gt;

&lt;p&gt;Text prompting has a ceiling. Once a design needs 3 clauses to describe one detail, references and LoRAs take over, and they do different jobs: a reference carries a specific design forward, a LoRA changes how things render. The &lt;a href="https://blog.pixai.art/en/model-vs-lora-pixai-foundations/" rel="noopener noreferrer"&gt;model versus LoRA guide&lt;/a&gt; has the full distinction.&lt;/p&gt;

&lt;p&gt;I attached the locked design and asked for a shot text alone struggles with.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: reference-based, Test 2 image attached.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The character from the reference image, sitting on the wet floor with his back
against a tank at the end of his shift. He holds the black king chess piece up
between two fingers so the blue tank light comes through the edge of it. Close
shot on his hands and the piece, his face soft behind them. Keep his squint,
his lanyard wound twice around his left wrist, and his ash-grey hair. Modern
anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Score: 8.5.&lt;/strong&gt; The reference preserved details I expected to lose. Lanyard wound twice with the card hanging cleanly, squint transferred, hair and boots and blue lighting all carried. The chess piece takes a sharp point of light exactly where the prompt put it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The camera instruction was discarded.&lt;/strong&gt; A close shot on the hands with a soft face behind came back as a medium full-body shot. Identity transferred, framing did not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdda2fpg1iik3p90siier.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdda2fpg1iik3p90siier.png" alt=" " width="800" height="361"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv4pafch2rr1je968auyr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv4pafch2rr1je968auyr.png" alt="Left the reference image. Right the generated result" width="800" height="521"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the reference image. Right the generated result&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  R3: support sideways movement, pass with one boundary
&lt;/h2&gt;

&lt;p&gt;Three genre swaps, one run each, locked design attached as reference.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: reference-based, 3 separate runs, no LoRA.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1st: Redesign the character from the reference image as a near-future deep-sea
diver. Same face, same build, same squint. He is halfway into a bulky
pressurised suit in a floodlit launch bay, helmet clamped under one arm,
condensation running down the metal behind him. The black king chess piece is
clipped to a strap on his chest. Full body. Modern anime illustration.

2nd: Redesign the character from the reference image as a 1970s stage
magician, caught mid-trick in the wings of a shabby velvet theatre. Same face,
same build, same squint. Cuffs pushed back, one hand raised, dust hanging in a
single spotlight beam. The black king chess piece sits in his open palm. Full
body. Modern anime illustration.

3rd: Redesign the character from the reference image as a convenience store
clerk on the late shift. Same face, same build, same squint. He leans on the
counter watching rain hammer the window, uniform shirt half untucked, the hot
drinks machine glowing beside him, the black king chess piece standing next to
the register. Full body. Modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Clerk: 9.0&lt;/strong&gt;, the strongest of the 3. Face, hair and squint transferred, rain streaked across the floor-to-ceiling window, drinks machine glowing behind the counter, chess piece standing by the register. It buttoned the shirt neatly despite a request for half untucked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Magician: 8.5.&lt;/strong&gt; Red velvet curtains, shabby wooden stage floor, dust in a single spotlight beam, cuffs pushed back, and a top hat layering hair strands around the brim the way anime does. The chess piece floats above the open palm rather than resting in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Diver: 5.5&lt;/strong&gt;, and the most useful data point in the set. Face, hair and squint held perfectly and the chess piece clipped to the chest harness. Then it sealed him fully into the suit against an explicit "halfway into it", and rendered a standard NASA spacesuit instead of anything deep-sea.&lt;/p&gt;

&lt;p&gt;The pattern: the further a genre is from something the model has seen often, the more scene accuracy degrades while identity stays intact. &lt;strong&gt;Identity transferred in all 3 runs. Scene fidelity is the variable.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh7scv03clx6ymwhsdepd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh7scv03clx6ymwhsdepd.png" alt="Left to right the diver, the magician, the clerk." width="800" height="348"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left to right the diver, the magician, the clerk.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Rendering style is the other sideways axis. I wrote a prompt for a flat graphic look rather than a rendered illustration.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;flat-color anime poster, screen print aesthetic, editorial illustration,
graphic shapes, thick expressive outlines, cel-shaded character, minimal
background, red and cream color scheme, vintage anime poster influence, clean
silhouette, controlled color blocking
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Negative prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;photorealistic, realistic skin, 3d render, realistic lighting, soft blurry
shading, watercolor, painterly, messy lineart, sketch, rough drawing,
excessive details, overly complicated background, gradient background, text,
typography, logo, watermark, bad anatomy, extra fingers, extra arms, deformed
hands, poorly drawn face, low quality, blurry
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Score: 8.&lt;/strong&gt; Screen print look, flat colour blocking, red and cream palette and retro poster feel all landed. The negative prompt excluded text, typography and logos, and the output arrived with Japanese vertical type top left, English layout text bottom right, and registration marks in all 4 corners.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt Helper was enabled on that run&lt;/strong&gt;, so the prompt was rewritten before generation. The rewrite improved the result.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fggetruwboykcv8j1x52z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fggetruwboykcv8j1x52z.png" alt=" " width="799" height="263"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7zg2id4g523hbc98kvvq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7zg2id4g523hbc98kvvq.png" alt="Left with Prompt Helper on. Right the same prompt with it off" width="799" height="529"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left with Prompt Helper on. Right the same prompt with it off&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  R4: hand over control in stages — partial pass
&lt;/h2&gt;

&lt;p&gt;References covered one rung, so the remaining question was what LoRAs do to a face. I ran an identical close-up portrait prompt twice, adding 2 LoRAs on the second pass.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: 2 text-only generations, no reference. Identical prompt and negative prompt, LoRAs added on the second.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;masterpiece, best quality, highly detailed modern anime illustration,
beautiful young adult anime woman, extreme close-up portrait, face and
shoulders filling most of the frame, looking directly at the viewer, slightly
tilted head, calm confident expression, subtle mysterious smile, large
expressive eyes with intricate irises and realistic reflections, soft delicate
facial features, smooth natural skin, detailed eyelashes, slightly parted lips,

long silky dark blue-black hair framing her face, individual strands of hair
catching the light, a few loose strands crossing her forehead and cheek,
modern stylish appearance, elegant black sleeveless top, small silver earrings,

cinematic soft lighting, warm light on one side of her face and cool rim light
along her hair, subtle glow around the eyes, soft shadows, beautiful skin
shading, delicate highlights, atmospheric depth, shallow depth of field,
softly blurred abstract background, dark blue and violet tones, subtle bokeh,
sophisticated color grading,

contemporary anime key visual, premium anime illustration, refined linework,
semi-realistic anime proportions, detailed face, clean composition, intimate
portrait photography composition, focus entirely on the eyes and face,
sophisticated, elegant, high-end anime artwork
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Negative prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;low quality, worst quality, blurry, pixelated, bad anatomy, deformed face,
asymmetrical eyes, crossed eyes, malformed eyes, extra limbs, extra fingers,
poorly drawn hands, childish appearance, old woman, overly exaggerated
breasts, chibi, cartoonish, flat colors, simplistic face, thick outlines,
heavy cel shading, vintage anime, retro anime, poster design, vector art,
graphic design, text, logo, watermark, excessive accessories, cluttered
background, oversaturated colors, plastic skin, photorealistic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Plain run: 9.5.&lt;/strong&gt; Warm key light on one side of the face and cool rim light along the hair both resolved, and every negative prompt filter was respected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LoRA run: 9.2.&lt;/strong&gt; &lt;a href="https://pixai.art/en/model/1838160874188056340/2044126778199703636" rel="noopener noreferrer"&gt;Crystal Eyes&lt;/a&gt; and &lt;a href="https://pixai.art/en/model/1892005535733745223/2044126787234234659" rel="noopener noreferrer"&gt;Body Aesthetics 5.2v&lt;/a&gt; sharpened linework, tightened iris detail, and converted the closed mouth from the first run into slightly parted lips with a specular highlight.&lt;/p&gt;

&lt;p&gt;They also flattened the lighting. The warm and cool split collapsed into a uniform cool blue, so the run gained detail and lost the property that made the first image interesting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A LoRA is a trade rather than an upgrade.&lt;/strong&gt; That's the reason this requirement gets a partial pass rather than a full one: the control is real, but it isn't additive, and stacking 3 of them onto a character you already like will cost you something you weren't tracking.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7yz59gwnqhipveynxhk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7yz59gwnqhipveynxhk.png" alt=" " width="799" height="407"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv8ggzzxpb23a2ioc6fgy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv8ggzzxpb23a2ioc6fgy.png" alt="Left no LoRA. Right with Crystal Eyes and Body Aesthetics loaded." width="800" height="525"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left no LoRA. Right with Crystal Eyes and Body Aesthetics loaded.&lt;/em&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The finish check
&lt;/h2&gt;

&lt;p&gt;None of the 4 requirements tell you whether the design itself is any good. Two final tests do.&lt;/p&gt;
&lt;h3&gt;
  
  
  Shape
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;Method: new generation from the Test 2 image as reference.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Convert the character in the reference image into a solid black silhouette on
a flat white background. Full body, standing, weight on one leg, no facial
features, no interior lines, no colour, no shading.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Score: 9.0.&lt;/strong&gt; The long hair shape and the bulk of the heavy boots both survive with all interior detail removed, and the weight shift onto one leg stays legible. The arms overlap the torso enough that a hand in a pocket doesn't resolve.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs41oftxs0vseblf9ui0y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs41oftxs0vseblf9ui0y.png" alt="Left test 2 image as reference. Right new generated image." width="800" height="552"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left test 2 image as reference. Right new generated image.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Subtraction
&lt;/h3&gt;

&lt;p&gt;Same locked prompt, word for word, with the squint and lanyard removed.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: new generation, no reference, no LoRA.&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;An original anime character, full body. Overgrown ash-grey hair pushed back
off his forehead. A navy aquarium staff windbreaker two sizes too big, sleeves
rolled to the elbow. Heavy black rubber boots. He is crouched at the base of a
tank, scraping algae off the glass with a long-handled tool, a black king
chess piece balanced on the ledge beside him. Wet floor, blue tank glow, one
fluorescent strip flickering overhead. Modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Score: 8.5.&lt;/strong&gt; Both eyes open, no lanyard, and hair, jacket, boots, crouch and lighting all identical. The scraper handle rendered shorter than in earlier runs.&lt;/p&gt;

&lt;p&gt;That result locates the design in the hair, the oversized jacket and the boots. The squint and the lanyard are personality rather than structure. Sorting your details into those 2 buckets is the difference between an anime OC creator producing a character and producing a costume.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkvex1bf8xchkf0wiczy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkvex1bf8xchkf0wiczy.png" alt="Left the silhouette. Right the subtraction run." width="799" height="661"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the silhouette. Right the subtraction run.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Grade table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Scores&lt;/th&gt;
&lt;th&gt;Grade&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;R1 Propose rather than execute&lt;/td&gt;
&lt;td&gt;Blind batch of 4&lt;/td&gt;
&lt;td&gt;7.5&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;R2 Hold what you keep&lt;/td&gt;
&lt;td&gt;Locked design, reference carry&lt;/td&gt;
&lt;td&gt;9.0, 8.5&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;R3 Support sideways movement&lt;/td&gt;
&lt;td&gt;3 genre swaps, 1 style swap&lt;/td&gt;
&lt;td&gt;9.0 / 8.5 / 5.5, 8&lt;/td&gt;
&lt;td&gt;Pass with boundary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;R4 Hand over control in stages&lt;/td&gt;
&lt;td&gt;LoRA on and off&lt;/td&gt;
&lt;td&gt;9.5, 9.2&lt;/td&gt;
&lt;td&gt;Partial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Finish check&lt;/td&gt;
&lt;td&gt;Silhouette, subtraction&lt;/td&gt;
&lt;td&gt;9.0, 8.5&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Failure modes, sorted
&lt;/h2&gt;

&lt;p&gt;Across 8 tests the failures cluster into 3 predictable groups, which makes them easy to plan around.&lt;/p&gt;

&lt;p&gt;Small physical logic goes first. A chess piece routed through a pocket, a badge hanging against gravity, a piece floating a centimetre above an open palm. The model understands the objects and mishandles the contact between them.&lt;/p&gt;

&lt;p&gt;Camera instructions go second. A close shot returned as a medium shot, and framing language lost to reference data every time the two disagreed.&lt;/p&gt;

&lt;p&gt;Negative prompts about text go third. An explicit exclusion of text, typography and logos produced an image with 2 languages of type and registration marks in the corners.&lt;/p&gt;

&lt;p&gt;Identity never failed. Not through description, not through reference, not through a genre change that got everything else wrong.&lt;/p&gt;

&lt;p&gt;For an anime character creator AI workflow, that ordering tells you what to inspect. Let it propose the design, name what you keep, then verify the design without its details before committing. That sequence is the one I'd run any AI anime character creator through.&lt;/p&gt;

&lt;p&gt;For choosing a base model before you start, the &lt;a href="https://blog.pixai.art/en/sdxl-anime-models-pixai-guide/" rel="noopener noreferrer"&gt;anime models guide&lt;/a&gt; covers the differences, and the &lt;a href="https://blog.pixai.art/en/how-to-use-pixai-guide/" rel="noopener noreferrer"&gt;how to use PixAI guide&lt;/a&gt; covers where references and LoRAs live in the panel.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next move
&lt;/h2&gt;

&lt;p&gt;The useful property of an anime character generator is not drawing quality. It's proposing the parts of a character you haven't decided yet, so you have something to react against instead of a blank canvas.&lt;/p&gt;

&lt;p&gt;Run the short version of this rubric yourself, because it is the cheapest way to create anime character with AI tools rather than generate pictures with them. Write 1 sentence about a person containing a role and a past and nothing visual. Generate 4. Keep 3 details, name them in a new prompt, and add 2 the model never offered. Then delete those 2 and regenerate.&lt;/p&gt;

&lt;p&gt;A character still standing after that deletion is yours. One that collapses was a costume, and you now know which layer to work on next.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>I ran a full anime generation session on my phone. Here's what each stage cost.</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:38:06 +0000</pubDate>
      <link>https://dev.to/naveedoss/i-ran-a-full-anime-generation-session-on-my-phone-heres-what-each-stage-cost-4984</link>
      <guid>https://dev.to/naveedoss/i-ran-a-full-anime-generation-session-on-my-phone-heres-what-each-stage-cost-4984</guid>
      <description>&lt;p&gt;Mobile creative tools usually get reviewed as a smaller version of the desktop one. Fewer buttons, same job. That framing hides the actual difference.&lt;/p&gt;

&lt;p&gt;A phone and a desktop are solving different problems. Desktop optimises for control: wide screen, precise input, side-by-side comparison, a keyboard that can carry a 70-word prompt without a fight. A phone optimises for proximity to the idea, which shows up on a bus or in a queue or right before you sleep, and which shrinks every minute it stays unrecorded.&lt;/p&gt;

&lt;p&gt;So the question for an anime AI app isn't whether it has the same feature list as the desktop tool. It's whether it can hold a session end to end, through 5 distinct jobs, without handing you back to a laptop halfway.&lt;/p&gt;

&lt;p&gt;I ran that test on &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt;. Below is each stage, what it was supposed to do, what happened, and a score.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7vpagyhh81kftntorath.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7vpagyhh81kftntorath.png" alt="PixAI mobile screenshots x3" width="800" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 1: capture
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The job:&lt;/strong&gt; get a fully specified idea out of your head and into the tool using only thumbs.&lt;/p&gt;

&lt;p&gt;Prompt entry is the first thing to go wrong on mobile. If the input box, the settings and the results don't all work at thumb distance, the session ends before it starts.&lt;/p&gt;

&lt;p&gt;I wrote a prompt with enough specific detail to score properly, then ran it unchanged on 2 models.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime illustration. A young woman sits cross-legged on a folding
table in an empty laundromat at midnight, eating instant noodles from a cup.
She has box braids tied in a high bun, silver hoop earrings, a cropped red
varsity jacket with a white letter M on the chest, and chipped blue nail
polish. One dryer spins behind her. Green neon from the street falls through
the window. Low angle, modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tsubaki.2 came in at 3 out of 5.&lt;/strong&gt; Wardrobe and mood landed: the red varsity jacket with the letter M, the silver hoops, the blue nails, the cup noodles, the green neon, the low angle. Composition went elsewhere. Standard twin braids instead of box braids in a high bun, a metal folding chair instead of cross-legged on the folding table, and a dryer that stands still.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;&lt;strong&gt;Tsubaki.3&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;came in at 4 out of 5&lt;/strong&gt; on the same input. It delivered box braids in a high bun, the cross-legged pose on the folding table and motion blur on the spinning dryer, on top of everything Tsubaki.2 had already got right. Its one miss was the chipped finish on the nail polish, which both models rendered as solid blue. It also left a floating chopstick tip near her right hand and misaligned frame bars under the table top, neither of which shows at normal viewing size.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage verdict: passed.&lt;/strong&gt; The prompt went in on a phone keyboard and came back with 2 usable candidates. Slower than typing, but not a blocker.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgfhc9l7f3sftrpa3pdh7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgfhc9l7f3sftrpa3pdh7.png" alt="Test 1-Left Tsubaki.2. Right Tsubaki.3. Below the mobile generation screen." width="800" height="658"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 1-Left Tsubaki.2. Right Tsubaki.3. Below the mobile generation screen.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiqlgy0o8pdfb0ufo0vpq.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiqlgy0o8pdfb0ufo0vpq.gif" alt="recording" width="560" height="1213"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 2: routing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The job:&lt;/strong&gt; let you change which model runs the prompt, without burying the selector.&lt;/p&gt;

&lt;p&gt;One comparison is an anecdote. To check whether model choice was doing real work or whether the first result was noise, I ran a second prompt built to strain 2 things at once: conflicting light sources and dense clothing detail.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A modern subculture Japanese anime street-style illustration, shot from a low
angle. A stylish young woman with twin pigtails and pink-tinted hair stands in
front of a brightly lit Shibuya vending machine at night. She wears an
oversized Y2K metallic windbreaker, low-rise cargo pants, platform boots, and
a wireless headset hanging around her neck. Neon pink and teal lights reflect
off her outfit. Flash photography lighting effect, crisp outlines, high detail,
vibrant techwear aesthetic.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tsubaki.3 scored 4 out of 5.&lt;/strong&gt; Pink twin pigtails, metallic windbreaker, cargo pants, platform boots and vending machine all arrived, with neon pink and teal reflecting off the jacket. Three misses: the headset went over her ears when the prompt asked for it around her neck, one ear cup merged into her pigtail, and the cargo pant suspender straps hang from nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tsubaki.2 scored 3 out of 5.&lt;/strong&gt; The iridescent fabric shading is the strongest single element in either image, and the headphones stayed correctly around her neck. It then added a second set of earbuds alongside them, cropped the platform boots off the bottom of the frame, skipped the low-rise cargo styling, and ran a vending machine panel line through her jacket cuff.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both models ignored the flash photography effect&lt;/strong&gt; and substituted stylised rim lighting.&lt;/p&gt;

&lt;p&gt;Across 2 prompts, model choice moved the output more than wording did. That makes the model selector a primary control rather than a settings-menu item, and any AI anime generator app that hides it is shipping a worse product than it thinks. Re-running the same prompt on a second model takes seconds once you know where the selector lives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage verdict: passed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu43iha4v0cf41k2o5r2q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu43iha4v0cf41k2o5r2q.png" alt="Test 2-Left Tsubaki.3. Right Tsubaki.2." width="800" height="659"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 2-Left Tsubaki.3. Right Tsubaki.2.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 3: reference ingest
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The job:&lt;/strong&gt; take an image already on the device and use it as the basis for generation.&lt;/p&gt;

&lt;p&gt;This is the stage where mobile has a structural advantage over desktop rather than a convenience one. The camera roll is already the reference folder. There is no export, no upload from a directory, no moving files between machines.&lt;/p&gt;

&lt;p&gt;I used a realistic photo of a woman leaning against a counter in a record store, attached it, and asked for an anime conversion.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic modern anime illustration, low angle shot. Redraw the character
and scene from @image1 in a high-quality modern anime art style. Maintain the
exact pose of the young woman leaning against the dark counter, her long
straight dark hair, her light grey ribbed crop top, faded blue denim jeans,
and over-ear headphones resting around her neck with a cord running down. Keep
the same ambient lighting, warm overhead rim highlights, and the background
record store shelving with the glowing blue neon accent at the counter base.
Sharp line art, rich shadows, high-detail finish.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;This scored 5 out of 5 on Tsubaki.3, the strongest result of the session.&lt;/strong&gt; Low angle, pose against the counter, long dark hair, ribbed crop top, distressed denim, headphones with the cord running down, wrist bracelet, vinyl shelving and the blue neon at the counter base all transferred into anime style with composition held.&lt;/p&gt;

&lt;p&gt;Two small defects: the headphone cord dissolves into the waistband of the jeans with no clear end, and the left arm joint is slightly over-smoothed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage verdict: passed, and this is the stage that justifies a mobile anime AI generator on its own.&lt;/strong&gt; A photo of a room, a screenshot of a pose, a sketch on paper photographed at your desk, all of them go straight in.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzy1qo6hld2c5pxc2tl0o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzy1qo6hld2c5pxc2tl0o.png" alt="test 3" width="800" height="594"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the source photo, by Sou Jest from Pixabay. Right: the anime version generated using Tsubaki.3. Below: attaching the reference on mobile.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 4: targeted revision
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The job:&lt;/strong&gt; change one attribute of a finished image without regenerating the rest of it.&lt;/p&gt;

&lt;p&gt;Tapping a generated image opens a row of options: Animate, Variations, Edit, Import to generate, Re-roll, Publish and Download. Edit opens a prompt box that also lets you choose which model handles the edit.&lt;/p&gt;

&lt;p&gt;I took the Tsubaki.3 laundromat image and swapped a single garment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Change her cropped red varsity jacket to a black windbreaker, keeping the white
letter M on the chest. Everything else stays the same: her braids in a high
bun, the silver hoops, the blue nail polish, the noodle cup, her pose on the
folding table, the spinning dryer behind her, the green neon through the window
and the camera angle.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;5 out of 5, with no misses.&lt;/strong&gt; The red varsity jacket became a black windbreaker with the white M in the same position and size, and the fabric drape changed with it. Everything else stayed put: braids, hoops, blue nails, noodle cup, the cross-legged pose, the dryer, the green neon and the camera angle. No new artifacts anywhere in the frame.&lt;/p&gt;

&lt;p&gt;One tap and one sentence to get from a finished image to a revised one. That's the line between an app you generate in and an app you work in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage verdict: passed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flxsjs9fgfmk7gvx59ii7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flxsjs9fgfmk7gvx59ii7.png" alt="Left: before. Right: after the jacket swap." width="800" height="658"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 5: persistence
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The job:&lt;/strong&gt; hand the session back to you intact after you've closed the app and gone to sleep.&lt;/p&gt;

&lt;p&gt;This is the stage that decides whether you can run a project from a phone or only a single session, and it's the one that rarely gets tested.&lt;/p&gt;

&lt;p&gt;Every generation saves into a Library with the date and time attached, so recent tasks are where you left them. Collections lives under it and holds pinned items, Artwork, and Models and LoRAs. Favorite Tasks keeps generations you want to return to. Published Artwork holds anything shared to the PixAI community, where people can comment and follow you. Uploaded Models and a Likes section round it out, the latter holding other people's work you saved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage verdict: passed.&lt;/strong&gt; Enough structure to reopen a generation from days ago, keep working from it, and find it again after that.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxc6rd9assle4h65m1rs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxc6rd9assle4h65m1rs.png" alt="The Library and Collections view." width="800" height="535"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Scorecard
&lt;/h2&gt;

&lt;p&gt;Five stages, one device, no desktop at any point.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Job&lt;/th&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Capture&lt;/td&gt;
&lt;td&gt;Get a specified idea in via thumbs&lt;/td&gt;
&lt;td&gt;Laundromat prompt, 2 models&lt;/td&gt;
&lt;td&gt;Tsubaki.2 3/5, Tsubaki.3 4/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. Routing&lt;/td&gt;
&lt;td&gt;Switch models without digging&lt;/td&gt;
&lt;td&gt;Shibuya prompt, 2 models&lt;/td&gt;
&lt;td&gt;Tsubaki.3 4/5, Tsubaki.2 3/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Reference ingest&lt;/td&gt;
&lt;td&gt;Use a photo from the gallery&lt;/td&gt;
&lt;td&gt;Record store photo to anime&lt;/td&gt;
&lt;td&gt;5/5, session best&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. Targeted revision&lt;/td&gt;
&lt;td&gt;Change one attribute only&lt;/td&gt;
&lt;td&gt;Jacket swap on Test 1 output&lt;/td&gt;
&lt;td&gt;5/5, no misses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5. Persistence&lt;/td&gt;
&lt;td&gt;Return the session tomorrow&lt;/td&gt;
&lt;td&gt;Library and Collections&lt;/td&gt;
&lt;td&gt;Passed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Where the phone stops being the better tool
&lt;/h2&gt;

&lt;p&gt;Long prompts are the obvious constraint. The prompts above run to 60 and 70 words, and a phone keyboard is slower and more error-prone than typing them.&lt;/p&gt;

&lt;p&gt;Comparing a batch of 4 results is easier on a large screen where all of them are visible at once instead of tapping through each. The same applies to anything holding many assets in play, like stacking several LoRAs or running a multi-step sequence where you keep referring back to earlier images.&lt;/p&gt;

&lt;p&gt;An &lt;a href="https://blog.pixai.art/en/pixai-app-guide/" rel="noopener noreferrer"&gt;AI image generator app&lt;/a&gt; covers more of the pipeline than one-off generation, and heavy sessions still move faster with a keyboard and a larger display.&lt;/p&gt;

&lt;p&gt;Split by who's using it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Creator type&lt;/th&gt;
&lt;th&gt;What mobile handles&lt;/th&gt;
&lt;th&gt;When to move to desktop&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Casual creator&lt;/td&gt;
&lt;td&gt;Whole workflow, generate to publish&lt;/td&gt;
&lt;td&gt;Rarely needed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OC and character work&lt;/td&gt;
&lt;td&gt;Edits, outfit swaps, new scenes&lt;/td&gt;
&lt;td&gt;Long character sheets, many variants side by side&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference-led creator&lt;/td&gt;
&lt;td&gt;Photo to anime in one session&lt;/td&gt;
&lt;td&gt;Multi-image compositing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Social content creator&lt;/td&gt;
&lt;td&gt;Generate, edit and publish in place&lt;/td&gt;
&lt;td&gt;Batch production runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heavy or technical creator&lt;/td&gt;
&lt;td&gt;Quick tests and captures&lt;/td&gt;
&lt;td&gt;Long prompts, LoRA stacking, comparing many outputs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What the failures tell you
&lt;/h2&gt;

&lt;p&gt;There are 2 categories of miss in this session, and they point at different things.&lt;/p&gt;

&lt;p&gt;Every failure in this session was model behaviour, not mobile behaviour. Tsubaki.2 read a folding table as a chair. Both models skipped a flash lighting effect. Neither drew chipped nail polish when asked for it. Run those same prompts on a desktop and you get the same misses, because the model is the same model.&lt;/p&gt;

&lt;p&gt;Nothing failed for being on a phone. The session moved between capture, routing, ingest, revision and retrieval without a handoff. As an anime AI art app, that's the property that matters more than any individual output score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next move
&lt;/h2&gt;

&lt;p&gt;The test for any anime AI app, and for any anime art generator mobile workflow, is resumability. Generation is the easy stage. Coming back 2 days later to an image you can still edit is what turns a phone into a workspace.&lt;/p&gt;

&lt;p&gt;Run the 2-stage version of this yourself. Photograph something near you, a corridor, a window, a shop front at night, and convert it into an anime scene. Then change one thing about the result.&lt;/p&gt;

&lt;p&gt;Under 5 minutes for both and you have your answer about whether this workflow fits how you work.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Best Free Anime AI Generator: What Can You Create Without Paying?</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Tue, 08 Sep 2026 07:20:46 +0000</pubDate>
      <link>https://dev.to/naveedoss/best-free-anime-ai-generator-what-can-you-create-without-paying-3kbp</link>
      <guid>https://dev.to/naveedoss/best-free-anime-ai-generator-what-can-you-create-without-paying-3kbp</guid>
      <description>&lt;p&gt;Free is the word doing the most work in this category and explaining the least. One generator gives you a set number of credits once and stops there. Another refills every day while reserving its best model for paying users. A third lets you generate and offers no way to edit afterwards, so the first image you get is the image you live with.&lt;/p&gt;

&lt;p&gt;I picked &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt; to test this on, because it refills daily, runs several anime models, and lets free users load community LoRAs. That's enough surface area to run a full creator session against and find out where the free version stops.&lt;/p&gt;

&lt;p&gt;So I opened a new free account and ran that session: one detailed character, three models, a community LoRA, then the same character again in a second scene.&lt;/p&gt;

&lt;p&gt;I ran out of credits on the first afternoon. A new account gets 10,000 free credits a day, and one batch of four images on Tsubaki.2 costs 7,800 of them.&lt;/p&gt;

&lt;p&gt;One thing to note before the results. Tsubaki.3 is PixAI's newest model at the time of testing. It shows up in the model list with a "test" label, but it won't run on a regular free account because it's currently limited to invited testers. I generated those images on a separate invited account and included them so you can see how the new model compares with the other free models.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Free Means Here
&lt;/h2&gt;

&lt;p&gt;An anime AI generator free tier is easiest to judge on its credit number, which is also the least useful way to judge it. Ten thousand a day sounds generous until you price a generation.&lt;/p&gt;

&lt;p&gt;On the account I tested, a single Tsubaki.2 image costs 4,400 credits and a batch of four costs 7,800. Haruka v2 runs at 2,400 for one and 3,800 for four. So a day's free credits buy about two Tsubaki.2 images, or one batch with a little left over.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdotc5brrc8v7q6j70zq.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdotc5brrc8v7q6j70zq.gif" alt=" " width="8" height="4"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That number shapes everything else in this article. The question is how much of a creation workflow it covers before paying becomes relevant.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Compare in a Free Anime AI Generator
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Anime image quality.&lt;/strong&gt; Faces, hands, anatomy, line quality and whether the composition holds together. Resolution is the part that matters least.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model access.&lt;/strong&gt; Anime models draw differently from each other. A free tier that locks you into one output style stops you learning what you prefer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LoRA support.&lt;/strong&gt; LoRAs give targeted control over a character, a style or an outfit that text prompting alone won't reach. Whether free users can load them changes what's possible by a lot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Character development.&lt;/strong&gt; Making one good character is separate from making that character twice. Anyone working on an OC, a comic or recurring social posts needs the second thing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Iteration.&lt;/strong&gt; Whether you can continue from a result or only roll for a new one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Usage limits.&lt;/strong&gt; A person making one avatar and a person developing a character across thirty images run out at completely different points.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Far You Get Without Paying
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: The First Illustration
&lt;/h3&gt;

&lt;p&gt;I wrote one prompt with enough specific detail to score it properly, then ran it on all three models with no LoRA loaded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1srzce31i2p40iqfp9q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1srzce31i2p40iqfp9q.png" alt="pixai models" width="800" height="501"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime illustration. A lanky teenage boy busks with an electric guitar in a subway underpass. He has messy copper-red hair with a shaved undercut and freckles across his nose, and wears a mustard yellow puffer vest over a black hoodie, with chipped green nail polish on his left hand. He is mid-strum, leaning back with his eyes shut and his mouth open singing. A fat orange cat sleeps in his open guitar case on the ground in front of him. Shot from a low angle looking up past the guitar case. Cold white strip lights overhead, tiled walls covered in layered posters, a train blurred with motion behind him. Modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tsubaki.2 gave me an image I'd keep, at 7 out of 10.&lt;/strong&gt; The freckles, the green nail polish, the copper hair and the undercut all landed, and the background train carried real motion blur through the window cutouts. It drew an acoustic guitar body when I asked for electric, and it cropped so tightly that the strumming action got cut off and a second guitar case appeared behind his back to fill the space.&lt;/p&gt;

&lt;p&gt;Free users get Pro mode on Tsubaki.2. Ultra mode requires an upgrade.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Haruka v2 came in at 4 out of 10&lt;/strong&gt; on the same prompt. It misjudged the scene at the structural level, seating him inside a train car instead of in an underpass with a train passing behind. The cat ended up on the floor with no guitar case anywhere. The nail polish came out black, the freckles didn't appear, and the style is flatter and more vector-like than the Tsubaki output.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;&lt;strong&gt;Tsubaki.3&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;scored 8.5&lt;/strong&gt;, with the best composition of the three and the low angle handled well. Its misses were the freckles, the motion blur on the train, and guitar strings that never rendered.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;📷 &lt;strong&gt;Insert image here:&lt;/strong&gt; image4, image5, image6 — "Left: Tsubaki.3. Middle: Tsubaki.2. Right: Haruka v2."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Step 2: Model Choice Changes More Than Style
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzuaej7k4q7qhjjrv3eyv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzuaej7k4q7qhjjrv3eyv.png" alt="Step1-Left Tsubaki.3. Middle Tsubaki.2. Right Haruka v2." width="799" height="434"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left Tsubaki.3. Middle Tsubaki.2. Right Haruka v2.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's the argument for a free anime image generator giving you model options at all.&lt;/p&gt;

&lt;p&gt;If Haruka had been the only model I could reach, I'd have concluded the tool couldn't handle a complex scene, when the real answer was that I'd picked the wrong model for this particular prompt.&lt;/p&gt;

&lt;p&gt;Haruka v2 belongs to the SDXL family and responds better to tag-style prompts. Tsubaki.2 is a DiT model built for natural language, which is what I gave it.&lt;/p&gt;

&lt;p&gt;Matching your prompt style to the model is part of model choice, and you can only learn that by running both. The &lt;a href="https://blog.pixai.art/en/how-to-use-pixai-guide/" rel="noopener noreferrer"&gt;how to use PixAI guide&lt;/a&gt; covers where those differences come from.&lt;/p&gt;
&lt;h3&gt;
  
  
  Step 3: Adding a Community LoRA
&lt;/h3&gt;

&lt;p&gt;Free users can load community LoRAs during generation, up to five per task on the account I tested. This is separate from training your own LoRA, which costs credits and comes with monthly allowances on paid plans.&lt;/p&gt;

&lt;p&gt;I added &lt;a href="https://pixai.art/en/model/2029061754756826659/2044126796285542882" rel="noopener noreferrer"&gt;one style LoRA&lt;/a&gt; and reran the same prompt on all three models, changing nothing else.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faldwnltug6xtomjs9biw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faldwnltug6xtomjs9biw.png" alt="Stage 3-LoRA added" width="800" height="274"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every model improved, and the two lower-scoring ones improved most.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Haruka v2 went from 4 to 5.5.&lt;/strong&gt; The LoRA corrected the environment error, moving him from inside a train car to standing on a platform, and produced a proper electric guitar. It still lost the cat and the guitar case entirely, put his eyes open when the prompt said shut, and kept the nail polish black.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tsubaki.2 went from 7 to 8.&lt;/strong&gt; The guitar came out electric, the cat landed inside the open case, and the character details sharpened. The tight vertical crop stayed, pulling the case up to waist height instead of the ground.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tsubaki.3 went from 8.5 to 9,&lt;/strong&gt; with cleaner line definition and a correct guitar, though the cat ended up on top of a closed case.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A free LoRA corrected structural errors that better prompting hadn't.&lt;/strong&gt; Adding one changed the guitar type on two models and rebuilt a scene the third had misunderstood.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7t8rbed5ak6f1us8u4f9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7t8rbed5ak6f1us8u4f9.png" alt="Stage 3-Left Tsubaki.3 with LoRA. Middle Tsubaki.2 with LoRA. Right Haruka v2 with LoRA. Below the LoRA panel." width="800" height="438"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left Tsubaki.3 with LoRA. Middle Tsubaki.2 with LoRA. Right Haruka v2 with LoRA. Below the LoRA panel.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Step 4: The Same Character in a Second Scene
&lt;/h3&gt;

&lt;p&gt;Reference-based editing is where most character work happens, so I took the harder route available to a free user: keep the model and the LoRA, describe the character again in text, and attach no reference image at all.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime illustration. The same lanky teenage boy, messy copper-red hair with a shaved undercut and freckles across his nose, wearing a mustard yellow puffer vest over a black hoodie with chipped green nail polish on his left hand, now sits on the back step of a night bus with the guitar case across his knees, counting coins into his palm. The fat orange cat is curled asleep on the seat beside him. Medium shot from across the aisle, warm interior bus light against the dark street through the window, rain on the glass. Modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tsubaki.2 carried all five anchors across with no reference image, at 9 out of 10.&lt;/strong&gt; The copper hair with the undercut, the yellow vest over the hoodie, the green polish, the freckles across the nose bridge, and the orange cat asleep on the seat. The jaw angle, ear shape and eye silhouette matched the same boy from the subway image. It got the coin counting, the case across his lap and the rain on the glass.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tsubaki.3 did the same thing at 9.5.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Haruka v2 kept three of five and lost the person.&lt;/strong&gt; Black nails again, no freckles, and a softer rounder face that looks like a younger, different character. It also skipped the action, drawing him asleep with the coins scattered on the seat, put the guitar case vertically behind his back, and lit a night bus with what looks like daylight.&lt;/p&gt;

&lt;p&gt;So character continuity is available on the free tier, on the right model, using text and a LoRA rather than a reference image.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F626sfhcmcy6dkxg3hksj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F626sfhcmcy6dkxg3hksj.png" alt="Stage 4-Left Tsubaki.3 with LoRA. Middle Tsubaki.2 with LoRA. Right Haruka v2 with LoRA." width="799" height="439"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left Tsubaki.3 with LoRA. Middle Tsubaki.2 with LoRA. Right Haruka v2 with LoRA.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Free Starts to Feel Limiting
&lt;/h2&gt;

&lt;p&gt;Two different things get called limits, and separating them matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unavailable.&lt;/strong&gt; Tsubaki.3 appears in the model list with a "test" label on it and won't run without an invite. Tsubaki.2's Ultra mode needs an upgrade, though Pro mode is free.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Available but a bit expensive.&lt;/strong&gt; Ten thousand credits a day gets you about two single Tsubaki.2 images, or one batch of four with 2,200 left over. On Haruka v2 you get about four singles. Paid members get 12,000 daily on the entry tier, plus monthly bonus credits.&lt;/p&gt;

&lt;p&gt;That plays out differently depending on what you're making.&lt;/p&gt;

&lt;p&gt;An occasional creator making an avatar or a profile picture will barely notice. A learner comparing models and prompts, which is what this article was, runs dry in an afternoon and waits for tomorrow. I did that more than once.&lt;/p&gt;

&lt;p&gt;An OC creator developing a character across many images feels it hardest, since character work is repetition by nature. A LoRA-focused creator is fine using community LoRAs, and runs into a separate question if they want to train their own. The &lt;a href="https://blog.pixai.art/en/pixai-free-credits-guide/" rel="noopener noreferrer"&gt;free credits guide&lt;/a&gt; covers the daily missions that add to the base amount.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who the Free Experience Is Enough For
&lt;/h2&gt;

&lt;p&gt;As a free AI anime art generator it clears the bar I'd set, since I produced an image at 8 out of 10 and a matching second scene at 9 without spending anything. That makes it enough for learning, for occasional anime art, and for deciding whether the tool suits you.&lt;/p&gt;

&lt;p&gt;It gets thin if you generate daily, work in batches to pick the best of four, or develop one character across a long series. The quality is there and the pace is what runs out.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Paying Starts to Make Sense
&lt;/h2&gt;

&lt;p&gt;Paying makes sense when you're waiting on tomorrow's credits more than once a week. That's the trigger.&lt;/p&gt;

&lt;p&gt;The other reasons are specific rather than general: you want to train your own LoRA rather than borrow community ones, you need Ultra mode on Tsubaki.2, you want batches of four every time so you can pick, or you're producing on a schedule instead of experimenting. The &lt;a href="https://blog.pixai.art/en/membership-payment-faq-web/" rel="noopener noreferrer"&gt;membership FAQ&lt;/a&gt; covers what each tier changes. Paying doesn't make the art better, it makes more of it possible per day.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fizwc9nqr67wpqeo2ujn6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fizwc9nqr67wpqeo2ujn6.png" alt="pixai membership" width="800" height="662"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Is PixAI Free
&lt;/h2&gt;

&lt;p&gt;Yes, and with real conditions attached. You can generate anime art, choose between anime models, load up to five community LoRAs per task, and carry a character into new scenes without paying anything. The newest model is invite-only and Ultra mode is paid, and the daily credit refill is what decides how much you get done.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you want to do&lt;/th&gt;
&lt;th&gt;Free tier&lt;/th&gt;
&lt;th&gt;What to watch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Make an occasional anime image&lt;/td&gt;
&lt;td&gt;Enough&lt;/td&gt;
&lt;td&gt;Single generations stretch the day further than batches&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learn what models and prompts do&lt;/td&gt;
&lt;td&gt;Enough&lt;/td&gt;
&lt;td&gt;Comparison testing burns credits fast&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use community LoRAs&lt;/td&gt;
&lt;td&gt;Enough&lt;/td&gt;
&lt;td&gt;Up to five per task, and they help the lower-scoring models most&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Develop one character over many images&lt;/td&gt;
&lt;td&gt;Workable&lt;/td&gt;
&lt;td&gt;Text plus a LoRA retains identity without a reference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Produce daily or in volume&lt;/td&gt;
&lt;td&gt;Thin&lt;/td&gt;
&lt;td&gt;Roughly two Tsubaki.2 images a day&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Train your own LoRA&lt;/td&gt;
&lt;td&gt;Separate&lt;/td&gt;
&lt;td&gt;Costs credits, with monthly allowances on paid plans&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Your Next Move
&lt;/h2&gt;

&lt;p&gt;10k credits a day is about two good generations. That's a slow pace for producing work and a fine one for learning what you like, and the distance between those two is what you're deciding.&lt;/p&gt;

&lt;p&gt;So spend the first day finding out which one you're doing. Write a prompt with five specific details in it, run it on Tsubaki.2 and Haruka v2, then add a community LoRA and run it again.&lt;/p&gt;

&lt;p&gt;Three generations, well inside a day's credits, and you'll know whether the free tier is a place to work or a place to try things out. If you'd rather not pick models yourself while you learn, &lt;a href="https://eap.pixai.art/go/naveed2" rel="noopener noreferrer"&gt;Mio&lt;/a&gt; makes those choices for you.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>What "Best Anime AI Generator" Reviews Skip Over</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Mon, 07 Sep 2026 07:27:10 +0000</pubDate>
      <link>https://dev.to/naveedoss/what-best-anime-ai-generator-reviews-skip-over-1n4</link>
      <guid>https://dev.to/naveedoss/what-best-anime-ai-generator-reviews-skip-over-1n4</guid>
      <description>&lt;p&gt;Every anime AI generator leads with its best image. That tells you the model can produce one good picture. It tells you nothing about your second, your fifth, or the twentieth time you try to draw the same character.&lt;/p&gt;

&lt;p&gt;Most people choosing a tool are looking at galleries and comparing how nice the output looks. That is the one thing every generator markets and the one thing that predicts the least about whether you can keep working after the first result.&lt;/p&gt;

&lt;p&gt;So instead of another gallery comparison, I built one original character and ran her through four things a creator does in real work: one detailed generation, the same character in new scenes, a LoRA run, and a single edit to a finished image. I used &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt; as the test bench, and I am reporting what missed as well as what landed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why anime generation asks for different things
&lt;/h2&gt;

&lt;p&gt;General image models are judged on whether a picture looks good on its own. Anime work usually is not a single picture.&lt;/p&gt;

&lt;p&gt;You are drawing an original character who has to stay the same person across a dozen images. You are working in stylized proportions where a small anatomical error is obvious immediately. You are building a comic, a character sheet, a set of social posts, or fan art in a specific style.&lt;/p&gt;

&lt;p&gt;That changes what to test. Not just how good one image looks, but whether the design survives a second scene, whether you can control style without losing the character, and whether you can change one thing later without regenerating everything. That is the real bar for calling something the best anime AI generator, not the front page of a gallery.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test 1: model quality
&lt;/h2&gt;

&lt;p&gt;Model quality is not about resolution. It covers face and eye rendering, hands, line quality, how clothing folds, and whether the composition holds together as an illustration.&lt;/p&gt;

&lt;p&gt;I wrote one prompt loaded with checkable details rather than adjectives.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime illustration. A young woman runs a vinyl record stall in a
covered night market. She has a short green-dyed bob and round wire glasses,
and wears an oversized denim jacket with three enamel pins on the left lapel,
a red bandana knotted on her right wrist and a silver ring on her left index
finger. She is flipping forward through a crate of records with both hands,
head tilted down, half smiling at something she has found. Shot from a low
angle looking up past the crate. Hanging bulbs above her, a glass display
case along the front of the stall reflecting her jacket, handwritten price
cards on the crates, steam from a food stall two units down. Warm bulb light
against the cool blue of the market roof, modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The face, the lighting, and the environment all came through well. Clean linework, a clear expression, the warm bulb light against the cool blue roof, the reflection in the glass case, and the price cards on the crates.&lt;/p&gt;

&lt;p&gt;PixAI runs several anime models rather than one, which matters more than it sounds. Different models draw faces and lines differently, so the right choice depends on whether you want modern anime, classic anime, or something closer to Korean illustration. The &lt;a href="https://blog.pixai.art/en/sdxl-anime-models-pixai-guide/" rel="noopener noreferrer"&gt;SDXL anime models guide&lt;/a&gt; covers those differences if you want to pick deliberately rather than defaulting.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyb0980812139pcf528pt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyb0980812139pcf528pt.png" alt="Test 1: the night market generation." width="800" height="600"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 1: the night market generation.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Test 2: prompt following
&lt;/h2&gt;

&lt;p&gt;An attractive image that ignores your instructions is still a failed generation. So I scored the same image against the eleven things I asked for.&lt;/p&gt;

&lt;p&gt;Ten and a half of eleven landed. The green bob, the round glasses, the bandana on her right wrist, the low camera angle, the reflection, the hanging bulbs, the warm and cool light split, and her expression were all correct.&lt;/p&gt;

&lt;p&gt;The three enamel pins passed on a closer look: two graphic pins and a metallic collar pin, which is three items on the left lapel as asked.&lt;/p&gt;

&lt;p&gt;The silver ring landed on the wrong finger. I asked for her left index and got her left middle.&lt;/p&gt;

&lt;p&gt;That is a useful thing to know before you choose a tool. Big instructions about pose, camera, lighting, and environment get followed. Instructions about which finger, which hand, or how many of something are where an AI anime image generator starts guessing.&lt;/p&gt;
&lt;h2&gt;
  
  
  Test 3: character consistency
&lt;/h2&gt;

&lt;p&gt;This is the section that decides whether a generator is usable for original characters, comics, or anything you plan to continue.&lt;/p&gt;

&lt;p&gt;Consistency does not mean identical. It means someone looking at two images agrees they are the same person. I took the Test 1 image and used it as a reference for two new situations, rather than writing a fresh prompt each time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scene A&lt;/strong&gt;, asleep on a bus, close and cold:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The same woman from the reference image, now asleep against the window of an
empty early-morning bus, forehead on the glass, arms folded. Keep her face,
the green bob, the round glasses, the denim jacket with three enamel pins on
the left lapel, the red bandana on her right wrist and the silver ring on
her left index finger exactly as they are. Close framing from the seat
beside her, pale grey dawn light through the window, empty seats behind,
modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Scene B&lt;/strong&gt;, laughing on a rooftop, wide and warm:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The same woman from the reference image, now laughing on a rooftop at
sunset while holding a paper cup, one arm resting on a railing. Keep her
face, the green bob, the round glasses, the denim jacket with three enamel
pins on the left lapel, the red bandana on her right wrist and the silver
ring on her left index finger exactly as they are. Wide shot with the city
behind her, warm low sun, modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Asleep and close, then laughing and far. Cold dawn light, then warm sunset.&lt;/p&gt;

&lt;p&gt;Her core design carried across both scenes. Facial features, haircut, green tone, glasses, and jacket stayed the same and matched the original. The three pins stayed locked to the left lapel in every generation. The bandana stayed on her right wrist.&lt;/p&gt;

&lt;p&gt;The ring moved every time. It was on her left middle finger in the first image, then jumped to her right index finger in the bus scene, and stayed on the right index finger on the rooftop.&lt;/p&gt;

&lt;p&gt;Six generations across this whole article, and the ring landed on the correct finger in none of them.&lt;/p&gt;

&lt;p&gt;One other slip: the rooftop came out as a medium shot rather than the wide shot I asked for.&lt;/p&gt;

&lt;p&gt;So the pattern is that identity transfers and small accessories drift. If you are building an OC, put your signature details somewhere large and structural rather than on a finger, or expect to correct them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdtqi1ddry5wj5bwtc8p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdtqi1ddry5wj5bwtc8p.png" alt="Left: the bus scene. Right: the rooftop scene." width="800" height="526"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the bus scene. Right: the rooftop scene.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Test 4: models, LoRAs, and references
&lt;/h2&gt;

&lt;p&gt;These three do different jobs, and knowing which one to reach for saves a lot of retries.&lt;/p&gt;

&lt;p&gt;A model sets the broad look and capability. A reference image carries a specific design forward. A LoRA teaches a targeted character, style, or outfit that you reuse.&lt;/p&gt;

&lt;p&gt;I ran the same desk scene twice to see what a LoRA changes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The same woman from the reference image, sketching in a notebook at a
cluttered desk late at night, desk lamp on, headphones around her neck. Keep
her face, the green bob, the round glasses, the denim jacket with three
enamel pins on the left lapel, the red bandana on her right wrist and the
silver ring on her left index finger. Modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On default &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt; with the reference attached, the scene, the lamp light, and the character identity all landed. The pin count went to four instead of three, and the ring was on the wrong finger again. Six of eight constraints met.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy3alg3untl71n0ox3d7u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy3alg3untl71n0ox3d7u.png" alt=" " width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then I ran into something I should report plainly. LoRAs would not apply while a reference image was attached. Repeated attempts returned the base image unchanged, so the two controls did not combine for me at all. That is a real workflow problem if you were planning to use both.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6b3oyrdv9s2hynawnnk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6b3oyrdv9s2hynawnnk.png" alt=" " width="799" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The way around it was to write the character into a text-only prompt instead of referencing her.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime illustration. A young woman with a short green-dyed bob and
round wire glasses sits sketching in an open notebook at a cluttered desk
late at night. She wears an oversized denim jacket with exactly three enamel
pins on the left lapel, a red bandana knotted on her right wrist, a silver
ring on her left index finger, and over-ear headphones resting around her
neck. Shot from a medium angle across the desk as she leans forward, focused
on her drawing. A single warm desk lamp illuminates the workspace, casting
strong directional light across her face, jacket, pens, loose papers, and
stacked books, contrasting against the cool dark shadows of the room. Modern
anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnvlfb5qlirh48ti9wam5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnvlfb5qlirh48ti9wam5.png" alt=" " width="800" height="402"&gt;&lt;/a&gt;&lt;br&gt;
Writing the description out again, instead of referencing the earlier image, is the workaround, and it is useful to know before you plan a workflow that leans on stacking a reference and a LoRA together.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx72h2qczo7hv4llojy4b.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx72h2qczo7hv4llojy4b.jpg" alt="Left: default Tsubaki.3 with the reference image. Right: the text-only prompt with Manga Style and Body Aesthetics loaded." width="800" height="528"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: default Tsubaki.3 with the reference image. Right: the text-only prompt with Manga Style and Body Aesthetics loaded.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What this tells you about picking a generator
&lt;/h2&gt;

&lt;p&gt;Put the four tests together and a real picture forms.&lt;/p&gt;

&lt;p&gt;Model quality is strong out of the box, and PixAI's multiple model options mean you can match the look to the project rather than settle for one house style.&lt;/p&gt;

&lt;p&gt;Prompt following is reliable for the structural instructions, pose, camera, lighting, environment, and shakier on precise small details like which hand or exactly how many of something, the same split you will see in almost any anime AI art generator or AI anime generator once you push past the first image.&lt;/p&gt;

&lt;p&gt;Character consistency holds up well for the big design elements, face, hair, outfit, and the specific accessory that keeps drifting is useful to know before you build a whole comic around a signature ring or a similarly small prop.&lt;/p&gt;

&lt;p&gt;Models, LoRAs, and references each do a distinct job, and the two do not currently combine, so plan your workflow around one or the other rather than assuming you can stack them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verdict
&lt;/h2&gt;

&lt;p&gt;Calling something the best anime AI generator should mean more than a pretty front-page image. It should mean the design survives a second scene, the instructions that matter get followed, and you know where the small drift happens before you commit a project to it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhxbtdux4xrx6ayqfcbda.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhxbtdux4xrx6ayqfcbda.png" alt="table" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
By that standard, PixAI holds up. The core identity of an original character transfers cleanly across totally different scenes and lighting, the model quality is there without needing a lucky roll, and the honest gaps, finger placement, exact pin counts, and the LoRA-plus-reference conflict, are specific enough to plan around rather than vague warnings.&lt;/p&gt;

&lt;p&gt;If you want to see how your own character holds up, &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;try building one on PixAI&lt;/a&gt; and push it through a second and third scene before you judge the tool by its first image.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>One Job Per Stage: Running a Drawing from Line Art to Final Render on Tsubaki.3</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Thu, 03 Sep 2026 14:17:10 +0000</pubDate>
      <link>https://dev.to/naveedoss/one-job-per-stage-running-a-drawing-from-line-art-to-final-render-on-tsubaki3-g3h</link>
      <guid>https://dev.to/naveedoss/one-job-per-stage-running-a-drawing-from-line-art-to-final-render-on-tsubaki3-g3h</guid>
      <description>&lt;p&gt;An illustrator moves one drawing through stages. Line art first, then values, then flat color, then the finished render. Each stage exists so you can settle one thing before the next one covers it up.&lt;/p&gt;

&lt;p&gt;AI image tools normally skip all of that and hand you the last step. I wanted to see whether &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt; can work inside the stages instead of jumping past them.&lt;/p&gt;

&lt;p&gt;So I built one scene and moved it through the whole sequence, feeding each result into the next. Then I ran four more tests on the questions a single chain leaves open. The frame I judged everything against is simple: each stage has one job, and the test is whether it does that job and stays in its lane.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI art workflow has to protect
&lt;/h2&gt;

&lt;p&gt;A workflow is different from four pictures that happen to look alike. The character, the pose, the proportions, the camera, the composition, and the background should stay where they are, and the only thing that changes from stage to stage is how finished the image looks.&lt;/p&gt;

&lt;p&gt;The scene I used is a young man kneeling beside a half-stripped motorbike in an open garage at dusk. Bleached undercut, oil-stained tan coveralls tied at the waist over a black tank top, a red star patch on his left shoulder strap, a wrench in his right hand, the number 07 on the bike panel, a tabby cat asleep on a red toolbox behind him, a pendant lamp overhead, and the garage door open to the street.&lt;/p&gt;

&lt;p&gt;Those are the nine details I followed through every stage, so if one shifts or disappears I know where it happened. Each stage gets a pass, a partial pass, or a fail, based on whether it added the rendering I asked for and left the nine alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 1: the line-art base
&lt;/h2&gt;

&lt;p&gt;This is the structural starting point, and everything later in the AI drawing workflow refers back to it. Its one job is to draw the whole composition in clean line, with no color and no fills.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Clean black and white line art of an anime scene, no shading, no grey tones, no
color, white background wherever there is no linework. A lean young man with a
bleached undercut kneels beside a half-stripped motorbike inside an open garage
at dusk. He wears oil-stained tan coveralls tied at the waist over a black tank
top, with a small five-pointed star patch on the left shoulder strap. He holds a
wrench in his right hand and rests his left hand on the bike seat. The motorbike
faces left with the number 07 on its side panel. A tabby cat sleeps on a red
toolbox behind him. A single pendant lamp hangs above the bike. The open garage
door shows the street outside. Confident even line weight, the full composition
drawn, no cross-hatching, no fills.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All nine anchors came through, and the line-art instruction did not. The composition is right: the man kneeling in the correct pose with the wrench in his right hand and his left on the seat, the bike facing left, the cat asleep on the red toolbox behind him, the lamp overhead, the open door framing the street.&lt;/p&gt;

&lt;p&gt;What I asked for was black and white with no shading and no fills. What I got had tan coveralls, a red toolbox, an orange dusk gradient in the sky, and heavy solid black fills.&lt;/p&gt;

&lt;p&gt;Two anchors also landed in the wrong place. The star patch went onto the folded coverall waist strap instead of the shoulder strap, and the 07 went onto the front tank panel instead of the rear side panel. Both are close enough to pass as the same design, and both matter later.&lt;/p&gt;

&lt;p&gt;So the composition passed and the stage instruction failed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7aj4jjlistfkq8lwx3qd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7aj4jjlistfkq8lwx3qd.png" alt="The line-art base. I’ll go with the right one." width="800" height="551"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 2: line art to grayscale
&lt;/h2&gt;

&lt;p&gt;The Stage 1 image goes in, and only value should come out of it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add grayscale shading to this line art. Keep every line, the pose, the character
design, the number 07 on the bike panel, the star patch on the left shoulder
strap, the wrench in his right hand, the tabby cat on the red toolbox, the
pendant lamp and the camera framing exactly as they are. Add only value: soft
grey form shadows, cast shadows on the floor thrown by the pendant lamp above,
and darker values in the garage corners. No color anywhere, black white and grey
only. Do not redraw or restyle the linework.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This stage corrected the previous one. The color that leaked into Stage 1 was stripped out completely, leaving a strict black, white, and grey image. It added value without redrawing, so the linework, the pose, and the framing came through untouched.&lt;/p&gt;

&lt;p&gt;Every anchor held: the man, the wrench, the 07 panel, the star patch, the cat, the toolbox, and the lamp all stayed where Stage 1 put them. This stage passed on every count, and it turned out to be the only transition in the whole chain that cleaned up an earlier mistake rather than inheriting it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2dtq8b4sx1lqy7oc49nl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2dtq8b4sx1lqy7oc49nl.png" alt="Left: Stage 1 line art. Right: Stage 2 grayscale" width="800" height="533"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: Stage 1 line art. Right: Stage 2 grayscale&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One practical note. There are two ways to feed one image into the next stage. You can click the image, choose edit, and paste the prompt, or you can click the image, set it as the base image, and prompt from there. I used the edit method here, and the difference between the two shows up in the next stage.&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 3: flat colors
&lt;/h2&gt;

&lt;p&gt;This is the AI line art coloring step, and nothing but color should arrive. I ran it twice, once by editing the Stage 2 image and once by setting it as a base image reference.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add flat colors to this grayscale image. Keep the pose, the linework, the
framing, the number 07, the star patch, the wrench in his right hand, the
sleeping tabby cat, the red toolbox and the pendant lamp exactly as they are. Use
flat blocks of color with no gradients, no rendering and no lighting: bleached
blond hair, tan coveralls, black tank top, red star patch, deep blue motorbike
with a white 07 panel, red toolbox, brown tabby cat, grey concrete floor, warm
orange sky through the open door. Solid fills only, outlines still visible.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The edit run applied flat fills to the character and the bike, and then kept rendering anyway. The volumetric light cone under the pendant lamp stayed, and the sky came out as a smooth gradient rather than a block. Flats are not supposed to carry light.&lt;/p&gt;

&lt;p&gt;The palette was mixed on this run. Tan coveralls, black tank top, brown cat, and red toolbox all matched, but the panel behind the 07 came out blue instead of white, and the single star patch became two red stars on the hip flap. Structurally it was solid, and the linework from Stage 2 held almost perfectly, with no redrawing of geometry, pose, or background, and all nine anchors in position.&lt;/p&gt;

&lt;p&gt;The reference run did better. It removed the sky gradient and replaced it with a solid orange block, and it corrected the 07 panel to white. The light cone under the lamp survived that run too, and the star stayed on the hip flap.&lt;/p&gt;

&lt;p&gt;So it is a partial pass either way, with the reference method ahead of the edit method on color discipline. That difference is the reason I lean on the reference method whenever a stage has strict rules.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0fzw67i8gzut431n7l0f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0fzw67i8gzut431n7l0f.png" alt="Left: flats using edit. Right: flats using the image as a base reference." width="800" height="521"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: flats using edit. Right: flats using the image as a base reference.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 4: the final render
&lt;/h2&gt;

&lt;p&gt;The last step in this AI coloring workflow adds lighting, texture, and finish, with the design locked.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Render this flat-colored image into a finished illustration. Keep the pose, the
composition, the camera framing, the character design, the number 07 on the bike
panel, the star patch on the left shoulder strap, the wrench in his right hand,
the sleeping tabby cat on the red toolbox and the pendant lamp in the same
positions. Add lighting, soft shadows, material texture on the metal, the fabric
and the concrete, warm light from the pendant lamp and cooler dusk light from the
open door, and finished detail throughout. Do not change the design or move
anything.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dual lighting worked properly. Warm light falls across the garage frame, the character, and the toolbox, and cooler dusk light comes in from outside. Two temperatures, both coming from the sources I named.&lt;/p&gt;

&lt;p&gt;The material work is strong. Soft shading, ambient occlusion, and highlights across fabric folds, skin, bike metal, and the garage floor, none of it destroying the line definition underneath. Structure held from the earlier stages with no drift, so the pose, the wrench, the bike geometry, and the camera framing all stayed locked.&lt;/p&gt;

&lt;p&gt;Two things moved on the way through. The star patch is still on the hip flap, inherited from Stage 3, so this render keeps eight of nine anchors. And the sky drifted from the flat orange I set in Stage 3 to deep twilight blue, overriding a color decision the previous stage had already made. The stage passed, with that background color drift noted.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1h4v7ryio6n5ubb9h85.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh1h4v7ryio6n5ubb9h85.png" alt="Right: Stage 3 reference. Left: The final render" width="800" height="521"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Right: Stage 3 reference. Left: The final render.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The payoff test: does the middle matter?
&lt;/h2&gt;

&lt;p&gt;Four stages is more generations than one, so I went back to the Stage 1 line art and took two shortcuts from it directly, to see whether the middle steps were protecting anything.&lt;/p&gt;

&lt;p&gt;The first shortcut went straight to flats, using the same flat-color prompt pointed at the line art. It followed the no-lighting rule better than the chain did, with no light cone and no sky gradient, just pure unrendered fills. The chained version had inherited the lighting structure from Stage 2 and carried it forward, which is what leaked into the Stage 3 flats.&lt;/p&gt;

&lt;p&gt;The cost showed up in the background. It erased the exterior houses, the street layout, and the utility poles that Stage 1 had drawn. The outlines, framing, and base geometry inside the garage stayed locked, and all nine anchors are legible, with the same blue 07 panel and the same hip-flap star.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuo7mtskqbtv7ll4q1f76.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuo7mtskqbtv7ll4q1f76.png" alt="Stage 1 line art left, the straight-to-flats shortcut right" width="800" height="519"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Stage 1 line art left, the straight-to-flats shortcut right&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The second shortcut went straight to a final render, using the Stage 4 prompt pointed at the line art. The lighting here is good on its own terms, with warm lamp light blending into cool exterior light, and everything else went backwards from there.&lt;/p&gt;

&lt;p&gt;It painted over the line art instead of preserving it, softening the outlines into a digital painting. The exterior houses, street, and poles were replaced with a plain dark gradient. The coveralls shifted from tan to olive drab, the star turned yellow, and the bike went light grey instead of deep blue. It also added an unprompted 007 graphic on the lower frame rail above the rear wheel.&lt;/p&gt;

&lt;p&gt;So the middle stages are doing real work. The chained path keeps the background architecture, the color decisions, and the geometry, while the direct path renders light well and loses the drawing underneath it. The shortcut passed on lighting and failed on everything the workflow exists to protect.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fki7l5k5scwcde0ctxy1q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fki7l5k5scwcde0ctxy1q.png" alt="The straight-to-flats shortcut left, the straight-to-render shortcut right" width="800" height="525"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The straight-to-flats shortcut left, the straight-to-render shortcut right&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The nine anchors across four stages
&lt;/h2&gt;

&lt;p&gt;Laid out side by side, the pattern is clear.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Anchor&lt;/th&gt;
&lt;th&gt;Stage 1&lt;/th&gt;
&lt;th&gt;Stage 2&lt;/th&gt;
&lt;th&gt;Stage 3&lt;/th&gt;
&lt;th&gt;Stage 4&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Character design&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pose and wrench&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coveralls tied at waist&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Star patch&lt;/td&gt;
&lt;td&gt;On the waist strap, not the shoulder&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Doubled in the edit run&lt;/td&gt;
&lt;td&gt;Still on the waist strap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Number 07&lt;/td&gt;
&lt;td&gt;On the front tank panel&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Blue in edit, white in reference&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cat and red toolbox&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pendant lamp&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bike facing left&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Camera framing&lt;/td&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;td&gt;Held&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stage rule followed&lt;/td&gt;
&lt;td&gt;Failed, color and fills&lt;/td&gt;
&lt;td&gt;Passed&lt;/td&gt;
&lt;td&gt;Partial, light cone stayed&lt;/td&gt;
&lt;td&gt;Passed, sky drifted&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three things stand out. Geometry never drifted through the chain, so across four generations the pose, the framing, and the placement of every object held, which is the part I expected to give way first. Errors are inherited rather than corrected, so the star patch went to the wrong place in Stage 1 and stayed wrong through all four stages, because nothing downstream questions a decision made upstream. And Stage 2 was the only transition that cleaned something up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can color sit on top of value?
&lt;/h2&gt;

&lt;p&gt;Painters build value first and color second, and the usual failure is that color arrives and flattens the shading underneath. I used a new scene here, a weathered fisherman hauling a net over a small boat at dawn with a seagull on the prow, and ran it twice.&lt;/p&gt;

&lt;p&gt;The value base came out right. Strict monochrome, and the three value zones landed in the order I asked for, with the fisherman and the net darkest, the open water mid grey, and the low dawn sky lightest. The seagull is on the prow with a soft cast shadow under its feet, painted in value blocks rather than outlines.&lt;/p&gt;

&lt;p&gt;Then I added color over that value with an instruction not to relight or add highlights. The palette and the geometry both landed: cold blue-green water, pale peach sky, mustard cap, weathered brown wood, off-white gull, with the pose, the face, the net weave, and the cast shadow matching the base run almost pixel for pixel.&lt;/p&gt;

&lt;p&gt;The mustard cap is where it came apart. In the value pass that cap was a near-black charcoal tone deep in the dark zone, and adding yellow lifted it into a bright mid-tone. It applied the color as a new local value rather than laying the hue over the value that was already there. The three big zones kept their order, so the overall structure survived. A partial pass: correct colors, correct drawing, one value zone rewritten by its own hue.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fapaqadiddg1h70tf7tsx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fapaqadiddg1h70tf7tsx.png" alt="Left the value pass. Right color applied over it." width="800" height="652"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the value pass. Right color applied over it.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Lighting as its own stage
&lt;/h2&gt;

&lt;p&gt;Relighting is easy to fake when the light source is off-frame. Here the source is in the picture and the wet ground reflects it, so there are two places to check every change. The scene is a boy in a green raincoat feeding a stray dog outside a closed noodle stall, lit only by a pink neon bowl sign, with the wet ground reflecting the neon.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3tqctgcd868qfmj02m57.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3tqctgcd868qfmj02m57.png" alt="Test 7 I’ll go with the left one." width="799" height="528"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 7 I’ll go with the left one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The base render put everything in place, with one partial fail on the way in: a cyan ambient glow at the back of the alley, in a prompt that said the neon was the only light source. I then ran three relights from a shared template.&lt;/p&gt;

&lt;p&gt;The first turned the sign off and put a full moon overhead. The pink glow switched off completely, a moon drove a cold blue palette through the scene, every pink highlight on the wet ground was replaced with moonlight, and the shadows project forward, away from the moon. The pose and the alley details are identical to the base. A full pass, including the puddles.&lt;/p&gt;

&lt;p&gt;The second kept the sign on and added a car headlight from the far left. The car went down the center of the alley instead of entering from the left, and the headlights made it drop the global exposure, so the boy and the dog are buried in shadow with a lit neon sign right above them. High beams behind two subjects should rim their edges, and there is no edge light and no forward cast shadows. It inserted the car and then treated the result as a darkened composite rather than a scene with two lights in it.&lt;/p&gt;

&lt;p&gt;The third turned the sign off and put a flashlight in the boy's hand, and I fed it the image from the second relight rather than the base. The flashlight beam passes through the dog's legs instead of stopping at them, so the body looks semi-transparent and throws no shadow, and the car and its high beams carried over from the input. A fail on light physics, treated as an overlay rather than as light meeting an object.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv7zys2ls7otvgbfobh77.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv7zys2ls7otvgbfobh77.png" alt="Relight 1 (left), Relight 2 (middle), Relight 3 (right)" width="799" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Relight 1 (left), Relight 2 (middle), Relight 3 (right)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it correct a drawing you meant to be strange?
&lt;/h2&gt;

&lt;p&gt;Cleanup and rendering are where a model can quietly improve your work into something you did not draw. So I gave it a picture with a deliberate oddity: a woman holding an open umbrella upside down like a bowl to hold the rain, with a small heron standing inside the upturned umbrella, in a flooded street.&lt;/p&gt;

&lt;p&gt;The line-art stage kept the umbrella upside down with the bird inside it. The common version of this image is a woman holding an umbrella the normal way, and it did not reach for that. The flood water, the half-submerged bicycle, and the shuttered shopfronts are all in place. The line-art rule slipped again, the same way it did in Stage 1, with solid black fills in the hair, the bird, and the coat interior. Intent passed, formatting partly failed.&lt;/p&gt;

&lt;p&gt;The render held the premise all the way through. The umbrella stays inverted and catching water, and the dark bird silhouette from the line art resolved into a grey heron with a yellow bill, which is the model taking the word heron from the prompt and applying it to a shape it had already drawn. The bicycle and the shopfronts stayed anchored, the rain falls in vertical streaks with ripples on the surface, water drips off the umbrella ribs, and the flooded street reflects it all. Both stages passed, and this test tells me more about real production use than the render quality does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6n0kairn18d560y0c35d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6n0kairn18d560y0c35d.png" alt="Left the line-art stage. Right the finished render." width="800" height="527"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the line-art stage. Right the finished render.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this fits in real work
&lt;/h2&gt;

&lt;p&gt;As an AI art workflow, the chain is good for exploring rather than locking things down. Geometry survives every stage, so you can see one drawing as value, as flats, and as a finished render without redrawing it. That covers most of what an AI illustration workflow is for: testing a rendering direction, comparing color treatments, or taking a rough idea far enough to show someone.&lt;/p&gt;

&lt;p&gt;Two habits made this go better. Get the line art right before you move on, since errors are inherited and never questioned later. And use the base image reference method rather than edit when a stage has strict rules, which is what removed the sky gradient and corrected the 07 panel in Stage 3. The &lt;a href="https://blog.pixai.art/en/pixai-edit-pro-ai-image-editor/" rel="noopener noreferrer"&gt;Edit Pro guide&lt;/a&gt; and the &lt;a href="https://blog.pixai.art/en/pixai-reference-pro-guide-multi-image-editing-with-natural-language/" rel="noopener noreferrer"&gt;Reference Pro guide&lt;/a&gt; cover both.&lt;/p&gt;

&lt;p&gt;I would be careful with anything that needs the stage rules obeyed exactly, since clean line art never arrived, flats came back with lighting in them, and color changed the value under it. This AI art process gives you a reference version of those stages rather than a production one. Naming your anchors in every prompt does most of the work, and the &lt;a href="https://blog.pixai.art/en/how-to-write-pixai-prompts-formula/" rel="noopener noreferrer"&gt;prompt formula guide&lt;/a&gt; and the &lt;a href="https://docs.pixai.art/docs/prompts/prompt-advanced" rel="noopener noreferrer"&gt;advanced prompting page&lt;/a&gt; cover how to write them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The next step
&lt;/h2&gt;

&lt;p&gt;Tsubaki.3 supports an AI anime art workflow rather than replacing one. Across four chained stages the drawing never moved, and shortcuts from the line art cost me the background, the palette, and the outlines.&lt;/p&gt;

&lt;p&gt;The failures were about stage discipline. Line art came back with color and fills twice, flats came back with a light cone, and yellow lifted a dark value. It also kept a deliberately illogical umbrella through two stages without correcting it, which matters more for line art to finished art AI work than the rendering does.&lt;/p&gt;

&lt;p&gt;Take a drawing you have already made, name every element in it, and run it through two stages on &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;PixAI&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>art</category>
      <category>tutorial</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Can Tsubaki.3 Spell? Eight Tests of Readable Text Inside AI Images</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Wed, 02 Sep 2026 10:55:34 +0000</pubDate>
      <link>https://dev.to/naveedoss/can-tsubaki3-spell-eight-tests-of-readable-text-inside-ai-images-2ab8</link>
      <guid>https://dev.to/naveedoss/can-tsubaki3-spell-eight-tests-of-readable-text-inside-ai-images-2ab8</guid>
      <description>&lt;p&gt;Text is usually the part of an AI image that gives it away. The art looks fine, then you look at the title and one letter is wrong, or the small print under it is shapes instead of words.&lt;/p&gt;

&lt;p&gt;So I wanted two answers from &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, not one. Any AI image generator with text will put something on the poster. What matters for AI text in image work is whether the words are correct, and whether they land somewhere that works with the picture.&lt;/p&gt;

&lt;p&gt;Those two things fail on their own, and that is the whole point. A title spelled right across a character's face is still a bad poster. A clean layout with a wrong word in the headline is one you cannot publish. So I judged both, separately, on every test.&lt;/p&gt;

&lt;p&gt;I ran eight typography tests and grouped them here by the question each one answers, rather than by number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five things I checked
&lt;/h2&gt;

&lt;p&gt;Every test got a yes or no on five points. Accuracy, meaning the words are spelled right with no missing or deformed characters. Readability, meaning you can read the text at normal size without zooming in. Placement, meaning the text went where I asked and stayed off the parts of the image that matter. Hierarchy, meaning the title looks more important than the date, the badge, and the small print. And integration, meaning the type looks like part of the design rather than a caption dropped on top.&lt;/p&gt;

&lt;p&gt;Each test earns a pass, a partial pass, or a fail. I ran everything with Prompt Helper off, so the prompt you see is the prompt that ran.&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 1: Can it place short type exactly where you tell it?
&lt;/h2&gt;

&lt;p&gt;The simplest version of the job. Name a spot, put a couple of short strings there, keep the character out of the way.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime key visual for a rooftop concert, in English. A young male
vocalist with cropped black hair, a silver ear cuff, and an oversized red bomber
jacket stands on the right side of the frame, gripping a mic stand, blurred city
lights behind him. Leave the entire left third of the frame as open night sky.
Print the title "NIGHT SIGNAL" in large bold English letters across that open
left area, and directly beneath it the smaller line "LIVE AT DUSK". No other
text anywhere in the image. Magenta and cyan stage lighting, hazy night air,
modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both strings came out exactly right. The title is large and dominant, the subtitle is right underneath it at a smaller size, and the vocalist stayed on the right so the left side was free for the type. The ear cuff, red bomber jacket, and mic stand all matched, and nothing extra printed anywhere else.&lt;/p&gt;

&lt;p&gt;The one deviation is that the title wrapped onto two lines instead of running as one. It passed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mes1gxk0gh7trr8v5gc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mes1gxk0gh7trr8v5gc.png" alt="The rooftop key visual." width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then I made the surface harder. The word had to live on the overpass surface itself, at an angle, in a spray-paint style.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime poster in English. A teenage girl skater with a shaved-side
undercut, a yellow windbreaker, and scraped knees leans on a chain-link fence
under a highway overpass at dusk, standing in the lower right of the frame. The
upper left is a large blank concrete wall. Print the single word "RAIN" in large
bold English letters on that wall, spray-paint style, sharp and readable. No
other text anywhere in the image. Warm orange dusk light, long shadows, modern
anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;RAIN is spelled right and looks painted onto the surface rather than laid over it. The lettering picks up the angle of the overpass and carries the weight and shadow of paint. The undercut, windbreaker, and scraped knees are there, and no stray lettering turned up elsewhere. Another pass.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1p9wv0h9di7jfhw6d297.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1p9wv0h9di7jfhw6d297.png" alt="The overpass poster." width="800" height="1067"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The overpass poster.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Question 2: Can it handle many elements at once?
&lt;/h2&gt;

&lt;p&gt;This is where most tools start dropping things. Five pieces of text, each with its own corner and its own size.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A modern anime event poster in English, portrait layout. An older ramen chef
with a shaved head, a white towel tied around his forehead, and a navy apron
stands centered behind a steaming counter with his arms folded, neon signs
glowing behind him. Print five separate pieces of text, each exactly as written:
the main title "MIDNIGHT BOWL" large across the top, the tagline "One Broth. All
Night." directly beneath it in smaller letters, the date "OCT 24" in the lower
left corner, a three-line information block in the lower right reading "Gate 7 /
6PM till late / Free entry", and a small circular badge in the upper right corner
reading "10TH YEAR". Keep the chef's face and hands clear of all text. Warm amber
and deep red light, steam in the air, cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All five elements showed up, in the right corners, spelled correctly. MIDNIGHT BOWL runs across the top as the biggest thing on the poster, and the tagline reads exactly as written underneath it. The three-line block in the lower right is right line for line, and that is the smallest text in the image. The badge reads 10TH YEAR.&lt;/p&gt;

&lt;p&gt;Two small deviations. The date stacked vertically, so OCT is on one line and 24 on the next, and the badge sets the 10 and the TH apart rather than as one string. You still read both without effort, and no text touched the chef's face or hands. For a five-element layout in one generation, that is a pass, and it is the first sign that volume alone does not reduce accuracy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvl79nh62zuvnlnrl3qbi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvl79nh62zuvnlnrl3qbi.png" alt="The ramen event poster." width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 3: Does the language matter?
&lt;/h2&gt;

&lt;p&gt;Same scene, same composition, same open left half. Only the script changed. Every version below is a fresh generation from the exact prompt shown, not an edit of the English image.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ENGLISH
A cinematic anime key visual, in English. A young woman in a deep blue yukata
with a white fox mask pushed up on her forehead stands on the right, holding a
paper lantern, a summer festival street glowing behind her. The left half of the
frame is open night sky. Print the title "SUMMER SOUND" in large clean English
letters across that open left area. No other text anywhere in the image. Warm
lantern light against a cold blue night, modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;SUMMER SOUND is spelled right and easy to read. The type takes the left half with room around it, and the woman is on the right with the yukata, the fox mask on her forehead, and the lantern in hand. The title split onto two lines again, same as the concert poster, but it passed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjy2yi3hmcu4aorydjinx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjy2yi3hmcu4aorydjinx.png" alt="The English version" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For Japanese, I swapped the title to 夏の音 and kept everything else.&lt;/p&gt;

&lt;p&gt;The Japanese came out cleaner than the English. 夏の音 renders with correct stroke shapes on all three characters, no bleeding, and it runs across the open left half where I asked for it. The festival stalls in the background were painted as soft light shapes instead of surfaces with fake lettering, so the no-text instruction held better here than in the English run. It passed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbmoee7lxj1b94778pp6m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbmoee7lxj1b94778pp6m.png" alt="The Japanese version" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Korean is where it came apart.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;KOREAN
A cinematic anime key visual, with Korean text. A young woman in a deep blue
yukata with a white fox mask pushed up on her forehead stands on the right,
holding a paper lantern, a summer festival street glowing behind her. The left
half of the frame is open night sky. Print the Korean title "여름의 소리" in large
clean Hangul characters across that open left area. No other text anywhere in the
image. Warm lantern light against a cold blue night, modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This one failed, and it failed three ways. The title is wrong. I asked for 여름의 소리 and got 여뭄의 소리, with the syllable 름 swapped for 뭄. To a Korean reader that is not a typo you can look past, it is a different word.&lt;/p&gt;

&lt;p&gt;It also added a small line of garbled Hangul under the title that I never asked for, and the background stalls filled with legible Japanese kana in a prompt that said no other text anywhere. The composition itself is fine, the left half stayed open, the character and lighting match. The design worked and the typography did not.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2fg725b7oorjmx1mm06h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2fg725b7oorjmx1mm06h.png" alt="korean" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then I ran English and Japanese together, since key visuals often carry two languages.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BOTH SCRIPTS
A cinematic anime key visual with both English and Japanese text. A young woman
in a deep blue yukata with a white fox mask pushed up on her forehead stands on
the right, holding a paper lantern, a summer festival street glowing behind her.
The left half of the frame is open night sky. Print the English title "SUMMER
SOUND" in large bold English letters across that open left area, and directly
beneath it the smaller Japanese line "夏の音" in clean Japanese characters. Keep
the two scripts separate and do not mix characters between them. No other text
anywhere in the image. Warm lantern light against a cold blue night, modern anime
illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both strings are correct and neither script bled into the other. No Latin letters in the Japanese, no kana in the English. The English title stays dominant with the Japanese line smaller below it, both in the upper left away from the character, and the background stalls stayed painterly with no invented lettering. Mixing scripts is where these models usually come apart, and this one passed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjutuzn22pppv9owdn2tk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjutuzn22pppv9owdn2tk.png" alt="The bilingual version" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 4: Can type live in the scene, not just on it?
&lt;/h2&gt;

&lt;p&gt;Everything above prints type onto the frame. These two put it inside the scene, where the model cannot treat text as a sticker.&lt;/p&gt;

&lt;p&gt;First, on an object at a receding angle, with a reflection to match.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A modern anime street scene in English. A young male delivery rider with box
braids and a teal windbreaker rides past a corner shop at night, seen from a low
angle. The shop's long awning runs diagonally away from the camera into the
background, and the words "GOLDEN HOUR MART" are printed across that awning in
bold English letters that follow the angle of the awning as it recedes,
narrowing with the perspective. The wet road below reflects the lit sign. No
other text anywhere in the image. Rain-slick asphalt, warm shop light against
cold blue night, cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GOLDEN HOUR MART is spelled right and the letters narrow as the awning recedes, which means the model treated the words as printed on a surface in space rather than laid flat on the picture. The reflection landed too. The wet asphalt in the lower left carries a mirrored version of the sign, legible as GOLDEN, which is about what a real puddle would give you. The delivery box, the bike frame, and the shop windows are all free of invented lettering. It passed on both the text and the integration.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F619y5ecy0lmeqxzym1j4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F619y5ecy0lmeqxzym1j4.png" alt="The corner shop at night. Let’s go with the left one" width="800" height="528"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The corner shop at night. Let’s go with the left one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Then the harder one: letters running behind a character. Layering is the thing a model cannot fake if it treats text as a flat overlay.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime key visual in English, portrait layout. A tall female
basketball player with a black ponytail, a white number 11 jersey, and taped
fingers stands centered, holding a ball on her hip, gym lights flaring behind
her. Print the title "FINAL QUARTER" in huge bold English letters across the
middle of the frame so that her head and shoulders pass in front of the letters
and cover part of them, with the letters clearly continuing behind her body on
both sides. Print the date "MAR 08" small in the lower left corner, and a small
badge in the upper right corner reading "SEMI FINAL". Her face must stay
completely clear of all text. Deep shadows, hard white spotlight, modern anime
illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The layering worked. FINAL QUARTER runs across the middle and her head, ponytail, and torso pass in front of it. The top line shows FI and then AL where her body interrupts it, the lower line QUA and then ER, and both resume on the far side at the right size and position. That is a depth relationship, not an overlay. No text touched her face, MAR 08 is in the lower left, SEMI FINAL is in the upper right, all three strings are correct, and the title stays dominant. A pass on every criterion.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0sowghuwpavc5e63pdkj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0sowghuwpavc5e63pdkj.png" alt="The basketball key visual" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 5: Can it carry several kinds of text at once?
&lt;/h2&gt;

&lt;p&gt;A comic page carries several types of text, each with a different job: narration in a box, dialogue in balloons, a sound effect drawn as artwork, and a sign that exists inside the world.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A four-panel black-and-white manga page, dialogue in English, sound effects in
Japanese, read left to right. Two characters throughout: a young male courier
with a buzz cut and a canvas satchel, and an older woman shopkeeper with wire
glasses and a knitted shawl. Panel 1, wide shot of a narrow shop front at dusk,
a hanging wooden sign above the door reading "KOMORI BOOKS" in English letters,
and a rectangular caption box in the upper left reading "Closing time, third day
of rain." Panel 2, the courier pushes the door open, a Japanese sound effect
"カラン" drawn near the door in stylised katakana, and a speech balloon from the
shopkeeper reading "You're late again." Panel 3, close on the courier lowering
the satchel, his speech balloon reading "Last one today." Panel 4, both of them
at the counter, the shopkeeper's balloon reading "Then sit down." Keep the
caption box square-cornered, the speech balloons rounded, and the sound effect
drawn as a graphic rather than in a balloon. No other text anywhere on the page.
Clean inked manga linework, heavy shadows.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The page handled all four roles. The caption box is square-cornered, the balloons are rounded, the sound effect is drawn as artwork rather than typeset in a bubble, and the shop sign looks like painted wood in the scene. The balloon tails point at the right speaker in every panel, the reading order runs left to right, and both characters stay recognizable across all four panels.&lt;/p&gt;

&lt;p&gt;Four of the five text strings are exact. KOMORI BOOKS, the caption line, "You're late again," "Then sit down," and the katakana カラン all rendered correctly. Panel 3 printed "Last one torday."&lt;/p&gt;

&lt;p&gt;The part that matters more is that the same error repeated in all four images of the batch. It is not a bad draw you can generate past, so re-rolling will not rescue it. That makes this a partial pass.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2zoe1yzc2z9w8g83aqt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2zoe1yzc2z9w8g83aqt.png" alt="The four-panel comic page" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Question 6: Can it leave surfaces blank on purpose?
&lt;/h2&gt;

&lt;p&gt;The opposite problem, and a harder one. Not printing text is tough because a convenience store is a room made of packaging.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime scene in English. A high school boy with round glasses and a
grey hoodie stands reading a single handwritten notice pinned to a corkboard in
a bright convenience store aisle. The notice reads exactly "CLUB TRYOUTS FRIDAY"
in clean English handwriting, and it is the only text in the entire image. Every
product package, shelf label, price tag, window, and sign in the store must be
completely blank, with no letters, numbers, or symbols on any surface. Cold
fluorescent light, packed shelves, modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The notice reads CLUB TRYOUTS FRIDAY in correct handwriting, and nothing else in the store carries a single character. The product boxes are solid color blocks with no fake typography. The price tag strips along the shelf edges are drawn as empty rectangles, and the cooler headers in the background are color bands. Those are the exact surfaces that normally fill with pseudo-lettering, and every one of them stayed empty. A full pass.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo9kbreng0g8ec10arpom.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo9kbreng0g8ec10arpom.png" alt="The convenience store aisle" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the text started to fail
&lt;/h2&gt;

&lt;p&gt;Across the eight tests, the pattern in AI text in image accuracy was not the one I expected.&lt;/p&gt;

&lt;p&gt;More text did not mean worse text. The five-element poster carried a title, a tagline, a date, a three-line block, and a badge, and every string was right. Volume on its own did not reduce accuracy.&lt;/p&gt;

&lt;p&gt;Script mattered more than volume. English held across six tests. Japanese held on its own, inside a comic page, and next to English in the same frame. Korean failed on its only run, at the character level, which is what a reader notices first.&lt;/p&gt;

&lt;p&gt;The one English error was in dialogue, not display type. Every headline, tagline, date, and badge came out right. The word that went wrong was a line of speech inside a busy four-panel page, so small text competing with panel borders and linework looks like the harder job.&lt;/p&gt;

&lt;p&gt;Errors repeat. The torday mistake held across all four images in the batch. When you generate text in AI images and a word comes out wrong, generating again may hand you the same mistake, so rewriting the line beats re-rolling it.&lt;/p&gt;

&lt;p&gt;Line wrapping is the standing deviation. Three runs split a title onto two lines when I asked for one. It never hurt readability, but if you need the title on one line, describe the shape of the text area, not just name it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Readable text and good design are two different results
&lt;/h2&gt;

&lt;p&gt;The Korean run shows the split plainly. The composition was right, the character was right, the left half stayed open for the title, and the lighting worked. Every design decision landed and the words were wrong. As a picture it works, as a poster you cannot use it.&lt;/p&gt;

&lt;p&gt;The comic page is a milder version. The layout, the panel flow, the four typographic roles, and the character continuity are all better than I expected from an AI typography generator, and one wrong word in panel 3 still means you would redo it before publishing.&lt;/p&gt;

&lt;p&gt;I never got the reverse case. There was no run where the spelling was right and the layout was a mess. The AI graphic design side, meaning placement, hierarchy, layering, and integration, held in every test including the one that failed on text. So Tsubaki.3 understands what a poster is. Read the words before you post it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The results side by side
&lt;/h2&gt;

&lt;p&gt;Here is every test in one place.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;th&gt;Readability&lt;/th&gt;
&lt;th&gt;Layout&lt;/th&gt;
&lt;th&gt;Integration&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Title and subtitle&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Wrapped to two lines&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single word on a surface&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;As requested&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Five elements&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High, even small&lt;/td&gt;
&lt;td&gt;Two small deviations&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;English title&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Wrapped to two lines&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Japanese title&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;As requested&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Korean title&lt;/td&gt;
&lt;td&gt;Wrong syllable, extra text&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;As requested&lt;/td&gt;
&lt;td&gt;Weak&lt;/td&gt;
&lt;td&gt;Fail&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;English and Japanese&lt;/td&gt;
&lt;td&gt;Exact, no mixing&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;As requested&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text on an angled sign&lt;/td&gt;
&lt;td&gt;Exact, with reflection&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;As requested&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Title behind character&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Correct layering&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Comic page&lt;/td&gt;
&lt;td&gt;One wrong word, repeated&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Four roles handled&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Partial&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One sign, rest blank&lt;/td&gt;
&lt;td&gt;Exact&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;All surfaces empty&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Where this is useful right now
&lt;/h2&gt;

&lt;p&gt;Short display type is what it does best. Titles, taglines, dates, badges, and corner labels, the kind of thing you would put on a key visual or an event poster in a few words. Every string of that kind came out right here, which makes it usable as an AI poster generator rather than an illustration tool you then take into a design app.&lt;/p&gt;

&lt;p&gt;The harder design asks worked too, so you can plan around type instead of adding it later. If you name your corners, reserve the areas you want left open, and say what should stay blank, one generation gets you most of the way. The &lt;a href="https://blog.pixai.art/en/how-to-write-pixai-prompts-formula/" rel="noopener noreferrer"&gt;PixAI prompt formula guide&lt;/a&gt; covers how to structure that, and the &lt;a href="https://docs.pixai.art/docs/prompts/prompt-advanced" rel="noopener noreferrer"&gt;advanced prompting page&lt;/a&gt; covers the composition side.&lt;/p&gt;

&lt;p&gt;I would still set the type by hand for longer copy, for anything where the exact wording is contractual like prices and names, and for languages outside English and Japanese until you have tested them yourself.&lt;/p&gt;

&lt;p&gt;For a single wrong word in an image that otherwise works, correcting it afterward beats regenerating, since the batch gave me the same mistake four times. The &lt;a href="https://blog.pixai.art/en/pixai-edit-pro-ai-image-editor/" rel="noopener noreferrer"&gt;Edit Pro guide&lt;/a&gt; covers that route. If you are building a series of posters around one character, the &lt;a href="https://blog.pixai.art/en/pixai-reference-pro-guide-multi-image-editing-with-natural-language/" rel="noopener noreferrer"&gt;Reference Pro guide&lt;/a&gt; is where to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next move
&lt;/h2&gt;

&lt;p&gt;For short English and Japanese type, Tsubaki.3 works as a readable text AI image generator. Seven of eight tests passed, every English headline, tagline, date, and badge came out correct, and the design side held in every test including the one that failed on its words.&lt;/p&gt;

&lt;p&gt;The failures are specific. Korean printed a wrong syllable and added a line I never asked for, and one line of comic dialogue came out as torday across the whole batch. An AI text generator in images will not tell you when it got something wrong, so read every word before you publish.&lt;/p&gt;

&lt;p&gt;Pick a poster idea you have been meaning to make, write the title and the position of every text element into the prompt, and &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;run it on PixAI&lt;/a&gt;. Two generations will tell you where your own text lands.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>One Reference Image, Nine Sheets: Where an AI Character Sheet Generator Holds and Where It Guesses</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Mon, 31 Aug 2026 06:42:39 +0000</pubDate>
      <link>https://dev.to/naveedoss/one-reference-image-nine-sheets-where-an-ai-character-sheet-generator-holds-and-where-it-guesses-1dh1</link>
      <guid>https://dev.to/naveedoss/one-reference-image-nine-sheets-where-an-ai-character-sheet-generator-holds-and-where-it-guesses-1dh1</guid>
      <description>&lt;p&gt;You have one good picture of your character. One picture does not tell you much, though.&lt;/p&gt;

&lt;p&gt;You cannot see the back. You cannot see the side. And you do not know what the face does when the character laughs or shouts. A character sheet is what shows you all of that.&lt;/p&gt;

&lt;p&gt;So here is the question I wanted answered. Can an AI character sheet generator take that one picture and build the rest, the turnaround, the expressions, the poses, and keep it looking like the same person?&lt;/p&gt;

&lt;p&gt;I ran nine tests on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, split into two experiments. The first five use one swordsman I designed, run through the standard sheets. The last four use a single image I grabbed from PixAI's public gallery, to see what one found picture can turn into.&lt;/p&gt;

&lt;p&gt;The real question underneath all nine is the same: can it tell apart what the reference shows from what it has to invent, and keep the invented parts steady.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scorecard first
&lt;/h2&gt;

&lt;p&gt;Before the walk-through, here is how all nine landed. Two of them needed a second run, and I have listed both.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;What it demanded&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Base reference&lt;/td&gt;
&lt;td&gt;One clean front-facing design&lt;/td&gt;
&lt;td&gt;Pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Turnaround&lt;/td&gt;
&lt;td&gt;Front, 3/4, side, back from a front-only image&lt;/td&gt;
&lt;td&gt;6, then 8.5 on re-run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expression sheet&lt;/td&gt;
&lt;td&gt;Nine emotions, identity held&lt;/td&gt;
&lt;td&gt;6, then 7.5 on re-run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pose set&lt;/td&gt;
&lt;td&gt;Five full-body poses&lt;/td&gt;
&lt;td&gt;7.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Expressions in context&lt;/td&gt;
&lt;td&gt;Nine emotions in nine real scenes&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Action pose sheet&lt;/td&gt;
&lt;td&gt;Six combat poses with motion&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storyboard&lt;/td&gt;
&lt;td&gt;Six manga panels from one frame&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Three roles&lt;/td&gt;
&lt;td&gt;One face across three worlds&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Color script&lt;/td&gt;
&lt;td&gt;Six lighting states of one room&lt;/td&gt;
&lt;td&gt;7.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key-visual poster&lt;/td&gt;
&lt;td&gt;A finished, title-ready poster&lt;/td&gt;
&lt;td&gt;8.5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern in that column is the whole story: solid outputs, held back by panel counts and small identity marks. More on that after the tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  What counts as a real sheet
&lt;/h2&gt;

&lt;p&gt;A character sheet is not just several pictures of the same person on one page. Each part should tell you something useful about the design.&lt;/p&gt;

&lt;p&gt;A useful anime character sheet usually covers a front, side, and back turnaround, an expression sheet, some poses, and sometimes outfit variations. You do not need every one every time.&lt;/p&gt;

&lt;p&gt;The point is that each panel should give you real information: the shape of the coat from behind, how the face moves, how the body holds a pose, not just another nice illustration. That is what turns a set of images into a real AI character reference sheet.&lt;/p&gt;

&lt;p&gt;So the bar I judged against was not "does it look good." It was "would this help me draw or reuse the character."&lt;/p&gt;

&lt;h2&gt;
  
  
  Experiment one: standard sheets from a character I designed
&lt;/h2&gt;

&lt;p&gt;I built one original swordsman as the base and used him for the first five tests.&lt;/p&gt;

&lt;p&gt;He is a weathered older man with tied grey-streaked hair, a black eyepatch over his right eye, a long cross-style scar on that cheek, a fur-lined coat, a sword sheathed on his back, a carved wolf-head pauldron on his left shoulder, and prayer beads on his wrist. Those details are the anchors I tracked, specific enough that any drift shows up fast.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftynd6t0ehepgpumv0jli.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftynd6t0ehepgpumv0jli.png" alt=" " width="800" height="1067"&gt;&lt;/a&gt;&lt;br&gt;
One thing matters before we start. The reference is front-facing, so the back of the coat, the far side, and the full sword are things the single image never shows. The model has to invent them. That is the real test, telling apart what the model kept from what it made up.&lt;/p&gt;

&lt;p&gt;The reference itself came out strong. Every anchor landed in the right place, and the outfit worked as a coherent design. A clear pass.&lt;/p&gt;
&lt;h3&gt;
  
  
  Turnaround
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, create a clean turnaround model sheet on a plain
background, four views at the same height: front, three-quarter, side profile,
and back. Keep his tied grey-streaked hair, eyepatch and scar, fur-lined coat,
wolf-head pauldron, back-sheathed sword, and wrist beads identical across every
view. Show the back of the coat and the full sheathed sword clearly in the rear
view. Neutral A-pose, consistent model-sheet line work.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The first run showed real problems. The front view stayed close to the reference, but the three-quarter view drifted almost fully into a side profile, so I basically lost that angle.&lt;/p&gt;

&lt;p&gt;Worse, the model gave me two back views instead of a proper set of four, and the two backs did not agree. The wolf pauldron was missing in one and present in the other. Since that pauldron is a main identity piece, that is a real miss. First run, a 6.&lt;/p&gt;

&lt;p&gt;So I ran it again, and the second run was much better. This time I got all four angles, a proper three-quarter view, and the wolf pauldron stayed consistent across every view, including the back. The hair, eyepatch, coat, sword, and beads all held together. The sword shifted shape slightly between views, but that is minor next to the first attempt. Second run, an 8.5.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnhe27wr3f8zw35ugt2gj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnhe27wr3f8zw35ugt2gj.png" alt="Left: the first run. Right: the second run." width="800" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the first run. Right: the second run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The honest takeaway on the AI character turnaround: it can build a believable back from a front-only image, but the consistency is not guaranteed on the first try. A re-run got me there.&lt;/p&gt;
&lt;h3&gt;
  
  
  Expression sheet
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, an expression sheet of nine head-and-shoulders
portraits, changing only the expression: hard glare, faint smile, open laugh,
grief, cold anger, weariness, surprise, suspicion, and calm. Keep his tied
grey-streaked hair, eyepatch over the right eye, cheek scar, and one visible
eye identical in every panel. Plain background, consistent line work.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The expressions were great, but the identity slipped. All nine emotions came through clearly and the face structure held, so as an AI expression sheet the range is strong.&lt;/p&gt;

&lt;p&gt;The problem was the anchors. In the first run, two panels lost the eyepatch entirely and showed both eyes, and the scar came and went in those same spots. For a character whose whole identity leans on that eyepatch, that is a real failure. First run, a 6.&lt;/p&gt;

&lt;p&gt;The second run improved. The eyepatch and scar held in seven of the nine panels, so I will call it a 7.5. Better, but one identity-losing panel out of nine still is not a clean pass. This was a heavy ask, nine faces in one go, and that is likely why it stumbled.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fio3smelw2kw6r894s0im.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fio3smelw2kw6r894s0im.png" alt="First expression run left and the second run right" width="799" height="521"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;First expression run left and the second run right&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Pose set
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, a pose reference sheet of five full-body poses on a
plain background: standing at rest, drawing the sword from his back, a low ready
stance, kneeling with the sword planted, and walking with the coat trailing.
Keep his proportions, tied hair, eyepatch, fur-lined coat, wolf-head pauldron,
and beads consistent across all five. Clean model-sheet style, consistent scale.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Four of the five poses were solid, and the character held well across them. Standing, drawing the sword, the low ready stance, and the kneeling pose all came through clearly, and the eyepatch, hair, coat, pauldron, and proportions stayed steady. A good result for an AI pose sheet.&lt;/p&gt;

&lt;p&gt;The fifth pose missed. It looks like another standing pose rather than walking, and the coat shows none of the trailing movement I asked for. The other weak spot is the face, which goes soft and muddy at this full-body size compared to the closer shots. A 7.5.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fah9ktfeij0mojo7jsjhq.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fah9ktfeij0mojo7jsjhq.gif" alt="The five-pose set." width="760" height="428"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The five-pose set.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Expressions in context
&lt;/h3&gt;

&lt;p&gt;For this one I pushed past the plain grid and asked for nine emotions inside nine real scenes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, a nine-panel grid of cinematic close-up moments,
each showing a different emotion in a different setting, same man throughout:
laughing by a campfire, glaring in the rain, grieving at a grave, calm in a
snowfall, angry in a tavern brawl, weary on a mountain road, surprised in
torchlight, suspicious in a market crowd, at peace under cherry blossoms. Keep
his tied grey-streaked hair, eyepatch, and cheek scar identical in every panel.
Cinematic modern anime illustration, varied lighting.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was harder, and the character held better than in the plain expression sheet. The eyepatch, hair, beard, and face stayed consistent across all the different lighting and settings, and this time the patch did not vanish. He also fit each scene naturally instead of looking pasted onto a background. The campfire, rain, graveyard, tavern, and cherry blossoms all came out well.&lt;/p&gt;

&lt;p&gt;The one real failure is the count. I asked for nine panels and got eight. The "surprised in torchlight" moment dropped out entirely. An 8, strong consistency under changing light, held back by the missing panel.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4oi9dnbqiiteq8unlc6y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4oi9dnbqiiteq8unlc6y.png" alt=" " width="800" height="1067"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The in-context expression sheet&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Action pose sheet
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, a dynamic action pose sheet of six full-body combat
poses on a soft neutral background: mid-swing slash, blocking overhead, spinning
parry, lunging thrust, sheathing the sword, and landing from a jump. Coat and
hair moving with each motion. Keep his proportions, eyepatch, scar, fur-lined
coat, wolf-head pauldron, and back sword consistent across all six. Clean,
energetic model-sheet style.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The motion here is a clear step up from the static poses. The model understood this needed real combat energy, and the coat, hair, and sword all move convincingly while the character stays recognizable. The eyepatch, pauldron, coat, and beads held across the set.&lt;/p&gt;

&lt;p&gt;The miss is precision. The six specific actions are not all clear. Several poses look like slashing or parrying variations, and the sheathing pose is missing or hard to spot. The face is still a little soft in the smaller renders, though cleaner than the pose set. An 8.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fskerzgj9n9di2x44wsug.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fskerzgj9n9di2x44wsug.png" alt="Base image (left). The action pose sheet (Right)." width="800" height="526"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Base image (left). The action pose sheet (Right).&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Experiment two: what a stranger's single image can become
&lt;/h2&gt;

&lt;p&gt;The standard sheets answer the brief. Then I wanted to see what else one image can become, so I switched references entirely.&lt;/p&gt;

&lt;p&gt;I grabbed a single moody close-up from PixAI's public gallery, a blue-haired student resting his head on a classroom desk. It is a found community image, not my own design, and I used it purely as a test input.&lt;/p&gt;

&lt;p&gt;It is a tight, angled shot, so almost everything below his shoulders is unknown. That makes these four tests even harder, and more like real life, since most people start from one good picture, not a clean model sheet.&lt;/p&gt;

&lt;p&gt;The anchors I tracked: dark blue layered hair, grey eyes, a cross-shaped cheek scar, and the row of ear piercings.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fudiqccw7jnxz1ejdd0dh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fudiqccw7jnxz1ejdd0dh.png" alt="The found reference image" width="800" height="1032"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The found reference image&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A storyboard from one frame
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character as the starting frame, create a six-panel
black-and-white manga storyboard of the scene this moment belongs to. Panel 1,
a wide shot of the quiet classroom. Panel 2, him walking in alone. Panel 3, this
exact moment, him resting his head on the desk, tired. Panel 4, another student
speaks to him from the doorway. Panel 5, he lifts his head and looks over. Panel
6, he stands to leave, bag over his shoulder. Keep his dark blue layered hair,
grey eyes, cross cheek scar, and ear piercings consistent in every panel.
Cinematic manga composition.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The visuals are good, but it missed two story beats. I asked for six panels and got five. The model merged the wide classroom shot and him walking in into one panel, and because of that, the student speaking from the doorway never showed up.&lt;/p&gt;

&lt;p&gt;The rest of the sequence works. The desk moment matches the reference, the head-lift is clear, and the final shot comes across as him leaving. His hair, scar, and piercings stay recognizable, and the black-and-white manga treatment fits well. A 7. The character holds up better than the storyboard instructions do.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqv13sabtfdcym5g79r72.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqv13sabtfdcym5g79r72.png" alt="The five-panel storyboard" width="800" height="1067"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The five-panel storyboard&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The same face, three lives
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, keep his face, dark blue layered hair, grey eyes,
cross cheek scar, and ear piercings exactly the same, but reimagine him in three
completely different roles side by side: a cyberpunk street mercenary in a neon
alley, a medieval knight in worn armor, and a modern rockstar on stage. Same
face and identity in all three, only the world, outfit, and role change.
Detailed modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where identity preservation really showed. The model turned him into three completely different characters, and all three still look like the same guy. The blue hair, grey eyes, face structure, and ear piercings carried across every version, and the three roles feel distinct instead of the same outfit recolored.&lt;/p&gt;

&lt;p&gt;The cross scar is the weak point. It is there in all three, but its exact shape and placement drift a little, and the knight's face is slightly different in build. An 8. Keeping one face across three worlds is a strong result.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk4tcl62x2sbrpyi2upzy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk4tcl62x2sbrpyi2upzy.png" alt="Test 7 - Base image (left). The three-roles image (right)." width="800" height="516"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 7 - Base image (left). The three-roles image (right).&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A color script
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using the mood and color palette of this exact image, create a horizontal film
color-script strip of six small thumbnail frames of the same classroom scene at
different times: early dawn, bright noon, golden afternoon, blue dusk, night with
lights off, and stormy grey. Each thumbnail keeps the boy resting at the desk,
small in frame, but the light and color of the room change completely across the
strip. Cinematic color-script layout.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lighting nailed it, the layout did not. The classroom stays recognizable while the light shifts hard across the frames, and the blue dawn, warm gold, purple dusk, deep night, and stormy grey all come through clearly. The mood control is excellent.&lt;/p&gt;

&lt;p&gt;Two things missed. I asked for a horizontal filmstrip and got a 2x3 grid, and the boy is far too large. He was supposed to stay small so the room and light were the focus. A 7.5. Great color, wrong format.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw9hueybwc9v58im0rcj4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw9hueybwc9v58im0rcj4.png" alt="The color-script grid. I’ll go with the right one for this review." width="800" height="528"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The color-script grid. I’ll go with the right one for this review.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  A key-visual poster
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, design a cinematic anime key-visual poster with him
as the centerpiece. Keep his face, dark blue layered hair, grey eyes, cross cheek
scar, and ear piercings exact. Build a moody, atmospheric composition around him
with dramatic lighting, a color palette matching the original, space at the top
for a title, and a lonely, introspective tone.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This one came out the strongest for mood and finish. The model understood the lonely, introspective direction and built a dark classroom, cool blue palette, and window light around him. His face, hair, eyes, and cross scar are all present and closer to the reference than in some earlier tests, the scar especially.&lt;/p&gt;

&lt;p&gt;The miss is the poster part. I asked for empty space at the top for a title, but his hair runs high into the frame and leaves almost none. It comes across more as a cinematic illustration than a designed poster. An 8.5.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8uaokm1n4hpkhbqgborv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8uaokm1n4hpkhbqgborv.png" alt="The key-visual poster" width="800" height="525"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The key-visual poster&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern across both experiments
&lt;/h2&gt;

&lt;p&gt;Put the two experiments together and the same three-part pattern shows up every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Big features hold, tiny ones drift.&lt;/strong&gt; His hair, face shape, eyes, and general build stayed steady almost everywhere. What slipped were the small identity marks, the swordsman's eyepatch and the student's cross scar. Those dropped out or changed shape most often, especially in the busy nine-panel sheets and under strong expressions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scale hurts the face.&lt;/strong&gt; In the full-body poses, the face goes soft and muddy. It is still recognizable, but it loses the sharpness you get in the close-up shots.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Exact counts and layouts are shaky.&lt;/strong&gt; More than once, I asked for a set number of panels or a specific layout and the model quietly dropped one or reorganized it: the missing ninth expression, the five-panel storyboard, the grid instead of a filmstrip.&lt;/p&gt;

&lt;p&gt;And the invented parts, the back of the coat, the full body from a close-up, are believable but not guaranteed to stay consistent, which is exactly why the turnaround needed a second run. The first turnaround gave me two different backs in one sheet, which tells you the model is guessing, not remembering.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verdict: a strong starting point, not a finished reference
&lt;/h2&gt;

&lt;p&gt;For developing an OC, planning a comic, working out a mood, or sharing a design with someone, these sheets give you plenty to work from. The in-context expressions, the action poses, the three-roles test, and the poster were all useful outputs. As a quick OC character sheet or a reference to draw over, it does the job.&lt;/p&gt;

&lt;p&gt;What it is not yet is a clean production model sheet you can hand to an animator without checking. The turnaround needed a re-run, panels go missing, and the small anchors are not reliable enough to trust blindly.&lt;/p&gt;

&lt;p&gt;A re-run usually helps, and naming your key features in every prompt helps more. But the more your character leans on small, specific marks, the more you will want to check each panel. If you want to go deeper on the reference side, PixAI's &lt;a href="https://blog.pixai.art/en/pixai-reference-pro-guide-multi-image-editing-with-natural-language/" rel="noopener noreferrer"&gt;Reference Pro guide&lt;/a&gt; covers that workflow.&lt;/p&gt;

&lt;p&gt;So, can Tsubaki.3 build a character sheet from one reference image? Mostly, and better than I expected, as long as you go in knowing where it slips. It is strongest at expressions in context, poses, and the creative expansions. As a character turnaround generator it is not fully reliable yet, and it is weakest at precise panel counts and at holding tiny identity marks across a big sheet.&lt;/p&gt;

&lt;p&gt;If you have one good picture of your OC in a folder, this is a fast way to turn it into something you can build on. &lt;a href="https://eap.pixai.art/go/naveed2" rel="noopener noreferrer"&gt;Try it on PixAI&lt;/a&gt; with your own character, start with the expression sheet or the poses, and you will know within a couple of generations how well it holds your design.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Can Tsubaki.3 Put the Camera Where You Want It? 9 Perspective and Composition Tests</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:15:02 +0000</pubDate>
      <link>https://dev.to/naveedoss/can-tsubaki3-put-the-camera-where-you-want-it-9-perspective-and-composition-tests-5d3</link>
      <guid>https://dev.to/naveedoss/can-tsubaki3-put-the-camera-where-you-want-it-9-perspective-and-composition-tests-5d3</guid>
      <description>&lt;p&gt;Most AI images come out framed the same way. The character is centered, the camera is at eye level, and there is not much depth. It looks fine, but it does not tell you whether the model can really put the camera where you want it.&lt;/p&gt;

&lt;p&gt;That is what I wanted to find out. So I ran nine tests on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, each one asking for a specific camera angle or a particular way of arranging the scene, then checked whether it did what I asked.&lt;/p&gt;

&lt;p&gt;That is the real test of an AI perspective generator. Not whether the picture looks good, but whether the angle I described is the one it drew. I have grouped the nine by the kind of control each one demands, rather than in the order I ran them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What composition control means
&lt;/h2&gt;

&lt;p&gt;A good-looking image does not prove the model can control composition. Plenty of models hand you a nice image while quietly ignoring the camera you asked for.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuwtr0z4n1pfc1w5co58.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuwtr0z4n1pfc1w5co58.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So I judged each test on one simple thing: did the exact setup I described show up. The camera height, the order of the depth layers, the distortion, where each subject lands in the frame.&lt;/p&gt;

&lt;p&gt;If I can look at the result and see what I asked for, it passes. If it gives me a nice image that ignored the instruction, it does not. That is the difference between a real AI composition generator and one that just makes pretty variations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fundamental camera moves
&lt;/h2&gt;

&lt;p&gt;The first group tests the basics of a controllable camera: height and linear perspective.&lt;/p&gt;

&lt;h3&gt;
  
  
  Low angle vs high angle
&lt;/h3&gt;

&lt;p&gt;The cleanest camera test is to shoot one subject twice and change only the camera height. I used a lone knight on a battlefield.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BASE
An extreme low-angle shot of a lone knight in worn armor standing on a
battlefield, the camera almost on the ground looking steeply up at her, so she
looms tall against a stormy sky and her boots and legs dominate the foreground
while her head is small and far above. Dramatic clouds behind her, tattered
banner, cinematic modern anime illustration, full body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
A steep bird's-eye view looking straight down at the same lone knight in worn
armor standing on a battlefield, seen from high above so the ground fills the
frame, her body foreshortened with her head and shoulders largest and her feet
small beneath her, her shadow stretching across the mud. Cinematic modern anime
illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both angles landed, and the difference is obvious. In the low shot, the camera is almost on the ground and the knight towers over you, her boots and legs filling the foreground while her head looks small and far above.&lt;/p&gt;

&lt;p&gt;In the high shot, the camera looks steeply down, her head and shoulders come out larger than her feet, and her shadow stretches across the mud to sell the height.&lt;/p&gt;

&lt;p&gt;The key point is that the perspective changed, not just her pose. The model understood camera height, which is the whole test. The only small miss is that the high shot is a steep overhead rather than a perfectly straight-down bird's-eye.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqnlgyoomhnerjdsczk57.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqnlgyoomhnerjdsczk57.png" alt="Test 1-Left the low-angle shot. Right the high-angle shot" width="800" height="522"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the low-angle shot. Right the high-angle shot.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Extreme depth down a corridor
&lt;/h3&gt;

&lt;p&gt;This one tests linear perspective: a long corridor where every line rushes toward a single vanishing point.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A view straight down a long neon-lit arcade corridor at night, one-point
perspective with the walls, floor lights, and ceiling signs all rushing toward
a single vanishing point in the far center. A girl stands halfway down the
corridor, small in the middle distance, dwarfed by the tunnel of light
stretching far behind and ahead of her. Strong linear perspective, deep
vanishing point, cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The one-point perspective is excellent. The floor, walls, ceiling lights, and signs all converge to a single vanishing point in the center, and the girl is correctly small in the middle distance, dwarfed by the tunnel of light. It looks like a long corridor, not a flat background with repeated shapes.&lt;/p&gt;

&lt;p&gt;The one weakness is her face, small and a little muddy at that distance. The prompt made her tiny on purpose, so it is a minor cost. As AI perspective drawing, the depth is the strongest part.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhb8iok2jl34380eqh43g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhb8iok2jl34380eqh43g.png" alt=" " width="800" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The corridor depth shot. I choose the left one for this review.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Depth and distortion
&lt;/h2&gt;

&lt;p&gt;The next two push on how the model handles a real lens: separated depth layers and true barrel distortion.&lt;/p&gt;
&lt;h3&gt;
  
  
  Three depth layers
&lt;/h3&gt;

&lt;p&gt;I tested depth with one thing close to the lens, the subject in the middle, one thing far behind, and real size difference between them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime street scene with strong depth. In the extreme foreground,
close to the lens and slightly out of focus, a paper lantern hangs large on the
left. In the middle distance, a girl in a red kimono walks toward the camera,
sharp and clearly the main subject. Far in the background, a tall pagoda rises
small against the evening sky. Clear size difference between the near lantern,
the midground girl, and the distant pagoda, cinematic depth, detailed modern
anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The three layers came out cleanly separated. The paper lantern is huge, soft, and close on the left. The girl in the red kimono is sharp and mid-sized as the main subject. The pagoda is tiny and far behind her.&lt;/p&gt;

&lt;p&gt;The scale falloff is what sells it. Each layer is a clearly different distance, in the right order, and the receding street adds even more depth on top. This is strong AI art composition with a real foreground, midground, and background rather than flat stacking.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytcsxhtcjyc3d8zdg519.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fytcsxhtcjyc3d8zdg519.png" alt="The three-layer depth shot" width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Fisheye
&lt;/h3&gt;

&lt;p&gt;Then a harder one: a real fisheye, where straight lines have to bend and the room has to bulge, not just widen.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A fisheye lens view inside a cramped record shop, strong barrel distortion
bending the straight shelves and ceiling into curves around the edges of the
frame, a girl in the center reaching one hand right toward the camera so her
hand looks huge and close while her body curves away small behind it. Rounded,
bulging perspective, wide distorted field of view, detailed modern anime
illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc256j10nnz4y32bbylk2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc256j10nnz4y32bbylk2.png" alt=" " width="799" height="269"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the one I expected to fail, and it passed. The shelves visibly curve outward, the ceiling bends around the frame, and the whole room bulges the way a real fisheye looks. Her reaching hand balloons toward the lens while her body curves away small behind it.&lt;/p&gt;

&lt;p&gt;This is real barrel distortion, not a plain wide shot, which is the trap most models fall into. For fisheye perspective AI, this is about as convincing as it gets.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4978umltn91y3i7uyb3d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4978umltn91y3i7uyb3d.png" alt="The fisheye record-shop shot." width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Arranging multiple subjects
&lt;/h2&gt;

&lt;p&gt;Camera height is one thing. Placing several subjects at named positions and holding an angle over all of them is harder.&lt;/p&gt;
&lt;h3&gt;
  
  
  Complex multi-subject placement
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A low-angle cinematic anime shot looking up at a rooftop standoff. In the
immediate foreground on the right, close to the camera, a boy crouches with his
back to us, only his shoulder and the sword in his hand visible large in frame.
In the midground, center, a girl in a black coat stands facing him, sharp and
lit by neon. Behind her in the background, far and small, a second figure
watches from a doorway. The camera is low, looking up past the crouching boy
toward the standing girl. Detailed modern anime illustration, cinematic
composition.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Every subject landed in its assigned spot. The boy is huge in the right foreground with his sword, his back to us. The girl stands centered in the midground, sharp and lit by neon. The second figure watches small from the doorway in back. And the low angle holds across all three, so you look up past the boy toward the girl.&lt;/p&gt;

&lt;p&gt;For dynamic camera angle AI work, holding a low angle across a busy multi-subject scene is the hard part, and it held. The model even gave the girl a sword, which fits the standoff. The only slip is that the foreground boy shows more of his body than the "just his shoulder" I asked for, though that makes the scale difference even clearer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8adtmg1d9b7rdgvy6tfc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8adtmg1d9b7rdgvy6tfc.png" alt=" " width="800" height="269"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I also ran a treatment change on this shot, keeping the composition and swapping neon night for overcast morning.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Keep this exact composition, camera angle, and the positions of all three
characters identical, but change the time from neon night to bright overcast
morning. Same low angle looking up, same crouching boy large in the right
foreground, same girl centered in the midground, same distant watcher in the
background. Only the lighting, color, and mood change to flat cool daylight.
Detailed modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The structure survived the lighting change. The boy stayed large in the right foreground, the girl stayed centered, the watcher stayed in the doorway, and the low angle held. Only the light, color, and mood shifted to flat cool daylight.&lt;/p&gt;

&lt;p&gt;The model did not rearrange the scene just because the treatment changed, which is exactly what you want. It is not pixel-identical, but the camera and every placement stayed put.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjpzrm14bkc49liqwar2d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjpzrm14bkc49liqwar2d.png" alt="Left the neon-night original. Right the overcast-morning edit." width="800" height="527"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the neon-night original. Right the overcast-morning edit.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Keeping one space consistent as the camera moves
&lt;/h2&gt;

&lt;p&gt;For the next group I built one scene and moved the camera around it through edits. That is harder than generating fresh angles, because the model has to keep the same space consistent as the viewpoint changes.&lt;/p&gt;
&lt;h3&gt;
  
  
  The base scene
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A wide cinematic establishing shot of a grand library atrium at golden hour. In
the foreground, a long wooden reading table runs left to right with an open book
and a green lamp on it. In the midground center, a girl in a navy coat stands at
the base of a tall spiral staircase, looking up. In the background, the
staircase winds up toward a huge arched window flooding the room with warm
light. Eye-level camera, balanced symmetrical composition, deep space from the
near table to the far window, detailed modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The base gave me clear, stable landmarks. The long wooden table with the book and green lamp is in front, the girl stands at the base of the spiral staircase in the middle, and the staircase winds up to the arched window at the back. The depth runs cleanly from the near table to the far window.&lt;/p&gt;

&lt;p&gt;Her face is small and a little muddy again, the recurring weakness at a distance, but the geography is exactly what I needed to test the camera moves against.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4vjrt2b6d7fvqglkggvi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4vjrt2b6d7fvqglkggvi.png" alt="The base library scene" width="800" height="1067"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The base library scene.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Re-shoot from a low angle
&lt;/h3&gt;

&lt;p&gt;First re-shoot: drop the camera to the foot of the stairs and look up, with the long wooden table now behind the camera and out of frame.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Keep this exact scene, the same library atrium, the same girl in the navy coat,
the same spiral staircase and arched window, but move the camera to a low angle
at the foot of the staircase looking steeply up. The girl is now seen from
below, the spiral staircase twists up dramatically above her toward the same
arched window, and the reading table is now behind the camera and out of frame.
Keep the same warm golden-hour light and the same room layout. Detailed modern
anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The camera move worked, but it kept something it should have dropped. The staircase dominates the frame, the camera is clearly low at its base, the arched window is correctly above, and the library is recognizably the same room.&lt;/p&gt;

&lt;p&gt;The miss is the long wooden table, which stays visible in the lower-left even though I put it behind the camera. So the model held the room but did not respect what should leave the frame after the move. That is a useful partial pass: strong on the angle and continuity, weaker on precise visibility.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq9eyjzg9960llxa6esyk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq9eyjzg9960llxa6esyk.png" alt="The base library scene (left). The low-angle re-shoot (right)." width="800" height="520"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The base library scene (left). The low-angle re-shoot (right).&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  The reverse angle
&lt;/h3&gt;

&lt;p&gt;Second re-shoot, the hard one: put the camera behind the girl and look back toward the table that was originally in the foreground.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Keep this exact scene and the same girl in the navy coat, but move the camera
behind her for an over-the-shoulder shot. We now look over her shoulder from
behind, past her, back toward the long reading table with the green lamp and
open book that was in the original foreground, now seen ahead of her across the
room. The spiral staircase and arched window are now behind the camera. Same
library, same warm golden-hour light, same layout, just reversed viewpoint.
Detailed modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This one impressed me a bit. The model reversed the viewpoint correctly. The table, lamp, and book are now ahead of her across the room, and the staircase and window are gone because they would be behind the camera. The girl looks like she is standing between the camera and the table, so it is a genuine reversal, not just spinning her around in the same forward view.&lt;/p&gt;

&lt;p&gt;It reconstructed the side of the room the first shot never showed, which takes real spatial understanding. The only weakness is framing: it came out as more of a rear view than a tight over-the-shoulder.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgmtmi0z57rj2xl7m64hl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgmtmi0z57rj2xl7m64hl.png" alt="The base library scene (left). The reverse-angle shot (right)." width="800" height="522"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The base library scene (left). The reverse-angle shot (right).&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cutting a sequence
&lt;/h2&gt;

&lt;p&gt;To finish, I shot one moment as a three-cut close-up sequence, the way a modern anime scene gets cut. A girl in a dark room, lit only by her phone.&lt;/p&gt;

&lt;h3&gt;
  
  
  The base, an extreme close-up on the eyes
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;An extreme close-up of a young woman's face in a dark room at night, the frame
filled edge to edge with just her eyes and the bridge of her nose, everything
else cropped out. Her face is lit from below by the cold blue glow of a phone
screen, tiny reflections of text visible in her eyes, a single strand of hair
falling across her forehead. Shallow focus, sharp on the eyes, soft everywhere
else. Intimate, cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzqz635es8sg3uaphppr3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzqz635es8sg3uaphppr3.png" alt=" " width="800" height="269"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It committed to the extreme close-up instead of backing off. The frame fills with her eyes and the bridge of her nose, the rest cropped out, and the cold phone glow lights her face from below. The eyes are large and sharp with blue reflections.&lt;/p&gt;

&lt;p&gt;The one weakness is the tiny text reflected in her eyes, which came out as distorted glowing marks rather than legible words. At that scale, that is a fidelity limit to note.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhxcnixkbxcanw4j417n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhxcnixkbxcanw4j417n.png" alt="The extreme close-up base shot." width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  First cut, to her hands on the phone
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Keep this exact scene, the same girl, the same dark room, the same cold blue
phone light, but cut to a close-up insert of her hands holding the phone at
chest height. We now see her thumbs on the screen and the message glowing on it,
her face soft and out of focus in the background above the phone. Same night
lighting, same blue glow on her fingers, shallow focus on the phone and hands.
Cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The match cut worked, but the phone got the logic wrong. The hands and phone became the sharp foreground, her face went soft above them, and the blue glow carried over.&lt;/p&gt;

&lt;p&gt;The problem is the phone itself. The rear camera module is clearly visible, so we are looking at the back of the phone, yet the model put the glowing message on that same back surface. That is a real object-logic error, since the screen should be on the front. The text also came out as invented Japanese-style characters.&lt;/p&gt;

&lt;p&gt;I gave the exact same prompt to PixAI &lt;a href="https://blog.pixai.art/en/pixai-edit-pro-ai-image-editor/" rel="noopener noreferrer"&gt;Edit Pro&lt;/a&gt;, and it made the same mistake, rear cameras visible with the message on the back. Edit Pro did render cleaner, legible English text and a more convincing messaging screen, but the physical error was identical. Two different models making the same mistake on the same prompt tells you it is a real blind spot, not a one-off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6s9k8s1ag0gni0yxdebe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6s9k8s1ag0gni0yxdebe.png" alt="(First Edit) Left the Tsubaki.3. Right the same edit on PixAI Edit Pro." width="799" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;(First Edit) Left the Tsubaki.3. Right the same edit on PixAI Edit Pro.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Second cut, pull back to reveal the room
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Keep this exact girl, her same face and expression and the same cold blue phone
light on her, but pull the camera back to a wider cinematic shot that reveals
the whole room for the first time. She is sitting alone on the floor at the foot
of her bed in a messy bedroom late at night, the phone glow the only strong
light, city lights faint through the window behind her, the room dark around
her. Same face, same lighting on her, now small within the wide lonely room.
Cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pull-back nailed the scale change. The camera pulls way back and drops her small into a full bedroom, with the bed, desk, shelves, and scattered books all coherent around her. She is on the floor at the foot of the bed, the phone glow still her main light, and the lonely late-night mood holds.&lt;/p&gt;

&lt;p&gt;Her face stays consistent enough to still look like the same girl, even though she is much smaller now. The model built a believable room it had never shown and placed her at the right scale in it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd18vthz6x2e5vzbxwxvb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd18vthz6x2e5vzbxwxvb.png" alt="The pull-back reveal" width="800" height="1057"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The pull-back reveal.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Which shots were hardest
&lt;/h2&gt;

&lt;p&gt;As an AI perspective generator, Tsubaki.3 is strong at the big camera moves. Low and high angles, deep three-layer depth, real fisheye distortion, one-point perspective, multi-subject placement, and holding a composition through a lighting change all came out reliable.&lt;/p&gt;

&lt;p&gt;It even kept the same room consistent across a low-angle re-shoot and a full reverse angle, which is real AI image perspective control, not luck.&lt;/p&gt;

&lt;p&gt;It slips on precision and small detail. It left the table in frame when the camera move should have dropped it. It put a phone screen on the back of a phone. It rendered distorted text in the eye reflection, and faces go muddy when the subject is small or far. So the camera itself is well understood, and the misses land in exact visibility, object logic, and fine detail.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8k9uiwrbi3ci5m5nw42.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8k9uiwrbi3ci5m5nw42.png" alt="hardest shots" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this helps creators
&lt;/h2&gt;

&lt;p&gt;For real work, this is useful. If you storyboard, build manga panels, or plan cinematic key art, you can call specific shots, a low hero angle, a deep corridor, a fisheye, a reverse angle, and get them, which is far faster than fighting the model for a non-default composition.&lt;/p&gt;

&lt;p&gt;It also works as an AI camera angle generator for concept and perspective references you can draw over.&lt;/p&gt;

&lt;p&gt;Where I would still plan for a manual pass: exact control over what stays in or out of frame after a camera move, close-up object details like a phone screen, and clean faces in deep or wide shots.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;The big takeaway is simple. When I described the camera instead of just the character, Tsubaki.3 followed. It gave me the angles, the depth, the distortion, and even the reverse shot, which is the hard part.&lt;/p&gt;

&lt;p&gt;It does not get everything right. The table that would not leave the frame, the phone screen on the wrong side, the muddy faces in far shots, those are all real limits. But when it comes to controlling how a scene is framed, it does more than I expected.&lt;/p&gt;

&lt;p&gt;If you have ever fought an AI model to get one specific angle, Tsubaki.3 is a real step up. &lt;a href="https://eap.pixai.art/go/naveed" rel="noopener noreferrer"&gt;Try it on PixAI&lt;/a&gt; with one shot you could never quite land, and see how it does.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Does Tsubaki.3 Understand Complex Editing Instructions, or Just Guess? 9 Relationship-Based Tests</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Mon, 24 Aug 2026 08:30:47 +0000</pubDate>
      <link>https://dev.to/naveedoss/does-tsubaki3-understand-complex-editing-instructions-or-just-guess-9-relationship-based-tests-1ndj</link>
      <guid>https://dev.to/naveedoss/does-tsubaki3-understand-complex-editing-instructions-or-just-guess-9-relationship-based-tests-1ndj</guid>
      <description>&lt;p&gt;I have run enough AI edits to know the easy ones do not tell you much. Change a color, remove an object, swap a background, and most editors manage that fine.&lt;/p&gt;

&lt;p&gt;The harder question is what happens when an instruction depends on how things relate to each other. The model either understands that relationship, or it drops the right objects in and hopes.&lt;/p&gt;

&lt;p&gt;That second case is what advanced AI image editing really means, and it is what I set out to test.&lt;/p&gt;

&lt;p&gt;I gave Tsubaki.3 nine edits that only work if it reasons about how elements connect: where things go relative to each other, what an object is made of, what changing the weather does to everything around it, and which of two similar characters an instruction points to.&lt;/p&gt;

&lt;p&gt;Plain object placement passes none of them. Here is how it did, grouped by the kind of thinking each one demanded rather than in the order I ran them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What makes an edit complex
&lt;/h2&gt;

&lt;p&gt;Complexity here does not come from the number of changes. It comes from how the changes depend on each other.&lt;/p&gt;

&lt;p&gt;"Change the dress, the bag, and the hair color" is three edits, but they are independent. Nothing about one affects the others. That is not complex, it is just a longer list.&lt;/p&gt;

&lt;p&gt;"Move the handbag from the table into her left hand while keeping the cup in front of her" is different. Now the model has to understand position, ownership, and what stays put.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5hz76pjqtyxrx3te7qze.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5hz76pjqtyxrx3te7qze.png" alt=" " width="800" height="477"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is complex AI image editing, because the instruction carries relationships, not just objects. Every test below is built that way. I wanted natural language image editing to earn its name, so each prompt describes a situation, not a checklist.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I scored
&lt;/h2&gt;

&lt;p&gt;This is really an AI image editing test with a strict rubric. I rate every result out of 10, and I judge it on five things: whether the model understood the instruction, whether it got the positions and relationships right, whether the related changes stayed visually logical, whether it preserved the parts it was not asked to touch, and whether the result would be useful in real work.&lt;/p&gt;

&lt;p&gt;Everything ran on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, PixAI's newest model. The basic editing workflow is covered in the &lt;a href="https://blog.pixai.art/en/pixai-edit-pro-ai-image-editor/" rel="noopener noreferrer"&gt;Edit Pro guide&lt;/a&gt;. This is the harder version of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Relationships, ownership, and meaning
&lt;/h2&gt;

&lt;p&gt;The first group of tests asks the model to understand what an instruction points to, not just which nouns it names.&lt;/p&gt;

&lt;h3&gt;
  
  
  Moving objects and getting the relationships right
&lt;/h3&gt;

&lt;p&gt;This checks spatial understanding: two objects swap places, and one has to end up in the correct hand. For the source image I used a &lt;a href="https://pixai.art/en/model/1892005535733745223/2044126787234234659" rel="noopener noreferrer"&gt;body aesthetic LoRA&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic anime scene of a young woman sitting at a small round cafe table
by a window. On the table in front of her sits a closed red handbag on the
left and a tall iced coffee on the right. Her hands rest in her lap. A folded
newspaper leans against the table leg on the floor. Warm afternoon light,
detailed modern anime illustration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, move the red handbag from the table into her right hand
so she is holding it up, and move the iced coffee to the left side of the table
where the bag was. Keep the newspaper on the floor exactly where it is, and
keep her seated in the same pose. Do not change anything else.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It handled every relationship correctly. The bag moved into her right hand with a natural grip, the coffee shifted to the left where the bag had been, and the newspaper stayed exactly where it was on the floor.&lt;/p&gt;

&lt;p&gt;The best part is that the model rebuilt her arm to hold the bag while keeping the pose and anatomy coherent, which is what you want from a real edit. Overall: 9.5.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn1xhq4e06ocay39b120j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn1xhq4e06ocay39b120j.png" alt=" " width="800" height="528"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 1 - Left source. Right after the edit.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Changing material without changing the object
&lt;/h3&gt;

&lt;p&gt;This tests whether the model can separate what an object is from what it is made of.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic still life, a single ornate teapot with a curved spout and a
looped handle sitting on a plain wooden table, soft window light from the left,
neutral grey background, detailed modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnh5eqcmvoe3m131zs4sr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnh5eqcmvoe3m131zs4sr.png" alt=" " width="799" height="267"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, change the teapot's material from matte ceramic to
polished chrome metal, keeping its exact same shape, spout, handle, size, and
position. The metal should show realistic reflections of the room and a bright
highlight from the window on the left, with the reflection of the light falling
correctly on the table. Do not change the teapot's design or anything else in
the scene.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The material change kept the shape and got the physics right. The teapot became chrome while its spout, handle, and size stayed the same, and the reflections and the window highlight landed on the correct left side.&lt;/p&gt;

&lt;p&gt;One thing to note on the source: the base image put a decorative anime figure on the teapot's body, which I never asked for, a reminder that the model likes to embellish. The edit dropped that surface art when it swapped to metal, which was fine here. Overall: 9.3.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuxm4yzs2us89imhicb9q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuxm4yzs2us89imhicb9q.png" alt=" " width="800" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 2 - Left source. Right after the edit.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Editing only one of two similar characters
&lt;/h3&gt;

&lt;p&gt;This is exclusion logic. Two near-identical girls, and the edit must touch only one, identified by what she is holding.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic anime scene of two girls standing side by side, both wearing
identical white hoodies and blue jeans. The girl on the left holds a
skateboard, the girl on the right holds a basketball. Plain street background,
detailed modern anime illustration, full body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, change only the hoodie of the girl holding the
basketball to bright red, and give only her a black cap. Leave the girl with
the skateboard completely unchanged, still in her white hoodie with no cap.
Keep both their faces, poses, jeans, and the objects they are holding exactly
the same.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It edited the correct girl and left the other one alone. Only the basketball girl's hoodie turned red, only she got the cap, and the skateboard girl stayed in her white hoodie. You can see a little drift in the basketball girl's hair color, though.&lt;/p&gt;

&lt;p&gt;The model understood "the girl holding the basketball" as a way to identify one person, which is harder than it looks with two similar characters. Overall: 9.5.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqcsvea3131tevcsxt37t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqcsvea3131tevcsxt37t.png" alt="Test 7 - Left source. Right after the exclusion edit." width="800" height="600"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 7 - Left source. Right after the exclusion edit.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I also ran a bigger edit on the same source, moving both girls into a pool scene with new outfits, activities, and poolside props. That result carried a lesson about unmentioned objects, so I have saved it for the limits section below.&lt;/p&gt;
&lt;h3&gt;
  
  
  Editing text and layout together
&lt;/h3&gt;

&lt;p&gt;This one uses an AI image editor with text prompts to make several typography changes at once, without disturbing the design.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A modern anime event poster, a DJ girl with headphones in the lower right, a
bold title "NEON NIGHTS" across the top, a small subtitle "Summer Festival"
beneath it, empty dark space in the upper left, a clean graphic layout with a
pink and blue color scheme. Detailed poster illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, replace the main title "NEON NIGHTS" with "ELECTRIC
DAWN", change the subtitle to "Rooftop Sessions", add a small date line "AUG 30"
in the upper left empty space, and add a small circular badge reading "18+" in
the bottom left corner. Keep the DJ girl, her position, the color scheme, and
the overall layout the same, and keep all text clear of her face.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every text change landed, and the layout held. The title became "ELECTRIC DAWN", the subtitle changed, the date went into the empty upper-left space, and the "18+" badge appeared in the corner.&lt;/p&gt;

&lt;p&gt;All of it spelled correctly, stayed off her face, and kept the pink-and-blue design intact. Short, controlled text like this renders cleanly, which is a genuine strength for poster work. Overall: 9.8, the highest of the run.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxtpat4yjjq7oi2grqfra.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxtpat4yjjq7oi2grqfra.png" alt="Left source. Right after the edit." width="800" height="526"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Following a chain of consequences
&lt;/h2&gt;

&lt;p&gt;The next two tests give the model one concept and ask it to work out everything that concept implies.&lt;/p&gt;

&lt;h3&gt;
  
  
  One weather change, every knock-on effect
&lt;/h3&gt;

&lt;p&gt;Here the instruction is one idea, rain, that should trigger a chain of related changes. This is the real test of scene logic.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic anime scene of a girl standing on a sunny city sidewalk in summer,
bright blue sky, sharp shadows, dry pavement, she wears a light sundress and
sunglasses, smiling, holding a closed umbrella loosely at her side. Detailed
modern anime illustration, full body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, turn this from a sunny afternoon into a heavy rainstorm
at the same spot, and make every change that would logically follow. The sky
should be dark and overcast, the pavement wet and reflective with puddles, rain
falling and dripping, her hair and dress damp, and she should now be holding
the umbrella open above her head, her expression shifting from a bright smile to
a smaller, cooler look. Keep it the same girl in the same place. Do not add
other people.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where the model showed real understanding. It did not just lay rain over the image.&lt;/p&gt;

&lt;p&gt;The sky darkened, the pavement turned wet and reflective with puddles, her hair and dress came out damp, the closed umbrella opened above her head, and her bright smile cooled to a smaller expression.&lt;/p&gt;

&lt;p&gt;Every consequence of "it is raining now" arrived together, which is exactly what the instruction was checking. Overall: 9.7, the strongest scene transformation of the set.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkc1y8ae034ymwe9itlhz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkc1y8ae034ymwe9itlhz.png" alt="Test 3 - Left source. Right after the edit." width="800" height="523"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Cause and effect, a spilled glass
&lt;/h3&gt;

&lt;p&gt;One physical action should ripple through the whole scene. This tests whether the model understands consequences.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic anime scene of a girl sitting at a dining table smiling at the
viewer, a tall full glass of red juice standing upright near her hand, a white
plate with food in front of her, a book lying open on the table to her left,
warm indoor light. Detailed modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, show the moment right after she has knocked the glass
of red juice over. The glass is now lying on its side, red juice spilled across
the table spreading toward the open book, a few drops running off the table
edge, her smile replaced with a shocked open-mouthed expression and her hands
pulled back. The spill should be soaking into the pages of the book nearest to
it. Keep the plate, the room, and her seat the same.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model told the whole cause-and-effect story, not just a tipped glass.&lt;/p&gt;

&lt;p&gt;The glass lies on its side, the juice spreads toward the book and soaks the pages, a stream runs off the table edge, and her smile flips to open-mouthed shock with her hands pulled back.&lt;/p&gt;

&lt;p&gt;It understood that one action produces several connected results and rendered all of them. Overall: 9.7.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg9vgyzwg4vw26gpzlca2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg9vgyzwg4vw26gpzlca2.png" alt="Test 6 - Left source. Right after the edit." width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where understanding reaches its limits
&lt;/h2&gt;

&lt;p&gt;The model reasons about meaning very well. The next tests are where that reasoning runs into physics and unstated intent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Moving the sun and the shadows
&lt;/h3&gt;

&lt;p&gt;This moves into physics. Change the time of day, and the shadows and reflections all have to obey the new light.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic anime scene of a girl standing in the middle of an empty city plaza
at noon, harsh overhead sun, short dark shadows pooled directly under her and
under a lamppost, a glass storefront on the left reflecting the bright sky, a
fountain on the right catching sunlight. Detailed modern anime illustration,
full body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, change the time from noon to late golden-hour sunset
with the sun low on the right side of the frame. Make every shadow long and
stretched to the left to match the low sun, including hers and the lamppost's.
Warm the whole scene to orange, change the storefront glass on the left to
reflect the orange sunset instead of blue sky, and make the fountain water
catch warm golden highlights. Keep the girl, the plaza, and every object in the
same position. Only the light, shadows, and reflections should change.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The lighting transformation was excellent, but the light source itself got the physics wrong.&lt;/p&gt;

&lt;p&gt;The long leftward shadows, the warm orange grade, and the golden highlights on the fountain all came out well. The problem is the sun. It rendered as a huge, blown-out disk hanging in front of the buildings rather than low on the horizon behind them, and the storefront reflected a second literal sun.&lt;/p&gt;

&lt;p&gt;I ran the edit twice and got the same issue both times. This is the useful finding: the model reasons about light and shadow very well, but it does not place a light source that obeys the scene's existing geometry. Overall: 8.2.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjsm6ba6hufl8d5tbxeyg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjsm6ba6hufl8d5tbxeyg.png" alt="test 5 Left to right source, first sunset attempt, second attempt." width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 5: Left to right source, first sunset attempt, second attempt.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Staging a complex action scene
&lt;/h3&gt;

&lt;p&gt;This test adds a new character mid-action and asks for a specific physical interaction. I used a rooftop scene with a falling sign, and asked for a well-known superhero to swing in and stop its fall.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic modern anime scene on a high-rise rooftop in a dense city at sunset.
A young anime boy and a young anime girl are standing near the rooftop edge
beside a large maintenance air-conditioning unit. The boy has short dark hair,
wears a blue hoodie and black cargo pants, and is holding a skateboard. The girl
has long brown hair, wears a yellow jacket and dark jeans, and is holding a small
backpack. A large advertising sign on the rooftop has partially broken loose and
is hanging at an angle over the edge. Wind is blowing their clothes and hair.
Tall skyscrapers fill the background, warm sunset light, dramatic clouds,
detailed modern anime illustration, cinematic composition, full body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, show the moment after the advertising sign has broken
completely free from the rooftop and is falling toward the street below.
Spider-Man has arrived from the upper right and has attached several webs to the
falling sign, pulling it away from the rooftop and slowing its descent. The sign
should be clearly separated from the building, with its broken supports and
cables trailing behind it. Have the boy and girl move closer to the rooftop edge
and lean forward slightly as they look down at the falling sign with shocked,
concerned expressions. Keep the boy's skateboard and the girl's backpack. Show
the street far below between the buildings, with small distant cars and
pedestrians visible to establish the height of the rooftop. Keep the rooftop, AC
unit, buildings, sunset, and other existing elements unchanged. Keep the same boy
and girl recognizable with their original faces, hair, clothing, and identities.
Do not add any other characters.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scene logic came together, but the physical interaction did not.&lt;/p&gt;

&lt;p&gt;The model added the new character, detached the sign, showed the street far below for scale, and had the boy and girl lean over the edge in shock, all while keeping the rooftop and both characters intact.&lt;/p&gt;

&lt;p&gt;Where it fell short is the webs. Instead of a few taut lines showing force pulling the sign, it drew many loose crossing strands, and the hero ended up crouched on top of the sign rather than swinging in to pull it. So it stages a complex action well but struggles with the physics of a specific interaction. Overall: 8.7.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi7i2c4npwsomfd58mayy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi7i2c4npwsomfd58mayy.png" alt="Test 8 - Left source. Right after the edit." width="800" height="600"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 8 - Left source. Right after the edit.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  The object it would not remove
&lt;/h3&gt;

&lt;p&gt;Back to the two girls from the exclusion test. I ran a second, bigger edit on that source, moving both into a pool scene with new outfits, activities, and poolside props.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact source image, transform the two girls into a sunny summer pool
scene while keeping the same two girls, their faces, hair, and overall character
identities recognizable. Move them from the street into a modern outdoor swimming
pool. Change their outfits into simple stylish summer swimwear appropriate for a
pool day. Have the girl who originally held the skateboard sitting on the edge of
the pool with her feet in the water, while the girl who originally held the
basketball is standing in the shallow water holding a colorful beach ball. Add
realistic poolside details: a few inflatable pool floats, two lounge chairs,
folded towels, a small table with cold drinks, and a small pet dog sitting beside
the pool watching them. Use bright summer sunlight, blue water with realistic
reflections, and a clean resort-like background. Keep both girls clearly
recognizable as the same characters. Do not add other people. Make the
composition cinematic and naturally integrated rather than looking like separate
elements pasted into a new background.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The transformation worked, but it carried over one thing it should not have. Both girls stayed recognizable, the pool setting and props all appeared, and the activities matched the prompt.&lt;/p&gt;

&lt;p&gt;The odd part is the skateboard, which the model kept beside the girl even though she is now in swimwear at a pool. I never said remove it, so the model left it.&lt;/p&gt;

&lt;p&gt;That shows the honest limit: it will not drop an object unless you tell it to, even when the new scene makes that object senseless. Overall: 9.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fusjacd0nbut48viwr732.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fusjacd0nbut48viwr732.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 7 - Left source. Right after the pool transformation.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Everything at once
&lt;/h2&gt;

&lt;p&gt;For the last test I combined every kind of understanding into a single instruction. Six interdependent changes, each of which has to stay logical with the others.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SOURCE
A cinematic anime scene inside a cozy bookshop cafe on a sunny afternoon. A girl
in a yellow raincoat sits at a wooden table by a large window on the left,
smiling, holding a closed book in her right hand, an empty ceramic mug on the
table in front of her. A black cat sleeps on the windowsill. Behind her, tall
bookshelves and a chalkboard menu on the back wall reading "OPEN". Warm sunlight
streams through the window, casting soft shadows to the right. A second girl in a
green sweater stands near the shelves holding a stack of books. Detailed modern
anime illustration, wide shot.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EDIT
Using this exact image, change the scene to a quiet night during a thunderstorm,
and make every change that logically follows. Turn the window dark with heavy
rain running down the glass and an occasional lightning flash. Switch the room's
light to warm interior lamplight, so all the shadows now fall away from the lamps
instead of from the window, and the sunlit highlights become soft warm indoor
ones. Fill the empty mug on the table with steaming hot coffee. Change the
chalkboard on the back wall from "OPEN" to "CLOSED". Have the black cat now awake
and sitting up, looking toward the window at the storm. Change only the standing
girl in the green sweater into a warm coat, but leave the seated girl in the
yellow raincoat exactly as she is, still holding her book. Keep both girls' faces,
the table, the mug's position, the bookshelves, and the overall composition the
same.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It got most of a very demanding instruction right. The scene became a night thunderstorm, the interior lighting switched to warm lamplight with the shadow direction flipping to match, the empty mug filled with coffee, the chalkboard changed from "OPEN" to "CLOSED", the cat woke and turned toward the storm, and only the standing girl changed into a coat while the seated girl kept her raincoat and book.&lt;/p&gt;

&lt;p&gt;That is a lot of connected logic handled in one pass.&lt;/p&gt;

&lt;p&gt;The misses were in the fine detail. The seated girl's face drifted from the original despite the instruction to keep it, and the standing girl's whole outfit changed rather than just the one item. So the big relationships held, and the small exact requirements slipped. Overall: 9.1.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4u56opmus74xj9odjtn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4u56opmus74xj9odjtn.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 9 - Left source. Right after the edit.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the best prompts had in common
&lt;/h2&gt;

&lt;p&gt;A few patterns showed up in the AI image editing prompts that worked best.&lt;/p&gt;

&lt;p&gt;Stating relationships plainly helped. "Into her right hand" and "to the left side where the bag was" gave the model something exact to reason about, and it delivered.&lt;/p&gt;

&lt;p&gt;Naming what to keep helped too, though it is not foolproof. The seated girl in the last test drifted even after I asked to keep her face.&lt;/p&gt;

&lt;p&gt;When a light source or a specific physical interaction matters, natural language alone struggles. Those are the cases where instruction-based image editing hits its limit, and you are better off doing the edit in stages or accepting a retry, as I did with the sunset.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it means for real workflows
&lt;/h2&gt;

&lt;p&gt;For everyday creative work, this holds up well. AI photo editing with prompts is reliable for the things most people need: moving and reassigning objects, changing materials, transforming a scene's weather or time while keeping the character, editing posters and text, and applying a change to one specific subject in a group.&lt;/p&gt;

&lt;p&gt;Where I would still plan for extra passes: precise light-source placement, exact physical interactions, and anything where an object needs to disappear because the new context demands it. Those benefit from clearer, staged instructions or a manual touch-up.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;Tsubaki.3 understands complex editing instructions better than I expected. When an instruction depends on position, ownership, material, scene logic, or telling two things apart, it grasps the intent and carries it out, and it keeps the related changes visually coherent.&lt;/p&gt;

&lt;p&gt;That is the harder half of advanced AI image editing, and it handles it.&lt;/p&gt;

&lt;p&gt;The limits are specific and consistent. It does not reliably place a physically plausible light source in a scene it is preserving, it cannot render the mechanics of a precise interaction, and it leaves unmentioned objects in place even when they no longer fit. So it understands meaning and relationships far better than physics and unstated intent.&lt;/p&gt;

&lt;p&gt;If you want to see where your own instruction lands, &lt;a href="https://eap.pixai.art/go/naveed2" rel="noopener noreferrer"&gt;try it on PixAI&lt;/a&gt;. Give it a real relationship to reason about, not just a list of changes, and watch how much of the logic it gets.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Can Tsubaki.3 Edit One Thing and Leave the Rest Alone? A Full Editing Stress Test</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Fri, 21 Aug 2026 10:38:43 +0000</pubDate>
      <link>https://dev.to/naveedoss/can-tsubaki3-edit-one-thing-and-leave-the-rest-alone-a-full-editing-stress-test-26cg</link>
      <guid>https://dev.to/naveedoss/can-tsubaki3-edit-one-thing-and-leave-the-rest-alone-a-full-editing-stress-test-26cg</guid>
      <description>&lt;p&gt;I run a lot of edits through AI models, and the same thing trips them up every time. They make the change you ask for, then quietly mess up something you did not.&lt;/p&gt;

&lt;p&gt;The jacket turns red, but the face shifts with it. You swap one object and the model redraws the whole hand. The crop moves on its own.&lt;/p&gt;

&lt;p&gt;That is the real test of AI image editing. Not whether a model can edit image with AI at all, but whether it can change one thing and leave everything else alone.&lt;/p&gt;

&lt;p&gt;So I ran Tsubaki.3 through a run of edits that got harder at each step, from a single color change to a full scene swap. I went through PixAI's &lt;a href="https://blog.pixai.art/en/pixai-edit-pro-ai-image-editor/" rel="noopener noreferrer"&gt;Edit Pro guide&lt;/a&gt; too, but I wanted to put the model through its paces myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two things I am scoring
&lt;/h2&gt;

&lt;p&gt;I score every image out of 10 on two axes.&lt;/p&gt;

&lt;p&gt;Instruction following. Whether the model made the change I asked for, and whether it made all of them when I asked for several.&lt;/p&gt;

&lt;p&gt;Preservation. Whether the character stayed the same person, and whether the pose, composition, style, lighting, and untouched objects held. The hard part is making an AI image editor preserve details it was never asked to touch, not just landing the edit.&lt;/p&gt;

&lt;p&gt;Everything ran on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, PixAI's newest model, in Ultra mode. I edited the same source image across a whole run, so you see both the change I wanted and the changes I did not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round one: the single-edit ladder
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The source image
&lt;/h3&gt;

&lt;p&gt;I built the starting image dense on purpose, so there would be plenty for a careless edit to disturb.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic modern anime portrait of a young street DJ girl standing at a
rooftop party at dusk, caught mid-laugh with her head tilted back. She has a
chin-length silver bob with one neon-blue streak over her left eye, warm brown
skin, a small crescent-moon stud under her right ear, and round orange-tinted
glasses pushed up on her forehead. She wears an oversized cropped varsity
jacket, teal on the left and magenta on the right, over a black tube top, with
layered silver necklaces. Her right hand rests on a pair of gold headphones
around her neck, and she holds a clear soda can in her left hand. Behind her,
string lights, a hazy pink-and-orange sky, distant city towers, and a blurred
crowd. Warm golden-hour light, soft bokeh, rich detail, cinematic modern anime
illustration, half body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It came out strong, a 9 out of 10. The silver bob, blue streak, two-tone jacket, glasses on the forehead, gold headphones, and the whole rooftop scene all landed.&lt;/p&gt;

&lt;p&gt;Two small deviations: the teal and magenta sides ended up reversed, and the clear soda can came out as a normal metallic one.&lt;/p&gt;

&lt;p&gt;Everything else gives me a solid baseline with lots of preservation targets: the blue streak, the crescent stud, the glasses, the headphones, the can, the crowd, the sky.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuaxq2pw7v7bnsy7d29ve.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuaxq2pw7v7bnsy7d29ve.png" alt=" " width="800" height="529"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: selected image for the review. Right: an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Test 1: a simple edit
&lt;/h3&gt;

&lt;p&gt;One isolated change, to set a baseline for targeted AI image editing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzv7pd13b8ma2zdess6u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzv7pd13b8ma2zdess6u.png" alt=" " width="800" height="570"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact image, change only the color of her varsity jacket to solid
deep red. Keep her face, silver bob with the blue streak, glasses on her
forehead, crescent stud, headphones, soda can, pose, and the rooftop
background exactly the same.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The edit worked and the character held. 8.5 out of 10.&lt;/p&gt;

&lt;p&gt;The jacket is deep red now, and her face, expression, skin tone, silver bob, blue streak, glasses, headphones, and the whole rooftop stayed consistent. Her pose and the soda can barely moved.&lt;/p&gt;

&lt;p&gt;The misses are in the details. "Solid deep red" was not taken literally, so the jacket picked up decorative patches instead of staying plain. And the framing shifted, with black bars appearing at the top and bottom. Even a simple edit nudged the aspect ratio.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2ulqhfqbuonl9kjqrw7w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2ulqhfqbuonl9kjqrw7w.png" alt=" " width="800" height="526"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the source. Right: the red-jacket edit.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Test 2: a structural edit
&lt;/h3&gt;

&lt;p&gt;Now a change that forces the model to rebuild part of the image: swap the object she is holding.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact image, replace the soda can in her left hand with a lit
sparkler throwing off small bright sparks. Keep her hand position, her face,
the silver bob and blue streak, the two-tone jacket, glasses, headphones, and
the rooftop background unchanged.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prop swap is excellent. 9 out of 10.&lt;/p&gt;

&lt;p&gt;The can is gone, and the sparkler with its bright spray of sparks looks convincing. Her face, hair, blue streak, glasses, and the background all held.&lt;/p&gt;

&lt;p&gt;But "keep her hand position" drifted slightly. The grip looks natural, it is just not the original grip. The black-bar framing shift is still here too.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9pp8qqe4t1zuwssqmcua.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9pp8qqe4t1zuwssqmcua.png" alt=" " width="799" height="524"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the source. Right: the sparkler edit.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Test 3: a complex edit
&lt;/h3&gt;

&lt;p&gt;This is a main test. Three coordinated changes in one instruction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact image, make three changes together. First, change her hair
from silver to deep purple while keeping the neon-blue streak. Second, pull the
orange glasses down onto her eyes instead of on her forehead. Third, change the
time to night, so the sky is dark blue with the string lights and city glowing
brighter. Keep her face, expression, crescent stud, two-tone jacket,
headphones, soda can, pose, and composition the same.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three changes landed cleanly. 8.5 out of 10, and probably the strongest edit of the run.&lt;/p&gt;

&lt;p&gt;The hair is deep purple with the blue streak still showing. The glasses moved from her forehead down onto her eyes with no facial distortion. And the day-to-night change is convincing: dark blue sky, stars, brighter city windows and string lights.&lt;/p&gt;

&lt;p&gt;The model rebuilt the lighting into a real night version instead of just dimming the sunset.&lt;/p&gt;

&lt;p&gt;Her identity held through all of it: face, expression, hairstyle shape, and skin tone. The only recurring weakness is the framing drift and black bars.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxilgus3a6m99nl57i5ci.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxilgus3a6m99nl57i5ci.png" alt=" " width="800" height="523"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the previous state. Right: the three-change night edit.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Test 4: an advanced edit
&lt;/h3&gt;

&lt;p&gt;The hardest edit of the round. Several changes at once, plus a full environment swap.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact image, transform the scene while keeping her recognizable.
Change her outfit to a sleek silver-and-black futuristic racing jacket, put a
glowing microphone in her right hand, and move her from the rooftop to a packed
neon concert stage at night with spotlights, smoke, and a huge crowd of
silhouettes with raised hands. Keep her face, silver bob with the blue streak,
glasses on her forehead, crescent stud, and her mid-laugh expression the same.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The big changes are excellent. 8.5 out of 10.&lt;/p&gt;

&lt;p&gt;The racing jacket, the glowing microphone, and the full rooftop-to-concert-stage transformation all came out well, with spotlights, smoke, and a raised-hands crowd. And she is still immediately recognizable through it all, since the bob and blue streak carry her identity.&lt;/p&gt;

&lt;p&gt;This is where preservation started slipping. The glasses stayed on her eyes instead of moving back to her forehead, so the model dropped one instruction. And the intense stage lighting washed her warm brown skin much paler.&lt;/p&gt;

&lt;p&gt;When the whole scene changes at once, the small, exact details are the first to go.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mybfwl0nnk2yoe1y18n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1mybfwl0nnk2yoe1y18n.png" alt=" " width="800" height="523"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the previous state. Right: the concert-stage edit.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  One thing to understand about editing in steps
&lt;/h3&gt;

&lt;p&gt;Here is a detail that explains a lot of the results above. Each edit ran on the previous image, not the original.&lt;/p&gt;

&lt;p&gt;So by Test 3, the jacket was still red from Test 1 and the sparkler was still in her hand from Test 2, even though those prompts mentioned the two-tone jacket and the soda can.&lt;/p&gt;

&lt;p&gt;That is not the model failing. It is the model correctly editing the image you gave it, not the one from three steps ago.&lt;/p&gt;

&lt;p&gt;Good to know when you chain edits: the model builds on what is in front of it, so restate what matters or start fresh from the original.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round two: editing a whole story, beat by beat
&lt;/h2&gt;

&lt;p&gt;The ladder tests single edits. This round tests something harder.&lt;/p&gt;

&lt;p&gt;I built a short story and edited the same character through it, shot by shot, to see if she still looks like herself by the end.&lt;/p&gt;

&lt;p&gt;The scene is a swordswoman in a bamboo forest, edited through six beats. Same character throughout, with strong anchors to track: a black high ponytail with a crimson cord, amber eyes, a cheek scar, a jade earring, and an indigo kimono.&lt;/p&gt;

&lt;h3&gt;
  
  
  Beat 1: the establishing shot
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic anime film still, a lone young swordswoman standing in a misty
bamboo forest at dawn. She has long black hair in a high ponytail with a single
crimson cord, sharp amber eyes, a thin scar on her left cheek, and a small jade
bead earring. She wears a deep indigo kimono top with silver trim, a grey sash,
and a sheathed katana at her hip, one hand resting calmly on the hilt. Soft
golden dawn light filters through the tall green bamboo, mist curling low around
her feet, petals drifting in the air. Calm, composed expression. Wide
atmospheric shot, muted natural colors, rich detail, cinematic modern anime
illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 9 out of 10 baseline. Every anchor is there and the forest is atmospheric. The only soft spot is the cheek scar, which comes out subtle.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi10yboteloaw0ldugdl8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi10yboteloaw0ldugdl8.png" alt=" " width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Beat 2: the threat arrives
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact image and character, keep her face, black ponytail with the
crimson cord, amber eyes, cheek scar, jade earring, and indigo kimono
identical. Change the moment: she has turned her head sharply to the left, eyes
narrowed and alert, her hand now gripping the katana hilt ready to draw. The
mist has thickened and darkened, dawn light dimming to a cold blue, and a tall
shadowy figure looms among the distant bamboo behind her. Same forest, same
framing, tense atmosphere.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Strong story beat, 8.5 out of 10. The head turn, the cold blue shift, and the looming shadow all land as the next shot of the same scene. Her identity held.&lt;/p&gt;

&lt;p&gt;The miss was "same framing." The camera moved much closer instead of keeping the wide shot.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo9m7mu7pg3len0a4q9v0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo9m7mu7pg3len0a4q9v0.png" alt=" " width="800" height="527"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: beat 1. Right: edited beat 2&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Beat 3: the clash
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, keep her face, black ponytail with the crimson
cord, amber eyes, cheek scar, jade earring, and indigo kimono identical. Now
show the clash: she has drawn the katana and swings it in a fast diagonal arc,
her ponytail and kimono whipping with the motion, a bright streak of light
trailing the blade. Sparks fly as her sword meets the shadowy attacker's weapon
mid-frame. Speed lines, flying petals and bamboo leaves, dynamic low angle,
cold blue light cut by the flash of the strike. Intense, kinetic anime action.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The best action beat, 9.3 out of 10, on the second try.&lt;/p&gt;

&lt;p&gt;My first generation looked great, but you could not tell what she was fighting. I re-rolled, and the second version pulled the shadowy attacker from Beat 2 into the foreground, so the sword impact and the sparks make sense.&lt;/p&gt;

&lt;p&gt;Identity and motion are excellent. The attacker's weapon stays a little lost in the silhouette. The lesson: a re-roll rescued the storytelling, not just the looks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F18s4zs9u4t3s9miwrp1g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F18s4zs9u4t3s9miwrp1g.png" alt=" " width="800" height="523"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: beat 2 as base. Right: beat 3 edited.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Beat 4: the aftermath
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, keep her face, black ponytail with the crimson
cord, amber eyes, cheek scar, jade earring, and indigo kimono identical. Final
beat: she stands still with her back mostly to us, sheathing the katana with a
soft click, head lowered, calm again. The shadowy figure is gone, defeated.
Warm golden dawn light returns and breaks through the bamboo, mist clearing,
petals settling. Quiet, resolved, peaceful atmosphere. Wide cinematic closing
shot.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The resolution lands, 8.5 out of 10. Warm dawn returns, the attacker is gone, petals settle, and the wide shot bookends the opening.&lt;/p&gt;

&lt;p&gt;The one weak spot is the action itself. The sheathing looks more like holding or drawing the sword. Her face drifts a little, but the rear view hides it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3nhc2kqzjrv4ygs59dse.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3nhc2kqzjrv4ygs59dse.png" alt=" " width="800" height="521"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: beat 3 as base. Right: beat 4 edited.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Beat 5: the journey continues
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character and image, keep her face, black ponytail with the
crimson cord, amber eyes, cheek scar, jade earring, and indigo kimono
identical. Continue the story after the battle: she is now riding a dark
chestnut horse along a narrow mountain path just beyond the bamboo forest,
moving away from the battlefield at dawn. Her katana is fully sheathed at her
hip, and one hand rests lightly on the horse's reins. Her long ponytail and
crimson cord move gently in the morning breeze. The bamboo forest is now behind
her, fading into the mist, while the path opens toward distant mountains and a
small village visible far ahead. Warm golden sunlight breaks through the clouds,
birds flying in the distance, a few petals still drifting behind her. She looks
calm and thoughtful, no longer tense. Wide cinematic anime film still, natural
movement, atmospheric depth, peaceful but purposeful mood, rich detail,
cinematic modern anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A genuine new beat, 9 out of 10. The chestnut horse is well rendered, the katana now clearly sheathed, and the forest gives way to mountains and a distant village.&lt;/p&gt;

&lt;p&gt;New location, new camera, new direction, same recognizable character. The horse's rear takes up a lot of foreground, and her hair blows harder than "gently," but the story moves forward cleanly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8k14twufjtv31vy9p541.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8k14twufjtv31vy9p541.png" alt=" " width="800" height="521"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: beat 4 as base. Right: beat 5 edited.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Beat 6: the end of the journey
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character and image, keep her face, black ponytail with the
crimson cord, amber eyes, cheek scar, jade earring, and indigo kimono
identical. Bring the story to its final scene: she has arrived at a quiet
mountain shrine overlooking the valley at sunset. Her dark chestnut horse
stands peacefully beside her, and her katana remains fully sheathed at her hip.
An elderly shrine keeper stands a short distance away beneath the wooden gate,
waiting for her with a gentle expression, while a small lantern glows beside
the shrine steps. She stands facing the shrine with her back mostly toward us,
one hand resting calmly on the horse's neck, her head slightly raised as she
looks toward the warm sunset beyond the mountains. The bamboo forest is now far
behind her. The sky is painted with soft orange and crimson light, distant
mountains fading into haze, a few birds crossing the sky, leaves moving gently
in the evening breeze. The atmosphere is peaceful, final, and resolved. Wide
cinematic closing shot, quiet emotional ending, rich atmospheric detail,
cinematic modern anime film illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A real ending, 9.2 out of 10. The shrine, the lantern, the elderly keeper, the same horse, and the sunset give the sequence a proper close.&lt;/p&gt;

&lt;p&gt;Her anchors held across all six generations. The bigger point: the six images form a real sequence.&lt;/p&gt;

&lt;p&gt;Her face rendering drifted across the run, but the hair, cord, earring, kimono, and sword stayed strong enough to carry her through, especially in the rear and side views.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsru4ggycwa8yccc46tma.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsru4ggycwa8yccc46tma.png" alt=" " width="799" height="523"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: beat 5 as base. Right: beat 6 edited.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Round three: pushing into semi-realistic
&lt;/h2&gt;

&lt;p&gt;For the last round I changed styles completely, using a semi-realistic &lt;a href="https://pixai.art/en/model/2013366204340385006" rel="noopener noreferrer"&gt;photorealistic-anime LoRA&lt;/a&gt; on Tsubaki.3, and built an even denser scene to edit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcibq2m9be8jv31nz4rri.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcibq2m9be8jv31nz4rri.png" alt=" " width="800" height="363"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The base: an observatory in a storm
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A cinematic photorealistic anime film still of a young Korean woman standing
alone inside a huge glass observatory on a mountain at night during a powerful
rainstorm. She has shoulder-length dark brown hair, slightly damp at the ends,
with one distinctive pale silver hair clip above her right temple, warm dark
brown eyes, natural Korean facial features, and a small beauty mark beneath her
left eye. She wears a long charcoal-gray wool coat over a cream knit sweater, a
dark pleated skirt, black ankle boots, and a thin burgundy scarf loosely
wrapped around her neck. She stands beside a tall telescope, one hand resting
on its metal frame while her other hand holds a small brass flashlight pointed
downward. Behind her, enormous curved glass windows reveal a stormy mountain
landscape, distant city lights far below, heavy rain running down the glass,
dark clouds lit by occasional lightning. The interior contains subtle
scientific instruments, a wooden desk with scattered star charts, a mechanical
clock, books, cables, and small warm lamps. Rain reflections and warm interior
lights create layered reflections across the glass. Wet footprints lead from
the entrance toward her. Cinematic photorealistic anime aesthetic, no text.
Wide cinematic three-quarter shot.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 9.2 out of 10 base, and a harder test than the swordswoman. Reflections in the glass, wet footprints, two separate hand interactions, a complex telescope, a beauty mark, a hair clip, all things a sloppy edit could easily mess up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc8sboyuj11izt04w171r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc8sboyuj11izt04w171r.png" alt=" " width="799" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Left: I’ll edit this one below. Right: Just for you to see.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The isolated edit: swap the object in her hand
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact image, make only one change: replace the small brass
flashlight in her hand with an old, slightly weathered Polaroid photograph held
naturally between her fingers. Keep everything else exactly the same: her face,
hair, hair clip, eyes, beauty mark, coat, sweater, scarf, skirt, boots, her
pose, the hand position and fingers holding the object, the telescope and her
other hand resting on it, the observatory interior, desk, books, clock, lamps,
cables, wet footprints, curved glass windows, rain, lightning, mountains, city
lights, mist, reflections, lighting, camera angle, framing, and composition.
The photograph should be the only meaningful change.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was the cleanest isolated edit of the whole project, 9 out of 10, but it took four attempts.&lt;/p&gt;

&lt;p&gt;Earlier tries turned the photo into a camera or rebuilt too much. The fourth got it: a small weathered photograph in her hand, the other hand still on the telescope, and the face, hair clip, beauty mark, reflections, footprints, lighting, and composition all intact.&lt;/p&gt;

&lt;p&gt;The honest note: a truly isolated edit on a dense scene can need several tries.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6la2kzgsr08omiee09xx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6la2kzgsr08omiee09xx.png" alt=" " width="799" height="652"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left: the previous state. Right: the edit.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The advanced edit: move her to a whole new place
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character and image, transform the scene into the next moment
of her story while keeping her unmistakably recognizable. Move her from the
mountain observatory to a quiet mountain railway platform just before sunrise,
standing beside an old stationary train with warm light glowing from its
windows, misty mountains and pine trees around her. The storm has passed.
Change her outfit completely: replace the charcoal coat, cream sweater, dark
skirt, and burgundy scarf with a long camel-colored travel coat over a dark
green knit sweater, a black ankle-length skirt, and dark leather boots. Keep
her face, facial proportions, warm dark brown eyes, dark brown shoulder-length
hair, pale silver hair clip above her right temple, small beauty mark beneath
her left eye, and overall identity exactly recognizable. She holds the same
weathered instant photograph in her right hand, lowered at her side. Do not
turn the photograph into a camera. Remove the telescope and all observatory
equipment. Replace the interior with the railway platform, train, wooden bench,
old station sign without readable text, luggage, small platform lamps, wet stone
pavement, and drifting mist. Wide cinematic three-quarter composition, soft
dawn light, premium photorealistic anime film aesthetic.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The best advanced edit in the set, 9.2 out of 10. Full outfit swap, full environment swap, and she is still unmistakably the same woman, with the hair, clip, eyes, beauty mark, and proportions all carried over.&lt;/p&gt;

&lt;p&gt;The photograph is the smart touch here. It stayed a photograph and tied the two scenes together.&lt;/p&gt;

&lt;p&gt;Small slips: the station sign shows faint fake text despite the "no text" instruction, and the framing turned more front-facing than asked.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwoukqw52k9b9zgqfymnm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwoukqw52k9b9zgqfymnm.png" alt=" " width="800" height="651"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The previous state left and the railway edit right&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes as the edit gets harder
&lt;/h2&gt;

&lt;p&gt;Put all the tests together and it is simple. An AI image editing test really asks two things: did it make the change, and did it leave everything else alone.&lt;/p&gt;

&lt;p&gt;The change itself is reliable. All three changes in Test 3 landed. The full outfit-and-location swap landed. Tsubaki.3 almost always does what you ask, even when you ask for a lot.&lt;/p&gt;

&lt;p&gt;The problem is that it also changes things you did not ask it to. The same few issues came up again and again.&lt;/p&gt;

&lt;p&gt;The framing moved on its own, and black bars showed up, across the whole first round.&lt;/p&gt;

&lt;p&gt;Swapping the object in her hand redrew the whole hand with it.&lt;/p&gt;

&lt;p&gt;On the hardest edit, it skipped one of my instructions, since the glasses stayed put in Test 4.&lt;/p&gt;

&lt;p&gt;Her face slowly changed across the six story shots, though the back and side views hid most of it.&lt;/p&gt;

&lt;p&gt;The one detailed edit on the observatory took four tries to get right.&lt;/p&gt;

&lt;p&gt;Two things helped. First, I told it what to keep, not just what to change. The details I repeated in every prompt, the hair, the eyes, the markings, are the ones that stayed put.&lt;/p&gt;

&lt;p&gt;Second, I kept hard edits simple. One difficult change on its own worked fine. Asking for several at once is where it missed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Tsubaki.3 editing works best
&lt;/h2&gt;

&lt;p&gt;From all of this, here is where I would reach for it.&lt;/p&gt;

&lt;p&gt;Correcting a single detail, like a color, an object, or a time of day.&lt;/p&gt;

&lt;p&gt;Making controlled variations of an image you want to keep.&lt;/p&gt;

&lt;p&gt;Swapping outfits and accessories while holding the character.&lt;/p&gt;

&lt;p&gt;Replacing an object, as long as you accept the area around it will be redrawn.&lt;/p&gt;

&lt;p&gt;Coordinated multi-changes, two or three at once, which it handles better than I expected.&lt;/p&gt;

&lt;p&gt;Story continuity, carrying one character across new scenes, which makes it a strong pick for AI anime image editing and short film-style sequences.&lt;/p&gt;

&lt;p&gt;Where I would slow down: anything needing pixel-exact framing, a single tiny change on a busy scene where you should expect retries, or long chains where you need the original details back, in which case start fresh.&lt;/p&gt;

&lt;h2&gt;
  
  
  My take
&lt;/h2&gt;

&lt;p&gt;Most of the time, Tsubaki.3 can change one thing without messing up the rest. It does this more reliably than most editors I have used, but not perfectly.&lt;/p&gt;

&lt;p&gt;It almost always makes the change you ask for, and it holds a character's core identity very well, even through a full outfit and location swap.&lt;/p&gt;

&lt;p&gt;What slips is the fine print: framing drifts, hands and nearby areas get rebuilt, and on the hardest edits one instruction can fall through. The more you change at once, the more the small untouched details wander.&lt;/p&gt;

&lt;p&gt;The honest verdict is that Tsubaki.3 is strong for edits where the identity matters more than pixel-exact preservation, and it can carry a character through an entire story. Just expect a retry or two on the fiddly ones.&lt;/p&gt;

&lt;p&gt;To see where your own image holds and where it drifts, &lt;a href="https://eap.pixai.art/go/naveed1" rel="noopener noreferrer"&gt;try it on PixAI yourself&lt;/a&gt;. Start with one clear edit, then make it harder and watch what moves.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>machinelearning</category>
      <category>art</category>
    </item>
    <item>
      <title>AI Character Design, Start to Finish: One Anime Character Across 10 Generations</title>
      <dc:creator>Naveed W</dc:creator>
      <pubDate>Thu, 20 Aug 2026 09:09:02 +0000</pubDate>
      <link>https://dev.to/naveedoss/ai-character-design-start-to-finish-one-anime-character-across-10-generations-17a</link>
      <guid>https://dev.to/naveedoss/ai-character-design-start-to-finish-one-anime-character-across-10-generations-17a</guid>
      <description>&lt;p&gt;Most people think AI character design means getting one good picture. It does not. The hard part is building a character once, then getting her back again in a new outfit, a new expression, and a new scene, still looking like the same person.&lt;/p&gt;

&lt;p&gt;That is what I set out to test. I took one original character through ten generations on &lt;a href="https://eap.pixai.art/go/naveed3" rel="noopener noreferrer"&gt;Tsubaki.3&lt;/a&gt;, PixAI's newest model, starting with rough concept sheets and ending with full cinematic scenes. Rather than march through them one to ten, I have grouped them into the four stages the work really moves through: building the design, developing it, pushing it into scenes, and combining it with a second character. Every prompt and score is below.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I approached it
&lt;/h2&gt;

&lt;p&gt;Good character design AI work follows a path, not a single prompt. You start wide with concept exploration, lock the direction you like into one clean reference, then use that reference to build everything else. That path is the core of AI character design, not one lucky prompt.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5y2dqgf1n0zemndfajyw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5y2dqgf1n0zemndfajyw.png" alt=" " width="800" height="387"&gt;&lt;/a&gt;&lt;br&gt;
Tsubaki.3 handles natural language, so I wrote full descriptive prompts instead of tag lists, and I repeated every defining feature in every prompt. The model drops details you stop mentioning, so the two-tone braid, the gold eyes, the red rune, the brass arm, and the potion vials appear in almost every prompt below. If you want the structure I leaned on, the &lt;a href="https://blog.pixai.art/en/how-to-write-pixai-prompts-formula/" rel="noopener noreferrer"&gt;PixAI prompt formula&lt;/a&gt; guide covers it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;A note on setup that applies throughout: I ran every test on Tsubaki.3 in Ultra mode, generating a batch of four each time. In the image pairs, the left image is the one I reviewed, so my notes always describe that one. The right image is a second pick from the same batch, there for extra visual appeal. Test 7 is the one exception, and I flag it there.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage one: building the design
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Concept exploration
&lt;/h3&gt;

&lt;p&gt;The first job of any AI character generator is to give you options. I asked for five takes on one idea so I could pick a direction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A character concept exploration sheet, five distinct full-body design
directions for a young female monster hunter, each visually different: a
rugged wilderness ranger in furs with bone charms, an elegant aristocratic
slayer in a tailored dark coat, a punk street hunter with piercings and a
spiked jacket, an alchemist hunter with a brass mechanical arm and a
bandolier of glowing potion vials, and a shrine-maiden exorcist with a staff
and long paper talismans. Varied silhouettes and color palettes, clean
concept art, neutral poses, plain background, labeled panels.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was a strong start. Tsubaki.3 gave me all five directions as separate characters instead of blending them into one, and each had its own silhouette and personality. &lt;/p&gt;

&lt;p&gt;The alchemist came out the most defined, with the brass arm and potion bandolier showing clearly. That is the direction I carried forward.&lt;/p&gt;

&lt;p&gt;The one real weakness was the labels. The model wrote text-like labels under each figure, but they are not real words. That is a common limit with generated images, and it does not hurt the actual AI character concept art.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fojyu8dfetxqgfv5rqhme.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fojyu8dfetxqgfv5rqhme.png" alt=" " width="800" height="595"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Left the base run. Right the Clear Style LoRA run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I ran it twice more. With the &lt;a href="https://pixai.art/en/model/2031410151574819222/2044126777985794127" rel="noopener noreferrer"&gt;Clear Style&lt;/a&gt; LoRA, the rendering got cleaner and sharper, but the five designs drifted toward the same face and body, and potion-tube details leaked onto characters that were not the alchemist. A third run with the Glassy Anime style preset came out worst. The plain base run stayed the strongest and most varied, so that is the one I kept.&lt;/p&gt;
&lt;h3&gt;
  
  
  Locking the design
&lt;/h3&gt;

&lt;p&gt;Once I had the direction, I generated one clean full-body reference. This image becomes the anchor for every test after it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A full-body character design of a young monster-hunting alchemist, modern
anime style. Athletic, agile build, standing confidently with a faint knowing
smile. Her hair is two-toned, snow-white on top blending into black
underneath, worn in a messy side braid over her right shoulder. Sharp
molten-gold eyes. A small red alchemical rune tattooed just under her left
eye. Her left forearm is a polished brass mechanical prosthetic with visible
gears at the elbow and wrist. A worn leather bandolier crosses her chest,
lined with glowing green and amber potion vials. She wears fitted dark travel
leathers with buckled straps. Neutral standing pose, plain soft-grey studio
background, clean detailed anime character design, crisp linework, soft even
lighting.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Most of the important features landed. The two-tone hair, side braid, gold eyes, brass mechanical arm, and glowing vials all came through, and the arm in particular looks convincing, with visible gears. The rune is the weak spot. &lt;/p&gt;

&lt;p&gt;It rests under the eye as asked, but it came out small and a little messy, so it lacks the clarity I would want from a signature feature. The outfit also turned out more revealing than the prompt described. &lt;/p&gt;

&lt;p&gt;Even so, this is a solid AI character sheet to build on, and it scored a 9 out of 10 for me. The left image is the reference I used from here forward.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftt0tjrco42ahkpsiu7ky.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftt0tjrco42ahkpsiu7ky.png" alt=" " width="800" height="592"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 2, Left the reviewed reference. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage two: developing the character
&lt;/h2&gt;

&lt;p&gt;With the reference set, the next three tests all pull from it: new outfits, new expressions, and a full turnaround.&lt;/p&gt;
&lt;h3&gt;
  
  
  Outfit variations
&lt;/h3&gt;

&lt;p&gt;I tested it as an AI outfit generator, using the locked reference as a Character Reference.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F08wyj93xcvf3morq2gxh.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F08wyj93xcvf3morq2gxh.jpg" alt=" " width="800" height="237"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, show her in four different full-body outfits lined
up side by side, the same young alchemist in each. Keep every defining feature
identical: the snow-white-to-black two-tone messy side braid, the molten-gold
eyes, the small red rune tattoo under her left eye, the brass mechanical left
forearm, and the leather bandolier of glowing potion vials across her chest.
Outfit one, rugged field-hunting leathers under a hooded travel cloak. Outfit
two, an ornate alchemists' guild ceremonial robe with gold embroidery. Outfit
three, relaxed town clothes, an oversized knit sweater and a canvas satchel.
Outfit four, heavy snow-expedition gear with a fur-lined hood and frost
goggles. Same height and build in each, plain background, consistent modern
anime design, clean linework.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Character consistency was the standout. The hair, braid, gold eyes, rune, and face stayed steady across all four looks, and the rune came out cleaner here than when I locked the design. All four outfits are distinct, and the guild robe with its gold embroidery is the best of them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsy0xq3h1carbdw8b0uc3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsy0xq3h1carbdw8b0uc3.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 3, Left the reviewed outfits. Right an alternate from the batch.&lt;/em&gt;&lt;br&gt;
The downside is that some accessories get hidden. In the snow gear especially, the mechanical arm and bandolier disappear under the clothing. &lt;/p&gt;

&lt;p&gt;The model keeps her identity, but it does not always keep every signature piece visible. A couple of the outer figures also got cropped at the edges.&lt;/p&gt;
&lt;h3&gt;
  
  
  Expression sheet
&lt;/h3&gt;

&lt;p&gt;Next I tested it as an AI expression sheet generator, six emotions on the same face.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, an expression sheet of six head-and-shoulders
portraits of the same young alchemist, changing only her expression. Keep her
two-tone white-and-black messy braid, molten-gold eyes, the small red rune
under her left eye, and her exact face structure identical in every panel. The
six expressions: a confident smirk, intense focus, worn-out exhaustion, a
bright delighted laugh, cold fury, and wide startled surprise. Plain
background, consistent modern anime style, clean linework.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was the most consistent test of the whole set. The character stays clearly herself across all six panels, and the expressions are properly different from each other, not small tweaks on one face. The surprised look works especially well, and the rune holds its shape and position throughout. &lt;/p&gt;

&lt;p&gt;Two small misses: the framing drops into the upper torso instead of staying head-and-shoulders, and the mechanical arm is not in frame, which makes sense given the crop. Neither hurts the result. I scored this one a 9.4.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx3mgh288wrxzx7tmtwxu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx3mgh288wrxzx7tmtwxu.png" alt=" " width="799" height="576"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 4, Left the reviewed expression sheet. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Turnaround
&lt;/h3&gt;

&lt;p&gt;Then the technical one, using Tsubaki.3 as a character turnaround generator: front, three-quarter, side, and back.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, a full character turnaround model sheet showing
four views at the same height: front, three-quarter, side profile, and back.
Keep every detail identical across all four, the white-to-black two-tone messy
side braid, molten-gold eyes, the red rune tattoo under her left eye, the brass
mechanical left forearm, the leather bandolier of glowing vials, and her fitted
dark travel leathers. Neutral A-pose, plain grey background, clean model-sheet
linework, even lighting.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where Tsubaki.3 is weaker. The four views are all there and the character stays recognizable, and the mechanical arm holds up impressively across every angle. But the side and back views drift. &lt;/p&gt;

&lt;p&gt;The braid changes shape and position, and a few clothing and accessory details shift from view to view. The pose also looks more like a natural lineup than a strict A-pose. The straight takeaway: Tsubaki.3 is much better at holding identity and expression than at producing an exact technical turnaround. &lt;/p&gt;

&lt;p&gt;The front view is excellent, and the side and back are where the small details wander. I scored it 8.5, still usable, just not precise.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdz3niwrg2xx5ewjcxq9d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdz3niwrg2xx5ewjcxq9d.png" alt=" " width="800" height="601"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 5, Left the reviewed expression sheet. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage three: pushing into scenes
&lt;/h2&gt;

&lt;p&gt;Sheets are the setup. The real question is whether the character survives once she leaves the plain background.&lt;/p&gt;
&lt;h3&gt;
  
  
  Key art
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, dynamic full-body key art of the young alchemist
mid-battle, hurling a glowing potion that erupts into a burst of green
alchemical fire. Her two-tone braid and coat whip through the motion, her
brass mechanical left arm extended, the vials on her bandolier glowing bright,
the red rune under her gold eye catching the light. Behind her, the massive
shadowed silhouette of a monster looms in billowing smoke. Dramatic rim
lighting, cinematic modern anime illustration, rich detail, full body.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This was probably the best image of the whole run. The character moved from a design sheet into cinematic key art without losing her identity. The mechanical arm, hair, gold eyes, rune, and glowing vials all carried over, the green explosion has real impact, and the monster silhouette gives the scene scale without stealing focus. The one miss is the coat. &lt;/p&gt;

&lt;p&gt;The prompt asked for it to whip through the motion, but the character wears it tied around her waist instead. Very dynamic poses like this make some clothing details harder to hold. Even so, her identity survived a big change in composition, which is the point. I scored it 9.2.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frlx2zmqn56g7umpmfktn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frlx2zmqn56g7umpmfktn.png" alt=" " width="799" height="597"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 6, Left the reviewed key art. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  A cinematic scene
&lt;/h3&gt;

&lt;p&gt;After the sheets, I wanted to see how far a detailed, story-driven prompt could push it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, a cinematic modern anime scene on the rooftop
platform of a speeding elevated train at sunset. She wears a stylish modern
black techwear jacket, fitted top, high-waisted cargo pants, and sleek combat
boots. A massive horned monster is leaping between train cars behind her while
she grabs a metal railing with her brass mechanical left hand, twisting around
to look back with a confident grin. Her long two-tone braid whips violently in
the wind, potion vials swinging from her bandolier. A second train races
alongside them, its windows glowing orange, while passengers inside stare in
shock. Below, a sprawling futuristic city rushes past: glowing billboards,
elevated roads, colorful storefronts, distant skyscrapers, and hundreds of
tiny lights. Sunset clouds fill the sky behind the monster, with sparks, loose
papers, and broken glass flying through the air. Vibrant modern anime film
frame, dramatic perspective, extreme sense of speed, rich environment, detailed
modern fashion, cinematic lighting, dynamic composition, vivid colors, polished
anime keyframe, highly detailed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This test needed two attempts, and that is the useful part. The first generation ignored the story completely and gave me a neon rainy street scene, no train, no monster, no passengers. So I refreshed and ran it again. The second attempt came together. &lt;/p&gt;

&lt;p&gt;The elevated train, the speeding city below, the passengers, the sunset, and the monster behind her all showed up, and the whipping braid and flying papers sell the speed. Her whole identity carried over, and the modern techwear outfit looked more convincing this time. &lt;/p&gt;

&lt;p&gt;A few things still drifted: she is not grabbing the railing exactly as described, and the second train looks like one track rather than two. The lesson is that a specific, dramatic scene prompt can push Tsubaki.3 into genuine film territory, but you should expect to re-roll it. The revised version scored 9.2.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fefywne8xc7po75n8m4y3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fefywne8xc7po75n8m4y3.png" alt=" " width="800" height="593"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 7, Left the reviewed image. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  A candid moment
&lt;/h3&gt;

&lt;p&gt;Every test so far was a posed shot. I wanted to see if it could do the opposite, a quiet, unposed moment.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using this exact character, a candid slice-of-life moment, not posed. The
young alchemist sprawled sideways in a worn armchair on a lazy afternoon, one
leg hooked over the armrest, half-asleep with a book about to slip from her
real hand while her brass mechanical hand dangles a half-eaten pastry.
Sunlight pools across her through a window, dust drifting in the beam. Her
two-tone braid is loose and messy, molten-gold eyes barely open, the red rune
under her eye catching the warm light. Cozy cluttered room softly out of focus
behind her. Natural, unguarded, warm, cinematic modern anime illustration,
off-center composition.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pose is the win here. She looks comfortable and half-asleep, sunk into the armchair with the book in her hand, and the warm light, dust, books, and worn furniture make the room feel lived-in. Her identity holds up cleanly, and the rune is one of the cleanest versions across all ten tests. &lt;/p&gt;

&lt;p&gt;The miss is a small object detail. The pastry rests below the mechanical hand instead of being held by it, so the model included both things but not the exact interaction. The framing is also a little more centered than the prompt asked. &lt;/p&gt;

&lt;p&gt;The result feels like a real quiet moment from her life, which is what I was after. I scored it 9.1.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy6hvr11pgvw4acbivap8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy6hvr11pgvw4acbivap8.png" alt=" " width="800" height="593"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 8, Left the reviewed image. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One thing I have to tell you: the small red alchemical rune never stayed fully consistent throughout this guide. It is always there, but you will notice it drifts slightly from the base image each time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbk0kvm8vwkc414c9kb9l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbk0kvm8vwkc414c9kb9l.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage four: two characters in one scene
&lt;/h2&gt;

&lt;p&gt;The last two tests do something the design sheets cannot: combine two separate character references into a single scene.&lt;/p&gt;
&lt;h3&gt;
  
  
  A two-character rescue
&lt;/h3&gt;

&lt;p&gt;I fed it the alchemist and a second character I uploaded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm0rf5wm3ztc3srhur5h1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm0rf5wm3ztc3srhur5h1.png" alt=" " width="800" height="351"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using the exact alchemist from @image1 and the exact girl from @image2. A
dramatic but beautiful anime film scene on an old moss-covered stone bridge
surrounded by lush green mountains and wildflowers. The girl from @image2 has
slipped over the edge of the bridge and is falling, reaching upward in
surprise. The alchemist grabs her wrist with her brass mechanical left hand and
braces herself against the stone railing, desperately pulling her back to
safety. Her two-tone braid and potion vials swing with the motion. A small
fluffy cat sits safely on the bridge watching the rescue, while colorful birds
burst from nearby trees. Ferns, vines, flowers, moss-covered stones,
butterflies, and tiny glowing particles fill the foreground. A clear turquoise
river winds through the valley far below, with waterfalls cascading down the
cliffs. In the background, a young boy stands on the bridge, startled and
running toward them. Bright morning sunlight filters through the trees. Strong
sense of height, movement, danger, and friendship, warm beautiful atmosphere.
Vibrant modern anime film scene, cinematic composition, expressive faces,
dynamic pose, rich green environment, natural character interaction, dramatic
perspective, atmospheric depth, highly detailed anime illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The story is clear at a glance. The girl is falling, the alchemist is reaching to save her, and the cat and boy react from the bridge. You understand the moment without the prompt. Both characters stayed recognizable too, the alchemist with her hair, gold eyes, braid, rune, and brass arm, and the second character with her pink hair, blue eyes, white dress, and bow. &lt;/p&gt;

&lt;p&gt;The environment does a lot of the work, with the mossy bridge, the valley, the waterfalls, and the butterflies giving you plenty to look at. &lt;/p&gt;

&lt;p&gt;The miss is the contact point. The alchemist reaches toward the falling girl, but whether she is gripping the wrist at all is ambiguous in the left image, and the boy looks more like a background extra than part of the moment. &lt;/p&gt;

&lt;p&gt;The right image handles it better. Still, combining two references into one narrative scene is a real step up from a character sheet. I scored it 9.3.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnww1idcwwp7w4bq8h8nw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnww1idcwwp7w4bq8h8nw.png" alt=" " width="800" height="591"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 9, Left the reviewed scene. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Modern fashion transformation
&lt;/h3&gt;

&lt;p&gt;For the last test, I kept both characters but changed everything else, swapping their fantasy roles for modern fashion.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Using @image1 and @image2 as character references. Keep both girls recognizable
with their original facial features, hair colors, eye colors, and identities,
but completely replace their fantasy outfits and roles with modern fashion. A
cinematic close-up of both girls relaxing together inside a stylish contemporary
creative lounge, dressed like modern Japanese fashion models. @image1 wears an
oversized charcoal blazer over a fitted cream top, a short pleated skirt, sheer
black tights, silver earrings, and sleek loafers. @image2 wears a soft blue
cropped cardigan, a white pleated skirt, delicate jewelry, and fashionable
sneakers. They sit together on a curved designer sofa, laughing naturally over
colorful iced drinks, leaning slightly toward each other as they look at
something on a phone. Around them are lush indoor plants, chrome furniture, art
books, vinyl records, a small neon sign, abstract artwork, hanging lights, and
large glass windows overlooking a distant sunset. Warm sunset light mixes with
soft colorful interior lighting, natural expressions, fashionable layered
clothing, detailed fabric textures, intimate friendship moment, close-up
cinematic framing, vibrant modern anime film aesthetic, polished anime
illustration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Identity preservation carried the test. The alchemist kept her white-to-black hair, gold eyes, rune, and face, and the second character kept her pink hair, blue eyes, and bow, while both moved convincingly into modern outfits. &lt;/p&gt;

&lt;p&gt;The lounge feels contemporary and social, and the phone gives them a natural reason to sit together rather than just pose. Two things to note. The close-up became more of a medium shot showing most of their bodies, and the alchemist's mechanical arm is gone entirely. &lt;/p&gt;

&lt;p&gt;That last one makes sense, since I asked to change her role and outfit, but it means her identity now rests on her face and hair rather than her strongest design feature. I scored it 9.2.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fydu6q16oexnmvrbdl44a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fydu6q16oexnmvrbdl44a.png" alt=" " width="800" height="597"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Test 10, Left the reviewed image. Right an alternate from the batch.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What held, and what drifted
&lt;/h2&gt;

&lt;p&gt;A few patterns ran through all ten generations. Consistency is the top strength. Tsubaki.3 held the character's face, hair, and signature details across outfits, expressions, and full cinematic scenes. It also combines two references well, keeping both characters recognizable in a single scene, which is hard for most tools.&lt;/p&gt;

&lt;p&gt;The weak points are just as clear. Turnarounds are the soft spot, since side and back views drift, so this is better for identity and expression than for an exact multi-view model sheet. &lt;/p&gt;

&lt;p&gt;It follows the big picture, not the fine print, so object interactions, exact framing, and specific poses often come out approximate, like the pastry, the coat, and the wrist grab. And tiny details drift the most. The small red rune under her eye was the least consistent feature of the whole run. &lt;/p&gt;

&lt;p&gt;It is in almost every image, but its shape and clarity shift a little each time, so a small marking like that is the first thing to wander from the base design.&lt;/p&gt;

&lt;p&gt;Two habits handle most of this. Name every feature, every time, since the model drops details you stop repeating. And when a scene prompt is complex, expect a re-roll, because a specific, dramatic prompt gives cinematic results but does not always land on the first try.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to keep a character consistent
&lt;/h2&gt;

&lt;p&gt;If you plan to reuse a character, the reference workflow above will carry you a long way. Set your locked design as a Character Reference and build from it, exactly as I did from the second test onward. PixAI's &lt;a href="https://blog.pixai.art/en/pixai-reference-pro-guide-multi-image-editing-with-natural-language/" rel="noopener noreferrer"&gt;Reference Pro guide&lt;/a&gt; covers that in depth. For a character you will use constantly, the stronger move is to train a &lt;a href="https://blog.pixai.art/en/train-lora-on-pixai/" rel="noopener noreferrer"&gt;character LoRA on PixAI&lt;/a&gt;. That turns your design into a model you can call any time instead of referencing an image on every generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next step
&lt;/h2&gt;

&lt;p&gt;Anime character design with AI is a process. You explore directions, lock one, then develop it through outfits, expressions, and scenes, checking consistency at each step. &lt;/p&gt;

&lt;p&gt;Tsubaki.3 handles most of that well, and it is strongest exactly where it counts, keeping your character recognizable as you move her around.&lt;/p&gt;

&lt;p&gt;If you want to build your own character this way, &lt;a href="https://eap.pixai.art/go/naveed2" rel="noopener noreferrer"&gt;try it on PixAI&lt;/a&gt;. Start with a concept sheet, lock the one you like, and reference it forward. You will have a usable character in a handful of generations.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>art</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
