I have run enough AI edits to know the easy ones do not tell you much. Change a color, remove an object, swap a background, and most editors manage that fine.
The harder question is what happens when an instruction depends on how things relate to each other. The model either understands that relationship, or it drops the right objects in and hopes.
That second case is what advanced AI image editing really means, and it is what I set out to test.
I gave Tsubaki.3 nine edits that only work if it reasons about how elements connect: where things go relative to each other, what an object is made of, what changing the weather does to everything around it, and which of two similar characters an instruction points to.
Plain object placement passes none of them. Here is how it did, grouped by the kind of thinking each one demanded rather than in the order I ran them.
What makes an edit complex
Complexity here does not come from the number of changes. It comes from how the changes depend on each other.
"Change the dress, the bag, and the hair color" is three edits, but they are independent. Nothing about one affects the others. That is not complex, it is just a longer list.
"Move the handbag from the table into her left hand while keeping the cup in front of her" is different. Now the model has to understand position, ownership, and what stays put.
That is complex AI image editing, because the instruction carries relationships, not just objects. Every test below is built that way. I wanted natural language image editing to earn its name, so each prompt describes a situation, not a checklist.
How I scored
This is really an AI image editing test with a strict rubric. I rate every result out of 10, and I judge it on five things: whether the model understood the instruction, whether it got the positions and relationships right, whether the related changes stayed visually logical, whether it preserved the parts it was not asked to touch, and whether the result would be useful in real work.
Everything ran on Tsubaki.3, PixAI's newest model. The basic editing workflow is covered in the Edit Pro guide. This is the harder version of that.
Relationships, ownership, and meaning
The first group of tests asks the model to understand what an instruction points to, not just which nouns it names.
Moving objects and getting the relationships right
This checks spatial understanding: two objects swap places, and one has to end up in the correct hand. For the source image I used a body aesthetic LoRA.
SOURCE
A cinematic anime scene of a young woman sitting at a small round cafe table
by a window. On the table in front of her sits a closed red handbag on the
left and a tall iced coffee on the right. Her hands rest in her lap. A folded
newspaper leans against the table leg on the floor. Warm afternoon light,
detailed modern anime illustration
EDIT
Using this exact image, move the red handbag from the table into her right hand
so she is holding it up, and move the iced coffee to the left side of the table
where the bag was. Keep the newspaper on the floor exactly where it is, and
keep her seated in the same pose. Do not change anything else.
It handled every relationship correctly. The bag moved into her right hand with a natural grip, the coffee shifted to the left where the bag had been, and the newspaper stayed exactly where it was on the floor.
The best part is that the model rebuilt her arm to hold the bag while keeping the pose and anatomy coherent, which is what you want from a real edit. Overall: 9.5.

Test 1 - Left source. Right after the edit.
Changing material without changing the object
This tests whether the model can separate what an object is from what it is made of.
SOURCE
A cinematic still life, a single ornate teapot with a curved spout and a
looped handle sitting on a plain wooden table, soft window light from the left,
neutral grey background, detailed modern anime illustration.
EDIT
Using this exact image, change the teapot's material from matte ceramic to
polished chrome metal, keeping its exact same shape, spout, handle, size, and
position. The metal should show realistic reflections of the room and a bright
highlight from the window on the left, with the reflection of the light falling
correctly on the table. Do not change the teapot's design or anything else in
the scene.
The material change kept the shape and got the physics right. The teapot became chrome while its spout, handle, and size stayed the same, and the reflections and the window highlight landed on the correct left side.
One thing to note on the source: the base image put a decorative anime figure on the teapot's body, which I never asked for, a reminder that the model likes to embellish. The edit dropped that surface art when it swapped to metal, which was fine here. Overall: 9.3.

Test 2 - Left source. Right after the edit.
Editing only one of two similar characters
This is exclusion logic. Two near-identical girls, and the edit must touch only one, identified by what she is holding.
SOURCE
A cinematic anime scene of two girls standing side by side, both wearing
identical white hoodies and blue jeans. The girl on the left holds a
skateboard, the girl on the right holds a basketball. Plain street background,
detailed modern anime illustration, full body.
EDIT
Using this exact image, change only the hoodie of the girl holding the
basketball to bright red, and give only her a black cap. Leave the girl with
the skateboard completely unchanged, still in her white hoodie with no cap.
Keep both their faces, poses, jeans, and the objects they are holding exactly
the same.
It edited the correct girl and left the other one alone. Only the basketball girl's hoodie turned red, only she got the cap, and the skateboard girl stayed in her white hoodie. You can see a little drift in the basketball girl's hair color, though.
The model understood "the girl holding the basketball" as a way to identify one person, which is harder than it looks with two similar characters. Overall: 9.5.

Test 7 - Left source. Right after the exclusion edit.
I also ran a bigger edit on the same source, moving both girls into a pool scene with new outfits, activities, and poolside props. That result carried a lesson about unmentioned objects, so I have saved it for the limits section below.
Editing text and layout together
This one uses an AI image editor with text prompts to make several typography changes at once, without disturbing the design.
SOURCE
A modern anime event poster, a DJ girl with headphones in the lower right, a
bold title "NEON NIGHTS" across the top, a small subtitle "Summer Festival"
beneath it, empty dark space in the upper left, a clean graphic layout with a
pink and blue color scheme. Detailed poster illustration.
EDIT
Using this exact image, replace the main title "NEON NIGHTS" with "ELECTRIC
DAWN", change the subtitle to "Rooftop Sessions", add a small date line "AUG 30"
in the upper left empty space, and add a small circular badge reading "18+" in
the bottom left corner. Keep the DJ girl, her position, the color scheme, and
the overall layout the same, and keep all text clear of her face.
Every text change landed, and the layout held. The title became "ELECTRIC DAWN", the subtitle changed, the date went into the empty upper-left space, and the "18+" badge appeared in the corner.
All of it spelled correctly, stayed off her face, and kept the pink-and-blue design intact. Short, controlled text like this renders cleanly, which is a genuine strength for poster work. Overall: 9.8, the highest of the run.
Following a chain of consequences
The next two tests give the model one concept and ask it to work out everything that concept implies.
One weather change, every knock-on effect
Here the instruction is one idea, rain, that should trigger a chain of related changes. This is the real test of scene logic.
SOURCE
A cinematic anime scene of a girl standing on a sunny city sidewalk in summer,
bright blue sky, sharp shadows, dry pavement, she wears a light sundress and
sunglasses, smiling, holding a closed umbrella loosely at her side. Detailed
modern anime illustration, full body.
EDIT
Using this exact image, turn this from a sunny afternoon into a heavy rainstorm
at the same spot, and make every change that would logically follow. The sky
should be dark and overcast, the pavement wet and reflective with puddles, rain
falling and dripping, her hair and dress damp, and she should now be holding
the umbrella open above her head, her expression shifting from a bright smile to
a smaller, cooler look. Keep it the same girl in the same place. Do not add
other people.
This is where the model showed real understanding. It did not just lay rain over the image.
The sky darkened, the pavement turned wet and reflective with puddles, her hair and dress came out damp, the closed umbrella opened above her head, and her bright smile cooled to a smaller expression.
Every consequence of "it is raining now" arrived together, which is exactly what the instruction was checking. Overall: 9.7, the strongest scene transformation of the set.
Cause and effect, a spilled glass
One physical action should ripple through the whole scene. This tests whether the model understands consequences.
SOURCE
A cinematic anime scene of a girl sitting at a dining table smiling at the
viewer, a tall full glass of red juice standing upright near her hand, a white
plate with food in front of her, a book lying open on the table to her left,
warm indoor light. Detailed modern anime illustration.
EDIT
Using this exact image, show the moment right after she has knocked the glass
of red juice over. The glass is now lying on its side, red juice spilled across
the table spreading toward the open book, a few drops running off the table
edge, her smile replaced with a shocked open-mouthed expression and her hands
pulled back. The spill should be soaking into the pages of the book nearest to
it. Keep the plate, the room, and her seat the same.
The model told the whole cause-and-effect story, not just a tipped glass.
The glass lies on its side, the juice spreads toward the book and soaks the pages, a stream runs off the table edge, and her smile flips to open-mouthed shock with her hands pulled back.
It understood that one action produces several connected results and rendered all of them. Overall: 9.7.
Where understanding reaches its limits
The model reasons about meaning very well. The next tests are where that reasoning runs into physics and unstated intent.
Moving the sun and the shadows
This moves into physics. Change the time of day, and the shadows and reflections all have to obey the new light.
SOURCE
A cinematic anime scene of a girl standing in the middle of an empty city plaza
at noon, harsh overhead sun, short dark shadows pooled directly under her and
under a lamppost, a glass storefront on the left reflecting the bright sky, a
fountain on the right catching sunlight. Detailed modern anime illustration,
full body.
EDIT
Using this exact image, change the time from noon to late golden-hour sunset
with the sun low on the right side of the frame. Make every shadow long and
stretched to the left to match the low sun, including hers and the lamppost's.
Warm the whole scene to orange, change the storefront glass on the left to
reflect the orange sunset instead of blue sky, and make the fountain water
catch warm golden highlights. Keep the girl, the plaza, and every object in the
same position. Only the light, shadows, and reflections should change.
The lighting transformation was excellent, but the light source itself got the physics wrong.
The long leftward shadows, the warm orange grade, and the golden highlights on the fountain all came out well. The problem is the sun. It rendered as a huge, blown-out disk hanging in front of the buildings rather than low on the horizon behind them, and the storefront reflected a second literal sun.
I ran the edit twice and got the same issue both times. This is the useful finding: the model reasons about light and shadow very well, but it does not place a light source that obeys the scene's existing geometry. Overall: 8.2.

Test 5: Left to right source, first sunset attempt, second attempt.
Staging a complex action scene
This test adds a new character mid-action and asks for a specific physical interaction. I used a rooftop scene with a falling sign, and asked for a well-known superhero to swing in and stop its fall.
SOURCE
A cinematic modern anime scene on a high-rise rooftop in a dense city at sunset.
A young anime boy and a young anime girl are standing near the rooftop edge
beside a large maintenance air-conditioning unit. The boy has short dark hair,
wears a blue hoodie and black cargo pants, and is holding a skateboard. The girl
has long brown hair, wears a yellow jacket and dark jeans, and is holding a small
backpack. A large advertising sign on the rooftop has partially broken loose and
is hanging at an angle over the edge. Wind is blowing their clothes and hair.
Tall skyscrapers fill the background, warm sunset light, dramatic clouds,
detailed modern anime illustration, cinematic composition, full body.
EDIT
Using this exact image, show the moment after the advertising sign has broken
completely free from the rooftop and is falling toward the street below.
Spider-Man has arrived from the upper right and has attached several webs to the
falling sign, pulling it away from the rooftop and slowing its descent. The sign
should be clearly separated from the building, with its broken supports and
cables trailing behind it. Have the boy and girl move closer to the rooftop edge
and lean forward slightly as they look down at the falling sign with shocked,
concerned expressions. Keep the boy's skateboard and the girl's backpack. Show
the street far below between the buildings, with small distant cars and
pedestrians visible to establish the height of the rooftop. Keep the rooftop, AC
unit, buildings, sunset, and other existing elements unchanged. Keep the same boy
and girl recognizable with their original faces, hair, clothing, and identities.
Do not add any other characters.
The scene logic came together, but the physical interaction did not.
The model added the new character, detached the sign, showed the street far below for scale, and had the boy and girl lean over the edge in shock, all while keeping the rooftop and both characters intact.
Where it fell short is the webs. Instead of a few taut lines showing force pulling the sign, it drew many loose crossing strands, and the hero ended up crouched on top of the sign rather than swinging in to pull it. So it stages a complex action well but struggles with the physics of a specific interaction. Overall: 8.7.

Test 8 - Left source. Right after the edit.
The object it would not remove
Back to the two girls from the exclusion test. I ran a second, bigger edit on that source, moving both into a pool scene with new outfits, activities, and poolside props.
EDIT
Using this exact source image, transform the two girls into a sunny summer pool
scene while keeping the same two girls, their faces, hair, and overall character
identities recognizable. Move them from the street into a modern outdoor swimming
pool. Change their outfits into simple stylish summer swimwear appropriate for a
pool day. Have the girl who originally held the skateboard sitting on the edge of
the pool with her feet in the water, while the girl who originally held the
basketball is standing in the shallow water holding a colorful beach ball. Add
realistic poolside details: a few inflatable pool floats, two lounge chairs,
folded towels, a small table with cold drinks, and a small pet dog sitting beside
the pool watching them. Use bright summer sunlight, blue water with realistic
reflections, and a clean resort-like background. Keep both girls clearly
recognizable as the same characters. Do not add other people. Make the
composition cinematic and naturally integrated rather than looking like separate
elements pasted into a new background.
The transformation worked, but it carried over one thing it should not have. Both girls stayed recognizable, the pool setting and props all appeared, and the activities matched the prompt.
The odd part is the skateboard, which the model kept beside the girl even though she is now in swimwear at a pool. I never said remove it, so the model left it.
That shows the honest limit: it will not drop an object unless you tell it to, even when the new scene makes that object senseless. Overall: 9.

Test 7 - Left source. Right after the pool transformation.
Everything at once
For the last test I combined every kind of understanding into a single instruction. Six interdependent changes, each of which has to stay logical with the others.
SOURCE
A cinematic anime scene inside a cozy bookshop cafe on a sunny afternoon. A girl
in a yellow raincoat sits at a wooden table by a large window on the left,
smiling, holding a closed book in her right hand, an empty ceramic mug on the
table in front of her. A black cat sleeps on the windowsill. Behind her, tall
bookshelves and a chalkboard menu on the back wall reading "OPEN". Warm sunlight
streams through the window, casting soft shadows to the right. A second girl in a
green sweater stands near the shelves holding a stack of books. Detailed modern
anime illustration, wide shot.
EDIT
Using this exact image, change the scene to a quiet night during a thunderstorm,
and make every change that logically follows. Turn the window dark with heavy
rain running down the glass and an occasional lightning flash. Switch the room's
light to warm interior lamplight, so all the shadows now fall away from the lamps
instead of from the window, and the sunlit highlights become soft warm indoor
ones. Fill the empty mug on the table with steaming hot coffee. Change the
chalkboard on the back wall from "OPEN" to "CLOSED". Have the black cat now awake
and sitting up, looking toward the window at the storm. Change only the standing
girl in the green sweater into a warm coat, but leave the seated girl in the
yellow raincoat exactly as she is, still holding her book. Keep both girls' faces,
the table, the mug's position, the bookshelves, and the overall composition the
same.
It got most of a very demanding instruction right. The scene became a night thunderstorm, the interior lighting switched to warm lamplight with the shadow direction flipping to match, the empty mug filled with coffee, the chalkboard changed from "OPEN" to "CLOSED", the cat woke and turned toward the storm, and only the standing girl changed into a coat while the seated girl kept her raincoat and book.
That is a lot of connected logic handled in one pass.
The misses were in the fine detail. The seated girl's face drifted from the original despite the instruction to keep it, and the standing girl's whole outfit changed rather than just the one item. So the big relationships held, and the small exact requirements slipped. Overall: 9.1.

Test 9 - Left source. Right after the edit.
What the best prompts had in common
A few patterns showed up in the AI image editing prompts that worked best.
Stating relationships plainly helped. "Into her right hand" and "to the left side where the bag was" gave the model something exact to reason about, and it delivered.
Naming what to keep helped too, though it is not foolproof. The seated girl in the last test drifted even after I asked to keep her face.
When a light source or a specific physical interaction matters, natural language alone struggles. Those are the cases where instruction-based image editing hits its limit, and you are better off doing the edit in stages or accepting a retry, as I did with the sunset.
What it means for real workflows
For everyday creative work, this holds up well. AI photo editing with prompts is reliable for the things most people need: moving and reassigning objects, changing materials, transforming a scene's weather or time while keeping the character, editing posters and text, and applying a change to one specific subject in a group.
Where I would still plan for extra passes: precise light-source placement, exact physical interactions, and anything where an object needs to disappear because the new context demands it. Those benefit from clearer, staged instructions or a manual touch-up.
My take
Tsubaki.3 understands complex editing instructions better than I expected. When an instruction depends on position, ownership, material, scene logic, or telling two things apart, it grasps the intent and carries it out, and it keeps the related changes visually coherent.
That is the harder half of advanced AI image editing, and it handles it.
The limits are specific and consistent. It does not reliably place a physically plausible light source in a scene it is preserving, it cannot render the mechanics of a precise interaction, and it leaves unmentioned objects in place even when they no longer fit. So it understands meaning and relationships far better than physics and unstated intent.
If you want to see where your own instruction lands, try it on PixAI. Give it a real relationship to reason about, not just a list of changes, and watch how much of the logic it gets.





Top comments (0)