DEV Community

Cover image for A Hands-On Walkthrough of Advanced AI Image Editing With Tsubaki.3
Abirami Vina
Abirami Vina

Posted on Originally published at Medium

A Hands-On Walkthrough of Advanced AI Image Editing With Tsubaki.3

Can an AI editor understand relationships, materials, and layout? See how Tsubaki.3 handled four advanced AI image editing tests, including where it slipped.


If you have used an AI image editor, you already know it can change a hair color or delete a stray object. Those edits are easy to request and easy to check. Advanced AI image editing picks up after that point. In-depth revisions usually involve position, ownership, material, atmosphere, or layout, and each depends on the model understanding not just what is in the image, but how the pieces fit together.

Take a simple anime café scene with two characters. One is working on a laptop while the other is using a phone.

Two anime characters sitting at a café table, one in a black hoodie and blue headphones using a laptop, the other in a blue trench coat using a phone.

The original café scene we used as a starting point for our advanced AI image editing tests.

Now ask an AI model to swap their devices, without changing either character's pose, position, expression, or anything else in the scene. That sounds like a straightforward edit.

But the model needs to understand who is using each object, where the objects should move, what other items are connected to them, and which parts of the image have to remain exactly the same. This is where advanced AI image editing gets interesting. Basic edits are easy to test, but more complex instructions show how well a model understands relationships, context, and the overall logic of a scene.

Sometimes, the edits may not go according to plan. For instance, clothing or character poses may change, or extra elements may be added, as in the example below. We tried to swap the devices the characters were using, but the model included an extra headphone in the wrong place.

Although it's not what we asked for, the edited image was still contextually accurate and usable. This was the prompt we used:

"Swap the gadgets and headphones between both characters: give the man in the black hoodie the phone and remove his blue headphones, while giving the man in the blue coat the laptop and the blue headphones; keep their faces, hairstyles, outfits, poses, expressions, positions, and the café setting exactly the same."

And this is the output we got.

Before-and-after comparison showing an anime-style character scene, with the original image above and the edited version below, where the characters’ devices have been swapped.

An Editing Scenario Where Characters' Devices Are Swapped Through Edits

The edit above and the rest of the edits we'll cover in this article were made using PixAI, an AI platform for creating and editing anime-style images. Every result came from Tsubaki.3, PixAI's latest model, designed for precise, instruction-based anime image generation and editing.

From here, we'll run Tsubaki.3 through several advanced AI image editing scenarios and see how well it reads complex editing intent, visual relationships, and logical dependencies. Let's get started!

What Makes AI Image Editing Prompts Complex?

Advanced AI image editing isn't necessarily an edit with a long prompt or a large number of requested changes. Complexity often comes from the relationships between different elements in an image.

For example, changing a dress, a bag, and a character's hair color may involve several edits, but each change can be handled easily when they are done independently. However, asking a model to move a handbag from a table into a character's left hand while keeping everything else in place is more complex.

Side-by-side comparison of an anime girl at a café table, with a tan handbag on the table in the original on the left and held in her hand in the edited version on the right.

The handbag moved from the table into the character's hand while the coffee cup and the rest of the scene stayed in place.

The model needs to understand where the bag was originally located, who it should belong to after the edit, which hand it should be placed in, and how the character should interact with it. The same applies to other scenarios involving AI photo editing with prompts.

For instance, changing a fabric material to leather requires the model to preserve the object's original shape while adjusting its texture, highlights, and reflections. Similarly, turning a sunny street into a rainy one can affect the lighting, atmosphere, and surfaces across the entire scene. Even editing text can become complex when the new words need to fit within the existing placement, spacing, and overall layout.

Next, we'll test out some of these edits using the Tsubaki.3 model and see how much of this relational logic it works out on its own, and where a clearer instruction helps it along.

Test 1: Spatial Relationships in Instruction-Based Image Editing

Our first test was to check whether Tsubaki.3 truly understands how objects and characters fit together in a scene. We pushed the model beyond basic object recognition by checking whether it grasps spatial and contextual relationships, like left-and-right positioning and how characters interact with their surroundings.

We generated two characters, placed them in a park, and asked the model to swap their positions from left to right. Here is the edit prompt we ran:

"Interchange the positions of the characters in the image. The man on the left is on the right, and vice versa."

And here's the output.

Before-and-after comparison showing the original image above and the edited image below, with the characters changing positions while the rest of the scene remains consistent.

Original Image (Above) And Edited Image (Below) With Characters Changing Positions.

The main goal was to see if the requested change happened accurately while everything else stayed still. The model changed the characters' positions, but one character's pose changed slightly. Although the pose changed, the scene still fits contextually.

A successful edit keeps the characters' appearances, poses, facial expressions, camera angles, and main background features identical. It should only change what was specifically asked for. At the same time, any modified objects need to maintain realistic scale, rotation, and placement so they blend naturally into the scene.

Overall, this advanced AI image editing test showed that Tsubaki.3 reads left and right positioning accurately and can move characters across a scene without losing their identities. Also, the handbag edit earlier in this article tested a different kind of spatial relationship, moving an object out of a scene and into a character's grip, which asks the model to work out ownership and contact rather than left and right. Tsubaki.3 handled that cleanly as well.

Test 2: Changing Materials Without Changing the Object

Our second test asks whether Tsubaki.3 can change what an object is made of without changing what it is. A good result keeps the shape, proportions, and design exactly as they were and changes only the texture and how the surface responds to light.

We used Tsubaki.3 to change a character's cotton trench coat to a leather trench coat while keeping the same clothing design, fit, and color. We wanted the edited coat to look like the same trench coat, but with a smooth leather texture, subtle sheen, and the characteristic way leather catches and reflects light.

Check out the edit prompt we used for this advanced AI image editing test:

"Change the fabric of the blue trench coat to leather. Make sure it's still blue, but in leather."

This was our result.

Side-by-side comparison showing the original image on the left and the edited image on the right, where the character’s blue cotton trench coat has been changed to a blue leather trench coat.

Original Image (Left) And Edited Image (Right) With Blue Leather Trench Coat.

Tsubaki.3 changed the material without redesigning the coat. The collar, lapels, belt, and length are all unchanged, while the surface now behaves like leather, with tighter highlights along the folds and a sheen that cotton wouldn't produce.

The one drift is color. We asked to keep it blue, and it stayed blue, but the shade deepened toward navy rather than holding the original lighter tone.

This could come down to how leather handles light. It reflects in sharper, narrower bands than woven fabric, which leaves a larger share of the surface in shadow and pushes the overall tone toward navy. If holding an exact shade is important to your project, Tsubaki.3's color palette feature lets you hand the model actual hex codes rather than describing a color in words.

Ultimately, Tsubaki.3 kept the object and swapped only its surface, which is what makes this kind of edit useful for testing finishes and textures without rebuilding a design.

Test 3: Semantic Scene Changes Through Natural Language Image Editing

Our third test looks at how well Tsubaki.3 can change the look and feel of an entire scene while keeping all the changes visually connected. To do this, the model needs to understand the overall scene and adjust multiple elements to match the new setting.

For example, changing a sunny afternoon park scene into a nighttime scene means updating the lighting, colors, atmosphere, and reflections so everything feels familiar but in a different environment. So we built that scene and ran it.

Starting from a sunny morning park, we gave Tsubaki.3 one line and deliberately kept it short to see how much it would work out on its own:

"Change the scene from daytime to nighttime."

The instruction names the time of day and nothing else. Here is the result.

Comparison showing a sunny daytime scene above and the same scene edited into a nighttime setting below using Tsubaki.3.

Sunny Daytime Image (Above) And Edited Nighttime Image (Below).

As you can see, Tsubaki.3 handles these connected changes naturally, even with a short and simple edit instruction. The nighttime setting affects more than the overall view. The lighting drops, the shadows shift, the atmosphere changes, and street lights appear along the path.

At the same time, the characters, their positions, and the overall layout stay consistent, so a large change to the environment doesn't redraw the people standing in it. The characters weren't left untouched by the change either. Both pick up the cooler cast of the night lighting, with the lamp glow catching along their shoulders and hair, so they read as standing in the new scene rather than pasted over it.

That last part is what makes the edit work. Tsubaki.3 understood which parts of the scene the new setting should touch and which parts it shouldn't.

One interesting change is that the original sunny image had only one street lamp, and the edited version added several more. Strictly speaking, that strays from the source. But it works here, since an unlit park at night would look wrong, and the model added what the scene needed rather than only what we asked for.

This is often the useful kind of drift. When you would rather hold it back, a short clause such as "change nothing else" keeps the model closer to the original.

Test 4: Typography and Layout in an AI Image Editor With Text Prompts

For our fourth test, we explored how well Tsubaki.3 can edit text within an existing composition without disrupting the surrounding design. We asked the model to replace, add, or modify typography while keeping the original placement, hierarchy, scale, and overall layout intact.

This test involves more than simply generating new words. We want to see whether Tsubaki.3 can treat text as part of the overall design and make localized changes without rebuilding the image unnecessarily. This means preserving details such as spacing, alignment, font style, relative sizing, and the relationship between the text and nearby visual elements.

To test it out, we first created a fictional ID card using one of our existing characters, then used the Tsubaki.3 model to edit a portion of the text within it.

We wanted to change the nationality on the ID card from 'Japanese' to 'American'. For this edit, we used the edit prompt below:

"Edit the nationality on the ID card to 'American' instead of 'Japanese'. Keep the font, color, and size of the text the same."

And this is the result we got.

Before-and-after ID card edit showing the nationality changed from Japanese in the top image to American in the bottom image.

ID Card With Nationality Edited From Japanese (Above) To American (Below).

As you can see from the result, the model accurately changed the text on the ID card to what we wanted. The label column, the colon alignment, the row spacing, the barcode, the signature block, and the hologram all came through untouched, which is the harder half of a typography edit.

Tsubaki.3 didn't rebuild the card to accommodate a new word. It found the one row that needed changing and left the design system around it intact, which is the difference between a text edit and a redesign.

The one drift is weight. "American" came back bold, while every other value on the card stayed regular. The model didn't resize the text block or reflow the layout. It emphasized the single word it had just replaced, as though the edit needed marking.

Our first instinct was to correct it by naming the property directly, so we asked for the new word in regular weight rather than bold. That made things worse. The word stayed bold, and several other values gained weight alongside it, which suggests that putting the word "bold" in the prompt at all was pulling the output toward it.

Simplifying the prompt, as shown below, worked better.

"Change the word 'Japanese' to 'American' on the nationality row. Change nothing else."

This was our result.

Side-by-side comparison of two edited ID cards, with the nationality reading in bold on the left and in regular weight on the right, each with the nationality row enlarged below.

The first attempt returned the word in bold, while a simpler instruction returned it at the same weight as the other values.

With the property left unmentioned, the word came back at the same weight as every other value on the card. The lesson here is the reverse of what you might expect. When a small detail comes out wrong, naming that detail in the retry can reinforce it, and a shorter instruction that simply scopes the edit is often the better fix.

Notably, the layout never wavered across any of these attempts. Even the run that pushed several other values into bold kept the rows, rules, and alignment exactly where they were. What moved was type weight, not structure.

Suppose you are working on a poster, a thumbnail, an advertisement, or a manga cover. Holding the composition while the words change is often the harder requirement, and this test suggests Tsubaki.3 treats a layout as something to work within rather than something to regenerate.

Which Complex AI Image Editing Prompts Did Tsubaki.3 Understand Best?

Across our four tests, Tsubaki.3 performed best when the edit involved a clear visual relationship and the changes could be applied consistently across the image.

The material test was one of the strongest examples. It changed the blue trench coat from cotton to leather while preserving its design, shape, and fit. The leather also responded naturally to the scene's lighting, with changes to its highlights, reflections, and surface appearance.

The scene-wide transformation also produced solid results. When we changed the sunny park scene into a nighttime setting, Tsubaki.3 adjusted the lighting, shadows, atmosphere, and street lights while keeping the characters, their positions, and the main composition intact. The model even added extra street lights to make the nighttime setting feel more natural.

Here are more example images we edited using Tsubaki.3. Instead of nighttime, we edited the earlier park scene to create autumn, winter, and a rainy atmosphere.

Examples of the same park scene edited with Tsubaki.3 to create autumn, winter, and rainy atmospheres using different prompts.

The same park scene edited into autumn, winter, and rainy versions, each from a single short prompt.

The typography and spatial test worked well too. Simply put, Tsubaki.3 showed a strong ability to understand and apply clear visual changes while preserving most of the original image. The main challenges appeared when an edit required very precise control over spatial relationships, poses, or typography details.

Next, we'll look at an example that shows a different kind of limit. In a café scene, Tsubaki.3 handled each part of the instruction accurately, and the result was still wrong, since following the words and understanding the intent behind them aren't the same thing.

When Do Complex AI Image Editing Prompts Need to Be More Explicit?

Across our tests, Tsubaki.3 did better when we were clear about what should change and what should stay. The more relationships an instruction ties together, the more it helps to spell both sides out.

We noticed this especially with scene transformations. We wanted the café scene to take place at night, and our first attempt aimed the instruction at the view through the window.

"Edit the image so it looks like night outside the café."

That wording sent the model after the café itself rather than the time of day. The back wall and counter disappeared entirely, replaced by an open night sky with distant city lights, so the two characters ended up looking like they were sitting outdoors rather than inside. So we tried again, describing the setting instead of the window.

"Edit to make the scene take place at nighttime."

Here is how the two results compare.

Comparison showing a distorted café scene after an initial nighttime edit prompt above and a corrected nighttime scene below after using a clearer, simpler instruction.

Distorted Image (Above) and Corrected Image (Below).

We got a better result by making the instruction more specific and simple: the scene should be set at night, rather than implying that the café's outside should look like night.

Preservation details also matter. When we specify which characters, positions, objects, or design elements shouldn't change, it gives the model a clearer understanding of the full editing intent.

In the café example, we didn't specify anything about preserving the characters, which left more room for unintended changes. This shows that for complex edits, being explicit about both what to change and what to preserve can lead to more consistent results.

What Does This Mean for Real Advanced AI Image Editing Workflows?

What our tests point to is that Tsubaki.3 can work with an image as a whole rather than as a list of objects, and that changes what you can do with a piece you have already finished. Instead of regenerating and hoping the good parts survive, you can revise the one thing that needs changing.

Here is where that fits into everyday creative work:

  • Character Refinement: Clothing, accessories, hairstyles, and expressions can all be updated without the character losing their identity or overall look.
  • Posters and Key Visuals: You can change titles, text, objects, or decorative elements while keeping the original composition and visual hierarchy.
  • Environment Changes: A scene can easily be adapted to a different time, weather, or atmosphere by updating lighting, shadows, reflections, and surrounding details.
  • Concept Art Revisions: Try out new locations, props, or architectural elements while the strongest parts of the original concept stay untouched.
  • Material Exploration: Testing textures, finishes, and surface properties becomes quick, since the object's shape and position hold steady through each version.
  • Manga and Visual Storytelling: Props, backgrounds, expressions, and weather can shift between panels while your characters and settings stay continuous.

Something to keep in mind is that retries are part of the process. Our typography test needed three attempts before the weight settled, and the café scene needed a rewritten instruction before it worked at all.

Also, the solution isn't always a longer prompt. A shorter one fixed our ID card. Naming what to preserve fixed the café scene. Splitting an edit across two passes fixes plenty of others, and when a change has to stay inside one exact region, tools like inpainting give you mask-level control that natural language can't, and some fixes are still faster by hand.

Does Tsubaki.3 Understand Complex AI Image Editing Intent?

After putting Tsubaki.3 through several complex editing tests, we have a clearer picture of where the model performs well and where it can still improve. The short answer is that it understands more than we expected, though not always in the way we phrased the request.

Our tests show that complex AI image editing is about more than making multiple changes. It requires the model to understand relationships, context, and what needs to remain consistent. Tsubaki.3 performed particularly well with material and scene-wide edits, while precise spatial and typography changes were sometimes less consistent.

If you want to test this yourself, skip the color swaps and give Tsubaki.3 something with a dependency in it, an object that has to change hands, a material that has to change without the object changing, or a scene that has to shift time of day.

You can create a free PixAI account and try Tsubaki.3 today.

Top comments (0)