DEV Community

Cover image for Testing Tsubaki.3 on Hard Edits: Spatial Reasoning, Materials, Weather, and Typography
Rehan Jamshed
Rehan Jamshed

Posted on

Testing Tsubaki.3 on Hard Edits: Spatial Reasoning, Materials, Weather, and Typography

AI image editing gets interesting the moment an instruction stops being a simple command.

Recolouring hair or deleting an object is easy enough. Real editing requests tend to be more specific than that: move an object into someone's hand without disturbing anything nearby, change what something is made of while its shape stays put, shift a scene into different weather, replace several pieces of text without wrecking the layout.

That's where advanced AI image editing gets hard. The problem stops being object recognition and becomes understanding how objects relate to each other and what ought to follow from a change.

I ran these tests on PixAI's Tsubaki.3 model. Rather than basic object replacement, this AI image editing test covered four harder categories: spatial relationships and object interaction, material and physical-property changes, semantic scene transformation, and typography and layout.

Could Tsubaki.3 grasp the intent behind an edit rather than just the words in it? Let's find out.

Editing workspace
Tsubaki.3 workspace interface featuring an image-guided edit prompt alongside model parameters and generation settings.

What Makes an Image Edit "Complex"?

Length doesn't make a prompt hard.

"Change the jacket to green, make the shoes white, and add a necklace" contains several instructions, but each one stands alone.

Compare that with: "Move the camera from the table into the character's hand, but keep the coffee cup in its original position and make sure her fingers naturally wrap around the camera."

Now several dependencies have to be tracked at once — identify the camera, decide who holds it, distinguish it from the cup, change the spatial relationship, make the interaction believable.

That's the gap between multiple edits and instruction-based image editing.

Complexity can arise from:

  • Relationships between multiple objects
  • Relative position
  • Character-object interaction
  • Foreground and background relationships
  • Material and physical properties
  • Environmental cause and effect
  • Lighting and atmosphere
  • Text hierarchy and placement
  • Several visual consequences caused by one instruction

Which is also why natural language image editing makes a useful test: people describe outcomes, not technical operations. Closing that gap is what these tests were built to check.

Test 1 — Spatial Relationships and Object Interaction

The opening test targeted spatial reasoning. An object had to change location, become associated with a character, and interact naturally with her hand — while a second nearby object stayed exactly put.

Original image:
Before and after comparison
Left: Source. Right: Edited Image.

AI image editing prompt:

"Move the small camera from the right side of the table directly into the character's right hand, wrapping her fingers naturally around it. Keep the coffee cup in front of her on the table untouched, maintaining the original character, seating, and café setting."

What I looked for:

  • Whether the right object moved
  • Whether the cup stayed fixed
  • Whether the fingers gripped the camera convincingly
  • Whether the rest of the composition held steady.

Evaluation: The camera landed in her right hand, and the coffee cup didn't budge. Clean pass on the core spatial swap. Tsubaki.3 picked a natural camera-ready grip with both hands raised toward the viewfinder rather than an awkward one-handed hold.

Café background and lighting stayed stable; the character's pose changed as part of the camera interaction.

Pose is where the model exercised its own judgement. Instead of a low resting hold near the table, she lifted the camera toward eye level as if about to shoot. A reasonable reading of "wrapping her fingers naturally around it" — but it shows the model deciding how an interaction should look, not only where an object belongs.

Test 2 — Change the Material, Keep the Object

Test two asked whether Tsubaki.3 could separate an object's identity from its material, using a hard-shell suitcase whose exterior had to change without the object being redesigned.

Original image:
Before and after comparison
Left: Source. Right: Edited Image.

AI image editing prompt:

"Change the exterior surface of the upright suitcase from smooth plastic to brushed aluminium with realistic horizontal metallic texture and subtle light reflections. Preserve the exact rectangular shape, wheels, handle, character, and train station background."

What I looked for:

  • Whether the silhouette held
  • Whether the wheels and handle stayed intact
  • Whether the metal looked physically plausible
  • Whether the character and background were left alone

Evaluation: The most consistent of the four. Rectangular shape, wheel placement and telescoping handle all held exactly. The surface switched convincingly to brushed metal with horizontal texture and platform-light reflections along the edges.

Character and train-station background were untouched. Tsubaki.3 didn't simply tint the suitcase silver — it added the sort of directional highlight a real aluminium shell picks up under station lighting, which suggests the model connected "material" to how light behaves.

Good evidence for practical AI photo editing with prompts. Material was treated as a property of an object, not a cue to redesign it.

Test 3 — Semantic Scene Editing

Test three moved past any single object. One broad scene-level transformation, and a check on whether the model could infer the visual changes that should logically follow.

Simple concept: turn a dry autumn afternoon into the moment a light rain shower has just begun.

Original image:
Before and after comparison
Left: Source. Right: Edited Image.

AI image editing prompt:

"Transform the dry autumn afternoon into the moment a light rain shower has just begun. Preserve the character, her clothing and features, the street layout, buildings, trees, camera angle and overall composition. Make the resulting scene visually coherent with the new weather and preserve the existing warm shop lighting where appropriate. Do not turn it into a storm or change the season."

What I looked for:

Arguably the most important test, because the prompt described a scene-level transformation without listing every visual consequence.

I wanted to see whether Tsubaki.3 would infer that a rain shower affects more than the sky.

Specifically:

  • Rain appearing naturally throughout the scene
  • Wet or reflective pavement
  • Water-related details such as puddles or ripples
  • Changes to the fallen leaves and surrounding surfaces
  • A believable shift in atmosphere and lighting
  • Preservation of the character, street layout and overall composition

The question was whether those changes would cohere as one weather transformation rather than reading as unrelated additions.

Evaluation: Strong advanced AI image editing here. Without being told to add individual weather effects, Tsubaki.3 turned the dry street into a convincing rainy scene. Rain throughout the image, pavement dark and reflective, puddles and ripples on the ground, fallen leaves matching the wet conditions. Atmosphere shifted cooler and softer while the warm shop lighting stayed visible.

Equally important, the character, buildings, trees, street layout and overall composition all stayed recognisable. The umbrella stayed closed in her hand — consistent with a prompt describing the moment rain had "just begun" rather than instructing her to open it.

The interesting part: most of those changes were never individually specified. Tsubaki.3 had to interpret what a light rain shower looks like across a scene. A stronger test of semantic scene editing than asking for rain, wet pavement and reflections one by one.

Test 4 — Typography and Layout Editing

Test four checked whether Tsubaki.3 could edit text inside an already-designed composition without breaking the layout — a festival poster with a large headline, a subtitle, a character and a small info block.

Original image:
Before and after comparison
Left: Source. Right: Edited Image.

AI image editing prompt:

"Edit the poster typography: replace the main headline with 'MOONLIGHT FESTIVAL' as the dominant title, change the date line below it to '18 OCTOBER 2026', and place a small rounded badge reading 'LIVE' in the top-right corner. Maintain the visual layout, character placement, and ensure text does not cover the character's face."

What I looked for:

  • Text accuracy
  • Whether the headline stayed dominant
  • Whether the date read as secondary
  • Whether the badge stayed subordinate
  • Whether the composition needed redesigning to fit the new text

Evaluation: This one fell slightly short on typographic fidelity. Tsubaki.3 replaced "SUMMER FESTIVAL" with "MOONLIGHT FESTIVAL", placed the updated date beneath, and pinned the rounded "LIVE" badge top-right without disrupting the character or composition. The typography itself drifted from the original design though: the font turned rounder, and the text moved from the original slate-blue treatment to a heavier charcoal. So text replacement, hierarchy and placement were understood — original typographic styling wasn't preserved exactly.

Which Types of Complex Instructions Did Tsubaki.3 Understand Best?

Instruction types

Overall testing result: Semantic scene transformation and material substitution were the most reliable; they maintained the key relationships and dependencies. Spatial interaction was close behind: the camera consistently reached the right hand while the cup stayed fixed, although Tsubaki.3 made its own judgement call about the exact pose. Typography was the least consistent, with the model introducing noticeable changes to the original font style and colour.

Which shows why visual polish alone can't judge complex AI image editing. The strongest outputs weren't necessarily those with the most impressive surface detail, but those where Tsubaki.3 read the relationships in the instruction correctly. Across the tests, the model generally understood what needed to change and what needed to stay connected to the original scene.

When Does a Complex Edit Need More Explicit Instructions?

Natural language carries surprisingly complex intentions, but ambiguity still costs accuracy.

The spatial edit worked from a fairly direct instruction, while the harder parts benefited from stating the dependency outright — keeping the cup in place while moving the camera.

The semantic scene test showed the reverse. A higher-level instruction can suffice when the concept has clear visual consequences. That rain prompt never asked for wet pavement, puddles, reflections or damp leaves, yet Tsubaki.3 produced all of them consistently.

No rigid formula for writing AI image editing prompts falls out of this. Natural language works well when the intended visual relationship is clear; explicit instructions earn their place when a particular detail must be preserved or changed exactly.

What Does This Mean for Real Editing Workflows?

These capabilities translate into practical work: an illustrator moving a prop into a character's hand, a concept artist testing an environment under different weather, a designer previewing an object in another material, a poster designer swapping copy without rebuilding the layout.

The rain test is especially relevant to concept art and environment work, since a creator can describe the atmosphere they want without specifying every resulting surface and lighting change.

That's where an AI image editor with text prompts genuinely saves time. These tests also show why manual editing hasn't become irrelevant, though. Where a relationship needs to be pixel-precise, or text has to be exactly right, build a second pass into the workflow. An imperfect first result may just mean one relationship or detail needs another go.

Final Verdict — Does Tsubaki.3 Understand Complex Editing Intent?

Tsubaki.3's showing here demonstrates why advanced AI image editing needs judging differently from basic object replacement. The question isn't whether an image contains the requested object, colour, material or text — it's whether the model understands why those elements need to change and how they relate to everything around them.

Across these tests, scene-level and object-level relationships were handled particularly well. The strongest results came where the instruction involved a clear visual dependency. Typography exposed a more specific limitation: layout and hierarchy handled well, original font treatment and colour noticeably altered.

For creators, complex AI image editing is already practical for higher-level revisions — provided the result gets a quick check before it ships, particularly where exact text styling or highly precise positioning matters.

To see how far it goes, try your own scenario in PixAI Tsubaki.3: give it an edit depending on two or more elements interacting, rather than asking it to change one object.

That's where the real test of an AI image editor begins.

Top comments (0)