DEV Community

Cover image for Can Tsubaki.3 Put the Camera Where You Want It? 9 Perspective and Composition Tests
Naveed W
Naveed W

Posted on

Can Tsubaki.3 Put the Camera Where You Want It? 9 Perspective and Composition Tests

Most AI images come out framed the same way. The character is centered, the camera is at eye level, and there is not much depth. It looks fine, but it does not tell you whether the model can really put the camera where you want it.

That is what I wanted to find out. So I ran nine tests on Tsubaki.3, each one asking for a specific camera angle or a particular way of arranging the scene, then checked whether it did what I asked.

That is the real test of an AI perspective generator. Not whether the picture looks good, but whether the angle I described is the one it drew. I have grouped the nine by the kind of control each one demands, rather than in the order I ran them.

What composition control means

A good-looking image does not prove the model can control composition. Plenty of models hand you a nice image while quietly ignoring the camera you asked for.

So I judged each test on one simple thing: did the exact setup I described show up. The camera height, the order of the depth layers, the distortion, where each subject lands in the frame.

If I can look at the result and see what I asked for, it passes. If it gives me a nice image that ignored the instruction, it does not. That is the difference between a real AI composition generator and one that just makes pretty variations.

The fundamental camera moves

The first group tests the basics of a controllable camera: height and linear perspective.

Low angle vs high angle

The cleanest camera test is to shoot one subject twice and change only the camera height. I used a lone knight on a battlefield.

BASE
An extreme low-angle shot of a lone knight in worn armor standing on a
battlefield, the camera almost on the ground looking steeply up at her, so she
looms tall against a stormy sky and her boots and legs dominate the foreground
while her head is small and far above. Dramatic clouds behind her, tattered
banner, cinematic modern anime illustration, full body.
Enter fullscreen mode Exit fullscreen mode
EDIT
A steep bird's-eye view looking straight down at the same lone knight in worn
armor standing on a battlefield, seen from high above so the ground fills the
frame, her body foreshortened with her head and shoulders largest and her feet
small beneath her, her shadow stretching across the mud. Cinematic modern anime
illustration.
Enter fullscreen mode Exit fullscreen mode

Both angles landed, and the difference is obvious. In the low shot, the camera is almost on the ground and the knight towers over you, her boots and legs filling the foreground while her head looks small and far above.

In the high shot, the camera looks steeply down, her head and shoulders come out larger than her feet, and her shadow stretches across the mud to sell the height.

The key point is that the perspective changed, not just her pose. The model understood camera height, which is the whole test. The only small miss is that the high shot is a steep overhead rather than a perfectly straight-down bird's-eye.

Test 1-Left the low-angle shot. Right the high-angle shot
Left the low-angle shot. Right the high-angle shot.

Extreme depth down a corridor

This one tests linear perspective: a long corridor where every line rushes toward a single vanishing point.

A view straight down a long neon-lit arcade corridor at night, one-point
perspective with the walls, floor lights, and ceiling signs all rushing toward
a single vanishing point in the far center. A girl stands halfway down the
corridor, small in the middle distance, dwarfed by the tunnel of light
stretching far behind and ahead of her. Strong linear perspective, deep
vanishing point, cinematic modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

The one-point perspective is excellent. The floor, walls, ceiling lights, and signs all converge to a single vanishing point in the center, and the girl is correctly small in the middle distance, dwarfed by the tunnel of light. It looks like a long corridor, not a flat background with repeated shapes.

The one weakness is her face, small and a little muddy at that distance. The prompt made her tiny on purpose, so it is a minor cost. As AI perspective drawing, the depth is the strongest part.


The corridor depth shot. I choose the left one for this review.

Depth and distortion

The next two push on how the model handles a real lens: separated depth layers and true barrel distortion.

Three depth layers

I tested depth with one thing close to the lens, the subject in the middle, one thing far behind, and real size difference between them.

A cinematic anime street scene with strong depth. In the extreme foreground,
close to the lens and slightly out of focus, a paper lantern hangs large on the
left. In the middle distance, a girl in a red kimono walks toward the camera,
sharp and clearly the main subject. Far in the background, a tall pagoda rises
small against the evening sky. Clear size difference between the near lantern,
the midground girl, and the distant pagoda, cinematic depth, detailed modern
anime illustration.
Enter fullscreen mode Exit fullscreen mode

The three layers came out cleanly separated. The paper lantern is huge, soft, and close on the left. The girl in the red kimono is sharp and mid-sized as the main subject. The pagoda is tiny and far behind her.

The scale falloff is what sells it. Each layer is a clearly different distance, in the right order, and the receding street adds even more depth on top. This is strong AI art composition with a real foreground, midground, and background rather than flat stacking.

The three-layer depth shot

Fisheye

Then a harder one: a real fisheye, where straight lines have to bend and the room has to bulge, not just widen.

A fisheye lens view inside a cramped record shop, strong barrel distortion
bending the straight shelves and ceiling into curves around the edges of the
frame, a girl in the center reaching one hand right toward the camera so her
hand looks huge and close while her body curves away small behind it. Rounded,
bulging perspective, wide distorted field of view, detailed modern anime
illustration.
Enter fullscreen mode Exit fullscreen mode

This is the one I expected to fail, and it passed. The shelves visibly curve outward, the ceiling bends around the frame, and the whole room bulges the way a real fisheye looks. Her reaching hand balloons toward the lens while her body curves away small behind it.

This is real barrel distortion, not a plain wide shot, which is the trap most models fall into. For fisheye perspective AI, this is about as convincing as it gets.
The fisheye record-shop shot.

Arranging multiple subjects

Camera height is one thing. Placing several subjects at named positions and holding an angle over all of them is harder.

Complex multi-subject placement

A low-angle cinematic anime shot looking up at a rooftop standoff. In the
immediate foreground on the right, close to the camera, a boy crouches with his
back to us, only his shoulder and the sword in his hand visible large in frame.
In the midground, center, a girl in a black coat stands facing him, sharp and
lit by neon. Behind her in the background, far and small, a second figure
watches from a doorway. The camera is low, looking up past the crouching boy
toward the standing girl. Detailed modern anime illustration, cinematic
composition.
Enter fullscreen mode Exit fullscreen mode

Every subject landed in its assigned spot. The boy is huge in the right foreground with his sword, his back to us. The girl stands centered in the midground, sharp and lit by neon. The second figure watches small from the doorway in back. And the low angle holds across all three, so you look up past the boy toward the girl.

For dynamic camera angle AI work, holding a low angle across a busy multi-subject scene is the hard part, and it held. The model even gave the girl a sword, which fits the standoff. The only slip is that the foreground boy shows more of his body than the "just his shoulder" I asked for, though that makes the scale difference even clearer.

I also ran a treatment change on this shot, keeping the composition and swapping neon night for overcast morning.

Keep this exact composition, camera angle, and the positions of all three
characters identical, but change the time from neon night to bright overcast
morning. Same low angle looking up, same crouching boy large in the right
foreground, same girl centered in the midground, same distant watcher in the
background. Only the lighting, color, and mood change to flat cool daylight.
Detailed modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

The structure survived the lighting change. The boy stayed large in the right foreground, the girl stayed centered, the watcher stayed in the doorway, and the low angle held. Only the light, color, and mood shifted to flat cool daylight.

The model did not rearrange the scene just because the treatment changed, which is exactly what you want. It is not pixel-identical, but the camera and every placement stayed put.

Left the neon-night original. Right the overcast-morning edit.
Left the neon-night original. Right the overcast-morning edit.

Keeping one space consistent as the camera moves

For the next group I built one scene and moved the camera around it through edits. That is harder than generating fresh angles, because the model has to keep the same space consistent as the viewpoint changes.

The base scene

A wide cinematic establishing shot of a grand library atrium at golden hour. In
the foreground, a long wooden reading table runs left to right with an open book
and a green lamp on it. In the midground center, a girl in a navy coat stands at
the base of a tall spiral staircase, looking up. In the background, the
staircase winds up toward a huge arched window flooding the room with warm
light. Eye-level camera, balanced symmetrical composition, deep space from the
near table to the far window, detailed modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

The base gave me clear, stable landmarks. The long wooden table with the book and green lamp is in front, the girl stands at the base of the spiral staircase in the middle, and the staircase winds up to the arched window at the back. The depth runs cleanly from the near table to the far window.

Her face is small and a little muddy again, the recurring weakness at a distance, but the geography is exactly what I needed to test the camera moves against.

The base library scene
The base library scene.

Re-shoot from a low angle

First re-shoot: drop the camera to the foot of the stairs and look up, with the long wooden table now behind the camera and out of frame.

Keep this exact scene, the same library atrium, the same girl in the navy coat,
the same spiral staircase and arched window, but move the camera to a low angle
at the foot of the staircase looking steeply up. The girl is now seen from
below, the spiral staircase twists up dramatically above her toward the same
arched window, and the reading table is now behind the camera and out of frame.
Keep the same warm golden-hour light and the same room layout. Detailed modern
anime illustration.
Enter fullscreen mode Exit fullscreen mode

The camera move worked, but it kept something it should have dropped. The staircase dominates the frame, the camera is clearly low at its base, the arched window is correctly above, and the library is recognizably the same room.

The miss is the long wooden table, which stays visible in the lower-left even though I put it behind the camera. So the model held the room but did not respect what should leave the frame after the move. That is a useful partial pass: strong on the angle and continuity, weaker on precise visibility.

The base library scene (left). The low-angle re-shoot (right).
The base library scene (left). The low-angle re-shoot (right).

The reverse angle

Second re-shoot, the hard one: put the camera behind the girl and look back toward the table that was originally in the foreground.

Keep this exact scene and the same girl in the navy coat, but move the camera
behind her for an over-the-shoulder shot. We now look over her shoulder from
behind, past her, back toward the long reading table with the green lamp and
open book that was in the original foreground, now seen ahead of her across the
room. The spiral staircase and arched window are now behind the camera. Same
library, same warm golden-hour light, same layout, just reversed viewpoint.
Detailed modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

This one impressed me a bit. The model reversed the viewpoint correctly. The table, lamp, and book are now ahead of her across the room, and the staircase and window are gone because they would be behind the camera. The girl looks like she is standing between the camera and the table, so it is a genuine reversal, not just spinning her around in the same forward view.

It reconstructed the side of the room the first shot never showed, which takes real spatial understanding. The only weakness is framing: it came out as more of a rear view than a tight over-the-shoulder.

The base library scene (left). The reverse-angle shot (right).
The base library scene (left). The reverse-angle shot (right).

Cutting a sequence

To finish, I shot one moment as a three-cut close-up sequence, the way a modern anime scene gets cut. A girl in a dark room, lit only by her phone.

The base, an extreme close-up on the eyes

An extreme close-up of a young woman's face in a dark room at night, the frame
filled edge to edge with just her eyes and the bridge of her nose, everything
else cropped out. Her face is lit from below by the cold blue glow of a phone
screen, tiny reflections of text visible in her eyes, a single strand of hair
falling across her forehead. Shallow focus, sharp on the eyes, soft everywhere
else. Intimate, cinematic modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

It committed to the extreme close-up instead of backing off. The frame fills with her eyes and the bridge of her nose, the rest cropped out, and the cold phone glow lights her face from below. The eyes are large and sharp with blue reflections.

The one weakness is the tiny text reflected in her eyes, which came out as distorted glowing marks rather than legible words. At that scale, that is a fidelity limit to note.
The extreme close-up base shot.

First cut, to her hands on the phone

Keep this exact scene, the same girl, the same dark room, the same cold blue
phone light, but cut to a close-up insert of her hands holding the phone at
chest height. We now see her thumbs on the screen and the message glowing on it,
her face soft and out of focus in the background above the phone. Same night
lighting, same blue glow on her fingers, shallow focus on the phone and hands.
Cinematic modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

The match cut worked, but the phone got the logic wrong. The hands and phone became the sharp foreground, her face went soft above them, and the blue glow carried over.

The problem is the phone itself. The rear camera module is clearly visible, so we are looking at the back of the phone, yet the model put the glowing message on that same back surface. That is a real object-logic error, since the screen should be on the front. The text also came out as invented Japanese-style characters.

I gave the exact same prompt to PixAI Edit Pro, and it made the same mistake, rear cameras visible with the message on the back. Edit Pro did render cleaner, legible English text and a more convincing messaging screen, but the physical error was identical. Two different models making the same mistake on the same prompt tells you it is a real blind spot, not a one-off.

(First Edit) Left the Tsubaki.3. Right the same edit on PixAI Edit Pro.
(First Edit) Left the Tsubaki.3. Right the same edit on PixAI Edit Pro.

Second cut, pull back to reveal the room

Keep this exact girl, her same face and expression and the same cold blue phone
light on her, but pull the camera back to a wider cinematic shot that reveals
the whole room for the first time. She is sitting alone on the floor at the foot
of her bed in a messy bedroom late at night, the phone glow the only strong
light, city lights faint through the window behind her, the room dark around
her. Same face, same lighting on her, now small within the wide lonely room.
Cinematic modern anime illustration.
Enter fullscreen mode Exit fullscreen mode

The pull-back nailed the scale change. The camera pulls way back and drops her small into a full bedroom, with the bed, desk, shelves, and scattered books all coherent around her. She is on the floor at the foot of the bed, the phone glow still her main light, and the lonely late-night mood holds.

Her face stays consistent enough to still look like the same girl, even though she is much smaller now. The model built a believable room it had never shown and placed her at the right scale in it.

The pull-back reveal
The pull-back reveal.

Which shots were hardest

As an AI perspective generator, Tsubaki.3 is strong at the big camera moves. Low and high angles, deep three-layer depth, real fisheye distortion, one-point perspective, multi-subject placement, and holding a composition through a lighting change all came out reliable.

It even kept the same room consistent across a low-angle re-shoot and a full reverse angle, which is real AI image perspective control, not luck.

It slips on precision and small detail. It left the table in frame when the camera move should have dropped it. It put a phone screen on the back of a phone. It rendered distorted text in the eye reflection, and faces go muddy when the subject is small or far. So the camera itself is well understood, and the misses land in exact visibility, object logic, and fine detail.

hardest shots

Where this helps creators

For real work, this is useful. If you storyboard, build manga panels, or plan cinematic key art, you can call specific shots, a low hero angle, a deep corridor, a fisheye, a reverse angle, and get them, which is far faster than fighting the model for a non-default composition.

It also works as an AI camera angle generator for concept and perspective references you can draw over.

Where I would still plan for a manual pass: exact control over what stays in or out of frame after a camera move, close-up object details like a phone screen, and clean faces in deep or wide shots.

My take

The big takeaway is simple. When I described the camera instead of just the character, Tsubaki.3 followed. It gave me the angles, the depth, the distortion, and even the reverse shot, which is the hard part.

It does not get everything right. The table that would not leave the frame, the phone screen on the wrong side, the muddy faces in far shots, those are all real limits. But when it comes to controlling how a scene is framed, it does more than I expected.

If you have ever fought an AI model to get one specific angle, Tsubaki.3 is a real step up. Try it on PixAI with one shot you could never quite land, and see how it does.

Top comments (0)