DEV Community

Cover image for Controlling Camera Angles With an AI Perspective Generator
Abirami Vina
Abirami Vina

Posted on Originally published at Medium

Controlling Camera Angles With an AI Perspective Generator

We tested Tsubaki.3 as an AI perspective generator across four composition challenges, from low-angle hero shots to fisheye interiors to layered scenes.


Anime illustrators can make the same character read as a hero or a target with nothing but a camera angle change. Drop the camera to ground level and look up, and that character towers over you. Climb above and look down, and the same figure in the same pose becomes small against everything around it.

Thanks to recent tech advancements, AI image generators already understand vocabulary related to camera angles. Low angle, bird's-eye, fisheye, and wide angle are all terms these models have seen. Whether an AI perspective generator can turn those words into the camera you asked for, or just an attractive image that happens to include your character, is what we set out to test.

The trouble is that "dynamic composition" is hard to check. A tilted horizon or some motion blur can make an image look dynamic even when the camera hasn't moved at all. So instead of judging whether a result looks striking, we wrote every prompt around something we could verify about the AI art composition, like a stated camera position, an object closest to the lens, one subject behind another, or straight edges that should bend if the lens is working.

We used Tsubaki.3, the latest model from PixAI, an anime-focused AI art generator, to run every test in this article. Tsubaki.3 is built around instruction control rather than one-off illustrations.

Meet Kaito, the character we tested with. He's a skydiver in his early twenties, generated in a single pass.

An anime skydiver in his early twenties with ash blue hair and goggles on his forehead hangs under an open parachute in an orange and black jumpsuit above a coastline with cliffs and breaking waves.

Kaito, generated in one pass with Tsubaki.3, is the baseline for our camera tests.

We picked the setting as carefully as the character. The cliff edges, clifftop road, and shoreline are all long lines that make the camera angle easy to check. Move the camera up, and they flatten. Move it down, and they rise. Bend the lens, and they curve. The parachute lines help too, fanning out toward the camera.

Over the next few sections, we'll put this AI perspective generator through four tests. We'll change the camera height, build a scene in layers of distance, push the lens into fisheye, and then put several subjects in one frame to see whether each lands where we asked. Let's get started!

Exploring What an AI Perspective Generator Actually Controls

Before testing specific camera instructions, let's take a closer look at what we're grading, because an image can follow a camera instruction and still look wrong, or ignore it completely and still look good.

What we're grading is composition control, which means how much say you have over the way a scene is built rather than what appears in it. Two images can contain the same character, the same outfit, and the same coastline while being completely different pictures, and the difference lives in a handful of decisions.

The camera has a height, an angle, and a distance from the subject. A scene divides into what sits closest to the lens, what sits in the middle, and what sits behind. Each subject lands somewhere specific in the frame; the lens bends what it sees by some amount, and the sizes of those elements relate to each other in a way that tells a viewer which thing is nearer.

An AI perspective generator can produce a striking image while getting most of that wrong, so we wrote every prompt around something we could verify afterward. The question throughout was whether the model built the scene we described or just made an attractive variation on it.

What an AI Perspective Generator Does When You Say Nothing

To see what the model reaches for unprompted, we fed Kaito's baseline image back in as a reference and kept the prompt to just the scene, with no camera words at all.

the boy is hanging under an open parachute, rugged coastline far below,
bright afternoon light, blue sky, anime style, masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode

This was our result.

An anime skydiver in an orange and black jumpsuit hangs centered under an orange parachute canopy at roughly eye level, with a cliff-lined coastline, a winding road, and turquoise sea filling the lower half of the frame

With no camera instruction in the prompt, Tsubaki.3 centered Kaito, held the camera at about his own height, and split the frame evenly between sky and coast.

Kaito sits dead center, the camera holds at his own height, and the canopy stays whole above him. The coast is detailed, with cliffs and a road winding along the clifftop, but it stays under a horizon parked near the middle. It's a good image that makes no spatial decisions.

Then we ran it again in a wide frame, changing nothing else.

A wide shot of an anime skydiver seen from above, with his parachute canopy cropped by the top edge of the frame and a cliff-lined coastline with turquoise water filling almost the entire image below him.

The same prompt in a wide frame lifted the camera above Kaito and tilted it down, pushing the horizon to the top edge and cropping the canopy on both sides.

The horizon climbed to the top edge, the coast went from half the frame to nearly all of it, and the camera rose above Kaito and angled down, cropping the canopy at the corners. He reads smaller now, and the cliffs are seen from above rather than side-on.

Nothing in the prompt or the reference mentioned that. The reference showed him at eye level under a whole canopy, and the model rebuilt the shot anyway rather than widening what it was given. The frame you pick is already a composition decision, before you write a word about the camera.

Test 1: High and Low Shots With an AI Camera Angle Generator

Camera height is the clearest factor of an AI perspective generator to test, because the result is either there or it isn't. Ask for a shot from below, and the near parts of the subject should grow, the far parts should shrink, and the horizon should drop toward the bottom edge. Ask for the reverse, and all three should invert.

We ran both from the same reference, using Kaito's baseline image so the prompts could spend their words on the camera instead of re-describing him.

the boy in @image1 at an extreme low angle shot from directly below him looking
straight up, the underside of the open canopy fills the top of the frame, his boots
closest to the camera, sky behind him, bright afternoon light, anime style,
masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode
extreme high angle shot of the boy in @image1 from the back, the coastline far below
filling most of the frame, bright afternoon light, anime style, masterpiece,
best quality, absurdres
Enter fullscreen mode Exit fullscreen mode

Here's the output we got.

Two anime skydiver images side by side, with a low-angle view showing enlarged boots and the underside of an orange canopy on the left, and a high-angle view from behind showing the character small above a cliff-lined coastline on the right.

The low angle put Kaito's boots nearest the lens with the canopy overhead, while the high angle moved the camera behind and above him to look down at the coast.

Both landed on the first attempt. In the low angle, his boots are enormous, his knees sit nearer the lens than his hips, and his legs compress rather than stretching to full length. The horizon drops to the bottom edge, with the coast reduced to a strip beneath his feet. The model redrew him for a camera under his feet instead of rotating a standing figure.

The high angle inverts every one of those. We're behind and above him, looking at his back and the top of his head, and the coast fills the frame with the horizon pushed to the top corner. His boots are now the farthest thing from the lens rather than the nearest. The canopy left the frame entirely, which is correct, since a camera above him would sit between him and it, and only the lines running upward remain.

One thing arrived unasked for. The coastline in the high-angle shot curves, bowing across the frame the way a wide lens would bend it, and nothing in the prompt mentioned distortion. It reads as height rather than as an error, but it's the model adding a lens characteristic on its own.

Then we pushed the camera further in both directions, first straight overhead and then onto Kaito himself.

Aerial view looking straight down at a rugged coastline from high above, the top
surface of an open parachute canopy seen from above in the upper part of the frame,
the boy in @image1 far below the camera, cliffs and turquoise water filling most of
the frame, bright afternoon light, anime style, masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode
looking down at ocean from the perspective of the boy, bright afternoon light,
anime style, masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode

It turned out better than expected.

Two anime skydiving images side by side, with an overhead view of an orange parachute canopy above cliffs and turquoise water on the left, and a first-person view down past the character's own orange sleeves and legs to the ocean below on the right.

The aerial view showed the top surface of the canopy from above, and the single-line POV prompt moved the camera onto Kaito to look down past his own hands and legs.

The aerial shot got most of the way there. We see the top surface of the canopy rather than its underside; the coast fills the frame, and Kaito reads small beneath it. It isn't quite straight down, since the cliff faces are still partly seen side-on and his face is turned up toward the lens, so the camera sits above and to one side rather than directly overhead.

The POV shot was the surprise. One line with no camera vocabulary in it produced a first-person view looking down past his own hands and forearms, with his legs foreshortened correctly beneath him and the cliff running down the frame. Asked for a perspective rather than a camera position, the model placed the lens at his eyes and worked out what he'd see from there.

It isn't a finished image. There are no risers or canopy lines in view, and his hands are open rather than gripping the toggles, so the pose reads as freefall rather than canopy flight. The camera went where we wanted, and the equipment didn't follow, which is the kind of gap a second pass with the missing details named would close.

Overall, every prompt here ran short because Kaito's appearance was already handled, which left the wording free to describe the shot. When your words aren't competing with a character description, camera instructions carry more weight.

Test 2: Building a Scene in Layers With an AI Composition Generator

Camera height changes where your character sits in a frame. Spatial layering changes how far into the picture you can see, and it's a harder request because the model has to hold various distances in one frame and draw each element at the right size.

So we asked for three things at three distances. Another jumper's canopy against the lens, Kaito in the middle, and the aircraft they both left far behind him. A sport canopy runs around nine meters across and a light aircraft around ten meters long, so if the plane came back anywhere near the size of the near canopy, the layering would fail.

We also avoided the word foreground. Models read it as a mood rather than a position, and the usual result is an object sitting politely whole in a corner. Instead, we described what a close object physically does to a frame, which is to overflow it. Then we ran the prompt twice with the sides swapped, to check whether the structure was real or a lucky default.

the boy in @image1 hanging under his parachute in the middle distance, a second
parachute canopy very close to the camera filling the left side of the frame and
cropped by the frame edge, a small light aircraft far in the distance behind him,
rugged coastline far below, clear afternoon light, anime style, masterpiece,
best quality, absurdres
Enter fullscreen mode Exit fullscreen mode
the boy in @image1 hanging under his parachute in the middle distance, a second
parachute canopy very close to the camera filling the right side of the frame and
cropped by the frame edge, a small light aircraft far in the distance behind him on
the left, rugged coastline far below, clear afternoon light, anime style,
masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode

This is what came back.

Two anime skydiving images side by side, each showing a large parachute canopy cropped by the frame edge, the character under his own canopy in the middle distance, and a tiny aircraft far behind him above a cliff-lined coastline.

The same three distances held when we mirrored the near canopy from left to right, though the camera climbed higher in the first run than the second.

The spatial order is correct in both runs. Two canopies of identical real size came back at wildly different scales, the near one overflowing the frame edge we named, and Kaito's reduced to a small curve above him. Nothing but distance produces that gap. The aircraft holds the far end, and Kaito stays readable in the middle rather than crowded by what sits in front of him.

The coastline survived both times as well. What moved instead was the camera, which climbed above Kaito in the first run and dropped closer to his level in the mirrored one. Neither prompt mentioned camera height.

However, two things gave way. The near canopy in the first run has radial seams fanning from a center point, closer to a parasol than a ram-air parachute. Kaito's own canopy also lost its top edge there, so the element we placed in the middle distance sits partly outside the frame.

Layering held across both runs. Something still has to make room for three distances in one frame, and the camera is what the model moved to find it.

Test 3: Pushing Fisheye Perspective AI to Its Limits

A fisheye angle is a tricky shot to generate because it has a signature. Straight lines bow outward at the middle of each edge, and the corners fall away. A vignette with a wide crop won't produce that.

It also needs something to bend, and open sky has no straight edges. So we put Kaito back in the aircraft at the open door and led the prompt with the lens rather than the character. Then we tested anatomy under the same pressure, with a hand pushed toward the camera.

fisheye lens shot, extreme wide angle, interior of a small aircraft cabin with the
jump door open, curved distortion at the frame edges, straight door frame and window
line bending outward, the boy in @image1 crouched at the open doorway, bright
daylight outside the door, anime style, masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode
the boy in @image1 reaching one hand directly toward the camera, his hand very large
and close to the lens, his head and body small and far behind it, strong perspective
distortion, blue sky and clouds behind him, anime style, masterpiece, best quality,
absurdres
Enter fullscreen mode Exit fullscreen mode

This was what our AI perspective generator returned.

Two anime skydiving images side by side, with a fisheye view inside an aircraft cabin showing curved window frames and an arched jump door on the left, and a character reaching one enlarged hand toward the camera against a cloudy sky on the right.

The cabin came back with real fisheye geometry, and the reaching hand held its scale gap without losing the arm behind it.

The geometry in the cabin came out great. Window frames curve on both walls, the floor plates bow away beneath him, and the rivet lines run as arcs rather than straight rows.

Meanwhile, the door opening arches on all four sides, and the corners fall into shadow the way a wide lens loses its edges. This is distortion built into the scene rather than a crop pretending to be one.

Kaito himself is barely bent, and that's correct. A curved lens distorts long straight edges far more than it distorts a body, so a warped character inside a warped cabin would have been the wrong answer.

The foreshortening test held too. His hand is enormous, his head and torso sit well behind it, and the scale gap reads as distance rather than as a badly sized hand. The blur across the near fingers reinforces it, since that's what a lens does to something inside its focal range.

Test 4: AI Art Composition With Multiple Subjects in One Frame

Next, we wanted to see whether the model could place several subjects rather than simply draw them, because spatial order gets harder to hold as elements add up.

We started with two jumpers, an object, and a stated camera position. The reach makes the front-to-back relationship checkable, since a hand either extends toward Kaito or points at nothing. Then we added a third jumper and named the ground directly to see whether the structure survived more elements.

Here are the prompts we used:

low angle shot from below looking up, the boy in @image1 in freefall closest to the
camera with his arms spread, a second skydiver in a green jumpsuit smaller and
further away above and behind him reaching one hand down toward him, an open orange
windsock on a pole at the left edge of the frame far below them, blue sky and
scattered clouds, anime style, masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode
low angle shot from below looking up at the sky, the boy in @image1 in freefall
closest to the camera with his arms spread, a second skydiver in a green jumpsuit
smaller and further away above and behind him reaching one hand down toward him, a
third skydiver in a white jumpsuit smallest and furthest away at the top right of
the frame, a small red and white striped hot air balloon far below them at the left
edge of the frame, green fields visible far below, blue sky with scattered clouds,
anime style, masterpiece, best quality, absurdres
Enter fullscreen mode Exit fullscreen mode

And the outputs we got.

Two anime skydiving images side by side, with two jumpers reaching toward each other above an orange flag on a pole on the left, and three jumpers at descending scales above a hot air balloon and green fields on the right.

The jumpers landed in the right order in both runs, but the windsock came back as a flag on a pole with no ground beneath it, while naming the fields gave the balloon something to sit above.

The distance order held in both attempts. Kaito is nearest and largest, the green jumper sits above and behind him at a clearly reduced scale, and in the second run the white jumper lands smallest at the top right exactly where we put him.

The reach also connects. The green jumper's hand extends down toward Kaito while Kaito's hand comes up to meet it, and the gap between them reads as distance rather than as a pose aimed at nothing. Adding a third jumper didn't disturb it, so the action stayed readable.

The object is where the first run broke. We asked for a windsock on a pole far below them, and the result put an orange flag at the lower left with open sky underneath it. The position is right, and the shape is wrong, but the real problem is that nothing in the prompt said what sat below, so the model gave a ground-mounted object no ground to stand on.

Naming the fields fixed it. In the second run, the balloon sits small and low over green farmland, tied to a place rather than floating at the same distance as the jumpers. That's the same object type in the same corner of the frame, and the difference is one clause about what's underneath.

Adding elements made the composition better rather than worse. The second prompt gave every item a distance, a side, and something to sit above, and all five constraints landed.

Holding the AI Art Composition Through a Lighting Change

Till now, we've seen the AI perspective generator move the camera on its own whenever we left it unspecified. The open question is what it does to a lens we did specify, so we went back to the fisheye cabin, changed the light to a warm sunset, and left every spatial word alone. Then we ran it again in a wide frame, changing nothing else.

Two fisheye views inside an aircraft cabin at sunset side by side, with a tall frame showing the character crouched in a distant doorway and a wide frame showing him close to the lens between two strongly curved windows.

The sunset landed without costing the lens, and switching to a wide frame produced the strongest curvature in the article while rebuilding the cabin around it.

The lighting landed, and the geometry survived it. Low sun pours through the opening, a long shadow stretches across the floor plates, and the door frame still arches while the corners fall into shadow. A strong lighting instruction doesn't automatically cost you the lens.

The wide frame changed two things, and only one was asked for. The distortion is the strongest yet, with the window frames bowing hard on both walls, the floor curving at its edges, and the corners dropping away much faster than in the tall version.

The composition changed as well. Kaito moved close to the lens with both hands gripping the frame, seats appeared behind him, and the doorway he was crouched inside is gone. The model didn't widen the previous shot. It built a different one to fit the new frame.

Both make sense together. Fisheye distortion grows with distance from the center of the lens, so a tall frame keeps most of the picture near the middle where bending is weakest, while a wide frame gives the same lens more edge to work across. The same rebuilding happened in the control shots, where a wide frame reorganized the whole scene rather than extending the tall one.

So aspect ratio is part of the camera instruction rather than a neutral factor. A tall frame limits how much of a wide lens you actually get, and switching to a wide one won't simply extend what you already had.

The Dynamic Camera Angle AI Instructions That Were Hardest to Follow

Across four tests, our AI perspective generator reached every camera position we asked for, and most of them on the first attempt. What separated the easy requests from the hard ones wasn't a shot's complexity. It was how completely we described the space.

Camera height was the most reliable instruction of all. The low angle, the high angle from behind, and the aerial view each landed once, with the near parts of the figure growing and the horizon moving exactly as each viewpoint demanded. The POV shot was the strongest result in the article, and it came from the shortest prompt we wrote, which asked for a perspective rather than naming a camera at all.

Fisheye was easier than expected. The cabin curved on all four sides on the first attempt and held that curve through a change of lighting. What it couldn't hold was fine anatomy under pressure. The reaching hand kept its arm and its scale gap, but the fingers came back soft and merged, and his far hand arrived small and clawed. Detail leaves the extremities first, so hands are where to look when checking a foreshortened result.

Spatial layering never failed. The order and scale were correct with the near canopy on the left, again when mirrored to the right, and again with a third jumper added.

What moved instead was everything around it. The camera climbed in one layering run and dropped in the other, and the wide frame rebuilt the cabin entirely. Something has to make room for three distances, and the model decides what gives.

Multi-subject scenes also held better than expected, since adding an element improved the result. Both runs got the jumpers in the right order and the reach connected in each. The difference was the object. In the first run, we never said what sat below, so a windsock arrived on a pole with open sky beneath it. In the second, we named green fields, and the balloon settled over them.

So the instructions most likely to fail weren't the crowded ones. They were the ones that left a spatial gap, whether that was a camera height we never mentioned, a frame shape we treated as neutral, or a ground we assumed the model would supply. Wherever we left that gap, it got filled with the most ordinary version of the shot.

When AI Perspective Drawing Becomes Useful for Creators

A dozen or so generations across four tests gave us a reasonable sense of what an AI perspective generator is good for.

A storyboard sheet generated with Tsubaki.3, showing an anime skydiver across eight panels that move from a low angle in the aircraft doorway to a wide shot of him falling above a coastline, a close-up under his open canopy, aerial views of cliffs and water, and first-person shots looking down past his own hands.

Tsubaki.3 works as an AI perspective generator across a range of creative uses, from storyboards to concept art.

Here's an overview of where you can use Tsubaki.3 for getting camera angles right:

  • Storyboards and manga panels: This is the strongest fit. A panel needs the camera in the right place and the subjects in the right order, and both were held across every run we made. Spatial layering never failed once, so a sequence of boards at varying camera heights is well within reach.
  • Concept art and key visuals: The fisheye cabin and the layered freefall scenes both came back usable on the first attempt. If you are establishing a location or a mood before anything gets drawn properly, generated compositions hand you options faster than sketching them one at a time.
  • Action scenes: The camera work holds up well here, and a limb thrust toward the lens keeps its scale and its arm. The fingers are the part that softens under that much compression, so a quick pass over the hands afterward gets you the rest of the way.
  • Perspective references for drawing: These are useful for working out what a scene looks like from a given camera height, and for how near and far elements scale against each other. They are less useful as an AI perspective drawing reference for the figure itself, since a character can come back convincingly framed and still be wrong at the extremities.
  • Multi-character scenes: These outputs are reliable once every element is given a distance and a side. Name what sits below your subjects as well as beside them, and everything lands, which is how our balloon settled neatly over green fields once we said the fields were there.

On the other hand, the weaker areas were narrower than we expected, and most of them come down to what the prompt left unsaid. A camera named without a view gets resolved toward the most familiar shot. Or, a subject named without a position gets placed wherever the model finds room.

Also, aspect ratio is part of the instruction rather than a container for it, so a wide lens needs a wide frame to show what it can do. And a single generation isn't evidence of a model's ceiling, since the same fisheye prompt gave us different degrees of curvature on different runs.

When a specific composition is already clear in your head, a rough thumbnail sketch or a reference photo will often get you there faster than rewording a prompt, and either one can then go in as a reference image to anchor the generation.

If you are new to the platform, the PixAI prompting guide covers the basics, and the same care with naming details applies to character consistency work.

How Much AI Image Perspective Control Tsubaki.3 Gives You

Tsubaki.3 gives you more perspective control than we expected going in.

Camera height was the strongest area. Low, high, aerial, and first-person all landed on the first attempt. Placement was next, holding three distances on the left, mirrored to the right, and again with a third jumper added. Fisheye held too, bowing on all four sides and curving harder in the wide frame.

Complex scenes stayed reliable as long as every element had a distance and a side. Where they slipped was on what we left unsaid, which is how a windsock ended up on a pole with no ground beneath it.

So the model moves well past centered framing. It just won't invent the parts you leave out.

Try it on a composition of your own. Pick a camera position, name what sits closest to the lens and what sits behind it, then check whether each thing landed where you put it. Open Tsubaki.3 in PixAI and see what your own framing survives.

Top comments (0)