DEV Community

Cover image for Four ways an AI storyboard fails, ranked by how long they take to spot
Naveed W
Naveed W

Posted on

Four ways an AI storyboard fails, ranked by how long they take to spot

A storyboard is closer to a data structure than a gallery.

Each frame has to carry information no other frame carries. The frames have to run in an order that means something. And any one of them has to be replaceable without invalidating the rest. Six anime images sharing a character satisfy none of those constraints.

So rather than walking through my AI anime storyboard in order, I've sorted it by failure class. Eleven scored runs on a six-shot scene, all on Tsubaki.3 inside PixAI, scored out of 5. The classes are ordered by how long each one takes to notice, because that ordering is what decides where you spend your review time.

The setup

The scene: a teenage courier carries a lit brass lantern across the rooftops before a storm reaches it.

Six shots, written as a list before any prompt existed: a wide establishing shot, low tracking on her boots, an insert close-up on the lantern, an over-the-shoulder at the drop, a wide low angle of the jump, and a close shot on the landing.

Two things repeat in every prompt. The character anchors, being her buzzcut, olive windbreaker, red delivery sash and scraped left knee. And the flame, which moves from steady to leaning to thin and finishes as a stub. The flame is the clock, so anyone reading the board knows how much time is left without being told.

One configuration note: aspect ratio is a size setting rather than a prompt instruction. I set Landscape 16:9 in the selector before generating instead of spending prompt tokens on it, which keeps every frame the same shape for free.

Method for every run below: new generation, Landscape 16:9 set in the size selector before prompting.


Class 1: unmade spatial decisions (visible immediately)

The fastest failures to spot, and the most common. Both of my low-scoring frames belong here, and both left a spatial question open that the model then answered badly.

Over the shoulder, 2.5 out of 5

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Over-the-shoulder shot from behind her buzzcut head and left
shoulder, looking forward along the roof ridge. A wide gap opens between this
building and the next, the far ledge sitting lower and to the right. The
lantern is visible in her lowered right hand. First rain streaks in the air.
Enter fullscreen mode Exit fullscreen mode

The framing is correct. Her head blocks the left foreground, the rain reads, the space between the buildings is there.

Then it drew a second version of her down in the alley holding the lantern. Rather than connecting the lantern to the hand of the character whose shoulder we occupy, it instantiated a new figure to hold the object. The roof slope also tilts up almost vertically, which flattens the depth between the foreground shoulder and the ground below.

The rule this implies: an over-the-shoulder frame that asks for both the shoulder and something held further down the same body gives the model two incompatible placements. It resolves the conflict by duplicating the character. Generate that shot twice, or reframe it.

Over the shoulder was the worst frame of the six, at 2.5

Wide establishing, 3 out of 5

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Wide establishing shot from high above a city of tiled rooftops at
dusk. A short stocky teenage girl with a buzzcut, an olive windbreaker and a
red delivery sash runs left to right along a narrow roof ridge, small in
frame, carrying a lit brass lantern in her right hand. A black storm wall
fills the left third of the sky behind her. Rooftops receding into haze.
Enter fullscreen mode Exit fullscreen mode

The storm, the roof ridge and the running pose all arrived, and the lantern glow on the tiles works.

It ignored "small in frame" and placed her centrally at medium distance, so the shot lands between a wide and a medium without committing to either.

The wide establishing shot went the other way.


Class 2: anchor drift (visible only side by side)

Slower to spot, because each frame passes on its own. You find these by laying the board out as a set.

The leap, 4.2 out of 5

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Wide low-angle shot from the alley far below, looking up. She is
mid-air over the gap between buildings, body stretched left to right, olive
windbreaker flared, red sash streaming, the brass lantern thrust out ahead of
her in her right hand. Rain falling hard through the frame. Two roof edges
silhouetted against a pale sky.
Enter fullscreen mode Exit fullscreen mode

The best drawing on the board. Strong silhouette, real depth between the dark roof edges and the pale sky, convincing fabric motion.

It also gave her medium-length flowing hair rather than the buzzcut. The prompt specifies buzzcut. The model traded the character design for a more dramatic motion silhouette, which is a substitution invisible in isolation and obvious in sequence.

The rule this implies: restate every anchor in every prompt, then review the board as a set. Frame-by-frame review will not surface this class at all.

Character Drift Across the Sequence

Class 3: state that doesn't carry (visible only in playback)

Slower still, because it needs you to read the frames in order at speed. The model evaluates each prompt independently, so nothing continuous between frames survives unless you write it in every time.

Reverse angle, 3.6 out of 5

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Reverse angle of the same moment, shot from the far rooftop looking
back towards her. She approaches the camera from the opposite building,
buzzcut, olive windbreaker, red delivery sash, lit brass lantern in her right
hand, the gap between the roofs in the foreground of the frame. First rain in
the air, storm wall behind her.
Enter fullscreen mode Exit fullscreen mode

Camera placement worked, the storm reads behind her, and the lantern glow reflecting off the wet roof does real atmospheric work.

She's walking. Five frames of a girl sprinting ahead of a storm, and the reverse angle has her strolling towards the camera. The drop in the foreground also rendered as a flat ledge rather than a fall between two buildings.

The rule this implies: pace, urgency and physical effort are anchors too. Momentum is not a property the model inherits from adjacent frames, so restate it alongside the costume and the hair.

shw is walking


Class 4: instruction leak (visible on close inspection, easy to shrug off)

Constraints that hold most of the time and not reliably. These are the ones you stop noticing.

Colour leaked past the greyscale instruction 3 times across the board. The landing frame returned warm orange in the flame. The inserted ledge-grab frame came back with a red sash and a brass lantern.

Hand anatomy belongs here too. The lantern insert scored 3.3 with the flame physics read well and wind streaks crossing the panes, but the hand gripping the handle came out inverted, fingers wrapping from the top with no structural logic. The inserted ledge-grab frame scored 3.6 with a convincing drop beneath her, and the gripping hand has 6 fingers.

The rule this implies: treat no-colour and hand geometry as probabilistic rather than guaranteed. On a board this is tolerable, since the frames are disposable. Carry the same assumption into finished art and it will cost you.

Frame 3, lantern insert close-up. Scored 3.3. This is the one with the inverted hand grip.
Frame 3, lantern insert close-up. Scored 3.3. This is the one with the inverted hand grip.

Revision 2, the inserted beat. Scored 3.6. This is the one with six fingers on the gripping hand and the red sash that leaked through the greyscale instruction.
Revision 2, the inserted beat. Scored 3.6. This is the one with six fingers on the gripping hand and the red sash that leaked through the greyscale instruction.


What passed, and the pattern underneath it

The successes cluster as tightly as the failures do.

Low tracking shot, 4.3 out of 5

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Low tracking shot at ankle height, camera moving with her, boots
striking loose roof gravel left to right. Her scraped left knee and the hem of
the olive windbreaker fill the upper frame. The brass lantern swings into the
right of frame, flame leaning backwards. Gravel kicking up.
Enter fullscreen mode Exit fullscreen mode

Highest of the 6. Camera height, boot strike, kicked gravel, scraped knee, jacket hem filling the upper frame and the lantern swinging in from the right all landed. Foreground blur supplies speed while the boot stays sharp and the roof edge behind softens.

Frame 2, low tracking shot. Scored 4.3, highest of the six.
Frame 2, low tracking shot. Scored 4.3, highest of the six.

Landing, 3.9 out of 5

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Close shot, low angle. She lands hard on the far rooftop on one
knee, the scraped left knee down on wet tile, buzzcut soaked, red sash
clinging. Her left hand cups around the mouth of the brass lantern, shielding
a flame reduced to a small stub. Rain bouncing off the tiles around her.
Enter fullscreen mode Exit fullscreen mode

Low angle, soaked buzzcut, knee on wet tile and water bouncing off the roof all work, with the character design matching earlier frames. It placed her left hand on the ground for balance rather than cupping the lantern, substituting the standard landing pose for the requested one.

Frame 6, close shot on the landing. Scored 3.9.
Frame 6, close shot on the landing. Scored 3.9.

The pattern across all 6: the tighter the frame, the better the result. The two close shots and the low angle scored 4.3, 4.2 and 4.1. The wide establishing shot and the over-the-shoulder scored 3 and 2.5.

Framing tightly removes spatial decisions from the model's hands, which is the same root cause as Class 1. An anime storyboard generator does what you specify, and a tight frame leaves it less room to specify anything itself. That's the whole of what anime storyboard AI planning has going for it, and an AI storyboard generator anime creators can plan with is one that keeps that property.


Revision: the property that makes the whole thing usable

A board that cannot be revised has no advantage over a gallery. Two revisions tested that.

Reframe, 4.1 out of 5. A wide establishing opener is a weak start if you want the audience close to her before anything happens, so I swapped it for a face.

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Tight close-up on her face filling the frame, buzzcut, jaw set,
eyes fixed on something off frame right, rain not yet falling. A hard sliver
of storm light across her cheek. The top edge of the brass lantern glows at
the bottom of the frame. Rooftop tiles blurred behind her.
Enter fullscreen mode Exit fullscreen mode

Tight framing, set jaw, gaze off frame right, blurred tiles and lantern glow cropped at the bottom edge all arrived, with the sliver of storm light rendering as a sharp high-contrast patch rather than general shading.

Inserted beat, 3.6 out of 5. The jump read too easily, so I added a frame between the leap and the landing.

Anime storyboard frame, rough greyscale marker sketch, loose confident lines,
no colour. Tight shot on her right hand and forearm slamming onto the wet edge
of the far rooftop, fingers gripping the tile lip, the red sash whipping past.
The brass lantern hangs from her other hand below the ledge, flame nearly out.
Rain running off the roof edge. Nothing else visible except sheer building
wall below.
Enter fullscreen mode Exit fullscreen mode

The ledge grab, the water splashing off the lip, the sash whipping past and the lantern hanging below the roofline all work, and the dark drop underneath sells the height. Its defects are the Class 4 ones listed above.

I’d give the reframe a 4.1.
I’d give the reframe a 4.1.


Two production modes

Generating the whole board in one pass is a different operation from generating frames individually, and the tradeoff is clean.

A six-panel anime storyboard sheet on white paper, arranged as two rows of
three panels, numbered 1 to 6, thin black panel borders with white gutters
between them, rough greyscale marker sketch style, no colour. Panel 1, wide
high shot, a girl with a buzzcut and a red sash running left to right along a
rooftop ridge carrying a lit lantern, storm wall behind. Panel 2, low shot on
her boots hitting roof gravel. Panel 3, close-up on the brass lantern, flame
leaning. Panel 4, over the shoulder looking at a gap between buildings. Panel
5, wide low angle from below, she is mid-air over the gap. Panel 6, close
shot, she lands on one knee shielding the dying flame. Six panels exactly,
numbered in order.
Enter fullscreen mode Exit fullscreen mode

Contact sheet: 3.8 out of 5. Six panels, two rows of three, numbered 1 to 6 with white gutters. The layout is clean and reads as a board immediately, with marker style consistent across all six.

Every individual panel came out less sharp than its single-frame version. Panel 4 dropped the space between buildings. Panel 5 centres her statically rather than showing the horizontal leap. Panel 3's hand grip warped.

Use the sheet to show someone the shape of a sequence. Use single frames when you need the shot itself.

The Whole Board in One Pass

The Whole Board in One Pass

The handoff

A board is a planning asset, so the terminal operation is taking a settled frame and finishing it.

A finished modern anime illustration, full colour, dramatic lighting. Wide
low-angle shot from the alley far below, looking up. A short stocky teenage
girl with a buzzcut, olive windbreaker and red delivery sash is mid-air over
the gap between two buildings, body stretched left to right, sash streaming, a
lit brass lantern thrust out ahead of her in her right hand. Rain falling hard
through the frame, storm sky behind.
Enter fullscreen mode Exit fullscreen mode

4.7, the highest score in the whole test. Warm rim light from the lantern on her face, her sleeve and the wall beside her, against cool storm tones. Building perspective holds, every colour anchor returned, and the buzzcut stayed this time.

That result is not an argument for skipping the board. The rough frames took seconds to read and surfaced two failures the finished render would have buried under good lighting. Once a frame is settled, it goes to illustration, a manga panel, or a starting frame for animation.

Finished illustration of the leap, full colour. Scored 4.7, highest number in the whole test.
Finished illustration of the leap, full colour. Scored 4.7, highest number in the whole test.


Requirements the tool had to satisfy

Four capabilities, all present. A size selector to lock 16:9 before generating so every frame matches. Single-frame generation for the shots that matter. A one-pass contact sheet for the overview. And per-frame regeneration that leaves the rest of the board untouched.

For anime scene planning AI work the last one carries the most load, since a board you can't revise is a gallery with ambitions. The PixAI Studio templates guide covers the Studio workflows available to clone, and the how to use PixAI guide covers the generation panel.

Your next step

Every frame that scored well came from a shot already decided on paper. Every frame that scored badly asked the model to make a spatial decision left open in the prompt. That's the whole finding, and it's the thing an AI storyboard workflow depends on.

To create storyboard with AI tools, write the shot list before opening anything. Six lines of text, one per frame, each naming the camera position and the single thing that happens. Generate, lay them out together, and hunt for the frame where your character changed or stopped moving.

That frame is the note a director would have given you. You got it a week earlier, for the cost of one generation.

Top comments (0)