DEV Community

Cover image for AI Animation From Idea to Film: Eight Small Jobs Instead of One Impossible Prompt
Amir Reza Dalir
Amir Reza Dalir

Posted on

AI Animation From Idea to Film: Eight Small Jobs Instead of One Impossible Prompt

You type one prompt. "A man wakes up, walks to the window, and looks at the city. Anime style." You press enter.

And you get a video. A real one. For a moment it feels like magic.

Then you watch it again. The man in the bed has black hair. The man at the window has brown hair. The clock was round, and now it is gone. The floor was wood in the first second and carpet in the third.

It is an animation. It is also garbage. ๐Ÿ—‘๏ธ

Not because the model is bad. Because the model does not know your world โ€” your man, your clock, your room. Every second it draws them again from nothing, and every time a little different. You asked for one thing and got a hundred small guesses, glued together.

Your next idea is a stronger prompt. Describe the man, the clock, the room, every scene, every camera. Try it. The prompt grows to a page, then three, and the hair still changes. A prompt that really pins down one man, one room, three props and six scenes is not three pages. It is a book. No model reads a book and keeps all of it in mind for every frame.

So you stop asking for one thing. You cut the impossible job into eight small jobs, and at every step you hand the model something to copy instead of something to imagine.

๐Ÿ“ The whole method is one folder

An AI animation is many small calls, and after a week you cannot remember which prompt made which image. So every step is one markdown file of prompts next to one folder of results:

story/
โ”œโ”€โ”€ story.md      # 1 ๐ŸŽฌ the six lines, the worlds, characters, props
โ”œโ”€โ”€ assets.md     # 2 ๐ŸŽจ one prompt per world, character, prop
โ”‚                 # 3 ๐ŸŽž๏ธ โ€ฆ and per raw scene
โ”œโ”€โ”€ scenes.md     # 4 ๐Ÿ” one prompt per first/end frame
โ”‚                 # 5 ๐Ÿ‘๏ธ โ€ฆ and per half-screen between scenes
โ”œโ”€โ”€ video.md      # 7 ๐ŸŽฅ one prompt per 4-second clip (6 ๐ŸŽ™๏ธ the sound is in it)
โ”œโ”€โ”€ film.txt      # 8 โœ‚๏ธ the edit: clip order and joins
โ”œโ”€โ”€ cut.sh        # 8 โœ‚๏ธ builds the film from film.txt
โ”œโ”€โ”€ assets/
โ”œโ”€โ”€ scenes/
โ”œโ”€โ”€ video/
โ””โ”€โ”€ final.mp4
Enter fullscreen mode Exit fullscreen mode

Every prompt has the same three-line header, then the prompt:

Header Says
Output the file this prompt makes
Attach which earlier files go with it, and why
Use where the result is needed later

The markdown is the project; the images and clips are its output. Read a .md top to bottom and you see the film before it exists. Change one prompt, make one file again, and nothing else moves. Put the folder in git and you see what you changed last Tuesday.

Prompts in files, not in chat history. The chat is gone next week. The file is not.

๐ŸŽฌ Step 1: A very short story

Before any tool, you need a story. And it must be small โ€” not because small is beautiful, but because AI is bad at keeping things the same.

Limit Max Why
๐ŸŒ Worlds 2 every place has its own light and colour
๐Ÿง Characters 2 each one must look the same in every scene
๐Ÿงฉ Props 3 they repeat, so they must match
๐ŸŽž๏ธ Scenes 6 every scene is a new chance to fail
  • A world is a place. A bedroom, a city street. One inside, one outside.
  • A character has a story behind it. A man, a dog, a robot. The viewer follows it.
  • A prop appears in more than one scene, so it must look the same each time. A clock, a window. Nobody asks what the clock wants โ€” but if it is round in scene 2 and square in scene 5, the video is broken.
  • A scene is one shot. One place, one action. If you need "and then", it is two scenes.

Everything else โ€” the wall, the sheet, the sky โ€” appears once and can change. Do not spend time on it.

Count what repeats. That is what the AI has to get right twice.

The story in this article: a man wakes up, walks to the window, and sees the city. Two worlds, one character, three props, six scenes:

# Scene
1 A man is asleep in his bed
2 The alarm clock on the bedside table rings
3 The man wakes up
4 He gets out of bed and walks to the window
5 He stands at the window and looks out
6 The city, as he sees it

Write your six lines. List and count your worlds, characters and props. No prompts yet. No tools yet.

๐ŸŽจ Step 2: Assets, not scenes

The next thing most people do is open an image model and type scene 1. Then scene 2, and the man is a different man.

Make assets first: one reference image of each world, character and prop, alone, on a plain background. Six assets, six prompts in assets.md. I use GPT Image 2.5 Sunburst, because it can look at attached images โ€” the whole method depends on that.

Four rules for every asset:

  1. ๐ŸŒ The world comes first, and it is empty. No characters, no props, not even the bed. If you draw the bed into the room now, you have two beds โ€” the one in the room and the one in bed.png โ€” and they will not match.
  2. ๐ŸŽจ Everything after attaches the world โ€” not to draw the room again, but so the colours and the light match. A man drawn alone comes out in whatever light the model likes. A man drawn with the room attached belongs in that room.
  3. ๐Ÿšซ No text. No labels, no speech balloons, no watermark. Think of a comic strip with the balloons removed. Words go over the clean image later, where you control them.
  4. ๐Ÿ“ 16:9, all the same size. Even the clock. The video model wants the exact size of the video for its first frame. One odd-sized asset and you find out three steps later.

The first prompt sets the style, and every later prompt copies this paragraph word for word. I will write [STYLE] for it from here on:

Flat cel colour fills, one hard shadow tone and a soft highlight,
crisp dark ink outlines, simplified shapes, limited palette,
hand-drawn anime TV series look. Not photorealistic, not a 3D render.
No text, no labels, no watermark.
Enter fullscreen mode Exit fullscreen mode

๐ŸŒ The world. Empty, with a fixed camera. You choose the camera once, and every scene in this room uses it.

## The room
- Output: assets/room.png
- Attach: nothing
- Use: background, scenes 1 to 5

Anime background illustration, 16:9, no characters, no furniture,
no objects. [STYLE]

A small empty bedroom in the early morning. Plain cream walls, a
wooden floor, a white ceiling. Pale morning light from the far wall
falls across the floor. No bed, no table, no window, no door.

Camera, fixed for this room: standing eye height, from the door,
looking toward the far wall. Left wall and far wall both visible,
the floor filling the bottom third of the frame.
Enter fullscreen mode Exit fullscreen mode

๐Ÿง The character. A sheet, not a picture: at least three views, full body from head to feet, nothing in the hands. In scene 4 the man walks across the room, and the model has to know his feet. If the sheet stops at the chest, the model guesses, and it guesses differently every time.

## The man
- Output: assets/man.png
- Attach: assets/room.png (colours and light only, do not draw the room)
- Use: character reference, every scene

Character reference sheet, 16:9, plain light-grey background.
The exact art style of the attached image. [STYLE]

Character only: no props, nothing in the hands, no background.
Three views side by side, same scale, each the full body from head
to feet, nothing cropped: front, side, three-quarter. Same face,
clothes and colours in all three.

A man around thirty, slim, light skin. Short messy black hair, tired
brown eyes, a day of stubble. Plain white T-shirt, grey pyjama
trousers, bare feet. Front: arms at his sides, eyes half open.
Side: standing straight. Three-quarter: one hand rubbing the back
of his neck, a small yawn.
Enter fullscreen mode Exit fullscreen mode

๐Ÿงฉ The props. The opposite: one view, in full detail. The viewer knows the man by his face. The viewer knows the clock by its bells, its red body, its black numbers. So name every part.

## The clock
- Output: assets/clock.png
- Attach: assets/room.png (colours and light only)
- Use: prop reference, scenes 1 to 3

Prop reference sheet, 16:9, plain light-grey background, no
characters, no hands. The exact art style of the attached image.
[STYLE]

One view only, large in the frame: three-quarter from slightly
above, so the face and the top are both visible.

A round red alarm clock, old style. Red metal body with a soft shine.
Two silver bells on top, a small silver hammer between them, a silver
ring handle behind. White face, black numbers 1 to 12, black hour and
minute hands, a thin red second hand. Two short black legs. No glow,
no digital display.
Enter fullscreen mode Exit fullscreen mode

Do the same for the bed and the window. Then the city, with nothing attached and its own fixed camera: from the window, looking out.

One prompt, one thing, alone. Scenes come later, and they only copy.

This is the cheapest place to be wrong: a bad asset costs one image. Fix the man's hair here and it is right in all six scenes. Open all the assets side by side โ€” same film, same size โ€” and do not move on until they match.

๐ŸŽž๏ธ Step 3: The scenes

Each of the six lines becomes one image. They go in assets/ too, numbered โ€” raw material for the frames in Step 4.

assets/
โ”œโ”€โ”€ room.png โ€ฆ window.png
โ”œโ”€โ”€ 01-asleep.png
โ”œโ”€โ”€ 02-alarm.png
โ”œโ”€โ”€ 03-awake.png
โ”œโ”€โ”€ 04-walk.png
โ”œโ”€โ”€ 05-window.png
โ””โ”€โ”€ 06-city.png
Enter fullscreen mode Exit fullscreen mode

Spend a minute on the names, because the same name travels through every step: 02-alarm.png โ†’ 02-alarm-first.png โ†’ 02-alarm.mp4. Two digits first, so files sort in story order. One word after, the thing the viewer sees. Lowercase, no spaces. When you are twenty files deep and a clip looks wrong, 04-walk tells you which prompt to open. IMG_0417 tells you nothing.

The prompts live in scenes.md. A scene attaches everything that appears in it:

## 02 ยท The alarm
- Output: assets/02-alarm.png
- Attach: assets/room.png ยท assets/man.png ยท assets/bed.png ยท assets/clock.png

Single illustration, 16:9, in the exact 2D anime style of the
attached images. [STYLE]

The attached images are the only source of truth. room.png is the
room: same walls, floor, light and camera. man.png is the man: same
face, hair and clothes. bed.png is the bed and clock.png is the clock,
exactly as drawn. Draw nothing that is not in the attached images or
described below. Only the poses and the action change.

Camera: the fixed camera of the room, from the door.

The bed against the left wall, the man asleep in it, on his side,
eyes closed. The bedside table next to it, the clock on it, ringing:
bells blurred with motion, three small motion lines on each side.
Enter fullscreen mode Exit fullscreen mode

Look at how little of this is about the scene. Style, copied. The camera, copied. A list of what each image is. The action is four lines.

Something I took a while to accept: image models read pictures better than words. Write "a round red alarm clock with two silver bells" and you get a different clock every time. Attach clock.png and say "this clock", and you get that clock. When a scene is hard, do not reach for a longer prompt. Reach for another picture.

Rule Why
โœ… Attach only what appears the city is not in scene 2, so city.png stays out โ€” extra images confuse it
๐Ÿ”ข Attach in order world, then characters, then props โ€” the first image is the base
๐Ÿท๏ธ Name every attachment "room.png is the room" โ€” the model does not know which picture is which
๐Ÿ“ท Same camera as the world five scenes from one camera look like a film; from five cameras, a mess

A scene prompt describes the action. The pictures describe everything else.

๐Ÿ“– The comic-strip test. When all six are done, put them in a row and read them like a comic strip with no words. If the pictures tell the story by themselves, your story and your scenes are right. If you reach one and think "wait, what happened here?", a scene is missing or shows the wrong moment. Do not fix it in the next step. Go back to the six lines, change them, and make that scene again. A hole here becomes a hole in the film.

๐Ÿ” Step 4: Two frames per scene

Here is the tricky part. Look at scene 2. The clock is ringing. Now imagine the clip. Is this image the first frame or the last?

It is the last. The clip starts with a quiet clock, then it rings.

A video model that works from images wants two: where the clip starts and where it ends. So every scene needs two frames, and you already have one. Decide which, then copy it into a new scenes/ folder with the answer as a suffix. The raw scene stays in assets/.

Scene What you have It is Still needed
01 ยท asleep the man asleep, the room still first end: he turns over in his sleep
02 ยท alarm the clock ringing end first: the clock still
03 ยท awake the man sitting up, eyes open end first: eyes closed, head on the pillow
04 ยท walk the man standing by the bed first end: the man at the window, his back to us
05 ยท window the man at the window first end: the same, the curtain moved by the wind
06 ยท city the city, wide first end: the same city, the camera a little closer
scenes/
โ”œโ”€โ”€ 01-asleep-first.png
โ”œโ”€โ”€ 02-alarm-end.png
โ”œโ”€โ”€ 03-awake-end.png
โ”œโ”€โ”€ 04-walk-first.png
โ”œโ”€โ”€ 05-window-first.png
โ””โ”€โ”€ 06-city-first.png
Enter fullscreen mode Exit fullscreen mode

Half the files are missing. To make each one, attach the frame you already have โ€” the finished scene itself โ€” plus only the assets involved in the change. The prompt is tiny, because you describe one difference:

## 04 ยท The walk, end frame
- Output: scenes/04-walk-end.png
- Attach: scenes/04-walk-first.png ยท assets/man.png ยท assets/window.png

Single illustration, 16:9, the same style as the attached scene.
04-walk-first.png is the frame this picture follows: same room,
camera, bed and light. man.png is the man, for his face, hair and
clothes. window.png is the window, exactly as drawn.

One change only: the man has crossed the room. He stands at the
window on the far wall, his back to the camera, one hand on the
curtain. The bed is empty, the blanket pushed back. Everything else
stays exactly where it is.
Enter fullscreen mode Exit fullscreen mode

The man moved, so his sheet is attached again, so his back is right. He touches the window now, so it is attached. The room and the bed come from the scene itself.

You do not describe a scene twice. You describe it once, then describe what changed.

๐Ÿ‘๏ธ Step 5: Half-screens (optional, recommended)

Put the clips in a row and watch. Scene 1 ends with the man turning in his sleep. Scene 2 starts with him still. Same room, but the arm moved, the blanket moved, and your eye catches the jump. Six scenes, five jumps.

My first fix was a clip for the gap itself: from the end of scene 1 to the first frame of scene 2, so nothing would ever cut. I spent a lot of time on this. It does not work. The model has to invent motion between two frames that were never meant to connect, and what it invents is a slow, strange morph. It looks worse than the jump.

Two honest choices:

  • โœ‚๏ธ Leave the cut. Every film is full of cuts. No fade, no effect. Nobody minds.
  • ๐Ÿ‘๏ธ Put a half-screen between them. Think of anime: between two scenes, a short still shot โ€” a character's eyes, a hand, a clock. Almost nothing moves. It holds for a second, then the next scene begins. I call it a half-screen. It makes the cut look like a choice.

A half-screen is one frame, no first and no end. Attach the world for the light, the character or prop it shows, and write a small prompt. Name it after the scene it follows, with -half:

## 02 ยท half-screen, the eyes
- Output: scenes/02-alarm-half.png
- Attach: assets/room.png (light only) ยท assets/man.png

Single illustration, 16:9, the exact style of the attached images.
[STYLE] man.png is the man: same face, hair and stubble.

Extreme close-up of the man's face, filling the frame, on his side on
the pillow, eyes closed. Morning light across his face from the
right. One eyebrow slightly raised, as if the ringing has just
reached him. No bed edge, no clock, no room.
Enter fullscreen mode Exit fullscreen mode
After scene Half-screen Small motion in the clip
01 ยท asleep the clock face, close the second hand ticks
02 ยท alarm the man's closed eyes the eyebrow lifts
03 ยท awake bare feet touching the wooden floor the toes curl
04 ยท walk his hand on the white curtain the curtain sways
05 ยท window his eyes, open, with light in them a slow blink

You do not need all five. Use one where the jump is ugly, a plain cut where it is not.

In the edit, a half-screen dissolves in and out, half a second to a second on each side, over the scene before and the scene after. So a 4-second half-screen shows alone for two to three seconds. The eye is on the close-up while the room changes underneath it โ€” that is what makes the jump disappear.

A half-screen is a cut that looks like it was planned.

๐ŸŽ™๏ธ Step 6: Do not make the voice

You may want to make the sound now โ€” a voice from a voice model, the alarm, the city โ€” and give it to the video model with the frames. You cannot. Not with the two frames.

Seedance 2.5 runs on many platforms. I use it inside ElevenLabs โ€” the same place I would make the voice โ€” and even there, a voice file and a first-and-end frame pair cannot go into the same request. On fal.ai there is no audio input at all. On ByteDance's own API you can attach an audio reference, but the moment you do, the first and end frames stop being first and end โ€” they become loose references, and the clip no longer runs from one to the other. I tried this more than once. Each time I got a good audio file and no place for it.

So the voice goes in the prompt. Write the line in quotes, describe the voice, say who speaks and when. The model renders the voice over the clip, with the mouth on the words, and makes the room sound too โ€” the ring, the sheets, the far city. I have tested this many times. It is not perfect, but it is good, and it sits exactly where the picture needs it, because the same model made both.

He stands at the window and says, in a low, tired voice, a man in
his thirties just awake: "Morning." His mouth moves with the word.
Only he speaks.
Sound: his voice, the curtain, the city far below. No music.
Enter fullscreen mode Exit fullscreen mode

Honest about quality: ElevenLabs' own voice model is better. But I cannot attach it next to my two frames, so it does not matter how good it is. Maybe a future version will take both. Until then, the best voice is the one you can actually put in the clip.

The voice you cannot attach is not a voice. It is a file.

One rule from here: ๐ŸŽต no music in the clips. Music goes over the whole film at the end, in one piece.

๐ŸŽฅ Step 7: The clips

One clip per scene and per half-screen, into video/, prompts in video.md. For a scene, attach the first and the end frame. For a half-screen, the single frame.

Three rules, and they all say keep it short:

Keep short How Why
โฑ๏ธ The clip 4 seconds the model is at its best in short clips; long ones drift
๐Ÿƒ The motion one thing moves the man walks, or the clock rings โ€” not both
โœ๏ธ The prompt a few lines a long prompt makes worse motion, not better

Why 4 seconds and not less? Because Seedance will not go lower. Many moments are shorter than that โ€” a clock starts ringing in one second โ€” but the clip is 4 seconds whether you need them or not. So put the motion in the middle of the clip and leave the first and last second quiet: still at the start, hold at the end. Those quiet seconds are what the edit fades over in Step 8. If the action starts on frame one, the fade eats it.

## 04 ยท The walk
- Output: video/04-walk.mp4
- Attach: scenes/04-walk-first.png (first) ยท scenes/04-walk-end.png (end) ยท 4 s

Image-to-video, 4 seconds, from the first frame to the end frame.
Keep the room, camera, man and bed exactly as drawn.
0-1 s: he stands by the bed, still. 1-3 s: he walks slowly to the
window, bare feet on wood. 3-4 s: he stops, back to us, his hand
reaches for the curtain. Hold the end frame.
Sound: soft footsteps on wood. No music.
Enter fullscreen mode Exit fullscreen mode
## 04 ยท half-screen, the curtain
- Output: video/04-walk-half.mp4
- Attach: scenes/04-walk-half.png (first frame only) ยท 4 s

Image-to-video, 4 seconds, from this single frame; the picture holds
to the end. Small motion only: the curtain sways once, the fingers
tighten on the cloth. No camera move.
Sound: the curtain, the city far away. No music.
Enter fullscreen mode Exit fullscreen mode

The temptation is to add the light, the mood, what the man feels. Every line you add, the model obeys by making the motion worse. Say what moves, say when, say what it sounds like, and stop.

Short clip, one motion, few words. The frames do the talking.

You will remake some clips. When one is wrong, check the frames first โ€” if the two frames do not agree, no prompt saves the clip. If the frames are right, cut the prompt, do not grow it.

โœ‚๏ธ Step 8: Cut them together

I use ffmpeg. It is free, it runs anywhere, and the edit becomes a text file you can read next month.

๐Ÿ“‹ The order. One clip per line, and after each name, how that clip comes in: cut or fade. A scene after a scene is a cut. Anything touching a half-screen is a fade. This file is the edit. Save it as film.txt:

01-asleep        cut
01-asleep-half   fade
02-alarm         fade
02-alarm-half    fade
03-awake         fade
03-awake-half    fade
04-walk          fade
04-walk-half     fade
05-window        fade
05-window-half   fade
06-city          fade
Enter fullscreen mode Exit fullscreen mode

๐Ÿ”— The join. A fade overlaps two clips by FADE seconds, picture and sound together โ€” one second is soft, half a second keeps more of the half-screen on screen. A cut is a one-frame crossfade: invisible, but it removes the click a hard cut leaves in the sound. This script reads the list and builds the chain:

#!/bin/sh
# cut.sh โ€” joins the clips in film.txt into film.mp4
# Every clip is LEN seconds. A fade overlaps two clips by FADE seconds.
set -e
LEN=4; FADE=1
set --
for n in $(awk '{print $1}' film.txt); do set -- "$@" -i "video/$n.mp4"; done
N=$(wc -l < film.txt | tr -d ' ')
FILTER=$(awk -v len="$LEN" -v fade="$FADE" '
  { n++; join[n] = $2 }
  END {
    v = "[0:v]"; a = "[0:a]"; t = len
    for (i = 2; i <= n; i++) {
      d = (join[i] == "fade") ? fade : 0.04
      printf "%s[%d:v]xfade=transition=fade:duration=%s:offset=%.2f[v%d];", v, i-1, d, t - d, i
      printf "%s[%d:a]acrossfade=d=%s[a%d];", a, i-1, d, i
      v = "[v" i "]"; a = "[a" i "]"; t += len - d
    }
  }' film.txt)
ffmpeg -y "$@" -filter_complex "${FILTER%;}" -map "[v$N]" -map "[a$N]" \
  -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac film.mp4
Enter fullscreen mode Exit fullscreen mode

Run sh cut.sh. Eleven clips of 4 seconds, ten fades of one second: 34 seconds of film, no clicks. Change a line in film.txt โ€” swap a fade for a cut, drop a half-screen โ€” and run it again.

๐ŸŽต The finish. A fade from black and to black, then the music over the whole film, quietly under the clips' own sound:

ffmpeg -y -i film.mp4 \
  -vf "fade=t=in:d=0.5,fade=t=out:st=33.5:d=0.5" \
  -af "afade=t=in:d=0.5,afade=t=out:st=33.5:d=0.5" \
  film-faded.mp4

ffmpeg -y -i film-faded.mp4 -i music.mp3 \
  -filter_complex "[1:a]volume=0.2,afade=t=out:st=30:d=4[m];[0:a][m]amix=inputs=2:duration=first" \
  -c:v copy final.mp4
Enter fullscreen mode Exit fullscreen mode

The edit is a text file. Change one line, run it again.

๐Ÿ’ธ What it costs

Two models, and you can count every call before you start. List prices, October 2026: Seedance 2.5 at 720p with sound is $0.473 per second on fal.ai (other platforms differ). OpenAI bills GPT Image 2.5 by tokens and publishes no per-image price; the closest published figure is $0.165 for a 1536ร—1024 high-quality image. Run one call, read the usage field, and put your own number in.

Step Model What Calls Each Cost
2 ๐ŸŽจ GPT Image 2.5 2 worlds, 1 character, 3 props 6 $0.17 $0.99
3 ๐ŸŽž๏ธ GPT Image 2.5 6 raw scenes 6 $0.17 $0.99
4 ๐Ÿ” GPT Image 2.5 6 second frames 6 $0.17 $0.99
5 ๐Ÿ‘๏ธ GPT Image 2.5 5 half-screens 5 $0.17 $0.83
7 ๐ŸŽฅ Seedance 2.5 11 clips ร— 4 s, at $0.473 / s 11 $1.89 $20.81
34 seconds of film 34 $24.61

About twenty-five dollars, if every call comes out right the first time. It will not. Plan for every second call needing a retry, and the film costs closer to forty.

All twenty-three images together cost about as much as two clips. So the expensive mistake is never a bad image โ€” it is a bad frame you only notice once it moves. That is why every step before 7 ends with "do not move on until they match." Spend the cheap calls. Save the expensive ones.

๐Ÿ Your turn

One folder, five text files, and a film at the bottom. When something is wrong โ€” and something will be โ€” you know which file to open.

  • โœ… six lines, two worlds, two characters, three props โ€” counted
  • โœ… the world empty; the character full body in three views; the prop in one detailed view
  • โœ… everything 16:9, the same size, no text
  • โœ… every scene names what it attaches, in order: world, characters, props
  • โœ… every frame is first or end, and its partner is made from it
  • โœ… every clip 4 seconds, one motion, the sound in the prompt

If you have been further than this โ€” a model that takes a voice reference, a better way to bridge two scenes โ€” tell me in the comments. I am still looking for the fix to the one step I could not make work.

Top comments (0)