Most people meet a video model with a sentence. "A robot walking through a neon city, cinematic" returns something that moves, but rarely something you can cut into a timeline. The gap is not the model's quality, it is that a prompt is not a shot. Write the shot instead.
Think in shot, subject, motion, camera
A usable prompt has four parts:
- Shot - framing and duration: close-up, medium, wide; one continuous take.
- Subject - who or what, with the two or three details that matter: a courier in a yellow raincoat.
- Motion - what changes across the clip. This is the part people leave out, and it is the part the model needs most.
- Camera - static, slow push in, handheld follow, drone pull back.
"Slow push in on a courier in a yellow raincoat waiting under an awning, rain streaking past the lens, one continuous take" is a prompt you can actually evaluate, because you know what a correct result looks like.
Image-to-video beats text-to-video for control
If the first frame matters, start from a still you generated - or a photograph you own - and let the model do only the motion. You keep composition, lighting and product details under your control, and the model gets a much smaller job. This is also how you keep a product or a face consistent across several clips.
Describe one motion
Two motions in one clip usually produce a smear. Pick the single change you want - a hand reaching, a door opening, steam rising - and let everything else hold still. Movement that the camera also performs makes it harder: either the camera moves, or the subject does, not both aggressively.
Keep them short
Generation quality falls off with duration and complexity. Aim for three to five seconds per clip and assemble the sequence in an editor. Ten short clips you can control beat one long clip you have to apologise for.
Iterate on one variable
Change the motion description, not five words at once. When a generation is close but wrong, the fix is usually a smaller change: slower, closer, one subject fewer. Save the prompts that worked along with the seed, so a re-render reproduces the shot instead of approximating it.
What to avoid in a first pass
- Crowds and busy backgrounds.
- Legible text on signs, labels or shirts.
- Complex hand interaction or overlapping limbs.
- Fast camera moves through occluded space.
Add those after the shot works without them.
Tooling
You can test the shot-description approach without a local pipeline: a browser studio such as Kling 4.0 takes a text or image start and lets you direct the motion, then preview and download the clip. Use it to find the framing and motion that work, then keep the prompts that did.
The habit
Write the shot. Start from a still when composition matters. Move one thing. Keep clips short. The model is a camera operator with no memory of your intent - the shot list is how you give it one.
Top comments (0)