Most AI video tools respond better to a shot list than to a paragraph of adjectives. Over the last few weeks I have been testing text-to-video generators for short product clips, and the single biggest improvement came from writing prompts the way a camera operator reads a call sheet.
1. One subject, one action
Start every prompt with who or what is on screen and the one thing it does: "a ceramic mug slowly rotating on a wooden table". Two actions in one clip usually produce a muddy result.
2. Name the camera move
Models understand basic film language. "Slow dolly in", "static wide shot" and "handheld follow" each give a very different feel. Pick one per clip.
3. Set light and mood in a few words
"Soft morning window light, calm" beats a long list of style tags. Short, concrete descriptions are easier for the model to keep consistent across frames.
4. Decide the format first
Choose 9:16 for Shorts and Reels, 16:9 for YouTube and landing pages, 1:1 for feeds. Changing the aspect ratio later often means regenerating anyway.
5. Iterate one variable at a time
Keep the subject and action fixed, then change only the camera move or only the lighting. You learn what the model responds to much faster this way.
A reusable template
[subject] [single action], [camera move], [lighting], [mood], [aspect ratio]
Example: "a paper airplane gliding across a classroom, slow tracking shot, warm afternoon light, playful, 16:9".
I have been trying this template with Kling 4.0, a browser-based AI video generator that accepts both text prompts and reference images. Learn more if you want to try the same workflow.
What prompt structure works best for you? I would love to compare notes in the comments.
Top comments (0)