DEV Community

Micheal Zh
Micheal Zh

Posted on

I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt Engineering for Video

When I started building CineGen, an AI text-to-video generator, I assumed prompt engineering for video would be "image prompting, plus the word moving." It is not. After hundreds of test renders and a lot of embarrassing outputs (a seagull with nine wings remains burned in my memory), I learned that prompting for video is closer to directing a 5-second film than describing a photograph.

Here are the lessons that actually changed the quality of my outputs. No hype, just what works.

1. A video prompt is a timeline, not a painting

The single biggest mindset shift: an image prompt describes one moment. A video prompt describes a sequence of moments. If your prompt only describes a static scene, the model has to invent the motion — and it will invent something weird.

Before (image thinking):

a beautiful sunset over the ocean, cinematic
Enter fullscreen mode Exit fullscreen mode

After (video thinking):

Wide aerial shot of an ocean at sunset. The camera slowly pans left
across the water as waves roll toward the shore. Orange light flickers
on the wave crests. A small sailboat crosses the frame from right to
left. Cinematic lighting, calm mood.
Enter fullscreen mode Exit fullscreen mode

The second prompt works because it answers three questions the model needs: what's in the frame, what's moving, and how does the shot evolve? A structure I keep coming back to is the three-beat arc: establish → action → resolve. Even in a 4-second clip, giving the model a beginning, middle, and end dramatically reduces the "slideshow of random frames" effect.

2. Camera language does half the work

This was the highest-leverage discovery. Generative video models respond strongly to cinematography vocabulary — much more than you'd expect. Naming the shot is often more effective than describing the scene in detail.

A short glossary that covers ~90% of what I use:

  • Shot scale: extreme close-up, close-up, medium shot, wide shot, aerial shot
  • Movement: slow dolly in, pan left/right, tilt up/down, tracking shot, static shot, orbit around
  • Lens feel: shallow depth of field, 35mm, handheld, smooth gimbal motion

Before:

a person walking through a forest
Enter fullscreen mode Exit fullscreen mode

After:

Tracking shot following a hiker from behind on a forest trail,
camera gliding smoothly at walking pace. Tall pine trees blur past
on both sides, morning fog drifting between trunks. Shallow depth
of field, natural light.
Enter fullscreen mode Exit fullscreen mode

The difference is night and day. "Tracking shot" tells the model how the camera behaves, which constrains the motion field and kills a huge class of artifacts where the background slides around unnaturally. If you take one thing from this article: direct the camera, not just the scene.

One caution: don't stack contradictory camera moves. dolly in + pan left + tilt up + orbit in one prompt is asking the model to solve an impossible motion puzzle. One primary camera move per clip.

3. Verbs beat adjectives

In image prompting, adjectives carry the load: beautiful, stunning, ultra-detailed. In video prompting, most adjectives are noise. What the model needs is motion specification — verbs with direction, speed, and rhythm.

Compare:

# Adjective-heavy (weak)
a stunning beautiful waterfall in a gorgeous lush forest, amazing

# Verb-heavy (strong)
Waterfall plunging down a mossy cliff into a pool below, mist
rising and drifting left. Ferns swaying gently in the foreground.
Camera holds a static wide shot.
Enter fullscreen mode Exit fullscreen mode

Notice the second prompt barely uses adjectives, yet produces a far better clip. My rule of thumb: every noun in the prompt should have a verb attached to it. If something is in the frame, say what it's doing — even if it's just "standing still" (which, by the way, is a legitimate and useful instruction: the cat sits perfectly still, only its tail flicks).

Speed words matter too: slowly, gently, rapidly, suddenly. Models genuinely differentiate these. "Walks slowly toward the camera" and "runs toward the camera" produce very different motion — use that dial deliberately.

4. Temporal consistency is the real boss fight

The hardest problem in AI video isn't making pretty frames — it's making frame 1 and frame 48 agree with each other. Faces morph, jackets change color, a coffee cup teleports between hands. Here's what actually helps:

Anchor the subject with specific, repeated attributes. Don't write "a woman"; write "a woman with short black hair in a red jacket." The more specific the anchor, the harder it is for the model to drift. Color anchors (red jacket) work especially well because color is one of the more stable features across frames.

One action per clip. This is the constraint I resisted longest and benefited from most. "She picks up the cup, drinks, sets it down, and waves" will break. "She lifts the cup and takes a sip" works. Complex multi-stage actions across a few seconds are where morphing artifacts breed. Chain short clips instead of cramming everything into one prompt.

Avoid mid-scene transformations. Prompts like "the car transforms into a robot" or "day turns to night" ask the model to do the single hardest thing in generative video: coherent metamorphosis. It will produce something, but it won't be what you pictured. Keep state changes out of the prompt; do them as separate clips and cut between them.

5. Negative prompts earn their keep in video

In image generation, negative prompts are optional polish. In video, they're load-bearing, because video has failure modes that images don't: flickering, morphing, limb duplication during motion, warping geometry.

A negative prompt block I reuse constantly:

negative prompt: morphing face, extra limbs, extra fingers, flickering,
warping background, distorted hands, text, watermark, sudden scene change,
deformed body, disappearing objects
Enter fullscreen mode Exit fullscreen mode

A few notes on this list:

  • flickering and warping background target specifically temporal artifacts — these do almost nothing for still images but matter enormously for video.
  • text, watermark — generated text in video is doubly cursed: not only is it usually gibberish, it writhes between frames. Unless readable text is the point, ban it.
  • sudden scene change suppresses the model's urge to "cut" mid-clip when it gets confused, which reads as a glitch rather than an edit.

Keep the negative list focused. A 40-item negative prompt dilutes into noise; 8–12 targeted terms beat a kitchen sink.

6. A prompt structure that actually works

After all this trial and error, I converged on a template. It's boring, and that's the point — boring is reproducible:

[SHOT] + [SUBJECT + ANCHORS] + [ACTION with verbs] + [ENVIRONMENT]
+ [LIGHTING / MOOD] + [CAMERA MOVEMENT] + [STYLE TAG]
Enter fullscreen mode Exit fullscreen mode

Filled in:

Medium shot of a barista with tied-back brown hair and a denim apron,
pouring steamed milk into a ceramic cup to form latte art. Warm morning
light through a cafe window, dust motes drifting in the sunbeam.
Camera slowly dollies in. Photorealistic, shallow depth of field.

Negative: morphing face, extra fingers, flickering, warping background,
text, watermark
Enter fullscreen mode Exit fullscreen mode

Every slot filled, one camera move, one action, anchored subject, targeted negatives. This structure won't win avant-garde awards, but it produces usable clips at a dramatically higher hit rate than freeform prose. When a render fails, the template also makes debugging easy: you know exactly which slot to change.

7. What reliably fails (so you stop wasting renders)

Some things just don't work yet, regardless of prompt craft. Knowing the walls saves you render credits and frustration:

  • Close-up hands doing precise things. Typing, playing piano, tying shoelaces — finger count and articulation fall apart. Keep hands small in frame or still.
  • Legible text of any kind. Signs, book pages, phone screens. It renders as writhing glyph-soup. Either ban text or keep it tiny and out of focus.
  • Fast, complex multi-object motion. Crowds running, confetti explosions, splashing water with people in it. Motion blur plus multiple agents equals mush. Slow it down or reduce the agent count.
  • Camera and subject moving fast simultaneously. A sprinting subject plus a whip-pan is beyond what current models can keep coherent. Move one or the other, not both.
  • Physics the model hasn't internalized. Liquids pouring, cloth in wind, objects colliding — the rough shape is right, the details are wrong. Wide shots hide this; close-ups expose it.

My workaround for all of these is the same: compose around the weakness. Can't do hands? Frame the shot so hands are out of view. Can't do text? Make the sign out of focus. Prompt engineering isn't just writing better prompts — it's choosing shots the model can actually execute.

8. Treat prompts like code: version and iterate

The workflow that finally made this sustainable:

  1. Draft short. Start with the template, minimal adjectives, one action. Render.
  2. Lock what works. When a render is 80% right, freeze the prompt and change one slot at a time — just the camera move, just the lighting. This is git-diff thinking applied to prompts.
  3. Keep a prompt log. I keep a running file of prompt → result notes. Patterns emerge fast: you'll learn your model's quirks (every model has them) within a few dozen renders.
  4. Vary one axis at a time. Changing the subject, camera, and lighting simultaneously teaches you nothing when the output changes. Scientific method applies.

The unglamorous truth: great AI video comes from iteration discipline, not from one magical prompt. The people getting the best results aren't better writers — they're better experimenters.


That's the honest version of what I learned building CineGen. Video prompting rewards directors, not poets: think in timelines, speak in camera moves, anchor everything, and iterate like an engineer.

If you want to put this into practice without wrestling with local GPU setups, CineGen is the text-to-video tool I built around exactly this workflow — describe your shot, iterate fast, and keep the clips that work. There's a Pro plan at $9.90/month and a $199 lifetime deal if you'd rather not do subscriptions. Either way, go make something weird — the nine-winged seagull era of your prompting journey is waiting.

Top comments (0)