<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Micheal Zh</title>
    <description>The latest articles on DEV Community by Micheal Zh (@micheal_zh_e114bbd789e4c3).</description>
    <link>https://dev.to/micheal_zh_e114bbd789e4c3</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146266%2F670cda3d-7b92-4bc1-b347-280f0b65af15.png</url>
      <title>DEV Community: Micheal Zh</title>
      <link>https://dev.to/micheal_zh_e114bbd789e4c3</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/micheal_zh_e114bbd789e4c3"/>
    <language>en</language>
    <item>
      <title>I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt Engineering for Video</title>
      <dc:creator>Micheal Zh</dc:creator>
      <pubDate>Mon, 28 Sep 2026 04:12:20 +0000</pubDate>
      <link>https://dev.to/micheal_zh_e114bbd789e4c3/i-built-an-ai-text-to-video-generator-heres-what-i-learned-about-prompt-engineering-for-video-4gll</link>
      <guid>https://dev.to/micheal_zh_e114bbd789e4c3/i-built-an-ai-text-to-video-generator-heres-what-i-learned-about-prompt-engineering-for-video-4gll</guid>
      <description>&lt;p&gt;When I started building &lt;a href="https://www.cine-gen.com" rel="noopener noreferrer"&gt;CineGen&lt;/a&gt;, an AI text-to-video generator, I assumed prompt engineering for video would be "image prompting, plus the word &lt;em&gt;moving&lt;/em&gt;." It is not. After hundreds of test renders and a lot of embarrassing outputs (a seagull with nine wings remains burned in my memory), I learned that prompting for video is closer to directing a 5-second film than describing a photograph.&lt;/p&gt;

&lt;p&gt;Here are the lessons that actually changed the quality of my outputs. No hype, just what works.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. A video prompt is a timeline, not a painting
&lt;/h2&gt;

&lt;p&gt;The single biggest mindset shift: an image prompt describes &lt;strong&gt;one moment&lt;/strong&gt;. A video prompt describes &lt;strong&gt;a sequence of moments&lt;/strong&gt;. If your prompt only describes a static scene, the model has to invent the motion — and it will invent something weird.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before (image thinking):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;a beautiful sunset over the ocean, cinematic
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;After (video thinking):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Wide aerial shot of an ocean at sunset. The camera slowly pans left
across the water as waves roll toward the shore. Orange light flickers
on the wave crests. A small sailboat crosses the frame from right to
left. Cinematic lighting, calm mood.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second prompt works because it answers three questions the model needs: &lt;strong&gt;what's in the frame, what's moving, and how does the shot evolve?&lt;/strong&gt; A structure I keep coming back to is the three-beat arc: &lt;em&gt;establish → action → resolve&lt;/em&gt;. Even in a 4-second clip, giving the model a beginning, middle, and end dramatically reduces the "slideshow of random frames" effect.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Camera language does half the work
&lt;/h2&gt;

&lt;p&gt;This was the highest-leverage discovery. Generative video models respond strongly to cinematography vocabulary — much more than you'd expect. Naming the shot is often more effective than describing the scene in detail.&lt;/p&gt;

&lt;p&gt;A short glossary that covers ~90% of what I use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Shot scale:&lt;/strong&gt; &lt;code&gt;extreme close-up&lt;/code&gt;, &lt;code&gt;close-up&lt;/code&gt;, &lt;code&gt;medium shot&lt;/code&gt;, &lt;code&gt;wide shot&lt;/code&gt;, &lt;code&gt;aerial shot&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Movement:&lt;/strong&gt; &lt;code&gt;slow dolly in&lt;/code&gt;, &lt;code&gt;pan left/right&lt;/code&gt;, &lt;code&gt;tilt up/down&lt;/code&gt;, &lt;code&gt;tracking shot&lt;/code&gt;, &lt;code&gt;static shot&lt;/code&gt;, &lt;code&gt;orbit around&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lens feel:&lt;/strong&gt; &lt;code&gt;shallow depth of field&lt;/code&gt;, &lt;code&gt;35mm&lt;/code&gt;, &lt;code&gt;handheld&lt;/code&gt;, &lt;code&gt;smooth gimbal motion&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Before:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;a person walking through a forest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;After:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tracking shot following a hiker from behind on a forest trail,
camera gliding smoothly at walking pace. Tall pine trees blur past
on both sides, morning fog drifting between trunks. Shallow depth
of field, natural light.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference is night and day. "Tracking shot" tells the model how the &lt;em&gt;camera&lt;/em&gt; behaves, which constrains the motion field and kills a huge class of artifacts where the background slides around unnaturally. If you take one thing from this article: &lt;strong&gt;direct the camera, not just the scene.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One caution: don't stack contradictory camera moves. &lt;code&gt;dolly in + pan left + tilt up + orbit&lt;/code&gt; in one prompt is asking the model to solve an impossible motion puzzle. One primary camera move per clip.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Verbs beat adjectives
&lt;/h2&gt;

&lt;p&gt;In image prompting, adjectives carry the load: &lt;em&gt;beautiful, stunning, ultra-detailed&lt;/em&gt;. In video prompting, most adjectives are noise. What the model needs is &lt;strong&gt;motion specification&lt;/strong&gt; — verbs with direction, speed, and rhythm.&lt;/p&gt;

&lt;p&gt;Compare:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Adjective-heavy (weak)&lt;/span&gt;
a stunning beautiful waterfall in a gorgeous lush forest, amazing

&lt;span class="gh"&gt;# Verb-heavy (strong)&lt;/span&gt;
Waterfall plunging down a mossy cliff into a pool below, mist
rising and drifting left. Ferns swaying gently in the foreground.
Camera holds a static wide shot.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the second prompt barely uses adjectives, yet produces a far better clip. My rule of thumb: &lt;strong&gt;every noun in the prompt should have a verb attached to it.&lt;/strong&gt; If something is in the frame, say what it's &lt;em&gt;doing&lt;/em&gt; — even if it's just "standing still" (which, by the way, is a legitimate and useful instruction: &lt;code&gt;the cat sits perfectly still, only its tail flicks&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Speed words matter too: &lt;code&gt;slowly&lt;/code&gt;, &lt;code&gt;gently&lt;/code&gt;, &lt;code&gt;rapidly&lt;/code&gt;, &lt;code&gt;suddenly&lt;/code&gt;. Models genuinely differentiate these. "Walks slowly toward the camera" and "runs toward the camera" produce very different motion — use that dial deliberately.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Temporal consistency is the real boss fight
&lt;/h2&gt;

&lt;p&gt;The hardest problem in AI video isn't making pretty frames — it's making frame 1 and frame 48 agree with each other. Faces morph, jackets change color, a coffee cup teleports between hands. Here's what actually helps:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Anchor the subject with specific, repeated attributes.&lt;/strong&gt; Don't write "a woman"; write "a woman with short black hair in a red jacket." The more specific the anchor, the harder it is for the model to drift. Color anchors (&lt;code&gt;red jacket&lt;/code&gt;) work especially well because color is one of the more stable features across frames.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One action per clip.&lt;/strong&gt; This is the constraint I resisted longest and benefited from most. "She picks up the cup, drinks, sets it down, and waves" will break. "She lifts the cup and takes a sip" works. Complex multi-stage actions across a few seconds are where morphing artifacts breed. Chain short clips instead of cramming everything into one prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Avoid mid-scene transformations.&lt;/strong&gt; Prompts like "the car transforms into a robot" or "day turns to night" ask the model to do the single hardest thing in generative video: coherent metamorphosis. It will produce &lt;em&gt;something&lt;/em&gt;, but it won't be what you pictured. Keep state changes out of the prompt; do them as separate clips and cut between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Negative prompts earn their keep in video
&lt;/h2&gt;

&lt;p&gt;In image generation, negative prompts are optional polish. In video, they're load-bearing, because video has failure modes that images don't: flickering, morphing, limb duplication during motion, warping geometry.&lt;/p&gt;

&lt;p&gt;A negative prompt block I reuse constantly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;negative prompt: morphing face, extra limbs, extra fingers, flickering,
warping background, distorted hands, text, watermark, sudden scene change,
deformed body, disappearing objects
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few notes on this list:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;flickering&lt;/code&gt; and &lt;code&gt;warping background&lt;/code&gt; target specifically temporal artifacts — these do almost nothing for still images but matter enormously for video.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;text, watermark&lt;/code&gt; — generated text in video is doubly cursed: not only is it usually gibberish, it &lt;em&gt;writhes&lt;/em&gt; between frames. Unless readable text is the point, ban it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sudden scene change&lt;/code&gt; suppresses the model's urge to "cut" mid-clip when it gets confused, which reads as a glitch rather than an edit.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep the negative list focused. A 40-item negative prompt dilutes into noise; 8–12 targeted terms beat a kitchen sink.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. A prompt structure that actually works
&lt;/h2&gt;

&lt;p&gt;After all this trial and error, I converged on a template. It's boring, and that's the point — boring is reproducible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[SHOT] + [SUBJECT + ANCHORS] + [ACTION with verbs] + [ENVIRONMENT]
+ [LIGHTING / MOOD] + [CAMERA MOVEMENT] + [STYLE TAG]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Filled in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Medium shot of a barista with tied-back brown hair and a denim apron,
pouring steamed milk into a ceramic cup to form latte art. Warm morning
light through a cafe window, dust motes drifting in the sunbeam.
Camera slowly dollies in. Photorealistic, shallow depth of field.

Negative: morphing face, extra fingers, flickering, warping background,
text, watermark
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every slot filled, one camera move, one action, anchored subject, targeted negatives. This structure won't win avant-garde awards, but it produces usable clips at a dramatically higher hit rate than freeform prose. When a render fails, the template also makes debugging easy: you know exactly which slot to change.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. What reliably fails (so you stop wasting renders)
&lt;/h2&gt;

&lt;p&gt;Some things just don't work yet, regardless of prompt craft. Knowing the walls saves you render credits and frustration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Close-up hands doing precise things.&lt;/strong&gt; Typing, playing piano, tying shoelaces — finger count and articulation fall apart. Keep hands small in frame or still.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legible text of any kind.&lt;/strong&gt; Signs, book pages, phone screens. It renders as writhing glyph-soup. Either ban text or keep it tiny and out of focus.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fast, complex multi-object motion.&lt;/strong&gt; Crowds running, confetti explosions, splashing water with people in it. Motion blur plus multiple agents equals mush. Slow it down or reduce the agent count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Camera and subject moving fast simultaneously.&lt;/strong&gt; A sprinting subject &lt;em&gt;plus&lt;/em&gt; a whip-pan is beyond what current models can keep coherent. Move one or the other, not both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Physics the model hasn't internalized.&lt;/strong&gt; Liquids pouring, cloth in wind, objects colliding — the rough shape is right, the details are wrong. Wide shots hide this; close-ups expose it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;My workaround for all of these is the same: &lt;strong&gt;compose around the weakness.&lt;/strong&gt; Can't do hands? Frame the shot so hands are out of view. Can't do text? Make the sign out of focus. Prompt engineering isn't just writing better prompts — it's choosing shots the model can actually execute.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Treat prompts like code: version and iterate
&lt;/h2&gt;

&lt;p&gt;The workflow that finally made this sustainable:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Draft short.&lt;/strong&gt; Start with the template, minimal adjectives, one action. Render.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lock what works.&lt;/strong&gt; When a render is 80% right, freeze the prompt and change &lt;em&gt;one slot at a time&lt;/em&gt; — just the camera move, just the lighting. This is git-diff thinking applied to prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a prompt log.&lt;/strong&gt; I keep a running file of prompt → result notes. Patterns emerge fast: you'll learn &lt;em&gt;your&lt;/em&gt; model's quirks (every model has them) within a few dozen renders.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vary one axis at a time.&lt;/strong&gt; Changing the subject, camera, and lighting simultaneously teaches you nothing when the output changes. Scientific method applies.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The unglamorous truth: great AI video comes from iteration discipline, not from one magical prompt. The people getting the best results aren't better writers — they're better experimenters.&lt;/p&gt;




&lt;p&gt;That's the honest version of what I learned building &lt;a href="https://www.cine-gen.com" rel="noopener noreferrer"&gt;CineGen&lt;/a&gt;. Video prompting rewards directors, not poets: think in timelines, speak in camera moves, anchor everything, and iterate like an engineer.&lt;/p&gt;

&lt;p&gt;If you want to put this into practice without wrestling with local GPU setups, CineGen is the text-to-video tool I built around exactly this workflow — describe your shot, iterate fast, and keep the clips that work. There's a Pro plan at $9.90/month and a $199 lifetime deal if you'd rather not do subscriptions. Either way, go make something weird — the nine-winged seagull era of your prompting journey is waiting.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>tutorial</category>
      <category>video</category>
    </item>
  </channel>
</rss>
