<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jun Heo</title>
    <description>The latest articles on DEV Community by Jun Heo (@heojunfoi).</description>
    <link>https://dev.to/heojunfoi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4142133%2F7069eb23-4362-46ce-bc10-d590fd23b744.png</url>
      <title>DEV Community: Jun Heo</title>
      <link>https://dev.to/heojunfoi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/heojunfoi"/>
    <language>en</language>
    <item>
      <title>I rerolled one AI video prompt 10 times. The prompt was the problem, not the dice.</title>
      <dc:creator>Jun Heo</dc:creator>
      <pubDate>Fri, 25 Sep 2026 02:22:42 +0000</pubDate>
      <link>https://dev.to/heojunfoi/i-rerolled-one-ai-video-prompt-10-times-the-prompt-was-the-problem-not-the-dice-5bbe</link>
      <guid>https://dev.to/heojunfoi/i-rerolled-one-ai-video-prompt-10-times-the-prompt-was-the-problem-not-the-dice-5bbe</guid>
      <description>&lt;p&gt;When an AI video clip comes out wrong, the reflex is to hit generate again. Sometimes that's right: video models are probabilistic, and some takes are just weaker than others.&lt;/p&gt;

&lt;p&gt;But when the same thing goes wrong ten times in a row, it's not the dice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment
&lt;/h2&gt;

&lt;p&gt;I wanted a 5-second vertical clip of a coffee spill in a café. One line:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A barista bumps a customer's elbow and a paper coffee cup slips and spills onto the café floor next to white sneakers, slow motion, handheld close-up, warm morning light.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I generated it 10 times with the same settings (9:16, 5 s).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5m944m9z4m9zhubygrkk.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5m944m9z4m9zhubygrkk.jpg" alt="Before: ten rerolls where the cup falls, rolls or lands upright, but the coffee barely spills. After: ten generations where the coffee splashes onto the sneaker every time" width="800" height="1227"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Top row (before):&lt;/strong&gt; the cup falls, rolls, lands upright. The spill, which was the whole point, barely happens in any of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom row (after):&lt;/strong&gt; the same app and settings, with a rewritten prompt for just the spill. The splash lands on the sneaker in 10 out of 10.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually wrong
&lt;/h2&gt;

&lt;p&gt;Nothing exotic. Two mistakes that show up in almost every failing video prompt:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Two actions fighting in one clip.&lt;/strong&gt; "A barista bumps an elbow" &lt;em&gt;and&lt;/em&gt; "the cup spills." In 5 seconds the model has to stage a bump, a slip, a fall and a splash. It picks some of that and drops the rest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The key action was the last clause.&lt;/strong&gt; The spill came after the barista, the elbow, the cup and the café. Models tend to weight what comes first. Mine kept delivering the setup and skipping the payoff.&lt;/p&gt;

&lt;p&gt;A real film crew would never shoot it that way. The director would break it into shots. The DP would frame the spill tight at floor level with a locked camera, because liquid physics is hard. The script supervisor would flag that one clip is trying to do two things.&lt;/p&gt;

&lt;p&gt;So the rewrite for the spill shot was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Extreme close-up at floor level, 50mm lens, locked-off camera, on clean white leather low-top sneakers standing on a polished grey concrete floor. Black coffee splashes down onto the sneaker toe and spreads across the concrete in a wide splatter, droplets bouncing. 2x slow motion. Warm morning sunlight raking from camera-left, hard highlights on the wet coffee. 35mm film look, fine grain.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;One subject. One action, stated first in its own sentence. One camera instruction. The bump becomes its own shot.&lt;/p&gt;

&lt;p&gt;It isn't perfect. In a few takes the coffee reads a bit syrupy, so the next change would be dropping slow motion to real time. But now I'm fixing texture, not whether the shot happens at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist
&lt;/h2&gt;

&lt;p&gt;These are the rules I now run every video prompt through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One action and one camera move per shot.&lt;/strong&gt; Two actions means two shots.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lead with the action.&lt;/strong&gt; Put the event you need first, with a clear start and end.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Describe the frame, not the feeling.&lt;/strong&gt; "Tense" isn't something a camera can see. "Jaw clenched, eyes on the door" is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name a lens and a light source.&lt;/strong&gt; "50mm, window light from camera-left" does more than "cinematic."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frame hard physics tight.&lt;/strong&gt; Pours, splashes, hands and small objects: close-up, simple background, locked camera.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Paste character descriptions verbatim in every shot.&lt;/strong&gt; Paraphrasing a character gives you a different person.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No readable text in frame.&lt;/strong&gt; Add it in the edit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delete empty quality words.&lt;/strong&gt; "8k, masterpiece, stunning" push out words that help.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;After a bad take, change one thing.&lt;/strong&gt; If you change five things you learn nothing from the reroll.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  I turned it into an agent skill
&lt;/h2&gt;

&lt;p&gt;Running that checklist by hand gets tedious, so I packaged it as an open-source agent skill: &lt;a href="https://github.com/HEOJUNFO/ai-film-crew" rel="noopener noreferrer"&gt;&lt;strong&gt;ai-film-crew&lt;/strong&gt;&lt;/a&gt;. It runs your idea past seven roles, each with one job:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Catches&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Director&lt;/td&gt;
&lt;td&gt;mood-only beats, slow openings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Production designer&lt;/td&gt;
&lt;td&gt;characters drifting between shots (writes a continuity bible)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DP&lt;/td&gt;
&lt;td&gt;combined camera moves that warp geometry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gaffer&lt;/td&gt;
&lt;td&gt;flat, evenly lit "AI look"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Editor&lt;/td&gt;
&lt;td&gt;shots longer than the model can generate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sound&lt;/td&gt;
&lt;td&gt;audio cues sent to silent models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Script supervisor&lt;/td&gt;
&lt;td&gt;continuity, feasibility, slop; can veto any shot&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It has three modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Plan:&lt;/strong&gt; idea in, shot list and one prompt per shot out&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fix:&lt;/strong&gt; paste a failing prompt, get the problems ranked and a rewrite&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review:&lt;/strong&gt; describe a bad clip, get the one change to make before rerolling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It works with Wan, LTX, HunyuanVideo, Kling, Veo, Seedance, Hailuo and Runway, and it runs in Claude Code, Codex, Cursor or anything else that reads &lt;code&gt;SKILL.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add HEOJUNFO/ai-film-crew
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then just ask for a video:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;gt; 15s vertical ad for a handmade ceramic mug, for Kling
&amp;gt; this prompt keeps failing: "a chef flips a pancake in slow motion, cinematic, 8k"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The coffee example and a full Kling ad plan are in the repo's &lt;code&gt;examples/&lt;/code&gt; folder.&lt;/p&gt;

&lt;h2&gt;
  
  
  Disclosure
&lt;/h2&gt;

&lt;p&gt;I work on &lt;a href="https://share.ludyte.com/junfoi" rel="noopener noreferrer"&gt;Ludyte&lt;/a&gt;, an AI video app, and both rows of clips above were generated there. The skill is MIT-licensed and model-agnostic; nothing in it depends on Ludyte.&lt;/p&gt;

&lt;p&gt;If an adapter for your model is out of date, or you have a before/after that proves a rule wrong, PRs are welcome. The per-model advice goes stale fast.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>video</category>
      <category>claude</category>
    </item>
  </channel>
</rss>
