DEV Community

chen mensen
chen mensen

Posted on

I Stopped Asking One Prompt to Build the Whole AI Video

TL;DR

For two-person AI video scenes, I now split the work into two passes: approve a still image first, then animate only that composition. It makes failures easier to diagnose and reduces wasted video generations.

๐ŸŽฌ Why the all-in-one prompt kept failing

I used to put everything into one prompt: two people, clothing, background, camera, props, gestures, and motion. When the result failed, I could not tell whether the problem came from the starting composition or the animation.

Two-person scenes multiply the failure points. A model has to keep two faces separate, preserve left/right positions, render hands, and leave enough empty space for movement.

๐Ÿ–ผ๏ธ Pass one: lock the still frame

I start with a simple scene: two distinct performers, a plain background, one central prop, and soft light. Before moving on, I check:

  • left/right order
  • separation between the performers
  • faces, hands, clothes, and props
  • room for motion
  • the final aspect ratio

If the still is wrong, I regenerate the still. Video credits do not need to be part of that debugging loop.

๐Ÿ•บ Pass two: describe only movement

The second prompt does less. I ask for subtle head movement, alternating hand gestures, natural breathing, and a steady camera. Fast turns, crossed arms, and complicated choreography can wait until the stable version works.

I tested this structure with an orange-studio duo workflow in Hotel Lobby AI. The site separates reference-image preparation from video generation, which fits the way I now debug these clips. It does not include an original song or guarantee exact choreography or lip sync, so I treat audio and final editing as separate steps too.

๐Ÿงช Keep a tiny generation log

For each attempt, I record the model, aspect ratio, duration, reference-image version, motion prompt, and the main defect. Then I change one variable. This is boring in the best way: after a few runs, successful setups become reusable instead of accidental.

Anyone else using a still-first workflow for multi-person image-to-video? What has helped you keep identities stable?

Top comments (0)