Multimodal video models are most useful when a prompt behaves less like a mood board and more like a compact production brief. The goal is not to pack in every visual adjective. It is to tell the model what should remain stable, what should move, how the camera observes it, and what the audience should hear.
I have been testing this structure with MiniMax H3 AI Video Studio, which supports text, image, video, and audio references in the same workflow. The same planning method also transfers well to other reference-driven video systems.
1. Start with an intent sentence
Before writing the full prompt, reduce the shot to one sentence:
A compact travel blender is demonstrated in one continuous commercial shot, ending on the finished drink.
This gives the generation a single job. If the intent sentence contains three locations, several time jumps, or unrelated actions, split it into separate clips.
A useful check is: could a camera crew understand what the finished shot is supposed to prove? For a product clip, the answer might be “the lid locks securely.” For a character clip, it might be “the character recognizes someone and relaxes.”
2. Separate invariants from changes
Reference-driven generation works better when the prompt says which details must not change.
Invariants can include:
- face, hairstyle, wardrobe, or body proportions
- product geometry, logo placement, and button layout
- room layout and lighting direction
- exact UI labels and their position
- the identity of each speaker
Changes describe the action:
- the camera moves from overhead to eye level
- the character turns toward the window
- the interface cards reorganize
- the liquid changes from separate ingredients to a smooth blend
This distinction is more precise than repeatedly asking for “consistency.”
3. Give references explicit roles
A reference should not be treated as a vague inspiration source. State what the model should copy and what it should replace.
For an image reference:
Preserve the character's face, hairstyle, jacket, and color palette from the image. Change the environment to a sunlit train carriage.
For a motion reference:
Use the reference video only for body timing, footwork, and camera rhythm. Replace the performer with the supplied robot character and preserve the robot's exact proportions and surface materials.
For an audio reference:
Keep the pacing and emotional rise of the reference audio, but generate new dialogue and new environmental sound for the scene.
This prevents the model from copying an unwanted background, camera angle, or subject identity along with the useful part of the reference.
4. Describe one camera path
Prompts often fail because they combine a drone shot, macro lens, handheld chase, orbit, and close-up in a few seconds. Pick one primary path.
Examples of camera language that is easy to reason about:
- slow push from medium shot to close-up
- clockwise orbit at eye level
- locked tripod frame with the subject crossing left to right
- overhead opening that lowers smoothly to countertop height
- gentle lateral move past foreground steam
Then add framing and speed. “Slow push, ending in a tight close-up” is more actionable than “cinematic camera movement.”
5. Treat audio as a timeline
“Add realistic sound” is underspecified. List sound layers and connect important effects to visible events.
A useful audio plan has three layers:
- Foreground: dialogue, a button click, footsteps, or an engine start.
- Environment: rain, room tone, traffic, wind, or a kitchen.
- Optional music: genre, energy, and whether it should sit under dialogue.
For stereo placement, simple spatial notes are enough:
Dialogue remains centered. Rain is wide at the sides. Kitchen activity sits softly behind the speakers. The lid click and motor start align exactly with the visible actions.
If music is not needed, say so. Removing a layer can make a short clip feel much more intentional.
6. Put constraints at the end
Constraints are easiest to review when they are grouped together:
Keep the logo centered and correctly spelled. Preserve the appliance shape and button layout. One continuous shot. No extra hands. No background music. 12 seconds.
Do not create a huge negative-prompt inventory. Focus on the few failures that would make the shot unusable.
7. A complete product-demo example
Here is the structure assembled into one prompt:
A pair of hands demonstrates a compact travel blender on a bright kitchen counter. Show three clear actions in one continuous shot: add fruit, lock the lid, then start blending. The camera begins overhead and smoothly lowers to eye level for the blending moment, ending on the finished drink. Preserve the appliance shape, button layout, material colors, and exact brand text from the reference image. Natural daylight, crisp commercial color, controlled reflections, and realistic liquid physics. Stereo audio: fruit pieces dropping into the cup, a precise lid click, the motor ramping up, and a final glass placement. Keep the logo legible. No extra hands, no cuts, no music, 15 seconds.
Notice that each sentence has a job: action, camera, invariants, visual treatment, audio, and constraints.
8. Iterate by changing one layer
When a result is close, do not rewrite everything. Change one layer at a time:
- If identity drifts, strengthen the invariant list.
- If the scene feels chaotic, reduce the number of actions.
- If the camera is unstable, replace several moves with one path.
- If sound feels generic, add event timing and spatial placement.
- If product text changes, specify exact spelling and screen position.
- If a reference dominates too much, narrow its role.
This makes each new generation a test of a specific hypothesis instead of another random attempt.
Final checklist
Before generating, verify that the prompt answers these questions:
- What is the single purpose of the shot?
- Which details must remain unchanged?
- What should be copied from each reference?
- What is the primary camera path?
- What are the beginning and ending actions?
- Which sounds happen at visible moments?
- Which two or three constraints would make or break the result?
A strong multimodal prompt is not necessarily long. It is organized. Once the production intent, references, motion, camera, audio, and constraints each have a clear role, iteration becomes faster and much easier to diagnose.
Top comments (0)