You give an AI video tool a clear portrait of the person you want to see. Then you add a short reference video because you like its movement or camera angle. The clip starts well, but halfway through, the face, clothing, or body shape begins drifting back toward the person in the original video.
It can look as if the main image was ignored. More often, two references are being asked to answer the same question.
A portrait says, “This is the person.” A video can quietly say the same thing through its actor, clothes, posture, framing, and movement. When both files show a visible person, the model has to decide which identity should lead. A prompt such as “replace the person with my character” may not be enough to resolve that conflict.
One Reference Can Carry More Than One Instruction
Imagine you have a portrait of a woman in a mustard-yellow coat. You also have a three-second subway clip of someone else walking past the camera. You want the woman from the portrait to perform the walking motion from the subway clip.
The intention is simple: keep the woman, use the motion.
But the subway clip contains much more than movement. It also includes another person’s face, hairstyle, coat shape, proportions, posture, lighting, and distance from the camera. It is a complete visual scene with its own identity cues.
That is why a result may shift back and forth. The still image and the video clip are supplying competing answers about who should appear on screen.
Give Every Input One Clear Job
Before generating, write one short sentence for each reference file.
For example:
- Portrait image: defines the character’s face, hairstyle, yellow coat, and overall appearance.
- Subway clip: provides only the walking pace, camera distance, and direction of movement.
- Text prompt: defines the new setting and the single action.
- Ending frame, if the chosen generator accepts one: defines where the character should stand at the end of the shot. This will not make every system separate those roles perfectly. But it can reveal a problem early: a motion clip with a clear, visible actor may contain too much identity information to work as a motion-only reference. When that happens, a longer prompt is usually not the first thing to try. Start with a simpler test instead. Run the Smallest Useful Test First Use the portrait image and one plain-text action. Keep the camera still, or give it only one gentle move. For the subway example, the test could be: The woman in the mustard-yellow coat walks forward for three steps in a bright subway station. Medium shot. The camera stays at the same distance. No cut and no other person enters the frame.
If the character changes in this simple version, check the portrait, the scene complexity, or the action before adding a reference video.
If the character stays stable, add the motion video in a second version. Compare the first frame, the middle of the movement, and the final frame. The moment when the character begins to drift is more useful than a general feeling that the full clip is “wrong.”
Ask Whether the Motion Reference Needs a Visible Person
Sometimes a reference video is useful only because of its camera motion. In that case, a clip containing another visible actor may be the wrong source material.
A simpler motion reference could show:
- an empty hallway with a camera moving forward;
- a close shot of footsteps without showing the upper body;
- a silhouette with no recognizable face or clothing;
- a simple object moving across the frame. Reducing unrelated identity information in a motion reference may make it easier for the model to follow the intended character reference. The most useful clip is not always the most detailed-looking one. It is the one that supplies the missing movement instruction without introducing a second subject. Do Not Solve Everything in One Generation A short test works best when it answers one question at a time:
- Can the portrait hold the intended identity during a simple action?
- Can the scene change without changing the person?
- Does the camera movement still work after the motion reference is added?
- Does the final frame match the planned result? If several things fail at once, remove one input instead of adding more instructions. A crowded prompt, two visible people, a scene change, and a moving camera can hide the real source of the conflict. If you are preparing this kind of test in Vidu Q3, treat the first result as a small diagnostic draft rather than a way to make conflicting references automatically agree. Give each reference one clear job, test the smallest useful motion first, and check whether the intended person stays stable before adding more controls. A Practical Check Before You Keep a Result Pause the clip at the beginning, middle, and end. Ask:
- Is the intended person still recognizable?
- Did one action happen clearly from start to finish?
- Did the motion reference add useful movement, or did it pull in the original actor’s appearance?
- Is there one input you can remove without losing the core idea? If the answer to the last question is yes, remove it for the next attempt. A reference image and a reference video can work together, but they should not both be asked to define the same person. Decide which input owns identity, give the other one a narrower role, and test the smallest version of the shot before building anything more complicated. When a motion reference includes another visible person, do you remove it first, crop it down, or keep it and rely on the prompt to separate the roles?
Top comments (0)