DEV Community

harris
harris

Posted on AI-assisted

How to Use Motion Transfer for Solo, Duo, and Crowd Videos

Motion Transfer uses an existing video for the movement, timing, camera, framing, and edits. Reference images provide the new cast, clothing, or scene. This lets you reuse one recorded performance with different visuals.

While organizing 97 Motion Transfer examples and prompts, I found that reliable results usually come from giving each input one clear job:

Input What it controls
@Video1 Movement order, timing, camera, framing, edits, and character positions
@Image1, @Image2... Face, body, clothing, product, or scene appearance
Prompt Who each image should replace, what should change, and what should stay consistent

Once these jobs are clear, the prompt can often stay short.

@Video1 means the first uploaded video. @Image1 means the first uploaded image, and the number increases with the upload order.

Choose the source video first

The source video sets the ceiling for the final result. Watch it from beginning to end and check:

  • Is the subject visible and the movement easy to read?
  • Do side, back, or full-body shots need extra reference images?
  • Are hand contact, crossing paths, and occlusion clear?
  • Do the number of people, clip length, and aspect ratio fit the target?
  • Do you have the right to use and publish the material?

If you want different movement, timing, or camera work, choose a better source clip first. Motion Transfer works best when the performance already fits the goal.

Prepare images for the people in the video

For a solo scene, start with one clear image of the person. Add matching side, back, or full-body images when the video includes those angles or large movements.

For two-person and crowd scenes, write down who each image should replace:

@Video1 = original performance, camera, and edit
@Image1 = left performer or lead
@Image2 = right performer or surrounding crowd
Enter fullscreen mode Exit fullscreen mode

Give each image one role whenever possible. Multiple images of the same person should keep the face, hair, clothing, and body shape consistent. A scene image should define the environment, lighting, color, and spatial layout.

What the prompt needs to explain

A prompt usually needs to answer three questions:

  1. Who should each image replace?
  2. Which movement, camera work, and positions should come from the source video?
  3. What needs extra protection at the most difficult moments?

A simple scene may need only one sentence. Two-person and crowd scenes need clear left and right positions, lead and supporting roles, and the expected number of people. The English prompts below are copied from the public community records without shortening or grammar edits.

Example 1: Replace one lead character

Urban Music Lead uses one image to replace the lead. The complete prompt in the community record is one sentence:

Swap the main character in the attached video with the attached character
Enter fullscreen mode Exit fullscreen mode

This is the smallest useful setup: one source video, one character image, and one clearly named role. After generation, check for face changes during turns, distortion when a hand passes over the face, grounded feet, and sudden clothing changes.

If the person changes appearance during the clip, add a matching side or full-body image before adding more instructions.

Example 2: Hotel Lobby duo

Hotel Lobby is a 23-second two-person template:

@Video1 = two-person performance, timing, camera, and edit
@Image1 = left performer
@Image2 = right performer
Enter fullscreen mode Exit fullscreen mode

The dedicated page already includes the motion reference and fixed instructions, so the user only needs to upload two character images. The complete prompt in the community record is:

The character from @Image1 @Image2 performs the exact same movements as the persons in @Video1, matching every motion, timing and rhythm smooth grounded movement, consistent lighting, camera and framing follow @Video1, no identity drift, no extra people, the left person should be replaced with @Image1  and right person with @Image2
Enter fullscreen mode Exit fullscreen mode

Review the moments most likely to fail: when the performers move close together, their arms cross, a hand passes over a face, or either person reappears after a cut. A clean opening frame does not prove that both people stay consistent throughout the clip.

View the Hotel Lobby motion reference · View the public source example

Example 3: STORM II lead and crowd

STORM II uses one lead and a group of synchronized performers:

@Video1 = formation, movement, camera, and final bow
@Image1 = foreground lead
@Image2 = appearance of the surrounding crowd
Enter fullscreen mode Exit fullscreen mode

Here, @Image1 should apply only to the lead, while @Image2 should apply to the other performers. The community record includes one image, although the original prompt refers to two character references. The dedicated STORM II page therefore asks for one lead image and one crowd image.

View the complete community prompt

Use the uploaded reference video as the exact motion, choreography, timing, composition, and camera reference.
MAIN CHARACTER:
Replace ONLY the single central main character with the person shown in Character Reference 1.
The central main character must exactly match Character Reference 1 throughout the entire video.
Preserve the exact facial identity, facial features, hairstyle, skin tone, body proportions, clothing, and overall appearance of Character Reference 1.
The central character must remain clearly distinguishable from all surrounding characters at all times.
SURROUNDING CHARACTERS:
Replace ALL surrounding background people with the characters shown in Character Reference 2.
Every surrounding person must consistently match Character Reference 2 in clothing, headwear, mask, colors, body proportions, and overall appearance.
Do NOT apply Character Reference 2 to the central main character.
Do NOT apply Character Reference 1 to any surrounding character.
ACTION AND CHOREOGRAPHY:
Reproduce the exact actions, body movements, choreography, poses, timing, rhythm, and synchronization from the original reference video.
The surrounding crowd must perform the same synchronized group movements and synchronized deep bowing actions shown in the reference video.
The central main character must perform only the movements of the original central character.
Do not invent new gestures, movements, interactions, or choreography.
IDENTITY LOCK:
Keep the identity of the central character completely consistent from the first frame to the final frame.
Do not change, morph, merge, distort, duplicate, or swap the central character's face.
Do not replace the central character with any background character.
Keep all background characters visually consistent throughout the entire sequence.
CROWD CONSISTENCY:
Preserve the original number, placement, spacing, formation, depth, and relative positions of the surrounding people.
Do not add or remove people.
Do not randomly move people between positions.
Do not merge bodies or duplicate characters.
MOTION LOCK:
Strictly preserve the original video's motion.
Preserve the exact body movements, movement direction, choreography, synchronization, pacing, rhythm, and action timing.
CAMERA AND FRAMING:
Keep the original camera angle, camera movement, camera position, focal length, perspective, depth of field, framing, composition, shot scale, and shot duration unchanged.
Follow every original camera movement exactly.
Do not introduce any new camera movement.
Do not re-frame, crop, zoom, rotate, tilt, or alter the original composition unless that exact movement exists in the reference video.
SCENE:
Preserve the original spatial arrangement, crowd formation, foreground-background relationship, and scene geometry.
Maintain natural lighting, realistic shadows, realistic human anatomy, and photorealistic textures.
TEMPORAL CONSISTENCY:
Maintain strong frame-to-frame consistency.
No flickering.
No face changes.
No clothing changes.
No disappearing characters.
No duplicated people.
No warped hands or limbs.
No sudden changes in body proportions.
FINAL REQUIREMENT:
The final video must follow the reference video's original motion, choreography, timing, crowd formation, camera work, and editing as closely as possible, while changing only the visual identities specified by Character Reference 1 and Character Reference 2.
Enter fullscreen mode Exit fullscreen mode

Review the whole formation: Does the lead appear more than once? Do the row count and spacing change? Does anyone disappear after a cut? Do the feet stay grounded? Does the final bow stay synchronized? Looking only at the lead's face will miss most crowd errors.

View the STORM II motion reference · View the public source example

Common problems and first fixes

Problem Common cause First fix
Movement becomes stiff The pose in the image competes with the video movement State that the image provides appearance and @Video1 provides movement
Side or back views become another person The matching angle is missing Add consistent side, back, or full-body images
Two people swap faces The prompt does not say who each image should replace Identify each person by left, right, or another clear position
The lead is duplicated in the crowd Lead and crowd images have unclear jobs State that the lead appears once and the crowd uses only its assigned image
The number of people, formation, or camera changes The prompt lets the model rearrange the frame Preserve the count, positions, framing, camera, and occlusion
Hands, feet, or contact become distorted The source interaction is difficult to read Shorten the clip or choose a source with clearer contact

Change one thing at a time

After each generation, note the timestamp of the clearest failure. Keep the source video and the job of each image unchanged. Replace one image or edit one prompt sentence, then compare the same timestamp again. Keep the strongest result as your next baseline.

You can start with a ready-made template on Motion Transfer or browse the GitHub example library for the full catalog, the job of each reference image, and copyable prompts. The public preview videos show source examples and target effects. Results will vary with the material, model, and settings.

Top comments (0)