The problem nobody warns you about
You train a model to turn two photos into a short music video. The first three seconds look great. Then, around second five, the guy on the left suddenly has the other guy's jaw. By second eight, their faces have traded places entirely.
If you have built any kind of face-to-video pipeline, you know this failure mode well. It has a name: identity drift. And for a two-person performance video — a rap duet, a duet cover, two friends singing at each other — it is the single hardest thing to get right.
This post is about why identity drift happens, and the engineering decisions you can make to hold two faces steady across a full clip. It is not a product pitch. It is a walk through the failure modes and the levers.
Why one face is already hard
Single-subject face video is a solved-ish problem because the pipeline can lean on strong priors. You detect a face, you extract an identity embedding, and you condition every generated frame on that embedding so the person stays the same. Tools like insightface give you a face detector and an embedding extractor that is fast and stable enough to run per-frame, and models like inswapper do the face swap itself.
The identity embedding is the load-bearing piece. It is a vector — a few hundred floats — that tries to encode "who this person is" while discarding "what expression they are making, what lighting they are in, what angle we are seeing." In practice those two things are not cleanly separable. That is where the drift comes from.
Why two people make it dramatically harder
A duet is not two single-face problems glued together. It is one problem with a new axis of failure: attribution.
Consider what the model has to decide at every frame:
- Which face is which. When both people are in frame, the pipeline must map "left person" and "right person" to the correct embeddings, every frame, even as heads turn and the camera moves.
- Whose motion belongs to whom. A rap duet has the two performers trading lines, leaning toward the microphone, overlapping in the frame. The model has to keep each person's gestures attached to their own identity, not bleed movement from one onto the other.
- Sustaining it over time. Thirty seconds is hundreds of frames. A tiny per-frame error compounds. By the end, two people can have converged toward a single averaged face — the classic "they merged" artifact.
If you have built a single-person pipeline, point 2 is the one that will surprise you. With one face, motion is basically free — any gesture is automatically "the person's." With two faces, motion attribution becomes a hard constraint, not a freebie.
The levers that actually matter
Through building and debugging this, a few decisions made the biggest difference. These are listed roughly in order of impact, not in the order you would build them.
1. Identity conditioning has to be per-person and explicit
Do not pass a single "image" to the generator and hope it figures out there are two people. Pass two separate, clearly tagged identity embeddings, and keep them spatially anchored (left vs. right) for as long as the frame composition allows. When the camera changes composition, re-anchor rather than letting the model guess.
2. Use a pose or layout prior, not just pixels
For a performance video, the two people have a scripted relationship to each other — one is "talking," one is "reacting." Encoding that structure as a pose or layout signal (keypoints, or a depth/segmentation layout) stops the model from inventing plausible-but-wrong interactions. A face-detection + keypoint pass per frame is cheap and removes a whole class of swap artifacts.
3. Constrain the visual style, because style drift feeds identity drift
Here is a non-obvious one. If the scene style is allowed to drift — lighting changes, camera angle wanders, background shifts — the face model gets a worse signal, and identity drift gets worse with it. Locking the style down hard (same color palette, same lighting, same framing for the whole clip) is not just an aesthetic choice. It is an identity-preservation technique. This is one reason a fixed, stylized "stage" look is so much easier to keep consistent than a free-form cinematic scene.
4. Accept that you are trading cost against consistency
Full identity consistency over long clips is expensive. The cheaper the model, the sooner faces start to drift. If you want a "set it and forget it" experience — drop two photos in, get a clip out — you are implicitly making a trade: you spend more compute on identity anchoring so the user does not have to fiddle with prompts or a timeline. Knowing where that trade sits is the difference between a demo that impresses for three seconds and a product that survives a full video.
What "good enough" looks like
You will know your pipeline has crossed the line when you can watch a full clip and never once think about the faces. That is the real acceptance test. The faces just are — two distinct people, clearly themselves from the first frame to the last, even when they trade lines, lean into frame, or sing together.
Reaching that point is less about any single model and more about the scaffolding around it: explicit per-person embeddings, a layout prior, a locked style, and a deliberate cost/consistency tradeoff.
One thing that has nothing to do with the code
A quick reminder before you go build something fun with this: if you are putting a real person's face into a generated video — a friend, a partner, a celebrity, anyone — get their consent first. Face-swap and identity-preservation tech is powerful, and the same pipeline that keeps two faces stable for a music video can just as easily put someone in a video they never agreed to. The technical skill is the easy part. The responsibility is on you.
Have you hit the "they merged into one person" bug? How did you fix it — better embeddings, a pose prior, or something else? I would like to hear what worked in your pipeline.
Top comments (0)