The "two photos in, one rap performance out" generators are everywhere right now. If you're a developer, the product is a curiosity — but the engineering underneath it is a genuinely interesting problem, and it maps onto a skill you'll reuse for a lot more than memes.
This post is a build guide, not an explainer. I'll walk through the architecture, the model choices at each stage, and a minimal prototype that gets you to "a face stays the same across a moving clip" without a PhD in diffusion.
The Hard Problem, Stated Plainly
Forget the orange stage. The core challenge in any personalized video generator is identity preservation: keep a specific face believable across dozens of frames while the body moves, the mouth animates, and the camera shifts.
A generic face can drift and nobody notices. Your user's face has to stay recognizably them for the whole clip, or the joke dies and the product looks broken. This is the axis everything else hangs on.
The Pipeline, as Something You'd Actually Build
A working prototype is five stages. Each one is a replaceable component, which matters because you'll swap models at every layer.
upload ──► face detect/align ──► identity embed ──► motion+scene ──► audio ──► composite/export
1. Face detection and alignment
You can't feed a raw selfie to a diffusion model. First you localize the face, crop it, and normalize it to a canonical orientation.
-
Detection: RetinaFace or SCRFD (both available via
insightface) are the standard picks — fast on CPU, accurate enough. - Alignment: fit five facial landmarks (eyes, nose, mouth corners) and apply a similarity transform to warp the face to a fixed 112×112 or 256×256 template.
from insightface.app import FaceAnalysis
app = FaceAnalysis(name="buffalo_l", providers=["CPUExecutionProvider"])
app.prepare(ctx_id=0, det_size=(640, 640))
img = load_image("person_a.jpg")
faces = app.get(img) # each face has .bbox, .kps, .normed_embedding
The normed_embedding on each detected face is the prize — it's the identity vector you'll use downstream.
2. Identity embedding
This is the step most tutorials skip. You need a compact vector that represents the person, independent of pose, lighting, and expression. Face recognition models trained on large datasets (ArcFace-family backbones) give you exactly that — a 512-dim embedding where "same person" = "high cosine similarity."
Keep the embedding for the source face (the user's photo). In generation, you'll pull the generated face back toward this vector.
3. Motion and scene
Here you have two real architectural choices, and they're not equivalent:
| Approach | How it works | Trade-off |
|---|---|---|
| Face-swap over a template | Generate or reuse a fixed performance clip, then swap in the target face frame-by-frame | Fast, cheap, identity locked by construction; less control over scene |
| Identity-conditioned video diffusion | Drive a video model with the identity embedding (IP-Adapter-style injection or LoRA) | More flexible scenes; harder to keep identity stable, higher compute |
For a two-person rap clip, the first approach is why these tools feel instant: the expensive part (a believable performance with choreography and camera) is solved once, at build time, and reused for every user. You're not generating a scene from scratch per request — you're compositing faces into a pre-rendered one. That's a cost and latency decision, not just a convenience.
If you want bespoke scenes, look at identity-injection via IP-Adapter + AnimateDiff, or a fine-tuned LoRA per recurring character. Budget 10–100× the compute.
4. Audio
Don't synthesize the copyrighted reference track. Generate an original beat plus lyrics and vocals — typically a music/lyrics LLM for the bars, then a TTS or singing-vocals model (Bark, or a commercial text-to-song API) for the performance. Keep audio as a separate pass and mux it at the end; coupling it to the video model only adds failure modes.
5. Composite and export
Stitch frames with ffmpeg, mux the audio, encode for social (vertical, H.264). Nothing exotic here, but get the frame rate and aspect ratio right before you scale.
A Minimal Face-Swap Prototype
If you want something working in a weekend, the face-swap route is the pragmatic choice. The modern standard is insightface's inswapper model:
import cv2, numpy as np
from insightface.app import FaceAnalysis
import onnxruntime as ort
# Load models once
swapper = insightface.model_zoo.get_model("inswapper_128.onnx",
providers=["CPUExecutionProvider"])
source = app.get(load_image("person_a.jpg"))[0] # identity to inject
video = cv2.VideoCapture("template_performance.mp4")
writer = setup_writer("out.mp4", fps=video.get(cv2.CAP_PROP_FPS))
while True:
ok, frame = video.read()
if not ok:
break
targets = app.get(frame)
for t in targets:
frame = swapper.get(frame, t, source, paste_back=True)
writer.write(frame)
writer.release()
This gives you a template performance with a stable face in under a hundred lines. Two people? Run the swap for each of two source identities, or use a template that already has two performers in fixed left/right positions and assign each face to its slot.
Where the Real Complexity Lives
The prototype gets you 80% there. The last 20% is where products are won or lost:
- Occlusion and motion blur. Hands, microphones, head turns — a raw swap breaks down. The good tools add a face-consistency loss or an identity-regularized diffusion pass on top of the swap.
- Multi-face assignment. Keeping face A on the left and face B on the right across 12 seconds, without drift or swap, is a tracking problem as much as a generation problem.
- Latency and cost. A reusable template keeps per-request cost near constant. Generating per user scales your GPU bill with your traffic — the difference between a hobby project and a dead one.
Build It or Use It?
If this is your product, build the pipeline — the identity-preservation problem is a genuine moat worth owning.
If you just need the output — say, a shareable two-person clip for a campaign or a quick test of the format — don't re-solve a solved problem. Something like AI Rap Duo already bundles the template, the identity preservation, and the original beat into a two-photo upload with no prompt. That's the right call when the goal is the video, not the architecture.
Either way, the interesting engineering question is the same one that makes any personalized AI product hard: keeping a real, specific face believable while everything around it moves. Nail that, and the rest is just plumbing.
Top comments (0)