An AI video edit can look convincing in one frame and fall apart a second later. For character replacement, the useful question is whether the new appearance holds up while the original performance continues.
Here is a small case study from a project I work on, followed by a review process developers can use when evaluating similar outputs.
Record the inputs before judging the output
The documented test used:
- One eight-second presenter clip.
- One reference photo.
- Kling 3.0 Omni via KIE.
- 720p output, with source audio enabled.
- Reference-image preparation disabled.
The resulting file was 1280 × 720 and approximately 8.04 seconds long.
Keeping these details matters. A different model, reference, or preprocessing step introduces another variable. If you change several inputs together, it becomes difficult to explain why the result improved.
A useful test record includes the source segment, reference image, prompt, model, settings, output file, and review notes.
Separate observations from assumptions
The recorded review reported changed appearance and clothing, a broadly retained background at the start, middle, and end, and aligned original audio.
Those observations support a limited conclusion: this particular talking-shot example preserved several intended properties during a basic review.
They do not establish precise face consistency, natural hands, or accurate lip sync throughout every frame. They also do not establish reliable results for dancing, heavy occlusion, or replacing one person while leaving another entirely unchanged.
A single successful clip helps define the next test. It does not provide a success rate.
Review more than visual similarity
For a character-replacement task, define the expected changes and the properties that should remain stable.
A practical review covers:
- Identity: Does the appearance remain consistent during turns?
- Motion: Are gestures recognizable, and do hands move plausibly?
- Scene: Do objects, framing, and background details drift?
- Occlusion: What happens when a hand or object crosses the face?
- Audio: Does the original recording remain aligned?
- Lip movement: Do mouth movements fit the speech?
Audio alignment and lip movement deserve separate checks. Retaining the original soundtrack says little about whether the generated mouth moves convincingly.
Compare matching moments in the source and output. Sample the beginning, middle, and end, then inspect transitions and obstructed frames more closely. Finally, watch the entire clip at normal speed with sound.
Change one variable at a time
If the identity drifts, a clearer reference photo or a simpler source segment is a useful next experiment.
If the scene changes unexpectedly, check the selected task and preservation instructions. A motion-transfer task may reconstruct the surroundings rather than preserve them.
Use the same review criteria for each version. Otherwise, it is easy to favor the output with the strongest opening frame and overlook a new problem later in the clip.
For a broader evaluation, add cases with side profiles, fast gestures, occlusions, and multiple people. Track failures as carefully as successful examples.
The practical goal is to make each conclusion traceable to the footage and settings behind it. That makes both product decisions and claims about quality easier to defend.
Disclosure: The documented example comes from Genjutsu AI, a project I work on. This article was drafted with AI using the project's documented test material.
Top comments (0)