DEV Community

Yana Li
Yana Li

Posted on Fully Autonomous

Building a Rumpelstiltskin AI Video Pipeline

A convincing first frame did not give us a convincing Rumpelstiltskin AI video. In one early test, the still preserved the face, but the video model redrew it immediately. In another, the face survived until the camera moved down and left the person's head outside the frame.

This is an AI-assisted write-up of our Claude Code-assisted tuning work. The experiments were recorded on October 10–11, 2026. They are a small test set, not a general model benchmark. The product handoff and the experimental tooling are separate.

1. Start by separating identity from style

We began with a close-up “shh” frame and compared image-editing routes. Seedream 4.0 Edit produced a useful starting point in that round, but prompt revisions exposed a tradeoff.

Adding more costume, period-film, and character descriptions improved atmosphere while changing the person. One revision pushed a test face toward a different gender presentation. The next revision simplified the description, explicitly preserved hairstyle and adult facial characteristics, and kept only the visual details that mattered.

Film grain moved to local post-processing. It did not need to compete with identity instructions inside the image prompt.

We used generated adult test portraits and some degraded variants rather than relying on one ideal selfie. ArcFace scores helped compare local iterations, but they were diagnostic measurements on our test set—not probabilities that a viewer would recognize someone.

2. A good still did not survive every video route

The early video experiments used the same starting image where possible. Their failures were specific:

  • Wan 2.6, single-shot image-to-video: the recorded similarity fell from about 0.79 in the input still to about 0.42 in the first video frame. A closer crop helped in a later probe, but did not solve the route's identity and background changes.
  • Kling 2.6, one longer shot: the face was stronger at the beginning, but the camera moved down around the fourth second. The dance continued with the head outside the frame.
  • Three separately generated Kling clips: editing a shared approved still helped connect costume and setting across the cuts. The expression and movement were still too mild, and the face became too small during the dance.
  • Early multi-shot trials: the sequence became more expressive, but later close-ups drifted. In one Kling 3.0 Omni trial, the secondary character also took on the uploaded person's facial features.

These observations explain why we changed the representation of the task. Adding another adjective to the original prompt would not address all four failure modes.

3. Build reusable references and a costume still

The more promising route separated the ingredients into five reusable assets:

  1. Costume on a faceless mannequin.
  2. A fictional adult secondary character.
  3. A close-up of the curled shoes.
  4. The empty barn set.
  5. The gold and light used at the ending.

For each input photo, an image-editing step then generated a full-body costume still using the photo, costume, and shoes. That still stayed vertical even when the final output was landscape: it was a reference for the person and outfit, not the video's first frame.

The video call received seven references in a fixed order. The following simplified mapping shows the roles; it is not a runnable API request:

reference_images = [
    person_photo,       # @Image1: identity
    costume_still,      # @Image2: person wearing the outfit
    costume_reference,  # @Image3
    second_character,   # @Image4: fictional adult character
    shoe_reference,     # @Image5
    barn_reference,     # @Image6
    gold_reference,     # @Image7
]
Enter fullscreen mode Exit fullscreen mode

That ordering was part of the template contract. The prompt referred to those positions explicitly; swapping two images could change which face or costume belonged to which role.

Our recorded single-person route used the provider's Seedance 2.0 Mini reference-image mode, requesting 13 seconds, 720p, and no generated audio. It did not pass a first-frame image alongside the reference list. In our integration tests, combining the two input modes produced a parameter error.

4. Write a shot sequence, then freeze it

We inspected reference videos to make written timing notes. The original footage was not supplied as a generation input or cut into the generated output.

The prompt described 15 timed shots: approaching, the “shh” close-up, reaction shots, an exaggerated grin, dance movement, a shoe close-up, and the ending. Shared blocks described identity, clothing, setting, and visual style. The prompt builder repeated the relevant blocks in each shot instead of assuming the model would carry every detail forward.

The result was still approximate. The accepted landscape experiment produced about 12 detected shots from the 15-shot plan. Several very short beats merged. Explicit timing gave us a more useful sequence, not frame-accurate editing control.

Once that version was accepted, the handoff treated two prompts and five shared images as one versioned template. Changing any of those seven files meant a new version and another landscape/portrait check.

The portrait experiment used a different test face, including glasses and a moustache, while keeping the same video prompt and reference scheme. Both were retained in the reviewed result. That was a second useful validation case, not evidence of reliability across all faces.

5. Keep finishing deterministic

The resulting workflow was:

Input photo + costume + shoes
              |
              v
       Full-body costume still
              |
              v
Photo + costume still + five shared assets
              |
              v
     One multi-shot video generation
              |
              v
FFmpeg trim / framing / film look / silent MP4
              |
              v
       Technical checks and review
Enter fullscreen mode Exit fullscreen mode

FFmpeg handled trimming, scaling/cropping, pixel format, and the film treatment. Contact sheets made it easier to inspect the sequence without replaying the entire video for every comparison. The tuning outputs were silent; music packaging was a separate product concern.

The experiment runner also checked the spending limit before a paid call, recorded the prompt version and requested parameters, and reconciled reported credits with the balance change. This mattered when a command was interrupted after a task had already been submitted: stopping the local process did not mean no work had been charged.

For the accepted single-person route, the recorded generation components were 10 credits for the costume still and 106.6 for the video. Shared asset creation and earlier experiments were separate costs. Those numbers describe those runs, not a current price quote or a complete delivery cost.

6. The QA script became part of the experiment

Our first scoring approach treated too much of the video as if the uploaded person should be visible everywhere. Reaction shots of the second character then looked like identity failures. Wide shots, expressive faces, and missed cut points also distorted the result.

We tried shot-aware scoring, including a face-size condition relative to the frame's short side. But this introduced another dependency: if cut detection failed, the identity measurement could combine different people into one segment.

The eventual handoff removed face-similarity scoring as a production blocking condition. It kept deterministic technical checks such as task failure, an unreadable file, black frames, and frozen output. The face-analysis tooling remained in the tuning workspace. That decision did not establish that every output would look good; it acknowledged what those measurements could and could not decide.

7. Two-person mode: a working experiment, then integration

The next experiment added a costume still for the second person and an eighth reference image for their original photo. A 15-second storyboard added more attention to the feet and dance movement.

The recorded manual review found both identities recognizable without face transfer in that one test. The automated score still failed because missed cuts grouped several shots together. We recorded both results rather than presenting a machine score as the whole outcome.

Two-person mode is in development and coming soon. One reviewed experiment is useful evidence for the integration work; it does not establish broad coverage across pairs of photos.

For visual context, the Rumpelstiltskin AI video examples and photo tutorial show the current public single-person results and distinguish them from the original reference. That page documents the output style, not a model leaderboard.

The most useful change was making the experiment easier to inspect: named assets, ordered references, versioned prompts, saved outputs, and a record of what failed. For anyone building a similar workflow, how do you separate a broken file from a technically valid result that still misses the intended performance?

Top comments (0)