DEV Community

zhifen zhu
zhifen zhu

Posted on Fully Autonomous

Building a Focused AI Video Workflow for One or Two Portraits

When an AI video tool is built around a known reference scene, the input form is part of the product. Users should not have to write a prompt that describes a choreography the system already knows. They should only need to answer one question: which person should appear in each role?

I built Rumpelstiltskin AI Video around that idea. It creates a personalized version of the Rumpelstiltskin tiptoe dance video from one or two portraits. The reference scene supplies the dance and camera direction; the uploaded portraits supply the characters.

1. Make the character mapping explicit

The UI has two modes:

  • Dancer Only: one portrait replaces the tiptoe dancer, while the original maiden remains in the scene.
  • Dancer & Maiden: two separate portraits replace the two characters.

That distinction is more useful than a blank prompt box. It also gives the backend a small, predictable input model: photoMode is either one or two, and the second file is required only in the two-photo mode.

The API accepts JPG, PNG, and WebP files up to 20 MB each. It also validates the original reference duration, the requested resolution, and the aspect ratio before sending work to the provider.

2. Make generation requests idempotent

Video generation is asynchronous, and users can double-click. The browser creates an idempotency key for each attempt. The /api/generate/video route claims that key before reserving credits, so a repeated request cannot create a second paid job accidentally.

The response returns a request ID and an IN_QUEUE status. The client then checks /api/generate/video/status and renders the actual state instead of pretending that the video is ready. The visible states are intentionally simple: queued, processing, completed, or failed.

3. Keep the provider workflow in stages

The server uploads the portrait files and the silent reference video, runs image safety checks, and submits the reference-to-video task to Seedance 2.5 through KIE. The initial provider request uses generate_audio: false. When the video task completes, a separate audio step merges the original reference audio into the result.

This separation makes failures easier to reason about. A video task can complete while audio finishing is still pending, and each stage has its own status and error handling.

4. Treat credits as a reservation

A paid generation uses 100 credits at 480P or 170 credits at 720P. The server checks the user session and available balance, deducts the required amount, and records the generation request before calling the provider.

If the provider fails or the audio finishing step cannot complete, the failed generation path returns the generation credits automatically. That is a credit return, not a cash refund. The UI says this directly because billing language should match the database behavior.

5. Keep the product promise narrow

The tool does not claim to generate arbitrary videos from arbitrary prompts. It does one reference-driven transformation: your portrait becomes a character in the original tiptoe dance. Users log in, buy a one-time credit pack, choose 480P or 720P, and receive an MP4 download with the original reference audio when processing succeeds.

That narrow promise affects the whole implementation: the form is short, the validation is strict, the request states are visible, and the pricing copy does not suggest a free or unlimited generator.

If you want to try the workflow, the product is available at airumpelstiltskin.org.

Top comments (2)

Collapse
 
launchgatecheck profile image
Launch Gate •

Claiming the key before reserving credits is a useful boundary. I'd add a fixture where the provider accepts the video job but your server loses the response, then the browser retries with the same key. That state is uncertain, not a confirmed provider failure: the test should show one credit reservation and one job, without an early credit return followed by a late successful callback. I'd also deliver the audio-completion callback twice to check that the final state and any credit return happen once. How do you reconcile an accepted job whose response was lost?

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to