DEV Community

Aleister
Aleister

Posted on AI-assisted

Two photos in, one duet out: building an AI video generator

I recently shipped LobbyBooth. You upload two photos, it returns a 10-second 720p clip of those two subjects performing side by side in an orange studio, under a single hanging microphone, lips moving to the same track.

There are six examples on the landing page — two people in yellow coveralls, a couple dancing, a cat and a golden retriever, raccoons in streetwear, anime characters, an elderly couple. Same room, same mic, same camera, same music every time. Only the performers change.

Three things were much harder than I expected:

  1. Consistency. Getting the model to keep the room and the characters instead of re-inventing both every run.
  2. Long jobs. A 10s 720p clip takes minutes to render — the normal case is under three, and I've watched one take thirteen. You can't hold an HTTP request open that long, and you can't lose the job — or charge for it twice — when the user refreshes.
  3. Audio. The vendor model doesn't preserve the reference track, so the music had to be muxed back afterwards. ffmpeg can't run in a Cloudflare Worker, so that piece lives somewhere less fashionable.

Here's what each one looks like in practice.


1. Consistency is a reference-video problem, not a prompt problem

The naive version is a text prompt: "two people singing in an orange studio with a hanging microphone." That gets you a different room every time — different orange, different mic height, camera drifting between runs. It's not a style problem you can talk your way out of.

What actually works: stop describing the scene and start handing it over.

I keep a plate — a 10-second clip of the booth with the camera move, the mic, the lighting and the music already locked in, stored in R2. At generation time the plate goes in as a video reference, and the user's two photos go in as image references. The model's job is reduced to "keep everything, replace the performers":

Keep everything the same in the video - the same deeply saturated orange studio
backdrop with its warm center glow, the same hanging microphone, the same camera
movement, lighting and music. Keep the orange rich and vivid rather than pale or
washed out. The two people from the two reference images are the performers, the
first image on the left and the second image on the right.
Enter fullscreen mode Exit fullscreen mode

That's the entire prompt. One string for the whole product — people, pets, anime characters, all of it. Request shape (Seedance 2.0 mini, through an aggregator API):

{
  "model": "seedance-2.0-mini",
  "prompt": "<the plate prompt above>",
  "duration": 10,
  "resolution": "720p",
  "size": "16:9",
  "generate_audio": false,
  "image_urls": ["<left photo>", "<right photo>"],
  "video_urls": ["<plate.mp4>"]
}
Enter fullscreen mode Exit fullscreen mode

Order is load-bearing: reference image #1 is the left performer, #2 is the right one.

A few things I learned the hard way, with dates, because each one cost real renders:

  • "Keep everything the same" is not filler. Without that sentence the model quietly drops the hanging mic and drifts the composition. With it, the plate holds.
  • Never describe the operation as "replacing faces." replace the faces / face swapping phrasing gets the task rejected by the content safety system — no output, and the vendor still bills for it. "Preserve both faces and outfits" passes fine. Same intent, opposite outcome, purely lexical.
  • Don't put third-party music in the reference audio. It gets blocked by an audio-copyright policy. Use a plate you made yourself.
  • The same prompt covers pets. Identity comes entirely from the reference images; there's no separate animal template.

What holds and what drifts (ten seconds of social video, not a portrait studio): outfits, hairstyles, colors and pose energy are stable. Pet fur markings survive. A flat 2D anime illustration comes back with 3D-ish shading — recognizable, but it is a style shift, and I say so on the site rather than pretending otherwise. Fine facial detail can drift on a fast head turn.

One more thing worth knowing if you build something similar: the vendor does not return the plate's audio. I measured the correlation between the reference track and the output audio at roughly 0.01 — it's a new, unrelated audio generation. So generate_audio: false and mux the music back in yourself. Which brings us to the least glamorous part of the stack.


2. Long jobs: leases, slots, and a unique index doing a queue's job

The architecture is small on purpose: the API submits to the vendor and persists a task row; the task row is the source of truth; the browser polls our endpoint, never the vendor's.

POST /api/ai/generate ──► consume credits ──► submit to vendor ──► ai_task row (pending)
                                                                        │
   browser polls GET /api/ai/task ◄────────────── our API ◄────────────┘
                                                        │  poll lease + singleflight
                                                        ▼
                                                   vendor GET /tasks/:id
Enter fullscreen mode Exit fullscreen mode

Three production problems, in the order they showed up:

a) Concurrency without adding a queue

Every user gets 10 concurrent slots. No Redis, no job runner — just a partial unique index:

uniqueIndex('uq_ai_task_user_active_slot').on(table.userId, table.activeSlot)
Enter fullscreen mode Exit fullscreen mode

Insert with slot 0…9, retry on conflict, and the database is the lock. What this buys is honesty under load: a user can be running ten renders and the eleventh gets a clear answer instead of a pile of unbounded vendor calls.

The failure mode is a ghost slot: a request that dies between claiming a slot and finishing leaves it taken forever. Any task that hasn't changed state in 20 minutes is expired and its slot released. (I have a test suite that simulates exactly this — ghost reclaim, lease renewal, tab takeover — because these are the bugs you can't reproduce by clicking around.)

b) Two tabs, one vendor task

With a browser doing the polling, a second tab (or a refresh) means two pollers hammering the same vendor task, doubling the API cost for a task you already paid for once.

Fix: a poll lease. A poller writes poll_lease_id + a 30-second expiry, and renews every 10 seconds while it's working. A second poller sees a live lease and returns the cached state instead of calling the vendor — singleflight, enforced in the database rather than in the client.

c) Charging people correctly

This is the part that keeps you up at night, because it's other people's money.

Credits are consumed when the job is accepted, FIFO across the user's credit batches (which have expirations). If the render fails, the consumption record is revoked — the user gets their credit back.

The exception is deliberate: content-policy refusals are not refunded, because the vendor bills me for them regardless. That rule is written into the code next to the refund call, because in six months it looks exactly like a bug.

And a related bug I did ship, briefly: credit cost used to depend on the request parameters (model + duration + resolution had to match exactly). One entry point passed a default 5s duration instead of the configured 10s, so the server charged a different tier — users paid for 10 seconds and got 5. The fix wasn't a patch on the arithmetic; it was decoupling cost from parameters entirely. One render = one credit, regardless of what's in the payload. When pricing depends on a parameter that a caller can silently default, the price is wrong and you find out from a support email.


3. ffmpeg does not belong in a Cloudflare Worker

The app runs on Workers, and Workers can't host ffmpeg. But every clip needs the same background music swapped in.

So there's a small mux worker on a plain server. It polls the app for work over two internal endpoints, muxes, and posts the result back. It holds no database credentials and no R2 keys — just a bearer secret for those two endpoints. If that box is compromised, the blast radius is "someone can mux videos."

The actual command is four lines and every flag matters:

ffmpeg -y -i source.mp4 -stream_loop -1 -i bgm.m4a \
  -map 0:v:0 -map 1:a:0 \
  -c:v copy -c:a aac -b:a 192k \
  -shortest -movflags +faststart out.mp4
Enter fullscreen mode Exit fullscreen mode
  • -stream_loop -1 on the BGM is required. Without it, -shortest ends the output at the length of the music — a 10-second music bed silently truncates a 30-second clip to 10 seconds. Looping the audio makes -shortest resolve to the video's length instead.
  • -map 0:v:0 -map 1:a:0 drops every source audio track, whether the model produced one, several, or none.
  • -c:v copy leaves the video bitstream untouched — no re-encode, no generation loss, and the whole pass takes under a second.
  • -movflags +faststart moves the moov atom to the front, which also fixes vendor MP4s that ship without it.

The lesson generalizes: serverless is great until you need a real binary. The fix isn't to bend the runtime — it's to put a very dumb, credential-free worker where the binary lives.


The stack, for the curious

  • TanStack Start (React 19, Vite) on Cloudflare Workers; D1 for data, R2 for media
  • Drizzle ORM — the schema is multi-dialect, production is D1
  • better-auth for accounts; an inlined payment layer with Stripe, PayPal, Creem, Alipay and WeChat Pay adapters; a FIFO credit ledger with expirations
  • Paraglide for i18n — English and Chinese, same routes, /zh prefix on one side only
  • Video: Seedance 2.0 mini via an aggregator (APIMart), 10s / 720p / 16:9, 1 credit per render
  • Pricing is deliberately boring: $4.90 for one video, $9.90 for three, $19.90 for ten

Things I'd tell myself at the start

  • The prompt is the smallest part. The plate (a fixed reference clip) does more for quality than any wording. If your format has a fixed scene — kitchen, car, gym mirror — record it once and reference it forever.
  • Don't build a queue until the database stops being enough. A unique index and a lease took this surprisingly far.
  • Keep the vendor layer thin. Model names and tiers change monthly; every provider gets one adapter that normalizes to a single task shape. Swapping models should be a config value, not a refactor.
  • Put the money rules in code comments. "Policy refusals are non-refundable" is a business decision that reads like a bug to anyone who didn't make it.

Try it, or tell me how you'd solve it better

Two photos in, one duet out: lobbybooth.com — the six examples are the best way to see what the pipeline actually produces before spending anything.

If you're working on character consistency for video models, I'd genuinely like to compare notes in the comments — especially if you've solved hair and accessory drift across longer clips, or found a better way to keep an illustration's original style instead of letting the model re-shade it.

Top comments (0)