<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Razi Kallayi</title>
    <description>The latest articles on DEV Community by Razi Kallayi (@razi_kallayi).</description>
    <link>https://dev.to/razi_kallayi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3615048%2F759ff408-8c02-4357-93a0-d0e4d168c666.png</url>
      <title>DEV Community: Razi Kallayi</title>
      <link>https://dev.to/razi_kallayi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/razi_kallayi"/>
    <language>en</language>
    <item>
      <title>What Broke Automating Film Production with Claude Code and Google Flow</title>
      <dc:creator>Razi Kallayi</dc:creator>
      <pubDate>Wed, 02 Sep 2026 23:59:23 +0000</pubDate>
      <link>https://dev.to/razi_kallayi/what-broke-automating-film-production-with-claude-code-and-google-flow-5hl9</link>
      <guid>https://dev.to/razi_kallayi/what-broke-automating-film-production-with-claude-code-and-google-flow-5hl9</guid>
      <description>&lt;h1&gt;
  
  
  What Broke Automating Film Production with Claude Code and Google Flow
&lt;/h1&gt;

&lt;h2&gt;
  
  
  I made my four-year-old and my two-year-old into superheroes. Here is every single thing that broke.
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flmkn2ruyszeq6mvq3gzi.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flmkn2ruyszeq6mvq3gzi.webp" alt="What broke automating film production with Claude Code and Google Flow — Comet and Boing beneath the moon in the tree" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Lia is four. For about a year now she has been directing rescue missions from the top bunk — she has watched enough PAW Patrol to have absorbed its story grammar whole, and she narrates her own episodes out loud at bedtime. Airik is two. He follows her around shouting the parts he can pronounce.&lt;/p&gt;

&lt;p&gt;One night I was doing the voices for the fourth time and thought: the tools are supposedly here now. Are they actually? Could one parent, one GPU and one month's subscription put those two children on screen, in a real film, with their own faces and a proper score and a narrator?&lt;/p&gt;

&lt;p&gt;This is the answer. It is called &lt;strong&gt;Star Patrol: The Moon in the Tree&lt;/strong&gt;. Twenty shots. Three minutes and twenty seconds. Comet can fly but is not strong; Boing is strong but cannot fly; the moon is stuck in a tree and neither of them can get it down alone.&lt;/p&gt;

&lt;p&gt;It works. It also broke in about nine different ways, several of which I have not seen written down anywhere, and that is the useful half of this post.&lt;/p&gt;

&lt;p&gt;One thing before you start: &lt;strong&gt;this is one path, not the path.&lt;/strong&gt; At every decision below there was a cheaper route, a local route, or a route I looked at and walked away from, and I have written those in alongside what I actually did — tagged &lt;strong&gt;tried&lt;/strong&gt;, &lt;strong&gt;tried and failed&lt;/strong&gt;, or &lt;strong&gt;not tried&lt;/strong&gt;, so you can tell the difference. The subscription in the table below bought me speed, not the film. If you have no subscription, no GPU, or a different language to work in, there is a way through and it is marked.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbjog7gmk0yk0ebmhxo57.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbjog7gmk0yk0ebmhxo57.webp" alt="Comet and Boing, the two heroes, in a full-height line-up against a plain studio background" width="800" height="1182"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The line-up image, &lt;code&gt;Together.jpg&lt;/code&gt;. This one file fixed two separate production disasters at once — and it was the third character image I generated, when it should have been the first. More on that below.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Watch it first
&lt;/h2&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Z35GG4hxVMU" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;h2&gt;
  
  
  What it cost and how long it took
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Runtime&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3:20 (200.04 s), 1920×1080, 24 fps, 20 shots × 10 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Direct cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One month of Google AI Pro. Nothing else — and this could equally have been made for nothing, on the free plan's 50 daily Google Flow credits.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Was the subscription necessary?&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No. It bought wall-clock, not capability — Flow's free tier makes the same film in about five days. See Making this with no subscription at all.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Flow credits spent&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;200 of the 1,000 in the monthly allowance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Marginal cost of everything else&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Zero — score, voice cloning, transcription and assembly all ran locally or on free tiers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;The real cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Time. Reference-image iteration, one long debugging session on Malayalam durations, and the narration edit. Video generation was the cheap part.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deliverables&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Six finished cuts — two Malayalam, three English, one music-only. Picture and score identical across all six; only the narration track differs.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The stack, end to end: &lt;strong&gt;Veo 3.1 Lite&lt;/strong&gt; via &lt;strong&gt;Google Flow&lt;/strong&gt; for picture and native audio, &lt;strong&gt;Nano Banana&lt;/strong&gt; (Gemini image generation) for reference art, &lt;strong&gt;MusicGen Medium&lt;/strong&gt; for the score, &lt;strong&gt;F5-TTS&lt;/strong&gt; and &lt;strong&gt;IndicF5&lt;/strong&gt; for voice cloning, &lt;strong&gt;ElevenLabs&lt;/strong&gt; for one stock-voice cut and for the Malayalam transcription, and &lt;strong&gt;FFmpeg&lt;/strong&gt; holding the whole thing together. All of it driven from &lt;strong&gt;Claude Code&lt;/strong&gt; running &lt;strong&gt;Opus&lt;/strong&gt;, which wrote every script named in this post — more on that below.&lt;/p&gt;

&lt;p&gt;Everything below is what actually happened.&lt;/p&gt;




&lt;h2&gt;
  
  
  What was automated, and what was not
&lt;/h2&gt;

&lt;p&gt;Before the story starts, here is who actually did what.&lt;/p&gt;

&lt;p&gt;Almost every mechanical step below was performed by an agent. Not &lt;em&gt;assisted by&lt;/em&gt; — performed. I described what I wanted in a terminal, argued with the first draft, and kept the version that survived contact with real output. The scripts in &lt;code&gt;05_TOOLS/&lt;/code&gt; are the result, and I did not type them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Who did it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Story, shot list, prompt assembly&lt;/td&gt;
&lt;td&gt;Agent drafted, I edited and approved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference art direction&lt;/td&gt;
&lt;td&gt;Me — every taste call, every re-roll kept or killed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Generating the twenty shots&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Me, by hand.&lt;/strong&gt; Flow has no API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Score generation and arrangement&lt;/td&gt;
&lt;td&gt;Agent wrote the pipeline; I picked between three options&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Voice cloning pipeline&lt;/td&gt;
&lt;td&gt;Agent — and I threw the output away&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cutting the recordings into shot slots&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transcription and alignment&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mixing, ducking, loudness&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Subtitle timing and .srt files, two languages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Agent&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Six deliverable builds from one master&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Drifting watermark&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;The narration you actually hear&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Two humans, one phone&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;No video editor was ever opened.&lt;/strong&gt; No Premiere Pro, no Final Cut Pro, no DaVinci Resolve, no CapCut, no InShot — no timeline at any point. The shots were generated by hand in Google Flow, and everything after that — the assembly, the mix, the subtitles, the watermark and six separate deliverables — was FFmpeg — and every one of those filter graphs was generated by an LLM in Claude Code rather than typed by me.&lt;/p&gt;

&lt;p&gt;The interesting rows are the bold ones, and they disagree with each other on purpose.&lt;/p&gt;

&lt;p&gt;The agent did far more than I expected — a sidechain chain I would not have found, a watermark that moves on two incommensurate sine periods, subtitle timing I never asked for. It also absorbed a day of environment damage that would otherwise have eaten a weekend.&lt;/p&gt;

&lt;p&gt;And it could not do the two things that mattered most: it could not click through Google Flow, and it could not make a voice sound like it loved anybody.&lt;/p&gt;

&lt;p&gt;If you only take one thing from this post, take that shape. The boring, expensive, repeatable middle of a creative pipeline is now genuinely automatable. The ends — taste at one end, feeling at the other — are still yours. There is a longer account of how this ran, and what became of it, in The pipeline had a fourth author.&lt;/p&gt;




&lt;h2&gt;
  
  
  Twenty shots, because the model can only count to eight
&lt;/h2&gt;

&lt;p&gt;The constraint came before the story.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/video" rel="noopener noreferrer"&gt;Veo 3.1&lt;/a&gt; in &lt;a href="https://labs.google/flow/about" rel="noopener noreferrer"&gt;Google Flow&lt;/a&gt; generates 8-second clips, and credits are consumed per clip regardless of what is in it. So the story had to be told in a whole number of short, self-contained beats — and if I wanted a round runtime with a bit of breathing room, each 8-second generation would sit in a 10-second slot on the timeline.&lt;/p&gt;

&lt;p&gt;Twenty shots. 200 seconds. That is the entire structural argument.&lt;/p&gt;

&lt;p&gt;Twenty beats turns out to be almost exactly the PAW Patrol episode skeleton, which is convenient, because it is what a four-year-old already has installed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Beats&lt;/th&gt;
&lt;th&gt;Function&lt;/th&gt;
&lt;th&gt;Shots&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Meet the squad&lt;/td&gt;
&lt;td&gt;Establish the two heroes and their catchphrases&lt;/td&gt;
&lt;td&gt;1–2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The alert&lt;/td&gt;
&lt;td&gt;A problem arrives from outside&lt;/td&gt;
&lt;td&gt;3–4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Roll out&lt;/td&gt;
&lt;td&gt;Vehicles, motion, joy&lt;/td&gt;
&lt;td&gt;5–7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The reveal&lt;/td&gt;
&lt;td&gt;See the problem at full scale&lt;/td&gt;
&lt;td&gt;8–10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The complication&lt;/td&gt;
&lt;td&gt;Each hero fails &lt;em&gt;alone&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;11–14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The idea&lt;/td&gt;
&lt;td&gt;The two abilities combine&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The rescue&lt;/td&gt;
&lt;td&gt;Payoff&lt;/td&gt;
&lt;td&gt;16–19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home&lt;/td&gt;
&lt;td&gt;Descend to sleep&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn265oems6hfbs01pq3ia.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn265oems6hfbs01pq3ia.webp" alt="Contact sheet showing every shot in the film as a grid of stills" width="799" height="572"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The whole film as a contact sheet. Read left to right, top to bottom, and the structure above is visible without a single word of dialogue: bedroom, alert, roll-out, the tree, the failures, the light, the moon going home, bed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two things were written specifically for a 2- and 4-year-old audience.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every emotional beat is physical.&lt;/strong&gt; Comet can fly but is not strong. Boing is strong but cannot fly. That is the whole theme — nobody can do it alone — and it is expressed entirely as two children failing to reach a thing in a tree. No dialogue is required to understand it. A two-year-old gets it from the picture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every line of dialogue is a shout.&lt;/strong&gt; "Comet — bright and ready!", "Star Patrol — on the go!", "SUPER SIBLINGS — GO!", "I SEE IT!", "Gotcha." Eleven dialogue lines across twenty shots, none longer than a breath. That was partly about attention span and partly a hedge against a problem I will come back to: Veo generates a &lt;em&gt;fresh voice on every clip&lt;/em&gt;, and short loud lines vary far less perceptibly than sustained conversation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;There isn't enough tone in a two-word shout for the ear to lock onto and then notice changing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The final shot list lives in a single production document (&lt;code&gt;01_SCRIPT/star-patrol-production-pack.md&lt;/code&gt;) holding the prose story, the credit budget, all twenty prompts, the reference-image prompts and the working rules. Everything downstream reads from it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — longer beats.&lt;/strong&gt; Twenty ten-second slots is a shape the &lt;em&gt;tool&lt;/em&gt; chose. Flow has an extend feature, and some other services generate ten seconds or more natively; either would let you write a real two-hander scene instead of eleven shouts. If your story needs conversation rather than physical comedy, don't inherit my structure — it was built around an eight-second ceiling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — no narrator at all.&lt;/strong&gt; One of the six cuts is music-only and it holds up. If the story is expressed physically, the narration track is a nice-to-have, and skipping it removes the entire voice-cloning half of this post from your pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — a different skeleton.&lt;/strong&gt; PAW Patrol's grammar was chosen because my four-year-old already had it installed. A bedtime &lt;em&gt;poem&lt;/em&gt; — twenty rhyming couplets, one per shot — would fit the same constraint and is arguably easier to narrate.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Turning two children into ingredients
&lt;/h2&gt;

&lt;p&gt;The third image I made should have been the first.&lt;/p&gt;

&lt;p&gt;That sentence is most of what this chapter cost me, and I am putting it at the top so you get it for nothing. What follows is how two real children became something a video model would accept, and the two disasters that one boring group shot fixed at a stroke.&lt;/p&gt;

&lt;p&gt;The kids had to be recognisably themselves. That meant starting from photographs and ending at something stylised enough that a video model would accept it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A note on the photographs.&lt;/strong&gt; The source photos of Lia and Airik are not in this post and never will be. They are family pictures of small children and they stay off the internet. Everything you see here is the &lt;em&gt;output&lt;/em&gt; — the stylised 3D characters — which is also the only thing that ever entered the video pipeline. That was a deliberate rule from the start, and it turned out to have a technical payoff too, discussed in the safety section below.&lt;/p&gt;

&lt;p&gt;Four candidate photos per child. The selection criteria were unglamorous: front-facing, sharp, neutral colour temperature. One of Airik's photos had the better grin but a heavy orange cast that would have propagated into the render as skin tone, so it lost to a duller, brighter, more neutral frame. For Lia the best face was a video still, which meant cropping off the phone UI — status bar, scrubber, share strip — before uploading, at the cost of some softness.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/image-generation" rel="noopener noreferrer"&gt;Gemini's image generation&lt;/a&gt; — Nano Banana — accepts multiple reference images, and attaching two photos per child with the instruction "use both as reference for the same child" produced a visibly better likeness than either alone.&lt;/p&gt;

&lt;p&gt;These were run in the &lt;strong&gt;Gemini app&lt;/strong&gt;, deliberately, because the app's free image quota is a separate pool from Flow's video credits. Six images cost nothing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnerq37y30tgudgxo9pta.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnerq37y30tgudgxo9pta.webp" alt="Comet, the older sister character, in orange pyjamas and a gold cape" width="800" height="1183"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfk6rln115ajxzrsiswd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfk6rln115ajxzrsiswd.webp" alt="Boing, the younger brother character, in blue pyjamas with coil-spring boots" width="800" height="1183"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Comet and Boing as isolated ingredients. The source photographs are withheld deliberately — these renders are the only version of these children that this project ever published, or uploaded anywhere.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here is the prompt that produced Comet:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Using the uploaded photo(s) as reference for the same child, create a 3D animated Pixar/DreamWorks-style character of this girl. Preserve her exact facial structure, features, proportions and expression from the reference photo — the only change is a fair, light skin tone. Keep her clearly recognizable: fine dark brown-black wavy hair pulled back into a small ponytail with soft wispy curly tendrils escaping around her face and temples and a light wispy fringe across her forehead, large round bright dark brown eyes with dark lashes, softly arched dark eyebrows, a round face with full cheeks and dimples, a small nose, and a wide joyful open smile showing small child's teeth. Stylize into soft appealing 3D animation with slightly larger expressive eyes while keeping her real facial structure. Full body head to toe, standing centered, facing camera, confident happy hero pose, arms relaxed at her sides. She wears bright orange pajamas with a gold star on the chest, a flowing golden-yellow cape, and a glowing gold five-pointed star badge, with a soft warm golden aura around her. Plain light-grey studio background, soft even studio lighting, sharp focus. No text, no watermark, no other characters.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three structural choices in there are worth isolating:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;"Preserve her exact facial structure… the only change is X."&lt;/strong&gt; Naming exactly one permitted deviation stops the model treating the photo as loose inspiration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The features are enumerated in prose anyway.&lt;/strong&gt; Hair, eyes, brows, face shape, nose, mouth. The reference image is not trusted to carry them on its own — and the spacesuit incident, later, shows exactly why that instinct was right.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full body, centered, plain light-grey background, studio lighting.&lt;/strong&gt; This is a &lt;em&gt;character sheet&lt;/em&gt;, not a picture. It exists to be cut out and re-lit inside twenty different scenes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Boing's prompt is the same skeleton with his own features and a deliberate note that his skin tone is "a touch warmer and very slightly deeper than his sister's." Sibling colour relationships have to be stated &lt;em&gt;relatively&lt;/em&gt; or the two renders drift apart.&lt;/p&gt;

&lt;p&gt;The same treatment produced four props:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7yhrcrnn212nf55jzfih.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7yhrcrnn212nf55jzfih.webp" alt="A horizontal strip of the four prop ingredients: gold star glider, blue spring buggy, sleepy-faced moon, small pink star" width="800" height="109"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;code&gt;GlideStar&lt;/code&gt;, &lt;code&gt;BounceBuggy&lt;/code&gt;, &lt;code&gt;Moon&lt;/code&gt; and &lt;code&gt;Twink&lt;/code&gt; — isolated, centred, plain background, no text, no other characters, every time. An ingredient image with scenery in it drags that scenery into your shot.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The image that saved the film
&lt;/h3&gt;

&lt;p&gt;The first pass produced &lt;code&gt;Comet.jpg&lt;/code&gt; and &lt;code&gt;Boing.jpg&lt;/code&gt;. Both head-to-toe. Both filling the frame.&lt;/p&gt;

&lt;p&gt;Which meant that &lt;strong&gt;neither image contained any height information.&lt;/strong&gt; Framed identically, a four-year-old and a two-year-old are exactly the same size.&lt;/p&gt;

&lt;p&gt;Veo drew them exactly the same size. In the eleven shots where both children appear, the entire visual joke of the film — the big sister and the small round brother — evaporated. In some renders the inversion went the other way, and the two-year-old came out taller than his sister.&lt;/p&gt;

&lt;p&gt;The fix was a third reference, the &lt;code&gt;Together.jpg&lt;/code&gt; line-up at the top of this post: one frame, both children, correct relative heights baked in, and — the part I did not anticipate — both rendered in the &lt;em&gt;same pass&lt;/em&gt;, so their styles could not diverge. Generating them separately had already allowed small drift in eye size and skin shading that a single joint generation eliminates by construction.&lt;/p&gt;

&lt;p&gt;That gave three character ingredients plus four props, and one rule that follows immediately from having all three: &lt;strong&gt;never attach &lt;code&gt;Comet&lt;/code&gt; + &lt;code&gt;Boing&lt;/code&gt; + &lt;code&gt;Together&lt;/code&gt; at once.&lt;/strong&gt; Three references of the same two children dilute each other. Use the two solos &lt;em&gt;or&lt;/em&gt; the line-up.&lt;/p&gt;

&lt;p&gt;The line-up also solves an arithmetic problem. Flow allows a &lt;strong&gt;maximum of three ingredients per prompt&lt;/strong&gt;, and shots 6 and 7 need both children &lt;em&gt;and&lt;/em&gt; a vehicle. &lt;code&gt;Together&lt;/code&gt; covers both kids in one slot, leaving room for &lt;code&gt;@GlideStar&lt;/code&gt; and &lt;code&gt;@BounceBuggy&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — generate the whole cast in one sheet and crop the solos out of it.&lt;/strong&gt; This is the same insight as &lt;code&gt;Together.jpg&lt;/code&gt;, applied from the first image instead of the third. One joint generation cannot drift in style, and every solo you cut out of it inherits the correct scale for free. If you take one thing from this section, take this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — local image models instead of Nano Banana.&lt;/strong&gt; SDXL or Flux with IP-Adapter or InstantID for face conditioning runs on the same GPU that made the score, costs nothing, and never sends a photograph anywhere. Expect to work harder for likeness: the Gemini app's "use both photos as reference for the same child" is doing a lot of quiet lifting, and the sibling-scale problem gets &lt;em&gt;worse&lt;/em&gt;, not better, when you have finer control and more knobs to get wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried, and deliberately so — training a LoRA on each child's face.&lt;/strong&gt; This is the technically strongest route to consistency, and I did not take it. A LoRA is a portable, redistributable model &lt;em&gt;of a real child's face&lt;/em&gt;, which is a categorically different artefact from four stylised PNGs sitting in a folder on my machine. The consistency was not worth creating that file. If your subjects are adults who have agreed to it, the calculus is yours to make.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — hand-drawn or commissioned character sheets.&lt;/strong&gt; The pipeline does not care where the reference image came from. An illustrator friend and a scanner produce a better ingredient than any of this, and the twenty prompts downstream are unchanged.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Eighty re-rolls beat one pretty render
&lt;/h2&gt;

&lt;p&gt;This is where the project's real budget lives, and it is worth being precise, because the difference between the right and wrong setting is a factor of ten.&lt;/p&gt;

&lt;p&gt;Flow's &lt;a href="https://labs.google/flow/about" rel="noopener noreferrer"&gt;free tier gives 50 credits a day&lt;/a&gt;. Google AI Pro gives &lt;strong&gt;1,000 credits a month with no daily cap&lt;/strong&gt;, which is what I used — it collapses the shoot from a two-week drip into one or two sittings.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Credits per 8 s clip&lt;/th&gt;
&lt;th&gt;All 20 shots&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Veo 3.1 Lite&lt;/strong&gt; ← baseline&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;200&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Veo 3.1 Fast&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Veo 3.1 Quality&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;2,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The plan was Lite for everything. 200 credits buys the film and leaves &lt;strong&gt;800 for re-rolls — eighty of them.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That trade is the single most important production decision in the project.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;On generative video, more takes beats higher resolution. A Quality render of a shot where the model misunderstood the action is worth nothing; a Lite render on the fourth attempt, where it finally got the timing right, is worth everything.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Quality on the three money shots (9, 16, 19) would have cost 300 more and still fit comfortably inside the budget. It was not needed. And 1080p upscaling is free on Pro, so every finished clip was upscaled anyway.&lt;/p&gt;

&lt;p&gt;I shot &lt;strong&gt;in story order&lt;/strong&gt;, 1 through 20. Re-rolls being cheap, there was no reason not to, and continuity problems surfaced while there was still budget to fix them.&lt;/p&gt;

&lt;h3&gt;
  
  
  The setting that quietly costs you a month
&lt;/h3&gt;

&lt;p&gt;Before generating anything, four agent settings need checking. One of them is a real trap:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Must be&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Confirm before generating&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Always&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;On &lt;em&gt;Never&lt;/em&gt;, the agent burns the monthly allocation without asking.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video model&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Veo 3.1 – Lite&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The default was &lt;strong&gt;Omni 1.1 Flash&lt;/strong&gt; — outside the Veo family, and outside what the free video credits cover. Materially more per clip.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video aspect&lt;/td&gt;
&lt;td&gt;16:9&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video count&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;×1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;×4 generates four variants per click. Re-roll deliberately.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The model default is the expensive one. Flow's interface offers Gemini Omni alongside Veo, and it was pre-selected. Nothing in the UI flags that the credit-eligible Veo tiers are a different menu entry. Check it before your first generation, not after your twentieth.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prompt anatomy
&lt;/h3&gt;

&lt;p&gt;Every shot prompt is four stacked blocks. The style block is pasted first, &lt;strong&gt;verbatim, never retyped&lt;/strong&gt; — small wording drifts cause visible style drift across shots:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Modern 3D animated family film, Pixar-quality rendering, soft warm cinematic lighting, shallow depth of field, rich saturated colors, cozy storybook mood, 16:9. No subtitles, no on-screen text, no captions.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then the scene, which is where the words should actually go. Shot 8, the one where Comet becomes the light source:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;They stop at the edge of a pitch-black forest. &lt;a class="mentioned-user" href="https://dev.to/boing"&gt;@boing&lt;/a&gt; looks up nervously. Then &lt;a class="mentioned-user" href="https://dev.to/comet"&gt;@comet&lt;/a&gt; rises into the air and BLAZES with golden light, and the whole forest path illuminates gold — every leaf and puddle shining like daytime. Wide cinematic reveal. Audio: a swelling magical chime, &lt;a class="mentioned-user" href="https://dev.to/boing"&gt;@boing&lt;/a&gt; whispering "whoa."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Note that audio direction is part of the prompt. Veo generates the soundtrack with the picture — spring boings, hiccups, orchestral swells and dialogue all come out of the same generation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhngw5a1d0td00hzih2rg.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhngw5a1d0td00hzih2rg.webp" alt="The two heroes standing under a giant oak with the moon wedged in its branches, lit blue by moonlight" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 9, the reveal. The whole film exists to earn this frame: the moon is not in the sky, it is stuck in a tree, and it is enormous, and they are very small. Chaining shot 8's final frame into shot 9's start image is what made the forest match across the cut.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Then, on the eleven two-shots, the sibling-scale line:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Comet is noticeably taller than Boing — he is about a head shorter, rounder and chubbier, with shorter limbs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And on any shot with dialogue, the voice block for each speaking character, copied verbatim and never reworded between shots:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;COMET speaks with a bright, clear little girl's voice, around five years old — warm, confident and cheerful, mid-high pitch, with an excited upward lilt.&lt;/p&gt;

&lt;p&gt;THE MOON speaks with an enormous, deep, slow, sleepy male voice — gentle and rumbling, like a kind old giant.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Veo invents a fresh voice on every generation. There is no voice-lock at any Flow tier. The repeated verbatim description is the only lever available — and it does work. It turns "two different children" into "the same child on a different day."&lt;/p&gt;

&lt;p&gt;Re-roll budget was concentrated on shots &lt;strong&gt;4&lt;/strong&gt; and &lt;strong&gt;15&lt;/strong&gt;. Shot 4 establishes both catchphrases and both voices and sets the reference the ear anchors to; shot 15 is "SUPER SIBLINGS — GO!", the emotional turn, and it pays off shot 4. Everything else got what it got. The Moon, incidentally, is the most reliable voice in the film — a distinctive deep slow rumble is an easy target for a model to hit repeatedly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.razi.pro/videos/star-patrol/rollout-launch-loop.mp4" rel="noopener noreferrer"&gt;Video: Animated GIF: the heroes launching out of the driveway, glider and spring buggy trailing gold sparks&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 6, the roll-out. Three ingredients in one prompt — &lt;code&gt;Together&lt;/code&gt; for both children, plus &lt;code&gt;@GlideStar&lt;/code&gt; and &lt;code&gt;@BounceBuggy&lt;/code&gt; — which is exactly the maximum Flow allows, and only possible because the line-up image collapses two characters into one slot.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One free continuity trick worth stealing: &lt;strong&gt;chain the shots.&lt;/strong&gt; Take the final frame of a shot you like and feed it as the start image for the next. It costs nothing and substantially improves the match across cuts.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tried — the free tier.&lt;/strong&gt; 50 credits a day is 5 clips a day at Lite rates. The same twenty shots take five sittings, or about ten days if you want a real re-roll budget. Nothing about the film changes; only the calendar does. The Pro month bought me two evenings instead of two weeks, and that is &lt;em&gt;all&lt;/em&gt; it bought.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — Quality on the money shots.&lt;/strong&gt; Shots 9, 16 and 19 at Veo 3.1 Quality would have cost 300 more credits and still fitted inside the allowance comfortably. I skipped it on judgment, not on budget, because the Lite renders were already the ones I wanted. If your film has three frames that carry it, that is where the credits belong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — ×4 variants instead of ×1.&lt;/strong&gt; Generating four variants per click trades credits for choice, and on a shot you already know is hard (my shots 4 and 15) it may well beat four sequential deliberate re-rolls. I turned it off to stop the allowance evaporating by accident, which is a different concern from whether it is a good idea.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — first-frame &lt;em&gt;and&lt;/em&gt; last-frame conditioning.&lt;/strong&gt; Chaining fixes the seam going forward. Pinning both ends of a shot, where the tool supports it, would have removed most of my continuity re-rolls outright.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — leaving Google entirely.&lt;/strong&gt; Open-weight video models — Wan, LTX-Video, HunyuanVideo, CogVideoX — run locally on a consumer GPU for nothing but electricity, and the hosted services (Kling, Hailuo, Pika, Luma) all have free daily allowances. The catch is specific and large: &lt;strong&gt;none of them generate the soundtrack with the picture.&lt;/strong&gt; Veo's native audio gave me spring boings, hiccups, orchestral swells and every line of dialogue in the same generation. Take that away and eleven dialogue lines and twenty shots of foley become your problem. Check each service's terms on children's likenesses before uploading anything, too — they differ, and they change.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Why my son turned into an astronaut
&lt;/h2&gt;

&lt;p&gt;This is the section I would have wanted to read before starting.&lt;/p&gt;

&lt;h3&gt;
  
  
  The reference image was right there
&lt;/h3&gt;

&lt;p&gt;Shot 8's first prompt referenced &lt;code&gt;@Boing&lt;/code&gt; and described what he does, but not what he looks like. The reasoning seemed sound: the ingredient image already carries the costume, so why retype it?&lt;/p&gt;

&lt;p&gt;It came back with Airik in a &lt;strong&gt;spacesuit and helmet&lt;/strong&gt;. Cape gone.&lt;/p&gt;

&lt;p&gt;The model had kept the coil springs and the star badge, correctly identified the scene as nocturnal and sky-adjacent, and free-associated its way to an astronaut.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Ingredients bias the look; they do not lock it. A reference image shifts the distribution — it does not constrain it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Where the prompt text is silent, the model fills the silence from scene context, and "dark sky, glowing, flying" is an extremely strong pull toward space.&lt;/p&gt;

&lt;p&gt;The fix is redundancy. Every shot a character appears in, describe them inline &lt;em&gt;as well as&lt;/em&gt; attaching the ingredient:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/comet"&gt;@comet&lt;/a&gt; — a girl in bright orange pajamas with a gold star on the chest and a flowing golden-yellow cape, bare head, dark hair in a small ponytail&lt;/p&gt;

&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/boing"&gt;@boing&lt;/a&gt; — a much smaller little boy in royal-blue pajamas with a sky-blue cape worn backwards over his chest, bare head with thick dark hair, and blue boots with coil springs&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Plus an explicit negative on every shot with a child in it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Both children have bare heads with visible hair. No helmets, no space suits, no antennae, no goggles, no masks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Naming the specific failure mode is far more effective than a generic "no costume changes." The model needs the token.&lt;/p&gt;

&lt;h3&gt;
  
  
  Refused. Twice.
&lt;/h3&gt;

&lt;p&gt;Shot 14 is the emotional low point of the film. Boing has bounced into a branch three times, sits down in the grass, and his lip wobbles. The original prompt read, in part:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;a small three-year-old boy … lower lip wobbling, right on the edge of tears&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It was refused. Twice. With no indication of what had tripped it.&lt;/p&gt;

&lt;p&gt;The classifier is not objecting to the &lt;em&gt;shot&lt;/em&gt;. It is objecting to the &lt;strong&gt;conjunction of an explicit child age with distress vocabulary&lt;/strong&gt; — a pattern that legitimately warrants scrutiny in the general case, and which also happens to describe a completely ordinary beat in every children's film ever made. Because the refusal message names no token, the mitigations have to be derived by reasoning about the classifier rather than by reading the error.&lt;/p&gt;

&lt;p&gt;Four fixes, in descending order of effectiveness:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Remove every age word.&lt;/strong&gt; "The smaller character", "the taller character". The height relationship is already carried by the sibling-scale line and the &lt;code&gt;Together&lt;/code&gt; reference, so nothing is lost. This alone cleared most refusals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Soften distress vocabulary.&lt;/strong&gt; "Glum", "deflated", "thoughtful". Never "tears", "crying", "upset". The performance you get back is nearly identical — a deflated two-year-old and a tearful two-year-old look the same on screen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Soften physical contact.&lt;/strong&gt; "An encouraging pat on the shoulder" rather than "puts an arm around him."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open with fiction context.&lt;/strong&gt; Begin the prompt "Two animated cartoon characters…" so the classifier has established this is animation before it reaches the action.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://www.razi.pro/videos/star-patrol/bounce-bonk-loop.mp4" rel="noopener noreferrer"&gt;Video: Animated GIF: Boing bouncing up toward the moon in the tree and bonking a branch&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 13 — the third failed bounce, which is the beat that leads into the refused shot 14. Rewritten to drop every age word, it generated first time.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And one rule adopted before the first generation and never broken: &lt;strong&gt;never upload the original photographs of the children to Flow.&lt;/strong&gt; Veo flags real faces, and routing the likeness through a stylised Gemini render keeps the actual photos out of the video pipeline entirely. It is the right call on privacy grounds first, and it happens to also be the smoother technical path. If Flow rejects a reference as too photorealistic, push it further into cartoon territory in Gemini ("more stylized, softer, more like a Pixar character, less photographic") and re-upload.&lt;/p&gt;

&lt;h3&gt;
  
  
  The one that was just my own fault
&lt;/h3&gt;

&lt;p&gt;Minor but persistent: retyping the style block instead of pasting it produced visible tonal drift between shots. Lighting warmth and colour saturation are sensitive to remarkably small wording changes. Paste. Never retype.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tried, failed — trusting the ingredient image to carry the costume.&lt;/strong&gt; That is the astronaut. Describe the character inline on every single shot, even when the reference is attached.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tried, failed — naming the children's ages in a prompt with a sad beat in it.&lt;/strong&gt; Refused twice, with no indication why. Every age word came out and it generated first time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — shooting the sad beat without the child in frame.&lt;/strong&gt; Shot 14 could have been the moon looking down, or a cape on the grass, or the buggy stopped and still. A refusal is sometimes telling you to find a better shot rather than a better prompt, and "the thing he dropped" is a more grown-up piece of filmmaking than "his lip wobbles" anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — a different service when a shot is refused.&lt;/strong&gt; I rewrote until Veo accepted. Moving one stubborn shot to another model is legitimate and I never needed to; the cost is that the style block does not travel, so that one shot will not match the other nineteen without grading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A caution on the workarounds above.&lt;/strong&gt; Removing age words and softening the vocabulary worked because the &lt;em&gt;shot itself&lt;/em&gt; was an ordinary beat from a children's film. These are techniques for describing an innocuous scene in language a classifier can recognise as innocuous. They are not a method for getting something past a filter that is right to stop you, and if you find yourself reaching for them to do that, the filter is working.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  A score that follows the story
&lt;/h2&gt;

&lt;p&gt;Every film you have ever loved lied to you about sound.&lt;/p&gt;

&lt;p&gt;You remember the pictures. What you actually felt was the score underneath them, doing quiet work you never noticed. I did not understand that properly until I had twenty finished shots that were technically correct and completely dead.&lt;/p&gt;

&lt;p&gt;Two of the three jobs here went to plan. The third taught me something I have not seen written down anywhere, and it is the only part of this project that briefly frightened me.&lt;/p&gt;

&lt;p&gt;Clip audio from Veo covers the diegetic sound — springs, wind, dialogue — but there is no continuous score, and twenty independently generated 8-second beds would not join.&lt;/p&gt;

&lt;p&gt;The score was generated locally with &lt;a href="https://huggingface.co/facebook/musicgen-medium" rel="noopener noreferrer"&gt;MusicGen Medium&lt;/a&gt; (1.5 B parameters) through &lt;a href="https://huggingface.co/docs/transformers/model_doc/musicgen" rel="noopener noreferrer"&gt;HuggingFace Transformers&lt;/a&gt;, rather than via &lt;a href="https://github.com/facebookresearch/audiocraft" rel="noopener noreferrer"&gt;AudioCraft&lt;/a&gt; directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scipy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;wavfile&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;wav&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;MusicgenForConditionalGeneration&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;AutoProcessor&lt;/span&gt;

&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;facebook/musicgen-medium&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;proc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AutoProcessor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MusicgenForConditionalGeneration&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float16&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;to&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cuda&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;sr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;audio_encoder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sampling_rate&lt;/span&gt;

&lt;span class="n"&gt;PROMPTS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
 &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gentle warm orchestral lullaby for a children&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s animated bedtime film, soft music box, harp and warm strings, tender magical and hopeful, slow calm tempo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
 &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;whimsical light orchestral score for a children&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s animated adventure, playful pizzicato strings, soft woodwinds, warm gentle and storybook, moderate tempo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
 &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;c&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dreamy cinematic lullaby, celesta and soft strings under a night sky, magical wonder, calm slow and emotional, no drums&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;TOKENS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt;   &lt;span class="c1"&gt;# 50 audio tokens per second -&amp;gt; 30 seconds
&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tag&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PROMPTS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;inputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;proc&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;padding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;return_tensors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;to&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cuda&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;no_grad&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;inputs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;do_sample&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;guidance_scale&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;3.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                               &lt;span class="n"&gt;max_new_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;TOKENS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;cpu&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;numpy&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;abs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="mf"&gt;1e-6&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.9&lt;/span&gt;
    &lt;span class="n"&gt;wav&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;music-option-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;tag&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;sr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;32767&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three 30-second cues, at 50 audio tokens per second. MusicGen's practical ceiling for coherent output is around 30 seconds; asking for 200 produces something that wanders.&lt;/p&gt;

&lt;p&gt;The naive next step — loop one 30-second cue twenty times — produces three minutes of unchanging texture that actively fights the film. The moon reveal and the goodnight need different music.&lt;/p&gt;

&lt;p&gt;So the bed is &lt;em&gt;arranged&lt;/em&gt;, using the two cues that survived audition (&lt;code&gt;b&lt;/code&gt;, the adventure spine; &lt;code&gt;c&lt;/code&gt;, wonder and emotion) laid out against the shot map:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Section&lt;/th&gt;
&lt;th&gt;Shots&lt;/th&gt;
&lt;th&gt;Cue&lt;/th&gt;
&lt;th&gt;Length&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Adventure&lt;/td&gt;
&lt;td&gt;1–8&lt;/td&gt;
&lt;td&gt;b&lt;/td&gt;
&lt;td&gt;82 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wonder + the quiet beat&lt;/td&gt;
&lt;td&gt;9–14&lt;/td&gt;
&lt;td&gt;c&lt;/td&gt;
&lt;td&gt;62 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Triumph&lt;/td&gt;
&lt;td&gt;15–18&lt;/td&gt;
&lt;td&gt;b&lt;/td&gt;
&lt;td&gt;42 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home&lt;/td&gt;
&lt;td&gt;19–20&lt;/td&gt;
&lt;td&gt;c&lt;/td&gt;
&lt;td&gt;20 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each section is built by crossfading a cue into itself — a hard loop point is audible, a 2-second triangular crossfade is not — and then the four sections are crossfaded together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;X&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2   &lt;span class="c"&gt;# crossfade seconds&lt;/span&gt;

&lt;span class="c"&gt;# seamless_loop &amp;lt;src&amp;gt; &amp;lt;copies&amp;gt; &amp;lt;trim_len&amp;gt; &amp;lt;out&amp;gt;&lt;/span&gt;
seamless_loop &lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt; &lt;span class="nv"&gt;copies&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$2&lt;/span&gt; &lt;span class="nv"&gt;len&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$3&lt;/span&gt; &lt;span class="nv"&gt;out&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$4&lt;/span&gt;
  &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt; &lt;span class="nb"&gt;fc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt; &lt;span class="nv"&gt;cur&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[0:a]"&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq &lt;/span&gt;0 &lt;span class="k"&gt;$((&lt;/span&gt;copies-1&lt;span class="k"&gt;))&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;&lt;span class="nv"&gt;inputs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$inputs&lt;/span&gt;&lt;span class="s2"&gt; -i &lt;/span&gt;&lt;span class="nv"&gt;$src&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
  for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq &lt;/span&gt;1 &lt;span class="k"&gt;$((&lt;/span&gt;copies-1&lt;span class="k"&gt;))&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    &lt;/span&gt;&lt;span class="nb"&gt;fc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fc&lt;/span&gt;&lt;span class="k"&gt;}${&lt;/span&gt;&lt;span class="nv"&gt;cur&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;[&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;i&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:a]acrossfade=d=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;X&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:c1=tri:c2=tri[x&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;i&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;];"&lt;/span&gt;
    &lt;span class="nv"&gt;cur&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[x&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;i&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;]"&lt;/span&gt;
  &lt;span class="k"&gt;done&lt;/span&gt;
  &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$FF&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="nv"&gt;$inputs&lt;/span&gt; &lt;span class="nt"&gt;-filter_complex&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fc&lt;/span&gt;&lt;span class="k"&gt;}${&lt;/span&gt;&lt;span class="nv"&gt;cur&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;atrim=0:&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;len&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;,asetpts=PTS-STARTPTS,aresample=48000[o]"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-map&lt;/span&gt; &lt;span class="s2"&gt;"[o]"&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt;:a pcm_s16le &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$out&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

seamless_loop music-option-b.mp3 3 82 _s1.wav   &lt;span class="c"&gt;# shots 1-8   adventure&lt;/span&gt;
seamless_loop music-option-c.mp3 3 62 _s2.wav   &lt;span class="c"&gt;# shots 9-14  wonder + the quiet beat&lt;/span&gt;
seamless_loop music-option-b.mp3 2 42 _s3.wav   &lt;span class="c"&gt;# shots 15-18 triumph&lt;/span&gt;
seamless_loop music-option-c.mp3 1 20 _s4.wav   &lt;span class="c"&gt;# shots 19-20 home&lt;/span&gt;

&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$FF&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; _s1.wav &lt;span class="nt"&gt;-i&lt;/span&gt; _s2.wav &lt;span class="nt"&gt;-i&lt;/span&gt; _s3.wav &lt;span class="nt"&gt;-i&lt;/span&gt; _s4.wav &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-filter_complex&lt;/span&gt; &lt;span class="s2"&gt;"[0][1]acrossfade=d=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;X&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:c1=tri:c2=tri[a];&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
[a][2]acrossfade=d=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;X&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:c1=tri:c2=tri[b];&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
[b][3]acrossfade=d=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;X&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:c1=tri:c2=tri[c];&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
[c]loudnorm=I=-22:TP=-3,atrim=0:200.05[out]"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-map&lt;/span&gt; &lt;span class="s2"&gt;"[out]"&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt;:a libmp3lame &lt;span class="nt"&gt;-b&lt;/span&gt;:a 192k music.mp3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The bed is normalised to &lt;strong&gt;−22 LUFS&lt;/strong&gt; — deliberately quiet, because it is going to sit under narration and clip audio in the final mix and get normalised again there.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — no score at all.&lt;/strong&gt; Veo's native audio already carries the film. The music-only cut exists because the score turned out well, not because the picture needed rescuing. If you are on a laptop with no GPU, cutting this entire section is a legitimate choice and costs you less than you would think.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — library music.&lt;/strong&gt; The YouTube Audio Library, the Free Music Archive and Incompetech are free, cleared, and vastly better recorded than anything MusicGen produces. What you give up is exactly the thing this section is about: a library track does not know where your reveal is. You would be crossfading someone else's cues against your beat map instead of your own, which is the same arrangement work with less control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — MusicGen's melody-conditioned and continuation modes.&lt;/strong&gt; Those are the routes to genuine through-composition instead of four crossfaded loops. My arrangement is the cheap eighty per cent: it changes when the story changes and it never repeats audibly, and that was enough. It is not the same thing as a score that develops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — Stable Audio Open, or MusicGen Small on CPU.&lt;/strong&gt; The Small model runs without a GPU at maybe a minute of compute per thirty seconds of audio. Three cues is a coffee break, not an obstacle.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — scoring it properly.&lt;/strong&gt; A person with a keyboard and a free DAW will beat all of the above in an afternoon, and I mention it because it is easy to forget that the non-AI option is still sitting there.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc72p9apwbezfk954ajmb.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc72p9apwbezfk954ajmb.webp" alt="Comet blazing with gold light beside the moon, high in the glowing oak" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 16, where cue &lt;code&gt;b&lt;/code&gt; comes back for the triumph section. The tree is lit by Comet, not by the moon — the payoff of a power established eight shots earlier.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The voice you are not allowed to clone
&lt;/h2&gt;

&lt;p&gt;The film works on Veo's own dialogue and the score alone. But a single constant narrator across all twenty shots gives the ear something stable to hold, and the drifting character voices become colour rather than the spine of the film. So I recorded a narration, and then cloned it.&lt;/p&gt;

&lt;h3&gt;
  
  
  English: F5-TTS
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/SWivid/F5-TTS" rel="noopener noreferrer"&gt;F5-TTS&lt;/a&gt; is a flow-matching TTS with zero-shot voice cloning: give it a reference clip and its transcript, and it generates arbitrary new text in that voice. The reference is one sentence from my own recording:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;REF_TEXT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;At the top of the puddle hills stood the biggest oak tree in the world, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;and in it was the moon.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;tts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;F5TTS&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;lines&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;tts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;infer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ref_file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;../04_AUDIO/vo_raw/_ref.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ref_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;REF_TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
              &lt;span class="n"&gt;gen_text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;file_wave&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;narr-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;remove_silence&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Narration is stored as a machine-readable &lt;code&gt;NN|text&lt;/code&gt; file keyed to shot number, so the same lines drive every language and every voice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;01|One night, the sky went dark. Very, very dark.
02|This is Liya. She is four. And this is Eye-rick. He is two. Every night, they turn into superheroes!
...
19|Up and up went the moon, all the way home. Click!
20|The moon said thank you, and tucked them into bed. Goodnight, Liya. Goodnight, Eye-rick.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the phonetic respellings — "Liya", "Eye-rick". The models mispronounce the real spellings, and this file is read by machines, not humans; the human reading scripts are separate files with modulation marks.&lt;/p&gt;

&lt;p&gt;Only fifteen of the twenty slots carry narration. Shots 3, 4, 5, 10 and 15 are left alone: those are the catchphrase and dialogue shots, and talking over them would be vandalism.&lt;/p&gt;

&lt;p&gt;Each generated line then gets fitted to its 10-second slot, with the time-stretch capped so it stays natural:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;d&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;ffprobe &lt;span class="nt"&gt;-v&lt;/span&gt; error &lt;span class="nt"&gt;-show_entries&lt;/span&gt; &lt;span class="nv"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;duration &lt;span class="nt"&gt;-of&lt;/span&gt; &lt;span class="nv"&gt;csv&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;p&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;t&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;python &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"d=&lt;/span&gt;&lt;span class="nv"&gt;$d&lt;/span&gt;&lt;span class="s2"&gt;; print(round(min(max(d/9.2,1.0),1.15),4))"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s2"&gt;"1.0"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; ffmpeg &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-filter&lt;/span&gt;:a &lt;span class="s2"&gt;"atempo=&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; _t.wav &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;mv&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; _t.wav &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A cap of 1.15× is the point past which a bedtime read starts sounding hurried. If a line will not fit at 1.15×, the correct fix is to shorten the line, not to speed it up further.&lt;/p&gt;

&lt;p&gt;Lines are then placed at their shot's start time — shot &lt;em&gt;N&lt;/em&gt; begins at (&lt;em&gt;N&lt;/em&gt;−1)×10 s — with a 0.4 s lead-in so the picture establishes before the voice starts, and mixed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;FILTER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$FILTER&lt;/span&gt;&lt;span class="s2"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;$IDX&lt;/span&gt;&lt;span class="s2"&gt;:a]aresample=48000,adelay=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;off&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;|&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;off&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;[d&lt;/span&gt;&lt;span class="nv"&gt;$IDX&lt;/span&gt;&lt;span class="s2"&gt;];"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The ElevenLabs finding
&lt;/h3&gt;

&lt;p&gt;I evaluated &lt;a href="https://elevenlabs.io" rel="noopener noreferrer"&gt;ElevenLabs&lt;/a&gt; as an alternative, and two things came back that shaped the project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice cloning is not on the free tier.&lt;/strong&gt; Instant Voice Cloning starts at the paid Starter plan; Professional Voice Cloning starts at Creator. The free tier gets text-to-speech with stock voices only. That is why one of the six cuts is narrated by "Sarah", a stock voice — it was the free-tier option, and it is genuinely good, but it is not anyone's parent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cloning a minor's voice is prohibited outright.&lt;/strong&gt; The &lt;a href="https://elevenlabs.io/use-policy" rel="noopener noreferrer"&gt;ElevenLabs use policy&lt;/a&gt; forbids replicating a person's voice "without consent or legal right", forbids material designed to impersonate a minor, and restricts service availability to minors. Parental consent does not open a door here; the prohibition is on the output.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The children's voices in this film are entirely Veo's invention. Only the adults' voices — mine and their mother's — were ever cloned.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is the right line, and I would have drawn it there myself, but it is worth knowing before you plan a pipeline around it. If you want a child's voice in a film: generate a fictional one, or record the child directly. Do not clone.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — free local cloning stacks.&lt;/strong&gt; F5-TTS is not the only one: XTTS-v2, OpenVoice and Chatterbox all do zero-shot cloning on consumer hardware. Read the licence before anything commercial — some of these ship permissive code with non-commercial weights, and the two are separately licensed. None of them change the rule above: &lt;strong&gt;do not clone a child's voice&lt;/strong&gt;, whichever tool makes it easy, and get an adult's explicit agreement before cloning theirs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — no cloning at all.&lt;/strong&gt; Piper runs on a CPU, has good stock voices and is genuinely instant; Kokoro is small and unusually natural for its size. Neither will sound like you. For a narrator, that may not matter — one of my six cuts is narrated by an ElevenLabs stock voice and it is perfectly good, it is just not anybody's parent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tried, and it won — recording it yourself.&lt;/strong&gt; The cheapest option in this entire pipeline is a phone and a quiet-ish room, and it produced the best voice in the film. See the grandmother section below; the numbers are not close. If you have any way at all to record a real person reading it, do that first and treat cloning as the fallback for when you cannot.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Malayalam: IndicF5
&lt;/h3&gt;

&lt;p&gt;Their mother's family speaks Malayalam, so two of the six cuts are Malayalam — one read by their grandmother, one in my cloned voice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://huggingface.co/ai4bharat/IndicF5" rel="noopener noreferrer"&gt;IndicF5&lt;/a&gt; from AI4Bharat is an F5-architecture model trained on 1,400+ hours across 11 Indian languages, Malayalam included, MIT-licensed. It ships as a HuggingFace repo with a bundled &lt;code&gt;f5_tts&lt;/code&gt; package.&lt;/p&gt;

&lt;p&gt;Getting it running produced the most interesting failure in the entire project.&lt;/p&gt;




&lt;h2&gt;
  
  
  Three times too long, and full of things I never wrote
&lt;/h2&gt;

&lt;p&gt;Every Malayalam line came out roughly three times longer than it should have been. And the surplus was not silence. It was not noise either.&lt;/p&gt;

&lt;p&gt;It was &lt;strong&gt;fluent, confident, entirely invented Malayalam&lt;/strong&gt; — the model calmly narrating sentences that appeared nowhere in the input, in my own cloned voice, to my children.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where it comes from
&lt;/h3&gt;

&lt;p&gt;F5-TTS estimates how much audio to generate from the length of the target text. The estimate is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For English, UTF-8 is 1 byte per character and this is a perfectly fine proxy for duration.&lt;/p&gt;

&lt;p&gt;Malayalam lives in the &lt;code&gt;U+0D00&lt;/code&gt; block. &lt;strong&gt;Three bytes per character.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So a 60-character Malayalam sentence measures as 180 "characters". Scaled against an English reference clip, the model is instructed to produce roughly 3× the audio the text actually needs — and a flow-matching TTS asked to fill a duration will fill it. It does not stop early. It generates plausible speech until the buffer is full.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is a byte-length-as-proxy-for-text-length bug — the same family as truncating a UTF-8 string at a byte offset. It is completely invisible until the script changes.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The fix
&lt;/h3&gt;

&lt;p&gt;Estimate from &lt;strong&gt;akshara count&lt;/strong&gt; instead — Malayalam's orthographic units. An akshara is a base character optionally followed by combining marks (vowel signs, virama), and the combining marks do not add duration. So: normalise to NFC, count characters that are neither combining marks nor whitespace.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;UNITS_PER_SEC&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;6.6&lt;/span&gt;   &lt;span class="c1"&gt;# Malayalam aksharas per second; lower = slower/longer
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;akshara_count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ch&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;unicodedata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NFC&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
               &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;unicodedata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;combining&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isspace&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;est_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;akshara_count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;UNITS_PER_SEC&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;0.35&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;6.6 aksharas per second was tuned by ear against a test line; the +0.35 s is trailing room so the final word is not clipped. The estimate is then passed explicitly as &lt;code&gt;fix_duration&lt;/code&gt;, computed as the reference clip's duration plus the target's:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;infer_batch_process&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ref_t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sr_in&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;ref_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vocoder&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;mel_spec_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vocos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;speed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;speed&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;fix_duration&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ref_dur&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;gen_sec&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  And then it came back
&lt;/h3&gt;

&lt;p&gt;Calling &lt;code&gt;infer_process()&lt;/code&gt; — the obvious high-level entry point — reintroduces the whole problem from a different direction. It re-chunks the input on a &lt;strong&gt;byte-based &lt;code&gt;max_chars&lt;/code&gt;&lt;/strong&gt; and then applies &lt;code&gt;fix_duration&lt;/code&gt; to &lt;em&gt;each&lt;/em&gt; resulting sub-batch, so a line that got split in two comes out at twice the intended length.&lt;/p&gt;

&lt;p&gt;The fix is to chunk it yourself, on sentence boundaries, sized by your own duration estimate, and call &lt;code&gt;infer_batch_process&lt;/code&gt; directly with exactly one batch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;split_sentences&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_sec&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;10.0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Split on Malayalam/Latin sentence enders, regrouping into &amp;lt;= max_sec chunks.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(?&amp;lt;=[.!?।])\s+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;cand&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;est_seconds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cand&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;max_sec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;
        &lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cand&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cur&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Chunks are concatenated with a 0.18 s gap and per-chunk silence trimming.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two more IndicF5 mechanics
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Namespace-package shadowing.&lt;/strong&gt; The IndicF5 mirror bundles its own &lt;code&gt;f5_tts&lt;/code&gt; package, which must take precedence over the pip-installed F5-TTS. Because &lt;code&gt;f5_tts&lt;/code&gt; is a namespace package, Python merrily merges both directories on &lt;code&gt;sys.path&lt;/code&gt; and you get a hybrid that imports without error and behaves wrongly — the worst possible failure mode. Prepending to &lt;code&gt;sys.path&lt;/code&gt; is not enough; you have to pin &lt;code&gt;__path__&lt;/code&gt; and assert it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;REPO&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;f5_tts&lt;/span&gt;
&lt;span class="n"&gt;f5_tts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__path__&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;REPO&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;f5_tts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;   &lt;span class="c1"&gt;# namespace pkg: pin to bundled copy
&lt;/span&gt;&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abspath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f5_tts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__path__&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;REPO&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrong f5_tts: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f5_tts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;__path__&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Checkpoint key prefixes.&lt;/strong&gt; The published &lt;code&gt;model.safetensors&lt;/code&gt; is a saved &lt;code&gt;torch.compile&lt;/code&gt;d wrapper. Its keys are prefixed &lt;code&gt;ema_model._orig_mod.&lt;/code&gt; and it also bundles the vocoder weights, while the F5-TTS loader expects bare EMA keys. Strip and filter once, cache the result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;safetensors.torch&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_file&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;save_file&lt;/span&gt;
&lt;span class="n"&gt;pre&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ema_model._orig_mod.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;sd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;load_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pre&lt;/span&gt;&lt;span class="p"&gt;):]:&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;contiguous&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;sd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pre&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no ema_model._orig_mod.* keys found&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="nf"&gt;save_file&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;And a platform note.&lt;/strong&gt; &lt;a href="https://pytorch.org/get-started/locally/" rel="noopener noreferrer"&gt;PyTorch&lt;/a&gt; on a Blackwell GPU (sm_120) needs the CUDA 12.8 wheels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; venv .venv-tts
.venv-tts/Scripts/python.exe &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install &lt;/span&gt;torch torchaudio &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--index-url&lt;/span&gt; https://download.pytorch.org/whl/cu128
.venv-tts/Scripts/python.exe &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install &lt;/span&gt;f5-tts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also, &lt;code&gt;torchaudio&lt;/code&gt;'s current I/O path goes through torchcodec, which wants FFmpeg ≤ 7 shared libraries; the FFmpeg 9 static build on this machine is incompatible. Rather than maintaining a second FFmpeg, a ten-line shim routes &lt;code&gt;torchaudio.load&lt;/code&gt;/&lt;code&gt;save&lt;/code&gt; through &lt;code&gt;soundfile&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# route torchaudio I/O through soundfile: torchcodec needs ffmpeg&amp;lt;=7 shared libs, we have 9 static
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;soundfile&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sf&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;torchaudio&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;float32&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;always_2d&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_numpy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ascontiguousarray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;T&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="n"&gt;sr&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sample_rate&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;detach&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;cpu&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;numpy&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ndim&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;T&lt;/span&gt;
    &lt;span class="n"&gt;sf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sample_rate&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;torchaudio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;load&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;torchaudio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;save&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_load&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_save&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — a forced aligner instead of a duration estimate.&lt;/strong&gt; The whole bug exists because the model &lt;em&gt;guesses&lt;/em&gt; how long the text should take. Montreal Forced Aligner and aeneas measure it against real audio instead. That does not help you generate speech, but it is the right tool for anything downstream that needs to know where words land, and it would have made the subtitle timing a solved problem rather than a clever one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — other Indic TTS.&lt;/strong&gt; AI4Bharat publish more than IndicF5, and the large cloud providers have Malayalam voices on free tiers. If your target language is not English, budget an evening for the possibility that the obvious model simply does not speak it well.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — skipping synthesis entirely for the second language.&lt;/strong&gt; The Malayalam cut that people actually prefer is the one their grandmother read into a phone. Cloning my own voice into Malayalam was the technically interesting path and the second-best result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The transferable rule, which cost me a day:&lt;/strong&gt; synthesise &lt;strong&gt;one line&lt;/strong&gt; in the target script and listen to it before you build anything around the model. Three bytes per character would have been obvious in five minutes and instead it surfaced across a full batch of fluent, invented, confidently-narrated Malayalam.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Fluent, and completely flat
&lt;/h3&gt;

&lt;p&gt;I fixed the bug. The lines came back the right length, in my own voice, saying exactly what I had written. It worked.&lt;/p&gt;

&lt;p&gt;It was also lifeless, and that is the part I want you to hear rather than take my word for.&lt;/p&gt;

&lt;p&gt;Fluent is not the same as good. A model trained to fill a duration gives you the right phonemes and almost no performance: even stress, even pace, no lift into a question, no warmth landing on a child's name. Measured against the real read, the clone covers &lt;strong&gt;5.1 semitones of pitch range to the recording's 10.0&lt;/strong&gt; — less than half the melody. It is a voice that is technically mine and expressively nobody's.&lt;/p&gt;

&lt;p&gt;It was slow, too. Every re-render was minutes rather than seconds, so tuning fifteen lines was an evening's work, not a loop you could spin quickly. I spent that evening chasing warmth that the method could not produce.&lt;/p&gt;

&lt;p&gt;Listen to the same line twice. This is shot 2, in both languages — the clone first, then the person.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Malayalam, shot 2&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ഇത് ലിയ. അവൾക്ക് നാല് വയസ്സ്. ഇത് ഐറിക്. അവന് രണ്ട് വയസ്സ്. രാത്രിയായാൽ അവരുടെ ലോകം സൂപ്പർ ഹീറോകളുടെ ലോകമാണ്.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is Lia. She is four. This is Airik. He is two. At night, their world becomes a world of superheroes.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;IndicF5 clone&lt;/strong&gt; — my voice, cloned into Malayalam:
&lt;a href="https://www.razi.pro/audio/blog/star-patrol/ml-02-ai-indicf5.mp3" rel="noopener noreferrer"&gt;https://www.razi.pro/audio/blog/star-patrol/ml-02-ai-indicf5.mp3&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real recording&lt;/strong&gt; — read into a phone, in a living room:
&lt;a href="https://www.razi.pro/audio/blog/star-patrol/ml-02-real-amma.mp3" rel="noopener noreferrer"&gt;https://www.razi.pro/audio/blog/star-patrol/ml-02-real-amma.mp3&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://www.razi.pro/blog/making-an-ai-animated-film-with-veo-and-flow#fluent-and-completely-flat" rel="noopener noreferrer"&gt;▶ Play both takes side by side on razi.pro&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;English, shot 2&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is Liya. She is four. And this is Eye-rick. He is two. Every night, they turn into superheroes!&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ElevenLabs stock voice&lt;/strong&gt; — the best synthetic take of the three I tried:
&lt;a href="https://www.razi.pro/audio/blog/star-patrol/en-02-ai-elevenlabs.mp3" rel="noopener noreferrer"&gt;https://www.razi.pro/audio/blog/star-patrol/en-02-ai-elevenlabs.mp3&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real recording&lt;/strong&gt; — one continuous read, cut up afterwards:
&lt;a href="https://www.razi.pro/audio/blog/star-patrol/en-02-real-dad.mp3" rel="noopener noreferrer"&gt;https://www.razi.pro/audio/blog/star-patrol/en-02-real-dad.mp3&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://www.razi.pro/blog/making-an-ai-animated-film-with-veo-and-flow#fluent-and-completely-flat" rel="noopener noreferrer"&gt;▶ Play both takes side by side on razi.pro&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The synthetic takes are not bad. They are competent, clear, and completely uninterested in the story. The recordings have someone in them who knows these two children.&lt;/p&gt;

&lt;p&gt;So the film ships human voices. Every clone I built became an alternate cut, and the two versions in the FINAL folder are both real reads — which quietly vindicated a working rule I had written on day one, before I understood why it mattered: &lt;strong&gt;record the narration yourself rather than relying on generated voices.&lt;/strong&gt; I had meant it as a note about consistency. It turned out to be a note about feeling.&lt;/p&gt;

&lt;p&gt;That is not an argument against voice cloning. It is an argument about what it is for. Cloning solved a problem I actually had — a Malayalam cut in my own voice, which I cannot perform — and it solved it well enough to ship as an option. What it could not do was care.&lt;/p&gt;




&lt;h2&gt;
  
  
  Six films from one master
&lt;/h2&gt;

&lt;p&gt;A film is not finished when it looks finished.&lt;/p&gt;

&lt;p&gt;Twenty clips existed. The score existed. What did not exist was a &lt;em&gt;file&lt;/em&gt; — something with fades and a mix and captions that a person could open on a phone in another country and understand. That gap took longer than the shooting.&lt;/p&gt;

&lt;p&gt;Six versions come out of this chapter, all frame-identical, differing only in narration. Almost everything below exists because six is not one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Picture
&lt;/h3&gt;

&lt;p&gt;The twenty clips are joined with &lt;a href="https://ffmpeg.org/ffmpeg-formats.html" rel="noopener noreferrer"&gt;FFmpeg's concat demuxer&lt;/a&gt;, which is stream-copy — no re-encode, no generation loss — producing &lt;code&gt;_master-assembly-nofades.mp4&lt;/code&gt; at 200.042667 s. Every deliverable is encoded from that one file, so all six cuts are frame-identical.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-f&lt;/span&gt; concat &lt;span class="nt"&gt;-safe&lt;/span&gt; 0 &lt;span class="nt"&gt;-i&lt;/span&gt; shots.txt &lt;span class="nt"&gt;-c&lt;/span&gt; copy _master-assembly-nofades.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This works because all twenty clips came out of the same model at the same settings and therefore share codec parameters exactly. It is the reason to keep the pipeline uniform.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.razi.pro/videos/star-patrol/mega-bounce-loop.mp4" rel="noopener noreferrer"&gt;Video: Animated GIF: the huge combined bounce, Boing launched high toward the moon&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 17, the mega bounce. &lt;code&gt;LRA=11&lt;/code&gt; in the final loudness normalisation exists specifically so this still feels big after everything else has been levelled.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The mix
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;build-final.sh&lt;/code&gt; takes the master, a music bed, and optionally a narration track, and produces one deliverable. The whole mix is a single &lt;code&gt;filter_complex&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Clip audio is attenuated by how much else is competing with it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt;   &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$NARR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="k"&gt;then &lt;/span&gt;&lt;span class="nv"&gt;CLIPVOL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.45     &lt;span class="c"&gt;# narration present: clip audio is texture&lt;/span&gt;
&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MUSIC&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then &lt;/span&gt;&lt;span class="nv"&gt;CLIPVOL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.85
&lt;span class="k"&gt;else                       &lt;/span&gt;&lt;span class="nv"&gt;CLIPVOL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1.0&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The music bed is looped to length, faded, and mixed under the clip audio:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;INPUTS+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="nt"&gt;-stream_loop&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MUSIC&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;IDX&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:a]volume=0.45,afade=t=in:st=0:d=3,afade=t=out:st=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;FOUT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:d=1.5,atrim=0:&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;DUR&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;,asetpts=PTS-STARTPTS[mus];"&lt;/span&gt;
FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[clip][mus]amix=inputs=2:duration=first:normalize=0[bed];"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;normalize=0&lt;/code&gt; on every &lt;code&gt;amix&lt;/code&gt; is important: by default &lt;code&gt;amix&lt;/code&gt; divides by the number of inputs, which would silently halve everything each time a layer is added. Levels here are set deliberately, so the automatic normalisation is switched off.&lt;/p&gt;

&lt;p&gt;Narration is the interesting part. The bed is &lt;strong&gt;sidechain-compressed against the narration&lt;/strong&gt;, so music and clip audio duck automatically whenever the narrator speaks and come back up when they stop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;INPUTS+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$NARR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;IDX&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:a]volume=1.3,aresample=48000[narr];"&lt;/span&gt;
FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;IDX&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:a]volume=1.3,aresample=48000[narrsc];"&lt;/span&gt;
FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BED&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;[narrsc]sidechaincompress=threshold=0.04:ratio=9:attack=15:release=350[duck];"&lt;/span&gt;
FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"[duck][narr]amix=inputs=2:duration=first:normalize=0[mixed];"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The narration is split into two identical branches — &lt;code&gt;[narr]&lt;/code&gt; goes into the mix, &lt;code&gt;[narrsc]&lt;/code&gt; is the sidechain key — because &lt;a href="https://ffmpeg.org/ffmpeg-filters.html" rel="noopener noreferrer"&gt;&lt;code&gt;sidechaincompress&lt;/code&gt;&lt;/a&gt; consumes its key input. The parameters: &lt;code&gt;ratio=9&lt;/code&gt; is aggressive, because the goal is intelligibility rather than subtlety; &lt;code&gt;attack=15&lt;/code&gt; ms ducks fast enough not to clip the narrator's first syllable; &lt;code&gt;release=350&lt;/code&gt; ms is slow enough that the music does not pump between words within a sentence.&lt;/p&gt;

&lt;p&gt;Finally, &lt;a href="https://ffmpeg.org/ffmpeg-filters.html" rel="noopener noreferrer"&gt;loudness normalisation&lt;/a&gt;, a safety limiter, and matched audio/video fades:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;FC+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BED&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;loudnorm=I=-16:TP=-1.5:LRA=11,alimiter=limit=0.97,afade=t=in:st=0:d=1.5,afade=t=out:st=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;FOUT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:d=1.5[aout]"&lt;/span&gt;

&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$FFMPEG&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;INPUTS&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-filter_complex&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$FC&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-map&lt;/span&gt; 0:v &lt;span class="nt"&gt;-map&lt;/span&gt; &lt;span class="s2"&gt;"[aout]"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="s2"&gt;"fade=t=in:st=0:d=1.5,fade=t=out:st=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;FOUT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:d=1.5"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt;:v libx264 &lt;span class="nt"&gt;-crf&lt;/span&gt; 18 &lt;span class="nt"&gt;-preset&lt;/span&gt; medium &lt;span class="nt"&gt;-pix_fmt&lt;/span&gt; yuv420p &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt;:a aac &lt;span class="nt"&gt;-b&lt;/span&gt;:a 192k &lt;span class="nt"&gt;-ar&lt;/span&gt; 48000 &lt;span class="nt"&gt;-movflags&lt;/span&gt; +faststart &lt;span class="nt"&gt;-t&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$DUR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$OUT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;−16 LUFS&lt;/strong&gt; with a −1.5 dBTP ceiling is the streaming-delivery target and the right choice for something played back on a tablet or a TV at bedtime. Video fade in and out are 1.5 s, matched to the audio fades, starting at 198.5 s.&lt;/p&gt;

&lt;p&gt;Building a deliverable is then one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;NARR_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;narration-ml-amma.wav &lt;span class="nv"&gt;NARR_OUT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"Star Patrol (Malayalam - Amma).mp4"&lt;/span&gt; ./build-final.sh
./build-final.sh    &lt;span class="c"&gt;# music-only version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six cuts, one script, one master. The script itself came out of a terminal session rather than my own hands — The pipeline had a fourth author.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — a timeline editor.&lt;/strong&gt; DaVinci Resolve's free edition and Shotcut both do everything in this section with a mouse — as would Premiere Pro, Final Cut Pro, CapCut or InShot — and for a &lt;em&gt;single&lt;/em&gt; deliverable they are unquestionably the faster route. The script pays for itself at cut number two: six frame-identical versions, in six languages and voices, from one command and no chance of a drifting edit between them. If you are only ever shipping one file, use the GUI and skip this whole section without guilt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — manual ducking instead of a sidechain.&lt;/strong&gt; Drawing volume automation under the narration by hand is what an editor would do, and it sounds better than any compressor because a human knows which word matters. It also has to be redone for every language. Sidechaining is the choice that scales; automation is the choice that sounds best.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tried — leaving the clip audio alone.&lt;/strong&gt; The music-only cut runs Veo's native audio at full level with the score under it, and it is the version I would show someone who has never seen the film. Not every deliverable needs the whole mix.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Subtitles nobody asked me to burn in
&lt;/h3&gt;

&lt;p&gt;Every deliverable ships with a &lt;code&gt;.srt&lt;/code&gt; beside it. None of them have the subtitles burned into the picture, and that was a decision rather than an omission.&lt;/p&gt;

&lt;p&gt;I did not write the subtitle timing either. The cue splitting, the character-per-second limits and the two-language variants all came out of the same terminal session as the rest of &lt;code&gt;05_TOOLS/&lt;/code&gt; — I specified the constraints and corrected the output against the film.&lt;/p&gt;

&lt;p&gt;Burned-in subtitles are permanent. They cannot be turned off, they cannot be restyled by a player that knows more about the viewer's screen than I do, they are baked at one resolution and re-scaled badly at every other, and they make the file untranslatable — the Malayalam cut and the English cut share a picture master precisely so that all six deliverables are frame-identical, and burning text in would fork the picture six ways. A sidecar &lt;code&gt;.srt&lt;/code&gt; costs two kilobytes and keeps every one of those options open.&lt;/p&gt;

&lt;p&gt;The timing is the part worth stealing. The obvious approach is to estimate: count characters, divide by a reading rate, hope. That is the same class of mistake as the byte-length duration bug, and it fails the same way — silently, and worse on the language you can't proofread.&lt;/p&gt;

&lt;p&gt;So nothing is estimated. Each narration line is anchored to its shot slot by the same arithmetic the audio uses — line for shot &lt;em&gt;N&lt;/em&gt; starts at &lt;code&gt;(N-1) * 10 + 0.4&lt;/code&gt; — and then every &lt;em&gt;word&lt;/em&gt; inside that line inherits a measured timestamp from the Scribe transcription of the actual read:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;slot&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SLOT&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;offset&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;slot&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;rd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;SHOT&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;rd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;LEAD_IN&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;s_src&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;s_src&lt;/code&gt; is where the line began in the raw take. The line was cut there and placed at its slot, so a word's position in the finished film is just its measured position in the source plus that constant offset. The subtitle and the voice cannot drift apart, because they are derived from the same measurement.&lt;/p&gt;

&lt;p&gt;One refinement that matters more than it sounds: the cue text comes from &lt;strong&gt;the script&lt;/strong&gt;, not from the transcriber. Scribe heard "Leah", "Iric" and "Quk"; the script says "Lia", "Airik" and "quick". Subtitling the ASR output would put the recogniser's spelling of my children's names on screen. Instead the script's tokens are aligned to the ASR's tokens with &lt;code&gt;difflib.SequenceMatcher&lt;/code&gt;, inherit their times, and anything unmatched is interpolated between its known neighbours.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Take the timing from the machine and the words from the human. Never both from the machine.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Cues are then wrapped to 42 characters and two lines (38 for Malayalam — the glyphs are wider), split at sentence ends where possible, capped at 6 seconds, and nudged so no cue ever overlaps the next by less than 40 ms.&lt;/p&gt;

&lt;p&gt;And the practical detail that makes all of it work: &lt;strong&gt;name the &lt;code&gt;.srt&lt;/code&gt; to match the &lt;code&gt;.mp4&lt;/code&gt; byte for byte, basename included.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Star Patrol (English - Dad FINAL).mp4
Star Patrol (English - Dad FINAL).srt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;VLC, mpv, Plex, Jellyfin and the Windows and macOS system players all load a sidecar subtitle automatically on that rule alone. There is no metadata, no muxing, no configuration. &lt;code&gt;Star Patrol (English - Dad FINAL) subs.srt&lt;/code&gt; does not load. The space before &lt;code&gt;subs&lt;/code&gt; is the entire difference between a subtitle that appears and one that does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The burn-in I tried anyway.&lt;/strong&gt; For completeness I did attempt a hard-subbed Malayalam cut, and it failed twice for two unrelated reasons, both worth knowing before you spend an evening on them.&lt;/p&gt;

&lt;p&gt;First, the filter path. FFmpeg's &lt;code&gt;subtitles&lt;/code&gt; filter takes a &lt;code&gt;fontsdir&lt;/code&gt;, and the filter-graph parser treats &lt;code&gt;:&lt;/code&gt; as the option separator — so a Windows path detonates on its own drive letter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;-vf "subtitles=subs.srt:fontsdir=C:/Windows/Fonts"
                                  ^ parsed as the end of the fontsdir option
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Escaping it (&lt;code&gt;C\\:/Windows/Fonts&lt;/code&gt;) works sometimes and depends on how many layers of shell and filter-graph quoting the string has already survived, which on Git Bash under Windows is more than you would like. The reliable fix is the one the watermark script below already uses: stage the font into the working directory and pass a plain relative path, because a relative path is the one form that no parser in the stack tries to rewrite.&lt;/p&gt;

&lt;p&gt;Second, and fatally, the glyphs. Malayalam is a complex script — conjuncts, reordered vowel signs, the lot — and the default font libass reaches for does not have it. You get boxes, or you get correctly-spaced garbage, which is worse because it looks like it worked. Malayalam on Windows needs &lt;strong&gt;Nirmala UI&lt;/strong&gt;, and it needs a shaper that will actually reorder the marks.&lt;/p&gt;

&lt;p&gt;Both problems are solvable. Neither is solvable &lt;em&gt;quickly&lt;/em&gt;, and both of them are arguments for the sidecar file, which hands the rendering to a player that already has a font stack and already knows the viewer's screen. The &lt;code&gt;.srt&lt;/code&gt; was not a compromise. It was the better answer that I only fully appreciated after trying the other one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — soft subtitles muxed into the MP4.&lt;/strong&gt; &lt;code&gt;-c:s mov_text&lt;/code&gt; puts the captions &lt;em&gt;inside&lt;/em&gt; the file as a track the viewer can switch off, which gets you one file to send someone without giving up any of the advantages of not burning in. It is the obvious middle road and I simply did not need it, because everything here is played from a folder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — letting YouTube do it.&lt;/strong&gt; Upload the &lt;code&gt;.srt&lt;/code&gt;, or upload nothing and edit the auto-captions in YouTube's own editor. Zero tooling, and for a film that lives on YouTube anyway it is hard to argue with. You lose the sidecar file for local playback, which is the only reason I did it myself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — an aligner instead of my trigram matching.&lt;/strong&gt; See the note in the Malayalam section. &lt;code&gt;difflib.SequenceMatcher&lt;/code&gt; against ASR tokens is a workaround for not having a forced aligner; it works, but a real aligner is the tool for this job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worth keeping whatever you choose:&lt;/strong&gt; take the timing from the machine and the words from the human. That rule survives every one of these routes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  A watermark that will not sit still
&lt;/h3&gt;

&lt;p&gt;Every deliverable carries a drifting &lt;code&gt;razi.pro&lt;/code&gt; watermark, applied during the build. This is the least glamorous section in the post and the one I would most want if I were shipping something to the internet.&lt;/p&gt;

&lt;p&gt;The whole mark is defined in exactly one place, &lt;code&gt;05_TOOLS/_watermark-vf.sh&lt;/code&gt;, which emits a filter chain on stdout and is called by both the builder and a standalone burner. Nothing else knows what the watermark looks like.&lt;/p&gt;

&lt;p&gt;Three decisions in it are non-obvious.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is &lt;code&gt;drawtext&lt;/code&gt;, not an image overlay.&lt;/strong&gt; The no-asset route: no PNG to keep in sync, no alpha channel to get wrong, and the size is an &lt;em&gt;expression&lt;/em&gt; rather than a number, so the mark scales with the frame instead of being authored for 1080p and shrinking to nothing on a 4K export:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;FONT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"05_TOOLS/.wm-work/wm-font.ttf"&lt;/span&gt;   &lt;span class="c"&gt;# staged: a relative path survives every parser&lt;/span&gt;
ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;.mp4 &lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"drawtext=fontfile=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;FONT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:text='razi.pro':fontsize=h/40:fontcolor=white&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
:shadowcolor=black@0.35:shadowx=1:shadowy=1:alpha='0.060'&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
:x=(W-tw)/2:y=(H-th)/2"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt;:v libx264 &lt;span class="nt"&gt;-crf&lt;/span&gt; 18 &lt;span class="nt"&gt;-preset&lt;/span&gt; medium &lt;span class="nt"&gt;-pix_fmt&lt;/span&gt; yuv420p &lt;span class="nt"&gt;-c&lt;/span&gt;:a copy out.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;fontsize=h/40&lt;/code&gt; is 27 px at 1080p and 54 px at 4K. The shadow is not decoration — it is what keeps a 6% white mark readable when a bright frame comes up underneath it, and it is scaled to the mark, because at that alpha a heavier shadow reads as a dark smudge with nothing inside it.&lt;/p&gt;

&lt;p&gt;The image-overlay equivalent is worth having in your pocket for when a logo is mandatory rather than a text mark. Roughly 8% of frame width, bottom-right with a margin, and the opacity applied to the asset rather than baked into the PNG:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;.mp4 &lt;span class="nt"&gt;-i&lt;/span&gt; logo.png &lt;span class="nt"&gt;-filter_complex&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"[1:v][0:v]scale2ref=w=iw*0.08:h=ow/mdar[wm][base];&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
   [wm]format=rgba,colorchannelmixer=aa=0.12[wmo];&lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
   [base][wmo]overlay=W-w-24:H-h-24"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-c&lt;/span&gt;:v libx264 &lt;span class="nt"&gt;-crf&lt;/span&gt; 18 &lt;span class="nt"&gt;-preset&lt;/span&gt; medium &lt;span class="nt"&gt;-pix_fmt&lt;/span&gt; yuv420p &lt;span class="nt"&gt;-c&lt;/span&gt;:a copy out.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;scale2ref&lt;/code&gt; rather than &lt;code&gt;scale&lt;/code&gt; is not a stylistic choice: plain &lt;code&gt;scale&lt;/code&gt; cannot see the main frame's dimensions, and reaching for &lt;code&gt;main_w&lt;/code&gt; inside it fails with &lt;em&gt;"Expressions with scale2ref variables are not valid in scale filter"&lt;/em&gt;. &lt;code&gt;scale2ref&lt;/code&gt; takes the video as a second input purely to measure it, so &lt;code&gt;iw*0.08&lt;/code&gt; means 8% of the &lt;em&gt;frame&lt;/em&gt;, not 8% of the logo — and the mark stays proportional at every export resolution.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;format=rgba,colorchannelmixer=aa=0.12&lt;/code&gt; is the part people get wrong. Exporting a pre-faded PNG bakes the opacity into the asset and you cannot change your mind without going back to the image editor; doing it in the filter graph makes opacity a parameter of the build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It moves, and it never repeats.&lt;/strong&gt; A watermark in a fixed corner is one crop away from gone. Two copies of the text drift on independent paths, each axis a sum of two sines with mutually incommensurate periods:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;AX&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"(W-tw)/2*(1+0.60*sin(2*PI*t/71.0)+0.36*sin(2*PI*t/29.3+2.1))"&lt;/span&gt;
&lt;span class="nv"&gt;AY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"(H-th)/2*(1+0.58*sin(2*PI*t/53.7+0.7)+0.38*sin(2*PI*t/23.1+3.9))"&lt;/span&gt;
&lt;span class="nv"&gt;AA&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"0.060+0.012*sin(2*PI*t/41.0)"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Position is a continuous function of &lt;code&gt;t&lt;/code&gt;, so it flows rather than teleporting between corners; the amplitudes sum to under 1.0, so it mathematically cannot clip out of frame; and because 71 and 29.3 share no common period inside 200 seconds, no fixed crop or blur region removes the mark for the whole runtime. The opacity breathes too, on a third period, so it never settles into something a static filter could subtract.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A watermark's job is not to be seen. It is to be expensive to remove. Motion on incommensurate periods is what turns a corner logo into a removal problem.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;It is applied at the final encode, and never to the master.&lt;/strong&gt; &lt;code&gt;_master-assembly-nofades.mp4&lt;/code&gt; is clean, and stays clean. The mark is appended to the video filter chain &lt;em&gt;after&lt;/em&gt; the fades, inside the same encode that does the mix — so it costs nothing extra, and it exists only in the distribution copy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;VF&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"fade=t=in:st=0:d=1.5,fade=t=out:st=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;FOUT&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:d=1.5"&lt;/span&gt;
&lt;span class="nv"&gt;WM&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;bash &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HERE&lt;/span&gt;&lt;span class="s2"&gt;/_watermark-vf.sh"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$WM&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; VF+&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;",&lt;/span&gt;&lt;span class="nv"&gt;$WM&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the rule that has no exceptions. A watermarked master is a master you can never re-grade, never re-cut, never hand to a festival, and never re-encode without stacking a second mark on top of the first. &lt;code&gt;WATERMARK=0&lt;/code&gt; builds a clean copy, and that is what a client or an archive gets. The archival copy is the one that has to survive decisions you have not made yet.&lt;/p&gt;

&lt;p&gt;Finally: &lt;strong&gt;check it on a dark shot.&lt;/strong&gt; Faint white over black is where a watermark disappears, and a value tuned against a daylight frame can vanish entirely at night. To confirm it survived the encode, lift the shadows rather than boosting contrast — &lt;code&gt;eq=contrast&lt;/code&gt; pivots at mid-grey and crushes exactly the range the mark lives in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; out.mp4 &lt;span class="nt"&gt;-vf&lt;/span&gt; &lt;span class="s2"&gt;"format=gray,lut=y='min(255&lt;/span&gt;&lt;span class="se"&gt;\,&lt;/span&gt;&lt;span class="s2"&gt;val*9)'"&lt;/span&gt; &lt;span class="nt"&gt;-frames&lt;/span&gt;:v 1 check.png
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — no watermark.&lt;/strong&gt; Genuinely an option, and for a film about your own children that mostly gets played on your own television, arguably the right one. Everything above is machinery for a problem you may not have. The rule that matters even then is the last one: &lt;strong&gt;never mark the master.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — Content Credentials (C2PA).&lt;/strong&gt; Cryptographically signed provenance metadata attached to the file rather than painted onto the picture. It answers "where did this come from" properly, which a drifting text mark does not, and it survives none of the things a visible mark survives. The two solve different problems and are not alternatives so much as complements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — invisible watermarking.&lt;/strong&gt; Steganographic marks that survive re-encoding and cropping, recoverable by a detector rather than by eye. If the actual goal is proving ownership after the fact, this is the serious answer and drawing text on the frame is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be honest about what the visible mark does.&lt;/strong&gt; It raises the cost of a lazy repost. It does not stop a determined one. I built the drifting version because a corner logo is one crop away from gone, not because I think it is protection.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjx5780ufggi4ky4yz0o4.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjx5780ufggi4ky4yz0o4.webp" alt="The moon rising back into the night sky on a gold trail, Comet flying alongside" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 19. Cue &lt;code&gt;c&lt;/code&gt; returns, the sidechain has nothing left to duck against, and the film starts its 1.5-second fade out.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Cutting a recording apart with a computer
&lt;/h2&gt;

&lt;p&gt;The best performance in the film was recorded on a phone, in a living room, by someone who had never recorded anything.&lt;/p&gt;

&lt;p&gt;That is the chapter. Everything below is the machinery for getting a real human read — retakes, fridge hum, corrupt frame and all — into fifteen shot-shaped pieces without flattening what made it worth keeping.&lt;/p&gt;

&lt;p&gt;The recorded narrations arrived as one continuous take per language: fifteen lines read top to bottom, three seconds of silence between them, fumbles simply repeated after a pause. That is the right way to record — no syncing to picture, no per-line file management — but it means the take has to be cut up afterwards, retakes and all.&lt;/p&gt;

&lt;h3&gt;
  
  
  Whisper vs Scribe on Malayalam
&lt;/h3&gt;

&lt;p&gt;First attempt used &lt;a href="https://github.com/SYSTRAN/faster-whisper" rel="noopener noreferrer"&gt;faster-whisper&lt;/a&gt; with &lt;a href="https://huggingface.co/openai/whisper-large-v3" rel="noopener noreferrer"&gt;&lt;code&gt;large-v3&lt;/code&gt;&lt;/a&gt;, which is the obvious local choice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;faster_whisper&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;WhisperModel&lt;/span&gt;
&lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;WhisperModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;large-v3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cpu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;compute_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;int8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cpu_threads&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;segs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transcribe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;voice-female-clean.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ml&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;vad_filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                       &lt;span class="n"&gt;vad_parameters&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;min_silence_duration_ms&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;350&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;speech_pad_ms&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                       &lt;span class="n"&gt;word_timestamps&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;beam_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It failed on the Malayalam recording. Whisper's training distribution is heavily English-weighted, and its documented tendency to hallucinate text that was not spoken gets much worse in low-resource languages; the output was not usable for alignment. &lt;strong&gt;ElevenLabs Scribe&lt;/strong&gt; transcribed the same file correctly — Malayalam is in its high-accuracy band, under 10% WER, across its 90+ supported languages.&lt;/p&gt;

&lt;p&gt;Worth stating plainly: for English, faster-whisper is excellent and free and there was no reason to use anything else. For Malayalam it was simply the wrong tool, and there is no amount of parameter tuning that fixes a model that does not know the language well enough.&lt;/p&gt;

&lt;h3&gt;
  
  
  Nine of thirteen segments started mid-word
&lt;/h3&gt;

&lt;p&gt;The English recording needed the same treatment, and my first attempt at it was bad: segments chopped at fixed durations to fit 10-second slots, with nine of thirteen starting or ending mid-word.&lt;/p&gt;

&lt;p&gt;The rebuild — again, written in-session rather than by me — uses a deliberately &lt;strong&gt;energy-first, not ASR-first&lt;/strong&gt; design:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Find utterance blocks by measured energy.&lt;/strong&gt; Compute an RMS envelope, take the 10th percentile as the noise floor, gate at 6 dB above it, merge pauses shorter than 250 ms, discard blips under 350 ms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use word timestamps only to decide which block is which line.&lt;/strong&gt; Whisper's word boundaries stretch across pauses; measured silence does not lie. Cutting on energy is immune to both that and to the flubbed retake sitting in the middle of the take.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Match blocks to script lines on character trigrams, not words.&lt;/strong&gt; Word-level overlap is too brittle — the transcript says "good night" and "mid air" where the script has "Goodnight" and "midair", which zeroes a word-overlap score on precisely the lines that most need placing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detect false starts by asymmetric containment.&lt;/strong&gt; A false start is a short fragment whose text all reappears in the retake that follows. Symmetric F1 sinks on that comparison because the lengths differ wildly; asymmetric containment — how much of the &lt;em&gt;small&lt;/em&gt; block reappears in the following blocks — catches it cleanly. Blocks are then assigned a drop cost: free if they contain no words (breath, page turn) or are a detected false start, expensive otherwise, which is what prevents the optimiser cheerfully deleting real narration.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every cut is then verified: if either edge of a cut sits more than 8 dB above the noise floor, it landed on speech, and the script exits non-zero.&lt;/p&gt;

&lt;p&gt;The retained segments get a "storyteller warm" enhancement chain. The order matters, and the comments record a mistake worth reproducing — a naive bass boost cost 4 dB of SNR, because this room's tone is low-frequency and sits exactly where the warmth goes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ENHANCE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;highpass=f=85:poles=2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                &lt;span class="c1"&gt;# rumble out (fundamental is 120-140Hz)
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;afftdn=nr=20:nf=-38&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                  &lt;span class="c1"&gt;# denoise at the measured floor
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agate=threshold=0.03:ratio=3:attack=10:release=200:knee=6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;equalizer=f=125:t=q:w=0.9:g=4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;        &lt;span class="c1"&gt;# chest weight at the fundamental
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bass=g=3:f=110:w=0.6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                 &lt;span class="c1"&gt;# low shelf for body
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;equalizer=f=330:t=q:w=1.6:g=-2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="c1"&gt;# de-box (narrow: a wide cut thins the voice)
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;equalizer=f=4000:t=q:w=1.2:g=3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;       &lt;span class="c1"&gt;# presence / intimacy
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;treble=g=3:f=7000&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                    &lt;span class="c1"&gt;# air
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deesser=i=0.35&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                       &lt;span class="c1"&gt;# tame the sibilance the HF lift adds
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acompressor=threshold=-18dB:ratio=2.5:attack=15:release=250:makeup=1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Clean the low end &lt;em&gt;before&lt;/em&gt; adding weight to it; expand &lt;em&gt;before&lt;/em&gt; compressing. After the fix, room tone sits 17.2 dB under the speech, better than the 14.2 dB of the untouched original.&lt;/p&gt;

&lt;p&gt;Assembly onto the timeline asserts non-overlap rather than assuming it — the previous version had &lt;code&gt;narr-16&lt;/code&gt; running 1.14 s into &lt;code&gt;narr-17&lt;/code&gt; at 02:40 — and time-fits a line only when it genuinely does not fit its available room, capped at 1.12×.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — record line by line.&lt;/strong&gt; Fifteen takes, fifteen files, &lt;code&gt;narr-01.wav&lt;/code&gt; through &lt;code&gt;narr-20.wav&lt;/code&gt;, and this entire section evaporates. There is no cut to make, no false start to detect, no energy gate to tune. The cost is studio discipline from someone who may not have any: stopping and starting fifteen times is a real imposition on a reader, and it flattens the performance because nobody builds momentum across a take they keep restarting. I chose one continuous read for the performance and paid for it in code. That was the right trade for me and might not be for you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — cutting it by hand.&lt;/strong&gt; Audacity, fifteen selections, fifteen exports. Twenty minutes for one language. My auto-cut is worth writing because there are two languages, six cuts and a near-certainty of re-recording — not because slicing a WAV is hard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tried, failed — chopping at fixed durations.&lt;/strong&gt; The first version cut to fit the ten-second slots and left nine of thirteen segments starting or ending mid-word. Cut on measured silence, then fit; never fit first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tried, failed — Whisper large-v3 on Malayalam.&lt;/strong&gt; No amount of parameter tuning fixes a model that does not know the language. For English it was excellent and free and I had no reason to use anything else.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  A grandmother, a phone, and a room with a fridge in it
&lt;/h3&gt;

&lt;p&gt;The Malayalam read did not come out of a booth. It came out of a phone, held at arm's length, in an ordinary living room, by someone who had never recorded anything before and was not going to be asked to do it twice.&lt;/p&gt;

&lt;p&gt;That is the realistic case, and it is worth treating as the default rather than the exception. The performance in that file is the best thing in the project — the real read measures 10.0 semitones of pitch range against the clone's 5.1 — and every problem with it is fixable.&lt;/p&gt;

&lt;p&gt;The first problem was not audio at all. The MP3 had a corrupt frame near the end of the file, and &lt;code&gt;libsndfile&lt;/code&gt; refused it outright, which is exactly the wrong response to a recording that cannot be re-made. FFmpeg decodes straight through the damage with a warning and loses a few milliseconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decode_source&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dst&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;ffmpeg, not soundfile: the mp3 has a corrupt frame near EOF that
    libsndfile chokes on but ffmpeg skips with a warning.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Decode with the most forgiving tool you have, not the most correct one. A strict decoder on an irreplaceable file is a bug, not a feature.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then the restoration chain, in this order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;highpass&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;f=85,&lt;/span&gt;
&lt;span class="py"&gt;afftdn&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;nf=-32:tn=1,&lt;/span&gt;
&lt;span class="err"&gt;adeclick,&lt;/span&gt;
&lt;span class="py"&gt;equalizer&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;f=250:t=q:w=1.0:g=-2.5,&lt;/span&gt;
&lt;span class="py"&gt;equalizer&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;f=3200:t=q:w=0.9:g=2.5,&lt;/span&gt;
&lt;span class="py"&gt;deesser&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;i=0.35,&lt;/span&gt;
&lt;span class="py"&gt;acompressor&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;threshold=-20dB:ratio=3:attack=8:release=120:makeup=2,&lt;/span&gt;
&lt;span class="py"&gt;loudnorm&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;I=-18:TP=-2:LRA=9&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eight stages, and the order is the whole argument. Repair first, then subtract, then add, then control dynamics, then set level — because every stage amplifies whatever the stages before it left behind.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;highpass=f=85&lt;/code&gt;.&lt;/strong&gt; Nothing in a human voice lives below 85 Hz. Everything else does: the fridge compressor, traffic through a window, the low thud of a hand shifting on a phone case, and any DC offset the phone's ADC introduced. It is not audible on laptop speakers, which is why people leave it in, and it eats headroom that the compressor at stage seven will later spend gain amplifying. Cut it before anything downstream can act on it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;afftdn=nf=-32:tn=1&lt;/code&gt;.&lt;/strong&gt; Spectral denoise, with &lt;code&gt;nf&lt;/code&gt; set to the room's &lt;em&gt;measured&lt;/em&gt; noise floor rather than a guess. −32 dBFS is what this room actually read between lines. Too low and the hiss survives; too high and you start subtracting the quiet tails of words, which is the underwater artefact everyone recognises and nobody can name. &lt;code&gt;tn=1&lt;/code&gt; turns on noise tracking so the filter follows the floor as it moves — a room is not stationary, and neither is someone holding a phone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;adeclick&lt;/code&gt;.&lt;/strong&gt; Impulse repair: lip smacks, a fingernail on the phone case, the decoder's own stitching over that corrupt frame. It runs &lt;em&gt;after&lt;/em&gt; the denoise deliberately. A click detector fed broadband hiss finds clicks everywhere; on a cleaned signal, a real transient stands out and gets repaired without the filter chewing on consonants.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;equalizer=f=250 … g=-2.5&lt;/code&gt;.&lt;/strong&gt; A phone held close to a face has proximity effect, and a small room piles its lowest modes into the same band. Both land around 200–300 Hz and both sound the same: boxy, muffled, close-but-dull. −2.5 dB is a small number on purpose. This band also carries the warmth that makes a grandmother's voice sound like a grandmother, and a 6 dB scoop takes the person out along with the room.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;equalizer=f=3200 … g=+2.5&lt;/code&gt;.&lt;/strong&gt; Consonant definition. This is where &lt;code&gt;t&lt;/code&gt;, &lt;code&gt;k&lt;/code&gt; and &lt;code&gt;s&lt;/code&gt; live, where intelligibility actually comes from, and where a phone microphone is weakest. Note that it is the &lt;em&gt;fifth&lt;/em&gt; stage, not the second: lifting presence before cutting mud means the compressor sees a signal with both problems still in it. Subtract, then add.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;deesser=i=0.35&lt;/code&gt;.&lt;/strong&gt; Payment for the previous stage. A presence lift is indiscriminate — it raises sibilance by exactly as much as it raises articulation. The de-esser gives back the harshness and keeps the clarity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;acompressor=threshold=-20dB:ratio=3&lt;/code&gt;.&lt;/strong&gt; An untrained reader varies by 15 dB or more within a single line — leaning in on the emphatic words, dropping away at the end of sentences. A 3:1 ratio at −20 dB is enough to put every syllable in the same neighbourhood without flattening the performance. This one is not cosmetic: the final mix ducks music and clip audio via &lt;code&gt;sidechaincompress&lt;/code&gt; keyed off the narration, and a sidechain key with 15 dB of internal variation ducks unevenly. Compressing here is what makes the ducking downstream predictable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;loudnorm=I=-18:TP=-2:LRA=9&lt;/code&gt;.&lt;/strong&gt; Deliver at a known number. −18 LUFS and not the −16 of the finished film, because this track gets normalised again in the final mix and arriving pre-loud only means arriving pre-squashed. &lt;code&gt;TP=-2&lt;/code&gt; leaves true-peak headroom for the AAC encode, which can overshoot the sample peak. &lt;code&gt;LRA=9&lt;/code&gt; is the one to argue about: it is deliberately generous, because a bedtime story that has been levelled to a constant loudness is no longer a bedtime story.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every stage in that chain exists to pay for the one before it. That is what makes the order load-bearing and not a matter of taste.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And then the finding that I did not expect and would put at the top of the post if it were not buried three thousand words in:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The phone beat the GPU. The best voice in this film cost nothing, needed no model, no CUDA wheels, no byte-length debugging and no licence review — 10.0 semitones of pitch range against the clone's 5.1, from someone who had never recorded anything before and did it once.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every other section of this post is a machine standing in for a person who was not available. If the person &lt;em&gt;is&lt;/em&gt; available, use the person. The restoration chain above exists to make an ordinary living room acceptable, and it is worth having precisely because it means you never have to say "we can't, we haven't got a studio."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — a $25 USB microphone.&lt;/strong&gt; Most of that eight-stage chain is compensating for a phone held at arm's length. Any cheap cardioid on a desk removes the room before you have to remove it in software.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — cleaning it up by ear in Audacity.&lt;/strong&gt; Noise reduction, a high-pass, a compressor, done in a GUI while listening. For one irreplaceable file that is the humane route; my chain exists because there were several files and I wanted the same treatment on all of them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A rule worth taking regardless:&lt;/strong&gt; decode an irreplaceable recording with the most &lt;em&gt;forgiving&lt;/em&gt; tool you have, not the most correct one. &lt;code&gt;libsndfile&lt;/code&gt; refused this file outright. FFmpeg walked through the damage and lost a few milliseconds.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The directory listing is the pipeline
&lt;/h2&gt;

&lt;p&gt;You can tell how a project actually ran by looking at its folders.&lt;/p&gt;

&lt;p&gt;Mine are numbered in pipeline order, and that is not tidiness — it is the working memory the agent read from and wrote back to for a month. This chapter is the shape of the thing, who really wrote it, and what the whole exercise cost.&lt;/p&gt;

&lt;p&gt;Post-production folders are numbered in pipeline order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;00_DELIVERABLES/   six finished .mp4 + matching .srt
01_SCRIPT/         production pack, narration lines (machine), reading scripts (human)
02_ART/            characters/  items/  reference_photos/  contact-sheet.jpg
03_FOOTAGE/        shot-01..20.mp4 + _master-assembly-nofades.mp4
04_AUDIO/          music/  vo_raw/  vo_lines/  vo_mixed/
05_TOOLS/          the scripts
99_ARCHIVE/        superseded builds, scratch — safe to delete wholesale
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three conventions carry most of the value:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;00&lt;/code&gt; is what you hand someone; &lt;code&gt;99&lt;/code&gt; is what you can delete.&lt;/strong&gt; No judgement call about what is disposable six months from now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subtitles sit beside their video with a matching basename&lt;/strong&gt;, so players load them automatically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero-padded shot numbers everywhere&lt;/strong&gt; — &lt;code&gt;shot-01&lt;/code&gt;, &lt;code&gt;narr-01&lt;/code&gt; — so lexical sort equals story order in every tool, shell glob included. &lt;code&gt;make-narration.sh&lt;/code&gt; iterates &lt;code&gt;narr-*.wav&lt;/code&gt; and derives each line's timeline offset arithmetically from its filename: &lt;code&gt;(N-1)*10 + 0.4&lt;/code&gt;. Filename discipline is what makes that a one-liner instead of a lookup table.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One more, learned the hard way and recorded in the README: &lt;strong&gt;do not move the virtualenv.&lt;/strong&gt; Python virtualenvs store absolute paths.&lt;/p&gt;




&lt;h2&gt;
  
  
  The pipeline had a fourth author
&lt;/h2&gt;

&lt;p&gt;Nothing in &lt;code&gt;05_TOOLS/&lt;/code&gt; was typed by me.&lt;/p&gt;

&lt;p&gt;The whole production ran out of a terminal: &lt;strong&gt;Claude Code&lt;/strong&gt;, on &lt;strong&gt;Opus&lt;/strong&gt;, sitting in the project directory with the numbered folders above as its working memory. No IDE, no notebook. The production pack in &lt;code&gt;01_SCRIPT/&lt;/code&gt; was the shared document — I edited it, the model read it, and everything downstream derived from it rather than restated it.&lt;/p&gt;

&lt;p&gt;What that actually bought me was the FFmpeg. Every filter graph in this post arrived as a first draft out of a conversation and was then argued into shape against real output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;make-music-bed.sh&lt;/code&gt;.&lt;/strong&gt; I asked for a 200-second score and rejected the obvious answer, because looping one 30-second cue six times is audible. What came back instead was a &lt;code&gt;seamless_loop&lt;/code&gt; function chaining copies through &lt;code&gt;acrossfade=d=2:c1=tri:c2=tri&lt;/code&gt;, then four sections crossfaded into a bed that follows the story — 82 s adventure, 62 s wonder, 42 s triumph, 20 s home.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;build-final.sh&lt;/code&gt;.&lt;/strong&gt; The sidechain chain two sections up, including the detail I would not have found alone: that &lt;code&gt;sidechaincompress&lt;/code&gt; &lt;em&gt;consumes&lt;/em&gt; its key input, so the narration has to be split into two identical branches. The watermark got folded into the same &lt;code&gt;-vf&lt;/code&gt; as the fades, so there is no second encode and no second generation of loss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;_watermark-vf.sh&lt;/code&gt;.&lt;/strong&gt; Two drifting &lt;code&gt;drawtext&lt;/code&gt; marks whose x and y are each a sum of two sines with incommensurate periods — 71.0 and 29.3 seconds on one axis, 61.3 and 19.7 on the other — so the marks never line up, with the amplitudes summing to under 1.0 so neither can wander out of frame.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And it absorbed the environment damage, which is the part that would otherwise have eaten a weekend. &lt;code&gt;torchaudio&lt;/code&gt; wanted FFmpeg 7's shared libraries; this machine has FFmpeg 9, static — so a shim monkey-patches &lt;code&gt;load&lt;/code&gt; and &lt;code&gt;save&lt;/code&gt; onto &lt;code&gt;soundfile&lt;/code&gt; and gets imported before anything else. IndicF5 ships a vendored &lt;code&gt;f5_tts&lt;/code&gt; that collides with the pip-installed one, so its &lt;code&gt;__path__&lt;/code&gt; is pinned and then asserted at startup. A Windows console in cp1252 cannot print Malayalam, so stdout is reconfigured to UTF-8 on the first line of the script. None of that is interesting. All of it is load-bearing.&lt;/p&gt;

&lt;p&gt;Where it was no help at all: Flow has no API, and pointing a browser automation at it was not something I was willing to do, so all twenty shots and all eighty re-rolls were clicked by hand. Every taste call was mine — which of the three MusicGen options, which re-roll to keep, whether the line-up actually looked like my children. And the arithmetic lies: I asked Veo for 8 seconds and got 10, which is why the master gets measured with &lt;code&gt;ffprobe&lt;/code&gt; before the fades are placed instead of calculated as twenty times anything.&lt;/p&gt;

&lt;p&gt;The honest sequence is that the pipeline was improvised, in-session, while the film was being made. It only became a &lt;em&gt;system&lt;/em&gt; afterwards. Once the six cuts were rendered I turned the whole run into six Claude Code skills — &lt;code&gt;film-00-new&lt;/code&gt; through &lt;code&gt;film-05-post&lt;/code&gt; — sharing one &lt;code&gt;film.json&lt;/code&gt; that holds the cast, the style block, the backend, and a &lt;code&gt;state&lt;/code&gt; object recording which shots are done, which references exist, and what has shipped, so a run can be resumed rather than restarted. Two places it stops and waits for a human: after the script, and after the character line-up. Both are checkpoints this film taught me to want.&lt;/p&gt;

&lt;p&gt;Everything in them is a mistake from this post, written down so the next film does not pay for it twice. Restate the costume in the text of every single shot, or you get an astronaut. Generate the group line-up first and cut the characters out of &lt;em&gt;it&lt;/em&gt;, or the four-year-old comes back shorter than the two-year-old. Set &lt;strong&gt;Confirm before generating&lt;/strong&gt; to Always. Shoot on Lite and spend the savings on re-rolls.&lt;/p&gt;

&lt;p&gt;The film cost a month of Pro. The skills are what I got to keep.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Other ways to do this&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything here is &lt;strong&gt;not tried&lt;/strong&gt; — I used one agent for the whole film and cannot tell you how the others behave under this particular load. The category moved fast enough during production that the honest note is "check what these are today, not what a blog said in 2026."&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — &lt;a href="https://github.com/openai/codex" rel="noopener noreferrer"&gt;Codex CLI&lt;/a&gt;.&lt;/strong&gt; OpenAI's terminal agent, Apache-2.0. The nearest like-for-like swap for what I did here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — &lt;a href="https://antigravity.google/" rel="noopener noreferrer"&gt;Antigravity CLI&lt;/a&gt;.&lt;/strong&gt; Google's agent platform spans an IDE, a desktop app, a CLI and an SDK; the CLI is the part comparable to this workflow. Worth knowing if you are on Gemini: from 18 June 2026 Gemini CLI stopped serving Google AI Pro/Ultra and free individual Code Assist users, who are directed to Antigravity CLI instead — Code Assist Standard and Enterprise are unaffected. Unlike the rest of this list it is not published under an open licence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — &lt;a href="https://opencode.ai" rel="noopener noreferrer"&gt;OpenCode&lt;/a&gt;.&lt;/strong&gt; Open source (MIT), model-agnostic, terminal-first. Note the name collision: the archived &lt;code&gt;opencode-ai/opencode&lt;/code&gt; repository is a different, discontinued project whose author moved on to Crush; the maintained one is &lt;a href="https://github.com/sst/opencode" rel="noopener noreferrer"&gt;&lt;code&gt;sst/opencode&lt;/code&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — &lt;a href="https://aider.chat" rel="noopener noreferrer"&gt;Aider&lt;/a&gt;.&lt;/strong&gt; Apache-2.0, model-agnostic, git-native. It is built as an interactive pair-programmer rather than an unattended runner, so treat it as a different shape of tool rather than a drop-in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not tried — &lt;a href="https://github.com/MoonshotAI/kimi-code" rel="noopener noreferrer"&gt;Kimi Code CLI&lt;/a&gt;.&lt;/strong&gt; Moonshot's terminal agent, MIT. Its predecessor &lt;code&gt;kimi-cli&lt;/code&gt; is being wound down. Separately, the Kimi Code API can be pointed at as the model backend for Claude Code, OpenCode and Codex, which is the cheaper experiment if you like your current agent and only want to change the model underneath it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The thing that actually mattered was not the brand.&lt;/strong&gt; It was that the agent lived in the project directory, could run FFmpeg and read its own output, and kept the numbered folders as working memory. Any tool with those three properties would have got most of the way here.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Costs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Google AI Pro (1,000 Flow credits/month)&lt;/td&gt;
&lt;td&gt;one month's subscription&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flow credits spent — 20 shots at Veo 3.1 Lite&lt;/td&gt;
&lt;td&gt;200 of 1,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flow credits available for re-rolls&lt;/td&gt;
&lt;td&gt;800 (≈80 re-rolls)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p upscaling, all 20 clips&lt;/td&gt;
&lt;td&gt;free on Pro&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference images — Gemini app free tier&lt;/td&gt;
&lt;td&gt;free (separate pool from Flow credits)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MusicGen Medium score&lt;/td&gt;
&lt;td&gt;free (local GPU)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;F5-TTS English cloning&lt;/td&gt;
&lt;td&gt;free (local GPU)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IndicF5 Malayalam cloning&lt;/td&gt;
&lt;td&gt;free (local GPU, MIT)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs stock-voice narration&lt;/td&gt;
&lt;td&gt;free tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs Scribe transcription&lt;/td&gt;
&lt;td&gt;free tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FFmpeg assembly, mixing, encoding&lt;/td&gt;
&lt;td&gt;free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Marginal cost beyond the subscription&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;zero&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Everything except the video generation runs locally on consumer hardware. For anyone on the free Flow tier — 50 credits a day — the same film is five days of shooting at Lite rates with a handful of re-rolls, or about ten days if you want a real re-roll budget. The section above, Making this with no subscription at all, maps every stage onto a free substitute.&lt;/p&gt;

&lt;p&gt;The genuine cost of this project was time: reference-image iteration, the Malayalam duration debugging, and the narration edit. Video generation was the cheap part.&lt;/p&gt;




&lt;h2&gt;
  
  
  Making this with no subscription at all
&lt;/h2&gt;

&lt;p&gt;The one line in the cost table that deserves an asterisk is the first one. A month of Google AI Pro bought me &lt;strong&gt;wall-clock, not capability&lt;/strong&gt;: 1,000 credits with no daily cap collapses the shoot from a two-week drip into two evenings. Nothing in the finished film requires it.&lt;/p&gt;

&lt;p&gt;Here is the whole pipeline with the paid step removed. Everything marked &lt;em&gt;not tried&lt;/em&gt; is exactly that — a route I believe works and did not walk, offered so you can start from a real map instead of from my one path.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;What I used&lt;/th&gt;
&lt;th&gt;Free substitute&lt;/th&gt;
&lt;th&gt;What it costs you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Video&lt;/td&gt;
&lt;td&gt;Veo 3.1 Lite, Google AI Pro&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Flow free tier&lt;/strong&gt;, 50 credits/day — &lt;em&gt;same model&lt;/em&gt; (&lt;strong&gt;tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Five days of shooting, ten if you want re-rolls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video, no Google&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Open weights locally: Wan, LTX-Video, HunyuanVideo, CogVideoX (&lt;strong&gt;not tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;No native audio.&lt;/strong&gt; Twenty shots of foley and eleven lines of dialogue become separate jobs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video, no GPU&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Free daily credits on Kling, Hailuo, Pika, Luma (&lt;strong&gt;not tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Small daily quotas; check each service's terms on children's likenesses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference art&lt;/td&gt;
&lt;td&gt;Nano Banana, Gemini app free tier&lt;/td&gt;
&lt;td&gt;SDXL or Flux with IP-Adapter / InstantID, local (&lt;strong&gt;not tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Likeness is harder to hold; the line-up trick matters more, not less&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Score&lt;/td&gt;
&lt;td&gt;MusicGen Medium, local GPU&lt;/td&gt;
&lt;td&gt;MusicGen Small on CPU; Stable Audio Open; YouTube Audio Library (&lt;strong&gt;not tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Slower, or cues that don't follow your beat map&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Score, at all&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Skip it — Veo's native audio carries the film (&lt;strong&gt;tried&lt;/strong&gt;: the music-only cut)&lt;/td&gt;
&lt;td&gt;Nothing, honestly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;English narration&lt;/td&gt;
&lt;td&gt;F5-TTS clone, local GPU&lt;/td&gt;
&lt;td&gt;Piper or Kokoro, CPU, stock voices (&lt;strong&gt;not tried&lt;/strong&gt;); or &lt;strong&gt;read it yourself into a phone&lt;/strong&gt; (&lt;strong&gt;tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Stock voices aren't anyone's parent. Your own voice is better than both&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Malayalam narration&lt;/td&gt;
&lt;td&gt;IndicF5, local GPU&lt;/td&gt;
&lt;td&gt;A family member and a phone (&lt;strong&gt;tried — and it won&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Nothing. This was the best audio in the project&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transcription&lt;/td&gt;
&lt;td&gt;ElevenLabs Scribe, free tier&lt;/td&gt;
&lt;td&gt;faster-whisper, local (&lt;strong&gt;tried&lt;/strong&gt;, English only)&lt;/td&gt;
&lt;td&gt;For Malayalam, nothing local worked for me&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cut, mix, encode, subtitle&lt;/td&gt;
&lt;td&gt;FFmpeg&lt;/td&gt;
&lt;td&gt;FFmpeg — or DaVinci Resolve free / Shotcut (&lt;strong&gt;not tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware&lt;/td&gt;
&lt;td&gt;RTX Blackwell, local&lt;/td&gt;
&lt;td&gt;Google Colab / HuggingFace Spaces free tiers (&lt;strong&gt;not tried&lt;/strong&gt;)&lt;/td&gt;
&lt;td&gt;Session limits and queueing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three honest caveats about the free route.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Leaving Veo is the expensive "free" decision.&lt;/strong&gt; Every open video model on that list generates a silent clip. Veo handed me the springs, the hiccups, the orchestral swell on the reveal and every spoken line inside the same generation I was paying for anyway. Rebuilding that by hand is more work than everything else in this post put together, and it is the reason I would tell someone on a strict budget to use Flow's free tier slowly rather than a local model quickly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The free tier is not a worse film. It is a longer calendar.&lt;/strong&gt; Same model, same weights, same prompts, five credits a day fewer. If you are not in a hurry, there is no argument for the subscription at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The thing that is never free is time,&lt;/strong&gt; and this post is mostly a list of the places I spent it. Reference-image iteration. A day on byte-length duration estimation. The narration edit. None of that gets cheaper on a paid plan.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;Lia can now recite the narration before it plays.&lt;/p&gt;

&lt;p&gt;That is the only metric that ever mattered here, and it is worth saying out loud before the corrections start. Everything below is what I would do differently — not because the film failed, but because most of these cost me a day each and they should cost you an afternoon between them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generate the two-character line-up first, not third.&lt;/strong&gt; &lt;code&gt;Together.jpg&lt;/code&gt; fixed the inverted sibling heights and the mismatched render styles in one image, and both problems only existed because the children were generated in separate passes. Generate the group shot first, then crop the solos out of it if you need them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the costume description into the prompt from shot one.&lt;/strong&gt; The spacesuit render cost credits and an afternoon. Ingredients bias the look; they do not lock it. Assume the reference contributes nothing the text does not also say, and you will never be surprised.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check the model selector before the first generation.&lt;/strong&gt; Flow defaulted to Omni 1.1 Flash, outside the credit-eligible Veo family and more expensive per clip. This is a thirty-second check that protects a month's allocation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prototype the non-English TTS on one line before building the pipeline.&lt;/strong&gt; The 3-bytes-per-character duration bug would have surfaced in five minutes on a single test sentence. Instead it surfaced across a full batch of hallucinated narration. &lt;code&gt;--test&lt;/code&gt; mode exists in &lt;code&gt;make-narration-ml.py&lt;/code&gt; now; it should have existed first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cut the narration recording properly the first time.&lt;/strong&gt; The initial version chopped at fixed durations to fit shot slots and left nine of thirteen segments starting or ending mid-word. The energy-first re-cut is objectively better &lt;em&gt;and&lt;/em&gt; it was less work than the manual fixing it replaced. Automate the cut before you automate anything downstream of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assert, don't assume, on the timeline.&lt;/strong&gt; Two narration lines overlapped by 1.14 seconds for several builds before anyone noticed. Non-overlap is a one-line assertion. Add it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decide the watermark policy before the first deliverable, not after.&lt;/strong&gt; The rule is one line — the mark goes on at the final encode and never touches the master — and it is trivially cheap to follow from the start and expensive to retrofit, because by then you have shipped files you can no longer prove are clean.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test the transcriber on the target language before designing around it.&lt;/strong&gt; Whisper large-v3 is the reflexive choice and it was the wrong one for Malayalam. Ten minutes of evaluation would have saved a day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost the free path before buying the subscription.&lt;/strong&gt; The Pro month bought two evenings instead of two weeks. That is a real thing to buy and I would buy it again — but I bought it before I knew that was all it was, which is a different and worse reason.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Record a human before you clone one.&lt;/strong&gt; The best voice in the film came out of a phone held at arm's length in a living room with a fridge in it, and it beat the GPU on every measure I could think of to check. I built the cloning pipeline first and discovered that second.&lt;/p&gt;




&lt;h2&gt;
  
  
  Reference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Link&lt;/th&gt;
&lt;th&gt;Licence / cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Google Flow&lt;/td&gt;
&lt;td&gt;Video generation front-end, ingredients, upscaling&lt;/td&gt;
&lt;td&gt;&lt;a href="https://labs.google/flow/about" rel="noopener noreferrer"&gt;labs.google/flow&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Free tier 50 credits/day; Google AI Pro 1,000/month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Veo 3.1&lt;/td&gt;
&lt;td&gt;The video model itself (picture + native audio)&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/video" rel="noopener noreferrer"&gt;ai.google.dev/gemini-api/docs/video&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Credit-metered; Lite = 10 credits / 8 s clip&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini API pricing&lt;/td&gt;
&lt;td&gt;Credit and token rates&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;ai.google.dev/gemini-api/docs/pricing&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana (Gemini image generation)&lt;/td&gt;
&lt;td&gt;Character sheets, prop ingredients&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/image-generation" rel="noopener noreferrer"&gt;ai.google.dev/gemini-api/docs/image-generation&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Free tier in the Gemini app; API metered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MusicGen Medium&lt;/td&gt;
&lt;td&gt;Score generation, 1.5 B params&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/facebook/musicgen-medium" rel="noopener noreferrer"&gt;huggingface.co/facebook/musicgen-medium&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Weights CC-BY-NC 4.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AudioCraft&lt;/td&gt;
&lt;td&gt;MusicGen's reference implementation&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/facebookresearch/audiocraft" rel="noopener noreferrer"&gt;github.com/facebookresearch/audiocraft&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Code MIT, weights CC-BY-NC 4.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HuggingFace Transformers&lt;/td&gt;
&lt;td&gt;How MusicGen was actually invoked&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/docs/transformers/model_doc/musicgen" rel="noopener noreferrer"&gt;huggingface.co/docs/transformers/model_doc/musicgen&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Apache 2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;F5-TTS&lt;/td&gt;
&lt;td&gt;English zero-shot voice cloning&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/SWivid/F5-TTS" rel="noopener noreferrer"&gt;github.com/SWivid/F5-TTS&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Open source, free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IndicF5&lt;/td&gt;
&lt;td&gt;Malayalam voice cloning, 11 Indian languages&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/ai4bharat/IndicF5" rel="noopener noreferrer"&gt;huggingface.co/ai4bharat/IndicF5&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs&lt;/td&gt;
&lt;td&gt;Stock-voice narration; cloning evaluated&lt;/td&gt;
&lt;td&gt;&lt;a href="https://elevenlabs.io" rel="noopener noreferrer"&gt;elevenlabs.io&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Free tier excludes cloning; Starter $6/mo up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs use policy&lt;/td&gt;
&lt;td&gt;Voice-cloning restrictions, minors&lt;/td&gt;
&lt;td&gt;&lt;a href="https://elevenlabs.io/use-policy" rel="noopener noreferrer"&gt;elevenlabs.io/use-policy&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ElevenLabs Scribe&lt;/td&gt;
&lt;td&gt;Malayalam transcription that worked&lt;/td&gt;
&lt;td&gt;&lt;a href="https://elevenlabs.io/speech-to-text" rel="noopener noreferrer"&gt;elevenlabs.io/speech-to-text&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Free tier available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;faster-whisper&lt;/td&gt;
&lt;td&gt;English word-level timestamps for auto-editing&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/SYSTRAN/faster-whisper" rel="noopener noreferrer"&gt;github.com/SYSTRAN/faster-whisper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Whisper large-v3&lt;/td&gt;
&lt;td&gt;ASR weights (failed on Malayalam)&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/openai/whisper-large-v3" rel="noopener noreferrer"&gt;huggingface.co/openai/whisper-large-v3&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Apache 2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PyTorch&lt;/td&gt;
&lt;td&gt;GPU runtime; cu128 wheels for Blackwell&lt;/td&gt;
&lt;td&gt;&lt;a href="https://pytorch.org/get-started/locally/" rel="noopener noreferrer"&gt;pytorch.org/get-started/locally&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;BSD-3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FFmpeg — filters&lt;/td&gt;
&lt;td&gt;sidechaincompress, loudnorm, acrossfade, afftdn, adeclick, deesser, drawtext, overlay, scale2ref&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ffmpeg.org/ffmpeg-filters.html" rel="noopener noreferrer"&gt;ffmpeg.org/ffmpeg-filters.html&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LGPL/GPL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FFmpeg — formats&lt;/td&gt;
&lt;td&gt;concat demuxer, lossless assembly&lt;/td&gt;
&lt;td&gt;&lt;a href="https://ffmpeg.org/ffmpeg-formats.html" rel="noopener noreferrer"&gt;ffmpeg.org/ffmpeg-formats.html&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;LGPL/GPL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;libass&lt;/td&gt;
&lt;td&gt;What burns an &lt;code&gt;.srt&lt;/code&gt; into picture, when you insist&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/libass/libass" rel="noopener noreferrer"&gt;github.com/libass/libass&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;ISC&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nirmala UI&lt;/td&gt;
&lt;td&gt;The Windows font with real Malayalam glyph coverage&lt;/td&gt;
&lt;td&gt;&lt;a href="https://learn.microsoft.com/en-us/typography/font-list/nirmala-ui" rel="noopener noreferrer"&gt;learn.microsoft.com/typography/font-list/nirmala-ui&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Ships with Windows&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Alternatives named in this post that I did not use
&lt;/h3&gt;

&lt;p&gt;Everything in this table is &lt;strong&gt;not tried&lt;/strong&gt; — it is the map I wish I'd had, not a list of endorsements. Check the licence yourself before anything commercial; several of these ship permissive code with restricted weights, and the two are licensed separately.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Would replace&lt;/th&gt;
&lt;th&gt;Link&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wan&lt;/td&gt;
&lt;td&gt;Veo, locally, no native audio&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/Wan-Video" rel="noopener noreferrer"&gt;github.com/Wan-Video&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LTX-Video&lt;/td&gt;
&lt;td&gt;Veo, locally, fast&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/Lightricks/LTX-Video" rel="noopener noreferrer"&gt;github.com/Lightricks/LTX-Video&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HunyuanVideo&lt;/td&gt;
&lt;td&gt;Veo, locally&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/Tencent-Hunyuan/HunyuanVideo" rel="noopener noreferrer"&gt;github.com/Tencent-Hunyuan/HunyuanVideo&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CogVideoX&lt;/td&gt;
&lt;td&gt;Veo, locally, modest VRAM&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/THUDM/CogVideo" rel="noopener noreferrer"&gt;github.com/THUDM/CogVideo&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FLUX&lt;/td&gt;
&lt;td&gt;Nano Banana, locally&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/black-forest-labs/flux" rel="noopener noreferrer"&gt;github.com/black-forest-labs/flux&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IP-Adapter&lt;/td&gt;
&lt;td&gt;Face/style conditioning for the character sheets&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/tencent-ailab/IP-Adapter" rel="noopener noreferrer"&gt;github.com/tencent-ailab/IP-Adapter&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;InstantID&lt;/td&gt;
&lt;td&gt;Identity-preserving character generation&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/InstantID/InstantID" rel="noopener noreferrer"&gt;github.com/InstantID/InstantID&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ComfyUI&lt;/td&gt;
&lt;td&gt;The whole local image/video front end&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/comfyanonymous/ComfyUI" rel="noopener noreferrer"&gt;github.com/comfyanonymous/ComfyUI&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stable Audio Open&lt;/td&gt;
&lt;td&gt;MusicGen&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/stabilityai/stable-audio-open-1.0" rel="noopener noreferrer"&gt;huggingface.co/stabilityai/stable-audio-open-1.0&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;YouTube Audio Library&lt;/td&gt;
&lt;td&gt;Generating a score at all&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.youtube.com/audiolibrary" rel="noopener noreferrer"&gt;youtube.com/audiolibrary&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Free Music Archive / Incompetech&lt;/td&gt;
&lt;td&gt;Cleared library music&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://freemusicarchive.org" rel="noopener noreferrer"&gt;freemusicarchive.org&lt;/a&gt; · &lt;a href="https://incompetech.com" rel="noopener noreferrer"&gt;incompetech.com&lt;/a&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Piper&lt;/td&gt;
&lt;td&gt;F5-TTS, CPU-only, no cloning&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/rhasspy/piper" rel="noopener noreferrer"&gt;github.com/rhasspy/piper&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kokoro&lt;/td&gt;
&lt;td&gt;F5-TTS, small and natural&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/hexgrad/Kokoro-82M" rel="noopener noreferrer"&gt;huggingface.co/hexgrad/Kokoro-82M&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;XTTS-v2&lt;/td&gt;
&lt;td&gt;F5-TTS, multilingual cloning — licence needs reading&lt;/td&gt;
&lt;td&gt;&lt;a href="https://huggingface.co/coqui/XTTS-v2" rel="noopener noreferrer"&gt;huggingface.co/coqui/XTTS-v2&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenVoice&lt;/td&gt;
&lt;td&gt;F5-TTS, tone-colour cloning&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/myshell-ai/OpenVoice" rel="noopener noreferrer"&gt;github.com/myshell-ai/OpenVoice&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chatterbox&lt;/td&gt;
&lt;td&gt;F5-TTS, expressive cloning&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/resemble-ai/chatterbox" rel="noopener noreferrer"&gt;github.com/resemble-ai/chatterbox&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Montreal Forced Aligner&lt;/td&gt;
&lt;td&gt;Duration estimation and subtitle alignment&lt;/td&gt;
&lt;td&gt;&lt;a href="https://montreal-forced-aligner.readthedocs.io" rel="noopener noreferrer"&gt;montreal-forced-aligner.readthedocs.io&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;aeneas&lt;/td&gt;
&lt;td&gt;Lighter-weight text↔audio alignment&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/readbeyond/aeneas" rel="noopener noreferrer"&gt;github.com/readbeyond/aeneas&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DaVinci Resolve (free)&lt;/td&gt;
&lt;td&gt;The FFmpeg assembly and mix&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.blackmagicdesign.com/products/davinciresolve" rel="noopener noreferrer"&gt;blackmagicdesign.com/products/davinciresolve&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Premiere Pro&lt;/td&gt;
&lt;td&gt;The FFmpeg assembly and mix, with a timeline&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.adobe.com/products/premiere.html" rel="noopener noreferrer"&gt;adobe.com/products/premiere&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Final Cut Pro&lt;/td&gt;
&lt;td&gt;The FFmpeg assembly and mix, on a Mac&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.apple.com/final-cut-pro/" rel="noopener noreferrer"&gt;apple.com/final-cut-pro&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CapCut&lt;/td&gt;
&lt;td&gt;The whole edit, on a phone or the web&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.capcut.com" rel="noopener noreferrer"&gt;capcut.com&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex CLI&lt;/td&gt;
&lt;td&gt;Claude Code, in the terminal — Apache-2.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/openai/codex" rel="noopener noreferrer"&gt;github.com/openai/codex&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Antigravity CLI&lt;/td&gt;
&lt;td&gt;Claude Code — Google's agent platform, not open source&lt;/td&gt;
&lt;td&gt;&lt;a href="https://antigravity.google" rel="noopener noreferrer"&gt;antigravity.google&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenCode&lt;/td&gt;
&lt;td&gt;Claude Code, open source and model-agnostic — MIT&lt;/td&gt;
&lt;td&gt;&lt;a href="https://opencode.ai" rel="noopener noreferrer"&gt;opencode.ai&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aider&lt;/td&gt;
&lt;td&gt;Claude Code, as an interactive pair-programmer — Apache-2.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://aider.chat" rel="noopener noreferrer"&gt;aider.chat&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi Code CLI&lt;/td&gt;
&lt;td&gt;Claude Code, from Moonshot — MIT&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/MoonshotAI/kimi-code" rel="noopener noreferrer"&gt;github.com/MoonshotAI/kimi-code&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shotcut&lt;/td&gt;
&lt;td&gt;The FFmpeg assembly, lighter&lt;/td&gt;
&lt;td&gt;&lt;a href="https://shotcut.org" rel="noopener noreferrer"&gt;shotcut.org&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audacity&lt;/td&gt;
&lt;td&gt;The narration cut and clean-up, by ear&lt;/td&gt;
&lt;td&gt;&lt;a href="https://www.audacityteam.org" rel="noopener noreferrer"&gt;audacityteam.org&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C2PA Content Credentials&lt;/td&gt;
&lt;td&gt;Provenance, where a visible watermark can't help&lt;/td&gt;
&lt;td&gt;&lt;a href="https://c2pa.org" rel="noopener noreferrer"&gt;c2pa.org&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Colab&lt;/td&gt;
&lt;td&gt;The local GPU&lt;/td&gt;
&lt;td&gt;&lt;a href="https://colab.research.google.com" rel="noopener noreferrer"&gt;colab.research.google.com&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  The acceptance test
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg79rrhazirbzckcvbjyy.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg79rrhazirbzckcvbjyy.webp" alt="The last shot: both children asleep in their bunk beds, moonlight through the window, capes hung on the bedpost" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shot 20. The capes are on the bedpost. The moon is back where it belongs. The film is over in 200.04 seconds, which is roughly the length of a bedtime story.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Lia has now watched it enough times to recite the narration ahead of it.&lt;/p&gt;

&lt;p&gt;Airik, who is two and does not yet care about ingredient drift or byte-length duration estimation, points at the screen and says his own name.&lt;/p&gt;

&lt;p&gt;That was the acceptance test. It passed.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;How this was made: the picture came from Veo 3.1 via Google Flow, the score from MusicGen, and the narration from two people reading into a phone. No video editor was ever opened — the film was assembled, mixed, subtitled and delivered entirely with FFmpeg scripts generated using Claude Code, running Opus. Every script named in this post lives in &lt;code&gt;05_TOOLS/&lt;/code&gt;, and the whole run has since been turned into six reusable Claude Code skills. The taste calls, the re-rolls and the voices are human.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>videoproduction</category>
      <category>tutorial</category>
      <category>automation</category>
    </item>
  </channel>
</rss>
