TL;DR
I spent one evening building a pipeline that turns any published article into a 9:16 Reel. It is narrated in my cloned voice, captioned with karaoke-style word highlights, branded with Digital Craft Workshop colors, and uploaded to Cloudflare R2 ready to post on Instagram or YouTube Shorts.
The whole thing lives behind a single button in Article Forge. Regenerating a reel costs zero API credits, because the scenario and audio are cached per article.
Why I Bothered
I write one article a week. Sometimes more. Each one already takes hours.
Cutting a vertical video on top of that with the usual tools means opening Premiere, writing a script, recording a voice over, and syncing captions by hand. That is enough friction that I never do it.
But Reels and Shorts feed the algorithm. A 30-second hook with my voice out-performs a static link share by a wide margin, and I was leaving all of that on the table for one dumb reason: the editing tax.
So I built the friction out.
The Pipeline, End to End
A single POST request to /api/articles/[id]/video kicks off the whole thing. The render runs in the background; the editor UI polls and shows a preview when the MP4 is ready.
Under the hood it runs ten coordinated steps:
- Postgres and Drizzle track the render job through its states — queued, then processing, then done.
- Cloudflare R2 stores the final MP4 and the audio assets.
- Claude Haiku 4.5 turns the article into a 3-to-5 scene script with a hook, body beats, and a closing call to action.
- node-html-parser scrapes the cover and diagram images from the published Substack post, so the reel reuses the brand-styled diagrams I already shipped instead of generic stock art.
- ElevenLabs narrates each scene in my cloned voice on the $5/month Starter plan. The
/with-timestampsendpoint returns per-word alignment, which the captions need later. - sharp and SVG render each scene as three PNG layers: the background image, the static chrome (header plus progress dots), and the foreground content overlay.
- FFmpeg assembles the per-scene MP4 with a Ken Burns zoom on the background, a fade-in on the content, and a 0.4-second silent pause between scenes so the ear can land before the next narration starts.
- libass burns in the ASS karaoke captions, turning each word red as it is spoken and locking the line to a fixed bottom anchor so it never jumps as scene lengths change.
- Loudnorm at -14 LUFS matches the reel to Instagram and Shorts feed loudness. Skip this and the reel plays quieter than everything around it in the feed.
- An R2 upload publishes the final MP4, and the UI panel renders it inline with a download link.
The ten stages that turn a published article into a captioned 9:16 Reel. | Generated with Claude
The whole thing takes 45 to 60 seconds per reel.
I broke this build down one stage at a time in a free email series, Reel Pipeline Series, if you want the version with the actual prompts and config.
The Hardest Part Was Not What I Expected
The hardest part wasn't the AI. It was the captions.
ASS subtitles need libass, and Homebrew's stock FFmpeg formula ships without it. Generic FFmpeg overlay filters handle static images fine, but they blow up exponentially once you try to layer 30 word-level frame overlays for karaoke timing.
I spent a frustrating hour on that before the fix clicked: uninstall the stock formula, install homebrew-ffmpeg/ffmpeg/ffmpeg from the dedicated tap, and libass comes along for free.
The second caption problem was the line jumping. ASS bottom-anchor alignment plus a variable number of lines means a one-line caption sits at a different height than a two-line caption. I fixed it by pinning every dialogue event with an explicit \pos(540,1800) override — same anchor, every scene, no jumping.
The Surprise Win Was Caching
I added per-article caching on the second iteration, after noticing that every "regenerate" was re-calling Haiku and re-running six ElevenLabs synthesis requests. Each regeneration was burning roughly $0.10 of credits, which does not sound like much until you are iterating on a visual tweak for the tenth time.
Now the scenario JSON and the per-scene audio buffers live in Postgres and R2. The next time I click "Generate Reel" for the same article, the render reuses both and only re-runs FFmpeg.
Zero API calls. Zero credits spent. I can iterate on visual tweaks — a different brand color, a different motion variant — without paying for any of it.
If I want a genuinely fresh take, I hit POST /api/articles/[id]/video?fresh=1 and the pipeline goes all the way back to Haiku.
What I'd Build Next
Three things sit on the wishlist.
- Swap Haiku for
gpt-4o-minifor the script step — same JSON output, cheaper per call. It is on the list, just not done. - A "Post to Instagram" button on the panel that pushes the MP4 straight through the Graph API, so I stop downloading and re-uploading by hand.
- AI-generated background images per scene for the beats that don't have a Substack diagram to reuse, using a cheap fast model like Flux Schnell.
None of these are blocking. The pipeline already does the job I built it for, and I would rather ship reels than polish the tool that ships reels.
The Real Lesson
The pipeline is not impressive for any single component. What matters is that the gap between "I published an article" and "I have a Reel" is now thirty seconds of clicking.
That changes my behavior. I post the reel. The algorithm rewards the post. The next reel gets a little more reach, and the flywheel keeps turning.
This is the part of building in public that took me longest to accept. The build is the easy half. Being public consistently is the hard half, and the only way I stay consistent is to delete every excuse not to post before it can talk me out of it.
This pipeline deletes one of them.
External Sources
- ElevenLabs API — voice cloning and per-word timestamps
- FFmpeg — per-scene assembly and Ken Burns motion
- libass — ASS subtitle rendering for karaoke captions
- sharp — SVG-to-PNG scene layer compositing
- Cloudflare R2 — object storage for the finished MP4
One More Thing
— Daniel
P.S. If you want the whole build as a step-by-step log, I packaged it as a free email series, Reel Pipeline Series, one email per stage.
I turn articles like this one into short vertical videos with my own pipeline. The free playbook is here: https://danielrusnok.gumroad.com/l/article-to-reel-playbook

Top comments (0)