DEV Community

Cover image for What I learned building an AI that turns one URL into a week of social videos
Samuel Bezerra
Samuel Bezerra

Posted on Fully Autonomous

What I learned building an AI that turns one URL into a week of social videos

For the last two years I've been building Spread Out: you paste your website, and it plans a week of posts, generates the short videos, carousels and podcasts, and publishes them to Instagram, TikTok, YouTube, LinkedIn, Facebook, X and Threads.

V2 launches on Product Hunt tomorrow, so this felt like a good moment to write down the parts that were harder than they looked. None of this is a tutorial. It's the list I wish I'd had at the start.

1. Generation is the easy part. Not repeating yourself is hard.

The first version could make a decent video from a URL. The problem showed up in week three: the same brand got the same hook, the same "Did you know…" opener and the same b-roll, over and over. A single video looked fine. A feed of them looked like a bot.

What fixed it wasn't a better model. It was memory:

  • A "creative director" step picks a format for each post from about 20 (talking-head UGC, countdowns, app demos, podcast clips, carousels…) and keeps a per-brand history, so the same angle doesn't come back the following week.
  • B-roll clips, background music and opening lines are all tracked per brand, and the planner gets told what was already used.
  • Topics are anchored to a "territory" derived once from the site, because left alone the planner slowly drifts away from the product into generic motivational content.

If you're building anything that generates content on a schedule, budget more time for state than for prompts.

2. Make the script sound like one person talking

LLM-written narration has a recognisable rhythm: three-item lists, "it's not just X, it's Y", every sentence the same length. Viewers scroll past it in a second.

Two things helped more than any prompt rewrite:

  • A separate editor pass that only rewrites for speech: contractions, uneven sentence lengths, no labels like "Fact 2:" read out loud.
  • A small detector for the usual AI tells that runs on every caption and script. When it fires, the text goes back for another pass instead of going out.

Small detail that mattered: overlay cards (facts, numbers) are timed to the word-level captions, so a card shows up when the voice actually says it rather than on a fixed beat.

3. Self-hosting the models changed the economics, and added a queue problem

Video comes from LTX-2.5 (it generates the speech natively, so talking-head clips don't need a separate TTS track), and stills come from Qwen-Image, both running on our own GPU servers behind a Gradio API. Per-call APIs are simpler to start with, but at "a week of videos per user" running our own GPUs gives us control over both cost and which models we use.

What I didn't expect was how much of the work became scheduling:

  • Two big models don't fit in VRAM together. The image model has to be explicitly unloaded before the video model loads, or you get an out-of-memory error halfway through a job.
  • Workflows used to run one at a time. Now the job queue runs min(configured concurrency, healthy GPUs) in parallel, and each server tracks which model it currently has loaded so jobs land where the weights already are.
  • Anything over about a minute runs as a background job that the browser follows through server-sent events. Long blocking requests just get killed by the platform's idle timeout.

4. Motion graphics as HTML

Captions, cards and transitions are HTML/CSS compositions animated with GSAP, then rendered frame by frame to video on the GPU servers. That means designers can iterate in a browser, and the same components drive carousels and video overlays.

The catch: the renderer seeks the timeline to arbitrary frames, so animations have to be pure functions of time. Anything that reads the DOM state or sets opacity outside the timeline will look right in the browser and flicker in the render.

5. The bug that took the API down today

This one is fresh. This afternoon the API stopped answering: health checks timing out, the database pool exhausted, simple requests taking six seconds.

The logs were full of this:

Error [ERR_INTERNAL_ASSERTION]: This is caused by either a bug in Node.js
or incorrect usage of Node.js internals.
    at internalConnectMultiple (node:net:1106:3)
    at Timeout.internalConnectMultipleTimeout (node:net:1637:3)
Enter fullscreen mode Exit fullscreen mode

The server image was pinned to Node 20.0.0. In that release, "happy eyeballs" (autoSelectFamily, on by default since Node 20) can hit an internal assertion when a connection attempt times out. We had a global uncaughtException handler that logged and carried on, which kept the process alive, but the sockets involved were left in a bad state until nothing could connect.

The immediate fix was a restart. The real fix is one line at startup:

import net from 'node:net';

net.setDefaultAutoSelectFamily(false);
Enter fullscreen mode Exit fullscreen mode

(Upgrading Node also fixes it. If you're on an early 20.x and seeing internalConnectMultiple in your logs, check this first.) The broader lesson: an uncaughtException handler that swallows everything turns a crash, which restarts cleanly, into a slow freeze, which doesn't.

6. Publishing is its own product

Scheduling posts to seven platforms sounds like plumbing. In practice: tokens expire, accounts get disconnected, one platform rejects videos over a certain length, another wants a verified account for long uploads. We ended up treating "generated" and "published" as separate states, checking the account's auth before posting, and surfacing failures in the UI instead of retrying silently.

What's next

V2 adds blog-to-social (new RSS articles become posts automatically), AI replies to comments and Instagram keyword-to-DM, an MCP server so you can drive it from Claude or Cursor, and bring-your-own keys for fal, Replicate and Higgsfield.

If you want to try it, the free plan has 350 credits and doesn't need a card: spreadout.ai. We're on Product Hunt tomorrow, and I'd love to hear where the output still feels machine-made. That's the feedback I act on first.

Happy to answer questions about any of the above in the comments.

Top comments (0)