DEV Community

javav
javav

Posted on AI-assisted

How I added a local AI to my free video app (no API key, runs on your PC)

I build AI Video Studio, a free and open-source Windows app that turns a script into a short video: stock footage, a neural voice and word-by-word captions.

The feedback I kept getting was that writing the script was the annoying part. Every line had to be in the app's own Visual: / Voice: format. So I wanted the app to write the script itself from plain text.

The obvious way is to call a hosted AI service. I didn't want that: the app already needs a free stock-footage key, and every extra key is another reason for someone to give up before making their first video. I wanted an AI that ships with the app and runs on the user's own computer.

Here is what I learned.

The setup

  • Model: Qwen3 4B, as a 4-bit GGUF file (2.5 GB), Apache 2.0.
  • Runtime: llama.cpp's llama-server, MIT licence, bundled with the app.
  • Delivery: the model is too big for the installer, so the installer downloads it once and checks its SHA-256. It lives outside the install folder, so app updates never download it again.

The app starts llama-server as a separate process, sends one request, and kills it straight after. That matters because rendering the video needs the memory back.

Lesson 1: small models are really small

I tested three sizes on the same texts.

Model Download What happened
Qwen2.5 0.5B 491 MB Ignored the length and put whole sentences where search words should go
Qwen2.5 1.5B 1.1 GB Mostly copied my sentences back to me, sometimes twice
Qwen3 4B 2.5 GB Actually rewrote the text and hit the length

I wanted the small one to work. It didn't. The 4B model was the smallest that did the job, so that is what ships.

Lesson 2: don't ask for a word count, fix the structure

My first prompt said "write about 80 words". Every model ignored it. Scripts came out anywhere from 20 to 140 words.

What fixed it was two things together:

  1. Force the output with a JSON schema: a list of scenes, each with visual (search words) and voice (one spoken line).
  2. Set the exact number of scenes in both the prompt and the schema (minItems equal to maxItems).

A model that can't count words can still be made to write exactly eight short sentences. Eight sentences of about ten words is a 30-second video.

Lesson 3: show one example

Adding one worked example as a past user/assistant exchange did more than any extra rule. It also stopped the model copying sentences, because the example shows a rewrite.

It has a side effect: the model imitates the example's size. That is another reason the fixed scene count is needed.

Lesson 4: always have a way out

A local model fails in ordinary ways: not enough memory, a slow PC, a strange answer. So every failure falls back to a plain sentence splitter with no AI, and the app tells the user why. Nobody ends up with nothing.

I also added a three-minute timeout and a free-memory check before starting.

Lesson 5: the GPU is worth one switch

llama.cpp has a Vulkan build that works on NVIDIA, AMD and Intel cards and only adds a few megabytes. I swapped to it and added a single switch, off by default.

On an RTX 4060, the same script takes about 3 seconds on the card and about 23 seconds on the processor.

One detail: I always pass the number of GPU layers explicitly, including 0. Otherwise the runtime may decide to use the card by itself, and "off" has to mean off. If loading on the card fails, the app retries on the processor.

What it can't do

  • It is a small model, and it sometimes changes a detail. "Every workday" became "every day" in one of my tests. The app tells users to read the script before rendering.
  • It barely changes its wording between audiences, so the audience choice mostly affects the voice and pacing.
  • Windows only, for now.

Try it

The app is free and open source:
https://github.com/bodrumundenizi-beep/free-ai-video-generator

A 30-second video made with it:
https://www.youtube.com/shorts/vBrTKgUT7FM

If you know a better small model for this kind of constrained rewriting, I'd like to hear about it.

Top comments (0)