<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: javav</title>
    <description>The latest articles on DEV Community by javav (@bodrumundenizibeep).</description>
    <link>https://dev.to/bodrumundenizibeep</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174179%2F8dc99f33-7196-4032-a313-f90313b22861.png</url>
      <title>DEV Community: javav</title>
      <link>https://dev.to/bodrumundenizibeep</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bodrumundenizibeep"/>
    <language>en</language>
    <item>
      <title>How I added a local AI to my free video app (no API key, runs on your PC)</title>
      <dc:creator>javav</dc:creator>
      <pubDate>Fri, 09 Oct 2026 20:22:27 +0000</pubDate>
      <link>https://dev.to/bodrumundenizibeep/how-i-added-a-local-ai-to-my-free-video-app-no-api-key-runs-on-your-pc-40dl</link>
      <guid>https://dev.to/bodrumundenizibeep/how-i-added-a-local-ai-to-my-free-video-app-no-api-key-runs-on-your-pc-40dl</guid>
      <description>&lt;p&gt;I build &lt;strong&gt;AI Video Studio&lt;/strong&gt;, a free and open-source Windows app that turns a script into a short video: stock footage, a neural voice and word-by-word captions.&lt;/p&gt;

&lt;p&gt;The feedback I kept getting was that writing the script was the annoying part. Every line had to be in the app's own &lt;code&gt;Visual:&lt;/code&gt; / &lt;code&gt;Voice:&lt;/code&gt; format. So I wanted the app to write the script itself from plain text.&lt;/p&gt;

&lt;p&gt;The obvious way is to call a hosted AI service. I didn't want that: the app already needs a free stock-footage key, and every extra key is another reason for someone to give up before making their first video. I wanted an AI that ships with the app and runs on the user's own computer.&lt;/p&gt;

&lt;p&gt;Here is what I learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; Qwen3 4B, as a 4-bit GGUF file (2.5 GB), Apache 2.0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runtime:&lt;/strong&gt; llama.cpp's &lt;code&gt;llama-server&lt;/code&gt;, MIT licence, bundled with the app.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delivery:&lt;/strong&gt; the model is too big for the installer, so the installer downloads it once and checks its SHA-256. It lives outside the install folder, so app updates never download it again.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The app starts &lt;code&gt;llama-server&lt;/code&gt; as a separate process, sends one request, and kills it straight after. That matters because rendering the video needs the memory back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 1: small models are really small
&lt;/h2&gt;

&lt;p&gt;I tested three sizes on the same texts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Download&lt;/th&gt;
&lt;th&gt;What happened&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen2.5 0.5B&lt;/td&gt;
&lt;td&gt;491 MB&lt;/td&gt;
&lt;td&gt;Ignored the length and put whole sentences where search words should go&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen2.5 1.5B&lt;/td&gt;
&lt;td&gt;1.1 GB&lt;/td&gt;
&lt;td&gt;Mostly copied my sentences back to me, sometimes twice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3 4B&lt;/td&gt;
&lt;td&gt;2.5 GB&lt;/td&gt;
&lt;td&gt;Actually rewrote the text and hit the length&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I wanted the small one to work. It didn't. The 4B model was the smallest that did the job, so that is what ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 2: don't ask for a word count, fix the structure
&lt;/h2&gt;

&lt;p&gt;My first prompt said "write about 80 words". Every model ignored it. Scripts came out anywhere from 20 to 140 words.&lt;/p&gt;

&lt;p&gt;What fixed it was two things together:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Force the output with a &lt;strong&gt;JSON schema&lt;/strong&gt;: a list of scenes, each with &lt;code&gt;visual&lt;/code&gt; (search words) and &lt;code&gt;voice&lt;/code&gt; (one spoken line).&lt;/li&gt;
&lt;li&gt;Set the &lt;strong&gt;exact number of scenes&lt;/strong&gt; in both the prompt and the schema (&lt;code&gt;minItems&lt;/code&gt; equal to &lt;code&gt;maxItems&lt;/code&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A model that can't count words can still be made to write exactly eight short sentences. Eight sentences of about ten words is a 30-second video.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 3: show one example
&lt;/h2&gt;

&lt;p&gt;Adding one worked example as a past user/assistant exchange did more than any extra rule. It also stopped the model copying sentences, because the example shows a rewrite.&lt;/p&gt;

&lt;p&gt;It has a side effect: the model imitates the example's size. That is another reason the fixed scene count is needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 4: always have a way out
&lt;/h2&gt;

&lt;p&gt;A local model fails in ordinary ways: not enough memory, a slow PC, a strange answer. So every failure falls back to a plain sentence splitter with no AI, and the app tells the user why. Nobody ends up with nothing.&lt;/p&gt;

&lt;p&gt;I also added a three-minute timeout and a free-memory check before starting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lesson 5: the GPU is worth one switch
&lt;/h2&gt;

&lt;p&gt;llama.cpp has a Vulkan build that works on NVIDIA, AMD and Intel cards and only adds a few megabytes. I swapped to it and added a single switch, off by default.&lt;/p&gt;

&lt;p&gt;On an RTX 4060, the same script takes about &lt;strong&gt;3 seconds&lt;/strong&gt; on the card and about &lt;strong&gt;23 seconds&lt;/strong&gt; on the processor.&lt;/p&gt;

&lt;p&gt;One detail: I always pass the number of GPU layers explicitly, including &lt;code&gt;0&lt;/code&gt;. Otherwise the runtime may decide to use the card by itself, and "off" has to mean off. If loading on the card fails, the app retries on the processor.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it can't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It is a small model, and it sometimes changes a detail. "Every workday" became "every day" in one of my tests. The app tells users to read the script before rendering.&lt;/li&gt;
&lt;li&gt;It barely changes its wording between audiences, so the audience choice mostly affects the voice and pacing.&lt;/li&gt;
&lt;li&gt;Windows only, for now.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The app is free and open source:&lt;br&gt;
&lt;a href="https://github.com/bodrumundenizi-beep/free-ai-video-generator" rel="noopener noreferrer"&gt;https://github.com/bodrumundenizi-beep/free-ai-video-generator&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A 30-second video made with it:&lt;br&gt;
&lt;a href="https://www.youtube.com/shorts/vBrTKgUT7FM" rel="noopener noreferrer"&gt;https://www.youtube.com/shorts/vBrTKgUT7FM&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you know a better small model for this kind of constrained rewriting, I'd like to hear about it.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>python</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
