<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lars Saleh</title>
    <description>The latest articles on DEV Community by Lars Saleh (@larssaleh).</description>
    <link>https://dev.to/larssaleh</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4136764%2F58b3bb67-7ed0-4be0-8b64-9b7dee600110.png</url>
      <title>DEV Community: Lars Saleh</title>
      <link>https://dev.to/larssaleh</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/larssaleh"/>
    <language>en</language>
    <item>
      <title>Open Generative AI GitHub Repo: A Self-Hosting Teardown (2026)</title>
      <dc:creator>Lars Saleh</dc:creator>
      <pubDate>Wed, 07 Oct 2026 03:13:24 +0000</pubDate>
      <link>https://dev.to/larssaleh/open-generative-ai-github-repo-a-self-hosting-teardown-2026-ajg</link>
      <guid>https://dev.to/larssaleh/open-generative-ai-github-repo-a-self-hosting-teardown-2026-ajg</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; If you search "open generative ai github" in 2026, the repo you land on is &lt;code&gt;Anil-matcha/Open-Generative-AI&lt;/code&gt;: an MIT-licensed Next.js and Electron studio for image, video, and lip sync generation, at roughly 29.7k stars and 5.4k forks as of October 6. The UI is open source and self-hostable. Most of the models behind it are not. Cloud generation goes through the Muapi.ai API with your own key, and truly local inference only exists in the desktop app, through two engines with real hardware requirements.&lt;/p&gt;

&lt;p&gt;Some context on why I care. I build media features into client apps, mostly backend work: queues, webhooks, storage. When a client wants text-to-video in production, I usually call a hosted API such as &lt;a href="https://aivideoapi.com/?utm_source=devto&amp;amp;utm_medium=ugc&amp;amp;utm_campaign=larssaleh&amp;amp;utm_content=open-generative-ai-github-intro" rel="noopener noreferrer"&gt;AI Video API&lt;/a&gt; rather than run video models on our own GPUs. Open Generative AI keeps coming up as the "free, self-hosted" alternative, so I read through the repo, the package scripts, and the release assets to see what you actually get.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What the Open Generative AI GitHub repo is
&lt;/h2&gt;

&lt;p&gt;The dev.to article that currently ranks for this search is a January 2025 list of three image projects: CompVis Stable Diffusion, DALL-E Mini, and StyleGAN3. Those are model repos. Open Generative AI is a different kind of thing: a front end that puts hundreds of hosted models behind one interface.&lt;/p&gt;

&lt;p&gt;From the README and &lt;code&gt;package.json&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stack:&lt;/strong&gt; Next.js 14 (App Router), React 18, Tailwind CSS v3, npm workspaces. A shared &lt;code&gt;packages/studio&lt;/code&gt; component library holds the studios, and &lt;code&gt;packages/studio/src/models.js&lt;/code&gt; is the single list of model definitions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Studios:&lt;/strong&gt; Image, Video, Audio, Lip Sync, Cinema (camera, lens, focal length, and aperture presets turned into prompt modifiers), Workflow (node-based pipelines), Agent, Design Agent, Marketing, AI Clipping, Body Swap, and a few more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distribution:&lt;/strong&gt; a web build (&lt;code&gt;npm run dev&lt;/code&gt;), an Electron desktop build, and a hosted copy on muapi.ai.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;License:&lt;/strong&gt; MIT.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same &lt;code&gt;packages/studio&lt;/code&gt; library powers the hosted version on muapi.ai, which explains why the model list moves so fast. Model updates land in one file and ship to both.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. What's open and what isn't
&lt;/h2&gt;

&lt;p&gt;This is the part the repo description ("Self-hosted, MIT licensed") glosses over.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Where it runs&lt;/th&gt;
&lt;th&gt;What you need&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;UI (all studios)&lt;/td&gt;
&lt;td&gt;Your machine or server&lt;/td&gt;
&lt;td&gt;Node.js 18+, or the desktop installer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cloud models (Flux, Kling, Veo, Sora, Seedance, Midjourney...)&lt;/td&gt;
&lt;td&gt;Muapi's API&lt;/td&gt;
&lt;td&gt;A Muapi access key, billed by Muapi&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sd.cpp local engine&lt;/td&gt;
&lt;td&gt;Your machine, desktop app only&lt;/td&gt;
&lt;td&gt;CPU works; Metal on Apple Silicon, CUDA/Vulkan/ROCm elsewhere&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wan2GP local engine&lt;/td&gt;
&lt;td&gt;Your GPU box, desktop app only&lt;/td&gt;
&lt;td&gt;Your own Wan2GP install on a CUDA or ROCm GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted web version&lt;/td&gt;
&lt;td&gt;muapi.ai&lt;/td&gt;
&lt;td&gt;A free account; always uses cloud APIs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The API flow is documented in the README. The app submits a job with &lt;code&gt;POST /api/v1/{model-endpoint}&lt;/code&gt;, then polls &lt;code&gt;GET /api/v1/predictions/{request_id}/result&lt;/code&gt; until the status is &lt;code&gt;completed&lt;/code&gt;, authenticating with an &lt;code&gt;x-api-key&lt;/code&gt; header. Your key sits in browser &lt;code&gt;localStorage&lt;/code&gt; and, per the README, is only sent to Muapi.&lt;/p&gt;

&lt;p&gt;So "free" means the code costs nothing. Generations on the big commercial models cost whatever Muapi charges for them. That's a perfectly reasonable design, a bring-your-own-key client, but you should budget for it before you tell a stakeholder it's free.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Running it from source
&lt;/h2&gt;

&lt;p&gt;The README is clear that most people should grab a prebuilt installer. If you want to hack on it, this is the path, and I checked each script name against &lt;code&gt;package.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# submodules are required for the workflow and agent packages&lt;/span&gt;
git clone &lt;span class="nt"&gt;--recurse-submodules&lt;/span&gt; https://github.com/Anil-matcha/Open-Generative-AI.git
&lt;span class="nb"&gt;cd &lt;/span&gt;Open-Generative-AI

&lt;span class="c"&gt;# "setup" = git submodule update + npm install + build all workspace packages&lt;/span&gt;
&lt;span class="c"&gt;# plain `npm install` is not enough&lt;/span&gt;
npm run setup

&lt;span class="c"&gt;# then ONE of these&lt;/span&gt;
npm run electron:dev   &lt;span class="c"&gt;# desktop app (Vite build, then Electron)&lt;/span&gt;
npm run dev            &lt;span class="c"&gt;# web version on http://localhost:3000&lt;/span&gt;

&lt;span class="c"&gt;# production web build&lt;/span&gt;
npm run build &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm run start
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If Next.js complains it can't find a &lt;code&gt;pages&lt;/code&gt; directory, you're either not in the repo root or you cloned without submodules. Re-run &lt;code&gt;npm run setup&lt;/code&gt; if &lt;code&gt;packages/Vibe-Workflow&lt;/code&gt; or &lt;code&gt;packages/agents&lt;/code&gt; are empty.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Local inference: two engines, very different costs
&lt;/h2&gt;

&lt;p&gt;Local generation is desktop-only. The hosted and web builds always call cloud APIs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;sd.cpp (bundled).&lt;/strong&gt; Built on stable-diffusion.cpp. You install it with one click under Settings, then Local Models. It handles image models only: Dreamshaper 8, Realistic Vision 5.1, and Anything v5 (SD 1.5, about 2.1 GB each), SDXL Base 1.0 (6.9 GB), and Z-Image Turbo and Base, which also need a 2.4 GB Qwen3-4B text encoder and a 335 MB FLUX VAE. The README warns that Z-Image is known to hang a base 8 GB M-series Mac and recommends 16 GB of RAM. On 8 GB, stick to SD 1.5.&lt;/p&gt;

&lt;p&gt;Weights default to Electron's app-data folder (&lt;code&gt;~/Library/Application Support/open-generative-ai/local-ai&lt;/code&gt; on macOS, &lt;code&gt;~/.config/open-generative-ai/local-ai&lt;/code&gt; on Linux). If you'd rather keep multi-GB weights on another drive, set &lt;code&gt;OPEN_GENERATIVE_AI_LOCAL_AI_DIR&lt;/code&gt; before launch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wan2GP (bring your own).&lt;/strong&gt; This is where video lives. The app doesn't bundle Python or weights. You install Wan2GP yourself on a CUDA or ROCm machine, and the app either runs &lt;code&gt;wgp.py&lt;/code&gt; from a local folder or talks to Wan2GP's MCP server on another box:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python wgp.py &lt;span class="nt"&gt;--mcp&lt;/span&gt; &lt;span class="nt"&gt;--mcp-api-version&lt;/span&gt; 1 &lt;span class="nt"&gt;--mcp-transport&lt;/span&gt; streamable-http &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--mcp-host&lt;/span&gt; 0.0.0.0 &lt;span class="nt"&gt;--mcp-port&lt;/span&gt; 7866
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You paste &lt;code&gt;http://&amp;lt;host&amp;gt;:7866/mcp&lt;/code&gt; into the settings panel. The Gradio UI on port 7860 won't work for this. Only bind to &lt;code&gt;0.0.0.0&lt;/code&gt; on a network you trust.&lt;/p&gt;

&lt;p&gt;The remote mode is the clever part. Wan2GP has no Apple Silicon path, so a Mac user can keep the desktop app and push inference to a Linux GPU box or a rented instance. Supported models include Wan 2.1 1.3B (480p, about 8 GB VRAM, the lightest option), Wan 2.2 14B T2V/I2V, Hunyuan Video 1.5, LTX-2 Distilled (720p), plus Flux.1 Dev and Qwen Image for stills.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Gotchas I'd check before relying on it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The numbers in the README disagree with each other.&lt;/strong&gt; The tagline says 400+ models, a features bullet says 200+, and the category table says 420+ "verified against models.js." Lip sync is "9 dedicated models" in one section and 15 in the category table. Count &lt;code&gt;models.js&lt;/code&gt; yourself if the number matters to you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The download table is stale.&lt;/strong&gt; It still links 1.0.9 installers, while the latest GitHub release is v2.0.0 (May 23, 2026), with &lt;code&gt;.dmg&lt;/code&gt; builds for arm64 and Intel, a Windows &lt;code&gt;.exe&lt;/code&gt;, an &lt;code&gt;.AppImage&lt;/code&gt;, and a &lt;code&gt;.deb&lt;/code&gt;. Go to the Releases page directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The binaries are unsigned.&lt;/strong&gt; On macOS you'll need &lt;code&gt;xattr -cr "/Applications/Open Generative AI.app"&lt;/code&gt; or the "Open Anyway" button; on Windows, SmartScreen's "Run anyway."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ubuntu 24.04 AppArmor.&lt;/strong&gt; The AppImage can die silently because of &lt;code&gt;apparmor_restrict_unprivileged_userns&lt;/code&gt;. The &lt;code&gt;.deb&lt;/code&gt; ships an AppArmor profile, so install that instead of flipping the sysctl.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"No content filters"&lt;/strong&gt; describes the UI. Your generations still run on Muapi and on the upstream model providers, so read their terms before you build a product on that promise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The README sells white-label hosting.&lt;/strong&gt; MuAPI's white-label plan starts at $49/mo. It's optional, but expect upsell links throughout the docs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  6. How it compares to the older GitHub picks
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Repo&lt;/th&gt;
&lt;th&gt;Stars (Oct 2026)&lt;/th&gt;
&lt;th&gt;Last push&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anil-matcha/Open-Generative-AI&lt;/td&gt;
&lt;td&gt;29.7k&lt;/td&gt;
&lt;td&gt;Oct 2026&lt;/td&gt;
&lt;td&gt;Multi-model studio UI, BYO key plus local engines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CompVis/stable-diffusion&lt;/td&gt;
&lt;td&gt;73.5k&lt;/td&gt;
&lt;td&gt;Jun 2024&lt;/td&gt;
&lt;td&gt;Original SD research code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;borisdayma/dalle-mini&lt;/td&gt;
&lt;td&gt;14.7k&lt;/td&gt;
&lt;td&gt;Nov 2023&lt;/td&gt;
&lt;td&gt;DALL-E Mini model code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NVlabs/stylegan3&lt;/td&gt;
&lt;td&gt;6.9k&lt;/td&gt;
&lt;td&gt;Sep 2023&lt;/td&gt;
&lt;td&gt;GAN research code, NVIDIA GPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deepbeepmeep/Wan2GP&lt;/td&gt;
&lt;td&gt;10.1k&lt;/td&gt;
&lt;td&gt;Oct 2026&lt;/td&gt;
&lt;td&gt;Local video/image runtime for modest GPUs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;leejet/stable-diffusion.cpp&lt;/td&gt;
&lt;td&gt;7.5k&lt;/td&gt;
&lt;td&gt;Oct 2026&lt;/td&gt;
&lt;td&gt;C++ diffusion inference, powers the sd.cpp engine&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 2025 list is still useful as history. If you want to generate something this week, the three bottom rows are where the activity is, and the top row ties them together behind a UI.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is Open Generative AI on GitHub really free?&lt;/strong&gt;&lt;br&gt;
The code is MIT licensed and free. Cloud generations go through Muapi with your own key, which is billed. Local generation through sd.cpp or Wan2GP costs only your hardware.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I run Open Generative AI fully offline?&lt;/strong&gt;&lt;br&gt;
Only in the desktop app, using sd.cpp for images or a Wan2GP install for video. The web build and the hosted version always use cloud APIs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What GPU do I need for local video?&lt;/strong&gt;&lt;br&gt;
Wan2GP needs a CUDA or ROCm GPU. Wan 2.1 1.3B at 480p is listed at about 8 GB VRAM. Larger models such as Wan 2.2 14B are noted as slow on consumer cards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Open Generative AI the same as Stable Diffusion?&lt;/strong&gt;&lt;br&gt;
No. Stable Diffusion is a model. Open Generative AI is a studio that can run SD 1.5 and SDXL locally and also calls hundreds of hosted models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I landed
&lt;/h2&gt;

&lt;p&gt;For prototyping and comparing models side by side, Open Generative AI is the most complete open-source front end I've looked at, and the Wan2GP-over-MCP setup is worth stealing even if you never use the UI. For a production feature where a user clicks "generate" and expects a video back through a webhook, I still prefer a narrower hosted service. For that job I reach for &lt;a href="https://aivideoapi.com/?utm_source=devto&amp;amp;utm_medium=ugc&amp;amp;utm_campaign=larssaleh&amp;amp;utm_content=open-generative-ai-github-outro" rel="noopener noreferrer"&gt;AI Video API&lt;/a&gt;, which handles text-to-video and image-to-video with callback webhooks, so my backend stays a queue and a handler. The repo is at &lt;a href="https://github.com/Anil-matcha/Open-Generative-AI" rel="noopener noreferrer"&gt;github.com/Anil-matcha/Open-Generative-AI&lt;/a&gt; if you want to read &lt;code&gt;models.js&lt;/code&gt; yourself.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>javascript</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
