<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: QVAC by Tether</title>
    <description>The latest articles on DEV Community by QVAC by Tether (qvac).</description>
    <link>https://dev.to/qvac</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F14103%2Fd035de70-5972-420d-9583-d364a8aa15af.jpg</url>
      <title>DEV Community: QVAC by Tether</title>
      <link>https://dev.to/qvac</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/qvac"/>
    <language>en</language>
    <item>
      <title>VisionPsy-Nano: state-of-the-art vision AI in its weight class, small enough to run on your phone</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Wed, 29 Jul 2026 12:14:46 +0000</pubDate>
      <link>https://dev.to/qvac/visionpsy-nano-state-of-the-art-vision-ai-in-its-weight-class-small-enough-to-run-on-your-phone-fid</link>
      <guid>https://dev.to/qvac/visionpsy-nano-state-of-the-art-vision-ai-in-its-weight-class-small-enough-to-run-on-your-phone-fid</guid>
      <description>&lt;p&gt;Tether AI Research* is releasing &lt;strong&gt;VisionPsy-Nano&lt;/strong&gt;, a family of ~460M-parameter vision-language models built to run on the device in your pocket.&lt;/p&gt;

&lt;p&gt;There are two variants: &lt;strong&gt;VisionPsy-Nano-460M&lt;/strong&gt; when accuracy is the priority, and &lt;strong&gt;VisionPsy-Nano-460M-Flash&lt;/strong&gt; when time-to-first-token, and memory are what matters most.&lt;/p&gt;

&lt;p&gt;Download VisionPsy-Nano now on HuggingFace: &lt;a href="https://huggingface.co/collections/qvac/visionpsy" rel="noopener noreferrer"&gt;https://huggingface.co/collections/qvac/visionpsy&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Best in its weight class
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8byctdno1i8w993zbcdy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8byctdno1i8w993zbcdy.png" alt=" " width="800" height="409"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Against every other ~0.5B vision model tested (LFM2.5-VL-450M, SmolVLM2-500M, and its own nanoVLM-460M-8k base), VisionPsy-Nano-460M leads on &lt;strong&gt;16 of 17 benchmarks&lt;/strong&gt;, with the highest overall normalized score in its class:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Overall normalized score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;VisionPsy-Nano-460M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;62.3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LFM2.5-VL-450M&lt;/td&gt;
&lt;td&gt;59.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;nanoVLM-460M-8k&lt;/td&gt;
&lt;td&gt;54.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SmolVLM2-500M&lt;/td&gt;
&lt;td&gt;52.5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It also leads &lt;strong&gt;all four capability categories&lt;/strong&gt;, not just the average: document understanding and OCR (73.9), visual perception (61.5), reasoning and knowledge (52.2), and instruction following and reliability (65.1). Its widest margins are the two that usually break small models: reasoning and knowledge (+7.4% over the next best ~0.5B model) and visual perception (+4.6%).&lt;/p&gt;

&lt;h2&gt;
  
  
  It punches above its size class
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Famffrfz1diqdv3xmaj53.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Famffrfz1diqdv3xmaj53.png" alt=" " width="799" height="463"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The more interesting result is what happens when you stop comparing it to its own weight class. On three full benchmarks, a 460M model beats every one of the 0.75B to 1B models tested outright, including Apple's FastVLM-0.5B (759M), Qwen3.5-0.8B (873M), and InternVL3.5-1B (1061M), which are 1.6x to 2.3x its size:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;VisionPsy-Nano-460M&lt;/th&gt;
&lt;th&gt;FastVLM-0.5B&lt;/th&gt;
&lt;th&gt;Qwen3.5-0.8B&lt;/th&gt;
&lt;th&gt;InternVL3.5-1B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ScienceQA&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;86.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;84.7&lt;/td&gt;
&lt;td&gt;71.6&lt;/td&gt;
&lt;td&gt;79.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MM-IFEval (instruction following)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;42.3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;21.1&lt;/td&gt;
&lt;td&gt;36.9&lt;/td&gt;
&lt;td&gt;33.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POPE (hallucination robustness)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;87.9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;86.5&lt;/td&gt;
&lt;td&gt;87.3&lt;/td&gt;
&lt;td&gt;86.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;MM-IFEval matters more than it looks: it measures whether the model actually does what you asked. Leading every larger model tested on that axis is the difference between a model you can build a product on and one you have to babysit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Flash: the first token, on a real phone
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyghdd3wiqw7t97eu77w0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyghdd3wiqw7t97eu77w0.png" alt=" " width="800" height="402"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The cost that decides whether on-device vision feels usable is not the model size, it is how long you wait after pointing the camera. That wait is driven by the number of visual tokens the model has to read. Flash uses &lt;strong&gt;64 visual tokens&lt;/strong&gt; where the full model uses 1088, and it keeps the quality: within about 1% of the full model (61.4 vs 62.3 normalized).&lt;/p&gt;

&lt;p&gt;The result, measured on real phones with quantized GGUF builds at 512x512:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Device&lt;/th&gt;
&lt;th&gt;Flash, time to first token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;iPhone 15&lt;/td&gt;
&lt;td&gt;0.3s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Galaxy S25 Ultra&lt;/td&gt;
&lt;td&gt;2.6s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Galaxy S23&lt;/td&gt;
&lt;td&gt;5.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pixel 9&lt;/td&gt;
&lt;td&gt;6.1s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is &lt;strong&gt;11.9x to 25.1x faster to the first token&lt;/strong&gt; than nanoVLM-460M and SmolVLM2-500M on the same runtime, and still &lt;strong&gt;1.3x to 2.4x faster than LFM2.5-VL-450M&lt;/strong&gt; and &lt;strong&gt;2.3x to 3.5x faster than Qwen3.5-0.8B&lt;/strong&gt; across all four devices. Peak memory drops 24% to 39% on Android and 41% to 65% on iPhone 15 against the full-token baselines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Built to run locally, not shrunk to fit
&lt;/h2&gt;

&lt;p&gt;VisionPsy-Nano is not a cloud model trimmed down until it could run on a phone. Running on the device was the design constraint from the start, and it shaped every choice: the parameter budget, the visual-token budget that Flash pulls on, the decision to optimize for a single image per query, and the quantized builds that ship next to the full-precision weights. A phone hands one app a small slice of RAM and no dedicated VRAM, so a model that is going to live there has to be small, quantizable, and fast to the first token. That is what it was built to be.&lt;/p&gt;

&lt;p&gt;What that buys is what local AI is for. The image is understood where it was taken, so it never has to leave the device to be read. The model keeps working with no signal. And it fits the situations it was scoped for: asking questions about what the camera is pointed at, reading documents, charts and diagrams, pulling text out of a scene, and following light visual instructions.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it got there: Failure-driven, not scale-driven
&lt;/h2&gt;

&lt;p&gt;The result did not come from scaling data. It came from finding out precisely where the model was wrong and fixing that.&lt;/p&gt;

&lt;p&gt;Tether AI Research built an agentic failure-discovery harness: a larger teacher model interrogates the small model with realistic questions on real images, grades every answer with the image in view, then clusters what it found into a weakness catalogue that steers the next round. Over 300 rounds and roughly 18,000 probes, it mapped concrete failure modes: misread prices and license plates, reading a chart value correctly but failing to sum a column, confusing left and right, and degenerate repetition loops.&lt;/p&gt;

&lt;p&gt;Those clusters became the training signal. A five-stage post-training pipeline then turns a good base model into a best-in-class one: broad supervised fine-tuning, capability-targeted fine-tuning, weakness-targeted fine-tuning on teacher-verified synthetic data, merging the resulting capability specialists into a single checkpoint with TIES so their strengths accumulate at no extra inference cost, and a final preference-alignment stage that kills the repetition loops.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get it
&lt;/h2&gt;

&lt;p&gt;Open weights under Apache 2.0, these models are intended to support researchers and for educational purposes, with three ways to run them today: full precision via Transformers, quantized GGUF builds for on-device inference through llama.cpp, and vLLM for high-throughput server inference. Support in the QVAC SDK is coming.&lt;/p&gt;

&lt;p&gt;Every model in the report was scored under one VLMEvalKit harness with the eval configs and judge setup published, so the benchmark numbers are reproducible.&lt;/p&gt;




&lt;p&gt;* References to Tether AI Research are references to Tether Data, S.A. de C.V.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;References:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Hugging Face VisionPsy-Nano Collection: &lt;a href="https://huggingface.co/collections/qvac/visionpsy" rel="noopener noreferrer"&gt;https://huggingface.co/collections/qvac/visionpsy&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hugging Face Blog Article: &lt;a href="https://huggingface.co/blog/qvac/visionpsy" rel="noopener noreferrer"&gt;https://huggingface.co/blog/qvac/visionpsy&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>vision</category>
    </item>
    <item>
      <title>Run OpenCode locally with QVAC: a coding agent with no cloud</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Wed, 29 Jul 2026 09:22:43 +0000</pubDate>
      <link>https://dev.to/qvac/run-opencode-locally-with-qvac-a-coding-agent-with-no-cloud-ok4</link>
      <guid>https://dev.to/qvac/run-opencode-locally-with-qvac-a-coding-agent-with-no-cloud-ok4</guid>
      <description>&lt;h2&gt;
  
  
  What OpenCode is
&lt;/h2&gt;

&lt;p&gt;OpenCode is an open-source coding agent. You give it a goal in plain language, and it plans, reads and writes files, runs commands, and reports back. It lives in your terminal or in a browser interface, and it works the way a teammate would: open the repo, make the change, run it, fix what broke.&lt;/p&gt;

&lt;p&gt;The detail that matters for this guide is how OpenCode talks to a model. It speaks to any endpoint that exposes an OpenAI-compatible API. That means it is not tied to a single cloud vendor. You point it at whatever model endpoint you want, including one running on your own machine.&lt;/p&gt;

&lt;p&gt;That is the whole idea here. OpenCode for the agent loop, a local model for the intelligence, both on hardware you control.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2069412641416843506-783" src="https://platform.twitter.com/embed/Tweet.html?id=2069412641416843506"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2069412641416843506-783');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2069412641416843506&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;h2&gt;
  
  
  Why local is the right default for a coding agent
&lt;/h2&gt;

&lt;p&gt;When a model only chats, where it runs is mostly a question of privacy. When a model drives a coding agent that reads your repository and runs commands, it becomes a question of control. The agent acts on your machine. If the model behind it lives in someone else's datacenter, every file it reads and every command it runs depends on a system you do not own: its uptime, its latency, its terms, and its access policy, any of which can change without notice.&lt;/p&gt;

&lt;p&gt;Running the model locally changes the property, not just the policy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your code stays on the machine.&lt;/strong&gt; The files the agent reads, the code it writes, the commands it runs: none of it leaves your disk to reach the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works offline.&lt;/strong&gt; Once the model is downloaded, you can disconnect entirely and keep working. Turn off the network and the agent still answers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No per-token bill.&lt;/strong&gt; A long coding session that would run up a metered API costs nothing locally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No silent model swaps.&lt;/strong&gt; The weights are a file on your disk. They do not change under you between sessions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a coding agent, that adds up to a simple property: the work, and the intelligence doing it, are both yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  QVAC, and the tools you already use
&lt;/h2&gt;

&lt;p&gt;QVAC is the piece that makes the local half practical. It is an open-source SDK and CLI from Tether that runs models on your own device across every major GPU backend (NVIDIA, AMD, and Intel through Vulkan, and Apple Silicon through Metal), and it ships an OpenAI-compatible server. That server is the bridge a coding agent plugs into.&lt;/p&gt;

&lt;p&gt;It is also not a one-tool trick. QVAC is a first-class local provider for OpenCode and OpenClaw, and because the server speaks the standard OpenAI API, it works with the other coding tools developers already use, including Cline, Aider, Continue, and Roo. For OpenCode specifically there is a dedicated plugin, &lt;code&gt;@qvac/opencode-plugin&lt;/code&gt;, that wires the whole thing up from one line of config (below). You bring your own model, running locally, to the agent you already like. QVAC is Apache 2.0, with no API keys, no rate limits, and no per-token cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like
&lt;/h2&gt;

&lt;p&gt;Running Qwen3.6-35B-A3B locally through QVAC (downloaded once, then with the network off), OpenCode handled three ordinary developer tasks back to back:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Build from a prompt.&lt;/strong&gt; From one sentence, it wrote a self-contained page animating eight hundred glowing particles that drift and react to the cursor, then opened it in the browser. No libraries, every line generated on the machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explain an unfamiliar repo.&lt;/strong&gt; Asked what a repository does, it read the code and explained it: a solar-system simulation with planets orbiting on circular paths. Opening the page confirmed it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Find and fix a bug.&lt;/strong&gt; Handed a shopping cart whose total rendered as garbled text instead of a number, it located the cause, edited the file, and the total added up correctly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of it ran locally, on an open and free stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which model to run
&lt;/h2&gt;

&lt;p&gt;Coding agent work (reading a repo, generating a file, fixing a bug) asks more of a model than a chat does, so it pays to run a capable one. You select a model by id in the plugin config (step 2 below), and that id flows straight through to OpenCode's model picker.&lt;/p&gt;

&lt;p&gt;That id is a QVAC model reference, not a path to a file you supply. On first run, QVAC downloads that model once into its own local store and serves it locally from then on, with no network. Because it pulls QVAC's own copy, a &lt;code&gt;.gguf&lt;/code&gt; you may already have from another tool is not reused.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Use case&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;RAM&lt;/th&gt;
&lt;th&gt;Disk&lt;/th&gt;
&lt;th&gt;Config id&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Recommended default, runs on most recent laptops&lt;/td&gt;
&lt;td&gt;Qwen3.5-4B (4-bit)&lt;/td&gt;
&lt;td&gt;~8 GB&lt;/td&gt;
&lt;td&gt;~3 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_5_4B_MULTIMODAL_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Balanced, the plugin's out-of-box default&lt;/td&gt;
&lt;td&gt;Qwen3.5-9B (4-bit)&lt;/td&gt;
&lt;td&gt;~12 GB&lt;/td&gt;
&lt;td&gt;~6 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_5_9B_MULTIMODAL_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Strongest coding, used in the demo&lt;/td&gt;
&lt;td&gt;Qwen3.6-35B-A3B (4-bit)&lt;/td&gt;
&lt;td&gt;~32 GB&lt;/td&gt;
&lt;td&gt;~21 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The demo uses Qwen3.6-35B-A3B: 35 billion parameters of knowledge, only about 3 billion active per token, so it stays fast despite its size. For most laptops, Qwen3.5-4B or the 9B default is the better starting point.&lt;/p&gt;

&lt;p&gt;A GPU helps but is not required. QVAC uses Apple Metal on Apple Silicon and Vulkan on NVIDIA, AMD, and Intel; on a CPU-only machine it still runs, just slower. To check what your hardware supports before you start, run &lt;code&gt;npx -y @qvac/cli doctor&lt;/code&gt; (no install needed).&lt;/p&gt;

&lt;h2&gt;
  
  
  Set it up
&lt;/h2&gt;

&lt;p&gt;No second terminal and no manual server. OpenCode reads the QVAC plugin from your project config, brings up a managed QVAC server by itself, points itself at it, and shuts it down when you quit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Requirement:&lt;/strong&gt; Node.js 22.17 or newer (&lt;code&gt;node --version&lt;/code&gt;, from nodejs.org). Disk and RAM depend on the model you pick: the demo's Qwen3.6-35B-A3B needs about 21 GB of disk and 32 GB of RAM; the lighter &lt;code&gt;qwen3.5-9b&lt;/code&gt; default needs much less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Install OpenCode.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; opencode-ai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Add the QVAC plugin to your project.&lt;/strong&gt; In the folder you want to code in, create an &lt;code&gt;opencode.json&lt;/code&gt;. The whole file is one line, using the plugin's default model (&lt;code&gt;qwen3.5-9b&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"plugin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"@qvac/opencode-plugin"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To match the demo with the strongest model instead (heavier: about 21 GB of disk and 32 GB of RAM), name it explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"plugin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="s2"&gt;"@qvac/opencode-plugin"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Run OpenCode.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;opencode web        &lt;span class="c"&gt;# browser UI, the most visual&lt;/span&gt;
opencode            &lt;span class="c"&gt;# terminal UI&lt;/span&gt;
opencode run &lt;span class="s2"&gt;"..."&lt;/span&gt;  &lt;span class="c"&gt;# one-shot&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole setup. On first run the plugin installs itself, spawns a managed QVAC server in the background (downloading the model once), wires OpenCode to it as the local &lt;code&gt;qvac&lt;/code&gt; provider, and reaps it when you quit. There is no &lt;code&gt;qvac serve&lt;/code&gt; to run yourself, no provider block to write, and no &lt;code&gt;toolsMode&lt;/code&gt; to set: the plugin handles all of it, and loads on startup whatever interface you use.&lt;/p&gt;

&lt;p&gt;A good first test: ask it whether it is running in the cloud or on your machine, then turn off your network and ask again. It keeps answering, because nothing was being sent anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to try next
&lt;/h2&gt;

&lt;p&gt;From here, the same setup handles the rest of what a coding agent does. Point OpenCode at a project folder and have it read and edit code, generate a small app from a prompt, explain an unfamiliar repository, or track down a bug. The agent loop is identical to the cloud experience. The only thing that changed is that the model behind it is yours, on your hardware, with nothing leaving the machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes and limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A model that fits on your machine is not a frontier cloud model. Expect it to need clearer, more focused instructions, and give it single-goal tasks rather than long open-ended ones.&lt;/li&gt;
&lt;li&gt;Local responses can take longer than a cloud API today. On a single local worker, a tool-using turn with the 35B model runs in roughly 20 to 30 seconds; a smaller model is snappier. That gap is closing fast as local hardware speeds up and small models improve.&lt;/li&gt;
&lt;li&gt;One QVAC worker runs per machine, and multiple OpenCode windows share it. If the OpenCode desktop app is open it can hold locks the terminal needs, so quit it when running &lt;code&gt;opencode&lt;/code&gt; from the terminal.&lt;/li&gt;
&lt;li&gt;QVAC is Apache 2.0 and free. The model is downloaded once and cached locally. Source and docs: &lt;a href="https://github.com/tetherto/qvac" rel="noopener noreferrer"&gt;github.com/tetherto/qvac&lt;/a&gt; and &lt;a href="https://docs.qvac.tether.io" rel="noopener noreferrer"&gt;docs.qvac.tether.io&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>coding</category>
    </item>
    <item>
      <title>How to (easily) run a fully local, private AI assistant with OpenClaw and QVAC</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Wed, 29 Jul 2026 09:22:29 +0000</pubDate>
      <link>https://dev.to/qvac/how-to-easily-run-a-fully-local-private-ai-assistant-with-openclaw-and-qvac-2bmp</link>
      <guid>https://dev.to/qvac/how-to-easily-run-a-fully-local-private-ai-assistant-with-openclaw-and-qvac-2bmp</guid>
      <description>&lt;h2&gt;
  
  
  What an agent harness is, and what OpenClaw does
&lt;/h2&gt;

&lt;p&gt;A language model on its own can only produce text. It cannot open a file, run a command, call an API, or remember what it did five minutes ago. An agent harness is the layer that closes that gap. It takes the model's text output, turns it into real actions (read this file, run this command, search this folder), feeds the results back to the model, and loops until the task is done. The model is the brain. The harness is the hands.&lt;/p&gt;

&lt;p&gt;OpenClaw is one of these harnesses, and right now it is the most widely used one. It crossed a large install base in a few months and sits at the top of agent usage charts. You give it a goal in plain language, and it plans, calls tools, and reports back. It speaks to any model that exposes an OpenAI-compatible API, which is the detail that matters for this guide: you are not locked to a single cloud vendor. You point it at whatever model endpoint you want, including one running on your own machine.&lt;/p&gt;

&lt;p&gt;That is the whole idea here. OpenClaw for the agent loop, QVAC for the model, both on hardware you control.&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2067624599270166813-984" src="https://platform.twitter.com/embed/Tweet.html?id=2067624599270166813"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2067624599270166813-984');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2067624599270166813&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;h2&gt;
  
  
  What people use OpenClaw for
&lt;/h2&gt;

&lt;p&gt;The harness is general, so the use cases are broad. The common ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Coding tasks.&lt;/strong&gt; Read a repo, write a function, run the tests, fix what broke. The agent edits files and runs commands directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;File and system chores.&lt;/strong&gt; Rename a batch of files, summarize a folder of documents, reorganize a directory, pull a number out of a spreadsheet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local research and drafting.&lt;/strong&gt; Read a set of notes and produce a summary, a draft, or a structured table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Glue work.&lt;/strong&gt; Chain a few steps together that would otherwise be a manual sequence: generate something, save it, open it, move it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Always-on assistants.&lt;/strong&gt; Run it on a small machine that stays on, and reach it from your laptop or phone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In every one of these, the agent is taking actions on real files and a real system. Which is exactly why where the model runs starts to matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why local AI is the right default for an agent
&lt;/h2&gt;

&lt;p&gt;When a model only chats, where it runs is mostly a question of privacy. When a model drives an agent that reads your files and runs commands, it becomes a question of control. The agent acts on your machine, and if the model behind it lives in someone else's datacenter, every file it reads and every command it runs depends on a system you do not own: its uptime, its latency, its terms, and its access policy, any of which can change without you. Running the model locally keeps the agent answerable to your hardware, not to a service that can throttle it, change it, or cut it off.&lt;/p&gt;

&lt;p&gt;Running the model locally changes the property, not just the policy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your data stays on the machine.&lt;/strong&gt; The files the agent reads, the code it writes, the commands it runs: none of it leaves your disk to reach the model. There is no prompt log on a server you cannot see.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works offline.&lt;/strong&gt; Once the model is downloaded, you can disconnect entirely and the agent keeps working. Enable airplane mode and ask it a question. It still answers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No per-token bill.&lt;/strong&gt; The model runs on hardware you already own. A long agent session that would cost real money against a metered API costs nothing locally. And no risk of seeing the bill inflate over time due to rising costs for serving inference at scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No silent model swaps.&lt;/strong&gt; The weights are a file on your disk. They do not change under you between sessions, a practice that some providers run to decrease their cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade is latency and raw capability: a model that fits on a laptop is smaller than a frontier cloud model, so it is slower and less sharp. For a large class of real work, that trade is worth it, and it keeps getting better as small models improve.&lt;/p&gt;

&lt;p&gt;QVAC is the piece that makes the local half practical. It is an open-source SDK and CLI from Tether that runs models on your own device across every major GPU backend (NVIDIA, AMD, Intel, Adreno (Qualcomm) and Mali on Linux, Windows and Android through Vulkan, and Apple Silicon through Metal), and it ships an OpenAI-compatible server. That server is the bridge OpenClaw plugs into.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it fits together
&lt;/h2&gt;

&lt;p&gt;Three pieces, two of them long-running:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The QVAC server&lt;/strong&gt; holds the model in memory and answers OpenAI-compatible requests on your local port of your own machine. You start it once and leave it running.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The OpenClaw gateway&lt;/strong&gt; runs the agent loop. You point it at the QVAC server during setup, then leave it running too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your commands&lt;/strong&gt; go to the gateway. You ask a question or hand it a task, and it works against the local model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nothing in this chain calls out to the internet for inference. The model file is downloaded once, then everything happens on the machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which model to run
&lt;/h2&gt;

&lt;p&gt;The setup uses Qwen3-8B, which is the best balance for most laptops. If you have less memory, drop to the 4B. If you have a workstation and want stronger results on multi-step coding tasks, step up to the Qwen3.6-27B multimodal model. To switch, change the model name in the config file (step 2 of the setup) to the constant in the last column.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Use case&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Recommended RAM&lt;/th&gt;
&lt;th&gt;Required storage&lt;/th&gt;
&lt;th&gt;Config name&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Light and fast, mostly conversation&lt;/td&gt;
&lt;td&gt;Qwen3-4B (4-bit)&lt;/td&gt;
&lt;td&gt;8 GB&lt;/td&gt;
&lt;td&gt;~3 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_4B_INST_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recommended default, good all-round balance&lt;/td&gt;
&lt;td&gt;Qwen3-8B (4-bit)&lt;/td&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;~6 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_8B_INST_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Heavier coding, multi-step agents, and vision&lt;/td&gt;
&lt;td&gt;Qwen3.6-27B multimodal (4-bit)&lt;/td&gt;
&lt;td&gt;48 GB&lt;/td&gt;
&lt;td&gt;~18 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_6_27B_MULTIMODAL_Q4_K_XL&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A GPU helps a lot but is not required. QVAC uses Apple Metal on Apple Silicon and Vulkan on NVIDIA, AMD, and Intel. On a CPU-only machine it still runs, just slower. Run &lt;code&gt;qvac doctor&lt;/code&gt; to see what your hardware supports.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set it up
&lt;/h2&gt;

&lt;p&gt;The walkthrough below is pure copy-paste. Run each step in order, and you will have a local coding agent in a few minutes. The only thing that takes real time is the first model download. The setup uses Qwen3-8B quantized at 4 bits (about 4.7 GB), which downloads once and is cached after that.&lt;/p&gt;

&lt;p&gt;There are two paths through it, and the steps below handle both:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If you do &lt;strong&gt;not&lt;/strong&gt; have OpenClaw yet, run the optional install step.&lt;/li&gt;
&lt;li&gt;If you &lt;strong&gt;already&lt;/strong&gt; have OpenClaw, skip that one step. Everything else is the same.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Requirement:&lt;/strong&gt; Node.js 22.17 or newer (check with &lt;code&gt;node --version&lt;/code&gt;, get it from nodejs.org). About 5 GB of free disk for the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Install the QVAC CLI&lt;/strong&gt; (all platforms)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @qvac/cli @qvac/sdk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Create a model config.&lt;/strong&gt; Makes a folder and writes one small config file.&lt;/p&gt;

&lt;p&gt;macOS and Linux:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/qvac-openclaw &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ~/qvac-openclaw
&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; qvac.config.json &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
{
  "plugins": ["@qvac/sdk/llamacpp-completion/plugin"],
  "serve": {
    "models": {
      "qwen3-8B-Q4-chat": {
        "model": "QWEN3_8B_INST_Q4_K_M",
        "type": "llamacpp-completion",
        "preload": true,
        "config": { "tools": true, "toolsMode": "static", "ctx_size": 16384, "gpu_layers": -1, "reasoning_budget": 0 }
      }
    }
  }
}
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Windows (PowerShell):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;mkdir&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="bp"&gt;$HOME&lt;/span&gt;&lt;span class="nx"&gt;\qvac-openclaw&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;cd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="bp"&gt;$HOME&lt;/span&gt;&lt;span class="nx"&gt;\qvac-openclaw&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;@'
{
  "plugins": ["@qvac/sdk/llamacpp-completion/plugin"],
  "serve": {
    "models": {
      "qwen3-8B-Q4-chat": {
        "model": "QWEN3_8B_INST_Q4_K_M",
        "type": "llamacpp-completion",
        "preload": true,
        "config": { "tools": true, "toolsMode": "static", "ctx_size": 16384, "gpu_layers": -1, "reasoning_budget": 0 }
      }
    }
  }
}
'@&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Set-Content&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Encoding&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;utf8&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;qvac.config.json&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Start the QVAC server&lt;/strong&gt; (all platforms). Run it in that folder and leave the terminal open. &lt;code&gt;qwen3-8B-Q4-chat&lt;/code&gt; is the alias you defined in step 2; it serves Qwen3-8B at 4-bit (Q4_K_M). First run downloads the model once (about 4.7 GB). Ready when you see &lt;code&gt;QVAC API server listening&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;qvac serve openai &lt;span class="nt"&gt;--model&lt;/span&gt; qwen3-8B-Q4-chat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;4. Install OpenClaw (optional).&lt;/strong&gt; Skip if you already have it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; openclaw
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;5. Point OpenClaw at QVAC&lt;/strong&gt; (new terminal). The first command connects OpenClaw to the local server. The next three keep the agent fast and reliable on a local model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw onboard &lt;span class="nt"&gt;--auth-choice&lt;/span&gt; custom-api-key &lt;span class="nt"&gt;--custom-base-url&lt;/span&gt; http://127.0.0.1:11434/v1 &lt;span class="nt"&gt;--custom-model-id&lt;/span&gt; qwen3-8B-Q4-chat &lt;span class="nt"&gt;--custom-api-key&lt;/span&gt; &lt;span class="s2"&gt;"qvac"&lt;/span&gt; &lt;span class="nt"&gt;--non-interactive&lt;/span&gt; &lt;span class="nt"&gt;--accept-risk&lt;/span&gt; &lt;span class="nt"&gt;--skip-channels&lt;/span&gt; &lt;span class="nt"&gt;--skip-daemon&lt;/span&gt; &lt;span class="nt"&gt;--skip-search&lt;/span&gt; &lt;span class="nt"&gt;--skip-ui&lt;/span&gt; &lt;span class="nt"&gt;--skip-skills&lt;/span&gt; &lt;span class="nt"&gt;--skip-health&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two values here tie back to step 2. &lt;code&gt;--custom-model-id qwen3-8B-Q4-chat&lt;/code&gt; must match the model alias in your config (the key under &lt;code&gt;serve.models&lt;/code&gt;). The &lt;code&gt;--custom-api-key "qvac"&lt;/code&gt; is only a placeholder: the local QVAC server does not require a key, but OpenClaw's setup needs the field filled, so any non-empty value works.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw config &lt;span class="nb"&gt;set &lt;/span&gt;tools.profile coding
openclaw config &lt;span class="nb"&gt;set &lt;/span&gt;tools.allow &lt;span class="s1"&gt;'["write","read","exec"]'&lt;/span&gt; &lt;span class="nt"&gt;--strict-json&lt;/span&gt;
openclaw config &lt;span class="nb"&gt;set &lt;/span&gt;models.providers.custom-127-0-0-1-11434.timeoutSeconds 600
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;6. Start the agent gateway&lt;/strong&gt; (new terminal, leave open). Ready when you see &lt;code&gt;ready&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw gateway run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;7. Talk to your local agent&lt;/strong&gt; (new terminal). Try turning off your network and asking again.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw agent &lt;span class="nt"&gt;--agent&lt;/span&gt; main &lt;span class="nt"&gt;--message&lt;/span&gt; &lt;span class="s2"&gt;"Are you running in the cloud or on my machine? Answer in one sentence."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;8. Have it build something (optional).&lt;/strong&gt; Asks the agent to write a small animated web page and open it. On a local model it takes a minute or two. The open command differs per platform: &lt;code&gt;open&lt;/code&gt; on macOS, &lt;code&gt;xdg-open&lt;/code&gt; on Linux, &lt;code&gt;start&lt;/code&gt; on Windows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw agent &lt;span class="nt"&gt;--agent&lt;/span&gt; main &lt;span class="nt"&gt;--message&lt;/span&gt; &lt;span class="s2"&gt;"Create an HTML file at ~/openclaw_lobster.html showing a large lobster emoji at 140px pulsing with a CSS scale animation on a dark #0f1410 background, with the text 'openclaw running on local with QVAC' in teal #16E3C1 monospace below it, fully visible immediately. Then run the shell command: open ~/openclaw_lobster.html"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What you just built, and what to try next
&lt;/h2&gt;

&lt;p&gt;You now have a coding agent running entirely on your own machine. A good first test is to confirm the obvious: ask it whether it is running in the cloud or on your machine, then turn off your network and ask again. It keeps answering, because nothing was ever being sent to a server in the first place.&lt;/p&gt;

&lt;p&gt;For something you can see, hand it a small build task and watch it write a file and open it. A coding task like generating a small web page runs in a minute or two on a typical laptop, fully offline, because the model is doing real work locally rather than streaming from a datacenter. The setup includes a ready-made example: it asks the agent to write a tiny animated web page and open it in your browser, with no further input from you.&lt;/p&gt;

&lt;p&gt;From here, the same setup handles the rest of what OpenClaw does. Point it at a project folder, ask it to read and edit code, have it summarize a directory, or give it a multi-step chore. The agent loop is identical. The only thing that changed is that the intelligence behind it is yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes and limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A model that fits on a laptop is not a frontier cloud model. Expect it to be slower and to need clearer instructions. For agent work, give it focused, single-goal tasks rather than long open-ended ones.&lt;/li&gt;
&lt;li&gt;Depending on your hardware, a local response can take longer than a cloud API today. Treat that as a temporary gap, not a permanent cost: local hardware keeps getting faster and capable models keep getting smaller while holding their quality, so the gap is closing fast. The physics also runs in local's favor. A request to a datacenter is bounded by the speed of light, a round trip no provider can engineer away. For anything that has to react in real time, a robot, a control loop, a live interface, that round trip is a hard floor, and local is the only option that stays reliable.&lt;/li&gt;
&lt;li&gt;QVAC is Apache 2.0 and free. The model you run is downloaded once and cached locally. Source and docs: &lt;a href="https://github.com/tetherto/qvac" rel="noopener noreferrer"&gt;github.com/tetherto/qvac&lt;/a&gt; and &lt;a href="https://docs.qvac.tether.io" rel="noopener noreferrer"&gt;docs.qvac.tether.io&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>openclaw</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>How to run Hermes, a self-improving personal AI agent, fully local with QVAC</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Fri, 24 Jul 2026 14:49:18 +0000</pubDate>
      <link>https://dev.to/qvac/how-to-run-hermes-a-self-improving-personal-ai-agent-fully-local-with-qvac-463d</link>
      <guid>https://dev.to/qvac/how-to-run-hermes-a-self-improving-personal-ai-agent-fully-local-with-qvac-463d</guid>
      <description>&lt;h2&gt;
  
  
  What Hermes is, and why it is different
&lt;/h2&gt;

&lt;p&gt;Most AI agents you have seen are task tools. You give them a job, they do it, they forget you. Hermes Agent, from Nous Research, is built on a different idea: an agent that is yours, that remembers you, and that gets better the longer you use it.&lt;/p&gt;

&lt;p&gt;Three things make it stand out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It has a persona.&lt;/strong&gt; A single file, &lt;code&gt;SOUL.md&lt;/code&gt;, defines how it talks. Edit it and the agent changes character on the next message.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It remembers.&lt;/strong&gt; Conversations, facts, and context persist across sessions in a local memory, so it is a companion, not a one-shot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It learns.&lt;/strong&gt; Hermes has a library of skills (research, notes, GitHub, documents, smart home, and many more) and a background curator that creates and refines skills from experience. The agent that finishes your week is not the one that started it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of that personality, memory, and learned skill is data about you. Which raises the obvious question: where does the model behind it run?&lt;/p&gt;

&lt;p&gt;&lt;iframe class="tweet-embed" id="tweet-2077391578310791482-199" src="https://platform.twitter.com/embed/Tweet.html?id=2077391578310791482"&gt;
&lt;/iframe&gt;

  // Detect dark theme
  var iframe = document.getElementById('tweet-2077391578310791482-199');
  if (document.body.className.includes('dark-theme')) {
    iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2077391578310791482&amp;amp;theme=dark"
  }



&lt;/p&gt;

&lt;h2&gt;
  
  
  Why local is the right default for a personal agent
&lt;/h2&gt;

&lt;p&gt;When a model only answers questions, where it runs is mostly about privacy. When a model is your personal agent, holding your persona, your memory, and skills it learned from your own habits, it becomes about ownership. If the brain lives in someone else's datacenter, then your assistant's memory of you, and everything it learns from you, flows through a system you do not control: its uptime, its terms, its access policy, all of which can change without you.&lt;/p&gt;

&lt;p&gt;Running the model locally changes the property, not just the policy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your context stays on the machine.&lt;/strong&gt; Your persona, your memory, the skills the agent builds: none of it has to leave your disk to reach the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works offline.&lt;/strong&gt; Once the model is downloaded, you can disconnect entirely and keep talking to your agent. Turn on airplane mode and it still answers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No per-token bill.&lt;/strong&gt; A personal agent you talk to all day, that runs scheduled jobs in the background, would meter up fast against a cloud API. Locally it costs nothing per token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No silent model swaps.&lt;/strong&gt; The weights are a file on your disk. Your agent's "brain" does not change under you between sessions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trade is latency and raw capability: a model that fits on a laptop is smaller than a frontier cloud model. For a personal agent that lives with your data, that trade is easy, and it keeps getting better as small models improve.&lt;/p&gt;

&lt;p&gt;QVAC is the piece that makes the local half practical. It is an open-source SDK and CLI from Tether that runs models on your own device across every major GPU backend (NVIDIA, AMD, and Intel through Vulkan, Apple Silicon through Metal), and it ships an OpenAI-compatible server. That server is the bridge Hermes plugs into.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it fits together
&lt;/h2&gt;

&lt;p&gt;Two long-running pieces and you:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The QVAC server&lt;/strong&gt; holds the model in memory and answers OpenAI-compatible requests on a local port. Start it once and leave it running.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hermes&lt;/strong&gt; runs the agent: persona, memory, skills, tools. You point it at the QVAC server once, then just talk to it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your messages&lt;/strong&gt; go to Hermes, which thinks using the local model and acts with its tools.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The model inference never leaves the machine. (Some skills, like web research, still use the internet for that specific tool, the same way a browser does. The brain stays on-device.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Which model to run
&lt;/h2&gt;

&lt;p&gt;Hermes is tool-heavy: it sends a large set of tools on every turn, so it pays to run a capable model. The setup below uses Qwen3.6-35B-A3B, a mixture-of-experts model: 35 billion parameters of knowledge, only about 3 billion active per token, so it stays fast for its size. To switch models, change the constant in the config file (step 2 of the setup).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Use case&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Recommended RAM&lt;/th&gt;
&lt;th&gt;Required storage&lt;/th&gt;
&lt;th&gt;Config name&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Lightest, runs almost anywhere&lt;/td&gt;
&lt;td&gt;Qwen3.5-4B (4-bit)&lt;/td&gt;
&lt;td&gt;8 GB&lt;/td&gt;
&lt;td&gt;~3 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_5_4B_MULTIMODAL_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Balanced&lt;/td&gt;
&lt;td&gt;Qwen3.5-9B (4-bit)&lt;/td&gt;
&lt;td&gt;16 GB&lt;/td&gt;
&lt;td&gt;~6 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_5_9B_MULTIMODAL_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recommended for the full agent&lt;/td&gt;
&lt;td&gt;Qwen3.6-35B-A3B MoE (4-bit)&lt;/td&gt;
&lt;td&gt;36 GB&lt;/td&gt;
&lt;td&gt;~22 GB&lt;/td&gt;
&lt;td&gt;&lt;code&gt;QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Hermes leans on tool calls, and smaller models are weaker at that, so on the 4B or 9B keep the active skill set lean (see "What to try next"). The 35B-A3B MoE handles the full toolset most reliably.&lt;/p&gt;

&lt;p&gt;Be realistic: these are small local models, and tool-heavy agent work is where they struggle most. The 4B is for trying the setup, the 9B mainly adds reliability over the 4B rather than smarts, and the 35B-A3B MoE is the only one that drives the full toolset dependably, so it is what we recommend. It is a capable on-device assistant, not a frontier cloud model.&lt;/p&gt;

&lt;p&gt;A GPU helps a lot but is not required. QVAC uses Apple Metal on Apple Silicon and Vulkan on NVIDIA, AMD, and Intel. Run &lt;code&gt;qvac doctor&lt;/code&gt; to see what your hardware supports. On the 35B-A3B, expect the first turn to take around 20 seconds (Hermes sends a large system prompt plus its tools), and faster turns after. For a snappier agent, narrow the active skills (see "What to try next").&lt;/p&gt;

&lt;h2&gt;
  
  
  Set it up
&lt;/h2&gt;

&lt;p&gt;The walkthrough is copy-paste. Run each step in order. The only slow part is the first model download, which is cached after that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Requirement:&lt;/strong&gt; Node.js 22.17 or newer (&lt;code&gt;node --version&lt;/code&gt;, from nodejs.org). Disk and RAM per the model table above.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Install the QVAC CLI&lt;/strong&gt; (all platforms)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @qvac/cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Create a model config.&lt;/strong&gt; Makes a folder and writes one small config file.&lt;/p&gt;

&lt;p&gt;macOS and Linux:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/qvac-hermes &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; ~/qvac-hermes
&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; qvac.config.json &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
{
  "plugins": ["@qvac/sdk/llamacpp-completion/plugin"],
  "serve": {
    "models": {
      "qwen3.6-moe": {
        "model": "QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M",
        "type": "llamacpp-completion",
        "default": true,
        "preload": true,
        "config": { "tools": true, "toolsMode": "static", "ctx_size": 32768, "gpu_layers": -1, "reasoning_budget": 0 }
      }
    }
  }
}
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Windows (PowerShell):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;mkdir&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="bp"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;\qvac-hermes"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Set-Location&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="bp"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;\qvac-hermes"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="sh"&gt;@'
{
  "plugins": ["@qvac/sdk/llamacpp-completion/plugin"],
  "serve": {
    "models": {
      "qwen3.6-moe": {
        "model": "QWEN3_6_35B_A3B_MULTIMODAL_Q4_K_M",
        "type": "llamacpp-completion",
        "default": true,
        "preload": true,
        "config": { "tools": true, "toolsMode": "static", "ctx_size": 32768, "gpu_layers": -1, "reasoning_budget": 0 }
      }
    }
  }
}
'@&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Out-File&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Encoding&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;utf8&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;qvac.config.json&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two settings that matter: &lt;code&gt;"tools": true&lt;/code&gt; and &lt;code&gt;"toolsMode": "static"&lt;/code&gt;. Static is what lets an external client like Hermes run the model's tool calls itself. (With "dynamic" the tool calls never execute, which looks like the model is broken.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Start the QVAC server&lt;/strong&gt; (all platforms). Run it in that folder and leave the terminal open. First run downloads the model once (about 22 GB for the MoE). Ready when you see &lt;code&gt;QVAC API server listening&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;qvac serve openai &lt;span class="nt"&gt;--model&lt;/span&gt; qwen3.6-moe
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It listens on &lt;code&gt;http://localhost:11434/v1&lt;/code&gt;. That is the whole command: &lt;code&gt;--model qwen3.6-moe&lt;/code&gt; is the alias you defined in step 2, and running from the &lt;code&gt;~/qvac-hermes&lt;/code&gt; folder lets QVAC auto-detect that &lt;code&gt;qvac.config.json&lt;/code&gt; (which is also what sets &lt;code&gt;toolsMode: static&lt;/code&gt;). Run it somewhere without that config and it will error that the alias is unknown.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Install Hermes&lt;/strong&gt; (new terminal). The installer sets up Python, the global &lt;code&gt;hermes&lt;/code&gt; command, and its dependencies.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://hermes-agent.nousresearch.com/install.sh | bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Windows, run that installer inside WSL (Windows Subsystem for Linux). The QVAC server from step 3 runs natively on Windows, and Hermes in WSL reaches it over &lt;code&gt;localhost&lt;/code&gt; either way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Point Hermes at QVAC.&lt;/strong&gt; Tell Hermes to use your local server as a custom OpenAI-compatible provider.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.provider custom
hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.base_url &lt;span class="s2"&gt;"http://localhost:11434/v1"&lt;/span&gt;
hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.default qwen3.6-moe
hermes config &lt;span class="nb"&gt;set &lt;/span&gt;model.api_key qvac
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;qwen3.6-moe&lt;/code&gt; must match the alias in your config (the key under &lt;code&gt;serve.models&lt;/code&gt;). The api-key value is a placeholder: the local QVAC server does not check it, but Hermes wants the field filled, so any non-empty value works. (You can also run the interactive &lt;code&gt;hermes model&lt;/code&gt; wizard and pick "custom" instead of these four commands.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Talk to your local agent.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That opens an interactive chat. Try turning off your network and asking it whether it runs in the cloud or on your machine. It keeps answering.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you just built, and what to try next
&lt;/h2&gt;

&lt;p&gt;You now have a personal AI agent running entirely on your own machine. A good first test is the obvious one: ask it whether it is in the cloud or local, then enable airplane mode and ask again. It keeps answering, because the inference was never leaving your device.&lt;/p&gt;

&lt;p&gt;From there, the things that make Hermes Hermes all run on the same local model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Give it a personality.&lt;/strong&gt; Edit &lt;code&gt;~/.hermes/SOUL.md&lt;/code&gt;, then send a message. The persona is loaded fresh each time, so the change is immediate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let it remember.&lt;/strong&gt; Tell it something about how you work, start a new session later, and it recalls it. The memory lives in &lt;code&gt;~/.hermes/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use a skill.&lt;/strong&gt; Run &lt;code&gt;hermes skills&lt;/code&gt; to see the library, then ask it for something real: take structured notes, organize a project on its task board, summarize a folder of documents. Pick local skills (notes, files, board) to stay fully offline; web skills will use the internet for that tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let it improve.&lt;/strong&gt; Over time the curator refines skills from what you do. That learning happens on your hardware and stays there.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For speed, narrow the active skills to what you actually use (&lt;code&gt;hermes skills&lt;/code&gt;, &lt;code&gt;hermes tools&lt;/code&gt;); a smaller toolset means a smaller prompt and faster turns on a local model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes and limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A model that fits on a laptop is not a frontier cloud model. Expect it to need clearer instructions and to be slower than a cloud API today. That gap is closing as local hardware speeds up and small models improve.&lt;/li&gt;
&lt;li&gt;"Local" here means the model runs on your device. A given skill may still call the internet for its own job (web research, for example), the same as any app would. The intelligence, your persona, and your memory stay on the machine.&lt;/li&gt;
&lt;li&gt;QVAC is Apache 2.0 and free. Hermes Agent is from Nous Research (github.com/NousResearch/hermes-agent). The model is downloaded once and cached locally. QVAC source and docs: &lt;a href="https://github.com/tetherto/qvac" rel="noopener noreferrer"&gt;github.com/tetherto/qvac&lt;/a&gt; and &lt;a href="https://docs.qvac.tether.io" rel="noopener noreferrer"&gt;docs.qvac.tether.io&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>hermes</category>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Meet MedPsy: a private medical AI model small enough for your phone</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Fri, 24 Jul 2026 07:05:58 +0000</pubDate>
      <link>https://dev.to/qvac/meet-medpsy-a-private-medical-ai-model-small-enough-for-your-phone-40lo</link>
      <guid>https://dev.to/qvac/meet-medpsy-a-private-medical-ai-model-small-enough-for-your-phone-40lo</guid>
      <description>&lt;p&gt;QVAC MedPsy is a free, open-source medical AI model that can run entirely on your own device, with no cloud and no data leaving your phone. It ships in two sizes, 1.7B and 4B. On standard medical benchmarks, the 4B version matches or beats Google's MedGemma-27B, a model nearly seven times larger, while the 1.7B version is small enough to run on an ordinary smartphone.&lt;/p&gt;

&lt;p&gt;The headline is simple: &lt;strong&gt;better training beat raw size&lt;/strong&gt;. MedPsy reaches the level of models two to seven times bigger, not by being huge, but by being trained more carefully on higher-quality medical data. And because you can run locally, the most sensitive data there is, your health, never has to travel to someone else's server.&lt;/p&gt;

&lt;h2&gt;
  
  
  A private medical AI model built to run on your device
&lt;/h2&gt;

&lt;p&gt;MedPsy is a family of small medical language models developed by Tether Data's AI Research group as part of the QVAC platform, Tether's open-source, local-AI ecosystem. It answers everyday medical questions and walks through clinical reasoning in clear, plain text. There are two sizes: &lt;strong&gt;MedPsy-1.7B&lt;/strong&gt; for phones, and &lt;strong&gt;MedPsy-4B&lt;/strong&gt; for high-end phones and laptops. Both are built on open Qwen3 models and released under the Apache 2.0 license for research and educational use. Both models are available on Hugging Face, in their full version and as smaller compressed versions that run efficiently with llama.cpp or the QVAC SDK. The QVAC SDK provides out-of-the-box optimal execution support, offering a simple library that runs locally on any device, from a phone to a laptop or server with seamless cross-platform behavior, delegated inference and peer-to-peer (P2P) capabilities.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclaimer: MedPsy is a general medical model for question-answering and clinical reasoning. It is not a mental health, therapy, or crisis-support tool, and has not been trained or validated for that.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Better training, not raw size
&lt;/h2&gt;

&lt;p&gt;In AI, how you train a model matters as much as how big it is. The QVAC team focused on the quality of the medical training data and a careful, multi-stage training process aimed squarely at medicine, instead of simply scaling up parameters. A model that specializes deeply in one field can outperform a much larger model that tries to know everything.&lt;/p&gt;

&lt;p&gt;The training data was built using &lt;a href="https://qvac.tether.io/dev/genesis/" rel="noopener noreferrer"&gt;QVAC Genesis II&lt;/a&gt; synthetic medical datasets, which cover diverse medical domains such as biology and chemistry along with publicly available open-source medical QA prompts. The datasets were used as seeds by a strong teacher model (Baichuan-M3-235B) to generate chain-of-thought traces, extended rationales, and decision-oriented answers to improve the model's quality. For more information about the training methodology, see &lt;a href="https://huggingface.co/blog/qvac/medpsy#2-data-methodology" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;There is a practical bonus. MedPsy-4B reaches its answers using &lt;strong&gt;about 3.2 times fewer tokens&lt;/strong&gt; than the general model it is based on. Fewer tokens means faster replies and less battery used, which is exactly what you want on a phone.&lt;/p&gt;

&lt;h2&gt;
  
  
  How MedPsy compares to Google's MedGemma
&lt;/h2&gt;

&lt;p&gt;On the average across closed-ended medical benchmarks, &lt;strong&gt;MedPsy-4B scores 70.54 and edges out MedGemma-27B at 69.95&lt;/strong&gt;, despite being nearly seven times smaller. It leads most clearly on the hardest, most realistic tests: HealthBench, HealthBench Hard and the expert-level MedXpertQA. On a few academic multiple-choice tests it runs a point or two behind. Here is the honest side-by-side.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Medical benchmark&lt;/th&gt;
&lt;th&gt;MedPsy-4B&lt;/th&gt;
&lt;th&gt;MedGemma-27B&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Average (closed-ended)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;70.54&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;69.95&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HealthBench (realistic care)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;74.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;65.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HealthBench Hard&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;58.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;42.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MedXpertQA (expert level)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;30.61&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;25.18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MedQA-USMLE&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84.39&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;83.29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PubMedQA&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;75.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.93&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MedMCQA&lt;/td&gt;
&lt;td&gt;72.15&lt;/td&gt;
&lt;td&gt;72.77&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AfriMedQA&lt;/td&gt;
&lt;td&gt;71.50&lt;/td&gt;
&lt;td&gt;73.07&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU Health&lt;/td&gt;
&lt;td&gt;89.70&lt;/td&gt;
&lt;td&gt;90.48&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MMLU-Pro Health&lt;/td&gt;
&lt;td&gt;70.45&lt;/td&gt;
&lt;td&gt;72.94&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The smaller model tells the same story. &lt;strong&gt;MedPsy-1.7B scores 62.62 on average and beats Google's MedGemma-1.5-4B (51.20) by more than 11 points&lt;/strong&gt;, while being less than half its size and small enough to run on a normal smartphone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why running on your device matters
&lt;/h2&gt;

&lt;p&gt;Health data is the most sensitive data most of us will ever generate. The usual way medical AI works is to send your question to a company's servers in the cloud. MedPsy flips that: it runs &lt;strong&gt;100% on your device&lt;/strong&gt;, so your symptoms, questions and notes stay with you. There is no account, no subscription, and nothing to leak.&lt;/p&gt;

&lt;p&gt;Running locally also makes it work where the cloud cannot. It keeps working with no internet, on a plane, in a remote clinic, or anywhere bandwidth is scarce. And because it fits on a phone people already own, it does not need an expensive server or a data center to be useful, which matters most in the places that have the least access to care.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try MedPsy
&lt;/h2&gt;

&lt;p&gt;The simplest path is the QVAC SDK. Install it, &lt;a href="https://huggingface.co/collections/qvac/medpsy" rel="noopener noreferrer"&gt;pull a MedPsy model from Hugging Face&lt;/a&gt;, and run it on your device:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; @qvac/sdk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No API key, no cloud bill. Get the models for free and get started.&lt;/p&gt;

&lt;h2&gt;
  
  
  MedPsy on Hugging Face
&lt;/h2&gt;

&lt;p&gt;All MedPsy models, GGUF files, quantized variants, and resources in one place: &lt;a href="https://huggingface.co/collections/qvac/medpsy" rel="noopener noreferrer"&gt;open the collection&lt;/a&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MedPsy-4B&lt;/strong&gt;: higher-quality edge model. Surpasses MedGemma-27B-text-it on closed-ended medical benchmarks at about 7x smaller. &lt;a href="https://huggingface.co/qvac/MedPsy-4B" rel="noopener noreferrer"&gt;Model card&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MedPsy-1.7B&lt;/strong&gt;: smartphone-class medical model. Beats MedGemma-1.5-4B-it by +11.42 points on closed-ended; matches Qwen3-4B-Thinking-2507. &lt;a href="https://huggingface.co/qvac/MedPsy-1.7B" rel="noopener noreferrer"&gt;Model card&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MedPsy-4B-GGUF&lt;/strong&gt;: GGUF repo with an unquantized BF16 export and seven quantized files. Q5_K_M (3.16 GB) adds a high-quality 5-bit tier; Q4_K_M (2.72 GB) remains the recommended size/quality trade-off. &lt;a href="https://huggingface.co/qvac/MedPsy-4B-GGUF" rel="noopener noreferrer"&gt;GGUF repo&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MedPsy-1.7B-GGUF&lt;/strong&gt;: smartphone-ready GGUF repo with an unquantized BF16 export and seven quantized files. Q5_K_M (1.47 GB) is nearly lossless; Q4_K_M (1.28 GB) is the best mobile trade-off. &lt;a href="https://huggingface.co/qvac/MedPsy-1.7B-GGUF" rel="noopener noreferrer"&gt;GGUF repo&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;Important:&lt;/strong&gt; MedPsy is not a substitute for professional medical judgment, diagnosis, or treatment, and must never be used in emergencies or as a sole decision-making tool. Like any language model, it can produce wrong or incomplete answers. It is text-only, English-only, and released for research and educational purposes. Always consult a qualified healthcare professional.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently asked questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is QVAC MedPsy free?
&lt;/h3&gt;

&lt;p&gt;Yes. MedPsy is open source under the Apache 2.0 license and free to download from Hugging Face, released for research and educational purposes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does my data leave my device when I use MedPsy?
&lt;/h3&gt;

&lt;p&gt;No. MedPsy is designed to run entirely on your own device through the standard llama.cpp inference engine and the QVAC SDK. There is no cloud call, so your prompts and any health information you type never leave the device.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is MedPsy a mental health or therapy AI?
&lt;/h3&gt;

&lt;p&gt;No. MedPsy is built for general medical question-answering and clinical reasoning. It is not trained or validated for mental health, therapy, or crisis support.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can MedPsy replace a doctor?
&lt;/h3&gt;

&lt;p&gt;No. MedPsy is not a substitute for professional medical judgment, diagnosis, or treatment, and it should never be used for emergencies. Its answers can contain mistakes and should always be checked with a qualified clinician.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does MedPsy compare to Google MedGemma?
&lt;/h3&gt;

&lt;p&gt;On the closed-ended medical benchmark average, MedPsy-4B scores 70.54 versus 69.95 for MedGemma-27B, despite being nearly seven times smaller. It leads clearly on the hardest realistic health tests, such as HealthBench (74 vs 65) and HealthBench Hard (58 vs 42).&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between MedPsy-1.7B and MedPsy-4B?
&lt;/h3&gt;

&lt;p&gt;MedPsy-1.7B is small enough to run on an ordinary smartphone and beats Google's MedGemma-1.5-4B by over 11 points. MedPsy-4B is more capable, targets high-end phones and laptops, and matches models nearly seven times its size.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/blog/qvac/medpsy" rel="noopener noreferrer"&gt;QVAC MedPsy launch post (Hugging Face blog)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/qvac/MedPsy-4B" rel="noopener noreferrer"&gt;MedPsy-4B model card&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/qvac/MedPsy-1.7B" rel="noopener noreferrer"&gt;MedPsy-1.7B model card&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/qvac/MedPsy-4B-GGUF" rel="noopener noreferrer"&gt;MedPsy-4B GGUF&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/qvac/MedPsy-1.7B-GGUF" rel="noopener noreferrer"&gt;MedPsy-1.7B GGUF&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.qvac.tether.io/" rel="noopener noreferrer"&gt;QVAC SDK documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://qvac.tether.io/dev/genesis/" rel="noopener noreferrer"&gt;QVAC Genesis dataset&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Nothing is more private than a thought: how QVAC runs a brain-to-text model fully on-device</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Fri, 24 Jul 2026 06:48:02 +0000</pubDate>
      <link>https://dev.to/qvac/nothing-is-more-private-than-a-thought-how-qvac-runs-a-brain-to-text-model-fully-on-device-1g15</link>
      <guid>https://dev.to/qvac/nothing-is-more-private-than-a-thought-how-qvac-runs-a-brain-to-text-model-fully-on-device-1g15</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnq3jvfsk08k2kudhkuze.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnq3jvfsk08k2kudhkuze.png" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;Nothing is more private than a thought. So when a model can decode speech directly from the brain, where it runs matters as much as how well it works. QVAC, Tether's open-source on-device AI stack, now runs exactly that kind of model fully on the device, with nothing sent to a cloud. The model is BrainWhisperer, built by Tether Evo's research team and now available in the QVAC SDK as a proof-of-concept capability. The version in the QVAC SDK is a small, end-to-end model: on real neural recordings it decodes more than 90% of words correctly for a single participant, an 8.7% word error rate, while running in under 2 GB of memory. A larger, more complex version of the same research placed 4th of 466 teams in an international challenge. It is early, but it puts a brain-to-text engine in developers' hands, locally, as a foundation for the products that could one day give people their voice back.&lt;/p&gt;

&lt;h2&gt;
  
  
  QVAC and the model
&lt;/h2&gt;

&lt;p&gt;QVAC is Tether's open-source stack for running AI on your own device, with no cloud. Most of what it runs is familiar: language models, speech, vision, image generation. Its newest and most experimental addition is a brain-computer interface: it can take neural signals and decode them into text, on-device, through a single SDK call.&lt;/p&gt;

&lt;p&gt;The model behind that capability is BrainWhisperer, built by Tether Evo's research team. It is integrated into the QVAC SDK at proof-of-concept level, and that is the important framing. It is not a product. It is a working engine that developers can run locally today, and a foundation for the assistive products that could be built on top of it tomorrow. The rest of this article is what it does, how well it does it, and why running it on-device is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hardest kind of silence
&lt;/h2&gt;

&lt;p&gt;Imagine knowing exactly what you want to say and not being able to say it.&lt;/p&gt;

&lt;p&gt;For people living with ALS and similar conditions, that is daily life. As the disease progresses, it takes control of the muscles: the arms, the hands, and eventually the face and the voice. Some can still try to form sounds, but the words come out unintelligible. Others cannot produce sound at all.&lt;/p&gt;

&lt;p&gt;What does not change is the mind. The thoughts are still there. The language is still there. The person still composes full sentences, still has things to tell the people they love. The connection between the brain and the world is what breaks, not the brain itself. The result is a particular kind of isolation: present, aware, and unable to be heard.&lt;/p&gt;

&lt;p&gt;For decades, the workarounds have been slow. Communication by tracking eye movements, letter by letter, is exhausting and error-prone. A full sentence can take minutes. That is the problem Tether Evo set out to change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the words from the source
&lt;/h2&gt;

&lt;p&gt;Traditionally, people who cannot speak communicate through eye-tracking, choosing letters one at a time, which is slow and error-prone. BrainWhisperer goes straight to the source and reads the brain directly.&lt;/p&gt;

&lt;p&gt;When a person tries to speak, a specific part of the brain, the speech motor cortex, lights up with electrical activity, the same way it would if the muscles were responding. A small implant placed there picks up that activity. BrainWhisperer takes those neural signals and translates them into text. You do not have to move your face. You think the words, and the model works as a translation layer, turning the language of the brain into English.&lt;/p&gt;

&lt;p&gt;The clever part is where the model comes from. BrainWhisperer is a redesign of Whisper, the same family of speech-recognition AI that powers modern transcription, trained on hundreds of thousands of hours of human speech. Tether Evo rebuilt it to take neural data as input instead of audio. So instead of learning to turn sound into text from scratch, with the tiny amount of brain data that exists, it inherits everything a large speech model already knows about how language works, and applies it to the brain.&lt;/p&gt;

&lt;p&gt;Once the words are decoded, they can go anywhere text can go: read aloud in a synthetic voice, or handed to an AI assistant to act on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the results actually show
&lt;/h2&gt;

&lt;p&gt;This is the part that matters, because it is measured, not promised.&lt;/p&gt;

&lt;p&gt;Tether Evo tested BrainWhisperer on real recordings from people implanted with a 256-channel array in the speech motor cortex, using a public research dataset, and then entered it into the Brain-to-Text '25 challenge.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It decoded a large majority of words correctly. On many sentences the transcription is exact, word for word: "i just bought a new house", "do you believe in the dallas cowboys", "that was the point", all decoded perfectly.&lt;/li&gt;
&lt;li&gt;Measured as word error rate, the standard accuracy metric, the version in the SDK reaches 8.7% for a single participant in an open vocabulary setting. That is well below the 10 percent mark researchers treat as the threshold for real-world usefulness, and it comes from a fully end-to-end model.&lt;/li&gt;
&lt;li&gt;A separate, more complex version of the model placed 4th out of 466 teams in the Brain-to-Text '25 challenge. It scores higher, but it is not end-to-end, so the SDK ships the smaller end-to-end model, which is the direction we want to build in.&lt;/li&gt;
&lt;li&gt;And it does this efficiently. The fast version runs in under 2 GB of memory with roughly 50 milliseconds of delay, compared to the hundreds of gigabytes of memory older approaches needed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point is not a detail. It is the difference between a system that lives in a data center and one that can run on the device in front of the person.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why local matters here more than anywhere
&lt;/h2&gt;

&lt;p&gt;A thought is the most private thing a person has.&lt;/p&gt;

&lt;p&gt;A speech decoder reads intention straight from the brain. If that has to travel to a server to be processed, you are sending the most intimate data a person has across a network. Because BrainWhisperer is light enough to run on a local device, the decoding can happen where the person is, with nothing leaving their hands. That is the role of QVAC, Tether's open-source on-device AI stack: it is what lets a model like this run anywhere, privately, with no cloud in the loop.&lt;/p&gt;

&lt;p&gt;The control stays with the person in another way too. The system only decodes speech the user is actively trying to produce. A person can simply choose not to, or think about something else, and there is nothing to read. Further ahead, ideas like "mental passwords" could ensure these systems always answer to their user and no one else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this stands today, honestly
&lt;/h2&gt;

&lt;p&gt;This is research, and it is important to be clear about that.&lt;/p&gt;

&lt;p&gt;BrainWhisperer is a proof of concept, now available in the QVAC SDK as an experimental brain-computer-interface capability. It shows two things: that intended speech can be decoded from the brain accurately, and that it can run on ordinary local hardware thanks to QVAC. What it is not, yet, is a product you can use. What it is, is a working engine a developer can run locally today and build toward those products with.&lt;/p&gt;

&lt;p&gt;The honest limits: the kind of brain implant this needs is rare, and getting one involves real medical, safety, and regulatory hurdles. The models are trained on very few people, because very few people have these implants and only pre-recorded research data was available. The main model was trained on 6 participants; the smaller version in the SDK on a single one. It generalizes across the participants it has seen, which is itself a milestone, but a brand-new person would still need an implant and a short calibration period to record their own data before the model could adapt to them. There is real work between here and a usable system.&lt;/p&gt;

&lt;p&gt;We would rather say that plainly than oversell it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The frontier ahead
&lt;/h2&gt;

&lt;p&gt;Today, the target is speech. But speech is one signal among many that the brain produces.&lt;/p&gt;

&lt;p&gt;The same idea, reading the brain's own activity and translating it into something we can share, points at a much wider horizon. The team sees a path toward decoding not just words, but the images a person pictures in their mind, the sounds they imagine, one day perhaps intended movement. We are not putting a date on any of it. It is the frontier we are walking toward, one careful result at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  A team at the edge of what is possible
&lt;/h2&gt;

&lt;p&gt;BrainWhisperer comes out of Tether Evo, Tether's research team working where technology, fundamental science, and medicine meet. Bringing their work into the QVAC SDK is what turns a research result into something a developer can run on a device today, and build the next assistive product on tomorrow. The thread is simple: use technology to make a life better, starting with the people who need it most.&lt;/p&gt;

&lt;p&gt;Read the research: &lt;a href="https://arxiv.org/abs/2603.13321" rel="noopener noreferrer"&gt;BrainWhisperer, arXiv 2603.13321&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is BrainWhisperer?
&lt;/h3&gt;

&lt;p&gt;A research model that decodes intended speech directly from brain activity and turns it into text. It was built by Tether Evo on top of a large speech-recognition model adapted to read neural signals, and is now available in the QVAC SDK as a proof-of-concept capability.&lt;/p&gt;

&lt;h3&gt;
  
  
  Who is it for?
&lt;/h3&gt;

&lt;p&gt;People who have lost the ability to speak, for example through ALS or severe paralysis, but whose ability to think and form language is intact.&lt;/p&gt;

&lt;h3&gt;
  
  
  How accurate is it?
&lt;/h3&gt;

&lt;p&gt;On real neural recordings, the model in the QVAC SDK reaches an 8.7% word error rate for a single participant in an open vocabulary setting, so more than 90% of words are decoded correctly. A separate, more complex version placed 4th of 466 teams in the Brain-to-Text '25 challenge; it scores higher but is not end-to-end, which is why the SDK ships the smaller end-to-end model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is my brain data sent to the cloud?
&lt;/h3&gt;

&lt;p&gt;No. The model is small enough to run on a local device, so decoding can happen without sending neural data anywhere, and the system only reads speech the user is actively trying to produce.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I use it today?
&lt;/h3&gt;

&lt;p&gt;As a developer, yes. The brain-to-text capability ships in the QVAC SDK as a proof of concept you can run and experiment with locally today. What it is not yet is a finished product for end users: real-world use by a patient still depends on a specialized brain implant, which carries significant medical and regulatory hurdles, plus a short per-person calibration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;BrainWhisperer reads intended speech from the brain and turns it into text, for people who cannot speak.&lt;/li&gt;
&lt;li&gt;It decodes more than 90% of words correctly for a single participant, an 8.7% word error rate. A separate, more complex version ranked 4th of 466 teams in an international challenge.&lt;/li&gt;
&lt;li&gt;It runs in under 2 GB on a local device, so the most private data a person has never has to leave their hands.&lt;/li&gt;
&lt;li&gt;It is research, not a product. The frontier ahead includes imagined images, sounds, and movement, with no fixed timeline.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;BrainWhisperer is a research model from Tether Evo, now a proof-of-concept capability in the QVAC SDK. QVAC is Tether's open-source on-device AI stack. Apache 2.0.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Quick-start: build your first local AI app with your coding agent</title>
      <dc:creator>Thomas</dc:creator>
      <pubDate>Wed, 22 Jul 2026 14:20:56 +0000</pubDate>
      <link>https://dev.to/qvac/quick-start-build-your-first-local-ai-app-with-your-coding-agent-1ff5</link>
      <guid>https://dev.to/qvac/quick-start-build-your-first-local-ai-app-with-your-coding-agent-1ff5</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;You do not need to learn an SDK by heart to build a local AI app. You need to understand what local AI is good and bad at, know what is possible, and hand your coding agent the right context so it writes the code for you. This is that orientation: the trade-offs of running AI on your own hardware, the 12 things the QVAC SDK can do, a set of app ideas, and how to point your agent so it builds correctly. Almost no code here, on purpose.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Building this way? &lt;strong&gt;QVAC.md&lt;/strong&gt; is the one file that makes your agent write QVAC code that runs the first time. &lt;a href="https://qvac.tether.io/blog/quick-start-build-your-first-local-ai-app-with-your-coding-agent/" rel="noopener noreferrer"&gt;Get it here.&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  How this guide works
&lt;/h2&gt;

&lt;p&gt;This is not a copy-paste tutorial. You will not be pasting large code blocks. So this guide gives you the three things a builder actually needs:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The trade-offs of local AI, so you pick the right project for it.&lt;/li&gt;
&lt;li&gt;What the SDK can do, so you know what is on the menu.&lt;/li&gt;
&lt;li&gt;How to set up your agent, so the code it writes actually works.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What "local AI" actually means
&lt;/h2&gt;

&lt;p&gt;Local AI means the model runs on your hardware (or your user's device), not behind a cloud API. That single fact creates real constraints and real advantages. Knowing both is how you choose good projects.&lt;/p&gt;

&lt;h3&gt;
  
  
  The constraints (know these going in)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your hardware sets the ceiling.&lt;/strong&gt; RAM and GPU decide which models you can run and how fast. A small model runs fine on a laptop CPU. Bigger and faster wants a GPU and more memory. As a rough guide, a 4B model needs a few GB; a 30B-class mixture-of-experts model needs around 20 GB or more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A local model is not a frontier model.&lt;/strong&gt; The largest cloud models are far bigger than anything that fits on a device. On the hardest reasoning, the broadest world-knowledge, and very long context, a local model is weaker. Use it for the jobs it is good at, not for everything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The first run downloads the model&lt;/strong&gt; (hundreds of MB to several GB), then it is cached and runs offline forever after.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The advantages (why it is worth it)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Private by construction.&lt;/strong&gt; After the download, nothing leaves the device. No prompt logging, no data shipped to a third party.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No API key, no per-token bill.&lt;/strong&gt; The cost is hardware you already own. Run it as much as you like.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Works offline.&lt;/strong&gt; On a plane, in a clinic, on a factory floor, in the field, with no signal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can run 24/7.&lt;/strong&gt; An always-on background agent, a daemon, a scheduled job. No metered API to drain your budget, no rate limits to throttle you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low latency.&lt;/strong&gt; No network round-trip to a data center.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You own it.&lt;/strong&gt; The weights are on your disk. Nothing changes under you.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The rule of thumb
&lt;/h3&gt;

&lt;p&gt;Reach for local AI when privacy, offline use, always-on operation, high volume, or cost are what matter. Reach for a frontier cloud model when you need maximum reasoning or breadth and the data is not sensitive. Many of the best products use both: local for the private, high-frequency path, cloud for the occasional hard question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can build: the 12 capabilities
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb6pfs01nrk6kg24o139n.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb6pfs01nrk6kg24o139n.jpeg" alt=" " width="800" height="303"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;QVAC is Tether's open-source local-AI SDK. One &lt;code&gt;npm install&lt;/code&gt; gives you twelve AI tasks behind one consistent API, all running on-device:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Text and chat&lt;/strong&gt; - a local LLM to assist, summarize, extract, or classify by prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vision&lt;/strong&gt; - ask questions about an image, describe a photo, read a chart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings&lt;/strong&gt; - turn text into vectors for search and similarity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAG&lt;/strong&gt; - search your own documents and answer with citations, locally.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speech to text&lt;/strong&gt; - transcribe audio and meetings, multilingual.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text to speech&lt;/strong&gt; - generate spoken audio in many languages, including voice cloning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Translation&lt;/strong&gt; - the Mozilla Bergamot engine, many language pairs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image generation&lt;/strong&gt; - text to image, on-device.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OCR&lt;/strong&gt; - read text out of scans and photos.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classification&lt;/strong&gt; - sort inputs into labels.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image upscaling&lt;/strong&gt; - enhance images on-device.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On-device fine-tuning&lt;/strong&gt; - adapt a small model to a person or domain, trained on the device itself.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The point is the consistency: your agent learns the pattern once and can reach for any of the twelve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example apps to spark ideas
&lt;/h2&gt;

&lt;p&gt;You do not build "an AI feature", you build a product. A few that play to local AI's strengths, each one you would simply describe to your coding agent or harness:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A private meeting transcriber and summarizer&lt;/strong&gt; (speech to text + chat). Recordings never leave the laptop. Good for lawyers, doctors, journalists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chat with your notes&lt;/strong&gt; (RAG + embeddings). A local "second brain" that answers from your own files, with citations, offline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A private medical Q&amp;amp;A assistant&lt;/strong&gt; (chat). Answer clinical questions on-device with &lt;a href="https://qvac.tether.io/blog/meet-medpsy-a-private-medical-ai-small-enough-for-your-phone/" rel="noopener noreferrer"&gt;MedPsy, a medical model small enough for a phone&lt;/a&gt; - nothing sent to a server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An offline travel translator&lt;/strong&gt; (translation + speech to text + text to speech). Works with no signal abroad.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An always-on document sorter&lt;/strong&gt; (OCR + classification). A daemon that files incoming scans around the clock, with no per-page API cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An on-device image studio&lt;/strong&gt; (image generation + upscaling). No cloud render bills, no usage caps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A kiosk or in-car voice assistant&lt;/strong&gt; (speech to text + chat + text to speech). Fully offline and private.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A personal-style writing model&lt;/strong&gt; (fine-tuning). Adapt a small model to your own voice, trained on your machine, never uploaded. (See how it works in &lt;a href="https://qvac.tether.io/blog/lora-fine-tuning-bitnet-b1-58-llms-on-heterogeneous-edge-gpus-via-qvac-fabric/" rel="noopener noreferrer"&gt;LoRA fine-tuning on edge GPUs via QVAC Fabric&lt;/a&gt;.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Want to see real builds on QVAC? &lt;a href="https://qvac.tether.io/blog/how-to-easily-run-a-fully-local-private-ai-assistant-with-openclaw-and-qvac/" rel="noopener noreferrer"&gt;OpenClaw runs a fully local, private AI assistant&lt;/a&gt;, and you can even &lt;a href="https://qvac.tether.io/blog/run-opencode-locally-with-qvac-a-coding-agent-with-no-cloud/" rel="noopener noreferrer"&gt;run your own coding agent on a local model with OpenCode&lt;/a&gt;. Pick an idea, describe its first slice to your agent, and iterate.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to build it with your agent
&lt;/h2&gt;

&lt;p&gt;The whole workflow is short:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start an empty project&lt;/strong&gt; and install the SDK (&lt;code&gt;npm install @qvac/sdk&lt;/code&gt;; you need Node 22.17 or newer, which is above the usual LTS).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give your agent QVAC.md.&lt;/strong&gt; This is the step that matters. Coding agents are unreliable on SDKs they have not seen much of: they invent functions that do not exist. QVAC.md is a single file you drop into your agent's rules (a &lt;code&gt;CLAUDE.md&lt;/code&gt;, an &lt;code&gt;AGENTS.md&lt;/code&gt;, a Cursor or Cline rules file) that gives it the real API. With it, the agent writes code that runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Describe one capability at a time.&lt;/strong&gt; For example: "Using the QVAC context, build a CLI that transcribes an audio file passed on the command line, fully locally."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run it and iterate&lt;/strong&gt; with your agent.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Cross-platform: what to keep in mind
&lt;/h2&gt;

&lt;p&gt;The same QVAC code runs on desktop (macOS, Linux, Windows via Node) and on mobile (iOS and Android, native, via Expo). A few things worth telling your agent up front:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Match the model to the target device.&lt;/strong&gt; A workstation can run a large model; a phone needs a small one. The model has to fit the device's memory, so choose per platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The GPU differs by platform&lt;/strong&gt; (Apple Metal, Vulkan, or a CPU fallback). The SDK picks the right backend automatically, but speed varies, so test on the real device, not just your laptop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The first download can be large.&lt;/strong&gt; On mobile especially, show a progress state so users are not staring at a frozen screen on first launch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node 22.17+ on desktop.&lt;/strong&gt; Below that, setup fails in confusing ways.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Get QVAC.md
&lt;/h2&gt;

&lt;p&gt;Everything above is the human's half. The agent's half is &lt;strong&gt;QVAC.md&lt;/strong&gt;: a single drop-in file with the real API surface, the model-type rules, the gotchas that stop hallucinations, and verified snippets. Paste it into your agent and it builds local AI apps that run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://qvac.tether.io/blog/quick-start-build-your-first-local-ai-app-with-your-coding-agent/" rel="noopener noreferrer"&gt;Get QVAC.md for your coding agent&lt;/a&gt;&lt;/strong&gt; on the original guide. Drop your email and we send you the drop-in file. Paste it into your agent and start building.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Which coding agent does this work with?
&lt;/h3&gt;

&lt;p&gt;Any of them. Cursor, Claude Code, Cline, Aider, OpenCode, or any harness that reads project rules or takes a system prompt. QVAC.md tells you where to paste it for each.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need a GPU?
&lt;/h3&gt;

&lt;p&gt;No. QVAC uses your GPU when one is available (Metal, NVIDIA, AMD, Intel) and falls back to CPU otherwise. Small models run fine on a laptop CPU; larger models are much faster with a GPU.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does a local model compare to ChatGPT or Claude?
&lt;/h3&gt;

&lt;p&gt;A model that fits on your device is smaller, so on the hardest reasoning and the broadest knowledge it is weaker than a frontier cloud model. In exchange it is private, offline, free to run, and always on. Pick the right tool for the job, and feel free to combine both.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is my data sent anywhere?
&lt;/h3&gt;

&lt;p&gt;No. After the model downloads once, inference happens entirely on-device. No API call, no prompt logging at run time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can this run on a phone?
&lt;/h3&gt;

&lt;p&gt;Yes. The same code runs natively on iOS and Android through Expo, on the device hardware. Use a smaller model to fit the phone's memory. &lt;a href="https://qvac.tether.io/blog/meet-medpsy-a-private-medical-ai-small-enough-for-your-phone/" rel="noopener noreferrer"&gt;MedPsy&lt;/a&gt; is one example of a model built to run there.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is it free? What is the license?
&lt;/h3&gt;

&lt;p&gt;The SDK is Apache 2.0 and free, including commercial use. It is JavaScript and TypeScript, on Node 22.17+, Bare 1.24+, and Expo 54+.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;You build local AI apps by directing your coding agent, not by memorizing an SDK.&lt;/li&gt;
&lt;li&gt;Local AI trades raw frontier capability for privacy, offline use, zero per-call cost, and 24/7 operation. Choose projects that want those.&lt;/li&gt;
&lt;li&gt;One &lt;code&gt;npm install&lt;/code&gt; gives you twelve on-device AI tasks behind one API.&lt;/li&gt;
&lt;li&gt;QVAC.md is what makes your agent write correct QVAC code.&lt;/li&gt;
&lt;li&gt;The same code runs on desktop and mobile; match the model to the device.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Next steps
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Grab &lt;a href="https://qvac.tether.io/blog/quick-start-build-your-first-local-ai-app-with-your-coding-agent/" rel="noopener noreferrer"&gt;QVAC.md&lt;/a&gt; and paste it into your agent.&lt;/li&gt;
&lt;li&gt;Pick one app idea above and describe its first slice to your agent.&lt;/li&gt;
&lt;li&gt;Read the SDK docs and the model registry when you want to go deeper.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a big commitment. You are not adopting a framework, you are giving your agent the right context and letting it build. The intelligence in the app you ship will not be rented from a server you do not control; it will run where your code runs, on hardware someone already owns. Pick one small idea, have it run entirely on-device, and you will feel the difference the first time it answers with the wifi off.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://qvac.tether.io/blog/quick-start-build-your-first-local-ai-app-with-your-coding-agent/" rel="noopener noreferrer"&gt;QVAC blog&lt;/a&gt;. QVAC is Tether's open-source local-AI SDK. Apache 2.0.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>local</category>
    </item>
  </channel>
</rss>
