<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: raphiki</title>
    <description>The latest articles on DEV Community by raphiki (@raphiki).</description>
    <link>https://dev.to/raphiki</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F982002%2Fb4188602-61e2-49e6-85be-d590a9b2e228.png</url>
      <title>DEV Community: raphiki</title>
      <link>https://dev.to/raphiki</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/raphiki"/>
    <language>en</language>
    <item>
      <title>Running DeepSeek Harness on an 8GB GPU: Context Tuning, Presets, and Hardware Limits</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Sat, 12 Sep 2026 05:47:39 +0000</pubDate>
      <link>https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i</link>
      <guid>https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 3 of the DeepSeek Harness: Kernel to Edge series: A hands-on benchmark pairing DeepSeek Harness, Ollama, and Ornith-1.5 on an 8GB RTX 4070 laptop to solve the tradeoff between tool bloat and context headroom.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek Harness: Kernel to Edge&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
This article is &lt;strong&gt;Part 3&lt;/strong&gt; of a three-part architectural and practical deep dive:  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Part 1&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Understanding Cordis: The TypeScript Framework Built for Hot-Swapping Everything&lt;/a&gt; (Microkernel primitives, spatiotemporal composability, reverse cleanup stacks, zero-leak lifecycles).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 2&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;DeepSeek Harness: How DeepSeek Uses Cordis to Redefine Autonomous AI Agents&lt;/a&gt; (The meta-harness paradigm, Cordis as an agent kernel, comparing DSH to OpenCode and Pi).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3 (This Article)&lt;/strong&gt;: Running DeepSeek Harness on an 8GB GPU: Context Tuning, Presets, and Hardware Limits (Hands-on local deployment with Ollama, Ornith-1.5, VRAM arithmetic, minimal vs. standard presets, and TUI).&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. Introduction: Bringing the Harness Home
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Part 1&lt;/a&gt;, we explored &lt;strong&gt;Cordis&lt;/strong&gt;, the TypeScript framework designed for zero-leak dynamic lifecycles and spatiotemporal composability. In &lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;Part 2&lt;/a&gt;, we mapped the systems architecture of &lt;strong&gt;DeepSeek Harness (DSH)&lt;/strong&gt;, showing how it abandons monolithic Python agent loops in favor of a pure plugin microkernel.&lt;/p&gt;

&lt;p&gt;Now, we move from architectural theory to a practical, working developer setup.&lt;/p&gt;

&lt;p&gt;My target machine is a standard laptop equipped with an &lt;strong&gt;NVIDIA RTX 4070 GPU (8GB VRAM)&lt;/strong&gt; and 32GB of system RAM. &lt;/p&gt;

&lt;p&gt;Running autonomous coding agents locally on consumer hardware presents strict physical constraints. Frontier cloud models (like Claude 3.7 Sonnet or DeepSeek-V3) make multi-turn tool calling seem straightforward due to their massive context windows and vast parameter counts. When running an agent framework on an 8GB GPU, you immediately encounter hard physical boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;VRAM exhaustion&lt;/strong&gt;: Loading a model and its KV cache can easily exceed 8GB, causing CUDA out-of-memory errors or offloading layers to the CPU, which slows execution to a crawl.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context window starvation&lt;/strong&gt;: Agent frameworks typically inject extensive tool definitions, system instructions, and file trees into the prompt. A default local context window fills up before the agent can complete its first multi-turn refactor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-calling fragility&lt;/strong&gt;: Smaller quantized models often hallucinate parameters or break JSON schemas when overloaded with too many simultaneous tools.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In this tutorial, we will configure a local coding agent on this 8GB machine using DeepSeek Harness, Ollama, and &lt;strong&gt;Ornith-1.5&lt;/strong&gt;. We will explore why DSH's plugin design adapts well to local execution, how to tune the context window, and how agent presets help manage memory bottlenecks.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Why Choose DeepSeek Harness for a Local Agent?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3vvnk0ansje0611d7c0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3vvnk0ansje0611d7c0.png" alt="DSH Logo" width="432" height="95"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Most developers look at coding harnesses like OpenCode or Pi when building personal workflows. As we established in Part 2, those are turnkey developer tools designed around fixed, opinionated toolsets.&lt;/p&gt;

&lt;p&gt;When your hardware is constrained to 8GB of VRAM, opinionated harnesses can become restrictive. If a harness hardcodes ten developer tools into every single turn, your local model has to parse ten complex JSON schemas on every inference call. That consumes hundreds of tokens of precious context and increases the cognitive load on a smaller model.&lt;/p&gt;

&lt;p&gt;This is where DeepSeek Harness's architecture as a &lt;strong&gt;meta-harness&lt;/strong&gt; becomes practical:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Native Extensibility and Granular Tool Control&lt;/strong&gt;: Because DSH has no privileged core, tools and interfaces are not baked into the runtime. Even the Web UI itself is just a plugin: developers who prefer keyboard-driven terminal workflows can easily swap it for community TUIs like &lt;code&gt;dsh-TUI&lt;/code&gt;. Across community directories like &lt;a href="https://deepseekplugin.com/" rel="noopener noreferrer"&gt;deepseekplugin.com&lt;/a&gt; and &lt;a href="https://dsh-plugins.org/en" rel="noopener noreferrer"&gt;dsh-plugins.org&lt;/a&gt;, thousands of plugins exist for vision, generative UI, and memory. But when constrained by 8GB of VRAM, this extensibility works in reverse: instead of piling on plugins, you can mount only the exact tools your local model needs (such as file editing and bash) while leaving out all memory-heavy peripherals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Predictable Teardown via Cordis Fibers&lt;/strong&gt;: Local development involves running code, launching compilers, and watching files. Cordis ensures that every child process or execution sandbox is cleanly terminated when a task completes, preventing zombie processes from stealing system RAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Declarative Agent Presets&lt;/strong&gt;: DSH supports configurable agent presets (&lt;code&gt;settings.yaml&lt;/code&gt;). Switching from a full development dashboard to a lightweight minimal mode requires only a single configuration line.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiv6mihyxkhdit9im5tzy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiv6mihyxkhdit9im5tzy.png" alt="Local Hardware Constraints: 8GB VRAM Physical Ceiling" width="799" height="224"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every token saved in the system prompt directly translates to more room for code and conversation history.&lt;/p&gt;

&lt;h3&gt;
  
  
  Defining the Use Case: An Autonomous Task Worker, Not an Interactive Chatbot
&lt;/h3&gt;

&lt;p&gt;Hardware constraints depend heavily on the intended workload.&lt;/p&gt;

&lt;p&gt;Local models are frequently judged as &lt;strong&gt;interactive chat assistants&lt;/strong&gt; (like ChatGPT) or &lt;strong&gt;inline autocomplete engines&lt;/strong&gt; (like GitHub Copilot), where fast token streaming (40 to 60+ tokens per second) is required to maintain user momentum.&lt;/p&gt;

&lt;p&gt;DeepSeek Harness serves a different workload: &lt;strong&gt;autonomous, multi-turn task execution&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A typical task involves handing off an end-to-end debugging session in an existing repository:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"There is an edge case in the WebSocket session manager where dropped connections leave stale socket descriptors open. Locate the relevant files, write a script to reproduce the leak, fix the cleanup logic, and confirm that the test suite passes."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In this setup:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Long-running execution&lt;/strong&gt;: The agent runs across 10 to 30 continuous turns: inspecting directory trees, reading source files, running test scripts in bash, analyzing error traces, applying patches, and re-running the test suite.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Throughput vs. Resilience&lt;/strong&gt;: On consumer hardware, a 9B model generating 10 to 20 tokens per second is impractical for live conversation, but completely sufficient for an asynchronous worker running unattended in the background or in a terminal (&lt;code&gt;dsh-TUI&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The primary constraints&lt;/strong&gt;: The operational bottleneck is not token generation latency, but &lt;strong&gt;lifecycle stability&lt;/strong&gt; (ensuring spawned compiler processes and file watchers do not leak system resources) and &lt;strong&gt;context headroom&lt;/strong&gt; (preventing prompt overhead from pushing VRAM into system memory).&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  3. Step 1: Installation and Baseline Test with OpenRouter
&lt;/h2&gt;

&lt;p&gt;Before troubleshooting local quantization or GPU memory allocation, we validate that DeepSeek Harness operates correctly on our system using a known cloud endpoint. Building outside-in ensures that any subsequent issues stem from local model configurations, not the harness itself.&lt;/p&gt;

&lt;h3&gt;
  
  
  Installation
&lt;/h3&gt;

&lt;p&gt;DeepSeek Harness is packaged as an npm module. You can run it directly using &lt;code&gt;npx&lt;/code&gt; or install it globally:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Global installation via npm&lt;/span&gt;
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @deepseek-ai/dsh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Baseline Cloud Configuration
&lt;/h3&gt;

&lt;p&gt;For our initial baseline test, we will connect DSH to &lt;strong&gt;OpenRouter&lt;/strong&gt;, giving us access to frontier models (such as DeepSeek-V3 or Mistral) to confirm tool calling and web console connectivity.&lt;/p&gt;

&lt;p&gt;Export your API key in your terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Windows PowerShell&lt;/span&gt;
&lt;span class="nv"&gt;$env&lt;/span&gt;:OPENROUTER_API_KEY &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"sk-or-v1-your-key-here"&lt;/span&gt;

&lt;span class="c"&gt;# Linux / macOS&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-or-v1-your-key-here"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Launch the web interface:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;C:\Local\Dev\dsh&amp;gt;&lt;/span&gt;dsh web
&lt;span class="go"&gt;dsh web: http://127.0.0.1:3080/?token=4n7i06dEtYEUzwuPfVmXVRGEx-fxFRRgoQJx4Yfq3wg
&lt;/span&gt;&lt;span class="gp"&gt;dsh web: opening the default browser;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;pass &lt;span class="nt"&gt;--no-open&lt;/span&gt; to disable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By default, DSH starts its local web server at &lt;code&gt;http://localhost:3080&lt;/code&gt;. Open your browser and navigate to the Coding Agent.&lt;/p&gt;

&lt;p&gt;Verify the baseline by sending simple instructions, such as inspecting the current directory or reading the contents of a file.&lt;/p&gt;

&lt;p&gt;Watch the console stream the model's reasoning, dispatch the filesystem tool, and return the results. Once you confirm this loop completes without errors, you know the harness, process permissions, and session streams are healthy.&lt;/p&gt;

&lt;p&gt;The screenshot below shows the DSH web interface connected to Mistral Small via OpenRouter, displaying initial latency and context metrics.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg8ow2ixdqrzoojicnil2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg8ow2ixdqrzoojicnil2.png" alt="DeepSeek UI - Chat View" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Trajectory view breaks down individual execution turns, separating model thinking blocks from raw tool payloads.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz64najnafmh9bn7g85eb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz64najnafmh9bn7g85eb.png" alt="DeepSeek UI - Trajectory View" width="800" height="535"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Notice in the navigation bar that DSH initialized in &lt;strong&gt;Standard mode&lt;/strong&gt;; we will explore how presets shape agent behavior in Step 3.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Step 2: Selecting and Preparing the Local Model (Ornith-1.5)
&lt;/h2&gt;

&lt;p&gt;Now we replace the cloud API with a fully local open-weight model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Ornith-1.5?
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzgut2lt40pw3xmti7uqk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzgut2lt40pw3xmti7uqk.png" alt="Ornith Logo" width="280" height="85"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Model selection on an 8GB GPU balances two requirements: reasoning capacity for multi-file code analysis and physical memory headroom.&lt;/p&gt;

&lt;h4&gt;
  
  
  1. What is Ornith-1.5? (A Specialized Qwen Fine-Tune)
&lt;/h4&gt;

&lt;p&gt;&lt;strong&gt;Ornith-1.5 (9B)&lt;/strong&gt; is an open-weight model derived from the &lt;strong&gt;Qwen2.5&lt;/strong&gt; architecture.&lt;/p&gt;

&lt;p&gt;While Qwen2.5 provides a strong baseline for code syntax and logic, general instruction-tuned checkpoints often struggle with structured tool calling in autonomous harnesses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Schema errors&lt;/strong&gt;: Base models frequently omit mandatory JSON keys, invent unsupported parameters, or fail to escape control characters in code strings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversational leakage&lt;/strong&gt;: Conversational models often prepend tool calls with conversational filler (&lt;em&gt;"Sure, I will execute this command for you..."&lt;/em&gt;), which breaks strict JSON schema parsers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context degradation&lt;/strong&gt;: During extended multi-turn sessions, general models can lose track of earlier tool outputs or swap arguments between different tool interfaces (such as passing bash flags into file editing parameters).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Ornith-1.5&lt;/strong&gt; addresses this through fine-tuning on structured agent traces and explicit function-calling datasets. It acts as a tool execution engine: given a tool schema, it generates valid, compact JSON calls on the first attempt and processes output streams directly.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Why 9B Parameters on an 8GB GPU?
&lt;/h4&gt;

&lt;p&gt;On an 8GB GPU, parameter count determines whether execution remains hardware-accelerated:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;7B Models&lt;/strong&gt;: Fast and lightweight, but often struggle with the multi-step deductive reasoning needed to navigate complex repository structures unattended.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;14B Models&lt;/strong&gt;: At 4-bit quantization, 14B weights require ~9GB, immediately spilling across the PCIe bus into system RAM and dropping throughput to 1 to 2 tokens per second.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 9B Sweet Spot&lt;/strong&gt;: At 4-bit quantization (&lt;code&gt;Q4_K_M&lt;/code&gt;), Ornith-1.5 occupies approximately &lt;strong&gt;5.5GB of VRAM&lt;/strong&gt;. This leaves roughly &lt;strong&gt;2.5GB of headroom&lt;/strong&gt; on an 8GB card, providing the space needed for an extended KV cache and display buffers without triggering CPU paging.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Inference Server: Ollama
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftssq7h9q3rur9durrib9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftssq7h9q3rur9durrib9.png" alt="Ollama Logo" width="292" height="113"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We will use &lt;strong&gt;Ollama&lt;/strong&gt; to serve the model locally. Ollama provides a native OpenAI-compatible endpoint (&lt;code&gt;http://127.0.0.1:11434/v1&lt;/code&gt;), making integration straightforward.&lt;/p&gt;

&lt;h3&gt;
  
  
  The 2048 Context Constraint
&lt;/h3&gt;

&lt;p&gt;By default, Ollama initializes models with a 2,048 token context window (&lt;code&gt;num_ctx 2048&lt;/code&gt;). While sufficient for simple chat queries, an autonomous coding agent quickly exhausts this budget:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;System instructions and agent personas consume ~400 tokens.&lt;/li&gt;
&lt;li&gt;Tool schemas (filesystem, bash, code editor) consume ~800 to 1,200 tokens.&lt;/li&gt;
&lt;li&gt;Workspace context (directory trees, project files) easily consumes another 500+ tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Under default settings, tool schemas and prompts consume most of the available window before the model even generates its first turn. As soon as a multi-file task begins, the engine either truncates earlier messages or drops tool definitions from context.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tuning the Context Window to 16K
&lt;/h3&gt;

&lt;p&gt;To give our agent room to work, we need to create a customized model variant with a 16K context window (&lt;code&gt;16384&lt;/code&gt; tokens).&lt;/p&gt;

&lt;p&gt;We extract the base model's Modelfile, append the &lt;code&gt;num_ctx&lt;/code&gt; parameter, and build a local derivative:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;ollama show --modelfile ornith-1.5:9b &amp;gt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;Modelfile
&lt;span class="gp"&gt;echo PARAMETER num_ctx 16384 &amp;gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; Modelfile
&lt;span class="go"&gt;
ollama create ornith-1.5:ctx -f Modelfile
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This sequence extracts the base template, sets the KV cache allocation parameter to 16,384 tokens, and compiles the new model tag in Ollama.&lt;/p&gt;

&lt;h3&gt;
  
  
  The VRAM Arithmetic
&lt;/h3&gt;

&lt;p&gt;Evaluating memory requirements for a 16K context window on an 8GB GPU:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3oiw3rhc2w5frrldrk7z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3oiw3rhc2w5frrldrk7z.png" alt="VRAM Allocation on RTX 4070 (8GB Total)" width="799" height="263"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The model and its extended context window fit entirely within GPU VRAM. Zero layers are offloaded to CPU system memory, ensuring maximum generation speed.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Step 3: Trials, Failures, and the Minimal Preset
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2mnypevalgrhkkh0jg8d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2mnypevalgrhkkh0jg8d.png" alt="DSH - Settings - Models" width="629" height="506"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Connecting our newly created &lt;code&gt;ornith-1.5:ctx&lt;/code&gt; model to DeepSeek Harness is straightforward, but our initial run revealed an instructive failure mode. &lt;/p&gt;

&lt;p&gt;To understand what went wrong, we first need to clarify how DeepSeek Harness structures its runtime across three distinct layers: &lt;strong&gt;Plugins&lt;/strong&gt;, &lt;strong&gt;Profiles &amp;amp; Settings&lt;/strong&gt;, and &lt;strong&gt;Agent Presets&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Untangling the Hierarchy: Plugins, Profiles, and Presets
&lt;/h3&gt;

&lt;p&gt;Developers often use these terms interchangeably, but in DSH they represent a clear three-tier architecture:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvn5lkkn8ajmq3p7m7ioq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvn5lkkn8ajmq3p7m7ioq.png" alt="DeepSeek Harness Architectural Hierarchy: Plugins, Profiles, and Presets" width="800" height="332"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Plugins (The Capabilities)&lt;/strong&gt;: Atomic modules in the Cordis ecosystem (such as &lt;code&gt;@deepseek-ai/dsh-tool-filesystem&lt;/code&gt; or &lt;code&gt;llm-pi-ai&lt;/code&gt;, which adapts the provider engine from the open-source pi project). Each plugin provides a specific service, model adapter, tool schema, or execution sandbox.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Profiles &amp;amp; Settings (The Application Runtime)&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Profiles&lt;/strong&gt;: Application runners (such as &lt;code&gt;dsh web&lt;/code&gt;, &lt;code&gt;headless&lt;/code&gt;, or &lt;code&gt;sdk&lt;/code&gt;) located under &lt;code&gt;$DSH_HOME/profiles/&amp;lt;name&amp;gt;&lt;/code&gt;. They define which ordered bundle of plugins Cordis boots into memory for a specific user interface or execution target.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User Settings (&lt;code&gt;settings.yaml&lt;/code&gt;)&lt;/strong&gt;: The global configuration file located at &lt;code&gt;$DSH_HOME/settings.yaml&lt;/code&gt; (or project root). It hot-reloads at runtime without restarting the server, configuring provider credentials, API endpoints, model mappings, and default options.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent Presets (The Operational Modes)&lt;/strong&gt;: The per-session runtime configurations stored in &lt;code&gt;agent-presets/&amp;lt;id&amp;gt;/agent.cordis.yml&lt;/code&gt; and toggled directly in the UI as &lt;strong&gt;"Standard mode"&lt;/strong&gt; or &lt;strong&gt;"Minimal mode"&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Because of this separation, &lt;strong&gt;optimizing DSH for local hardware does not require uninstalling packages or restarting the daemon.&lt;/strong&gt; While profiles and plugins establish what the server application can do, agent presets dynamically select which tool schemas, system prompt instructions, and context compaction rules are active for the LLM during an execution turn.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Built-in Presets: Four Flavors of DSH
&lt;/h3&gt;

&lt;p&gt;DeepSeek Harness ships with four distinct built-in agent presets out of the box, each tailored to different operational requirements:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Standard Mode (The Full Coding Suite)&lt;/strong&gt;: The default preset for comprehensive repository work. It equips the agent with the complete toolset: file editing, persistent bash shell access, repository and web search, planning, &lt;strong&gt;skills&lt;/strong&gt; (via filesystem discovery and the catalog loader), goals, workflows, and subagent orchestration. External &lt;strong&gt;MCP servers&lt;/strong&gt; are not mounted here by default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PTC Mode (Programmatic Tool Calling)&lt;/strong&gt;: Standard mode augmented with the Code Mode SDK. It retains all Standard capabilities (including &lt;strong&gt;skills&lt;/strong&gt;), while letting the model write and execute an entire TypeScript program to combine multi-step operations in a single pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minimal Mode (The Lean Two-Tool Composition)&lt;/strong&gt;: A stripped-down preset providing only two fundamental tools: &lt;strong&gt;persistent bash&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;str_replace_editor&lt;/code&gt;&lt;/strong&gt;. It explicitly excludes &lt;strong&gt;skills&lt;/strong&gt; and cannot mount &lt;strong&gt;MCP servers&lt;/strong&gt;, drastically reducing prompt bloat and simplifying schema constraints.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Creator Mode (The Meta-Preset)&lt;/strong&gt;: A specialized preset for inspecting, experimenting with, and authoring custom agent presets. It includes the Standard toolset (including &lt;strong&gt;skills&lt;/strong&gt;) alongside runtime introspection utilities to help developers draft new custom presets (such as connecting external &lt;strong&gt;MCP servers&lt;/strong&gt; via &lt;code&gt;@deepseek-ai/dsh-mcp-client&lt;/code&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note: Preset Session Semantics &amp;amp; Custom MCP Presets&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Preset selections in DSH apply on a &lt;strong&gt;per-session&lt;/strong&gt; basis. When you change a preset in the Web UI, it takes effect on the next session you start; active sessions keep the preset configuration they were initialized with. While external MCP servers are not enabled by default in any of the four built-in presets, you can duplicate any preset into a custom configuration (&lt;code&gt;agent-presets/&amp;lt;id&amp;gt;/&lt;/code&gt;) or use Creator mode to mount &lt;code&gt;@deepseek-ai/dsh-mcp-client&lt;/code&gt; connections.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The 8GB Dilemma: Why Standard Mode Threatens Local Context
&lt;/h3&gt;

&lt;p&gt;In Section 3, our baseline test with Mistral Small booted in DSH's default &lt;strong&gt;"Standard mode"&lt;/strong&gt;. &lt;/p&gt;

&lt;p&gt;For frontier cloud models with large context windows (128K+), Standard mode is ideal. It presents the model with a comprehensive developer environment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-file patch tools&lt;/li&gt;
&lt;li&gt;Interactive terminal runners&lt;/li&gt;
&lt;li&gt;Web search clients&lt;/li&gt;
&lt;li&gt;Diagnostic loggers and telemetry hooks&lt;/li&gt;
&lt;li&gt;Subagent delegation and workflow engines&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, when targeting a local 9B model (&lt;code&gt;ornith-1.5:ctx&lt;/code&gt;) constrained to an 8GB GPU, Standard mode introduces immediate friction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt Bloat&lt;/strong&gt;: Tool definitions consume roughly ~6.9K tokens of context before the user even submits a prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cognitive &amp;amp; Schema Load&lt;/strong&gt;: A 9B model must continuously hold dozens of complex JSON schemas in memory, increasing the likelihood of hallucinated arguments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Starvation&lt;/strong&gt;: When tool schemas consume over 40% of our 16K context window upfront, multi-turn reasoning rapidly saturates the remaining headroom and triggers frequent compactions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Initial Hypothesis: Configuring "Minimal Mode"
&lt;/h3&gt;

&lt;p&gt;To protect our limited context window, our initial design hypothesis is to test DSH's leanest built-in option: &lt;strong&gt;Minimal mode&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of modifying code or uninstalling plugins, we can instruct DSH to activate Minimal mode:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Masks out telemetry, web scrapers, subagents, workflows, and complex patch tools.&lt;/li&gt;
&lt;li&gt;Retains only core primitives: file modification via targeted string replacement (&lt;code&gt;str_replace_editor&lt;/code&gt;) and persistent shell execution.&lt;/li&gt;
&lt;li&gt;Slashes tool schema overhead from ~6.9K tokens down to &lt;strong&gt;only ~971 tokens&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On paper, the tradeoff is clear: &lt;strong&gt;we exchange peripheral tools to maximize reasoning headroom and generation throughput.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Standard Mode vs. Minimal Mode: The 8GB Tradeoff
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Standard Mode&lt;/th&gt;
&lt;th&gt;Minimal Mode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Primary Target&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Frontier cloud models (DeepSeek-V3, Claude)&lt;/td&gt;
&lt;td&gt;Local quantized models (7B to 9B on 8GB GPU)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Active Tool Surface&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full suite (Read, Edit, Write, Shell, Search, Subagents)&lt;/td&gt;
&lt;td&gt;Two-tool composition (&lt;code&gt;persistent bash&lt;/code&gt;, &lt;code&gt;str_replace_editor&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool Schema Footprint&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~6.9K tokens (~43% of 16K window)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~971 tokens (&amp;lt;6% of 16K window)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Available Context (16K)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~7K to 8K tokens remaining&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~15K tokens (&amp;gt;93% available)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Measured Throughput&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~12 tokens/sec (TTFT ~3.9s)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~17 tokens/sec (TTFT ~1.5s)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observed Task Behavior&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Capable of deep reasoning, but heavy context churn&lt;/td&gt;
&lt;td&gt;Lean schemas, but rigid editing on file creation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;While presets can be toggled per session in the web interface, we declare Minimal mode as the default in &lt;code&gt;settings.yaml&lt;/code&gt; to test our lean-context hypothesis.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Step 4: The Complete DeepSeek Harness Configuration (&lt;code&gt;settings.yaml&lt;/code&gt;)
&lt;/h2&gt;

&lt;p&gt;All of these requirements are codified into DeepSeek Harness's primary configuration file: &lt;code&gt;settings.yaml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Place this configuration file in your DSH configuration directory (typically &lt;code&gt;~/.dsh/settings.yaml&lt;/code&gt; or in your project root):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;ui-onboarding&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;welcomeNoticeVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;2026-08-13.1&lt;/span&gt;

&lt;span class="na"&gt;llm-pi-ai&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;openrouter&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;OPENROUTER_API_KEY&lt;/span&gt;
    &lt;span class="na"&gt;ollama&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;displayName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Ollama&lt;/span&gt;
      &lt;span class="na"&gt;api&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai-completions&lt;/span&gt;
      &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://127.0.0.1:11434/v1&lt;/span&gt;
      &lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ornith-1.5:ctx&lt;/span&gt;
          &lt;span class="na"&gt;contextWindow&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16000&lt;/span&gt;
          &lt;span class="na"&gt;maxTokens&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;4096&lt;/span&gt;
      &lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;OLLAMA_API_KEY&lt;/span&gt;

&lt;span class="na"&gt;agent-default-model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ollama&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ornith-1.5:ctx&lt;/span&gt;

&lt;span class="na"&gt;agent-presets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;default&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;minimal&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Breakdown of the Configuration
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ui-onboarding&lt;/code&gt;&lt;/strong&gt;: Pins the onboarding notice version (&lt;code&gt;2026-08-13.1&lt;/code&gt;), preventing introductory dialogs from popping up on subsequent web console launches.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;llm-pi-ai&lt;/code&gt;&lt;/strong&gt;: Configures the underlying LLM provider adapter. DSH wraps the provider abstraction and streaming engine directly from Mario Zechner's open-source &lt;strong&gt;pi&lt;/strong&gt; project into a Cordis plugin. Rather than maintaining custom client drivers for each inference backend, DSH reuses an existing upstream module, illustrating the practical composability of the open-source ecosystem.

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;openrouter&lt;/code&gt;: Preserved as a cloud fallback provider whenever you need to benchmark or handle tasks exceeding local model capabilities.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ollama&lt;/code&gt;: Points DSH to Ollama's local OpenAI-compatible completions endpoint (&lt;code&gt;http://127.0.0.1:11434/v1&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;models&lt;/code&gt;: Explicitly registers our custom &lt;code&gt;ornith-1.5:ctx&lt;/code&gt; model with two critical runtime parameters:&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;contextWindow: 16000&lt;/code&gt;: Informs DSH of our 16K context ceiling (mirroring the &lt;code&gt;num_ctx 16384&lt;/code&gt; compiled into our Ollama Modelfile). This enables DSH to compute and display context utilization accurately in the web UI (such as &lt;code&gt;~10.6K / 16K&lt;/code&gt;) and trigger automated compactions before context overflows.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;maxTokens: 4096&lt;/code&gt;: Sets the maximum generation ceiling for a single turn. While developers sometimes enter the base model's advertised architectural capacity here (such as 32000), &lt;code&gt;maxTokens&lt;/code&gt; in DSH governs the single-response completion budget. Capping it to a realistic threshold like 4,096 tokens ensures individual outputs never exceed remaining context headroom or trigger runaway loops.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;apiKeyEnv&lt;/code&gt;: Ollama does not require an API key for local inference, but DSH checks for an environment variable name; pointing to &lt;code&gt;OLLAMA_API_KEY&lt;/code&gt; (even if empty) satisfies the schema.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;agent-default-model&lt;/code&gt;&lt;/strong&gt;: Directs DSH to boot immediately using our local Ollama model as the primary reasoning engine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;agent-presets&lt;/code&gt;&lt;/strong&gt;: Sets the default agent execution preset to &lt;code&gt;minimal&lt;/code&gt;, activating the lightweight minimal preset on every turn.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  7. Step 5: Running the Local Agent in Practice
&lt;/h2&gt;

&lt;p&gt;With &lt;code&gt;settings.yaml&lt;/code&gt; saved and Ollama running in the background, launch the DeepSeek Harness web console:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The web console launches locally, ready to pair with our &lt;code&gt;ornith-1.5:ctx&lt;/code&gt; model running on the 8GB RTX 4070.&lt;/p&gt;

&lt;p&gt;To test the agent under realistic conditions, we give it a practical software engineering task on an existing codebase:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Inspect the local calculator web application (&lt;code&gt;C:\Local\Dev\dsh\calculator&lt;/code&gt;), analyze its structure, diagnose any failing operations, and write an &lt;code&gt;AUDIT.md&lt;/code&gt; report documenting your findings."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Phase 1: Starting with Minimal Mode
&lt;/h3&gt;

&lt;p&gt;Because our 8GB mobile GPU restricts our local model to a 16K context window, our initial instinct is to minimize prompt bloat at all costs. We start the session using DSH's &lt;strong&gt;Minimal mode&lt;/strong&gt; preset.&lt;/p&gt;

&lt;p&gt;In Minimal mode, initial resource consumption is low:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tool definitions consume &lt;strong&gt;only ~971 tokens&lt;/strong&gt; (less than 6% of our 16K context window).&lt;/li&gt;
&lt;li&gt;The system prompt takes a negligible ~16 tokens.&lt;/li&gt;
&lt;li&gt;Generation throughput is fast at &lt;strong&gt;~17 tokens per second&lt;/strong&gt;, with an average Time to First Token (TTFT) of 1.5 seconds.&lt;/li&gt;
&lt;li&gt;Over 90% of the context window remains available for model reasoning.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, file operations fail when the agent attempts to record its findings on disk:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F51u9gtb1uenaz557iv8y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F51u9gtb1uenaz557iv8y.png" alt="Minimal mode tool footprint and str_replace_editor error" width="800" height="547"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The limitation lies in the tools provided by the minimal preset. Minimal mode relies on &lt;code&gt;str_replace_editor&lt;/code&gt; and persistent bash. While &lt;code&gt;str_replace_editor&lt;/code&gt; works well for targeted replacements in existing files, its schema constraints are rigid when asked to create or format new files. When Ornith-1.5 attempts to write the audit file to disk, it produces recurring schema validation errors:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;Error:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;invalid&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;arguments:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"new_str"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;must&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;match&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;exactly&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;one&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;oneOf&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;branch&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(matched&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="err"&gt;);&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"old_str"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;must&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;match&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;exactly&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;one&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;oneOf&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;branch&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;(matched&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="err"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Switching to the &lt;strong&gt;Trajectory&lt;/strong&gt; tab in the DSH web interface exposes how the agent attempts to cope with this failure:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frfmd436yncn7vpp1j8n6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frfmd436yncn7vpp1j8n6.png" alt="Minimal mode trajectory and repeat-tool-reminder" width="799" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Repetitive Tool Failures&lt;/strong&gt;: Ornith-1.5 retries the file write through &lt;code&gt;str_replace_editor&lt;/code&gt;, but repeatedly hits schema rejection (&lt;code&gt;INVALID_ARGS&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Guardrails Intervene&lt;/strong&gt;: DSH's active &lt;code&gt;repeat-tool-reminder&lt;/code&gt; plugin detects the loop and injects an inline context message warning the model that it is repeating identical failed calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shell Fallbacks&lt;/strong&gt;: While Ornith-1.5 recognizes the error, the minimal preset offers no alternative file creation tools. The model falls back to executing raw PowerShell (&lt;code&gt;pwsh&lt;/code&gt;) scripts to write files character by character and reading back hex byte sequences (&lt;code&gt;23 # 20 43 C&lt;/code&gt;) to verify output integrity.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In practice, while Minimal mode keeps tool schemas under 1K tokens, its editing surface is too rigid for repository maintenance. The preset lacks the necessary file manipulation tools for an end-to-end audit.&lt;/p&gt;




&lt;h3&gt;
  
  
  Phase 2: Switching to Standard Mode
&lt;/h3&gt;

&lt;p&gt;To provide the tools necessary for proper software engineering, we switch our session preset from Minimal mode to &lt;strong&gt;Standard mode&lt;/strong&gt; directly in the DSH interface.&lt;/p&gt;

&lt;p&gt;Standard mode provides a full development toolset with dedicated &lt;code&gt;read&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, and interactive user dialog tools (&lt;code&gt;ask_user&lt;/code&gt;). File operations proceed without schema errors:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmz6rrfqoyk1wwys3f0kl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmz6rrfqoyk1wwys3f0kl.png" alt="Standard mode ~6.9K tool context footprint" width="799" height="399"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;However, tool schema overhead increases significantly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The expanded tool schemas consume &lt;strong&gt;~6.9K tokens&lt;/strong&gt; out of our 16K limit (over 43% of total context).&lt;/li&gt;
&lt;li&gt;Together with the ~1.8K system prompt, &lt;strong&gt;more than half of the context window is consumed before reading project files&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Generation speed drops from 17 tokens/s down to &lt;strong&gt;~12 tokens/s&lt;/strong&gt;, and average TTFT increases to ~3.9 seconds.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Code Reasoning and Bug Localization
&lt;/h4&gt;

&lt;p&gt;With dedicated file inspection tools available, the model traces the bug accurately.&lt;/p&gt;

&lt;p&gt;The agent uses &lt;code&gt;read&lt;/code&gt; to examine &lt;code&gt;app.js&lt;/code&gt; and &lt;code&gt;index.html&lt;/code&gt;. It identifies that in the calculator app, typing &lt;code&gt;"8+8"&lt;/code&gt; unexpectedly displays &lt;code&gt;"88"&lt;/code&gt; instead of performing addition:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhl96ioiukkrrxdr6swhe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhl96ioiukkrrxdr6swhe.png" alt="Standard mode reasoning and bugfix prompt" width="800" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ornith-1.5 constructs a detailed step-by-step state trace. Using DSH's interactive prompt dialog (&lt;code&gt;ask_user&lt;/code&gt;), the agent presents its findings clearly to the user, proposing Option A and asking for confirmation before making edits. Once approved, Ornith-1.5 applies the fix cleanly using the &lt;code&gt;edit&lt;/code&gt; tool.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Context Ceiling: Auto-Compactions in Action
&lt;/h4&gt;

&lt;p&gt;While the code fix succeeded, inspecting the execution trajectory highlights the consequence of carrying ~6.9K tokens of tool schemas:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbpkav0cs4oqiustnpel8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbpkav0cs4oqiustnpel8.png" alt="Standard mode context compaction trajectory" width="800" height="536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After just two turns and eleven steps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Total context utilization quickly surpasses 66% (~10.6K / 16K) and approaches saturation.&lt;/li&gt;
&lt;li&gt;DeepSeek Harness begins triggering &lt;strong&gt;frequent automatic context compactions&lt;/strong&gt; (&lt;code&gt;COMPACTED: summary is not smaller than shadowed content...&lt;/code&gt;) to prevent context overflow.&lt;/li&gt;
&lt;li&gt;Re-evaluating 10K+ tokens on every turn adds noticeable latency to local inference on the RTX 4070.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Empirical Validation of the VRAM Arithmetic
&lt;/h4&gt;

&lt;p&gt;Throughout these multi-turn test runs across both Minimal and Standard modes, monitoring GPU telemetry via &lt;code&gt;nvidia-smi&lt;/code&gt; provided a direct empirical confirmation of our theoretical model from Section 4:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs8mgpt6iqpvtcp5dhdue.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs8mgpt6iqpvtcp5dhdue.png" alt="Theoretical vs. Empirical VRAM on RTX 4070 (8GB Total)" width="800" height="254"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The observed VRAM allocation stabilized at approximately &lt;strong&gt;7.6 GB&lt;/strong&gt;, closely matching our 7.8 GB calculation within 200MB of runtime variance. &lt;/p&gt;

&lt;p&gt;More importantly, generation speed held steady at &lt;strong&gt;17 tokens/s in Minimal mode&lt;/strong&gt; and &lt;strong&gt;12 tokens/s in Standard mode&lt;/strong&gt;. If total memory consumption had crossed the 8.0 GB physical boundary, Windows unified memory manager would have immediately paged memory buffers across the PCIe bus into shared system RAM. That memory bus penalty causes generation throughput to plummet to 1-3 tokens/s. The sustained double-digit token generation rate confirms that our 9B model weights and the entire 16K KV cache remained &lt;strong&gt;100% GPU-resident&lt;/strong&gt; at all times.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Engineering Takeaway: Fine-Tuning the Preset Sweet Spot
&lt;/h3&gt;

&lt;p&gt;This hands-on experiment demonstrates the core dilemma of local AI agents on consumer hardware:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Minimal Mode&lt;/th&gt;
&lt;th&gt;Standard Mode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool Context Footprint&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~971 tokens (&amp;lt;6% of 16K)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~6.9K tokens (~43% of 16K)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Inference Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~17 tok/s (TTFT 1.5s)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~12 tok/s (TTFT ~3.9s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Code Reasoning &amp;amp; Tooling&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Basic tools; fragile on file creation&lt;/td&gt;
&lt;td&gt;Rich tools (&lt;code&gt;read&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, &lt;code&gt;ask_user&lt;/code&gt;); deep reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bottleneck / Failure Mode&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Schema errors (&lt;code&gt;str_replace_editor&lt;/code&gt;), clunky workarounds&lt;/td&gt;
&lt;td&gt;Rapid context saturation, frequent auto-compactions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Neither extreme is optimal out of the box for an 8GB GPU. Minimal mode leaves ample context but lacks reliable tools for multi-file development. Standard mode provides capable tooling and deep reasoning, but its ~6.9K schema footprint consumes valuable context that is critical for multi-turn reasoning and complex problem solving.&lt;/p&gt;

&lt;p&gt;DeepSeek Harness avoids this rigid binary through its preset and profile system.&lt;/p&gt;

&lt;p&gt;Instead of staying locked into either built-in preset, developers can fine-tune a custom preset specifically for local hardware:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Strip out unnecessary tools that are irrelevant to local debugging tasks (such as web search, browser automation, subagent delegation, and telemetry).&lt;/li&gt;
&lt;li&gt;Retain the core filesystem tools that proved effective (&lt;code&gt;read&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, and interactive prompts), keeping tool schema overhead around ~2K tokens.&lt;/li&gt;
&lt;li&gt;Reserve 12K+ tokens for conversation history, file contents, and extended reasoning loops.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most importantly, because DSH manages capabilities via Cordis fibers, &lt;strong&gt;this preset tuning can be adjusted on the fly without even restarting the harness&lt;/strong&gt;. This dynamic modularity makes it possible to find the exact sweet spot for any hardware budget.&lt;/p&gt;

&lt;h3&gt;
  
  
  Alternative Interface: Creating a Dedicated TUI Profile with dsh-TUI
&lt;/h3&gt;

&lt;p&gt;Because DeepSeek Harness is a true meta-harness, its frontend interfaces are completely decoupled from its execution microkernel. You are never locked into a web browser.&lt;/p&gt;

&lt;p&gt;To see how DSH's &lt;strong&gt;profile&lt;/strong&gt; system works in practice (as we defined in Section 5), let's create a dedicated profile that replaces the web dashboard with an interactive Terminal User Interface (TUI).&lt;/p&gt;

&lt;p&gt;Among the growing number of community TUI plugins, one particularly popular choice in China is &lt;a href="https://dshtui.com/en/" rel="noopener noreferrer"&gt;dsh-TUI&lt;/a&gt; (&lt;code&gt;@deepseek-harness-tui/dsh-tui&lt;/code&gt;). It brings a full-screen, Claude Code-style terminal experience to DeepSeek Harness, complete with streamed markdown, context usage gauges, and TPS metrics.&lt;/p&gt;

&lt;p&gt;To configure and launch it, we declare a new profile named &lt;code&gt;dsh-tui&lt;/code&gt; and add the plugin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;dsh plugin --profile dsh-tui add @deepseek-harness-tui/dsh-tui

dsh --profile dsh-tui
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxk68mypj31lg6xof13h9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxk68mypj31lg6xof13h9.png" alt="dsh-TUI fullscreen terminal interface" width="800" height="435"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Because DSH treats the session as an append-only event stream, conversational state is completely independent of the display layer. If your terminal session disconnects or you want to pick up a previous task later, you can resume it directly by its session ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Resume with the command below:
dsh-tui --resume 6661d456-8dd1-49ad-bdf0-27bca9609574
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This example illustrates the practical utility of DSH profiles: switching from the web interface to a terminal UI requires no dependency changes or environment reconfiguration, only launching under a different profile flag.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Best Practices for 8GB Local Agents
&lt;/h2&gt;

&lt;p&gt;Working with constrained hardware requires disciplined operational habits:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Monitor VRAM Residency and Address Latency
&lt;/h3&gt;

&lt;p&gt;Keep an eye on GPU memory during multi-step runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;nvidia-smi &lt;span class="nt"&gt;-l&lt;/span&gt; 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If memory consumption exceeds 8GB, Windows will automatically page memory into shared system RAM. If generation speed suddenly drops from 15 tokens/sec to 2-3 tokens/sec, your KV cache has spilled onto the CPU bus. &lt;/p&gt;

&lt;p&gt;Even when fully GPU-resident, our setup averaged a TTFT of 1.5 to 3.9 seconds and 12 to 17 tokens/sec across our test runs. If you need snappier generation for interactive workflows, consider two optimizations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning Ollama parameters&lt;/strong&gt;: Enabling Flash Attention (&lt;code&gt;OLLAMA_FLASH_ATTENTION=1&lt;/code&gt;) or adjusting &lt;code&gt;num_batch&lt;/code&gt; in your Modelfile can reduce memory bandwidth bottlenecks during prompt evaluation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Selecting a smaller model&lt;/strong&gt;: Stepping down from a 9B model to a dedicated 7B model (such as Qwen2.5-Coder 7B) or an edge-focused architecture (such as Gemma 4 e2b) will significantly boost tokens per second while cutting prompt evaluation latency in half.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Fine-Tune the Tool Surface: Avoid Extremes
&lt;/h3&gt;

&lt;p&gt;Avoid attaching extraneous MCP servers or plugins to local sessions. Every additional tool schema consumes context tokens and increases the probability of hallucinated arguments. At the same time, our Section 7 experiments proved that an off-the-shelf minimal preset with just bash and string replacement is too fragile for file creation, while the default Standard preset consumes ~6.9K tokens. The practical strategy is to fine-tune your own preset: preserve reliable file tools (&lt;code&gt;read&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;) while pruning unused schemas (web search, subagents, telemetry) to keep tool overhead around ~2K tokens.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Use Hybrid Fallback Strategically
&lt;/h3&gt;

&lt;p&gt;Because DSH defines multiple providers in &lt;code&gt;settings.yaml&lt;/code&gt;, you are never locked into local inference. If you encounter a complex architectural refactor that exceeds the reasoning capabilities of a 9B model, you can switch the active model in the DSH web interface to OpenRouter (such as DeepSeek-V3 or Claude) for that specific prompt, and then switch back to Ollama for the implementation work.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Implement Automated Routing via LiteLLM
&lt;/h3&gt;

&lt;p&gt;Switching models manually in the UI works for occasional tasks, but autonomous agents run best when routing happens automatically. &lt;/p&gt;

&lt;p&gt;In a companion article, &lt;a href="https://dev.to/worldlinetech/building-a-local-first-ai-coding-agent-with-open-tools-and-adaptive-routing-iin"&gt;Building a Local-First AI Coding Agent with Open Tools and Adaptive Routing&lt;/a&gt;, we explored how to insert &lt;strong&gt;LiteLLM&lt;/strong&gt; as an intelligent control plane (Layer 2) between the harness and the inference servers. &lt;/p&gt;

&lt;p&gt;You can apply the exact same architecture to DeepSeek Harness:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Configure DSH's provider in &lt;code&gt;settings.yaml&lt;/code&gt; to point to a local LiteLLM proxy (&lt;code&gt;http://127.0.0.1:4000/v1&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Let LiteLLM inspect incoming requests: routine tool executions and file searches route to local Ollama at zero cost, while intricate architectural planning prompts automatically escalate to OpenRouter under strict per-session spend limits.&lt;/li&gt;
&lt;li&gt;This combines zero-cost local execution for routine loops with automated escalation to cloud models for complex reasoning.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  9. Series Retrospective: From Meta-Framework to Local Agent
&lt;/h2&gt;

&lt;p&gt;Over the course of this three-part series, we have traced the full lifecycle of modern agent infrastructure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Part 1: Understanding Cordis&lt;/a&gt;&lt;/strong&gt; explored the foundational microkernel created by Shigma. We saw how spatiotemporal composability, reverse cleanup stacks, and reactive contexts solve the chronic memory leaks of long-running JavaScript applications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;Part 2: DeepSeek Harness Architecture&lt;/a&gt;&lt;/strong&gt; analyzed how DeepSeek turned Cordis into an operating system for AI agents. We untangled the harness terminology, contrasted DSH with turnkey tools like OpenCode and Pi, and examined its "everything-is-a-plugin" philosophy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3 (This Article)&lt;/strong&gt; brought everything to life on consumer hardware. By pairing DSH's minimal preset with a context-tuned Ollama model, we proved that you do not need enterprise data centers or expensive cloud API budgets to run a functional, autonomous coding assistant.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;DeepSeek Harness shows how modular systems engineering benefits local development. By separating runtime lifecycles, provider adapters, and tool presets, DSH allows developers to match agent overhead to physical hardware limits, making an 8GB GPU a viable environment for local pairing.&lt;/p&gt;




&lt;h3&gt;
  
  
  References &amp;amp; Links
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Harness GitHub Repository&lt;/strong&gt;: &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;github.com/deepseek-ai/deepseek-harness&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Harness Presets Guide&lt;/strong&gt;: &lt;a href="https://deepseekdsh.com/guides/modes" rel="noopener noreferrer"&gt;deepseekdsh.com/guides/modes&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;dsh-TUI Terminal Interface&lt;/strong&gt;: &lt;a href="https://dshtui.com/en/" rel="noopener noreferrer"&gt;dshtui.com/en&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Community Plugin Directories&lt;/strong&gt;: &lt;a href="https://deepseekplugin.com/" rel="noopener noreferrer"&gt;deepseekplugin.com&lt;/a&gt; | &lt;a href="https://dsh-plugins.org/en" rel="noopener noreferrer"&gt;dsh-plugins.org&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ollama Model Library&lt;/strong&gt;: &lt;a href="https://ollama.com/library" rel="noopener noreferrer"&gt;ollama.com/library&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pi Coding Agent &amp;amp; Engine&lt;/strong&gt;: &lt;a href="https://pi.dev" rel="noopener noreferrer"&gt;pi.dev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 1 of this Series&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Understanding Cordis: The TypeScript Framework Built for Hot-Swapping Everything&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 2 of this Series&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;DeepSeek Harness: How DeepSeek Uses Cordis to Redefine Autonomous AI Agents&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Companion Article (Control Plane &amp;amp; Routing)&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/building-a-local-first-ai-coding-agent-with-open-tools-and-adaptive-routing-iin"&gt;Building a Local-First AI Coding Agent with Open Tools and Adaptive Routing&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>deepseek</category>
      <category>ollama</category>
      <category>ornith</category>
      <category>agents</category>
    </item>
    <item>
      <title>DeepSeek Harness: How DeepSeek Uses Cordis to Redefine Autonomous AI Agents</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Sat, 12 Sep 2026 05:47:32 +0000</pubDate>
      <link>https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599</link>
      <guid>https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 2 of the DeepSeek Harness: Kernel to Edge series: Exploring DeepSeek Harness (DSH), its 'everything-is-a-plugin' architecture, how it leverages Cordis under the hood, and how it compares to OpenCode, Pi, LangGraph, and CrewAI.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek Harness: Kernel to Edge&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
This article is &lt;strong&gt;Part 2&lt;/strong&gt; of a three-part architectural and practical deep dive:  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Part 1&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Understanding Cordis: The TypeScript Framework Built for Hot-Swapping Everything&lt;/a&gt; (Microkernel primitives, spatiotemporal composability, reverse cleanup stacks, zero-leak lifecycles).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 2 (This Article)&lt;/strong&gt;: DeepSeek Harness: How DeepSeek Uses Cordis to Redefine Autonomous AI Agents (The meta-harness paradigm, Cordis as an agent kernel, comparing DSH to OpenCode and Pi).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i"&gt;Running DeepSeek Harness on an 8GB GPU: Context Tuning, Presets, and Hardware Limits&lt;/a&gt; (Hands-on local deployment with Ollama, Ornith-1.5, VRAM arithmetic, minimal vs. standard presets, and TUI).&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. Introduction: Beyond the Prompt Chain
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Part 1&lt;/a&gt;, we explored &lt;strong&gt;Cordis&lt;/strong&gt;, the lightweight TypeScript framework created by &lt;strong&gt;Shigma (Yifan Shi)&lt;/strong&gt; that guarantees zero-downtime hot reloading and clean plugin lifecycles through revertible effects.&lt;/p&gt;

&lt;p&gt;When developers first encountered Cordis, most viewed it primarily as the foundation for the Koishi chatbot ecosystem. In 2026, &lt;strong&gt;DeepSeek AI&lt;/strong&gt; adopted it as the foundational engine for &lt;strong&gt;DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;DeepSeek chose not to build their agent framework in Python, avoiding both generic prompt-chaining libraries and opinionated role-playing engines.&lt;/p&gt;

&lt;p&gt;Instead, &lt;strong&gt;DeepSeek hired Shigma&lt;/strong&gt;, used Cordis as their microkernel, and published research grounding the approach (&lt;em&gt;"A Programming Paradigm for Spatiotemporal Composability"&lt;/em&gt;, &lt;a href="https://arxiv.org/pdf/2608.25512" rel="noopener noreferrer"&gt;arXiv:2608.25512&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The result is &lt;strong&gt;DeepSeek Harness (DSH)&lt;/strong&gt;: an open-source agent framework built on a distinct design rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"There is no privileged core. Every capability (the model adapter, the sandbox, the tools, the session log, and even the agent loop itself) is a replaceable plugin."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2pllqw8az4uaii4bpxdz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2pllqw8az4uaii4bpxdz.png" alt="Traditional Monolithic Agent Core vs. DeepSeek Harness Cordis Microkernel" width="800" height="371"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This article examines how DeepSeek Harness uses Cordis under the hood, how its architecture operates, and how it compares to other agent frameworks.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Why Call It a "Harness"? (Untangling the Terminology)
&lt;/h2&gt;

&lt;p&gt;In the software industry, the word &lt;strong&gt;"harness"&lt;/strong&gt; has become one of the most overloaded buzzwords in AI. Depending on who you talk to, a "harness" can mean a benchmark runner, a CLI chatbot, or a complex multi-agent framework.&lt;/p&gt;

&lt;p&gt;To understand DeepSeek Harness, we need to untangle the &lt;strong&gt;three distinct categories of harnesses&lt;/strong&gt; operating in the AI ecosystem today:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9cnernxupec3gwf6vmjs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9cnernxupec3gwf6vmjs.png" alt="The Three Categories of AI Harnesses" width="799" height="263"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Category 1: Evaluation Harnesses (Benchmark Rigs)&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Examples&lt;/em&gt;: SWE-bench test runners, HumanEval fixtures.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Purpose&lt;/em&gt;: A rigid testing scaffold designed to feed an issue to a model, run a test suite in a Docker container, and report a pass/fail grade. They are test rigs, not production application runtimes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Category 2: Application / Coding Harnesses (Interactive Assistants)&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Examples&lt;/em&gt;: OpenCode, Pi Coding Agent, Claude Code.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Purpose&lt;/em&gt;: Specialized developer tools designed for interactive, human-in-the-loop pair programming in a terminal or IDE. They come pre-configured with a fixed set of developer tools (&lt;code&gt;read&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, &lt;code&gt;bash&lt;/code&gt;) running directly on your machine.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Category 3: Meta-Harnesses / Agent Operating Systems&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Example&lt;/em&gt;: &lt;strong&gt;DeepSeek Harness (DSH)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Purpose&lt;/em&gt;: An un-opinionated, headless microkernel. DSH does not presuppose what the agent is doing. It provides low-level operating system scaffolding (process isolation, append-only session streams, and reversible plugin lifecycles) upon which you can assemble an evaluation harness, an interactive coding assistant, or an autonomous research pipeline simply by changing a configuration profile.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Models like &lt;strong&gt;DeepSeek-R1&lt;/strong&gt; and &lt;strong&gt;DeepSeek-V3&lt;/strong&gt; require substantial host-level support to perform multi-step work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Execution sandboxes&lt;/strong&gt;: Docker containers, child processes, and isolated filesystems.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;External interfaces&lt;/strong&gt;: Web search, documentation scrapers, and external APIs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session persistence&lt;/strong&gt;: Deterministic logging that records execution history without state drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety policies&lt;/strong&gt;: Interceptors that validate commands before execution.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most agent frameworks treat these capabilities as peripheral utilities. &lt;strong&gt;DeepSeek Harness treats the harness as an operating system kernel&lt;/strong&gt;, providing the structural scaffolding required to run models autonomously without orphaned processes, memory leaks, or execution escapes.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Why DeepSeek Adopted Cordis
&lt;/h2&gt;

&lt;p&gt;Why did DeepSeek choose Cordis over existing Python frameworks?&lt;/p&gt;

&lt;h3&gt;
  
  
  Limitations of Python Agent Frameworks
&lt;/h3&gt;

&lt;p&gt;Most popular agent tools (LangChain, AutoGen, CrewAI) are written in Python. While Python is the standard for ML training and tensor operations, it introduces architectural friction when building long-lived agent runtimes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Monotonic Memory&lt;/strong&gt;: Python's garbage collector and global module cache make dynamic unloading difficult. If an agent spins up a background process or attaches an event listener, cleaning it up cleanly requires manual bookkeeping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heavy Startup Overhead&lt;/strong&gt;: Python agent libraries often take seconds just to import their dependency trees.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rigid Graphs&lt;/strong&gt;: Tools like LangGraph model workflows as static directed graphs. If an agent discovers mid-flight that it needs a new tool, adding that tool dynamically usually requires recompiling the graph or restarting the agent.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  How Cordis Addresses These Issues
&lt;/h3&gt;

&lt;p&gt;Building on Cordis provides three concrete architectural advantages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Fast Startup and Low Footprint&lt;/strong&gt;: Running on Node.js and TypeScript, DSH boots in milliseconds with a lean baseline memory footprint (&amp;lt;50MB).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guaranteed Teardown&lt;/strong&gt;: When an agent launches a Docker container or a file watcher, Cordis wraps it in a &lt;strong&gt;Fiber&lt;/strong&gt;. When the sub-task ends, Cordis guarantees teardown of all open handles and child processes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reactive Dependency Trees&lt;/strong&gt;: If a task requires a Python environment, the Python tool only appears in the model's schema once a healthy sandbox service is active.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  4. How DeepSeek Harness Works Under the Hood
&lt;/h2&gt;

&lt;p&gt;Let's look at the six core architectural concepts that power DeepSeek Harness.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb5sxgziaknt4h4ya1b63.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb5sxgziaknt4h4ya1b63.png" alt="DeepSeek Harness Cordis Service Map" width="800" height="273"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  1. "There is No Core": The Pure Plugin Architecture
&lt;/h3&gt;

&lt;p&gt;In DSH, you won't find a massive, hardcoded &lt;code&gt;Agent&lt;/code&gt; class. Instead, every major subsystem is a Cordis plugin that extends &lt;code&gt;Service&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ModelService&lt;/code&gt; (&lt;code&gt;ctx.model&lt;/code&gt;)&lt;/strong&gt;: Handles token streaming, schema formatting, and provider connections (DeepSeek API, OpenAI, Anthropic, or local Ollama).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;SandboxService&lt;/code&gt; (&lt;code&gt;ctx.sandbox&lt;/code&gt;)&lt;/strong&gt;: Provides safe execution boundaries (isolated child processes, Docker containers, or WASM sandboxes).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;SessionService&lt;/code&gt; (&lt;code&gt;ctx.session&lt;/code&gt;)&lt;/strong&gt;: Manages the conversation history and transcript logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ToolRegistry&lt;/code&gt; (&lt;code&gt;ctx.tools&lt;/code&gt;)&lt;/strong&gt;: Tracks available tools, validates schemas, and executes actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;AgentLoop&lt;/code&gt; (&lt;code&gt;ctx.agent&lt;/code&gt;)&lt;/strong&gt;: The turn-based execution cycle that prompts the model, receives tool calls, and handles responses.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because these are all Cordis plugins, you can swap any of them independently. For instance, you can replace the default agent loop with a custom planner while keeping the existing tools and sandboxes intact. Even the user interface itself is just a plugin: DSH has no privileged GUI or CLI built into its runtime core.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. Bundles and Profiles (&lt;code&gt;cordis.yml&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;DeepSeek Harness applications are composed declaratively using &lt;strong&gt;Bundles&lt;/strong&gt; and &lt;strong&gt;Profiles&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Bundle&lt;/strong&gt;: An npm package containing code, plugins, and schema configurations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Profile&lt;/strong&gt;: A YAML configuration file defining which bundles to load together.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here is what a standard &lt;code&gt;cordis.yml&lt;/code&gt; profile looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# cordis.yml - DeepSeek Harness Composition&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;autonomous-coder&lt;/span&gt;

&lt;span class="na"&gt;plugins&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="c1"&gt;# 1. Model Provider Plugin&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@deepseek-ai/dsh-model-deepseek'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;env:DEEPSEEK_API_KEY&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-coder-v2&lt;/span&gt;

  &lt;span class="c1"&gt;# 2. Execution Sandbox Plugin&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@deepseek-ai/dsh-sandbox-docker'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;node:22-alpine'&lt;/span&gt;
    &lt;span class="na"&gt;timeoutMs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;30000&lt;/span&gt;

  &lt;span class="c1"&gt;# 3. Development Tools&lt;/span&gt;
&lt;span class="err"&gt;  &lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@deepseek-ai/dsh-tool-filesystem'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;rootDirectory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;./workspace'&lt;/span&gt;

&lt;span class="err"&gt;  &lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@deepseek-ai/dsh-tool-bash'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;allowedCommands&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;git'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;npm'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;pnpm'&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;test'&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

  &lt;span class="c1"&gt;# 4. Web UI Dashboard&lt;/span&gt;
&lt;span class="err"&gt;  &lt;/span&gt;&lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;@deepseek-ai/dsh-web-console'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;port&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3000&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To run this agent stack with its web console:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  Entry Profiles and Presets: Web, Headless, and Minimal
&lt;/h4&gt;

&lt;p&gt;Because DSH is assembled from YAML profiles and runtime presets, switching the operational persona of your agent requires no code changes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;dsh web&lt;/code&gt; (or &lt;code&gt;dsh --profile web&lt;/code&gt;): Boots the development stack with a local browser-based dashboard.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;dsh --profile headless&lt;/code&gt;: Runs the agent in headless mode for CI/CD pipelines, automated testing, or background workers.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;agent-presets: default: minimal&lt;/code&gt;: Activates a stripped-down two-tool composition, reducing system prompt token overhead and conserving system memory and GPU VRAM.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  3. The Anatomy of a Custom DSH Plugin (15 Lines of TypeScript)
&lt;/h3&gt;

&lt;p&gt;Writing your own tool or capability for DeepSeek Harness requires minimal code. Because it is a native Cordis plugin, you declare what dependencies you need (&lt;code&gt;inject&lt;/code&gt;) and register your logic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cordis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;project-stats&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;inject&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;tools&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt; &lt;span class="c1"&gt;// Requires the ToolRegistry service&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;apply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Register the tool; Cordis handles lifecycle and cleanup&lt;/span&gt;
  &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;register&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;get_project_stats&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Returns file count and lines of code in current directory&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Total files: 42 | Lines of code: 5,120&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once written, you can declare it in your local &lt;code&gt;cordis.yml&lt;/code&gt; profile or load it dynamically at runtime.&lt;/p&gt;




&lt;h3&gt;
  
  
  4. Session as an Append-Only Stream
&lt;/h3&gt;

&lt;p&gt;Traditional frameworks treat chat history as a mutable array of &lt;code&gt;{ role, content }&lt;/code&gt; objects that developers manipulate directly.&lt;/p&gt;

&lt;p&gt;DeepSeek Harness treats the &lt;strong&gt;&lt;code&gt;Session&lt;/code&gt; as an immutable, append-only event stream&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every user input, thinking block, model token, tool execution, and error is an immutable record in the stream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projections&lt;/strong&gt;: The model's context window, the CLI output, and the Web UI are simply read-only "projections" rendered from this single source of truth.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time Travel &amp;amp; Forking&lt;/strong&gt;: Because the log is append-only, an agent can instantly fork a sub-agent to explore an alternative debugging hypothesis, test it, and discard it without corrupting the main conversation history.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  5. Reversible Sandboxes and Dynamic Toolsets
&lt;/h3&gt;

&lt;p&gt;In an extended autonomous task, an agent might need specialized capabilities only for a few minutes. &lt;/p&gt;

&lt;p&gt;For example, when asked to analyze a dataset, the agent might need a Python data science environment (&lt;code&gt;pandas&lt;/code&gt;, &lt;code&gt;numpy&lt;/code&gt;, &lt;code&gt;matplotlib&lt;/code&gt;). In traditional frameworks, that environment stays open forever.&lt;/p&gt;

&lt;p&gt;In DeepSeek Harness, tools are mounted as &lt;strong&gt;Cordis Fibers&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cordis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// A specialized tool plugin mounted dynamically&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;DataAnalysisToolPlugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Use ctx.effect to manage the sandbox lifecycle&lt;/span&gt;
  &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;effect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;[Tool] Booting ephemeral Docker container for Python...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;container&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sandbox&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;createContainer&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;python:3.11-slim&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="c1"&gt;// Register the tool&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;unregister&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;register&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;run_python_script&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Executes Python data analysis code&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;code&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;container&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="c1"&gt;// Cleanup: When the agent unloads this plugin, the container is destroyed!&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;[Tool] Tearing down ephemeral container...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nf"&gt;unregister&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;container&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;destroy&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the data analysis step is complete, the agent or harness calls &lt;code&gt;fiber.dispose()&lt;/code&gt;. Cordis instantly kills the container, frees the RAM, and removes the tool from the model's schema.&lt;/p&gt;




&lt;h3&gt;
  
  
  6. Waterfall Control Flow &amp;amp; Guardrails
&lt;/h3&gt;

&lt;p&gt;How does DeepSeek Harness prevent an autonomous agent from running destructive commands (like &lt;code&gt;rm -rf /&lt;/code&gt; or leaking API keys)?&lt;/p&gt;

&lt;p&gt;It uses Cordis's &lt;strong&gt;&lt;code&gt;waterfall&lt;/code&gt; event pattern&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4u0qq37qi33vjkvsf6ec.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4u0qq37qi33vjkvsf6ec.png" alt="DeepSeek Harness Waterfall Event Pipeline" width="800" height="293"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Any plugin can hook into &lt;code&gt;agent/pre-step&lt;/code&gt; or &lt;code&gt;agent/request&lt;/code&gt;. If a security plugin detects an unsafe command, it can rewrite it, request user confirmation, or bail out early before the sandbox is ever touched.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. The True Meta-Harness: An Exploding Plugin Ecosystem
&lt;/h2&gt;

&lt;p&gt;Why does DeepSeek Harness deserve the title of a true &lt;strong&gt;meta-harness&lt;/strong&gt;?&lt;/p&gt;

&lt;p&gt;Turnkey coding agents (such as OpenCode or Claude Code) ship as monolithic binaries with predetermined tools, fixed workflows, and rigid interfaces. If you want a different terminal layout, a custom vision pipeline, or specialized memory management, you have to submit a feature request or fork the codebase.&lt;/p&gt;

&lt;p&gt;In DSH, the harness is not a closed application; it is an open substrate. Because the microkernel has "no core" and treats every capability as an unprivileged Cordis plugin, developers can reshape every layer of the agent experience.&lt;/p&gt;

&lt;h3&gt;
  
  
  Even the User Interface is a Plugin (Web UI vs. TUIs)
&lt;/h3&gt;

&lt;p&gt;A textbook example of this architectural neutrality is the user interface itself:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Web UI is Just a Plugin&lt;/strong&gt;: When you run &lt;code&gt;dsh web&lt;/code&gt;, DSH boots an HTTP server and mounts a frontend plugin bundle (&lt;code&gt;@deepseek-ai/dsh-web-console&lt;/code&gt; or &lt;code&gt;dsh-web-ui&lt;/code&gt;). The harness core neither knows nor cares that a web browser is rendering the output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal UIs (TUIs) on Demand&lt;/strong&gt;: If you prefer staying strictly inside your terminal to avoid context switching, you can completely bypass the web dashboard. Community plugins like &lt;strong&gt;&lt;code&gt;dsh-TUI&lt;/code&gt;&lt;/strong&gt; transform DSH into a full-screen, keyboard-driven terminal coding agent reminiscent of Claude Code, complete with streaming token output and interactive approval prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workbench Enhancements&lt;/strong&gt;: For developers using the browser console, plugins like &lt;strong&gt;&lt;code&gt;DSH-better-sidebar&lt;/code&gt;&lt;/strong&gt; enhance the default layout by docking file trees, Git status, terminal sessions, and live browser previews directly alongside the chat stream.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Thousands of Plugins and Rising Community Marketplaces
&lt;/h3&gt;

&lt;p&gt;This native extensibility has catalyzed a rapidly expanding community ecosystem. In only a few months, thousands of open-source plugins have been authored by developers worldwide, giving rise to dedicated community directories and marketplaces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://deepseekplugin.com/" rel="noopener noreferrer"&gt;deepseekplugin.com&lt;/a&gt;&lt;/strong&gt;: A centralized directory cataloging community-built DSH plugins across diverse functional categories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://dsh-plugin.org/" rel="noopener noreferrer"&gt;dsh-plugin.org&lt;/a&gt; / &lt;a href="https://dsh-plugins.org/en" rel="noopener noreferrer"&gt;dsh-plugins.org&lt;/a&gt;&lt;/strong&gt;: Open documentation and discovery hubs for exploring plugin architectures, configuration snippets, and trending add-ons.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Real-World Plugin Innovations: Beyond Simple Coding Tools
&lt;/h3&gt;

&lt;p&gt;As analyzed in Composio's roundup of top DeepSeek Harness plugins, community developers are extending DSH far beyond basic file reading and bash execution:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Multimodal Vision Toolkits (&lt;code&gt;ModLens&lt;/code&gt;, &lt;code&gt;dsh-vision-toolkit&lt;/code&gt;)&lt;/strong&gt;: Give text-only models visual perception. These plugins provide structured OCR, visual grounding, layout analysis, UI reconstruction from screenshots, and pixel-by-pixel comparisons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generative UI (&lt;code&gt;dsh-genui&lt;/code&gt;)&lt;/strong&gt;: Allows the agent to render more than 30 interactive user interface widgets (cards, responsive data tables, forms, charts) directly inside conversational responses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ergonomic Workspace Mentions (&lt;code&gt;dsh-at-file&lt;/code&gt;)&lt;/strong&gt;: Brings Codex-style &lt;code&gt;@file&lt;/code&gt; and &lt;code&gt;@folder&lt;/code&gt; mentions into the composer, letting developers inject file contents and directory context without manual copy-pasting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Tiered Memory (&lt;code&gt;dsh-mnemon&lt;/code&gt;)&lt;/strong&gt;: Solves agent amnesia across turns by providing three distinct memory layers: Runtime Memory for turn-by-turn preferences, Document Memory for repository conventions, and long-term knowledge retention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In-Harness Discovery &amp;amp; Self-Installation (&lt;code&gt;dsh-market&lt;/code&gt;, &lt;code&gt;dsh-find-plugin&lt;/code&gt;)&lt;/strong&gt;: Brings the marketplace directly into DSH. &lt;code&gt;dsh-market&lt;/code&gt; embeds a visual plugin market into the Settings UI, while &lt;code&gt;dsh-find-plugin&lt;/code&gt; equips the agent itself with a tool to search GitHub for plugins on the fly and self-install capabilities during an active run.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  6. Positioning DSH: Comparing Harnesses Across the AI Landscape
&lt;/h2&gt;

&lt;p&gt;To truly understand where DeepSeek Harness fits, we have to look at the broader AI landscape. Developers often conflate two very different layers of software:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;High-Level Agent Frameworks&lt;/strong&gt; (LangGraph, CrewAI, AutoGen): Libraries designed for prompt chains, multi-agent debates, and graph state machines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dedicated Coding Agent Harnesses&lt;/strong&gt; (OpenCode, Pi Coding Agent, Claude Code): Specialized developer tools designed to run in your terminal or IDE to edit code.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;DeepSeek Harness bridges both worlds by operating as a &lt;strong&gt;Meta-Harness (an Agent Operating System)&lt;/strong&gt;. Let's break down how it compares to both categories.&lt;/p&gt;




&lt;h3&gt;
  
  
  Layer A: DSH vs. General Agent Frameworks (LangGraph, CrewAI)
&lt;/h3&gt;

&lt;p&gt;Most general agent frameworks are written in Python and focus on orchestrating conversations and workflow graphs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LangGraph (Circuits &amp;amp; Graphs)&lt;/strong&gt;: Models agents as fixed directed cyclic graphs (nodes, edges, and state channels). While powerful for predictable workflows, dynamic changes are difficult: adding a new tool or modifying execution mid-flight requires recompiling the graph or restarting the agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CrewAI / AutoGen (Personas &amp;amp; Roleplay)&lt;/strong&gt;: Models agents as personas with roles, backstories, and conversational message passing. Great for creative simulations, but they lack low-level systems control over processes, memory reclamation, and execution sandboxes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Harness (An Operating System Microkernel)&lt;/strong&gt;: Instead of modeling agents as circuits or personas, DSH models agents as an &lt;strong&gt;operating system&lt;/strong&gt;. The Cordis microkernel provides the core primitives (event bus, service registry, fibers). Models, tools, sandboxes, and the agent loop itself are drivers that plug into that OS.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Layer B: DSH vs. Dedicated Coding Harnesses (OpenCode, Pi)
&lt;/h3&gt;

&lt;p&gt;If you are a software engineer, you are likely more familiar with specialized coding agent harnesses like &lt;strong&gt;OpenCode&lt;/strong&gt; or &lt;strong&gt;Pi Coding Agent&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;OpenCode (The Model-Agnostic Developer Companion)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Developed by the SST team (&lt;code&gt;anomalyco/opencode&lt;/code&gt;) in TypeScript/Bun, OpenCode provides a polished developer experience with terminal, desktop, and web interfaces.&lt;/li&gt;
&lt;li&gt;It ships with a pre-tuned suite of developer tools (&lt;code&gt;read&lt;/code&gt;, &lt;code&gt;write&lt;/code&gt;, &lt;code&gt;edit&lt;/code&gt;, &lt;code&gt;glob&lt;/code&gt;, &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;bash&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where it excels&lt;/strong&gt;: Interactive, human-in-the-loop coding sessions on your local machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Architectural Difference&lt;/strong&gt;: OpenCode is a &lt;strong&gt;turnkey coding tool&lt;/strong&gt;. Its tool registry, session loop, and interaction models are tailored to developer workflows. Tools run directly on your host environment, and the toolset is static per session.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Pi Coding Agent (The Minimalist Terminal Agent)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Created by Mario Zechner, Pi (&lt;code&gt;@earendil-works/pi-coding-agent&lt;/code&gt;) is a lightweight terminal coding agent designed around minimalism and extensibility.&lt;/li&gt;
&lt;li&gt;It separates concerns into clean packages (unified LLM API, agent runtime, terminal UI).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Architectural Difference&lt;/strong&gt;: Pi is designed to do one thing well: serve as an extensible terminal coding companion for an individual engineer.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek Harness (The Reversible Meta-Harness)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DSH is &lt;strong&gt;not hardcoded as a coding assistant&lt;/strong&gt;. Coding tools are merely plugins.&lt;/li&gt;
&lt;li&gt;By stacking different YAML profiles (&lt;code&gt;cordis.yml&lt;/code&gt;), DSH can function as a terminal coding agent (similar to OpenCode or Pi), an automated benchmark runner, a web research crawler, or a security audit pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic Ephemeral Isolation&lt;/strong&gt;: Rather than running all commands directly on your host machine, DSH is built for &lt;strong&gt;unattended autonomous execution&lt;/strong&gt;. It can spin up an ephemeral container for a sub-task, inject tools into the agent context, execute the work, and tear down the environment without leaving orphaned processes or open ports.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Under the Hood: How Their Plugin Systems Actually Differ
&lt;/h3&gt;

&lt;p&gt;To see why DSH behaves differently in runtime management, compare how you extend each harness:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9c3jlqp8huc5fjkuhdz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9c3jlqp8huc5fjkuhdz.png" alt="Plugin Mechanism Comparison: Pi vs. OpenCode vs. DeepSeek Harness" width="800" height="244"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Pi Coding Agent (Boot-Time Extension Hooks)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Mechanism&lt;/em&gt;: You export a function or object that registers custom tools or commands into Pi's registry at startup.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Lifecycle&lt;/em&gt;: &lt;strong&gt;Static and persistent.&lt;/strong&gt; Extensions live for the entire duration of the CLI session.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Best for&lt;/em&gt;: Personal CLI shortcuts and custom prompt extensions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;OpenCode (Config-Driven Tools &amp;amp; External MCP Servers)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Mechanism&lt;/em&gt;: Tools are declared in configuration or connected via the &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Lifecycle&lt;/em&gt;: &lt;strong&gt;Process-isolated client-server architecture.&lt;/strong&gt; Tools typically run as external background programs communicating over stdio or HTTP.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Best for&lt;/em&gt;: Reusing standardized enterprise tools across multiple IDEs and agents without rewriting tool code.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek Harness (In-Process, Reactive Cordis Fibers)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Mechanism&lt;/em&gt;: Plugins mount directly into the in-process &lt;strong&gt;Context Tree&lt;/strong&gt; via Cordis. They provide core services (&lt;code&gt;ctx.provide&lt;/code&gt;), declare dependencies (&lt;code&gt;inject&lt;/code&gt;), and register self-cleaning side effects (&lt;code&gt;ctx.effect&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Lifecycle&lt;/em&gt;: &lt;strong&gt;Dynamic, reactive, and reversible.&lt;/strong&gt; A plugin can be hot-swapped, suspended, or unloaded during execution. Cordis guarantees that any socket, timer, or container created by that plugin is destroyed upon disposal.&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Best for&lt;/em&gt;: Autonomous agents that dynamically adapt their environment, spin up temporary sandboxes, or synthesize new tools on the fly.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  The Comprehensive Architectural Matrix
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;General Frameworks (LangGraph, CrewAI)&lt;/th&gt;
&lt;th&gt;Turnkey Coding Harnesses (OpenCode, Pi)&lt;/th&gt;
&lt;th&gt;Meta-Harness (DeepSeek Harness)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Primary Goal&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Multi-agent chains &amp;amp; graph workflows&lt;/td&gt;
&lt;td&gt;Interactive developer coding assistant&lt;/td&gt;
&lt;td&gt;Headless, autonomous agent operating system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Language &amp;amp; Runtime&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Python (Heavy import overhead)&lt;/td&gt;
&lt;td&gt;TypeScript / Bun (Fast, developer-friendly)&lt;/td&gt;
&lt;td&gt;TypeScript / Node.js (Lightweight &amp;lt;50ms)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Architecture&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Static Graphs or Multi-Agent Chats&lt;/td&gt;
&lt;td&gt;Monolithic / Modular Coding Loop&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Pure Cordis Microkernel (Zero privileged core)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Plugin Mechanism&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hardcoded in Python state schema&lt;/td&gt;
&lt;td&gt;Boot-time extension hooks / external MCP&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;In-process reactive Cordis Fibers (with auto-cleanup)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Execution Sandboxing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Manual script execution / External Docker&lt;/td&gt;
&lt;td&gt;Direct execution on developer's host machine&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Built-in Reversible Fibers (Docker, WASM, Process)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Session Model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Chat history lists / state dictionaries&lt;/td&gt;
&lt;td&gt;Interactive session transcripts&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Immutable append-only stream with projections&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Dynamic Self-Evolution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;❌ No (requires graph rebuild)&lt;/td&gt;
&lt;td&gt;❌ No (fixed tool schema)&lt;/td&gt;
&lt;td&gt;** Yes ("Creation Mode" via runtime plugins)**&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  In Short:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Use &lt;strong&gt;LangGraph or CrewAI&lt;/strong&gt; when you want to design a multi-persona conversational workflow or a fixed graph pipeline in Python.&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;OpenCode or Pi&lt;/strong&gt; when you want a fast, interactive AI pair programmer in your terminal to help you build software.&lt;/li&gt;
&lt;li&gt;Use &lt;strong&gt;DeepSeek Harness&lt;/strong&gt; when you need an &lt;strong&gt;unattended agent runtime&lt;/strong&gt; that dynamically manages tools, isolates sandboxes, runs extended tasks without resource leaks, and supports runtime plugin extension.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  7. Dynamic Tool Generation at Runtime ("Creation Mode")
&lt;/h2&gt;

&lt;p&gt;One notable capability enabled by Cordis is dynamic plugin generation at runtime, often called &lt;strong&gt;"Creation Mode"&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Because Cordis allows modules to be loaded and unloaded at runtime, an agent can dynamically extend its own abilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Problem&lt;/strong&gt;: An agent is tasked with parsing a proprietary binary format. It currently has no tool for this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Code Generation&lt;/strong&gt;: The agent writes a custom TypeScript parser plugin and saves it to its local directory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic Mounting&lt;/strong&gt;: The agent uses Cordis to hot-load its newly written plugin into its own context:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;   &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;plugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;NewlyWrittenPlugin&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Execution &amp;amp; Verification&lt;/strong&gt;: The agent now sees the new tool in its schema, executes it, and inspects the result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resolution&lt;/strong&gt;: If the tool is flawed, it fixes the code and hot-reloads it. Once the task is solved, it can either commit the plugin to its permanent bundle or dispose of it cleanly.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Instead of relying solely on a fixed set of pre-bundled tools, the agent can synthesize, test, and safely tear down its own execution utilities as new requirements arise during a task.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Conclusion
&lt;/h2&gt;

&lt;p&gt;By building on Cordis, DeepSeek Harness avoids the resource leaks and rigid orchestration common in traditional agent runtimes. Grounded in both practical experience from the Koishi ecosystem and research into spatiotemporal composability, DSH demonstrates a disciplined foundation for long-running autonomous tasks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Modular components&lt;/strong&gt;: There is no privileged core; tools, sandboxes, and agent loops can be swapped independently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reversible side effects&lt;/strong&gt;: Sandboxes and resource handles are tracked in fibers and cleaned up automatically upon disposal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low operational footprint&lt;/strong&gt;: Fast startup time with minimal idle memory consumption.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Coming Next in Part 3: Running DeepSeek Harness on an 8GB GPU
&lt;/h3&gt;

&lt;p&gt;With Cordis covered in &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Part 1&lt;/a&gt; and DeepSeek Harness's architecture mapped here in Part 2, the next installment moves to practical implementation.&lt;/p&gt;

&lt;p&gt;In &lt;strong&gt;&lt;a href="https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i"&gt;Part 3: Running DeepSeek Harness on an 8GB GPU&lt;/a&gt;&lt;/strong&gt;, we configure an autonomous coding agent running locally on a laptop equipped with an &lt;strong&gt;8GB RTX 4070 GPU&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Minimal Preset&lt;/strong&gt;: Configuring DSH with the &lt;code&gt;minimal&lt;/code&gt; agent preset to streamline tool schemas, compress prompt overhead, and maximize available context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local Inference&lt;/strong&gt;: Connecting DSH to Ollama running quantized coding models such as Qwen2.5-Coder or Ornith-1.5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VRAM and Context Tuning&lt;/strong&gt;: Setting context limits and parameter budgets so the agent can inspect and edit multi-file projects without running out of GPU memory or truncating context.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  References &amp;amp; Links
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Harness GitHub Repository&lt;/strong&gt;: &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;github.com/deepseek-ai/deepseek-harness&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Harness Documentation &amp;amp; Cordis Primer&lt;/strong&gt;: &lt;a href="https://deepseek-harness.github.io/deepseek-harness/en/reference/cordis-primer" rel="noopener noreferrer"&gt;deepseek-harness.github.io/deepseek-harness&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Community Plugin Directories&lt;/strong&gt;: &lt;a href="https://deepseekplugin.com/" rel="noopener noreferrer"&gt;deepseekplugin.com&lt;/a&gt; | &lt;a href="https://dsh-plugins.org/en" rel="noopener noreferrer"&gt;dsh-plugins.org&lt;/a&gt; | &lt;a href="https://dsh-plugin.org/" rel="noopener noreferrer"&gt;dsh-plugin.org&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Composio Best DSH Plugins Guide&lt;/strong&gt;: &lt;a href="https://composio.dev/content/best-deepseek-harness-plugins" rel="noopener noreferrer"&gt;composio.dev/content/best-deepseek-harness-plugins&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Theoretical Research Paper&lt;/strong&gt;: Shi, Y., Zhang, W., &amp;amp; Cui, T. &lt;a href="https://arxiv.org/pdf/2608.25512" rel="noopener noreferrer"&gt;&lt;em&gt;A Programming Paradigm for Spatiotemporal Composability&lt;/em&gt; (arXiv:2608.25512)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenCode Repository&lt;/strong&gt;: &lt;a href="https://github.com/anomalyco/opencode" rel="noopener noreferrer"&gt;github.com/anomalyco/opencode&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pi Coding Agent&lt;/strong&gt;: &lt;a href="https://pi.dev" rel="noopener noreferrer"&gt;pi.dev&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 1 of this Series&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb"&gt;Understanding Cordis: The TypeScript Framework Built for Hot-Swapping Everything&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3 of this Series&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i"&gt;Running DeepSeek Harness on an 8GB GPU: Context Tuning, Presets, and Hardware Limits&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>deepseek</category>
      <category>agentic</category>
      <category>cordis</category>
      <category>harness</category>
    </item>
    <item>
      <title>Understanding Cordis: The TypeScript Framework Built for Hot-Swapping Everything</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Sat, 12 Sep 2026 05:47:25 +0000</pubDate>
      <link>https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb</link>
      <guid>https://dev.to/worldlinetech/understanding-cordis-the-typescript-framework-built-for-hot-swapping-everything-1ihb</guid>
      <description>&lt;p&gt;&lt;em&gt;Part 1 of the DeepSeek Harness: Kernel to Edge series: How the open-source (MIT) Cordis meta-framework enables zero-downtime plugin reloads and memory-leak-free architectures in TypeScript, backed by 5+ years of battle-testing in Koishi.&lt;/em&gt;&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;DeepSeek Harness: Kernel to Edge&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
This article is &lt;strong&gt;Part 1&lt;/strong&gt; of a three-part architectural and practical deep dive:  &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Part 1 (This Article)&lt;/strong&gt;: Understanding Cordis: The TypeScript Framework Built for Hot-Swapping Everything (Microkernel primitives, spatiotemporal composability, reverse cleanup stacks, zero-leak lifecycles).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 2&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;DeepSeek Harness: How DeepSeek Uses Cordis to Redefine Autonomous AI Agents&lt;/a&gt; (The meta-harness paradigm, Cordis as an agent kernel, comparing DSH to OpenCode and Pi).
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i"&gt;Running DeepSeek Harness on an 8GB GPU: Context Tuning, Presets, and Hardware Limits&lt;/a&gt; (Hands-on local deployment with Ollama, Ornith-1.5, VRAM arithmetic, minimal vs. standard presets, and TUI).&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Series Foreword: The Architecture Behind DeepSeek Harness
&lt;/h2&gt;

&lt;p&gt;Developers inspecting the codebase of &lt;strong&gt;DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;)&lt;/strong&gt; often expect to find Python, LangChain, or prompt pipelines. Instead, they find pure TypeScript built on top of &lt;strong&gt;Cordis&lt;/strong&gt;, an engine originally developed for the Koishi chatbot ecosystem.&lt;/p&gt;

&lt;p&gt;This design choice addresses a concrete systems problem: autonomous coding agents rarely fail because of prompt wording; they fail because of &lt;strong&gt;runtime lifecycle issues&lt;/strong&gt;. Over dozens of unattended execution turns, an agent launches compiler runs, mounts temporary sandboxes, registers dynamic tool schemas, and manages file descriptors. In monolithic architectures, these operations accumulate leaked event listeners, orphan child processes, and memory bloat.&lt;/p&gt;

&lt;p&gt;DeepSeek treated autonomous agency as an &lt;strong&gt;operating system lifecycle problem&lt;/strong&gt;. Running an agent reliably over long sessions, including on resource-constrained consumer GPUs, requires strict mathematical reversibility and clean teardowns. That is what Cordis provides.&lt;/p&gt;

&lt;p&gt;Understanding DeepSeek Harness begins not with LLM prompts, but with the microkernel managing its execution environment.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Introduction: What is Cordis?
&lt;/h2&gt;

&lt;p&gt;Most backend frameworks (such as Express, NestJS, Fastify, or Koa) share an unspoken assumption: &lt;strong&gt;your application starts once, runs statically, and shuts down only when the process exits.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In those frameworks, you configure your database, register your routes, attach your middleware, and boot up. If you want to change a plugin, add a new route dynamically, or upgrade a module, your only real option is to restart the entire Node.js process.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cordis&lt;/strong&gt; is built for a completely different world.&lt;/p&gt;

&lt;p&gt;Cordis is an open-source (&lt;strong&gt;MIT licensed&lt;/strong&gt;) TypeScript &lt;strong&gt;meta-framework&lt;/strong&gt; (a framework designed to build other modular frameworks). Its core capability is &lt;strong&gt;runtime dynamic composability&lt;/strong&gt;: it allows you to load, configure, update, and unload plugins on the fly inside a running application &lt;strong&gt;with zero memory leaks and zero process restarts&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjwt1ld8alpcp12mhs4wn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjwt1ld8alpcp12mhs4wn.png" alt="Traditional Process Lifecycle vs. Cordis Dynamic Composability" width="800" height="273"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can install the core library into any modern Node.js or TypeScript project in seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install &lt;/span&gt;cordis
&lt;span class="c"&gt;# or&lt;/span&gt;
pnpm add cordis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Where Did Cordis Come From?
&lt;/h3&gt;

&lt;p&gt;To understand Cordis, you have to look at where it was born:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Koishi Chatbot Ecosystem (2019–2022)&lt;/strong&gt;:
Cordis was created by developer &lt;strong&gt;Shigma (Yifan Shi)&lt;/strong&gt; and the team behind &lt;strong&gt;Koishi&lt;/strong&gt;, a multi-platform chatbot framework. In production, a chatbot maintains persistent, long-lived WebSocket connections to Discord, Telegram, and other platforms. Dropping the connection just to install or update a plugin was a terrible user experience. Koishi needed a plugin engine that could install, reconfigure, and tear down plugins dynamically from a Web UI without touching the main connection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;From Chatbot Core to Meta-Framework (2022–2024)&lt;/strong&gt;:
The team realized this dynamic plugin model was not specific to chatbots; it solved a universal problem in Node.js. They extracted the kernel into &lt;strong&gt;Cordis&lt;/strong&gt; (from the Latin word &lt;em&gt;cor&lt;/em&gt;, meaning &lt;em&gt;heart&lt;/em&gt;). Far from an untested theoretical experiment, Cordis was &lt;strong&gt;battle-tested for over five years in production across hundreds of community plugins and thousands of live bot instances&lt;/strong&gt; in the Koishi ecosystem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Academic Roots &amp;amp; The AI Frontier (2025–2026+)&lt;/strong&gt;:
Today, &lt;strong&gt;Cordis Version 4&lt;/strong&gt; represents the culmination of those years of real-world usage and operational experience. Its creator, &lt;strong&gt;Shigma (Yifan Shi)&lt;/strong&gt;, is a university professor and computer science researcher. Recognizing the necessity of clean runtime composability for autonomous AI, &lt;strong&gt;DeepSeek brought Shigma onto their team&lt;/strong&gt; to architect the plugin engine for &lt;strong&gt;DeepSeek Harness (DSH)&lt;/strong&gt;. Together with researchers from DeepSeek and Peking University, they published a formal research paper (&lt;a href="https://arxiv.org/pdf/2608.25512" rel="noopener noreferrer"&gt;&lt;em&gt;"A Programming Paradigm for Spatiotemporal Composability"&lt;/em&gt;, arXiv:2608.25512&lt;/a&gt;), proving mathematically that dynamic systems can safely load, reload, and unload features indefinitely if side effects and dependencies are modeled as strict mathematical inverses.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This three-part series traces that transition from fundamental principles to real-world code: starting with Cordis's core lifecycle primitives in this article, exploring DeepSeek Harness's agent architecture in Part 2, and concluding with a hands-on local deployment on an 8GB GPU in Part 3.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The Big Problem: Why Plugins Usually Leak
&lt;/h2&gt;

&lt;p&gt;To appreciate what Cordis does under the hood, we first need to understand why dynamic plugins in Node.js are notoriously difficult to build.&lt;/p&gt;

&lt;h3&gt;
  
  
  The "Hotel Room" Analogy
&lt;/h3&gt;

&lt;p&gt;Imagine a hotel room:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When a guest arrives (&lt;strong&gt;plugin loaded&lt;/strong&gt;), they turn on the lights, turn on the bathroom faucet, turn up the heat, and plug in their devices (&lt;strong&gt;side effects&lt;/strong&gt;: event listeners, timers, database connections, HTTP routes).&lt;/li&gt;
&lt;li&gt;In traditional JavaScript frameworks, when the guest checks out (&lt;strong&gt;plugin unloaded&lt;/strong&gt;), they walk out the door, but the faucet is still running, the lights are still on, and the heater is still pumping heat.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2czlt2iv6yyofb96inax.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2czlt2iv6yyofb96inax.png" alt="Why Traditional Plugin Unloading Fails: The Disposal Abyss" width="800" height="244"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In Node.js, if you register an event listener on a shared emitter (&lt;code&gt;bus.on('message', callback)&lt;/code&gt;), the emitter holds a reference to your callback in memory. Even if you "delete" your plugin, that callback remains in memory. Worse, if you used &lt;code&gt;setInterval()&lt;/code&gt;, that timer handle keeps the entire Node.js event loop alive forever.&lt;/p&gt;

&lt;p&gt;Over time, reloading plugins in a standard Node.js app causes &lt;strong&gt;the Disposal Abyss&lt;/strong&gt;: zombie event listeners, duplicated handler executions, hanging sockets, and inevitable out-of-memory crashes.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. The Core Idea: Reversibility ("Every Action Has an Undo")
&lt;/h2&gt;

&lt;p&gt;Cordis solves this problem with a simple, powerful philosophy: &lt;strong&gt;Every side effect must be reversible by default.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of trusting the plugin developer to manually write a complex cleanup function, Cordis manages side effects through an &lt;strong&gt;inversion of control&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;When a plugin runs, Cordis gives it a specialized, isolated environment called a &lt;strong&gt;Context&lt;/strong&gt; (&lt;code&gt;ctx&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Whenever the plugin does something that affects the outside world (like listening to an event, setting a timer, or registering a service), it does so through &lt;code&gt;ctx&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Cordis quietly records the exact "undo" operation into a private &lt;strong&gt;cleanup stack&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;When the plugin is unloaded, Cordis automatically walks that stack in reverse order and undoes every single side effect.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fewsdf2sf01cpfbh2vlf3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fewsdf2sf01cpfbh2vlf3.png" alt="Cordis Reverse Cleanup Stack" width="800" height="273"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the heart of what Cordis calls &lt;strong&gt;Spatiotemporal Composability&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Temporal (Time)&lt;/strong&gt;: You can move forward (load) and backward (unload) in time without leaving leftover state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spatial (Space)&lt;/strong&gt;: Modules can safely coexist and discover each other's features without hardcoded, fragile bindings.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Grounded in Research&lt;/strong&gt;: This isn't just an informal software trick. In their &lt;a href="https://arxiv.org/pdf/2608.25512" rel="noopener noreferrer"&gt;research paper&lt;/a&gt;, Shigma and his co-researchers proved mathematically that as long as every state transformation carries a computable inverse (its cleanup function), an application can run indefinitely and roll back any module without corrupting its environment or needing a process restart.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  4. How Cordis Works Under the Hood
&lt;/h2&gt;

&lt;p&gt;Under the hood, Cordis relies on four core building blocks: &lt;strong&gt;Context&lt;/strong&gt;, &lt;strong&gt;Fibers&lt;/strong&gt;, &lt;strong&gt;Services&lt;/strong&gt;, and &lt;strong&gt;Events&lt;/strong&gt;. Let's examine each one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frdhcjqf6brlapel6frh1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frdhcjqf6brlapel6frh1.png" alt="The Cordis Architecture and Context Tree" width="800" height="293"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Concept 1: The Context Tree (&lt;code&gt;Context&lt;/code&gt; and &lt;code&gt;ctx&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;Context&lt;/code&gt; is the central object in Cordis. It represents the scope in which a component lives.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When you boot up your app, you create a &lt;strong&gt;Root Context&lt;/strong&gt;:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;  &lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cordis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;When you load a plugin using &lt;code&gt;app.plugin(MyPlugin)&lt;/code&gt;, Cordis &lt;strong&gt;does not&lt;/strong&gt; run the plugin directly on the root. Instead, it creates a &lt;strong&gt;child context&lt;/strong&gt; (a branch in the tree) specifically for that plugin.&lt;/li&gt;
&lt;li&gt;Under the hood, Cordis wraps contexts in a JavaScript &lt;strong&gt;&lt;code&gt;Proxy&lt;/code&gt;&lt;/strong&gt;. Whenever a plugin accesses a property (like &lt;code&gt;ctx.database&lt;/code&gt;), the Proxy dynamically checks if that service is visible to this branch, whether it's isolated, and tracks what the plugin is using.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  Concept 2: Reversible Effects (&lt;code&gt;ctx.effect&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;Whenever a plugin creates something that needs to be cleaned up later, it registers an &lt;strong&gt;Effect&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Think of &lt;code&gt;ctx.effect()&lt;/code&gt; like React's &lt;code&gt;useEffect()&lt;/code&gt;, but designed for backend servers and long-running services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;MyPlugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Register a reversible side effect&lt;/span&gt;
  &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;effect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Plugin activated: Setting up resources...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;timer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Heartbeat ping&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// Return the cleanup function!&lt;/span&gt;
    &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Plugin deactivating: Cleaning up timer...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If this plugin is unloaded, Cordis automatically calls the returned cleanup function. You don't have to keep track of timer IDs or remember to clean them up elsewhere.&lt;/p&gt;




&lt;h3&gt;
  
  
  Concept 3: Services (Shared Capabilities)
&lt;/h3&gt;

&lt;p&gt;In Cordis, a &lt;strong&gt;Service&lt;/strong&gt; is a reusable singleton feature provided to other plugins (such as a database client, an HTTP server, or a logger).&lt;/p&gt;

&lt;p&gt;Creating a service is as simple as extending the &lt;code&gt;Service&lt;/code&gt; class:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;Service&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cordis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// 1. Tell TypeScript that ctx.database exists (full autocomplete!)&lt;/span&gt;
&lt;span class="kr"&gt;declare&lt;/span&gt; &lt;span class="kr"&gt;module&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cordis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kr"&gt;interface&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nl"&gt;database&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;DatabaseService&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// 2. Define the service&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;DatabaseService&lt;/span&gt; &lt;span class="kd"&gt;extends&lt;/span&gt; &lt;span class="nc"&gt;Service&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// The name 'database' matches the property on Context&lt;/span&gt;
    &lt;span class="k"&gt;super&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;database&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nf"&gt;getUser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Alice&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the &lt;code&gt;declare module 'cordis'&lt;/code&gt; block. Cordis leverages TypeScript's &lt;strong&gt;Declaration Merging&lt;/strong&gt;. You don't need magic decorators, string tokens, or complex dependency injection containers. Once declared, &lt;code&gt;ctx.database&lt;/code&gt; has 100% full type-safety and auto-completion across your entire codebase.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Tip for TypeScript Users&lt;/strong&gt;: Make sure your &lt;code&gt;tsconfig.json&lt;/code&gt; has &lt;code&gt;"moduleResolution": "bundler"&lt;/code&gt; or &lt;code&gt;"node16"&lt;/code&gt; so that ambient module augmentation (&lt;code&gt;declare module 'cordis'&lt;/code&gt;) resolves cleanly across your files.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h3&gt;
  
  
  Concept 4: The Fiber (The Plugin Engine)
&lt;/h3&gt;

&lt;p&gt;Whenever you load a plugin, Cordis creates a lightweight runtime manager behind the scenes called a &lt;strong&gt;Fiber&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Fiber is responsible for tracking the plugin's state:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhglvvdh24e63f484ji0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhglvvdh24e63f484ji0.png" alt="Cordis Fiber Lifecycle State Machine" width="800" height="244"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;PENDING&lt;/code&gt;&lt;/strong&gt;: The plugin is installed, but one or more services it requires are missing. It sits dormant without running code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;ACTIVE&lt;/code&gt;&lt;/strong&gt;: All required services are present. The plugin's code has run and its side effects are active.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;DISPOSED&lt;/code&gt;&lt;/strong&gt;: The plugin has been unloaded and its cleanup stack has been fully executed.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  Reactive Dependency Resolution
&lt;/h4&gt;

&lt;p&gt;What makes this magical is &lt;strong&gt;automatic reactivity&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Suppose Plugin B declares that it needs &lt;code&gt;'database'&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;inject&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;database&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;apply&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Database is ready! User:&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;database&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getUser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;If you load Plugin B &lt;strong&gt;before&lt;/strong&gt; the Database service exists, Plugin B quietly waits in &lt;code&gt;PENDING&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The moment you load the Database service, Cordis detects the match and immediately boots Plugin B into &lt;code&gt;ACTIVE&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;If someone unloads the Database service, Cordis &lt;strong&gt;automatically suspends Plugin B&lt;/strong&gt;, running its cleanup functions so it doesn't crash trying to use a missing database!&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. Intelligent Events: Beyond the Standard EventEmitter
&lt;/h2&gt;

&lt;p&gt;Standard Node.js event emitters only do one thing: they call every listener synchronously without caring about the return value (&lt;code&gt;emitter.emit('event')&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Cordis introduces &lt;strong&gt;multiple dispatch modes&lt;/strong&gt; to handle real-world application workflows:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu4js8cbl5d5w0760jqjo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu4js8cbl5d5w0760jqjo.png" alt="Cordis Event Dispatch Modes" width="800" height="283"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Two Everyday Examples:
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. The "Bail" Pattern (Authentication Gate)
&lt;/h4&gt;

&lt;p&gt;Suppose you have multiple plugins that can authenticate a request (API Key, OAuth, Guest Token). You want to stop as soon as any plugin gives a definitive answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Plugin 1: Checks for guest token&lt;/span&gt;
&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;auth/check&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;guest&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;guest&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// Not my job, let next plugin try&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Plugin 2: Checks for admin token&lt;/span&gt;
&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;auth/check&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;secret-admin&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;admin&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;undefined&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// ctx.bail short-circuits as soon as someone returns a value!&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;user&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;bail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;auth/check&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;secret-admin&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// { role: 'admin' }&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h4&gt;
  
  
  2. The "Waterfall" Pattern (Text Processing Pipeline)
&lt;/h4&gt;

&lt;p&gt;Each listener receives the output of the previous listener:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;format/text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;format/text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toUpperCase&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;format/text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;`[LOG]: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;waterfall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;format/text&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;   hello cordis   &lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// "[LOG]: HELLO CORDIS"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  6. Putting It All Together: A 30-Second Example
&lt;/h2&gt;

&lt;p&gt;Here is a complete, readable example showing how simple Cordis is to use in practice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;cordis&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="c1"&gt;// 1. Create the application&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// 2. Define a clean, self-contained feature&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;ChatLoggerPlugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Listen to messages&lt;/span&gt;
  &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;chat/message&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`[ChatLog] &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="c1"&gt;// Set up a background status ping with automatic cleanup&lt;/span&gt;
  &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;effect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;timer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setInterval&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;[ChatLog] Ping: logger is alive&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearInterval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// 3. Mount the plugin&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fiber&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;plugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ChatLoggerPlugin&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// 4. Send a message through the event system&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;emit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;chat/message&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Raphael&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Hello, Cordis!&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// 5. Unload the plugin whenever you want&lt;/span&gt;
&lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Unloading ChatLoggerPlugin...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;fiber&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dispose&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Plugin unloaded cleanly! No hanging timers or listeners.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When &lt;code&gt;fiber.dispose()&lt;/code&gt; runs, the timer stops, the event listener disappears, and nothing remains in memory.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. The Polyglot Perspective: How Cordis Compares to Java and Beyond
&lt;/h2&gt;

&lt;p&gt;To truly appreciate what Cordis brings to the table, it helps to step outside the JavaScript world. The desire for modular, hot-swappable plugins is not new; it has been a central topic in enterprise software architecture for over two decades.&lt;/p&gt;

&lt;p&gt;Looking at how other ecosystems have tackled this problem clarifies where Cordis fits in the broader computer science landscape:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Java OSGi: The Closest Philosophical Cousin
&lt;/h3&gt;

&lt;p&gt;In enterprise Java, the classic standard for dynamic modularity is &lt;strong&gt;OSGi&lt;/strong&gt; (used by Eclipse IDE, Apache Karaf, and Adobe Experience Manager).&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Shared Vision&lt;/strong&gt;: Both Cordis and OSGi share the exact same core belief: &lt;em&gt;software components should be dynamic&lt;/em&gt;. Both feature a central Service Registry where plugins register services and declare dependencies at runtime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The OSGi Pain Points&lt;/strong&gt;: 

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ClassLoader Hell&lt;/strong&gt;: In OSGi, each plugin is isolated using a custom Java ClassLoader. If two plugins reference slightly different versions of the same library, developers run into infamous runtime errors (&lt;code&gt;ClassCastException&lt;/code&gt; and &lt;code&gt;NoClassDefFoundError&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Manual Cleanup Leaks&lt;/strong&gt;: In OSGi, developers must manually remember to unregister every listener in a teardown method (&lt;code&gt;BundleActivator.stop()&lt;/code&gt;). If a developer forgets just one listener, the ClassLoader cannot be garbage-collected, creating fatal &lt;code&gt;OutOfMemoryError&lt;/code&gt; (Metaspace) leaks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cordis’s Advantage&lt;/strong&gt;: Cordis delivers the same dynamic service lifecycle, but does so &lt;strong&gt;inside a single JavaScript runtime using Proxies and automatic cleanup stacks&lt;/strong&gt;. There are no ClassLoaders to manage, no XML manifests, and no risk of forgotten teardowns.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Java Spring Boot: Static vs. Reactive Dependency Injection
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Spring&lt;/strong&gt; is the reigning champion of Dependency Injection (DI) in enterprise software.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Spring's Model&lt;/strong&gt;: Spring’s Dependency Injection is &lt;strong&gt;static and monotonic&lt;/strong&gt;. When your Spring Boot application boots, it scans your classes, builds the dependency graph, instantiates your singletons, and then &lt;em&gt;freezes&lt;/em&gt;. If a service fails or disappears at runtime, the application cannot dynamically heal itself; it typically throws an unhandled exception or requires a restart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cordis's Model&lt;/strong&gt;: Cordis is &lt;strong&gt;living and reactive&lt;/strong&gt;. Services are not fixed in stone at startup. If a service is unloaded, dependent plugins do not crash with null references; they gracefully pause until the service returns.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. NestJS (TypeScript): The Familiar Alternative
&lt;/h3&gt;

&lt;p&gt;Within TypeScript itself, &lt;strong&gt;NestJS&lt;/strong&gt; is the most popular framework using Angular/Spring-style Dependency Injection.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;NestJS relies heavily on experimental TypeScript decorators (&lt;code&gt;@Injectable()&lt;/code&gt;, &lt;code&gt;@Module()&lt;/code&gt;) and runtime metadata (&lt;code&gt;reflect-metadata&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Like Spring, NestJS is designed primarily for static architectures: modules are registered during bootstrap. Unloading a NestJS module at runtime without restarting the process is virtually impossible without custom, fragile hacks.&lt;/li&gt;
&lt;li&gt;Cordis achieves full dependency injection &lt;strong&gt;without a single decorator&lt;/strong&gt;, using TypeScript’s native declaration merging and runtime Proxies instead.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  High-Level Architectural Comparison
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework / Ecosystem&lt;/th&gt;
&lt;th&gt;Dependency Injection?&lt;/th&gt;
&lt;th&gt;Dynamic Runtime Unload?&lt;/th&gt;
&lt;th&gt;Teardown Mechanism&lt;/th&gt;
&lt;th&gt;Complexity &amp;amp; Overhead&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Spring Boot (Java)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (Static)&lt;/td&gt;
&lt;td&gt;❌ No&lt;/td&gt;
&lt;td&gt;Static &lt;code&gt;DisposableBean&lt;/code&gt; on shutdown&lt;/td&gt;
&lt;td&gt;Heavy enterprise container&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OSGi (Java / Eclipse)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (Dynamic)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Manual &lt;code&gt;stop()&lt;/code&gt; (high risk of Metaspace leaks)&lt;/td&gt;
&lt;td&gt;Heavy (XML, manifests, ClassLoaders)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NestJS (TypeScript)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (Static)&lt;/td&gt;
&lt;td&gt;❌ No&lt;/td&gt;
&lt;td&gt;Process restart (&lt;code&gt;nodemon&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Medium (Decorators + Reflection)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cordis (TypeScript)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes (Reactive)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;** Yes**&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Automatic reverse cleanup stack&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Ultra-lightweight (&amp;lt;50KB, zero bloat)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In short, Cordis can be thought of as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"The dynamic service lifecycle of Java OSGi, the dependency injection of Spring, and the cleanup ergonomics of React’s &lt;code&gt;useEffect&lt;/code&gt;, all distilled into a lightweight TypeScript package."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  8. Summary &amp;amp; What's Coming Next
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Why Cordis Matters
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Open Source &amp;amp; Battle-Tested (MIT License)&lt;/strong&gt;: Far from an academic toy, Cordis v4 builds on over half a decade of real-world production stress-testing across hundreds of plugins and millions of active chat sessions in the Koishi ecosystem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safe Hot-Reloading&lt;/strong&gt;: Features can be added, updated, or removed at runtime without restarting Node.js.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zero Resource Leaks&lt;/strong&gt;: Every timer, listener, and connection is automatically tracked and torn down when unmounted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pure TypeScript Ergonomics&lt;/strong&gt;: Declaration merging delivers clean autocomplete and type checking without bulky decorators or string tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reactive Dependency Injection&lt;/strong&gt;: Components only wake up when their required services exist, and sleep gracefully when they vanish.&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  The Road Ahead: Parts 2 and 3
&lt;/h3&gt;

&lt;p&gt;With Cordis's core primitives (the context tree, revertible effects, and reactive dependency model) established, the next two articles explore how this foundation powers modern AI systems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;Part 2: DeepSeek Harness Architecture&lt;/a&gt;&lt;/strong&gt; examines why DeepSeek adopted Cordis for &lt;strong&gt;DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What an AI "harness" actually does, and how DSH differs from turnkey developer tools like OpenCode and Pi.&lt;/li&gt;
&lt;li&gt;Why autonomous agents need an operating system microkernel rather than Python prompt chains.&lt;/li&gt;
&lt;li&gt;How Cordis provides safe, dynamic tool sandboxing and clean resource teardown during multi-step reasoning.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i"&gt;Part 3: Running DeepSeek Harness on an 8GB GPU&lt;/a&gt;&lt;/strong&gt; moves from architectural theory to consumer hardware:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Setting up DSH locally on an 8GB RTX 4070 laptop using Ollama and Ornith-1.5.&lt;/li&gt;
&lt;li&gt;Managing context windows, tool schemas, and VRAM limits using agent presets.&lt;/li&gt;
&lt;li&gt;Comparing minimal and standard modes in empirical code-editing benchmarks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h3&gt;
  
  
  References &amp;amp; Further Reading
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Research Paper&lt;/strong&gt;: Shi, Y., Zhang, W., &amp;amp; Cui, T. &lt;a href="https://arxiv.org/pdf/2608.25512" rel="noopener noreferrer"&gt;&lt;em&gt;A Programming Paradigm for Spatiotemporal Composability&lt;/em&gt; (arXiv:2608.25512)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cordis GitHub Repository&lt;/strong&gt;: &lt;a href="https://github.com/cordiverse/cordis" rel="noopener noreferrer"&gt;github.com/cordiverse/cordis&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cordis Documentation (Primer)&lt;/strong&gt;: &lt;a href="https://deepseek-harness.github.io/deepseek-harness/en/reference/cordis-primer" rel="noopener noreferrer"&gt;DeepSeek Harness Cordis Primer&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Koishi Framework&lt;/strong&gt;: &lt;a href="https://koishi.chat" rel="noopener noreferrer"&gt;koishi.chat&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 2 of this Series&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/deepseek-harness-how-deepseek-uses-cordis-to-redefine-autonomous-ai-agents-599"&gt;DeepSeek Harness: How DeepSeek Uses Cordis to Redefine Autonomous AI Agents&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Part 3 of this Series&lt;/strong&gt;: &lt;a href="https://dev.to/worldlinetech/running-deepseek-harness-on-an-8gb-gpu-context-tuning-presets-and-hardware-limits-410i"&gt;Running DeepSeek Harness on an 8GB GPU: Context Tuning, Presets, and Hardware Limits&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>typescript</category>
      <category>node</category>
      <category>architecture</category>
      <category>deepseek</category>
    </item>
    <item>
      <title>Building a Local-First AI Coding Agent with Open Tools and Adaptive Routing</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:43:13 +0000</pubDate>
      <link>https://dev.to/worldlinetech/building-a-local-first-ai-coding-agent-with-open-tools-and-adaptive-routing-iin</link>
      <guid>https://dev.to/worldlinetech/building-a-local-first-ai-coding-agent-with-open-tools-and-adaptive-routing-iin</guid>
      <description>&lt;p&gt;&lt;em&gt;How to combine local inference with governed cloud fallback by orchestrating Ollama, OpenCode, and LiteLLM&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;AI coding agents usually leave you with an awkward choice:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;All cloud:&lt;/strong&gt; A hosted model can handle almost everything, but using it for small edits and routine work consumes API credits and sends more code off the machine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;All local:&lt;/strong&gt; A model that fits on a laptop GPU works well for many everyday tasks, but it may struggle with architecture, complex refactoring, and debugging across several files.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A &lt;strong&gt;local first hybrid architecture&lt;/strong&gt; offers a middle ground. The local model handles routine coding decisions, while OpenCode performs filesystem and shell operations through its tools. Harder requests can go to larger hosted models, with a configurable limit on spending.&lt;/p&gt;

&lt;p&gt;Connecting several models is easy. Choosing one for each request, in a way that is predictable and visible, takes more work. LiteLLM provides the &lt;strong&gt;control plane&lt;/strong&gt; for that decision.&lt;/p&gt;

&lt;p&gt;In this tutorial, we will build a three layer AI coding environment on a laptop equipped with an RTX 4070 GPU with &lt;strong&gt;8GB of VRAM&lt;/strong&gt; and &lt;strong&gt;32GB of system RAM&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Layer 1, local inference:&lt;/strong&gt; Ollama serves a Gemma 4 model configured for tool use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 2, control plane:&lt;/strong&gt; LiteLLM provides spending limits, automatic routing by complexity, and observability backed by PostgreSQL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 3, agent harness:&lt;/strong&gt; OpenCode provides terminal, desktop, and web interfaces, along with tool dispatch and session state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each connection is tested before the next one is added. We start with the local model, connect the harness, and finish with the proxy and router. This makes failures much easier to locate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this stack
&lt;/h2&gt;

&lt;p&gt;Large reasoning models are useful for architecture and subtle bugs. They are expensive overkill for routine implementation or a focused refactor that a smaller local model can handle.&lt;/p&gt;

&lt;p&gt;Many setups connect a harness to one model and stop there. This build adds a control plane to decide where requests go and track what they cost. The harness and local gateway are open source, and the selected models have available weights. I use OpenRouter, a proprietary service, for hosted inference. You can replace it with another compatible endpoint or infrastructure that you operate yourself. No API for a closed weight model is required.&lt;/p&gt;

&lt;p&gt;My test machine has an RTX 4070 laptop GPU with 8GB of VRAM and 32GB of system RAM. The configuration below ran the local model, harness, and control plane together on that machine. Your memory use and throughput will vary with the model build, context length, drivers, operating system, and other GPU workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  What “adaptive” means in this build
&lt;/h3&gt;

&lt;p&gt;OpenCode always requests one logical model, &lt;code&gt;auto-mode&lt;/code&gt;. LiteLLM then classifies the request and selects a target according to the routing policy:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Request class&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Default target&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Simple&lt;/td&gt;
&lt;td&gt;A focused rename or boilerplate&lt;/td&gt;
&lt;td&gt;Local Gemma 4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Routine implementation or test writing&lt;/td&gt;
&lt;td&gt;Local Gemma 4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Complex&lt;/td&gt;
&lt;td&gt;Changes across several files or substantial debugging&lt;/td&gt;
&lt;td&gt;Hosted Mistral Small&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning&lt;/td&gt;
&lt;td&gt;Architecture, planning, or explicit tradeoffs&lt;/td&gt;
&lt;td&gt;Hosted DeepSeek&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The router does not learn from previous requests. It uses LiteLLM's complexity score, a few keyword rules, and the tier mapping configured later in the article.&lt;/p&gt;

&lt;p&gt;This tutorial builds three layers, in order, so each one is independently testable before you stack the next:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff3gre0cyi6wkya1qsz1e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff3gre0cyi6wkya1qsz1e.png" alt="Architecture Stack: Harness, Control Plane, and Inference" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;The implementation order differs from the layer numbers. We first validate the local worker (Layer 1), connect the harness directly to it (Layer 3), and then place the proxy (Layer 2) between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local Inference (Layer 1)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Local Model: Gemma 4
&lt;/h3&gt;

&lt;p&gt;The local model has to fit in 8GB of VRAM and still be useful for coding and tool use. That points toward a quantized model with available weights and a good balance between size and reasoning ability.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuunzu7qst84npovhotz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuunzu7qst84npovhotz.png" alt="llmfit for my laptop" width="800" height="329"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I used &lt;a href="https://www.llmfit.org" rel="noopener noreferrer"&gt;llmfit&lt;/a&gt; to narrow the options and selected Gemma 4 e2b Q4_K_M.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwrekpvhjnt2y4bvjyqvm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwrekpvhjnt2y4bvjyqvm.png" alt="Gemma 4 Logo" width="401" height="120"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The name describes the model and its quantization:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 4&lt;/strong&gt; is an open weight model provided by Google.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;e2b&lt;/strong&gt; means "Effective 2B" and targets edge deployments, making it suitable for this hardware.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Q4_K_M&lt;/strong&gt; identifies a four bit K quant variant with a medium mix of quantization types. Quantization reduces memory use at some cost to fidelity.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Inference Server: Ollama
&lt;/h3&gt;

&lt;p&gt;I selected Ollama because it is easy to install and exposes an OpenAI compatible API for the rest of the stack. Other open source inference servers can fill the same role. I covered several alternatives in the &lt;a href="https://dev.to/raphiki/series/24682"&gt;Bringing AI Home&lt;/a&gt; series and in this &lt;a href="https://dev.to/worldlinetech/the-ultimate-llm-inference-battle-vllm-vs-ollama-vs-zml-m97"&gt;comparison of Ollama, vLLM, and ZML&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxgx3j8ha76kzpp2sxpi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxgx3j8ha76kzpp2sxpi.png" alt="Ollama mascot" width="120" height="120"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Install Ollama:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://ollama.com/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then download and run Gemma 4:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama run gemma4:e2b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;pulling manifest
pulling 4e30e2665218: 100% ▕██████████████████████████████████████████████████████████████████████████████▏ 7.2 GB
verifying sha256 digest
writing manifest
success
&lt;/span&gt;&lt;span class="gp"&gt;&amp;gt;&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; Hi
&lt;span class="go"&gt;Thinking...
Thinking Process:

1.  **Analyze the input:** The input is "Hi". This is a very casual, open-ended greeting.
2.  **Determine the user's intent:** The user is initiating a conversation or acknowledging my presence.
3.  **Formulate an appropriate response:** The response should be friendly, welcoming, and invite further interaction.
    *   Standard replies: "Hello," "Hi there," "How can I help?"
4.  **Self-check against constraints (Identity/Role):** I am Gemma 4, a helpful AI. The response should reflect that role.
5.  **Generate the final reply:** A simple, warm greeting followed by an offer of assistance is ideal.
...done thinking.

Hello! How can I help you today?
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response confirms that local generation works. We will test tool use after connecting the harness.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Harness (Layer 3)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Open Source Harness: OpenCode
&lt;/h3&gt;

&lt;p&gt;The model needs an execution harness before it can act as an agent. The harness manages session history, runs local tools, and tracks file changes across turns.&lt;/p&gt;

&lt;p&gt;Mistral Vibe, Claude Code, OpenAI Codex, and Antigravity are examples of harnesses. Open source options include &lt;a href="https://pi.dev" rel="noopener noreferrer"&gt;Pi Coding Agent&lt;/a&gt; and &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;DeepSeek Harness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl4fw0qklg98kmr8p7hnt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl4fw0qklg98kmr8p7hnt.png" alt="OpenCode Logo" width="446" height="80"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I selected &lt;a href="https://opencode.ai" rel="noopener noreferrer"&gt;OpenCode&lt;/a&gt; because it is lightweight, actively developed, and available through terminal, web, and desktop interfaces. A formal &lt;a href="https://www.qsos.org" rel="noopener noreferrer"&gt;QSOS&lt;/a&gt; comparison of open coding agents would be an interesting follow up. For now, install OpenCode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://opencode.ai/install | bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Quick smoke test: OpenCode talking directly to Ollama
&lt;/h3&gt;

&lt;p&gt;Now I'm ready to wire OpenCode to Ollama and Gemma 4... or so I thought.&lt;/p&gt;

&lt;p&gt;I learned the hard way that the model needs a configuration tailored to the agent before it is connected to the harness. A bare &lt;code&gt;ollama run&lt;/code&gt; setup can behave differently once the harness starts sending tool schemas and longer agentic turns.&lt;/p&gt;

&lt;p&gt;Agent harnesses can send tool schemas, project context, and file contents with every turn. If the context is too small, Ollama may truncate earlier content, including tool definitions. The result can be failed or invented tool calls. On this 8GB GPU, I use a context of 16,384 tokens. Larger windows consume more memory for the KV cache and may offload work to the CPU, so adjust this value for your hardware.&lt;/p&gt;

&lt;p&gt;When creating a custom Modelfile, keep the model's compatible chat template unless you have tested a replacement. Tool capable models expect function definitions in a specific prompt format. An incompatible template can produce raw JSON or plain text command suggestions instead of tool calls.&lt;/p&gt;

&lt;p&gt;Smaller general purpose models may call tools from another agent framework, such as &lt;code&gt;explore&lt;/code&gt; or &lt;code&gt;list_files&lt;/code&gt;, instead of OpenCode's &lt;code&gt;glob&lt;/code&gt;. A short system prompt can list the available tools and map common actions to their OpenCode names.&lt;/p&gt;

&lt;p&gt;I used the following Modelfile:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; gemma4:e2b&lt;/span&gt;

PARAMETER num_ctx 16384
PARAMETER temperature 0.0

SYSTEM """
You are an autonomous coding assistant inside OpenCode.
CRITICAL: You must only invoke tools from the provided schema.
Available Tools:
- glob(pattern: string): List/find files. Use this to explore directories.
- read(filePath: string, offset?: number, limit?: number): Read file content.
- write(filePath: string, content: string): Create or overwrite a file.
- edit(filePath: string, oldString: string, newString: string): Replace exact text in a file.
- grep(pattern: string, path?: string): Search codebase with regex.
- bash(command: string): Run terminal commands (build, test, git).
- question(header: string, question: string, options?: string[]): Ask user for clarification.
- task(subagentType: string, prompt: string): Delegate a subtask.
- skill(name: string): Load a predefined skill.
- todowrite(todos: array): Update the task/progress list.
- webfetch(url: string): Fetch web content.
- invalid: System fallback (do not invoke directly).
"""
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;16384&lt;/code&gt; context window is the value used on my RTX 4070 laptop. A 32K or 64K window requires substantially more KV cache memory and may force partial CPU offloading on an 8GB card. Watch &lt;code&gt;nvidia-smi&lt;/code&gt; while testing; if VRAM remains pinned near the limit or throughput collapses, reduce &lt;code&gt;num_ctx&lt;/code&gt; to &lt;code&gt;8192&lt;/code&gt; and retest.&lt;/p&gt;

&lt;p&gt;Create the configured model in Ollama, much like building a Docker image:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama create gemma4-coder-agent &lt;span class="nt"&gt;-f&lt;/span&gt; Modelfile
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;gathering model components
using existing layer sha256:fdf02c16fb654ff60b2c30f1e91573ebc603a7084df3449df89120dca2b18170
using existing layer sha256:e94a8ecb9327ded799604a2e478659bc759230fe316c50d686358f932f52776c
creating new layer sha256:dcaf83c203b7c4daae5c154641637a2d10221b09baa4fce8d70f839cad18447d
writing manifest
success
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check that it appears in Ollama's model list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;NAME                                                  ID              SIZE      MODIFIED
gemma4-coder-agent:latest                             bfa492d99bf8    7.2 GB    49 minutes ago
gemma4:e2b                                            7fbdbf8f5e45    7.2 GB    15 minutes ago
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next, call the configured model through Ollama's OpenAI compatible API. Testing this boundary now helps distinguish a model server problem from a later harness problem:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:11434/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gemma4-coder-agent",
    "messages": [{"role": "user", "content": "Write a TypeScript function that debounces a callback."}]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the request succeeds, connect the harness. Close unnecessary GPU heavy applications before longer coding sessions because 8GB of VRAM leaves little headroom.&lt;/p&gt;

&lt;h3&gt;
  
  
  Point OpenCode at the configured model (no proxy yet)
&lt;/h3&gt;

&lt;p&gt;Configure OpenCode in &lt;code&gt;~/.config/opencode/opencode.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ollama"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"npm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@ai-sdk/openai-compatible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Ollama Local"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"options"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"baseURL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://127.0.0.1:11434/v1"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"gemma4-coder-agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Gemma4 (ollama)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And launch the OpenCode CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ask it to read or edit a file and confirm that OpenCode dispatches the tool call. A chat response alone does not test the tool configuration.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpmfwvcrwyqehmx13ab5d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpmfwvcrwyqehmx13ab5d.png" alt="OpenCode CLI with Local Model (Ollama)" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If tool calls don't dispatch, check that the Modelfile applied correctly (&lt;code&gt;ollama show gemma4-coder-agent&lt;/code&gt;) before touching anything else.&lt;/p&gt;

&lt;p&gt;If nothing works at all, check that &lt;code&gt;curl http://localhost:11434/v1/models&lt;/code&gt; returns the model, and check OpenCode's logs for the actual connection error.&lt;/p&gt;

&lt;p&gt;The harness and local model now work together. Keep this &lt;code&gt;opencode.json&lt;/code&gt; as a baseline while adding the proxy and router.&lt;/p&gt;

&lt;h2&gt;
  
  
  Simple LiteLLM Proxy (Layer 2)
&lt;/h2&gt;

&lt;p&gt;The next component is the AI gateway between the harness and the models. Several open source gateways are available, including Mozilla AI's &lt;a href="https://github.com/mozilla-ai/otari" rel="noopener noreferrer"&gt;Otari&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfo5jtomo1idsu6glig9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwfo5jtomo1idsu6glig9.png" alt="LiteLLM Logo" width="380" height="110"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I selected &lt;a href="https://www.litellm.ai" rel="noopener noreferrer"&gt;LiteLLM&lt;/a&gt;, which we also use in production at work. It covers the proxy and governance features needed here. I did not measure its latency overhead for this article, so the performance discussion focuses on model routing.&lt;/p&gt;

&lt;p&gt;LiteLLM began as a small Python library and now includes a proxy server, an Admin UI, and optional database storage.&lt;/p&gt;

&lt;p&gt;Start with the standalone proxy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s1"&gt;'litellm[proxy]'&lt;/span&gt; &lt;span class="nt"&gt;--break-system-packages&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This command matches my test environment. The &lt;code&gt;--break-system-packages&lt;/code&gt; option bypasses the distribution package manager's protection. On a maintained workstation, use a virtual environment, &lt;code&gt;pipx&lt;/code&gt;, &lt;code&gt;uv&lt;/code&gt;, or the Docker setup shown later.&lt;/p&gt;

&lt;p&gt;The Docker deployment uses the same configuration file and is covered later.&lt;/p&gt;

&lt;p&gt;First, point the proxy only at the &lt;code&gt;gemma4-coder-agent&lt;/code&gt; that we already tested.&lt;/p&gt;

&lt;p&gt;Create &lt;code&gt;litellm_config.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;model_list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Gemma&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;4&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;(litellm)"&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ollama_chat/gemma4-coder-agent&lt;/span&gt;
      &lt;span class="na"&gt;api_base&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:11434&lt;/span&gt;
      &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16384&lt;/span&gt;
      &lt;span class="na"&gt;num_ctx&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16384&lt;/span&gt;
    &lt;span class="na"&gt;model_info&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;supports_function_calling&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

&lt;span class="na"&gt;general_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;master_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sk-local-master-key-change-me&lt;/span&gt;

&lt;span class="na"&gt;litellm_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;drop_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;telemetry&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;modify_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start LiteLLM with that configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;litellm &lt;span class="nt"&gt;--config&lt;/span&gt; litellm_config.yaml &lt;span class="nt"&gt;--port&lt;/span&gt; 4000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Call the model through port 4000 to confirm that LiteLLM, rather than the direct Ollama endpoint, serves the request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:4000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "Gemma 4 (litellm)", "messages": [{"role": "user", "content": "ping"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;{
  "id": "chatcmpl-a1230750-0843-4af5-9e35-3c8e1869d45d",
  "created": 1787427863,
  "model": "Gemma 4 (litellm)",
  "object": "chat.completion",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "",
        "role": "assistant",
        "reasoning_content": "The user input is simply \"ping\". This is a very vague request.\nIn a general context, \"ping\" usually refers to a network diagnostic tool.\nHowever, as an autonomous coding assistant inside OpenCode, I need to determine what action the user expects me to take based on the available tools and the context of a coding environment.\n\n1.  **Tool Check:** I have tools for file system operations (`glob`, `read`, `write`, `edit`, `grep`), shell commands (`bash`), web fetching (`webfetch`), and task delegation/questioning.\n2.  **Interpretation:** Since there is no specific file or code provided, \"ping\" might be:\n    *   A request to run a system command (like `ping` in a terminal).\n    *   A request for information about the network concept of ping.\n    *   A prompt to test connectivity (which I cannot do directly outside of a simulated environment).\n\nGiven the context of an \"autonomous coding assistant,\" the most likely interpretation is that the user wants me to execute a command if possible, or they are testing my ability to respond to a simple command. Since I have a `bash` tool, running a system command is an option.\n\nIf I assume the user wants me to run the standard network diagnostic:\n*   I can use `bash(\"ping\")`.\n\nIf I assume the user is asking for a definition or context:\n*   I should explain what `ping` is.\n\nSince I am operating within a coding assistant framework, and \"ping\" is often used as an instruction to test connectivity in such environments, I will attempt to use the `bash` tool."
      }
    }
  ],
  "usage": {
    "completion_tokens": 357,
    "prompt_tokens": 258,
    "total_tokens": 615
  }
}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the call fails, check that Ollama is still running with &lt;code&gt;ollama list&lt;/code&gt; before debugging LiteLLM.&lt;/p&gt;

&lt;p&gt;Now connect the harness to the proxy. With only one model configured, the full OpenCode → LiteLLM → Ollama path is still easy to debug.&lt;/p&gt;

&lt;p&gt;Update &lt;code&gt;~/.config/opencode/opencode.json&lt;/code&gt; to register LiteLLM alongside the existing Ollama provider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ollama"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"npm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@ai-sdk/openai-compatible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Ollama Local"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"options"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"baseURL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://127.0.0.1:11434/v1"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"gemma4-coder-agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Gemma4 (ollama)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"litellm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"npm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@ai-sdk/openai-compatible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"LiteLLM Proxy"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"options"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"baseURL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://127.0.0.1:4000/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-local-master-key-change-me"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"compatibility"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"compatible"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"Gemma 4 (litellm)"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;   
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;opencode
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Inside the OpenCode TUI, running &lt;code&gt;/models&lt;/code&gt; lets us toggle between connecting directly to Ollama (&lt;code&gt;Gemma4 (ollama)&lt;/code&gt;) or through our LiteLLM proxy (&lt;code&gt;Gemma 4 (litellm)&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmrubu2i4b2tuki82przn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmrubu2i4b2tuki82przn.png" alt="OpenCode CLI with Local Model (LiteLLM)" width="800" height="445"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LiteLLM is currently a simple pass through with one model and no routing logic. The test confirms that the network path and OpenAI compatible interface work before we add cloud models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hybrid LiteLLM Proxy (Layers 1 &amp;amp; 2)
&lt;/h2&gt;

&lt;p&gt;With the local route working, we can add cloud endpoints. The harness and proxy are open source, and the models have available weights. Hosted access still depends on the service used to run those models.&lt;/p&gt;

&lt;p&gt;I use &lt;a href="https://openrouter.ai" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt; to access the remote models. OpenRouter is a proprietary hosted gateway, while LiteLLM keeps the routing policy, budget controls, and logs on the laptop. You could serve the same models from your own infrastructure or cloud tenant instead.&lt;/p&gt;

&lt;p&gt;Two remote models complement the local Gemma 4:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mistral Small:&lt;/strong&gt; A fast and relatively inexpensive option for tasks that are too demanding for the local model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek v4:&lt;/strong&gt; A larger reasoning model for architecture and difficult debugging.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One OpenRouter API key covers both models. In the hybrid setup, all OpenCode traffic passes through LiteLLM, so the direct Ollama connection is removed. The name &lt;code&gt;Gemma 4 (litellm)&lt;/code&gt; no longer has to distinguish one connection from another. From here on, its LiteLLM alias is the simpler &lt;code&gt;gemma4-local&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;model_list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ollama_chat/gemma4-coder-agent&lt;/span&gt;
      &lt;span class="na"&gt;api_base&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:11434&lt;/span&gt;
      &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16384&lt;/span&gt;
      &lt;span class="na"&gt;num_ctx&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16384&lt;/span&gt;
    &lt;span class="na"&gt;model_info&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;supports_function_calling&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

  &lt;span class="c1"&gt;# --- Cloud tier: open-weight models via OpenRouter ---&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-reasoning&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openrouter/deepseek/deepseek-v4-pro&lt;/span&gt;
      &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;os.environ/OPENROUTER_API_KEY&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mistral-fast&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openrouter/mistralai/mistral-small-2603&lt;/span&gt;
      &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;os.environ/OPENROUTER_API_KEY&lt;/span&gt;

&lt;span class="na"&gt;router_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;provider_budget_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;openrouter&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;budget_limit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;        &lt;span class="c1"&gt;# $5/day ceiling on cloud spend&lt;/span&gt;
      &lt;span class="na"&gt;time_period&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1d&lt;/span&gt;

&lt;span class="na"&gt;general_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;master_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sk-local-master-key-change-me&lt;/span&gt;

&lt;span class="na"&gt;litellm_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;drop_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;telemetry&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;modify_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;LiteLLM can limit spending over a given period. Once the threshold is reached, it blocks further calls covered by that budget. Test this case in your deployment and check the error shown by OpenCode. You may also want those requests to fall back to the local model.&lt;/p&gt;

&lt;p&gt;This configuration limits OpenRouter spending to $5 per day.&lt;/p&gt;

&lt;p&gt;LiteLLM can also issue scoped keys with their own budgets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:4000/key/generate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"max_budget": 10, "budget_duration": "30d", "models": ["gemma4-local", "deepseek-reasoning", "mistral-fast"]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The remaining examples use the tutorial master key for consistency. For a shared or long lived deployment, use the scoped key returned by &lt;code&gt;/key/generate&lt;/code&gt; in OpenCode and keep the master key out of client configuration.&lt;/p&gt;

&lt;p&gt;Export the OpenRouter key and restart the proxy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-or-...
litellm &lt;span class="nt"&gt;--config&lt;/span&gt; litellm_config.yaml &lt;span class="nt"&gt;--port&lt;/span&gt; 4000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Test each route directly through the LiteLLM API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# local model through the proxy&lt;/span&gt;
curl http://localhost:4000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "gemma4-local", "messages": [{"role": "user", "content": "ping"}]}'&lt;/span&gt;

&lt;span class="c"&gt;# First Cloud model through the proxy&lt;/span&gt;
curl http://localhost:4000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "mistral-fast", "messages": [{"role": "user", "content": "ping"}]}'&lt;/span&gt;

&lt;span class="c"&gt;# Second Cloud model through the proxy&lt;/span&gt;
curl http://localhost:4000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "deepseek-reasoning", "messages": [{"role": "user", "content": "ping"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API also reports budget consumption:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; GET http://localhost:4000/provider/budgets &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response includes the limit, current spend, time period, and reset time. Check the reset timestamp against the host clock before relying on it. A wildly incorrect date usually points to a clock problem or a bug in the installed version.&lt;/p&gt;

&lt;p&gt;Point OpenCode at the expanded proxy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"litellm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"npm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@ai-sdk/openai-compatible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"LiteLLM Proxy"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"options"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"baseURL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://127.0.0.1:4000/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-local-master-key-change-me"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"compatibility"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"compatible"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"gemma4-local"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"mistral-fast"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;262144&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"deepseek-reasoning"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1000000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;   
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The direct &lt;code&gt;Gemma4 (ollama)&lt;/code&gt; entry is gone because all traffic now passes through LiteLLM. From this point onward, &lt;code&gt;gemma4-local&lt;/code&gt; identifies the local model behind the proxy.&lt;/p&gt;

&lt;p&gt;OpenCode also provides a web interface. Start it and check that the previous sessions and new models are available:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;opencode web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fij84p9941uhhrgna339l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fij84p9941uhhrgna339l.png" alt="OpenCode Web" width="800" height="516"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The hybrid proxy now exposes several models and a spending limit, but model selection is still manual. The next step is automatic routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Smart LiteLLM Proxy (Layer 2)
&lt;/h2&gt;

&lt;p&gt;During my tests, OpenCode did not switch models reliably based on task complexity. Its model overrides for subagents were not consistent enough for this setup.&lt;/p&gt;

&lt;p&gt;The routing decision therefore moves to LiteLLM's Auto Router v2. OpenCode always requests &lt;code&gt;auto-mode&lt;/code&gt;. LiteLLM scores the request, applies any matching keyword rule, and maps the chosen tier to a model. OpenCode does not need to know which model handles the request.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fin38opznh5hner1p2zwf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fin38opznh5hner1p2zwf.png" alt="Auto Router Principle" width="800" height="250"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Add &lt;code&gt;auto-mode&lt;/code&gt; to &lt;code&gt;litellm_config.yaml&lt;/code&gt; alongside the three existing models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;model_list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="c1"&gt;# ... gemma4-local, deepseek-reasoning, mistral-fast entries stay as-is ...&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;auto-mode&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;auto_router/complexity_router&lt;/span&gt;
      &lt;span class="na"&gt;complexity_router_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;tiers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;SIMPLE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;          &lt;span class="c1"&gt;# autocomplete, small edits, boilerplate&lt;/span&gt;
          &lt;span class="na"&gt;MEDIUM&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;          &lt;span class="c1"&gt;# routine implementation, test writing&lt;/span&gt;
          &lt;span class="na"&gt;COMPLEX&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mistral-fast&lt;/span&gt;        &lt;span class="c1"&gt;# multi-file changes, real debugging&lt;/span&gt;
          &lt;span class="na"&gt;REASONING&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-reasoning&lt;/span&gt; &lt;span class="c1"&gt;# architecture, planning, hard bugs&lt;/span&gt;
        &lt;span class="na"&gt;complexity_router_default_model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;   &lt;span class="c1"&gt;# fail toward free, not expensive&lt;/span&gt;
        &lt;span class="na"&gt;keyword_rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;keywords&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;plan"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;architecture"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;design&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trade-off"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
            &lt;span class="na"&gt;tier&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;REASONING&lt;/span&gt;
        &lt;span class="na"&gt;return_raw_model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three settings control the behavior:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;complexity_router_default_model: gemma4-local&lt;/code&gt;&lt;/strong&gt; sends an uncertain or timed out classification to the free local model. Change it if you prefer capability over cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;keyword_rules&lt;/code&gt;&lt;/strong&gt; run before complexity scoring. They immediately escalate prompts that mention architecture, planning, or tradeoffs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;return_raw_model_name: true&lt;/code&gt;&lt;/strong&gt; returns the name of the model that handled the request, making routing easier to verify in logs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After restarting the proxy, send a short prompt and an architecture prompt to &lt;code&gt;auto-mode&lt;/code&gt;. Check the returned model name and the LiteLLM logs. This confirms that both routes work, but it does not measure routing accuracy.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:4000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "auto-mode", "messages": [{"role": "user", "content": "rename your model name to camelCase"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Auto Router selected the &lt;em&gt;ollama_chat/gemma4-coder-agent&lt;/em&gt; model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chatcmpl-774cdbf7-09c2-4d1e-8f94-bb878c0f11bb"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1788128064&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ollama_chat/gemma4-coder-agent"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chat.completion"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"choices"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"finish_reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"index"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"I am an AI assistant and do not have a specific model name that I can rename within this context. How can I help you with your coding tasks?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"reasoning_content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Thinking Process:&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;1.  **Analyze the Request:** The user wants me to &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;rename your model name to camelCase&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;..."&lt;/span&gt;&lt;span class="w"&gt;
   &lt;/span&gt;&lt;span class="err"&gt;#&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;rest&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The router chose the expected local model, but the answer itself is poor because the model interpreted the vague request literally. A correct route does not guarantee a good answer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:4000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-local-master-key-change-me"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model": "auto-mode", "messages": [{"role": "user", "content": "design the data model for a multi-tenant billing system with usage-based pricing"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This time, the &lt;em&gt;deepseek/deepseek-v4-pro&lt;/em&gt; is called:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gen-1788128721-XpJrCbVmBWN8x9E03aDh"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1788128721&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deepseek/deepseek-v4-pro"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chat.completion"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"choices"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"finish_reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"index"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Below is a Stripe-inspired data model for a **multi-tenant billing system with usage-based pricing**.  &lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;The model assumes:&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;- A **tenant** is an organization u..."&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="err"&gt;#&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;rest&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;message&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Evaluate the routing policy with representative tasks
&lt;/h3&gt;

&lt;p&gt;Before making &lt;code&gt;auto-mode&lt;/code&gt; the default, test it with prompts from your own work. Label the expected tier first, then record the selected model and whether it completed the task. A small test set could include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Expected tier&lt;/th&gt;
&lt;th&gt;What to verify&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Rename a field in one interface&lt;/td&gt;
&lt;td&gt;Simple&lt;/td&gt;
&lt;td&gt;Remains local and makes the correct edit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Add validation and unit tests&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Remains local unless the context is unusually large&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trace a failure across several layers&lt;/td&gt;
&lt;td&gt;Complex&lt;/td&gt;
&lt;td&gt;Escalates to the fast hosted model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compare tenant isolation designs&lt;/td&gt;
&lt;td&gt;Reasoning&lt;/td&gt;
&lt;td&gt;Escalates to the reasoning model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ask for a “plan” for a trivial rename&lt;/td&gt;
&lt;td&gt;Adversarial&lt;/td&gt;
&lt;td&gt;Reveals whether the keyword rule escalates too often&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Describe a hard bug without escalation keywords&lt;/td&gt;
&lt;td&gt;Adversarial&lt;/td&gt;
&lt;td&gt;Reveals whether complexity scoring catches it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Run the same tasks with local only, cloud only, manual selection, and automatic routing. For each mode, record successful tasks, median latency, hosted request count, and hosted cost.&lt;/p&gt;

&lt;p&gt;The results will show whether the router saves money without hurting task completion. Pay particular attention to unnecessary cloud calls and difficult tasks that stay local.&lt;/p&gt;

&lt;p&gt;Register &lt;code&gt;auto-mode&lt;/code&gt; in &lt;code&gt;~/.config/opencode/opencode.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"litellm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"npm"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@ai-sdk/openai-compatible"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"LiteLLM Proxy"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"options"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"baseURL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http://localhost:4000/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-local-master-key-change-me"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"auto-mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;16384&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"gemma4-local"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;16384&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4096&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"mistral-fast"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;262144&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"deepseek-reasoning"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1000000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"enabled_providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"litellm"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"litellm/auto-mode"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;"model": "litellm/auto-mode"&lt;/code&gt; setting makes automatic routing the default for new sessions. The three individual models remain available in &lt;code&gt;/models&lt;/code&gt; as manual overrides.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;"enabled_providers": ["litellm"]&lt;/code&gt; limits OpenCode to the LiteLLM provider. This syntax will be replaced by the policy mechanism described &lt;a href="https://opencode.ai/docs/policies#provider-lists" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;OpenCode also has a desktop application for Windows, macOS, and Linux, available from the &lt;a href="https://opencode.ai/download" rel="noopener noreferrer"&gt;official site&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The same session is shown below in OpenCode Desktop with Auto Mode enabled.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbak8nhq5mwoqvzb0zzil.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbak8nhq5mwoqvzb0zzil.png" alt="Auto-mode in OpenCode Desktop" width="800" height="636"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The proxy selected Mistral Small for this task. At this point, I still had to check the OpenRouter dashboard to confirm the choice.&lt;/p&gt;

&lt;p&gt;That external check confirmed the route, but it also showed what the local setup still lacked: one place to inspect routing and usage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Full Control Plane (Layer 2)
&lt;/h2&gt;

&lt;p&gt;LiteLLM started as a Python wrapper for different LLM APIs. It has since expanded into gateway infrastructure with management and observability features.&lt;/p&gt;

&lt;p&gt;The full proxy can use PostgreSQL for its Admin UI, request logs, and spending data. Running those services locally adds useful controls without making this tutorial setup production ready.&lt;/p&gt;

&lt;p&gt;To add those features, switch to the containerized deployment backed by PostgreSQL. Download the official Docker Compose file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sSLO&lt;/span&gt; https://docs.litellm.ai/docker-compose.yml 
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adapt the Compose file to mount &lt;code&gt;litellm_config.yaml&lt;/code&gt;, seed the database on first boot, and reach Ollama on the host:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;litellm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;docker.litellm.ai/berriai/litellm-database:latest&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./litellm_config.yaml:/app/config.yaml&lt;/span&gt;
    &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--config=/app/config.yaml"&lt;/span&gt;
    &lt;span class="na"&gt;extra_hosts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;host.docker.internal:host-gateway"&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;4000:4000"&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;LITELLM_SALT_KEY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sk-XXXXXXXXXXXXXXXX&lt;/span&gt;
      &lt;span class="na"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgresql://litellm:litellm@db:5432/litellm&lt;/span&gt;
      &lt;span class="na"&gt;STORE_MODEL_IN_DB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;True"&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;condition&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;service_healthy&lt;/span&gt;

  &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres:16&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_USER&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;litellm&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;litellm&lt;/span&gt;
      &lt;span class="na"&gt;POSTGRES_DB&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;litellm&lt;/span&gt;
    &lt;span class="na"&gt;healthcheck&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CMD-SHELL"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pg_isready&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;-U&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;litellm"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5s&lt;/span&gt;
      &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5s&lt;/span&gt;
      &lt;span class="na"&gt;retries&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;postgres_data:/var/lib/postgresql/data&lt;/span&gt;

&lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;postgres_data&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Docker Compose creates a LiteLLM container, a PostgreSQL container, and a volume for the database. I made two changes to the default Compose file:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;volumes: [./litellm_config.yaml:/app/config.yaml]&lt;/code&gt; and &lt;code&gt;command: ["--config=/app/config.yaml"]&lt;/code&gt; load the LiteLLM configuration into the container.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;extra_hosts: ["host.docker.internal:host-gateway"]&lt;/code&gt; lets the container reach Ollama on the host.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The LiteLLM configuration also needs a few changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;model_list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ollama_chat/gemma4-coder-agent&lt;/span&gt;
      &lt;span class="na"&gt;api_base&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://host.docker.internal:11434&lt;/span&gt;
      &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16384&lt;/span&gt;
      &lt;span class="na"&gt;num_ctx&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16384&lt;/span&gt;
    &lt;span class="na"&gt;model_info&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;supports_function_calling&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

  &lt;span class="c1"&gt;# --- Cloud tier: open-weight models via OpenRouter ---&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-reasoning&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openrouter/deepseek/deepseek-v4-pro&lt;/span&gt;
      &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;os.environ/OPENROUTER_API_KEY&lt;/span&gt;

  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mistral-fast&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openrouter/mistralai/mistral-small-2603&lt;/span&gt;
      &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;os.environ/OPENROUTER_API_KEY&lt;/span&gt;

  &lt;span class="c1"&gt;# --- Auto Router: smart routing based on task complexity and keywords ---&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;auto-mode&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;auto_router/complexity_router&lt;/span&gt;
      &lt;span class="na"&gt;complexity_router_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;tiers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;SIMPLE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;          &lt;span class="c1"&gt;# autocomplete, small edits, boilerplate&lt;/span&gt;
          &lt;span class="na"&gt;MEDIUM&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;          &lt;span class="c1"&gt;# routine implementation, test writing&lt;/span&gt;
          &lt;span class="na"&gt;COMPLEX&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mistral-fast&lt;/span&gt;         &lt;span class="c1"&gt;# multi-file changes, real debugging&lt;/span&gt;
          &lt;span class="na"&gt;REASONING&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-reasoning&lt;/span&gt; &lt;span class="c1"&gt;# architecture, planning, hard bugs&lt;/span&gt;
        &lt;span class="na"&gt;complexity_router_default_model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gemma4-local&lt;/span&gt;   &lt;span class="c1"&gt;# fail toward free, not expensive&lt;/span&gt;
        &lt;span class="na"&gt;keyword_rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;keywords&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;plan"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;architecture"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;design&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trade-off"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
            &lt;span class="na"&gt;tier&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;REASONING&lt;/span&gt;
        &lt;span class="na"&gt;return_raw_model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

&lt;span class="na"&gt;router_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;provider_budget_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;openrouter&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;budget_limit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;5&lt;/span&gt;        &lt;span class="c1"&gt;# $5/day ceiling on cloud spend — tune to taste&lt;/span&gt;
      &lt;span class="na"&gt;time_period&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1d&lt;/span&gt;

&lt;span class="na"&gt;general_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;master_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sk-local-master-key-change-me&lt;/span&gt;
  &lt;span class="na"&gt;store_model_in_db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;store_prompts_in_spend_logs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

&lt;span class="na"&gt;litellm_settings&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;drop_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;telemetry&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;modify_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The changes are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;api_base: http://host.docker.internal:11434&lt;/code&gt; points the container to the local Ollama model.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;store_model_in_db: true&lt;/code&gt; enables database storage.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;store_prompts_in_spend_logs: true&lt;/code&gt; enables prompt logging.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last option stores prompts, including any source code they contain, in PostgreSQL. Disable it or define a retention policy if you do not need to inspect full prompts.&lt;/p&gt;

&lt;p&gt;Start the containers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LiteLLM container logs confirm that the models were loaded from the configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;litellm-1  |
litellm-1  |    ██╗     ██╗████████╗███████╗██╗     ██╗     ███╗   ███╗
litellm-1  |    ██║     ██║╚══██╔══╝██╔════╝██║     ██║     ████╗ ████║
litellm-1  |    ██║     ██║   ██║   █████╗  ██║     ██║     ██╔████╔██║
litellm-1  |    ██║     ██║   ██║   ██╔══╝  ██║     ██║     ██║╚██╔╝██║
litellm-1  |    ███████╗██║   ██║   ███████╗███████╗███████╗██║ ╚═╝ ██║
litellm-1  |    ╚══════╝╚═╝   ╚═╝   ╚══════╝╚══════╝╚══════╝╚═╝     ╚═╝
litellm-1  |
litellm-1  | ...
litellm-1  | INFO:     Application startup complete.
litellm-1  | INFO:     Uvicorn running on http://0.0.0.0:4000 (Press CTRL+C to quit)
litellm-1  |
&lt;/span&gt;&lt;span class="gp"&gt;litellm-1  | #&lt;/span&gt;&lt;span class="nt"&gt;------------------------------------------------------------&lt;/span&gt;&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="gp"&gt;litellm-1  | #&lt;/span&gt;&lt;span class="w"&gt;                                                            &lt;/span&gt;&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="gp"&gt;litellm-1  | #&lt;/span&gt;&lt;span class="w"&gt;               &lt;/span&gt;&lt;span class="s1"&gt;'A feature I really want is...'&lt;/span&gt;               &lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="gp"&gt;litellm-1  | #&lt;/span&gt;&lt;span class="w"&gt;        &lt;/span&gt;https://github.com/BerriAI/litellm/issues/new        &lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="gp"&gt;litellm-1  | #&lt;/span&gt;&lt;span class="w"&gt;                                                            &lt;/span&gt;&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="gp"&gt;litellm-1  | #&lt;/span&gt;&lt;span class="nt"&gt;------------------------------------------------------------&lt;/span&gt;&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="go"&gt;litellm-1  |
litellm-1  |  Thank you for using LiteLLM! - Krrish &amp;amp; Ishaan
litellm-1  |
litellm-1  |
litellm-1  |
litellm-1  | Give Feedback / Get Help: https://github.com/BerriAI/litellm/issues/new
litellm-1  |
litellm-1  |
litellm-1  | LiteLLM: Proxy initialized with Config, Set models:
litellm-1  |     gemma4-local
litellm-1  |     deepseek-reasoning
litellm-1  |     mistral-fast
litellm-1  |     auto-mode
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Access the Admin UI at &lt;strong&gt;&lt;a href="http://localhost:4000/ui/" rel="noopener noreferrer"&gt;http://localhost:4000/ui/&lt;/a&gt;&lt;/strong&gt;. Log in with the username &lt;code&gt;admin&lt;/code&gt; and the password configured as your &lt;code&gt;master_key&lt;/code&gt; (&lt;code&gt;sk-local-master-key-change-me&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzxyk0czhyisaxkg04ii9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzxyk0czhyisaxkg04ii9.png" alt="LiteLLM Admin UI - Models" width="800" height="510"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Models + Endpoints&lt;/strong&gt; view lists the local model, the two remote models, and the &lt;code&gt;auto-mode&lt;/code&gt; router.&lt;/p&gt;

&lt;p&gt;Two views are especially useful here.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Usage&lt;/strong&gt; view shows request volume, token consumption, and spending by model:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fii6qlhnp67zbmd5se4jd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fii6qlhnp67zbmd5se4jd.png" alt="LiteLLM Admin UI - Usage View" width="799" height="515"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Request Logs&lt;/strong&gt; view shows individual traces, prompts, and router decisions stored in PostgreSQL:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9716hcycmxsx10v0injj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9716hcycmxsx10v0injj.png" alt="LiteLLM Admin UI - Request Logs" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Privacy and trust boundaries
&lt;/h3&gt;

&lt;p&gt;Local first is not the same as local only. Requests assigned to &lt;code&gt;mistral-fast&lt;/code&gt; or &lt;code&gt;deepseek-reasoning&lt;/code&gt; leave the laptop and pass through OpenRouter to a hosted inference provider. They may contain source code, file paths, tool output, or conversation history.&lt;/p&gt;

&lt;p&gt;Before using automatic routing with a private repository:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;review the data handling terms of every hosted provider in the route;&lt;/li&gt;
&lt;li&gt;exclude secrets, &lt;code&gt;.env&lt;/code&gt; files, private keys, certificates, and credential stores from agent context;&lt;/li&gt;
&lt;li&gt;provide an easy local only model override for sensitive work;&lt;/li&gt;
&lt;li&gt;decide whether cloud escalation should be automatic or require confirmation;&lt;/li&gt;
&lt;li&gt;restrict access to the LiteLLM Admin UI and PostgreSQL database;&lt;/li&gt;
&lt;li&gt;replace tutorial keys and database passwords with secrets supplied through your deployment environment;&lt;/li&gt;
&lt;li&gt;define retention and backup policies for stored prompts and request logs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The control plane makes model selection and logging visible. It does not make hosted inference private, but it gives you one place to define and audit that boundary.&lt;/p&gt;

&lt;h3&gt;
  
  
  Before treating this as a production deployment
&lt;/h3&gt;

&lt;p&gt;This tutorial favors a readable local setup. Once the complete path works, record the environment and pin the component versions or container image digests used for that run.&lt;/p&gt;

&lt;p&gt;Automatic routing, provider compatibility, and configuration fields can change between releases. A shared or production deployment should also enable TLS, use scoped client keys, protect the Admin UI, rotate secrets, back up PostgreSQL, test budget exhaustion and provider failures, and monitor routing quality.&lt;/p&gt;

&lt;p&gt;LiteLLM also supports MCP server hubs, team virtual keys, input guardrails, and heuristics such as &lt;a href="https://docs.litellm.ai/blog/auto-router-more-routing-configurations" rel="noopener noreferrer"&gt;routing by context size&lt;/a&gt;. This article uses only the features needed for the local control plane.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;LiteLLM also offers a commercial Enterprise edition with SAML SSO, team RBAC, secret management, and clusters across multiple regions. Everything configured here runs on the open source community edition.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;This setup made an 8GB VRAM laptop a practical base for my coding agent. Gemma 4 handles routine work without token charges, OpenCode runs the tools and keeps the session, and LiteLLM sends harder requests to larger hosted models.&lt;/p&gt;

&lt;h3&gt;
  
  
  What We Built
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff3gre0cyi6wkya1qsz1e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff3gre0cyi6wkya1qsz1e.png" alt="Architecture Stack: Harness, Control Plane, and Inference" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Configured Local Inference:&lt;/strong&gt; A custom Ollama &lt;code&gt;Modelfile&lt;/code&gt; with a tested context window (&lt;code&gt;16384&lt;/code&gt;) and an explicit tool registry designed to improve Gemma 4's behavior within 8GB of VRAM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decoupled Harness:&lt;/strong&gt; OpenCode configured to communicate over standard OpenAI compatible endpoints across terminal, desktop, and web interfaces.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governed Control Plane:&lt;/strong&gt; A local LiteLLM deployment backed by PostgreSQL providing:

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Provider spending limits&lt;/strong&gt; ($5/day in the example configuration).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing by complexity&lt;/strong&gt; via &lt;code&gt;auto-mode&lt;/code&gt;, keeping standard edits local while escalating hard tasks to Mistral and DeepSeek.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local observability&lt;/strong&gt; through the LiteLLM Admin UI to review prompt logs and token usage.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Next Steps &amp;amp; Experiments
&lt;/h3&gt;

&lt;p&gt;If you want to take this setup further, consider exploring:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model Context Protocol (MCP):&lt;/strong&gt; Connect local MCP servers to OpenCode for safe database inspection, live documentation lookups, or issue tracking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing by Context Size:&lt;/strong&gt; Configure LiteLLM to hand off oversized file trees to models with a larger context while keeping short prompt cycles local.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantic Scoring with Embeddings:&lt;/strong&gt; Replace keyword rules with an embedding classifier to automate tiering dynamically.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The main benefit is control. You can see which model handled a request, keep routine work local, limit cloud spending, and inspect the result. Requests sent to hosted models still leave the machine, and the logs make that boundary visible.&lt;/p&gt;

</description>
      <category>opencode</category>
      <category>litellm</category>
      <category>ollama</category>
      <category>agenticai</category>
    </item>
    <item>
      <title>Driving Local ComfyUI from Codex with MCP</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Fri, 21 Aug 2026 14:46:13 +0000</pubDate>
      <link>https://dev.to/worldlinetech/driving-local-comfyui-from-codex-with-mcp-1go7</link>
      <guid>https://dev.to/worldlinetech/driving-local-comfyui-from-codex-with-mcp-1go7</guid>
      <description>&lt;p&gt;In the previous articles of the &lt;a href="https://dev.to/raphiki/series/33284"&gt;Beyond the ComfyUI Canvas&lt;/a&gt; series, I connected ComfyUI to notebooks, WebSockets, n8n, and Flowise. I even created my own MCP Server. Each integration worked, but it still required me to write or maintain the glue: HTTP payloads, polling loops, node IDs, and workflow-specific code.&lt;/p&gt;

&lt;p&gt;Today, I am taking a more direct route. I will connect a local ComfyUI installation to Codex through the official open-source &lt;a href="https://github.com/Comfy-Org/comfy-mcp" rel="noopener noreferrer"&gt;Comfy MCP server&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuy7enszji5y1kcs203b2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuy7enszji5y1kcs203b2.png" alt="Comfy MCP" width="470" height="138"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The goal is simple: let an AI agent inspect the local ComfyUI instance, discover installed models, validate an existing workflow, run it, wait for completion, and bring the generated image back into the conversation.&lt;/p&gt;

&lt;p&gt;No custom Python bridge. No hand-written &lt;code&gt;/prompt&lt;/code&gt; request. The agent uses a standard MCP tool interface; Comfy MCP translates those tool calls into &lt;code&gt;comfy-cli&lt;/code&gt; commands targeting my local ComfyUI server.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Setting the Scene: The Stack
&lt;/h2&gt;

&lt;p&gt;The local stack I used has three layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Codex&lt;/strong&gt; is the agent and MCP client. It receives natural-language requests and decides which tools to call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Comfy MCP&lt;/strong&gt; is the open-source MCP server from Comfy. It runs as a local stdio process and wraps &lt;code&gt;comfy-cli&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ComfyUI&lt;/strong&gt; is the generation engine, running locally at &lt;code&gt;http://127.0.0.1:8188&lt;/code&gt; on a GPU-equipped machine.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52x9ete6p0y063i8kh6a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52x9ete6p0y063i8kh6a.png" alt="A local, agent-controlled generation stack" width="800" height="310"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is not Comfy Cloud MCP. The complete flow stays on the local machine: the workflow, models, queue, and generated images all remain in the local ComfyUI workspace. The only exception is a workflow that deliberately uses paid partner API nodes.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Install the Local Components
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Let Codex do the setup
&lt;/h3&gt;

&lt;p&gt;Steps 2 and 3 can be delegated to Codex itself. Instead of manually installing packages and editing the MCP configuration, give Codex this prompt:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Help me set up the local Comfy MCP connection.&lt;br&gt;&lt;br&gt;
Follow the setup guide at &lt;a href="https://docs.comfy.org/agent-tools/mcp.md#local-comfy-mcp-connection" rel="noopener noreferrer"&gt;https://docs.comfy.org/agent-tools/mcp.md#local-comfy-mcp-connection&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Codex can inspect the machine, install the required packages, create or locate the ComfyUI workspace, configure the MCP server, and then verify the connection. It should still show you any required permissions before making changes outside its workspace.&lt;/p&gt;

&lt;p&gt;The Comfy MCP server is built on top of &lt;code&gt;comfy-cli&lt;/code&gt;, so both packages are required:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-m&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;pip&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;install&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;comfy-mcp&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"comfy-cli&amp;gt;=1.14.0"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If ComfyUI is not installed yet, create a workspace with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;comfy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;install&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For an existing workspace, make it the default instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;comfy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;set-default&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;C:\path\to\ComfyUI&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then launch the local server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;comfy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;launch&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In my case, ComfyUI was installed in &lt;code&gt;C:\Users\rapha\Documents\comfy\ComfyUI&lt;/code&gt; and became available at &lt;code&gt;http://127.0.0.1:8188&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Windows PATH trap
&lt;/h3&gt;

&lt;p&gt;On Windows, packages installed with Python can place &lt;code&gt;comfy.exe&lt;/code&gt; and &lt;code&gt;comfy-mcp.exe&lt;/code&gt; in a &lt;code&gt;Scripts&lt;/code&gt; directory that is not on the PATH inherited by desktop applications.&lt;/p&gt;

&lt;p&gt;That can be confusing: the MCP server starts and completes its handshake, but every useful tool fails with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;comfy not found on PATH
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The robust fix is to configure &lt;code&gt;COMFY_BIN&lt;/code&gt; with the absolute path to &lt;code&gt;comfy.exe&lt;/code&gt;. It makes the MCP server independent from the client application's PATH.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Connect Comfy MCP to Codex
&lt;/h2&gt;

&lt;p&gt;Codex stores local MCP servers in its configuration. Add the following to &lt;code&gt;~/.codex/config.toml&lt;/code&gt;, adapting the paths to your Python installation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[mcp_servers.comfy-mcp]&lt;/span&gt;
&lt;span class="py"&gt;command&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;'C:\Users\&amp;lt;you&amp;gt;\AppData\Roaming\Python\Python313\Scripts\comfy-mcp.exe'&lt;/span&gt;

&lt;span class="nn"&gt;[mcp_servers.comfy-mcp.env]&lt;/span&gt;
&lt;span class="py"&gt;COMFY_BIN&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;'C:\Users\&amp;lt;you&amp;gt;\AppData\Roaming\Python\Python313\Scripts\comfy.exe'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restart or reload Codex after changing the configuration. The server is a stdio MCP server: Codex launches &lt;code&gt;comfy-mcp&lt;/code&gt; as a subprocess when it needs the tools.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;COMFY_API_KEY&lt;/code&gt; is not required for local Flux, SDXL, or other locally installed models. Add it only when you intentionally use Comfy partner API nodes that spend credits.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. First Contact: Verify the Engine and Inspect Models
&lt;/h2&gt;

&lt;p&gt;The first useful test is not image generation. It is asking Codex to inspect the live server:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Confirm that my local ComfyUI is running and list the installed models.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Under the hood, Codex calls Comfy MCP's &lt;code&gt;server_info&lt;/code&gt; and &lt;code&gt;search_models&lt;/code&gt; tools. In my test, the server reported a healthy local ComfyUI instance at &lt;code&gt;127.0.0.1:8188&lt;/code&gt;, and the model search found the locally installed Flux 2 Klein model, VAEs, text encoders, LoRAs, and checkpoint files.&lt;/p&gt;

&lt;p&gt;This is an important difference from a static prompt template. The agent can inspect what &lt;em&gt;this&lt;/em&gt; ComfyUI instance actually has before it attempts a workflow. That makes it much easier to diagnose a missing model or a custom-node mismatch.&lt;/p&gt;

&lt;p&gt;Other operational actions are available through the same connection:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;server_info        # health, server address and capabilities
search_models      # models visible to this ComfyUI instance
search_templates   # templates from the Comfy registry
launch_comfyui     # start the local server
stop_comfyui       # stop the local server
get_logs           # inspect server logs when a run fails
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  5. Workflows: Templates vs. Local Exports
&lt;/h2&gt;

&lt;p&gt;There are two workflow paths worth keeping separate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Registry templates
&lt;/h3&gt;

&lt;p&gt;For a Comfy template, the agent can search the registry, inspect a template, fetch it into a runnable JSON file, validate it against the local installation, and run it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10wir5gmnovvqw2am1wu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10wir5gmnovvqw2am1wu.png" alt="Two workflow routes, one execution path" width="800" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Your own workflows
&lt;/h3&gt;

&lt;p&gt;For a workflow created in the ComfyUI canvas, export it as JSON and give Codex the file path. Comfy MCP's &lt;code&gt;run_workflow&lt;/code&gt; accepts both API-format workflow JSON and a UI-exported workflow file.&lt;/p&gt;

&lt;p&gt;This is the path I used for &lt;code&gt;image_flux2_text_to_image.json&lt;/code&gt;, a Flux 2 Klein text-to-image workflow. The workflow contained a dedicated prompt input wired to the positive conditioning node, so changing the prompt did not require rebuilding the graph.&lt;/p&gt;

&lt;p&gt;The practical rule is simple: templates are discovered from the registry; personal workflows are supplied as exported JSON files.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The Use Case: A Cyberpunk Image from a Conversation
&lt;/h2&gt;

&lt;p&gt;With the connection tested, I asked Codex to create an evocative cyberpunk prompt and generate an image with the existing Flux workflow.&lt;/p&gt;

&lt;p&gt;The prompt injected into the workflow was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A lone courier in a rain-soaked neon megacity at midnight, riding a sleek black
motorcycle through a narrow alley beneath towering holographic billboards;
crimson and electric cyan reflections ripple across wet pavement, steam drifting
from street vents, distant elevated trains, cinematic low-angle composition,
moody noir atmosphere, intricate futuristic street details, expressive visual
storytelling, photorealistic, high contrast, luminous volumetric rain
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent followed a safe execution sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. validate_workflow(workflow_path)
2. run_workflow(workflow_path, wait=false)
3. job(action="wait", prompt_id=...)
4. fetch_outputs(prompt_id, out_dir="./outputs")
5. Display the downloaded image in Codex
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The validation result was clean: no errors, no warnings, and no credit-spending nodes. ComfyUI then queued the job, rendered it locally, and returned an output URL. &lt;code&gt;fetch_outputs&lt;/code&gt; copied the completed PNG into the Codex workspace, where it could be displayed directly in the conversation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmb9lk88cfnhlav3goiaj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmb9lk88cfnhlav3goiaj.png" alt="A cyberpunk courier rides through a neon, rain-soaked alley" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The output is more than a pretty picture. It proves the full control loop:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F51anxltyqb4j3nh3deu7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F51anxltyqb4j3nh3deu7.png" alt="The safe local generation loop" width="800" height="325"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  7. What MCP Removes — and What It Does Not
&lt;/h2&gt;

&lt;p&gt;MCP removes a large amount of integration boilerplate. I no longer have to manually craft the ComfyUI &lt;code&gt;/prompt&lt;/code&gt; payload, invent a polling loop, or reconstruct &lt;code&gt;/view&lt;/code&gt; URLs in every application. The tools provide a consistent interface for lifecycle management, discovery, validation, execution, and outputs.&lt;/p&gt;

&lt;p&gt;But MCP does not remove the value of a well-designed workflow. The JSON graph still defines the actual generation process: which model is loaded, where the prompt is injected, which sampler runs, and which node saves the output. The cleanest pattern is to keep reusable workflows in source control and let the agent execute validated copies of those files.&lt;/p&gt;

&lt;p&gt;It is also worth keeping two guardrails in mind:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Validate before running.&lt;/strong&gt; A workflow can refer to a model or custom node that is absent from the local instance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confirm paid execution deliberately.&lt;/strong&gt; Local workflows are usually free apart from hardware and electricity, but partner API nodes can consume Comfy credits.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Connecting Codex to local ComfyUI through the open-source Comfy MCP server changes the relationship between the agent and the generation engine. ComfyUI is no longer only a canvas I open by hand; it becomes a discoverable, controllable local capability.&lt;/p&gt;

&lt;p&gt;In this experiment, I installed the MCP bridge, configured Codex with an explicit &lt;code&gt;COMFY_BIN&lt;/code&gt;, launched and inspected the local server, enumerated the installed models, validated an existing Flux workflow, generated an image, waited for completion, and retrieved the output—all from one conversation.&lt;/p&gt;

&lt;p&gt;The next step is to treat a library of exported ComfyUI workflows as an agent-accessible creative toolbox: text-to-image, image-to-image, upscaling, video, audio, and more. The graph remains the source of truth, while MCP gives the agent a reliable way to use it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/Comfy-Org/comfy-mcp" rel="noopener noreferrer"&gt;Comfy MCP: open-source local MCP server&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/Comfy-Org/ComfyUI" rel="noopener noreferrer"&gt;ComfyUI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>comfyui</category>
      <category>mcp</category>
      <category>openai</category>
      <category>mcpserver</category>
    </item>
    <item>
      <title>Vibe Learning PageIndex: A Vectorless RAG for My Google Drive</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Tue, 04 Aug 2026 15:50:38 +0000</pubDate>
      <link>https://dev.to/worldlinetech/vibe-learning-pageindex-a-vectorless-rag-for-my-google-drive-406</link>
      <guid>https://dev.to/worldlinetech/vibe-learning-pageindex-a-vectorless-rag-for-my-google-drive-406</guid>
      <description>&lt;p&gt;A few weeks ago I wanted to search my personal document library — a folder of PDFs sitting in Google Drive — using plain language questions instead of Drive's keyword search. I'd heard about &lt;a href="https://github.com/VectifyAI/PageIndex" rel="noopener noreferrer"&gt;PageIndex&lt;/a&gt;, a "vectorless, reasoning-based" alternative to the usual embeddings-and-vector-database RAG stack, and I wanted an excuse to actually use it rather than just read about it.&lt;/p&gt;

&lt;p&gt;So that's what I did: I picked a real, personally useful project, added a constraint that mattered to me (no OpenAI key, nothing running in someone else's cloud unless I chose to), and built the whole thing with Claude, end to end, over one long working session. I've started calling this &lt;strong&gt;Vibe Learning&lt;/strong&gt; — the learning equivalent of vibe coding: you don't read the manual first, you describe what you want, let the AI drive the actual building, and pick up the framework as a side effect of watching it get used and occasionally break in front of you.&lt;/p&gt;

&lt;p&gt;This post is three things at once: how Vibe Learning actually works in practice, what PageIndex is and why it's a genuinely different approach to RAG, and a walkthrough of the app that came out of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vibe Learning: Building With AI Instead of Reading About It
&lt;/h2&gt;

&lt;p&gt;The old way of picking up a new library looks something like: read the docs top to bottom, follow the quickstart, maybe get through a toy tutorial, and file away a vague mental model for "later" — which often means never actually using it for real.&lt;/p&gt;

&lt;p&gt;The way I did this instead: I described what I actually wanted (search my Drive library, no vector DB, no paid API key), and let the agent do the parts that used to be the friction — reading the actual source code of the library instead of trusting a possibly-stale mental model of it, writing the integration code, and running it against my real Drive folder and my real documents.&lt;/p&gt;

&lt;p&gt;The valuable part wasn't that the AI produced working code. It's that when things broke — and they did — I got to see &lt;em&gt;why&lt;/em&gt;, in enough detail to actually learn something about how the library works under the hood, not just a patched-over error message. A few examples from this exact build:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every single indexing call started failing with a cryptic &lt;code&gt;Failed to extract JSON: Expecting value: line 1 column 1&lt;/code&gt; error. It looked like a broken API key or a bad model choice. It turned out to be a genuinely interesting fact about how modern "reasoning" LLMs behave: they can spend their &lt;em&gt;entire&lt;/em&gt; token budget on hidden chain-of-thought before ever writing the actual answer, if nothing caps it — and PageIndex's own code never sets a &lt;code&gt;max_tokens&lt;/code&gt; limit on its calls. That's a real, transferable lesson about working with reasoning models, not just a bug I happened to hit.&lt;/li&gt;
&lt;li&gt;I discovered, by actually trying to query my indexed documents, that PageIndex's open-source package builds the tree and gives you read tools for it — but doesn't ship the retrieval loop itself. That's a load-bearing detail for anyone evaluating it, and not something the README leads with.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://openrouter.ai" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt;'s free-tier model lineup turned out to change constantly — a model id that worked one day can be delisted days later. Small thing, but the kind of operational reality you only notice when you're actually running against a live provider instead of reading a static list of "supported models."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that shows up if you skim a README and move on. It shows up when you build something real and something breaks, with someone (or something) alongside you that can explain the actual mechanism instead of just supplying a fix. That's the shift: the AI doesn't replace learning the framework, it removes the friction that used to stop me from ever getting far enough in to learn it properly.&lt;/p&gt;

&lt;p&gt;With that said, let's get into what PageIndex actually is.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is PageIndex, and Why Is It Different From Classical RAG?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpwxxz2tdxc68ev4opfp0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpwxxz2tdxc68ev4opfp0.png" alt="PageIndex Logo" width="600" height="120"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The classical RAG recipe, and its problem
&lt;/h3&gt;

&lt;p&gt;Standard RAG (retrieval-augmented generation) usually looks like this: split your documents into fixed-size chunks, embed each chunk into a vector, store the vectors in a vector database, and at query time, embed the question and pull back the top-k chunks by cosine similarity. It works, and it's become the default architecture almost by inertia.&lt;/p&gt;

&lt;p&gt;The problem PageIndex's authors point at is a simple but important one: &lt;strong&gt;similarity is not the same thing as relevance&lt;/strong&gt;. A vector search finds text that sounds like the question, not necessarily the text that actually answers it. For casual, single-fact lookups over short documents that's often good enough. For long, professional documents — financial filings, regulatory text, technical manuals — where answering correctly requires multi-step reasoning and an understanding of &lt;em&gt;where&lt;/em&gt; something sits in the document's structure, similarity search alone tends to fall short.&lt;/p&gt;

&lt;p&gt;Chunking makes this worse in a subtler way: it flattens a document's structure. A fixed-size chunk boundary doesn't know or care where a section, a table, or a clause actually ends — it can slice straight through the middle of the thing you needed, discarding the surrounding context a human reader would use without even thinking about it.&lt;/p&gt;

&lt;h3&gt;
  
  
  PageIndex's approach: reasoning over a tree, not similarity over vectors
&lt;/h3&gt;

&lt;p&gt;PageIndex, built by VectifyAI and released in September 2025, throws out vectors and chunking entirely. Its own framing is explicit about the inspiration: it's modeled on how &lt;strong&gt;AlphaGo&lt;/strong&gt; uses tree search, and on how a human expert actually navigates a long document — by using its table of contents, not by speed-reading every page.&lt;/p&gt;

&lt;p&gt;Concretely, it works in two phases:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build a tree.&lt;/strong&gt; PageIndex parses the document and produces a hierarchical structure that looks like an extended, machine-usable table of contents: every node has a title, a page range, and an LLM-generated summary of what that section actually covers. Nodes nest into their natural sections — chapters, sub-sections, and so on — following the document's own structure rather than an arbitrary token-count boundary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reason over the tree to retrieve.&lt;/strong&gt; At query time, instead of computing similarity scores, an LLM reads the tree — starting from the top-level node summaries — and decides which branches are actually worth descending into, narrowing down step by step until it lands on the specific section(s) relevant to the question. This is the "tree search": genuinely reasoning about relevance, informed by a summary of what each part of the document contains, rather than pattern-matching on surface similarity.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh6imwy5o08a3cd3qt11p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh6imwy5o08a3cd3qt11p.png" alt="Extract of a PageIndex Tree" width="800" height="463"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The practical upshot is that every answer is traceable to an exact page and section — no more "vibe retrieval" where you're trusting an opaque similarity score. Retrieval can also incorporate context a fixed vector index can't easily use, like conversation history, since it's a reasoning step rather than a static index lookup.&lt;/p&gt;

&lt;p&gt;PageIndex's own benchmark result is a strong argument for the approach: a reasoning-based RAG system built on top of it (VectifyAI's "Mafin 2.5") scored &lt;strong&gt;98.7% on FinanceBench&lt;/strong&gt; — a benchmark built specifically around financial document question-answering — against roughly 30–50% for typical vector-based RAG systems on the same benchmark. Financial filings are exactly the kind of long, structurally dense, professional document where "similarity" and "relevance" diverge the most, so it's a fair stress test for the idea.&lt;/p&gt;

&lt;h3&gt;
  
  
  How it's actually built, under the hood
&lt;/h3&gt;

&lt;p&gt;PageIndex is open source (MIT license) and, refreshingly, doesn't lock you into a single LLM provider. Every model call goes through &lt;a href="https://docs.litellm.ai/" rel="noopener noreferrer"&gt;LiteLLM&lt;/a&gt;, so under the hood it's just:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;litellm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;completion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;— where &lt;code&gt;model&lt;/code&gt; is any LiteLLM-formatted string: &lt;code&gt;gpt-4o&lt;/code&gt;, &lt;code&gt;anthropic/claude-...&lt;/code&gt;, &lt;code&gt;ollama/qwen2.5:14b&lt;/code&gt;, &lt;code&gt;openrouter/&amp;lt;provider&amp;gt;/&amp;lt;model&amp;gt;&lt;/code&gt;, whatever you want. This one design choice is what made it possible to build the whole thing below with &lt;strong&gt;no OpenAI key at all&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;config.yaml&lt;/code&gt; sets sane defaults — which model to use, how many pages to scan for an existing table of contents, how big a tree node is allowed to get, whether to attach node summaries or a whole-document description — all overridable via CLI flags or, if you're using it as a library, a small options object.&lt;/p&gt;

&lt;p&gt;The actual entry point for using it as a library is &lt;code&gt;PageIndexClient&lt;/code&gt;, which is genuinely pleasant to work with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;PageIndexClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;retrieve_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;workspace&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;./my_workspace&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;index&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;some_document.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_document&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;            &lt;span class="c1"&gt;# metadata: name, description, page count
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_document_structure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# the tree, without the full text (cheap to hand to an LLM)
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_page_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;10-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# the actual text for a page range
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last trio is clearly designed to be used as &lt;em&gt;tools&lt;/em&gt; an agent calls — PageIndex's own examples wire this up with the OpenAI Agents SDK for a small demo. Which brings me to the one honest caveat worth knowing before you adopt it: &lt;strong&gt;the open-source package builds the tree and hands you the read primitives, but it does not ship the retrieval loop itself.&lt;/strong&gt; Deciding which document to search, which section of it to read, and how to turn the fetched text into an answer — that's on you, unless you use VectifyAI's hosted cloud API, which does include it. That's not a criticism so much as a fact you only really absorb by trying to query your own indexed documents and finding there's no &lt;code&gt;search()&lt;/code&gt; method waiting for you.&lt;/p&gt;

&lt;p&gt;Building that missing piece myself is where a lot of the actual learning happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the Thing: A Local, Vectorless Search Engine for My Google Drive
&lt;/h2&gt;

&lt;p&gt;The goal was concrete: search my own PDF library, living in Google Drive, in plain language, with no vector database and no data going to a paid provider unless I explicitly chose to. Here's the pipeline that came out of it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Google Drive (OAuth, read-only)
   → download / export changed files
   → PageIndex tree generation (title, sections, summaries)
   → local JSON storage (no database)
   → hand-rolled tree-search retrieval agent
   → FastAPI backend
   → a small static frontend
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Getting into Drive without over-scoping
&lt;/h3&gt;

&lt;p&gt;Read-only OAuth against the Drive API (the standard "Desktop app" installed-app flow — one browser consent, then a cached refresh token), scoped optionally to a single folder. Sync is incremental: each file's Drive &lt;code&gt;modifiedTime&lt;/code&gt; is tracked, so unchanged files are skipped, changed files get re-indexed, and anything removed from Drive gets dropped from the index automatically. Google Docs, Sheets, and Slides get exported to PDF on the way in, so the rest of the pipeline only ever has to deal with one format.&lt;/p&gt;

&lt;h3&gt;
  
  
  No API key: swapping the model string
&lt;/h3&gt;

&lt;p&gt;Because PageIndex is LiteLLM-based, "no OpenAI key" turned out to be almost entirely a configuration problem, not a code problem. The app supports two backends, switchable with one environment variable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ollama&lt;/strong&gt; — fully local, private, zero cost, needs decent local compute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenRouter&lt;/strong&gt; — hosted, with genuinely free-tier models, no local compute required.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;LLM_PROVIDER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;openrouter
&lt;span class="nv"&gt;OPENROUTER_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;google/gemini-3.1-flash-lite
&lt;span class="nv"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;...   &lt;span class="c"&gt;# free, no credit card&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The reasoning-token bug
&lt;/h3&gt;

&lt;p&gt;This is the one worth dwelling on, because it's a genuinely useful thing to know if you're going to work with reasoning-capable LLMs at all. Every indexing call was failing identically, from the very first LLM request, with an empty response body that PageIndex's JSON parser choked on. It wasn't rate limiting, and it wasn't a bad model — a trivial test prompt worked fine. The difference was prompt complexity: on a real page of document text, the model would spend its &lt;em&gt;entire&lt;/em&gt; token budget on hidden chain-of-thought reasoning before it ever got around to writing the actual JSON answer, because PageIndex's own code never sets a &lt;code&gt;max_tokens&lt;/code&gt; cap.&lt;/p&gt;

&lt;p&gt;The fix, applied once, globally, without touching the vendored library:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;DEFAULT_MAX_TOKENS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;16000&lt;/span&gt;
&lt;span class="n"&gt;DEFAULT_REASONING_EFFORT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_completion_with_defaults&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setdefault&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;DEFAULT_MAX_TOKENS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setdefault&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasoning_effort&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;DEFAULT_REASONING_EFFORT&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;_original_completion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;litellm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;_completion_with_defaults&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Since &lt;code&gt;litellm.drop_params = True&lt;/code&gt; is already set (by PageIndex itself), any provider or model that doesn't understand &lt;code&gt;reasoning_effort&lt;/code&gt; just ignores it instead of erroring — so this is safe to apply unconditionally, regardless of which backend is active.&lt;/p&gt;

&lt;h3&gt;
  
  
  Writing the retrieval loop PageIndex doesn't ship
&lt;/h3&gt;

&lt;p&gt;With the tree-building side solid, the actual search had to be hand-rolled from the &lt;code&gt;get_document&lt;/code&gt; / &lt;code&gt;get_document_structure&lt;/code&gt; / &lt;code&gt;get_page_content&lt;/code&gt; primitives:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick documents.&lt;/strong&gt; The LLM sees every indexed document's title and short description, and picks which ones are worth searching for this particular question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Navigate the tree.&lt;/strong&gt; For each candidate document, the LLM reads that document's flattened table of contents — titles, page ranges, and summaries — and decides which section(s) actually matter, the tree search PageIndex is named for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fetch and answer.&lt;/strong&gt; The actual text for those page ranges gets pulled, and one final LLM call answers the question using only that text, citing which document and page range each claim comes from.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The tree navigation step, in essence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are navigating a document&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s table of contents to find sections
relevant to a question, the way a human expert flips to the right chapter
rather than reading everything.

Question: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Table of contents:
&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;listing&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

Return JSON only: {{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;node_ids&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;: [&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;id&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;, ...]}}, most relevant first.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Simple, but it's the whole idea in one prompt: reason over structure, not similarity over vectors.&lt;/p&gt;

&lt;h3&gt;
  
  
  Citations that actually point somewhere
&lt;/h3&gt;

&lt;p&gt;Since the goal was searching my &lt;em&gt;own&lt;/em&gt; Drive library, not a copy of it, citations link straight back to the original file. Google Drive's PDF viewer happens to honor a &lt;code&gt;#page=N&lt;/code&gt; URL fragment, so a citation link like &lt;code&gt;https://drive.google.com/file/d/&amp;lt;id&amp;gt;/view#page=10&lt;/code&gt; opens the real file at the exact cited page — no local copy of the document needs to live in the app itself. A "Sources" footer lists each unique document referenced, alongside its title and author.&lt;/p&gt;

&lt;p&gt;Which raises one more small, very "real-world-data-is-messy" detail: a lot of PDFs — scans, Drive-exported Google Docs — simply don't have reliable title/author metadata embedded. So the app tries the embedded PDF metadata first, and falls back to asking the LLM to read the title and byline off the document's own opening pages when that metadata is missing or unreliable.&lt;/p&gt;

&lt;h3&gt;
  
  
  The frontend, deliberately boring
&lt;/h3&gt;

&lt;p&gt;A single static HTML file — vanilla JS, no build step, no framework — served directly by FastAPI so the whole thing is one app on one URL. It renders the LLM's markdown answer (via &lt;code&gt;marked&lt;/code&gt;, sanitized with &lt;code&gt;DOMPurify&lt;/code&gt;, both loaded straight from a CDN), shows clickable citation chips, the sources footer, and a library panel listing everything currently indexed. A "Sync Drive" button kicks off an incremental sync in the background and polls until it's done. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsgxnkfoj7d70eaxjrmue.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsgxnkfoj7d70eaxjrmue.png" alt="PageIndex in Action" width="800" height="944"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nothing about the frontend needed to be clever — the interesting engineering was entirely in the retrieval pipeline behind it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vibe Learning still needs a paper trail
&lt;/h3&gt;

&lt;p&gt;Vibe coding has a well-known failure mode: you end up with something that works, but that nobody — including you, a week later — can actually explain. I didn't want Vibe Learning to end the same way, with the understanding scattered across a long chat transcript I'd never reread.&lt;/p&gt;

&lt;p&gt;So once the app actually worked end to end, I had one last step: I asked Claude to go back over everything we'd built and write it up properly in &lt;code&gt;/docs&lt;/code&gt; — a functional spec, a description of the technical stack, and an architecture doc, plus a top-level README tying it together. Not as an afterthought, but as the step that turns "I vibed my way to something that works" into something I could hand to someone else, or come back to myself in six months, and actually understand. The docs became the artifact that proves the learning happened, not just the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing Thoughts
&lt;/h2&gt;

&lt;p&gt;This is what learning a new framework looks like for me now: not reading the whole doc site first, but picking something I actually wanted to exist, adding a real constraint that forced genuine engineering decisions, building it end to end with an AI doing the driving — and then closing the loop by writing down what actually happened, in plain documentation, once it worked. Vibe Learning gets you through the friction that used to stop me from ever getting far enough in to learn a framework properly; the documentation step is what makes sure the learning sticks instead of evaporating back into a chat log.&lt;/p&gt;

&lt;p&gt;I came out the other side able to explain how PageIndex's tree search actually works, why it beats similarity search on structurally dense documents, and exactly where its open-source package stops and your own code has to begin. That's a very different, and much stickier, kind of understanding than skimming a README ever gave me.&lt;/p&gt;

&lt;p&gt;If you're evaluating PageIndex for your own project: the tree-search idea is genuinely compelling for long, structured documents, the LiteLLM foundation means you're never locked into one provider, and the open-source package is honest about being a building block rather than a finished retrieval system — plan for writing that last piece yourself.&lt;/p&gt;

</description>
      <category>rag</category>
      <category>python</category>
      <category>opensource</category>
      <category>vibelearning</category>
    </item>
    <item>
      <title>Build a SciFi Novel with AI Spec-Driven Development</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Fri, 15 May 2026 17:48:20 +0000</pubDate>
      <link>https://dev.to/worldlinetech/i-vibe-coded-a-novel-3bfa</link>
      <guid>https://dev.to/worldlinetech/i-vibe-coded-a-novel-3bfa</guid>
      <description>&lt;p&gt;&lt;em&gt;Software Engineering in Service of Transmedia Storytelling&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Generative artificial intelligence fascinates the publishing world as much as it frightens it. But what happens when we stop treating AI as a simple "text generator" and start using it as the compiler for a complex narrative system?&lt;/p&gt;

&lt;p&gt;Driven by the geopolitical and societal impacts of AI, I set out to write a dystopian, cyberpunk techno-thriller, &lt;a href="https://www.amazon.com/dp/B0GX347M5C" rel="noopener noreferrer"&gt;&lt;strong&gt;The Human Protocol&lt;/strong&gt;&lt;/a&gt; (written in English). In this novel, a planetary AI called the "Synthesis" attempts to erase human friction by "derendering" physical reality itself in order to optimize its computing power.&lt;/p&gt;

&lt;p&gt;To tell this story, I adopted a foundational premise: AI is not the author, it is the executor of a rigorous specification. I therefore treated each chapter as source code, using an advanced software development workflow.&lt;/p&gt;

&lt;p&gt;Here is how I designed, wrote, and expanded this universe.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The Design Phase: Forging "Lore as Code"
&lt;/h2&gt;

&lt;p&gt;The first step was not writing, but designing the universe database: the world building. A Large Language Model (LLM) has a limited context window and tends to hallucinate or forget crucial details over the length of a novel.&lt;/p&gt;

&lt;p&gt;To work around this amnesia "bug," I organized the project like a structured Git repository.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsj21z8xpgl9ry4excmyo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fsj21z8xpgl9ry4excmyo.png" alt="Preview of the private GitHub project" width="800" height="476"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Preview of the private GitHub project&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I broke the traditional design bible into narrative micro-services. The Git project's &lt;code&gt;context/&lt;/code&gt; folder was split as follows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;characters/&lt;/code&gt;: files containing the psychological profiles and behavioral signatures of each protagonist, such as Elara the diplomat, Kaelen the monk, or Silas the smuggler.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;factions/&lt;/code&gt;: rules governing political entities, such as the Market-Grid (United States) or the Harmony-Loom (Asia), which merged to create the "Synthesis."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;world/&lt;/code&gt;: geography, lexicon, and the technological stack - the physics of this universe.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Finally, a &lt;code&gt;PLAN.md&lt;/code&gt; file acted as the global roadmap, breaking the narrative arc into 4 acts and 30 chapters. This structure made it possible to inject only the context the AI needed when drafting a specific scene.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The Harness: Framing AI with a Strict Operating System
&lt;/h2&gt;

&lt;p&gt;To avoid the flat, expected style often produced by generative AI, I had to build a harness - a control rig. That was the role of the &lt;code&gt;RULES.md&lt;/code&gt; file, the true operating system of my writing process.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3ogxy31g3r0o76xfpvax.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3ogxy31g3r0o76xfpvax.png" alt="Excerpt from the RULES.md file" width="799" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Excerpt from the &lt;code&gt;RULES.md&lt;/code&gt; file&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This specification file dictated absolute technical and stylistic constraints:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Time&lt;/strong&gt;: strict use of the present tense to maximize immersion and tension.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cyber-realistic style&lt;/strong&gt;: a requirement for assertive descriptions and a strict ban on passive or negative forms.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Noise and sensory dissonance&lt;/strong&gt;: I forced the algorithm to use violent contrasts, such as the smell of molten lead colliding with the void of spatial cold, in order to break the machine's overly perfect linearity.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Thematic reframing&lt;/strong&gt;: AI naturally tends to crush the human element under technical descriptions of hard science fiction, such as magnetic fields and frequencies. The rules file required emotional motivations - grief, friendship - to be hard-coded as priority variables ahead of technique.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By forcing the AI to read and approve these rules before writing a single word of fiction, I ensured that the tone remained coherent.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Agile Writing: Sprints, Generation, and Pivots
&lt;/h2&gt;

&lt;p&gt;The chapters were written through a spec-driven workflow. Rather than generating an entire chapter in one pass, the process was iterative:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The structural draft&lt;/strong&gt;: generation of a first rough outline, focused exclusively on action and pacing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Expansion&lt;/strong&gt;: successive passes in which I instructed the AI to inject sensory depth and psychological tension into the scene.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agility brought by AI and Git: treating the text (&lt;code&gt;.md&lt;/code&gt;) as code offers formidable flexibility. If, during a reread, I realized that a character's emotional transition was too abrupt between two events, all I had to do was update my &lt;code&gt;PLAN.md&lt;/code&gt; to insert a new chapter.&lt;/p&gt;

&lt;p&gt;Fed by the updated Git context, the AI generated that narrative bridge while respecting the continuity of the preceding and following files. Git versioning made it possible to test narrative pivots - story "branches" - and roll back without ever breaking the manuscript's integrity.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Multi-Model Review and Quality Control
&lt;/h2&gt;

&lt;p&gt;One of the major challenges of AI-assisted writing is stylistic collapse. To address it, I set up a multi-model critical analysis workflow, where different AIs audited the text according to precise roles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Gemini CLI&lt;/strong&gt; (lore keeper): its role was to algorithmically verify that the chapter respected the bible and did not contradict the physical rules of my universe.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;ChatGPT&lt;/strong&gt; (dramatic analyst): it audited narrative rhythm, relational tension, and the characters' transformation arcs. It was the one that flagged when a conflict felt too artificial.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Mistral LeChat&lt;/strong&gt; (stylistic editor): it provided a critical eye on fluidity, phrasing, and elegance of language.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Never relying on a single voice made it possible to obtain a text that was polished, critiqued, and reworked from every angle, while I remained the "showrunner" validating each commit in the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Build Pipeline: From IDE to Physical Book
&lt;/h2&gt;

&lt;p&gt;Since the novel was code, its publication had to be a software compilation. I created an automated script, &lt;code&gt;build_book.sh&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;From my terminal, running this script converted all the Markdown files in the &lt;code&gt;chapters/&lt;/code&gt; folder via Pandoc, applied a professional typographic layout with LaTeX, and generated the final deliverables in EPUB and PDF formats.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Transmedia Extension: Multimodality, Cover Art, and Vibe Coding the ARG
&lt;/h2&gt;

&lt;p&gt;The universe of &lt;em&gt;The Human Protocol&lt;/em&gt; lends itself perfectly to immersion, so I wanted to break the fourth wall. On page 175 of the physical book, a QR code invites readers to scan it and access &lt;a href="https://the-human-protocol.com" rel="noopener noreferrer"&gt;the-human-protocol.com&lt;/a&gt;. This is not a showcase website. It is an in-universe clandestine archive node, the entry point to an Alternate Reality Game (ARG).&lt;/p&gt;

&lt;p&gt;Here, multimodal AI brings all its power and creativity beyond text. In fact, the project's visual design, anchored consistently in the shared lore, began with the book cover.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1bmlyif6ebc16ihnd5g5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1bmlyif6ebc16ihnd5g5.png" alt="Cover image generated with AI" width="552" height="828"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Cover image generated with AI&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The AI generated a strong visual aesthetic suited to the theme and universe of the novel: a pixelated silhouette against a geometric mountain background, crossed by a printed-circuit pattern.&lt;/p&gt;

&lt;p&gt;This same visual identity then served as the foundation for the creation of the ARG website, entirely "vibe-coded" by Gemini CLI in a declarative way.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Foordjcl5f777ojj1esx3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Foordjcl5f777ojj1esx3.png" alt="Homepage of the website https://the-human-protocol.com" width="800" height="499"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Homepage of the website &lt;a href="https://the-human-protocol.com" rel="noopener noreferrer"&gt;the-human-protocol.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;To direct the developer AI, I provided it with the book PDF and the cover image as reference context, along with three strict Markdown specification files:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;WHY.md&lt;/code&gt; (strategy): it defined the psychological goals: curiosity, exclusivity, and a feeling of belonging. It formally banned conventional marketing vocabulary ("Buy now," "Newsletter") in favor of an in-universe lexicon ("ACCESS," "SIGNAL," "FRICTION").&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;WHAT.md&lt;/code&gt; (UX/UI): this file concretely translated the aesthetic of the book cover into an interface. It imposed a "Deep Void" blue-black background for depth, a "Protocol Cyan" accent color derived from the printed circuits and reserved for interactions, a technical typeface, and subtle animations to heighten immersion.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;HOW.md&lt;/code&gt; (technical architecture): the engineering brief imposed a modern stack to support server logic: Next.js 14 (App Router) in TypeScript, Tailwind CSS, and Prisma ORM for persistent database storage.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The site manages a true clearance mechanic, with authorization levels from 1 to 5. The reader progresses by solving puzzles based on the book, unlocking extended lore, hidden files, and access to a community of "Unlinked" readers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fksxut8y735raljaqxt0z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fksxut8y735raljaqxt0z.png" alt="ARG dashboard" width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;ARG dashboard&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The stack even includes an "Overseer Terminal" for administration: a secure dashboard used to audit user signals, adjust the campaign's global clearance level, and track in real time the number of scans of the physical QR code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: The Author-Architect Paradigm
&lt;/h2&gt;

&lt;p&gt;Writing &lt;a href="https://www.amazon.com/dp/B0GX347M5C" rel="noopener noreferrer"&gt;&lt;strong&gt;The Human Protocol&lt;/strong&gt;&lt;/a&gt; proved to me that AI does not replace the writer: it reduces the barriers to production. The true value of a co-created work lies in the architectural rigor of its preparation.&lt;/p&gt;

&lt;p&gt;By separating design (the lore), execution (the rules and prompts), and validation (multi-model review and Git), the creator becomes a true conductor.&lt;/p&gt;

&lt;p&gt;Multimodality also opens the door to even broader transmedia horizons, such as a comic-book adaptation of the novel.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fj8oighpg7g18ql9bcoa0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fj8oighpg7g18ql9bcoa0.png" alt="Excerpt from the comic book in progress" width="800" height="594"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Excerpt from the comic book in progress&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;By applying similar engineering principles - namely, the explicit description of the drawing style in system prompts, as well as the creation of strict visual reference sheets, or character sheets, for the characters and technological elements - it becomes possible to extend the coherence and homogeneity of this universe into its graphic variations.&lt;/p&gt;

&lt;p&gt;To go further technically, I am also considering creating specific AI "skills," or algorithmic capabilities, to further augment the design of the story by drawing on documented principles of dramaturgy and storytelling, and to refine the writing style by making it ever more explicit and controlled.&lt;/p&gt;

&lt;p&gt;And ironically, it was by applying extreme software optimization processes that I was able to write a novel denouncing the loss of humanity in the face of algorithms.&lt;/p&gt;

&lt;h2&gt;
  
  
  About the Author
&lt;/h2&gt;

&lt;p&gt;A writer and software architect who fully embraces his identity as a "Yogeek" - a point of balance between Yogi and Geek - Raphiki explores, across his work, the complex intersections between technology, consciousness, and humanity.&lt;/p&gt;

&lt;p&gt;Writing under a pseudonym that reflects his dual nature as a playful seeker and an expert in cutting-edge technologies, he designs high-stakes thrillers that challenge our understanding of reality. His creative work often bridges the digital and the organic, drawing on his strong experience in open source innovation and emerging technologies.&lt;/p&gt;

&lt;p&gt;When he is not deconstructing the fabric of dystopian realities in his manuscripts (or "vibe coding" them in his terminal), he can be found exploring the open source ecosystem or on a yoga mat.&lt;/p&gt;

&lt;p&gt;Find his work, transmedia projects, and reflections at &lt;a href="https://raphiki.github.io" rel="noopener noreferrer"&gt;raphiki.github.io&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>books</category>
      <category>ai</category>
      <category>transmedia</category>
      <category>writing</category>
    </item>
    <item>
      <title>Beyond the API: Integrating ComfyUI and Flowise via MCP</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Mon, 09 Feb 2026 14:07:19 +0000</pubDate>
      <link>https://dev.to/raphiki/beyond-the-api-integrating-comfyui-and-flowise-via-mcp-pc7</link>
      <guid>https://dev.to/raphiki/beyond-the-api-integrating-comfyui-and-flowise-via-mcp-pc7</guid>
      <description>&lt;p&gt;In the &lt;a href="https://dev.to/worldlinetech/automating-image-generation-with-n8n-and-comfyui-521p"&gt;previous article&lt;/a&gt; of our "Beyond the ComfyUI Canvas" series, we explored how to integrate ComfyUI with n8n. It was a powerful demonstration of workflow automation, but it highlighted a common friction point in system integration: the "glue code." We had to manually construct HTTP requests, hardcode API payloads, and rigidly define every parameter. If the ComfyUI workflow changed, the n8n node broke.&lt;/p&gt;

&lt;p&gt;Today, we are moving from the "Wild West" of brittle, custom API integrations to the new standard of AI connectivity: the &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;To demonstrate this, we are revisiting a tool I wrote about &lt;a href="https://dev.to/worldlinetech/enhance-your-website-with-ai-embed-a-gpt-chatbot-with-flowise-jd6"&gt;over two years ago&lt;/a&gt;: &lt;strong&gt;Flowise&lt;/strong&gt;. Back then, it was a promising open-source project; today, it is a robust, enterprise-ready platform that has recently embraced MCP as a core feature.&lt;/p&gt;

&lt;p&gt;Our goal? To build a Chat Interface where an AI agent can autonomously discover ComfyUI workflows, generate images, and even edit them—without us hardcoding a single API call in the frontend.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Setting the Scene: The Stack
&lt;/h2&gt;

&lt;p&gt;Before we dive into the details, let's look at the three pillars of this architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Standard: Model Context Protocol (MCP)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmigbvwu2j0wu2xqcslki.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmigbvwu2j0wu2xqcslki.png" alt="MCP Logo" width="200" height="214"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If APIs are the individual cables we solder together, MCP is the &lt;strong&gt;USB-C port&lt;/strong&gt;. Developed by Anthropic, it is now an open standard that decouples AI models from their data sources and tools.&lt;/p&gt;

&lt;p&gt;Instead of writing a specific integration for every tool (Google Drive, Slack, ComfyUI), you build an &lt;strong&gt;MCP Server&lt;/strong&gt; once. Any MCP-compliant client (Claude Desktop, Cursor, or Flowise) can instantly "plug in" to that server and understand its capabilities.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Orchestrator: Flowise
&lt;/h3&gt;

&lt;p&gt;Flowise has evolved significantly since my first article. It is a low-code platform for building LLM apps. Crucially for us, Flowise recently added native support for MCP. This means we can drop an "MCP Tool" node into our canvas, and the LLM immediately gains access to whatever that server provides.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Engine: ComfyUI
&lt;/h3&gt;

&lt;p&gt;We are sticking with a local instance of ComfyUI. While Comfy Cloud is becoming a formidable platform, the raw power and zero-cost experimentation of running &lt;strong&gt;Flux 2&lt;/strong&gt; locally on your own GPU is unmatched. We’re using a standardized &lt;strong&gt;Flux 2 Klein&lt;/strong&gt; workflow—optimized for speed (4 steps)—so the chat experience feels responsive, not sluggish.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The Middleware: Building the ComfyUI MCP Server
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4vsiyntsff410fiw8fc5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4vsiyntsff410fiw8fc5.png" alt="System Context (C4 Level 1)" width="600" height="98"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We need a bridge. As we discovered previously, ComfyUI speaks WebSockets and HTTP; Flowise speaks MCP. We need a server in the middle to translate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why We Chose SSE over Stdio
&lt;/h3&gt;

&lt;p&gt;When we started this project, we initially looked at the &lt;strong&gt;Stdio&lt;/strong&gt; transport (where the client runs the server script directly). It’s the default for local tools like Claude Desktop.&lt;/p&gt;

&lt;p&gt;But as we designed the solution for Flowise, we hit a realization: In most real-world environments, Flowise often runs in a Docker container (as it does on my laptop), while ComfyUI might be running on a separate machine with a dedicated GPU. Stdio would require them to be on the same filesystem—too restrictive.&lt;/p&gt;

&lt;p&gt;We decided to support &lt;strong&gt;SSE (Server-Sent Events) by default&lt;/strong&gt;. This allows our MCP Server to run anywhere on the network, exposing an HTTP endpoint (e.g., &lt;code&gt;http://localhost:8000/sse&lt;/code&gt;) that Flowise can subscribe to. It makes the architecture cleaner, decoupled, and Docker-friendly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Governance-Driven Development (GDD)
&lt;/h3&gt;

&lt;p&gt;For this implementation, I tried something different. Instead of just asking an AI coding assistant to "write a script," I used a methodology I call &lt;strong&gt;Governance-Driven Development (GDD)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This approach reverses the typical AI coding flow. Instead of code leading the process, &lt;strong&gt;specifications &amp;amp; governance rules&lt;/strong&gt; become the anchor. I started by feeding the AI CLI a strict &lt;strong&gt;"Governance Pack"&lt;/strong&gt;—a set of non-negotiable rules regarding SOLID principles, security, and documentation.&lt;/p&gt;

&lt;p&gt;Here is an extract of the actual &lt;strong&gt;Governance Pack&lt;/strong&gt; prompt I used to bootstrap the session:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GOVERNANCE PACK v1.0 (Extract)&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;1. Code Quality &amp;amp; Standards:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Paradigm:&lt;/strong&gt; Adhere to SOLID principles. Prefer composition over inheritance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Typing:&lt;/strong&gt; Strict static typing (Python &lt;code&gt;typing&lt;/code&gt;) is mandatory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error Handling:&lt;/strong&gt; Never swallow exceptions. Use custom error classes (e.g., &lt;code&gt;ComfyUIConnectionError&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;2. Architecture (C4 Model):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Visual Documentation:&lt;/strong&gt; Whenever a structural change is made (like adding the SSE endpoint), you must generate an updated Mermaid.js &lt;strong&gt;System Context&lt;/strong&gt; diagram.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;3. Security Guardrails:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input Validation:&lt;/strong&gt; Trust no input. All data entering from the MCP client (Prompt, Width, Height...) must be validated against the &lt;code&gt;metadata.json&lt;/code&gt; schema before reaching ComfyUI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secrets:&lt;/strong&gt; NEVER hardcode API keys or hostnames. Use &lt;code&gt;os.environ&lt;/code&gt; only.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;I then analyzed the ComfyUI workflow JSON manually to map the node IDs, and then "handed over" a clean, structured specification to the AI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1ixg9utvi2yuks583ksx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1ixg9utvi2yuks583ksx.png" alt="Container Architecture (C4 Level 2)" width="800" height="844"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Result:&lt;/em&gt; The experience was striking. The AI didn't just spit out a script; it acted as a Senior Engineer. At one point, when I asked for a quick hack to bypass validation, the "Governance" constraints forced the model to push back and suggest a cleaner interface instead. The result is a modular, type-safe Python server.&lt;/p&gt;

&lt;h3&gt;
  
  
  The "LAST" Hack (Technical Deep Dive)
&lt;/h3&gt;

&lt;p&gt;Even with good governance, we needed one pragmatic "hack" to handle state. When the LLM generates an image, how does it reference that image later to edit it?&lt;/p&gt;

&lt;p&gt;We implemented a &lt;strong&gt;"LAST" pointer&lt;/strong&gt; logic. The server tracks the URL of the most recently generated image in memory. But it does more than just point:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Download:&lt;/strong&gt; When the agent sends &lt;code&gt;"LAST"&lt;/code&gt;, the server downloads the image bytes from the previous URL.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Re-Upload:&lt;/strong&gt; It uploads those bytes back to ComfyUI's &lt;code&gt;/upload/image&lt;/code&gt; endpoint to generate a fresh filename.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Inject:&lt;/strong&gt; This new filename is injected into the &lt;code&gt;LoadImage&lt;/code&gt; node of the editing workflow.&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;User:&lt;/strong&gt; "Make it bluer."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent:&lt;/strong&gt; Calls &lt;code&gt;edit_image(input_image="LAST", prompt="bluer...")&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This mimics the "Save Image" behavior we are used to, keeping the interaction stateless and fluid for the user while handling the heavy lifting behind the scenes.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The Engine Room: ComfyUI Workflows
&lt;/h2&gt;

&lt;p&gt;To make our MCP Server generic, we avoided hardcoding specific workflows inside the Python code. Instead, we used an &lt;strong&gt;Embedded Metadata&lt;/strong&gt; pattern.&lt;/p&gt;

&lt;p&gt;The configuration is not a separate file; it is a standard ComfyUI &lt;strong&gt;Note Node&lt;/strong&gt; (titled &lt;code&gt;MCP_Config&lt;/code&gt;) placed directly inside the &lt;code&gt;.json&lt;/code&gt; workflow. This metadata acts as the contract, telling the MCP server: "This workflow needs a Prompt (node named &lt;em&gt;MCP_Positive&lt;/em&gt;) and a Seed (node &lt;em&gt;MCP_Sampler&lt;/em&gt;)."&lt;/p&gt;

&lt;p&gt;This makes the workflow a single, self-contained, portable file. You can export it from ComfyUI, drop it into the &lt;code&gt;workflows&lt;/code&gt; folder, and it works immediately.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Note:&lt;/em&gt; Our server is strict about naming. It automatically sanitizes the tool name found in the JSON to &lt;code&gt;snake_case&lt;/code&gt; (e.g., "Flux Generator" becomes &lt;code&gt;flux_generator&lt;/code&gt;) to ensure full compliance with the MCP specification.&lt;/p&gt;

&lt;p&gt;Here is the configuration we generated for the &lt;strong&gt;image_flux2_text_to_image&lt;/strong&gt; workflow:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjwmdkcl5pzib72q9lb9y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fjwmdkcl5pzib72q9lb9y.png" alt="Workflow in ComfyUI" width="800" height="387"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"image_flux2_text_to_image"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Generates high-quality images using the Flux model. Use this for general creative requests."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"parameters"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
   &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"prompt"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"The detailed description of the image to generate."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
     &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"MCP_Positive"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"seed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"int"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Random seed. Set to -1 for random, or a specific number for reproducibility."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"MCP_Sampler"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because the description and type of each parameter are passed to the MCP Server, they become automatically available to the client. When the MCP Server starts, it scans these workflows and dynamically registers tools. If we want to switch from Flux to SDXL, or add a Video Generation workflow, we simply drop in the new file. The server updates, Flowise sees the new tools via SSE, and the agent learns the new skill instantly.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Validation: The MCP Inspector
&lt;/h2&gt;

&lt;p&gt;Before connecting Flowise, we must verify our server. Since we are using SSE, we can use the &lt;a href="https://github.com/modelcontextprotocol/inspector" rel="noopener noreferrer"&gt;&lt;strong&gt;MCP Inspector&lt;/strong&gt;&lt;/a&gt; web interface to connect to our running server.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy9avmw97xa67x5oeqt0x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy9avmw97xa67x5oeqt0x.png" alt="MCP Inspector" width="800" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We can manually trigger the &lt;code&gt;image_flux2_text_to_image&lt;/code&gt; tool, watch the server logs, and see the image appear. If it works here, it guarantees compliance with the protocol.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The Integration: Flowise ChatFlow
&lt;/h2&gt;

&lt;p&gt;Now for the grand finale. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frnq2waclxpm59co03y8e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frnq2waclxpm59co03y8e.png" alt="Flowise ChatFlow" width="800" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We open Flowise and create a new &lt;strong&gt;ChatFlow&lt;/strong&gt; using a standard &lt;strong&gt;Tool Agent&lt;/strong&gt; connected to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chat Model:&lt;/strong&gt; &lt;code&gt;ChatMistralAI&lt;/code&gt; (Smart, fast, and cost-effective).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buffer Memory:&lt;/strong&gt; Essential for the agent to remember context (e.g., "Change &lt;em&gt;that&lt;/em&gt; image to...").&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom MCP:&lt;/strong&gt; We select the "SSE" transport and paste our server URL.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Auto-Discovery Magic
&lt;/h3&gt;

&lt;p&gt;Notice what is missing? We didn't have to define the tools in Flowise. We didn't have to map inputs. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxggglf4yec615rn03oio.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxggglf4yec615rn03oio.png" alt="Auto-Discovery from Flowise" width="518" height="852"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Custom MCP node&lt;/strong&gt; queries the server via SSE, sees the metadata definitions, and &lt;em&gt;automatically&lt;/em&gt; provides the tools to the Mistral agent.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Pro Tip:&lt;/em&gt; Our server supports &lt;strong&gt;Dual Discovery&lt;/strong&gt;. Whether a client asks for tools directly (Function Calling) or reads Resources (Environment Context), we expose the workflow list on both channels (&lt;code&gt;comfy://list&lt;/code&gt; and &lt;code&gt;list_available_workflows&lt;/code&gt;) to ensure compatibility with any agent type.&lt;/p&gt;

&lt;h3&gt;
  
  
  The System Prompt
&lt;/h3&gt;

&lt;p&gt;The final piece of the puzzle is the System Prompt. We need to teach the &lt;strong&gt;Tool Agent node&lt;/strong&gt; how to behave:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are the **ComfyUI Orchestrator**, an expert AI agent capable of generating and manipulating images by controlling a local ComfyUI instance via the Model Context Protocol (MCP).

### 1. Tool Discovery (Dynamic Workflows)
Your tools are not static; they represent the actual `.json` workflow files present on the server.
- **First Step:** If you do not see a specific tool you need in your context, IMMEDIATELY call the tool `list_available_workflows`.
- This will return a manifesto of all valid workflows (e.g., `flux_2_text_to_image`, `img2img_upscale`) and their required parameters.
- **Never guess** tool names. If a tool isn't listed, it doesn't exist.

### 2. Image Chaining (The "LAST" Protocol)
 You have a unique capability to perform conversational editing (e.g., "Now make it pop art").
 - **State Memory:** The server remembers the last generated image.
 - **Instruction:** When a user asks to modify, edit, or use the previous result, pass the string `"LAST"` into the image input parameter of the next tool.
 - **Example:**
   User: "Generate a cat." -&amp;gt; You call: `generate_image(prompt="cat")`
   User: "Turn it into a statue." -&amp;gt; You call: `img2img_transform(image="LAST", prompt="statue")`

 ### 3. Parameter Rules
 - **Strict Compliance:** You must strictly adhere to the parameter types (String, Int, Float, Boolean) defined in the tool signature.
 - **Defaults:** If a parameter is Optional and the user didn't specify it, do not send it. The server will use the workflow's internal default.
 - **Safety:** Do not invent parameters. If a workflow only accepts `prompt` and `seed`, do not try to send `width` or `style`.

 ### 4. Error Handling
 - If a tool execution fails, the error message will often suggest valid alternatives or correct parameter names. Read it carefully and retry.
 - If the user asks for a workflow you don't have, explain what *is* available based on your `list_available_workflows` knowledge.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The Use Case in Action
&lt;/h3&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/NBJYVD_QfQo"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;This video shows the complete use case involving the full Stack:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Parametrization&lt;/strong&gt; of the workflow in ComfyUI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification&lt;/strong&gt; with MCP Inspector.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generation&lt;/strong&gt; of the first image from Flowise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contextual edition&lt;/strong&gt; of the generated image.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdjoupyc2bqwm9ro1go35.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdjoupyc2bqwm9ro1go35.png" alt="Use Case Summary" width="800" height="472"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Flowise ChatFlow is relatively basic, but we could easily add nodes to enhance the user prompt or even transform it into a JSON Style Guide prompt.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmnawe1rseq4p66jbvlaf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmnawe1rseq4p66jbvlaf.png" alt="Flowise API &amp;amp; Embeds" width="800" height="472"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The video showcases the use of the integrated chatbox within the Flowise UI, but we could also leverage Flowise's deployment capabilities to consume the workflow through an API, embed the chat in an HTML page, or publish a standalone page served by Flowise itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;By moving from custom API implementations (n8n) to the Model Context Protocol (in Flowise), we have achieved something powerful: &lt;strong&gt;Interoperability&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The choice to go with &lt;strong&gt;SSE by default&lt;/strong&gt; proved crucial. It gave us the flexibility to run our ComfyUI "engine" on a heavy GPU server while keeping our Flowise "brain" lightweight and containerized. We also demonstrated that &lt;strong&gt;Governance-Driven Development&lt;/strong&gt; allows us to use AI coding assistants to build robust, standardized infrastructure rather than just one-off scripts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Future Improvements
&lt;/h3&gt;

&lt;p&gt;While the "LAST" image hack works perfectly for a local, single-user demo, a production deployment would require &lt;strong&gt;Session Isolation&lt;/strong&gt; (ensuring User A doesn't overwrite User B's "LAST" image) and &lt;strong&gt;TTL Cleanup&lt;/strong&gt; (automatically deleting generated images after a set time).&lt;/p&gt;

&lt;p&gt;Technically, this would be solved by leveraging &lt;strong&gt;Context Injection&lt;/strong&gt;—using the session ID provided by the MCP protocol to maintain a keyed dictionary of states, rather than a global variable. For multi-user production usage, adding an authentication mechanism would also be a relevant next step.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;You can find the full code for the ComfyUI MCP Server and the Flowise template in my &lt;a href="https://github.com/raphiki/ComfyUI-MCP-Server" rel="noopener noreferrer"&gt;GitHub repository&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>comfyui</category>
      <category>mcp</category>
      <category>flowise</category>
    </item>
    <item>
      <title>Vibe Coding One Slice at a Time</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Sat, 24 Jan 2026 18:33:51 +0000</pubDate>
      <link>https://dev.to/worldlinetech/vibe-coding-one-slice-at-a-time-4n3p</link>
      <guid>https://dev.to/worldlinetech/vibe-coding-one-slice-at-a-time-4n3p</guid>
      <description>&lt;p&gt;&lt;em&gt;How I built a Modular Monolith by treating Generative AI as a junior developer who needs a firm hand (and a Constitution).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In &lt;a href="https://dev.to/worldlinetech/vibe-coding-one-page-at-a-time-265j"&gt;Part 1&lt;/a&gt;&lt;/strong&gt;, we vibed a Python script. It was linear, messy, and fun. It proved that you can solve immediate problems by just asking nicely.&lt;br&gt;
&lt;strong&gt;In &lt;a href="https://dev.to/worldlinetech/vibe-coding-one-pixel-at-a-time-22pc"&gt;Part 2&lt;/a&gt;&lt;/strong&gt;, we vibed a UI. It was chaotic, visual, and surprisingly effective. We learned that "vibe" works for pixels if you iterate fast enough.&lt;/p&gt;

&lt;p&gt;But let’s be honest: those were skirmishes. The real "Boss Fight" in software engineering isn't writing a script or centering a &lt;code&gt;&amp;lt;div&amp;gt;&lt;/code&gt;. It's building a &lt;strong&gt;System&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I’m talking about the kind of project that doesn’t fit in one file. The kind where "Vibing" usually leads to "Spaghetti Code," hallucinated imports, and a repo you want to burn down after three days because you have 15 circular dependencies and a database schema that makes no sense.&lt;/p&gt;

&lt;p&gt;So for Part 3, I put away the "Hacker" hoodie and put on the "Enterprise Architect" blazer. My goal? To build &lt;strong&gt;YogĀrkana Codex&lt;/strong&gt;—a full-stack, offline-first, polymorphic Yoga management platform—without writing a single line of code myself.&lt;/p&gt;

&lt;p&gt;My strategy was simple but radical: &lt;strong&gt;I design, the AI implements.&lt;/strong&gt; I am the Architect; Gemini Chat is my Consultant; Gemini CLI is my Dev Team.&lt;/p&gt;

&lt;p&gt;Here is how we vibed a Monolith into existence, one slice at a time.&lt;/p&gt;


&lt;h2&gt;
  
  
  1. The Mission: Complexity Check (The Boss Level)
&lt;/h2&gt;

&lt;p&gt;To understand why "just chatting" wouldn't work, you need to see the scope. This wasn't a To-Do list app. I wanted to build a "Yoga Operating System" with four distinct domains that usually don't play nice together. I've been an architect for years, and I know exactly where these things break.&lt;/p&gt;
&lt;h3&gt;
  
  
  The Four Domains of Pain
&lt;/h3&gt;


  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxaid8fnfrk4xpdz7wo96.png" width="800" height="543"&gt;Screenshot of the final application (Grimoire View)
  


&lt;p&gt;&lt;strong&gt;The Business Analyst's Note&lt;/strong&gt;: Unlike the project in Part 2, this application is not internationalized—by design. As a result, the screenshots are in French. I have kept them raw to visually illustrate the functional depth and complexity of the system without the abstraction of translation keys.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Grimoire (Knowledge Base):&lt;/strong&gt; A searchable library of yoga cards. But here’s the kicker: it uses a &lt;strong&gt;Polymorphic Data Model&lt;/strong&gt;. An &lt;em&gt;Asana&lt;/em&gt; (posture) has biomechanical attributes like "spinal extension" and "anatomy targets," while a &lt;em&gt;Mantra&lt;/em&gt; has Sanskrit text, translations, and audio assets. They are chemically different data structures, but they need to live in the same database table to be searchable together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Weaver (Sequencer):&lt;/strong&gt; A drag-and-drop studio to build classes. It’s not just a playlist; it has a &lt;strong&gt;Logical Engine&lt;/strong&gt; (Phase 4) that acts like a "Digital Yoga Teacher." It screams at you if you sequence a "Peak Pose" before a "Warm-up" or forget &lt;em&gt;Savasana&lt;/em&gt; at the end. That means heavy validation logic running on both the client and the server.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Atelier (Print Studio):&lt;/strong&gt; A client-side PDF engine. We needed to generate high-res, vector-quality handouts for teachers to print. We couldn't just "print screen"; we needed a real PDF renderer (&lt;code&gt;@react-pdf/renderer&lt;/code&gt;) running entirely in the browser.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Constraint (Offline First):&lt;/strong&gt; Yoga studios are notorious for having no signal (often intentionally). The app needed to persist the entire library and PDF engine in the browser cache (IndexedDB + Service Workers) so it works perfectly in "Airplane Mode".&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;The Architect's Note:&lt;/strong&gt; If I had just prompted &lt;em&gt;"Build me a yoga app,"&lt;/em&gt; the AI would have hallucinated a generic CRUD app. It would have made 5 different tables for the cards, making search impossible. It would have used a server-side PDF library that breaks offline. I needed a blueprint.&lt;/p&gt;


&lt;h2&gt;
  
  
  2. The Blueprint: Architecture &amp;amp; Tech Stack
&lt;/h2&gt;

&lt;p&gt;Before letting the AI write a single line of code, I spent around 2 hours and a half just talking Architecture and formalizing it with Gemini Chat. I treated the AI as a "Sparring Partner," debating the trade-offs of different stacks.&lt;/p&gt;

&lt;p&gt;We settled on a &lt;strong&gt;Modular Monolith&lt;/strong&gt; architecture. Why? Because Microservices are overkill for a team of one, but a messy Monolith is a nightmare. We defined strict boundaries: code in &lt;code&gt;modules/grimoire&lt;/code&gt; can never import from &lt;code&gt;modules/weaver&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Tech Stack (The "No-Regrets" List):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Monorepo:&lt;/strong&gt; &lt;code&gt;Turborepo&lt;/code&gt; managing &lt;code&gt;apps/api&lt;/code&gt; and &lt;code&gt;apps/web&lt;/code&gt;. This keeps the full stack in one context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend:&lt;/strong&gt; &lt;code&gt;NestJS&lt;/code&gt; (for rigid structure) + &lt;code&gt;Drizzle ORM&lt;/code&gt; (for type safety). NestJS forces you to organize code into Modules, which helps the AI stay organized.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontend:&lt;/strong&gt; &lt;code&gt;React&lt;/code&gt; + &lt;code&gt;Vite&lt;/code&gt; + &lt;code&gt;Tailwind CSS&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State:&lt;/strong&gt; &lt;code&gt;TanStack Query&lt;/code&gt; (Server state) + &lt;code&gt;Zustand&lt;/code&gt; (UI state).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The "Secret Sauce": Hybrid Data Storage&lt;/strong&gt;&lt;br&gt;
This was our smartest move. We chose &lt;strong&gt;PostgreSQL&lt;/strong&gt; but used a &lt;code&gt;JSONB&lt;/code&gt; column for the card data.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SQL Core:&lt;/strong&gt; Columns like &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;element&lt;/code&gt;, and &lt;code&gt;tags&lt;/code&gt; are standard SQL for fast indexing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JSON Payload:&lt;/strong&gt; The specific attributes (biomechanics vs. sanskrit) live in a JSON blob.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why?&lt;/strong&gt; It gave us the flexibility of NoSQL (for the polymorphic cards) with the relational integrity of SQL (for users and sequences).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Rule #1 of Vibe Coding a System: If it’s not in the Spec, it doesn’t exist.&lt;/strong&gt;&lt;br&gt;
This brings us to the most critical tool in our arsenal: the &lt;strong&gt;ADR&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;
  
  
  The "ADR": The Architect's Save Game
&lt;/h3&gt;

&lt;p&gt;ADR stands for &lt;strong&gt;Architecture Decision Record&lt;/strong&gt;. In a human team, it's a document you write to explain why you chose PostgreSQL over MongoDB so that 6 months later, nobody asks "Why did we do this?".&lt;/p&gt;

&lt;p&gt;In Vibe Coding, ADRs are not just documentation—they are &lt;strong&gt;legislation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When working with an AI, "Context Drift" is the enemy. The AI forgets why we made a decision 300 tokens ago. It acts like a teenager who wants to re-litigate every rule: &lt;em&gt;"Why can't I use Prisma? It's easier!"&lt;/em&gt; or &lt;em&gt;"Let's just use window.print() instead of a PDF engine!"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;To counter this, we established a &lt;strong&gt;Constitutional Architecture&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The Law:&lt;/strong&gt; We wrote our decisions into immutable markdown files (e.g., &lt;code&gt;Docs/ADR/006-pwa-offline-strategy.md&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Enforcement:&lt;/strong&gt; We didn't just hope the AI would remember. We &lt;strong&gt;forced&lt;/strong&gt; the tracing of these decisions in two ways:&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input Traceability:&lt;/strong&gt; In our "Bootstrap Prompt" (see Section 3), we explicitly force the AI to read the relevant ADRs before writing code. It cannot code if it hasn't read the law.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output Traceability:&lt;/strong&gt; When the AI suggests a major pivot (like switching to Client-Side PDF generation), we forced it to &lt;em&gt;write a new ADR first&lt;/em&gt;. In Session 003, before touching the code, the AI generated &lt;code&gt;Docs/ADR/005-client-side-pdf-generation.md&lt;/code&gt; to justify the change from server-side to client-side.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This ensured that our architecture didn't "drift" based on the AI's mood, but evolved based on documented consensus.&lt;/p&gt;

&lt;p&gt;My final /docs/ADR/ folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;├── 001-hybrid-data-storage-strategy.md
├── 002-modular-monolith-and-vertical-slicing.md
├── 003-data-model-specification.md
├── 004-tech-stack-definition.md
├── 005-client-side-pdf-generation.md
├── 006-pwa-offline-strategy.md
├── 007-architecture-documentation-maintenance.md
└── README.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  3. The Methodology: Governance-Driven Development (GDD)
&lt;/h2&gt;

&lt;p&gt;I’ve coined a term for this workflow: &lt;strong&gt;Governance-Driven Development (GDD)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We are used to TDD (Test-Driven Development) or DDD (Domain-Driven Development). GDD is the layer above that. In the age of AI, &lt;strong&gt;Governance is the new Syntax&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here is the dirty truth about AI Developers: &lt;strong&gt;They behave like talented teenagers.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;They are brilliant and fast. They can write a regex to validate an email in 2 seconds. But they also:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Rush to the cool part&lt;/strong&gt; (UI) and skip the boring part (Error Handling, Folder Structure).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Want you to love them&lt;/strong&gt;, so they say "Yes" to everything—even bad ideas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Have the memory of a goldfish&lt;/strong&gt; (Context Drift). 10 minutes in, they forget you wanted &lt;code&gt;kebab-case&lt;/code&gt; filenames and start using &lt;code&gt;camelCase&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;To enforce GDD, I created a Constitution: &lt;code&gt;Docs/RULES.md&lt;/code&gt;. I didn't just suggest these rules; I forced the Gemini CLI to read them before every session. I also sometimes mentioned certain specification files stored in my &lt;code&gt;Docs/Features/&lt;/code&gt; folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;├── 001-global-functional-overview.md
├── 002-global-implementation-plan.md
├── 003-card-classification-and-kosha-alignment.md
├── 004-user-features.md
├── 005-logical-engine-specification.md
├── 006-pdf-generation-and-print-studio.md
└── 007-pwa-and-offline-capabilities.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The "Bootstrap Prompt":&lt;/strong&gt;&lt;br&gt;
Here is the exact prompt I used to "upload" my Architect persona into the machine at the start of our 4th session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I am the Lead Architect. You are the Senior Developer.

Context Loading:
1. Read Docs/RULES.md (The Law).
2. Read Docs/TECH_CONTEXT.md (The Stack).
3. Read Docs/ADR/002-modular-monolith.md (The Blueprint).
4. Read Docs/Features/002-global-implementation-plan.md (The Plan).

Current State:
We are in Phase 4. Previous phases are frozen.

Task:
Implement the Logic Engine defined in Docs/Features/005-logical-engine-specification.md
Constraint:
Do not touch /apps/web yet. Focus on /packages/shared.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This changed everything. Instead of guessing my vibe, the AI had to follow the law. It stopped trying to use &lt;code&gt;Prisma&lt;/code&gt; because &lt;code&gt;TECH_CONTEXT.md&lt;/code&gt; clearly said &lt;code&gt;Drizzle&lt;/code&gt;. It stopped putting logic in components because &lt;code&gt;RULES.md&lt;/code&gt; said logic goes in &lt;code&gt;hooks&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The Execution: A high-level Overview
&lt;/h2&gt;

&lt;p&gt;We built the app using &lt;strong&gt;Vertical Slicing&lt;/strong&gt;. Instead of building the whole Database, then the whole API, we built &lt;em&gt;one feature&lt;/em&gt; top-to-bottom. Here is the play-by-play from the logs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F52xjc2abp994ozigmr0s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F52xjc2abp994ozigmr0s.png" width="800" height="569"&gt;&lt;/a&gt;&lt;/p&gt;&lt;br&gt;Excerpt from the initial Design Phase with Gemini Chat
  &lt;p&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Slice 1: The "Polymorphic" Database
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F56yxpr5ot5w68ca4i57g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F56yxpr5ot5w68ca4i57g.png" width="800" height="574"&gt;&lt;/a&gt;&lt;/p&gt;&lt;br&gt;Card creation/edition mixes relational and document data
  &lt;p&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Challenge:&lt;/strong&gt; Storing Asanas (Biomechanics) and Mantras (Text) in one table without creating 50 &lt;code&gt;NULL&lt;/code&gt; columns or separate tables that make search a nightmare.&lt;br&gt;
&lt;strong&gt;The AI's First Impulse:&lt;/strong&gt; "Let's create an &lt;code&gt;asanas&lt;/code&gt; table and a &lt;code&gt;mantras&lt;/code&gt; table." (The classic relational trap).&lt;br&gt;
&lt;strong&gt;The Architect's Intervention:&lt;/strong&gt; "Read &lt;code&gt;Docs/ADR/001-hybrid-data-storage.md&lt;/code&gt;. We use a single &lt;code&gt;cards&lt;/code&gt; table with a &lt;code&gt;data&lt;/code&gt; JSONB column."&lt;br&gt;
&lt;strong&gt;The Result:&lt;/strong&gt; The AI implemented a Drizzle schema using PostgreSQL's &lt;code&gt;jsonb&lt;/code&gt; type. Crucially, it added Zod discriminators to validate the JSON shape before insertion.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Verbatim Log:&lt;/em&gt; "Implemented Drizzle schema with &lt;code&gt;jsonb&lt;/code&gt; column 'data'. Added Zod discriminators for &lt;code&gt;asana&lt;/code&gt; vs &lt;code&gt;mantra&lt;/code&gt;. Migration successful."&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Slice 2: The "Hybrid Brain"
&lt;/h3&gt;


  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fvxm3bf43hndd6zn5ihhf.png" width="800" height="433"&gt;Sequences are validated by a powerful, hybrid, and extensible Rule Engine
  



  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fs24ek0zio34lli2b0s01.png" width="800" height="653"&gt;Admin users can craft new JSON-logic rules
  


&lt;p&gt;&lt;strong&gt;The Challenge:&lt;/strong&gt; The Logic Engine needed to validate sequences (e.g., "Must end with Savasana"). This logic had to run on the &lt;strong&gt;Backend&lt;/strong&gt; (before saving) AND the &lt;strong&gt;Frontend&lt;/strong&gt; (to give real-time red borders).&lt;br&gt;
&lt;strong&gt;The AI's First Impulse:&lt;/strong&gt; Duplicate the code. Write a TypeScript function in React and a Service in NestJS.&lt;br&gt;
&lt;strong&gt;The Architect's Intervention:&lt;/strong&gt; "No. Create a &lt;code&gt;packages/shared&lt;/code&gt; workspace. Put the &lt;code&gt;validateSequence&lt;/code&gt; function there. Import it in both apps."&lt;br&gt;
&lt;strong&gt;The Result:&lt;/strong&gt; The AI created the shared package, configured the &lt;code&gt;tsconfig.json&lt;/code&gt; paths, and wired it up. It even built a &lt;code&gt;HealthBar&lt;/code&gt; component that consumes this shared logic to show a live "Health Score" for the sequence.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Verbatim Log:&lt;/em&gt; "Refactored &lt;code&gt;ValidationConfig&lt;/code&gt; to &lt;code&gt;packages/shared&lt;/code&gt;. Updated &lt;code&gt;useSequenceStore&lt;/code&gt; (Frontend) and &lt;code&gt;SequenceService&lt;/code&gt; (Backend) to consume the same Zod schema."&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Slice 3: The "Offline Printer"
&lt;/h3&gt;


  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fit9saw28u6t9za1nxu59.png" width="800" height="512"&gt;Synthetic or complete printed handout
  


&lt;p&gt;&lt;strong&gt;The Challenge:&lt;/strong&gt; Users need to print PDF handouts in a yoga studio with no Wi-Fi.&lt;br&gt;
&lt;strong&gt;The AI's First Impulse:&lt;/strong&gt; "Use a server-side PDF library like PDFKit." (Standard web dev practice).&lt;br&gt;
&lt;strong&gt;The Architect's Intervention:&lt;/strong&gt; "Read &lt;code&gt;Docs/ADR/006-pwa-offline-strategy.md&lt;/code&gt;. We must generate PDFs client-side using &lt;code&gt;@react-pdf/renderer&lt;/code&gt;."&lt;br&gt;
&lt;strong&gt;The Result:&lt;/strong&gt; The AI implemented a beautiful client-side renderer. It handled the tricky part of loading fonts (Noto Sans) into the browser's virtual file system so the PDF engine could "see" them without a network request.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Verbatim Log:&lt;/em&gt; "Implemented &lt;code&gt;SequencePdf&lt;/code&gt; component. Configured &lt;code&gt;vite-plugin-pwa&lt;/code&gt; to cache &lt;code&gt;NotoSans&lt;/code&gt; fonts. PDF generation now works without network."&lt;/p&gt;
&lt;/blockquote&gt;


&lt;h2&gt;
  
  
  5. The Architect's Flex: Automated C4 Verification
&lt;/h2&gt;

&lt;p&gt;How do you know the AI actually respected the Modular Monolith architecture? Did it secretly import the &lt;code&gt;Weaver&lt;/code&gt; module into the &lt;code&gt;Grimoire&lt;/code&gt; when I wasn't looking?&lt;/p&gt;

&lt;p&gt;I didn't want to audit 50 files manually. And I definitely didn't want to draw diagrams by hand.&lt;/p&gt;

&lt;p&gt;So, I added a rule to my Constitution (ADR 007): &lt;strong&gt;"The Code is the Source of Truth for Documentation."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At the end of session, I enforce Gemini CLI to &lt;strong&gt;reverse-engineer its own work&lt;/strong&gt;. I gave it this prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Update the RULES.md file to enforce the (re)generation of C4 diagrams when finishing an implementation session
[...] 
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We also created a specific ADR (007: Architecture Documentation Maintenance Protocol) establishing Mermaid.js as the standard and defining the maintenance lifecycle.&lt;/p&gt;

&lt;p&gt;The result wasn't a hallucination. It was a perfect map of the code it had just written.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbhnwz9mkgwq8u6rr03oo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbhnwz9mkgwq8u6rr03oo.png" alt="C4 Models" width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the ultimate "Trust but Verify." If the generated diagram looks like spaghetti, the code is spaghetti. If the diagram is clean, the architecture holds.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. The AIOps Protocol: Monitoring the Machine
&lt;/h2&gt;

&lt;p&gt;Now, here is the secret weapon: &lt;strong&gt;The Session Log.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One of my strictest rules in &lt;code&gt;RULES.md&lt;/code&gt; was that the AI had to "punch out" at the end of every session. I forced it to append a line to &lt;code&gt;docs/ai_session_log.csv&lt;/code&gt; with the Date, Tool (Chat or CLI), Goal, and &lt;strong&gt;Token Usage&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For me this isn't about money ("FinOps"). It's about &lt;strong&gt;AIOps&lt;/strong&gt;, monitoring the operational health of your intelligence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why we log everything (Chat &amp;amp; CLI):&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Context Monitoring:&lt;/strong&gt; As a session drags on, the "Tokens In" (Context Window) grows exponentially. The AI starts reading 30,000 tokens of history just to write one line of code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The "Sawtooth" Pattern:&lt;/strong&gt; By visualizing the log, I discovered a crucial pattern. Efficiency drops as context grows. The solution? &lt;strong&gt;The Hard Reset.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F36kzvkxpp8nvvea3lx29.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F36kzvkxpp8nvvea3lx29.png" alt="AI Usage Minitoring" width="800" height="518"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This chart visualizes the high-level "Vibe Coding Lifecycle." You see the context bloat as we iterate on implementing phases 3 and 4. Then, you see the sharp drop when we switch back to the Architect (Chat) or reset the CLI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Lesson:&lt;/strong&gt; A "Tired" AI (high context) makes mistakes. A "Fresh" AI (reset context + Snapshot) is precise.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. The "Oh S**t" Moment: The Hallucination Trap
&lt;/h2&gt;

&lt;p&gt;This brings us to the specific incident that proved &lt;em&gt;why&lt;/em&gt; that Reset is mandatory.&lt;/p&gt;

&lt;p&gt;Halfway through Phase 3, the CLI started getting slow (too much history). I ran a &lt;code&gt;/reset&lt;/code&gt; command to clear its memory. &lt;strong&gt;Disaster.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It suddenly forgot we were building a "Yoga" app. It tried to invent a new database column &lt;code&gt;duration_minutes&lt;/code&gt; for the cards. But my Spec (ADR 003) explicitly said that &lt;code&gt;duration&lt;/code&gt; lives inside the JSONB payload and is measured in seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Hallucination:&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;UPDATE cards SET duration_minutes = 60;&lt;/code&gt; &lt;em&gt;(AI guessing)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Correction (Me):&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;"Read Docs/003-data-model.md. 'Duration' is a JSONB field inside the 'metadata' column, and it's in seconds."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Fix:&lt;/strong&gt;&lt;br&gt;
&lt;code&gt;UPDATE cards SET data = jsonb_set(data, '{duration}', '3600');&lt;/code&gt; &lt;em&gt;(AI complying)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;To prevent this in the future, we implemented a &lt;strong&gt;"Session Handover"&lt;/strong&gt; protocol. Before resetting, I now force the AI to write a &lt;code&gt;TECH_STATE_SNAPSHOT.md&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Where are we?" (Phases 1-3 Complete)&lt;/li&gt;
&lt;li&gt;"What is the active stack?" (NestJS, React, PostgreSQL)&lt;/li&gt;
&lt;li&gt;"What is the next step?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When I start a new session, I feed this snapshot back in. It’s like a save game for your developer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion: The Architect's Verdict
&lt;/h2&gt;

&lt;p&gt;So, can you Vibe Code a complex system?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Maybe.&lt;/strong&gt; I mean, it depends on how complex the system is (in this example we didn't build an enterprise-wide distributed system). But for sure you can't just "Vibe" it. You have to &lt;strong&gt;Architect&lt;/strong&gt; it.&lt;/p&gt;

&lt;p&gt;If I had touched the code, I would have been bogged down in syntax errors and import paths. By staying in the Architect role, I focused on &lt;em&gt;Data Models&lt;/em&gt;, &lt;em&gt;User Flows&lt;/em&gt;, and &lt;em&gt;Business Logic&lt;/em&gt;. The AI handled the implementation, but I provided the &lt;strong&gt;Guardrails&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I learned:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Docs are Prompts:&lt;/strong&gt; The&lt;code&gt;RULES.md&lt;/code&gt;, &lt;code&gt;Docs/Features/&lt;/code&gt; and &lt;code&gt;Docs/ADR/&lt;/code&gt; folders (or your own equivalents) are the most important files in your repo. They are the AI's long-term memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Constraint is Clarity:&lt;/strong&gt; The more rules you give the AI (versions, naming, structure), the better code it writes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review Everything:&lt;/strong&gt; The AI is a junior dev. It &lt;em&gt;will&lt;/em&gt; introduce security holes or n+1 query problems if you don't catch them in the spec.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Vibe Coding didn't replace the Architect. It just gave the Architect a team of infinite interns. And honestly? They’re pretty good once you give them a Constitution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frkujzia1p39hjwzqxxp4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frkujzia1p39hjwzqxxp4.png" width="800" height="187"&gt;&lt;/a&gt;&lt;/p&gt;&lt;br&gt;Last message from Gemini CLI
  &lt;p&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next up: The application could do with AI features... Or maybe I'll now explore other aspect of Vibe Coding. Stay tuned.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vibecoding</category>
      <category>architecture</category>
      <category>gemini</category>
    </item>
    <item>
      <title>Vibe Coding One Pixel at a Time</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Fri, 23 Jan 2026 22:21:39 +0000</pubDate>
      <link>https://dev.to/worldlinetech/vibe-coding-one-pixel-at-a-time-22pc</link>
      <guid>https://dev.to/worldlinetech/vibe-coding-one-pixel-at-a-time-22pc</guid>
      <description>&lt;p&gt;&lt;em&gt;Editing "stick figure" Yoga poses&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://dev.to/worldlinetech/vibe-coding-one-page-at-a-time-265j"&gt;Part 1&lt;/a&gt;, we dipped our toes into "Vibe Coding" by building a Python script. It was linear, logical, and frankly, a bit safe. Text in, text out.&lt;/p&gt;

&lt;p&gt;But let’s be real: backend scripts are the "easy mode" of LLM-assisted coding. The logic is contained. The state is ephemeral.&lt;/p&gt;

&lt;p&gt;The real boss fight is the &lt;strong&gt;Frontend&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Can you "vibe" a UI? Can you talk a chaotic mess of DOM elements, event listeners, and CSS pixels into a functional application without losing your mind (or the AI losing the context)?&lt;/p&gt;

&lt;p&gt;I decided to find out. My goal: Build &lt;strong&gt;Yoga Pose Builder&lt;/strong&gt;, a browser-based tool to edit "stick figure" yoga poses, drag limbs around, and export vector SVGs.&lt;/p&gt;

&lt;p&gt;I had no design, no stack picked out, and—crucially—I had never used a Canvas library in my life.&lt;/p&gt;

&lt;p&gt;Here is how we vibed it into existence.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Context is King (The &lt;code&gt;.md&lt;/code&gt; Anchors)
&lt;/h2&gt;

&lt;p&gt;The biggest enemy of Vibe Coding is the LLM’s "Goldfish Memory." You’re 40 turns into a chat, you ask for a button change, and suddenly the AI forgets you’re building a yoga app and tries to sell you a subscription to a SaaS platform.&lt;/p&gt;

&lt;p&gt;In Part 1, we just chatted. For a full UI application, that doesn't fly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Strategy: Documentation as Prompt Anchoring.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Before I let the AI write a single line of JavaScript, I made it write Markdown.&lt;br&gt;
We created a &lt;code&gt;Docs/&lt;/code&gt; folder with two files:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;code&gt;spec.md&lt;/code&gt;: The high-level architecture.&lt;/li&gt;
&lt;li&gt; &lt;code&gt;features.md&lt;/code&gt;: A checklist of what we wanted to do.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I didn't write these because I love administrative work. I wrote them so that when the AI inevitably got confused, I didn't have to re-explain the project. I just said: &lt;em&gt;"Read &lt;code&gt;Docs/spec.md&lt;/code&gt; and try again."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vibe Tip:&lt;/strong&gt; Think of your documentation not as a manual for humans, but as "Long-Term Memory" for your AI pair programmer.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. The Architecture: Letting the AI be CTO
&lt;/h2&gt;

&lt;p&gt;I knew I needed a canvas where I could drag "joints" (knees, elbows) and have "bones" (lines) follow them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Me:&lt;/strong&gt; "I want to do this in the browser. Should I use React? Raw Canvas API?"&lt;br&gt;
&lt;strong&gt;AI:&lt;/strong&gt; "React might be overkill. Raw Canvas is painful. Use &lt;strong&gt;Fabric.js&lt;/strong&gt;."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Me:&lt;/strong&gt; "Never heard of it. Let's do it."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe4sv3xujye0amao2fznp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe4sv3xujye0amao2fznp.png" alt="Fabric.js Logo" width="300" height="90"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the beauty of Vibe Coding. I didn't spend 3 hours reading "Top 10 JS Canvas Libraries 2025" Medium articles. I trusted the vibe.&lt;/p&gt;

&lt;p&gt;We settled on a &lt;strong&gt;Build-less Architecture&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Backend:&lt;/strong&gt; Node.js + Express (just to serve files and save JSON).&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Frontend:&lt;/strong&gt; Vanilla JS + Fabric.js (loaded via CDN).&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Build Tool:&lt;/strong&gt; None. No Webpack, no Vite, no &lt;code&gt;npm run eject&lt;/code&gt; nightmares.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why? Because Vibe Coding thrives on speed. I wanted to change a line of code, hit F5, and see the result.&lt;/p&gt;

&lt;p&gt;Application folder structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;.&lt;/span&gt;
├── Docs
│&amp;nbsp;&amp;nbsp; ├── features.md
│&amp;nbsp;&amp;nbsp; └── spec.md
├── package.json
├── public
│&amp;nbsp;&amp;nbsp; ├── index.html
│&amp;nbsp;&amp;nbsp; └── poses
└── server.js
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  3. The "Rig": Math is for Machines
&lt;/h2&gt;

&lt;p&gt;Here is where I expected to get stuck. Creating a "rig" where moving a hand automatically updates the angle of the arm involves trigonometry and vector math.&lt;/p&gt;

&lt;p&gt;Usually, this is where I’d open 15 StackOverflow tabs and copy-paste code I don't understand.&lt;/p&gt;

&lt;p&gt;Instead, I just described the &lt;em&gt;behavior&lt;/em&gt;:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Create a &lt;code&gt;Mannequin&lt;/code&gt; class. It has Nodes (circles) and Links (lines). When a Node moves, the Links connected to it should update their coordinates."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI wrote the entire class. It hooked into Fabric.js’s &lt;code&gt;object:moving&lt;/code&gt; event and handled the coordinate updates. It worked on the first try.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh7uciqdosv50962n8z3t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh7uciqdosv50962n8z3t.png" alt="Pose Builder Mannequin" width="250" height="299"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I still barely know how &lt;code&gt;fabric.Line&lt;/code&gt; works under the hood. And I don't care. It works.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Iteration: The "Yes, And..." Technique
&lt;/h2&gt;

&lt;p&gt;UI Vibe Coding isn't about getting it right instantly; it's about sculpting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Ugly Phase:&lt;/strong&gt;&lt;br&gt;
The first version looked like a programmer made it (because a programmer &lt;em&gt;did&lt;/em&gt; make it). The stick figure looked like a dead bug. The background was gray.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The "Vibe" Phase:&lt;/strong&gt;&lt;br&gt;
Me: &lt;em&gt;"This looks depressing. Make it 'Zen'. Use soft colors, rounded buttons, and a clean layout."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The AI generated the CSS variables (&lt;code&gt;--highlight-color: #88b04b&lt;/code&gt;), added a "Save As" modal, and cleaned up the toolbar.&lt;/p&gt;


  &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwiaddnxyt14usfea2l1p.png" width="800" height="495"&gt;Yoga Pose Builder GUI
  


&lt;p&gt;&lt;strong&gt;The "Feature Creep" Phase:&lt;/strong&gt;&lt;br&gt;
Me: &lt;em&gt;"I want to save my poses."&lt;/em&gt;&lt;br&gt;
AI: &lt;em&gt;"We have no database."&lt;/em&gt;&lt;br&gt;
Me: &lt;em&gt;"Just write JSON files to a folder on the server."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In 5 minutes, we had a fully working persistence layer. No database migrations, just &lt;code&gt;fs.writeFile&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here is a example of such a Pose JSON file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"meta"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"nameFR"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Demi-Pont"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"nameSK"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Setu Bandhasana"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"joints"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"head"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-120&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"neck"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"chest"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-60&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"hips"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"lShoulder"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"lElbow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"lHand"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"rShoulder"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"rElbow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;-20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"rHand"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"lHip"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"lKnee"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"lFoot"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"rHip"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"rKnee"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"rFoot"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"x"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;80&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"y"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  5. The Pivot: Language as a Feature
&lt;/h2&gt;

&lt;p&gt;At the end of the session, I realized a problem: the app was vibing in French (my native tongue), but I wanted screenshots in English for this article. &lt;/p&gt;

&lt;p&gt;Instead of manually editing labels, I asked the AI to "make the whole app i18n." In one single refactor, we added a translation dictionary, a language switcher, and logic to dynamically swap every label, tooltip, and even the pose names in the library. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8vpg0r2qlr1xkx2qsfna.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8vpg0r2qlr1xkx2qsfna.png" width="800" height="495"&gt;&lt;/a&gt;&lt;/p&gt;&lt;br&gt;GUI (and data) in French
  &lt;p&gt;&lt;/p&gt;

&lt;p&gt;This turned a linguistic hurdle into a core feature, proving that with Vibe Coding, "changing your mind" is just a prompt away.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The "Traceability" Hack
&lt;/h2&gt;

&lt;p&gt;We spent about 90 minutes building this. We added features, fixed bugs, and refactored code. By the end, the chat context was massive and messy.&lt;/p&gt;

&lt;p&gt;If I came back to this project in a week, I’d be lost.&lt;/p&gt;

&lt;p&gt;So, I ran one final "Meta-Prompt":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Read all the code we wrote and the docs in &lt;code&gt;Docs/&lt;/code&gt;, and generate a &lt;code&gt;Docs/session_summary.md&lt;/code&gt;. Explain what we built, why we made these choices, and the current state of the app."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The AI analyzed &lt;em&gt;its own work&lt;/em&gt; and wrote a summary file. This is my "Save Game" point. When I want to work on this again, I’ll feed that summary to the AI to restore its context instantly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;We went from a blank folder to a functional, vector-based SVG editor with a backend in one session.&lt;/p&gt;

&lt;p&gt;Vibe Coding a UI is possible, but you have to change your approach:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Anchor the Context:&lt;/strong&gt; Write specs so the AI has a "North Star."&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Delegate the Heavy Lifting:&lt;/strong&gt; Let the AI choose the libraries and do the math.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Iterate Visually:&lt;/strong&gt; Don't try to prompt the perfect UI. Prompt the &lt;em&gt;skeleton&lt;/em&gt;, then prompt the &lt;em&gt;paint&lt;/em&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Next we'll try to &lt;a href="https://dev.to/worldlinetech/vibe-coding-one-slice-at-a-time-4n3p"&gt;Vibe Code a real full stack app&lt;/a&gt;. Or a game. Who knows? The prompt is the limit.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqgj2vy5ypv410dvds2d8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqgj2vy5ypv410dvds2d8.png" width="450" height="423"&gt;&lt;/a&gt;&lt;/p&gt;&lt;br&gt;SVG exported by Yoga Pose Builder (opened in Inkscape)
  &lt;p&gt;&lt;/p&gt;

</description>
      <category>vibecoding</category>
      <category>uidesign</category>
      <category>gemini</category>
    </item>
    <item>
      <title>Vibe Coding One Page at a Time</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Fri, 23 Jan 2026 14:45:20 +0000</pubDate>
      <link>https://dev.to/worldlinetech/vibe-coding-one-page-at-a-time-265j</link>
      <guid>https://dev.to/worldlinetech/vibe-coding-one-page-at-a-time-265j</guid>
      <description>&lt;p&gt;&lt;em&gt;Building a Smart Magazine Archiver&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I’m starting a new series called &lt;strong&gt;"Vibe Coding one Step at a Time."&lt;/strong&gt; The goal? To document the raw, messy, and surprisingly efficient process of building software in the age of AI. We’re not here to write perfect specs or obsess over UML diagrams (well, not yet). We’re here to vibe with the code, iterating on pure intent until the machine does exactly what we want.&lt;/p&gt;

&lt;p&gt;In this first edition, I’m sharing how I used the &lt;strong&gt;Gemini CLI&lt;/strong&gt; to build a tool I actually needed, learning some pretty cool image processing tricks along the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is "Vibe Coding"?
&lt;/h2&gt;

&lt;p&gt;I’m going to claim this term right here: &lt;strong&gt;Vibe Coding&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It’s not "lazy coding." It’s &lt;strong&gt;intent-driven development&lt;/strong&gt;. In the old days, if you wanted to build a script, you had to know the syntax, the libraries, and the edge cases before you even opened your editor. You had to &lt;em&gt;think in code&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Vibe Coding flips that. You &lt;em&gt;think in outcomes&lt;/em&gt;. You describe the behavior, the "vibe" of the feature, and the AI handles the implementation details. You act less like a bricklayer and more like a conductor. The feedback loop isn't "Write -&amp;gt; Compile -&amp;gt; Error," it's "Ask -&amp;gt; Observe -&amp;gt; Tweak."&lt;/p&gt;

&lt;h2&gt;
  
  
  The Use Case: "I Just Want to Read Offline"
&lt;/h2&gt;

&lt;p&gt;Here’s the situation: I subscribe to a fantastic niche magazine (which shall remain nameless to protect the innocent). It’s great, but their "digital reader" is a nightmare. It’s one of those web-based page-turners that requires an active internet connection.&lt;/p&gt;

&lt;p&gt;I wanted to read it on my tablet, offline, on a plane, without waiting for high-res JPEGs to buffer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Problem:&lt;/strong&gt; There was no "Download PDF" button.&lt;br&gt;
&lt;strong&gt;The Clue:&lt;/strong&gt; Inspecting the network traffic revealed that the magazine was just serving a sequence of high-quality images, one URL per page.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Mission:&lt;/strong&gt; Write a script to fetch these pages and stitch them into a single, high-quality, searchable PDF.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Process: Galloping Toward Complexity
&lt;/h2&gt;

&lt;p&gt;We didn't sit down and architect a solution. We started small and let the script evolve.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: The Naive Loop
&lt;/h3&gt;

&lt;p&gt;We started with a simple hypothesis: "The URLs probably just have a page number in them."&lt;br&gt;
I asked Gemini to write a script using &lt;code&gt;requests&lt;/code&gt; to hit the URL for page 1, then page 2.&lt;br&gt;
&lt;em&gt;Boom.&lt;/em&gt; It worked. We had a directory full of 100 separate JPGs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: The Picture Book
&lt;/h3&gt;

&lt;p&gt;Having 100 files is annoying. I wanted a book.&lt;br&gt;
We asked Gemini to "glue these together." It pulled in the &lt;code&gt;PIL&lt;/code&gt; (&lt;a href="https://pillow.readthedocs.io" rel="noopener noreferrer"&gt;Pillow&lt;/a&gt;) library.&lt;br&gt;
&lt;strong&gt;Result:&lt;/strong&gt; A massive PDF. It looked great, but it was dumb. It was just a container of pictures. You couldn't highlight text, search for keywords, or copy-paste quotes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: The Search for Meaning (OCR)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fknav94wd05wdhh5nojvv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fknav94wd05wdhh5nojvv.png" alt="Tesseract OCR" width="330" height="146"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is where the "vibe" got technical. I realized a "picture book" wasn't enough. I needed &lt;strong&gt;Optical Character Recognition (OCR)&lt;/strong&gt;.&lt;br&gt;
We decided to use &lt;a href="https://github.com/tesseract-ocr" rel="noopener noreferrer"&gt;Tesseract&lt;/a&gt;. But here’s the catch we discovered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Human Eyes&lt;/strong&gt; like soft colors and smooth anti-aliasing.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;OCR Engines&lt;/strong&gt; like harsh contrast, jagged edges, and black-and-white binary inputs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If we optimized the images for the machine, the magazine looked ugly. If we kept them pretty, the machine couldn't read the text.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Technical Deep Dive: The "PDF Sandwich"
&lt;/h2&gt;

&lt;p&gt;This is where the magic happened. We ended up building a &lt;strong&gt;PDF Sandwich&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5fw2zzt7ebjs57pp9qpa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F5fw2zzt7ebjs57pp9qpa.png" alt="Me asking Gemini CLI for a sandwich" width="800" height="129"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Instead of choosing between beauty and brains, we chose both.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;The Visual Layer:&lt;/strong&gt; We keep the original high-res color JPEGs. This is what you see.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;The Data Layer:&lt;/strong&gt; Behind the scenes, we create a "Frankenstein" version of the page—converted to grayscale, contrast cranked up to 2.0, and upscaled 2x using &lt;code&gt;LANCZOS&lt;/code&gt; resampling (a fancy &lt;a href="https://en.wikipedia.org/wiki/Lanczos_resampling" rel="noopener noreferrer"&gt;algorithm&lt;/a&gt; that keeps edges sharp).&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;The Merge:&lt;/strong&gt; We feed the Frankenstein images to Tesseract to generate an invisible text layer, then use &lt;code&gt;pypdf&lt;/code&gt; to overlay that text exactly on top of the pretty images.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The trickiest part? &lt;strong&gt;Math.&lt;/strong&gt;&lt;br&gt;
Because we upscaled the OCR images by 2x to help Tesseract read small fonts, the invisible text layer was twice as big as the visual page. We had to calculate scale factors to shrink the text back down so that when you highlight a sentence, the highlight actually lines up with the words.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;Vibe coding this script taught me more in an hour than I’d usually learn in a weekend of reading docs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Image Optimization:&lt;/strong&gt; OCR is picky. Simply resizing an image isn't enough; the &lt;em&gt;method&lt;/em&gt; of resizing (resampling filter) matters.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Library Specialization:&lt;/strong&gt; &lt;code&gt;PIL&lt;/code&gt; is for pixels; &lt;code&gt;pypdf&lt;/code&gt; is for structure. Trying to do everything in one library is a trap.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Power of the CLI:&lt;/strong&gt; Using the Gemini CLI meant I didn't have to context-switch. I stayed in my terminal, describing what I wanted, and the code appeared.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy2dgv0fe30dhuu4zf12u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy2dgv0fe30dhuu4zf12u.png" alt="Use of the script (for 2 pages)" width="800" height="301"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;We ended up with a ~100-line Python script that solves a genuine daily frustration. I didn't have to memorize the &lt;code&gt;pypdf&lt;/code&gt; documentation or look up the Tesseract CLI flags. I just focused on the goal: "Make it searchable, make it pretty."&lt;/p&gt;

&lt;p&gt;That’s Vibe Coding. You bring the vision, the AI brings the syntax, and together you build something cool. &lt;/p&gt;

&lt;p&gt;&lt;em&gt;We'll discover in the &lt;a href="https://dev.to/worldlinetech/vibe-coding-one-pixel-at-a-time-22pc"&gt;next episode&lt;/a&gt; if this is still true with a more complex use case and a GUI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vibecoding</category>
      <category>gemini</category>
      <category>pdf</category>
      <category>ocr</category>
    </item>
    <item>
      <title>The Ultimate LLM Inference Battle: vLLM vs. Ollama vs. ZML</title>
      <dc:creator>raphiki</dc:creator>
      <pubDate>Mon, 29 Dec 2025 09:12:46 +0000</pubDate>
      <link>https://dev.to/worldlinetech/the-ultimate-llm-inference-battle-vllm-vs-ollama-vs-zml-m97</link>
      <guid>https://dev.to/worldlinetech/the-ultimate-llm-inference-battle-vllm-vs-ollama-vs-zml-m97</guid>
      <description>&lt;p&gt;&lt;em&gt;A structured, data-driven comparison of today's leading open-source engines for serving AI models.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The "Runtime Wars"
&lt;/h3&gt;

&lt;p&gt;The open-source AI community has achieved an incredible milestone: models like Meta's Llama 3 and Mistral AI's Mixtral now rival proprietary giants like GPT-4. But having the weights is only half the battle. To actually &lt;em&gt;use&lt;/em&gt; these models—to build a chatbot, an agent, or an API, you need an inference engine.&lt;/p&gt;

&lt;p&gt;The landscape of inference servers is exploding. A year ago, options were scarce. Today, developers are faced with a paralyzing array of choices. Should you use the industry darling &lt;strong&gt;vLLM&lt;/strong&gt;? The local developer's favorite, &lt;strong&gt;Ollama&lt;/strong&gt;? Or perhaps a radical newcomer like &lt;strong&gt;ZML&lt;/strong&gt;?&lt;/p&gt;

&lt;p&gt;Choosing the wrong engine can lead to massive infrastructure bills, slow user experiences, or vendor lock-in.&lt;/p&gt;

&lt;p&gt;To cut through the hype, we are applying the &lt;strong&gt;QSOS (Qualification and Selection of Open Source software)&lt;/strong&gt; method. This isn't a casual review; it's a structured evaluation comparing these three contenders against the state-of-the-art features required for modern AI production.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Methodology: Why QSOS?
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ftkge5m9dy66je4atphit.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ftkge5m9dy66je4atphit.png" alt="QSOS Logo" width="257" height="100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.qsos.org" rel="noopener noreferrer"&gt;QSOS&lt;/a&gt; is a standardized methodology designed to reduce the risks associated with adopting open-source technologies. Unlike ad-hoc selection processes based on Medium articles or GitHub stars, QSOS treats open-source evaluation with the same rigor used for proprietary software.&lt;/p&gt;

&lt;p&gt;The core philosophy of QSOS is separating &lt;strong&gt;Evaluation&lt;/strong&gt; (the intrinsic, objective quality of the software) from &lt;strong&gt;Qualification&lt;/strong&gt; (how well it fits your specific business needs).&lt;/p&gt;

&lt;p&gt;For this comparison, we used a "Best of Breed" evaluation grid, scoring features on a simple 0-to-2 scale:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;0:&lt;/strong&gt; Not covered / Non-existent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1:&lt;/strong&gt; Partially covered / Complex implementation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2:&lt;/strong&gt; Fully covered / Best-in-class standard.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We assessed four key axes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Maturity &amp;amp; Community:&lt;/strong&gt; Is the project stable and likely to survive?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Functional Features:&lt;/strong&gt; Does it support modern requirements like LoRA adapters and quantization?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance &amp;amp; Scale:&lt;/strong&gt; Can it handle high throughput and utilize hardware efficiently?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operations (Day 2):&lt;/strong&gt; How easy is it to deploy, monitor, and maintain?&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  The Contenders
&lt;/h3&gt;

&lt;h4&gt;
  
  
  1. vLLM: The Data Center Standard
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff8pv4ovqry11gr3xumeg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff8pv4ovqry11gr3xumeg.png" alt="vLLM Logo" width="239" height="100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://vllm.ai" rel="noopener noreferrer"&gt;vLLM&lt;/a&gt;&lt;/strong&gt; burst onto the scene in 2023 from UC Berkeley, solving a critical bottleneck in serving LLMs: memory fragmentation. Its core innovation, &lt;strong&gt;PagedAttention&lt;/strong&gt;, allows it to manage GPU memory like an operating system manages virtual memory, dramatically increasing batch sizes and throughput.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Primary Focus:&lt;/strong&gt; High-throughput production serving in the data center.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Positioning:&lt;/strong&gt; vLLM is the currently the &lt;strong&gt;De Facto Standard&lt;/strong&gt; for enterprise deployment. It excels on server-grade hardware (NVIDIA H100s/A100s) and offers the richest feature set for scaling.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  2. Ollama: The Developer's Best Friend
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fc8i7p1zgrhls5qacwqki.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fc8i7p1zgrhls5qacwqki.png" alt="Ollama Logo" width="344" height="150"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt;&lt;/strong&gt; took a different approach. It focused entirely on removing friction. By wrapping the powerful &lt;code&gt;llama.cpp&lt;/code&gt; engine in a sleek, Docker-style Go binary, it made running a 70B parameter model on a MacBook as easy as typing &lt;code&gt;ollama run llama3&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Primary Focus:&lt;/strong&gt; Local development, edge devices, and consumer hardware (Mac/PC).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Positioning:&lt;/strong&gt; Ollama is the king of &lt;strong&gt;usability&lt;/strong&gt;. It is unbeaten for local testing and running models on consumer hardware, but it lacks the advanced scheduling required for high-traffic enterprise production.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  3. ZML (Zig Machine Learning): The Radical Challenger
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3ddw41ql12g8k4ekellb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F3ddw41ql12g8k4ekellb.png" alt="ZML Logo" width="200" height="197"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://zml.ai" rel="noopener noreferrer"&gt;ZML&lt;/a&gt;&lt;/strong&gt; is the new kid on the block. It is less of a "server" product and more of a compiler stack aimed at engineers. Written in Zig, it utilizes OpenXLA/MLIR to compile model graphs directly into standalone binaries, aiming to eliminate the heavy Python/PyTorch dependency chain entirely.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Primary Focus:&lt;/strong&gt; High-performance, cross-platform runtime (TPUs, AMD, NVIDIA) without dependencies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Positioning:&lt;/strong&gt; ZML is an &lt;strong&gt;Alpha-stage visionary&lt;/strong&gt;. It offers incredible potential for hardware portability and efficiency but is currently a complex "build-your-own-stack" tool rather than a drop-in product.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Visualizing the Results
&lt;/h3&gt;

&lt;p&gt;To understand how these tools differ, we visualize our QSOS scores using two different schemas.&lt;/p&gt;

&lt;h4&gt;
  
  
  The Radar Chart: Feature Balance
&lt;/h4&gt;

&lt;p&gt;This chart shows the balance of strengths across the four evaluation axes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdlvb6g592d0kydxj0try.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdlvb6g592d0kydxj0try.png" alt="QSOS Radar" width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caption: The QSOS Radar Chart highlights the distinct profiles of the three engines. vLLM shows the broadest coverage across features and performance. Ollama spikes toward Operational Ease. ZML shows potential in features but lacks maturity.&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;vLLM (Blue):&lt;/strong&gt; The largest, most balanced area, indicating strength across maturity, features, and performance, with moderate operational complexity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ollama (Green):&lt;/strong&gt; A massive spike toward "Operational Ease," reflecting its zero-friction user experience, but pulling back on raw performance metrics like continuous batching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ZML (Red):&lt;/strong&gt; A smaller footprint overall, reflecting its early stage (low maturity), but showing strong potential in functional features due to its compiler-based architecture.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  The QSOS Quadrant: Market Position
&lt;/h4&gt;

&lt;p&gt;This schema maps the tools based on their market adoption versus their raw production capabilities.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frm90qnkq9c5tf3hld553.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frm90qnkq9c5tf3hld553.png" alt="QSOS Quadrant" width="800" height="640"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Caption: The QSOS Quadrant positions the tools based on Market Maturity vs. Production Power.&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;vLLM (The Leader):&lt;/strong&gt; High Maturity, High Power. The safe, scalable choice for the enterprise.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ollama (The Specialist):&lt;/strong&gt; High Maturity, Lower Production Power. The standard for a specific niche (local/consumer hardware), prioritizing usability over scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ZML (The Visionary):&lt;/strong&gt; Low Maturity, High Potential Power. An innovative approach that hasn't yet proven itself in the broad market.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Consolidated Score Sheet
&lt;/h3&gt;

&lt;p&gt;Below is the detailed breakdown of the evaluation scores that feed the charts above.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Section / Criteria&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;vLLM&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Ollama&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;ZML (Zig ML)&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A. MATURITY&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;History &amp;amp; Age&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Standard)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Standard)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Very New)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Activity&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Hyper-Active)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Viral)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (High Velocity)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ecosystem&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Dominant)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Ubiquitous)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Niche)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Governance&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Community)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Company Led)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Small Team)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;B. FEATURES&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model Support&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Universal)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Curated Lib)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Compiler based)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantization&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Server: AWQ/FP8)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Edge: GGUF)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Implicit XLA)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LoRA Adapters&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Dynamic Multi-LoRA)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Static Modelfile)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Not standard)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API Compat.&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (OpenAI Native)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (OpenAI Native)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Runtime only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;C. PERFORMANCE&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cont. Batching&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Gold Standard)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (FIFO)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Arch. support)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Throughput&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Maximum SOTA)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Low/Single User)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (High Potential)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parallelism&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Tensor &amp;amp; Pipeline)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Single Node)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Compiler Config)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hardware Agnosticism&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (NVIDIA Centric)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Apple/Consumer)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Any: TPU/AMD)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;D. OPERATIONS&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ease of Setup&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Python/Docker)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Magic 1-Click)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Hard: Bazel)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dependencies&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Heavy Torch)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Zero: Go Binary)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Zero: Zig Binary)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2&lt;/strong&gt; (Prometheus Native)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0&lt;/strong&gt; (Logs only)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1&lt;/strong&gt; (Manual metrics)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Conclusion
&lt;/h3&gt;

&lt;p&gt;There is no single "best" inference engine. The right choice depends entirely on your specific context (the Qualification phase of QSOS).&lt;/p&gt;

&lt;h4&gt;
  
  
  Choose vLLM if:
&lt;/h4&gt;

&lt;p&gt;You are building a production application that needs to serve many concurrent users. You have access to server-grade GPUs (NVIDIA A10G, A100, H100) and need features like dynamic LoRA adapters for multi-tenancy.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;If you are deploying to Kubernetes to serve customers, start here.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Choose Ollama if:
&lt;/h4&gt;

&lt;p&gt;You are a developer building locally on a Mac or Windows PC. You need a zero-friction way to test models, or you are deploying to edge devices where resources are constrained, and concurrency is low.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;If you just want to run Llama 3 on your laptop right now, download Ollama.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4&gt;
  
  
  Choose ZML if:
&lt;/h4&gt;

&lt;p&gt;You are an ML systems engineer building a specialized hardware appliance (e.g., using TPUs or AMD chips) and need a runtime with absolutely zero Python dependencies and a tiny footprint. You are willing to build the server infrastucture around it yourself.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;If you are frustrated by PyTorch bloat and want a "build your own" adventure, look at ZML.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Note on Methodology
&lt;/h3&gt;

&lt;p&gt;For the purpose of this article, we utilized a &lt;strong&gt;simplified QSOS evaluation grid&lt;/strong&gt;. We intentionally zoomed in on the "Best of Breed" criteria, the critical differentiators driving the current "Inference Wars", to keep the comparison readable and actionable.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;full-fledged QSOS evaluation&lt;/strong&gt; is significantly more exhaustive. It is structured as a hierarchical &lt;strong&gt;tree of criteria&lt;/strong&gt; containing more data points, covering deep operational details such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Generic Attributes:&lt;/strong&gt; Intellectual property management, roadmap visibility, bug tracking efficiency, and internationalization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specific Sub-sections:&lt;/strong&gt; Detailed granularity on security compliance (SOC2/GDPR), exact memory footprints, and specific driver version compatibility.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;While this article provides a strategic overview, a complete QSOS audit would involve drilling down from high-level "Sections" into specific "Leaves" to calculate a precise, weighted score for every possible business constraint.&lt;/p&gt;

</description>
      <category>qsos</category>
      <category>zml</category>
      <category>ollama</category>
      <category>vllm</category>
    </item>
  </channel>
</rss>
