<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Amariah Kamau</title>
    <description>The latest articles on DEV Community by Amariah Kamau (@abishaiama).</description>
    <link>https://dev.to/abishaiama</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3394455%2F903b6008-0825-4cf9-b88b-bdaba42c4258.png</url>
      <title>DEV Community: Amariah Kamau</title>
      <link>https://dev.to/abishaiama</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/abishaiama"/>
    <language>en</language>
    <item>
      <title>Add a Coding Agent to Zed and JetBrains IDEs With the Agent Client Protocol</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:21:41 +0000</pubDate>
      <link>https://dev.to/abishaiama/add-a-coding-agent-to-zed-and-jetbrains-ides-with-the-agent-client-protocol-4o08</link>
      <guid>https://dev.to/abishaiama/add-a-coding-agent-to-zed-and-jetbrains-ides-with-the-agent-client-protocol-4o08</guid>
      <description>&lt;p&gt;Zed and JetBrains IDEs both speak the Agent Client Protocol (ACP): an open protocol between an editor and a coding agent, so the editor provides the chat, the diffs and the approvals, and any agent that implements the protocol can sit behind them. No plugin is involved. Atlarix implements it with one command, &lt;code&gt;atlarix acp&lt;/code&gt;. Here is how to add it to each editor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install the Atlarix CLI
&lt;/h2&gt;

&lt;p&gt;ACP editors start the agent as a program, so first install the CLI: &lt;code&gt;npm i -g atlarix&lt;/code&gt;, &lt;code&gt;brew install amariahak/atlarix/atlarix&lt;/code&gt;, or the one-line installers on &lt;a href="https://www.atlarix.dev/cli" rel="noopener noreferrer"&gt;the CLI page&lt;/a&gt;. Check it with &lt;code&gt;atlarix --version&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zed
&lt;/h2&gt;

&lt;p&gt;Open Zed's settings (&lt;code&gt;~/.config/zed/settings.json&lt;/code&gt;) and add Atlarix as an agent server: &lt;code&gt;"agent_servers": { "Atlarix": { "type": "custom", "command": "atlarix", "args": ["acp"] } }&lt;/code&gt;. Then open the Agent Panel, start a new thread and choose Atlarix. Sign in when it asks; it opens your browser once.&lt;/p&gt;

&lt;h2&gt;
  
  
  JetBrains IDEs
&lt;/h2&gt;

&lt;p&gt;In IntelliJ IDEA, PyCharm, WebStorm, GoLand or any other JetBrains IDE with AI Assistant, open the chat's agent menu and choose &lt;strong&gt;More Agents → Add Custom Agent&lt;/strong&gt;. That opens &lt;code&gt;~/.jetbrains/acp.json&lt;/code&gt;; add &lt;code&gt;"agent_servers": { "Atlarix": { "command": "/path/to/atlarix", "args": ["acp"] } }&lt;/code&gt; with the full path from &lt;code&gt;which atlarix&lt;/code&gt; (macOS, Linux) or &lt;code&gt;where atlarix&lt;/code&gt; (Windows). You do not need a JetBrains AI subscription to use an ACP agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you get in the editor
&lt;/h2&gt;

&lt;p&gt;Atlarix's modes — Build, Explore, Plan, Debug, Review — and its models appear in the editor's own mode and model pickers. Edits and commands come to you as the editor's permission prompts before anything changes. In Plan mode the plan shows as the editor's plan list, and building it switches the mode picker to Build. File references in replies open at the line. Past chats load from the editor's history.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signing in
&lt;/h2&gt;

&lt;p&gt;Atlarix offers two sign-in methods to an ACP editor: a browser sign-in the editor can start for you, and &lt;code&gt;atlarix login&lt;/code&gt; in a terminal. Either signs in the CLI on that computer, and the editor's chats then use your account, your models and your own keys.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same agent everywhere
&lt;/h2&gt;

&lt;p&gt;A chat started in Zed or a JetBrains IDE is an Atlarix session like any other, so it shows up in the Atlarix app and CLI, and other Atlarix sessions can message it. The editor is just one more place to type.&lt;/p&gt;

&lt;p&gt;We have also submitted Atlarix to the ACP registry, which Zed and JetBrains read to offer agents with one click; until that is merged, the settings above are the way in. For VS Code, Cursor, Kiro and Windsurf, which do not speak ACP, Atlarix has its own extension — see the &lt;a href="https://www.atlarix.dev/docs/guides/editors" rel="noopener noreferrer"&gt;editors guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.atlarix.dev/blogs/add-coding-agent-zed-jetbrains-agent-client-protocol" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>productivity</category>
      <category>programming</category>
    </item>
    <item>
      <title>How to Use Your Own API Key With an AI Coding Agent in VS Code</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:20:51 +0000</pubDate>
      <link>https://dev.to/abishaiama/how-to-use-your-own-api-key-with-an-ai-coding-agent-in-vs-code-19bg</link>
      <guid>https://dev.to/abishaiama/how-to-use-your-own-api-key-with-an-ai-coding-agent-in-vs-code-19bg</guid>
      <description>&lt;p&gt;A coding agent on your own API key means three things: you choose the model, you pay the provider's price with no markup, and your requests go to the provider you picked. This is how to set that up in VS Code — or Cursor, Kiro or Windsurf — with the Atlarix extension, start to finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Install the extension
&lt;/h2&gt;

&lt;p&gt;In your editor's Extensions view, search for &lt;strong&gt;Atlarix&lt;/strong&gt; and install it — it comes from the &lt;a href="https://marketplace.visualstudio.com/items?itemName=atlarix.atlarix" rel="noopener noreferrer"&gt;Visual Studio Marketplace&lt;/a&gt; in VS Code and from &lt;a href="https://open-vsx.org/extension/atlarix/atlarix" rel="noopener noreferrer"&gt;Open VSX&lt;/a&gt; in Cursor, Kiro and Windsurf. It brings its own copy of the agent, so there is no CLI or runtime to install separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Sign in once
&lt;/h2&gt;

&lt;p&gt;Open the Atlarix view (the ant in the activity bar, or ⌘⌥A / Ctrl+Alt+A) and sign in. An Atlarix account is free; it is what keeps your chats, settings and keys the same across the editor, the desktop app and the CLI.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Connect your provider
&lt;/h2&gt;

&lt;p&gt;Open the ⋯ menu at the top of the chat and choose &lt;strong&gt;Account, usage and billing&lt;/strong&gt;. Under &lt;strong&gt;Your own keys&lt;/strong&gt;, press &lt;strong&gt;Connect a provider…&lt;/strong&gt;, pick the provider — OpenAI, Anthropic, Google, DeepSeek, OpenRouter, or another from the list — and paste the key. The prompt links to the page where that provider issues keys. The key is stored by Atlarix on your computer, and the Atlarix app and CLI on the same computer use it too.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Pick the model
&lt;/h2&gt;

&lt;p&gt;The model picker sits just under the message box. Your key's models now appear in it alongside Atlarix Auto and the Atlarix Core models; choose one and it applies to this chat. Different chats can use different models — a fast cheap one for questions, a stronger one for a refactor.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Work as normal
&lt;/h2&gt;

&lt;p&gt;Every mode works on your own key: Build, Explore, Plan, Debug and Review. Edits open in the editor's diff for your Yes or No, commands wait for approval, and plans tick off as they are built. The account panel shows what this chat has cost on your key so far.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and where requests go
&lt;/h2&gt;

&lt;p&gt;Atlarix adds nothing to your own key's usage: you pay the provider exactly what they charge. Requests on your key go from your computer to the provider, not through Atlarix. If the provider rejects the key — revoked, mistyped or out of credit — the chat says so, so you know to fix the key rather than retry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which model to start with
&lt;/h2&gt;

&lt;p&gt;There is no single answer, and the honest advice is to try two. Frontier models from OpenAI, Anthropic and Google are the strongest at long multi-step changes; DeepSeek and GLM are far cheaper per token and very capable for most everyday work; OpenRouter gives you one key for many providers. Because the picker is per chat, comparing them on your own code costs one extra chat.&lt;/p&gt;

&lt;p&gt;The same key works in Zed and JetBrains IDEs, which run Atlarix over the Agent Client Protocol — see the &lt;a href="https://www.atlarix.dev/docs/guides/editors" rel="noopener noreferrer"&gt;editors guide&lt;/a&gt;. And if you would rather not manage keys at all, Atlarix Auto in the same picker chooses a model for each step for you.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.atlarix.dev/blogs/use-your-own-api-key-ai-coding-agent-vs-code" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vscode</category>
      <category>ai</category>
      <category>tutorial</category>
      <category>llm</category>
    </item>
    <item>
      <title>Atlarix in VS Code, Cursor, Kiro and Windsurf: An AI Coding Agent That Uses Any Model</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:19:52 +0000</pubDate>
      <link>https://dev.to/abishaiama/atlarix-in-vs-code-cursor-kiro-and-windsurf-an-ai-coding-agent-that-uses-any-model-3p8i</link>
      <guid>https://dev.to/abishaiama/atlarix-in-vs-code-cursor-kiro-and-windsurf-an-ai-coding-agent-that-uses-any-model-3p8i</guid>
      <description>&lt;p&gt;Until this release, using Atlarix meant a window beside your editor: the desktop app, or since 15.0 the CLI in a terminal. With 15.4.0 the agent moves into the editor itself. One extension for VS Code and the editors built on it — Cursor, Kiro, Windsurf, VSCodium — carries the same agent as the app and the CLI, and Zed and JetBrains IDEs run it through the Agent Client Protocol. This post is what you get and how to start.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install it
&lt;/h2&gt;

&lt;p&gt;Search for &lt;strong&gt;Atlarix&lt;/strong&gt; in your editor's Extensions view, or install it from the &lt;a href="https://marketplace.visualstudio.com/items?itemName=atlarix.atlarix" rel="noopener noreferrer"&gt;Visual Studio Marketplace&lt;/a&gt; (VS Code) or &lt;a href="https://open-vsx.org/extension/atlarix/atlarix" rel="noopener noreferrer"&gt;Open VSX&lt;/a&gt; (Cursor, Kiro, Windsurf, VSCodium). There are builds for macOS, Windows and Linux on Intel and ARM, and the extension carries its own copy of the Atlarix CLI, so there is nothing else to install. Open the Atlarix view (the ant in the activity bar, or ⌘⌥A / Ctrl+Alt+A) and sign in the first time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every mode, any model
&lt;/h2&gt;

&lt;p&gt;The message box has the mode: &lt;strong&gt;Build&lt;/strong&gt; changes code and runs commands, &lt;strong&gt;Explore&lt;/strong&gt; reads and answers without changing anything, &lt;strong&gt;Plan&lt;/strong&gt; writes a plan you can read before anything happens, &lt;strong&gt;Debug&lt;/strong&gt; chases a failure to its cause, &lt;strong&gt;Review&lt;/strong&gt; reads a change and reports what is wrong with it. Just under it is the model: Atlarix Auto, which picks a model for each step, the Atlarix Core models, or any model from your own API key — OpenAI, Anthropic, Google, DeepSeek, OpenRouter and the rest. The picker is the same list the app and CLI show.&lt;/p&gt;

&lt;h2&gt;
  
  
  Edits you approve in the editor's own diff
&lt;/h2&gt;

&lt;p&gt;When the agent wants to change a file, the editor opens its own side-by-side diff with the proposed version before anything is written. &lt;strong&gt;Yes&lt;/strong&gt; applies it, &lt;strong&gt;No&lt;/strong&gt; leaves the file alone, and &lt;strong&gt;Always&lt;/strong&gt; stops asking for that kind of change in this workspace. Commands wait for you the same way. You are reading the change in the tool you already trust for reading changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plans you can watch being built
&lt;/h2&gt;

&lt;p&gt;In Plan mode the plan appears in its own panel with a checklist of steps. &lt;strong&gt;Build it&lt;/strong&gt; switches to Build and carries it out, and the panel stays on screen while it works: each step goes from ☐ to ◐ to ☑, the header counts them off, and it says Built at the end. You can see where the work is without reading the transcript.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context without copy-paste
&lt;/h2&gt;

&lt;p&gt;Select code and press ⌘⌥L (Ctrl+Alt+L) to add it to the chat with its file and lines; type @ to attach a file; the + button attaches images. The model's thinking streams into a fold above its answer, so you can watch it reason and close it when the answer starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your Chrome, your other sessions
&lt;/h2&gt;

&lt;p&gt;With the Atlarix Chrome extension installed, the agent in your editor can open pages in your own Chrome — signed in as you — to test what it just built. Every Atlarix session can message every other one: the app, the CLI, Chrome or another editor can hand your editor a task, and it arrives as a card and is answered there. &lt;code&gt;/editor&lt;/code&gt; in the CLI moves a terminal chat into your editor; the ⋯ menu in the editor moves it back to the app or a terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your account in the panel
&lt;/h2&gt;

&lt;p&gt;⋯ → &lt;strong&gt;Account, usage and billing&lt;/strong&gt; shows your plan, credit and this chat's usage, with Get Pro, Add credit and Sign out; your own provider keys; and the MCP servers in this folder, which you can test, switch off or add. Come back after a few minutes away and one line recaps where the chat stands; turn that off in the same menu.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zed and JetBrains IDEs
&lt;/h2&gt;

&lt;p&gt;Zed and JetBrains IDEs speak the Agent Client Protocol natively, so they run Atlarix's CLI as their agent with no plugin: &lt;code&gt;atlarix acp&lt;/code&gt;. Building a plan switches the editor's mode picker, the plan shows as the editor's own plan list, and file links open at the line. The &lt;a href="https://www.atlarix.dev/docs/guides/editors" rel="noopener noreferrer"&gt;editors guide&lt;/a&gt; has the settings for both.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;Installing the extension is free. Your own API keys cost nothing through Atlarix — you pay your provider directly. Atlarix Core models are paid from Pro or pay-as-you-go credit, the same balance as the app and CLI.&lt;/p&gt;

&lt;p&gt;Chats you start in the editor show up in the app and the CLI, and the other way round: it is one agent with one history, wherever you happen to be typing. Install it from the &lt;a href="https://marketplace.visualstudio.com/items?itemName=atlarix.atlarix" rel="noopener noreferrer"&gt;Marketplace&lt;/a&gt; or &lt;a href="https://open-vsx.org/extension/atlarix/atlarix" rel="noopener noreferrer"&gt;Open VSX&lt;/a&gt;, and see the &lt;a href="https://www.atlarix.dev/docs/guides/editors" rel="noopener noreferrer"&gt;editors guide&lt;/a&gt; for everything else.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.atlarix.dev/blogs/atlarix-vs-code-cursor-kiro-windsurf-ai-coding-agent-any-model" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vscode</category>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Atlarix CLI 15.0.1: Sub-Agents in the Terminal, Coloured Code and a Plan You Can Read</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:06:10 +0000</pubDate>
      <link>https://dev.to/abishaiama/atlarix-cli-1501-sub-agents-in-the-terminal-coloured-code-and-a-plan-you-can-read-2hpo</link>
      <guid>https://dev.to/abishaiama/atlarix-cli-1501-sub-agents-in-the-terminal-coloured-code-and-a-plan-you-can-read-2hpo</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.atlarix.dev/blogs/atlarix-cli-sub-agents-coloured-code-readable-plan" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Atlarix CLI shipped in 15.0.0 as the same agent as the desktop app, in a terminal. We then used it for real work: an actual pull request to a well-known open-source project, start to finish, from the terminal. That surfaced the bugs and rough edges a demo never does. Atlarix 15.0.1 is what came out of it, and this post walks through what changed and why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sub-agents work, and they sit under the prompt box
&lt;/h2&gt;

&lt;p&gt;Sub-agents are how an agent explores in parallel: one reads the tests while another traces a function. In 15.0.0 every sub-agent the CLI started failed before it ran and then showed as running forever. Two separate bugs were behind it, and a third turned up while fixing them: outside Atlarix's own repository the CLI could not find the prompts it ships with, so every turn quietly ran without its mode's instructions and every sub-agent failed to start. All three are fixed, with tests that run the CLI from a folder that is not ours.&lt;/p&gt;

&lt;p&gt;While Atlarix works, the space under the prompt box now lists what is running: the main agent and what it is doing, then one row per sub-agent with its current step, how long it has run and its tokens. Finished ones drop off, and background commands share one row. Press ↓ and enter on a sub-agent to watch it live, its tool calls and what it is writing; ← and → step between sub-agents and esc backs out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code is coloured, for the languages you actually read
&lt;/h2&gt;

&lt;p&gt;Answers and edit diffs are now syntax highlighted, and inline code stands out in lavender rather than a dim grey chip. The interesting part is how languages are handled. Atlarix does not ship a list of grammars. The first time a language shows up, its grammar is found by name: one already on your machine from another terminal app built on the same toolkit is reused, otherwise the language's own grammar project is fetched once. Only the languages you read ever land on disk, and offline the code simply stays plain.&lt;/p&gt;

&lt;h2&gt;
  
  
  A plan you can read while you decide
&lt;/h2&gt;

&lt;p&gt;Plan mode used to end with a popup that asked you to approve a plan you could not see. Now the plan stays pinned above the prompt box while it waits to be built. You read it, type what to change underneath, and watch it redraw as Atlarix revises it. Build it, or /build, starts the work; esc hides the plan and /plan brings it back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sending is instant
&lt;/h2&gt;

&lt;p&gt;The first message of a session used to freeze the terminal for about a second before it appeared. The cause was token counting, which rebuilt its tokenizer every time it counted, dozens of times on a first turn. It is built once now. The desktop app got faster from the same fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The small things that add up
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The context meter shows how full the context is as a share and as tokens of the model's window, instead of a running total that could pass the window and look like a contradiction.&lt;/li&gt;
&lt;li&gt;A message you send while Atlarix works is drawn where the agent read it, not under the whole reply.&lt;/li&gt;
&lt;li&gt;/compact typed mid-answer waits for the answer to finish, and messages typed after it still go first.&lt;/li&gt;
&lt;li&gt;Pasting the same text again expands the pasted-text chip so you can edit it.&lt;/li&gt;
&lt;li&gt;The approval dialog's diff is coloured like the transcript's.&lt;/li&gt;
&lt;li&gt;The terminal tab pulses while Atlarix works, so a busy session is visible from another tab.&lt;/li&gt;
&lt;li&gt;Long session names are shortened, the prompt box sits a step above the page, and checklists read on one line.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to get it
&lt;/h2&gt;

&lt;p&gt;If you already have the CLI it updates itself. Otherwise: &lt;code&gt;npm install -g atlarix&lt;/code&gt;, &lt;code&gt;brew install amariahak/atlarix/atlarix&lt;/code&gt;, or the one-line installer on &lt;a href="https://www.atlarix.dev/cli" rel="noopener noreferrer"&gt;atlarix.dev/cli&lt;/a&gt;. The CLI shares the desktop app's sessions, models, keys and MCP servers, and runs on your own API keys, local models or Atlarix Core.&lt;/p&gt;

&lt;p&gt;Most of 15.0.1 came from using the CLI on a real task and fixing what got in the way. It also gave us the next one: Atlarix in Slack, which works on your repositories from a message and is coming soon. The full list of changes is in the &lt;a href="https://www.atlarix.dev/changelog" rel="noopener noreferrer"&gt;Atlarix changelog&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cli</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Atlarix Auto: A Free AI Coding Agent That Picks the Model for Each Step</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Wed, 23 Sep 2026 20:57:54 +0000</pubDate>
      <link>https://dev.to/abishaiama/atlarix-auto-a-free-ai-coding-agent-that-picks-the-model-for-each-step-4g1i</link>
      <guid>https://dev.to/abishaiama/atlarix-auto-a-free-ai-coding-agent-that-picks-the-model-for-each-step-4g1i</guid>
      <description>&lt;p&gt;Most AI coding tools ask you to make two decisions before you write a line: which model, and whose API key. Atlarix Auto removes both. It is one option in the model picker — the one new installs start on — and instead of running your whole session on a single model, it chooses a model for each step of the work. On an account with no credit, Auto in the Atlarix desktop app is free. On Atlarix Pro or credit, it is paid and private.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Atlarix Auto is
&lt;/h2&gt;

&lt;p&gt;Atlarix is a desktop coding agent: it plans, edits your code, runs commands in a per-OS sandbox, and runs your tests before it calls anything done, with every change waiting for your approval. Auto is the model setting that decides which language model does each piece of that work. You pick Auto and give it a task.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Auto picks a model, step by step
&lt;/h2&gt;

&lt;p&gt;A coding session is not one kind of work. Reading a codebase, writing a change, planning a refactor and reviewing a diff need very different amounts of reasoning. Auto routes on what each request is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What the step is&lt;/th&gt;
&lt;th&gt;What Auto uses&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exploring and building&lt;/td&gt;
&lt;td&gt;A fast, affordable model that is strong at code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Planning and code review&lt;/td&gt;
&lt;td&gt;A stronger reasoning model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deep thinking switched on&lt;/td&gt;
&lt;td&gt;Every step moves up one level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A step with a screenshot, image or PDF&lt;/td&gt;
&lt;td&gt;A model that can actually read it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sub-agents, session titles, summaries&lt;/td&gt;
&lt;td&gt;The cheapest capable model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A model is rate-limited&lt;/td&gt;
&lt;td&gt;The same level from another provider&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Routing happens per request, not per session: a long task that suddenly needs to look at a screenshot gets a vision-capable model for that one step, then returns to the cheaper one. If a step fails, Atlarix retries it; a "Try with a stronger model" button reruns a hard turn one level up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free Auto, and the trade that makes it free
&lt;/h2&gt;

&lt;p&gt;On an account with no credit, Auto in the desktop app costs nothing. It runs on OpenAI models through OpenAI's data-sharing program: OpenAI provides free daily capacity for traffic it may use to improve its models. That is the whole reason it is free. Before your first free message, Atlarix shows a one-time notice explaining this, and you choose whether to go ahead. Free Auto is available in the desktop app (and to Atlarix Reviewer for public repositories only); the Chrome extension runs on paid Auto.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the free daily allowance works
&lt;/h2&gt;

&lt;p&gt;Each account gets a fair daily share, sized from what is left and how many people are using it, with a floor so everyone gets real work done and a ceiling so no single account drains it. When your share runs out, Auto falls back to a free model rather than cutting you off; only past that does it pause until the daily reset — shown on your own clock.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stays private
&lt;/h2&gt;

&lt;p&gt;With Atlarix Pro or credit, Auto is paid and private: nothing it handles is shared or used to train models. Your own API keys work as before. Local models through Ollama or LM Studio never leave your machine. Atlarix itself never trains on your code, in any mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs when you pay
&lt;/h2&gt;

&lt;p&gt;The app is free — every feature, no gating. Paid inference is Atlarix Pro ($19/month, including Core usage across the desktop app, the Chrome extension and Atlarix Reviewer) or pay-as-you-go credit, with the cost of every reply shown.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Download Atlarix from &lt;a href="https://www.atlarix.dev" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;. New installs start on Auto; if you picked a model before, switch to Auto in the model picker, then give it a real task on your own code.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.atlarix.dev/blogs/atlarix-auto-free-ai-coding-agent-no-api-key" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>tooling</category>
    </item>
    <item>
      <title>The harness, not the model: how to make a weak or local model reliable enough to ship</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Fri, 11 Sep 2026 19:41:05 +0000</pubDate>
      <link>https://dev.to/abishaiama/the-harness-not-the-model-how-to-make-a-weak-or-local-model-reliable-enough-to-ship-2plp</link>
      <guid>https://dev.to/abishaiama/the-harness-not-the-model-how-to-make-a-weak-or-local-model-reliable-enough-to-ship-2plp</guid>
      <description>&lt;p&gt;If you've tried to point a local model — Qwen, a quantized Llama, whatever fits on your GPU — at a real task in a real repo, you already know the feeling. It starts confidently. It edits three files. It announces it's done. And then you run the tests and half of them are red, one of the files it "edited" is byte-for-byte unchanged, and the function it swore it added isn't there.&lt;/p&gt;

&lt;p&gt;The usual conclusion is: the model is too weak. Get a bigger model, or rent a frontier API, and the problem goes away.&lt;/p&gt;

&lt;p&gt;That conclusion is mostly wrong, and it's expensive. A large share of what looks like model weakness is actually &lt;strong&gt;the absence of a system around the model&lt;/strong&gt; — the scaffolding that a frontier model partially compensates for on its own and a weak model does not. If you build that system, a much weaker model becomes usable for work you'd have assumed it couldn't touch. The industry has a name for that system: the &lt;strong&gt;harness&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I've spent the last several months building one — a coding harness called Atlarix that's designed to run any model, including small local ones, and get real work out of them. Along the way it's produced code that got merged into projects like Remix, Caddy, Traefik, and Valkey (more on that, honestly, at the end). This post isn't a pitch for it. It's the set of principles I learned building it, written so you can apply them to your &lt;em&gt;own&lt;/em&gt; agent, whatever it's built on. Atlarix is just the reference implementation I'll point at to prove I actually did each thing rather than just theorizing.&lt;/p&gt;

&lt;p&gt;Here's the core claim, stated plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A weak model doesn't need a bigger prompt. It needs a system that catches its mistakes instead of trusting them.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Four mechanisms do most of the work. None of them require a better model.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Verify edits against reality, not against the model's word
&lt;/h2&gt;

&lt;p&gt;The single most common failure mode of a weak model in a coding loop is the &lt;strong&gt;confident false completion&lt;/strong&gt;: it reports success for work it didn't actually do. The edit didn't apply. The file didn't change. The function it described isn't on disk.&lt;/p&gt;

&lt;p&gt;A frontier model does this too — it just does it less often, so you get away with trusting it. With a weak model you cannot trust the report at all. So don't. The fix is a principle I'd now build into any agent from day one:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Never let the model be the judge of whether its own change succeeded. Check the world.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Concretely, after every edit, the harness re-reads the file from disk and confirms the change is actually present. If the edit didn't land, that fact goes &lt;strong&gt;back to the model as a tool result&lt;/strong&gt; — "the file is unchanged" — instead of forward to the user as "done." The model gets a chance to notice and retry, in the same turn, before anything reaches you.&lt;/p&gt;

&lt;p&gt;The same principle extends to commands. When the agent runs the project's own checks — &lt;code&gt;tsc&lt;/code&gt;, &lt;code&gt;eslint&lt;/code&gt;, &lt;code&gt;ruff&lt;/code&gt;, &lt;code&gt;mypy&lt;/code&gt;, &lt;code&gt;pytest&lt;/code&gt;, whatever the repo uses — a non-zero exit code is not something the agent gets to narrate its way past. In Atlarix the turn is &lt;em&gt;held open&lt;/em&gt;: a command that exits non-zero, or an edit that didn't verify on disk, is routed back to the agent rather than surfaced as a finished result. The agent literally cannot declare a task done while the project's own tests are failing.&lt;/p&gt;

&lt;p&gt;That one rule — &lt;strong&gt;the agent can't mark work complete while the repo's checks are red&lt;/strong&gt; — eliminates the most damaging class of weak-model errors, because the most damaging errors aren't wrong code. They're wrong code &lt;em&gt;reported as correct&lt;/em&gt;. Wrong-but-flagged is recoverable. Wrong-but-confident is what ships bugs.&lt;/p&gt;

&lt;p&gt;You can implement this in any loop. Run the checks. Read the exit code. If it's non-zero, feed the failure back as the next observation instead of returning to the caller. It's not clever. It's just refusing to take the model's word for anything you can verify yourself.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Enforce control in the harness, not in the prompt
&lt;/h2&gt;

&lt;p&gt;There's a strong temptation to solve agent misbehavior with more instructions. &lt;em&gt;"Always wait for the command to finish before continuing. Never claim completion prematurely. Always ask before running destructive commands."&lt;/em&gt; You write a longer and longer system prompt, and a frontier model mostly follows it, and a weak model mostly doesn't — because following a paragraph of procedural instructions &lt;em&gt;is itself a capability&lt;/em&gt; that weak models lack.&lt;/p&gt;

&lt;p&gt;So stop asking the model to behave and make the behavior structural.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If a rule matters, the harness should enforce it mechanically, so a model that "forgets" the rule physically can't break it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tool approvals, background command handling, wait-states, sub-agent orchestration — in a prompt-driven agent these are all things you &lt;em&gt;ask&lt;/em&gt; the model to handle correctly. In a harness-driven agent they're enforced by the system around the model. A weak model can't emit a premature "task complete" if completion is gated on verification it doesn't control (see #1). It can't skip an approval if the approval is a queue the execution path &lt;em&gt;must&lt;/em&gt; pass through, not a politeness the model chooses to observe.&lt;/p&gt;

&lt;p&gt;The practical test: for every "always/never" line in your system prompt, ask &lt;em&gt;"what happens if the model ignores this?"&lt;/em&gt; If the answer is "something bad happens," that rule doesn't belong in the prompt — it belongs in the harness, as a gate the model routes through whether it wants to or not. The prompt is for guidance. The harness is for guarantees. Weak models need guarantees.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Give it a sandbox, so a mistake is contained instead of catastrophic
&lt;/h2&gt;

&lt;p&gt;A weak model &lt;em&gt;will&lt;/em&gt; try to run something it shouldn't — a command scoped too broadly, a write outside the project, a destructive operation it didn't reason through. If your safety story is "the model is careful," you don't have a safety story.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Contain execution at the OS level, so the blast radius of a bad call is bounded by the system, not by the model's judgment.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Atlarix confines command execution per-OS — Landlock on Linux, AppContainer on Windows, Seatbelt on macOS — so the agent physically cannot write outside the project you opened. There's an approval queue with hunk-level accept/reject on every diff, and a danger gate on destructive operations. The point isn't the specific primitives; it's the principle: &lt;strong&gt;the model's mistakes are caught by the system, not trusted on faith.&lt;/strong&gt; When containment is structural, you can let a weaker, less-trustworthy model act, because the cost of it being wrong is bounded.&lt;/p&gt;

&lt;p&gt;This is also what makes local-first viable at all. If the model runs on your machine and the execution is sandboxed to your project, "the AI touched my system" stops being a leap of faith. The containment is what earns the autonomy.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Feed it structure on demand, not the whole repo
&lt;/h2&gt;

&lt;p&gt;Weak models have smaller effective context and degrade faster as you fill it. The instinct is to stuff the repo into the prompt so the model "has everything." This is exactly backwards — you drown the small model in tokens and it performs &lt;em&gt;worse&lt;/em&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Retrieve narrowly and on demand. Let the model pull what it needs when it needs it, instead of pre-loading everything it might need.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Atlarix searches the repo on demand — bundled &lt;code&gt;ripgrep&lt;/code&gt; for &lt;code&gt;grep&lt;/code&gt;/&lt;code&gt;glob&lt;/code&gt;, no index to build, no background watcher — and keeps durable notes on what it learns about the codebase so it doesn't re-derive the same structure every turn. The model asks for what it needs; it isn't forced to hold the whole tree in its head.&lt;/p&gt;

&lt;p&gt;The general principle for any agent: on-demand, tool-driven retrieval beats context-stuffing, and it beats it &lt;em&gt;more&lt;/em&gt; the weaker your model is. A frontier model can afford to waste context. A local model can't. Give it a way to look things up, and a way to remember what it found, and you've freed its limited context for actual reasoning.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this matters beyond saving on API bills
&lt;/h2&gt;

&lt;p&gt;The obvious payoff is cost — run a local model, pay nothing per token, keep your code on your own machine. That's real. But the deeper payoff is &lt;em&gt;trust&lt;/em&gt;. Every one of these mechanisms replaces a place where you were trusting the model with a place where the system checks. Verified edits replace trusting the completion report. Harness-enforced control replaces trusting the model to follow rules. The sandbox replaces trusting it not to do damage. On-demand retrieval replaces trusting it to juggle the whole repo.&lt;/p&gt;

&lt;p&gt;Stack them and something surprising happens: &lt;strong&gt;the model's reliability stops being the ceiling.&lt;/strong&gt; The system's guarantees become the floor, and a weaker model operating above a solid floor beats a stronger model operating on trust. That's the whole thesis. The harness, not the model.&lt;/p&gt;




&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;I said I'd be straight about the merges, so here it is.&lt;/p&gt;

&lt;p&gt;Atlarix, running on open-weight and local models, has produced code that maintainers merged into real projects — Remix (merged by its co-creator), Caddy, Traefik, Valkey, and others. Every one is a public, verifiable pull request; the links are on &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt; if you want to check them, and you should.&lt;/p&gt;

&lt;p&gt;But I want to be precise about what that does and doesn't mean, because a technical audience deserves it and will figure it out anyway. These were &lt;strong&gt;human-directed&lt;/strong&gt;. I drove the agent — chose the target, steered the work, reviewed the diff before it went out — and the agent appears as a &lt;em&gt;co-author&lt;/em&gt; on the commits, not the sole author. A maintainer of Remix or Caddy reviewed the change and merged it after dealing with me as the human behind it. What the harness did was get an open-weight model's output to the point where it could clear that bar: correct, verified against the project's own checks, and clean enough to survive a real maintainer's review.&lt;/p&gt;

&lt;p&gt;That's the claim, and it's a narrower and more honest one than "AI merged code into Remix." It's: a well-built harness plus a modest model, driven by a human who knows what they want, can produce work that passes the same gate a human contributor's work passes. That's not magic. It's the four mechanisms above, doing their job.&lt;/p&gt;

&lt;p&gt;If you're building your own agent, take the principles and leave the tool. If you want to see the reference implementation, it's &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;Atlarix&lt;/a&gt; — a private AI workstation that runs any model, local or hosted, with your code staying on your machine. Either way: stop blaming the model. Build the harness.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I'm Amariah Abishai, a self-taught engineer in Nairobi building AI developer tooling at NorahLabs. If you want the deeper version of the retrieval design, there's a &lt;a href="https://doi.org/10.5281/zenodo.20381860" rel="noopener noreferrer"&gt;published paper&lt;/a&gt; on an earlier structural-retrieval approach I tried, measured, and eventually moved away from — which is its own lesson for another post.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>localagents</category>
      <category>agents</category>
    </item>
    <item>
      <title>I shipped two PRs into Alibaba's qwen-code using only open-weight models. Here's the honest version.</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Fri, 03 Jul 2026 16:28:04 +0000</pubDate>
      <link>https://dev.to/abishaiama/i-shipped-two-prs-into-alibabas-qwen-code-using-only-open-weight-models-heres-the-honest-version-3clk</link>
      <guid>https://dev.to/abishaiama/i-shipped-two-prs-into-alibabas-qwen-code-using-only-open-weight-models-heres-the-honest-version-3clk</guid>
      <description>&lt;p&gt;I build &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;Atlarix&lt;/a&gt; — a desktop coding harness that runs whatever model you point it at: managed, your own API key, or fully local via Ollama or LM Studio. The whole thesis is that the gap between a weak open-weight model and a frontier one closes when the harness does the heavy lifting. This post is a test of exactly that.&lt;/p&gt;

&lt;p&gt;So I decided to test that thesis in the least forgiving place I could think of: contributing real code to a frontier lab's &lt;em&gt;own&lt;/em&gt; production repo, using nothing but open-weight models to write it.&lt;/p&gt;

&lt;p&gt;Two pull requests are now merged into &lt;a href="https://github.com/QwenLM/qwen-code" rel="noopener noreferrer"&gt;QwenLM/qwen-code&lt;/a&gt; — Alibaba's open-source coding agent. Here's the honest version, including the part that didn't go well and the constraint that shaped how I worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first attempt got closed — correctly
&lt;/h2&gt;

&lt;p&gt;My first swing was ambitious: a full always-on scheduled-task daemon. Tasks that fire on a cron schedule without an interactive session open, system service installation, webhook triggers, the works. Around 5,000 lines across four rollout phases.&lt;/p&gt;

&lt;p&gt;A maintainer closed it. And they were right to.&lt;/p&gt;

&lt;p&gt;The problem wasn't that the code didn't work — it was that I'd built a &lt;em&gt;second, parallel&lt;/em&gt; daemon with its own storage format and lifecycle, when the repo already had a long-running daemon (&lt;code&gt;qwen serve&lt;/code&gt;) and a durable scheduler I should have extended instead. On top of that, four phases in one PR is simply too much to review well.&lt;/p&gt;

&lt;p&gt;The maintainer's feedback was blunt and generous at the same time: reuse the existing infrastructure, make the change incremental, split it into reviewable pieces. What could have been ~300 lines of extension, I'd written as thousands of lines of duplication.&lt;/p&gt;

&lt;p&gt;That stung. But it was the most useful thing that happened in the whole process, because it taught me the lesson every open-source contributor eventually learns: &lt;strong&gt;read the existing architecture before you build, and keep your PRs small enough to review.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  So I did the right-sized thing instead
&lt;/h2&gt;

&lt;p&gt;Instead of fighting to rebuild the giant feature immediately, I went looking for a small, well-scoped problem — the kind of change that's easy to review, easy to revert, and unlikely to break anything.&lt;/p&gt;

&lt;p&gt;I found one in the web-shell's model picker.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/QwenLM/qwen-code/pull/6209" rel="noopener noreferrer"&gt;PR #6209&lt;/a&gt; — vision model support in the web-shell UI.&lt;/strong&gt; The CLI could select a vision model; the web-shell daemon UI couldn't. I added it, following the exact pattern the codebase already used for other model modes. Small, mechanical, convention-matching. It merged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/QwenLM/qwen-code/pull/6236" rel="noopener noreferrer"&gt;PR #6236&lt;/a&gt; — a real data-loss fix.&lt;/strong&gt; This one mattered more than its size suggests. When a user selected a vision model from the web-shell picker, they'd see a success toast — but their choice was silently discarded. The picker stored the model ID in one format (&lt;code&gt;modelId(authType)&lt;/code&gt;, ACP-style) while the core resolver expected another (&lt;code&gt;authType:modelId&lt;/code&gt;). The mismatch meant the stored value never resolved, and the system quietly fell back to auto-select. The settings page still &lt;em&gt;showed&lt;/em&gt; the value, which masked the failure completely. The user's explicit choice had no effect, and nothing told them.&lt;/p&gt;

&lt;p&gt;The fix re-encodes the format before persisting, plus type-safe dispatch to replace some fragile ternary chains, plus the missing English and Chinese i18n keys. It merged too.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part: the models didn't write perfect code
&lt;/h2&gt;

&lt;p&gt;This is the bit I actually care about, because it's where the hype usually lies.&lt;/p&gt;

&lt;p&gt;Both PRs went through several rounds of genuine maintainer review. And the maintainers — Alibaba collaborators, plus an automated reviewer running on Qwen's own models — found &lt;em&gt;real&lt;/em&gt; bugs. Not style nitpicks. Actual security and lifecycle issues: an HMAC check computed over the wrong input, a timing-attack-vulnerable token comparison, a timer that could fire on a stopped process, an auth config silently stripped from a webhook path.&lt;/p&gt;

&lt;p&gt;I fixed and re-verified each one before merge. On #6236, a maintainer even built the PR locally with real browser tests and screenshots to confirm the fix worked end-to-end before approving.&lt;/p&gt;

&lt;p&gt;That review loop is the entire point. The claim isn't "open-weight models wrote flawless code." They didn't. The claim is that &lt;strong&gt;open-weight models, driven by a good harness, could take architectural feedback and iterate to something a senior maintainer at a frontier lab was willing to merge.&lt;/strong&gt; That's a much more interesting and much more honest result than "the AI one-shotted it."&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint nobody tells you about: cost
&lt;/h2&gt;

&lt;p&gt;Here's a detail I think is worth being transparent about, because it's the reality of building solo.&lt;/p&gt;

&lt;p&gt;PR #6209 was built almost entirely on Qwen (3.6 Plus, 3.7 Plus/Max) via OpenRouter. But partway through PR #6236, I started running low on OpenRouter credits. As a solo founder trying to conserve runway for actual users, I switched to using DeepSeek API credits instead. So #6236 ended up roughly a 50/50 mix of Qwen and DeepSeek.&lt;/p&gt;

&lt;p&gt;I could have hidden that and claimed "100% Qwen" across the board. But two things: first, it wouldn't be true, and the whole value of a post like this is that it's honest. Second — it actually makes the result &lt;em&gt;broader&lt;/em&gt;, not weaker. The thesis was never "Qwen specifically." It was "open-weight models, in a harness built for them, can do real work." Making that work across &lt;em&gt;two&lt;/em&gt; different open-weight labs is stronger evidence than making it work with one.&lt;/p&gt;

&lt;p&gt;No frontier model wrote either PR. Just open-weight models — Qwen and DeepSeek — running in Atlarix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters (to me, at least)
&lt;/h2&gt;

&lt;p&gt;I'm a self-taught developer building in Nairobi. The models I can afford to run at scale are open-weight ones. Atlarix exists because I needed a way to make those models genuinely productive — not "good enough for a demo," but good enough to ship code into a repo maintained by the people who &lt;em&gt;train&lt;/em&gt; the models.&lt;/p&gt;

&lt;p&gt;Two merged PRs in a Tier-1 lab's production repo, written by open-weight models in the harness, reviewed and merged by the lab's own maintainers, is the clearest proof of that thesis I've been able to produce.&lt;/p&gt;

&lt;p&gt;The gap between open-weight and frontier is real. But a lot of it lives in the harness, not the weights. Close that gap, and a model you can run yourself can punch well above its weight.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Atlarix is at &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;. If you're building with open-weight models or thinking about model-agnostic tooling, I'd genuinely like to compare notes.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>openweight</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Atlarix vs opencode on Terminal-Bench 2.0 — same model, only the harness changes (k=1, receipts included)</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Mon, 29 Jun 2026 19:22:30 +0000</pubDate>
      <link>https://dev.to/abishaiama/atlarix-vs-opencode-on-terminal-bench-20-same-model-only-the-harness-changes-k1-receipts-54nk</link>
      <guid>https://dev.to/abishaiama/atlarix-vs-opencode-on-terminal-bench-20-same-model-only-the-harness-changes-k1-receipts-54nk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Update — treat this as a historical data point, not a current claim.&lt;/strong&gt; Two things have superseded it. Terminal-Bench has moved to &lt;strong&gt;4.0&lt;/strong&gt;, which recalibrated task resources, dropped saturated tasks and standardised an 8-hour agent timeout — 2.x scores are not comparable to it. And the headless runner that produced these numbers dispatched reasoning &lt;strong&gt;one notch below&lt;/strong&gt; each model's declared ceiling, while published leaderboard entries run the ceiling. Both are fixed; the run has not been repeated.&lt;/p&gt;

&lt;p&gt;The caveat in the post below still stands and matters more than the result: a 3-task gap at k=1 is within noise, and this &lt;strong&gt;does not&lt;/strong&gt; show Atlarix is ahead of opencode. &lt;strong&gt;There is also no Atlarix SWE-bench number&lt;/strong&gt; — if you have seen one attributed to Atlarix, it is invented. Current results and methodology: &lt;strong&gt;&lt;a href="https://www.atlarix.dev/benchmark" rel="noopener noreferrer"&gt;atlarix.dev/benchmark&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;a href="https://atlarix.dev/benchmark" rel="noopener noreferrer"&gt;Benchmarks&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;I build &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;Atlarix&lt;/a&gt;, an agent workstation for open-weight models. The core claim behind it is that the harness — retrieval, tool surface, control loop — is what lets an open-weight model perform, not just the model's raw weights. This post is me trying to falsify that claim with a controlled run, and publishing every output file so you can check it.&lt;/p&gt;

&lt;p&gt;Short version: on Terminal-Bench 2.0, single attempt, &lt;strong&gt;Atlarix resolved 42/89 and opencode resolved 39/89&lt;/strong&gt; on the same model. That 3-task gap is &lt;strong&gt;within k=1 noise&lt;/strong&gt; — I'm not claiming a win. What it shows is that the harness isn't bottlenecking the model. Details and caveats below; raw files at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment
&lt;/h2&gt;

&lt;p&gt;The only variable is the harness. Everything else is pinned identical across both agents.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Benchmark:&lt;/strong&gt; &lt;code&gt;terminal-bench/terminal-bench-2&lt;/code&gt; — all 89 tasks, one isolated container each, automated verifiers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; &lt;code&gt;minimax/minimax-m3&lt;/code&gt;, routed through OpenRouter, pinned to a single provider at &lt;strong&gt;fp8&lt;/strong&gt; — identical for both harnesses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infrastructure:&lt;/strong&gt; &lt;a href="https://github.com/harbor-framework/terminal-bench" rel="noopener noreferrer"&gt;Harbor&lt;/a&gt; on Modal (&lt;code&gt;-e modal&lt;/code&gt;), one container per task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Attempts:&lt;/strong&gt; single attempt, &lt;code&gt;-k 1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Timeout:&lt;/strong&gt; native, &lt;code&gt;--timeout-multiplier 1&lt;/code&gt; (same for both).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retries:&lt;/strong&gt; &lt;code&gt;--max-retries 3&lt;/code&gt; (same for both).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool calling:&lt;/strong&gt; native function-calling forced, no text-tool shim.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Commands
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Atlarix harness&lt;/span&gt;
harbor run &lt;span class="nt"&gt;-d&lt;/span&gt; terminal-bench/terminal-bench-2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-m&lt;/span&gt; openai/minimax/minimax-m3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-n&lt;/span&gt; 24 &lt;span class="nt"&gt;-k&lt;/span&gt; 1 &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="nt"&gt;--timeout-multiplier&lt;/span&gt; 1 &lt;span class="nt"&gt;--max-retries&lt;/span&gt; 3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-e&lt;/span&gt; modal &lt;span class="nt"&gt;--agent-import-path&lt;/span&gt; atlarix_tb:AtlarixAgent

&lt;span class="c"&gt;# opencode harness (same model + provider + infra)&lt;/span&gt;
harbor run &lt;span class="nt"&gt;-d&lt;/span&gt; terminal-bench/terminal-bench-2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-m&lt;/span&gt; bench/minimax/minimax-m3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-n&lt;/span&gt; 24 &lt;span class="nt"&gt;-k&lt;/span&gt; 1 &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="nt"&gt;--timeout-multiplier&lt;/span&gt; 1 &lt;span class="nt"&gt;--max-retries&lt;/span&gt; 3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-e&lt;/span&gt; modal &lt;span class="nt"&gt;--agent-import-path&lt;/span&gt; atlarix_tb.opencode_proxy:BenchOpenCodeAgent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(&lt;code&gt;-n 24&lt;/code&gt; is concurrency — how many containers run in parallel — not a task count. All 89 tasks run.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Harness&lt;/th&gt;
&lt;th&gt;Resolved&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Atlarix&lt;/td&gt;
&lt;td&gt;42 / 89&lt;/td&gt;
&lt;td&gt;47%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;opencode&lt;/td&gt;
&lt;td&gt;39 / 89&lt;/td&gt;
&lt;td&gt;44%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Read this before you read the table
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;k=1 means one sample per task.&lt;/strong&gt; The official Terminal-Bench leaderboard requires &lt;strong&gt;k=5&lt;/strong&gt; specifically to measure run-to-run variance. A 3-task difference at k=1 is inside that noise band. So this is &lt;strong&gt;not&lt;/strong&gt; a leaderboard result and not a claim that Atlarix beats opencode. The honest takeaway: an open-weight model performs about as well under Atlarix as under a strong existing harness — the harness isn't holding it back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;~25% of tasks timed out — for both harnesses.&lt;/strong&gt; At native timeout (×1), roughly a quarter of tasks hit &lt;code&gt;AgentTimeoutError&lt;/code&gt; on each side and count as unresolved. So the sub-50% absolute scores aren't all capability failures; a meaningful share are wall-clock on heavy tasks. The timeout ceiling is identical for both agents, so the comparison stays fair — but that's why neither number is higher.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one config change (full disclosure)
&lt;/h2&gt;

&lt;p&gt;Atlarix's desktop app asks for human approval before every file write and command — a core safety feature. Benchmarks run unattended, so I grant that approval once via an explicit operator flag (&lt;code&gt;ATLARIX_AUTONOMOUS_DANGER=1&lt;/code&gt;). Without it, any task needing an install or privileged command is blocked and fails.&lt;/p&gt;

&lt;p&gt;This is &lt;strong&gt;not&lt;/strong&gt; an advantage over opencode — every agent auto-approves to run an automated benchmark; it's inherent to running unattended. Stating it for full transparency. The flag is off by default; the interactive app always asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;

&lt;p&gt;The exact Atlarix bundle I ran is a public, Electron-free headless build: &lt;code&gt;atlarix-headless-linux-amd64.tar.gz&lt;/code&gt;. The benchmark is the open-source Harbor framework. The raw Harbor result files — per-task pass/fail for both harnesses — are published unedited. Nothing is hand-typed.&lt;/p&gt;

&lt;p&gt;Everything (raw &lt;code&gt;result.json&lt;/code&gt; for both sides, &lt;code&gt;summary.csv&lt;/code&gt;, exact bundle, full setup): &lt;strong&gt;&lt;a href="https://atlarix.dev/benchmark" rel="noopener noreferrer"&gt;atlarix.dev/benchmark&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;More open-weight models, so no claim rests on one.&lt;/li&gt;
&lt;li&gt;The official Terminal-Bench (k=5) submission — on the roadmap.&lt;/li&gt;
&lt;li&gt;More benchmarks beyond terminal tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you spot something wrong in the result files, that's the point — tell me.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Built in Nairobi.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>productivity</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Show dev: I built an AI agent workstation in Nairobi for open-weight and local models</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Fri, 19 Jun 2026 09:10:36 +0000</pubDate>
      <link>https://dev.to/abishaiama/show-dev-i-built-an-ai-agent-workstation-in-nairobi-for-deepseek-qwen-kimi-minimax-lnb</link>
      <guid>https://dev.to/abishaiama/show-dev-i-built-an-ai-agent-workstation-in-nairobi-for-deepseek-qwen-kimi-minimax-lnb</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Historical.&lt;/strong&gt; Blueprint, ctags, ast-grep and SQLite FTS5 were all removed from Atlarix in &lt;strong&gt;v14.9.0 (10 July 2026)&lt;/strong&gt; — the index leaked file descriptors and did not beat plain search. &lt;strong&gt;Atlarix has no index of any kind today&lt;/strong&gt;; retrieval is bundled ripgrep, no embeddings, no graph. Current docs: &lt;strong&gt;&lt;a href="https://www.atlarix.dev/docs" rel="noopener noreferrer"&gt;atlarix.dev/docs&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;Atlarix — a ~400MB AI agent workstation that sits &lt;em&gt;beside&lt;/em&gt; your IDE (VS Code, IntelliJ, Vim) instead of replacing it. A native harness that runs whatever model you point it at — managed, your own API key across the live models.dev catalogue, or fully local via Ollama and LM Studio. &lt;em&gt;(This post originally named four open-weight labs as "the point"; that framing no longer holds — the managed lineup is configured remotely and changes, and the harness is deliberately model-agnostic.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Built solo under NorahLabs, in Nairobi, Kenya.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;I was running open-weight models for actual agentic work — multi-file edits, terminal commands, codebase exploration. Every tool I tried was built around Claude or GPT, with my models bolted on as a BYOK option. Context-window assumptions tuned for a different model. System prompts and tool-calling shaped around closed-model behavior. Retrieval that either dumps the whole repo or leans on a cloud vector DB that doesn't even exist on an offline machine.&lt;/p&gt;

&lt;p&gt;The models are frontier-class now. The tooling around them isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The approach
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Blueprint — structural retrieval, no embeddings &lt;em&gt;(removed in v14.9.0)&lt;/em&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Universal Ctags symbols + ast-grep edges, backed by SQLite FTS5 *(all removed in v14.9.0 — see the note at the top)*
- grep results reranked by structural relevance
- each result annotated with its enclosing function/class
- no vector DB, constant memory at any repo size
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The thesis: for &lt;em&gt;code&lt;/em&gt;, lexical + structural retrieval plus the model's own reasoning beats a vector index — which is also why Claude Code and opencode carry no embedding index. On my own large multi-repo workspace, a "find the signup code" query dropped from ~63K to ~26K turn tokens with exact &lt;code&gt;file:line&lt;/code&gt; citations. (That's my own workspace, not a published benchmark — a reproducible eval is what I'm building next.)&lt;/p&gt;

&lt;h3&gt;
  
  
  Verified edit loop
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- write → re-read from disk → compare to intended
- zero tokens on the happy path
- blocks "task complete" if an edit didn't actually land
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Live model catalog
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- managed model IDs fetched from a hosted config at startup
- new drops from the four labs appear automatically
- swapping a model is a config change, not an app rebuild
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Per-OS sandboxing
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- macOS: Seatbelt
- Linux: bubblewrap
- Windows: AppContainer
- every file write + command approval-gated, per-hunk diff review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Tech stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Electron + React + TypeScript&lt;/li&gt;
&lt;li&gt;SQLite FTS5 for retrieval&lt;/li&gt;
&lt;li&gt;Universal Ctags + ast-grep for structural indexing&lt;/li&gt;
&lt;li&gt;Rust helper for the Windows sandbox&lt;/li&gt;
&lt;li&gt;Node.js 24+&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where it's at, honestly
&lt;/h2&gt;

&lt;p&gt;v13.9.0, shipped and working. I'm one developer, so I'll be straight about the early edges: no published head-to-head benchmark yet, macOS is notarized but Windows builds are currently unsigned (signing's coming), and the retrieval numbers above are from my own use, not a controlled eval. The honest pitch is "the first workstation built &lt;em&gt;around&lt;/em&gt; open-weight models instead of just accepting them — here's exactly how," not "this beats everything."&lt;/p&gt;

&lt;h2&gt;
  
  
  Feedback wanted
&lt;/h2&gt;

&lt;p&gt;If you're running open-weight models for agentic work — what's your current setup, and where does it break? That's the feedback that actually shapes this.&lt;/p&gt;

&lt;p&gt;🌐 atlarix.dev&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>opensource</category>
      <category>showdev</category>
    </item>
    <item>
      <title>We Tried to Reduce LLM Context Usage in a Multi-Repo Codebase. The AI Used More Tokens, Not Less. Here's Why That's Correct.</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Mon, 25 May 2026 17:46:41 +0000</pubDate>
      <link>https://dev.to/abishaiama/we-tried-to-reduce-llm-context-usage-in-a-multi-repo-codebase-the-ai-used-more-tokens-not-less-3gje</link>
      <guid>https://dev.to/abishaiama/we-tried-to-reduce-llm-context-usage-in-a-multi-repo-codebase-the-ai-used-more-tokens-not-less-3gje</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Update — this experiment ended the way the numbers suggested it would.&lt;/strong&gt; Blueprint was removed from Atlarix in &lt;strong&gt;v14.9.0 (10 July 2026)&lt;/strong&gt;, along with the whole index: FTS5, BM25, ctags, the reranker and the file watcher, about 5,246 lines. The 54% result below was part of why. The watcher also turned out to hold one file descriptor per watched file, reaching roughly 12,365 on a large repository.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Retrieval is now purely lexical&lt;/strong&gt; — bundled ripgrep, no index, no embeddings, nothing in the background. This post is the evidence for that decision, not a description of a shipping feature.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;When we set out to build Blueprint — Atlarix's structural codebase retrieval system — the hypothesis was simple: give the AI a map of the codebase upfront, and it will need to read fewer files. Fewer files means less context. Less context means lower cost and faster responses.&lt;/p&gt;

&lt;p&gt;We ran a controlled benchmark. The AI with Blueprint used &lt;strong&gt;54% more context&lt;/strong&gt; than the AI without it.&lt;/p&gt;

&lt;p&gt;Here's why that's not a failure.&lt;/p&gt;




&lt;h3&gt;
  
  
  The problem with how most AI coding tools handle context
&lt;/h3&gt;

&lt;p&gt;If you've used Cursor, Claude Code, or GitHub Copilot on a large codebase, you've hit this wall: the AI either reads too much (dumping raw files into context until you hit the limit) or reads too little (making confident wrong assumptions about files it hasn't seen).&lt;/p&gt;

&lt;p&gt;The root cause is navigation. Without a structural map of the codebase, the AI is exploring blind — making guesses about which files matter, following import chains manually, or relying on whatever files happen to be open. In a multi-repository workspace, this gets worse fast. You might have 25 separate projects with thousands of files. The AI has no idea where it is.&lt;/p&gt;

&lt;p&gt;The standard solutions are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Raw file injection&lt;/strong&gt; — dump everything relevant into context upfront. Expensive and doesn't scale past a few hundred files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dense vector search&lt;/strong&gt; — embed the codebase and retrieve by semantic similarity. Loses structural relationships (call chains, import graphs, HTTP routes).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic search&lt;/strong&gt; — let the model call search/read tools and figure it out. Works, but slow and token-hungry as the model searches blindly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We built Blueprint to try a fourth approach: give the model a symbolic structural graph before it starts exploring.&lt;/p&gt;




&lt;h3&gt;
  
  
  What Blueprint actually is
&lt;/h3&gt;

&lt;p&gt;Blueprint is a four-layer index:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1 — Universal Ctags (symbol index)&lt;/strong&gt;&lt;br&gt;
Extracts every function, class, type, and method across 18 languages. Line-accurate positions. Cached to &lt;code&gt;.atlarix/symbols.json&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 2 — ast-grep (structural edges)&lt;/strong&gt;&lt;br&gt;
AST-level pattern matching for import edges, call edges, and HTTP route edges. Express &lt;code&gt;app.get&lt;/code&gt;, Fastify &lt;code&gt;fastify.post&lt;/code&gt;, Next.js &lt;code&gt;export async function GET&lt;/code&gt; — all become first-class nodes in the graph.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 3 — BM25 (semantic symbol ranking)&lt;/strong&gt;&lt;br&gt;
Ranks ctags symbols by concept query. "Authentication middleware" finds the right functions without requiring an exact name match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 4 — ripgrep (text fallback)&lt;/strong&gt;&lt;br&gt;
Exact string search for when you know precisely what you're looking for.&lt;/p&gt;

&lt;p&gt;The output is a compact Markdown slice — rooms (directory-scoped clusters), beacons (individual symbols), and edges (structural relationships). Section-scoped: the agent requests one folder at a time, not the whole workspace.&lt;/p&gt;




&lt;h3&gt;
  
  
  The benchmark
&lt;/h3&gt;

&lt;p&gt;We ran two arms of the same task on a production multi-repository workspace:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;25 sections&lt;/strong&gt; across the workspace root&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3,250 tracked files&lt;/strong&gt; total&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Target section:&lt;/strong&gt; a TypeScript CLI package, 99 files, ~9,500 lines&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Task:&lt;/strong&gt; Trace an event-driven HTTP-ingress-to-webhook-reply pipeline. Both arms had identical deliverables — narrative of the flow, key file paths, Mermaid sequence diagram.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arm A (with Blueprint):&lt;/strong&gt; Prescribed tool order — explore folder → &lt;code&gt;get_blueprint&lt;/code&gt; → text search → &lt;code&gt;read_file&lt;/code&gt; on 2-3 central files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arm B (without Blueprint):&lt;/strong&gt; Same task, no &lt;code&gt;get_blueprint&lt;/code&gt; — only explore folder → text search → &lt;code&gt;read_file&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model:&lt;/strong&gt; Kimi K2.6 (268K context window) via OpenRouter. Same model, same provider, both arms.&lt;/p&gt;




&lt;h3&gt;
  
  
  The results
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;With Blueprint&lt;/th&gt;
&lt;th&gt;Without Blueprint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Blueprint slice&lt;/td&gt;
&lt;td&gt;~6,500 tokens&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Final billed input&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;63,541 tokens&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;41,327 tokens&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens&lt;/td&gt;
&lt;td&gt;2,671&lt;/td&gt;
&lt;td&gt;2,534&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task completion&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Blueprint arm used &lt;strong&gt;54% more total context.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Context growth per turn:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;With Blueprint:&lt;/strong&gt; 8,661 → 13,966 → 24,771 → 25,012 → 31,717 → 54,188 → 63,541&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Without Blueprint:&lt;/strong&gt; 2,253 → 3,567 → 8,629 → 13,934 → 14,175 → 37,876 → 41,327&lt;/p&gt;

&lt;p&gt;The Blueprint arm took six tool-call turns. The no-Blueprint arm took five.&lt;/p&gt;




&lt;h3&gt;
  
  
  Why this is the correct result
&lt;/h3&gt;

&lt;p&gt;Here's what we found in the qualitative output comparison:&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Blueprint arm&lt;/strong&gt; named 7 specific internal functions by exact identifier — the auth validator, mention detector, memory clamp, post-processor, card builder, and two others. It surfaced a section-specific post-processor module not explicitly requested.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;no-Blueprint arm&lt;/strong&gt; found a client module in an &lt;code&gt;eval/&lt;/code&gt; subdirectory that Blueprint's section scope hadn't included. It named specific environment variables and API constants the text search found directly.&lt;/p&gt;

&lt;p&gt;Both arms completed the task correctly. But the &lt;em&gt;type&lt;/em&gt; of knowledge was different.&lt;/p&gt;

&lt;p&gt;Blueprint gave the model a symbol-level map before any file was read. With that map, the model knew which files were worth reading and went deeper — more function names, more architectural detail, more thorough coverage. Without the map, the model explored more conservatively: followed fewer paths, read fewer files, stopped sooner.&lt;/p&gt;

&lt;p&gt;The no-Blueprint arm used fewer tokens partly because it was less certain about what to look for next.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For a read-only exploration task, "explored less" isn't obviously worse.&lt;/strong&gt; Both arms got the answer. But for write tasks — bug fixes, refactors, feature implementation — a model that stops exploring because it's navigationally lost is not saving tokens. It's missing dependencies, and those missing dependencies become production bugs.&lt;/p&gt;




&lt;h3&gt;
  
  
  The real finding: structural understanding and execution context are separable problems
&lt;/h3&gt;

&lt;p&gt;The honest framing isn't "Blueprint reduces total context." It's that these are two different problems:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structural understanding cost&lt;/strong&gt; — how many tokens does it take to know where you are in the codebase?&lt;/p&gt;

&lt;p&gt;With Blueprint: &lt;strong&gt;~6,500 tokens&lt;/strong&gt;, regardless of section complexity, in ~3 seconds.&lt;br&gt;
Without Blueprint: amortised across many search/read tool calls over multiple turns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Execution context&lt;/strong&gt; — how many tokens accumulate as the model actually does the work?&lt;/p&gt;

&lt;p&gt;This is determined by exploration depth — how many files the model reads, how many tool calls it makes. Blueprint increases this by making the model more confident. But it's bounded and manageable.&lt;/p&gt;

&lt;p&gt;We address the execution context problem with a separate mechanism: post-turn tool-result summarisation. After each turn, large tool outputs in the persisted transcript are rewritten by a fast compaction model — keeping paths, symbol names, and key values, dropping JSON noise and repetition. In the benchmark runs, individual &lt;code&gt;read_file&lt;/code&gt; results compressed from 2,500–3,500 tokens to 60–110 tokens. ~95–98% reduction per qualifying block.&lt;/p&gt;

&lt;p&gt;Two mechanisms, two layers, two different problems.&lt;/p&gt;




&lt;h3&gt;
  
  
  What this means if you're building on top of LLMs
&lt;/h3&gt;

&lt;p&gt;If you're building an AI coding tool, an agentic system, or anything that needs to navigate a large codebase:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't chase "total context reduction" as a single metric.&lt;/strong&gt; It conflates structural overhead (knowable upfront, bounded by your retrieval design) with execution noise (determined by task complexity and model confidence).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Give the model a map before it explores.&lt;/strong&gt; Not raw files — a structural graph. The model will use more total context because it will explore more thoroughly. That's the right trade for write tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compress history, not retrieval.&lt;/strong&gt; Post-turn summarisation on tool outputs is more effective than trying to cram less information into the initial retrieval. The model needs the full file during the turn. Future turns don't.&lt;/p&gt;




&lt;h3&gt;
  
  
  The full paper
&lt;/h3&gt;

&lt;p&gt;This benchmark is documented in a technical paper published on Zenodo with full methodology, exact prompts, provider-billed token counts, and an honest discussion of limitations:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Blueprint: Section-Scoped Structural Graph Retrieval and Post-Turn Compression for Agentic LLM Coding in Multi-Repository Workspaces&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://zenodo.org/records/20381860" rel="noopener noreferrer"&gt;zenodo.org/records/20381860&lt;/a&gt; · DOI: 10.5281/zenodo.20381860&lt;/p&gt;

&lt;p&gt;Atlarix is available at &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;. The MCP server registry is open-source at &lt;a href="https://github.com/AmariahAK/atlarix-mcps" rel="noopener noreferrer"&gt;github.com/AmariahAK/atlarix-mcps&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built in Nairobi. Questions or thoughts? Drop them in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Build an AI Agent That Actually Understands Your Codebase (Without Switching Editors)</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Mon, 18 May 2026 13:04:51 +0000</pubDate>
      <link>https://dev.to/abishaiama/build-an-ai-agent-that-actually-understands-your-codebase-without-switching-editors-6d9</link>
      <guid>https://dev.to/abishaiama/build-an-ai-agent-that-actually-understands-your-codebase-without-switching-editors-6d9</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Historical.&lt;/strong&gt; Blueprint, ctags, ast-grep and SQLite FTS5 were all removed from Atlarix in &lt;strong&gt;v14.9.0 (10 July 2026)&lt;/strong&gt; — the index leaked file descriptors and did not beat plain search. &lt;strong&gt;Atlarix has no index of any kind today&lt;/strong&gt;; retrieval is bundled ripgrep, no embeddings, no graph. Current docs: &lt;strong&gt;&lt;a href="https://www.atlarix.dev/docs" rel="noopener noreferrer"&gt;atlarix.dev/docs&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;I use Neovim. I'm not switching.&lt;/p&gt;

&lt;p&gt;But I also want an agent that can refactor across 50 files, run tests, debug failures, and come back with a working PR — not just suggest the next line.&lt;/p&gt;

&lt;p&gt;The problem: every AI coding tool that does serious agentic work wants to be your editor. Cursor, Windsurf, GitHub Copilot Workspace — all VS Code. If you use anything else, you're a second-class citizen.&lt;/p&gt;

&lt;p&gt;So I built something different. An agent that has its own workspace &lt;em&gt;beside&lt;/em&gt; my editor, not inside it. I stay in Neovim. The agent gets a live map of the codebase, a terminal, file tools, and an approval queue.&lt;/p&gt;

&lt;p&gt;This post walks through how it works, how to set it up, and what it actually looks like in practice.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Core Problem with Raw Code Injection
&lt;/h2&gt;

&lt;p&gt;Before getting into the setup, it's worth understanding why most agents struggle with large codebases.&lt;/p&gt;

&lt;p&gt;The standard approach is to dump files into the context window:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here is your codebase:
[file 1 - 500 lines]
[file 2 - 800 lines]
[file 3 - 1200 lines]
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This works at small scale. At 20+ files it starts to break down. The model spends most of its reasoning budget reconstructing architecture from raw text instead of actually solving the problem.&lt;/p&gt;

&lt;p&gt;The better approach: give the agent a &lt;em&gt;structured map&lt;/em&gt; of the codebase and let it query specific parts on demand. Think of it like the difference between handing someone a stack of printed pages vs giving them a searchable database with a good schema.&lt;/p&gt;

&lt;p&gt;That's the core idea behind Atlarix's Live Code Map — your repo gets parsed into a node/edge graph that the agent navigates instead of reading raw files linearly.&lt;/p&gt;




&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Download and Install
&lt;/h3&gt;

&lt;p&gt;Grab the installer from &lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;macOS&lt;/strong&gt;: &lt;code&gt;.dmg&lt;/code&gt;, notarized, installs to Applications&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Linux&lt;/strong&gt;: &lt;code&gt;.AppImage&lt;/code&gt; (auto-updates), &lt;code&gt;.deb&lt;/code&gt;, or &lt;code&gt;.rpm&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Windows&lt;/strong&gt;: unsigned &lt;code&gt;.exe&lt;/code&gt; (code signing coming)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Install the CLI
&lt;/h3&gt;

&lt;p&gt;Open Atlarix → Settings → General → Install CLI.&lt;/p&gt;

&lt;p&gt;This drops an &lt;code&gt;atlarix&lt;/code&gt; binary into your PATH. After that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Open current directory as a workspace&lt;/span&gt;
atlarix &lt;span class="nb"&gt;.&lt;/span&gt;

&lt;span class="c"&gt;# Open a specific path&lt;/span&gt;
atlarix ~/projects/my-api
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same muscle memory as &lt;code&gt;code .&lt;/code&gt;. Atlarix opens in the background, your terminal returns immediately.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Connect a Model
&lt;/h3&gt;

&lt;p&gt;Go to Settings → AI. You have two options:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;BYOK (Bring Your Own Key)&lt;/strong&gt; — paste in an API key for any of the supported providers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAI, Anthropic, Google Gemini, Groq, Together AI,
Mistral, xAI, OpenRouter, AWS Bedrock
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Local models via Ollama or LM Studio&lt;/strong&gt; — set the base URL and pick your model. No API key needed. Works on the free Solo tier.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ollama base URL: http://localhost:11434
Model: qwen2.5-coder:7b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For local models I've had good results with &lt;code&gt;qwen2.5-coder:7b&lt;/code&gt; and &lt;code&gt;deepseek-coder-v2:16b&lt;/code&gt; on an M2 MacBook. The structured context from the code map means you don't need a massive model to get useful results.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model Bridge (managed)&lt;/strong&gt; — Atlarix's own managed tier (Speed/Standard/Deep) if you don't want to manage keys. Speed tier is included in the free Solo plan.&lt;/p&gt;




&lt;h2&gt;
  
  
  Opening a Workspace
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ~/projects/my-saas-api
atlarix &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Atlarix opens and loads your workspace. The first thing it does is a lightweight repo scan via &lt;code&gt;git ls-files&lt;/code&gt; — fast, respects your &lt;code&gt;.gitignore&lt;/code&gt;, and doesn't run any heavy parsing upfront.&lt;/p&gt;

&lt;p&gt;The Live Code Map was built on-demand when the agent first called &lt;code&gt;get_blueprint&lt;/code&gt;. &lt;em&gt;(The &lt;code&gt;get_blueprint&lt;/code&gt; tool, the Blueprint tab and the map itself were all removed in v14.9.0 — see the note at the top.)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For a 500-file TypeScript repo, the initial parse takes 3–5 seconds. After that it's cached and incremental.&lt;/p&gt;




&lt;h2&gt;
  
  
  Your First Agent Task
&lt;/h2&gt;

&lt;p&gt;Let's say I want to add rate limiting to my auth routes. Here's what that actually looks like.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Start in Explore Mode
&lt;/h3&gt;

&lt;p&gt;Pick &lt;strong&gt;Explore&lt;/strong&gt; from the mode selector. This is read-only — the agent can query the code map, read files, and search, but can't write anything. Good for orientation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Me: Map the auth flow. Where does a login request go from the 
    route handler to the database?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent called &lt;code&gt;get_blueprint&lt;/code&gt; to pull the architectural graph, then navigated from the auth route to the middleware stack to the database layer. &lt;em&gt;(Today it does this with ripgrep and reasoning over ranked results — no graph involved.)&lt;/em&gt; It comes back with a precise answer and which files are involved — without reading every file in the codebase.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Switch to Plan Mode
&lt;/h3&gt;

&lt;p&gt;Now switch to &lt;strong&gt;Plan&lt;/strong&gt; mode. The agent can now draft a plan but still can't execute writes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Me: I want to add rate limiting to the login and registration 
    routes. Use express-rate-limit. Failed attempts should be 
    tracked by IP. After 5 failures in 10 minutes, lock out 
    for 30 minutes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent knows the architecture from the previous explore session. It drafts a step-by-step plan:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;Plan: Add rate limiting to auth routes

[ ] Install express-rate-limit
[ ] Create rate limiter config in src/middleware/rateLimiter.ts
[ ] Apply limiter to POST /auth/login in src/routes/auth.ts
[ ] Apply limiter to POST /auth/register in src/routes/auth.ts
[ ] Add error response handler for 429 status
[ ] Update integration tests in tests/auth.test.ts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Review the plan. If it looks right, switch to &lt;strong&gt;Build&lt;/strong&gt; mode.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Build — Approval Queue in Action
&lt;/h3&gt;

&lt;p&gt;In Build mode, the agent starts executing. For every file write, you get an approval prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Write&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;middleware&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="nx"&gt;rateLimiter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;

&lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;rateLimit&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;express-rate-limit&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;loginRateLimiter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;rateLimit&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;windowMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// 10 minutes&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;max&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Too many login attempts. Try again later.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;standardHeaders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;legacyHeaders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;registerRateLimiter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;rateLimit&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;windowMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// 1 hour&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;max&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Too many registration attempts.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;standardHeaders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt;   &lt;span class="na"&gt;legacyHeaders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;Approve&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;Reject&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You approve. The file gets written. The agent continues through the plan.&lt;/p&gt;

&lt;p&gt;For terminal commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Terminal: npm &lt;span class="nb"&gt;install &lt;/span&gt;express-rate-limit

&lt;span class="o"&gt;[&lt;/span&gt;Approve] &lt;span class="o"&gt;[&lt;/span&gt;Reject]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You approve. It runs, shows you the output, continues.&lt;/p&gt;

&lt;p&gt;This isn't friction — it's the same review cycle as a PR, but live. You're watching the work happen and approving as it goes rather than reviewing after the fact.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Review in Your Editor
&lt;/h3&gt;

&lt;p&gt;The agent writes the files. You review them in Neovim, VS Code, IntelliJ, or wherever. The agent doesn't care what you're using to review — it just cares about the approval queue.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Meanwhile in my terminal&lt;/span&gt;
nvim src/middleware/rateLimiter.ts
&lt;span class="c"&gt;# Looks good, approve the rest in Atlarix&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 5: Tests
&lt;/h3&gt;

&lt;p&gt;If the agent runs tests as part of the plan and they fail, it reads the failure output and iterates. Self-correction is built in — up to a configurable number of retry attempts before it stops and asks you what to do.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;Test run: npm test
FAIL tests/auth.test.ts
  ● Auth › POST /login › should return 429 after 5 attempts
    Expected: 429
    Received: 200

Diagnosing failure...
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It reads the test, traces the failure to a missing &lt;code&gt;@types/express-rate-limit&lt;/code&gt; dev dependency, installs it, re-runs. Pass.&lt;/p&gt;




&lt;h2&gt;
  
  
  Modes Reference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What the agent can do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Explore&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Read files, query code map, search, web search. No writes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Plan&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Everything in Explore + draft plans, create &lt;code&gt;.atlarix/ATLARIX_PLAN.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Build&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Full tool access — file writes, terminal, tests, MCP calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fix&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Focused on diagnosing and fixing specific errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Review&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Read-only analysis, code quality feedback, architectural suggestions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;You can switch modes mid-session. The agent's context persists across the switch.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Works Well With Local Models
&lt;/h2&gt;

&lt;p&gt;If you're running Ollama, here's what I've found works well and what doesn't:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Works great:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Explore mode — querying the code map, finding files, understanding architecture. Even a 7B model does this well when it has the graph.&lt;/li&gt;
&lt;li&gt;Simple, scoped Build tasks — "add a field to this schema and update the related API endpoint"&lt;/li&gt;
&lt;li&gt;Fix mode — diagnosing TypeScript errors with LSP output injected into context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Works better with larger models:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-step autonomous plans across many files&lt;/li&gt;
&lt;li&gt;Complex refactors with non-obvious dependency chains&lt;/li&gt;
&lt;li&gt;Tasks that require reasoning about edge cases and side effects&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical threshold I've found: for anything touching more than 10 files or involving significant architectural decisions, I switch from my local 7B to a Standard tier cloud model. For everything else, local is fast and free.&lt;/p&gt;




&lt;h2&gt;
  
  
  Tips
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Use &lt;code&gt;.atlarix/ATLARIX.md&lt;/code&gt; for persistent context.&lt;/strong&gt; This file gets injected into every session for this workspace. Put your tech stack, conventions, and any context the agent should always know.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Project Context&lt;/span&gt;

Stack: Node.js, Express, TypeScript, PostgreSQL, Prisma
Auth: JWT with refresh tokens, stored in httpOnly cookies
Testing: Jest + Supertest
Conventions: 
&lt;span class="p"&gt;-&lt;/span&gt; All database queries go through service layer, never directly in routes
&lt;span class="p"&gt;-&lt;/span&gt; Error handling via custom AppError class in src/lib/errors.ts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Start tasks in Explore mode.&lt;/strong&gt; Even for tasks you think are simple, a quick explore turn first means the agent navigates correctly from the start rather than making wrong assumptions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reject and explain, don't just reject.&lt;/strong&gt; When you reject an approval queue item, add a reason. The agent uses it to replan. "Reject — use the existing &lt;code&gt;AppError&lt;/code&gt; class for error handling, not a raw &lt;code&gt;throw&lt;/code&gt;" gets you a better next attempt than a silent reject.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check the Blueprint canvas on new codebases.&lt;/strong&gt; &lt;em&gt;(Removed — there is no canvas.)&lt;/em&gt; The visual graph in the Blueprint tab is useful for understanding unfamiliar repos. Filter by file type, zoom into a module, and let the graph show you the dependency shape before you start prompting.&lt;/p&gt;




&lt;h2&gt;
  
  
  Free Tier Summary
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Solo (Free)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Workspaces&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local models (Ollama, LM Studio)&lt;/td&gt;
&lt;td&gt;✓ Unlimited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model Bridge Speed tier&lt;/td&gt;
&lt;td&gt;✓ Included usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Core tools (file, terminal, search, blueprint)&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP marketplace&lt;/td&gt;
&lt;td&gt;✗ (1 manual MCP)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent Behaviors marketplace&lt;/td&gt;
&lt;td&gt;✗&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a solo developer using local models, the free tier is genuinely unlimited. No token caps on Ollama usage.&lt;/p&gt;




&lt;h2&gt;
  
  
  Getting Started
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Download from atlarix.dev&lt;/span&gt;
&lt;span class="c"&gt;# 2. Install CLI from Settings → General&lt;/span&gt;
&lt;span class="c"&gt;# 3. Open your project&lt;/span&gt;
atlarix &lt;span class="nb"&gt;.&lt;/span&gt;

&lt;span class="c"&gt;# 4. Connect Ollama (or paste an API key)&lt;/span&gt;
&lt;span class="c"&gt;# 5. Start in Explore mode, describe your codebase&lt;/span&gt;
&lt;span class="c"&gt;# 6. Switch to Build when you're ready to execute&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The learning curve is less steep than it sounds. The hardest part is resisting the urge to jump straight into Build mode — a quick Explore turn first makes everything else go smoother.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://atlarix.dev" rel="noopener noreferrer"&gt;atlarix.dev&lt;/a&gt; — macOS + Linux, free Solo tier&lt;/p&gt;

&lt;p&gt;Open source: &lt;a href="https://github.com/AmariahAK/atlarix-skills" rel="noopener noreferrer"&gt;atlarix-skills&lt;/a&gt; · &lt;a href="https://github.com/AmariahAK/atlarix-mcps" rel="noopener noreferrer"&gt;atlarix-mcps&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Questions? Drop them in the comments — happy to go deep on any part of this.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>typescript</category>
      <category>devtools</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How I reduced AI codebase context from 100K to 5K tokens using a graph-based RAG — and why I deleted it</title>
      <dc:creator>Amariah Kamau</dc:creator>
      <pubDate>Thu, 30 Apr 2026 18:43:26 +0000</pubDate>
      <link>https://dev.to/abishaiama/how-i-reduced-ai-codebase-context-from-100k-to-5k-tokens-using-a-graph-based-rag-4dbd</link>
      <guid>https://dev.to/abishaiama/how-i-reduced-ai-codebase-context-from-100k-to-5k-tokens-using-a-graph-based-rag-4dbd</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Historical — the architecture below was removed from Atlarix in v14.9.0 (10 July 2026).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Blueprint, the parser, the SQLite graph, BM25 node scoring and the file watcher were all deleted — roughly 5,246 lines — after the watcher was found to hold one file descriptor per watched file and reach about 12,365 on a large repository.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Atlarix today has no index of any kind.&lt;/strong&gt; Retrieval is purely lexical: bundled ripgrep (&lt;code&gt;grep&lt;/code&gt; and &lt;code&gt;glob&lt;/code&gt;), no embeddings, no vector store, no graph, nothing in the background. Current architecture: &lt;strong&gt;&lt;a href="https://www.atlarix.dev/docs" rel="noopener noreferrer"&gt;atlarix.dev/docs&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;Most AI coding tools lie to you about context.&lt;/p&gt;

&lt;p&gt;They say "I understand your codebase." What they actually do is dump as many files as possible into the context window and hope the model figures it out. That works on a 3-file project. It falls apart on anything real.&lt;/p&gt;

&lt;p&gt;Here's the problem I kept hitting: a medium-sized production codebase hits ~100K tokens when you try to feed it to an LLM. That's expensive, slow, and surprisingly lossy — models start hallucinating relationships between files that don't exist, missing the ones that do.&lt;br&gt;
So I built a different approach inside Atlarix. Here's exactly how it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Parse the codebase into a typed graph (Round-Trip Engineering)
&lt;/h2&gt;

&lt;p&gt;When you open a project in Atlarix, it runs a parser over every file using Tree-sitter AST. Instead of storing raw text, it extracts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every function, class, interface, and type&lt;/li&gt;
&lt;li&gt;Import/export relationships between files&lt;/li&gt;
&lt;li&gt;Call relationships between functions&lt;/li&gt;
&lt;li&gt;File-level dependency edges&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This got stored as a typed node/edge graph in SQLite — what we called the Blueprint. &lt;em&gt;(Removed in v14.9.0. There is no SQLite graph and no Blueprint today.)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Query the graph, not the files
&lt;/h2&gt;

&lt;p&gt;When you asked Atlarix something, it didn't re-read your files — it queried the Blueprint graph. &lt;em&gt;(Today it does re-read files, deliberately: ripgrep over the workspace, no index in between.)&lt;/em&gt;&lt;br&gt;
We use BM25 scoring to rank nodes by relevance to the query. Only the top-scoring nodes get passed to the LLM. Everything else stays in SQLite.&lt;br&gt;
Result: instead of ~100K tokens, the average query uses ~5K tokens. That's a 95% reduction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Hierarchical context for complex queries
&lt;/h2&gt;

&lt;p&gt;For simple queries, BM25 node retrieval is enough. For complex ones — multi-file refactors, architecture questions — we use a three-layer hierarchy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mermaid diagram&lt;/strong&gt; — a persistent high-level map of the whole codebase, always in context&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BM25-scored Blueprint nodes&lt;/strong&gt; — targeted retrieval for the specific query&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fast-model compression&lt;/strong&gt; — if context hits 70% capacity, a fast model compresses the least relevant nodes before passing to the main model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This keeps context tight regardless of project size.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Provider-agnostic by design
&lt;/h2&gt;

&lt;p&gt;The Blueprint RAG layer sits beneath every AI provider. Whether you're using GPT-4o, Claude, Gemini, Groq, or a local Ollama model — the same 5K token context gets served. You're not paying for the model to re-read your whole codebase on every message.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters beyond token cost
&lt;/h2&gt;

&lt;p&gt;The token reduction is the headline number. But the real win is accuracy.&lt;br&gt;
When you give an LLM 100K tokens, it attends to all of it roughly equally. The file you actually care about is competing with 200 other files for attention. With 5K targeted tokens, the model is working with exactly what's relevant. Responses are more precise, edits land in the right place, and hallucinated file paths basically disappear.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Atlarix v7 added parallel agents named Research, Architect, Builder, Reviewer and Debugger on top of this foundation. &lt;strong&gt;Those names are long gone and are a common source of stale descriptions of Atlarix.&lt;/strong&gt; The work modes today are Explore, Plan, Build, Debug and Review, and the &lt;code&gt;task&lt;/code&gt; tool spawns scoped workers whose edits come back as reviewable proposals rather than writes.&lt;br&gt;
We also just shipped Windows support today — so Atlarix now runs on Mac, Windows, and Linux.&lt;br&gt;
If you're building something where AI context management is a bottleneck, I'd love to compare notes. Try it at atlarix.dev or ask me anything in the comments&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>buildinpublic</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
