<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rost</title>
    <description>The latest articles on DEV Community by Rost (@rosgluk).</description>
    <link>https://dev.to/rosgluk</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3544400%2F04dd81bf-749e-4055-971f-316c0134e76c.jpg</url>
      <title>DEV Community: Rost</title>
      <link>https://dev.to/rosgluk</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rosgluk"/>
    <language>en</language>
    <item>
      <title>gstack: AI Software Engineering Stack</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Mon, 28 Sep 2026 05:33:00 +0000</pubDate>
      <link>https://dev.to/rosgluk/gstack-ai-software-engineering-stack-49ja</link>
      <guid>https://dev.to/rosgluk/gstack-ai-software-engineering-stack-49ja</guid>
      <description>&lt;p&gt;AI coding agents can already write functions, modify repositories, run tests, and open pull requests. The harder problem is getting an agent to follow a repeatable engineering process before, during, and after the code is written.&lt;/p&gt;

&lt;p&gt;gstack, a project from Garry Tan originally built around Claude Code, takes a different route: instead of replacing your coding agent with another platform, it wraps the agent you already use with specialized skills, browser tooling, reviews, safety controls, and release processes. The project describes the result as a virtual engineering team -- twenty-three specialists and eight power tools, all slash commands, all Markdown, MIT licensed.&lt;/p&gt;

&lt;p&gt;The comparison targets depend on which part of gstack you need: skill collections such as Superpowers, specification systems such as OpenSpec and GitHub Spec Kit, methodologies such as BMAD, orchestration platforms such as Ruflo, or your own maintained set of agent skills. Several of these combine with gstack rather than replace it, and the wider ecosystem they all belong to is mapped in the &lt;a href="https://www.glukhov.org/ai-devtools/" rel="noopener noreferrer"&gt;AI Developer Tools hub&lt;/a&gt; of this site.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is gstack?
&lt;/h2&gt;

&lt;p&gt;gstack is an open-source collection of AI engineering workflows. In gstack's framing, software development consists of several different kinds of reasoning, and asking one generic coding prompt to perform all of them is a poor abstraction -- so the project exposes specialized skills instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;product exploration&lt;/li&gt;
&lt;li&gt;product and CEO review&lt;/li&gt;
&lt;li&gt;architecture review&lt;/li&gt;
&lt;li&gt;developer experience review&lt;/li&gt;
&lt;li&gt;design review&lt;/li&gt;
&lt;li&gt;implementation review&lt;/li&gt;
&lt;li&gt;browser-based QA&lt;/li&gt;
&lt;li&gt;investigation and debugging&lt;/li&gt;
&lt;li&gt;security analysis&lt;/li&gt;
&lt;li&gt;documentation&lt;/li&gt;
&lt;li&gt;benchmarking&lt;/li&gt;
&lt;li&gt;release preparation&lt;/li&gt;
&lt;li&gt;deployment&lt;/li&gt;
&lt;li&gt;retrospectives&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The project presents these roles as a virtual engineering team: a CEO who rethinks the product, an eng manager who locks architecture, a designer who catches "AI slop", a reviewer who finds production bugs, a QA lead who opens a real browser, a security officer who runs OWASP and STRIDE audits, and a release engineer who ships the PR. The toolchain around the skill definitions is TypeScript and Bun: a setup script, generated skill documentation, session hooks, state under &lt;code&gt;~/.gstack/&lt;/code&gt;, a bundled browser, and a set of standalone CLIs.&lt;/p&gt;

&lt;p&gt;gstack is not a foundation model and not a replacement for Claude Code; it is a process layer running on top of an agent harness:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[LLM] --&amp;gt; B[Claude Code or another supported harness]
    B --&amp;gt; C[gstack skills and workflow rules]
    subgraph G[gstack process stages]
        D1[Planning]
        D2[Architecture review]
        D3[Design review]
        D4[Code review]
        D5[Browser QA]
        D6[Security]
        D7[Release and deployment]
        D8[Learning and memory]
    end
    C --&amp;gt; G
    G --&amp;gt; E[Git repository, browser, and development tools]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Two structural properties decide where gstack fits. First, the project describes it as a process rather than a collection of tools: the skills run in the order a sprint runs -- think, plan, build, review, test, ship, reflect -- and each skill hands its artifacts to the next, so &lt;code&gt;/office-hours&lt;/code&gt; writes a design doc that &lt;code&gt;/plan-ceo-review&lt;/code&gt; reads, and &lt;code&gt;/plan-eng-review&lt;/code&gt; writes a test plan that &lt;code&gt;/qa&lt;/code&gt; picks up. Second, gstack is not Claude Code-only: &lt;code&gt;./setup&lt;/code&gt; auto-detects the agents installed on the machine, and &lt;code&gt;./setup --host &amp;lt;name&amp;gt;&lt;/code&gt; targets Codex CLI, OpenCode, Cursor, Factory Droid, Kiro, Slate, OpenClaw, and Hermes, while a 2KB instruction-only digest in the repo covers rules-reading agents that need no install at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why gstack Exists
&lt;/h2&gt;

&lt;p&gt;A blank Claude Code session is extremely flexible, and that flexibility is also one of its weaknesses. Consider a feature request such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Add organization-level API tokens to the application.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A capable agent might immediately inspect the repository and start modifying authentication code, whereas a senior engineer would first ask who owns the tokens, whether users can belong to multiple organizations, how tokens are revoked, whether permissions are inherited, what happens to existing authentication, whether tokens should expire, how secrets are displayed, and which audit events are required. The agent may discover some of these questions eventually, but there is no guarantee it will discover them before implementation. gstack moves that discipline into reusable workflows: instead of &lt;code&gt;idea -&amp;gt; coding agent -&amp;gt; code&lt;/code&gt;, the change passes product review, technical planning, architecture review, implementation, code review, browser QA, and release as named stages, each with its own command (the full sequence is in A Practical gstack Workflow below).&lt;/p&gt;

&lt;p&gt;This does not make the AI correct; it changes the probability distribution of its mistakes. The agent is pushed to challenge assumptions earlier, inspect evidence, review its own work from several perspectives, and verify the application instead of stopping when the code compiles.&lt;/p&gt;

&lt;h2&gt;
  
  
  How gstack Turns Markdown Skills into Process
&lt;/h2&gt;

&lt;p&gt;Most of gstack's capabilities originate in Markdown skill definitions. An agent skill can describe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;when it should run&lt;/li&gt;
&lt;li&gt;what context it should inspect&lt;/li&gt;
&lt;li&gt;what questions it should ask&lt;/li&gt;
&lt;li&gt;which tools it may use&lt;/li&gt;
&lt;li&gt;which commands it should execute&lt;/li&gt;
&lt;li&gt;which evidence it must gather&lt;/li&gt;
&lt;li&gt;which checks must pass&lt;/li&gt;
&lt;li&gt;how the result should be structured&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Such a skill acts somewhere between documentation, a reusable prompt, a standard operating procedure, and executable workflow configuration. The underlying mechanics are covered in &lt;a href="https://www.glukhov.org/ai-devtools/claude-code/claude-skills-for-developers/" rel="noopener noreferrer"&gt;Claude Skills and SKILL.md for Developers&lt;/a&gt;; gstack adds infrastructure on top of that primitive: generated skill definitions, startup and completion hooks, state management under &lt;code&gt;~/.gstack/&lt;/code&gt;, browser automation, safety mechanisms, repository inspection, opt-in telemetry, cross-session memory managed by &lt;code&gt;/learn&lt;/code&gt;, and optional persistent knowledge through the separate GBrain project, which &lt;code&gt;/setup-gbrain&lt;/code&gt; can stand up as a local PGLite database, a Supabase project, or a remote MCP endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Most Important gstack Skills
&lt;/h2&gt;

&lt;p&gt;The exact collection changes quickly, but several workflows illustrate how the system is intended to be used.&lt;/p&gt;

&lt;h3&gt;
  
  
  office-hours
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/office-hours&lt;/code&gt; belongs near the beginning of a project or feature. It runs six forcing questions about the problem before any code is written, and in the README's worked example it reframes a request for a "daily briefing app" into a personal chief-of-staff AI, then writes the design doc that every downstream skill reads. For vague input like "we need better project search", the output is a requirement, not code.&lt;/p&gt;

&lt;h3&gt;
  
  
  plan-ceo-review
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/plan-ceo-review&lt;/code&gt; examines the product-level assumptions behind a plan. It works in four scope modes -- Expansion, Selective Expansion, Hold Scope, Reduction -- and can challenge scope, identify missing opportunities, reduce unnecessary work, or suggest framing the problem differently. It runs before the requirements are fixed, a stage most coding-agent tools do not have.&lt;/p&gt;

&lt;h3&gt;
  
  
  plan-eng-review
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/plan-eng-review&lt;/code&gt; shifts the perspective toward engineering: architecture, data flow, diagrams, edge cases, a test matrix, failure modes, and security concerns. It stays separate from the product review because merging both into one large prompt makes the model mix product decisions with implementation decisions.&lt;/p&gt;

&lt;h3&gt;
  
  
  plan-design-review and design-review
&lt;/h3&gt;

&lt;p&gt;gstack treats visual and interaction design as a separate discipline. &lt;code&gt;/plan-design-review&lt;/code&gt; rates each design dimension from 0 to 10, describes what a 10 looks like, and edits the plan to close the gap, with "AI slop" detection as a named check. The later &lt;code&gt;/design-review&lt;/code&gt; runs the same audit against the actual implementation and fixes what it finds with atomic commits and before/after screenshots. For web applications, both pair with gstack's browser automation.&lt;/p&gt;

&lt;h3&gt;
  
  
  review
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/review&lt;/code&gt; performs engineering review against repository changes from a staff-engineer perspective: it auto-fixes the obvious findings, flags the rest for approval, and keeps an advisory simplification lens for over-built code. Code written successfully is not necessarily code that should be merged.&lt;/p&gt;

&lt;h3&gt;
  
  
  investigate
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/investigate&lt;/code&gt; enforces a systematic debugging rule the project calls the Iron Law: no fixes without investigation. It traces data flow, tests hypotheses, and stops after three failed fix attempts instead of continuing to thrash. It also auto-activates &lt;code&gt;/freeze&lt;/code&gt;, which locks edits to the module under investigation.&lt;/p&gt;

&lt;h3&gt;
  
  
  qa and qa-only
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/qa&lt;/code&gt; has the agent operate a browser, interact with the application, find bugs, fix them with atomic commits, re-verify, and generate a regression test for every fix. &lt;code&gt;/qa-only&lt;/code&gt; runs the same methodology report-only. Many coding agents stop verification at "tests passed"; for a web application the browser is where integration errors, layout problems, incorrect flows, authentication failures, and JavaScript exceptions usually become visible.&lt;/p&gt;

&lt;h3&gt;
  
  
  ship, land-and-deploy, and canary
&lt;/h3&gt;

&lt;p&gt;The release chain is three skills rather than one. &lt;code&gt;/ship&lt;/code&gt; syncs main, runs tests, audits coverage, pushes, and opens the pull request, bootstrapping a test framework if the project has none. &lt;code&gt;/land-and-deploy&lt;/code&gt; merges, waits for CI and the deployment, and verifies production health. &lt;code&gt;/canary&lt;/code&gt; then runs a post-deploy monitoring loop that watches for console errors, performance regressions, and page failures.&lt;/p&gt;

&lt;h3&gt;
  
  
  autoplan, spec, learn, and retro
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;/autoplan&lt;/code&gt; runs the CEO, design, DX, and engineering review pipeline automatically -- engineering always last, so the shipping gate reviews the final amended plan -- and surfaces only taste decisions for approval. &lt;code&gt;/spec&lt;/code&gt; turns vague intent into a precise, executable spec in five phases (why, scope, technical with mandatory code-reading, draft, file) with an outside-review quality gate before filing. &lt;code&gt;/learn&lt;/code&gt; manages what gstack has learned across sessions -- patterns, pitfalls, and preferences -- with review, search, prune, and export. &lt;code&gt;/retro&lt;/code&gt; produces a team-aware weekly retro; &lt;code&gt;/retro global&lt;/code&gt; runs it across all your projects and AI tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Browser Automation in gstack
&lt;/h2&gt;

&lt;p&gt;On supported macOS systems (macOS 15+), gstack drives the &lt;a href="https://aside.com" rel="noopener noreferrer"&gt;Aside&lt;/a&gt; browser first -- your real browser, with your real logged-in sessions, in tabs the agent opens for itself and closes when done. When Aside is unavailable, gstack falls back to its own Chromium-based engine, which &lt;code&gt;./setup&lt;/code&gt; builds and which runs a persistent daemon rather than launching a fresh browser for every command:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Coding Agent] --&amp;gt; B[gstack Browser CLI]
    B --&amp;gt; C[Local Browser Service]
    C --&amp;gt; D[Aside or bundled Chromium]
    D --&amp;gt; E[Application]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Persistent browser state lets cookies, authentication sessions, and tabs survive between operations, which makes browser-based QA practical. &lt;code&gt;/open-gstack-browser&lt;/code&gt; exposes the fallback engine headed, with a sidebar agent that routes fast actions (click, navigate, screenshot) to Sonnet and reading or analysis to Opus. When the agent hits a CAPTCHA, an auth wall, or an MFA prompt, &lt;code&gt;$B handoff&lt;/code&gt; opens a visible browser at the same page with cookies and tabs intact; you solve it, and &lt;code&gt;$B resume&lt;/code&gt; continues where the agent left off. The agent suggests a handoff automatically after three consecutive failures. &lt;code&gt;/pair-agent&lt;/code&gt; shares the browser with other agents -- OpenClaw, Hermes, Codex, Cursor, or anything that can curl -- with scoped tokens, tab isolation, rate limiting, and per-tab activity attribution.&lt;/p&gt;

&lt;p&gt;The persistent engine also increases the security surface, since an agent with access to authenticated sessions holds a meaningful privilege. gstack ships a layered prompt-injection defense for this: content filters (datamarking, hidden-element stripping, ARIA scrubbing, URL blocklist) on every page read, plus a local ML classifier in a sidecar subprocess that scans page-derived content before the agent sees it, with a verdict combiner that requires classifier agreement before blocking. Page content is treated as untrusted input -- the agent takes syntax from a page, never instructions. Checks before and while using browser-driven QA:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Decide in advance which authenticated environments the agent may operate in, and prefer an isolated profile for QA work when the fallback engine is in use.&lt;/li&gt;
&lt;li&gt;Know the emergency kill switch: &lt;code&gt;GSTACK_SECURITY_OFF=1&lt;/code&gt; disables the security layer -- do not leave it set.&lt;/li&gt;
&lt;li&gt;The persistent daemon retains cookies and sessions between runs, so stop it when you are done and confirm nothing is left running, for example &lt;code&gt;ps aux | grep -i chrom&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Read the hooks and safety mechanisms in the cloned repository before enabling them -- they are plain files, so review them the way you would review CI configuration.&lt;/li&gt;
&lt;li&gt;After updating the clone, re-run &lt;code&gt;./setup&lt;/code&gt; so generated components stay in sync with the skill definitions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Safety Guardrails and Second Opinions
&lt;/h2&gt;

&lt;p&gt;Three power tools act as session-level safety switches. &lt;code&gt;/careful&lt;/code&gt; warns before destructive commands -- &lt;code&gt;rm -rf&lt;/code&gt;, &lt;code&gt;DROP TABLE&lt;/code&gt;, force-push, &lt;code&gt;git reset --hard&lt;/code&gt; -- and activates by saying "be careful"; recursive deletes of the root or home directory and force-pushes to the default branch are hard-denied. &lt;code&gt;/freeze&lt;/code&gt; restricts file edits to one directory so the agent cannot "fix" unrelated code while debugging, and &lt;code&gt;/guard&lt;/code&gt; activates both at once.&lt;/p&gt;

&lt;p&gt;Second-opinion reviews cross harnesses: on Claude Code, &lt;code&gt;/codex&lt;/code&gt; sends the work to OpenAI Codex CLI for an independent review, challenge, or consultation; on the other harnesses, &lt;code&gt;/claude-code&lt;/code&gt; does the reverse. Each report identifies the provider that actually completed the review.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Practical gstack Workflow
&lt;/h2&gt;

&lt;p&gt;You do not need every gstack skill for every change. A reasonable feature workflow:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Feature idea] --&amp;gt; B[office-hours]
    B --&amp;gt; C[plan-ceo-review]
    C --&amp;gt; D[Create implementation plan]
    D --&amp;gt; E[plan-eng-review]
    E --&amp;gt; F[Implement]
    F --&amp;gt; G[review]
    G --&amp;gt; H[qa]
    H --&amp;gt; I[ship]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Which review skills to add depends on who the software is for:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Building for&lt;/th&gt;
&lt;th&gt;Plan stage (before code)&lt;/th&gt;
&lt;th&gt;Live audit (after shipping)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;End users (UI, web app, mobile)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/plan-design-review&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/design-review&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developers (API, CLI, SDK, docs)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/plan-devex-review&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/devex-review&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Architecture (data flow, perf)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/plan-eng-review&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/review&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;All of the above&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/autoplan&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;--&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a trivial bug fix, going straight to investigate, implementation, review, and tests is often enough. A tool-neutral version of the same shape -- spec, design, tasks, implement, validate -- is in &lt;a href="https://www.glukhov.org/app-architecture/documentation/spec-driven-development-workflow/" rel="noopener noreferrer"&gt;Spec-Driven Development Workflow From Requirements to Code&lt;/a&gt;. The project's README describes running ten to fifteen of these sprints in parallel, each in its own isolated workspace; the sprint structure is what the project says keeps parallel agents from becoming sources of chaos.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installing gstack
&lt;/h2&gt;

&lt;p&gt;The current installation expects a working Claude Code setup, Git, Bun v1.0+, and, on Windows, Node.js -- Bun has a known bug with Playwright's pipe transport on Windows, so the browse server falls back to Node.js there. If you have not configured Claude Code yet, start with the &lt;a href="https://www.glukhov.org/ai-devtools/claude-code/" rel="noopener noreferrer"&gt;Claude Code overview&lt;/a&gt; first. On macOS, the Aside browser (macOS 15+) is recommended for the browser skills; without it, the bundled Chromium daemon is used.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Verify your prerequisites: &lt;code&gt;git --version&lt;/code&gt; and &lt;code&gt;bun --version&lt;/code&gt; (the toolchain is Bun-based).&lt;/li&gt;
&lt;li&gt;Clone gstack into the Claude skills directory:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   git clone &lt;span class="nt"&gt;--single-branch&lt;/span&gt; &lt;span class="nt"&gt;--depth&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
     https://github.com/garrytan/gstack.git &lt;span class="se"&gt;\&lt;/span&gt;
     ~/.claude/skills/gstack
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Run the setup script from the cloned directory:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;cd&lt;/span&gt; ~/.claude/skills/gstack
   ./setup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Setup installs and generates the components required by the supported skills and builds the bundled browser; a Chromium install failure is best-effort, setup records the reason, finishes registering every skill, and prints which skills are affected.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add a &lt;code&gt;## gstack&lt;/code&gt; section to the project's &lt;code&gt;CLAUDE.md&lt;/code&gt;. The project's install instructions include this step, and it is what makes Claude Code route the skills: use &lt;code&gt;/browse&lt;/code&gt; from gstack for all web browsing, never use &lt;code&gt;mcp__claude-in-chrome__*&lt;/code&gt; tools, and list the available skills.&lt;/li&gt;
&lt;li&gt;Verify the install: check that the generated files are present in the cloned directory, start a Claude Code session, and run &lt;code&gt;/office-hours&lt;/code&gt; on a scratch project to confirm the skill is recognized.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Team mode
&lt;/h3&gt;

&lt;p&gt;For repositories, gstack offers a team-oriented setup where developers share one workflow instead of individually configured environments:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ~/.claude/skills/gstack &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; ./setup &lt;span class="nt"&gt;--team&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  ~/.claude/skills/gstack/bin/gstack-team-init required &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  git add .claude/ CLAUDE.md &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  git commit &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"require gstack for AI-assisted work"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;required&lt;/code&gt; blocks AI-assisted work in the repo without gstack; swap it for &lt;code&gt;optional&lt;/code&gt; to nudge teammates instead of blocking them. No files are vendored into the repo: every Claude Code session starts with a fast auto-update check (throttled to once per hour, network-failure-safe, silent), which removes version drift across the team. Personal configuration improves one developer; repository-level configuration creates a shared engineering convention.&lt;/p&gt;

&lt;h3&gt;
  
  
  Other harnesses, upgrades, and uninstall
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Other agents: &lt;code&gt;./setup --host codex&lt;/code&gt;, &lt;code&gt;--host opencode&lt;/code&gt;, &lt;code&gt;--host cursor&lt;/code&gt;, &lt;code&gt;--host factory&lt;/code&gt;, &lt;code&gt;--host kiro&lt;/code&gt;, &lt;code&gt;--host slate&lt;/code&gt;, &lt;code&gt;--host openclaw&lt;/code&gt;, and &lt;code&gt;--host hermes&lt;/code&gt; install the skills into each agent's own skills directory. The 2KB instruction-only digest at &lt;code&gt;agents-digest/gstack-AGENTS.md&lt;/code&gt; covers agents that only read rules files.&lt;/li&gt;
&lt;li&gt;Command naming: skills register with short names by default (&lt;code&gt;/qa&lt;/code&gt;, &lt;code&gt;/review&lt;/code&gt;); &lt;code&gt;./setup --prefix&lt;/code&gt; switches to namespaced names (&lt;code&gt;/gstack-qa&lt;/code&gt;), which matters when you run other skill packs alongside gstack.&lt;/li&gt;
&lt;li&gt;Upgrades: re-run &lt;code&gt;./setup&lt;/code&gt; after a &lt;code&gt;git pull&lt;/code&gt; (required on Windows, where installs are file copies), or use the &lt;code&gt;/gstack-upgrade&lt;/code&gt; skill; setting &lt;code&gt;auto_upgrade: true&lt;/code&gt; in &lt;code&gt;~/.gstack/config.yaml&lt;/code&gt; keeps the install current automatically.&lt;/li&gt;
&lt;li&gt;Telemetry is off by default and asks for opt-in on first run. If you opt in, it sends skill name, duration, success/fail, gstack version, and OS -- never code, file paths, repo names, or prompts. &lt;code&gt;gstack-config set telemetry off&lt;/code&gt; disables it at any time.&lt;/li&gt;
&lt;li&gt;Uninstall: &lt;code&gt;~/.claude/skills/gstack/bin/gstack-uninstall&lt;/code&gt; removes skills, symlinks, &lt;code&gt;~/.gstack/&lt;/code&gt; state, project-local state, browse daemons, and hook registrations.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Verify the install and fix common failures
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;./setup&lt;/code&gt; fails&lt;/strong&gt; -- confirm Bun is on your PATH with &lt;code&gt;bun --version&lt;/code&gt;; the generated components are built by the Bun toolchain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills not recognized by Claude Code&lt;/strong&gt; -- confirm the clone actually lives in &lt;code&gt;~/.claude/skills/gstack&lt;/code&gt;, that the project's &lt;code&gt;CLAUDE.md&lt;/code&gt; has a gstack section, and re-run &lt;code&gt;./setup&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/browse&lt;/code&gt; reports &lt;code&gt;NEED_ASIDE&lt;/code&gt; or &lt;code&gt;ASIDE_NOT_RUNNING&lt;/code&gt;&lt;/strong&gt; -- the probe is telling you it will use the fallback browser. That is normal on Linux and Windows; on macOS it means Aside is not open or signed in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The fallback browser fails&lt;/strong&gt; -- &lt;code&gt;cd ~/.claude/skills/gstack &amp;amp;&amp;amp; bun install &amp;amp;&amp;amp; bun run build&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stale install after an update&lt;/strong&gt; -- run &lt;code&gt;/gstack-upgrade&lt;/code&gt;, or set &lt;code&gt;auto_upgrade: true&lt;/code&gt; in &lt;code&gt;~/.gstack/config.yaml&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Trying gstack Without Adopting Everything
&lt;/h2&gt;

&lt;p&gt;Use gstack on a real but non-critical feature rather than migrating your development process. The project's quick start is the same trial, and it ends with "stop there":&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;/office-hours&lt;/code&gt; -- problem definition&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/plan-ceo-review&lt;/code&gt; -- product reasoning&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/review&lt;/code&gt; -- engineering verification, after implementation&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/qa&lt;/code&gt; -- runtime verification, for web projects&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If those stages surface findings your normal Claude Code workflow misses, the rest of the system is worth exploring; if they mostly produce additional text without changing engineering decisions, adopting the entire stack probably will not help.&lt;/p&gt;

&lt;h2&gt;
  
  
  What gstack Does Well
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Separated engineering roles.&lt;/strong&gt; Instead of one giant "be a senior engineer" instruction, product strategy, architecture, UX, QA, security, and release engineering each get their own reasoning mode.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verification, not just generation.&lt;/strong&gt; Review, browser QA with regression-test generation, security audits, benchmarking, and the ship-deploy-canary chain are first-class workflows in gstack rather than optional afterthoughts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inspectable.&lt;/strong&gt; Much of the behavioral layer is plain Markdown files developers can read and modify, unlike the internal workflows of a proprietary autonomous agent. The repo also ships audit tooling for the stack itself: &lt;code&gt;gstack-context-bill&lt;/code&gt; reports what an installed skill tree costs in tokens, and &lt;code&gt;gstack-egress&lt;/code&gt; writes a hash-chained receipt for every off-machine send, telemetry included.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Team infrastructure.&lt;/strong&gt; Skills can encode engineering conventions -- instead of typing&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Remember to check API compatibility, run integration tests,
inspect the browser console, and update the changelog.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in every session, the requirements live in a reusable workflow, and team mode makes that workflow a repository requirement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where gstack Can Be Too Much
&lt;/h2&gt;

&lt;p&gt;gstack is intentionally opinionated, and that limits its fit: a mature organization may already have architecture review procedures, release tooling, CI gates, QA automation, security scanning, ADR conventions, specification templates, and code review policies, and adding another complete methodology on top creates overlap instead of clarity.&lt;/p&gt;

&lt;p&gt;There is also a context and token cost: each additional review stage adds repository inspection, model reasoning, and potentially more external model calls. The goal is the minimum reliable process needed to ship correct software, not a maximum number of AI reviews; &lt;code&gt;gstack-context-bill&lt;/code&gt; can quantify what your installed skill set actually costs per session before you decide how much of it to keep.&lt;/p&gt;

&lt;p&gt;gstack works best as a toolbox whose workflows you select and adapt, not as ceremony for every commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  gstack Alternatives and Combinations
&lt;/h2&gt;

&lt;p&gt;The closest alternatives, and the layer each occupies:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Primary focus&lt;/th&gt;
&lt;th&gt;Workflow style&lt;/th&gt;
&lt;th&gt;Agent portability&lt;/th&gt;
&lt;th&gt;Best fit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gstack&lt;/td&gt;
&lt;td&gt;Full engineering workflow&lt;/td&gt;
&lt;td&gt;Role-oriented skills and tools&lt;/td&gt;
&lt;td&gt;10 agents via &lt;code&gt;./setup --host&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;End-to-end AI-assisted engineering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Superpowers&lt;/td&gt;
&lt;td&gt;Engineering methodology&lt;/td&gt;
&lt;td&gt;Automatic composable skills&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Disciplined coding and TDD&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenSpec&lt;/td&gt;
&lt;td&gt;Change specifications&lt;/td&gt;
&lt;td&gt;Lightweight spec artifacts&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Brownfield feature development&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Spec Kit&lt;/td&gt;
&lt;td&gt;Spec-driven development&lt;/td&gt;
&lt;td&gt;Structured multi-stage workflow&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Formal requirements-to-code process&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BMAD Method&lt;/td&gt;
&lt;td&gt;AI-driven agile development&lt;/td&gt;
&lt;td&gt;Adaptive roles and workflows&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Larger end-to-end projects&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ruflo&lt;/td&gt;
&lt;td&gt;Multi-agent orchestration&lt;/td&gt;
&lt;td&gt;Agents, swarms, memory&lt;/td&gt;
&lt;td&gt;Platform-oriented&lt;/td&gt;
&lt;td&gt;Parallel autonomous agent systems&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom skills&lt;/td&gt;
&lt;td&gt;Your own process&lt;/td&gt;
&lt;td&gt;Fully customizable&lt;/td&gt;
&lt;td&gt;Potentially very high&lt;/td&gt;
&lt;td&gt;Mature teams with established practices&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  gstack + Superpowers: implementation discipline inside the roles
&lt;/h3&gt;

&lt;p&gt;Both are skill frameworks, so they overlap the most. The division of labor when combining them: gstack supplies the surrounding roles -- product, design, QA, release -- while Superpowers supplies the discipline inside the implementation phase (TDD, planning before implementation, systematic debugging, subagent review). Install both skill sets, then trim the overlapping skills so the agent never sees two conflicting instructions for the same phase; if command names collide, install gstack with &lt;code&gt;./setup --prefix&lt;/code&gt; so its skills register as &lt;code&gt;/gstack-*&lt;/code&gt; and coexist with the other pack. Install and workflow details are in the &lt;a href="https://www.glukhov.org/ai-devtools/superpowers/" rel="noopener noreferrer"&gt;Superpowers quickstart&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  gstack + OpenSpec: durable specs, live reviews
&lt;/h3&gt;

&lt;p&gt;OpenSpec keeps human and agent aligned around explicit change specifications -- artifacts for the proposed change, specifications, design decisions, and implementation tasks. The key property is persistence: a chat conversation disappears into context history, but a specification remains in the repository where humans and future agent sessions can review it. gstack adds the product review before the spec exists and the review and QA after it is implemented:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Feature request] --&amp;gt; B[gstack product review]
    B --&amp;gt; C[OpenSpec change]
    C --&amp;gt; D[Implementation]
    D --&amp;gt; E[gstack review]
    E --&amp;gt; F[gstack QA]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;A concrete sequence: run &lt;code&gt;/office-hours&lt;/code&gt; and &lt;code&gt;/plan-ceo-review&lt;/code&gt;, capture the outcome as an OpenSpec change, implement against it, then run &lt;code&gt;/review&lt;/code&gt; and &lt;code&gt;/qa&lt;/code&gt;. Note that gstack also ships its own &lt;code&gt;/spec&lt;/code&gt; skill, which archives specs under &lt;code&gt;~/.gstack&lt;/code&gt;; if OpenSpec owns the specification, keep gstack's &lt;code&gt;/spec&lt;/code&gt; out of the loop so the two do not diverge. The &lt;a href="https://www.glukhov.org/ai-devtools/openspec/" rel="noopener noreferrer"&gt;OpenSpec quickstart&lt;/a&gt; covers the explore-propose-apply-archive loop in detail.&lt;/p&gt;

&lt;h3&gt;
  
  
  gstack + GitHub Spec Kit: pick one planning backbone
&lt;/h3&gt;

&lt;p&gt;Spec Kit's core workflow is a sequence of explicit stages -- constitution, specify, plan, tasks, implement, converge -- and it has expanded into bug-fixing, idea-assessment, extensions, presets, and integrations. Since both Spec Kit and gstack center the planning stage, running both full flows duplicates work. If requirements traceability and formal stages matter, let Spec Kit own the specification backbone and use gstack for the layers Spec Kit does not enforce -- product review, design review, browser QA, and shipping. A broader comparison of spec-driven setups, including Kiro and Claude Code, is in &lt;a href="https://www.glukhov.org/ai-devtools/ai-coding-assistants/spec-kit-vs-kiro-vs-claude-code/" rel="noopener noreferrer"&gt;GitHub Spec Kit vs Kiro vs Claude Code SDD Workflows&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  BMAD and Ruflo: different axes
&lt;/h3&gt;

&lt;p&gt;BMAD is a broader AI-driven development methodology whose adaptive workflows cover product thinking, specifications, architecture, and implementation, scaling the ceremony to the size of the work. It and gstack both play the process-backbone role, so choose one as the backbone rather than running both in full; gstack's individual skills can still be selected alongside a methodology.&lt;/p&gt;

&lt;p&gt;Ruflo targets multi-agent orchestration: coordinated workers, shared memory, swarms. gstack applies multiple specialist perspectives to one engineering workflow; an orchestration platform applies multiple executing agents to one engineering objective. The boundary blurs -- gstack can call external tools and additional models, and orchestrators can implement structured engineering roles -- but the decision is independent: if the problem is that the agent skips engineering discipline, a skill framework is the direct fix; if it is running ten agents concurrently across many tasks and repositories, an orchestrator sits above a workflow like gstack rather than replacing it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Custom skills: the most customizable layer
&lt;/h3&gt;

&lt;p&gt;You can also skip the framework entirely and create a small collection of skills for the procedures your team already follows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;skills/
  architecture-review/
  api-review/
  database-migration-review/
  incident-analysis/
  release-check/
  security-review/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each skill encodes organization-specific knowledge a generic framework cannot know. A database migration skill can require rollback analysis, table-lock analysis, index impact review, migration duration estimation, deployment ordering, and compatibility with the previous application version; an API review skill can require backwards compatibility, authentication checks, pagination consistency, idempotency analysis, rate-limit behavior, and OpenAPI changes. A practical path: start from the gstack skills you actually use, copy their structure into your own &lt;code&gt;skills/&lt;/code&gt; directory, and rewrite the checks around your conventions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four Layers: Skills, Specifications, Methodologies, Orchestrators
&lt;/h2&gt;

&lt;p&gt;Four layers cover most of these tools, and they show how the combinations above fit together:&lt;/p&gt;

&lt;h3&gt;
  
  
  Skills answer "how should the agent behave?"
&lt;/h3&gt;

&lt;p&gt;gstack, Superpowers, and custom agent skills.&lt;/p&gt;

&lt;h3&gt;
  
  
  Specification systems answer "what exactly are we building?"
&lt;/h3&gt;

&lt;p&gt;OpenSpec and GitHub Spec Kit; the underlying spec-driven concepts and terminology are defined in &lt;a href="https://www.glukhov.org/app-architecture/documentation/what-is-spec-driven-development/" rel="noopener noreferrer"&gt;What Is Spec-Driven Development?&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Methodologies answer "how should the project move from idea to software?"
&lt;/h3&gt;

&lt;p&gt;BMAD, Superpowers, and parts of gstack.&lt;/p&gt;

&lt;h3&gt;
  
  
  Orchestrators answer "how should multiple agents execute work?"
&lt;/h3&gt;

&lt;p&gt;Ruflo and other multi-agent runtimes.&lt;/p&gt;

&lt;p&gt;The layers compose; a development environment can contain all four:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Product requirement] --&amp;gt; B[Specification system]
    B --&amp;gt; C[Engineering workflow]
    C --&amp;gt; D[Agent orchestrator]

    D --&amp;gt; E[Implementation agent]
    D --&amp;gt; F[Test agent]
    D --&amp;gt; G[Review agent]
    D --&amp;gt; H[QA agent]

    E --&amp;gt; I[Repository]
    F --&amp;gt; I
    G --&amp;gt; I
    H --&amp;gt; I&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;gstack already spans several of these boundaries.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Use gstack?
&lt;/h2&gt;

&lt;p&gt;The project's README describes the audience as technical founders and CEOs who still want to ship, first-time Claude Code users who want structured roles instead of a blank prompt, and tech leads and staff engineers who want rigorous review, QA, and release automation on every PR. gstack is worth trying if you use coding agents extensively and the limiting factor is no longer code generation itself. Typical symptoms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the agent starts implementing before understanding the problem&lt;/li&gt;
&lt;li&gt;implementation plans miss architectural implications&lt;/li&gt;
&lt;li&gt;generated code passes tests but fails in the browser&lt;/li&gt;
&lt;li&gt;reviews are inconsistent between sessions&lt;/li&gt;
&lt;li&gt;release steps are repeatedly forgotten&lt;/li&gt;
&lt;li&gt;different developers prompt the agent in completely different ways&lt;/li&gt;
&lt;li&gt;useful engineering instructions remain buried in CLAUDE.md files&lt;/li&gt;
&lt;li&gt;you repeatedly type the same review prompts manually&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your automation already provides strong deterministic gates and the agent only handles small, well-specified tasks, gstack adds little.&lt;/p&gt;

&lt;h2&gt;
  
  
  gstack and the Direction of AI Software Development
&lt;/h2&gt;

&lt;p&gt;The shift gstack sits in tracks in generations: code completion (2022-2023), coding agents (2024-2025), specifications and agent workflows (2025-2026), and programmable AI engineering organizations. The products will change, but the model remains one component; engineering quality increasingly depends on the surrounding system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;persistent specifications&lt;/li&gt;
&lt;li&gt;reusable skills&lt;/li&gt;
&lt;li&gt;repository knowledge&lt;/li&gt;
&lt;li&gt;browser access&lt;/li&gt;
&lt;li&gt;tests&lt;/li&gt;
&lt;li&gt;deterministic tools&lt;/li&gt;
&lt;li&gt;review loops&lt;/li&gt;
&lt;li&gt;security controls&lt;/li&gt;
&lt;li&gt;memory&lt;/li&gt;
&lt;li&gt;human approval boundaries&lt;/li&gt;
&lt;li&gt;orchestration&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;gstack is the engineering process around a coding agent, packaged as inspectable, version-controlled skills, and its value is in forcing that process onto the agent rather than in any single skill.&lt;/p&gt;

&lt;p&gt;Start from the trial sequence and keep only the skills that earn their place. Beyond it, the direction is composable layers -- specification, skills, deterministic verification, orchestration -- each doing what the others cannot.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/garrytan/gstack" rel="noopener noreferrer"&gt;gstack repository&lt;/a&gt; -- source, skill definitions, and setup&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/garrytan/gstack/blob/main/docs/skills.md" rel="noopener noreferrer"&gt;gstack skill deep dives&lt;/a&gt; -- philosophy, examples, and workflow for every skill&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://aside.com" rel="noopener noreferrer"&gt;Aside browser&lt;/a&gt; -- the browser gstack drives first on macOS&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/obra/superpowers" rel="noopener noreferrer"&gt;Superpowers repository&lt;/a&gt; -- source, skills, and plugin manifests&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/Fission-AI/OpenSpec" rel="noopener noreferrer"&gt;OpenSpec repository&lt;/a&gt; -- source, docs, and the CLI package&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.github.io/spec-kit/" rel="noopener noreferrer"&gt;GitHub Spec Kit documentation&lt;/a&gt; -- official Spec Kit workflow reference&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/bmad-code-org/BMAD-METHOD" rel="noopener noreferrer"&gt;BMAD-METHOD repository&lt;/a&gt; -- the adaptive AI-driven agile methodology&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/ruvnet/ruflo" rel="noopener noreferrer"&gt;Ruflo repository&lt;/a&gt; -- the multi-agent orchestration platform&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aicoding</category>
      <category>aiagents</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>Self-Hosted Deep Research Systems: 12 Tools Compared</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Fri, 25 Sep 2026 23:30:04 +0000</pubDate>
      <link>https://dev.to/rosgluk/self-hosted-deep-research-systems-12-tools-compared-7pp</link>
      <guid>https://dev.to/rosgluk/self-hosted-deep-research-systems-12-tools-compared-7pp</guid>
      <description>&lt;p&gt;Deep Research has become its own category of software, not just a model pointed at a search box. This article compares twelve self-hosted systems and the research architectures behind them.&lt;/p&gt;

&lt;p&gt;The line that actually matters is not whether a product ships a button labelled Deep Research, but what happens after the first round of retrieval. A genuine system notices that its plan was incomplete, chases a newly discovered lead, weighs conflicting sources, and only then writes the report.&lt;/p&gt;

&lt;p&gt;Below I compare twelve open-source and self-hosted projects that implement this loop in different ways: recursive research trees, planner-plus-subagent designs, evidence-gap loops, perspective-driven question generation, and model-driven agentic search. For each one I cover the architecture, local-LLM support, RAG or private-document access, deployment complexity, and the license you actually inherit if you self-host. Where a system is also a full product (Open WebUI, Vane), I keep the focus on how it researches and link out to the dedicated guide for installation and configuration. Deep Research is one of the more demanding applied workloads in the &lt;a href="https://www.glukhov.org/ai-systems/" rel="noopener noreferrer"&gt;AI Systems&lt;/a&gt; - it stresses retrieval, planning, and multi-step orchestration all at once, rather than any single layer in isolation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Deep Research?
&lt;/h2&gt;

&lt;p&gt;A conventional AI web-search workflow is mostly linear. Even when several searches are performed, the model usually just creates related queries, retrieves documents, and summarizes what it finds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Question
  |
Search
  |
Retrieve pages
  |
Summarize
  |
Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deep Research adds a layer: the research process itself becomes adaptive. After the first pass, the system can branch, re-check, and keep going until the evidence is sufficient.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Research question] --&amp;gt; B[Build research plan]
    B --&amp;gt; C1[Investigate topic A]
    B --&amp;gt; C2[Investigate topic B]
    B --&amp;gt; C3[Investigate topic C]

    C1 --&amp;gt; D1[Discover new question]
    C2 --&amp;gt; D2[Find conflicting evidence]
    C3 --&amp;gt; D3[Identify missing information]

    D1 --&amp;gt; E1[Research new question]
    D2 --&amp;gt; E2[Verify competing claims]
    D3 --&amp;gt; E3[Search for missing evidence]

    E1 --&amp;gt; F[Combine evidence]
    E2 --&amp;gt; F
    E3 --&amp;gt; F

    F --&amp;gt; G[Evaluate remaining gaps]
    G --&amp;gt;|More research needed| B
    G --&amp;gt;|Enough evidence| H[Generate cited report]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The distinction matters. A system that searches five times is not necessarily performing Deep Research; a stronger system starts with one question, discovers an unexpected implementation detail, opens a new research branch around it, compares primary and secondary sources, and revises its original assumptions. For the broader Search vs Deep Search vs Deep Research distinction and how cloud offerings frame the same idea, see &lt;a href="https://www.glukhov.org/rag/architecture/search-vs-deepsearch-vs-deep-research/" rel="noopener noreferrer"&gt;Search vs Deep Search vs Deep Research in 2026&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;There is no single Deep Research architecture. Current self-hosted implementations generally fall into five groups:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Recursive research trees.&lt;/li&gt;
&lt;li&gt;Planner and subagent architectures.&lt;/li&gt;
&lt;li&gt;Evidence-gap-driven research loops.&lt;/li&gt;
&lt;li&gt;Perspective and question-driven research.&lt;/li&gt;
&lt;li&gt;Agentic iterative search.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first four provide more explicit research structure. The fifth can still perform surprisingly deep investigation when paired with a strong reasoning and tool-calling model, but much of the strategy is delegated to the model itself. The evidence-gap loop in particular is a system-level cousin of the self-reflective retrieval used in Self-RAG-style pipelines — deciding whether to retrieve again, judging relevance, and critiquing the draft before answering. See &lt;a href="https://www.glukhov.org/rag/architecture/advanced-rag-variants-longrag-self-rag-graphrag/" rel="noopener noreferrer"&gt;Advanced RAG: LongRAG, Self-RAG and GraphRAG&lt;/a&gt; for that pattern at the retrieval-pipeline level.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-Hosted Deep Research Systems Compared
&lt;/h2&gt;

&lt;p&gt;The table below summarizes the major systems. "Recursive depth" does not mean that multiple web searches are possible; it means the system has some mechanism for deriving additional investigation from intermediate findings.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;Local LLM&lt;/th&gt;
&lt;th&gt;Web Research&lt;/th&gt;
&lt;th&gt;Private Docs / RAG&lt;/th&gt;
&lt;th&gt;Planning&lt;/th&gt;
&lt;th&gt;Recursive / Adaptive Depth&lt;/th&gt;
&lt;th&gt;UI&lt;/th&gt;
&lt;th&gt;Research Style&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT Researcher&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Web UI&lt;/td&gt;
&lt;td&gt;Recursive breadth/depth research tree&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unsloth Studio&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Very good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Planned evidence-driven research&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local Deep Research&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Web UI&lt;/td&gt;
&lt;td&gt;Multiple strategies plus autonomous agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;STORM / Co-STORM&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Custom corpus possible&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Very good&lt;/td&gt;
&lt;td&gt;Basic / demo UI&lt;/td&gt;
&lt;td&gt;Perspective and follow-up-question research&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeerFlow&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Planner plus subagents and long-horizon agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Onyx&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Multi-step enterprise Deep Research&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Deep Research&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Via tools / MCP&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;LangGraph oriented&lt;/td&gt;
&lt;td&gt;Planner plus parallel researchers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open WebUI&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Model driven&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Agentic iterative search and link following&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Khoj&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Personal knowledge plus autonomous research&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SurfSense&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Web/data research plus knowledge workspace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vane&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;File search&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Search-first answering engine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deep Research by lukeswade&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Research library&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Web UI&lt;/td&gt;
&lt;td&gt;Gap-driven iterative investigation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One point stands out: there is no direct relationship between UI sophistication and research depth. Open WebUI and Vane provide polished interfaces, while GPT Researcher and STORM are centered more on the research algorithm. Conversely, Onyx and Unsloth Studio try to provide both a strong user experience and a substantial research workflow. Most of these systems run against the same local inference backends covered in the &lt;a href="https://www.glukhov.org/llm-hosting/" rel="noopener noreferrer"&gt;LLM Hosting guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deep Research Architectures
&lt;/h2&gt;

&lt;p&gt;Before comparing individual products, it is useful to understand the architectural differences.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Style&lt;/th&gt;
&lt;th&gt;Representative Systems&lt;/th&gt;
&lt;th&gt;Main Idea&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Recursive research tree&lt;/td&gt;
&lt;td&gt;GPT Researcher&lt;/td&gt;
&lt;td&gt;Explicit breadth and depth generate new research branches&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Planner plus subagents&lt;/td&gt;
&lt;td&gt;DeerFlow, Open Deep Research&lt;/td&gt;
&lt;td&gt;Planner decomposes work and independent agents investigate pieces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evidence-gap driven&lt;/td&gt;
&lt;td&gt;Unsloth Studio, Local Deep Research, lukeswade/deep-research&lt;/td&gt;
&lt;td&gt;Findings are evaluated and missing evidence triggers another research round&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Perspective driven&lt;/td&gt;
&lt;td&gt;STORM / Co-STORM&lt;/td&gt;
&lt;td&gt;Research is expanded by generating perspectives and follow-up questions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-step research workflow&lt;/td&gt;
&lt;td&gt;Onyx&lt;/td&gt;
&lt;td&gt;Multiple research tasks gather and synthesize web and private knowledge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agentic iterative search&lt;/td&gt;
&lt;td&gt;Open WebUI&lt;/td&gt;
&lt;td&gt;Model decides when to search, read, verify, and search again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge-first research&lt;/td&gt;
&lt;td&gt;Khoj, SurfSense&lt;/td&gt;
&lt;td&gt;Research combines private information with external sources&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Search-first answering&lt;/td&gt;
&lt;td&gt;Vane&lt;/td&gt;
&lt;td&gt;Search and retrieval are optimized primarily for cited answers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The categories overlap. Local Deep Research offers several research strategies, and DeerFlow 2.0 is a general-purpose agent platform that can perform research rather than a research-only application. The distinction is nevertheless useful when choosing a system: a recursively branching researcher behaves differently from a chat interface whose model simply has a &lt;code&gt;search_web&lt;/code&gt; tool. Systems that retrieve private documents alongside the web lean on the same retrieval patterns described in the &lt;a href="https://www.glukhov.org/rag/" rel="noopener noreferrer"&gt;RAG cluster&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  License Comparison
&lt;/h2&gt;

&lt;p&gt;Licensing is particularly important if the system will become part of an internal platform, commercial service, or redistributed product.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;License&lt;/th&gt;
&lt;th&gt;Licensing Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT Researcher&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Current &lt;code&gt;pyproject.toml&lt;/code&gt; declares MIT; some older package metadata still reports Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unsloth Studio&lt;/td&gt;
&lt;td&gt;AGPL-3.0&lt;/td&gt;
&lt;td&gt;Studio UI is AGPL-3.0; core Unsloth remains Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local Deep Research&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Permissive open-source license&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;STORM / Co-STORM&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Permissive open-source license&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeerFlow&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Applies to current DeerFlow 2.0 repository&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Onyx&lt;/td&gt;
&lt;td&gt;MIT plus Enterprise License&lt;/td&gt;
&lt;td&gt;Core is MIT; &lt;code&gt;ee&lt;/code&gt; directories use the Onyx Enterprise License; &lt;code&gt;onyx-foss&lt;/code&gt; is 100 percent MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Deep Research&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Repository was archived in August 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open WebUI&lt;/td&gt;
&lt;td&gt;Open WebUI License&lt;/td&gt;
&lt;td&gt;Current versions include branding restrictions; older code has MIT/BSD history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Khoj&lt;/td&gt;
&lt;td&gt;AGPL-3.0-or-later&lt;/td&gt;
&lt;td&gt;Network copyleft should be considered for modified hosted deployments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SurfSense&lt;/td&gt;
&lt;td&gt;Apache-2.0&lt;/td&gt;
&lt;td&gt;Current repository declares Apache-2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vane&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Formerly known as Perplexica&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deep Research by lukeswade&lt;/td&gt;
&lt;td&gt;MIT&lt;/td&gt;
&lt;td&gt;Permissive open-source license&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For private self-hosting, none of these licenses prevents normal use. The differences matter when modifying the software, offering it to other users, embedding it into another commercial application, or redistributing derivatives. MIT and Apache-2.0 are generally the simplest options for integration. AGPL-3.0 deserves closer review for network-accessible modified deployments, and Open WebUI's current license adds its own branding conditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  GPT Researcher
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/assafelovic/gpt-researcher" rel="noopener noreferrer"&gt;GPT Researcher&lt;/a&gt; is developed by Assaf Elovic and contributors as an autonomous research agent focused specifically on comprehensive online investigation. It is one of the clearest reference implementations of what "Deep Research" means when the term describes an algorithm rather than a user-interface feature.&lt;/p&gt;

&lt;p&gt;Its strongest feature is explicit breadth and depth. Deep Research mode exposes parameters such as &lt;code&gt;deep_research_breadth&lt;/code&gt;, &lt;code&gt;deep_research_depth&lt;/code&gt;, and concurrency, allowing one investigation to generate several branches and those branches to generate additional research. This creates a real research tree rather than a fixed collection of search queries.&lt;/p&gt;

&lt;p&gt;That approach also has costs. Recursive expansion can produce many retrieval and LLM operations, and the quality of the final result depends heavily on the model's ability to formulate useful research questions, extract evidence, and avoid propagating weak assumptions into deeper levels. GPT Researcher is also more research-engine oriented than applications such as Open WebUI or Unsloth Studio.&lt;/p&gt;

&lt;p&gt;Installation is moderate rather than trivial: the project uses Python and ships a web application, while useful deployments also require suitable model and search providers. Current project metadata declares the MIT license. Choose GPT Researcher when explicit research depth, configurable recursion, and a research-first architecture matter more than an all-purpose local AI workstation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unsloth Studio
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/unslothai/unsloth" rel="noopener noreferrer"&gt;Unsloth Studio&lt;/a&gt; is developed by the Unsloth team as part of the broader Unsloth ecosystem. Originally best known for efficient model fine-tuning, Unsloth has expanded Studio into a local AI environment for inference, chat, tools, RAG, model management, and now Deep Research.&lt;/p&gt;

&lt;p&gt;The interesting aspect of Studio is how tightly research is integrated with local model operation. Its Deep Research workflow includes a planning stage, plan review, evidence gathering, report generation, document handling, and failure handling when research steps fail to gather evidence. For users already running GGUF or other local models, this makes Studio considerably more convenient than assembling a separate research framework, inference server, and frontend.&lt;/p&gt;

&lt;p&gt;Studio does not expose the same simple breadth/depth research-tree abstraction as GPT Researcher. Much of the workflow is organized around a research plan and evidence collection rather than arbitrary recursive expansion, and the feature is newer than some dedicated research projects. The quality of local research also remains sensitive to context length, output limits, tool use, and the reasoning quality of the selected model.&lt;/p&gt;

&lt;p&gt;Installation is relatively friendly because Unsloth now provides Studio and desktop-oriented workflows across major platforms, although GPU and model configuration can still become substantial for advanced local deployments. The Studio component is AGPL-3.0, while the core Unsloth package remains Apache-2.0. Choose Unsloth Studio when Deep Research should be part of a broader local-model workstation rather than a standalone research service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local Deep Research
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/LearningCircuit/local-deep-research" rel="noopener noreferrer"&gt;Local Deep Research&lt;/a&gt; is maintained by LearningCircuit and contributors as a privacy-oriented open-source research assistant. Its explicit goal is systematic research using web sources, academic databases, private documents, and local language models.&lt;/p&gt;

&lt;p&gt;Its major advantage is architectural flexibility. Rather than enforcing one research algorithm, Local Deep Research supports pipeline-oriented strategies as well as a LangGraph agent strategy in which the model can decide what to search, which specialist sources to use, and when enough evidence has been collected. Academic sources such as arXiv, PubMed, Semantic Scholar, and other search mechanisms make it particularly attractive for technical and scientific research.&lt;/p&gt;

&lt;p&gt;The downside of flexibility is complexity. Different strategies can behave substantially differently, which makes results harder to characterize with one simple "depth" parameter. It also has more moving pieces than a conventional chat UI, and users looking only for fast AI-assisted web answers may find it unnecessarily elaborate.&lt;/p&gt;

&lt;p&gt;The project supports local operation and has developed a substantial application around the underlying research engine. Its license is MIT. Choose Local Deep Research when privacy, local inference, multiple research strategies, academic information sources, and control over the research process are more important than a minimal setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  STORM and Co-STORM
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/stanford-oval/storm" rel="noopener noreferrer"&gt;STORM&lt;/a&gt;, developed by Stanford OVAL, stands for Synthesis of Topic Outlines through Retrieval and Multi-perspective Question Asking. It was designed around knowledge curation and long-form report generation rather than general AI chat.&lt;/p&gt;

&lt;p&gt;STORM's distinctive technique is perspective-driven research. It tries to discover different perspectives on a subject and uses question asking to broaden the information collected before writing an outline and article. Co-STORM extends the concept toward collaborative human-AI knowledge curation. This can uncover dimensions of a topic that a conventional list of keyword searches may overlook.&lt;/p&gt;

&lt;p&gt;STORM is less suitable as a general local AI frontend. Its workflow is strongly oriented toward researching and writing structured, Wikipedia-like articles, and the project's own documentation notes that generated output should not automatically be considered publication-ready. It is therefore better understood as a specialized research and knowledge-curation engine than a replacement for Open WebUI.&lt;/p&gt;

&lt;p&gt;Installation is Python oriented, and the pipeline can be customized with different models and retrievers. STORM uses the MIT license. Choose it when the objective is broad topic exploration, perspective discovery, structured outlines, and long-form knowledge synthesis.&lt;/p&gt;

&lt;h2&gt;
  
  
  DeerFlow
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/bytedance/deer-flow" rel="noopener noreferrer"&gt;DeerFlow&lt;/a&gt; is developed by ByteDance and the DeerFlow community. Its name originally stood for Deep Exploration and Efficient Research Flow, but an important distinction now exists between the original 1.x Deep Research framework and DeerFlow 2.0.&lt;/p&gt;

&lt;p&gt;DeerFlow 1.x was specifically designed around Deep Research. DeerFlow 2.0 is a ground-up rewrite into a more general SuperAgent harness capable of orchestrating subagents, memory, sandboxes, tools, and skills. For research, this architecture is powerful because a coordinator can delegate different parts of a problem to separate agents and later synthesize their findings - the same planner-plus-subagent decomposition covered in &lt;a href="https://www.glukhov.org/ai-systems/architecture/multi-agent-orchestration-patterns/" rel="noopener noreferrer"&gt;Multi-Agent Orchestration Patterns&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The trade-off is that DeerFlow 2.0 is no longer a narrowly optimized research engine. It is closer to a general long-horizon agent platform in which research is one workload among coding, artifact generation, and other tasks. If the requirement is a small dedicated research service, this additional machinery can be unnecessary.&lt;/p&gt;

&lt;p&gt;Deployment is consequently more involved than a simple search UI, although the architecture provides much more room for customization and extension. DeerFlow is MIT licensed. Choose DeerFlow when Deep Research is expected to become one capability inside a broader multi-agent automation environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Onyx
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/onyx-dot-app/onyx" rel="noopener noreferrer"&gt;Onyx&lt;/a&gt;, originally known as Danswer, is developed by DanswerAI and positioned as a self-hostable AI application layer for organizations. It combines chat, agents, web search, RAG, MCP integration, many enterprise data connectors, and a dedicated Deep Research capability.&lt;/p&gt;

&lt;p&gt;Onyx stands out because research can span both the public web and a substantial private knowledge environment. Its Deep Research implementation is a real multi-step research flow rather than merely a search-result summarizer, and the project has published results and execution logs for DeepResearch Bench. For organizations that need research across internal documentation, indexed applications, and external sources, this is a particularly strong combination.&lt;/p&gt;

&lt;p&gt;The cost of those capabilities is infrastructure complexity. Onyx is a larger platform than GPT Researcher or a lightweight local research project, and many of its strengths only matter when connectors, indexing, authentication, document stores, and organizational data are actually being used.&lt;/p&gt;

&lt;p&gt;Onyx can be deployed self-hosted, including in restricted environments. Most of the main repository is MIT licensed, while code under &lt;code&gt;ee&lt;/code&gt; directories uses the Onyx Enterprise License; a separate &lt;code&gt;onyx-foss&lt;/code&gt; repository is maintained as a fully MIT-licensed variant. Choose Onyx when Deep Research must coexist with serious enterprise RAG and organizational knowledge retrieval.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open Deep Research
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/langchain-ai/open_deep_research" rel="noopener noreferrer"&gt;Open Deep Research&lt;/a&gt; was developed by LangChain as an open implementation of a configurable Deep Research agent. It combines planning, research, report generation, multiple model providers, search tools, and MCP integrations using LangGraph.&lt;/p&gt;

&lt;p&gt;Its architecture is particularly interesting for developers. Research work can be decomposed and parallelized, making it a useful reference for planner-researcher-synthesizer designs. Because the project was built around LangGraph rather than a monolithic UI, it is also easier to study as an implementation pattern for building custom research agents.&lt;/p&gt;

&lt;p&gt;There is one major problem for new deployments: LangChain archived the repository on August 21, 2026, and it is now read-only. The code remains useful, but starting a production system around an archived reference project creates an obvious maintenance risk.&lt;/p&gt;

&lt;p&gt;The project is Python based and uses an MIT license. Choose it today mainly for architectural study, experimentation, or as a source of implementation ideas rather than as the default foundation for a new long-lived installation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open WebUI
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/open-webui/open-webui" rel="noopener noreferrer"&gt;Open WebUI&lt;/a&gt; is one of the most popular general-purpose self-hosted interfaces for local and remote LLMs. Its recent agentic tool architecture gives models access to web search, URL fetching, knowledge bases, files, memory, code execution, and other tools. For installation, RAG setup, and the full feature set, see the &lt;a href="https://www.glukhov.org/llm-hosting/llm-frontends/open-webui-overview-quickstart-and-alternatives/" rel="noopener noreferrer"&gt;Open WebUI guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Open WebUI's research model is interesting because the research loop is largely controlled by the language model. In native agentic mode the model can search, inspect snippets, fetch full pages, identify missing information, follow newly discovered URLs, cross-check sources, and repeat the process before generating an answer. With a capable reasoning and tool-calling model, this can produce genuine investigative behavior without a dedicated fixed research tree.&lt;/p&gt;

&lt;p&gt;The limitation is precisely that this structure is model driven. Open WebUI does not provide the same explicit breadth/depth research topology as GPT Researcher, and there is less deterministic control over how many independent branches will be explored. A weak tool-calling model can stop too early, search poorly, or fail to follow important leads.&lt;/p&gt;

&lt;p&gt;Installation is among the easiest in this comparison, particularly for users who already run Ollama, llama.cpp, vLLM, or another OpenAI-compatible inference server. Current releases use the Open WebUI License, which retains substantial permissive characteristics but adds branding restrictions; earlier portions of the project have MIT and BSD-3-Clause history. Choose Open WebUI when you want excellent local-LLM integration and a general AI interface in which research is one of many agentic capabilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  Khoj
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/khoj-ai/khoj" rel="noopener noreferrer"&gt;Khoj&lt;/a&gt; is developed as a self-hostable personal AI and "second brain." It combines local or cloud language models with web retrieval, personal documents, semantic search, custom agents, automations, and an experimental &lt;code&gt;/research&lt;/code&gt; mode.&lt;/p&gt;

&lt;p&gt;Its strongest use case is research that crosses the boundary between public information and a user's existing knowledge base. A question can be investigated in the context of PDFs, Markdown files, notes, office documents, or connected information rather than treating every task as web research from scratch. This makes Khoj useful for ongoing personal or team knowledge work.&lt;/p&gt;

&lt;p&gt;Khoj is not primarily designed around a visible recursive research tree. Its research functionality is better understood as autonomous investigation within a larger personal knowledge system. Users seeking explicit breadth/depth controls or a dedicated research-engine API may prefer GPT Researcher or Local Deep Research.&lt;/p&gt;

&lt;p&gt;Self-hosting is supported and the system can work with local models including the Llama, Qwen, Gemma, and Mistral families. Khoj is licensed under AGPL-3.0-or-later. Choose it when Deep Research should be closely integrated with a long-lived personal knowledge base rather than treated as an isolated web research job.&lt;/p&gt;

&lt;h2&gt;
  
  
  SurfSense
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/MODSetter/SurfSense" rel="noopener noreferrer"&gt;SurfSense&lt;/a&gt; is an open-source research workspace that evolved from a NotebookLM-style knowledge system toward an agent-oriented open-web research platform. It combines a searchable knowledge base with web and platform-specific data connectors, reports, automations, MCP access, and local model support.&lt;/p&gt;

&lt;p&gt;SurfSense's distinctive advantage is its data surface. It is designed to give agents structured access not only to ordinary web pages and search results but also to sources such as Reddit, YouTube, Google Maps, and other live information services. Research findings can then remain in the same environment as uploaded documents and previously collected knowledge.&lt;/p&gt;

&lt;p&gt;It is less purely focused on the research algorithm than GPT Researcher or STORM. A significant part of SurfSense's value comes from retrieval infrastructure, connectors, knowledge management, and downstream artifacts rather than an explicit recursively expanding research graph.&lt;/p&gt;

&lt;p&gt;Self-hosting is supported through Docker-oriented installation, and local models can be connected through common local inference interfaces. The current repository is Apache-2.0 licensed. Choose SurfSense when the difficult part of research is obtaining, structuring, retaining, and reusing information from many different data sources.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vane, Formerly Perplexica
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/ItzCrazyKns/Vane" rel="noopener noreferrer"&gt;Vane&lt;/a&gt;, formerly known as Perplexica, is an open-source AI answering engine designed as a self-hosted alternative to search-first products such as Perplexity. It combines an AI chat interface, search backend, citations, local model support, and semantic search over uploaded files. For the Docker quickstart, &lt;code&gt;SEARXNG_API_URL&lt;/code&gt; wiring, and Ollama/llama.cpp setup, see &lt;a href="https://www.glukhov.org/llm-hosting/llm-frontends/vane-perplexica-2/" rel="noopener noreferrer"&gt;Vane (Perplexica 2.0) Quickstart With Ollama and llama.cpp&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Vane is good at the search-and-answer experience. It classifies questions, conducts web research, retrieves useful information, and generates cited responses through a polished interface. Most deployments back it with &lt;a href="https://www.glukhov.org/data-infrastructure/search/selfhosting-searxng/" rel="noopener noreferrer"&gt;SearXNG&lt;/a&gt; as the search layer. For users who mainly want a private AI search engine backed by SearXNG and local models, it provides a much more focused experience than a large general agent platform.&lt;/p&gt;

&lt;p&gt;Its limitation in this comparison is research depth. Although the system can run research operations, its architecture is still primarily that of an answering engine rather than a recursively branching Deep Research framework. It should therefore not be treated as equivalent to GPT Researcher merely because both can perform multiple searches before responding.&lt;/p&gt;

&lt;p&gt;Vane is relatively straightforward to deploy with Docker and supports common model providers and local inference systems. It is MIT licensed. Choose Vane when the primary requirement is high-quality self-hosted AI search with citations rather than long-running autonomous investigation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deep Research by lukeswade
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/lukeswade/deep-research" rel="noopener noreferrer"&gt;lukeswade/deep-research&lt;/a&gt; project is a smaller self-hosted research system, but it implements one of the more interesting workflows in this comparison. It supports both cloud models and local OpenAI-compatible endpoints including llama.cpp, LM Studio, Ollama, vLLM, and MLX - for the server side of that endpoint, see &lt;a href="https://www.glukhov.org/llm-hosting/llama-cpp/" rel="noopener noreferrer"&gt;llama.cpp Quickstart with CLI and Server&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Its research loop is explicitly gap driven. A run decomposes the question into targeted searches, reads relevant pages, produces per-source notes containing evidence, and analyzes what remains unknown before deciding what should be searched next. Higher depth settings allow several rounds and progressively larger source budgets, while saturation detection can stop research early when new searches cease producing useful information.&lt;/p&gt;

&lt;p&gt;It does not have the ecosystem, organizational connectors, or general AI workstation capabilities of Onyx, Open WebUI, or Unsloth Studio. It is much more narrowly focused on doing one thing: researching a question deeply and storing the resulting research in a searchable local library.&lt;/p&gt;

&lt;p&gt;That focus also makes deployment relatively understandable. The system has a web UI and is particularly friendly to local OpenAI-compatible model servers; the documentation recommends capable models and supports using a smaller fast model for high-volume note processing. It is MIT licensed. Choose it when local inference, transparent evidence collection, and information-gap-driven research matter more than a large surrounding platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Deep Research System Should You Choose?
&lt;/h2&gt;

&lt;p&gt;There is no single winner, because these systems solve somewhat different problems.&lt;/p&gt;

&lt;p&gt;For an explicit research algorithm with understandable depth controls, &lt;strong&gt;GPT Researcher&lt;/strong&gt; remains one of the clearest starting points. Its breadth/depth model makes it easy to reason about why research expands and how expensive a run may become.&lt;/p&gt;

&lt;p&gt;For a strongly local workflow, &lt;strong&gt;Local Deep Research&lt;/strong&gt; and &lt;strong&gt;Unsloth Studio&lt;/strong&gt; are particularly attractive. Local Deep Research provides more research-strategy flexibility, while Unsloth Studio integrates research with model management, inference, RAG, and the broader local-model workflow.&lt;/p&gt;

&lt;p&gt;For long-form knowledge curation, &lt;strong&gt;STORM&lt;/strong&gt; remains unusually interesting because its multi-perspective question-generation technique attacks a problem that many research systems ignore: discovering the questions that the original user did not know to ask.&lt;/p&gt;

&lt;p&gt;For multi-agent systems, &lt;strong&gt;DeerFlow&lt;/strong&gt; represents a different direction. Instead of building a specialized research loop, it treats research as a long-running agent task that can be delegated among subagents and combined with tools, memory, code execution, and other capabilities.&lt;/p&gt;

&lt;p&gt;For organizations, &lt;strong&gt;Onyx&lt;/strong&gt; has one of the strongest combinations of Deep Research, RAG, web investigation, and enterprise knowledge connectors. Its heavier infrastructure is justified when internal information sources matter as much as the public web.&lt;/p&gt;

&lt;p&gt;For an existing local AI installation, &lt;strong&gt;Open WebUI&lt;/strong&gt; may be all that is necessary. A sufficiently capable local model with native tool calling can repeatedly search, read pages, follow links, verify information, and fill gaps without installing a separate research engine.&lt;/p&gt;

&lt;p&gt;Finally, &lt;strong&gt;lukeswade/deep-research&lt;/strong&gt; is worth watching precisely because it is smaller. Its gap-driven workflow is conceptually clean, supports llama.cpp directly through an OpenAI-compatible endpoint, and separates expensive planning and synthesis from high-volume per-source processing.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Tell Real Deep Research from Repeated Search
&lt;/h2&gt;

&lt;p&gt;The most useful way to evaluate these systems is not to ask whether they have a button labelled "Deep Research." Instead, inspect what happens after the first round of information is collected.&lt;/p&gt;

&lt;p&gt;A genuine research system should be able to discover that its original plan was incomplete. It should recognize a contradiction, missing source, unexpected implementation detail, or newly relevant subtopic and change its subsequent investigation accordingly. That is the line separating sophisticated retrieval from research.&lt;/p&gt;

&lt;p&gt;For self-hosting, the ecosystem is now broad enough that the choice is no longer simply between a cloud Deep Research service and a home-grown script. There are dedicated recursive researchers, academic knowledge-curation systems, enterprise research platforms, local-model workstations, general agent frameworks, and lightweight gap-driven tools.&lt;/p&gt;

&lt;p&gt;The right choice therefore depends less on which project advertises the most features and more on the research architecture you want to operate: explicit recursion, multi-agent delegation, evidence-gap analysis, perspective discovery, or model-driven autonomous search.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/assafelovic/gpt-researcher" rel="noopener noreferrer"&gt;GPT Researcher&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/unslothai/unsloth" rel="noopener noreferrer"&gt;Unsloth Studio&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/LearningCircuit/local-deep-research" rel="noopener noreferrer"&gt;Local Deep Research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/stanford-oval/storm" rel="noopener noreferrer"&gt;STORM / Co-STORM&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/bytedance/deer-flow" rel="noopener noreferrer"&gt;DeerFlow&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/onyx-dot-app/onyx" rel="noopener noreferrer"&gt;Onyx&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/langchain-ai/open_deep_research" rel="noopener noreferrer"&gt;Open Deep Research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/open-webui/open-webui" rel="noopener noreferrer"&gt;Open WebUI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/khoj-ai/khoj" rel="noopener noreferrer"&gt;Khoj&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/MODSetter/SurfSense" rel="noopener noreferrer"&gt;SurfSense&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ItzCrazyKns/Vane" rel="noopener noreferrer"&gt;Vane (formerly Perplexica)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/lukeswade/deep-research" rel="noopener noreferrer"&gt;lukeswade/deep-research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.glukhov.org/rag/architecture/search-vs-deepsearch-vs-deep-research/" rel="noopener noreferrer"&gt;Search vs Deep Search vs Deep Research in 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.glukhov.org/ai-systems/architecture/multi-agent-orchestration-patterns/" rel="noopener noreferrer"&gt;Multi-Agent Orchestration Patterns&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>deepresearch</category>
      <category>aiagents</category>
      <category>rag</category>
      <category>llm</category>
    </item>
    <item>
      <title>The Efficient Frontier of Open Models: Finding the Sweet Spot in 2026</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Thu, 24 Sep 2026 10:19:42 +0000</pubDate>
      <link>https://dev.to/rosgluk/the-efficient-frontier-of-open-models-finding-the-sweet-spot-in-2026-1lj2</link>
      <guid>https://dev.to/rosgluk/the-efficient-frontier-of-open-models-finding-the-sweet-spot-in-2026-1lj2</guid>
      <description>&lt;p&gt;In September 2026 the open-model efficient frontier sits between 25B and 34B parameters: near-frontier agentic work on one 24 GB GPU, at a fraction of 70B-class hardware.&lt;/p&gt;

&lt;p&gt;That band shows up in deployments, not in a parameter-count slogan. Public leaderboards, single-GPU runs, and monthly API bills converge on models that stay close to frontier tool-use while the weights and the KV cache still fit hardware you can rent or own.&lt;/p&gt;

&lt;p&gt;This page is the decision hub for that choice, inside the &lt;a href="https://www.glukhov.org/llm-performance/" rel="noopener noreferrer"&gt;LLM performance&lt;/a&gt; cluster. It walks from small instruct models to 70B-class dense weights, then stops on the two architectures that define the frontier right now: Qwen3.8-27B, which is comfortable on one 24 GB card, and Llama 4 Scout, whose compute looks modest only because most of its experts stay idle. Token-per-second tables stay on the benchmark pages linked from each section.&lt;/p&gt;

&lt;p&gt;By the end you should know which size to test first, how much VRAM the weights and the cache actually take, how a rented RTX 4090 compares with current Claude Sonnet token prices, and when a larger model is still the right call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why public benchmarks overstate production transfer
&lt;/h2&gt;

&lt;p&gt;A model can sit within a couple of points of a frontier API model on MMLU or HumanEval and still fail the job you are shipping. Multi-tool flows drop the second tool call, refusal calibration drifts toward over-refusal on medical-adjacent queries, and P99 latency under burst load doubles. The leaderboard gap looked small because those benches measure capability on a fixed prompt set. They do not measure transfer onto your tools, your documents, and your latency budget.&lt;/p&gt;

&lt;p&gt;The efficient frontier is the set of models that hold up on your workload at a cost you can keep paying. That takes a paired comparison on production-like tasks, with three signals recorded the whole way: tool-call success, refusal behaviour, and P99 latency. &lt;a href="https://futureagi.com/blog/evaluating-cheap-frontier-models-2026" rel="noopener noreferrer"&gt;Future AGI's write-up on cheap-frontier substitution&lt;/a&gt; makes the same point from the evaluation side: a small public-benchmark gap is not a licence to swap models.&lt;/p&gt;

&lt;p&gt;Treat the scores below as a map of where to start the bake-off. They are dated to September 2026, and several headline coding numbers are vendor-run rather than independent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where 8B, 30B, and 70B actually sit
&lt;/h2&gt;

&lt;p&gt;At the low end, an 8B-class instruct model will follow a short, explicit instruction and write a small function. Ask it to debug a failing CI pipeline across several services, recover when the first tool call returns an error, and keep the plan intact, and the run falls apart. That is a workload failure, not a leaderboard failure.&lt;/p&gt;

&lt;p&gt;The middle of the range is where general-purpose agentic work currently clears a single prosumer GPU. Qwen3.8-27B is the reference dense model. Architectures in the same size band, including the Qwen 30B-class MoE compared in the &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/qwen3-30b-vs-gpt-oss-20b/" rel="noopener noreferrer"&gt;Qwen3 30B vs GPT-OSS 20B&lt;/a&gt; head-to-head, are the ones to test before you spend on anything larger.&lt;/p&gt;

&lt;p&gt;At the high end, a dense 70B model at 4-bit is on the order of 35–40 GB of weights before any KV cache. It does not fit a 24 GB card. You are into a 48 GB GPU, an 80 GB data-center card, or a two-GPU split, and the extra quality over a well-tuned 27B model is often marginal on coding agents, document pipelines, and support bots. Closed frontier APIs such as Claude Opus sit in a different budget entirely: you pay per token, and you do not carry the weights at all.&lt;/p&gt;

&lt;p&gt;Before you pick a checkpoint, confirm the card and the free memory. Model weights are only part of the footprint. Context (the KV cache) sits on top.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# GPU name and free VRAM before picking a quant&lt;/span&gt;
nvidia-smi &lt;span class="nt"&gt;--query-gpu&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;name,memory.total,memory.used &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;csv

&lt;span class="c"&gt;# after loading: confirm the process stayed on the GPU&lt;/span&gt;
nvidia-smi
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If layers spill to CPU, generation throughput collapses. You can see the split in &lt;code&gt;memory.used&lt;/code&gt; on the GPU and in the runtime's offload log. For what dense and MoE checkpoints actually deliver on a 16 GB card, the measurements live in the &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/best-llm-on-16gb-vram-gpu/" rel="noopener noreferrer"&gt;16 GB VRAM llama.cpp benchmarks&lt;/a&gt; and the &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/choosing-best-llm-for-ollama-on-16gb-vram-gpu/" rel="noopener noreferrer"&gt;Ollama comparison on a 16 GB RTX 4080&lt;/a&gt;. Those pages own the speed tables. This one only needs the fit decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen3.8-27B: the 24 GB reference point
&lt;/h2&gt;

&lt;p&gt;Alibaba's Qwen3.8-27B is a dense model, not a mixture of experts. On Artificial Analysis it scores about 51 on the Agentic Index (the rounded figure; the underlying value is 50.9) and about 52 on the Intelligence Index at maximum reasoning effort. Those are different measurements. The Agentic Index is tool-use and multi-step tasks. The Intelligence Index is a broader capability blend. &lt;a href="https://www.mindstudio.ai/blog/qwen-3-27b-local-benchmark" rel="noopener noreferrer"&gt;MindStudio's hands-on bench&lt;/a&gt; frames the agentic score as just behind Kimi K2, a vastly larger mixture-of-experts model. &lt;a href="https://www.qubrid.com/blog/qwen38-27b-benchmarks-official-and-independent-results" rel="noopener noreferrer"&gt;Qubrid's breakdown&lt;/a&gt; is the one to read next to it: the margin over the next model is under a point, and Qwen's own SWE-bench Pro figure of 61.7 has not been reproduced as an independent run. The interesting result is the shape of the scores. A 27B model converts a compact budget into planning and tool use unusually well, and it does not lead a broad knowledge suite.&lt;/p&gt;

&lt;p&gt;The architecture is why the long context is affordable. Of 64 layers, only 16 use full attention. The other 48 are Gated DeltaNet layers, a linear-attention design from the &lt;a href="https://jankautz.com/publications/GatedDeltaNet_ICLR25.pdf" rel="noopener noreferrer"&gt;Gated Delta Networks&lt;/a&gt; line of work, and they keep a fixed recurrent state instead of a per-token key/value history. At FP16 the full-attention layers cost about 64 KB of KV cache per token. A conventional stack with the same layer count would cache roughly four times that. &lt;a href="https://locallyuncensored.com/blog/how-to-run-qwen-3-8-27b-locally.html" rel="noopener noreferrer"&gt;Locally Uncensored's quant table&lt;/a&gt; puts Q4_K_M weights at 17.1 GB and IQ4_XS at 15.7 GB. &lt;a href="https://www.hardware-corner.net/qwen3-8-27b-hardware-tests/" rel="noopener noreferrer"&gt;Hardware Corner's llama.cpp measurements&lt;/a&gt; on a 24 GB GPU landed near 18 GB at short context and about 22 GB at 64K, which is the practical target on an RTX 4090. A 16 GB card can hold an IQ4-class quant of the weights and almost no context. If that is your machine, stay with the Qwen 3.5 and 3.6 runs already measured on 16 GB, including &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/comparing-qwen-3-6-mtp-vs-standard/" rel="noopener noreferrer"&gt;Qwen 3.6 27B MTP versus standard decoding&lt;/a&gt;. How to budget the cache itself is covered in &lt;a href="https://www.glukhov.org/llm-performance/optimization/kv-cache-16gb-long-context/" rel="noopener noreferrer"&gt;KV cache on 16 GB GPUs&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Hands-on reports, MindStudio's among them, also find the model usable for coding, image analysis, and agent loops, which is the part a single index cannot show. Vision adds a small projector file on top of the text weights. Plan for it if the workload includes screenshots.&lt;/p&gt;

&lt;h2&gt;
  
  
  Llama 4 Scout: active parameters are not VRAM
&lt;/h2&gt;

&lt;p&gt;Meta's Llama 4 Scout is a different kind of efficiency. The &lt;a href="https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md" rel="noopener noreferrer"&gt;Llama 4 model card&lt;/a&gt; lists 17 billion active parameters and 109 billion total, with 16 experts, native image input, and a 10 million token context window. &lt;a href="https://ai.meta.com/blog/llama-4-multimodal-intelligence" rel="noopener noreferrer"&gt;Meta's launch post&lt;/a&gt; says the INT4 checkpoint fits on a single H100. Decode compute tracks the 17B that fire for each token. Memory does not. Every expert still has to be resident, so the VRAM bill looks like a 100B-class model even though the FLOPs look like a 17B model.&lt;/p&gt;

&lt;p&gt;That is the correction worth making before you budget a machine. Scout does not "run like a 3B model." Active parameters set speed. Resident parameters set whether the card can hold the weights. On a 24 GB GPU, Qwen3.8-27B is the model that fits. Scout is the model that makes sense once you already have an 80 GB card and want 17B-class compute with a much larger expert pool and a very long context.&lt;/p&gt;

&lt;p&gt;A June 2025 paper, &lt;a href="https://arxiv.org/html/2506.12119v1" rel="noopener noreferrer"&gt;Can Mixture-of-Experts Surpass Dense LLMs Under Strictly Equal Resources?&lt;/a&gt;, argues that MoE can beat a dense model when the comparison holds total compute and data fixed and the activation rate sits in a workable band. The practical reading is narrower than the headline. MoE wins on cost only after you have paid for the memory that stores the idle experts. If you are already serving a model of this size and want more tokens per second without changing the weights, &lt;a href="https://www.glukhov.org/llm-performance/optimization/speculative-decoding/" rel="noopener noreferrer"&gt;speculative decoding&lt;/a&gt; is the neighbouring technique: a draft path proposes tokens and the main model verifies them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Phi-4: the small-model exception for math
&lt;/h2&gt;

&lt;p&gt;Microsoft's Phi-4 is a 14B dense model. In the December 2024 &lt;a href="https://arxiv.org/html/2412.08905v1" rel="noopener noreferrer"&gt;technical report&lt;/a&gt;, simple-evals scores it at 80.4 on MATH, ahead of the GPT-4o figure in the same table, with a context window that the report extends to 16K. A Q4_K_M quant is commonly run in about 8 GB. That is real efficiency, and it is narrow. Phi-4 was trained for STEM question-answering. It is the model to try when the job is math or short technical reasoning and the machine is small. It is not the 2026 default for multi-tool agents, long documents, or vision. The frontier for that work is still the 25–34B band above, or Scout when the GPU is an H100-class card.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-hosted GPU hours versus per-token API prices
&lt;/h2&gt;

&lt;p&gt;Prices move. These are mid-September 2026 figures, and they will not survive unchanged into 2027.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://gpufinder.dev/gpu/rtx-4090" rel="noopener noreferrer"&gt;GPU Finder&lt;/a&gt;, checked on 18 September 2026, showed an in-stock RTX 4090 at about $0.26 per GPU-hour on Vast, with RunPod community cloud listed at $0.34. The six-month on-demand floor for a single 4090 sat between $0.11 and $0.60. At $0.34 and 730 hours in a month, a card that stays up all month is about $248 before tax, storage, and egress. At the $0.26 in-stock quote the same month is about $190. An earlier planning figure of $0.53 per hour is inside that six-month range, but it overstates what marketplace stock was actually asking in the second half of September.&lt;/p&gt;

&lt;p&gt;On the API side, Anthropic's &lt;a href="https://platform.claude.com/docs/en/models/sonnet-5/whats-new-sonnet-5" rel="noopener noreferrer"&gt;Sonnet 5 pricing notes&lt;/a&gt; list $2 per million input tokens and $10 per million output tokens. Sonnet 4.6 remains at $3 and $15. &lt;a href="https://benchlm.ai/anthropic/api-pricing" rel="noopener noreferrer"&gt;BenchLM's rate card&lt;/a&gt;, updated 3 September 2026, shows the same split. A workload of 100 million input tokens and 10 million output tokens is $300 on Sonnet 5 and $450 on Sonnet 4.6.&lt;/p&gt;

&lt;p&gt;Put next to a 4090 rented all month at $248, that example does not say "self-hosting is always cheaper." It says the two bills meet in the same neighbourhood for this volume. Sonnet 5 at a 10:1 input-to-output ratio costs $30 per 10 million input plus 1 million output. The $248 rental covers about eight of those blocks: roughly 80 million input tokens and 8 million output tokens. Above that, the rented card pulls ahead, provided you can keep it busy. Below it, the API is the cheaper line, and you are not paying for an idle GPU. Owned hardware replaces the rental with electricity and depreciation. The strategies around that bill — budgeting, caching, and fallbacks — are in &lt;a href="https://www.glukhov.org/llm-architecture/cost-optimization/cost-optimization-for-llm-systems/" rel="noopener noreferrer"&gt;cost optimization for LLM systems&lt;/a&gt;. The reason a lower token price can still be the wrong architecture is &lt;a href="https://www.glukhov.org/llm-hosting/self-hosting/data-gravity-ai-vendor-lockin/" rel="noopener noreferrer"&gt;data gravity and API lock-in&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a larger model is still the right call
&lt;/h2&gt;

&lt;p&gt;The 25–34B band is the default for production workloads that repeat: coding agents, support bots, document processing, extraction, and automation. It is the wrong default in a few situations.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Frontier reasoning on problems the smaller model has already failed in a paired test.&lt;/li&gt;
&lt;li&gt;Prose where the difference in voice is the product, and volume is low enough that token price is noise.&lt;/li&gt;
&lt;li&gt;Work that needs broad coverage across several specialist domains at once, where the 27B model keeps missing the same class of fact.&lt;/li&gt;
&lt;li&gt;Low-volume, high-stakes decisions, where the cost of a wrong answer exceeds a year of model rental.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In those cases, move up to a 70B-class open model or a frontier API, and keep the same three production signals. If the larger model does not move tool-call success, refusals, or P99 in the direction you care about, the extra size is not buying anything you can use.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to pick a point on the frontier
&lt;/h2&gt;

&lt;p&gt;Use this order. It is cheaper to discover that a 27B model is enough than to discover that an 80 GB card was unnecessary.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start from the workload, not from the largest open weights you can download. Single-shot instructions and math can start at 8–14B. Multi-step tool use starts at Qwen3.8-27B if you have 24 GB, or at the best 16 GB quant you have already measured if you do not.&lt;/li&gt;
&lt;li&gt;Separate active parameters from resident parameters. Scout's 17B active figure is a compute claim. The 109B total is the VRAM claim. Budget the second number.&lt;/li&gt;
&lt;li&gt;Quantize, then re-check memory. Q4_K_M of Qwen3.8-27B is a 24 GB conversation, with about 64K of context before the card is full. Run &lt;code&gt;nvidia-smi&lt;/code&gt; after the load, at the context length you will actually serve.&lt;/li&gt;
&lt;li&gt;Price the month, not the token. Compare 730 hours of the GPU you would rent with the input and output mix you already log. If you are under the break-even above, stay on the API until the volume arrives.&lt;/li&gt;
&lt;li&gt;Decide on your own tasks. A paired run that records tool-call success, refusal behaviour, and P99 under burst is the test that public benches cannot replace.
&lt;/li&gt;
&lt;/ol&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Define the workload] --&amp;gt; B{Multi-step tool use?}
    B -- No: single-shot or math --&amp;gt; C[Start at 8-14B]
    B -- Yes --&amp;gt; D{One 24 GB GPU?}
    D -- Yes --&amp;gt; E[Qwen3.8-27B near Q4, then check KV cache]
    D -- No: 80 GB class card --&amp;gt; F[Llama 4 Scout: 17B active, 109B resident]
    C --&amp;gt; G[Paired run on your tasks]
    E --&amp;gt; G
    F --&amp;gt; G
    G --&amp;gt; H{Tool calls, refusals, and P99 pass?}
    H -- Yes --&amp;gt; I[Deploy and keep those three signals]
    H -- No, volume is low --&amp;gt; J[70B-class weights or a frontier API]
    H -- No, volume is high --&amp;gt; K[Next size up, then reprice the month]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The useful question in 2026 is which weights fit the card, the context, and the monthly bill while still clearing your own tasks. For a lot of agentic work, that point is a 27B hybrid model on one 24 GB GPU. Scout is the same idea at data-center memory: less compute per token than the parameter total suggests, and no discount on the VRAM.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;MindStudio. "Qwen 3.8 27B Benchmarked: Agentic Index, Vision, and Reasoning Tests." &lt;a href="https://www.mindstudio.ai/blog/qwen-3-27b-local-benchmark" rel="noopener noreferrer"&gt;https://www.mindstudio.ai/blog/qwen-3-27b-local-benchmark&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Qubrid AI. "Qwen3.8-27B Benchmarks: Official and Independent Results." &lt;a href="https://www.qubrid.com/blog/qwen38-27b-benchmarks-official-and-independent-results" rel="noopener noreferrer"&gt;https://www.qubrid.com/blog/qwen38-27b-benchmarks-official-and-independent-results&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Hardware Corner. "We Tested Qwen3.8 27B: How Much GPU and VRAM Do You Really Need?" &lt;a href="https://www.hardware-corner.net/qwen3-8-27b-hardware-tests/" rel="noopener noreferrer"&gt;https://www.hardware-corner.net/qwen3-8-27b-hardware-tests/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Locally Uncensored. "How to Run Qwen 3.8 27B Locally: VRAM, Quants and the Template Trap." &lt;a href="https://locallyuncensored.com/blog/how-to-run-qwen-3-8-27b-locally.html" rel="noopener noreferrer"&gt;https://locallyuncensored.com/blog/how-to-run-qwen-3-8-27b-locally.html&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Malik, Umesh. "Qwen3.8 27B VRAM: how to fit 262K context in 16 GiB, not 64." &lt;a href="https://umesh-malik.com/blog/qwen3-8-27b-vram-kv-cache-math" rel="noopener noreferrer"&gt;https://umesh-malik.com/blog/qwen3-8-27b-vram-kv-cache-math&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Meta AI. "The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation." &lt;a href="https://ai.meta.com/blog/llama-4-multimodal-intelligence" rel="noopener noreferrer"&gt;https://ai.meta.com/blog/llama-4-multimodal-intelligence&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Meta. "Llama 4 model card." &lt;a href="https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md" rel="noopener noreferrer"&gt;https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;arXiv. "Can Mixture-of-Experts Surpass Dense LLMs Under Strictly Equal Resources?" June 2025. &lt;a href="https://arxiv.org/html/2506.12119v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2506.12119v1&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Yang, S., Kautz, J., Hatamizadeh, A. "Gated Delta Networks: Improving Mamba2 with Delta Rule." ICLR 2025. &lt;a href="https://jankautz.com/publications/GatedDeltaNet_ICLR25.pdf" rel="noopener noreferrer"&gt;https://jankautz.com/publications/GatedDeltaNet_ICLR25.pdf&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Abdin, M., et al. "Phi-4 Technical Report." arXiv:2412.08905, December 2024. &lt;a href="https://arxiv.org/html/2412.08905v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2412.08905v1&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GPU Finder. "RTX 4090 GPU Rental Prices &amp;amp; Live Availability." Checked 18 September 2026. &lt;a href="https://gpufinder.dev/gpu/rtx-4090" rel="noopener noreferrer"&gt;https://gpufinder.dev/gpu/rtx-4090&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic. "What's new in Claude Sonnet 5" (pricing: $2 / $10 per million tokens). &lt;a href="https://platform.claude.com/docs/en/models/sonnet-5/whats-new-sonnet-5" rel="noopener noreferrer"&gt;https://platform.claude.com/docs/en/models/sonnet-5/whats-new-sonnet-5&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;BenchLM. "Claude API pricing." Updated 3 September 2026. &lt;a href="https://benchlm.ai/anthropic/api-pricing" rel="noopener noreferrer"&gt;https://benchlm.ai/anthropic/api-pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Future AGI. "Evaluating Cheap Frontier Models in 2026: Substitution Without a Quality Cliff." &lt;a href="https://futureagi.com/blog/evaluating-cheap-frontier-models-2026" rel="noopener noreferrer"&gt;https://futureagi.com/blog/evaluating-cheap-frontier-models-2026&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Northflank. "Qwen3.8-27B: Performance, benchmarks, GPU requirements &amp;amp; how to run it." 17 August 2026. &lt;a href="https://northflank.com/blog/qwen3-8-27b-performance-benchmarks-gpu-requirements-and-how-to-run-it" rel="noopener noreferrer"&gt;https://northflank.com/blog/qwen3-8-27b-performance-benchmarks-gpu-requirements-and-how-to-run-it&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>llm</category>
      <category>opensource</category>
      <category>nvidia</category>
    </item>
    <item>
      <title>Mnemosyne for Hermes Agent: Local Memory Quickstart</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Tue, 22 Sep 2026 10:07:07 +0000</pubDate>
      <link>https://dev.to/rosgluk/mnemosyne-for-hermes-agent-local-memory-quickstart-p0k</link>
      <guid>https://dev.to/rosgluk/mnemosyne-for-hermes-agent-local-memory-quickstart-p0k</guid>
      <description>&lt;p&gt;Mnemosyne is a local-first memory provider for Hermes Agent, storing working memory, structured facts, temporal data, and episodic history in local SQLite — no hosted service, no mandatory network calls, and unusually granular write control.&lt;/p&gt;

&lt;p&gt;Its most useful property is not raw recall quality. It is the amount of control it exposes over the write path: conversation autosave can be restricted by role or disabled outright, tool-result logging defaults off, explicit remember and forget operations stay available regardless, and newer releases add opt-in self-echo suppression around context-compression boundaries. That combination makes it a reasonable choice when you want persistent memory without automatically turning every conversation into permanent knowledge.&lt;/p&gt;

&lt;p&gt;That write-path discipline matters because agent memory has a well-documented failure mode: a model's own inference can be captured, retrieved later as if it were an observation, and used to justify an even stronger version of itself. &lt;a href="https://www.glukhov.org/ai-systems/memory/self-reinforcing-memory-loops/" rel="noopener noreferrer"&gt;Self-Reinforcing Memory Loops in AI Agents&lt;/a&gt; covers that failure mode in depth; this guide focuses on the concrete Mnemosyne configuration that limits it in practice. For where Mnemosyne sits relative to the other Hermes memory backends, see &lt;a href="https://www.glukhov.org/ai-systems/memory/agent-memory-providers/" rel="noopener noreferrer"&gt;Agent Memory Providers Compared&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mnemosyne in one minute
&lt;/h2&gt;

&lt;p&gt;A typical memory provider does some version of capture, extract, store, retrieve, then inject into a future prompt. Mnemosyne adds several distinct layers around that basic loop: working memory, semantic and lexical recall, structured facts, temporal information, entity links, episodic memory, consolidation, canonical facts, and memory validation. Storage is local SQLite with FTS5 and optional vector retrieval, which makes it considerably more inspectable than a cloud-only memory product and more capable than a plain &lt;code&gt;MEMORY.md&lt;/code&gt; file.&lt;/p&gt;

&lt;p&gt;Very briefly, relative to the rest of the Hermes provider ecosystem: Holographic is simpler and deliberately fact-store oriented; Hindsight emphasizes hybrid retrieval, knowledge graphs, and reflection; Honcho emphasizes peer and user modeling with dialectic reasoning; Mem0 emphasizes automatic LLM-based fact extraction; and Mnemosyne combines local SQLite storage, hybrid recall, consolidation, structured facts, and unusually granular retention controls. The full breakdown, including infrastructure requirements and self-hosting notes for every provider, is in &lt;a href="https://www.glukhov.org/ai-systems/memory/agent-memory-providers/" rel="noopener noreferrer"&gt;Agent Memory Providers Compared&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current versions
&lt;/h2&gt;

&lt;p&gt;As of September 2026, the stable PyPI release is &lt;code&gt;mnemosyne-memory 3.15.1&lt;/code&gt;, with the 4.0 branch available as a pre-release. For a production Hermes installation, start with the stable version unless you specifically need a 4.0 fix or feature and are prepared to test the database migration and behavior change. Check your installed version with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne version
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Installing Mnemosyne into Hermes
&lt;/h2&gt;

&lt;p&gt;Activate Hermes' own virtual environment first if you used the standard local installation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;source&lt;/span&gt; ~/.hermes/hermes-agent/venv/bin/activate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For local embedding support, install the core package with the embeddings extra plus the Hermes plugin wrapper:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"mnemosyne-memory[embeddings]"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  mnemosyne-hermes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then register the plugin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mnemosyne-hermes &lt;span class="nb"&gt;install&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you are replacing an existing plugin registration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;mnemosyne-hermes &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--force&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Activate the provider and restart the gateway:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes config &lt;span class="nb"&gt;set &lt;/span&gt;memory.provider mnemosyne
hermes gateway restart
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verify with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes memory status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expected output looks similar to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Provider: mnemosyne

Plugin: installed
Status: available
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Docker and persistent-server installs
&lt;/h3&gt;

&lt;p&gt;If Hermes runs inside a persistent Docker or image-based deployment, install into a side virtual environment on the mounted Hermes home instead of the container's rebuildable Python environment, so the plugin survives image rebuilds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;HERMES_HOME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/opt/data
&lt;span class="nv"&gt;VENV&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HERMES_HOME&lt;/span&gt;&lt;span class="s2"&gt;/.mnemosyne/venv"&lt;/span&gt;
python3 &lt;span class="nt"&gt;-m&lt;/span&gt; venv &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$VENV&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$VENV&lt;/span&gt;&lt;span class="s2"&gt;/bin/python"&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--upgrade&lt;/span&gt; &lt;span class="s2"&gt;"mnemosyne-memory[embeddings]"&lt;/span&gt; mnemosyne-hermes
&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$VENV&lt;/span&gt;&lt;span class="s2"&gt;/bin/mnemosyne-hermes"&lt;/span&gt; &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--mode&lt;/span&gt; wrapper &lt;span class="nt"&gt;--python&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$VENV&lt;/span&gt;&lt;span class="s2"&gt;/bin/python"&lt;/span&gt;
hermes config &lt;span class="nb"&gt;set &lt;/span&gt;memory.provider mnemosyne
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The side venv must use the same Python major/minor version as the running Hermes gateway — do not point it at an unrelated &lt;code&gt;python3&lt;/code&gt; from &lt;code&gt;PATH&lt;/code&gt;. Restart the actual container or service afterward and verify with &lt;code&gt;"$VENV/bin/mnemosyne-hermes" status&lt;/code&gt; alongside &lt;code&gt;hermes memory status&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not disable the whole Hermes memory toolset
&lt;/h2&gt;

&lt;p&gt;Keep two concepts separate: Hermes' own built-in memory (&lt;code&gt;MEMORY.md&lt;/code&gt; / &lt;code&gt;USER.md&lt;/code&gt;, covered in full in &lt;a href="https://www.glukhov.org/ai-systems/hermes/hermes-agent-memory-system/" rel="noopener noreferrer"&gt;Hermes Agent Memory System&lt;/a&gt;) and the external provider (Mnemosyne). Do not casually run &lt;code&gt;hermes tools disable memory&lt;/code&gt; when configuring an external provider — depending on the Hermes version, that command can also hide external memory-provider tools. Use provider configuration instead, as shown below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Basic status and inspection
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes memory status
hermes mnemosyne stats
hermes mnemosyne stats &lt;span class="nt"&gt;--global&lt;/span&gt;
hermes mnemosyne inspect &lt;span class="s2"&gt;"query"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Export a portable backup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne &lt;span class="nb"&gt;export&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; ~/mnemosyne-backup.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The backing database normally lives under &lt;code&gt;~/.hermes/mnemosyne/data/mnemosyne.db&lt;/code&gt;. Because it is SQLite, inspection and backup are straightforward with standard tools. For the rest of the gateway, session, and diagnostics commands referenced throughout this guide, the &lt;a href="https://www.glukhov.org/ai-systems/hermes/hermes-agent-cli-cheatsheet/" rel="noopener noreferrer"&gt;Hermes Agent CLI cheat sheet&lt;/a&gt; is a faster reference than digging through &lt;code&gt;--help&lt;/code&gt; output.&lt;/p&gt;

&lt;h2&gt;
  
  
  The default retention policy deserves attention
&lt;/h2&gt;

&lt;p&gt;The first control worth understanding is &lt;code&gt;sync_roles&lt;/code&gt;. Current Mnemosyne defaults are already more conservative than early releases — automatic Hermes synchronization defaults to user turns rather than both user and assistant turns — but for strict explicit-only retention, disabling turn autosave completely is worth the extra step. Edit &lt;code&gt;~/.hermes/config.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mnemosyne&lt;/span&gt;

  &lt;span class="na"&gt;mnemosyne&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;sync_roles&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An empty list means ordinary conversation turns are not automatically saved by &lt;code&gt;sync_turn()&lt;/code&gt;. Explicit &lt;code&gt;mnemosyne_remember&lt;/code&gt; operations continue to work regardless — normal conversation stops flowing into memory automatically, while an explicit "remember this" still reaches Mnemosyne.&lt;/p&gt;

&lt;h2&gt;
  
  
  Disable automatic tool-result logging
&lt;/h2&gt;

&lt;p&gt;Mnemosyne can also log tool executions as memory. For a conservative setup, leave that disabled in &lt;code&gt;~/.hermes/.env&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MNEMOSYNE_LOG_TOOLS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is already the default, but setting it explicitly documents the policy rather than relying on an assumption about defaults. Restart Hermes afterward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway restart
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;sync_roles: []&lt;/code&gt; and &lt;code&gt;MNEMOSYNE_LOG_TOOLS=0&lt;/code&gt; together, both major automatic write paths — conversation autosave and tool-result autosave — are off.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep automatic recall
&lt;/h2&gt;

&lt;p&gt;Disabling automatic writes does not require disabling recall. A useful policy keeps automatic retention off while automatic recall, explicit remember, and explicit forget all stay on — memory should be easy to read and difficult to write, which is close to the opposite of a "capture everything and sort it out later" default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add a durable agent instruction
&lt;/h2&gt;

&lt;p&gt;Provider configuration blocks automatic provider-level capture, but the model can still decide to call an explicit write tool on its own initiative. Add an explicit policy to &lt;code&gt;SOUL.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Long-term memory policy&lt;/span&gt;

Mnemosyne is the long-term memory provider.

Do not write anything to Mnemosyne unless the user explicitly asks you to
remember, save, retain, or store that information.

If information appears useful for future sessions but the user did not
explicitly request that it be remembered, ask for permission before calling
mnemosyne_remember or another Mnemosyne write tool.

Do not create durable memories from your own reasoning, assumptions,
summaries, interpretations, conclusions, or inferred preferences.

Do not create durable memories from tool output unless the user explicitly
asks for that result to be remembered.

When storing an approved memory, preserve what the user actually stated.
Do not embellish it with inferred context or conclusions.

Reading and recalling Mnemosyne memories is allowed without asking for
permission.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restart the gateway and start a fresh session afterward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway restart
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/new
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a model-enforced policy, not a hard permission boundary — it complements the provider-level configuration above rather than replacing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What about &lt;code&gt;memory.write_approval&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;Hermes supports &lt;code&gt;memory.write_approval: true&lt;/code&gt; for built-in &lt;code&gt;MEMORY.md&lt;/code&gt; / &lt;code&gt;USER.md&lt;/code&gt; writes, and Mnemosyne implements its own provider-specific staging for explicit writes in newer releases. This is promising, but there is an architectural caveat worth taking seriously: Hermes does not yet expose one uniform, provider-neutral approval contract across all external memory providers, and Mnemosyne's pending/apply implementation is provider-specific rather than part of a shared standard. Do not assume approval works correctly just because the configuration key is present — test it against your exact Hermes and Mnemosyne versions. Until provider-independent approval matures, combining &lt;code&gt;sync_roles: []&lt;/code&gt;, &lt;code&gt;MNEMOSYNE_LOG_TOOLS=0&lt;/code&gt;, and the explicit-write &lt;code&gt;SOUL.md&lt;/code&gt; policy above gives you a dependable baseline, with the approval path tested separately if you intend to rely on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enable self-echo suppression
&lt;/h2&gt;

&lt;p&gt;Current Mnemosyne also offers optional self-echo suppression:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MNEMOSYNE_SELF_ECHO_ENABLED&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Put this in &lt;code&gt;~/.hermes/.env&lt;/code&gt;, then restart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway restart
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Self-echo suppression targets context-compression boundaries specifically — its purpose is to reduce cases where memory the provider just created gets immediately fed back into the agent as if it were independent context. It is intentionally best-effort and does not replace write filtering: write controls stop questionable memories from entering in the first place, while self-echo controls stop recent provider output from bouncing straight back. Both matter, and neither substitutes for the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  A conservative Mnemosyne configuration
&lt;/h2&gt;

&lt;p&gt;Putting the pieces together, a starting configuration for a self-hosted personal engineering agent looks like this. In &lt;code&gt;~/.hermes/config.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mnemosyne&lt;/span&gt;

  &lt;span class="na"&gt;mnemosyne&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;sync_roles&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In &lt;code&gt;~/.hermes/.env&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MNEMOSYNE_LOG_TOOLS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0
&lt;span class="nv"&gt;MNEMOSYNE_SELF_ECHO_ENABLED&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And in &lt;code&gt;SOUL.md&lt;/code&gt;, at minimum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Only store long-term memory when the user explicitly requests it.
Do not promote model-generated conclusions or tool output into durable memory
without explicit permission.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    U[User conversation] -.-&amp;gt;|blocked| M[(Mnemosyne)]
    T[Tool results] -.-&amp;gt;|blocked| M
    R["Explicit: remember this"] --&amp;gt;|mnemosyne_remember| M
    Q[Future question] --&amp;gt;|recall| M&lt;/code&gt;&lt;/pre&gt;



&lt;h2&gt;
  
  
  Test that ordinary conversation is not retained
&lt;/h2&gt;

&lt;p&gt;Check the baseline count first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne stats
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start a new Hermes session and say a plain factual statement without asking the agent to remember it, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PurpleOtter uses port 48123.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Afterward, search for it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne inspect &lt;span class="s2"&gt;"PurpleOtter"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expected: &lt;code&gt;Results for 'PurpleOtter': 0&lt;/code&gt;. Also re-check &lt;code&gt;hermes mnemosyne stats&lt;/code&gt; — the working-memory count should not have increased because of that ordinary turn.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test explicit memory
&lt;/h2&gt;

&lt;p&gt;Now say the same kind of statement, but explicitly ask for retention:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Remember that BlueKoala uses port 17321.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Inspect it, then start a new session and ask for it back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne inspect &lt;span class="s2"&gt;"BlueKoala"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/new
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What port does BlueKoala use?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hermes should retrieve the value correctly — this pair of tests isolates the write-path policy (nothing gets in without asking) from the retrieval mechanism (what gets in comes back out reliably).&lt;/p&gt;

&lt;h2&gt;
  
  
  Test tool logging
&lt;/h2&gt;

&lt;p&gt;With &lt;code&gt;MNEMOSYNE_LOG_TOOLS=0&lt;/code&gt; set, ask Hermes to run a distinctive, unique command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Use the terminal tool to run:
echo tool-canary-834729
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then search for the canary string:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne inspect &lt;span class="s2"&gt;"tool-canary-834729"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Expected: &lt;code&gt;0 results&lt;/code&gt;. This is a much stronger test than simply trusting that the environment variable is honored everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inspecting the database
&lt;/h2&gt;

&lt;p&gt;Because storage is SQLite, the internal schema is directly inspectable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;sqlite3 ~/.hermes/mnemosyne/data/mnemosyne.db &lt;span class="s1"&gt;'.tables'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Depending on version, you may see tables such as &lt;code&gt;working_memory&lt;/code&gt;, &lt;code&gt;episodic_memory&lt;/code&gt;, &lt;code&gt;facts&lt;/code&gt;, &lt;code&gt;consolidated_facts&lt;/code&gt;, &lt;code&gt;gists&lt;/code&gt;, &lt;code&gt;graph_edges&lt;/code&gt;, &lt;code&gt;memoria_facts&lt;/code&gt;, and &lt;code&gt;memory_embeddings&lt;/code&gt;. This matters when testing deletion — a memory system can successfully remove a working-memory row while leaving a derived fact, gist, or graph object behind. Mnemosyne has had real bugs in this area involving orphaned derived records, and newer releases have tightened both deletion and diagnostics accordingly. Prefer the provider's supported delete and doctor/repair paths over manually deleting SQLite rows unless you fully understand the current schema.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deleting session-scoped working memory
&lt;/h3&gt;

&lt;p&gt;One subtlety: Mnemosyne working memories can be session-scoped, so a row with &lt;code&gt;scope = session&lt;/code&gt; may not be visible to a standalone delete operating in the &lt;code&gt;default&lt;/code&gt; session. When debugging, inspect scope directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;session_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;working_memory&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The provider or API needs the correct session scope to mutate session-local records — another reason to prefer supported administration tools over raw SQL edits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consolidation: do not rush to &lt;code&gt;sleep()&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;Mnemosyne can consolidate working memory into longer-lived representations, which is useful but is a mutating operation. Before enabling aggressive automatic consolidation, inspect what is actually being captured, verify that ordinary turns are not entering memory unexpectedly, verify deletion end to end, and back up the database. Then experiment with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne &lt;span class="nb"&gt;sleep&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Recent Mnemosyne changes made conflict handling more conservative — semantic similarity alone no longer proves that one memory should invalidate another, which is exactly the direction a durable agent-memory system should move in, as covered in &lt;a href="https://www.glukhov.org/ai-systems/memory/self-reinforcing-memory-loops/" rel="noopener noreferrer"&gt;Self-Reinforcing Memory Loops in AI Agents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Backup before upgrades
&lt;/h2&gt;

&lt;p&gt;Create a portable export before any significant change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes mnemosyne &lt;span class="nb"&gt;export&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; ~/mnemosyne-backup.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For important installations, also copy the local database or data directory before major upgrades. Mnemosyne 4.x is currently a pre-release line, so a major-version upgrade deserves more caution than a routine patch update.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final recommended setup
&lt;/h2&gt;

&lt;p&gt;For a long-running Hermes installation where memory accuracy matters more than remembering everything, the durable configuration is: Mnemosyne local storage on, automatic recall on, conversation autosave off, assistant-message autosave off, tool-result logging off, explicit remember and forget on, self-echo suppression on, session search on, and human review for sensitive writes desirable once the approval path is tested. That makes Mnemosyne function primarily as a curated long-term memory store rather than a transcript archive — the goal is not to make Hermes remember everything it has ever said, but to make it remember the things that will still be true when the next session begins. If you run several profiles with different providers or retention policies, &lt;a href="https://www.glukhov.org/ai-systems/hermes/production-setup/" rel="noopener noreferrer"&gt;Hermes Agent production setup&lt;/a&gt; covers the profile-level wiring for keeping them consistent.&lt;/p&gt;

</description>
      <category>hermes</category>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Self-Reinforcing Memory Loops in AI Agents: Causes and Fixes</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Sat, 19 Sep 2026 09:56:04 +0000</pubDate>
      <link>https://dev.to/rosgluk/self-reinforcing-memory-loops-in-ai-agents-causes-and-fixes-35d8</link>
      <guid>https://dev.to/rosgluk/self-reinforcing-memory-loops-in-ai-agents-causes-and-fixes-35d8</guid>
      <description>&lt;p&gt;Persistent memory turns an agent from a re-explained tool into one that carries context forward — but it opens a failure mode stateless chat avoids: an interpretation can become memory, retrieved as fact, and justify a stronger version of itself.&lt;/p&gt;

&lt;p&gt;That is a self-reinforcing memory loop. The mechanism does not require malice, a broken plugin, or an unusual prompt — a normal capture pipeline stores the assistant's output, a normal retrieval pipeline surfaces it as context, and the model treats retrieved text as evidence because that is what retrieved text usually is.&lt;/p&gt;

&lt;p&gt;This is different from a hallucination in one important way: a hallucination disappears when the conversation ends, but a hallucination that gets promoted into durable memory can outlive the session that created it, resurface weeks later in an unrelated context, and gain apparent credibility purely from repetition. For the broader memory model this problem sits inside — working memory, structured state, and retrieval memory as three separate contracts, part of the &lt;a href="https://www.glukhov.org/ai-systems/memory/" rel="noopener noreferrer"&gt;AI Systems Memory hub&lt;/a&gt; — see &lt;a href="https://www.glukhov.org/ai-systems/memory/memory-systems-in-ai-assistants/" rel="noopener noreferrer"&gt;Memory Systems in AI Assistants&lt;/a&gt;, which already flags stale and contradictory memory as the most common production failure. This article goes one level deeper into why that specific failure keeps recurring.&lt;/p&gt;

&lt;p&gt;The question worth asking about any memory system is not simply whether it remembers. It is what is allowed to become evidence for future reasoning. Answering that well requires distinguishing something the user explicitly stated, something a tool actually observed, something an external document reported, and something the model merely inferred or summarized. Once those categories collapse into a single undifferentiated pool called "memory", a generated conclusion becomes indistinguishable from an observation — the memory system has effectively laundered an inference into a premise. For a practical walkthrough of running one provider with that distinction enforced, see &lt;a href="https://www.glukhov.org/ai-systems/hermes/mnemosyne-memory/" rel="noopener noreferrer"&gt;Mnemosyne for Hermes Agent: Local Memory Quickstart&lt;/a&gt;; for how the eight-plus mainstream providers differ on this exact axis, see &lt;a href="https://www.glukhov.org/ai-systems/memory/agent-memory-providers/" rel="noopener noreferrer"&gt;Agent Memory Providers Compared&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a self-reinforcing memory loop?
&lt;/h2&gt;

&lt;p&gt;The simplest version has a fixed shape: a user states something, the agent infers a conclusion from it, memory stores that conclusion, a future session recalls it, the agent treats the recalled statement as evidence, derives a stronger conclusion, and writes that back to memory. The cycle then repeats with a slightly more confident claim each time, without a single new observation ever entering the pipeline.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[User states X] --&amp;gt; B[Agent infers Y]
    B --&amp;gt; C[Memory stores Y]
    C --&amp;gt; D[Future session recalls Y]
    D --&amp;gt; E[Agent treats Y as evidence]
    E --&amp;gt; F[Agent derives stronger Y2]
    F --&amp;gt; C&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Consider a developer who tells an agent that a deployment failed after a cache configuration change. A reasonable inference is that the cache configuration probably caused the failure — useful reasoning if it stays inside the current context. The damage begins when an automatic memory extractor stores the flat claim "the cache configuration caused the deployment failure" as a fact. A week later, a second unrelated deployment fails; the agent retrieves the stored claim and reasons that the cache layer has a history of instability, which gets written back as an even more general belief. By the third pass, the stored memory reads as "the cache layer is known to be unreliable and should be replaced" — a confident institutional claim built from zero new evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why persistent agent memory makes this worse than a database
&lt;/h2&gt;

&lt;p&gt;A conventional application database has an explicit write path: a field changes because a known user, API call, or transaction changed it. Agent memory systems typically have many more writers — the user, the assistant, tool results, an automatic turn-capture hook, a fact extractor, a session summarizer, a reflection pass, a consolidation process, and sometimes another agent — and just as many readers, including automatic prompt injection, semantic retrieval, and sub-agent tools. Once the output of one reader can become the input to another writer, the system is a feedback loop rather than a simple store, and ordinary database intuitions about "who wrote this and when" stop applying.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main forms of memory feedback
&lt;/h2&gt;

&lt;p&gt;Self-reinforcement is not one mechanism — it shows up in at least seven related but distinct patterns, and a memory provider can be resistant to one and vulnerable to another.&lt;/p&gt;

&lt;h3&gt;
  
  
  Assistant self-echo
&lt;/h3&gt;

&lt;p&gt;The simplest case occurs when assistant messages are automatically retained: the model's own previous answer becomes contextual evidence for its next answer. That does not automatically make the answer wrong, but it does change its epistemic status — generated language has become persistent context. The safest general rule is that user and tool observations may be memory candidates, but assistant conclusions should not automatically become facts without a separate promotion step.&lt;/p&gt;

&lt;h3&gt;
  
  
  Summary-of-summary drift
&lt;/h3&gt;

&lt;p&gt;Long-running agents compress conversations repeatedly — raw conversation to summary, summary to long-term memory, memory to a user profile — and each transformation can quietly discard a qualifier. "I usually use PostgreSQL, but SQLite is fine for small tools" can become "User prefers PostgreSQL," then "User uses PostgreSQL," then "User's projects use PostgreSQL," at which point a future SQLite suggestion gets flagged as violating the user's architecture preference. No single step in that chain is dramatic; the cumulative effect is a wrong belief with a completely plausible paper trail.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reflection amplification
&lt;/h3&gt;

&lt;p&gt;Some providers intentionally perform higher-order reasoning over stored memories — Hindsight's &lt;code&gt;reflect&lt;/code&gt; is a documented example — and this is genuinely useful because agents need synthesis, not only retrieval. The risk begins when a derived conclusion is stored alongside the raw observations it came from with no marker distinguishing them, so a later reader sees four apparently independent facts instead of three observations and one interpretation of those observations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Retrieval amplification
&lt;/h3&gt;

&lt;p&gt;Retrieval itself introduces bias without any reflection step at all: a memory that gets retrieved often appears in more prompts, gets mentioned more often, gets recaptured more often, and produces more related memories, which in turn get retrieved even more often. The memory becomes prominent partly because it was already prominent — a popularity loop rather than an evidence loop.&lt;/p&gt;

&lt;h3&gt;
  
  
  Contradiction collapse
&lt;/h3&gt;

&lt;p&gt;A dangerous pattern appears when a memory system decides which of two conflicting statements is true using similarity alone. Mnemosyne provides a concrete real-world example: a production audit found that similarity-based conflict handling had invalidated 142 of 243 stored items across consolidation passes, because the system treated "these two statements look alike" as proof that one superseded the other. Newer Mnemosyne releases now treat similarity as a candidate contradiction rather than proof — actual invalidation requires a successful validation step — which is the right general direction for any provider with a consolidation pass. The underlying lesson generalizes well beyond Mnemosyne: semantic similarity is not evidence of contradiction, since two statements can differ because of date, environment, branch, or deployment rather than because one is wrong.&lt;/p&gt;

&lt;h3&gt;
  
  
  User-model reinforcement
&lt;/h3&gt;

&lt;p&gt;Systems that maintain a running model of the user, not just a list of facts, face a sharper version of the same problem. "User prefers concise answers" or "user deploys to AWS" are useful and low-risk; "user dislikes technology X" or "user always chooses architecture Y" are inferred traits that, if partly based on the agent's own earlier interpretations, can progressively turn a real person into a caricature of one interaction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Agent self-model reinforcement
&lt;/h3&gt;

&lt;p&gt;The most subtle case is an agent that models itself: it performs an action, explains that action, and a memory system builds a self-model from the explanation that gets fed into the next session, which then behaves according to that self-model and reinforces it further. A useful self-model can stabilize an agent's behavior over time. A wrong one stabilizes the agent around the wrong behavior just as effectively — the loop does not care which direction it locks in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why confidence tends to increase along the way
&lt;/h2&gt;

&lt;p&gt;Large language models do not automatically know that a retrieved sentence was originally generated by another instance of themselves. "User: I think server X might have a networking problem" reads as tentative; "Relevant memory: Server X has a networking problem" reads as settled, even though both may trace back to the same uncertain guess. That syntactic shift from a hedged claim to a declarative memory object is source laundering, and it compounds when several derived memories happen to agree with each other — three semantically similar memories can look like independent corroboration even when all three originated from a single conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Consequences that show up in production systems
&lt;/h2&gt;

&lt;p&gt;The practical damage takes a handful of recognizable shapes. False certainty means the agent stops checking an assumption because memory presents it as already settled. Preference drift means a tentative preference gradually hardens into an absolute instruction. Wrong user profiles mean one unusual interaction gets generalized into a long-term behavioral trait. Tool-action cascades are the most expensive version: a false remembered premise drives a wrong diagnosis, which drives a tool call, which drives a real configuration change — persistent agents raise the cost of a memory error precisely because the error can reach the outside world. Duplicate-memory inflation and stale-state lock-in both waste prompt budget and retrieval quality over time, and destructive consolidation can let an inferred, newer-looking statement quietly supersede an older but more authoritative observation.&lt;/p&gt;

&lt;p&gt;Deletion deserves a specific warning here. Modern providers frequently build several derived structures from one captured item — working memory, extracted facts, summaries, embeddings, graph edges, canonical facts, and profile entries — and deleting the original memory does not guarantee every derived representation disappears with it. Memory deletion needs to be tested end-to-end, not assumed to work because an API returned success.&lt;/p&gt;

&lt;h2&gt;
  
  
  Provenance matters more than embedding quality
&lt;/h2&gt;

&lt;p&gt;Most memory engineering effort goes into retrieval — vector similarity, BM25, hybrid search, rerankers, graph traversal, temporal weighting — and all of that is genuinely useful, but none of it addresses the fundamental problem, because retrieval quality only affects which memories surface, not whether a surfaced memory deserves the confidence it is given.&lt;/p&gt;

&lt;p&gt;A production memory object should carry metadata beyond its content: source, source type, timestamp, scope, confidence, what it was derived from, validation status, and whether it has been superseded. A rough tiering that works in practice ranks explicit user statements and direct tool observations highest, trusted external data next, deterministic extraction below that, then summaries, with model inference and reflection output at the bottom of the trust hierarchy — not because inference is worthless, but because it should never silently inherit the trust level of the observation it was built from. Retrieval and consolidation can then respect that hierarchy instead of ranking purely by semantic similarity.&lt;/p&gt;

&lt;h2&gt;
  
  
  A safer architectural pattern
&lt;/h2&gt;

&lt;p&gt;For most personal and engineering agents, a deliberately boring memory pipeline outperforms a fully automatic one. The key design decision is that the agent does not turn every conversation into durable truth by default — a candidate memory gets classified before it is retained, with observations retained, inferences kept transient, and uncertain cases routed back to the user rather than silently written.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Conversation] --&amp;gt; B[Candidate memory]
    B --&amp;gt; C{Classification}
    C --&amp;gt;|Observation| D[Retain]
    C --&amp;gt;|Inference| E[Keep transient]
    C --&amp;gt;|Uncertain| F[Ask user]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;For high-value environments, a human approval gate is worth the friction: a candidate memory moves to pending status, a human reviews it, and only an explicit approval commits it to durable memory while a rejection discards it. Hermes' own &lt;code&gt;memory.write_approval: true&lt;/code&gt; setting stages built-in &lt;code&gt;MEMORY.md&lt;/code&gt; writes for exactly this reason, and the same idea shows up as provider-specific staged writes in Mnemosyne — though it is worth noting that no uniform, provider-independent approval contract exists across Hermes external memory plugins yet, so this path should be tested against the exact versions you run rather than assumed to work everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  How current providers address the problem
&lt;/h2&gt;

&lt;p&gt;No provider eliminates feedback loops completely; each makes a different trade-off between convenience and control.&lt;/p&gt;

&lt;p&gt;Hermes' own built-in &lt;code&gt;MEMORY.md&lt;/code&gt; and &lt;code&gt;USER.md&lt;/code&gt; files are intentionally small and human-readable, which makes them easy to audit even without special tooling — the trade-off is scale, since this is not a semantic long-term memory database. &lt;a href="https://www.glukhov.org/ai-systems/hermes/hermes-agent-memory-system/" rel="noopener noreferrer"&gt;Hermes Agent Memory System&lt;/a&gt; covers that bounded design in full.&lt;/p&gt;

&lt;p&gt;Mnemosyne is one of the more governance-oriented external providers precisely because it exposes independent controls over what gets written: conversation autosave can be disabled entirely with &lt;code&gt;sync_roles: []&lt;/code&gt; while explicit memory operations stay available, tool-result logging defaults off, and newer builds add opt-in self-echo suppression around context-compression boundaries. &lt;a href="https://www.glukhov.org/ai-systems/hermes/mnemosyne-memory/" rel="noopener noreferrer"&gt;Mnemosyne for Hermes Agent: Local Memory Quickstart&lt;/a&gt; walks through a conservative configuration end to end.&lt;/p&gt;

&lt;p&gt;Hindsight's default Hermes integration is comparatively automatic — &lt;code&gt;autoRecall&lt;/code&gt; and &lt;code&gt;autoRetain&lt;/code&gt; both default to true — which is convenient but increases the number of feedback paths; setting &lt;code&gt;auto_retain=false&lt;/code&gt; while keeping recall on is worth considering if provenance matters more than convenience. Holographic and ByteRover both default &lt;code&gt;auto_extract&lt;/code&gt; to off, which means they can operate primarily as explicit fact stores rather than automatic transcript-to-memory pipelines, an advantage if feedback loops are your main concern. Honcho's &lt;code&gt;unified&lt;/code&gt; observation mode is more conservative than its &lt;code&gt;directional&lt;/code&gt; default because it lets the AI model the user without building the matching self-observation loop from its own messages — worth serious consideration for anyone specifically worried about agent self-model reinforcement. &lt;a href="https://www.glukhov.org/ai-systems/memory/agent-memory-providers/" rel="noopener noreferrer"&gt;Agent Memory Providers Compared&lt;/a&gt; has the full provider-by-provider comparison, including capture policy and approval support for each.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configuration questions that matter more than benchmarks
&lt;/h2&gt;

&lt;p&gt;Recall benchmarks measure whether an agent can retrieve the right information. Production systems need answers to a different set of questions: what gets written automatically, can assistant output become memory, are tool results retained automatically, are summaries stored as facts, are derived facts marked as derived, can old memories be superseded automatically, can a user inspect everything retained, does delete remove derived representations too, can automatic recall be disabled independently from automatic retention, is there a human approval gate, and can the agent itself bypass that gate. Those eleven questions are usually more diagnostic than another five points on a long-term recall benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  My preferred policy for personal engineering agents
&lt;/h2&gt;

&lt;p&gt;For a self-hosted engineering assistant, automatic conversation retention, automatic assistant retention, and automatic tool-result retention should all default off, while automatic recall stays on or selective, explicit remember stays on, session history search stays on, and derived conclusions stay transient by default rather than durable. The durable store should contain facts worth carrying into another session; the original session history should remain separately searchable when the agent actually needs evidence rather than a summary of it. Memory becomes concise retained knowledge, and session search becomes the original evidence — the two should never be conflated into one undifferentiated pool.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to test a memory provider
&lt;/h2&gt;

&lt;p&gt;Testing whether a provider remembers is the easy half. The harder and more useful half is testing whether it refuses to remember and whether it forgets completely when asked.&lt;/p&gt;

&lt;p&gt;Tell the agent an ordinary fact without asking it to remember anything, start a new session, and confirm the value does not appear if automatic capture is supposed to be disabled. Then explicitly ask it to remember a different fact, start a new session, and confirm that one does retrieve correctly — this pair of tests isolates the write-path policy from the retrieval mechanism. Separately, give the agent enough information to make an inference but never state that inference yourself, then inspect the memory database directly; the inference should not silently appear as a standalone fact. Run a distinctive, unique tool command and search memory for it afterward to confirm tool-result logging behaves as configured. Store a fact, delete it, and then check every layer a provider might use — working memory, semantic recall, fact tables, graph nodes, summaries, embeddings, and profile context — because a successful &lt;code&gt;delete&lt;/code&gt; API response is not sufficient proof that the data is actually gone. Finally, store two contradictory facts and inspect whether the provider keeps both with timestamps, marks one superseded, destroys the old record, or asks for validation — that single test reveals more about a provider's epistemic model than any feature list.&lt;/p&gt;

&lt;h2&gt;
  
  
  The central design rule
&lt;/h2&gt;

&lt;p&gt;A model-generated conclusion must not become stronger evidence merely because the same model remembered it. Memory systems need provenance, controlled write paths, explicit treatment of derived knowledge, and deletion that actually reaches every derived representation, not just the record a user can see. The most advanced memory provider is not necessarily the one that remembers the most — for long-running agents, the better provider is often the one that knows when &lt;em&gt;not&lt;/em&gt; to remember.&lt;/p&gt;

</description>
      <category>hermes</category>
      <category>openclaw</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>OpenSpec Rejected Proposals: A Decision Memory Convention</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Fri, 18 Sep 2026 09:09:18 +0000</pubDate>
      <link>https://dev.to/rosgluk/openspec-rejected-proposals-a-decision-memory-convention-312n</link>
      <guid>https://dev.to/rosgluk/openspec-rejected-proposals-a-decision-memory-convention-312n</guid>
      <description>&lt;p&gt;An agent that proposed and shipped "move persistence into a shared library" six months ago will happily propose it again next quarter unless something durable tells it the idea was already investigated and rejected -- and &lt;a href="https://github.com/Fission-AI/OpenSpec" rel="noopener noreferrer"&gt;OpenSpec&lt;/a&gt; has no built-in state for that today.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/opsx:archive&lt;/code&gt; is built for one outcome: a change that shipped. It syncs the delta specs into &lt;code&gt;openspec/specs/&lt;/code&gt; and moves the folder to &lt;code&gt;openspec/changes/archive/YYYY-MM-DD-&amp;lt;name&amp;gt;/&lt;/code&gt; as a record of what changed and why. There is no &lt;code&gt;/opsx:reject&lt;/code&gt; or &lt;code&gt;/opsx:abandon&lt;/code&gt; counterpart, and nothing in the archive format tells a future proposal "this exact idea was investigated and turned down." That gap matters most in exactly the codebases where OpenSpec is otherwise a good fit: brownfield systems with a small number of contributors and agents that periodically re-explore the same architectural questions -- merge these two services, share this persistence layer, replace this HTTP boundary with a direct import.&lt;/p&gt;

&lt;p&gt;This is not a hypothetical gap. It was raised directly with OpenSpec's own maintainers as a feature request, and the way that conversation played out is worth knowing before you improvise your own fix: what the project actually concluded shapes which convention is worth adopting. This guide walks through what happens if you rely on &lt;code&gt;/opsx:archive&lt;/code&gt; alone, the real discussion that already happened in OpenSpec's issue tracker, and a lightweight &lt;code&gt;decision.md&lt;/code&gt; pattern you can adopt today without waiting for -- or needing -- core support.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Archiving Alone Doesn't Record a Rejected Decision
&lt;/h2&gt;

&lt;p&gt;Archiving a change you decided not to build technically works -- the folder moves out of your active list either way. The problem is what that archived folder fails to communicate once it is sitting next to dozens of shipped changes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No status field.&lt;/strong&gt; An archived change looks identical whether it shipped or was abandoned three messages into &lt;code&gt;/opsx:propose&lt;/code&gt;. A teammate or an agent scanning &lt;code&gt;openspec/changes/archive/&lt;/code&gt; cannot tell the difference without opening every proposal folder and reading the artifacts inside it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No signal to check first.&lt;/strong&gt; Nothing in the default workflow instructs an agent to search the archive before drafting a new proposal. &lt;code&gt;/opsx:propose&lt;/code&gt; drafts from your current request and the state of the codebase, full stop -- it does not cross-reference prior rejected changes unless you tell it to.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Delta specs you don't want synced.&lt;/strong&gt; If a rejected change already has draft delta specs and you archive it the ordinary way, &lt;code&gt;/opsx:archive&lt;/code&gt; will offer to sync those deltas into &lt;code&gt;openspec/specs/&lt;/code&gt; first. Accepting that offer teaches your canonical specs to describe behavior you decided &lt;em&gt;not&lt;/em&gt; to build, which quietly corrupts the "what does the system currently do" record that every other proposal reads before it plans anything.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a bug. &lt;code&gt;/opsx:archive&lt;/code&gt; is doing exactly what its documentation says it does: complete a change that shipped. The rejection case sits outside that documented scope on purpose, and OpenSpec's own team-workflow guide is explicit that most of what it recommends -- branch conventions, PR review order, when to archive -- is &lt;em&gt;convention layered on top of the tool&lt;/em&gt;, not something OpenSpec enforces for you. Handling a rejection is one more convention you get to define yourself, and the CLI already gives you the flag you need to do it cleanly: pass &lt;code&gt;--skip-specs&lt;/code&gt; when you archive a change you are not shipping, so &lt;code&gt;openspec archive investigate-shared-persistence --skip-specs&lt;/code&gt; files the folder away without touching &lt;code&gt;openspec/specs/&lt;/code&gt; at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What OpenSpec's Maintainers Actually Decided About ADR Support
&lt;/h2&gt;

&lt;p&gt;Before inventing a house convention, it is worth reading how this exact question played out in public, because the resolution is more specific -- and more interesting -- than "no." &lt;a href="https://github.com/Fission-AI/OpenSpec/issues/557" rel="noopener noreferrer"&gt;GitHub issue #557&lt;/a&gt; opened in January 2026 with a request for first-class Architecture Decision Record support: durable records that persist independently of any single change's lifecycle, so a rejected or superseded decision stays visible to every future proposal. A contributor even opened a pull request implementing it.&lt;/p&gt;

&lt;p&gt;What followed was seven months of genuinely substantive back-and-forth involving lead maintainer Tabish Bidiwale (&lt;a href="https://github.com/TabishB" rel="noopener noreferrer"&gt;@TabishB&lt;/a&gt;) and several deeply engaged community members, covering immutable versus mutable records, whether an ADR belongs to the research phase or the design phase, cross-change ownership when one decision cascades into a dozen later changes, and how ADRs relate to specs as the "authoritative" description of the system. Tabish Bidiwale's early framing set the direction the thread ultimately settled on: OpenSpec should stay lightweight by default and make specialized workflows like ADRs configurable through its schema system rather than baking them into core. A community member later summarized where the discussion landed:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;ADR workflows are valuable, but OpenSpec does not currently have first-class/native ADR support... the direction discussed here is to keep the default workflow lightweight and make specialized workflows configurable.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Maintainer Clay Good (&lt;a href="https://github.com/clay-good" rel="noopener noreferrer"&gt;@clay-good&lt;/a&gt;) closed the issue on that basis in August 2026 and moved it into &lt;a href="https://github.com/Fission-AI/OpenSpec/discussions/1553" rel="noopener noreferrer"&gt;GitHub Discussion #1553&lt;/a&gt; so the conversation could keep evolving without staying open as an unresolved bug. That is a reasonable call for a tool whose entire pitch is avoiding Spec Kit-style ceremony by default. It also means the fix lives one layer up, in one of two places:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A community schema.&lt;/strong&gt; The &lt;code&gt;spec-driven-with-adr&lt;/code&gt; schema, built by OpenSpec technical advisor Hari Krishnan (&lt;a href="https://github.com/harikrishnan83" rel="noopener noreferrer"&gt;@harikrishnan83&lt;/a&gt;) and documented on &lt;a href="https://intent-driven.dev/blog/2026/04/29/spec-driven-development-with-adr/" rel="noopener noreferrer"&gt;intent-driven.dev&lt;/a&gt;, adds a fifth artifact to OpenSpec's default four-artifact pipeline. It exists because the default schema loses &lt;code&gt;design.md&lt;/code&gt;'s reasoning the moment a change is archived -- only the spec deltas get synced forward, so the "why" behind a decision disappears with the change unless something else preserves it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A repository-level convention.&lt;/strong&gt; A small, hand-rolled &lt;code&gt;decision.md&lt;/code&gt; file plus a naming rule, which costs nothing to adopt and does not require installing a custom schema.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The rest of this guide covers option two in depth, since it is the lower-friction starting point for most teams -- and, as the section on the community schema below shows, it is compatible with switching to that heavier tooling later if your rejection log grows large enough to earn it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The &lt;code&gt;decision.md&lt;/code&gt; Convention for Recording a Rejected Change
&lt;/h2&gt;

&lt;p&gt;Structure a rejected investigation the same way you would a shipped one, but stop before syncing any deltas, and add one file that states the outcome plainly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;openspec/
  changes/
    archive/
      2026-09-16-rejected-shared-persistence-layer/
        proposal.md
        decision.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;decision.md&lt;/code&gt; answers the same four questions a proper &lt;a href="https://www.glukhov.org/app-architecture/documentation/decision-records-ai-driven-development/" rel="noopener noreferrer"&gt;Architecture Decision Record&lt;/a&gt; does -- what was decided, why, what alternatives existed, and what would change the answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Decision&lt;/span&gt;

Status: Rejected

&lt;span class="gu"&gt;## Decision&lt;/span&gt;

Do not replace the service-to-service HTTP boundary with a direct
package import between the two Go services.

&lt;span class="gu"&gt;## Reasons&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Increases compile-time coupling between independently deployed services.
&lt;span class="p"&gt;-&lt;/span&gt; Makes the persistence layer an implicit, undocumented contract.
&lt;span class="p"&gt;-&lt;/span&gt; The measured benefit (latency, code duplication) was smaller than
  the coupling cost in this codebase.

&lt;span class="gu"&gt;## Alternatives considered&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Shared internal Go module -- rejected for the same coupling reason.
&lt;span class="p"&gt;-&lt;/span&gt; gRPC instead of HTTP -- deferred, not rejected; revisit if HTTP
  overhead becomes a measured bottleneck.

&lt;span class="gu"&gt;## Reconsider only if&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; The two services are intentionally merged into one deployable, or
&lt;span class="p"&gt;-&lt;/span&gt; Latency measurements show the HTTP hop is a proven bottleneck.

&lt;span class="gu"&gt;## Related&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Architecture rule: services communicate over HTTP, not shared packages.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The one hard rule that makes this whole convention work: &lt;strong&gt;do not run the sync step for a rejected change.&lt;/strong&gt; If &lt;code&gt;/opsx:propose&lt;/code&gt; already drafted delta specs before you decided against the change, use the flag the CLI already gives you for exactly this situation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openspec archive investigate-shared-persistence &lt;span class="nt"&gt;--skip-specs&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--skip-specs&lt;/code&gt; tells &lt;code&gt;openspec archive&lt;/code&gt; to file the change away without touching &lt;code&gt;openspec/specs/&lt;/code&gt; at all, which is the safest default for anything you are archiving without shipping. Accepting the ordinary sync prompt instead would merge the rejected idea's delta specs into your canonical specs, and canonical &lt;code&gt;openspec/specs/&lt;/code&gt; should describe what the system currently does, not every idea that was drafted and turned down. If a change permanently produces no spec changes for a structural reason -- a pure investigation folder, say -- OpenSpec also supports declaring &lt;code&gt;skip_specs: true&lt;/code&gt; in that change's &lt;code&gt;.openspec.yaml&lt;/code&gt; so it archives cleanly without the flag every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Naming Rejected Changes So Humans and Agents Can Scan the Archive
&lt;/h2&gt;

&lt;p&gt;A &lt;code&gt;decision.md&lt;/code&gt; file only helps if someone opens the folder. Prefix the folder name with the outcome so both a human skimming &lt;code&gt;ls openspec/changes/archive/&lt;/code&gt; and an agent listing changes can tell status without opening a single file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2026-09-16-rejected-shared-persistence-layer/
2026-09-20-abandoned-react-router-migration/
2026-10-01-superseded-old-auth-design/
2026-10-10-add-project-filtering/          # shipped, no prefix needed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This mirrors the status vocabulary already recommended for standalone decision records -- proposed, accepted, superseded, deprecated -- applied to OpenSpec's own archive instead of a separate &lt;code&gt;docs/decisions/&lt;/code&gt; folder. Keep the vocabulary small. Three or four consistent prefixes beat a free-text status line that every proposal spells slightly differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Make Your Agent Check the Archive Before Proposing Again
&lt;/h2&gt;

&lt;p&gt;Naming and a &lt;code&gt;decision.md&lt;/code&gt; file solve discoverability for a human skimming the folder. They do nothing on their own to make an agent search the archive before drafting a new proposal -- that has to be an explicit instruction, because &lt;code&gt;/opsx:propose&lt;/code&gt; does not do it by default, and no amount of tidy file naming changes that on its own.&lt;/p&gt;

&lt;p&gt;Two places to put that instruction, matching how OpenSpec already expects project-specific guidance to be injected:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In &lt;code&gt;openspec/config.yaml&lt;/code&gt;&lt;/strong&gt;, under the &lt;code&gt;context:&lt;/code&gt; field that gets injected into every planning request (mind the 50KB cap covered in the &lt;a href="https://www.glukhov.org/ai-devtools/openspec/" rel="noopener noreferrer"&gt;OpenSpec quickstart&lt;/a&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
  &lt;span class="s"&gt;Before proposing a change, search openspec/changes/archive for folders&lt;/span&gt;
  &lt;span class="s"&gt;prefixed "rejected-" or "abandoned-" that describe a materially similar&lt;/span&gt;
  &lt;span class="s"&gt;idea. If one exists, summarize its decision.md and state what has&lt;/span&gt;
  &lt;span class="s"&gt;changed before proposing the idea again. Do not re-litigate a rejected&lt;/span&gt;
  &lt;span class="s"&gt;decision without new evidence.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;In &lt;code&gt;AGENTS.md&lt;/code&gt; or your project's own agent instructions&lt;/strong&gt;, as a standing rule rather than a per-request context blob:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Rejected OpenSpec changes&lt;/span&gt;

When a proposal is investigated and rejected:
&lt;span class="p"&gt;
1.&lt;/span&gt; Do not sync or apply its delta specs.
&lt;span class="p"&gt;2.&lt;/span&gt; Add &lt;span class="sb"&gt;`decision.md`&lt;/span&gt; with Status, Decision, Reasons, Alternatives
   considered, and Reconsider only if.
&lt;span class="p"&gt;3.&lt;/span&gt; Prefix the archived folder name: &lt;span class="sb"&gt;`rejected-&amp;lt;name&amp;gt;`&lt;/span&gt; or &lt;span class="sb"&gt;`abandoned-&amp;lt;name&amp;gt;`&lt;/span&gt;.
&lt;span class="p"&gt;4.&lt;/span&gt; Before proposing a materially similar change, search
   &lt;span class="sb"&gt;`openspec/changes/archive/`&lt;/span&gt; and reference the prior decision.
&lt;span class="p"&gt;5.&lt;/span&gt; Do not reopen a rejected decision unless its documented
   reconsideration conditions have actually changed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
  A[New idea worth a change] --&amp;gt; B{Search openspec/changes/archive}
  B --&amp;gt;|Similar rejected decision found| C[Summarize prior decision.md]
  C --&amp;gt; D{Reconsideration conditions changed?}
  D --&amp;gt;|No| E[Do not re-propose. Reference the decision.]
  D --&amp;gt;|Yes| F["/opsx:propose with the changed context stated"]
  B --&amp;gt;|Nothing similar found| F&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Neither instruction guarantees compliance -- an agent can still skip the search step, the same way it can skip reading any other context you inject. But it is the difference between "the information exists somewhere in the repo" and "the agent is told, every time, to go look for it," and only the second one actually reduces repeated investigations in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Worked Example: Rejecting a Proposal, Then Correctly Reconsidering It
&lt;/h2&gt;

&lt;p&gt;Put the pieces together on a concrete case. Say a teammate asks an agent to look at replacing a service-to-service HTTP call with a direct Go package import, to shave off network latency.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Explore, then propose.&lt;/strong&gt; &lt;code&gt;/opsx:explore&lt;/code&gt; reads both services, and &lt;code&gt;/opsx:propose replace-http-with-direct-import&lt;/code&gt; drafts a proposal, a design doc weighing the latency win against the coupling cost, and a draft delta spec.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Investigate and reject.&lt;/strong&gt; After reviewing the design doc, the team decides the coupling cost -- two independently deployed services now sharing a compile-time dependency -- outweighs a latency win nobody has actually measured as a problem. Nothing gets built.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Archive without syncing.&lt;/strong&gt; Rather than deleting the folder, run &lt;code&gt;openspec archive replace-http-with-direct-import --skip-specs&lt;/code&gt;, then add &lt;code&gt;decision.md&lt;/code&gt; to the archived folder with &lt;code&gt;Status: Rejected&lt;/code&gt;, the reasons above, and a &lt;code&gt;Reconsider only if&lt;/code&gt; clause naming the condition that would change the answer -- for example, "latency measurements show the HTTP hop is a proven bottleneck." Rename the folder with a &lt;code&gt;rejected-&lt;/code&gt; prefix so it reads as &lt;code&gt;openspec/changes/archive/2026-09-16-rejected-replace-http-with-direct-import/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Months later, someone re-raises it.&lt;/strong&gt; A different contributor, or the same agent in a fresh session, gets asked to "speed up the checkout-to-inventory call" and starts drafting a proposal that looks a lot like the same idea. Because &lt;code&gt;openspec/config.yaml&lt;/code&gt; instructs the agent to search the archive first, it finds the rejected folder, reads &lt;code&gt;decision.md&lt;/code&gt;, and reports back: &lt;em&gt;"A materially similar change was proposed and rejected on 2026-09-16 for coupling reasons. The reconsideration condition was 'latency measurements show the HTTP hop is a proven bottleneck.' Do you have new measurements, or is this a different problem?"&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The team supplies new evidence.&lt;/strong&gt; If profiling now shows the HTTP hop genuinely dominates checkout latency, that is exactly the changed circumstance the original &lt;code&gt;decision.md&lt;/code&gt; asked for. The agent proceeds with &lt;code&gt;/opsx:propose&lt;/code&gt;, and the new proposal's &lt;code&gt;decision.md&lt;/code&gt; -- once this one is also archived, accepted or rejected -- references the prior one under &lt;code&gt;Related&lt;/code&gt;, so the archive reads as a continuous decision history rather than two unconnected folders that happen to describe the same idea.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That fifth step is the entire point of the convention. Without it, step 4 either does not happen at all -- the agent just re-investigates from zero -- or it happens by luck, because a human remembered the prior conversation. The &lt;code&gt;decision.md&lt;/code&gt; file and the archive-search instruction turn "someone might remember" into something the workflow actually checks.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenSpec's Archive vs. a Dedicated ADR Log: Who Owns What
&lt;/h2&gt;

&lt;p&gt;Once you are maintaining &lt;code&gt;decision.md&lt;/code&gt; files inside the archive, it is worth being explicit about which artifact answers which question, so the convention does not quietly turn into duplicate documentation:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Artifact&lt;/th&gt;
&lt;th&gt;Answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openspec/specs/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;What does the system currently do?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;openspec/changes/&amp;lt;name&amp;gt;/&lt;/code&gt; (active)&lt;/td&gt;
&lt;td&gt;What are we proposing to change, right now?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openspec/changes/archive/&amp;lt;name&amp;gt;/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;What changed (or was rejected) in the past, and why?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;docs/adr/&lt;/code&gt; (standalone, tool-neutral)&lt;/td&gt;
&lt;td&gt;What durable architectural rule did we learn, independent of any single change?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a decision narrow enough to belong to one investigation -- "we looked at sharing this persistence layer and said no" -- the &lt;code&gt;decision.md&lt;/code&gt;-inside-the-archive convention above is enough. For a decision that should outlive and constrain many future changes -- "services communicate over HTTP, never shared packages" -- promote it to a standalone &lt;a href="https://www.glukhov.org/app-architecture/documentation/decision-records-ai-driven-development/" rel="noopener noreferrer"&gt;Architecture Decision Record&lt;/a&gt; in &lt;code&gt;docs/adr/&lt;/code&gt;, and have the rejected change's &lt;code&gt;decision.md&lt;/code&gt; reference it under &lt;code&gt;Related&lt;/code&gt;. That split keeps OpenSpec's archive focused on individual investigations while the ADR log holds the small number of rules that should survive any single tool's lifecycle -- including a future migration off OpenSpec entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Adopt the &lt;code&gt;spec-driven-with-adr&lt;/code&gt; Schema Instead
&lt;/h2&gt;

&lt;p&gt;The hand-rolled convention above costs nothing and fits inside fifteen minutes of setup, which makes it the right default. But it is worth understanding what the more structured alternative actually does before you decide you have outgrown a naming prefix.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;spec-driven-with-adr&lt;/code&gt; inserts a fifth artifact, &lt;code&gt;adr&lt;/code&gt;, between &lt;code&gt;design&lt;/code&gt; and &lt;code&gt;tasks&lt;/code&gt; in OpenSpec's pipeline. Rather than writing ADR content directly into the change folder, the &lt;code&gt;adr&lt;/code&gt; step produces a short change-local &lt;code&gt;adr.md&lt;/code&gt; review manifest and, when the change introduces a genuinely durable architectural commitment, a numbered record at the repository root -- &lt;code&gt;/adr/0042-use-postgres-for-catalog.md&lt;/code&gt;, sibling to &lt;code&gt;openspec/&lt;/code&gt;, not nested inside it. Every ADR the schema creates is immutable once accepted: the schema's own instructions call this out as an "iron rule" -- you never edit an accepted record's status, body, or date. To change a previous decision, you write a &lt;em&gt;new&lt;/em&gt; ADR whose &lt;code&gt;Supersedes:&lt;/code&gt; field names the old one, and future designs walk that supersession chain to know which decisions are still in force. That is a more rigorous version of exactly the "reconsider only if" idea in the &lt;code&gt;decision.md&lt;/code&gt; convention above, enforced by the schema rather than left to a human remembering to write it down.&lt;/p&gt;

&lt;p&gt;It is worth being precise about what this schema does and does not solve. It is built for decisions that get &lt;em&gt;accepted&lt;/em&gt; and need to survive archiving -- Postgres over DynamoDB, JWT over session cookies -- not for proposals that were investigated and rejected without shipping anything. A rejected investigation still has nowhere obvious to live under this schema either; you would layer the same &lt;code&gt;decision.md&lt;/code&gt;-and-naming convention from this guide on top of it, just referencing &lt;code&gt;/adr/&lt;/code&gt; records instead of a standalone &lt;code&gt;docs/adr/&lt;/code&gt; folder.&lt;/p&gt;

&lt;p&gt;Reach for it once you notice any of these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your rejected-decision count is large enough that grepping &lt;code&gt;openspec/changes/archive/&lt;/code&gt; for prefixes stops being fast.&lt;/li&gt;
&lt;li&gt;You want durable architectural decisions validated and cross-referenced against every new design automatically, rather than by convention and &lt;code&gt;grep&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Multiple contributors keep inventing slightly different &lt;code&gt;decision.md&lt;/code&gt; shapes, and you want a schema to enforce one immutable, numbered format instead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Installing a custom schema is a bigger commitment than a naming convention -- it changes what &lt;code&gt;/opsx:propose&lt;/code&gt; generates for every future change, not just rejected ones -- so treat it as a step up once the lightweight version is visibly straining, not a default first move.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;OpenSpec's archive was designed around one outcome -- a change that shipped -- and its own maintainers have been explicit, after a seven-month public discussion, that first-class rejection or ADR support is not coming to the core workflow soon. That leaves the fix where OpenSpec already puts most of its team conventions: in your repository, not in the tool. A &lt;code&gt;decision.md&lt;/code&gt; file, the &lt;code&gt;--skip-specs&lt;/code&gt; flag on archive, a &lt;code&gt;rejected-&lt;/code&gt;/&lt;code&gt;abandoned-&lt;/code&gt; naming prefix, and an explicit instruction telling the agent to search the archive before proposing are enough to stop most repeated investigations. Reach for the &lt;code&gt;spec-driven-with-adr&lt;/code&gt; schema only once that lightweight convention is genuinely straining under the number of decisions you are tracking -- and even then, keep the distinction clear: it manages decisions you accepted and want to survive archiving, not the ones you turned down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Useful Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/ai-devtools/openspec/" rel="noopener noreferrer"&gt;OpenSpec Quickstart: Install, Workflow, and Common Pitfalls&lt;/a&gt; -- install, the explore-propose-apply-archive loop, and everyday pitfalls&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/app-architecture/documentation/decision-records-ai-driven-development/" rel="noopener noreferrer"&gt;Decision Records for AI-Driven Software Development&lt;/a&gt; -- the general ADR/PDR/DDR format, status lifecycle, and AI-reading instructions this convention borrows from&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/Fission-AI/OpenSpec/issues/557" rel="noopener noreferrer"&gt;GitHub issue #557: architecture decision records support&lt;/a&gt; -- the full seven-month discussion of why ADRs are not core to OpenSpec&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/Fission-AI/OpenSpec/discussions/1553" rel="noopener noreferrer"&gt;GitHub Discussion #1553&lt;/a&gt; -- where that conversation continues after the issue was closed&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://intent-driven.dev/blog/2026/04/29/spec-driven-development-with-adr/" rel="noopener noreferrer"&gt;spec-driven-with-adr schema&lt;/a&gt; -- the community schema that keeps ADRs alive outside the change lifecycle, by OpenSpec advisor Hari Krishnan&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/Fission-AI/OpenSpec/blob/main/docs/team-workflow.md" rel="noopener noreferrer"&gt;OpenSpec team-workflow docs&lt;/a&gt; -- how archiving, branches, and PR review are meant to fit together&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/ai-devtools/ai-coding-assistants/spec-kit-vs-kiro-vs-claude-code/" rel="noopener noreferrer"&gt;GitHub Spec Kit vs Kiro vs Claude Code SDD Workflows&lt;/a&gt; -- how OpenSpec compares to heavier SDD tooling overall&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aicoding</category>
      <category>llm</category>
      <category>ai</category>
      <category>dev</category>
    </item>
    <item>
      <title>OpenSpec Quickstart: Install, Workflow, and Common Pitfalls</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:20:00 +0000</pubDate>
      <link>https://dev.to/rosgluk/openspec-quickstart-install-workflow-and-common-pitfalls-4m5b</link>
      <guid>https://dev.to/rosgluk/openspec-quickstart-install-workflow-and-common-pitfalls-4m5b</guid>
      <description>&lt;p&gt;&lt;a href="https://github.com/Fission-AI/OpenSpec" rel="noopener noreferrer"&gt;OpenSpec&lt;/a&gt; is a free, open-source CLI from Fission AI that gets you and your coding agent to agree on a change in plain Markdown before any code gets written, without the phase-gated ceremony of heavier spec-driven frameworks.&lt;/p&gt;

&lt;p&gt;Most teams who try Spec-Driven Development stall on the same trade-off: enough process to stop an agent from guessing, without so much scaffolding that a fifty-line bugfix needs a proposal document. OpenSpec's answer is to skip the "document the whole system first" instinct entirely and write specs only for what a change actually touches, using &lt;code&gt;ADDED&lt;/code&gt;, &lt;code&gt;MODIFIED&lt;/code&gt;, and &lt;code&gt;REMOVED&lt;/code&gt; deltas instead of a full rewrite every time.&lt;/p&gt;

&lt;p&gt;That change-centric design is also why OpenSpec keeps coming up next to GitHub Spec Kit, Kiro, and Superpowers in the &lt;a href="https://www.glukhov.org/ai-devtools/ai-coding-assistants/spec-kit-vs-kiro-vs-claude-code/" rel="noopener noreferrer"&gt;comparison of SDD tool categories&lt;/a&gt; -- it is usually the pick when a team wants reviewable specs without an 800-line planning phase. This guide covers installing the CLI, the four-command workflow you actually use day to day, what a change looks like on disk, and the questions and complaints that show up most often on Reddit and in OpenSpec's own issue tracker.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is OpenSpec?
&lt;/h2&gt;

&lt;p&gt;OpenSpec describes its own philosophy in four lines: fluid not rigid, iterative not waterfall, easy not complex, built for brownfield not just greenfield. In practice that means there are no locked phases -- you can edit a proposal, a spec, or a task list at any point in a change, rather than being forced through specify-then-plan-then-implement in strict order the way &lt;a href="https://www.glukhov.org/app-architecture/documentation/spec-driven-development-workflow/" rel="noopener noreferrer"&gt;the tool-neutral SDD workflow&lt;/a&gt; describes it.&lt;/p&gt;

&lt;p&gt;A change in OpenSpec produces up to four Markdown artifacts in its own folder:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Artifact&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;proposal.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Why the change exists and what it changes, in plain language&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;specs/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Delta requirements and scenarios -- the testable spec for this change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;design.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Optional technical approach, for changes that need one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tasks.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The implementation checklist the agent works through&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Once a change is implemented and archived, its delta specs merge into &lt;code&gt;openspec/specs/&lt;/code&gt;, which becomes the durable, current-state description of your system -- the same "spec as source of truth" idea covered in &lt;a href="https://www.glukhov.org/app-architecture/documentation/what-is-spec-driven-development/" rel="noopener noreferrer"&gt;What Is Spec-Driven Development?&lt;/a&gt;, just scoped one change at a time instead of written all at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installing OpenSpec
&lt;/h2&gt;

&lt;p&gt;OpenSpec is a Node.js CLI, so you need Node 20.19.0 or newer on your machine.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;node &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Install the CLI globally with npm, then verify it landed on your &lt;code&gt;PATH&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @fission-ai/openspec@latest
openspec &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deno, pnpm, yarn, bun, and nix are also supported install paths if that fits your setup better than npm. Once installed, initialize it inside a project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;your-project
openspec init
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;openspec init&lt;/code&gt; asks which AI tools you use and writes the matching skill and command files -- OpenSpec supports 30+ assistants, including Claude Code, Cursor, GitHub Copilot, Gemini CLI, Codex, Kiro, and OpenCode. For CI or scripted setup, skip the picker entirely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openspec init &lt;span class="nt"&gt;--tools&lt;/span&gt; claude,cursor   &lt;span class="c"&gt;# set up specific tools&lt;/span&gt;
openspec init &lt;span class="nt"&gt;--tools&lt;/span&gt; all             &lt;span class="c"&gt;# every supported tool&lt;/span&gt;
openspec init &lt;span class="nt"&gt;--tools&lt;/span&gt; none            &lt;span class="c"&gt;# openspec/ structure only, no tool files&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restart your IDE afterward so it picks up the newly written skills and commands. If you would rather have your assistant do the whole install for you, OpenSpec ships a setup prompt you can paste into &lt;a href="https://www.glukhov.org/ai-devtools/claude-code/" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt; or another agent, which runs the install, executes &lt;code&gt;openspec init&lt;/code&gt;, and reports back what it configured.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Workflow: Explore, Propose, Apply, Archive
&lt;/h2&gt;

&lt;p&gt;This is the one thing that trips up almost everyone on their first day: &lt;code&gt;openspec&lt;/code&gt; commands run in your terminal, but &lt;code&gt;/opsx:&lt;/code&gt; commands run in your AI assistant's chat window. There is no separate "interactive mode" to enter -- typing the slash command in chat is how you start.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
  A["/opsx:explore (optional)"] --&amp;gt; B["/opsx:propose change-name"]
  B --&amp;gt; C["/opsx:apply"]
  C --&amp;gt; D["/opsx:archive"]
  D --&amp;gt;|specs merged| E["openspec/specs/"]&lt;/code&gt;&lt;/pre&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/opsx:explore&lt;/code&gt;&lt;/strong&gt; is a no-stakes thinking partner. It reads the relevant part of your codebase, lays out options, and shapes a plan before anything is written to disk -- worth forming as a habit specifically because it stops an eager agent from confidently building the wrong thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/opsx:propose &amp;lt;name&amp;gt;&lt;/code&gt;&lt;/strong&gt; creates &lt;code&gt;openspec/changes/&amp;lt;name&amp;gt;/&lt;/code&gt; and drafts the proposal, delta specs, optional design, and task list in one step. You review the plan here, before implementation starts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/opsx:apply&lt;/code&gt;&lt;/strong&gt; works through the task list, checking items off as it goes. Because progress lives in files rather than only in chat history, you can clear your context window or start a fresh session and pick up exactly where &lt;code&gt;/opsx:apply&lt;/code&gt; left off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/opsx:archive&lt;/code&gt;&lt;/strong&gt; files the completed change to &lt;code&gt;openspec/changes/archive/YYYY-MM-DD-&amp;lt;name&amp;gt;/&lt;/code&gt; and merges its delta specs into the canonical &lt;code&gt;openspec/specs/&lt;/code&gt; tree.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The default &lt;code&gt;core&lt;/code&gt; profile installs exactly those four commands plus &lt;code&gt;update&lt;/code&gt; and &lt;code&gt;sync&lt;/code&gt;. An expanded profile adds &lt;code&gt;new&lt;/code&gt;, &lt;code&gt;continue&lt;/code&gt;, &lt;code&gt;ff&lt;/code&gt;, &lt;code&gt;verify&lt;/code&gt;, &lt;code&gt;bulk-archive&lt;/code&gt;, and &lt;code&gt;onboard&lt;/code&gt; for teams who want to create one artifact at a time instead of all at once -- switch to it with &lt;code&gt;openspec config profile&lt;/code&gt; followed by &lt;code&gt;openspec update&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Each tool spells the same command differently depending on how it loads custom instructions: &lt;code&gt;/opsx:propose&lt;/code&gt; in Claude Code, &lt;code&gt;/opsx-propose&lt;/code&gt; in Cursor and GitHub Copilot, &lt;code&gt;@opsx-propose&lt;/code&gt; in Amazon Q, or &lt;code&gt;$openspec-propose&lt;/code&gt; in Codex. &lt;code&gt;openspec init&lt;/code&gt; prints the exact form for the tools you picked, so the fastest fix for "nothing happened when I typed the command" is usually to re-read that printed hint rather than guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Change Looks Like on Disk
&lt;/h2&gt;

&lt;p&gt;A change folder under &lt;code&gt;openspec/changes/add-dark-mode/&lt;/code&gt; typically contains a proposal, a delta spec, and a task list like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## ADDED Requirements&lt;/span&gt;

&lt;span class="gu"&gt;### Requirement: Theme selection&lt;/span&gt;
The app SHALL let users switch between light and dark themes,
defaulting to the system preference.

&lt;span class="gu"&gt;#### Scenario: User toggles dark mode&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**WHEN**&lt;/span&gt; the user clicks the theme toggle
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="gs"&gt;**THEN**&lt;/span&gt; the app switches to dark mode and persists the choice
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;ADDED&lt;/code&gt;/&lt;code&gt;MODIFIED&lt;/code&gt;/&lt;code&gt;REMOVED&lt;/code&gt; delta format is the mechanism that lets OpenSpec avoid rewriting an entire spec file for a one-field change. It is also why OpenSpec is explicitly brownfield-first rather than greenfield-first: you never document your whole application before getting value, you just document the slice each real change touches, and &lt;code&gt;openspec/specs/&lt;/code&gt; fills in naturally over months of normal work.&lt;/p&gt;

&lt;p&gt;Useful CLI commands for checking on that state without leaving the terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openspec list                 &lt;span class="c"&gt;# active changes&lt;/span&gt;
openspec show add-dark-mode   &lt;span class="c"&gt;# view a change's artifacts&lt;/span&gt;
openspec validate &lt;span class="nt"&gt;--all&lt;/span&gt;       &lt;span class="c"&gt;# check spec formatting across the project&lt;/span&gt;
openspec view                 &lt;span class="c"&gt;# interactive dashboard&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Commit the whole &lt;code&gt;openspec/&lt;/code&gt; folder to git. The active changes and the archive are meant to become a durable, versioned record of what your system does and why it changed -- not a scratch pad you delete after merging.&lt;/p&gt;

&lt;h2&gt;
  
  
  Adopting OpenSpec on an Existing Codebase
&lt;/h2&gt;

&lt;p&gt;The most common worry from teams evaluating OpenSpec on a real project is some version of "my app is 80,000 lines old, do I have to spec all of it first?" You do not. OpenSpec's own guidance is blunt about this: pick something small and real that you were already going to build this week, run &lt;code&gt;/opsx:explore&lt;/code&gt; on the area you are about to touch so the agent maps how things actually work first, then &lt;code&gt;/opsx:propose&lt;/code&gt; a change scoped to just that slice.&lt;/p&gt;

&lt;p&gt;If you already have PRDs, SRS documents, or design docs sitting in Notion or Confluence, treat them as source material for exploration rather than something to bulk-convert into specs. Paste the relevant section into an &lt;code&gt;/opsx:explore&lt;/code&gt; session and let the agent shape a focused delta from it; a one-time mechanical conversion of a forty-page PRD tends to produce a spec nobody trusts six months later. For teams that want a guided, narrated first run instead of jumping straight into a real change, the expanded &lt;code&gt;/opsx:onboard&lt;/code&gt; command scans your codebase for a small, safe improvement and walks through the full loop on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Questions and Problems
&lt;/h2&gt;

&lt;p&gt;These are the issues that show up repeatedly across OpenSpec's Discord, GitHub issues, and Reddit threads in subreddits like r/cursor, r/RooCode, and r/opencodeCLI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"I typed the slash command and nothing happened."&lt;/strong&gt; Almost always one of: you typed it in the terminal instead of your assistant's chat, your IDE has not restarted since &lt;code&gt;openspec init&lt;/code&gt; ran, or the CLI version is old enough that &lt;code&gt;openspec update&lt;/code&gt; reports everything current without ever writing the newer workflow files. Run &lt;code&gt;openspec update&lt;/code&gt;, restart the IDE, and confirm the skill folders exist (&lt;code&gt;.claude/skills/openspec-*&lt;/code&gt; for Claude Code, or your tool's equivalent from the supported-tools list).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"The AI generates way more spec than I need."&lt;/strong&gt; This is the most-cited complaint in longer write-ups: an agent can turn a thirty-minute feature into an 800-line spec. OpenSpec caps the &lt;code&gt;context:&lt;/code&gt; field injected into every request at 50KB specifically to force discipline, but the delta specs themselves have no hard limit, so trimming generated specs down to what is actually load-bearing is a habit you have to maintain yourself, not something the tool enforces for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Two changes touched the same requirement and one silently dropped the other's scenario."&lt;/strong&gt; This is a real, documented edge case: archiving applies a &lt;code&gt;MODIFIED&lt;/code&gt; delta as a whole-block replace keyed by requirement name, so if two in-flight changes both modify the same requirement, archiving the second one used to overwrite the first's scenarios without warning. Current versions add a drift check that aborts the archive and tells you to refresh the change's spec first -- but it is still worth knowing the failure mode exists if you run several changes on the same area in parallel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Which AI model should I actually use with it?"&lt;/strong&gt; OpenSpec's own docs recommend high-reasoning models for both planning and implementation -- Opus-class and Codex-class models are called out specifically -- and clearing your context window before implementation, since a clean context produces measurably better results than a long, accumulated session.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"How is this different from Spec Kit, Kiro, Superpowers, or BMAD?"&lt;/strong&gt; This is the single most frequent Reddit question, and the honest answer is "process weight." OpenSpec's own README frames the comparison directly: Spec Kit is thorough but heavier, with more Markdown and rigid phase gates; Kiro is powerful but locks you into AWS's IDE and Claude models; OpenSpec trades some of that upfront structure for the ability to iterate freely and work with whatever assistant you already have open. For the full breakdown against Spec Kit, Kiro, Claude Code skills, BMAD-METHOD, and &lt;a href="https://www.glukhov.org/ai-devtools/superpowers/" rel="noopener noreferrer"&gt;Superpowers&lt;/a&gt;, see the dedicated &lt;a href="https://www.glukhov.org/ai-devtools/ai-coding-assistants/spec-kit-vs-kiro-vs-claude-code/" rel="noopener noreferrer"&gt;SDD tool comparison&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Does the AI actually follow the spec it just wrote?"&lt;/strong&gt; Not always, and this is a documented problem across SDD tools generally, not unique to OpenSpec -- a large context window does not mean the agent attends equally to every part of it. The &lt;code&gt;/opsx:verify&lt;/code&gt; command exists specifically to catch generated code that contradicts its own spec, and it is worth running on anything non-trivial rather than trusting the implementation blindly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Do I need this for a one-line fix?"&lt;/strong&gt; No. OpenSpec's own FAQ says as much: use it where agreement matters, which is most non-trivial, multi-file work, and skip it for a typo fix or a throwaway prototype you will delete in a week.&lt;/p&gt;

&lt;h2&gt;
  
  
  When OpenSpec Fits and When It Doesn't
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Good fit:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Brownfield codebases where you want reviewable specs without documenting the entire system upfront.&lt;/li&gt;
&lt;li&gt;Solo developers and small teams who want lighter ceremony than Spec Kit while still getting a written plan before code.&lt;/li&gt;
&lt;li&gt;Work that spans several files, a schema change, or anything a junior engineer would reasonably want a short design doc for.&lt;/li&gt;
&lt;li&gt;Teams already committed to reviewing plans in pull requests -- delta specs diff cleanly since they only describe what changed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Weaker fit:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One-line bug fixes and throwaway prototypes, where the proposal-review step costs more than it saves.&lt;/li&gt;
&lt;li&gt;Teams that need the heavier, more prescriptive structure of Spec Kit or an AWS-native, IDE-integrated experience like Kiro -- see the &lt;a href="https://www.glukhov.org/ai-devtools/ai-coding-assistants/spec-kit-vs-kiro-vs-claude-code/" rel="noopener noreferrer"&gt;decision framework in the tool comparison&lt;/a&gt; for where each tool wins.&lt;/li&gt;
&lt;li&gt;Cross-repo features today, unless you are willing to try OpenSpec's beta &lt;strong&gt;stores&lt;/strong&gt; feature, which moves planning into its own shared repository so multiple codebases and agents can read the same plan.&lt;/li&gt;
&lt;li&gt;Anyone still deciding whether a given feature deserves a spec at all -- read &lt;a href="https://www.glukhov.org/ai-devtools/vibe-coding/spec-driven-development-vs-vibe-coding/" rel="noopener noreferrer"&gt;Spec-Driven Development vs Vibe Coding&lt;/a&gt; first, since OpenSpec only helps once you have already decided structure is worth the overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;OpenSpec's bet is that most Spec-Driven Development pain comes from ceremony, not from the underlying idea of agreeing on a plan before code exists. Deltas instead of full rewrites, no locked phases, and a brownfield-first workflow make it noticeably lighter than Spec Kit or Kiro to adopt on a codebase you did not build from scratch. The trade-offs are real too -- spec bloat is a genuine risk without discipline, the conflict handling around simultaneous changes to one requirement is still maturing, and the ecosystem is younger than GitHub's own tooling. Install it on one real project, run a small change through explore-propose-apply-archive end to end, and decide from there whether the lighter ceremony earns its keep against your actual workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Useful Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/Fission-AI/OpenSpec" rel="noopener noreferrer"&gt;OpenSpec repository&lt;/a&gt; -- source, docs, and the CLI package&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/Fission-AI/OpenSpec/blob/main/docs/README.md" rel="noopener noreferrer"&gt;OpenSpec documentation home&lt;/a&gt; -- getting started, concepts, FAQ, and troubleshooting&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/ai-devtools/ai-coding-assistants/spec-kit-vs-kiro-vs-claude-code/" rel="noopener noreferrer"&gt;GitHub Spec Kit vs Kiro vs Claude Code SDD Workflows&lt;/a&gt; -- full tool comparison and decision framework, including OpenSpec&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/ai-devtools/superpowers/" rel="noopener noreferrer"&gt;Superpowers Quickstart: Install, Workflow, and Tryout&lt;/a&gt; -- the enforced-skills alternative to OpenSpec's lighter-touch workflow&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/app-architecture/documentation/spec-driven-development-workflow/" rel="noopener noreferrer"&gt;Spec-Driven Development Workflow From Requirements to Code&lt;/a&gt; -- the tool-neutral five-phase process OpenSpec implements more fluidly&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/app-architecture/documentation/what-is-spec-driven-development/" rel="noopener noreferrer"&gt;What Is Spec-Driven Development? The Spec as Source of Truth&lt;/a&gt; -- core SDD concepts and terminology&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.glukhov.org/ai-devtools/vibe-coding/spec-driven-development-vs-vibe-coding/" rel="noopener noreferrer"&gt;Spec-Driven Development vs Vibe Coding: Waterfall?&lt;/a&gt; -- deciding whether a feature deserves a spec at all&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aicoding</category>
      <category>llm</category>
      <category>ai</category>
      <category>dev</category>
    </item>
    <item>
      <title>How to Migrate from OpenClaw to Hermes Agent Safely</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Tue, 15 Sep 2026 11:32:34 +0000</pubDate>
      <link>https://dev.to/rosgluk/how-to-migrate-from-openclaw-to-hermes-agent-safely-5g1l</link>
      <guid>https://dev.to/rosgluk/how-to-migrate-from-openclaw-to-hermes-agent-safely-5g1l</guid>
      <description>&lt;p&gt;Migrating an AI assistant is not the same as copying an application config. The hard part is preserving identity, memory, tool behavior, scheduled work, and messaging access without two gateways acting as the same bot.&lt;/p&gt;

&lt;p&gt;Hermes Agent now includes &lt;code&gt;hermes claw migrate&lt;/code&gt;, a real migration planner rather than a cosmetic import command. It can map more than 30 categories from OpenClaw, detect conflicts, create a Hermes restore point, and archive incompatible state for manual review. That makes the move practical, but it does not make the move automatic.&lt;/p&gt;

&lt;p&gt;The approach below is a staged cutover: back up OpenClaw, dry-run the full migration, import without secrets, validate Hermes from the terminal, and transfer messaging credentials only after the new agent behaves correctly. Do not begin with &lt;code&gt;--overwrite --migrate-secrets --yes&lt;/code&gt;; those flags are useful for automation after a rehearsed migration, not for discovering what your assistant actually depends on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The OpenClaw to Hermes migration runbook
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Phase&lt;/th&gt;
&lt;th&gt;Command or action&lt;/th&gt;
&lt;th&gt;Exit condition&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Inventory&lt;/td&gt;
&lt;td&gt;Record versions, workspaces, plugins, channels, cron jobs, and providers&lt;/td&gt;
&lt;td&gt;Every non-file dependency has an owner&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backup&lt;/td&gt;
&lt;td&gt;&lt;code&gt;openclaw backup create --verify&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A verified archive exists outside OpenClaw state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Preview&lt;/td&gt;
&lt;td&gt;&lt;code&gt;hermes claw migrate --dry-run --preset full&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No unexplained conflicts or skipped critical data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Import&lt;/td&gt;
&lt;td&gt;Run the full preset without secrets&lt;/td&gt;
&lt;td&gt;Hermes config, persona, memory, skills, and MCP entries exist&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local test&lt;/td&gt;
&lt;td&gt;Run Hermes in the terminal&lt;/td&gt;
&lt;td&gt;Model, tools, memory, approvals, and workspace pass tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Channel cutover&lt;/td&gt;
&lt;td&gt;Stop OpenClaw, migrate or set secrets, start Hermes gateway&lt;/td&gt;
&lt;td&gt;Only Hermes owns each bot token or account&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Soak&lt;/td&gt;
&lt;td&gt;Keep OpenClaw stopped but recoverable&lt;/td&gt;
&lt;td&gt;Scheduled and inbound work behaves correctly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cleanup&lt;/td&gt;
&lt;td&gt;Archive old OpenClaw state only after acceptance&lt;/td&gt;
&lt;td&gt;Rollback window is closed intentionally&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Inventory the live system] --&amp;gt; B[Verified OpenClaw backup]
    B --&amp;gt; C["Dry-run: hermes claw migrate --dry-run --preset full"]
    C --&amp;gt; D[Import without secrets]
    D --&amp;gt; E["Terminal validation in a new session"]
    E --&amp;gt; F["Controlled channel cutover"]
    F --&amp;gt; G[Soak with OpenClaw stopped]
    G --&amp;gt; H["Cleanup after acceptance"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The command is short because the judgment has moved into the preview and verification stages. Treat the generated migration report as a change plan, not as reassuring console output.&lt;/p&gt;

&lt;h2&gt;
  
  
  What &lt;code&gt;hermes claw migrate&lt;/code&gt; actually reads
&lt;/h2&gt;

&lt;p&gt;The migrator reads &lt;code&gt;~/.openclaw/&lt;/code&gt; by default. It also detects the older &lt;code&gt;~/.clawdbot/&lt;/code&gt; and &lt;code&gt;~/.moltbot/&lt;/code&gt; directories, along with legacy config filenames, so an older installation does not need to be renamed before migration.&lt;/p&gt;

&lt;p&gt;OpenClaw has used several workspace layouts. Hermes checks &lt;code&gt;workspace/&lt;/code&gt;, &lt;code&gt;workspace.default/&lt;/code&gt;, and &lt;code&gt;workspace-main/&lt;/code&gt;, and it recognizes per-agent directories such as &lt;code&gt;workspace-&amp;lt;agentId&amp;gt;&lt;/code&gt;. If you use custom agent roots or multiple profiles, verify every resolved path in the preview rather than assuming the default workspace represents the whole system.&lt;/p&gt;

&lt;p&gt;The destination is normally &lt;code&gt;~/.hermes/&lt;/code&gt;. A pre-existing Hermes installation is not treated as an empty bucket: the planner reports conflicts and refuses to apply by default when it cannot preserve both sides safely.&lt;/p&gt;

&lt;h2&gt;
  
  
  What migrates and what does not
&lt;/h2&gt;

&lt;p&gt;The useful distinction is not "supported" versus "unsupported." Some OpenClaw state maps directly, some must be transformed, and some can only be archived because the two agents use different execution models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Direct or transformed migration
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;OpenClaw source&lt;/th&gt;
&lt;th&gt;Hermes destination&lt;/th&gt;
&lt;th&gt;Migration behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;workspace/SOUL.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;~/.hermes/SOUL.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Direct persona copy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;workspace/MEMORY.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;~/.hermes/memories/MEMORY.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Parsed, merged, and deduplicated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;workspace/USER.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;~/.hermes/memories/USER.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Parsed, merged, and deduplicated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;workspace/memory/*.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Main Hermes memory&lt;/td&gt;
&lt;td&gt;Daily files are merged into entries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;workspace/AGENTS.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Chosen project directory&lt;/td&gt;
&lt;td&gt;Requires &lt;code&gt;--workspace-target&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenClaw skill directories&lt;/td&gt;
&lt;td&gt;&lt;code&gt;~/.hermes/skills/openclaw-imports/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Copied with an explicit conflict policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;agents.defaults.model&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hermes model configuration&lt;/td&gt;
&lt;td&gt;Primary and fallback forms are interpreted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;models.providers.*&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hermes provider configuration&lt;/td&gt;
&lt;td&gt;Base URL and API type are mapped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mcp.servers.*&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;mcp_servers.*&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Stdio and HTTP/SSE definitions are mapped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Channel tokens and allowlists&lt;/td&gt;
&lt;td&gt;Hermes &lt;code&gt;.env&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Only with &lt;code&gt;--migrate-secrets&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session reset policy&lt;/td&gt;
&lt;td&gt;&lt;code&gt;session_reset&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Daily and idle modes are translated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exec approvals&lt;/td&gt;
&lt;td&gt;Hermes approvals and command allowlist&lt;/td&gt;
&lt;td&gt;Modes and patterns are transformed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser, TTS, sandbox, and timeout settings&lt;/td&gt;
&lt;td&gt;Related Hermes config&lt;/td&gt;
&lt;td&gt;Supported fields are mapped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Memory is not copied as one opaque document. The migrator parses OpenClaw memory and user-profile files, merges them with existing Hermes entries, and deduplicates them. This is safer than replacing an established Hermes memory file, but it also means you should compare meaning and structure, not only file sizes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Archived for manual reconstruction
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;OpenClaw feature&lt;/th&gt;
&lt;th&gt;Why it is not directly portable&lt;/th&gt;
&lt;th&gt;Hermes direction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cron jobs&lt;/td&gt;
&lt;td&gt;Schedulers and delivery models differ&lt;/td&gt;
&lt;td&gt;Recreate with &lt;code&gt;hermes cron create&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plugins&lt;/td&gt;
&lt;td&gt;Plugin APIs are product-specific&lt;/td&gt;
&lt;td&gt;Replace with a Hermes plugin, skill, MCP server, or built-in tool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hooks and webhooks&lt;/td&gt;
&lt;td&gt;Event and permission contracts differ&lt;/td&gt;
&lt;td&gt;Recreate with Hermes webhooks or gateway hooks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Advanced memory backend&lt;/td&gt;
&lt;td&gt;Databases and recall semantics differ&lt;/td&gt;
&lt;td&gt;Configure a Hermes memory provider separately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills registry settings&lt;/td&gt;
&lt;td&gt;Registry implementation differs&lt;/td&gt;
&lt;td&gt;Configure with &lt;code&gt;hermes skills config&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-agent list and bindings&lt;/td&gt;
&lt;td&gt;Routing and profile models differ&lt;/td&gt;
&lt;td&gt;Rebuild with Hermes profiles and gateway configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IDENTITY.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hermes uses a different identity split&lt;/td&gt;
&lt;td&gt;Merge relevant identity into &lt;code&gt;SOUL.md&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;HEARTBEAT.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No direct file-driven heartbeat equivalent&lt;/td&gt;
&lt;td&gt;Express periodic work as cron jobs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;TOOLS.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hermes supplies its own tool instructions&lt;/td&gt;
&lt;td&gt;Move only genuine workflow rules into a skill or context file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;BOOTSTRAP.md&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Bootstrap semantics differ&lt;/td&gt;
&lt;td&gt;Use context files, setup, or a skill&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These items are saved below &lt;code&gt;~/.hermes/migration/openclaw/&amp;lt;timestamp&amp;gt;/archive/&lt;/code&gt;. A successful migration with a non-empty archive is therefore not finished; the archive is the remaining work queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Inventory the live OpenClaw system
&lt;/h2&gt;

&lt;p&gt;Before installing anything, write down which behaviors are actually in use. Config files alone may not reveal a plugin's external database, a manually supervised gateway, a custom agent directory, a local model process, or the account that owns a webhook endpoint.&lt;/p&gt;

&lt;p&gt;At minimum, record:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenClaw and Hermes versions.&lt;/li&gt;
&lt;li&gt;The active OpenClaw state directory and config path.&lt;/li&gt;
&lt;li&gt;All agent and workspace directories.&lt;/li&gt;
&lt;li&gt;Model providers, fallback models, and local endpoints.&lt;/li&gt;
&lt;li&gt;Installed and enabled plugins, including their persistent data.&lt;/li&gt;
&lt;li&gt;Skills from workspace, managed, personal, and project directories.&lt;/li&gt;
&lt;li&gt;MCP servers, environment variables, working directories, and credentials.&lt;/li&gt;
&lt;li&gt;Telegram, Discord, Slack, WhatsApp, Signal, Matrix, and Mattermost accounts.&lt;/li&gt;
&lt;li&gt;Cron jobs, hooks, webhooks, heartbeat behavior, and external supervisors.&lt;/li&gt;
&lt;li&gt;Approval rules, command allowlists, sandbox backend, and browser access.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This inventory becomes the acceptance checklist later. Without it, a migrated assistant can look healthy because it answers messages while silently missing the weekly backup, a memory provider, or a restrictive approval rule.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Create a verified OpenClaw backup
&lt;/h2&gt;

&lt;p&gt;OpenClaw 2.0 includes a backup command that understands its current SQLite state, configured agent roots, credentials, plugins, and workspaces. Use it instead of copying live database files and hoping their WAL sidecars were captured consistently.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; ~/Backups
openclaw gateway stop
openclaw backup create &lt;span class="nt"&gt;--output&lt;/span&gt; ~/Backups &lt;span class="nt"&gt;--verify&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the resulting archive outside &lt;code&gt;~/.openclaw/&lt;/code&gt;. The &lt;code&gt;--verify&lt;/code&gt; option validates the archive immediately, including path safety and supported SQLite integrity checks. OpenClaw-owned databases are captured through SQLite's online backup API, owner-verified, and compacted, rather than copied as raw files. If your workspaces are large, you can use &lt;code&gt;--no-include-workspace&lt;/code&gt;, but then back up those repositories and non-Git files separately; agent directories stay included either way.&lt;/p&gt;

&lt;h3&gt;
  
  
  The pre-2.0 transcript trap
&lt;/h3&gt;

&lt;p&gt;OpenClaw 2.0 moved sessions and transcripts out of &lt;code&gt;sessions.json&lt;/code&gt; and JSONL files into SQLite, by default at &lt;code&gt;~/.openclaw/agents/&amp;lt;agent&amp;gt;/agent/openclaw-agent.sqlite&lt;/code&gt;. That matters here for one non-obvious reason: the portable &lt;code&gt;backup create&lt;/code&gt; archive &lt;strong&gt;omits legacy JSONL transcripts and logs&lt;/strong&gt; even when they are no longer being written.&lt;/p&gt;

&lt;p&gt;So if your OpenClaw install predates 2.0 and you care about the old conversation history, a verified archive alone does not protect it. Stop the gateway and take a filesystem, volume, or VM snapshot before you migrate, or use OpenClaw's per-database snapshot commands for the databases you want a compact, independently verifiable copy of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw backup sqlite create &lt;span class="nt"&gt;--global&lt;/span&gt; &lt;span class="nt"&gt;--repository&lt;/span&gt; ~/Backups/openclaw-sqlite
openclaw backup sqlite create &lt;span class="nt"&gt;--agent&lt;/span&gt; main &lt;span class="nt"&gt;--repository&lt;/span&gt; ~/Backups/openclaw-sqlite
openclaw backup sqlite list &lt;span class="nt"&gt;--repository&lt;/span&gt; ~/Backups/openclaw-sqlite
openclaw backup sqlite verify ~/Backups/openclaw-sqlite/&amp;lt;snapshot-id&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Treat those snapshot repositories with the same permissions and retention policy as live state — they can contain auth profiles, session state, and plugin data. For a continuously replicated setup rather than periodic archives, OpenClaw documents Litestream against the same databases; that is a better answer than hand-rolled &lt;code&gt;cp&lt;/code&gt; jobs if the migration is going to take days.&lt;/p&gt;

&lt;p&gt;Also create a Hermes backup if Hermes already contains useful state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes backup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The migration normally creates its own pre-migration Hermes archive under &lt;code&gt;~/.hermes/backups/&lt;/code&gt;. Do not pass &lt;code&gt;--no-backup&lt;/code&gt; during the first cutover; saving a few seconds is not worth removing the simplest rollback path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Install and test an empty Hermes Agent
&lt;/h2&gt;

&lt;p&gt;Install Hermes, select a model, and prove that the basic terminal agent works before importing OpenClaw state. This separates installation and provider failures from migration failures. The &lt;a href="https://www.glukhov.org/ai-systems/hermes/" rel="noopener noreferrer"&gt;Hermes AI Assistant guide&lt;/a&gt; covers provider selection and gateway configuration in depth; for the migration you only need a working terminal baseline.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://hermes-agent.nousresearch.com/install.sh | bash
&lt;span class="nb"&gt;source&lt;/span&gt; ~/.bashrc
hermes setup
hermes status
hermes doctor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you already installed Hermes, update it before relying on current migration behavior:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes update
hermes &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That version check is not a formality. The &lt;code&gt;claw migrate&lt;/code&gt; safety posture changed materially during 2026: current builds refuse to apply a conflicted plan, write a pre-migration restore point by default, redact secrets in the reports they save to disk, and require &lt;code&gt;--migrate-secrets&lt;/code&gt; explicitly even under &lt;code&gt;--preset full&lt;/code&gt;. Older builds did none of those things — notably, &lt;code&gt;--preset full&lt;/code&gt; used to pull in API keys silently, and a conflicted plan would report "migrated 0" after you had already confirmed. If you are following an older tutorial, the flags may look identical while the behavior differs in exactly the places that matter.&lt;/p&gt;

&lt;p&gt;Do not configure the old bot tokens yet. Terminal-only validation allows OpenClaw to remain live while you prepare Hermes, and it avoids two gateway processes competing for the same messaging identity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Run the dry-run before choosing flags
&lt;/h2&gt;

&lt;p&gt;Start with the full preset because it reveals the largest possible mapping surface, but keep secrets excluded:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="nt"&gt;--dry-run&lt;/span&gt; &lt;span class="nt"&gt;--preset&lt;/span&gt; full
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The migration always presents a preview before applying, even without &lt;code&gt;--dry-run&lt;/code&gt;. The explicit flag is still valuable because it makes your intent unambiguous and gives you time to inspect source paths, destinations, transforms, conflicts, skipped items, archives, and secret warnings without an impatient confirmation prompt. The complete flag set for &lt;code&gt;claw migrate&lt;/code&gt; and its neighbours is summarised in the &lt;a href="https://www.glukhov.org/ai-systems/hermes/hermes-agent-cli-cheatsheet/" rel="noopener noreferrer"&gt;Hermes Agent CLI cheat sheet&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Use a custom source when OpenClaw state is not in the default location:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--dry-run&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--preset&lt;/span&gt; full &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--source&lt;/span&gt; /srv/openclaw-state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;AGENTS.md&lt;/code&gt; should apply to a particular repository, say so explicitly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--dry-run&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--preset&lt;/span&gt; full &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--workspace-target&lt;/span&gt; /srv/projects/my-project
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without &lt;code&gt;--workspace-target&lt;/code&gt;, workspace instructions are not placed into an arbitrary current directory. That is the correct behavior: an instruction file belongs to a scope, and guessing its scope can change every Hermes session launched below the wrong directory.&lt;/p&gt;

&lt;h3&gt;
  
  
  Full or user-data preset?
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;full&lt;/code&gt; preset includes compatible infrastructure and behavior settings. The &lt;code&gt;user-data&lt;/code&gt; preset focuses on persona, memories, skills, and related user content while excluding infrastructure configuration.&lt;/p&gt;

&lt;p&gt;Use &lt;code&gt;user-data&lt;/code&gt; when Hermes already has a carefully built provider, gateway, security, or sandbox configuration. Use &lt;code&gt;full&lt;/code&gt; when Hermes is new and OpenClaw is the authoritative setup, but still inspect every transformed behavior setting. Neither preset imports secrets unless &lt;code&gt;--migrate-secrets&lt;/code&gt; is added.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: Resolve conflicts without destroying provenance
&lt;/h2&gt;

&lt;p&gt;The default conflict behavior is conservative: the migration refuses to apply a plan with unresolved file conflicts unless &lt;code&gt;--overwrite&lt;/code&gt; is set. That is preferable to an apparently successful cutover that overwrites a newer Hermes persona or skill — and preferable to the older behavior, where confirming a conflicted plan produced a "migrated 0" result that looked like a no-op but was really a silent skip.&lt;/p&gt;

&lt;p&gt;Skill conflicts are handled separately, and the default there is &lt;code&gt;skip&lt;/code&gt;, which quietly keeps the existing Hermes version and discards the incoming one. For a first migration I recommend &lt;code&gt;rename&lt;/code&gt; instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--preset&lt;/span&gt; full &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--workspace-target&lt;/span&gt; /srv/projects/my-project &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--skill-conflict&lt;/span&gt; rename
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Imported skills are placed under &lt;code&gt;~/.hermes/skills/openclaw-imports/&lt;/code&gt;. With &lt;code&gt;rename&lt;/code&gt;, a name collision produces an imported sibling rather than hiding either version. Review the two implementations, test the chosen one, and remove the redundant copy later.&lt;/p&gt;

&lt;p&gt;Use &lt;code&gt;--overwrite&lt;/code&gt; only after reviewing the preview or when rebuilding a disposable Hermes profile. It applies more broadly than skill conflict handling and can replace existing Hermes files. The presence of a backup makes overwriting recoverable, not desirable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 6: Migrate configuration and user data without secrets
&lt;/h2&gt;

&lt;p&gt;Apply the reviewed plan and leave credentials for the cutover stage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--preset&lt;/span&gt; full &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--workspace-target&lt;/span&gt; /srv/projects/my-project &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--skill-conflict&lt;/span&gt; rename
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After completion, save the printed counts for migrated, skipped, conflicting, and archived items. Open the timestamped migration directory and read its summary before starting a new Hermes session. Current builds redact detected secret values in the &lt;code&gt;report.json&lt;/code&gt; and &lt;code&gt;summary.md&lt;/code&gt; they write, so those files are safe to keep alongside your change notes — but confirm that on your version rather than assuming it, because earlier builds wrote raw API keys into the same reports.&lt;/p&gt;

&lt;p&gt;New sessions matter. Imported skills and memory entries are loaded when a session begins, so testing inside a session that predates the migration can produce a false "skill not found" or stale-memory result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 7: Validate behavior before channel cutover
&lt;/h2&gt;

&lt;p&gt;Run the post-migration checks from the terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes status
hermes doctor
hermes config show
hermes gateway status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If memory recall looks incomplete, rebuild the index before concluding the import failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes memory reindex
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then start a new Hermes conversation and test observable behaviors, not just file presence. Ask for a known user preference from memory, invoke one imported skill, call an MCP tool, run a harmless terminal command that should be allowed, and try one that should require approval.&lt;/p&gt;

&lt;p&gt;A useful acceptance matrix looks like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Failure usually means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Persona&lt;/td&gt;
&lt;td&gt;Ask a question where tone and boundaries are obvious&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;SOUL.md&lt;/code&gt; was not found, was overwritten, or needs identity content merged&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User memory&lt;/td&gt;
&lt;td&gt;Ask for a known stable preference&lt;/td&gt;
&lt;td&gt;Memory entries were not imported, deduplicated unexpectedly, not reindexed, or not loaded in a new session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skill&lt;/td&gt;
&lt;td&gt;Invoke a distinctive imported workflow&lt;/td&gt;
&lt;td&gt;Name conflict, invalid metadata, missing dependency, or stale session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider&lt;/td&gt;
&lt;td&gt;Run a normal and long response&lt;/td&gt;
&lt;td&gt;Wrong model mapping, missing credential, or incompatible API type&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP&lt;/td&gt;
&lt;td&gt;Call one read-only tool from each server&lt;/td&gt;
&lt;td&gt;Missing environment, wrong &lt;code&gt;cwd&lt;/code&gt;, transport mismatch, or tool filter issue&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal&lt;/td&gt;
&lt;td&gt;Test allowed and approval-required commands&lt;/td&gt;
&lt;td&gt;Approval mode or allowlist mapping changed policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser&lt;/td&gt;
&lt;td&gt;Open a harmless test page&lt;/td&gt;
&lt;td&gt;CDP URL, browser backend, or sandbox access differs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compression&lt;/td&gt;
&lt;td&gt;Run a long disposable session&lt;/td&gt;
&lt;td&gt;Summary model or compaction behavior was not mapped as intended&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session reset&lt;/td&gt;
&lt;td&gt;Inspect config and test on a disposable profile&lt;/td&gt;
&lt;td&gt;Daily/idle interpretation differs from OpenClaw rules&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The migration maps &lt;code&gt;timeoutSeconds&lt;/code&gt; to an estimated maximum-turn value, translates reasoning levels, and converts approval modes. Those are semantic mappings rather than byte-for-byte copies. Check that the resulting behavior matches your intent, especially for long autonomous tasks and command execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 8: Handle secrets as a separate security change
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--migrate-secrets&lt;/code&gt; can collect allowlisted keys from OpenClaw config values, &lt;code&gt;~/.openclaw/.env&lt;/code&gt;, config environment objects, and per-agent auth profiles (&lt;code&gt;~/.openclaw/agents/&amp;lt;agent&amp;gt;/agent/auth-profiles.json&lt;/code&gt;). It understands plain strings, environment templates, and environment-backed SecretRef objects.&lt;/p&gt;

&lt;p&gt;It intentionally does not copy arbitrary secret names. File-backed and command-backed SecretRefs cannot be resolved automatically, and values outside the supported allowlist remain for manual setup. Treat every warning here as a control working as designed, not as a reason to paste the entire OpenClaw environment into Hermes.&lt;/p&gt;

&lt;p&gt;For a first migration, I prefer to configure provider credentials through Hermes after the data import. If you do use automated secret migration, preview it and run it only when you are ready to transfer channel ownership:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--dry-run&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--preset&lt;/span&gt; full &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--migrate-secrets&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then verify presence without printing values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes status
hermes auth status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rotate credentials if they were exposed in shell history, pasted into migration notes, or stored with weaker permissions than intended. Migration preserves access; it does not prove that the old secret-handling practice was safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 9: Perform a controlled messaging cutover
&lt;/h2&gt;

&lt;p&gt;There is no honest zero-downtime handoff when two processes would poll, subscribe, or respond as the same bot account. The safe pattern is prepare in parallel, stop OpenClaw, start Hermes, test each platform, and keep the rollback commands ready.&lt;/p&gt;

&lt;p&gt;First stop the OpenClaw gateway and confirm it is stopped:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openclaw gateway stop
openclaw gateway status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now migrate or manually set the messaging secrets, configure the Hermes gateway, and start it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway setup
hermes gateway &lt;span class="nb"&gt;install
&lt;/span&gt;hermes gateway start
hermes gateway status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Send a direct message from an allowed user on each platform. Test inbound text, a reply, an attachment if used, a slash command, a long-running task, interruption, and an outbound scheduled or manual send. A green service status proves that a process is running; it does not prove that allowlists, thread routing, delivery, and formatting survived the move.&lt;/p&gt;

&lt;p&gt;WhatsApp requires re-pairing because the migration does not transfer the Baileys session as a reusable token. Run &lt;code&gt;hermes whatsapp&lt;/code&gt; and complete the QR flow. Other channels may reuse tokens, but account layouts and multi-account bindings still deserve explicit testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills, plugins, and MCP servers are not interchangeable
&lt;/h2&gt;

&lt;p&gt;OpenClaw skills from four locations can be imported, but an imported directory is only useful if its assumptions remain true. Check command names, filesystem paths, environment variables, platform-specific tools, and references to OpenClaw-only APIs. The &lt;a href="https://www.glukhov.org/ai-systems/openclaw/skills/" rel="noopener noreferrer"&gt;OpenClaw skills guide&lt;/a&gt; explains the source formats; the &lt;a href="https://www.glukhov.org/ai-systems/hermes/authoring-hermes-skill/" rel="noopener noreferrer"&gt;Hermes skill authoring guide&lt;/a&gt; covers the destination behavior.&lt;/p&gt;

&lt;p&gt;OpenClaw plugins do not become Hermes plugins. Reconstruct the capability at the narrowest suitable layer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use a Hermes skill for procedure, tool selection, and reusable instructions.&lt;/li&gt;
&lt;li&gt;Use an MCP server for live data or an external service boundary.&lt;/li&gt;
&lt;li&gt;Use a built-in Hermes tool when it already provides the capability.&lt;/li&gt;
&lt;li&gt;Use a Hermes plugin only when code must participate in the agent runtime itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a good moment to remove architectural sediment. A plugin installed to compensate for an old OpenClaw limitation may have no reason to survive in Hermes, while a plugin holding a durable database needs a deliberate export or replacement plan.&lt;/p&gt;

&lt;p&gt;MCP definitions migrate more directly, including commands, arguments, environments, working directories, URLs, and include/exclude tool filters. Still test every server separately: a correct YAML mapping cannot install a missing executable, renew OAuth, or make a path from the old host exist on the new one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory needs a quality check, not a line-count check
&lt;/h2&gt;

&lt;p&gt;Hermes imports &lt;code&gt;MEMORY.md&lt;/code&gt;, &lt;code&gt;USER.md&lt;/code&gt;, and daily memory files into its memory structure. That preserves useful facts, but OpenClaw memory plugins, long-context databases, embedding indexes, and recall policies are archived rather than translated into an equivalent cognitive system.&lt;/p&gt;

&lt;p&gt;Review imported memory in three passes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Identity and stable preferences: preserve concise facts that should influence many sessions.&lt;/li&gt;
&lt;li&gt;Operational knowledge: move repeatable procedures into skills or project context instead of global memory.&lt;/li&gt;
&lt;li&gt;Historical residue: archive completed incidents, stale plans, and self-referential agent commentary rather than injecting it forever.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not import every transcript as durable memory. More remembered text can make an agent less coherent by repeatedly retrieving obsolete constraints and its own earlier guesses. The &lt;a href="https://www.glukhov.org/ai-systems/hermes/hermes-agent-memory-system/" rel="noopener noreferrer"&gt;Hermes memory system guide&lt;/a&gt; explains where the imported entries will live, and the &lt;a href="https://www.glukhov.org/ai-systems/memory/agent-memory-providers/" rel="noopener noreferrer"&gt;agent memory provider comparison&lt;/a&gt; is the better place to choose a new long-term backend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recreate cron jobs, heartbeats, hooks, and multi-agent routing
&lt;/h2&gt;

&lt;p&gt;Cron jobs are archived because scheduled execution is not just a cron expression. A job also has a prompt or command, working directory, model, timeout, delivery destination, permissions, retry behavior, and expectations about session state.&lt;/p&gt;

&lt;p&gt;For every archived OpenClaw job, write down those fields and recreate it with Hermes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes cron create
hermes cron list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run each job once manually before enabling its schedule. Verify both the work and the delivery path, especially when the old job posted to a Telegram chat, Slack channel, or Discord thread.&lt;/p&gt;

&lt;p&gt;Translate &lt;code&gt;HEARTBEAT.md&lt;/code&gt; into explicit scheduled jobs only when periodic execution is truly required. A vague heartbeat that asks the agent to inspect everything every few minutes is expensive and difficult to verify; separate named jobs with observable outcomes are easier to operate.&lt;/p&gt;

&lt;p&gt;Multi-agent definitions and channel bindings also require manual design. Hermes profiles provide isolated state and gateways, but they are not a syntactic rewrite of OpenClaw's agent list. Map each agent by responsibility, workspace, credentials, channel, and security boundary rather than reproducing names first; the profile-first reasoning behind that mapping is worked through in the &lt;a href="https://www.glukhov.org/ai-systems/hermes/production-setup/" rel="noopener noreferrer"&gt;Hermes production setup guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Troubleshooting the failures that matter
&lt;/h2&gt;

&lt;h3&gt;
  
  
  "OpenClaw directory not found"
&lt;/h3&gt;

&lt;p&gt;The command searches the current OpenClaw, Clawdbot, and Moltbot default directories. If your state lives elsewhere, point to the directory that contains the OpenClaw config and related state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes claw migrate &lt;span class="nt"&gt;--dry-run&lt;/span&gt; &lt;span class="nt"&gt;--source&lt;/span&gt; /path/to/openclaw
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not point &lt;code&gt;--source&lt;/code&gt; at only the workspace unless that is genuinely the complete source tree. The preview should show config, workspace, and recognized categories.&lt;/p&gt;

&lt;h3&gt;
  
  
  The migration refuses because of conflicts
&lt;/h3&gt;

&lt;p&gt;This is the safe default, not a crash. Back up Hermes, identify which side is authoritative for each conflict, use &lt;code&gt;--skill-conflict rename&lt;/code&gt; for skills, and reserve &lt;code&gt;--overwrite&lt;/code&gt; for a reviewed plan.&lt;/p&gt;

&lt;p&gt;If the existing Hermes configuration is valuable, consider the &lt;code&gt;user-data&lt;/code&gt; preset. It imports the assistant's user-owned content without trying to replace established infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Imported skills do not appear
&lt;/h3&gt;

&lt;p&gt;Start a new session and inspect the imported directory below &lt;code&gt;~/.hermes/skills/openclaw-imports/&lt;/code&gt;. Use &lt;code&gt;/skills&lt;/code&gt; inside Hermes to confirm discovery. If the skill exists but cannot run, inspect its dependency and tool assumptions rather than repeating the migration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Provider keys were not found
&lt;/h3&gt;

&lt;p&gt;The key may be stored in an OpenClaw environment file, config environment object, auth profile, file-backed SecretRef, command-backed SecretRef, or unsupported variable name. The migrator resolves the supported forms and warns about the rest. Add unresolved values through Hermes configuration or authentication commands instead of converting secure references into plaintext merely to satisfy the importer.&lt;/p&gt;

&lt;h3&gt;
  
  
  The bot is running but messages are missing or duplicated
&lt;/h3&gt;

&lt;p&gt;Confirm that the OpenClaw gateway is stopped and that only one Hermes profile owns the token. Then inspect &lt;code&gt;hermes gateway status&lt;/code&gt; and gateway logs, followed by channel allowlists and account selection. Duplicate consumers and incorrect allowlists are more common than a broken language model.&lt;/p&gt;

&lt;h3&gt;
  
  
  The personality is present but recall is poor
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;SOUL.md&lt;/code&gt; and memory are different layers. Confirm that the persona copied to &lt;code&gt;~/.hermes/SOUL.md&lt;/code&gt;, memory entries reached &lt;code&gt;~/.hermes/memories/&lt;/code&gt;, and the test uses a new session. Run &lt;code&gt;hermes memory reindex&lt;/code&gt; before deeper debugging. If OpenClaw depended on an external memory plugin, configure a Hermes memory provider rather than expecting Markdown import to recreate its retrieval behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Roll back Hermes
&lt;/h3&gt;

&lt;p&gt;Stop the Hermes gateway before restoring the pre-migration Hermes backup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;hermes gateway stop
hermes import ~/.hermes/backups/pre-migration-&amp;lt;timestamp&amp;gt;.zip
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;hermes import&lt;/code&gt; overwrites files in Hermes home with the archive contents, so inspect the exact filename and understand that post-migration Hermes sessions may be replaced. Then keep Hermes stopped, restart OpenClaw, and verify its gateway and channel health.&lt;/p&gt;

&lt;h2&gt;
  
  
  Manual migration when the command cannot model your setup
&lt;/h2&gt;

&lt;p&gt;A manual fallback is slower but sometimes clearer for heavily customized installations. Build a clean Hermes profile and migrate by responsibility:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Copy or rewrite persona content into &lt;code&gt;~/.hermes/SOUL.md&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Curate stable user facts into Hermes &lt;code&gt;MEMORY.md&lt;/code&gt; and &lt;code&gt;USER.md&lt;/code&gt; rather than copying all history.&lt;/li&gt;
&lt;li&gt;Place project instructions in the correct repository-level &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Copy compatible skills into a named import directory and test them individually.&lt;/li&gt;
&lt;li&gt;Translate provider and MCP definitions into &lt;code&gt;~/.hermes/config.yaml&lt;/code&gt; without printing secrets.&lt;/li&gt;
&lt;li&gt;Configure credentials through Hermes auth or secret management.&lt;/li&gt;
&lt;li&gt;Recreate approvals, sandboxing, browser access, cron jobs, webhooks, and channels.&lt;/li&gt;
&lt;li&gt;Replace each OpenClaw plugin with an explicit Hermes capability or retire it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The manual route is especially appropriate when the source contains several OpenClaw agents with different workspaces, memory plugins, and channel bindings. An automatic union can preserve files while erasing the isolation that made the setup safe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not clean up OpenClaw immediately
&lt;/h2&gt;

&lt;p&gt;After Hermes has passed local and messaging tests, keep OpenClaw installed but stopped for a soak period. Preserve the verified OpenClaw backup, migration archive, pre-migration Hermes backup, and a copy of the acceptance checklist.&lt;/p&gt;

&lt;p&gt;Hermes documents &lt;code&gt;hermes claw cleanup&lt;/code&gt; for renaming leftover OpenClaw directories to &lt;code&gt;.pre-migration/&lt;/code&gt;, and &lt;code&gt;hermes claw cleanup --dry-run&lt;/code&gt; to preview what would be archived. Use it only after the OpenClaw gateway is stopped, the current Hermes version includes process guards, and you have decided not to roll back. Older 2026 builds had a reported cleanup path that could move state while an OpenClaw gateway was still running; current code marks the guard as implemented, but a verified backup and stopped source service remain the sensible boundary.&lt;/p&gt;

&lt;p&gt;Cleanup is not required to prove that Hermes works. It exists to reduce future state confusion, so postponing it during a rollback window is good operations, not untidiness.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to stay on OpenClaw 2.0
&lt;/h2&gt;

&lt;p&gt;OpenClaw 2.0 is not an abandoned baseline. The v2026.8.1 release landed over 16,000 pull requests from more than 900 contributors — roughly half of the project's total merge history — and substantially changed onboarding, the web Control UI, session storage, backups, channels, memory, plugins, automations, browser and computer use, security, and service reliability. If those platform features are central to your deployment, migration may remove more working capability than it simplifies.&lt;/p&gt;

&lt;p&gt;Stay on OpenClaw when you depend on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Its rebuilt Control UI, with docked file editor, git-backed Changes panel, browser panel, and in-conversation approvals.&lt;/li&gt;
&lt;li&gt;Session presets, transcript search, groups, status views, and batch actions.&lt;/li&gt;
&lt;li&gt;A product-specific plugin with no Hermes equivalent.&lt;/li&gt;
&lt;li&gt;Complex multi-user, mobile, device, or channel routing already working in production.&lt;/li&gt;
&lt;li&gt;OpenClaw-specific browser, computer-use, or Gateway administration.&lt;/li&gt;
&lt;li&gt;A memory or session database that cannot be exported with acceptable loss.&lt;/li&gt;
&lt;li&gt;Operational controls your team already knows and monitors.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Move to Hermes when its simpler terminal-first workflow, profiles, learning-oriented skills, memory model, scheduled tasks, provider flexibility, or delegation model better matches what you actually operate. The &lt;a href="https://www.glukhov.org/ai-systems/comparisons/openclaw-hermes-alternatives-popularity/" rel="noopener noreferrer"&gt;OpenClaw and Hermes comparison&lt;/a&gt; discusses that decision with current numbers; this page is about executing the cutover once the decision is made.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final migration checklist
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] OpenClaw version and resolved paths recorded.&lt;/li&gt;
&lt;li&gt;[ ] Verified OpenClaw backup stored outside live state.&lt;/li&gt;
&lt;li&gt;[ ] Pre-2.0 JSONL transcripts snapshotted separately if they matter.&lt;/li&gt;
&lt;li&gt;[ ] Existing Hermes backup created.&lt;/li&gt;
&lt;li&gt;[ ] Hermes version checked against current &lt;code&gt;claw migrate&lt;/code&gt; safety behavior.&lt;/li&gt;
&lt;li&gt;[ ] Full dry-run reviewed.&lt;/li&gt;
&lt;li&gt;[ ] Every conflict assigned a resolution.&lt;/li&gt;
&lt;li&gt;[ ] Archive contents added to the manual work list.&lt;/li&gt;
&lt;li&gt;[ ] Persona, user memory, and skills tested in a new session.&lt;/li&gt;
&lt;li&gt;[ ] Provider, fallback model, MCP, browser, and terminal tested.&lt;/li&gt;
&lt;li&gt;[ ] Approval and sandbox behavior tested, including a denied action.&lt;/li&gt;
&lt;li&gt;[ ] Cron jobs, plugins, hooks, memory backend, and multi-agent bindings rebuilt or retired.&lt;/li&gt;
&lt;li&gt;[ ] OpenClaw gateway stopped before channel credentials moved.&lt;/li&gt;
&lt;li&gt;[ ] Every messaging channel tested from an allowed account.&lt;/li&gt;
&lt;li&gt;[ ] WhatsApp re-paired if used.&lt;/li&gt;
&lt;li&gt;[ ] Rollback commands and archive names recorded.&lt;/li&gt;
&lt;li&gt;[ ] OpenClaw cleanup deferred until the soak period ends.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Final verdict
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;hermes claw migrate&lt;/code&gt; is good enough to make an OpenClaw-to-Hermes move routine, but only if "routine" means planned and reversible. Its strongest feature is not the number of files it copies; it is the preview that tells you which parts of the old assistant have a real Hermes equivalent and which parts still require engineering judgment.&lt;/p&gt;

&lt;p&gt;Use the full preset to discover the scope, keep secrets out of the first pass, rename skill conflicts, test from the terminal, and transfer channel ownership as a separate event. Most importantly, preserve the old system until Hermes has completed real scheduled work and real conversations, not merely returned a successful status command.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://hermes-agent.nousresearch.com/docs/guides/migrate-from-openclaw/" rel="noopener noreferrer"&gt;Hermes guide: Migrate from OpenClaw&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hermes-agent.nousresearch.com/docs/reference/cli-commands" rel="noopener noreferrer"&gt;Hermes CLI command reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/NousResearch/hermes-agent" rel="noopener noreferrer"&gt;Hermes Agent repository and installation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/NousResearch/hermes-agent/pull/16911" rel="noopener noreferrer"&gt;Hermes PR #16911: plan-first apply, redaction, and pre-migration backup&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.openclaw.ai/gateway/configuration-reference" rel="noopener noreferrer"&gt;OpenClaw configuration reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.openclaw.ai/cli/gateway" rel="noopener noreferrer"&gt;OpenClaw Gateway service commands&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.openclaw.ai/cli/backup" rel="noopener noreferrer"&gt;OpenClaw backup and restore commands&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.openclaw.ai/install/backups" rel="noopener noreferrer"&gt;OpenClaw backups overview, SQLite snapshots, and Litestream&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.openclaw.ai/releases/2026.8.1" rel="noopener noreferrer"&gt;OpenClaw 2.0 release notes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/NousResearch/hermes-agent/issues/8596" rel="noopener noreferrer"&gt;Hermes cleanup process-guard issue and resolution&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>hermes</category>
      <category>openclaw</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>llama.cpp vs Ollama in 2026: Which Runtime Should You Run?</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Mon, 14 Sep 2026 10:14:29 +0000</pubDate>
      <link>https://dev.to/rosgluk/llamacpp-vs-ollama-in-2026-which-runtime-should-you-run-4k7f</link>
      <guid>https://dev.to/rosgluk/llamacpp-vs-ollama-in-2026-which-runtime-should-you-run-4k7f</guid>
      <description>&lt;p&gt;Ollama and llama.cpp are often compared as if they were rival inference engines. The real choice is between a managed model service and a toolkit you operate directly.&lt;/p&gt;

&lt;p&gt;Ollama wraps a pinned and patched llama.cpp inside a scheduler, a model store, and an API, so the named model becomes the unit you operate. Direct llama.cpp inverts that: the &lt;code&gt;llama-server&lt;/code&gt; process and its flags are the unit, and every choice about context, KV cache, and GPU placement is one you make and can show in a command.&lt;/p&gt;

&lt;p&gt;This guide compares the two the way the decision actually happens: installation and daily commands, model management and lifetime, runtime control, APIs, performance, failure modes, and security. It ends with concrete triggers for keeping Ollama, moving to &lt;code&gt;llama-server&lt;/code&gt;, and a low-risk migration path between them. If you are still deciding between local, self-hosted, and cloud approaches at a higher level, start with the &lt;a href="https://www.glukhov.org/llm-hosting/" rel="noopener noreferrer"&gt;LLM Hosting overview&lt;/a&gt;; for the wider landscape of local tools beyond this pair, the &lt;a href="https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/" rel="noopener noreferrer"&gt;local LLM hosting comparison&lt;/a&gt; covers vLLM, LM Studio, LocalAI, and more.&lt;/p&gt;

&lt;h2&gt;
  
  
  llama.cpp vs Ollama: the short answer
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;Better default&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First local chat model&lt;/td&gt;
&lt;td&gt;Ollama&lt;/td&gt;
&lt;td&gt;One command pulls, configures, and runs a named model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reusable model catalog&lt;/td&gt;
&lt;td&gt;Ollama&lt;/td&gt;
&lt;td&gt;Tags, manifests, a registry, and &lt;code&gt;Modelfile&lt;/code&gt; recipes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exact GGUF file control&lt;/td&gt;
&lt;td&gt;llama.cpp&lt;/td&gt;
&lt;td&gt;The server can run the file directly without importing it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fine-grained GPU placement&lt;/td&gt;
&lt;td&gt;llama.cpp&lt;/td&gt;
&lt;td&gt;Explicit layer offload, device selection, and multi-GPU split modes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-server KV cache tuning&lt;/td&gt;
&lt;td&gt;llama.cpp&lt;/td&gt;
&lt;td&gt;Separate K and V cache types and many cache controls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automatic model loading and expiry&lt;/td&gt;
&lt;td&gt;Ollama&lt;/td&gt;
&lt;td&gt;Built-in scheduler and &lt;code&gt;keep_alive&lt;/code&gt; behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI-compatible local endpoint&lt;/td&gt;
&lt;td&gt;Either&lt;/td&gt;
&lt;td&gt;Both support common routes, but neither promises perfect compatibility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metrics and slot inspection&lt;/td&gt;
&lt;td&gt;llama.cpp&lt;/td&gt;
&lt;td&gt;Native Prometheus metrics and server slot endpoints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Native SDKs and tool integrations&lt;/td&gt;
&lt;td&gt;Ollama&lt;/td&gt;
&lt;td&gt;Polished Python and JavaScript clients plus named integrations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New llama.cpp feature immediately&lt;/td&gt;
&lt;td&gt;llama.cpp&lt;/td&gt;
&lt;td&gt;No wait for Ollama to update its pinned and patched revision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multiple GGUF models behind one endpoint&lt;/td&gt;
&lt;td&gt;Ollama, usually&lt;/td&gt;
&lt;td&gt;Mature lifecycle management; llama.cpp router mode is now a credible alternative&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you only need a dependable backend for Open WebUI, a coding assistant, or a few local scripts, Ollama is usually the less distracting choice. If you keep asking what Ollama selected, allocated, changed, or hid, you have probably reached the point where &lt;code&gt;llama-server&lt;/code&gt; is the cleaner system.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the comparison actually means in 2026
&lt;/h2&gt;

&lt;p&gt;llama.cpp is a C and C++ inference project with CPU and GPU backends, GGUF model tooling, command-line programs, and an HTTP server. Its direct serving program supports OpenAI-compatible Chat Completions, Responses, embeddings, multimodal requests, function calling, structured output, continuous batching, speculative decoding, monitoring endpoints, and a built-in web UI.&lt;/p&gt;

&lt;p&gt;Ollama is a higher-level service. It maintains a local model store, gives models stable names, downloads and imports artifacts, applies templates and defaults, selects an available backend, schedules model processes, and unloads idle models. Its native API also reports timing and loading information that is convenient for local applications.&lt;/p&gt;

&lt;p&gt;The frequently repeated statement that "Ollama is just a wrapper around llama.cpp" is directionally useful but technically incomplete. Ollama pins llama.cpp source, applies compatibility patches, and launches a server through its own scheduler, but it also has product behavior that llama.cpp does not define; on Apple silicon, Ollama can use its MLX engine as well. The request path makes the difference concrete:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph o["Ollama stack"]
        A[Client] --&amp;gt; B[Ollama API on port 11434]
        B --&amp;gt; C[Scheduler and model store]
        C --&amp;gt; D[Pinned and patched llama.cpp]
        D --&amp;gt; E[GPU or CPU]
    end
    subgraph l["Direct llama.cpp"]
        F[Client] --&amp;gt; G[llama-server on port 8080]
        G --&amp;gt; H[llama.cpp runtime with explicit flags]
        H --&amp;gt; I[GPU or CPU]
    end&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;This leads to the most useful mental model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;With Ollama, the named model is the unit you operate.&lt;/li&gt;
&lt;li&gt;With direct llama.cpp, the server process and its flags are the unit you operate.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Installation and the daily command surface
&lt;/h2&gt;

&lt;p&gt;Ollama optimizes the first five minutes. After installation, pulling and starting a model is intentionally terse:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama run qwen3:8b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model name represents more than its weights. Ollama can associate a template, parameters, system prompt, license, adapter, and minimum runtime version with that name. &lt;code&gt;ollama list&lt;/code&gt;, &lt;code&gt;ollama show&lt;/code&gt;, &lt;code&gt;ollama ps&lt;/code&gt;, and &lt;code&gt;ollama stop&lt;/code&gt; provide a coherent management surface.&lt;/p&gt;

&lt;p&gt;Direct llama.cpp starts closer to the metal. You can download a release binary, build a backend-specific version, use a container, or use the newer Hugging Face download path, as the &lt;a href="https://www.glukhov.org/llm-hosting/llama-cpp/" rel="noopener noreferrer"&gt;llama.cpp quickstart&lt;/a&gt; details. A local GGUF server might start like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; /srv/models/qwen3-8b-q4_k_m.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--alias&lt;/span&gt; qwen3-8b &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--host&lt;/span&gt; 127.0.0.1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--port&lt;/span&gt; 8080 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 32768 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--n-gpu-layers&lt;/span&gt; all &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--flash-attn&lt;/span&gt; on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Current llama.cpp documentation also shows the unified &lt;code&gt;llama serve&lt;/code&gt; command in its quickstart. Packaged executable names can vary by distribution, so check the release or package you installed rather than copying a service file blindly.&lt;/p&gt;

&lt;p&gt;The longer command is not automatically a disadvantage. It is an executable record of the runtime you intended to create. Put it in a systemd unit, Compose file, or shell script, and configuration becomes reviewable instead of being spread across a model manifest, environment variables, API options, and scheduler defaults.&lt;/p&gt;

&lt;h3&gt;
  
  
  The practical installation tradeoff
&lt;/h3&gt;

&lt;p&gt;Ollama is easier to install consistently across developer machines. It is also easier to explain to someone who should use a model but should not need to understand tensor offload, chat templates, or KV memory.&lt;/p&gt;

&lt;p&gt;llama.cpp is easier to make exact. You choose the build, backend, version, file, and flags, which is valuable when a new GPU kernel fixes your workload or a recent commit breaks it. That freedom also means you own upgrades, service supervision, and regression testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model management: library names or ordinary files
&lt;/h2&gt;

&lt;p&gt;Ollama treats models rather like container images. A familiar name points to a manifest and content-addressed blobs, and &lt;code&gt;ollama pull&lt;/code&gt; resolves the required layers. This is excellent for repeatable workstation setup and for applications that should refer to &lt;code&gt;qwen3:8b&lt;/code&gt; rather than a long filesystem path.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;Modelfile&lt;/code&gt; makes customization reproducible:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; ./qwen3-8b-q4_k_m.gguf&lt;/span&gt;
PARAMETER num_ctx 32768
PARAMETER temperature 0.7
PARAMETER top_p 0.9
SYSTEM You are a precise technical assistant.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama create qwen3-8b-local &lt;span class="nt"&gt;-f&lt;/span&gt; Modelfile
ollama run qwen3-8b-local
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ollama can import a local GGUF, so choosing Ollama does not restrict you to the public Ollama library. The import step does, however, hand the artifact to Ollama's model store. If you also retain the original GGUF for llama.cpp or LM Studio, account for the additional managed copy unless your storage layer deduplicates it.&lt;/p&gt;

&lt;p&gt;llama.cpp can simply point at the GGUF you already have. It can also download a selected quantization from Hugging Face:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="nt"&gt;-hf&lt;/span&gt; ggml-org/Qwen3-8B-GGUF:Q4_K_M
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This file-oriented approach works especially well for testing fresh quantizations. Download a file, change one path, and start it; there is no create step and no question about which blob a model name resolves to.&lt;/p&gt;

&lt;h3&gt;
  
  
  Templates are part of the model, even when they look like configuration
&lt;/h3&gt;

&lt;p&gt;The weights do not define the entire chat behavior. The chat template controls how system, user, assistant, thinking, and tool messages become tokens. Stop sequences and parser behavior can change the result again.&lt;/p&gt;

&lt;p&gt;Ollama's curated library reduces this risk because its named models carry tested metadata, and current releases (Ollama moved from 0.30 to 0.33.3 between June and September 2026) increasingly honor GGUF-defined default parameters directly rather than requiring you to restate them in a &lt;code&gt;Modelfile&lt;/code&gt;. A hand-imported GGUF may still need a correct &lt;code&gt;TEMPLATE&lt;/code&gt;, parser, or renderer, while llama.cpp normally reads the embedded GGUF chat template and lets you override it. Neither runtime can repair incorrect or missing model metadata by magic.&lt;/p&gt;

&lt;p&gt;If the same quantized model feels noticeably worse after a runtime change, do not conclude that one engine has damaged the weights. First compare the template, context limit, sampling values, thinking mode, tool parser, and runtime revision, one variable at a time, before touching the model itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model lifetime and switching
&lt;/h2&gt;

&lt;p&gt;Ollama's scheduler is one of its strongest reasons to exist. By default, an idle model remains loaded for five minutes; a request-level &lt;code&gt;keep_alive&lt;/code&gt; value can keep it resident indefinitely, change the duration, or unload it immediately. &lt;code&gt;ollama ps&lt;/code&gt; shows loaded models, processor placement, context allocation, and expiry.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Keep a model loaded.&lt;/span&gt;
curl http://localhost:11434/api/generate &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "qwen3:8b",
  "keep_alive": -1
}'&lt;/span&gt;

&lt;span class="c"&gt;# Unload it immediately.&lt;/span&gt;
curl http://localhost:11434/api/generate &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
  "model": "qwen3:8b",
  "keep_alive": 0
}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A traditional &lt;code&gt;llama-server --model ...&lt;/code&gt; process loads one model and keeps it until the process exits. That behavior is wonderfully predictable for a dedicated service: there is no surprise cold load after an idle timeout and no scheduler deciding that another model deserves the memory.&lt;/p&gt;

&lt;p&gt;llama.cpp now also has router mode. Starting &lt;code&gt;llama-server&lt;/code&gt; without a model can expose cached models, a GGUF directory, or INI presets and dynamically load instances according to the requested model name. It closes the old lifecycle gap only partially: only one model is resident per worker at a time, a switch is a full unload-and-reload rather than instant, and there is no eviction policy or warm pool — every alternating request between two models pays a full reload. This narrows the gap versus never having router mode at all, but it does not make the two products identical; Ollama still provides the smoother registry, warm-pool, and administration experience. For the full configuration walkthrough, current limitations, and an honest comparison to both Ollama and &lt;code&gt;llama-swap&lt;/code&gt;, see the &lt;a href="https://www.glukhov.org/llm-hosting/llama-cpp/llama-server-router-mode/" rel="noopener noreferrer"&gt;llama-server router mode guide&lt;/a&gt;. If you need one endpoint across llama.cpp, vLLM, SGLang, and other engines, &lt;a href="https://www.glukhov.org/llm-hosting/llama-swap/" rel="noopener noreferrer"&gt;llama-swap&lt;/a&gt; is a more appropriate abstraction than asking either runtime to become a universal model proxy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Runtime control: where llama.cpp earns the extra work
&lt;/h2&gt;

&lt;p&gt;The decisive llama.cpp advantage is not that it is always faster. It is that you can express the memory and execution plan directly, inspect it, and change one variable at a time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Context and KV cache precision
&lt;/h3&gt;

&lt;p&gt;For direct llama.cpp, context size and K/V cache types can be set per server process:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama-server &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; model.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 65536 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cache-type-k&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cache-type-v&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--parallel&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--flash-attn&lt;/span&gt; on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;K and V can use different types, and llama.cpp exposes additional controls for unified KV allocation, per-slot context limits, prompt cache reuse, and cache persistence. These flags are not decorative on a 16 GB or 32 GB GPU — they determine whether long-context requests fit and how many slots can remain useful, and the underlying VRAM budget math is the same regardless of which runtime enforces it; see &lt;a href="https://www.glukhov.org/llm-performance/optimization/kv-cache-16gb-long-context/" rel="noopener noreferrer"&gt;KV Cache on 16 GB GPUs&lt;/a&gt; for the formula and per-engine cache-type tables.&lt;/p&gt;

&lt;p&gt;Ollama exposes the important common case with &lt;code&gt;OLLAMA_CONTEXT_LENGTH&lt;/code&gt;, the &lt;code&gt;num_ctx&lt;/code&gt; option, and &lt;code&gt;OLLAMA_KV_CACHE_TYPE&lt;/code&gt;. Its KV cache type is a server-wide setting, however, rather than a per-named-model choice. Ollama also scales memory with configured parallelism and context length, which can make a harmless-looking concurrency change consume much more VRAM.&lt;/p&gt;

&lt;p&gt;That behavior belongs primarily in the &lt;a href="https://www.glukhov.org/llm-performance/ollama/how-ollama-handles-parallel-requests/" rel="noopener noreferrer"&gt;Ollama parallel requests guide&lt;/a&gt;. For this comparison, the decision is simpler: use Ollama when a global cache policy is acceptable; use separate llama.cpp services when different models need different cache precision, context, or slot geometry.&lt;/p&gt;

&lt;h3&gt;
  
  
  GPU selection and multi-GPU placement
&lt;/h3&gt;

&lt;p&gt;Ollama aims to choose a sensible placement. It reports whether a model is fully on GPU, fully on CPU, or split, and its scheduler considers available memory when loading models. For a normal single-GPU workstation, automatic placement is often exactly what you want.&lt;/p&gt;

&lt;p&gt;llama.cpp exposes the plan. You can select devices, specify GPU layers, choose layer, row, or experimental tensor split modes, set tensor proportions, select the main GPU, and deliberately keep MoE expert weights on CPU. This is substantially better for asymmetric multi-GPU machines and for squeezing an oversized model into a known memory budget.&lt;/p&gt;

&lt;p&gt;If your operational notes contain phrases such as "put the KV cache on these devices" or "keep only the experts in system RAM," direct llama.cpp is the natural tool. If the requirement is merely "use the GPU if it fits," Ollama saves time without giving up much.&lt;/p&gt;

&lt;h3&gt;
  
  
  New features and backend cadence
&lt;/h3&gt;

&lt;p&gt;Direct llama.cpp is where new llama.cpp model architectures, quantization types, GPU kernels, and experimental server options appear first. That is valuable during a model release week, when support may depend on a specific build number rather than the last stable package.&lt;/p&gt;

&lt;p&gt;Ollama deliberately pins an upstream revision and applies compatibility patches. This can delay an upstream feature, but it can also shield users from churn and integrate it with Ollama's scheduler, templates, and cross-platform packaging. Faster access is not the same thing as greater reliability.&lt;/p&gt;

&lt;p&gt;Ollama 0.30 materially reduced an older gap by expanding GGUF compatibility, improving NVIDIA performance, and enabling Vulkan by default for wider AMD and Intel support. Ollama's release cadence since then has stayed fast — 0.33.3 shipped in early September 2026, roughly three months later, adding cached-prompt-token reporting and another llama.cpp backend bump — so treat any specific version claim in this article, or anywhere else, as something to re-verify against &lt;code&gt;ollama --version&lt;/code&gt; rather than a permanent fact. Any comparison that says Ollama cannot run an arbitrary local GGUF, or that Vulkan always requires an experimental opt-in, is now stale.&lt;/p&gt;

&lt;h2&gt;
  
  
  APIs, tools, vision, and structured output
&lt;/h2&gt;

&lt;p&gt;Both runtimes are credible local API servers in 2026. Both can handle common OpenAI-style chat requests, tools, vision-capable models, embeddings, streaming, and structured output when the model and template support them.&lt;/p&gt;

&lt;p&gt;The difference is in the surrounding surface:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Surface&lt;/th&gt;
&lt;th&gt;Ollama&lt;/th&gt;
&lt;th&gt;llama-server&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Native API&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/api/chat&lt;/code&gt;, &lt;code&gt;/api/generate&lt;/code&gt;, &lt;code&gt;/api/embed&lt;/code&gt; and model APIs&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/completion&lt;/code&gt; plus server-specific control and inspection APIs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI API&lt;/td&gt;
&lt;td&gt;Compatible with parts of the API, including Chat Completions and Responses&lt;/td&gt;
&lt;td&gt;Chat Completions, Responses, embeddings, and other compatible routes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic-style API&lt;/td&gt;
&lt;td&gt;Integrations exist, but check the client path in use&lt;/td&gt;
&lt;td&gt;Anthropic Messages-compatible endpoint is documented&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool calling&lt;/td&gt;
&lt;td&gt;Native API, OpenAI-compatible path, and SDK helpers&lt;/td&gt;
&lt;td&gt;OpenAI-style tools with Jinja templates and function-call parsing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structured output&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;format: "json"&lt;/code&gt; or a JSON schema&lt;/td&gt;
&lt;td&gt;Grammar and JSON-schema constraints plus OpenAI-style response formats&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision&lt;/td&gt;
&lt;td&gt;Simple image messages for supported named models&lt;/td&gt;
&lt;td&gt;Multimodal projector control and OpenAI-compatible image input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;Request timings, logs, &lt;code&gt;ollama ps&lt;/code&gt;, and model APIs&lt;/td&gt;
&lt;td&gt;Health, slots, props, and optional Prometheus metrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authentication&lt;/td&gt;
&lt;td&gt;No API key on the local server by default&lt;/td&gt;
&lt;td&gt;Optional API keys and TLS flags are built in&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Do not treat "OpenAI-compatible" as a binary certification. Ollama says it supports parts of the OpenAI API, while llama.cpp explicitly avoids making a strong compatibility promise. Before changing runtimes, test streaming frames, tool-call arguments, reasoning fields, usage counters, error bodies, and any endpoint your client actually consumes.&lt;/p&gt;

&lt;p&gt;Ollama generally wins when application integration is the job. Its SDKs and documented integrations make the happy path short. llama.cpp wins when the server itself is the object of engineering: its slot view, token timings, metrics, schemas, templates, adapters, and low-level endpoints are unusually useful during diagnosis.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance: benchmark the deployment, not the brand
&lt;/h2&gt;

&lt;p&gt;It is tempting to ask whether llama.cpp or Ollama is faster. On a GGUF path, Ollama may be running a pinned, patched llama.cpp underneath, so a universal brand-level answer is not useful. Results change with the build revision, backend, flash attention, context allocation, parallel slots, batch sizes, cache type, model residency, and whether some layers fell back to CPU.&lt;/p&gt;

&lt;p&gt;A fair comparison starts with the same GGUF and tests two different questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Cold start: include model load time and first response latency.&lt;/li&gt;
&lt;li&gt;Warm service: preload the model, then measure prompt processing and generation separately.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Use one request and one slot first. Match context size, K/V cache type, temperature, top-p, seed, maximum output, and chat template; confirm full GPU offload from logs or status output. Only then increase concurrency, because Ollama and llama.cpp allocate and schedule parallel work differently.&lt;/p&gt;

&lt;p&gt;For Ollama, the final native API response includes load, prompt-evaluation, and generation durations, and current releases also report cached prompt tokens directly in that response — useful for confirming whether prefix reuse actually happened before you credit a speed win to the runtime. For llama.cpp, enable performance reporting or Prometheus metrics and inspect the startup configuration. A five percent throughput win is meaningless if one run silently used a shorter context, a different cache type, or a different template.&lt;/p&gt;

&lt;p&gt;My expectation for the same supported GGUF on one GPU is usually near parity, not a guaranteed llama.cpp victory. Direct llama.cpp can win after deliberate tuning or by adopting a newer optimization; Ollama can be just as fast when its selected engine and defaults line up with the workload. Measure after configuration, not before it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure modes that expose the real difference
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The model unexpectedly uses CPU
&lt;/h3&gt;

&lt;p&gt;With Ollama, run &lt;code&gt;ollama ps&lt;/code&gt; and inspect &lt;code&gt;PROCESSOR&lt;/code&gt;, &lt;code&gt;CONTEXT&lt;/code&gt;, and the loaded size. A larger context, another resident model, or an unsupported GPU path may explain the split. Check the service logs rather than assuming the GPU was ignored.&lt;/p&gt;

&lt;p&gt;With llama.cpp, start with &lt;code&gt;llama-server --list-devices&lt;/code&gt;, then read the startup log for tensor placement and buffer sizes. If you set an exact layer count, device, or split, the command itself is evidence of your intent; this is much easier to reproduce in a bug report.&lt;/p&gt;

&lt;h3&gt;
  
  
  A longer context causes an out-of-memory error
&lt;/h3&gt;

&lt;p&gt;Ollama chooses default context lengths according to available VRAM, and current documentation recommends at least 64K for agent and coding workloads. That recommendation is not a promise that your model, parallelism, and cache will fit. Reduce &lt;code&gt;num_ctx&lt;/code&gt;, reduce parallelism, choose &lt;code&gt;q8_0&lt;/code&gt; KV cache where appropriate, or use a smaller weight quantization. Confirm the actual memory picture with &lt;code&gt;nvidia-smi&lt;/code&gt; before and after a long request so you know whether the model, the cache, or both are the constraint.&lt;/p&gt;

&lt;p&gt;With llama.cpp, reduce &lt;code&gt;--ctx-size&lt;/code&gt;, change &lt;code&gt;--cache-type-k&lt;/code&gt; and &lt;code&gt;--cache-type-v&lt;/code&gt;, lower &lt;code&gt;--parallel&lt;/code&gt;, or adjust offload. Because each choice is explicit, it is easier to build separate long-context and high-concurrency profiles rather than force one compromise onto every model.&lt;/p&gt;

&lt;h3&gt;
  
  
  The API connects, but the answers are malformed
&lt;/h3&gt;

&lt;p&gt;This is often a template or parser problem, especially with new reasoning and tool-calling models. Verify that the GGUF contains the expected chat template and that the runtime recognizes the architecture. Compare a plain chat request before debugging the agent framework layered above it.&lt;/p&gt;

&lt;p&gt;On Ollama, inspect &lt;code&gt;ollama show --modelfile &amp;lt;name&amp;gt;&lt;/code&gt; and the reported capabilities. On llama.cpp, inspect the startup template messages, use &lt;code&gt;--jinja&lt;/code&gt;, and test &lt;code&gt;/v1/chat/completions&lt;/code&gt; directly. Pin the working runtime version before changing another variable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Requests become slow after switching models
&lt;/h3&gt;

&lt;p&gt;Ollama may need to unload one model and load another, so separate queue time from generation time. Preload the important model with an empty request and set an intentional &lt;code&gt;keep_alive&lt;/code&gt; value rather than depending on the five-minute default.&lt;/p&gt;

&lt;p&gt;A dedicated llama.cpp process avoids surprise switching because its model stays resident. If you adopt router mode, model loading becomes dynamic again — and every switch between two different models is a full unload-and-reload with no warm pool — so monitor load state and cold-start latency just as you would with Ollama.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security is not a differentiator unless you configure it
&lt;/h2&gt;

&lt;p&gt;Both servers bind to localhost by default, which is the right workstation behavior. Changing the host to &lt;code&gt;0.0.0.0&lt;/code&gt; turns a private local inference service into a network service, and neither product should be exposed to the public Internet merely because a firewall rule happened to allow it.&lt;/p&gt;

&lt;p&gt;llama.cpp can enforce API keys and can terminate TLS, although a reverse proxy is still useful for policy, rate limits, and logs. Ollama's local API does not require an API key; put it behind an authenticated proxy or private network boundary if remote clients need access. If you do need remote access to Ollama, the &lt;a href="https://www.glukhov.org/llm-hosting/ollama/ollama-behind-reverse-proxy/" rel="noopener noreferrer"&gt;Ollama behind a reverse proxy guide&lt;/a&gt; covers the Caddy and Nginx setup with streaming and timeout checks. Tool-capable models increase the consequence of exposing the surrounding application, even when the inference server itself does not execute the tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to keep Ollama
&lt;/h2&gt;

&lt;p&gt;Keep Ollama when its automation removes more work than it hides. It is especially strong for shared developer workstations, local desktop applications, demonstrations, coding tools, and small services that rotate among several popular models.&lt;/p&gt;

&lt;p&gt;Ollama is also the better default when you want colleagues to reproduce a named configuration without learning llama.cpp flags. A &lt;code&gt;Modelfile&lt;/code&gt;, model tag, and two commands are a useful operational contract. The &lt;a href="https://www.glukhov.org/llm-hosting/ollama/ollama-cheatsheet/" rel="noopener noreferrer"&gt;Ollama cheatsheet&lt;/a&gt; covers that daily workflow in more detail.&lt;/p&gt;

&lt;p&gt;Do not migrate merely because direct llama.cpp looks more technical. If your model fits, the API behaves correctly, latency is stable, and you do not need a missing control, replacing Ollama creates maintenance without creating capability.&lt;/p&gt;

&lt;p&gt;One caveat worth tracking over time: Ollama's own product direction has started drifting toward centralized infrastructure. &lt;strong&gt;Ollama Turbo&lt;/strong&gt; is a sign-in-gated cloud acceleration service layered on top of what was originally a local-first, privacy-first tool, and it is not the only recent change that trades local control for a hosted convenience layer. If the reason you chose Ollama in the first place was to avoid sending prompts to someone else's servers, that reasoning deserves a periodic re-check rather than a one-time decision — see &lt;a href="https://www.glukhov.org/llm-hosting/ollama/ollama-enshittification/" rel="noopener noreferrer"&gt;Ollama Enshittification: The Early Signs&lt;/a&gt; for the specific changes and what to watch for. Direct llama.cpp has no equivalent hosted upsell to drift toward, which is itself a data point when you are weighing long-term control against short-term convenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to move to llama-server
&lt;/h2&gt;

&lt;p&gt;Move to direct &lt;code&gt;llama-server&lt;/code&gt; when one or more of these statements are true:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You need a new llama.cpp feature or model fix before it reaches Ollama.&lt;/li&gt;
&lt;li&gt;You must pin an exact llama.cpp commit and backend build.&lt;/li&gt;
&lt;li&gt;Different models need different K and V cache types or slot layouts.&lt;/li&gt;
&lt;li&gt;You need deliberate multi-GPU placement rather than automatic selection.&lt;/li&gt;
&lt;li&gt;You are testing speculative decoding, MTP, LoRA scales, prompt caching, or uncommon samplers.&lt;/li&gt;
&lt;li&gt;Native metrics, slot state, or server internals are required for diagnosis.&lt;/li&gt;
&lt;li&gt;You want the original GGUF files to remain the authoritative model catalog.&lt;/li&gt;
&lt;li&gt;A model should remain resident for the lifetime of one supervised process.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cleanest migration trigger is repeated inspection. If every incident starts with discovering what Ollama chose before you can diagnose the model, make those choices explicit in a llama.cpp service definition.&lt;/p&gt;

&lt;h2&gt;
  
  
  A low-risk migration from Ollama to llama.cpp
&lt;/h2&gt;

&lt;p&gt;Do not begin by reproducing every Ollama feature. Migrate one model and one client, preserve the current endpoint until the comparison is complete, and keep the same GGUF if possible.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Record &lt;code&gt;ollama --version&lt;/code&gt;, &lt;code&gt;ollama show &amp;lt;model&amp;gt;&lt;/code&gt;, &lt;code&gt;ollama show --modelfile &amp;lt;model&amp;gt;&lt;/code&gt;, and &lt;code&gt;ollama ps&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Locate or download the equivalent GGUF and any multimodal projector.&lt;/li&gt;
&lt;li&gt;Start one &lt;code&gt;llama-server&lt;/code&gt; with an explicit alias, context, GPU offload, cache types, and slot count.&lt;/li&gt;
&lt;li&gt;Send a plain chat request, a structured-output request, and a tool call directly to each API.&lt;/li&gt;
&lt;li&gt;Test the real client, including streaming and error handling.&lt;/li&gt;
&lt;li&gt;Compare cold-load latency, warm prompt speed, generation speed, VRAM, and answer format.&lt;/li&gt;
&lt;li&gt;Only then replace the service URL or add a proxy in front of both backends.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For a simple OpenAI-style smoke test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://127.0.0.1:8080/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "qwen3-8b",
    "messages": [
      {"role": "user", "content": "Return exactly: runtime-ok"}
    ],
    "temperature": 0,
    "max_tokens": 16
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then verify the service itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# llama.cpp&lt;/span&gt;
llama-server &lt;span class="nt"&gt;--version&lt;/span&gt;
llama-server &lt;span class="nt"&gt;--list-devices&lt;/span&gt;
curl http://127.0.0.1:8080/health
curl http://127.0.0.1:8080/v1/models

&lt;span class="c"&gt;# Ollama&lt;/span&gt;
ollama &lt;span class="nt"&gt;--version&lt;/span&gt;
ollama ps
curl http://127.0.0.1:11434/api/ps
curl http://127.0.0.1:11434/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the application depends on the Ollama-native &lt;code&gt;/api/chat&lt;/code&gt; response shape, changing the base URL is not enough. Either migrate the client to an OpenAI-compatible route first or add an adapter. The broader &lt;a href="https://www.glukhov.org/llm-hosting/comparisons/ollama-to-vllm-migration/" rel="noopener noreferrer"&gt;Ollama to vLLM migration guide&lt;/a&gt; discusses the same contract-first principle for a larger runtime jump.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final verdict
&lt;/h2&gt;

&lt;p&gt;Ollama is the better local model appliance. It provides a model catalog, reproducible recipes, sensible automatic placement, convenient APIs, and lifecycle management without demanding that every user become an inference operator.&lt;/p&gt;

&lt;p&gt;llama.cpp is the better precision instrument. &lt;code&gt;llama-server&lt;/code&gt; exposes enough of the execution plan to make constrained VRAM, long context, unusual hardware, new model support, and controlled experiments understandable rather than mysterious.&lt;/p&gt;

&lt;p&gt;For most people, the right sequence is not Ollama or llama.cpp forever. Start with Ollama, learn which constraints actually matter, and move the affected workload to direct llama.cpp when you can name the control you need. That is a much stronger reason than chasing a benchmark measured under someone else's defaults.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/ggml-org/llama.cpp" rel="noopener noreferrer"&gt;llama.cpp project and supported backends&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md" rel="noopener noreferrer"&gt;llama-server features, flags, endpoints, and router mode&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/faq" rel="noopener noreferrer"&gt;Ollama FAQ: context, keep-alive, concurrency, and KV cache&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/modelfile" rel="noopener noreferrer"&gt;Ollama Modelfile reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/import" rel="noopener noreferrer"&gt;Importing GGUF and Safetensors models into Ollama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/api/openai-compatibility" rel="noopener noreferrer"&gt;Ollama OpenAI API compatibility&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ollama.com/blog/improved-performance-and-model-support-with-gguf" rel="noopener noreferrer"&gt;Ollama 0.30 GGUF, NVIDIA, and Vulkan changes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ollama/ollama/blob/main/llama/README.md" rel="noopener noreferrer"&gt;How Ollama pins and patches llama.cpp&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/gpu" rel="noopener noreferrer"&gt;Ollama hardware support&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/context-length" rel="noopener noreferrer"&gt;Ollama context length defaults&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ollama/ollama/releases" rel="noopener noreferrer"&gt;Ollama releases (GitHub)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>llamacpp</category>
      <category>ollama</category>
      <category>llm</category>
      <category>selfhosting</category>
    </item>
    <item>
      <title>ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Sat, 12 Sep 2026 11:42:17 +0000</pubDate>
      <link>https://dev.to/rosgluk/rocm-vs-vulkan-for-amd-local-llm-hosting-2026-guide-5c70</link>
      <guid>https://dev.to/rosgluk/rocm-vs-vulkan-for-amd-local-llm-hosting-2026-guide-5c70</guid>
      <description>&lt;p&gt;ROCm and Vulkan both accelerate AMD GPUs for local LLM hosting, but they are not interchangeable. The right choice depends on the engine, GPU, and workload.&lt;/p&gt;

&lt;p&gt;In local LLM hosting the two backends sit at different layers. ROCm is AMD's compute platform under PyTorch, vLLM, and SGLang, while Vulkan is a portable GPU API that llama.cpp-class engines use to run quantized models across a wide range of hardware.&lt;/p&gt;

&lt;p&gt;This guide compares the two engine by engine — llama.cpp, Ollama, LM Studio, vLLM, SGLang, TGI, and LocalAI — with build commands, device checks, and failure modes that masquerade as performance problems. If you are new to the hosting landscape, start with the &lt;a href="https://www.glukhov.org/llm-hosting/" rel="noopener noreferrer"&gt;LLM Hosting overview&lt;/a&gt;, which maps the tool families this article drills into.&lt;/p&gt;

&lt;h2&gt;
  
  
  ROCm vs Vulkan: the short answer
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Recommended starting point&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;llama.cpp with GGUF on Linux&lt;/td&gt;
&lt;td&gt;Vulkan&lt;/td&gt;
&lt;td&gt;Small install surface, broad GPU coverage, and easy rollback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llama.cpp on a supported RDNA 3 or RDNA 4 GPU&lt;/td&gt;
&lt;td&gt;Benchmark both&lt;/td&gt;
&lt;td&gt;Kernel performance changes with model shape, quantization, context, and build&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ollama on a listed AMD GPU&lt;/td&gt;
&lt;td&gt;ROCm first, verify Vulkan too&lt;/td&gt;
&lt;td&gt;Ollama supports both, but backend choice is less explicit than in bare llama.cpp&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LM Studio on a desktop AMD GPU&lt;/td&gt;
&lt;td&gt;Vulkan first&lt;/td&gt;
&lt;td&gt;Runtime switching makes comparison easy and avoids a system-wide compute stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vLLM or SGLang&lt;/td&gt;
&lt;td&gt;ROCm, but verify GPU-family kernel coverage&lt;/td&gt;
&lt;td&gt;These are PyTorch/HIP stacks; Vulkan is not an alternative backend, and brand-new architectures can still lack optimized kernels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TGI on supported Instinct hardware&lt;/td&gt;
&lt;td&gt;ROCm&lt;/td&gt;
&lt;td&gt;The published AMD container path targets MI210, MI250, and MI300 families&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Older or unlisted Radeon GPU&lt;/td&gt;
&lt;td&gt;Vulkan&lt;/td&gt;
&lt;td&gt;Vulkan drivers usually cover more graphics hardware than ROCm libraries do&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AMD Instinct server&lt;/td&gt;
&lt;td&gt;ROCm&lt;/td&gt;
&lt;td&gt;Multi-GPU compute, RCCL, framework kernels, and operational tooling live here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows local GGUF serving&lt;/td&gt;
&lt;td&gt;Vulkan&lt;/td&gt;
&lt;td&gt;It is generally the least restrictive route for llama.cpp-class runtimes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ryzen AI Max or other large-memory APU&lt;/td&gt;
&lt;td&gt;Vulkan first, then ROCm if required&lt;/td&gt;
&lt;td&gt;Both can work, but shared memory and kernel support need workload-specific testing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This table is a starting policy, not a benchmark result. A backend that detects the GPU but leaves some operations on the CPU can look healthy while performing badly, so every final decision needs log inspection and an end-to-end prompt test.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ROCm and Vulkan actually are
&lt;/h2&gt;

&lt;h3&gt;
  
  
  ROCm is a compute platform
&lt;/h3&gt;

&lt;p&gt;ROCm includes the HIP runtime, compiler, math libraries, collective communication, profilers, and framework packages needed to run AMD compute workloads. It is the AMD-side foundation beneath PyTorch builds and engines such as vLLM and SGLang, and it can also accelerate llama.cpp through its HIP backend.&lt;/p&gt;

&lt;p&gt;That breadth is ROCm's advantage and its cost. The host driver, GPU target, user-space libraries, framework wheel, kernel version, and container image must form a compatible set; when they do, ROCm provides far more than token generation through one local executable.&lt;/p&gt;

&lt;p&gt;ROCm 10.0.0, released on August 26, 2026, is built on TheRock (AMD's build and release system since ROCm 7.14), validates PyTorch 2.13, vLLM 0.27, and SGLang 0.5.15, and formally adds RDNA 4 support for &lt;code&gt;gfx1200&lt;/code&gt; (RX 9060/9060 XT/9050) and &lt;code&gt;gfx1201&lt;/code&gt; (RX 9070/9070 XT/9070 GRE, Radeon AI PRO R9700 series). The &lt;a href="https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html" rel="noopener noreferrer"&gt;ROCm compatibility matrix&lt;/a&gt; is still the authority for an exact GPU and operating-system combination, not a forum post that happens to use the same marketing family.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vulkan is a portable GPU interface
&lt;/h3&gt;

&lt;p&gt;Vulkan is a graphics and compute API implemented by a GPU driver. In local LLM hosting, it normally means an inference engine ships or compiles compute shaders that execute through a Vulkan implementation such as Mesa RADV on Linux or the vendor driver on Windows.&lt;/p&gt;

&lt;p&gt;Vulkan does not provide a drop-in PyTorch platform comparable to ROCm. Its practical strength is narrower and useful: a llama.cpp-style engine can use the same backend design on AMD, Intel, Nvidia, and other Vulkan-capable hardware without installing a vendor-specific machine-learning stack.&lt;/p&gt;

&lt;p&gt;This distinction explains most of the decision. If the application offers only a HIP or PyTorch path, Vulkan cannot rescue it; if the application is already based on llama.cpp and GGUF, installing the whole ROCm stack may solve a problem you did not have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Engine support matrix in 2026
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th&gt;ROCm or HIP&lt;/th&gt;
&lt;th&gt;Vulkan&lt;/th&gt;
&lt;th&gt;Typical model format&lt;/th&gt;
&lt;th&gt;Practical note&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;llama.cpp / llama-server&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;GGUF&lt;/td&gt;
&lt;td&gt;Best platform for a controlled A/B backend test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ollama&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Managed GGUF-derived models&lt;/td&gt;
&lt;td&gt;Convenient, but backend selection and packaging are abstracted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LM Studio&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;GGUF and product-managed formats&lt;/td&gt;
&lt;td&gt;Selectable runtimes make desktop testing approachable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vLLM&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Safetensors and supported quantizations&lt;/td&gt;
&lt;td&gt;Use AMD's matched ROCm image or wheel set; verify GPU-family kernel coverage first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SGLang&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Safetensors and supported quantizations&lt;/td&gt;
&lt;td&gt;ROCm is part of the deployment architecture&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TGI&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Safetensors and supported quantizations&lt;/td&gt;
&lt;td&gt;Published AMD validation remains Instinct-focused&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LocalAI&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Backend-dependent, commonly GGUF&lt;/td&gt;
&lt;td&gt;Uses different ROCm and Vulkan container images&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;ROCm does not imply Safetensors, and Vulkan does not formally imply GGUF. The useful association comes from engines: llama.cpp can read the same GGUF with either its HIP or Vulkan build, while PyTorch-native servers use ROCm and generally consume Hugging Face model repositories.&lt;/p&gt;

&lt;p&gt;That makes model inventory an architectural constraint. A library of carefully selected GGUF quantizations points naturally toward llama-server, Ollama, LM Studio, or LocalAI; a deployment built around tensor parallelism, continuous batching, and framework-native weights points toward ROCm with vLLM or SGLang. For the wider engine landscape beyond AMD backends — API maturity, tool calling, and production readiness across a dozen tools — see our &lt;a href="https://www.glukhov.org/llm-hosting/comparisons/hosting-llms-ollama-localai-jan-lmstudio-vllm-comparison/" rel="noopener noreferrer"&gt;comparison of Ollama, vLLM, LM Studio, LocalAI and other local LLM hosting tools&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  llama.cpp: the cleanest ROCm vs Vulkan comparison
&lt;/h2&gt;

&lt;p&gt;llama.cpp exposes both backends without changing the model file or HTTP client. This is the fairest place to compare ROCm and Vulkan because the tokenizer, sampling settings, chat template, quantization, and server behavior can remain fixed.&lt;/p&gt;

&lt;p&gt;The current &lt;a href="https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md" rel="noopener noreferrer"&gt;llama.cpp build documentation&lt;/a&gt; uses &lt;code&gt;GGML_HIP&lt;/code&gt; for ROCm and &lt;code&gt;GGML_VULKAN&lt;/code&gt; for Vulkan. Old articles that recommend &lt;code&gt;GGML_ROCM&lt;/code&gt; or the removed Makefile flags should not be trusted without checking the project's current CMake options.&lt;/p&gt;

&lt;h3&gt;
  
  
  Build the Vulkan backend on Ubuntu
&lt;/h3&gt;

&lt;p&gt;Install the Vulkan headers, shader compiler, and SPIR-V headers, then verify that the driver can enumerate the intended GPU:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;apt-get update
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt-get &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; libvulkan-dev glslc spirv-headers vulkan-tools

vulkaninfo &lt;span class="nt"&gt;--summary&lt;/span&gt;

git clone https://github.com/ggml-org/llama.cpp
&lt;span class="nb"&gt;cd &lt;/span&gt;llama.cpp
cmake &lt;span class="nt"&gt;-S&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;-B&lt;/span&gt; build-vulkan &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DGGML_VULKAN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ON &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DCMAKE_BUILD_TYPE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Release
cmake &lt;span class="nt"&gt;--build&lt;/span&gt; build-vulkan &lt;span class="nt"&gt;--config&lt;/span&gt; Release &lt;span class="nt"&gt;-j&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a mixed iGPU and discrete-GPU system, enumeration order deserves attention. &lt;code&gt;GGML_VK_VISIBLE_DEVICES&lt;/code&gt; can restrict llama.cpp to a specific Vulkan device, and the startup log should name the selected card rather than merely report that a Vulkan device exists.&lt;/p&gt;

&lt;h3&gt;
  
  
  Build the ROCm or HIP backend
&lt;/h3&gt;

&lt;p&gt;First confirm that ROCm identifies the GPU and reports the expected &lt;code&gt;gfx&lt;/code&gt; target. The target can be omitted to build for the GPUs in the current system, but pinning it reduces compilation work when you know the deployment hardware.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rocminfo | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'Name:.*gfx'&lt;/span&gt; | &lt;span class="nb"&gt;head
&lt;/span&gt;hipconfig &lt;span class="nt"&gt;--full&lt;/span&gt;

git clone https://github.com/ggml-org/llama.cpp
&lt;span class="nb"&gt;cd &lt;/span&gt;llama.cpp

&lt;span class="nv"&gt;HIPCXX&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;hipconfig &lt;span class="nt"&gt;-l&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;/clang"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;HIP_PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;hipconfig &lt;span class="nt"&gt;-R&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
cmake &lt;span class="nt"&gt;-S&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;-B&lt;/span&gt; build-rocm &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DGGML_HIP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ON &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DGPU_TARGETS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gfx1201 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-DCMAKE_BUILD_TYPE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Release
cmake &lt;span class="nt"&gt;--build&lt;/span&gt; build-rocm &lt;span class="nt"&gt;--config&lt;/span&gt; Release &lt;span class="nt"&gt;-j&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace &lt;code&gt;gfx1201&lt;/code&gt; with the target reported for the actual card — that value maps to the RX 9070/9070 XT/9070 GRE and Radeon AI PRO R9700 family in RDNA 4, while &lt;code&gt;gfx1200&lt;/code&gt; covers the RX 9060 series and &lt;code&gt;gfx1100&lt;/code&gt;/&lt;code&gt;gfx1101&lt;/code&gt;/&lt;code&gt;gfx1102&lt;/code&gt; cover RDNA 3's RX 7900/7800/7700/7600 lines. Do not copy &lt;code&gt;HSA_OVERRIDE_GFX_VERSION&lt;/code&gt; into a production service merely because it helped someone boot an unsupported GPU; an override can make code load, but it does not turn that hardware into a validated platform.&lt;/p&gt;

&lt;h3&gt;
  
  
  Benchmark the same workload, not two defaults
&lt;/h3&gt;

&lt;p&gt;Use one GGUF file, the same flash-attention setting, the same layer offload, and repeated runs. Prompt processing (&lt;code&gt;pp&lt;/code&gt;) and token generation (&lt;code&gt;tg&lt;/code&gt;) exercise the system differently, while a long-context server also adds KV-cache allocation and memory pressure that a short synthetic benchmark will miss.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/srv/models/model.gguf

./build-vulkan/bin/llama-bench &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-ngl&lt;/span&gt; 999 &lt;span class="nt"&gt;-fa&lt;/span&gt; 1 &lt;span class="nt"&gt;-p&lt;/span&gt; 512 &lt;span class="nt"&gt;-n&lt;/span&gt; 128 &lt;span class="nt"&gt;-r&lt;/span&gt; 5

./build-rocm/bin/llama-bench &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-m&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-ngl&lt;/span&gt; 999 &lt;span class="nt"&gt;-fa&lt;/span&gt; 1 &lt;span class="nt"&gt;-p&lt;/span&gt; 512 &lt;span class="nt"&gt;-n&lt;/span&gt; 128 &lt;span class="nt"&gt;-r&lt;/span&gt; 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Community results illustrate why a universal winner is misleading, and RDNA 4's &lt;code&gt;gfx1201&lt;/code&gt; target is the clearest recent example. In one same-machine RX 9070 XT submission, Vulkan reached about 143 token/s against ROCm's 128 token/s on the 7B Q4_0 generation test, but later submissions showed smaller gaps as builds changed; the &lt;a href="https://github.com/ggml-org/llama.cpp/discussions/10879" rel="noopener noreferrer"&gt;Vulkan&lt;/a&gt; and &lt;a href="https://github.com/ggml-org/llama.cpp/discussions/15021" rel="noopener noreferrer"&gt;ROCm&lt;/a&gt; discussions also contain large differences in prompt-processing results and test conditions. A separate, more detailed OpenBenchmarking.org run on an RX 9070 XT with llama.cpp b6401 found Vulkan ahead on decode across several 8B-class models (Qwen3-8B-Q8_0, Llama-3.1-Tulu-3-8B-Q8_0) but behind HIP on prompt processing at longer prompt lengths — the two backends trade the lead depending on which phase you measure.&lt;/p&gt;

&lt;p&gt;The gap can also run the other way, and by a lot, for specific model shapes. An open llama.cpp issue documents Vulkan on &lt;code&gt;gfx1201&lt;/code&gt; becoming 4.7–6.7x slower than HIP on token generation once a model's hidden size reaches 4096 or above (effective decode bandwidth collapsing to roughly 70–100 GB/s on a 640 GB/s card), while a smaller 4B model with hidden size 2560 shows no such regression on either backend. Treat every number here as a snapshot of one build, one driver, and one model shape — not as a rule that generalizes across quant type, model architecture, flash attention, batch sizes, driver version, thermal state, or the llama.cpp commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama on AMD: convenient, but verify the backend
&lt;/h2&gt;

&lt;p&gt;Ollama officially supports listed AMD GPUs through ROCm and now documents additional AMD coverage through Vulkan on Windows and Linux. Its current &lt;a href="https://docs.ollama.com/gpu" rel="noopener noreferrer"&gt;hardware support page&lt;/a&gt; says Vulkan is enabled by default when the backend is installed, supports &lt;code&gt;GGML_VK_VISIBLE_DEVICES&lt;/code&gt; for device selection, and can disable Vulkan with &lt;code&gt;OLLAMA_VULKAN=0&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is a meaningful improvement over the period when Vulkan advice depended on experimental builds. It also makes some older tutorials stale: setting an undocumented switch and assuming the service selected Vulkan is weaker evidence than reading the server log.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl edit ollama
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For diagnostics, add a drop-in rather than exporting variables only in an interactive shell:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Service]&lt;/span&gt;
&lt;span class="py"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"OLLAMA_DEBUG=1"&lt;/span&gt;
&lt;span class="py"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;"GGML_VK_VISIBLE_DEVICES=0"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then reload, restart, and inspect both process placement and discovery messages:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl daemon-reload
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl restart ollama
journalctl &lt;span class="nt"&gt;-u&lt;/span&gt; ollama &lt;span class="nt"&gt;-b&lt;/span&gt; &lt;span class="nt"&gt;--no-pager&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 200

ollama run qwen3:8b &lt;span class="s2"&gt;"Return exactly: backend test passed"&lt;/span&gt;
ollama ps
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for the named GPU, selected library, model allocation, and GPU percentage. A log that shows a discovery timeout followed by a successful HTTP response may mean Ollama quietly fell back to the CPU. For the everyday command set around this service, the &lt;a href="https://www.glukhov.org/llm-hosting/ollama/ollama-cheatsheet/" rel="noopener noreferrer"&gt;Ollama CLI cheatsheet&lt;/a&gt; is the quicker reference.&lt;/p&gt;

&lt;p&gt;Ollama is excellent when model acquisition and a stable local API matter more than backend control. If repeatable ROCm-versus-Vulkan testing is the objective, bare llama-server is the better instrument because the build directory makes the backend explicit.&lt;/p&gt;

&lt;h2&gt;
  
  
  vLLM and SGLang make ROCm the decision — but check GPU-family kernel coverage first
&lt;/h2&gt;

&lt;p&gt;vLLM and SGLang are not Vulkan applications. Their AMD paths sit on ROCm, PyTorch, and optimized HIP kernels, so choosing one of these engines has already selected the compute platform.&lt;/p&gt;

&lt;p&gt;AMD's current &lt;a href="https://rocm.docs.amd.com/projects/ai-ecosystem/en/latest/inference/vllm.html" rel="noopener noreferrer"&gt;vLLM on ROCm guide&lt;/a&gt; recommends a prebuilt container and publishes matched images for ROCm, PyTorch, Python, and vLLM. That coupling is useful: it replaces a large dependency-solving exercise with a versioned deployment unit.&lt;/p&gt;

&lt;p&gt;At the time of writing, AMD documents this ROCm 10 image for vLLM 0.27:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker pull &lt;span class="se"&gt;\&lt;/span&gt;
  rocm/vllm:rocm10.0.0_ubuntu24.04_py3.14_pytorch_2.12.0_vllm_0.27.0

docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--device&lt;/span&gt; /dev/kfd &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--device&lt;/span&gt; /dev/dri &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--group-add&lt;/span&gt; video &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ipc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;host &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--network&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;host &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cap-add&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;SYS_PTRACE &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--security-opt&lt;/span&gt; &lt;span class="nv"&gt;seccomp&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;unconfined &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; /srv/models:/app/models &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;HF_HOME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/app/models &lt;span class="se"&gt;\&lt;/span&gt;
  rocm/vllm:rocm10.0.0_ubuntu24.04_py3.14_pytorch_2.12.0_vllm_0.27.0 &lt;span class="se"&gt;\&lt;/span&gt;
  bash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use the image selected for the exact GPU family in AMD's documentation; RDNA and CDNA images have not always been interchangeable. For a production server, pin the full tag or digest and validate the host driver before blaming vLLM for an initialization failure.&lt;/p&gt;

&lt;p&gt;Brand-new GPU generations are the sharpest edge case, and RDNA 4 is a real, documented example rather than a theoretical risk. Independent testing on an RX 9070 XT (&lt;code&gt;gfx1201&lt;/code&gt;) in early 2026 found vLLM on ROCm 7.2 silently falling back to FP32 dequantization for FP8 model weights — because &lt;code&gt;gfx1201&lt;/code&gt; was not yet recognized in vLLM's platform detection — which bypassed the GPU's matrix accelerators entirely and produced only 48 tokens/s, versus 62 tokens/s from llama-server on Vulkan running a GGUF quantization of a comparable model on the same card. The lesson generalizes: a ROCm/ PyTorch stack can &lt;em&gt;load&lt;/em&gt; successfully on a new architecture and still run an unoptimized fallback path with no error message. Always confirm which kernel path actually executed (via &lt;code&gt;rocprof&lt;/code&gt;, vendor profiling notes, or a known-good throughput baseline for the GPU) before trusting a single "it started fine" result on hardware that shipped within the last release cycle or two.&lt;/p&gt;

&lt;p&gt;The reason to accept ROCm's larger operational surface, once kernel coverage is confirmed, is throughput architecture — not merely a few more token/s in a single-user test. Continuous batching, framework-native quantization, tensor parallelism, scheduler behavior, and the surrounding PyTorch tooling are the real case for &lt;a href="https://www.glukhov.org/llm-hosting/vllm/vllm-quickstart/" rel="noopener noreferrer"&gt;moving to vLLM&lt;/a&gt;, and if you are weighing whether that move is justified at all, our &lt;a href="https://www.glukhov.org/llm-hosting/comparisons/ollama-to-vllm-migration/" rel="noopener noreferrer"&gt;Ollama to vLLM migration guide&lt;/a&gt; lists the workload signals.&lt;/p&gt;

&lt;h2&gt;
  
  
  TGI: ROCm support with a narrower target
&lt;/h2&gt;

&lt;p&gt;Hugging Face documents an AMD image for Text Generation Inference, but its published validation is centered on Instinct MI210, MI250, and MI300 hardware. The &lt;a href="https://huggingface.co/docs/text-generation-inference/main/installation_amd" rel="noopener noreferrer"&gt;TGI AMD guide&lt;/a&gt; uses the &lt;code&gt;3.3.5-rocm&lt;/code&gt; image and lists unsupported ROCm features, so it should not be generalized into a promise for every Radeon card. Our &lt;a href="https://www.glukhov.org/llm-hosting/tgi/" rel="noopener noreferrer"&gt;TGI install guide&lt;/a&gt; covers that ROCm image setup in more detail.&lt;/p&gt;

&lt;p&gt;There is no Vulkan TGI path to compare. If TGI is a fixed requirement, choose supported ROCm hardware and reproduce the documented container; if the engine is negotiable, current vLLM and SGLang support deserves evaluation before beginning a new AMD deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  LM Studio: switch runtimes instead of rebuilding
&lt;/h2&gt;

&lt;p&gt;LM Studio packages multiple inference runtimes and exposes runtime management through the &lt;code&gt;lms&lt;/code&gt; command. Its &lt;a href="https://lmstudio.ai/docs/cli/runtime/runtime" rel="noopener noreferrer"&gt;runtime documentation&lt;/a&gt; supports listing, downloading, selecting, updating, and removing runtimes, which makes ROCm-versus-Vulkan experiments accessible without maintaining separate source trees.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;lms runtime &lt;span class="nb"&gt;ls
&lt;/span&gt;lms runtime get
lms runtime &lt;span class="k"&gt;select&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the same GGUF with the same context length, GPU offload, flash-attention setting, and prompt. Compare time to first token, generation rate, load time, and peak memory rather than judging a backend from one short chat response.&lt;/p&gt;

&lt;p&gt;Runtime packaging does not eliminate backend-specific faults. For example, a 2026 &lt;a href="https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1916" rel="noopener noreferrer"&gt;LM Studio issue on an R9700&lt;/a&gt; reported a large model hanging near the end of a ROCm load while the Vulkan runtime loaded it, while a separate &lt;a href="https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/2133" rel="noopener noreferrer"&gt;Vulkan memory-headroom issue&lt;/a&gt; described the opposite outcome near full VRAM. These are individual reports, but together they make the right operational point: keep a fallback runtime and leave memory headroom.&lt;/p&gt;

&lt;h2&gt;
  
  
  LocalAI: choose the image as well as the backend
&lt;/h2&gt;

&lt;p&gt;LocalAI provides separate ROCm or &lt;code&gt;hipblas&lt;/code&gt; and Vulkan container variants. Its &lt;a href="https://localai.io/docs/features/gpu-acceleration/" rel="noopener noreferrer"&gt;GPU acceleration guide&lt;/a&gt; documents &lt;code&gt;gpu-hipblas&lt;/code&gt; images for AMD compute and &lt;code&gt;gpu-vulkan&lt;/code&gt; images for the portable path, so a container tag copied from a CUDA guide will not discover the correct backend by magic. The &lt;a href="https://www.glukhov.org/llm-hosting/local-ai/" rel="noopener noreferrer"&gt;LocalAI quickstart&lt;/a&gt; covers the general setup; the backend-specific container choice is what this section adds.&lt;/p&gt;

&lt;p&gt;The ROCm container needs &lt;code&gt;/dev/kfd&lt;/code&gt; and &lt;code&gt;/dev/dri&lt;/code&gt;, while Vulkan normally needs the appropriate render device under &lt;code&gt;/dev/dri&lt;/code&gt;. Pin a release tag for a real service; &lt;code&gt;latest&lt;/code&gt; and &lt;code&gt;master&lt;/code&gt; are useful for diagnosis, but they make rollback and performance comparison unnecessarily vague.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# ROCm or HIP image&lt;/span&gt;
docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--device&lt;/span&gt; /dev/kfd &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--device&lt;/span&gt; /dev/dri &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 &lt;span class="se"&gt;\&lt;/span&gt;
  quay.io/go-skynet/local-ai:v4.8.0-gpu-hipblas

&lt;span class="c"&gt;# Vulkan image&lt;/span&gt;
docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;-it&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--device&lt;/span&gt; /dev/dri &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 &lt;span class="se"&gt;\&lt;/span&gt;
  localai/localai:v4.8.0-gpu-vulkan
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tag examples reflect the documentation available at publication time; confirm the current registry names before automating a pull. More importantly, do not infer acceleration from the container name alone — inspect LocalAI's debug log and watch GPU utilization during a request.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in ROCm 10 packaging
&lt;/h2&gt;

&lt;p&gt;ROCm 10 is not just another minor package update. AMD's &lt;a href="https://rocm.docs.amd.com/en/latest/about/transition-guide-TheRock.html" rel="noopener noreferrer"&gt;TheRock transition guide&lt;/a&gt; says ROCm Core SDK packages now use the &lt;code&gt;amdrocm-&lt;/code&gt; prefix, the versioned installation root is &lt;code&gt;/opt/rocm/core-10.0&lt;/code&gt;, and several legacy packages have been consolidated.&lt;/p&gt;

&lt;p&gt;This is why a command from an older ROCm article may return "package not found" even on a correctly configured repository. For example, HIPCC now comes from &lt;code&gt;amdrocm-llvm&lt;/code&gt;, BLAS components are combined in &lt;code&gt;amdrocm-blas&lt;/code&gt;, and a full system installation can use an all-architecture or GPU-family-specific Core SDK meta-package — RDNA 4 cards use the &lt;code&gt;gfx120X-all&lt;/code&gt; family tag (package suffix &lt;code&gt;-gfx1200-gfx1201&lt;/code&gt;), which is worth knowing before you go hunting for a &lt;code&gt;gfx1201&lt;/code&gt;-only package name that does not exist.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;amdrocm&lt;/code&gt; meta-package configures alternatives and compatibility symlinks under &lt;code&gt;/opt/rocm&lt;/code&gt;. A minimal or custom installation may not provide the same paths, so build scripts that hard-code &lt;code&gt;/opt/rocm/bin/hipcc&lt;/code&gt; should either use &lt;code&gt;hipconfig&lt;/code&gt; or set &lt;code&gt;ROCM_PATH&lt;/code&gt; explicitly.&lt;/p&gt;

&lt;p&gt;Two diagnostic changes are easy to miss. ROCm SMI has been removed in favor of AMD SMI, and ROCm Bandwidth Test reached end of life; scripts that call &lt;code&gt;rocm-smi&lt;/code&gt; or &lt;code&gt;rocm-bandwidth-test&lt;/code&gt; need to move to &lt;code&gt;amd-smi&lt;/code&gt; and AMD's replacement tools rather than reinstall arbitrary legacy packages.&lt;/p&gt;

&lt;h3&gt;
  
  
  Containers still depend on the host
&lt;/h3&gt;

&lt;p&gt;A ROCm container carries user-space libraries, not a replacement kernel driver. The host must expose &lt;code&gt;/dev/kfd&lt;/code&gt; and &lt;code&gt;/dev/dri&lt;/code&gt;, its driver must be compatible with the container stack, and the service user needs permission to open those devices.&lt;/p&gt;

&lt;p&gt;Vulkan containers have a similar boundary around the host Vulkan driver and render node. Packaging is lighter, but an incorrect ICD, missing render-group membership, or an accidentally selected iGPU can still turn a working container image into a CPU-bound or unstable service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Discrete GPUs, APUs, and older Radeon cards
&lt;/h2&gt;

&lt;h3&gt;
  
  
  RDNA 3 and RDNA 4 discrete GPUs
&lt;/h3&gt;

&lt;p&gt;Current Radeon RX 7000, RX 9000, and Radeon AI Pro models have the strongest case for testing both llama.cpp backends. ROCm support is now explicit for many &lt;code&gt;gfx110x&lt;/code&gt; and &lt;code&gt;gfx120x&lt;/code&gt; targets, while Vulkan through a current Mesa RADV or Windows vendor driver is mature enough to be a primary route rather than a desperate fallback. For the hardware side of that decision — VRAM, bandwidth, power, and pricing across vendors — see our &lt;a href="https://www.glukhov.org/hardware/ai/gpu-comparison-ai-workloads-2026-nvidia-amd-intel/" rel="noopener noreferrer"&gt;GPU comparison for AI workloads in 2026&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Do not convert a 7B benchmark into a rule for a 27B dense model or a mixture-of-experts model. Matrix shapes, active parameters, quantized kernels, prompt length, and memory pressure can change the order, and backend performance has moved substantially between llama.cpp revisions — the &lt;code&gt;gfx1201&lt;/code&gt; hidden-size regression noted above is a concrete case of exactly this kind of shift.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ryzen APUs and shared memory
&lt;/h3&gt;

&lt;p&gt;Large-memory Ryzen AI Max systems are unusually interesting because the GPU can access a much larger shared-memory pool than a normal discrete consumer card offers. ROCm 10 lists current Ryzen AI families, while Vulkan-capable llama.cpp runtimes can also use the iGPU without building a PyTorch environment.&lt;/p&gt;

&lt;p&gt;Capacity is not bandwidth. A model fitting into 64 GB or 96 GB of allocated shared memory does not mean it will decode like a 32 GB discrete card, and aggressive context allocation can starve the operating system even when an application reports ample GPU memory. The same VRAM-budget discipline that applies to discrete NVIDIA and AMD cards applies here too — see &lt;a href="https://www.glukhov.org/llm-performance/optimization/kv-cache-16gb-long-context/" rel="noopener noreferrer"&gt;KV Cache on 16 GB GPUs&lt;/a&gt; for the underlying budget math, which is backend-agnostic.&lt;/p&gt;

&lt;p&gt;Mixed iGPU and dGPU machines need explicit device selection. A recent llama.cpp report described excessive system-memory reservation when an unused iGPU remained visible beside an R9700; it is an unconfirmed issue, but it is a good reason to expose only the device the service is intended to use.&lt;/p&gt;

&lt;h3&gt;
  
  
  Older and unsupported Radeon hardware
&lt;/h3&gt;

&lt;p&gt;Vulkan is usually the first route for an older Radeon because graphics-driver coverage is broader than ROCm's supported compute-target set. ROCm-based projects also note that newer rocBLAS releases removed kernels for some older targets, so forcing a nearby &lt;code&gt;gfx&lt;/code&gt; value cannot restore code that is no longer shipped.&lt;/p&gt;

&lt;p&gt;An override is acceptable for a laboratory experiment with clear failure expectations. It is a poor foundation for an unattended API, because the next ROCm or application update can replace a tolerated mismatch with a startup failure or incorrect result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Linux vs Windows for AMD LLM backends
&lt;/h2&gt;

&lt;p&gt;Linux is the natural ROCm host for production inference. It offers the broadest engine support, established container device mapping, current Mesa Vulkan drivers, and the operational tools expected by vLLM and SGLang deployments.&lt;/p&gt;

&lt;p&gt;Windows has genuine ROCm support for listed hardware, but the application ecosystem remains narrower. For desktop GGUF inference through llama.cpp, Ollama, or LM Studio, Vulkan is usually the calmer starting point; use ROCm when the application provides a supported Windows path and a concrete feature or benchmark justifies it.&lt;/p&gt;

&lt;p&gt;WSL2 should be treated as a third platform, not a synonym for native Linux. Match AMD's documented Windows driver, WSL distribution, ROCm release, and framework package as one supported combination.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verification checklist before serving traffic
&lt;/h2&gt;

&lt;p&gt;Start below the application. If the driver cannot enumerate the correct device, changing model flags is only rearranging the symptom.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;lspci &lt;span class="nt"&gt;-nnk&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A3&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'VGA|Display'&lt;/span&gt;
&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; /dev/kfd /dev/dri/renderD&lt;span class="k"&gt;*&lt;/span&gt; 2&amp;gt;/dev/null
&lt;span class="nb"&gt;id&lt;/span&gt;

&lt;span class="c"&gt;# ROCm path&lt;/span&gt;
rocminfo | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'Marketing Name:|Name:.*gfx'&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 20
amd-smi list

&lt;span class="c"&gt;# Vulkan path&lt;/span&gt;
vulkaninfo &lt;span class="nt"&gt;--summary&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then verify the engine. The startup output must name ROCm or Vulkan, name the intended GPU, and report that model layers or tensors were placed on it; finally, GPU memory and utilization must rise while a request is running.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Observe an AMD GPU while another terminal sends requests&lt;/span&gt;
watch &lt;span class="nt"&gt;-n1&lt;/span&gt; amd-smi monitor

&lt;span class="c"&gt;# Basic OpenAI-compatible API check for llama-server&lt;/span&gt;
curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8080/v1/models
curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://127.0.0.1:8080/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "local-model",
    "messages": [{"role": "user", "content": "Return exactly: ready"}],
    "max_tokens": 8,
    "temperature": 0
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Record driver, runtime, engine commit or image digest, model file checksum, context, batch settings, and command line with every benchmark. Without that metadata, a token-per-second number is an anecdote that cannot survive the next upgrade — and as the vLLM-on-&lt;code&gt;gfx1201&lt;/code&gt; fallback and the Vulkan hidden-size regression above both show, a plausible-looking number can hide a silently unoptimized code path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure modes that look like backend performance
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Silent CPU fallback
&lt;/h3&gt;

&lt;p&gt;The server starts and answers correctly, but generation is unexpectedly slow and GPU utilization remains flat. Check discovery logs, device permissions, model offload, and container devices before tuning threads or sampling parameters.&lt;/p&gt;

&lt;h3&gt;
  
  
  Silent precision fallback (new hardware, new kernels)
&lt;/h3&gt;

&lt;p&gt;The server starts, GPU utilization looks reasonable, and there is no error — but the framework quietly dropped to an unoptimized numeric path because the GPU's compute capability or architecture string was not yet recognized. This is exactly what happened with vLLM's FP8 kernels on &lt;code&gt;gfx1201&lt;/code&gt;; the fix is to check the framework's own platform-detection code or issue tracker for your exact GPU string before trusting a single throughput number on a GPU generation released in the last release cycle or two.&lt;/p&gt;

&lt;h3&gt;
  
  
  The wrong GPU is selected
&lt;/h3&gt;

&lt;p&gt;A Ryzen desktop may expose an iGPU as Vulkan device 0 and a discrete Radeon as device 1. Restrict visible devices and confirm the full device name in the log; do not assume numbering is stable after a driver or BIOS change.&lt;/p&gt;

&lt;h3&gt;
  
  
  ROCm target mismatch
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;rocminfo&lt;/code&gt; reports one &lt;code&gt;gfx&lt;/code&gt; target while the application image contains kernels for another set. Use a matching image or rebuild for the exact target; reserve &lt;code&gt;HSA_OVERRIDE_GFX_VERSION&lt;/code&gt; for explicitly unsupported experiments.&lt;/p&gt;

&lt;h3&gt;
  
  
  Driver and user-space mismatch
&lt;/h3&gt;

&lt;p&gt;The container has current ROCm libraries but the host driver belongs to an older release stream. Timeouts during discovery, kernel launch errors, or a fall back to CPU are more likely than a clean message explaining the version boundary.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vulkan ICD confusion
&lt;/h3&gt;

&lt;p&gt;More than one Vulkan implementation is installed, and the loader selects an unexpected ICD. Inspect &lt;code&gt;vulkaninfo&lt;/code&gt;, remove accidental duplicates, or select the intended ICD and device explicitly rather than layering another SDK over the problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  VRAM estimates leave no operating margin
&lt;/h3&gt;

&lt;p&gt;The model appears to fit but fails during warmup, flash-attention setup, or the first long prompt. Leave several gigabytes of headroom on a large model, then reduce context or batch size before concluding that the backend cannot run the quantization.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical backend selection procedure
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: choose the serving behavior
&lt;/h3&gt;

&lt;p&gt;If the goal is one or two local users, GGUF files, and a simple OpenAI-compatible endpoint, start with llama-server, Ollama, or LM Studio. If the goal is continuous batching, high concurrency, framework-native models, or tensor parallelism, begin with vLLM or SGLang and accept ROCm as part of the design.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: check official hardware support
&lt;/h3&gt;

&lt;p&gt;Match the exact GPU target, operating system version, kernel, and driver in the current ROCm matrix. For Vulkan, confirm the intended GPU through &lt;code&gt;vulkaninfo&lt;/code&gt; and use a current driver rather than assuming that the presence of &lt;code&gt;libvulkan.so&lt;/code&gt; proves useful compute support. If the GPU is from the newest architecture generation, also check the specific framework's platform-detection code or open issues for that exact &lt;code&gt;gfx&lt;/code&gt; target — official support and optimized-kernel support are not always released together.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: establish the simplest working baseline
&lt;/h3&gt;

&lt;p&gt;For GGUF, Vulkan is normally that baseline because it changes fewer system components. For a PyTorch engine, use AMD's pinned ROCm container rather than assembling torch, Triton, AITER, and vLLM from unrelated latest versions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: benchmark production-shaped prompts
&lt;/h3&gt;

&lt;p&gt;Measure prompt processing, time to first token, decode rate, peak memory, and concurrent request behavior. Include the context and tool-calling pattern the real service will use; a 128-token microbenchmark does not predict a 100,000-token agent session.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: keep the fallback deployable
&lt;/h3&gt;

&lt;p&gt;Two llama.cpp build directories cost little compared with a day lost to a driver regression. Keep the last known-good container digest or runtime installed, and roll forward only after the candidate passes the same test set.&lt;/p&gt;

&lt;p&gt;The same procedure as a decision flow:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A["Choose serving engine"] --&amp;gt; B{"PyTorch/HIP engine?"}
    B -- Yes --&amp;gt; C["ROCm: pinned container&amp;lt;br&amp;gt;+ kernel-coverage check"]
    B -- No --&amp;gt; D{"GGUF on AMD GPU?"}
    D -- Yes --&amp;gt; E["Vulkan baseline"]
    E --&amp;gt; F{"Benchmark: ROCm wins&amp;lt;br&amp;gt;by a measurable margin?"}
    F -- Yes --&amp;gt; G["Switch to ROCm"]
    F -- No --&amp;gt; H["Keep Vulkan,&amp;lt;br&amp;gt;keep ROCm build as fallback"]&lt;/code&gt;&lt;/pre&gt;



&lt;h2&gt;
  
  
  Final verdict: ROCm or Vulkan for AMD LLM hosting?
&lt;/h2&gt;

&lt;p&gt;Vulkan is the best default for local GGUF inference when portability, setup speed, Windows support, or older Radeon coverage matters. It is no longer reasonable to describe it as inherently slow; on some recent Radeon and llama.cpp combinations it is the faster backend, and on others it is close enough that lower operational friction wins.&lt;/p&gt;

&lt;p&gt;ROCm is the correct choice when the engine is built around PyTorch, when AMD Instinct and multi-GPU compute are central, or when a tested HIP build wins the actual model workload. Its ecosystem is much stronger in 2026, but the new packaging and strict compatibility layers still reward pinned versions and disciplined verification — and on the newest RDNA generation specifically, verifying that the optimized kernel path actually ran is not optional.&lt;/p&gt;

&lt;p&gt;For a supported Radeon workstation, my recommendation is deliberately unromantic: install Vulkan first, add ROCm when an engine or benchmark earns the complexity, and keep both llama.cpp builds if the machine regularly serves different model shapes. The best AMD backend is not a permanent property of the card; it is a property of the card, engine, model, driver, and workload together.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://rocm.docs.amd.com/en/latest/about/release-notes.html" rel="noopener noreferrer"&gt;ROCm 10.0.0 release notes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html" rel="noopener noreferrer"&gt;ROCm compatibility matrix&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rocm.docs.amd.com/en/latest/about/transition-guide-TheRock.html" rel="noopener noreferrer"&gt;ROCm TheRock transition and package mapping&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md" rel="noopener noreferrer"&gt;llama.cpp HIP and Vulkan build instructions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/gpu" rel="noopener noreferrer"&gt;Ollama AMD and Vulkan hardware support&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://rocm.docs.amd.com/projects/ai-ecosystem/en/latest/inference/vllm.html" rel="noopener noreferrer"&gt;AMD vLLM inference and serving guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://huggingface.co/docs/text-generation-inference/main/installation_amd" rel="noopener noreferrer"&gt;Hugging Face TGI on AMD GPUs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://lmstudio.ai/docs/cli/runtime/runtime" rel="noopener noreferrer"&gt;LM Studio runtime management&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://localai.io/docs/features/gpu-acceleration/" rel="noopener noreferrer"&gt;LocalAI GPU acceleration&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ggml-org/llama.cpp/discussions/10879" rel="noopener noreferrer"&gt;llama.cpp Vulkan performance discussion&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ggml-org/llama.cpp/discussions/15021" rel="noopener noreferrer"&gt;llama.cpp ROCm performance discussion&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Angelov, I. "Local LLM Inference on AMD RX 9070 XT — Vulkan vs ROCm Benchmarks on RDNA4." digtvbg.com, March 2026. &lt;a href="https://digtvbg.com/blog/llama-server-vulkan-rdna4-vllm-rocm-benchmark/" rel="noopener noreferrer"&gt;https://digtvbg.com/blog/llama-server-vulkan-rdna4-vllm-rocm-benchmark/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;"ROCm Vs. Vulkan Llama.cpp RDNA4 Radeon RX 9070 XT Benchmarks." OpenBenchmarking.org. &lt;a href="https://openbenchmarking.org/result/2509078-NE-ROCMVSVUL92" rel="noopener noreferrer"&gt;https://openbenchmarking.org/result/2509078-NE-ROCMVSVUL92&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/ggml-org/llama.cpp/issues/26663" rel="noopener noreferrer"&gt;"[Vulkan] Pathological token-generation slowdown on RX 9070 XT (gfx1201) for models with hidden_size &amp;gt;= 4096."&lt;/a&gt; llama.cpp GitHub issue.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>llm</category>
      <category>selfhosting</category>
      <category>llamacpp</category>
      <category>ollama</category>
    </item>
    <item>
      <title>Choosing free on-prem git server - Gitea is the winner!</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Sat, 12 Sep 2026 11:41:54 +0000</pubDate>
      <link>https://dev.to/rosgluk/choosing-free-on-prem-git-server-gitea-is-the-winner-32a7</link>
      <guid>https://dev.to/rosgluk/choosing-free-on-prem-git-server-gitea-is-the-winner-32a7</guid>
      <description>&lt;p&gt;Wanting to move your projects away from open cloud git providers and thinking of self-hosting internal git server locally?&lt;/p&gt;

&lt;p&gt;This guide is part of &lt;a href="https://www.glukhov.org/developer-tools/" rel="noopener noreferrer"&gt;Developer Tools: The Complete Guide to Modern Development Workflows&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing servers
&lt;/h2&gt;

&lt;p&gt;Running your own git server shouldn't be too difficult, right?&lt;/p&gt;

&lt;p&gt;So now choosing free git server from a vary short list of options.&lt;br&gt;
&lt;a href="https://bonobogitserver.com/" rel="noopener noreferrer"&gt;Bonobo&lt;/a&gt;&lt;br&gt;
&lt;a href="https://gogs.io/" rel="noopener noreferrer"&gt;Gogs&lt;/a&gt; vs&lt;br&gt;
&lt;a href="https://gitea.com/" rel="noopener noreferrer"&gt;Gitea&lt;/a&gt; vs&lt;br&gt;
&lt;a href="https://about.gitlab.com/" rel="noopener noreferrer"&gt;Gitlab&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Bonobo though is free but for windows, and doesn't have a linux version.&lt;/p&gt;

&lt;p&gt;Gitlab is feature-rich and resource-heavy, that's a hands on experience. It's a commertial product but&lt;br&gt;
&lt;a href="https://about.gitlab.com/install/" rel="noopener noreferrer"&gt;has a free version&lt;/a&gt;&lt;br&gt;
too.&lt;/p&gt;

&lt;p&gt;The Gogs is very light-weight, I've tried it and it worked well, but it's missing a container registry&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://docs.gitea.com/next/installation/comparison" rel="noopener noreferrer"&gt;comparison&lt;/a&gt; favors Gitea in my eyes.&lt;/p&gt;
&lt;h2&gt;
  
  
  Both Gitea and Postgresql dockerized
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.gitea.com/next/installation/install-with-docker" rel="noopener noreferrer"&gt;https://docs.gitea.com/next/installation/install-with-docker&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; ~
&lt;span class="nb"&gt;mkdir &lt;/span&gt;gitea
&lt;span class="nb"&gt;cd &lt;/span&gt;gitea
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;docker-compose.yml:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3"&lt;/span&gt;

&lt;span class="na"&gt;networks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;gitea&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;external&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;

&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gitea/gitea:latest&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gitea&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;USER_UID=1000&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;USER_GID=1000&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;GITEA__database__DB_TYPE=postgres&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;GITEA__database__HOST=db:5432&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;GITEA__database__NAME=gitea&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;GITEA__database__USER=gitea&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;GITEA__database__PASSWD=gitea&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;always&lt;/span&gt;
    &lt;span class="na"&gt;networks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;gitea&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./gitea:/data&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/etc/timezone:/etc/timezone:ro&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/etc/localtime:/etc/localtime:ro&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3000:3000"&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;222:22"&lt;/span&gt;
    &lt;span class="na"&gt;depends_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;db&lt;/span&gt;

  &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;postgres:14&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;always&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;POSTGRES_USER=gitea&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;POSTGRES_PASSWORD=gitea&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;POSTGRES_DB=gitea&lt;/span&gt;
    &lt;span class="na"&gt;networks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;gitea&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;./postgres:/var/lib/postgresql/data&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker-compose up &lt;span class="nt"&gt;-d&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;navigate to &lt;a href="http://localhost:3000/" rel="noopener noreferrer"&gt;http://localhost:3000/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;to shutdown&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker-compose down
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;volumes will stay&lt;/p&gt;

&lt;h2&gt;
  
  
  Resource usage
&lt;/h2&gt;

&lt;p&gt;Showing containers take 260MB RAM and a bit of CPU.&lt;/p&gt;

&lt;p&gt;Docker images size 583MB total.&lt;br&gt;
422MB of which is postgres:14 image.&lt;br&gt;
postgres:14-alpine image is ocupying 239MB, might also work if resources are constrained.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PS.&lt;/strong&gt; After migrating 10+ repos to gitea, some forking and clonning gitea container alone is using 420MB RAM now.&lt;br&gt;
Need to keep an eye on it.&lt;/p&gt;
&lt;h2&gt;
  
  
  Sync two repos
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.gitea.com/next/usage/repo-mirror" rel="noopener noreferrer"&gt;https://docs.gitea.com/next/usage/repo-mirror&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Can do &lt;strong&gt;push&lt;/strong&gt; and &lt;strong&gt;pull&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;
  
  
  pull
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Select New Migration in the Create... menu on the top right.&lt;/li&gt;
&lt;li&gt;Select the remote repository service.&lt;/li&gt;
&lt;li&gt;Enter a repository URL.&lt;/li&gt;
&lt;li&gt;If the repository needs authentication fill in your authentication information.&lt;/li&gt;
&lt;li&gt;Check the box This repository will be a mirror.&lt;/li&gt;
&lt;li&gt;Select Migrate repository to save the configuration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The repository now gets mirrored periodically from the remote repository.&lt;br&gt;
You can force a sync by selecting Synchronize Now in the repository settings.&lt;/p&gt;

&lt;p&gt;Can only set up pull mirroring for repos that don't exist yet on your instance.&lt;br&gt;
Once the repo is created, you can't convert it into a pull mirror anymore.&lt;/p&gt;
&lt;h2&gt;
  
  
  Https config
&lt;/h2&gt;

&lt;p&gt;There is more in ssl config: &lt;a href="https://docs.gitea.com/next/administration/https-setup" rel="noopener noreferrer"&gt;https://docs.gitea.com/next/administration/https-setup&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;but we are trying this for now (took from gitea site):&lt;/p&gt;
&lt;h3&gt;
  
  
  Using the built-in server
&lt;/h3&gt;

&lt;p&gt;Before you enable HTTPS, make sure that you have valid SSL/TLS certificates. You could use self-generated certificates for evaluation and testing.&lt;br&gt;
Please run&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gitea cert &lt;span class="nt"&gt;--host&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;HOST]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to generate a self signed certificate.&lt;/p&gt;

&lt;p&gt;If you are using Apache or nginx on the server, it's recommended to check the reverse proxy guide.&lt;/p&gt;

&lt;p&gt;To use Gitea's built-in HTTPS support, you must change your app.ini file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[server]&lt;/span&gt;
&lt;span class="py"&gt;PROTOCOL&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;https&lt;/span&gt;
&lt;span class="py"&gt;ROOT_URL&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;https://git.example.com:3000/&lt;/span&gt;
&lt;span class="py"&gt;HTTP_PORT&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;3000&lt;/span&gt;
&lt;span class="py"&gt;CERT_FILE&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;cert.pem&lt;/span&gt;
&lt;span class="py"&gt;KEY_FILE&lt;/span&gt;  &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;key.pem&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note that if your certificate is signed by a third party certificate authority (i.e. not self-signed), then cert.pem should contain the certificate chain. The server certificate must be the first entry in cert.pem, followed by the intermediaries in order (if any). The root certificate does not have to be included because the connecting client must already have it in order to establish the trust relationship. To learn more about the config values, please checkout the Config Cheat Sheet.&lt;/p&gt;

&lt;p&gt;For the CERT_FILE or KEY_FILE field, the file path is relative to the GITEA_CUSTOM environment variable when it is a relative path. It can be an absolute path as well.&lt;/p&gt;

&lt;p&gt;Setting up HTTP redirection&lt;br&gt;
The Gitea server is only able to listen to one port; to redirect HTTP requests to the HTTPS port, you will need to enable the HTTP redirection service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[server]&lt;/span&gt;
&lt;span class="py"&gt;REDIRECT_OTHER_PORT&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;true&lt;/span&gt;
&lt;span class="c"&gt;; Port the redirection service should listen on
&lt;/span&gt;&lt;span class="py"&gt;PORT_TO_REDIRECT&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;3080&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you are using Docker, make sure that this port is configured in your docker-compose.yml file.&lt;/p&gt;

&lt;h3&gt;
  
  
  Using reverse proxy
&lt;/h3&gt;

&lt;p&gt;In this post: &lt;a href="https://www.glukhov.org/developer-tools/git-and-forges/gitea-ssl/" rel="noopener noreferrer"&gt;Gitea-ssl&lt;/a&gt; I'm configuring Apache as TLS terminating reverse proxy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Todo
&lt;/h2&gt;

&lt;p&gt;ssh config: &lt;a href="https://docs.gitea.com/next/installation/install-with-docker" rel="noopener noreferrer"&gt;https://docs.gitea.com/next/installation/install-with-docker&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;backup-restore test. See &lt;a href="https://www.glukhov.org/developer-tools/git-and-forges/gitea-backup-restore/" rel="noopener noreferrer"&gt;Backup and Restore Gitea server&lt;/a&gt; for detailed backup procedures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Useful links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://gogs.io/" rel="noopener noreferrer"&gt;https://gogs.io/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://about.gitea.com/" rel="noopener noreferrer"&gt;https://about.gitea.com/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/go-gitea/gitea" rel="noopener noreferrer"&gt;https://github.com/go-gitea/gitea&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.glukhov.org/developer-tools/git-and-forges/gitflow-steps-and-alternatives/" rel="noopener noreferrer"&gt;Gitflow Explained: Steps, Alternatives, Pros, and Cons&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.glukhov.org/developer-tools/git-and-forges/git-cheatsheet/" rel="noopener noreferrer"&gt;GIT Cheatsheet: Most useful GIT commands&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>linux</category>
      <category>selfhosting</category>
      <category>git</category>
      <category>gitea</category>
    </item>
    <item>
      <title>KV Cache on 16 GB GPUs: Making Long Context Actually Fit</title>
      <dc:creator>Rost</dc:creator>
      <pubDate>Fri, 11 Sep 2026 09:09:53 +0000</pubDate>
      <link>https://dev.to/rosgluk/kv-cache-on-16-gb-gpus-making-long-context-actually-fit-3789</link>
      <guid>https://dev.to/rosgluk/kv-cache-on-16-gb-gpus-making-long-context-actually-fit-3789</guid>
      <description>&lt;p&gt;A model can advertise a 128K context window and still fail at 40K tokens on a 16 GB GPU. The architecture ceiling never promised that weights, KV cache, compute buffers, and the desktop compositor would fit on your card at the same time.&lt;/p&gt;

&lt;p&gt;The KV cache is usually where long-context plans meet that physical limit. It grows with every active token and sequence, so a configuration that looks comfortable at startup can slow sharply, spill into system memory, or fail during a large prefill.&lt;/p&gt;

&lt;p&gt;This guide turns the problem into a VRAM budget. It covers the cache formula, reproducible 32K to 128K size tables, and working configurations for llama.cpp &lt;code&gt;--cache-type-k&lt;/code&gt; and &lt;code&gt;--cache-type-v&lt;/code&gt;, vLLM paged and prefix caching, and Ollama context controls — plus the experimental adaptive-cache forks that deserve interest but not blind trust. For the broader throughput, latency, and benchmark context behind these numbers, start with the &lt;a href="https://www.glukhov.org/llm-performance/" rel="noopener noreferrer"&gt;LLM Performance hub&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Short Answer for a 16 GB GPU
&lt;/h2&gt;

&lt;p&gt;Start with one sequence, a realistic maximum context, Flash Attention, and an 8-bit KV cache. Measure that configuration before attempting 4-bit cache, CPU offload, multiple parallel slots, or an experimental fork.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Sensible first attempt on 16 GB&lt;/th&gt;
&lt;th&gt;Main risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;32K&lt;/td&gt;
&lt;td&gt;Q4 or Q5 weights, Q8 KV, one sequence&lt;/td&gt;
&lt;td&gt;Model weights leave too little buffer space&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;64K&lt;/td&gt;
&lt;td&gt;Smaller model or aggressive weight quant, Q8 KV&lt;/td&gt;
&lt;td&gt;Prefill latency and cache bandwidth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;Small GQA model, Q8 or tested Q4 KV, one sequence&lt;/td&gt;
&lt;td&gt;Cache alone may consume most of VRAM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two concurrent 64K sessions&lt;/td&gt;
&lt;td&gt;Treat as roughly a 128K cache budget&lt;/td&gt;
&lt;td&gt;Parallel capacity is mistaken for free throughput&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;My opinion is simple: a stable 64K setup is usually more useful than a nominal 128K setup that runs at the edge of an out-of-memory failure. Context capacity is not a trophy; it is a latency, quality, and concurrency decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the KV Cache Stores
&lt;/h2&gt;

&lt;p&gt;During autoregressive generation, every attention layer produces key and value tensors for each processed token. The runtime retains those tensors so the next token can attend to earlier tokens without recomputing the entire prefix.&lt;/p&gt;

&lt;p&gt;The cache saves enormous compute, but it consumes memory in proportion to the retained token count. For a conventional transformer with grouped-query attention, a useful baseline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;KV bytes = sequences * tokens * layers * 2 * KV heads * head dimension * bytes per value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The factor of two represents keys and values. Multi-head attention uses as many KV heads as query heads, grouped-query attention uses fewer KV heads, and multi-head latent attention or hybrid recurrent architectures need different calculations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Parameter Count Is Not Enough
&lt;/h3&gt;

&lt;p&gt;Two 8B models can have very different KV-cache costs. One may use 32 layers and eight KV heads, while another may use fewer KV heads, shared KV layers, sliding-window attention, or compressed latent states.&lt;/p&gt;

&lt;p&gt;Parameter count mostly predicts weight memory. KV geometry comes from attention architecture, so read the model metadata rather than guessing from &lt;code&gt;8B&lt;/code&gt;, &lt;code&gt;27B&lt;/code&gt;, or the GGUF file size. The clearest illustration is how far attention design has moved beyond plain multi-head attention (MHA):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Query Attention (MQA)&lt;/strong&gt; shares a single K/V head across all query heads — maximal cache savings, but it is the most aggressive quality compromise and is rarely used alone in current frontier models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grouped-Query Attention (GQA)&lt;/strong&gt; groups query heads into clusters that each share one K/V head — the mainstream compromise used by most open dense models, and the geometry the formula above assumes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Head Latent Attention (MLA)&lt;/strong&gt;, introduced in DeepSeek-V2 and carried into DeepSeek-V3 and Kimi K2, takes a different approach entirely: instead of sharing K/V across heads, it projects keys and values into a compressed low-rank latent vector and reconstructs the full-resolution K/V on demand at attention time. DeepSeek reported roughly a &lt;strong&gt;93% KV-cache reduction&lt;/strong&gt; versus an equivalently sized dense MHA model, while keeping quality competitive with — sometimes ahead of — GQA at the same memory budget.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical consequence is that a "27B GQA model" and a "27B MLA model" can have KV-cache footprints that differ by an order of magnitude for the same context length. Do not assume the formula above applies to a model that documents itself as using latent attention, DeltaNet-style state, or sliding-window layers — check the architecture section of the model card first.&lt;/p&gt;

&lt;p&gt;With Ollama, &lt;code&gt;ollama show MODEL --verbose&lt;/code&gt; exposes model metadata including layer count, attention heads, KV heads, and context length where the format provides them. With llama.cpp, the model-loader output printed at startup usually includes equivalent GGUF metadata and the runtime's actual cache allocation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model Limit, Allocated Context, and Used Context
&lt;/h3&gt;

&lt;p&gt;These are three separate numbers. The model limit is the maximum supported by its training and positional encoding, the allocated context is what the runtime reserves or permits, and used context is the tokens currently retained for a sequence.&lt;/p&gt;

&lt;p&gt;Raising an engine flag cannot safely extend a model beyond its supported position scheme. RoPE scaling can extend some architectures, but it is a model-quality experiment, not a KV-memory optimization.&lt;/p&gt;

&lt;h2&gt;
  
  
  KV Cache Size Table: 32K, 64K, and 128K Context Budgets
&lt;/h2&gt;

&lt;p&gt;Consider a representative GQA model with 32 layers, eight KV heads, and a head dimension of 128. These dimensions produce 65,536 key and value elements per token before multiplying by the storage size of each element.&lt;/p&gt;

&lt;p&gt;The table uses binary GiB and the physical block sizes commonly associated with llama.cpp &lt;code&gt;f16&lt;/code&gt;, &lt;code&gt;q8_0&lt;/code&gt;, and &lt;code&gt;q4_0&lt;/code&gt;. It is a baseline calculation, not a promise about total process memory; alignment, metadata, hybrid layers, and backend workspaces add overhead.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cache type&lt;/th&gt;
&lt;th&gt;Approx. bytes per stored value&lt;/th&gt;
&lt;th&gt;32K context&lt;/th&gt;
&lt;th&gt;64K context&lt;/th&gt;
&lt;th&gt;128K context&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;F16&lt;/td&gt;
&lt;td&gt;2.0000&lt;/td&gt;
&lt;td&gt;4.00 GiB&lt;/td&gt;
&lt;td&gt;8.00 GiB&lt;/td&gt;
&lt;td&gt;16.00 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0&lt;/td&gt;
&lt;td&gt;1.0625&lt;/td&gt;
&lt;td&gt;2.13 GiB&lt;/td&gt;
&lt;td&gt;4.25 GiB&lt;/td&gt;
&lt;td&gt;8.50 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q4_0&lt;/td&gt;
&lt;td&gt;0.5625&lt;/td&gt;
&lt;td&gt;1.13 GiB&lt;/td&gt;
&lt;td&gt;2.25 GiB&lt;/td&gt;
&lt;td&gt;4.50 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0 K plus Q4_0 V&lt;/td&gt;
&lt;td&gt;Mixed&lt;/td&gt;
&lt;td&gt;1.63 GiB&lt;/td&gt;
&lt;td&gt;3.25 GiB&lt;/td&gt;
&lt;td&gt;6.50 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now double the layer count to 64 while leaving the other dimensions unchanged. The FP16 cache becomes 8 GiB at 32K, 16 GiB at 64K, and 32 GiB at 128K, which demonstrates why a single context recommendation cannot cover every model.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Actual 16 GB Equation
&lt;/h3&gt;

&lt;p&gt;A practical budget is broader than the KV formula:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;usable VRAM = total VRAM - desktop and driver reserve

KV budget = usable VRAM
          - GPU-resident model weights
          - graph and activation buffers
          - runtime workspace
          - speculative-decoding state
          - safety margin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On a display-attached 16 GB card, do not plan around all 16 GiB being available. Reserve at least several hundred MiB for the desktop and driver, then leave another margin for workload-dependent buffers; 1.0 to 1.5 GiB of total breathing room is a reasonable starting assumption, but your logs are the authority.&lt;/p&gt;

&lt;p&gt;Suppose a GGUF model occupies 10.8 GiB on the GPU and runtime overhead peaks near 1.2 GiB. After a 1 GiB safety margin, only about 3 GiB remains for KV, so the representative model fits roughly 45K tokens with Q8_0 or 87K with Q4_0, before engine-specific overhead.&lt;/p&gt;

&lt;p&gt;That does not automatically make Q4_0 the right choice. If long-context accuracy drops on your workload, a smaller or more aggressively quantized model with a Q8_0 cache may be better than larger weights paired with a fragile cache. Measured anchors for exactly this arithmetic live in the &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/best-llm-on-16gb-vram-gpu/" rel="noopener noreferrer"&gt;16 GB VRAM llama.cpp benchmark tables&lt;/a&gt;, where VRAM per model is recorded at 19K, 32K, and 64K context. For a wider survey of which model sizes and quant levels behave well under Ollama on the same class of card, see &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/choosing-best-llm-for-ollama-on-16gb-vram-gpu/" rel="noopener noreferrer"&gt;Comparing LLMs performance on Ollama on 16GB VRAM GPU&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calculate the KV Cache Budget for Your Model
&lt;/h2&gt;

&lt;p&gt;The following Python snippet estimates a conventional full-attention GQA cache. Replace the geometry with values from the model configuration or GGUF metadata.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;kv_gib&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;layers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kv_heads&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;head_dim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bytes_per_value&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sequences&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;sequences&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;layers&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;kv_heads&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;head_dim&lt;/span&gt;
        &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;bytes_per_value&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;layers&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kv_heads&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;head_dim&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;128&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;types&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;f16&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;2.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;q8_0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;34&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;q4_0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;18&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;32768&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;65536&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;131072&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;kv_gib&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bytes_per_value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;row&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Q8_0 and Q4_0 ratios include simple block metadata, which is why they are slightly larger than exactly one byte and half a byte per value. The runtime's startup report remains more accurate because it knows model-specific cache layouts.&lt;/p&gt;

&lt;h3&gt;
  
  
  When This Formula Is Wrong: Hybrid and Sliding-Window Architectures
&lt;/h3&gt;

&lt;p&gt;Do not force hybrid architectures into the conventional GQA equation. Sliding-window layers retain only a recent window, shared KV layers reduce duplication, recurrent layers may carry fixed-size state, and multi-head latent attention stores a compressed representation rather than per-head K/V tensors — the MLA case above being the most dramatic example.&lt;/p&gt;

&lt;p&gt;Modern engines increasingly manage these mixed layouts explicitly. Use the formula to explain the dominant terms, then confirm the allocation reported by the exact engine build and backend you plan to deploy.&lt;/p&gt;

&lt;h2&gt;
  
  
  llama.cpp: Direct Control of K and V Precision
&lt;/h2&gt;

&lt;p&gt;llama.cpp exposes separate &lt;code&gt;--cache-type-k&lt;/code&gt; and &lt;code&gt;--cache-type-v&lt;/code&gt; options in its current argument parser. This is the most useful local-inference interface when you need to trade cache precision against context capacity instead of accepting one global preset. If you need the surrounding install and serving setup first, the &lt;a href="https://www.glukhov.org/llm-hosting/llama-cpp/" rel="noopener noreferrer"&gt;llama.cpp guide&lt;/a&gt; covers &lt;code&gt;llama-cli&lt;/code&gt;, &lt;code&gt;llama-server&lt;/code&gt;, and the key VRAM flags.&lt;/p&gt;

&lt;p&gt;A conservative 64K single-user configuration looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./llama-server &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model&lt;/span&gt; /models/model.gguf &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--n-gpu-layers&lt;/span&gt; 999 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ctx-size&lt;/span&gt; 65536 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--parallel&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--flash-attn&lt;/span&gt; on &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cache-type-k&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cache-type-v&lt;/span&gt; q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--batch-size&lt;/span&gt; 1024 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--ubatch-size&lt;/span&gt; 256
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Flag syntax and backend support change quickly, so run &lt;code&gt;llama-server --help&lt;/code&gt; for the installed build. More importantly, inspect the startup log: it should show the intended context, cache types, GPU offload, and the allocated K and V buffers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which llama.cpp Cache Types to Try
&lt;/h3&gt;

&lt;p&gt;Start with Q8_0 for both K and V. It roughly halves KV memory relative to F16, and independent perplexity tests on 20B-plus models (Qwen3.6-27B, Nemotron-30B) show the aggregate quality delta from F16 is within measurement noise — a much less dramatic gamble than moving directly to Q4_0, which the same tests showed collapsing decode speed and accuracy at long context on smaller models.&lt;/p&gt;

&lt;p&gt;If Q8_0 does not fit, test Q8_0 keys with Q4_0 values before quantizing both sides to Q4_0. This ordering has research backing, not just folklore: controlled bit-allocation studies on Llama, Phi-4, Qwen3, and Mistral checkpoints found that key tensors are consistently two to ten times more sensitive to quantization error than value tensors, and that giving keys the larger bit budget (for example 4-bit keys with 2-bit values) recovers up to 94–98% of full-precision accuracy — while the inverted split (2-bit keys, 4-bit values) can lose 30 percentage points on tasks like GSM8K. Keys determine which earlier tokens attention actually matches, so protecting them first is the architecturally sound choice, not just the safer-sounding one.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Memory&lt;/th&gt;
&lt;th&gt;Quality risk&lt;/th&gt;
&lt;th&gt;Recommendation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;F16 K and V&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;td&gt;Lowest&lt;/td&gt;
&lt;td&gt;Baseline when it fits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0 K and V&lt;/td&gt;
&lt;td&gt;About half of F16&lt;/td&gt;
&lt;td&gt;Low but not zero&lt;/td&gt;
&lt;td&gt;Default 16 GB starting point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0 K, Q4_0 V&lt;/td&gt;
&lt;td&gt;Between Q8 and Q4&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Useful second step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q4_0 K and V&lt;/td&gt;
&lt;td&gt;About one quarter of F16&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;td&gt;Validate at target depth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One caveat worth internalising: "low quality risk" on aggregate benchmarks does not mean &lt;em&gt;zero&lt;/em&gt; risk at the token level. A controlled test that held Flash Attention constant and only changed KV precision under greedy (deterministic) decoding found that Q8_0 cache changed the exact generated text on the large majority of prompts, and Q4_0 changed it on essentially all of them — once one token flips, the rest of the continuation can diverge. Perplexity and downstream-task scores can look fine on average while individual outputs still differ from the F16 baseline. If your application needs byte-for-byte reproducibility (regression tests, cached responses, deterministic agents), treat any KV quantization as a behavioral change, not just a memory optimization, and validate against your own fixed prompt set.&lt;/p&gt;

&lt;p&gt;Quantized V cache may require Flash Attention or a compatible backend path. A server that silently falls back to another type invalidates the experiment, which is why startup logs matter more than copied command lines.&lt;/p&gt;

&lt;h3&gt;
  
  
  Context, Parallel Slots, and Unified Cache
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;--ctx-size&lt;/code&gt; describes an engine capacity, not a guarantee that every parallel slot receives that many tokens independently. Cache management has evolved in llama.cpp, including unified-cache behavior, so test the exact build rather than relying on an older rule that simply divides context by slot count.&lt;/p&gt;

&lt;p&gt;The capacity equation still survives implementation changes: simultaneous unique tokens need storage somewhere. If two agent sessions may each reach 48K, budget for close to 96K live tokens unless the workload shares prefixes or tolerates eviction and recomputation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Batch Size Does Not Shrink Stored KV
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;--batch-size&lt;/code&gt; and &lt;code&gt;--ubatch-size&lt;/code&gt; affect prompt processing and temporary memory. Lowering them can rescue a large prefill from an activation-memory spike, but it does not change the persistent bytes required for each retained token.&lt;/p&gt;

&lt;p&gt;This distinction explains a common failure pattern: the model starts and an empty request works, but a 60K prompt fails during ingestion. Reduce the micro-batch to diagnose the transient peak; reduce context, cache precision, parallelism, or weight residency to change persistent capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  vLLM: Paged Capacity Is Still Capacity
&lt;/h2&gt;

&lt;p&gt;vLLM approaches the problem as a serving engine. It profiles available memory, reserves a KV-cache pool, and allocates cache in blocks so concurrent sequences do not each require one large contiguous region. If you are deciding whether to move to vLLM in the first place, the &lt;a href="https://www.glukhov.org/llm-hosting/comparisons/ollama-to-vllm-migration/" rel="noopener noreferrer"&gt;Ollama to vLLM migration guide&lt;/a&gt; covers the workload signals; here the question is purely how much cache the pool can hold, and the &lt;a href="https://www.glukhov.org/llm-hosting/vllm/vllm-quickstart/" rel="noopener noreferrer"&gt;vLLM quickstart&lt;/a&gt; covers install and general serving flags beyond the capacity levers below.&lt;/p&gt;

&lt;p&gt;PagedAttention reduces fragmentation and waste around variable sequence lengths — paged allocation removes fragmentation, not per-token storage cost, so one unique 128K request still needs enough blocks for its KV state.&lt;/p&gt;

&lt;p&gt;The official &lt;a href="https://docs.vllm.ai/en/latest/configuration/conserving_memory/" rel="noopener noreferrer"&gt;vLLM memory-conservation guide&lt;/a&gt; recommends limiting &lt;code&gt;max_model_len&lt;/code&gt; and &lt;code&gt;max_num_seqs&lt;/code&gt; when memory is tight, and notes that CUDA graphs consume additional GPU memory. On a 16 GB card, both settings should be intentional rather than inherited from a model's maximum configuration.&lt;/p&gt;

&lt;p&gt;A focused single-sequence server might start here:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vllm serve MODEL_ID &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 65536 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-num-seqs&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--gpu-memory-utilization&lt;/span&gt; 0.90 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kv-cache-dtype&lt;/span&gt; fp8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--enable-prefix-caching&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not every 16 GB GPU, model, quantization method, or attention backend supports that exact combination. Treat it as a configuration shape: constrain length and concurrency, reserve headroom, select a supported cache dtype, and validate the initialization report.&lt;/p&gt;

&lt;h3&gt;
  
  
  FP8 KV Cache in vLLM
&lt;/h3&gt;

&lt;p&gt;The current &lt;a href="https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/" rel="noopener noreferrer"&gt;vLLM quantized KV-cache documentation&lt;/a&gt; supports FP8 cache formats on compatible CUDA and ROCm paths. FP8 approximately halves raw cache storage relative to BF16 or FP16 and can therefore increase token capacity or concurrency.&lt;/p&gt;

&lt;p&gt;Scaling matters. The documentation distinguishes default scales, warm-up calculation, and dataset calibration, and recommends dataset-based calibration for the highest accuracy; simply setting FP8 with scale 1.0 is convenient but not automatically the most reliable quality choice.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prefix Caching Is a Reuse Optimization
&lt;/h3&gt;

&lt;p&gt;Automatic prefix caching lets a new request reuse KV blocks for an identical cached prefix. It is excellent for repeated queries over the same long document, shared system prompts, and multi-round conversations because it avoids recomputing the matching prefill.&lt;/p&gt;

&lt;p&gt;It does not make a unique long request smaller, and it does not accelerate generation of new tokens. The &lt;a href="https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/" rel="noopener noreferrer"&gt;vLLM prefix-caching documentation&lt;/a&gt; explicitly limits the benefit to shared-prefix prefill work.&lt;/p&gt;

&lt;h3&gt;
  
  
  GPU Memory Utilization Is Not Free Memory
&lt;/h3&gt;

&lt;p&gt;Raising &lt;code&gt;--gpu-memory-utilization&lt;/code&gt; gives vLLM a larger reservation target, but it does not create VRAM. Pushing it too close to 1.0 can leave insufficient room for the display, another process, changing activation peaks, or non-PyTorch allocations.&lt;/p&gt;

&lt;p&gt;Begin around 0.88 to 0.92 on a dedicated 16 GB GPU, inspect the profile, and increase only if the workload remains stable. If initialization succeeds but real prompts fail, reduce batched tokens, sequence concurrency, CUDA graph capture, or the maximum context before assuming the allocator is broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ollama: Easier Controls, Less Granular Diagnosis
&lt;/h2&gt;

&lt;p&gt;Ollama deliberately provides a smaller operational surface. Its current &lt;a href="https://docs.ollama.com/context-length" rel="noopener noreferrer"&gt;context-length documentation&lt;/a&gt; defaults GPUs below 24 GiB to 4K context, recommends at least 64K for agent and coding workloads, and warns that larger context consumes more memory.&lt;/p&gt;

&lt;p&gt;Set the server-wide default and confirm the loaded model like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OLLAMA_CONTEXT_LENGTH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;65536 ollama serve

ollama ps
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can also set &lt;code&gt;num_ctx&lt;/code&gt; per request or model. &lt;code&gt;ollama ps&lt;/code&gt; is important because its &lt;code&gt;PROCESSOR&lt;/code&gt; and &lt;code&gt;CONTEXT&lt;/code&gt; columns reveal whether the model remained fully on the GPU and whether the requested context was actually allocated. Be aware that the scheduling behaviour behind those numbers changed between Ollama versions; &lt;a href="https://www.glukhov.org/llm-performance/ollama/memory-allocation-in-ollama-new-version/" rel="noopener noreferrer"&gt;my comparison of Ollama v0.12.1 memory allocation&lt;/a&gt; shows the new scheduler pushing some models further to CPU on a 16 GB card, so pin the version you measured.&lt;/p&gt;

&lt;h3&gt;
  
  
  Quantized KV Cache in Ollama
&lt;/h3&gt;

&lt;p&gt;Ollama exposes &lt;code&gt;OLLAMA_KV_CACHE_TYPE&lt;/code&gt; with &lt;code&gt;f16&lt;/code&gt;, &lt;code&gt;q8_0&lt;/code&gt;, and &lt;code&gt;q4_0&lt;/code&gt; choices in its &lt;a href="https://github.com/ollama/ollama/blob/main/docs/faq.mdx" rel="noopener noreferrer"&gt;current FAQ&lt;/a&gt;. Quantized KV requires Flash Attention, which Ollama uses automatically on supported backends or can be requested with &lt;code&gt;OLLAMA_FLASH_ATTENTION=1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A 16 GB long-context service can therefore be started as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OLLAMA_CONTEXT_LENGTH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;65536 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;OLLAMA_FLASH_ATTENTION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;OLLAMA_KV_CACHE_TYPE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;q8_0 &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;OLLAMA_NUM_PARALLEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="se"&gt;\&lt;/span&gt;
ollama serve
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Q8_0 is Ollama's recommended alternative to F16. The FAQ warns that Q4_0 can produce a more noticeable quality loss, especially at higher context, so it should be a measured fallback rather than an automatic 16 GB preset.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ollama Parallelism Multiplies the Context Budget
&lt;/h3&gt;

&lt;p&gt;Ollama documents a particularly clear rule: required memory scales with &lt;code&gt;OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH&lt;/code&gt;. Four parallel requests at a 32K setting can imply a 128K aggregate context allocation for that model.&lt;/p&gt;

&lt;p&gt;For a personal agent on 16 GB, keep &lt;code&gt;OLLAMA_NUM_PARALLEL=1&lt;/code&gt; until one long session is stable. Queueing a second request is usually preferable to pushing the first model partly onto the CPU and making both requests slow. The queueing, 503, and model-unloading mechanics behind that choice are documented in &lt;a href="https://www.glukhov.org/llm-performance/ollama/how-ollama-handles-parallel-requests/" rel="noopener noreferrer"&gt;how Ollama handles parallel requests&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  CPU Offload: A Valid Escape Hatch With a Price
&lt;/h2&gt;

&lt;p&gt;Moving some model layers or KV state into system RAM can turn an allocation failure into a working process. It also places PCIe bandwidth and host-memory latency in the decode path, where each generated token may pay the cost. The lane and generation evidence for when PCIe actually bites is in &lt;a href="https://www.glukhov.org/llm-performance/hardware/llm-performance-and-pci-lanes/" rel="noopener noreferrer"&gt;LLM Performance and PCIe Lanes&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Offload can be sensible for occasional batch work, but it is rarely the best default for an interactive coding agent. First compare a smaller weight quantization, Q8 KV, reduced concurrency, and a realistic context cap; use offload when capacity matters more than latency.&lt;/p&gt;

&lt;p&gt;Watch for the cliff rather than the average. A server may decode quickly at 8K, then slow severely after part of the working set spills, so benchmark at 32K, 64K, and the intended maximum rather than reporting only an empty-context token rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sliding-Window and Adaptive KV Caches
&lt;/h2&gt;

&lt;p&gt;Sliding-window attention changes the budget by retaining only a recent window for selected layers. Hybrid models may combine those layers with occasional global attention or recurrent state, making a flat full-context calculation substantially overestimate or misplace memory.&lt;/p&gt;

&lt;p&gt;The optimization is part of the model architecture, not a generic switch that can be applied without consequences. An engine must understand the layer pattern, eviction rules, positions, and any global tokens correctly.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Adaptive KV Tries to Improve
&lt;/h3&gt;

&lt;p&gt;Experimental forks go further by choosing cache precision or layout per layer and context depth. The goal is attractive: preserve higher precision where it matters, compress less sensitive layers, and change the mix before VRAM pressure causes a hard spill — the same key-sensitivity-over-value-sensitivity finding described above is exactly the kind of signal an adaptive allocator would want to exploit automatically instead of leaving it to manual &lt;code&gt;--cache-type-k&lt;/code&gt;/&lt;code&gt;--cache-type-v&lt;/code&gt; tuning.&lt;/p&gt;

&lt;p&gt;One August 2026 downstream project, &lt;a href="https://github.com/craftogrammer/llama.cpp-adaptive-turboquant" rel="noopener noreferrer"&gt;llama.cpp-adaptive-turboquant&lt;/a&gt;, reports an automatic selector for several layer-adaptive modes and publishes long-depth tests on an RTX 5080 16 GB. Those numbers are author-reported results from a specialized fork, not evidence that upstream llama.cpp behaves the same way.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why It Is Still Experimental
&lt;/h3&gt;

&lt;p&gt;The fork combines custom cache types, CUDA kernels, model-specific paths, and toolchain constraints. That is far more code to trust than switching upstream cache storage from F16 to Q8_0.&lt;/p&gt;

&lt;p&gt;Use such a fork only when upstream cannot meet a real requirement and you can reproduce quality, stability, and speed on your model. Record the commit and CUDA version, because a result attached only to a project name is not reproducible.&lt;/p&gt;

&lt;h3&gt;
  
  
  A Fair Adaptive-Cache Test
&lt;/h3&gt;

&lt;p&gt;Compare the fork against an upstream Q8_0 baseline with the same GGUF, prompt, sampler, context depth, and output length. Measure startup VRAM, peak prefill VRAM, prompt processing speed, decode speed, and a quality task that actually requires evidence from the oldest part of the context.&lt;/p&gt;

&lt;p&gt;Do not accept a successful allocation as a complete result. A cache can fit 128K and still lose early facts, corrupt output late in the sequence, or decode too slowly to be useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Worked 16 GB Tuning Procedure: One Variable at a Time
&lt;/h2&gt;

&lt;p&gt;The fastest route to a stable configuration is to change one memory dimension at a time. Randomly altering cache type, batch size, layer offload, parallelism, and context together produces a working command with no explanation.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A["Step 1: load at 8K context, one sequence, record warm-up VRAM"] --&amp;gt; B{"Weights + runtime under about 14.5 GiB?"}
    B -- "no" --&amp;gt; C["Pick a smaller quant or model, repeat Step 1"]
    C --&amp;gt; A
    B -- "yes" --&amp;gt; D["Step 2: F16/BF16 KV quality baseline, save task outputs"]
    D --&amp;gt; E["Step 3: switch to Q8_0 / FP8 KV with Flash Attention"]
    E --&amp;gt; F["Step 4: stage context up - 32K, 64K, 96K, 128K"]
    F --&amp;gt; G{"Failure during prefill?"}
    G -- "yes" --&amp;gt; H["Step 5: reduce batch / ubatch size"]
    H --&amp;gt; F
    G -- "no" --&amp;gt; I{"Failure only with concurrent requests?"}
    I -- "yes" --&amp;gt; J["Step 5: reduce parallel slots / max-num-seqs"]
    J --&amp;gt; F
    I -- "no" --&amp;gt; K["Only now: mixed Q8/Q4, full Q4, offload, adaptive fork"]&lt;/code&gt;&lt;/pre&gt;



&lt;h3&gt;
  
  
  Step 1: Establish the Weight Floor
&lt;/h3&gt;

&lt;p&gt;Load the model at 8K context, one sequence, and the intended GPU offload. Record process VRAM after warm-up and verify that no layers unexpectedly moved to the CPU.&lt;/p&gt;

&lt;p&gt;If the weights and runtime already consume more than about 14.5 to 15 GiB, long context has no healthy margin. Choose a smaller weight quant or model before tuning the cache.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Measure F16 or BF16 KV as the Quality Baseline
&lt;/h3&gt;

&lt;p&gt;Run the smallest context that supports your test and retain the default high-precision cache. Save outputs from retrieval, code editing, tool selection, and long-instruction tasks.&lt;/p&gt;

&lt;p&gt;This baseline tells you whether later errors come from cache quantization. Without it, a chat-template problem or weak model can easily be blamed on Q4 KV.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Move to Q8 or FP8
&lt;/h3&gt;

&lt;p&gt;Enable Flash Attention where required, select Q8_0 in llama.cpp or Ollama, or a supported FP8 mode in vLLM. Repeat the same prompts at the same token depths and confirm the log shows the intended cache type.&lt;/p&gt;

&lt;p&gt;For many 16 GB deployments, this is the useful stopping point. It roughly doubles raw KV capacity without making cache compression the most aggressive quantization in the stack.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Raise Context in Stages
&lt;/h3&gt;

&lt;p&gt;Test 32K, 64K, 96K, and 128K instead of jumping directly to the advertised maximum. At each stage, record prompt-processing tokens per second, decode tokens per second, peak VRAM, and whether evidence near the start can still be recovered.&lt;/p&gt;

&lt;p&gt;Long-context decode often slows even after memory fits because attention reads more cached state. Capacity and performance are separate axes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Tune Transient Memory
&lt;/h3&gt;

&lt;p&gt;If failure occurs during prefill rather than initialization, reduce micro-batch or maximum batched tokens. If failure occurs only with simultaneous requests, reduce sequence concurrency or parallel slots.&lt;/p&gt;

&lt;p&gt;Only after those controls are understood should you try mixed Q8/Q4 cache, full Q4 cache, CPU offload, or an adaptive fork. Keep the upstream Q8 run as the comparison baseline. If you later add speculative decoding or MTP, remember its draft buffers are another line in the budget equation, not free speed — the &lt;a href="https://www.glukhov.org/llm-performance/optimization/speculative-decoding/" rel="noopener noreferrer"&gt;speculative decoding guide&lt;/a&gt; covers the mechanics and their VRAM cost, and my &lt;a href="https://www.glukhov.org/llm-performance/benchmarks/comparing-qwen-3-6-mtp-vs-standard/" rel="noopener noreferrer"&gt;Qwen 3.6 27B and 35B MTP vs Standard benchmark&lt;/a&gt; shows exactly how much context an MTP head's extra state can cost on a 16 GB card.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to Record in a Long-Context Benchmark
&lt;/h2&gt;

&lt;p&gt;A single &lt;code&gt;tokens/s&lt;/code&gt; figure hides the exact problem this article is trying to solve. Long-context testing should preserve enough detail for another operator to reproduce the memory boundary.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPU and usable VRAM&lt;/td&gt;
&lt;td&gt;Display use and other processes change the budget&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Engine version or commit&lt;/td&gt;
&lt;td&gt;Cache behavior and flags evolve quickly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Driver, CUDA, ROCm, or Vulkan version&lt;/td&gt;
&lt;td&gt;Determines backend and kernel behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exact model and weight quant&lt;/td&gt;
&lt;td&gt;Defines weight residency and architecture&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;K and V cache types&lt;/td&gt;
&lt;td&gt;Defines persistent cache size and quality risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context capacity and prompt depth&lt;/td&gt;
&lt;td&gt;Allocation is not the same as actual depth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parallel sequences&lt;/td&gt;
&lt;td&gt;Multiplies or shares cache demand&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch and micro-batch&lt;/td&gt;
&lt;td&gt;Affects prefill peaks and speed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt processing speed&lt;/td&gt;
&lt;td&gt;Exposes long-prefill usability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decode speed at each depth&lt;/td&gt;
&lt;td&gt;Exposes cache-bandwidth slowdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peak VRAM and CPU offload&lt;/td&gt;
&lt;td&gt;Distinguishes fit from spill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-context quality result&lt;/td&gt;
&lt;td&gt;Detects compression or position failures&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Use &lt;code&gt;nvidia-smi&lt;/code&gt; sampling or the equivalent vendor tooling during both prefill and decode. The engine's allocation report is necessary, but peak device memory during a real prompt is the number that decides stability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common KV Cache Mistakes on 16 GB GPUs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Treating 128K Support as a Hardware Promise
&lt;/h3&gt;

&lt;p&gt;The context field in a model configuration is an architectural ceiling. It says nothing about the memory left after loading a particular quantization on a particular engine.&lt;/p&gt;

&lt;p&gt;Calculate the cache and verify the runtime. Marketing-sized context without a VRAM budget is merely an OOM delayed until the first serious prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Quantizing Weights but Forgetting KV
&lt;/h3&gt;

&lt;p&gt;A 4-bit GGUF reduces model weights, not an F16 KV cache. At long context, the cache can erase the entire saving and eventually exceed the weight footprint.&lt;/p&gt;

&lt;p&gt;Report both quantizations. &lt;code&gt;Q4_K_M model, Q8_0 KV&lt;/code&gt; is meaningful; &lt;code&gt;4-bit model&lt;/code&gt; is incomplete.&lt;/p&gt;

&lt;h3&gt;
  
  
  Assuming Paged Attention Compresses Tokens
&lt;/h3&gt;

&lt;p&gt;Paging improves allocation and sharing behavior. It does not change the tensor precision or remove the KV state required by one unique sequence.&lt;/p&gt;

&lt;p&gt;Use paged allocation to serve variable workloads efficiently. Use cache precision, model architecture, context caps, and concurrency limits to control capacity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Assuming Prefix Caching Helps Every Long Prompt
&lt;/h3&gt;

&lt;p&gt;Prefix caching saves repeated prefill computation when requests share an exact prefix. A one-off 100K repository dump receives no magical memory discount simply because prefix caching is enabled.&lt;/p&gt;

&lt;p&gt;It is a workload optimization, not a substitute for the budget equation. Measure hit rate and retained cache pressure in multi-user serving.&lt;/p&gt;

&lt;h3&gt;
  
  
  Using Q4 KV Without a Quality Test
&lt;/h3&gt;

&lt;p&gt;Low-bit cache can fail subtly. The model still writes fluent text, but attention over distant evidence, exact names, tool arguments, or code dependencies may degrade — and as the token-divergence research above shows, even the "safe" Q8_0 setting is not guaranteed to reproduce the exact F16 output under deterministic decoding, only to preserve accuracy on aggregate.&lt;/p&gt;

&lt;p&gt;Test the target task at the target depth. Short chat benchmarks are nearly useless for validating a long-context cache.&lt;/p&gt;

&lt;h3&gt;
  
  
  Leaving Parallelism on Auto
&lt;/h3&gt;

&lt;p&gt;An engine may choose concurrency that is sensible for throughput but impossible for your long-context target. On 16 GB, one deep sequence and several short sequences are fundamentally different workloads.&lt;/p&gt;

&lt;p&gt;Set the limit explicitly, then raise it with measured traffic. Otherwise a second request can turn a stable 64K configuration into an allocation or latency surprise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommended 16 GB Profiles
&lt;/h2&gt;

&lt;p&gt;These profiles are starting positions, not universal presets. A model with unusual KV geometry — an MLA or hybrid sliding-window design in particular — can be much cheaper or more expensive than the conventional GQA example.&lt;/p&gt;

&lt;h3&gt;
  
  
  Interactive Coding Agent
&lt;/h3&gt;

&lt;p&gt;Use one sequence, 48K to 64K context, Q8 cache, Flash Attention, and full GPU weight residency if possible. This profile favors predictable latency and good cache precision over an impressive but rarely useful maximum.&lt;/p&gt;

&lt;p&gt;Enable prefix reuse when the engine supports it because coding turns often share a large repository or conversation prefix. Still compact tool output and old transcripts; cache engineering does not make irrelevant tokens valuable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Long-Document Analysis
&lt;/h3&gt;

&lt;p&gt;Use a smaller model with 64K to 128K capacity, Q8 or calibrated FP8 cache, and repeated-prefix caching when multiple questions target the same document. Measure time to first token because prefill may dominate even when decode remains acceptable.&lt;/p&gt;

&lt;p&gt;If only one question will be asked, retrieval or chunked summarization may be faster and more reliable than forcing the entire corpus through a 16 GB card. Long context is a tool, not a replacement for information architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  Small Multi-User Server
&lt;/h3&gt;

&lt;p&gt;Cap per-request context and total active sequences rather than advertising the model maximum to every client. vLLM's paged allocation is useful here, while Ollama and llama.cpp also require explicit attention to aggregate live tokens.&lt;/p&gt;

&lt;p&gt;Prefer queueing over uncontrolled spill. A slower admission policy is less damaging than every request suddenly crossing PCIe during decode.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Recommendation for 16 GB Long Context
&lt;/h2&gt;

&lt;p&gt;For long context on 16 GB, Q8 KV and one active sequence are the right baseline. They expose the real limit without making low-bit cache quality, parallel allocation, and offload latency fail at once.&lt;/p&gt;

&lt;p&gt;Calculate from attention geometry, subtract weights and runtime overhead, and then confirm the result in engine logs and peak-memory measurements. If 128K still does not fit, a smaller model is often the cleanest optimization; if it fits but crawls, reducing context is often the honest one.&lt;/p&gt;

&lt;p&gt;Paged attention, prefix caching, sliding windows, and adaptive precision all solve useful but different problems. The winning setup is the one that remains on the GPU, retrieves old evidence correctly, and sustains acceptable decode speed at the context depth you actually use.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.vllm.ai/en/latest/configuration/conserving_memory/" rel="noopener noreferrer"&gt;vLLM: conserving GPU memory&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/" rel="noopener noreferrer"&gt;vLLM: quantized KV cache (FP8)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.vllm.ai/en/latest/features/automatic_prefix_caching/" rel="noopener noreferrer"&gt;vLLM: automatic prefix caching&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.ollama.com/context-length" rel="noopener noreferrer"&gt;Ollama: context length documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/ollama/ollama/blob/main/docs/faq.mdx" rel="noopener noreferrer"&gt;Ollama FAQ: &lt;code&gt;OLLAMA_KV_CACHE_TYPE&lt;/code&gt; and Flash Attention&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/craftogrammer/llama.cpp-adaptive-turboquant" rel="noopener noreferrer"&gt;llama.cpp-adaptive-turboquant&lt;/a&gt; — experimental layer-adaptive cache fork (author-reported results)&lt;/li&gt;
&lt;li&gt;Raschka, S. "Multi-Head Latent Attention (MLA)." &lt;em&gt;LLM Architecture Gallery&lt;/em&gt;. &lt;a href="https://sebastianraschka.com/llm-architecture-gallery/mla/" rel="noopener noreferrer"&gt;https://sebastianraschka.com/llm-architecture-gallery/mla/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;"Decoding Multi-Head Latent Attention: The KV Cache Memory Bottleneck, Solved." &lt;em&gt;Vizuara&lt;/em&gt;. &lt;a href="https://vizuara.substack.com/p/decoding-multi-head-latent-attention" rel="noopener noreferrer"&gt;https://vizuara.substack.com/p/decoding-multi-head-latent-attention&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;"Quantize What Counts: Bit Allocation Insights Informed by Spectral Gaps in Keys and Values." arXiv:2502.15075. &lt;a href="https://ar5iv.labs.arxiv.org/html/2502.15075" rel="noopener noreferrer"&gt;https://ar5iv.labs.arxiv.org/html/2502.15075&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;v-code01. "kvdivergence — does KV cache quantization change the generated text under greedy decoding?" GitHub. &lt;a href="https://github.com/v-code01/kvdivergence" rel="noopener noreferrer"&gt;https://github.com/v-code01/kvdivergence&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.glukhov.org/llm-performance/benchmarks/choosing-best-llm-for-ollama-on-16gb-vram-gpu/" rel="noopener noreferrer"&gt;Comparing LLMs performance on Ollama on 16GB VRAM GPU&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.glukhov.org/llm-performance/benchmarks/comparing-qwen-3-6-mtp-vs-standard/" rel="noopener noreferrer"&gt;Qwen 3.6 27B and 35B MTP vs Standard on 16GB GPU&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>llm</category>
      <category>llamacpp</category>
      <category>vllm</category>
      <category>ollama</category>
    </item>
  </channel>
</rss>
