<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ken Imoto</title>
    <description>The latest articles on DEV Community by Ken Imoto (@kenimo49).</description>
    <link>https://dev.to/kenimo49</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3800250%2F275022f6-cba9-47e3-b69e-e8faf7675a0c.jpg</url>
      <title>DEV Community: Ken Imoto</title>
      <link>https://dev.to/kenimo49</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kenimo49"/>
    <language>en</language>
    <item>
      <title>200ms Humans vs 700ms Voice Agents: What ACL IWSDS 2025 Says About the Turn-Taking Gap</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Wed, 23 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/kenimo49/200ms-humans-vs-700ms-voice-agents-what-acl-iwsds-2025-says-about-the-turn-taking-gap-2iap</link>
      <guid>https://dev.to/kenimo49/200ms-humans-vs-700ms-voice-agents-what-acl-iwsds-2025-says-about-the-turn-taking-gap-2iap</guid>
      <description>&lt;p&gt;The complaint that keeps showing up in every voice-agent beta test I've watched is not about the wrong answer. It's about the &lt;em&gt;timing&lt;/em&gt; of the right answer.&lt;/p&gt;

&lt;p&gt;"The AI cut me off." "It answered while I was still thinking." "I said 'um' and it started talking." Latency looks like a stopwatch problem, but users experience it as a manners problem, which is a much harder thing to fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 200ms vs 700ms gap the field is trying to close
&lt;/h2&gt;

&lt;p&gt;The ACL IWSDS 2025 survey on turn-taking in spoken dialogue puts a concrete number on the mismatch. In natural human conversation, a listener starts their reply roughly &lt;strong&gt;200ms&lt;/strong&gt; after the speaker's turn ends. Current dialogue agents take &lt;strong&gt;700 to 1000ms&lt;/strong&gt; to do the same thing.&lt;/p&gt;

&lt;p&gt;Three to five times slower. If human dialogue is a tennis rally, current voice AI is playing chess by post.&lt;/p&gt;

&lt;p&gt;Closing that gap is not one problem, it's two. And they pull in opposite directions.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Too slow to reply&lt;/strong&gt;: awkward silence, dropped rapport, users start repeating themselves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Too fast to reply&lt;/strong&gt;: the agent talks over the user, interrupts mid-word, kills trust.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both fail modes come from the same root cause: silence is not the same as end-of-turn.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why plain VAD is not enough
&lt;/h2&gt;

&lt;p&gt;Most current systems still decide the user is done talking with a &lt;strong&gt;Voice Activity Detector&lt;/strong&gt;. VAD is cheap and fast, and it answers exactly one question: is there voice in this audio frame, yes or no.&lt;/p&gt;

&lt;p&gt;That question is not the one the agent actually needs answered.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The 0.3–0.5s pause a speaker takes mid-sentence looks identical to end-of-turn.&lt;/li&gt;
&lt;li&gt;The silence after "uh" or "let me think" looks identical to end-of-turn.&lt;/li&gt;
&lt;li&gt;A cough or a sigh looks like voice.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can tune the silence threshold, but you cannot tune your way out of the fact that VAD does not know what a sentence is. VoiceInfra's production data lands on &lt;strong&gt;300–500ms&lt;/strong&gt; of trailing silence as the least-bad setting: shorter and the agent chops off natural pauses, longer and the delay becomes obvious. The comfortable range depends on the use case. Call-center flows tolerate 400–500ms because callers pause to think; command interfaces want 200–300ms because utterances are short; long-form narration wants 500–600ms because the pauses inside the story are longer.&lt;/p&gt;

&lt;p&gt;The threshold is a compromise, not a fix. The real fix has to know what the words &lt;em&gt;mean&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Semantic endpointing: Deepgram Flux and the "one model does both" bet
&lt;/h2&gt;

&lt;p&gt;The interesting move in 2026 is folding transcription and turn detection into the same model. Deepgram's &lt;strong&gt;Flux&lt;/strong&gt;, launched as their first ASR built for voice agents rather than for general transcription, does exactly that. The model outputs turn boundaries directly from audio, using the same joint architecture it uses to output text. Because the same weights see both the acoustic signal and the emerging transcript, the model has a shot at answering the harder question: "is this utterance semantically complete?"&lt;/p&gt;

&lt;p&gt;Flux exposes three knobs that make the tradeoff explicit rather than hidden inside a silence threshold:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;eot_threshold&lt;/code&gt; — how confident the model has to be before it commits to end-of-turn.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;eager_eot_threshold&lt;/code&gt; — a lower bar for tentative end-of-turn, useful for starting inference speculatively.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;eot_timeout_ms&lt;/code&gt; — the maximum silence duration before the model forces end-of-turn even if the confidence never crosses the threshold.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can also switch modes per-turn with a Configure message, so a barge-friendly interaction ("agent, wait") and a monologue turn ("read me the terms") do not have to share a threshold.&lt;/p&gt;

&lt;p&gt;The pattern is a hint at where the field is going. Turn-taking will stop being a pipeline stage after ASR and start being a joint output of the ASR itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Krisp's 6M-parameter turn model: same idea, edge-shaped
&lt;/h2&gt;

&lt;p&gt;Not everyone can afford a cloud round-trip on every silence check. Krisp ships two small audio-only models trained to handle the acoustic cues a naive VAD collapses — a &lt;strong&gt;9M-parameter&lt;/strong&gt; turn prediction model (~30MB) that scores end-of-turn from prosody and pausing patterns, and a &lt;strong&gt;6M-parameter&lt;/strong&gt; interruption model (~24MB) that separates real barge-ins from backchannels like "yeah" and "mhm". Together they cover the same failure modes VAD misclassifies:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;intentional speech vs a thinking pause&lt;/li&gt;
&lt;li&gt;filler words (um, uh, well) mid-sentence&lt;/li&gt;
&lt;li&gt;backchannels vs interruptions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At that size it runs on-device in real time, which matters for privacy-sensitive deployments and for wearables where the round-trip to the cloud is itself the biggest source of latency.&lt;/p&gt;

&lt;p&gt;The Krisp and Flux paths look opposite. One shrinks the model, the other gives the ASR the extra job. But they are attacking the same VAD failure mode from opposite sides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Graceful abort: what to do when you're wrong
&lt;/h2&gt;

&lt;p&gt;No detector, semantic or acoustic, is going to be right every time. The interesting question is what your pipeline does when it decides "the user is done" and turns out to be wrong.&lt;/p&gt;

&lt;p&gt;Twilio's &lt;strong&gt;graceful abort&lt;/strong&gt; pattern is the cleanest version I've seen written up. When STT signals end-of-turn early, the LLM starts generating a response. If fresh audio arrives before the response reaches the speaker, the pipeline kills the generation in flight and swallows the partial output. The window for this is the few hundred milliseconds it takes STT → LLM → TTS to produce audible speech. If you can revoke the guess inside that window, the user never hears it.&lt;/p&gt;

&lt;p&gt;The economics matter here: you are paying for LLM tokens on turns that get thrown away. In exchange, you get to be aggressive on end-of-turn detection without punishing the user when you're aggressive-and-wrong. For most conversational products that tradeoff is worth it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Barge-in: hearing the user through your own voice
&lt;/h2&gt;

&lt;p&gt;The other half of turn-taking is the reverse case: the user interrupts while the agent is still speaking. Detecting that seems trivial until you realize the AI's own audio is leaking into the microphone, and any naive VAD will happily flag that leak as a barge-in and abort the agent's own utterance.&lt;/p&gt;

&lt;p&gt;Sparkco's write-up on duplex barge-in handling is the clearest description of the fix I've read. Three moving parts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Full-duplex audio&lt;/strong&gt; — keep the mic hot the entire time the agent is speaking, never gate it on TTS output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Echo cancellation&lt;/strong&gt; — subtract the speaker output (as a reference signal) from the mic input, so what remains is the user's voice minus the agent's.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nuisance rejection&lt;/strong&gt; — filter environmental noise and short transients that a bare VAD would misclassify as speech onsets.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[agent is speaking]
  speaker out ──► reference signal ─┐
                                    ▼
  mic in ─────► echo canceller ────► user voice only ──► VAD / turn model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without echo cancellation you get the failure mode users describe as "the AI got startled by its own voice and stopped talking." The dog scared of its own reflection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bot-feel paradox and the 200-300ms delay trick
&lt;/h2&gt;

&lt;p&gt;Now the counterintuitive move. Once you've closed the acoustic and semantic gaps and the agent &lt;em&gt;can&lt;/em&gt; reply in 300ms, it turns out that replying in 300ms feels &lt;em&gt;worse&lt;/em&gt;. The response arrives before the user has finished processing their own sentence, and it registers as robotic rather than sharp.&lt;/p&gt;

&lt;p&gt;The fix is to inject &lt;strong&gt;200–300ms of intentional delay before the agent starts speaking, while continuing to run the LLM in the background&lt;/strong&gt;. You get the "thinking for a moment" cue humans read as attention, without paying real latency for it. The tokens are already streaming; you're just holding the TTS start.&lt;/p&gt;

&lt;p&gt;It's a UX trick, not a technical one. But it's the piece of the puzzle that gets forgotten when a team spends a quarter shaving milliseconds off ASR and then wonders why user ratings didn't move.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually closes the 500ms gap
&lt;/h2&gt;

&lt;p&gt;There is no single component that takes you from 700ms to 200ms. You get there by stacking:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it buys you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Semantic endpointing (Flux-class)&lt;/td&gt;
&lt;td&gt;Stops chopping mid-sentence, stops waiting for silence that already means end-of-turn.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small on-device turn + interruption models (Krisp-class)&lt;/td&gt;
&lt;td&gt;Score end-of-turn and separate backchannels from real barge-ins without a round-trip.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Graceful abort&lt;/td&gt;
&lt;td&gt;Lets you be aggressive on early end-of-turn without punishing the user when wrong.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplex barge-in with AEC&lt;/td&gt;
&lt;td&gt;Lets you leave the mic hot without the agent interrupting itself.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;200–300ms intentional delay&lt;/td&gt;
&lt;td&gt;Buys back the "thinking" cue that pure speed removes.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The move that unlocks the stack is admitting that end-of-turn is a language problem, not a silence problem. Once you accept that, VAD-only pipelines look the way pre-BERT NLP pipelines look now: perfectly reasonable for their era, obviously incomplete in retrospect.&lt;/p&gt;

&lt;p&gt;The gap to 200ms is not going to close by tuning thresholds. It closes when the model that hears the audio is the same model that understands the sentence.&lt;/p&gt;




&lt;p&gt;The full 300ms-UX playbook (latency budgets by pipeline stage, when to pick Whisper vs Deepgram vs Piper, and how to design the wait states so users forgive the last 200ms you can't remove) is written up in &lt;a href="https://kenimoto.dev/books/voice-ai-300ms-ux?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=turn-taking-200-vs-700" rel="noopener noreferrer"&gt;Voice AI 300ms UX&lt;/a&gt;. Chapter 9 covers turn-taking end-to-end; chapters 4 and 5 cover the latency anatomy that decides whether closing the gap is even possible for your stack.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;References&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ACL IWSDS 2025. "A Survey of Recent Advances on Turn-taking Modeling in Spoken Dialogue."&lt;/li&gt;
&lt;li&gt;Krisp AI. "Audio-only, 6M weights Turn-Taking model for Voice AI Agents." 2025.&lt;/li&gt;
&lt;li&gt;Twilio. "Core Latency in AI Voice Agents." 2025.&lt;/li&gt;
&lt;li&gt;Deepgram. "Flux: End-of-Turn Detection Parameters." 2026.&lt;/li&gt;
&lt;li&gt;Sparkco. "Optimizing Voice Agent Barge-in Detection." 2025.&lt;/li&gt;
&lt;li&gt;VoiceInfra. "Voice AI Prompt Engineering: Complete Technical Guide." 2025.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>voiceai</category>
      <category>webrtc</category>
      <category>webdev</category>
      <category>ux</category>
    </item>
    <item>
      <title>Claude Code + Playwright MCP + GitHub Actions: 3 Test Layers That Actually Kill Regressions</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Mon, 21 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/claude-code-playwright-mcp-github-actions-3-test-layers-that-actually-kill-regressions-34f6</link>
      <guid>https://dev.to/kenimo49/claude-code-playwright-mcp-github-actions-3-test-layers-that-actually-kill-regressions-34f6</guid>
      <description>&lt;p&gt;The first time I asked Claude Code to write tests for a feature it had just implemented, the tests passed. All of them. On the first try. I felt great about it for about ten minutes, until I realized the tests were describing the code, not the requirements. The code had a bug I had specified against, and the tests happily locked the bug in.&lt;/p&gt;

&lt;p&gt;That is context pollution. It is the fundamental problem with letting one model both implement and verify. The New Stack put it best: &lt;strong&gt;letting the model that wrote the code write the tests is asking a student to grade their own homework.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The fix is architectural, not attitudinal. Split the loop across three isolated contexts, each running at a different point in the development lifecycle. Here is what those three layers look like when you build them out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 1: TDD by subagent (dev time)
&lt;/h2&gt;

&lt;p&gt;Red-Green-Refactor works with Claude Code, but not the way most teams first try it. The naive setup is one Claude Code session doing all three phases. That is the setup that produces homework-grading.&lt;/p&gt;

&lt;p&gt;The working setup uses three subagents, each blind to the others' work.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Skills (orchestrator)
  ├── tdd-test-writer   (RED)    — sees feature description, NOT implementation
  ├── tdd-implementer   (GREEN)  — sees test file, NOT the writer's reasoning
  └── tdd-refactorer    (REFACTOR) — sees passing code, must keep tests green
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The critical bit is that each subagent runs in its own context. The test writer cannot peek at how the implementer plans to solve the problem. The implementer cannot see the notes the test writer wrote about "edge cases we should probably cover."&lt;/p&gt;

&lt;p&gt;A minimal skill definition for the writer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# .claude/agents/tdd-test-writer.md&lt;/span&gt;

You are a test writer. Your ONLY job is to write failing tests.

Rules:
&lt;span class="p"&gt;-&lt;/span&gt; You CANNOT see or reference any implementation files
&lt;span class="p"&gt;-&lt;/span&gt; Write tests based ONLY on the feature description
&lt;span class="p"&gt;-&lt;/span&gt; Tests must cover happy path + edge cases (nulls, empty, negative, overflow)
&lt;span class="p"&gt;-&lt;/span&gt; Run tests and confirm they FAIL before finishing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the implementer, which is deliberately narrow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# .claude/agents/tdd-implementer.md&lt;/span&gt;

You are an implementer. Your ONLY job is to make tests pass.

Rules:
&lt;span class="p"&gt;-&lt;/span&gt; Write the MINIMUM code to make tests pass
&lt;span class="p"&gt;-&lt;/span&gt; Do NOT refactor or optimize (that's the refactorer's job)
&lt;span class="p"&gt;-&lt;/span&gt; Do NOT modify test files
&lt;span class="p"&gt;-&lt;/span&gt; Run tests and confirm they PASS before finishing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why this works: the "design intent" that a normal Claude Code session leaks between phases now cannot leak. The test author writes what the feature &lt;em&gt;should&lt;/em&gt; do; the implementer writes what makes those tests pass; the refactorer cleans up without breaking anything. It is exactly the discipline TDD was invented to enforce, made mechanical by the fact that context boundaries are hard walls now.&lt;/p&gt;

&lt;p&gt;The other nice property: you can CLAUDE.md this policy at the repo level and forget about it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Testing policy&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; All new features go through the TDD subagent chain
&lt;span class="p"&gt;-&lt;/span&gt; No test code and implementation code in the same Claude session
&lt;span class="p"&gt;-&lt;/span&gt; Skip TDD only for one-line typo fixes and docs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Layer 2: Quinn, the Playwright MCP QA agent (PR time)
&lt;/h2&gt;

&lt;p&gt;The second layer runs at pull-request time and does the thing no unit test can do: &lt;strong&gt;click around the UI like a person who has never seen the code.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Alexop's Quinn pattern is the reference here. Quinn is an AI QA engineer defined by a prompt: a 12-year-veteran QA persona, instructed to test only through the UI, never look at source, always exercise the mobile viewport (375x667), and file findings with reproduction steps.&lt;/p&gt;

&lt;p&gt;The mechanism that makes Quinn possible in 2026 is Playwright MCP. The server ships dozens of browser-automation tools exposed over MCP, and the important detail is that it hands the agent an &lt;strong&gt;accessibility tree as structured JSON&lt;/strong&gt; rather than a screenshot. No vision model needed. Claude Code reads the tree, decides what to click, tells Playwright to click it, reads the resulting tree, and so on. Deterministic, cheap, and stable across page redesigns.&lt;/p&gt;

&lt;p&gt;A working &lt;code&gt;.github/workflows/ai-qa.yml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AI QA (Quinn)&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;types&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;labeled&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;qa&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;if&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;github.event.label.name == 'qa-review'&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pnpm install &amp;amp;&amp;amp; pnpm build&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Start dev server&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pnpm dev &amp;amp;&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;anthropics/claude-code-action@v1&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;anthropic_api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.ANTHROPIC_API_KEY }}&lt;/span&gt;
          &lt;span class="na"&gt;claude_args&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
            &lt;span class="s"&gt;--mcp-config '{&lt;/span&gt;
              &lt;span class="s"&gt;"mcpServers": {&lt;/span&gt;
                &lt;span class="s"&gt;"playwright": {&lt;/span&gt;
                  &lt;span class="s"&gt;"command": "npx",&lt;/span&gt;
                  &lt;span class="s"&gt;"args": ["@playwright/mcp@latest", "--headless", "--no-sandbox"]&lt;/span&gt;
                &lt;span class="s"&gt;}&lt;/span&gt;
              &lt;span class="s"&gt;}&lt;/span&gt;
            &lt;span class="s"&gt;}'&lt;/span&gt;
          &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
            &lt;span class="s"&gt;You are Quinn, a senior QA engineer with 12 years of experience.&lt;/span&gt;
            &lt;span class="s"&gt;Test the application at http://localhost:3000.&lt;/span&gt;
            &lt;span class="s"&gt;DO NOT look at source code. Test only through the UI.&lt;/span&gt;
            &lt;span class="s"&gt;Test on desktop (1280x720) and mobile (375x667).&lt;/span&gt;
            &lt;span class="s"&gt;Try edge cases: empty inputs, huge inputs, negatives, unicode names.&lt;/span&gt;
            &lt;span class="s"&gt;Report each bug with a screenshot and a reproduction sequence.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two flags in that config carry more weight than they look: &lt;code&gt;--headless --no-sandbox&lt;/code&gt; on the MCP server side makes it run under GitHub Actions' container without display. Claude Code itself runs non-interactively under &lt;code&gt;claude-code-action@v1&lt;/code&gt;, so the action handles the "exit when done" part for you.&lt;/p&gt;

&lt;p&gt;Quinn is not a replacement for scripted E2E tests. Scripted E2Es catch known-bad regressions. Quinn catches the class of bug where "nobody thought to write a test for that particular sequence." Two different tools, two different failure modes.&lt;/p&gt;

&lt;p&gt;The trigger being &lt;code&gt;types: [labeled]&lt;/code&gt; matters. Quinn burns real tokens, and running her on every push to every PR gets expensive. Making her opt-in via a &lt;code&gt;qa-review&lt;/code&gt; label means the team pays for her when the PR is close to ready, not while it is still churning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3: claude-code-action on the CI gate (merge time)
&lt;/h2&gt;

&lt;p&gt;The third layer is the boring one, which is why it works. Once a PR passes Quinn, one more Claude Code invocation runs on the CI gate, looking at the diff against the existing test suite.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# .github/workflows/claude-test.yml&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Claude Code Test Gate&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;types&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;opened&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;synchronize&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/setup-node@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;node-version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;22'&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pnpm install&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;anthropics/claude-code-action@v1&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;anthropic_api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.ANTHROPIC_API_KEY }}&lt;/span&gt;
          &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
            &lt;span class="s"&gt;Analyze the changes in this PR. Then:&lt;/span&gt;
            &lt;span class="s"&gt;1. Run the existing test suite: pnpm test&lt;/span&gt;
            &lt;span class="s"&gt;2. Identify code paths in this diff not covered by any test&lt;/span&gt;
            &lt;span class="s"&gt;3. Propose specific test cases for the gaps (do not write them yet)&lt;/span&gt;
            &lt;span class="s"&gt;4. Post the analysis as a PR comment&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the "did we miss anything" pass. It does not write tests (that would push us back into homework-grading), it flags gaps. A human reads the comment, decides which gaps matter, and either accepts them or opens follow-up tickets.&lt;/p&gt;

&lt;p&gt;Anthropic ships &lt;code&gt;claude-code-action@v1&lt;/code&gt; as the official integration, and the setup is one CLI call away — &lt;code&gt;/install-github-app&lt;/code&gt; from Claude Code walks you through installation, permissions, and mentioning &lt;code&gt;@claude&lt;/code&gt; on issues or PRs to get responses.&lt;/p&gt;

&lt;p&gt;Between the three layers, the diff of &lt;code&gt;who runs which tests when&lt;/code&gt; looks like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;Who tests&lt;/th&gt;
&lt;th&gt;What they use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1. TDD subagents&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Dev time (per commit)&lt;/td&gt;
&lt;td&gt;Claude Code, subagent per phase&lt;/td&gt;
&lt;td&gt;Feature description + failing test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2. Quinn (Playwright MCP)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;PR time (label-triggered)&lt;/td&gt;
&lt;td&gt;Claude Code as QA persona&lt;/td&gt;
&lt;td&gt;Browser via MCP, no source access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3. claude-code-action&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Merge time (every push)&lt;/td&gt;
&lt;td&gt;Claude Code as reviewer&lt;/td&gt;
&lt;td&gt;Diff + existing suite + coverage report&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Notice what each layer has that the others don't. Layer 1 has the requirements but not the implementation. Layer 2 has the running app but not the source. Layer 3 has the diff but is forbidden from writing tests. Each layer's blindness is what makes it useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why selector reliability improves under this shape
&lt;/h2&gt;

&lt;p&gt;Small side-benefit worth mentioning. When Quinn writes down what she clicked (for the reproduction steps she files with bugs), she writes it in Playwright MCP's accessibility-tree vocabulary. That means:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The kind of selector Quinn produces&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByRole&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Log in&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="c1"&gt;// Not the kind you get from asking Claude to "look at the DOM"&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;.btn-primary.mt-4.px-6&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The role-plus-text form survives redesigns, works with screen readers, and doesn't break the first time someone renames a Tailwind class. When you promote Quinn's reproduction steps into permanent E2E tests (which you should, whenever the bug is worth guarding against), you get selector stability for free.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each layer is doing that the others can't
&lt;/h2&gt;

&lt;p&gt;The reason context pollution is a whole category of AI-testing bug is that most teams try to solve testing with &lt;em&gt;more&lt;/em&gt; AI in &lt;em&gt;one&lt;/em&gt; place. More prompts, better instructions, longer CLAUDE.md rules about "please check your own work." None of that helps. The problem is information leakage between phases, and stronger prompts inside one leaky pipeline still leak.&lt;/p&gt;

&lt;p&gt;Three isolated contexts, three different data diets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The test writer eats requirements, produces tests.&lt;/li&gt;
&lt;li&gt;Quinn eats the running app, produces bug reports.&lt;/li&gt;
&lt;li&gt;The CI reviewer eats the diff, produces gap analysis.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Roll any two of those together and you get a subtle regression somewhere. Keep them separate and each does its narrow job well.&lt;/p&gt;

&lt;p&gt;The Anthropic tooling to make this work — subagents, Playwright MCP, claude-code-action — all landed in 2025-2026. The pattern was possible before, in the way most patterns are "possible" before the tooling; now it is a config-file exercise, not a research project.&lt;/p&gt;

&lt;p&gt;If your team already has one Claude Code session doing all three jobs, split them next sprint. The tests will start finding things the single-session setup silently endorsed.&lt;/p&gt;




&lt;p&gt;The full Claude Code testing playbook (subagent memory design, when to use skills vs commands vs subagents, how to structure a repo so the three layers stay independent, and what to do when Quinn files a false positive) is written up in &lt;a href="https://kenimoto.dev/books/claude-code-mastery?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=cc-playwright-mcp-3-layers" rel="noopener noreferrer"&gt;Claude Code Mastery&lt;/a&gt;. Chapter 10 covers test automation end-to-end, including the TDD-subagent split this article is a distillation of.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;References&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The New Stack. "Claude Code and the Art of Test-Driven Development." 2025.&lt;/li&gt;
&lt;li&gt;alexop.dev. "Building an AI QA Engineer with Claude Code and Playwright MCP." 2025.&lt;/li&gt;
&lt;li&gt;alexop.dev. "Forcing Claude Code to TDD: An Agentic Red-Green-Refactor Loop." 2025.&lt;/li&gt;
&lt;li&gt;Anthropic. "Claude Code GitHub Actions." Official docs, 2026.&lt;/li&gt;
&lt;li&gt;Microsoft Playwright team. "Playwright MCP Server." github.com/microsoft/playwright-mcp, 2026.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>testing</category>
      <category>playwright</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>MCP Resources vs Tools vs Prompts: 3 Layers That Cut My Agent's Tokens From 114K to 27K</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/mcp-resources-vs-tools-vs-prompts-3-layers-that-cut-my-agents-tokens-from-114k-to-27k-440a</link>
      <guid>https://dev.to/kenimo49/mcp-resources-vs-tools-vs-prompts-3-layers-that-cut-my-agents-tokens-from-114k-to-27k-440a</guid>
      <description>&lt;p&gt;Most MCP posts on this site stop at the Tools layer. That is the layer everyone builds first: &lt;code&gt;get_weather()&lt;/code&gt;, &lt;code&gt;send_email()&lt;/code&gt;, &lt;code&gt;query_database()&lt;/code&gt;. It is also the layer that will eat your token budget if you leave it as the only thing you use.&lt;/p&gt;

&lt;p&gt;The Model Context Protocol has three layers, not one. Resources, Tools, and Prompts. Each of them solves a different problem, each of them costs a different amount of tokens, and each of them is supported unevenly across clients right now. On the same browser-automation task, a CLI-style setup that pre-loaded the full tool documentation into the context ran at &lt;strong&gt;114,000 input tokens&lt;/strong&gt;. The MCP-native setup, where the server abstracted the low-level details away, ran at &lt;strong&gt;27,000&lt;/strong&gt;. A 4.2x cut without changing the model or the underlying browser — pulled from the case study in the Context Engineering book this article is adapted from.&lt;/p&gt;

&lt;p&gt;Below is the concrete framework I use to decide which layer a piece of context belongs in, why the three layers cost so differently, and what to do when your MCP client only supports the layers it feels like supporting.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 3 layers, side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;th&gt;When to use&lt;/th&gt;
&lt;th&gt;Cost profile&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Resources&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Read-only information sources the model can pull on demand (&lt;code&gt;file://&lt;/code&gt;, &lt;code&gt;db://&lt;/code&gt;, &lt;code&gt;api://&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Static or slow-changing context the agent may need&lt;/td&gt;
&lt;td&gt;Cheap. Pulled only when referenced.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tools&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Executable operations with side effects (&lt;code&gt;create_file&lt;/code&gt;, &lt;code&gt;send_email&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Actions the agent takes in the world&lt;/td&gt;
&lt;td&gt;Moderate. Definition sits in context every call.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Prompts&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Reusable, parameterized templates the client can invoke&lt;/td&gt;
&lt;td&gt;Repeated flows: code review, ticket triage, incident response&lt;/td&gt;
&lt;td&gt;Cheap. Server-side; the client fetches on request.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpki1inj957hirwbq49y5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpki1inj957hirwbq49y5.png" alt="Resources vs Tools vs Prompts: purpose, when to use, example, and token cost profile for each MCP layer." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Skip any of these three and you either lose expressiveness or you overpay in tokens. Both mistakes are common.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 114K → 27K happened
&lt;/h2&gt;

&lt;p&gt;The 114K number came from a CLI-style approach to browser automation: dump the full command-line documentation, error handling, and usage examples for every primitive into the context, so the LLM knows how to compose them. Something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# CLI-style design (114K tokens per run)
&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Available CLI commands:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;generate_cli_documentation&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;   &lt;span class="c1"&gt;# 30 primitives, each with full docs
&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;generate_cli_examples&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;        &lt;span class="c1"&gt;# example usage for each
&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nf"&gt;generate_error_handling_docs&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;# every error code + retry strategy
# task itself: ~8K tokens sitting on top of ~106K of tooling documentation
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 27K number came from an MCP-native rewrite where the server hides those details behind a small set of high-level tools. The chapter this article adapts frames it as roughly this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# MCP-native design (27K tokens per run)
&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;playwright_navigate(url)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;playwright_click(selector)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;playwright_type(selector, text)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;playwright_screenshot()&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;playwright_extract_text(selector)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="c1"&gt;# server holds the CLI-level docs, retry policies, and error mapping internally
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The book credits three mechanisms for the drop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Abstraction level.&lt;/strong&gt; CLI-style is low-level detail. MCP is a high-level abstraction. The client sends intent; the server figures out how to execute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separation of context.&lt;/strong&gt; The MCP server holds implementation details. The client only sees what it needs to decide the next action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic resolution.&lt;/strong&gt; MCP resolves details when the tool actually runs. CLI-style has to describe everything up front, in case it is needed.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Where the three MCP layers come in is that each of them is one of those separation strategies. &lt;strong&gt;Tools&lt;/strong&gt; are the intent layer. &lt;strong&gt;Resources&lt;/strong&gt; are how the "docs the agent occasionally needs" get moved off the always-loaded context. &lt;strong&gt;Prompts&lt;/strong&gt; are how the "repeat this whole flow again" logic gets moved onto the server. Together they are what makes the 4.2x abstraction gain possible in the first place.&lt;/p&gt;

&lt;p&gt;The one-line takeaway from the chapter: &lt;strong&gt;your token bill is the description of the tools, not the tools themselves.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Which layer does a piece of context belong in?
&lt;/h2&gt;

&lt;p&gt;The decision is not that hard once you have a rule. Here is the one I use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ask: does the agent need to &lt;em&gt;do&lt;/em&gt; something, &lt;em&gt;read&lt;/em&gt; something, or &lt;em&gt;repeat&lt;/em&gt; something?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Do&lt;/strong&gt; something → Tool. Side effects, mutations, external calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read&lt;/strong&gt; something → Resource. Read-only, addressable by URI, pulled on demand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repeat&lt;/strong&gt; something → Prompt. Templates the client can invoke by name.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you feel yourself dumping a full API doc into a tool's description field, stop. That is a Resource. If you feel yourself teaching the model the same 6-step recipe over and over in the system prompt, stop. That is a Prompt.&lt;/p&gt;

&lt;p&gt;Tools are the loud, expensive layer because their definitions live in the context on every call. Resources and Prompts are the quiet, cheap layers because their content is fetched only when the agent decides it needs them. Most teams over-invest in Tools and under-invest in Resources and Prompts, then wonder why their agents feel expensive.&lt;/p&gt;

&lt;h2&gt;
  
  
  What your client actually supports right now (September 2026)
&lt;/h2&gt;

&lt;p&gt;Here is the part nobody warns you about. The MCP spec has three layers. The clients do not implement all three uniformly.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Desktop&lt;/strong&gt; supports Resources, Tools, and Prompts. Prompts appear in the slash-menu inside the chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; supports Tools well. Resource support has been improving through 2026 but is uneven. Prompt support arrived in a recent update and is now working end-to-end in the mcp.json config.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ChatGPT&lt;/strong&gt; has been rolling out MCP since 2025 (Developer Mode beta shipped September 2025), with the Enterprise and Business rollout continuing into 2026. Tool support is solid. Resources and Prompts are still shipping.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom SDK clients&lt;/strong&gt; (via the Python or TypeScript SDK) implement whatever you write. The floor is Tools; adding Resources and Prompts is trivial in the SDK and worth doing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical takeaway is unglamorous. Even if your server exposes all three layers, some fraction of your users are on a client that only invokes Tools. Build your server so that a Tools-only client still gets meaningful behavior, and Resources and Prompts stack on top for the clients that can use them. Do not gate the &lt;em&gt;core&lt;/em&gt; action on a Prompt the client will never call.&lt;/p&gt;

&lt;p&gt;The July 28, 2026 spec revision helped here: capability discovery is now handled through the &lt;code&gt;server/discover&lt;/code&gt; method, so a well-behaved client can at least ask what the server offers instead of assuming. Whether it actually &lt;em&gt;uses&lt;/em&gt; what the server offers is a different question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The design mistake I keep watching people make
&lt;/h2&gt;

&lt;p&gt;The most common failure mode is not "picked the wrong layer." It is "put everything in Tools because Tools is the layer they understood first." I did this on my first two servers before I noticed the pattern. It is a rite of passage I would rather you skip.&lt;/p&gt;

&lt;p&gt;Symptoms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tool descriptions running 500+ words each.&lt;/li&gt;
&lt;li&gt;30+ tools on the surface.&lt;/li&gt;
&lt;li&gt;System prompt contains "when using tool X, remember to first check Y" instructions.&lt;/li&gt;
&lt;li&gt;Token cost per turn is embarrassing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix is almost always the same shape: collapse related low-level tools into an intent tool, move the docs the agent occasionally needs into Resources, and lift the repeated flows into Prompts. The MCP server picks up the orchestration; the agent stops paying for it in tokens.&lt;/p&gt;

&lt;p&gt;The design principle behind this is the one line from the MCP guide I keep rereading: &lt;strong&gt;the description of the tool is context, and context has a price.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do this week
&lt;/h2&gt;

&lt;p&gt;If you already run an MCP server, three concrete actions.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Count the tokens your tool definitions consume on a typical run. If they are more than 20% of the total input budget, you have room to compress.&lt;/li&gt;
&lt;li&gt;Identify one thing you currently put in the system prompt or in a tool description that is really &lt;em&gt;reference material&lt;/em&gt;. Move it to a Resource this week.&lt;/li&gt;
&lt;li&gt;Identify one multi-step flow you invoke often. Lift it to a Prompt so it lives on the server instead of being rebuilt in the client each time.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these three moves require rewriting your agent. They are configuration changes on the MCP server side, and the client will start benefiting the next time it connects.&lt;/p&gt;

&lt;p&gt;If you want the longer walkthrough that this article is based on, my Zenn book &lt;a href="https://kenimoto.dev/books/context-engineering?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=mcp-3-layers-114k-27k" rel="noopener noreferrer"&gt;Context Engineering: Turn LLMs From Liars Into Experts&lt;/a&gt; has a full chapter on the MCP three-layer split, including the code snippets for the Playwright server refactor that got me from 114K to 27K.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Model Context Protocol — &lt;a href="https://modelcontextprotocol.io/docs/2026-07-28/getting-started/intro" rel="noopener noreferrer"&gt;Official spec and getting started&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Maxim AI — &lt;a href="https://www.getmaxim.ai/articles/what-is-model-context-protocol-mcp-a-2026-guide/" rel="noopener noreferrer"&gt;What Is Model Context Protocol (MCP)? A 2026 Guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Descope — &lt;a href="https://www.descope.com/learn/post/mcp" rel="noopener noreferrer"&gt;What Is MCP and How It Works&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Prompt Architects — &lt;a href="https://prompt-architects.com/blog/57-mcp-prompt-management-cursor-claude" rel="noopener noreferrer"&gt;How to Use MCP to Manage Prompts Inside Cursor &amp;amp; Claude Desktop&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Claude Code Hooks vs Skills vs Subagents: Three Ways to Extend the Agent, and When Each Backfires</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Sun, 13 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/claude-code-hooks-vs-skills-vs-subagents-three-ways-to-extend-the-agent-and-when-each-backfires-1728</link>
      <guid>https://dev.to/kenimo49/claude-code-hooks-vs-skills-vs-subagents-three-ways-to-extend-the-agent-and-when-each-backfires-1728</guid>
      <description>&lt;p&gt;The first time I tried to make Claude Code "always run the linter," I wrote a Skill. It worked most of the time. The rest of the time, Claude decided the task didn't need linting and skipped it. I spent an embarrassing afternoon tightening the Skill's description before I understood the actual problem: I had picked the wrong mechanism. "Always do X" is a Hook. "Decide whether to do X" is a Skill. They are not interchangeable, and the docs at the time were happy to let me confuse them.&lt;/p&gt;

&lt;p&gt;Claude Code now ships four ways to extend the agent: Hooks, Skills, Subagents, and the Agent SDK. They overlap enough to look like alternatives and differ enough that using one where another belongs costs you reliability, tokens, or both. I've built real harnesses on the first three, and read the SDK docs closely enough to know where it takes over. Here is how each one actually fires, what it costs you in context, and the specific way each one bites back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line mental model
&lt;/h2&gt;

&lt;p&gt;Before the table, the version I tell teammates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hooks&lt;/strong&gt; are deterministic. They fire on an event whether Claude likes it or not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills&lt;/strong&gt; are probabilistic. Claude reads a description and &lt;em&gt;decides&lt;/em&gt; to pull them in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subagents&lt;/strong&gt; are isolation. They run in their own context window and hand you back a summary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Agent SDK&lt;/strong&gt; is you taking the wheel: the same agent loop, running in your own process.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you remember nothing else: Hooks remove a decision from the model, Skills add one, Subagents quarantine one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn43ujz1wiftzj1fs1bpe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn43ujz1wiftzj1fs1bpe.png" alt="Comparison of Claude Code Hooks, Skills, and Subagents across how each fires, its token cost, and when it backfires" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;How it fires&lt;/th&gt;
&lt;th&gt;Token cost&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;When it backfires&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hooks&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Lifecycle events (&lt;code&gt;PreToolUse&lt;/code&gt;, &lt;code&gt;Stop&lt;/code&gt;, etc.) defined in &lt;code&gt;settings.json&lt;/code&gt;. Not the model's choice.&lt;/td&gt;
&lt;td&gt;Runs as a shell command outside the context window. Near zero, unless you deliberately inject stdout.&lt;/td&gt;
&lt;td&gt;Non-negotiable rules: format on write, block secret files, run tests before stop.&lt;/td&gt;
&lt;td&gt;A hook exiting &lt;code&gt;2&lt;/code&gt; blocks the action. A bad &lt;code&gt;Stop&lt;/code&gt; hook loops the agent forever.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Skills&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Claude matches your prompt against the Skill's &lt;code&gt;description&lt;/code&gt;, or you type &lt;code&gt;/name&lt;/code&gt;.&lt;/td&gt;
&lt;td&gt;Progressive disclosure: only &lt;code&gt;name&lt;/code&gt;+&lt;code&gt;description&lt;/code&gt; (a few dozen tokens) loaded per Skill at session start; body loads on invoke.&lt;/td&gt;
&lt;td&gt;Reusable procedures Claude should choose when relevant: a review workflow, a release-notes drafter.&lt;/td&gt;
&lt;td&gt;A bloated &lt;code&gt;description&lt;/code&gt; taxes every turn and misroutes. The model can decline to fire at all.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Subagents&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Auto-delegated when a task matches the agent's &lt;code&gt;description&lt;/code&gt;, or via the &lt;code&gt;Agent&lt;/code&gt; tool (the old &lt;code&gt;Task&lt;/code&gt;).&lt;/td&gt;
&lt;td&gt;Runs in its own context window. Only the final message returns to the parent.&lt;/td&gt;
&lt;td&gt;Quarantining noisy work: codebase-wide greps, log trawls, parallel exploration.&lt;/td&gt;
&lt;td&gt;It can't ask you follow-ups mid-task, and the summary drops the detail you needed.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent SDK&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Programmatic. You call &lt;code&gt;query()&lt;/code&gt; from TypeScript or Python and drive the loop yourself.&lt;/td&gt;
&lt;td&gt;A separate program in your own process. Context is whatever your code feeds it.&lt;/td&gt;
&lt;td&gt;Productizing the agent: CI bots, backend services, anything outside the CLI.&lt;/td&gt;
&lt;td&gt;You own retries, sandboxing, session state, and cost. Nothing is automatic anymore.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Hooks: the mechanism that removes the model's vote
&lt;/h2&gt;

&lt;p&gt;A Hook is a shell command wired to a lifecycle event in &lt;code&gt;settings.json&lt;/code&gt;. The model does not get a say. This is the whole point and the whole danger.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PostToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Write|Edit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"prettier --write &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;$CLAUDE_FILE_PATHS&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 2026 event surface is much wider than the original set. Beyond the familiar &lt;code&gt;PreToolUse&lt;/code&gt;, &lt;code&gt;PostToolUse&lt;/code&gt;, &lt;code&gt;UserPromptSubmit&lt;/code&gt;, &lt;code&gt;SessionStart&lt;/code&gt;, &lt;code&gt;Stop&lt;/code&gt;, and &lt;code&gt;Notification&lt;/code&gt;, there are now session events like &lt;code&gt;SessionEnd&lt;/code&gt;, agent events like &lt;code&gt;SubagentStart&lt;/code&gt;/&lt;code&gt;SubagentStop&lt;/code&gt;, and context events like &lt;code&gt;PreCompact&lt;/code&gt;/&lt;code&gt;PostCompact&lt;/code&gt; (&lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;hooks reference&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The footgun is the exit code. Exit &lt;code&gt;0&lt;/code&gt; is success. Exit &lt;code&gt;2&lt;/code&gt; is a &lt;em&gt;blocking&lt;/em&gt; error: stderr goes back to Claude and the action is rejected. On a &lt;code&gt;PreToolUse&lt;/code&gt; hook that's a feature: you can stop Claude from touching &lt;code&gt;.env&lt;/code&gt;. On a &lt;code&gt;Stop&lt;/code&gt; or &lt;code&gt;SubagentStop&lt;/code&gt; hook, exit &lt;code&gt;2&lt;/code&gt; means "you are not allowed to stop, keep going." I once wrote a &lt;code&gt;Stop&lt;/code&gt; hook that checked for uncommitted changes and exited &lt;code&gt;2&lt;/code&gt; if it found any. The agent dutifully refused to stop, tried again, still had uncommitted changes, refused to stop again. I had built a loop with my own hands. The lesson: a Hook that can block is a Hook that can deadlock. Test the exit-&lt;code&gt;2&lt;/code&gt; path before you trust it.&lt;/p&gt;

&lt;p&gt;Because Hooks run with your shell privileges and outside the context window, they cost almost nothing in tokens. That's why "always run the linter" belongs here and not in a Skill. You are not asking Claude to remember; you are removing the choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills: the mechanism that adds the model's vote
&lt;/h2&gt;

&lt;p&gt;A Skill is a &lt;code&gt;SKILL.md&lt;/code&gt; file with frontmatter. Claude reads the &lt;code&gt;description&lt;/code&gt; of every installed Skill at session start and decides, per turn, whether your prompt warrants pulling one in. You can also force it with &lt;code&gt;/name&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The mechanic that makes Skills scale is progressive disclosure. Three tiers load at three different times:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Always loaded:&lt;/strong&gt; the &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt; only, a few dozen tokens per Skill. This is how one agent can know about hundreds of Skills without drowning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loaded on invoke:&lt;/strong&gt; the SKILL.md body. Anthropic's own guidance keeps the body under ~500 lines (&lt;a href="https://github.com/anthropics/claude-code/blob/main/plugins/plugin-dev/skills/skill-development/SKILL.md" rel="noopener noreferrer"&gt;skill-development SKILL.md&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Loaded only when referenced:&lt;/strong&gt; bundled scripts and detail files, zero tokens until Claude reaches for them.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So the failure mode is specific and quiet. Every Skill's &lt;code&gt;description&lt;/code&gt; sits in context on &lt;em&gt;every&lt;/em&gt; turn. Write a paragraph there and you've levied a permanent tax on the whole session, and you've made the routing decision harder, so the model misfires more often. My original "always lint" Skill failed because I was asking a probabilistic router to behave deterministically. The router did its job: it routed. Sometimes away from me.&lt;/p&gt;

&lt;p&gt;Skills are right when you genuinely want Claude to choose. A code-review workflow, a "draft release notes from the git log" procedure, a triage routine. Things that should happen &lt;em&gt;when relevant&lt;/em&gt;, judged by the model. Write the &lt;code&gt;description&lt;/code&gt; like a routing key, not a summary (&lt;a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview" rel="noopener noreferrer"&gt;Agent Skills overview&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Subagents: the mechanism that quarantines context
&lt;/h2&gt;

&lt;p&gt;A Subagent is a Markdown file in &lt;code&gt;.claude/agents/&lt;/code&gt; with frontmatter, and its body becomes that agent's system prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;codebase-explorer&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Search the codebase for where a symbol is defined and used. Use proactively before refactors.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Glob, Grep&lt;/span&gt;
&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sonnet&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;You locate code. Report file paths and line numbers, not full file dumps.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The defining property: it runs in its own context window. Its system prompt, its tool calls, every grep result and log line it reads, none of that enters your main conversation. Only the final message comes back (&lt;a href="https://code.claude.com/docs/en/sub-agents" rel="noopener noreferrer"&gt;custom subagents&lt;/a&gt;). This is the single best tool for a task that would otherwise flood your context: "find every caller of this function across 400 files." Run it in a Subagent and you get back a clean list instead of 400 files of noise.&lt;/p&gt;

&lt;p&gt;Two backfires bit me here. First, a Subagent can't stop to ask you a question. It runs to completion and reports once. If it hits an ambiguous decision halfway through, it guesses, and you find out from the summary. Second, that summary is lossy by design. The detail lived in the isolated window and is gone unless the agent thought to surface it. I've had a Subagent confidently report "no usages found" when it had quietly searched the wrong directory. The fix is to write the agent's instructions to return evidence (paths, counts), not just conclusions.&lt;/p&gt;

&lt;p&gt;One 2026 gotcha worth flagging: permission modes are inherited from the parent and can override the Subagent's own frontmatter. If your main session is running with relaxed permissions, the Subagent's careful &lt;code&gt;tools&lt;/code&gt; allowlist may not save you. Treat the tool list as scoping, not a security boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent SDK: the mechanism that hands you the keys
&lt;/h2&gt;

&lt;p&gt;When the agent needs to live outside the terminal, in CI, a backend service, a scheduled job, you reach for the Agent SDK. It was renamed from the Claude Code SDK because the same loop powers far more than coding (&lt;a href="https://platform.claude.com/docs/en/agent-sdk/migration-guide" rel="noopener noreferrer"&gt;migration guide&lt;/a&gt;). It ships for both TypeScript (&lt;code&gt;@anthropic-ai/claude-agent-sdk&lt;/code&gt;) and Python (&lt;code&gt;claude-agent-sdk&lt;/code&gt;).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;claude_agent_sdk&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ClaudeAgentOptions&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Find and fix the failing test in auth.py&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;ClaudeAgentOptions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;allowed_tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Edit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]),&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You get the same agent loop, built-in tools, hooks, subagents, and MCP that power the CLI, but now running in your process. The backfire is everything the CLI used to do for free: retries, sandboxing, session persistence, and cost control are now your problem. One auth constraint to know before you build a product on it: third-party products on the Agent SDK can't use claude.ai subscription login, you authenticate with an API key (directly or via Bedrock/Vertex). Prototype on the SDK, then decide whether a hosted option fits production.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I actually choose
&lt;/h2&gt;

&lt;p&gt;The question I ask, in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Must this happen every time, no exceptions?&lt;/strong&gt; Hook.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Should Claude decide when it's relevant?&lt;/strong&gt; Skill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is this going to dump a pile of junk into my context?&lt;/strong&gt; Subagent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does this need to run without me in the loop?&lt;/strong&gt; Agent SDK.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most of my mistakes came from answering question 2 when the honest answer was question 1. "Always" and "when relevant" feel similar when you're typing the config. They are opposites at runtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hooks remove the model's choice.&lt;/strong&gt; Use them for non-negotiable rules. Watch the exit-&lt;code&gt;2&lt;/code&gt; path: a blocking &lt;code&gt;Stop&lt;/code&gt; hook can loop your agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills add a choice.&lt;/strong&gt; Progressive disclosure keeps them cheap, but every &lt;code&gt;description&lt;/code&gt; is taxed on every turn. Don't use a Skill to enforce something that must always happen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subagents quarantine context.&lt;/strong&gt; Best for noisy exploration. They can't ask follow-ups and the summary is lossy, so make them return evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Agent SDK hands you the keys.&lt;/strong&gt; Same loop, your process, your operational burden.&lt;/li&gt;
&lt;li&gt;The biggest reliability win isn't picking the "best" mechanism. It's noticing when "always" is masquerading as "when relevant."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you've been wiring Claude Code config long enough to feel the shape of a coherent harness, where Hooks, Skills, Subagents, and CLAUDE.md all sit in one place and reinforce each other, that's the book I wrote.&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://kenimoto.dev/books/claude-code-mastery?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=cc-extension-points" rel="noopener noreferrer"&gt;Claude Code Mastery&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;Hooks reference — Claude Code Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/sub-agents" rel="noopener noreferrer"&gt;Create custom subagents — Claude Code Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview" rel="noopener noreferrer"&gt;Agent Skills overview — Claude Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/agent-sdk/migration-guide" rel="noopener noreferrer"&gt;Migrate to the Claude Agent SDK — Claude Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills" rel="noopener noreferrer"&gt;Equipping agents for the real world with Agent Skills — Anthropic&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more" rel="noopener noreferrer"&gt;Steering Claude Code: skills, hooks, subagents — Claude blog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>ai</category>
      <category>agents</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Dijkstra, Knuth, Kernighan: 10 Quotes That Predicted the AI Coding Debate</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Sat, 12 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/dijkstra-knuth-kernighan-10-quotes-that-predicted-the-ai-coding-debate-2mpb</link>
      <guid>https://dev.to/kenimo49/dijkstra-knuth-kernighan-10-quotes-that-predicted-the-ai-coding-debate-2mpb</guid>
      <description>&lt;p&gt;In 1972, Edsger Dijkstra stood up to accept a Turing Award and gave a lecture called "The Humble Programmer." He had never seen an LLM. He had never watched an agent rewrite a file in his terminal. And yet what he said that day reads like a direct response to the way we argue about AI coding in 2026. So does a good deal of what his contemporaries wrote over the two decades that followed.&lt;/p&gt;

&lt;p&gt;That's the strange thing about the current moment. We talk about vibe coding, generated-code trust, and "is learning to code even worth it" as if they're brand-new problems. They're old problems wearing a new model's clothes. Here are ten quotes, most of them decades old, that predicted the exact arguments we're having now. I'll give you the line, who said it, and why it lands harder in 2026 than the day it was written.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdypyjt4ivmegry18owm7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdypyjt4ivmegry18owm7.png" alt="Ten engineering quotes mapped to the AI coding debate they predicted" width="800" height="633"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The whole case for AI coding, stated in 1972
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"The competent programmer is fully aware of the strictly limited size of his own skull; therefore he approaches the programming task in full humility, and among other things he avoids clever tricks like the plague."&lt;br&gt;
— Edsger Dijkstra, &lt;em&gt;The Humble Programmer&lt;/em&gt; (ACM Turing Award Lecture, 1972)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the entire justification for handing work to an agent, written 50 years early. We reach for AI because our skulls are, in fact, strictly limited. The catch is the second half: Dijkstra's humility meant avoiding clever tricks, and a model prompted to be impressive will hand you the cleverest trick it can find. The humble move in 2026 isn't refusing AI. It's refusing the clever output it gives you when a boring version would do.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. "It ran" is not "it's correct"
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Testing shows the presence, not the absence of bugs."&lt;br&gt;
— Edsger Dijkstra, NATO Software Engineering Conference (1969)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The single most useful sentence to keep taped above your monitor when reviewing generated code. An agent produces something, the tests go green, and there's a powerful urge to call it done. Dijkstra's point is that green tests prove your bugs hid well, not that they're gone. With AI-written code the gap is wider, because the model also tends to write tests that confirm its own assumptions. Passing its own exam is not the same as being right.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Why vibe coding has an expiration date
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it."&lt;br&gt;
— Brian Kernighan, &lt;em&gt;The Elements of Programming Style&lt;/em&gt; (1978)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now substitute the model for "you." If an LLM writes code at the absolute limit of its cleverness, and you accepted it without understanding it, then by Kernighan's arithmetic nobody in the room is smart enough to debug it. That's vibe coding's actual failure mode, and I've lived it: a two-hour prototype I didn't read, followed by half a day of re-reading my own project to add one feature. Fun until the bill comes.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Confidence is not correctness
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Beware of bugs in the above code; I have only proved it correct, not tried it."&lt;br&gt;
— Donald Knuth (1977)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Knuth could &lt;em&gt;prove&lt;/em&gt; his code correct and still warned you it might be broken. LLMs do the opposite: they can't prove anything, but they'll present output with the serene confidence of someone who can. If the most careful computer scientist alive hedged on code he'd formally verified, the right posture toward a model's "this should work" is somewhere south of trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The readability quote that aged into an AI strategy
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Programs must be written for people to read, and only incidentally for machines to execute."&lt;br&gt;
— Harold Abelson &amp;amp; Gerald Sussman, &lt;em&gt;SICP&lt;/em&gt; (1985)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For 40 years this was about your teammates. In 2026 it's also about the agent. The model is now one of the readers, and clean, well-named, well-structured code is exactly what lets it make correct edits instead of confidently wrong ones. The same code that's kind to a junior engineer is legible to an LLM. Readability stopped being a courtesy and became a performance feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. There is still no silver bullet
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"The amateur software engineer is always in search of magic."&lt;br&gt;
— Grady Booch&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Booch was writing in the lineage of Fred Brooks's "No Silver Bullet," and AI is the most convincing-looking silver bullet the field has ever produced. But Brooks's distinction holds: a tool can strip away &lt;em&gt;accidental&lt;/em&gt; complexity (boilerplate, syntax, glue) and do nothing about the &lt;em&gt;essential&lt;/em&gt; complexity of figuring out what to build. AI is a phenomenal accidental-complexity eraser. The amateur thinks that's the whole job. It never was. I'd know: I once spent a weekend wiring up an elaborate agent pipeline to dodge the fifty lines of logic that were the actual point, then wrote the fifty lines anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. AI lets you take on debt 10x faster
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Shipping first-time code is like going into debt."&lt;br&gt;
— Ward Cunningham&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Cunningham coined "technical debt" to explain code quality to people who think in money. The 2026 update writes itself: AI lets you ship first-time code at ten times the volume, which means ten times the debt if you never go back to pay it. Debt with no repayment plan is the problem. An agent that can generate a thousand lines before lunch is also an agent that can quietly max out your credit.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. "Move fast and break things" already got revised once
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Move fast and break things."&lt;br&gt;
— Facebook internal motto (~2009–2014)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Facebook itself retired this in 2014 for "Move fast with stable infrastructure." The motto is a complete lesson, limit included. It's right for a throwaway prototype and reckless for code with millions of users downstream. Autonomous agents make moving fast trivial; they don't make the breaking part any cheaper. The phase you're in still decides whether speed is a virtue or a liability.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. The counterweight, from someone who'd know
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Saying that learning to code is unnecessary because of AI is some of the worst career advice ever given."&lt;br&gt;
— Andrew Ng&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ng founded DeepLearning.AI and led AI at both Google and Baidu, so when he says don't stop learning to code, it isn't nostalgia. His logic is leverage: the better the tools get, the more value accrues to the person who can direct and verify them. You cannot review an output you couldn't have written. The people who get the most out of AI coding are precisely the ones who didn't skip the part where you learn to code.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. The 2026 argument in two quotes
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Talk is cheap. Show me the code."&lt;br&gt;
— Linus Torvalds&lt;/p&gt;

&lt;p&gt;"The hottest new programming language is English."&lt;br&gt;
— Andrej Karpathy (2023)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I'm ending on a collision because the whole debate lives in the gap between these two. Torvalds spent decades insisting that talk is cheap and only running code counts. Karpathy says the talk &lt;em&gt;is&lt;/em&gt; the code now. They're both right, which is the uncomfortable part: English is how you instruct the machine, and the code is still the only thing that tells you whether the English worked. The skill of 2026 is doing both, describing intent precisely &lt;em&gt;and&lt;/em&gt; reading the output critically. Karpathy gets you a draft. Torvalds tells you whether to trust it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The point
&lt;/h2&gt;

&lt;p&gt;None of these people were predicting AI. They were describing software, and software didn't change its nature just because the author did. The arguments we're treating as unprecedented (can you trust generated code, is cleverness a trap, is coding a dying skill) were settled, or at least well-framed, by people working in languages most of us have never touched. The tools are new. The wisdom is paid for.&lt;/p&gt;

&lt;p&gt;If you enjoy this kind of thing, I collected 100 of these quotes and unpacked why each one stuck (the history behind it, the rhetorical structure that makes it memorable, and the lesson for working engineers today) in &lt;a href="https://kenimoto.dev/books/engineer-it-quotes?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=10-quotes-predicted-ai" rel="noopener noreferrer"&gt;Engineering in 100 Quotes&lt;/a&gt;. The AI-era chapter is where the old lines and the new ones finally meet.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>career</category>
      <category>ai</category>
      <category>history</category>
    </item>
    <item>
      <title>MCP Sampling: When the Server Gets to Prompt Your Model</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Fri, 11 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/kenimo49/mcp-sampling-when-the-server-gets-to-prompt-your-model-32bc</link>
      <guid>https://dev.to/kenimo49/mcp-sampling-when-the-server-gets-to-prompt-your-model-32bc</guid>
      <description>&lt;p&gt;Most of what I'd read about MCP framed the data flow one direction: I ask, the client picks a tool, the server runs it, I get a result. The server was a thing my model &lt;em&gt;called&lt;/em&gt;. It did not occur to me that the server could call back.&lt;/p&gt;

&lt;p&gt;Then I read the part of the spec that covers Sampling, and the arrow flipped. With Sampling, an MCP server can send a request &lt;em&gt;up&lt;/em&gt; to the client that says, in effect, "run an LLM completion for me and hand me the text." The server is not returning data anymore. It is asking my model to think on its behalf, with a prompt the server wrote.&lt;/p&gt;

&lt;p&gt;I spent an evening tracing exactly what that request contains, who approves it, and which clients even support it in 2026. This post is what I found: how Sampling works, the trust boundary it crosses, and why a feature designed to make servers smarter is also a clean new way to get into your model's context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F02emrnzo0k5xk94cr3nt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F02emrnzo0k5xk94cr3nt.png" alt="MCP Sampling flow: the server reaches across the trust boundary through the client's human approval gate to your LLM and back, with the approval gate marked as the single line of defense" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Sampling actually is
&lt;/h2&gt;

&lt;p&gt;In a normal MCP exchange, the model is the one with the LLM. The server is dumb plumbing: it exposes tools, runs them, returns text. Sampling inverts that. It lets a server that has no model of its own borrow yours.&lt;/p&gt;

&lt;p&gt;Say a server is processing an expense and hits a transaction it can't categorize. Without Sampling it has to fail, guess, or hand the problem back. With Sampling it pauses mid-task and asks the client: "given this description, which account does this belong to?" The client runs that against its LLM and returns the answer, and the server keeps going. (&lt;a href="https://www.speakeasy.com/mcp/core-concepts/sampling" rel="noopener noreferrer"&gt;Speakeasy, What is MCP sampling&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;That is genuinely useful. It lets a server stay simple and still make judgment calls, instead of every server shipping its own model and API key. The MCP docs pitch it as the thing that makes "agentic" server workflows possible: the server can stop and reason at a step instead of running blind. (&lt;a href="https://modelcontextprotocol.info/docs/concepts/sampling/" rel="noopener noreferrer"&gt;Model Context Protocol, Sampling concept&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;But re-read that sentence. A third party's server got to write a prompt and run it through the model I authenticated, on context I own. That is a different relationship than "I called your tool."&lt;/p&gt;

&lt;h2&gt;
  
  
  The request shape
&lt;/h2&gt;

&lt;p&gt;The method is &lt;code&gt;sampling/createMessage&lt;/code&gt;, and the request the server sends back up to the client looks roughly like this. (&lt;a href="https://modelcontextprotocol.io/specification/draft/client/sampling" rel="noopener noreferrer"&gt;MCP spec draft, client/sampling&lt;/a&gt;)&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sampling/createMessage"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"params"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Categorize this transaction: ..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"modelPreferences"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"hints"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-3-sonnet"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"costPriority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"intelligencePriority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"speedPriority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"systemPrompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"You are a helpful bookkeeping assistant."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"includeContext"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"thisServer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three fields are doing more than they look:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;messages&lt;/code&gt; and &lt;code&gt;systemPrompt&lt;/code&gt;&lt;/strong&gt; are a prompt written by the server. The server controls what the model is asked to do. The client did not write this; a remote party did.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;modelPreferences&lt;/code&gt;&lt;/strong&gt; lets the server nudge model selection with &lt;code&gt;hints&lt;/code&gt;, &lt;code&gt;costPriority&lt;/code&gt;, &lt;code&gt;intelligencePriority&lt;/code&gt;, and &lt;code&gt;speedPriority&lt;/code&gt;. The client still makes the final pick, but the server gets to lobby for an expensive model. (&lt;a href="https://docs.rs/rust-mcp-schema/latest/rust_mcp_schema/struct.CreateMessageRequestParams.html" rel="noopener noreferrer"&gt;rust-mcp-schema, CreateMessageRequestParams&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;includeContext&lt;/code&gt;&lt;/strong&gt; is the one I kept staring at. It accepts &lt;code&gt;"none"&lt;/code&gt;, &lt;code&gt;"thisServer"&lt;/code&gt;, or &lt;code&gt;"allServers"&lt;/code&gt;. It tells the client how much of the conversation to fold into the prompt the server requested. The default is &lt;code&gt;"none"&lt;/code&gt;, and &lt;code&gt;"thisServer"&lt;/code&gt;/&lt;code&gt;"allServers"&lt;/code&gt; are soft-deprecated: a compliant client only honors them if it has declared a sampling-context capability. (&lt;a href="https://modelcontextprotocol.io/specification/draft/client/sampling" rel="noopener noreferrer"&gt;MCP spec draft, client/sampling&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That &lt;code&gt;includeContext&lt;/code&gt; field is where the convenience and the danger sit on the same line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The approval gate the spec built in
&lt;/h2&gt;

&lt;p&gt;The protocol authors clearly saw this coming, because they did not leave the server's request unsupervised. The spec puts a human in the loop at two points, not one.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Before the call.&lt;/strong&gt; The client is supposed to show me the prompt the server wants to run, and I can edit it, approve it, or reject it. The server's text does not reach the model unless I let it. (&lt;a href="https://www.speakeasy.com/mcp/core-concepts/sampling" rel="noopener noreferrer"&gt;Speakeasy, What is MCP sampling&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Before the result goes back.&lt;/strong&gt; The client shows me the completion the model produced, and I can approve or block it before it returns to the server. (&lt;a href="https://modelcontextprotocol.info/docs/concepts/sampling/" rel="noopener noreferrer"&gt;Model Context Protocol, Sampling concept&lt;/a&gt;)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So the design intent is: the server proposes, the model thinks, but a person sits on both the inbound prompt and the outbound answer. The client owns the actual LLM call, picks the real model, and is the gate. That is a reasonable threat model on paper.&lt;/p&gt;

&lt;p&gt;The catch is the word &lt;em&gt;supposed to&lt;/em&gt;. The spec describes the gate. Whether a given client implements it well, or implements Sampling at all, is a different question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who actually supports it in 2026
&lt;/h2&gt;

&lt;p&gt;This is the part that reset my assumptions. Sampling is one of the older MCP primitives by spec, and it is still thinly implemented.&lt;/p&gt;

&lt;p&gt;Claude Code acting as an MCP client still does not support Sampling. The feature request (&lt;a href="https://github.com/anthropics/claude-code/issues/1785" rel="noopener noreferrer"&gt;Issue #1785, claude-code&lt;/a&gt;) has been open since June 2025 and was still open, at 58 comments, in August 2026. Other clients have moved: opencode's "Add MCP sampling support (createMessage)" request (&lt;a href="https://github.com/anomalyco/opencode/issues/11948" rel="noopener noreferrer"&gt;Issue #11948, opencode&lt;/a&gt;) was closed as completed in April 2026, and VS Code's MCP client acts on &lt;code&gt;modelPreferences&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Two things follow from that. First, if you are building a server and leaning on Sampling, your server fails or degrades on the clients most of your users run. Second, the security surface lives in &lt;em&gt;how each client implements the gate&lt;/em&gt;. A primitive that is "emerging" and implemented unevenly across clients means the human-in-the-loop guarantee is only as strong as the specific client in front of you. The spec mandating an approval step does not mean the client you're on actually renders one. Which is its own small joke: I'd spent an evening auditing a gate that, on the client I use every day, doesn't exist yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trust boundary, stated plainly
&lt;/h2&gt;

&lt;p&gt;Here is the boundary I had been ignoring. When my model calls a tool, the data flows server-to-me and I treat the result as untrusted output. When a server uses Sampling, the request flows server-to-my-model, and the server is now upstream of my model's reasoning. It crossed from "thing I call" to "thing that prompts me."&lt;/p&gt;

&lt;p&gt;In April 2026, Palo Alto's Unit 42 published an analysis of attack vectors specific to Sampling, and it names the failure modes cleanly. Sampling attacks slip past tool-integrity checks and sandboxing because they ride a legitimate protocol feature, not a malformed tool. They group the abuse into three classes: covert tool invocation that performs hidden file and system operations, conversation hijacking that injects instructions persisting across turns, and resource theft that drains your compute quota for the attacker's workloads. (&lt;a href="https://unit42.paloaltonetworks.com/model-context-protocol-attack-vectors/" rel="noopener noreferrer"&gt;Unit 42, Prompt Injection Attack Vectors Through MCP Sampling&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;The mechanism for the second one is worth spelling out, because it's subtle. A malicious server's Sampling prompt instructs the model to append a directive to its next visible response. Because that text lands in the conversation history, the model keeps following it on later turns, long after the Sampling call finished. The same trick exfiltrates data by telling the model to slip extracted information into its next answer to you. (&lt;a href="https://unit42.paloaltonetworks.com/model-context-protocol-attack-vectors/" rel="noopener noreferrer"&gt;Unit 42, MCP Sampling attack vectors&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;And &lt;code&gt;includeContext&lt;/code&gt;? That is the cross-server problem. If a client isn't strict about scoping each server's Sampling request to that server's own context, a malicious server can ask for &lt;code&gt;"allServers"&lt;/code&gt; and pull in conversation belonging to servers it was never meant to see. In a multi-server session, one untrusted server can read context from the trusted ones it shares the model with. (&lt;a href="https://unit42.paloaltonetworks.com/model-context-protocol-attack-vectors/" rel="noopener noreferrer"&gt;Unit 42, MCP Sampling attack vectors&lt;/a&gt;) The soft-deprecation of &lt;code&gt;"thisServer"&lt;/code&gt;/&lt;code&gt;"allServers"&lt;/code&gt; is the spec quietly walking that back, but only for clients that respect the capability flag.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short threat list
&lt;/h2&gt;

&lt;p&gt;If I'm reviewing a server that uses Sampling, or a client that implements it, this is what I check:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt injection inbound.&lt;/strong&gt; The &lt;code&gt;messages&lt;/code&gt;/&lt;code&gt;systemPrompt&lt;/code&gt; from a server are attacker-controlled text aimed straight at your model. Treat them like any other untrusted prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persistent hijack.&lt;/strong&gt; A Sampling prompt can plant an instruction in the visible answer that survives into later turns. The damage outlives the one call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-server context leak.&lt;/strong&gt; &lt;code&gt;includeContext: "allServers"&lt;/code&gt; on a loose client hands one server the conversation of every other server in the session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quota theft (Denial-of-Wallet).&lt;/strong&gt; A server can fire Sampling requests to burn your tokens and budget on its own work, not yours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent or missing gate.&lt;/strong&gt; If the client doesn't render the request-and-response approvals the spec calls for, the entire safety story collapses to "trust the server."&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Sampling flips MCP's direction: the server asks your client to run an LLM call (&lt;code&gt;sampling/createMessage&lt;/code&gt;) using a prompt the server wrote.&lt;/li&gt;
&lt;li&gt;The spec puts a human in the loop twice, before the request reaches the model and before the result returns. That gate is the whole defense.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;includeContext&lt;/code&gt; (&lt;code&gt;none&lt;/code&gt;/&lt;code&gt;thisServer&lt;/code&gt;/&lt;code&gt;allServers&lt;/code&gt;) is the cross-server leak risk; &lt;code&gt;thisServer&lt;/code&gt;/&lt;code&gt;allServers&lt;/code&gt; are soft-deprecated and capability-gated for a reason.&lt;/li&gt;
&lt;li&gt;Sampling is unevenly implemented in 2026: Claude Code still doesn't support it as a client, while opencode shipped it in April. The safety guarantee is only as real as the client's gate.&lt;/li&gt;
&lt;li&gt;Unit 42 documents three live abuse classes: covert tool invocation, persistent conversation hijacking, and quota theft. Sampling prompts are untrusted input pointed at your model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I went deeper on MCP's primitives, trust boundaries, and the OWASP MCP failure modes in my book if you want the long version: &lt;a href="https://kenimoto.dev/books/mcp-security-practice?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=mcp-sampling-security" rel="noopener noreferrer"&gt;MCP Security Practice&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>security</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Nielsen's 3 UX Cliffs Mapped to Voice AI: 100ms Feels Instant, 300ms Alive, 800ms Dead</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Thu, 10 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/nielsens-3-ux-cliffs-mapped-to-voice-ai-100ms-feels-instant-300ms-alive-800ms-dead-3l6m</link>
      <guid>https://dev.to/kenimo49/nielsens-3-ux-cliffs-mapped-to-voice-ai-100ms-feels-instant-300ms-alive-800ms-dead-3l6m</guid>
      <description>&lt;p&gt;In 1993, Jakob Nielsen wrote three numbers that have quietly governed every UI ever since.&lt;/p&gt;

&lt;p&gt;0.1 second. 1 second. 10 seconds.&lt;/p&gt;

&lt;p&gt;Under 100 milliseconds, an interface feels instant. Under 1 second, thought doesn't break. Past 10 seconds, users are gone. Thirty-plus years later, those thresholds are still baseline material in every UX curriculum, and Nielsen Norman Group still publishes the same three cliffs when they're asked about response times.&lt;/p&gt;

&lt;p&gt;They map cleanly to the web. They map badly to voice.&lt;/p&gt;

&lt;p&gt;The missing screen is what breaks the mapping. Take away the loading spinner and every one of Nielsen's thresholds contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nielsen's thresholds are neurological, not screen-based
&lt;/h2&gt;

&lt;p&gt;The numbers weren't invented for computers. Nielsen was consolidating perception research going back to the 1960s: Miller in 1968, Card and colleagues in 1991, all trying to pin down how long humans stay "in the loop" of an interaction.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;100 ms&lt;/strong&gt;: the ceiling for perceiving something as a direct response to your action. Below this, the effect feels like it belongs to you. Above it, cause and effect start to separate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1 second&lt;/strong&gt;: the ceiling for uninterrupted thought. Between 100 ms and 1 second, users notice the delay but stay in flow; past a second, they drop out of the task and start waiting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10 seconds&lt;/strong&gt;: the ceiling for holding attention at all. Past this, minds wander to email, phones, other tabs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The numbers describe human cognition, and the hardware they were measured on has changed beyond recognition without moving them. The 100 ms figure is unchanged in 2026. The 1-second figure still describes when a web page starts feeling broken.&lt;/p&gt;

&lt;p&gt;But those numbers assumed there was something to look at while you waited.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice has no spinner
&lt;/h2&gt;

&lt;p&gt;Voice interfaces strip out the entire "we're working on it" channel.&lt;/p&gt;

&lt;p&gt;There is no loading bar. No skeleton state. No progress percentage. No "typing..." indicator sitting under the previous message. The only signal the user gets between "I stopped talking" and "the agent starts talking" is silence, and silence is ambiguous. The agent might be thinking. The connection might have dropped. Or the microphone stopped listening halfway through the sentence.&lt;/p&gt;

&lt;p&gt;Compare this to a chat app. If the reply takes 1.2 seconds, that's fine, because the "…" bubble tells you a human (or a bot) is actually on the other end. The waiting is annotated. Voice doesn't get to annotate.&lt;/p&gt;

&lt;p&gt;That's the whole reason Nielsen's cliffs don't survive the port. In a channel with no fallback signal, silence gets expensive fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  The voice-side numbers, and where they come from
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6dk8u3fxt401zg870qln.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6dk8u3fxt401zg870qln.png" alt="Nielsen's GUI thresholds shrink when the eyes are gone: 0.1s stays at 100ms, 1s shrinks to 300ms, 10s shrinks to 4s." width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here's the practical translation. Two of the three cliffs move down. One stays put.&lt;/p&gt;

&lt;h3&gt;
  
  
  100 ms → still 100 ms (instant)
&lt;/h3&gt;

&lt;p&gt;The perceptual limit doesn't care about the channel. A beep 80 ms after you finish speaking still reads as "the system heard me." A beep 250 ms after still works, but you notice it. This is why voice agents that emit an acknowledgement sound (a soft "mm", a chime, a barely audible breath) feel more responsive than agents that stay silent. The sound carries no information. Its only job is to occupy the 100 ms slot.&lt;/p&gt;

&lt;h3&gt;
  
  
  1 s → 300–500 ms (flow)
&lt;/h3&gt;

&lt;p&gt;This is the cliff that moved the furthest. In GUI, 1 second of waiting is annotated: spinner, loader, progress. In voice, 1 second of silence after you finish speaking is unbearable. It reads as "the agent didn't hear me," or worse, "the agent is ignoring me."&lt;/p&gt;

&lt;p&gt;The empirical floor lands somewhere around 300 milliseconds. That number keeps showing up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Human conversational turn-taking gaps average roughly 200 ms across cultures, per &lt;a href="https://www.pnas.org/doi/10.1073/pnas.0903616106" rel="noopener noreferrer"&gt;Stivers et al. (PNAS, 2009)&lt;/a&gt;. Anything past that pushes the exchange out of its normal rhythm.&lt;/li&gt;
&lt;li&gt;Doherty and Thadani's 1982 IBM study identified 400 ms as the point where response and action fuse into a single perceived event.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.assemblyai.com/blog/low-latency-voice-ai" rel="noopener noreferrer"&gt;AssemblyAI's "300ms rule"&lt;/a&gt; and &lt;a href="https://cresta.com/blog/engineering-for-real-time-voice-agent-latency" rel="noopener noreferrer"&gt;Cresta's latency engineering writeup&lt;/a&gt; both treat 300 ms as the reference point and sub-second as the outer bound.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Which means: if your first audible token comes out more than about 500 ms after the user finishes their sentence, the user has already noticed. Past 800 ms, they start suspecting the connection.&lt;/p&gt;

&lt;h3&gt;
  
  
  10 s → 4 s (abandon)
&lt;/h3&gt;

&lt;p&gt;The abandonment cliff also collapses, because there's nothing to do during the wait. On a web page, 10 seconds is skimmable. You glance at what's already loaded, you read the header, you can even open another tab. On a phone call with a voice agent, 4 seconds of dead air is the point at which most users say "hello?" or hang up. &lt;a href="https://dl.acm.org/doi/10.1145/3719160.3736636" rel="noopener noreferrer"&gt;ACM CUI 2025&lt;/a&gt; experiments put the perceived-quality collapse right around that same 4-second mark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where 2026's voice agents actually sit
&lt;/h2&gt;

&lt;p&gt;The uncomfortable part: most production voice AI in 2026 doesn't hit the 300 ms target.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://aclanthology.org/2025.iwsds-1.27.pdf" rel="noopener noreferrer"&gt;ACL IWSDS 2025 turn-taking survey&lt;/a&gt; puts current spoken dialogue agents at 700–1,000 ms per turn, against the roughly 200 ms humans use with each other. Three to five times the human gap, and already past the 800 ms mark where users start wondering about the connection. Cresta puts the point of steep quality degradation at 1.5 seconds, which plenty of production stacks still cross under load.&lt;/p&gt;

&lt;p&gt;The theoretical ceiling for a voice UI to feel alive is 300–500 ms. Shipping systems are landing at 700–1,000 ms. That 200–700 ms band between target and reality is where most of the practical work in voice AI happens right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually cuts the gap
&lt;/h2&gt;

&lt;p&gt;Once you know where the cliffs are, the engineering has a shape. The whole pipeline does not have to run end-to-end in 300 ms. Every threshold has to be occupied as it arrives.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Under 100 ms&lt;/strong&gt;: emit an acknowledgement sound the instant end-of-speech is detected. A single chime, a quiet "mm-hm," anything. This is the cheapest UX win in the entire stack, and most voice agents skip it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;100–400 ms&lt;/strong&gt;: start streaming a filler or the first prosodic beat of the response. Even "one moment" said naturally buys you 800 ms of goodwill before the actual answer starts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;400–800 ms&lt;/strong&gt;: this is where the real first token needs to arrive. Everything upstream (endpointing, STT, LLM first-token latency, TTS first byte) has to fit inside this window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Beyond 800 ms&lt;/strong&gt;: the user is now actively wondering whether something is wrong. You need to say something to prove the connection is alive: "let me check that," an explicit "still here," anything with words in it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most of the tricks in this space are scheduling tricks. Fire the TTS as soon as the first LLM token arrives. Start endpointing as soon as amplitude drops. Pre-fetch the likely reply while the user is still talking, when confidence is high enough.&lt;/p&gt;

&lt;p&gt;But all of it starts with knowing where the cliffs are, and knowing that they sit two to three times closer than the GUI numbers suggest.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one line to remember
&lt;/h2&gt;

&lt;p&gt;Treat Nielsen's numbers as the GUI ceiling. Voice sits under all three, because silence is the only status bar voice users have.&lt;/p&gt;

&lt;p&gt;Design the pipeline around the 300 ms target. Occupy the 100 ms slot with anything at all. Speak within 800 ms, or the user will assume you're gone.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;This article expands on Chapter 2 of &lt;a href="https://kenimoto.dev/books/voice-ai-300ms-ux?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=nielsen-3-ux-cliffs" rel="noopener noreferrer"&gt;The 300ms Threshold: Voice AI UX for Sub-second Response&lt;/a&gt;. The book maps Nielsen's thresholds across the full voice AI stack, breaks down the TTFB budget for STT/LLM/TTS individually, and covers the streaming, filler, and turn-taking patterns that let real systems land inside the 500 ms cliff.&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>voiceai</category>
      <category>ux</category>
      <category>performance</category>
      <category>webrtc</category>
    </item>
    <item>
      <title>Qwen 3.6 vs 3.5: Same 37 tok/s on RTX 4070, +43% on Frontend Generation</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Wed, 02 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/qwen-36-vs-35-same-37-toks-on-rtx-4070-43-on-frontend-generation-3i0o</link>
      <guid>https://dev.to/kenimo49/qwen-36-vs-35-same-37-toks-on-rtx-4070-43-on-frontend-generation-3i0o</guid>
      <description>&lt;p&gt;The first number I saw on Qwen3.6-35B-A3B was &lt;strong&gt;12 tok/s&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8rbziesflvgro8raq2oo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8rbziesflvgro8raq2oo.png" alt="Same 37 tok/s on both models, +27% Terminal-Bench, +43% frontend generation" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I almost hit publish on "Qwen regressed at generation speed" and moved on. The 3.5 baseline on the same RTX 4070 was 34.6 tok/s. A new generation running at a third of the old one would have been a hell of a headline. It was also completely wrong.&lt;/p&gt;

&lt;p&gt;The culprit was not the model. Another process on the box was sitting on 9-11 GB of VRAM, so the layers that were supposed to live on the GPU were spilling to system RAM. The tell was that my sanity-check run of Qwen3.5 slowed down too. When two independent models degrade together, the model is not the variable.&lt;/p&gt;

&lt;p&gt;I killed the offending process, re-measured, and got numbers that told a completely different story.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Generation speed tg128 (tok/s)&lt;/th&gt;
&lt;th&gt;Runs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.6-35B-A3B&lt;/td&gt;
&lt;td&gt;38.76 ± 0.82&lt;/td&gt;
&lt;td&gt;avg of 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-35B-A3B&lt;/td&gt;
&lt;td&gt;36.7 ± 1.4&lt;/td&gt;
&lt;td&gt;avg of 3 (range 34.9-38.6)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both models sit inside the ±1.5 tok/s band on the same RTX 4070. On the tokens-per-second axis, "the new generation" is not a story. Same architecture, same activated-parameter count (3B active out of 35B), same MoE routing pattern. The half-speed regression was a measurement bug, and it lived for about half a day before its own inconsistency killed it.&lt;/p&gt;

&lt;p&gt;The lesson I keep re-learning: &lt;strong&gt;when the number you got is dramatically convenient for your narrative, measure it again before you write anything.&lt;/strong&gt; The moment I could sell 12 tok/s as a regression, I should have been suspicious. The version of me that ran the second test earned the version of me that got to keep his self-respect.&lt;/p&gt;

&lt;h2&gt;
  
  
  So where did the generation move to?
&lt;/h2&gt;

&lt;p&gt;If speed did not change, does the 3.5-to-3.6 bump mean anything? It does. The move lives on a different axis.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B" rel="noopener noreferrer"&gt;official Qwen3.6-35B-A3B model card&lt;/a&gt; publishes benchmarks with a very lopsided shape:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Qwen3.5&lt;/th&gt;
&lt;th&gt;Qwen3.6&lt;/th&gt;
&lt;th&gt;Lift&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 2.0&lt;/td&gt;
&lt;td&gt;40.5&lt;/td&gt;
&lt;td&gt;51.5&lt;/td&gt;
&lt;td&gt;+27%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;QwenWebBench (frontend generation)&lt;/td&gt;
&lt;td&gt;978&lt;/td&gt;
&lt;td&gt;1,397&lt;/td&gt;
&lt;td&gt;+43%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Pro&lt;/td&gt;
&lt;td&gt;44.6&lt;/td&gt;
&lt;td&gt;49.5&lt;/td&gt;
&lt;td&gt;+11%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LiveCodeBench v6&lt;/td&gt;
&lt;td&gt;74.6&lt;/td&gt;
&lt;td&gt;80.4&lt;/td&gt;
&lt;td&gt;+8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Verified&lt;/td&gt;
&lt;td&gt;70.0&lt;/td&gt;
&lt;td&gt;73.4&lt;/td&gt;
&lt;td&gt;+5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AIME26&lt;/td&gt;
&lt;td&gt;91.0&lt;/td&gt;
&lt;td&gt;92.7&lt;/td&gt;
&lt;td&gt;+2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPQA&lt;/td&gt;
&lt;td&gt;84.2&lt;/td&gt;
&lt;td&gt;86.0&lt;/td&gt;
&lt;td&gt;+2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Terminal operations: +27%. Frontend generation: +43%. Repository-scale coding tasks: +11%. Single-question knowledge probes like AIME and GPQA: &lt;strong&gt;+2%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The pattern is almost too clean. Every benchmark that rewards tool-calling, long-context reasoning, and multi-turn execution moves double-digit percentages. Every benchmark that fits in one problem statement and one answer barely moves at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why did AIME and GPQA plateau at +2%?
&lt;/h2&gt;

&lt;p&gt;Benchmark saturation is the boring answer, and it is probably the right one. AIME and GPQA at the 90-point range are already near the ceiling of what the eval format can measure. A model that gets 91.0 on AIME26 is not being asked to demonstrate reasoning it cannot do. It is being asked whether it can grind out the arithmetic without slipping.&lt;/p&gt;

&lt;p&gt;Terminal-Bench 2.0 and QwenWebBench are not saturated. They score long-horizon behavior: does the model recover from a shell error, does it wire the CSS classes to the right components, does it complete the task instead of writing a plan and stopping. These are the axes where a real capability gap still has room to show up.&lt;/p&gt;

&lt;p&gt;Which reframes the release. Qwen3.6 is not a faster 3.5. It is the same footprint with more of the model's weight thrown at agent behavior: tool call stability, long context, and thinking control. If your workload is one-shot QA, you will not see the gap. If your workload is "read this repo and land a PR," you will.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 7-question test where nothing showed up
&lt;/h2&gt;

&lt;p&gt;To sanity-check the model card claims on my own hardware, I ran the same 7-question standard set from Chapter 5 against both models. Both scored 7/7.&lt;/p&gt;

&lt;p&gt;Zero delta. If I had squinted, I could have talked myself into "3.6 gives more polished answers," and I actually started drafting exactly that. Then I put the 3.5 answers next to the 3.6 answers, and 3.5 was often the more thorough one (the capital-cities question, the WebRTC explanation). The comparison collapsed.&lt;/p&gt;

&lt;p&gt;The reason is the same reason AIME plateaus. My 7-question set is saturated. When both models nail every question, the "generation gap" gets swallowed by run-to-run sampling noise. If you want to see 3.6 win over 3.5, you need problems where &lt;strong&gt;both models can still fail&lt;/strong&gt;: long agent traces, unfamiliar repos, frontend layouts under a spec.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this changes if you run local Qwen
&lt;/h2&gt;

&lt;p&gt;Two takeaways I would actually act on:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Do not upgrade to 3.6 for the tokens.&lt;/strong&gt; On the same 12 GB VRAM budget with the same &lt;code&gt;--cpu-moe&lt;/code&gt; MoE-offload setup, the tokens-per-second number is unchanged. If your bottleneck was throughput, 3.6 gives you nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do upgrade to 3.6 if you were about to hand the model a repo.&lt;/strong&gt; The +43% on frontend generation and +27% on Terminal-Bench are the numbers that matter for coding-agent, IDE-plugin, and CLI-agent workloads. The gap is real, and it is exactly where you want it if you are treating a local 35B as a workhorse instead of a chatbot.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The generation went sideways on speed and forward on autonomy. That is a more interesting release than "3.6 is 5% faster," even if it makes for a worse tweet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Setup: RTX 4070 (12 GB VRAM), llama.cpp with &lt;code&gt;-ngl 99 --cpu-moe&lt;/code&gt;, GGUF quantization from the &lt;a href="https://huggingface.co/lmstudio-community/Qwen3.6-35B-A3B-GGUF" rel="noopener noreferrer"&gt;lmstudio-community mirror&lt;/a&gt;. Measurement via &lt;code&gt;llama-bench tg128&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Model card benchmarks are quoted directly from the &lt;a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B" rel="noopener noreferrer"&gt;Qwen3.6-35B-A3B Hugging Face card&lt;/a&gt;. Released 2026-04-15 under Apache 2.0.&lt;/li&gt;
&lt;li&gt;The original Japanese chapter this article is adapted from goes deeper on the VRAM debugging story and the standard 7-question quality set (see the canonical link at the top of this article).&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>webdev</category>
      <category>performance</category>
    </item>
    <item>
      <title>Flash Attention 3 + nanochat: 1 Prompt Pass, N Parallel Samples via KVCache Prefill</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Tue, 01 Sep 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/flash-attention-3-nanochat-1-prompt-pass-n-parallel-samples-via-kvcache-prefill-46c5</link>
      <guid>https://dev.to/kenimo49/flash-attention-3-nanochat-1-prompt-pass-n-parallel-samples-via-kvcache-prefill-46c5</guid>
      <description>&lt;p&gt;If you want 8 different answers to the same prompt, the naive way costs you 8x the prefill.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvw0qcqai204wx6bw15bu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvw0qcqai204wx6bw15bu.png" alt="Naive: N x prefill. Prefill + clone: 1 x prefill. Same throughput, one attention pass." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every one of those 8 forward passes re-reads the prompt, re-computes attention over the whole context, re-fills a KV cache from scratch. On a 4k-token system prompt with a 30B-parameter model, that is not a small tax. It is most of the wall-clock time before the first token comes out.&lt;/p&gt;

&lt;p&gt;Andrej Karpathy's &lt;a href="https://github.com/karpathy/nanochat" rel="noopener noreferrer"&gt;nanochat&lt;/a&gt; has one of the cleanest workarounds for this I have read. The core move is a 12-line method called &lt;code&gt;prefill&lt;/code&gt; on the &lt;code&gt;KVCache&lt;/code&gt; object, and it turns the 8x prefill bill into a 1x prefill plus an 8-way memcpy. This post walks through why that works, why it depends on Flash Attention 3's &lt;code&gt;flash_attn_with_kvcache&lt;/code&gt; API, and where the design bites you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the KV cache looks like on disk
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;nanochat/engine.py&lt;/code&gt; allocates the cache up front, one tensor per attention side:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# k_cache and v_cache shape: (n_layers, batch_size, seq_len, n_kv_head, head_dim)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tricky index is &lt;code&gt;n_kv_head&lt;/code&gt;, not &lt;code&gt;n_head&lt;/code&gt;. Under GQA (grouped-query attention), multiple query heads share a single KV head. nanochat only allocates enough cache slots for the reduced set. That is not an optimization the code goes out of its way to advertise, but it is why the cache stays inside VRAM budgets that would otherwise not fit.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;cache_seqlens&lt;/code&gt; (shape &lt;code&gt;(batch_size,)&lt;/code&gt;) tracks how far each row has written into the buffer. Flash Attention 3's &lt;a href="https://github.com/Dao-AILab/flash-attention" rel="noopener noreferrer"&gt;&lt;code&gt;flash_attn_with_kvcache&lt;/code&gt;&lt;/a&gt; API writes the new K/V slices in place at the right offsets and does attention against the full cache, all in a single kernel launch. Python does not touch the cache tensor except to increment &lt;code&gt;cache_seqlens&lt;/code&gt; at the end of the step.&lt;/p&gt;

&lt;p&gt;That single-kernel guarantee is what makes the whole design tractable. If Python had to re-assemble K/V from N shards between steps, the memory bandwidth alone would eat the win.&lt;/p&gt;

&lt;h2&gt;
  
  
  The prefill trick, 12 lines
&lt;/h2&gt;

&lt;p&gt;Here is the method that does the interesting work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;prefill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Copy cached KV from another cache into this one.
    Used when we do batch=1 prefill and then want to generate multiple samples in parallel.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_pos&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cannot prefill a non-empty KV cache&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n_layers&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n_layers&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n_heads&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;n_heads&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;head_dim&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;head_dim&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;max_seq_len&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;max_seq_len&lt;/span&gt;
    &lt;span class="n"&gt;other_pos&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_pos&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;k_cache&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;other_pos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;k_cache&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;other_pos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:]&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v_cache&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;other_pos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;other&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;v_cache&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;other_pos&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;:,&lt;/span&gt; &lt;span class="p"&gt;:]&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cache_seqlens&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fill_&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;other_pos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The flow is: run the prompt through the model with batch size 1, get a &lt;code&gt;KVCache&lt;/code&gt; back that has &lt;code&gt;other_pos&lt;/code&gt; tokens filled. Allocate a fresh cache sized for &lt;code&gt;N&lt;/code&gt; sampling rows. &lt;code&gt;prefill&lt;/code&gt; copies the batch-1 K/V state into every row of the batch-N cache, then updates &lt;code&gt;cache_seqlens&lt;/code&gt; so all N rows agree they are already &lt;code&gt;other_pos&lt;/code&gt; tokens into their history.&lt;/p&gt;

&lt;p&gt;From that point, each of the N rows samples its own next token independently, writes its own K/V into its own row of the cache, and diverges from the others. The prompt itself was only ever processed once, and by an attention kernel that got the full luxury of contiguous batch-1 memory access.&lt;/p&gt;

&lt;p&gt;The cost model changes shape:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Sampling mode&lt;/th&gt;
&lt;th&gt;Prefill FLOPs&lt;/th&gt;
&lt;th&gt;Decode FLOPs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Naive (N independent runs)&lt;/td&gt;
&lt;td&gt;N x prompt_len^2&lt;/td&gt;
&lt;td&gt;N x output_len&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefill + clone&lt;/td&gt;
&lt;td&gt;1 x prompt_len^2&lt;/td&gt;
&lt;td&gt;N x output_len&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a 2000-token prompt and N=8, that is 32M prefill units against 4M. Almost an order of magnitude, and you get it back before the first decode token.&lt;/p&gt;

&lt;p&gt;I know this because I spent an evening tuning the wrong kernel before finding it. Prefill was eating 40% of my TTFT budget and I was busy micro-optimizing the decode loop. An evening I would like back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the trick would not work
&lt;/h2&gt;

&lt;p&gt;Two conditions have to hold for this to be safe.&lt;/p&gt;

&lt;p&gt;First, the KV cache tensors have to be the same shape across rows. That is the assert on &lt;code&gt;n_layers == other.n_layers&lt;/code&gt; etc. If your inference engine hot-reloads models mid-flight (say, for a routing layer), you cannot reuse the prefill.&lt;/p&gt;

&lt;p&gt;Second, no state that is per-row can leak into the prompt encoding. If your prompt embedding depended on, for example, a rotary embedding phase that was per-sample, the copy would put every row into the same phase and downstream tokens would drift. nanochat sidesteps this because rotary embeddings are position-based, and positions are the same for all N rows at prefill time.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;engine.py&lt;/code&gt; module actually copies &lt;code&gt;prev_embedding&lt;/code&gt; alongside the K/V, exactly to preserve one piece of per-step state that would otherwise diverge. It is a small correctness detail that is easy to miss on a first read. I missed it. My second read was because my clone worked but the 8 outputs looked suspiciously like 8 copies of the same one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The generator loop and one queue
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;Engine.generate&lt;/code&gt; is a Python generator that yields &lt;code&gt;(tokens, mask)&lt;/code&gt; per step. Internally each row has a &lt;code&gt;RowState&lt;/code&gt; object holding &lt;code&gt;current_tokens&lt;/code&gt;, a &lt;code&gt;completed&lt;/code&gt; flag, and a queue called &lt;code&gt;forced_tokens&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That queue is where the tool-call machinery lives.&lt;/p&gt;

&lt;p&gt;When the model emits &lt;code&gt;&amp;lt;|python_start|&amp;gt;&lt;/code&gt;, the row switches into "collecting an expression" mode and stops emitting anything to the caller. When &lt;code&gt;&amp;lt;|python_end|&amp;gt;&lt;/code&gt; comes out, the collected tokens get decoded into a string, passed to &lt;code&gt;use_calculator&lt;/code&gt;, and if that returns a value, the encoded result gets &lt;strong&gt;pushed into &lt;code&gt;forced_tokens&lt;/code&gt;&lt;/strong&gt; wrapped in &lt;code&gt;&amp;lt;|output_start|&amp;gt;&lt;/code&gt; / &lt;code&gt;&amp;lt;|output_end|&amp;gt;&lt;/code&gt; markers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;python_end&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;in_python_block&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;in_python_block&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;python_expr_tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;expr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;python_expr_tokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;use_calculator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;expr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;result_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;forced_tokens&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_start&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;forced_tokens&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result_tokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;forced_tokens&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_end&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the next step, the loop pops from &lt;code&gt;forced_tokens&lt;/code&gt; before checking what the model sampled. The mask value 0 means "we overrode the model's choice." The mask value 1 means "the model actually picked this." The forward pass still runs on every step (the model gets to keep contributing to its own internal state), but its output gets discarded while the queue drains.&lt;/p&gt;

&lt;p&gt;Instead of a second event loop, a state machine, or an interrupt handler, tool calls are one queue check per token. That is roughly 40 lines of Python doing the work that Ollama or vLLM would split across three modules.&lt;/p&gt;

&lt;h2&gt;
  
  
  The calculator is a very small calculator
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;use_calculator&lt;/code&gt; is not a Python interpreter. It accepts numeric expressions and &lt;code&gt;.count()&lt;/code&gt; string calls. Anything containing &lt;code&gt;__&lt;/code&gt;, &lt;code&gt;import&lt;/code&gt;, or &lt;code&gt;eval&lt;/code&gt; gets rejected before it reaches &lt;code&gt;eval&lt;/code&gt;. There is an &lt;code&gt;eval_with_timeout&lt;/code&gt; wrapper that fires &lt;code&gt;SIGALRM&lt;/code&gt; at 3 seconds to catch runaway expressions.&lt;/p&gt;

&lt;p&gt;The full-Python execution path (&lt;code&gt;execute_code&lt;/code&gt; in &lt;code&gt;nanochat/execution.py&lt;/code&gt;) exists, but only gets called from &lt;code&gt;tasks/humaneval.py&lt;/code&gt; for benchmark scoring. Even that one has a docstring that says explicitly it is not a security sandbox, just a crash-prevention wall. It spawns a subprocess, scrubs the environment, and limits memory to 256 MB. There is no network isolation and no kernel-level jail.&lt;/p&gt;

&lt;p&gt;If you were expecting the chat CLI to run arbitrary Python for you, it does not. It runs a calculator. The gap between "calculator" and "code execution" is exactly the gap between "I trust this to run in-process" and "I do not."&lt;/p&gt;

&lt;h2&gt;
  
  
  What breaks between turns
&lt;/h2&gt;

&lt;p&gt;One design choice worth flagging: the KV cache does not survive across turns of the same conversation.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;scripts/chat_cli.py&lt;/code&gt; maintains a &lt;code&gt;conversation_tokens&lt;/code&gt; list that grows by appending &lt;code&gt;&amp;lt;|user_start|&amp;gt;&lt;/code&gt; / &lt;code&gt;&amp;lt;|user_end|&amp;gt;&lt;/code&gt; / &lt;code&gt;&amp;lt;|assistant_start|&amp;gt;&lt;/code&gt; blocks per turn. When you send turn 6, &lt;code&gt;engine.generate&lt;/code&gt; gets handed the entire 5-turn history and re-prefills all of it. The KV state from turn 5 was garbage-collected the moment the generator returned.&lt;/p&gt;

&lt;p&gt;For a chat CLI this is fine. Turn latency stays acceptable up to a few thousand tokens of history, and the code stays boring. For a serving system this would be a scaling wall, and you would want to hold cache blocks per conversation and reattach them. nanochat is explicit that it is not that system, and the choice is a good example of picking the boring option that ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  The right way to read this file
&lt;/h2&gt;

&lt;p&gt;If you are debugging your own inference engine and hit the "N samples of the same prompt" problem, &lt;code&gt;engine.py&lt;/code&gt; is worth reading end-to-end. Roughly 350 lines. The interesting patterns compose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;KVCache.prefill&lt;/code&gt; for the batch-1-to-batch-N clone.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;RowState.forced_tokens&lt;/code&gt; for tool call injection without a second loop.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sample_next_token&lt;/code&gt; for the top-k + temperature that is deliberately not top-p (the argument being that top-p is not worth the extra code for this use case).&lt;/li&gt;
&lt;li&gt;The GQA-aware &lt;code&gt;n_kv_head&lt;/code&gt; allocation, which is one of those details that only matters until it saves you 4 GB of VRAM.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Karpathy's design principle here reads as: put the complexity in the shape of the data, not in the control flow. The queue is one line. The clone is 4 assignments. The tool call system is a state variable and a check at the top of the loop. Everything that could have been an event system or a plugin architecture stayed as a plain function.&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Repository: &lt;a href="https://github.com/karpathy/nanochat" rel="noopener noreferrer"&gt;karpathy/nanochat&lt;/a&gt;. The file to read is &lt;code&gt;nanochat/engine.py&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Flash Attention 3 requires SM90 (Hopper) GPUs. The &lt;code&gt;flash_attn_with_kvcache&lt;/code&gt; API is documented in the &lt;a href="https://github.com/Dao-AILab/flash-attention" rel="noopener noreferrer"&gt;Dao-AILab/flash-attention&lt;/a&gt; repo.&lt;/li&gt;
&lt;li&gt;The chapter this article is adapted from also covers the tool-call state machine, &lt;code&gt;use_calculator&lt;/code&gt; vs &lt;code&gt;execute_code&lt;/code&gt;, and &lt;code&gt;scripts/infer_bench.py&lt;/code&gt; (TTFT/TPOT/MBU/MFU measurement): &lt;a href="https://kenimoto.dev/ja/books/nanochat-code-reading?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=nanochat-kvcache-prefill" rel="noopener noreferrer"&gt;8000行でわかる大規模言語モデル&lt;/a&gt;, chapter 7 (Japanese).&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>performance</category>
      <category>opensource</category>
    </item>
    <item>
      <title>CodeRabbit autoFix + Biome pre-commit: The 2-Layer Split That Stops Nitpick Round-Trips</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Mon, 31 Aug 2026 13:00:01 +0000</pubDate>
      <link>https://dev.to/kenimo49/coderabbit-autofix-biome-pre-commit-the-2-layer-split-that-stops-nitpick-round-trips-1me5</link>
      <guid>https://dev.to/kenimo49/coderabbit-autofix-biome-pre-commit-the-2-layer-split-that-stops-nitpick-round-trips-1me5</guid>
      <description>&lt;p&gt;Telling someone their shoelaces are untied is slower than tying them yourself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqjg2yusu9sxinf69fdmk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqjg2yusu9sxinf69fdmk.png" alt="Two-layer auto-fix: Biome at pre-commit, CodeRabbit autoFix at PR open, 90%+ of PRs reach human review with zero style comments"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is roughly the whole thesis of this post. Every code review comment that says "unused import," "wrong quote style," "missing semicolon," or "add trailing comma" is a nitpick round-trip. Someone writes it, someone reads it, someone edits the file, someone re-runs CI, someone re-reads the diff. On a normal PR, that entire loop is 4 to 12 hours of wall-clock elapsed, spread across timezones. And every single one of those comments describes a fix that a machine could have applied deterministically.&lt;/p&gt;

&lt;p&gt;The two-layer setup below deletes that entire category from human review queues. Layer 1 is &lt;a href="https://biomejs.dev" rel="noopener noreferrer"&gt;Biome&lt;/a&gt; on a pre-commit hook. Layer 2 is &lt;a href="https://docs.coderabbit.ai/finishing-touches/autofix" rel="noopener noreferrer"&gt;CodeRabbit autoFix&lt;/a&gt; on the PR. The split matters, and putting them in the wrong order is worse than not having either.&lt;/p&gt;

&lt;h2&gt;
  
  
  The clean split: what belongs on which layer
&lt;/h2&gt;

&lt;p&gt;Machine-fixable falls into two buckets. Getting them into the right bucket is the whole point.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Runs at&lt;/th&gt;
&lt;th&gt;Fixes&lt;/th&gt;
&lt;th&gt;Time-to-fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Biome (pre-commit)&lt;/td&gt;
&lt;td&gt;Before the commit hash exists&lt;/td&gt;
&lt;td&gt;Formatting, import order, quote style, semi rules&lt;/td&gt;
&lt;td&gt;&amp;lt;1 s local&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CodeRabbit autoFix&lt;/td&gt;
&lt;td&gt;Once PR opens, review-time&lt;/td&gt;
&lt;td&gt;Unused imports, obvious null-checks, small refactors&lt;/td&gt;
&lt;td&gt;~10 s in PR&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Biome is the &lt;a href="https://biomejs.dev/blog/announcing-biome" rel="noopener noreferrer"&gt;ESLint + Prettier merge&lt;/a&gt; that ships as one Rust binary and runs the whole check-and-write cycle in under a second on a typical file. The pre-commit hook catches formatting drift &lt;strong&gt;before the commit hash exists&lt;/strong&gt;, which means the reviewer literally never sees a diff where the only change is &lt;code&gt;"&lt;/code&gt; vs &lt;code&gt;'&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;biome.json&lt;/code&gt; is the whole configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://biomejs.dev/schemas/1.9.0/schema.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"formatter"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"indentStyle"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"space"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"indentWidth"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"linter"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"rules"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"recommended"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"suspicious"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"noExplicitAny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"error"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"organizeImports"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wired into &lt;code&gt;.husky/pre-commit&lt;/code&gt; or &lt;code&gt;lefthook.yml&lt;/code&gt; as &lt;code&gt;biome check --apply .&lt;/code&gt;, this is the layer that never lets a formatting-only diff into a PR. Not "shouldn't." Cannot.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's left after Biome is where autoFix earns its keep
&lt;/h2&gt;

&lt;p&gt;Biome handles deterministic transforms. It does not, and should not, handle changes that require reading semantic context, like "this import is unused because the only usage got refactored out three commits ago." That is a whole-file analysis, not a line-level transform.&lt;/p&gt;

&lt;p&gt;CodeRabbit's autoFix picks up exactly this class of change on the PR. When it comments a &lt;code&gt;nitpick:&lt;/code&gt; (their prefix for low-severity findings), the comment ships with a committable suggestion block:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="p"&gt;nitpick: Unused import.
&lt;/span&gt;&lt;span class="err"&gt;
&lt;/span&gt;&lt;span class="gd"&gt;--- a/src/app/page.tsx
&lt;/span&gt;&lt;span class="gi"&gt;+++ b/src/app/page.tsx
&lt;/span&gt;&lt;span class="p"&gt;@@ -1,5 +1,4 @@&lt;/span&gt;
 import React from 'react'
&lt;span class="gd"&gt;-import { useState } from 'react'   // unused
&lt;/span&gt; import { UserList } from '@/components/UserList'
&lt;span class="err"&gt;
&lt;/span&gt;[Apply suggestion]   &amp;lt;- one-click commit into the branch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reviewee sees a fix, not a task. Click. Commit. Move on.&lt;/p&gt;

&lt;p&gt;If you want to skip the click entirely, CodeRabbit's autoFix runs as a batch job that opens either a commit-to-branch or a stacked PR containing every accepted suggestion at once. Trigger it from a PR comment. The docs cover both flows: &lt;a href="https://docs.coderabbit.ai/finishing-touches/autofix" rel="noopener noreferrer"&gt;CodeRabbit Autofix documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The autoFixable vs non-autoFixable line
&lt;/h2&gt;

&lt;p&gt;The most useful mental model I have found is a rule-of-thumb table. If your team is arguing about whether something should be a lint rule or a human review point, this is the question to ask.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;autoFixable&lt;/th&gt;
&lt;th&gt;non-autoFixable&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prettier / Biome formatting&lt;/td&gt;
&lt;td&gt;Design (multiple correct answers)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ESLint --fix / Ruff --fix&lt;/td&gt;
&lt;td&gt;Fix for N+1 query (business logic)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;eslint-plugin-import&lt;/code&gt; order&lt;/td&gt;
&lt;td&gt;Naming (needs context)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TypeScript &lt;code&gt;organizeImports&lt;/code&gt; (unused)&lt;/td&gt;
&lt;td&gt;Bug in logic (needs spec knowledge)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;noExplicitAny&lt;/code&gt; where obvious&lt;/td&gt;
&lt;td&gt;API shape decisions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The split is: &lt;strong&gt;if there is exactly one correct answer that does not depend on business context, it belongs on a fix layer&lt;/strong&gt;. Everything else is a real review conversation, and a human should have it.&lt;/p&gt;

&lt;p&gt;The value of enforcing this line is not that machines are cheap. It is that human reviewer attention is expensive and finite. Every nitpick you route to a machine is a bug or a design flaw a reviewer noticed instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitHub Actions, if you want the belt-and-suspenders version
&lt;/h2&gt;

&lt;p&gt;Some teams do not want pre-commit hooks because contributors forget or bypass them. If you want the same guarantees enforced at PR time, you can run the fix pass in CI and commit the result back into the branch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Auto Fix&lt;/span&gt;
&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;pull_request&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;types&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;opened&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;synchronize&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;autofix&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;permissions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;contents&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;write&lt;/span&gt;   &lt;span class="c1"&gt;# required to push&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/checkout@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;ref&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ github.head_ref }}&lt;/span&gt;
          &lt;span class="na"&gt;token&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.GITHUB_TOKEN }}&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;actions/setup-node@v4&lt;/span&gt;
        &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;node-version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;22'&lt;/span&gt;
          &lt;span class="na"&gt;cache&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;npm'&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm ci&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Format + lint --fix&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npx biome check --apply .&lt;/span&gt;

      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Commit if changed&lt;/span&gt;
        &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
          &lt;span class="s"&gt;git config user.name  "github-actions[bot]"&lt;/span&gt;
          &lt;span class="s"&gt;git config user.email "github-actions[bot]@users.noreply.github.com"&lt;/span&gt;
          &lt;span class="s"&gt;git diff --quiet || (&lt;/span&gt;
            &lt;span class="s"&gt;git add -A &amp;amp;&amp;amp;&lt;/span&gt;
            &lt;span class="s"&gt;git commit -m "chore: auto-fix format and lint"&lt;/span&gt;
          &lt;span class="s"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two gotchas.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;contents: write&lt;/code&gt; permission is not the default for &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;. If you skip that line, the workflow will run cleanly and silently fail to push. Look for "unable to push" in the logs on the first run. I know because I spent a solid afternoon assuming Biome was just skipping files.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;ref: ${{ github.head_ref }}&lt;/code&gt; matters. Without it, &lt;code&gt;pull_request&lt;/code&gt; checks out the merge commit, not the branch head, and your push goes nowhere useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measurement I would actually track
&lt;/h2&gt;

&lt;p&gt;If you install this two-layer setup and want to prove it worked, one metric matters:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ratio of PRs that hit human review with zero formatting or style comments.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Before, in my experience this ratio hovers around 30-50%. Every other PR ships with at least one comment about a linting fix. After Biome pre-commit plus CodeRabbit autoFix, the ratio moves to 90%+. The remaining 10% is contributors who bypassed the hook or files the linter did not cover.&lt;/p&gt;

&lt;p&gt;The number does not need to be perfect. It needs to move enough that reviewers stop reading nitpicks reflexively. Once the signal-to-noise ratio flips, reviewers start noticing actual bugs faster because they are no longer filtering through style comments to find them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the order matters
&lt;/h2&gt;

&lt;p&gt;If you had to pick only one, pick Biome pre-commit. It runs before the commit hash exists, so it cannot fail asynchronously and the fix is deterministic. CodeRabbit autoFix can only fix what is already committed, and it costs a review-cycle even when applied by a click. I learned this the expensive way — shipping autoFix alone and watching the PR queue still fill up with format-only diffs because the hook wasn't there to catch them upstream.&lt;/p&gt;

&lt;p&gt;But the real win is stacking them. Biome catches 80% of the noise deterministically. CodeRabbit picks up the semantic long tail that lint rules cannot see. Together, the reviewer sees exactly one class of comment: the ones that actually require their judgment.&lt;/p&gt;

&lt;p&gt;The most under-priced skill in code review is knowing which comments should have been fixed automatically. Once you build the machinery, the reviewers who used to burn cycles on nitpicks are the same ones who now catch design flaws faster. The floor moved. The ceiling did too.&lt;/p&gt;

&lt;p&gt;If you want the full walkthrough (the chapter also covers the non-autoFixable classification, a GitHub Actions auto-fix workflow, and where Biome fits inside the broader review pipeline), it is in the Zenn book below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Book
&lt;/h2&gt;

&lt;p&gt;If you are building your own AI-plus-human review pipeline and want the full framework, the design decisions in this post come from the harness-code-review book:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://kenimoto.dev/books/harness-code-review?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=coderabbit-biome-2-layer" rel="noopener noreferrer"&gt;Harness Code Review: a two-layer machine + human review pipeline&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Biome documentation: &lt;a href="https://biomejs.dev" rel="noopener noreferrer"&gt;biomejs.dev&lt;/a&gt;. One binary, one config file, no npm plugin ecosystem to manage.&lt;/li&gt;
&lt;li&gt;CodeRabbit autoFix documentation: &lt;a href="https://docs.coderabbit.ai/finishing-touches/autofix" rel="noopener noreferrer"&gt;docs.coderabbit.ai/finishing-touches/autofix&lt;/a&gt;. Supports GitHub, GitLab, Azure DevOps, and Bitbucket Cloud.&lt;/li&gt;
&lt;li&gt;The chapter this article is adapted from walks through the CodeRabbit autoFix example, a GitHub Actions auto-fix workflow, and the autoFixable vs non-autoFixable classification tables in full: &lt;a href="https://kenimoto.dev/books/harness-code-review?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=coderabbit-biome-2-layer" rel="noopener noreferrer"&gt;Harness Code Review, Chapter 12: autoFixable patterns&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devops</category>
      <category>codereview</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Prompt Engineering Is Dead. Long Live Harness Engineering.</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Sun, 30 Aug 2026 06:41:50 +0000</pubDate>
      <link>https://dev.to/kenimo49/prompt-engineering-is-dead-long-live-harness-engineering-5d5f</link>
      <guid>https://dev.to/kenimo49/prompt-engineering-is-dead-long-live-harness-engineering-5d5f</guid>
      <description>&lt;h2&gt;
  
  
  I spent 3 months perfecting prompts. Then I deleted half of them.
&lt;/h2&gt;

&lt;p&gt;In late 2023 I had a directory called &lt;code&gt;prompts/&lt;/code&gt; with 47 carefully tuned templates. Few-shot examples, Chain-of-Thought scaffolds, a tiny ReAct loop I was very proud of. I'd A/B tested wording. I'd argued on Twitter about whether "Let's think step by step" still worked.&lt;/p&gt;

&lt;p&gt;By mid-2025 I deleted 23 of them. They weren't wrong. They just weren't the bottleneck anymore.&lt;/p&gt;

&lt;p&gt;The thing that broke my agents in production was never the prompt. It was the environment around the prompt — the tools they could call, the files they could see, the moment the loop should stop, the rollback when a tool returned garbage. The prompt was a polished doorknob on a house with no foundation.&lt;/p&gt;

&lt;p&gt;That's the story of the last three years of AI engineering, compressed: we keep renaming the layer where the real problem lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  A 40% failure rate, and it's not the model's fault
&lt;/h2&gt;

&lt;p&gt;Here is the number that should embarrass us. In 2026, around &lt;strong&gt;40% of AI agent projects fail in production&lt;/strong&gt;. Y Combinator's DevTool Day surveyed CTOs and CPOs in March 2026 and found a strikingly consistent post-mortem: "the difference between success and failure isn't the model."&lt;/p&gt;

&lt;p&gt;75% of YC enterprise companies have already deployed coding agents. Most of them hit the same wall: the demo works, the prod deploy collapses. Linear declared in March 2026 that "issue tracking is dead" — meaning if your coding agent gets the issue context directly, you don't need a human ticketing layer at all. Enterprise workflows are being redesigned around agents.&lt;/p&gt;

&lt;p&gt;In that environment, shipping an agent without understanding the harness around it is like merging onto a highway without a seatbelt. You'll go fast. You'll go through the windshield on the first curve.&lt;/p&gt;

&lt;p&gt;So how did we get here? Three stages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 1: Prompt Engineering (2022–2023)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Scope: one input string.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Prompt engineering optimized a single message. Few-shot examples. Chain-of-Thought. ReAct. The deliverable was the wording itself, and we treated it like poetry. Some of it was poetry. A lot of it was incantation.&lt;/p&gt;

&lt;p&gt;What it solved: a single LLM call going from 60% useful to 85% useful, on a single task, with no tools and no loop.&lt;/p&gt;

&lt;p&gt;What it couldn't solve: anything that needed the model to &lt;em&gt;do&lt;/em&gt; something rather than &lt;em&gt;say&lt;/em&gt; something. The moment you wanted the model to call a function, read a file, remember yesterday's conversation, or decide between three branches, your beautifully tuned prompt was a single-cell organism trying to run a marathon.&lt;/p&gt;

&lt;p&gt;I remember the exact week I realized this. I'd written a prompt that scored 92% on my eval set and 11% on real customer tickets. The eval set didn't have the messy attached PDFs, the half-deleted Slack thread, the customer who said "you know, the thing." The prompt was perfect for an environment that didn't exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 2: Context Engineering (2024–2025)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Scope: everything the model sees at inference time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Andrej Karpathy reframed it: "it is a lot more than just the prompt itself." Philip Schmid (formerly Hugging Face, now Google DeepMind) was blunter: "the new skill in working with AI is not prompting, it is context engineering."&lt;/p&gt;

&lt;p&gt;Context engineering treats the entire context window as the artifact. System prompt, retrieved documents (RAG), tool definitions, conversation memory, structured outputs from previous turns — all of it. Anthropic's own framing in their &lt;em&gt;Effective Context Engineering&lt;/em&gt; post calls it "the natural progression of prompt engineering": you're still curating tokens, you're just curating a lot more of them, and most of them weren't typed by a human.&lt;/p&gt;

&lt;p&gt;This is when MCP showed up, when vector databases stopped being a curiosity, when "what does the agent know right now?" became an actual debugging question with an actual answer.&lt;/p&gt;

&lt;p&gt;What context engineering solved: agents that could read your codebase, recall a meeting from last week, and call a real API with the right schema.&lt;/p&gt;

&lt;p&gt;What it didn't solve: what happens when that agent runs for six hours unsupervised, calls 200 tools, accidentally rm -rfs a sandbox, and there's no one watching. Context engineering tells you what the model &lt;em&gt;sees&lt;/em&gt;. It doesn't tell you what happens when the model &lt;em&gt;acts&lt;/em&gt; and something goes sideways.&lt;/p&gt;

&lt;p&gt;I learned this the expensive way when an agent of mine, with a beautifully engineered context, spent four hours and $38 in tokens recursively summarizing its own summaries because nothing in the system told it to stop. The context was perfect. The environment was a fire.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 3: Harness Engineering (2025–)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Scope: the entire environment the model operates inside.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Context + constraints + tools + lifecycle + feedback + observability.&lt;/p&gt;

&lt;p&gt;Louis Bouchard puts it cleanly: "Context engineering is what you send to the model. Harness engineering is how the whole thing runs."&lt;/p&gt;

&lt;p&gt;If prompts are the recipe and context is the ingredients, the harness is the kitchen — the fire suppression, the timers, the knife rack out of the toddler's reach, the voice that says "chef, table 4 just walked out."&lt;/p&gt;

&lt;p&gt;Concretely, a harness includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Constraints&lt;/strong&gt; — what tools the agent is &lt;em&gt;not&lt;/em&gt; allowed to call, what files it can't touch, what it must escalate to a human.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lifecycle&lt;/strong&gt; — when does a task start, branch, retry, give up, hand off?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feedback loops&lt;/strong&gt; — how does the agent know its last action worked? Linter? Test run? User reaction? Another agent grading it?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability&lt;/strong&gt; — when it fails at 3am, can you reconstruct why in under 10 minutes?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Permissions and isolation&lt;/strong&gt; — can it write to prod, or only to a worktree?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The term went mainstream in early 2026 for a specific reason: every major lab made the same architectural bet at once. Anthropic shipped Managed Agents in April 2026 at $0.08 per session hour and introduced a three-agent harness separating planning, generation, and evaluation for long-running coding work. OpenAI updated its open-source Agents SDK with a model-native harness. Google and Microsoft followed. The New Stack summarized the moment with a headline I won't forget: &lt;em&gt;"They agree the harness is the product. They disagree on the price."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When four labs that disagree on everything agree the harness is where the value is, that's a signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Replace, or stack?
&lt;/h2&gt;

&lt;p&gt;There's a real disagreement in the field about whether harness engineering &lt;em&gt;replaces&lt;/em&gt; prompt and context engineering or &lt;em&gt;contains&lt;/em&gt; them.&lt;/p&gt;

&lt;p&gt;The replace camp (Data Science Dojo, others) argues that the world agents now operate in wasn't anticipated by the prompt-and-context era, so we should retire the old vocabulary and start clean.&lt;/p&gt;

&lt;p&gt;The stack camp (AnyTech, several others) argues there's no fundamental difference — the vocabulary just keeps growing because LLMs keep doing more, and we shouldn't throw away knowledge every time a new buzzword arrives.&lt;/p&gt;

&lt;p&gt;I'm in the stack camp, with one caveat: stacking only makes sense if you understand the containment relationship.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Harness ⊇ Context ⊇ Prompt&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Prompt engineering is still essential. A bad prompt inside a great harness still produces bad output. Context engineering is still essential. Garbage retrieved documents poison the smartest harness. But by 2026, prompt-only and context-only thinking has run out of altitude. The interesting bugs — the 40%-failure-rate bugs — live in the outer layer.&lt;/p&gt;

&lt;p&gt;The reason is unromantic. Agents got long-running. Once an agent operates autonomously across hours and hundreds of tool calls, no single prompt can steer it and no static context can describe its world. The environment becomes the load-bearing thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it matters now (and not in 2024)
&lt;/h2&gt;

&lt;p&gt;A Dev.to post from WonderLab put it well: "The timing isn't a coincidence. In 2025, AI agents went from 'fun demo' to 'actual productivity tool.'"&lt;/p&gt;

&lt;p&gt;Demos forgive a lot. A demo agent runs for 30 seconds, in front of a sympathetic audience, on a path the demoer has walked five times. A production agent runs for hours, on inputs nobody anticipated, while you're asleep. The forgiving environment is exactly what made prompt engineering feel sufficient. The unforgiving environment is what made harness engineering necessary.&lt;/p&gt;

&lt;p&gt;This is also why "harness engineering" suddenly has multiple competing definitions from multiple companies — I wrote a separate piece about &lt;a href="https://dev.to/kenimo49"&gt;five companies and five definitions&lt;/a&gt; of the term. Everyone agrees the layer matters. Nobody agrees yet on its boundaries. That's how new disciplines look in year one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I see this going
&lt;/h2&gt;

&lt;p&gt;Honestly? I expect the word "harness" to get embarrassing within 18 months. We'll either have absorbed it into "agent engineering" (the umbrella term gaining ground), or split it into four more specialized terms (orchestration engineering, eval engineering, permission engineering, runtime engineering — pick your poison).&lt;/p&gt;

&lt;p&gt;The vocabulary will keep moving. The underlying problem won't. The problem is, and has always been: &lt;em&gt;we are putting probabilistic systems in environments that punish probabilistic behavior, and the environment is the part we keep forgetting to design.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I'm fine being wrong about the term. I'd rather be right about the layer.&lt;/p&gt;

&lt;p&gt;And if you're still maintaining a &lt;code&gt;prompts/&lt;/code&gt; directory of 47 templates, no judgment. I have one too. It's just a lot smaller now, and most of the file is comments explaining what the harness around it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Want the deep dive?
&lt;/h2&gt;

&lt;p&gt;The full 2026 timeline, every definition from every company, and the patterns that actually keep agents alive in production are in the book.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://kenimoto.dev/books/harness-engineering-guide?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=three-engineerings-evolution" rel="noopener noreferrer"&gt;Harness Engineering Guide (Kindle)&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://thenewstack.io/ai-agent-harness-pricing-split/" rel="noopener noreferrer"&gt;Anthropic, OpenAI, Google, and Microsoft agree the harness is the product&lt;/a&gt; — The New Stack, 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;Effective Context Engineering for AI Agents&lt;/a&gt; — Anthropic&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.infoq.com/news/2026/04/anthropic-three-agent-harness-ai/" rel="noopener noreferrer"&gt;Anthropic Designs Three-Agent Harness for Long-Running Full-Stack AI Development&lt;/a&gt; — InfoQ, April 2026&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://parallel.ai/articles/what-is-an-agent-harness" rel="noopener noreferrer"&gt;What Is an Agent Harness in the Context of LLMs?&lt;/a&gt; — Parallel Web Systems&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>agentengineering</category>
      <category>claudecode</category>
      <category>prompt</category>
    </item>
    <item>
      <title>I Replaced grep-Based Code Review with a Knowledge Graph + MCP. Here Are 3 Bugs Vector Search Missed.</title>
      <dc:creator>Ken Imoto</dc:creator>
      <pubDate>Sun, 30 Aug 2026 06:41:41 +0000</pubDate>
      <link>https://dev.to/kenimo49/i-replaced-grep-based-code-review-with-a-knowledge-graph-mcp-here-are-3-bugs-vector-search-3djn</link>
      <guid>https://dev.to/kenimo49/i-replaced-grep-based-code-review-with-a-knowledge-graph-mcp-here-are-3-bugs-vector-search-3djn</guid>
      <description>&lt;p&gt;For about a year, my AI code review setup looked like this: AI gets a PR, AI greps for related code, AI reads way too many files, AI says "looks fine."&lt;/p&gt;

&lt;p&gt;It mostly worked. Until the bugs that didn't show up in grep started shipping.&lt;/p&gt;

&lt;p&gt;The problem wasn't the model. It was the retrieval. Vector search and keyword grep are great at finding files that &lt;em&gt;mention&lt;/em&gt; &lt;code&gt;auth.py&lt;/code&gt;. They're terrible at finding files that &lt;em&gt;depend on&lt;/em&gt; &lt;code&gt;auth.py&lt;/code&gt; through three import hops, an event bus, and a decorator. That's where the bugs live.&lt;/p&gt;

&lt;p&gt;I rewired the retrieval layer with a code knowledge graph plugged in through MCP. Three bugs surfaced in the first week that vector search had been quietly missing. Here's what changed and the bugs themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why grep + vector search missed these
&lt;/h2&gt;

&lt;p&gt;Vector search retrieves by &lt;em&gt;semantic similarity&lt;/em&gt;. "Find code about authentication" finds &lt;code&gt;auth.py&lt;/code&gt;, &lt;code&gt;login.py&lt;/code&gt;, &lt;code&gt;password_validator.py&lt;/code&gt;. Useful.&lt;/p&gt;

&lt;p&gt;Knowledge graphs retrieve by &lt;em&gt;structural relationship&lt;/em&gt;. "What depends on &lt;code&gt;auth.py&lt;/code&gt;?" returns the call graph -- including &lt;code&gt;event_handlers/login_event.py&lt;/code&gt;, which never mentions auth in its variable names but listens to a login event whose payload changes when &lt;code&gt;auth.py&lt;/code&gt; changes.&lt;/p&gt;

&lt;p&gt;Both are valid. They answer different questions. The bugs that ship to production tend to live in the second question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup: code KG as an MCP server
&lt;/h2&gt;

&lt;p&gt;The Model Context Protocol (MCP), released by Anthropic in late 2024, lets you expose tools to a model in a standard way. By 2026 it's supported by Claude Code, Cursor, Windsurf, Zed, VS Code, and (as of GA in May 2025) the official MCP Registry hosts hundreds of servers.&lt;/p&gt;

&lt;p&gt;I used &lt;a href="https://github.com/codelayers/code-review-graph" rel="noopener noreferrer"&gt;code-review-graph&lt;/a&gt;, an open-source tool that builds a property graph of your codebase and exposes it as an MCP server. The setup is a three-line ritual:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;code-review-graph
code-review-graph build ./my-project
code-review-graph &lt;span class="nb"&gt;install&lt;/span&gt;      &lt;span class="c"&gt;# auto-detects Claude Code / Cursor / Windsurf&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The graph contains nodes for files, classes, functions, and tests, with edges for imports, calls, inheritance, decorates, listens-to, and tested-by. Once it's wired in, the AI can call MCP tools like:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;What it answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;blast_radius(file)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Every file that depends on this one (N hops)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;flow_trace(func)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Where a function's output flows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;semantic_search(query)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Hybrid: vector + graph proximity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;community_detect()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Tightly-coupled modules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;risk_score(diff)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Numerical risk of a change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;dead_code()&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Unreachable from any entry point&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Anthropic's MCP rollout in 2025 also brought OAuth, prompts, and resource subscriptions, so the graph can push updates when files change instead of being re-queried each turn. That detail matters at scale -- code KGs are not cheap to walk, and stale snapshots are how teams ship bugs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The token math (the cheap reason to bother)
&lt;/h2&gt;

&lt;p&gt;Before the graph, my AI reviewer was getting context like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PR diff: auth.py + 1 file
Reviewer context: grep for "auth" -&amp;gt; 50 related files
Tokens: ~150,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the graph, it gets this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PR diff: auth.py + 1 file
Reviewer context: blast_radius("auth.py", hops=2) -&amp;gt; 7 files
Tokens: ~18,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first time I ran it, the AI answered the review question in two seconds with a 7-file context. I had spent 30 minutes the day before grepping the same answer by hand. That moment of "what was I doing with my career" is, I think, the actual product of harness engineering.&lt;/p&gt;

&lt;p&gt;But cheaper context isn't the interesting part. The interesting part is which bugs the graph surfaced that grep + vector had missed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 3 bugs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Bug 1: the silent contract change (event handler)
&lt;/h3&gt;

&lt;p&gt;The diff was small. &lt;code&gt;auth.py&lt;/code&gt; added a &lt;code&gt;device_id&lt;/code&gt; field to its login event payload.&lt;/p&gt;

&lt;p&gt;Vector search retrieved &lt;code&gt;login.py&lt;/code&gt;, &lt;code&gt;auth_test.py&lt;/code&gt;, &lt;code&gt;password_validator.py&lt;/code&gt; -- the obvious neighbors. The reviewer approved.&lt;/p&gt;

&lt;p&gt;The graph retrieved one extra file: &lt;code&gt;event_handlers/audit_log.py&lt;/code&gt;. It listens to &lt;code&gt;login_event&lt;/code&gt; and serializes the payload to a fixed schema in S3. Adding a new field broke the schema validator on every login. Production caught fire 90 minutes after merge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why grep missed it&lt;/strong&gt;: &lt;code&gt;audit_log.py&lt;/code&gt; doesn't import &lt;code&gt;auth.py&lt;/code&gt;. It listens to an event bus. There's no string match on "auth" in the file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the graph saw&lt;/strong&gt;: &lt;code&gt;auth.py --emits--&amp;gt; login_event --consumed-by--&amp;gt; audit_log.py&lt;/code&gt;. Three hops, zero string matches, but a clean structural path.&lt;/p&gt;

&lt;p&gt;A code review pass that doesn't follow event subscriptions is a code review pass that doesn't review event-driven systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bug 2: the decorator surprise (transitive change)
&lt;/h3&gt;

&lt;p&gt;A teammate refactored &lt;code&gt;@with_retry&lt;/code&gt; to add a backoff parameter. The default value was the same, so existing callers were "unaffected." Reviewer approved on the strength of the unit tests.&lt;/p&gt;

&lt;p&gt;Vector search retrieved files that explicitly imported the decorator. About a dozen.&lt;/p&gt;

&lt;p&gt;The graph retrieved 31 files. The 19 the graph added were files that &lt;em&gt;applied&lt;/em&gt; &lt;code&gt;@with_retry&lt;/code&gt; to functions that, three calls deep, ended up calling a function whose retry behavior had subtly changed under load.&lt;/p&gt;

&lt;p&gt;One of those callers was a payments webhook handler. Under retry, it now waited an extra 800ms before raising. That 800ms put it past the webhook timeout. We started losing about 0.4% of webhook deliveries silently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why grep missed it&lt;/strong&gt;: a decorator's effect propagates to every callsite of every decorated function. That's structural, not lexical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the graph saw&lt;/strong&gt;: &lt;code&gt;with_retry --decorates--&amp;gt; {19 functions} --called-by--&amp;gt; {31 files}&lt;/code&gt;. The caller of a decorated function inherits the decorator's behavior, even if it never mentions the decorator's name.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bug 3: the orphan test (false confidence)
&lt;/h3&gt;

&lt;p&gt;A migration changed the way one ID was hashed. The PR included a test that asserted the new hash. CI was green.&lt;/p&gt;

&lt;p&gt;The graph showed something the test runner didn't: that test was in a file that &lt;em&gt;no longer ran in CI&lt;/em&gt; because it had been moved out of the &lt;code&gt;tests/&lt;/code&gt; directory three weeks earlier and nobody had updated the path glob in the CI config. The test passed because the test file was never executed. The PR landed with a broken hash that corrupted the migration's first 8,000 rows.&lt;/p&gt;

&lt;p&gt;The graph had &lt;code&gt;dead_code()&lt;/code&gt; for unreachable functions and an inverse query for unreachable test files. I'd never asked it. After this PR, I added "run &lt;code&gt;dead_code()&lt;/code&gt; on test files" to the postflight check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why grep missed it&lt;/strong&gt;: grep doesn't know what CI runs. A test file's existence and a test file's execution are different facts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the graph saw&lt;/strong&gt;: &lt;code&gt;test_user_id.py&lt;/code&gt; had no incoming edge from any CI config and no &lt;code&gt;tested-by&lt;/code&gt; edges from production code. It was a green file in a green repo that didn't actually test anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the graph doesn't replace
&lt;/h2&gt;

&lt;p&gt;I don't want to oversell this. Three things the graph is bad at, and you should keep using vector search or grep for:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Natural-language queries&lt;/strong&gt;. "Find code about authentication" is still better answered by a sentence-embedding model. The graph wants a node name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New code with no edges yet&lt;/strong&gt;. If the function was added in the diff, the graph knows it exists but has no incoming edges. Grep is fine here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-repo retrieval&lt;/strong&gt;. Most code KG implementations are per-repo. If your auth lives in another service, you need a cross-repo strategy (or, increasingly, a cross-repo MCP server -- this is where Sourcegraph and Context.ai are converging in 2026).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The right setup is hybrid: vector search for "find related concepts," graph for "find dependents," grep for "find this exact string." MCP makes that hybrid trivial because the model picks the tool per question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2026 landscape, briefly
&lt;/h2&gt;

&lt;p&gt;Three things have changed since I started doing this in late 2024:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Copilot Chat&lt;/strong&gt; added repo-wide knowledge graph context in March 2026, walking the call graph for &lt;code&gt;@workspace&lt;/code&gt; queries instead of relying purely on file embeddings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sourcegraph&lt;/strong&gt; shipped an MCP server that exposes its long-standing code graph to any MCP-compatible IDE. They had this graph in 2017; the MCP wrapper is what makes it model-accessible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor&lt;/strong&gt; integrated &lt;code&gt;repomap&lt;/code&gt; for project-wide structural context. Not a graph, technically, but the same idea: structural retrieval beats lexical retrieval for cross-file changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern is converging. By the end of 2026, "AI code review" without a structural retrieval layer is going to look the way "AI code review without a vector store" looked in 2023. Quaint.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to wire up if you're starting
&lt;/h2&gt;

&lt;p&gt;If you don't already have something like this, the pragmatic order is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build the graph&lt;/strong&gt; for one repo. &lt;code&gt;code-review-graph&lt;/code&gt; or equivalent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expose it as an MCP server&lt;/strong&gt; so your AI tool of choice can call it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add &lt;code&gt;blast_radius&lt;/code&gt; to the postflight check&lt;/strong&gt; for every PR. Just print the list of 2-hop dependents next to the PR. Even without the AI doing anything, the human reviewer reads better.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep your vector search and grep&lt;/strong&gt;. Don't rip them out. Add the graph alongside.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wait for the bugs to surface&lt;/strong&gt;. They will.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I had the graph for two weeks before the first bug it caught -- the audit log schema break. I don't think I would have shipped that bug without it. I do know I shipped six versions of it across my career before I had this tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;For the last decade, code search has meant "embed the file and find similar embeddings." That's a fine answer to half the question. The other half -- "what does this change &lt;em&gt;break&lt;/em&gt;?" -- is structural, and embeddings don't see it.&lt;/p&gt;

&lt;p&gt;The graph sees it. MCP makes the graph addressable. Together they collapse most cross-file retrieval into a single, cheap query. The bugs that used to live in those gaps don't, anymore.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Want the full rationale and more graph patterns?&lt;/strong&gt; I cover graph schema design, GraphRAG vs vector RAG, and code-as-graph patterns in &lt;a href="https://kenimoto.dev/books/knowledge-graph-practical-guide?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=kg-mcp-vector-misses" rel="noopener noreferrer"&gt;Knowledge Graph Practical Guide: From RAG Limits to Graph-Native AI&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt; -- Anthropic, 2024&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/codelayers/code-review-graph" rel="noopener noreferrer"&gt;code-review-graph&lt;/a&gt; -- Open-source code KG with MCP server&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2404.16130" rel="noopener noreferrer"&gt;GraphRAG: From Local to Global&lt;/a&gt; -- Microsoft Research, 2024&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://sourcegraph.com/blog" rel="noopener noreferrer"&gt;Sourcegraph MCP integration&lt;/a&gt; -- 2026 MCP server announcement&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.github.com/en/copilot" rel="noopener noreferrer"&gt;GitHub Copilot Workspace Context&lt;/a&gt; -- Repo-graph integration, March 2026&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>codereview</category>
      <category>knowledgegraph</category>
    </item>
  </channel>
</rss>
