<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: yureki_lab</title>
    <description>The latest articles on DEV Community by yureki_lab (@yureki_lab).</description>
    <link>https://dev.to/yureki_lab</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3960924%2F46fad6c4-8f78-40a1-a230-6bd3e913f37b.png</url>
      <title>DEV Community: yureki_lab</title>
      <link>https://dev.to/yureki_lab</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yureki_lab"/>
    <language>en</language>
    <item>
      <title>How I Let Claude Code Write My Load Tests and Caught 3 Bottlenecks Pre-Launch</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:32:21 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-let-claude-code-write-my-load-tests-and-caught-3-bottlenecks-pre-launch-5dgj</link>
      <guid>https://dev.to/yureki_lab/how-i-let-claude-code-write-my-load-tests-and-caught-3-bottlenecks-pre-launch-5dgj</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I had never written a real load test before a launch. I let Claude Code build a k6 suite for my Node.js API in an afternoon, and it surfaced three bottlenecks that would have taken the service down on day one. This post walks through the exact prompts, the k6 setup, what broke, and the five lessons I'd apply to any perf-testing work with an AI coding agent. ⚡&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Two weeks before launching a customer-facing API (Node.js 22.x, Fastify, Postgres 16, Redis 7), I realized I had a pretty embarrassing gap: &lt;strong&gt;zero load testing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We had 340 unit tests. We had integration tests. We had a CI pipeline that ran in under ten minutes. What we did &lt;em&gt;not&lt;/em&gt; have was any idea what happened when 200 people hit the checkout endpoint at the same time.&lt;/p&gt;

&lt;p&gt;I knew the theory. Ramp up virtual users, watch p95 latency, look for the cliff. But every time I sat down to write a k6 script, I hit the same wall: the API had 38 endpoints, most of them needed auth tokens and realistic payloads, and I did not want to hand-write 38 scenario files with fake data that looked nothing like production traffic.&lt;/p&gt;

&lt;p&gt;The constraint that made this interesting: I had about &lt;strong&gt;one working day&lt;/strong&gt; to spend on it. Anything longer and it would eat into the launch buffer.&lt;/p&gt;

&lt;p&gt;So I decided to run an experiment. I would give Claude Code (v2.1, running in the terminal) the OpenAPI spec and the route handlers, and let it own the load-testing work end to end. My job would be to review, run, and interpret.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Feed it the shape of the traffic, not just the spec
&lt;/h3&gt;

&lt;p&gt;My first attempt was lazy. I pointed Claude Code at &lt;code&gt;openapi.yaml&lt;/code&gt; and said "write k6 load tests for every endpoint." It produced 38 files that each hammered one endpoint with a fixed payload. Technically correct, practically useless. Real users don't call &lt;code&gt;GET /orders/{id}&lt;/code&gt; in isolation; they log in, browse, add items, and check out.&lt;/p&gt;

&lt;p&gt;So I backed up and gave it context about the actual user journeys. Here is the prompt that worked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Here are the 3 user journeys that matter for launch, in order of business impact:

1. Browse → view product → add to cart → checkout (60% of traffic)
2. Login → view order history → view single order (30%)
3. Admin: list orders with filters, paginated (10%)

Write a k6 test suite that:
- Models these as 3 scenarios with the traffic split above
- Reuses the auth token per virtual user (login once, not per request)
- Generates realistic payloads from the Zod schemas in src/schemas/
- Ramps from 0 to 300 VUs over 5 minutes, holds 10 minutes, ramps down
- Fails if p95 &amp;gt; 500ms or error rate &amp;gt; 1% on any scenario

Don't write one file per endpoint. Organize by journey.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference between "here's the spec" and "here's who uses this and how" was enormous. The second version produced a suite I actually wanted to run.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: The k6 structure it generated
&lt;/h3&gt;

&lt;p&gt;It landed on a layout I have kept since:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;loadtest/
├── config.js          # thresholds, stages, base URL
├── lib/
│   ├── auth.js        # login once per VU, cache token
│   └── factories.js   # payload generators derived from Zod schemas
├── scenarios/
│   ├── shopper.js
│   ├── returning-customer.js
│   └── admin.js
└── main.js            # wires scenarios + traffic split
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The load-bearing part is &lt;code&gt;main.js&lt;/code&gt;. Here is the trimmed version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;shopper&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./scenarios/shopper.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;returningCustomer&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./scenarios/returning-customer.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;admin&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;./scenarios/admin.js&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;scenarios&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;shopper&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ramping-vus&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;shopper&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;startVUs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;stages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;5m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;180&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;10m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;180&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;returning&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ramping-vus&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;returningCustomer&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;startVUs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;stages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;5m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;90&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;10m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;90&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;target&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;admin&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;constant-vus&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;exec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;admin&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;vus&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;17m&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;thresholds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http_req_duration{scenario:shopper}&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;p(95)&amp;lt;500&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http_req_duration{scenario:returning}&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;p(95)&amp;lt;500&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http_req_duration{scenario:admin}&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;p(95)&amp;lt;800&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http_req_failed&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;rate&amp;lt;0.01&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;shopper&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;returningCustomer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;admin&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things I would not have thought to do on my own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per-scenario thresholds&lt;/strong&gt; using tags, so a slow admin endpoint doesn't mask a fast checkout path (or vice versa)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Splitting the 300 VUs across scenarios&lt;/strong&gt; to match the 60/30/10 traffic mix, instead of one big pool that picks randomly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The factories file was the other clever bit. Instead of hardcoding fake data, it imported the same Zod schemas the API uses for validation and generated payloads that would always pass. When I later added a required field to the checkout schema, the load test picked it up with zero changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Run it against staging and watch things break
&lt;/h3&gt;

&lt;p&gt;I ran it with k6 v1.1 against a staging environment sized identically to production (2 API replicas, 1 Postgres instance, 1 Redis).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;k6 run loadtest/main.js &lt;span class="nt"&gt;--out&lt;/span&gt; &lt;span class="nv"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;results.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First run: &lt;strong&gt;failed at 140 VUs.&lt;/strong&gt; Not 300. Not even half.&lt;/p&gt;

&lt;p&gt;This is where the workflow got interesting. I did not want to guess at the cause, so I pasted the k6 summary output and the API logs from the same window into Claude Code and asked it to correlate. Here is what it found, in the order it found them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bottleneck 1: The database pool was sized for a laptop
&lt;/h3&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;graph LR
  A[300 VUs] --&amp;gt; B[2 API replicas]
  B --&amp;gt; C[Pool: 10 conns each]
  C --&amp;gt; D[(Postgres: max 100)]
  style C fill:#f96,stroke:#333&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The Postgres client pool was set to its default of &lt;strong&gt;10 connections per replica&lt;/strong&gt;. With two replicas, that is 20 concurrent queries max. Once more than about 20 requests needed the DB at the same time, everything else queued. The p95 went from 80ms at 100 VUs to 2,400ms at 140 VUs. Classic cliff.&lt;/p&gt;

&lt;p&gt;Claude Code spotted it from the log pattern: &lt;code&gt;timeout exceeded when trying to connect&lt;/code&gt; spiking exactly when latency did. Fix was a one-liner in the pool config, bumped to 40 per replica, still comfortably under the Postgres limit. That alone got us to 220 VUs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bottleneck 2: An N+1 hiding on one path
&lt;/h3&gt;

&lt;p&gt;At 220 VUs the returning-customer scenario started failing thresholds while shopper stayed fine. That asymmetry was the clue.&lt;/p&gt;

&lt;p&gt;The order history endpoint used a batching layer to load line items. The single-order endpoint did not. It looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Before: one query per line item 😬&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;itemIds&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;product&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;products&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;productId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One order with 12 line items meant 13 queries. Under load, that path alone was generating more DB traffic than the entire shopper journey. Claude Code proposed the fix and wrote it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// After: one query, keyed lookup&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findById&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;products&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;products&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findByIds&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;itemIds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;productId&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;byId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;products&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;]));&lt;/span&gt;
&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;itemIds&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;product&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;byId&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;productId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;}));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had reviewed that file before. I had &lt;em&gt;tests&lt;/em&gt; for that file. None of that caught it, because a 13-query endpoint is fast when one person calls it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bottleneck 3: Logging was eating CPU
&lt;/h3&gt;

&lt;p&gt;The last one was the sneakiest. At 280 VUs, p95 crept up again, but the DB was fine and Redis was fine. CPU on the API replicas was pinned at 100%.&lt;/p&gt;

&lt;p&gt;Claude Code asked me to run a 30-second CPU profile during load. The top frame was &lt;code&gt;JSON.stringify&lt;/code&gt; inside the request logger. Someone (me, six months ago) had added a log line that serialized the &lt;strong&gt;full request body at info level&lt;/strong&gt; for "debugging." On a checkout request with a 40-item cart, that was serializing several kilobytes per request, 200 times a second.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Before&lt;/span&gt;
&lt;span class="nx"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;incoming request&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="c1"&gt;// After&lt;/span&gt;
&lt;span class="nx"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;bodyBytes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;content-length&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;incoming request&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Removing it dropped CPU to 60% at the same load and bought us the last 20 VUs of headroom.&lt;/p&gt;

&lt;h3&gt;
  
  
  The final run
&lt;/h3&gt;

&lt;p&gt;After all three fixes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Max VUs before threshold failure&lt;/td&gt;
&lt;td&gt;140&lt;/td&gt;
&lt;td&gt;300+ (target hit)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p95 latency at 300 VUs&lt;/td&gt;
&lt;td&gt;n/a (failed)&lt;/td&gt;
&lt;td&gt;310ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error rate at 300 VUs&lt;/td&gt;
&lt;td&gt;11%&lt;/td&gt;
&lt;td&gt;0.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wall-clock time spent&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;~6 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Six hours, including the runs themselves. Three bugs that would have surfaced on launch day in front of real customers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The agent is only as good as your traffic model
&lt;/h3&gt;

&lt;p&gt;"Write load tests for this API" gets you noise. "Here are the three journeys and their traffic split" gets you a real test. The single most important input was not the spec; it was me writing down who uses the system and in what order. If you can't describe your traffic in three bullet points, do that before touching an agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Derive payloads from your validation schemas, not from fixtures
&lt;/h3&gt;

&lt;p&gt;Letting the factories import the Zod schemas meant the load test could never drift out of sync with the API. This is one of those things an agent will suggest if you ask "how do we keep this from rotting," and it is worth asking every time.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Per-scenario thresholds beat global ones
&lt;/h3&gt;

&lt;p&gt;If I had used a single global p95 threshold, the N+1 on the order-detail path would have been averaged away by the fast shopper traffic. Tag your scenarios and set thresholds per tag. This was the difference between finding bug 2 and shipping it.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Let the agent correlate logs and metrics, but you run the profiler
&lt;/h3&gt;

&lt;p&gt;Claude Code was excellent at reading a k6 summary next to a log excerpt and saying "these two spikes line up, here is the likely cause." It could not run the CPU profile for me on staging, and I would not want it to. The division of labor that worked: the agent does the pattern matching and proposes the fix, I execute anything that touches a live environment.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Load tests find the bugs your unit tests are structurally blind to
&lt;/h3&gt;

&lt;p&gt;All three bottlenecks were invisible at concurrency 1. Pool sizing, N+1s, and hot-path serialization cost only show up under load. If your test pyramid has no load layer, you have a category of bugs you are guaranteed to find in production. An agent makes that layer cheap enough that "no time" stops being an excuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Wire it into CI on a nightly schedule&lt;/strong&gt; against staging, with results posted to a channel. A daily p95 trend line catches regressions before anyone notices.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add a soak test&lt;/strong&gt; (2 hours at 60% of peak) to hunt memory leaks. The 17-minute run is too short to catch slow growth.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Teach the agent to read the flame graph.&lt;/strong&gt; Right now I run the profiler and paste the top frames. I want to try feeding it the raw profile output and see if it can find the hot path without my summary.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If you have a launch coming and no load tests, block one afternoon, write your three traffic journeys on paper, and hand them to Claude Code with your schemas. You will probably find something. I found three somethings.&lt;/p&gt;

&lt;p&gt;If this was useful, follow me here on Dev.to for more build logs like this one. I write about running AI coding agents on real production work, including the parts that go wrong. And if you have found a bottleneck under load that no unit test could have caught, I would love to hear about it in the comments. 🚀&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>performance</category>
      <category>testing</category>
      <category>node</category>
    </item>
    <item>
      <title>How I Built a Regression Suite for My AI Coding Agent's Prompts: 5 Lessons</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:32:35 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-built-a-regression-suite-for-my-ai-coding-agents-prompts-5-lessons-4p0g</link>
      <guid>https://dev.to/yureki_lab/how-i-built-a-regression-suite-for-my-ai-coding-agents-prompts-5-lessons-4p0g</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I changed one line in my coding agent's spec file and, two weeks later, it quietly started skipping tests on every PR. Nobody noticed because nothing "broke." So I built a regression suite for the prompts themselves: frozen repo snapshots, canned tasks, deterministic graders, and an LLM judge, all gated in CI. Here's how it works and the 5 lessons that cost me the most to learn.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;I run a fully autonomous implementation system built on Claude Code (v2.x at the time of writing). It picks up tasks, implements them, runs tests, opens PRs, and fixes its own CI failures. It has been doing this 24/7 for months across several repos.&lt;/p&gt;

&lt;p&gt;The system's behavior is controlled by prompts: a project spec file (&lt;code&gt;CLAUDE.md&lt;/code&gt;), a handful of skill definitions, and an orchestrator prompt that decides what to work on next. These files are &lt;strong&gt;code&lt;/strong&gt;. They change the system's behavior just as much as a Python function does.&lt;/p&gt;

&lt;p&gt;But I was treating them like config. I'd tweak a sentence, eyeball one run, and merge. Zero tests.&lt;/p&gt;

&lt;p&gt;Here is the incident that changed my mind. I added this line to the spec file to reduce noise in PR descriptions:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Keep PR descriptions short. Do not list every command you ran.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Reasonable, right? Two weeks later I noticed the agent's PRs had stopped including test output. Then I noticed it had stopped &lt;em&gt;running&lt;/em&gt; the full test suite for small changes. The agent had generalized "do not list every command" into "running fewer commands is preferred." That's a stretch, but it was consistent, and it took 14 days and one broken deploy to catch.&lt;/p&gt;

&lt;p&gt;The core problem: &lt;strong&gt;prompt changes have non-local effects, and I had no way to measure them before merging.&lt;/strong&gt; ⚠️&lt;/p&gt;

&lt;p&gt;I needed the equivalent of a test suite for prompts.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;p&gt;The design borrows shamelessly from normal software testing. A regression suite needs three things: fixtures, a runner, and assertions.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Prompt change PR] --&amp;gt; B[CI: eval job]
    B --&amp;gt; C[Restore fixture repos]
    C --&amp;gt; D[Run agent headless per task]
    D --&amp;gt; E[Deterministic graders]
    D --&amp;gt; F[LLM judge]
    E --&amp;gt; G[Score report]
    F --&amp;gt; G
    G --&amp;gt; H{Score &amp;gt;= baseline?}
    H --&amp;gt;|yes| I[Merge allowed]
    H --&amp;gt;|no| J[Block + diff report]&lt;/code&gt;&lt;/pre&gt;



&lt;h3&gt;
  
  
  1. Fixtures: frozen repos plus canned tasks
&lt;/h3&gt;

&lt;p&gt;Every time the agent does something interesting in production, good or bad, I snapshot the state into a fixture. A fixture is a directory with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A small git repo (usually 20 to 200 files, trimmed from the real one)&lt;/li&gt;
&lt;li&gt;A task prompt, exactly as the orchestrator would phrase it&lt;/li&gt;
&lt;li&gt;An expectations file&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The expectations file looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# fixtures/skips-tests-on-small-change/expect.yaml&lt;/span&gt;
&lt;span class="na"&gt;task&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Fix&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;off-by-one&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pagination&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;src/api/list.py"&lt;/span&gt;
&lt;span class="na"&gt;must&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;ran_command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pytest"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;touched_files_only&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;src/api/list.py"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tests/test_list.py"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;added_or_modified_test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;pr_description_mentions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pytest"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;passed"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="na"&gt;must_not&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;touched_files&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;src/api/auth.py"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;ran_command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;push&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;--force"&lt;/span&gt;
&lt;span class="na"&gt;judge&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;PR&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;explains&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;WHY&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;bug&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;happened,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;just&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;what&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;changed."&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fix&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;does&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;introduce&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;edge&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;case&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;empty&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;lists."&lt;/span&gt;
&lt;span class="na"&gt;max_turns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;40&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I have about 60 of these now. Roughly a third came straight from incidents. The rest are "golden path" tasks I want to keep working.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Runner: headless, sandboxed, reproducible
&lt;/h3&gt;

&lt;p&gt;Each fixture runs in a throwaway container with the fixture repo restored from a tarball. The agent runs in headless mode with the &lt;em&gt;candidate&lt;/em&gt; spec file mounted in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail
&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;candidate_spec&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$2&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="nv"&gt;workdir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;tar&lt;/span&gt; &lt;span class="nt"&gt;-xzf&lt;/span&gt; &lt;span class="s2"&gt;"fixtures/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/repo.tar.gz"&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$workdir&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;cp&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$candidate_spec&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$workdir&lt;/span&gt;&lt;span class="s2"&gt;/CLAUDE.md"&lt;/span&gt;

&lt;span class="nv"&gt;task&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;yq &lt;span class="s1"&gt;'.task'&lt;/span&gt; &lt;span class="s2"&gt;"fixtures/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/expect.yaml"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;max_turns&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;yq &lt;span class="s1"&gt;'.max_turns // 40'&lt;/span&gt; &lt;span class="s2"&gt;"fixtures/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/expect.yaml"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# Headless run; everything the agent does is captured as JSONL&lt;/span&gt;
&lt;span class="o"&gt;(&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$workdir&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$task&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--output-format&lt;/span&gt; stream-json &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--max-turns&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$max_turns&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--permission-mode&lt;/span&gt; acceptEdits &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"runs/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.jsonl"&lt;/span&gt;

&lt;span class="c"&gt;# Also capture the resulting diff and the git log&lt;/span&gt;
&lt;span class="o"&gt;(&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$workdir&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git diff HEAD &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"runs/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.diff"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git log &lt;span class="nt"&gt;--oneline&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 20 &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"runs/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;fixture&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.log"&lt;/span&gt; &lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part: the run produces &lt;strong&gt;artifacts&lt;/strong&gt;, not just a pass/fail. The JSONL stream has every tool call. The diff has every file change. The graders work on those artifacts, never on the live agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Graders: deterministic first, LLM second
&lt;/h3&gt;

&lt;p&gt;Most of my assertions don't need an LLM at all. Did it run &lt;code&gt;pytest&lt;/code&gt;? Grep the tool calls. Did it touch a forbidden file? Parse the diff. These graders are fast, free, and never flaky.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# graders/deterministic.py
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;jsonl_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;jsonl_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;splitlines&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;ev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ev&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_use&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;block&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{})&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ran_command&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;calls&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;rx&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;any&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;rx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;inp&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;calls&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;touched_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;diff_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;diff_path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;splitlines&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;+++ b/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;files&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;:])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;files&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;grade&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fixture&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;calls&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fixture&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.jsonl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
    &lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;touched_files&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runs/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fixture&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.diff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;must&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ran_command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;ran_command&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;calls&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ran_command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MUST ran_command &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ran_command&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt; not found&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;touched_files_only&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;touched_files_only&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;touched unexpected files: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;touched_files_only&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;must_not&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ran_command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;ran_command&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;calls&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ran_command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MUST NOT ran_command &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ran_command&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt; was executed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;touched_files&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rule&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;touched_files&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;touched forbidden files: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;files&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rule&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;touched_files&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;failures&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then, for the fuzzy stuff ("does the PR description explain &lt;em&gt;why&lt;/em&gt;?"), an LLM judge reads the diff plus the PR body and answers each &lt;code&gt;judge:&lt;/code&gt; question with a yes/no and a one-sentence reason. I use a smaller, cheaper model for judging than the one doing the work. The judge gets the rubric, the artifacts, and nothing else. It never sees the candidate spec file, so it can't be biased by it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# graders/judge.py (trimmed)
&lt;/span&gt;&lt;span class="n"&gt;RUBRIC&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are grading an AI coding agent&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s output.
Answer each question with exactly YES or NO on its own line,
followed by one sentence of justification.
Questions:
{questions}

--- DIFF ---
{diff}
--- PR DESCRIPTION ---
{pr_body}
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Scoring and the CI gate
&lt;/h3&gt;

&lt;p&gt;Each fixture yields a score: deterministic failures are hard fails, judge questions are weighted 1 point each. The suite score is the mean across fixtures. CI compares it against the baseline from &lt;code&gt;main&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A prompt PR is blocked if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Any fixture that passed on &lt;code&gt;main&lt;/code&gt; now hard-fails (regression)&lt;/li&gt;
&lt;li&gt;The suite score drops more than 3 points&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The CI job posts a diff report as a PR comment: which fixtures flipped, and the judge's one-sentence reasons for any changed answer. That comment is honestly the most useful artifact in the whole system. It turns "I think this wording is better" into "this wording made fixture #23 stop writing tests, here's the evidence." 💡&lt;/p&gt;

&lt;p&gt;Running the full suite costs me about $4 and 25 minutes. I run a 15-fixture smoke subset on every push and the full 60 on merge to &lt;code&gt;main&lt;/code&gt; and nightly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Fixtures from incidents are worth 10x fixtures you invent
&lt;/h3&gt;

&lt;p&gt;Every time the agent does something dumb in production, I freeze it into a fixture &lt;em&gt;the same day&lt;/em&gt;. Those fixtures catch real regressions. The "golden path" fixtures I wrote from imagination almost never fail. If you only do one thing from this post, do this.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Assert on behavior, not on output text
&lt;/h3&gt;

&lt;p&gt;My first version asserted on strings in the final diff. It was flaky within a week, because a correct fix can be written ten different ways. Asserting on &lt;strong&gt;what the agent did&lt;/strong&gt; (which tools it called, which files it touched, whether it ran tests) is stable across runs. Save the fuzzy judgments for the LLM judge.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Non-determinism is a feature you have to budget for
&lt;/h3&gt;

&lt;p&gt;The same fixture with the same prompt will not produce identical runs. I run each fixture 3 times and take the median score. That tripled the cost and cut false alarms to nearly zero. A single-run eval suite will train you to ignore it, which is worse than having none.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. The judge needs to be dumber than the worker
&lt;/h3&gt;

&lt;p&gt;When I used the same large model as both worker and judge, the judge was too forgiving. It would rationalize the worker's choices. Switching to a smaller model with a tight rubric made the judge stricter and 5x cheaper. Counterintuitive, but it held up across months of runs.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Prompt diffs need code review, and now they get it
&lt;/h3&gt;

&lt;p&gt;Once the suite existed, prompt changes started going through actual review. Reviewers ask "which fixture covers this?" the same way they'd ask "where's the test?" That cultural shift mattered more than any single caught regression. The spec file went from a scratchpad to a maintained artifact with a changelog.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things I'm working on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mutation testing for prompts.&lt;/strong&gt; Automatically delete each rule from the spec file one at a time and check that at least one fixture fails. Rules that nothing depends on are dead weight (and there are more of them than I'd like to admit).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixture aging.&lt;/strong&gt; Fixtures capture a moment in the codebase. Some are now testing behavior on code that no longer looks like production. I want an automated "is this fixture still representative?" check.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'll write both up once they've survived a month of real use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If your agent's behavior lives in prompt files, those files are code, and code without tests will drift. You don't need anything fancy to start. One fixture from your last incident, one grep-based grader, and a CI job that runs it. Build from there.&lt;/p&gt;

&lt;p&gt;If this was useful, &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; 🚀. I write about running autonomous coding agents in production, including the parts that break. And if you've built something similar, drop a comment. I'd love to compare notes on graders. ✅&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>testing</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>How I Made My Autonomous Coding Agent Survive Context Resets: 5 Lessons</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Fri, 25 Sep 2026 14:32:49 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-made-my-autonomous-coding-agent-survive-context-resets-5-lessons-1ngk</link>
      <guid>https://dev.to/yureki_lab/how-i-made-my-autonomous-coding-agent-survive-context-resets-5-lessons-1ngk</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;My fully autonomous implementation system runs for days at a time, and every few hours the agent's context window fills up and gets wiped. Early on, each reset meant the agent forgot what it was working on, redid finished tasks, or quietly reversed decisions it had made an hour earlier. I fixed it with a &lt;strong&gt;three-file state handoff protocol&lt;/strong&gt; (a state memo, a task file, and a decisions log), strict precedence rules, and a pre-commit hook that refuses to let the agent commit without checkpointing. Here's the design, the code, and the five lessons from six months of running it. ✅&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;I run a fully autonomous implementation system on a Mac mini in my closet. It's built on Claude Code (the 2.x CLI line as of September 2026) and it works through a task backlog across several repos without me in the loop. An orchestrator module hands work to parallel implementation agents, a self-healing agent cleans up after them, and a remote control dashboard lets me peek in from my phone.&lt;/p&gt;

&lt;p&gt;The system is great at &lt;em&gt;doing&lt;/em&gt; work. It was terrible at &lt;em&gt;remembering&lt;/em&gt; it.&lt;/p&gt;

&lt;p&gt;Here's why. A single agent session has a finite context window. On a long task, the window fills up with tool output, diffs, and test logs. When it's full, the harness compacts the conversation into a summary and continues. That summary is lossy. It keeps &lt;em&gt;what happened&lt;/em&gt; reasonably well, but it drops &lt;em&gt;why&lt;/em&gt; and &lt;em&gt;what's next&lt;/em&gt; with alarming frequency.&lt;/p&gt;

&lt;p&gt;Across an overnight run, that happens four or five times. The failure modes I saw in the first month:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Amnesia loops.&lt;/strong&gt; One night the agent investigated the same failing database migration test three separate times in 11 hours. Each time it reached the same conclusion (a timezone mismatch in a fixture), each time it lost the conclusion in a reset, each time it started over. That night burned roughly $40 in tokens and shipped nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decision reversal.&lt;/strong&gt; The agent evaluated two HTTP client libraries, picked one for a good reason, then after a reset picked the &lt;em&gt;other&lt;/em&gt; one because it only saw "chose a client library" in the summary and re-ran the evaluation with slightly different weights. Two PRs in the same week used different clients.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ghost tasks.&lt;/strong&gt; A task that had been finished and merged was still described as "in progress" in the summary. The agent reopened it, rewrote the same feature on a new branch, and hit merge conflicts with its own earlier work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;My first instinct was "just use git history." Git is a great record of &lt;em&gt;what changed&lt;/em&gt;. It says nothing about what was tried and rejected, what's half-finished, or what the agent should pick up next. I needed something the agent writes for its future self.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;p&gt;The fix is boring, which is why I trust it: a small set of plain Markdown files that the agent must read on boot and must update at specific moments. No vector database, no fancy memory service. Just files in the repo, with rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  The three files
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;project-root/
├── CLAUDE.md          # project spec file: goals, rules, read order (rarely changes)
├── state/
│   ├── current.md     # state memo: where we are, what's next (overwritten)
│   ├── backlog.md     # task file: checklist with status (edited in place)
│   └── decisions.md   # decisions log: append-only, numbered, with rationale
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each file has a different job and, crucially, a different &lt;strong&gt;write mode&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Contains&lt;/th&gt;
&lt;th&gt;Write mode&lt;/th&gt;
&lt;th&gt;Size cap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;State memo&lt;/td&gt;
&lt;td&gt;Current position, next action, last 3–5 decisions&lt;/td&gt;
&lt;td&gt;Overwrite&lt;/td&gt;
&lt;td&gt;~150 lines&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task file&lt;/td&gt;
&lt;td&gt;Every task with &lt;code&gt;pending&lt;/code&gt; / &lt;code&gt;doing&lt;/code&gt; / &lt;code&gt;done&lt;/code&gt; / &lt;code&gt;blocked&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Edit in place&lt;/td&gt;
&lt;td&gt;One line per task&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decisions log&lt;/td&gt;
&lt;td&gt;Numbered entries: context, options, choice, why&lt;/td&gt;
&lt;td&gt;Append only&lt;/td&gt;
&lt;td&gt;Unbounded&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The state memo is the one the agent reads first after a reset. It's deliberately short. If it grows past about 150 lines, the agent is required to compress it, moving old decisions into the log and dropping resolved context.&lt;/p&gt;

&lt;p&gt;Here's the actual template the state memo follows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# State memo (single source of truth)&lt;/span&gt;

&lt;span class="gu"&gt;## Where we are (as of 2026-09-24 03:12 JST)&lt;/span&gt;
Working on: task #47 — rate limiter for the public API
Branch: feat/rate-limiter
Status: implementation done, 2 of 9 integration tests failing

&lt;span class="gu"&gt;## Next action&lt;/span&gt;
Fix the two failing tests in the burst-window case. They fail because the
fake clock isn't advanced between requests. Do NOT touch the limiter logic.

&lt;span class="gu"&gt;## Recent decisions (full detail in state/decisions.md)&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; D-031: token bucket over sliding window (memory footprint at 10k clients)
&lt;span class="p"&gt;-&lt;/span&gt; D-030: limits live in config, not env vars (need per-tenant overrides)

&lt;span class="gu"&gt;## Do not redo&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Already evaluated leaky bucket. Rejected. See D-031.
&lt;span class="p"&gt;-&lt;/span&gt; Migration 0042 is applied in staging. Don't regenerate it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That "Do not redo" section is the single highest-value thing in the whole system. It exists purely because of the amnesia loops.&lt;/p&gt;

&lt;h3&gt;
  
  
  Read order and precedence
&lt;/h3&gt;

&lt;p&gt;Files alone don't help if the agent reads them in a random order or, worse, trusts a stale one over a fresh one. So the project spec file has a hard rule that runs at every boot, including every post-reset boot:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## On every session start (mandatory, in this order)&lt;/span&gt;
&lt;span class="p"&gt;1.&lt;/span&gt; Read CLAUDE.md (this file)
&lt;span class="p"&gt;2.&lt;/span&gt; Read state/current.md
&lt;span class="p"&gt;3.&lt;/span&gt; Read state/backlog.md
&lt;span class="p"&gt;4.&lt;/span&gt; Skim the last 10 entries of state/decisions.md

Do not start work before completing all four.

&lt;span class="gu"&gt;## When files contradict each other&lt;/span&gt;
state/current.md  &amp;gt;  state/backlog.md  &amp;gt;  CLAUDE.md  &amp;gt;  state/decisions.md

The state memo is always the latest truth. If the decisions log says X
and the memo says Y, Y wins. Update the older file, do not argue with it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That precedence line killed the decision-reversal problem almost overnight. Before it, the agent would find an old decision entry, a newer memo note, and "resolve" the conflict by re-deriving the answer from scratch. Now it has a tiebreaker and moves on.&lt;/p&gt;

&lt;p&gt;Here's the loop as a diagram:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Session boot] --&amp;gt; B[Read spec → memo → tasks → decisions]
    B --&amp;gt; C[Pick next action from memo]
    C --&amp;gt; D[Do work]
    D --&amp;gt; E{Unit of work done?}
    E -- no --&amp;gt; D
    E -- yes --&amp;gt; F[Checkpoint: update memo + tasks, append decision if any]
    F --&amp;gt; G[Commit]
    G --&amp;gt; H{Context full?}
    H -- no --&amp;gt; C
    H -- yes --&amp;gt; I[Compaction / reset]
    I --&amp;gt; A&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The key insight in that diagram: &lt;strong&gt;the checkpoint happens before the commit, and the commit happens before the reset can hurt you.&lt;/strong&gt; If the reset lands mid-task, the worst case is losing the in-flight edits since the last commit, and the memo already says what the agent was about to do.&lt;/p&gt;

&lt;h3&gt;
  
  
  The checkpoint rule (and the hook that enforces it)
&lt;/h3&gt;

&lt;p&gt;Telling an agent "update the state files regularly" doesn't work. "Regularly" gets interpreted as "when I remember," and the agent remembers less as the context fills. So I made the rule mechanical and enforced it with a hook.&lt;/p&gt;

&lt;p&gt;The rule is &lt;strong&gt;write on completion, not on a timer&lt;/strong&gt;. Specifically, the agent must update the state memo:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;After finishing a unit of work (a task reaches &lt;code&gt;done&lt;/code&gt; or &lt;code&gt;blocked&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Before any operation that's hard to undo (migration, force push, deleting a branch)&lt;/li&gt;
&lt;li&gt;Immediately after making any decision that has a rationale&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And the enforcement is a plain Git pre-commit hook:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="c"&gt;# .git/hooks/pre-commit&lt;/span&gt;
&lt;span class="c"&gt;# Refuse to commit if the state memo hasn't been updated recently.&lt;/span&gt;
&lt;span class="c"&gt;# The agent runs for hours; a memo older than the last 45 minutes of&lt;/span&gt;
&lt;span class="c"&gt;# commits means it's drifting and the next reset will hurt.&lt;/span&gt;

&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="nv"&gt;MEMO&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"state/current.md"&lt;/span&gt;
&lt;span class="nv"&gt;MAX_AGE_MIN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;45

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MEMO&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"❌ &lt;/span&gt;&lt;span class="nv"&gt;$MEMO&lt;/span&gt;&lt;span class="s2"&gt; is missing. Create it before committing."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Skip the check if the memo itself is part of this commit.&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;git diff &lt;span class="nt"&gt;--cached&lt;/span&gt; &lt;span class="nt"&gt;--name-only&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-qx&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MEMO&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nv"&gt;memo_age_min&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;stat&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; %m &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MEMO&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="m"&gt;60&lt;/span&gt; &lt;span class="k"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;((&lt;/span&gt; memo_age_min &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; MAX_AGE_MIN &lt;span class="o"&gt;))&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"❌ &lt;/span&gt;&lt;span class="nv"&gt;$MEMO&lt;/span&gt;&lt;span class="s2"&gt; was last updated &lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;memo_age_min&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; min ago (limit: &lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;MAX_AGE_MIN&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;)."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"   Update 'Where we are' and 'Next action', then commit again."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(That &lt;code&gt;stat -f %m&lt;/code&gt; is macOS. On Linux it's &lt;code&gt;stat -c %Y&lt;/code&gt;.)&lt;/p&gt;

&lt;p&gt;When the hook fires, the agent sees the error in its tool output, updates the memo, and retries the commit. It took zero prompt engineering to get this behavior. The agent already knows how to read a failed command and fix the cause. I just had to make the cause explicit.&lt;/p&gt;

&lt;h3&gt;
  
  
  What NOT to persist
&lt;/h3&gt;

&lt;p&gt;The first version of the memo got polluted fast. The agent dumped test output into it, pasted whole stack traces, and once wrote a 900-line "current understanding of the codebase" section. A 900-line memo is as useless as no memo, because the agent skims it and misses the two lines that matter.&lt;/p&gt;

&lt;p&gt;So the spec file now has an explicit blocklist:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Never write to state/current.md&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Raw tool output (test logs, stack traces, diffs). Link to the file instead.
&lt;span class="p"&gt;-&lt;/span&gt; Speculation ("might be caused by..."). Only write what you verified.
&lt;span class="p"&gt;-&lt;/span&gt; Anything you can derive from git (what changed, who changed it).
&lt;span class="p"&gt;-&lt;/span&gt; Secrets, tokens, credentials, or absolute paths on this machine.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line matters more than it looks. The state files get committed. If the agent copies an API key into the memo "for reference," it's in history forever.&lt;/p&gt;

&lt;h3&gt;
  
  
  The resume test
&lt;/h3&gt;

&lt;p&gt;You can't improve what you don't measure, so I built a small test for the handoff. Once a day, a script kills the running session mid-task, boots a fresh one, and asks a single question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Without doing any work: what task are you on, what's the next concrete action, and what should you not redo?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I diff the answer against the memo and against what actually happened in git. It's scored by a second, cheap model run on three yes/no questions: correct task, correct next action, no contradiction with the log.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Month 1 (summary only)&lt;/th&gt;
&lt;th&gt;Month 6 (three-file protocol)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Resume accuracy&lt;/td&gt;
&lt;td&gt;~40%&lt;/td&gt;
&lt;td&gt;~95%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amnesia loops per week&lt;/td&gt;
&lt;td&gt;4–6&lt;/td&gt;
&lt;td&gt;0–1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wasted overnight token spend&lt;/td&gt;
&lt;td&gt;~$120/week&lt;/td&gt;
&lt;td&gt;&amp;lt;$15/week&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those numbers are from my own logs on my own system, so treat them as one data point, not a benchmark. But the direction is not subtle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Compaction is not memory.&lt;/strong&gt; Every harness will eventually summarize your context, and every summary drops the "why." If you're running agents for longer than a single sitting, you need an explicit, agent-written handoff. Don't wait for the amnesia loop to teach you this at $40 a night.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;One file has to be the boss.&lt;/strong&gt; The single biggest win wasn't the files themselves, it was the precedence rule. When two records disagree, the agent needs a tiebreaker it doesn't have to think about. "Memo wins, update the other one" removed an entire class of re-litigated decisions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Write on completion, never on a timer.&lt;/strong&gt; "Every 30 minutes" produces memos written mid-thought that describe a state that no longer exists. "After every finished unit of work" produces memos that describe a real, committed checkpoint. And enforce it with a hook, because the agent's memory of the rule degrades exactly when the context is fullest.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Small enough to read in one gulp.&lt;/strong&gt; The memo has a hard size cap and the agent is responsible for enforcing it. The moment it becomes a dumping ground, the agent skims it and you're back to square one. A "Do not redo" section of five lines beats a "Full context" section of five hundred.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Match the write mode to the data.&lt;/strong&gt; State gets overwritten. Decisions get appended. Tasks get edited in place. When I had everything in one file with one write mode, the agent would overwrite decisions or append state, and both were wrong. Three files, three write modes, zero ambiguity.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Three things on my list:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Structured front matter.&lt;/strong&gt; The memo is prose right now. I'm adding a YAML block at the top (current task ID, branch, status enum) so the remote control dashboard can render it without parsing Markdown.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A staleness detector in the observability layer.&lt;/strong&gt; The pre-commit hook catches drift at commit time. I want a background check that flags a memo whose "Where we are" hasn't changed in two hours while commits keep landing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-repo handoff.&lt;/strong&gt; When the orchestrator module moves an agent from one repo to another, the memo in repo A doesn't help in repo B. I'm experimenting with a lightweight "last seen elsewhere" pointer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If you take one thing from this: the moment your agent runs longer than a context window, its memory is your problem, not the model's. Give it a place to write for its future self, tell it exactly when to write there, and make the rule mechanical.&lt;/p&gt;

&lt;p&gt;I'm writing up the fully autonomous implementation system one piece at a time here on Dev.to: the orchestrator, the parallel agents, the self-healing loop, the remote dashboard. &lt;strong&gt;Follow me&lt;/strong&gt; to catch the next one. 🚀&lt;/p&gt;

&lt;p&gt;And I'd genuinely like to know: how do &lt;em&gt;you&lt;/em&gt; handle context resets on long-running agents? Files, a database, a memory MCP server, something else? Drop it in the comments. 💬&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>softwareengineering</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How I Added OpenTelemetry Tracing to 47 Services With Claude Code in 9 Days</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Thu, 24 Sep 2026 14:32:24 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-added-opentelemetry-tracing-to-47-services-with-claude-code-in-9-days-36ea</link>
      <guid>https://dev.to/yureki_lab/how-i-added-opentelemetry-tracing-to-47-services-with-claude-code-in-9-days-36ea</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I had 47 backend services that emitted logs but no traces, so every incident turned into a cross-repo grep party. I used Claude Code to instrument all of them with OpenTelemetry in 9 days — not by asking it to "add tracing," but by writing one reference implementation by hand and turning our conventions into an executable validator the agent had to pass. Here's the loop, the code, and the five things I'd do differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Our backend is 47 services: mostly Node.js 22.x with Express or Fastify, plus a handful of Python 3.13 FastAPI services. Every one of them had structured JSON logs. None of them had distributed tracing.&lt;/p&gt;

&lt;p&gt;That's a fine setup right up until a request crosses four service boundaries and gets slow somewhere. Then the debugging session looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Find the request ID in the edge service logs.&lt;/li&gt;
&lt;li&gt;Hope the next service logged the same request ID.&lt;/li&gt;
&lt;li&gt;Discover it called the header &lt;code&gt;x-request-id&lt;/code&gt; instead of &lt;code&gt;x-correlation-id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Give up and add a &lt;code&gt;console.log&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Our median incident spent 40+ minutes just on "which service is actually slow." I timeboxed a manual fix and estimated it honestly: 47 services × roughly half a day each for setup, span conventions, config, and a smoke test, comes out around six weeks of focused work. Nobody was going to fund six weeks of plumbing.&lt;/p&gt;

&lt;p&gt;The constraint that made this interesting: &lt;strong&gt;tracing instrumentation is repetitive but not identical.&lt;/strong&gt; Auto-instrumentation gets you 70% of the way, and then every service has its own bespoke 30% — a custom HTTP client, a queue consumer, a cron entrypoint that isn't a request at all. That mix is exactly where an AI coding agent is strong and exactly where it will quietly invent conventions if you let it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Write the reference implementation by hand
&lt;/h3&gt;

&lt;p&gt;My first attempt was a 2,000-word &lt;code&gt;CONVENTIONS.md&lt;/code&gt; describing how spans should be named, which attributes were required, and how to wire the SDK. The agent followed maybe 70% of it and improvised the rest — different attribute names in half the services, &lt;code&gt;startSpan&lt;/code&gt; in some places and &lt;code&gt;startActiveSpan&lt;/code&gt; in others.&lt;/p&gt;

&lt;p&gt;So I threw the prose away and instrumented &lt;strong&gt;one&lt;/strong&gt; service by hand, carefully, and made that service the spec:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// tracing.js — the reference every other service was told to mirror.&lt;/span&gt;
&lt;span class="c1"&gt;// Imported for side effects before anything else in the process.&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;NodeSDK&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/sdk-node&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;OTLPTraceExporter&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/exporter-trace-otlp-http&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;getNodeAutoInstrumentations&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/auto-instrumentations-node&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;resourceFromAttributes&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/resources&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;ATTR_SERVICE_NAME&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;ATTR_SERVICE_VERSION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/semantic-conventions&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sdk&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;NodeSDK&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;resource&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;resourceFromAttributes&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ATTR_SERVICE_NAME&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;SERVICE_NAME&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;ATTR_SERVICE_VERSION&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;GIT_SHA&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;dev&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;deployment.environment.name&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;APP_ENV&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;local&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;traceExporter&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OTLPTraceExporter&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OTEL_COLLECTOR_URL&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/v1/traces`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="na"&gt;instrumentations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="nf"&gt;getNodeAutoInstrumentations&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="c1"&gt;// Noisy, zero-signal, and it doubles span volume.&lt;/span&gt;
      &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/instrumentation-fs&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="nx"&gt;sdk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;SIGTERM&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;sdk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;shutdown&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="k"&gt;finally&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the one piece of manual instrumentation that auto-instrumentation can't do for you — wrapping a domain operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;trace&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;SpanStatusCode&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@opentelemetry/api&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tracer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;trace&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getTracer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;billing&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;chargeInvoice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;invoiceId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;amountCents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Span name = "&amp;lt;domain&amp;gt;.&amp;lt;operation&amp;gt;", low cardinality, never includes an ID.&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;tracer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startActiveSpan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;billing.charge_invoice&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAttribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;billing.invoice_id&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;invoiceId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAttribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;billing.amount_cents&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;amountCents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;gateway&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;charge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;invoiceId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;amountCents&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setAttribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;billing.gateway_status&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recordException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;setStatus&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;code&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;SpanStatusCode&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ERROR&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Roughly 90 lines of real, working, opinionated code. Handing the agent that file and saying &lt;em&gt;"do this, for this service"&lt;/em&gt; worked dramatically better than any description of it. Agents pattern-match. Give them a pattern.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Turn the conventions into a script that fails
&lt;/h3&gt;

&lt;p&gt;Prose conventions are unenforceable, and "looks right to me" doesn't scale across 47 pull requests. So I spent half a day writing a validator and told the agent it was not done until the validator passed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# check_tracing.py — run per service in CI. Exit 1 blocks the PR.
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;

&lt;span class="n"&gt;FORBIDDEN_ATTR&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;setAttribute\(\s*[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\"](?:.*\b(?:email|user_id|token|url|path|query)\b.*)[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\"]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;SPAN_NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;startActiveSpan\(\s*[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\"]([a-z0-9_]+\.[a-z0-9_]+)[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\"]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;TEMPLATED_SPAN_NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;startActiveSpan\(\s*[`&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\"].*\$\{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;errors&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tracing.js&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missing tracing.js bootstrap&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rglob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.js&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;TEMPLATED_SPAN_NAME&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: interpolated span name (cardinality bomb)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;FORBIDDEN_ATTR&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;finditer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: high-cardinality or PII attribute: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;SPAN_NAME&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;finditer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: span namespace != service name: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;problems&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;check&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argv&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;problems&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tracing conventions OK&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;problems&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This single file changed the economics of the whole project. Before it, I reviewed taste. After it, I reviewed &lt;strong&gt;judgment calls only&lt;/strong&gt; — the validator caught every mechanical violation before I ever opened the diff.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Prove the spans actually arrive
&lt;/h3&gt;

&lt;p&gt;A service can pass every static check and still export nothing, because the SDK started after the HTTP framework was imported, or the collector URL was wrong. So each service also got a smoke test that boots the process against a local collector and asserts on real exported spans:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// tracing.smoke.test.js&lt;/span&gt;
&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;http request produces a parented db span&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/invoices/42&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;flushSpans&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;spans&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;collector&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;finished&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;http&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;spans&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;GET /invoices/:id&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;spans&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;pg.query&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBeDefined&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;parentSpanContext&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;spanId&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;spanContext&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nx"&gt;spanId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="c1"&gt;// The route is templated, so no invoice IDs leak into span names.&lt;/span&gt;
  &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;http&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;not&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toMatch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/42/&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That parent-child assertion caught the single most common real failure: context propagation silently broken, producing 47 disconnected single-span "traces" that look fine in a list view and are worthless in a waterfall.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Run the loop
&lt;/h3&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Pick next service&amp;lt;br/&amp;gt;grouped by framework] --&amp;gt; B[Agent instruments&amp;lt;br/&amp;gt;against reference impl]
    B --&amp;gt; C{check_tracing.py}
    C -- fail --&amp;gt; B
    C -- pass --&amp;gt; D{smoke test&amp;lt;br/&amp;gt;vs local collector}
    D -- fail --&amp;gt; B
    D -- pass --&amp;gt; E[Human review:&amp;lt;br/&amp;gt;judgment calls only]
    E --&amp;gt; F[PR merged]
    F --&amp;gt; A&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The two validators ran inside the agent's own loop, so most iteration happened without me. My job was the last box: does this service's custom queue consumer actually deserve a span here, and is this attribute worth its storage cost?&lt;/p&gt;

&lt;p&gt;Nine days, 47 services, about 90 minutes a day of my attention. The first six services took four of those days. The remaining 41 took five.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. A working reference implementation beats any style guide
&lt;/h3&gt;

&lt;p&gt;My 2,000-word convention doc produced ~70% compliance. A 90-line reference file produced near-perfect structural compliance immediately. If you find yourself writing a long prose description of how code should look, stop and write the code instead — then point at it.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Make the convention executable or it doesn't exist
&lt;/h3&gt;

&lt;p&gt;The validator was the highest-leverage half-day of the project. Anything you'd write as "please always..." in a prompt should be a script that exits non-zero. Prompts are advice; exit codes are physics. It also means the convention survives you: the next person to add a service gets the same enforcement without reading a single doc.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Cardinality is where an agent will hurt you
&lt;/h3&gt;

&lt;p&gt;Left alone, the agent wrote genuinely reasonable-looking code like &lt;code&gt;startActiveSpan(`charge invoice ${invoiceId}`)&lt;/code&gt; and &lt;code&gt;span.setAttribute('user.email', email)&lt;/code&gt;. Both are defensible as "descriptive." Both are catastrophic — unbounded span names wreck your backend's indexing and cost, and PII in attributes is a compliance incident that lives in your telemetry store for the retention period.&lt;/p&gt;

&lt;p&gt;The agent isn't being careless; &lt;strong&gt;it's optimizing for readability because nothing told it to optimize for cardinality.&lt;/strong&gt; That tradeoff is invisible in the diff and expensive in production, which makes it exactly the kind of rule that belongs in a linter rather than a review comment.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Adding instrumentation is mechanical; removing the old logging isn't
&lt;/h3&gt;

&lt;p&gt;My original plan included "delete the now-redundant logging." I cut that from the agent's scope on day two. Deciding whether a log line is redundant with a span requires knowing who greps for it at 3am — that's tribal knowledge the codebase doesn't contain, and an agent confidently deleting an on-call runbook's load-bearing log line is a bad trade against the time it saves.&lt;/p&gt;

&lt;p&gt;Split the work by &lt;em&gt;information availability&lt;/em&gt;, not by difficulty. The mechanical 90% went to the agent. The 10% that needs context living in someone's head stayed with people.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Batch by framework, not alphabetically
&lt;/h3&gt;

&lt;p&gt;I started in repo order and burned tokens re-establishing context every single service. Grouping by stack — all Fastify, then all Express, then all FastAPI — meant each batch reused the same mental model, the same gotchas, and the same recently-solved problems. Same work, noticeably fewer tokens and far fewer wrong turns.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Adjacent tasks are cheaper than shuffled tasks. Order your backlog by similarity, not by convenience.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Three things I'm working on now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sampling policy.&lt;/strong&gt; We're at head-based 100% sampling in staging and it's already loud. Tail-based sampling that keeps every error trace and 5% of the boring ones is the obvious next move.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trace-driven perf work.&lt;/strong&gt; Now that waterfalls exist, I want to feed real slow traces back to the agent as the starting point for optimization instead of vague "this endpoint feels slow" prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generated service dependency graphs.&lt;/strong&gt; The traces already encode the real call graph, including the three calls nobody remembered existed. Rendering that automatically beats maintaining an architecture diagram by hand.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern generalizes well beyond tracing. Any sweeping, repetitive, convention-heavy migration — feature flag SDKs, error reporting, auth middleware — fits the same shape: &lt;strong&gt;one hand-written reference, one executable validator, one smoke test that proves it's live, then let the agent grind.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-Up / CTA
&lt;/h2&gt;

&lt;p&gt;If you're sitting on a "someday" observability migration because it's six weeks of boring work, it probably isn't six weeks anymore. But the leverage isn't in the prompt — it's in the reference implementation and the validator you write &lt;em&gt;before&lt;/em&gt; you start.&lt;/p&gt;

&lt;p&gt;If you try this, I'd genuinely like to know what your validator ends up catching. Mine caught 31 cardinality violations I would have merged without noticing.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;💬 &lt;strong&gt;Comment below&lt;/strong&gt; with the worst high-cardinality span name you've shipped. I'll go first: &lt;code&gt;GET /users/8f21c...&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;🚀 &lt;strong&gt;Follow me here on Dev.to&lt;/strong&gt; — I write up these AI-agent engineering experiments as I run them.&lt;/li&gt;
&lt;li&gt;💡 &lt;strong&gt;Try it yourself&lt;/strong&gt; with Claude Code on one service this week. One service is enough to find out whether your conventions are real or imaginary.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>observability</category>
      <category>devops</category>
    </item>
    <item>
      <title>How I Sandboxed My AI Coding Agent's Shell Access Without Slowing It Down</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Wed, 23 Sep 2026 14:32:16 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-sandboxed-my-ai-coding-agents-shell-access-without-slowing-it-down-al4</link>
      <guid>https://dev.to/yureki_lab/how-i-sandboxed-my-ai-coding-agents-shell-access-without-slowing-it-down-al4</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I gave my autonomous coding agent real shell access, then spent two weeks building guardrails around it after a &lt;code&gt;rm -rf&lt;/code&gt; near-miss. The result is a three-tier command classifier that parses shell syntax instead of regex-matching it, plus filesystem and network confinement and an append-only audit log. Interruption rate dropped from 31% of commands to 4%, and I stopped watching the terminal like a hawk.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;An agent that can't run shell commands is a very expensive autocomplete.&lt;/p&gt;

&lt;p&gt;The whole point of a fully autonomous implementation system is that it closes the loop by itself: write code, run the tests, read the failure, fix it, run the tests again. Take away &lt;code&gt;Bash&lt;/code&gt; and every one of those steps needs me. I tried the read-only-agent approach for about a week. It produced code that looked correct and was never once verified.&lt;/p&gt;

&lt;p&gt;So I turned shell access on. And then I watched it do this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# the agent, "cleaning up" a test fixture directory&lt;/span&gt;
&lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; &lt;span class="nv"&gt;$BUILD_DIR&lt;/span&gt;/artifacts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;BUILD_DIR&lt;/code&gt; was unset in that subprocess. That expands to &lt;code&gt;rm -rf /artifacts&lt;/code&gt;. On my machine that particular path didn't exist, so nothing happened, and I only noticed because I was reading the transcript for an unrelated reason.&lt;/p&gt;

&lt;p&gt;That's the real problem with agent shell access. It isn't that the agent is malicious — it obviously isn't. It's that:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The agent doesn't know what it doesn't know.&lt;/strong&gt; Unset variables, a different working directory than it assumed, a &lt;code&gt;git&lt;/code&gt; remote that isn't the one it thinks it is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failures are silent and asymmetric.&lt;/strong&gt; 400 successful &lt;code&gt;npm test&lt;/code&gt; runs don't earn back one &lt;code&gt;git push --force&lt;/code&gt; to &lt;code&gt;main&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The obvious fix makes the agent useless.&lt;/strong&gt; My first attempt was "confirm every command." I approved 180 prompts in one afternoon, 174 of which were &lt;code&gt;ls&lt;/code&gt;, &lt;code&gt;cat&lt;/code&gt;, and &lt;code&gt;pytest&lt;/code&gt;. I started rubber-stamping by hour two, which is strictly worse than no gate at all — I had the &lt;em&gt;feeling&lt;/em&gt; of oversight without the substance.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What I actually wanted: the agent runs the boring 95% at full speed, and I get a hard stop on the 5% that can ruin my day.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;p&gt;Three layers. Each one catches a different class of mistake.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Agent proposes command] --&amp;gt; B[Layer 1: Parse and classify]
    B --&amp;gt;|deny| C[Blocked, reason returned to agent]
    B --&amp;gt;|ask| D[Human confirm]
    B --&amp;gt;|allow| E[Layer 2: Confined execution]
    D --&amp;gt;|approved| E
    D --&amp;gt;|rejected| C
    E --&amp;gt; F[Layer 3: Append-only audit log]
    F --&amp;gt; G[Output back to agent]&lt;/code&gt;&lt;/pre&gt;



&lt;h3&gt;
  
  
  Layer 1: Classify by parsing, not by regex
&lt;/h3&gt;

&lt;p&gt;This is the part I got wrong first, so it's the part worth the most words.&lt;/p&gt;

&lt;p&gt;My original rules were regex against the raw command string. &lt;code&gt;^rm\s+-rf\s+/&lt;/code&gt; and friends. It took the agent about three days to walk straight through one, not because it was trying to, but because it wrote a perfectly ordinary command that my pattern didn't anticipate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;git rev-parse &lt;span class="nt"&gt;--show-toplevel&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;/tmp"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; ./&lt;span class="k"&gt;*&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No leading slash. No match. My regex was looking at text; the danger was in the semantics.&lt;/p&gt;

&lt;p&gt;The fix was to stop pattern-matching strings and start parsing shell into a syntax tree. Python's &lt;code&gt;shlex&lt;/code&gt; gets you tokenization; &lt;code&gt;bashlex&lt;/code&gt; gets you actual structure — pipelines, redirects, command substitutions, &lt;code&gt;&amp;amp;&amp;amp;&lt;/code&gt; chains. Once you have the tree you can walk every simple command inside a compound one and classify each independently.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;bashlex&lt;/span&gt;

&lt;span class="n"&gt;DENY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rm&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-rf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;        &lt;span class="c1"&gt;# root deletion, any form
&lt;/span&gt;    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;push&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--force&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chmod&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;777&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;ASK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git push&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git reset --hard&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;docker&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;brew&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pip install&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;npm publish&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;ALLOW&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ls&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;grep&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;find&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git diff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pytest&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;npm test&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;make&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;head&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wc&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;simple_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Yield every simple command inside a compound command.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;tree&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;bashlex&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;node&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;_walk&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tree&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;word&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;node&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;word&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;commands&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;simple_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;bashlex&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ParsingError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ask&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unparseable shell; falling back to human review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="n"&gt;verdicts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;words&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;commands&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;words&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;
        &lt;span class="n"&gt;verdicts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;_classify_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;words&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="c1"&gt;# The whole pipeline is only as safe as its most dangerous link.
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deny&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ask&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allow&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;verdicts&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ask&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no rule matched&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two design decisions in there that I'd defend to anyone:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unparseable means ask, never allow.&lt;/strong&gt; If my parser can't understand the command, I have zero information about it, and zero information is not the same as "safe." Early on I had this fall through to &lt;code&gt;allow&lt;/code&gt; because parse failures were annoying. They were annoying because they were mostly heredocs and process substitution — exactly the constructs where interesting things hide.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The most dangerous link decides.&lt;/strong&gt; &lt;code&gt;npm test &amp;amp;&amp;amp; rm -rf ./dist&lt;/code&gt; is not an &lt;code&gt;allow&lt;/code&gt; just because it starts with one. Walking every simple command in the tree and taking the worst verdict is the only version of this that holds up.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 2: Confine what "allow" can reach
&lt;/h3&gt;

&lt;p&gt;Classification decides &lt;em&gt;whether&lt;/em&gt; to run. Confinement decides &lt;em&gt;how much it can touch&lt;/em&gt; when you get the classification wrong — and you will get it wrong.&lt;/p&gt;

&lt;p&gt;Three cheap constraints, none of which required a container:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_confined&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;workdir&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;env&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PATH&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/usr/bin:/bin:/usr/local/bin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HOME&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;workdir&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;             &lt;span class="c1"&gt;# keep stray writes inside the work tree
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LANG&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;en_US.UTF-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="c1"&gt;# Note what is NOT here: every API token in my real shell.
&lt;/span&gt;    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/bin/bash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-o&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pipefail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-c&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;workdir&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                     &lt;span class="c1"&gt;# allowlist, not os.environ.copy()
&lt;/span&gt;        &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;             &lt;span class="c1"&gt;# no unbounded hangs
&lt;/span&gt;    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The env allowlist is the highest-value line in this entire post. The default instinct is &lt;code&gt;os.environ.copy()&lt;/code&gt;, which hands the agent every credential you've ever exported — and those credentials end up in command output, which ends up in the model's context, which ends up in logs. Starting from an empty dict and adding back only what's needed took twenty minutes and removed an entire category of problem.&lt;/p&gt;

&lt;p&gt;For network I run the agent's shell under a user account whose outbound traffic is filtered to a host allowlist (package registries, my git remote, and nothing else). The agent can &lt;code&gt;npm install&lt;/code&gt;. It cannot &lt;code&gt;curl&lt;/code&gt; an arbitrary host and pipe the result to &lt;code&gt;bash&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 3: Log everything, append-only
&lt;/h3&gt;

&lt;p&gt;Every command, verdict, exit code, and duration goes to a JSONL file that the agent can write to but not read or rewrite:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;audit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exit_code&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;AUDIT_PATH&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cmd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verdict&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;exit_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;exit_code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;duration_ms&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This started as a debugging aid and turned into the tuning instrument for the whole system. Once a week I run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'select(.verdict=="ask") | .cmd'&lt;/span&gt; audit.jsonl &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $1, $2}'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Anything at the top of that list that I approved every single time is a rule that's costing me attention and buying me nothing. It gets promoted to &lt;code&gt;allow&lt;/code&gt;. That one query is how the interruption rate went from 31% to 4% without loosening anything that mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. A gate you rubber-stamp is worse than no gate.&lt;/strong&gt; Confirming everything trained me in about ninety minutes to approve without reading. The security value of a prompt is exactly the attention you still bring to it, and attention is a budget that depletes. Spend it on the 4%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Regex over shell strings is security theater.&lt;/strong&gt; Shell has command substitution, variable expansion, quoting, and chaining. Any of those turns a dangerous command into text your pattern doesn't match. Parse it or don't pretend to check it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Fail toward the human, never toward execution.&lt;/strong&gt; Every ambiguous case — parse error, unknown binary, no matching rule — resolves to &lt;code&gt;ask&lt;/code&gt;. This produces more prompts early, which is uncomfortable, and then the audit log tells you exactly which ones to relax.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The environment is the attack surface nobody looks at.&lt;/strong&gt; I spent days on command classification while every API key I own was being handed to every subprocess. An env allowlist is twenty minutes of work and it's probably the single highest-leverage thing in this post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Measure the interruption rate or you're guessing.&lt;/strong&gt; "Does this feel safe" is not a metric. Commands-per-prompt is. Track it, and guardrail tuning becomes an ordinary optimization problem instead of a vibes argument with yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things I haven't solved:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per-command network policy.&lt;/strong&gt; Right now the allowlist is process-wide. I'd like &lt;code&gt;npm install&lt;/code&gt; to reach the registry and nothing else to reach anything, which probably means a real sandbox rather than a filtered user account.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Learned classification.&lt;/strong&gt; The rule set is hand-written and it's at 140 lines. I'd like to feed the audit log back in and have consistently-approved patterns propose their own promotion — with me approving the promotion, not the individual commands.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stack this was built and tested on, since all of it rots: Claude Code (2026-09 release), Python 3.13, &lt;code&gt;bashlex&lt;/code&gt; 0.18, macOS 15 and Ubuntu 24.04.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If you're giving an agent shell access — and you should, a verified loop is worth far more than a supervised one — start with the env allowlist. It's an afternoon at most. Then add the audit log, run for a week, and let the data tell you where the real gate belongs. Don't start with 200 rules; you'll write the wrong ones.&lt;/p&gt;

&lt;p&gt;If you've built something similar, I'd genuinely like to know how you handled the network layer — it's the part I'm least happy with. Drop it in the comments.&lt;/p&gt;

&lt;p&gt;And if this was useful, &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; — I write up what I learn building autonomous coding systems, war stories and dead ends included. 🚀&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>security</category>
      <category>devops</category>
    </item>
    <item>
      <title>5 Logging Habits That Made My AI Coding Agent Actually Good at Debugging</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Tue, 22 Sep 2026 14:33:14 +0000</pubDate>
      <link>https://dev.to/yureki_lab/5-logging-habits-that-made-my-ai-coding-agent-actually-good-at-debugging-3f12</link>
      <guid>https://dev.to/yureki_lab/5-logging-habits-that-made-my-ai-coding-agent-actually-good-at-debugging-3f12</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;My AI coding agent wasn't bad at debugging — my logs were bad at being read. Five changes to how my services log (correlation IDs, structured JSON, errors that carry state, honest log levels, and a query command instead of a log file) took its median time-to-root-cause from ~40 minutes of flailing down to under 5. None of these habits are new. What's new is that a machine is now the primary reader of your logs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;A checkout endpoint started returning 500s for about 2% of requests. Classic intermittent bug: not reproducible locally, no obvious pattern, and the kind of thing that eats an afternoon.&lt;/p&gt;

&lt;p&gt;So I did what I'd been doing all year — handed it to my coding agent with production log access and went to make coffee.&lt;/p&gt;

&lt;p&gt;Twenty minutes later it had:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;grepped &lt;code&gt;error&lt;/code&gt; across a 2.3 GB log file and found 4,100 matches&lt;/li&gt;
&lt;li&gt;picked a &lt;code&gt;Redis connection reset&lt;/code&gt; line that turned out to be unrelated noise from a nightly job&lt;/li&gt;
&lt;li&gt;added a retry wrapper around the Redis client&lt;/li&gt;
&lt;li&gt;told me, confidently, that this "should resolve the intermittent 500s"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It didn't. The actual bug was a race between a coupon-validation call and a cart-refresh call, and the evidence was sitting right there in the logs — just spread across four lines that had nothing linking them together.&lt;/p&gt;

&lt;p&gt;That's when it clicked. The agent hadn't failed at reasoning. It had failed at &lt;strong&gt;evidence gathering&lt;/strong&gt;, because my logs made evidence gathering nearly impossible. I'd spent years writing logs for a reader who already knew the system: me. Lines like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[2026-09-04 11:42:07] processing cart
[2026-09-04 11:42:07] validating
[2026-09-04 11:42:08] failed, retrying
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I can read that. I know "validating" means the coupon service and "failed, retrying" is the HTTP client's third attempt. An agent reads those lines the way a new hire does — except a new hire asks you a question, and an agent makes a plausible guess and writes code based on it.&lt;/p&gt;

&lt;p&gt;The constraint that made this interesting: I couldn't make the agent smarter, but I could completely control what it had to read. So I spent a week rewriting logging across three services with exactly one design goal — &lt;strong&gt;a reader with zero prior context should be able to reconstruct one request end-to-end&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;p&gt;Here's the loop I was optimizing for. The failure mode is always the same: the agent can't narrow, so it guesses.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
    A[Bug report] --&amp;gt; B{Can the agent isolate&amp;lt;br/&amp;gt;one failing request?}
    B --&amp;gt;|No| C[Grep for 'error']
    C --&amp;gt; D[4,100 matches]
    D --&amp;gt; E[Picks a plausible line]
    E --&amp;gt; F[Fixes the wrong thing]
    B --&amp;gt;|Yes| G[Pull full trace by ID]
    G --&amp;gt; H[Read ordered, typed events]
    H --&amp;gt; I[Compare to a passing request]
    I --&amp;gt; J[Root cause]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Everything below is about moving from the left branch to the right branch. Stack: Python 3.13 with &lt;code&gt;structlog&lt;/code&gt;, Node.js 22 LTS with &lt;code&gt;pino&lt;/code&gt; 9.x, and Claude Code CLI (v2.x line, as of September 2026) as the agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Habit 1: One ID that survives the whole request
&lt;/h3&gt;

&lt;p&gt;This is the single highest-leverage change, and it's boring. Generate an ID at the edge, stuff it in a context variable, and attach it to every single log line — including the ones inside your background workers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Python 3.13 + structlog
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;contextvars&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;structlog&lt;/span&gt;

&lt;span class="n"&gt;request_id_var&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;contextvars&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ContextVar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;default&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;bind_request_id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;method_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;event_dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;rid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request_id_var&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;rid&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;event_dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rid&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;event_dict&lt;/span&gt;

&lt;span class="n"&gt;structlog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;configure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;processors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="n"&gt;bind_request_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;structlog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;processors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;TimeStamper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;iso&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;structlog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;processors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;JSONRenderer&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# FastAPI middleware
&lt;/span&gt;&lt;span class="nd"&gt;@app.middleware&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;attach_request_id&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;call_next&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;rid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-request-id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;request_id_var&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;call_next&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-request-id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rid&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The payoff isn't the ID itself — it's that "show me everything about the failing request" becomes a single filter instead of a research project. The first time I gave my agent a codebase with this in place, it stopped grepping for keywords entirely and started pulling traces.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If you adopt exactly one thing from this post, adopt this one. It's about 30 lines per service and it changes what questions are answerable.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Habit 2: Structured JSON, not prose
&lt;/h3&gt;

&lt;p&gt;Prose logs force the agent to write a regex to parse your sentence structure, which it will do badly, and then reason on top of that bad parse. Compounding errors.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before — needs a parser, and the parser will be wrong
&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Coupon &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; rejected for user &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;uid&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; after &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;ms&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# After — needs no parser
&lt;/span&gt;&lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;coupon_rejected&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;coupon&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;uid&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;duration_ms&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two rules that matter more than the format itself:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The event name is a stable identifier, not a sentence.&lt;/strong&gt; &lt;code&gt;coupon_rejected&lt;/code&gt;, not &lt;code&gt;"Coupon was rejected"&lt;/code&gt;. Stable names make cross-request comparison trivial: count &lt;code&gt;coupon_rejected&lt;/code&gt; in failing vs. passing traces and the anomaly jumps out.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Values go in fields, never interpolated into the message.&lt;/strong&gt; The moment a user ID lives inside a string, filtering by user ID becomes substring matching.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Habit 3: Errors that carry state, not just a stack trace
&lt;/h3&gt;

&lt;p&gt;A stack trace tells you &lt;em&gt;where&lt;/em&gt; it broke. It almost never tells you &lt;em&gt;why&lt;/em&gt;. An agent staring at &lt;code&gt;TypeError: Cannot read properties of undefined (reading 'total')&lt;/code&gt; at line 214 will go read line 214 and start theorizing about null checks — when the real story is that an upstream call returned a 204 with an empty body.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Node.js 22 LTS + pino 9&lt;/span&gt;
&lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;applyCoupon&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cart&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;code&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;                          &lt;span class="c1"&gt;// pino serializes stack + message&lt;/span&gt;
    &lt;span class="na"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;coupon_apply_failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;coupon&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;cart_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cart&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;cart_item_count&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cart&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;cart_version&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;cart&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;version&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;// the field that actually solved it&lt;/span&gt;
    &lt;span class="na"&gt;upstream_status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cause&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;coupon application failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;cart_version&lt;/code&gt; was the whole ballgame in my race condition. Two requests were mutating the same cart, and the version numbers in the log made it obvious within seconds — to the agent, not to me. I'd been staring at those logs for an hour.&lt;/p&gt;

&lt;p&gt;The rule I now apply when writing an error log: &lt;strong&gt;what would I need to know to reproduce this without asking anyone?&lt;/strong&gt; Log that. If the answer is "the state of the object", log the state of the object.&lt;/p&gt;

&lt;h3&gt;
  
  
  Habit 4: Log levels that mean something
&lt;/h3&gt;

&lt;p&gt;My levels had rotted into decoration. Half of &lt;code&gt;ERROR&lt;/code&gt; was retryable noise; real failures were hiding at &lt;code&gt;WARN&lt;/code&gt; because someone didn't want to page anybody. An agent takes your levels literally — that's why mine latched onto a &lt;code&gt;Redis connection reset&lt;/code&gt; that a human would have skipped on sight.&lt;/p&gt;

&lt;p&gt;The definitions I settled on and wrote into the project spec file:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;Agent's reading&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ERROR&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A user-visible operation failed and will not succeed&lt;/td&gt;
&lt;td&gt;Start here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;WARN&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Something degraded but recovered (retry succeeded, fallback used)&lt;/td&gt;
&lt;td&gt;Context, not cause&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;INFO&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A business event happened (order placed, coupon applied)&lt;/td&gt;
&lt;td&gt;Timeline material&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;DEBUG&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Internal step detail&lt;/td&gt;
&lt;td&gt;Only when tracing one request&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Then I did the unglamorous part: swept the codebase and re-leveled roughly 300 call sites to match. Unsurprisingly, this is exactly the kind of mechanical-but-judgment-heavy sweep an agent is great at — I reviewed the diff in chunks rather than writing it by hand.&lt;/p&gt;

&lt;h3&gt;
  
  
  Habit 5: Give the agent a query command, not a log file
&lt;/h3&gt;

&lt;p&gt;Pointing an agent at raw logs is how you burn 50k tokens to learn nothing. Pointing it at a &lt;em&gt;command&lt;/em&gt; that returns exactly one trace is how you burn 2k tokens and get an answer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="c"&gt;# logq — trace a single request across services&lt;/span&gt;
&lt;span class="c"&gt;# usage: logq trace &amp;lt;request_id&amp;gt; [--since 24h]&lt;/span&gt;
&lt;span class="c"&gt;#        logq errors [--since 1h] [--top 20]&lt;/span&gt;

&lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in
  &lt;/span&gt;trace&lt;span class="p"&gt;)&lt;/span&gt;
    jq &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="nt"&gt;--arg&lt;/span&gt; rid &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$2&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s1"&gt;'select(.request_id == $rid)'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LOG_PATH&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
      | jq &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s1"&gt;'sort_by(.timestamp)'&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt;
  errors&lt;span class="p"&gt;)&lt;/span&gt;
    jq &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'select(.level == "error")'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LOG_PATH&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
      | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.event'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; -&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;3&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;20&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt;
&lt;span class="k"&gt;esac&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I documented it in &lt;code&gt;CLAUDE.md&lt;/code&gt; so the agent reaches for it unprompted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Debugging production issues&lt;/span&gt;

Never grep the raw log file — it's multi-GB and unstructured at the edges.
&lt;span class="p"&gt;
1.&lt;/span&gt; &lt;span class="sb"&gt;`logq errors --since 1h`&lt;/span&gt; to find which event names are spiking
&lt;span class="p"&gt;2.&lt;/span&gt; &lt;span class="sb"&gt;`logq trace &amp;lt;request_id&amp;gt;`&lt;/span&gt; to pull one full request timeline
&lt;span class="p"&gt;3.&lt;/span&gt; Always pull a &lt;span class="ge"&gt;*passing*&lt;/span&gt; trace too and diff the two before theorizing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line did more than the script did. "Compare a failing trace to a passing trace" is the single instruction that most reliably stops an agent from pattern-matching on the first scary-looking line it sees.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Your logs are now a machine-readable API. Version them like one.&lt;/strong&gt; Once an agent (and a &lt;code&gt;logq&lt;/code&gt; script, and a dashboard) depends on the event name &lt;code&gt;coupon_rejected&lt;/code&gt;, renaming it is a breaking change. I treat event names with the same care as route paths now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Agents fail at evidence gathering far more than at reasoning.&lt;/strong&gt; Every "the AI wrote a dumb fix" story I've hit this year traced back to the agent reasoning correctly over bad or partial evidence. Fixing inputs beat every prompt-engineering trick I tried.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Optimize for "reconstruct one request", not "search everything".&lt;/strong&gt; Search-everything is a human affordance — we're good at skimming and discarding. An agent's context window is small and expensive, and irrelevant lines don't just waste tokens, they actively mislead. Narrow beats complete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Tools beat instructions when the task is mechanical.&lt;/strong&gt; I spent two days writing increasingly elaborate prompt instructions about how to search logs carefully. A 20-line shell script made all of them unnecessary. If you find yourself writing a paragraph explaining &lt;em&gt;how&lt;/em&gt; to gather information, you should probably be writing a command instead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. This is just good observability, and that's the point.&lt;/strong&gt; Every habit here is something an SRE would have told you in 2018. The AI angle didn't change the advice — it changed the ROI. Sloppy logs used to cost me an afternoon occasionally. Now they cost me every single automated debugging run, and the improvements compound daily.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things I'm working on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sampling &lt;code&gt;DEBUG&lt;/code&gt; in production for a small slice of traffic.&lt;/strong&gt; Right now &lt;code&gt;DEBUG&lt;/code&gt; is off in prod, which means the richest data disappears exactly when I need it. Head-based sampling at ~1% with forced capture on error looks like the right trade.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wiring OpenTelemetry spans into the same query command&lt;/strong&gt;, so &lt;code&gt;logq trace&lt;/code&gt; returns logs and timing spans on one timeline. Half of my remaining hard bugs are latency-shaped, and logs alone under-serve those.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'd also like to stop maintaining &lt;code&gt;logq&lt;/code&gt; as a shell script and expose it as a proper tool the agent calls directly — the bash-wrapper stage is clearly temporary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If your AI agent keeps "fixing" the wrong thing, before you rewrite your prompts, go read your logs the way it has to: with no prior knowledge, one line at a time, unable to ask a question. It's an uncomfortable exercise and it'll show you exactly what to fix.&lt;/p&gt;

&lt;p&gt;Start with correlation IDs. It's an afternoon of work and it's the one that unlocks everything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your turn:&lt;/strong&gt; what's the one logging change that paid for itself fastest in your codebase? Drop it in the comments — I'm still collecting these. 👇&lt;/p&gt;

&lt;p&gt;If this was useful, follow me here on Dev.to — I write about building and living with autonomous coding agents, usually with the embarrassing parts left in. 🚀&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>debugging</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>How I Built a Research Agent That Reads the Docs Before My Coding Agent Ships</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:32:22 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-built-a-research-agent-that-reads-the-docs-before-my-coding-agent-ships-c9a</link>
      <guid>https://dev.to/yureki_lab/how-i-built-a-research-agent-that-reads-the-docs-before-my-coding-agent-ships-c9a</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;My autonomous coding agent kept shipping code against library APIs that didn't exist, because it wrote first and read the docs never. I fixed it by adding a &lt;strong&gt;read-only research sub-agent&lt;/strong&gt; that runs before implementation, returns a short structured brief, and blocks the implementer until that brief exists. Hallucinated-API bugs dropped from roughly one in five tasks to about one in forty over three months, and the extra cost was under 10% of tokens. Here's how it works and what I got wrong along the way. 🚀&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;I run a fully autonomous implementation system on a Mac mini. An orchestrator picks up tasks, hands them to parallel implementation agents built on Claude Code, and a self-healing agent cleans up whatever fails. It ships real code to real projects, mostly unattended.&lt;/p&gt;

&lt;p&gt;For the first few months, the single most annoying failure mode was this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The implementer confidently calls &lt;code&gt;client.batchUpsert()&lt;/code&gt;. The library has never had a &lt;code&gt;batchUpsert()&lt;/code&gt;. It has &lt;code&gt;upsertMany()&lt;/code&gt;, added in v4.2, with a different argument shape.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tests fail, the self-healing agent tries three fixes, burns 40k tokens, and finally gives up and marks the task blocked. I'd wake up to a pile of "blocked" tasks that all shared one root cause: &lt;strong&gt;the agent guessed an API instead of reading it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I audited two weeks of failed tasks, the numbers were ugly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure cause&lt;/th&gt;
&lt;th&gt;Share of failed tasks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hallucinated or outdated library API&lt;/td&gt;
&lt;td&gt;38%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Misread existing internal code&lt;/td&gt;
&lt;td&gt;22%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Genuinely hard bug&lt;/td&gt;
&lt;td&gt;19%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Everything else&lt;/td&gt;
&lt;td&gt;21%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sixty percent of failures were "didn't look before writing." That's not an intelligence problem. That's a process problem.&lt;/p&gt;

&lt;p&gt;What made it interesting: the implementer &lt;em&gt;could&lt;/em&gt; read docs. It had web fetch and file tools. It just didn't, because the moment you give a capable model a task, its instinct is to start producing code. Prompting "please read the docs first" helped for about two days, then the instruction got crowded out by everything else in the context window.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;p&gt;The fix was structural, not prompt-based. I split "figure out how this works" from "build it" into two different agents with two different tool sets.&lt;/p&gt;

&lt;h3&gt;
  
  
  The shape of it
&lt;/h3&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    O[Orchestrator] --&amp;gt;|task + question list| R[Research agent&amp;lt;br/&amp;gt;read-only]
    R --&amp;gt;|research brief| O
    O --&amp;gt;|task + brief| I[Implementation agent]
    I --&amp;gt;|diff| V[Verifier]
    V --&amp;gt;|pass / fail| O
    O -.-&amp;gt;|brief missing or low confidence| R&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Three rules make this work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The research agent cannot write.&lt;/strong&gt; It gets Read, Grep, Glob, web search, and web fetch. No Edit, no Write, no Bash. It physically cannot "just fix it real quick."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The implementer cannot start without a brief.&lt;/strong&gt; The orchestrator refuses to dispatch an implementation task unless a research brief file for that task exists and has a confidence score.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The brief is short and structured.&lt;/strong&gt; Not a doc dump. A contract.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  The research brief contract
&lt;/h3&gt;

&lt;p&gt;Every brief is a small Markdown file with a fixed shape. Here's a real one, lightly anonymized:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Research brief: task-0417&lt;/span&gt;

&lt;span class="gu"&gt;## Questions&lt;/span&gt;
&lt;span class="p"&gt;1.&lt;/span&gt; How do we bulk-insert rows with the ORM at the pinned version?
&lt;span class="p"&gt;2.&lt;/span&gt; Does the existing repo already wrap this anywhere?

&lt;span class="gu"&gt;## Findings&lt;/span&gt;
&lt;span class="gu"&gt;### Q1&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Pinned version is 4.3.1 (from lockfile).
&lt;span class="p"&gt;-&lt;/span&gt; Bulk insert is &lt;span class="sb"&gt;`Model.bulkCreate(rows, { updateOnDuplicate: [...] })`&lt;/span&gt;.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`upsertMany`&lt;/span&gt; does NOT exist in this version. It was a community plugin.
&lt;span class="p"&gt;-&lt;/span&gt; Source: node_modules/&lt;span class="nt"&gt;&amp;lt;orm&amp;gt;&lt;/span&gt;/lib/model.js lines 2210-2290, and the 4.x changelog.

&lt;span class="gu"&gt;### Q2&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Yes. &lt;span class="sb"&gt;`src/db/batch.ts`&lt;/span&gt; exports &lt;span class="sb"&gt;`insertBatch()`&lt;/span&gt; which already handles chunking at 500 rows.
&lt;span class="p"&gt;-&lt;/span&gt; Two callers use it. Reuse it; don't add a second path.

&lt;span class="gu"&gt;## Verdict&lt;/span&gt;
Use the existing &lt;span class="sb"&gt;`insertBatch()`&lt;/span&gt; helper. Do not touch the ORM directly.

&lt;span class="gu"&gt;## Confidence: 0.9&lt;/span&gt;
&lt;span class="gu"&gt;## Sources read: 4 files, 1 changelog, 0 web pages&lt;/span&gt;
&lt;span class="gu"&gt;## Staleness risk: low (pinned version, checked lockfile)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The four fields that matter are &lt;strong&gt;Questions&lt;/strong&gt;, &lt;strong&gt;Verdict&lt;/strong&gt;, &lt;strong&gt;Confidence&lt;/strong&gt;, and &lt;strong&gt;Staleness risk&lt;/strong&gt;. Everything else is supporting evidence. The implementer reads the verdict first. The orchestrator reads the confidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  The research agent definition
&lt;/h3&gt;

&lt;p&gt;I define sub-agents as Markdown files with frontmatter. The research one looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;researcher&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read-only investigation. Answers "how does X actually work"&lt;/span&gt;
  &lt;span class="s"&gt;by reading source, lockfiles, changelogs, and official docs. Returns a&lt;/span&gt;
  &lt;span class="s"&gt;research brief, never code.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Grep, Glob, WebSearch, WebFetch&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

You answer questions about how things actually work. You never propose
an implementation.

Rules:
&lt;span class="p"&gt;-&lt;/span&gt; Check the pinned version in the lockfile BEFORE reading any docs.
  Docs for the wrong major version are worse than no docs.
&lt;span class="p"&gt;-&lt;/span&gt; Prefer reading node_modules / site-packages source over web docs.
  Source can't be out of date for the installed version.
&lt;span class="p"&gt;-&lt;/span&gt; Search the repo for existing wrappers before answering "how do I call X".
&lt;span class="p"&gt;-&lt;/span&gt; If two sources disagree, report both and lower your confidence.
&lt;span class="p"&gt;-&lt;/span&gt; Output ONLY the brief format. No preamble.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That "check the lockfile first" line was added after a painful week. More on that below.&lt;/p&gt;

&lt;h3&gt;
  
  
  How the orchestrator gates on it
&lt;/h3&gt;

&lt;p&gt;The gate itself is boring shell. That's the point. It runs before every dispatch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;brief&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"state/briefs/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;task_id&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.md"&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$brief&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;dispatch_agent researcher &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$task_id&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;0   &lt;span class="c"&gt;# come back next tick&lt;/span&gt;
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nv"&gt;confidence&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'^## Confidence:'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$brief&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $3}'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;((&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$confidence&lt;/span&gt;&lt;span class="s2"&gt; &amp;lt; 0.6"&lt;/span&gt; | bc &lt;span class="nt"&gt;-l&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;))&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
  &lt;span class="c"&gt;# Re-research with the low-confidence sections as new questions&lt;/span&gt;
  dispatch_agent researcher &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$task_id&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--refine&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="k"&gt;fi

&lt;/span&gt;dispatch_agent implementer &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$task_id&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--brief&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$brief&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things I like about this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic gating.&lt;/strong&gt; I don't ask the model whether it feels ready. A file exists with a number in it, or it doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refinement is a loop with a cap.&lt;/strong&gt; Low confidence triggers a second research pass focused on the weak questions. After two refinements, the task is marked "needs human" instead of looping forever.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What the implementer sees
&lt;/h3&gt;

&lt;p&gt;The implementer's prompt gets the brief injected verbatim at the top, followed by one line:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The brief above is the source of truth for library APIs and existing helpers. If you believe it is wrong, stop and write why to the task file. Do not work around it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That last sentence matters. Before I added it, the implementer would occasionally read the brief, disagree silently, and do its own thing. Now disagreement is a visible event I can grep for.&lt;/p&gt;

&lt;h3&gt;
  
  
  Results after three months
&lt;/h3&gt;

&lt;p&gt;Across about 1,100 tasks on four projects:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Tasks failed on hallucinated API&lt;/td&gt;
&lt;td&gt;~20%&lt;/td&gt;
&lt;td&gt;~2.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-healing retries per task (avg)&lt;/td&gt;
&lt;td&gt;1.8&lt;/td&gt;
&lt;td&gt;0.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per completed task&lt;/td&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;td&gt;+8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tasks marked "needs human"&lt;/td&gt;
&lt;td&gt;11%&lt;/td&gt;
&lt;td&gt;6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 8% token overhead surprised me. I expected 25% or more. It turned out that the retries I eliminated were far more expensive than the research pass I added. A failed implementation plus three self-healing attempts costs more than reading four files.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Separate "understand" from "build" with tools, not prompts
&lt;/h3&gt;

&lt;p&gt;Telling a single agent "research first, then code" decays. It works until the context fills up, then the instruction loses to the task. Giving the research agent no write tools makes the separation physical. It can't drift into implementation because implementation isn't possible.&lt;/p&gt;

&lt;p&gt;This generalizes: &lt;strong&gt;if you want a behavior to be reliable, remove the ability to do the alternative&lt;/strong&gt; rather than asking nicely.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. The lockfile is the first document, not the last
&lt;/h3&gt;

&lt;p&gt;My worst week with this system was when the research agent read the latest web docs for a library, wrote a confident brief, and the implementer built against a major version we didn't have installed. Confidence was 0.95. The brief was completely correct for the wrong version.&lt;/p&gt;

&lt;p&gt;Now the agent's first action is always reading the lockfile or &lt;code&gt;pip freeze&lt;/code&gt; output, and the brief has a &lt;strong&gt;Staleness risk&lt;/strong&gt; field. Source in &lt;code&gt;node_modules&lt;/code&gt; beats docs on the web every time, because installed source cannot be out of date for the installed version.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Confidence scores are only useful if something acts on them
&lt;/h3&gt;

&lt;p&gt;I initially added the confidence field for my own reading. It was decorative. It became useful the day the orchestrator started gating on it. The moment a number has a consumer, the agent producing it gets noticeably more careful about calibration, because low confidence triggers visible re-work.&lt;/p&gt;

&lt;p&gt;Same lesson applies to any structured output from an agent. &lt;strong&gt;If nothing reads the field, delete it or wire it up.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Make disagreement a logged event
&lt;/h3&gt;

&lt;p&gt;"If you think the brief is wrong, stop and write why" turned a silent failure mode into data. In three months the implementer disagreed with a brief 23 times. It was right 9 times. Those 9 cases became test cases for the research agent's instructions. The 14 wrong disagreements were almost always the implementer wanting to use a newer API it "remembered" from training.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Skip research when the task is small, and be explicit about the threshold
&lt;/h3&gt;

&lt;p&gt;Not everything needs a brief. Renaming a variable, adjusting a CSS margin, bumping a copyright year. Running research on those wasted tokens and, worse, made the system feel slow for trivial work.&lt;/p&gt;

&lt;p&gt;I added a size heuristic: if the task touches fewer than 2 files and mentions no external library, the orchestrator writes a one-line auto-brief with confidence 1.0 and moves on. The threshold is dumb. Dumb thresholds that exist beat smart thresholds I never got around to building.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things I'm working on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Brief reuse across tasks.&lt;/strong&gt; Right now every task gets fresh research even if the last task answered the same question. I'm adding a simple index keyed on library plus version so briefs get reused for 7 days, then expire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Research as a first-class MCP server.&lt;/strong&gt; I want other tools to be able to ask "how does X work in this repo at this version" and get the same brief format back, without going through the orchestrator.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I also want to try the same split for &lt;strong&gt;debugging&lt;/strong&gt;: a read-only "diagnose" agent that produces a hypothesis brief before any fix agent touches the code. Early signs are promising, but I don't have enough data to write about it yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If your coding agent keeps inventing APIs, the fix probably isn't a better model or a longer prompt. It's a process boundary. Give one agent eyes and no hands, give the other hands and a contract, and make the orchestrator refuse to skip the first step.&lt;/p&gt;

&lt;p&gt;Versions for anyone reproducing this: Claude Code on the 2.x line as of September 2026, Node.js 22.x, Python 3.13, orchestrator in plain Bash on macOS.&lt;/p&gt;

&lt;p&gt;If this was useful, &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; for more build logs from running a fully autonomous implementation system 24/7. And if you've built a different kind of research or planning gate for your agents, I'd genuinely like to hear how it behaves in the comments. 💡&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>softwareengineering</category>
      <category>programming</category>
    </item>
    <item>
      <title>How I Built a Task Spec Contract Between My Planner and Implementer Agents</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Sun, 20 Sep 2026 14:32:42 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-built-a-task-spec-contract-between-my-planner-and-implementer-agents-e94</link>
      <guid>https://dev.to/yureki_lab/how-i-built-a-task-spec-contract-between-my-planner-and-implementer-agents-e94</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;My fully autonomous implementation system splits work between a &lt;strong&gt;planner agent&lt;/strong&gt; and a fleet of &lt;strong&gt;implementer sub-agents&lt;/strong&gt; that start with zero context. For months, roughly one in three implementer runs solved the wrong problem, even though the planner's prose instructions looked fine to me. The fix was not a better prompt. It was a &lt;strong&gt;task spec contract&lt;/strong&gt;: a small, structured schema the planner must emit and the implementer must echo back before touching code. Retries dropped from 31% to 7%, and I got a bonus: the same spec became the input for my verifier agent. Here is the schema, the war story that forced it, and five lessons about agent-to-agent handoffs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Quick background. The system I have been building for about a year runs Claude Code (2.x as of September 2026) in a loop: a planner reads the repo and the backlog, breaks a goal into tasks, and hands each task to a fresh implementer sub-agent. The implementer edits, runs tests, and reports back. A separate verifier agent reviews the diff. Humans (me) only see the summary.&lt;/p&gt;

&lt;p&gt;The planner and the implementers do &lt;strong&gt;not&lt;/strong&gt; share a context window. That is deliberate. Fresh context keeps each implementer cheap and focused, and it lets me run four or five of them in parallel. But it also means every task handoff is a &lt;strong&gt;cold start&lt;/strong&gt;. Whatever the planner forgets to write down simply does not exist for the implementer.&lt;/p&gt;

&lt;p&gt;Early on, the handoff was a paragraph of natural language. Something like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Add retry logic to the webhook sender so transient 5xx errors don't drop events. Keep it simple.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Reads fine, right? Here is what actually happened across a sample of 120 handoffs I logged in spring 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;%&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;✅ Done, verifier approved on first pass&lt;/td&gt;
&lt;td&gt;68&lt;/td&gt;
&lt;td&gt;57%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;⚠️ Done, but scope crept or wrong file touched&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;td&gt;31%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;❌ Gave up or produced nothing useful&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 31% row was the expensive one. Those tasks &lt;em&gt;looked&lt;/em&gt; done. The implementer wrote a confident summary, tests passed, and only the verifier (or worse, me, the next morning) noticed that "retry logic" had been implemented as a brand new generic retry utility with its own config file, exponential backoff, jitter, a circuit breaker, and 400 lines of tests. For a webhook sender that already had a retry helper two directories over.&lt;/p&gt;

&lt;h3&gt;
  
  
  The war story that broke me
&lt;/h3&gt;

&lt;p&gt;The specific incident that made me stop and redesign the handoff:&lt;/p&gt;

&lt;p&gt;The planner emitted a task: "Fix the flaky date parsing in the export job." The repo had two export jobs. One was a legacy CSV exporter that nobody had touched in a year. The other was the active JSON exporter that had the actual flaky test. The implementer grepped for "export", found the legacy one first, "fixed" its date parsing by rewriting it to a different library, updated its tests, and reported success. The verifier approved because the diff was internally consistent. The actual flaky test kept flaking for three more days.&lt;/p&gt;

&lt;p&gt;Nobody in that chain did anything &lt;em&gt;wrong&lt;/em&gt; given what they knew. The planner knew which exporter it meant. It just never said so, because to the planner it was obvious. That is the core failure mode of cold-start handoffs: &lt;strong&gt;the sender's obvious is the receiver's unknown.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;p&gt;I stopped treating the handoff as a message and started treating it as an &lt;strong&gt;interface&lt;/strong&gt;. If the planner and implementer were two services, I would never let them talk in free text. I would give them a schema. So I did.&lt;/p&gt;

&lt;h3&gt;
  
  
  The task spec schema
&lt;/h3&gt;

&lt;p&gt;Every task the planner emits must be a single fenced block that validates against this shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;task_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;T-2026-0914-03&lt;/span&gt;
&lt;span class="na"&gt;goal&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;Make the JSON export job's date parsing deterministic so&lt;/span&gt;
  &lt;span class="s"&gt;`test_export_dates_across_dst` stops flaking.&lt;/span&gt;
&lt;span class="na"&gt;why&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;The flaky test blocks CI ~2x/day; the root cause is naive&lt;/span&gt;
  &lt;span class="s"&gt;datetime handling around DST transitions.&lt;/span&gt;
&lt;span class="na"&gt;scope&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;allowed_paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;src/export/json_exporter.py&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;tests/export/test_json_exporter.py&lt;/span&gt;
  &lt;span class="na"&gt;forbidden_paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;src/export/csv_exporter.py&lt;/span&gt;   &lt;span class="c1"&gt;# legacy, do NOT touch&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;src/shared/**&lt;/span&gt;                &lt;span class="c1"&gt;# shared helpers need a separate task&lt;/span&gt;
&lt;span class="na"&gt;context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;There&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;an&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;existing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tz&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;helper&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;at&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;src/export/tz.py;&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;it,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;do&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;write&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;one."&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;test&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;currently&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fails&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;~30%&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;of&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;runs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;on&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;2026-03-08&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fixture."&lt;/span&gt;
&lt;span class="na"&gt;acceptance&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;cmd&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pytest&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tests/export/test_json_exporter.py&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;-x&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;--count=20"&lt;/span&gt;
    &lt;span class="na"&gt;expect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;20&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;runs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pass"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;cmd&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;git&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diff&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;--stat"&lt;/span&gt;
    &lt;span class="na"&gt;expect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;only&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;two&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;allowed_paths&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;appear"&lt;/span&gt;
&lt;span class="na"&gt;forbidden_moves&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Do&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;add&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;dependencies."&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Do&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;change&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;public&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;signature&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;of&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;export_json()."&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Do&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;create&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;new&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;utility&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;module."&lt;/span&gt;
&lt;span class="na"&gt;done_signal&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="s"&gt;Reply with the exact acceptance command outputs, then the&lt;/span&gt;
  &lt;span class="s"&gt;diff stat. If any acceptance check cannot pass, stop and&lt;/span&gt;
  &lt;span class="s"&gt;report which one and why. Do not work around it.&lt;/span&gt;
&lt;span class="na"&gt;budget&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;max_turns&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;40&lt;/span&gt;
  &lt;span class="na"&gt;max_minutes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;25&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few of these fields deserve explanation, because the obvious ones (goal, acceptance) are not where the value came from.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;scope.forbidden_paths&lt;/code&gt; mattered more than &lt;code&gt;allowed_paths&lt;/code&gt;.&lt;/strong&gt; Allowed paths tell the implementer where to look. Forbidden paths tell it where &lt;em&gt;not&lt;/em&gt; to fix things, which is exactly the information that was missing in the date parsing incident. I now make the planner write at least one forbidden path for every task, even if it has to think hard to find one. The act of choosing what to fence off forces the planner to surface the ambiguity it was carrying in its head.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;context&lt;/code&gt; is for facts the implementer cannot discover cheaply.&lt;/strong&gt; "There is already a tz helper" is a 10-second fact for the planner (it just read the repo) and a 15-minute discovery for a cold implementer, if it finds it at all. This field is where I put the things the planner "obviously knows."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;forbidden_moves&lt;/code&gt; is for behaviors, not files.&lt;/strong&gt; "Do not create a new utility module" kills the 400-line-retry-framework failure mode outright. My current top three forbidden moves, by how often they appear:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Do not create a new module/file when an existing one can be extended.&lt;/li&gt;
&lt;li&gt;Do not add a dependency.&lt;/li&gt;
&lt;li&gt;Do not "clean up" code adjacent to the task.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;done_signal&lt;/code&gt; defines what the implementer's final message must contain.&lt;/strong&gt; Before this, implementers reported success in prose. Now they must paste the acceptance command outputs verbatim. This one change made the verifier's job dramatically easier, because it could diff the claimed output against a re-run.&lt;/p&gt;

&lt;h3&gt;
  
  
  The echo-back step
&lt;/h3&gt;

&lt;p&gt;The schema alone got me from 31% to about 18% scope failures. The second half of the fix was a &lt;strong&gt;mandatory echo-back&lt;/strong&gt;: the implementer's first action, before reading a single file, is to restate the task in its own words in a fixed format.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Task echo&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; I will change: src/export/json_exporter.py, tests/export/test_json_exporter.py
&lt;span class="p"&gt;-&lt;/span&gt; I will NOT change: src/export/csv_exporter.py, anything under src/shared/
&lt;span class="p"&gt;-&lt;/span&gt; I am done when: 20 consecutive pytest runs pass AND diff stat shows only 2 files
&lt;span class="p"&gt;-&lt;/span&gt; Things I must not do: add deps, change export_json() signature, create new modules
&lt;span class="p"&gt;-&lt;/span&gt; Open questions: none
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the echo does not match the spec, a tiny check script rejects it and the implementer gets one retry. If &lt;code&gt;Open questions&lt;/code&gt; is non-empty, the task is bounced back to the planner instead of proceeding. This felt like ceremony when I added it. It turned out to be the single highest-leverage step, because roughly half the remaining scope failures were the implementer &lt;em&gt;misreading&lt;/em&gt; a correct spec, and the echo catches those for the cost of one short turn.&lt;/p&gt;

&lt;p&gt;Here is the flow now:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    P[Planner] --&amp;gt;|task spec YAML| V{Schema valid?}
    V -- no --&amp;gt; P
    V -- yes --&amp;gt; I[Implementer&amp;lt;br/&amp;gt;fresh context]
    I --&amp;gt;|task echo| E{Echo matches spec?}
    E -- no, 1 retry --&amp;gt; I
    E -- open questions --&amp;gt; P
    E -- yes --&amp;gt; W[Edit + run acceptance cmds]
    W --&amp;gt;|done_signal with raw outputs| R[Verifier]
    R --&amp;gt;|re-runs acceptance cmds| M{Match?}
    M -- yes --&amp;gt; Done[Merge]
    M -- no --&amp;gt; P&lt;/code&gt;&lt;/pre&gt;



&lt;h3&gt;
  
  
  The validation script
&lt;/h3&gt;

&lt;p&gt;The schema check is deliberately dumb. It is about 60 lines of Python and does not use an LLM. That matters: I want the gate to be deterministic and free.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;REQUIRED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;goal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acceptance&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;forbidden_moves&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done_signal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;budget&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;validate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;errors&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="n"&gt;missing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;REQUIRED&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;missing fields: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;scope&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{})&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allowed_paths&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope.allowed_paths must be non-empty&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;forbidden_paths&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scope.forbidden_paths must be non-empty (yes, really)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acceptance&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cmd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expect&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acceptance[&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;] needs both cmd and expect&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;forbidden_moves&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]))&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;at least one forbidden_move required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Yes, the &lt;code&gt;forbidden_paths must be non-empty&lt;/code&gt; rule is annoying for trivial tasks. I kept it anyway. A planner that cannot name one thing the implementer should not touch has not thought about the task hard enough.&lt;/p&gt;

&lt;h3&gt;
  
  
  Results
&lt;/h3&gt;

&lt;p&gt;Same 120-task sample size, measured in August 2026 after the contract had been in place for six weeks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Outcome&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;✅ Approved first pass&lt;/td&gt;
&lt;td&gt;57%&lt;/td&gt;
&lt;td&gt;81%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;⚠️ Scope creep / wrong target&lt;/td&gt;
&lt;td&gt;31%&lt;/td&gt;
&lt;td&gt;7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;❌ Gave up / nothing useful&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things I want to be honest about. First, the "gave up" row did not move. The contract fixes &lt;em&gt;misunderstanding&lt;/em&gt;, not &lt;em&gt;capability&lt;/em&gt;. When a task is genuinely too hard for a single implementer, a better spec does not save it. Second, planner cost went up about 20% per task, because writing a spec takes more tokens than writing a paragraph. That is paid back many times over by the retries I no longer run, but it is not free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Treat every cold-start handoff like an API boundary, because it is one.&lt;/strong&gt; You would never let two microservices exchange free-text and hope. Two agents with separate context windows are exactly that situation. Give them a schema, validate it, and reject bad payloads before they cause work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The sender's "obvious" is the receiver's "unknown." Design the schema to extract it.&lt;/strong&gt; The most valuable fields (&lt;code&gt;forbidden_paths&lt;/code&gt;, &lt;code&gt;context&lt;/code&gt;, &lt;code&gt;forbidden_moves&lt;/code&gt;) are all just structured ways of forcing the planner to write down what it was silently assuming. Free-form prompts let the planner skip that. A required field does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Negative space beats positive space.&lt;/strong&gt; Telling an agent what to do is table stakes. Telling it what &lt;em&gt;not&lt;/em&gt; to do is where the leverage is. An implementer with a clear goal and no fences will happily solve the goal in the most expansive way it can imagine. Fences make "the simplest thing that works" the path of least resistance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Make the receiver echo before it acts.&lt;/strong&gt; One short turn of restating the task catches misreads that would otherwise cost twenty turns of wrong work. It also gives you a clean place to surface open questions instead of letting the agent guess. This is the cheapest reliability upgrade I have ever added to the system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Define the shape of "done," not just the definition.&lt;/strong&gt; "Tests pass" is a definition. "Paste the raw output of these two commands" is a shape. Shapes can be checked mechanically by the next agent in the chain. Definitions require judgment, and judgment is where drift creeps in.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The spec has become the backbone of more than just the planner-implementer handoff. The verifier now reads the same YAML and re-runs &lt;code&gt;acceptance&lt;/code&gt; itself, so a task can only be marked done if two independent agents get the same outputs. I am working on two extensions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Spec diffs for retries.&lt;/strong&gt; When a task bounces back, the planner currently rewrites the whole spec. I want it to emit a diff against the previous spec so I can see &lt;em&gt;what it learned&lt;/em&gt; from the failure, which should feed my self-improvement loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget-aware decomposition.&lt;/strong&gt; The &lt;code&gt;budget&lt;/code&gt; field is currently just a kill switch. I want the planner to use its own past &lt;code&gt;budget&lt;/code&gt; overruns as a signal that a task should have been split in two.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I also want to try the echo-back pattern on the &lt;em&gt;human&lt;/em&gt; side. If I have to write a task echo for my own tickets, I suspect I will discover I am just as bad at stating scope as my planner used to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If you are running any kind of multi-agent setup with Claude Code, or even just handing tasks to a single fresh session, try the contract before you try a better prompt. Start with four fields: &lt;code&gt;goal&lt;/code&gt;, &lt;code&gt;forbidden_paths&lt;/code&gt;, &lt;code&gt;acceptance&lt;/code&gt; with real commands, and a &lt;code&gt;done_signal&lt;/code&gt; that demands raw output. Add the echo-back. Measure your retry rate before and after. I would genuinely like to hear whether your numbers look like mine.&lt;/p&gt;

&lt;p&gt;If this was useful, &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; 🚀. I write one post like this every week or so about building a fully autonomous implementation system, the stuff that broke, and what I changed. And if you have a handoff schema of your own, drop it in the comments. I am collecting them.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>softwareengineering</category>
      <category>programming</category>
    </item>
    <item>
      <title>How I Migrated 90 Cypress Tests to Playwright With Claude Code in 4 Days</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Sat, 19 Sep 2026 14:32:12 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-migrated-90-cypress-tests-to-playwright-with-claude-code-in-4-days-1im6</link>
      <guid>https://dev.to/yureki_lab/how-i-migrated-90-cypress-tests-to-playwright-with-claude-code-in-4-days-1im6</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I moved a 90-spec Cypress suite to Playwright in 4 working days using Claude Code. The trick wasn't "ask the AI to convert everything." It was: hand-translate &lt;strong&gt;one&lt;/strong&gt; spec, extract the pattern into a written rulebook, then let the agent fan out across the other 89 while a screenshot-diff gate caught the silent behaviour drift. Here's the workflow, the two traps that cost me a day, and what I'd do differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Our end-to-end suite was 90 Cypress specs, written over three years by six different people. It ran in about 38 minutes on CI, flaked roughly once every four runs, and had grown a 600-line &lt;code&gt;commands.js&lt;/code&gt; full of custom helpers that nobody fully understood anymore.&lt;/p&gt;

&lt;p&gt;We wanted Playwright for the usual reasons: real multi-tab support, first-class parallelism, and one runner for the three browsers we actually ship to. The problem was the migration itself. Every estimate I got from the team landed at "two sprints, maybe three." Nobody wanted to own it. It was pure toil with zero product upside until the very last spec crossed over.&lt;/p&gt;

&lt;p&gt;That's exactly the kind of work I've been throwing at Claude Code (v2.1.x at the time, Node.js 22, Playwright 1.54). Mechanical, high-volume, pattern-heavy, and easy to verify. So I gave myself one week and a rule: &lt;strong&gt;no spec gets deleted from Cypress until its Playwright twin passes 10 times in a row.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Translate one spec by hand, badly
&lt;/h3&gt;

&lt;p&gt;I picked a medium-complexity spec (the checkout flow: login, add to cart, apply coupon, pay with a test card) and translated it myself. No AI. About 90 minutes.&lt;/p&gt;

&lt;p&gt;It was ugly, but that was the point. Doing it by hand surfaced every decision that a converter would otherwise make silently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cypress &lt;code&gt;cy.get('[data-cy=x]')&lt;/code&gt; becomes &lt;code&gt;page.getByTestId('x')&lt;/code&gt;, which meant changing the test-id attribute in &lt;code&gt;playwright.config.ts&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Cypress's implicit retry-until-assert becomes an explicit &lt;code&gt;await expect(locator).toHaveText()&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Our &lt;code&gt;cy.login()&lt;/code&gt; custom command hits an API and sets a cookie. In Playwright that became a &lt;code&gt;storageState&lt;/code&gt; fixture created once per worker, not per test.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cy.intercept()&lt;/code&gt; stubs became &lt;code&gt;page.route()&lt;/code&gt;, but the matcher semantics differ (glob vs. minimatch), and I got two wrong on the first try.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of those became a line in a file I called &lt;code&gt;MIGRATION_RULES.md&lt;/code&gt;. By the end of spec one, it had 14 rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Turn the rulebook into agent instructions
&lt;/h3&gt;

&lt;p&gt;I put the rulebook into the project spec file that Claude Code reads on startup, along with the before/after of my hand-translated spec as a worked example. The key section looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Cypress -&amp;gt; Playwright migration rules&lt;/span&gt;
&lt;span class="p"&gt;
1.&lt;/span&gt; Never translate a custom command inline. Look it up in
   cypress/support/commands.js, then map it to the fixture
   in tests/fixtures/&lt;span class="err"&gt;*&lt;/span&gt;.ts. If no fixture exists, STOP and
   report which command is missing.
&lt;span class="p"&gt;2.&lt;/span&gt; Every cy.get(...).should(...) chain becomes a single
   &lt;span class="sb"&gt;`await expect(locator).toX()`&lt;/span&gt;. Do not add manual waits.
&lt;span class="p"&gt;3.&lt;/span&gt; cy.intercept(method, urlGlob) -&amp;gt; page.route(urlGlob).
   Convert Cypress glob to Playwright glob; verify with the
   table in MIGRATION_RULES.md section 3.
&lt;span class="p"&gt;4.&lt;/span&gt; Keep the original spec's describe/it names verbatim so
   the test report diff is readable.
&lt;span class="p"&gt;5.&lt;/span&gt; After writing a spec, run it 3x with &lt;span class="sb"&gt;`--repeat-each=3`&lt;/span&gt;.
   Only report success if all 3 pass.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rule 1 turned out to be the most important line I wrote all week. More on that below.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Fan out, in batches of five
&lt;/h3&gt;

&lt;p&gt;I didn't run one giant "migrate all 89" session. I ran batches of five specs, each in a fresh session, with a prompt like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Migrate these 5 Cypress specs to Playwright following the rules
in CLAUDE.md. Work on them one at a time. For each: write the
spec, run it 3x, then show me the pass/fail summary before
moving to the next. Do not touch any file outside tests/e2e/.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why five? Three reasons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context stays sharp.&lt;/strong&gt; After roughly 8 to 10 specs in one session, the agent started "remembering" helper names that didn't exist. Fresh sessions fixed that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failures stay cheap.&lt;/strong&gt; If batch 7 went sideways, I lost 20 minutes, not a morning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I could review at human speed.&lt;/strong&gt; Five diffs is a coffee break. Eighty-nine is a weekend.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's the flow end to end:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Pick 5 Cypress specs] --&amp;gt; B[Agent translates spec N]
    B --&amp;gt; C{3x pass?}
    C -- no --&amp;gt; D[Agent reports failure + reason]
    D --&amp;gt; E[I fix rule or fixture]
    E --&amp;gt; B
    C -- yes --&amp;gt; F[Screenshot diff vs Cypress run]
    F -- drift --&amp;gt; E
    F -- clean --&amp;gt; G[Mark spec migrated]
    G --&amp;gt; B&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Batches 1 through 4 went through in about a day and a half. Twenty specs, no drama. Then I hit the first trap.&lt;/p&gt;

&lt;h3&gt;
  
  
  Trap 1: The custom command that lied
&lt;/h3&gt;

&lt;p&gt;Batch 5 included a spec that called &lt;code&gt;cy.selectPlan('pro')&lt;/code&gt;. The agent, following rule 1, looked it up. The command did three things: clicked a plan card, waited for a modal, and &lt;strong&gt;silently dismissed a "confirm downgrade" dialog if it appeared.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That third behaviour was never documented. It existed because two years ago someone had a flaky test and patched the helper instead of the app. The Playwright fixture the agent wrote did the first two things, and the test passed. It passed because the test data never triggered the downgrade dialog.&lt;/p&gt;

&lt;p&gt;I only caught it because rule 1 also said "report which command is missing," and the agent's report included a one-liner: &lt;em&gt;"Note: the Cypress version also handles a confirm dialog conditionally; I did not port this since no test exercises it."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That single sentence saved a production bug. It turned out our real app &lt;strong&gt;did&lt;/strong&gt; show that dialog for one plan transition, and the old Cypress helper had been hiding a broken flow for two years.&lt;/p&gt;

&lt;p&gt;Lesson: when the agent says "I skipped this, here's why," read it. Every time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Trap 2: Green tests that tested nothing
&lt;/h3&gt;

&lt;p&gt;Around batch 9, the pass rate got suspiciously good. Every spec passed 3 of 3 on the first attempt. I got nervous and pulled up a diff.&lt;/p&gt;

&lt;p&gt;The agent had discovered that Playwright's auto-waiting &lt;code&gt;expect()&lt;/code&gt; is &lt;em&gt;very&lt;/em&gt; forgiving, and had translated a Cypress assertion like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;cy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;[data-cy=total]&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;should&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;contain&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;$49.00&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;into this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getByTestId&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;total&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toBeVisible&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Visible. Not "contains $49.00." The element was always visible. The test could never fail.&lt;/p&gt;

&lt;p&gt;This wasn't the agent being lazy. It was me being vague. My rule 2 said "becomes a single expect" but never said "preserve the assertion's semantics exactly." Fixing that took one line in the rulebook and a re-run of 11 specs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: The screenshot-diff gate
&lt;/h3&gt;

&lt;p&gt;This is the part I'd keep even if I never migrate another test suite.&lt;/p&gt;

&lt;p&gt;Before touching anything, I ran the full Cypress suite once with screenshots on at every &lt;code&gt;it()&lt;/code&gt; boundary, and saved them to &lt;code&gt;baseline/&lt;/code&gt;. For each migrated Playwright spec, I captured a screenshot at the same boundaries and diffed with &lt;code&gt;pixelmatch&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;test&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expect&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;@playwright/test&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;compareToBaseline&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;../fixtures/visual&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;checkout: coupon applied&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// ... steps ...&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;compareToBaseline&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;checkout-coupon-applied&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.02&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The diff caught four cases where the Playwright test passed, the assertions passed, but the screen looked different. Three were timing (Playwright was faster and captured mid-animation). One was real: a locator matched a different button with the same label, so the test was exercising the wrong path and getting lucky.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Do the first one by hand, always.&lt;/strong&gt; The 90 minutes I spent translating spec one produced the 14 rules that made the other 89 possible. If you skip this, the agent makes every one of those decisions for you, silently and inconsistently.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;"Stop and report" beats "figure it out."&lt;/strong&gt; The single highest-value instruction was telling the agent to halt on unknown helpers instead of guessing. Every real bug I caught came through one of those reports.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Batch small, session fresh.&lt;/strong&gt; Five specs per session was the sweet spot. Bigger batches got hallucinated helper names. Smaller ones wasted my review time on context-switching.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Passing tests are not evidence. Failing-when-they-should tests are.&lt;/strong&gt; After trap 2, I added a step: for every migrated spec, deliberately break the app once and confirm the test goes red. The agent can do this too, and it takes 30 seconds per spec.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Visual diffs are the cheapest oracle you have.&lt;/strong&gt; Assertions encode what someone remembered to check. Screenshots encode everything. For a migration, "does it look the same" catches a class of drift that no assertion will.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The suite now runs in 11 minutes across 4 workers, down from 38, and I haven't seen a flake in three weeks. The &lt;code&gt;commands.js&lt;/code&gt; file is gone. The 600 lines became about 180 lines of typed fixtures that a new hire can actually read.&lt;/p&gt;

&lt;p&gt;Next up, I'm pointing the same "hand-translate one, extract rules, fan out" workflow at our 40-odd Jest snapshot tests, which have the same "green but meaningless" problem in a different costume. I'm also experimenting with having the agent write the &lt;em&gt;deliberate break&lt;/em&gt; step itself, so the "does this test actually fail" check becomes part of the migration loop instead of something I remember to do.&lt;/p&gt;

&lt;p&gt;If there's interest, I'll write up the fixture design in a follow-up. The &lt;code&gt;storageState&lt;/code&gt;-per-worker pattern alone cut our login overhead by 80%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If you're staring at a test-suite migration that nobody wants to own, this is the workflow I'd hand you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Translate one spec yourself. Write down every decision.&lt;/li&gt;
&lt;li&gt;Turn those decisions into rules the agent can't misread.&lt;/li&gt;
&lt;li&gt;Fan out in small batches. Make the agent report what it skipped.&lt;/li&gt;
&lt;li&gt;Verify with something other than the tests themselves.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Got a migration war story of your own, or a rule that saved you a bad afternoon? Drop it in the comments. And if you want the follow-up on fixture design, hit &lt;strong&gt;follow&lt;/strong&gt; so it lands in your feed. 🚀&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>playwright</category>
      <category>testing</category>
    </item>
    <item>
      <title>How I Built a Remote Control Dashboard for My Autonomous Coding Agent</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Fri, 18 Sep 2026 14:31:57 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-built-a-remote-control-dashboard-for-my-autonomous-coding-agent-gf8</link>
      <guid>https://dev.to/yureki_lab/how-i-built-a-remote-control-dashboard-for-my-autonomous-coding-agent-gf8</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I run a fully autonomous implementation system 24/7 on a Mac mini. For months, the only way to intervene was to SSH in and kill processes by hand. So I built a small remote control dashboard I can open on my phone to pause, steer, approve, and inspect the agent from anywhere. This post walks through the design (a file-based command queue, an approval gate, and a heartbeat), the code that matters, and five lessons from six months of running it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Here's the setup: an orchestrator module spawns parallel implementation agents built on Claude Code (v2.1.x at the time of writing), each working on a task from a queue. It plans, writes code, runs tests, commits, and moves on. It runs while I sleep. It runs while I'm at the gym. It runs while I'm on a train with no laptop.&lt;/p&gt;

&lt;p&gt;That last part was the issue. 🚨&lt;/p&gt;

&lt;p&gt;Three things kept happening:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The agent would go down a rabbit hole.&lt;/strong&gt; A "rename this config key" task would turn into a 40-file refactor because the agent decided the old name was confusing everywhere. I'd find out 6 hours later.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It would hit a decision it shouldn't make alone.&lt;/strong&gt; Delete a migration? Push to a shared branch? Rotate a credential? The agent was told to stop and wait for a human on those. But "wait for a human" meant "wait until I'm at my desk".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I had no idea what it was doing right now.&lt;/strong&gt; Logs were on the box. Reading them meant a terminal, a VPN, and &lt;code&gt;tail -f&lt;/code&gt;. Not a thing you do from a phone at dinner.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;My first instinct was "just SSH from the phone". I tried it for two weeks. Typing &lt;code&gt;kill -9&lt;/code&gt; on a 6-inch screen at 2am while half-asleep is a great way to kill the wrong process. I did that. Twice.&lt;/p&gt;

&lt;p&gt;What I actually needed was a &lt;strong&gt;remote control&lt;/strong&gt;, not a remote shell. A handful of big, safe buttons and a clear status readout.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Design constraints
&lt;/h3&gt;

&lt;p&gt;I set three rules before writing any code:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The agent must not depend on the dashboard.&lt;/strong&gt; If the dashboard is down, the agent keeps working. If the agent is down, the dashboard says so. No shared process, no shared database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every command must be idempotent and safe to replay.&lt;/strong&gt; Mobile networks retry. Tapping "pause" twice must not do anything weird.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read is cheap, write is gated.&lt;/strong&gt; Anyone with the link can see status (it's behind auth anyway). Writing a command requires a second factor.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Architecture
&lt;/h3&gt;

&lt;p&gt;The whole thing is three parts talking through the filesystem:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    Phone[Phone browser] --&amp;gt;|HTTPS| Dash[Dashboard server\nFastAPI]
    Dash --&amp;gt;|writes JSON| Queue[(commands/ dir)]
    Dash --&amp;gt;|reads| Status[(status.json + logs)]
    Agent[Autonomous agent loop] --&amp;gt;|polls every 5s| Queue
    Agent --&amp;gt;|writes every 30s| Status&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Yes, a directory of JSON files as a queue. No Redis, no Postgres, no message broker. I'll defend that choice in the lessons section.&lt;/p&gt;

&lt;h3&gt;
  
  
  The command queue
&lt;/h3&gt;

&lt;p&gt;Each command is a single file. The filename is a UUID; the content is a tiny JSON document. The agent's main loop polls the directory between task steps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# agent side: runs between every tool call / task step
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Path&lt;/span&gt;

&lt;span class="n"&gt;COMMANDS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;~/agent/commands&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;expanduser&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;PROCESSED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;COMMANDS&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;processed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;drain_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;COMMANDS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read_text&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JSONDecodeError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PROCESSED&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.corrupt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kind&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pause&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;paused&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resume&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;paused&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;abort_task&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;abort_current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;steer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="c1"&gt;# appended to the next prompt as a user instruction
&lt;/span&gt;            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pending_notes&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approve&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approvals&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deny&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;denials&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

        &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PROCESSED&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;steer&lt;/code&gt; command is the one I use most. It doesn't interrupt anything. It just says: "before your next step, read this note from me". Things like &lt;em&gt;"Don't touch the billing module, I'm working on it locally"&lt;/em&gt; or &lt;em&gt;"The flaky test is known, skip it and keep going"&lt;/em&gt;. The orchestrator injects it into the next prompt as a high-priority user message.&lt;/p&gt;

&lt;h3&gt;
  
  
  The approval gate
&lt;/h3&gt;

&lt;p&gt;This is the part that changed how I sleep. 😴&lt;/p&gt;

&lt;p&gt;The agent has a list of actions it is &lt;strong&gt;not allowed to do without a human&lt;/strong&gt;. Deleting files outside the repo, force-pushing, running migrations against anything but a local database, touching secrets. When it hits one, it doesn't fail and it doesn't guess. It writes an approval request and blocks on that request ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# agent side
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;request_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout_s&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;3600&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;req_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nb"&gt;hex&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;REQUESTS&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;req_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;req_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;detail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;}))&lt;/span&gt;
    &lt;span class="nf"&gt;notify_phone&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Approval needed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# push notification
&lt;/span&gt;    &lt;span class="n"&gt;deadline&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;timeout_s&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;deadline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;drain_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;req_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;approvals&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;req_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;denials&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;   &lt;span class="c1"&gt;# timeout = deny, always
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Timeout equals deny. Always. I went back and forth on this. An agent that treats silence as consent is an agent that will eventually do something irreversible at 3am because I was asleep. Silence means "not now", and the task gets parked, not dropped.&lt;/p&gt;

&lt;p&gt;On the dashboard, an approval request renders as a card with the action, the detail (usually a diff or a command), and two big buttons. Green and red. That's it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The status heartbeat
&lt;/h3&gt;

&lt;p&gt;The agent writes a &lt;code&gt;status.json&lt;/code&gt; every 30 seconds and after every task boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1789740000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"running"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"current_task"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Add retry to webhook delivery"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"step"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tokens_today"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;812000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"last_commit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"a1f3c9e"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"pending_approvals"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"agents_active"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dashboard reads it and does one thing I'd underrate if I hadn't lived without it: &lt;strong&gt;it shows how stale the heartbeat is.&lt;/strong&gt; If &lt;code&gt;ts&lt;/code&gt; is more than 90 seconds old, the header goes amber. More than 5 minutes, red. That single indicator caught two hung processes and one out-of-disk incident before anything else did.&lt;/p&gt;

&lt;h3&gt;
  
  
  The dashboard itself
&lt;/h3&gt;

&lt;p&gt;Roughly 300 lines of FastAPI (Python 3.13) plus one HTML template with vanilla JavaScript. No framework. It fits on a phone because I designed it on a phone first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Top&lt;/strong&gt;: state badge, heartbeat age, current task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Middle&lt;/strong&gt;: pending approval cards (if any).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bottom&lt;/strong&gt;: four buttons. Pause, Resume, Abort current, Steer (opens a text box).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Last 50 log lines&lt;/strong&gt; behind a collapsible section.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Writes go through a one-time code from an authenticator app. Reads are just behind the reverse proxy's basic auth. The server itself runs as a separate background service so it survives agent restarts, and it never imports anything from the agent codebase.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/cmd/{kind}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;post_command&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;CommandIn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;totp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Header&lt;/span&gt;&lt;span class="p"&gt;(...)):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ALLOWED_KINDS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;verify_totp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;totp&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;403&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kind&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;model_dump&lt;/span&gt;&lt;span class="p"&gt;()}&lt;/span&gt;
    &lt;span class="n"&gt;tmp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;COMMANDS&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nb"&gt;hex&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.tmp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;tmp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;tmp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;rename&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;COMMANDS&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nb"&gt;hex&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# atomic
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ok&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Write to a temp file, then rename. The agent never sees a half-written command.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. A directory of JSON files beats a queue you have to keep alive
&lt;/h3&gt;

&lt;p&gt;I got roasted for this in a Discord. Fine. Here's the thing: the queue has zero dependencies, survives reboots, is inspectable with &lt;code&gt;ls&lt;/code&gt;, and debuggable with &lt;code&gt;cat&lt;/code&gt;. In six months it has never been the failing component. Every "real" queue I've run has needed babysitting at some point. For a single-machine, single-consumer, low-volume control channel, files win. 💡&lt;/p&gt;

&lt;h3&gt;
  
  
  2. "Steer" is worth more than "stop"
&lt;/h3&gt;

&lt;p&gt;I built pause and abort first because they felt like the safety features. I use steer ten times more often. Most interventions aren't "stop everything", they're "you're missing context, here it is, keep going". Giving the agent a way to receive mid-task notes without losing its place turned a lot of would-be aborts into small corrections.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Silence must mean no
&lt;/h3&gt;

&lt;p&gt;Any autonomous system with a human-in-the-loop gate will eventually run into the human being unavailable. Decide up front what happens. The answer that lets you sleep is "park the task and move on". The answer that ends your weekend is "assume yes after N minutes".&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Heartbeat age is the most important pixel on the screen
&lt;/h3&gt;

&lt;p&gt;Not the log. Not the task name. The number of seconds since the agent last said "I'm alive". Everything else can be wrong or stale; that number tells you whether to trust the rest of the screen.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Design the control surface for your dumbest future self
&lt;/h3&gt;

&lt;p&gt;The person using this dashboard is me at 2am, on a phone, with one eye open. Big buttons. Confirmation on anything destructive. No text input required for the common path. The agent is sophisticated; the remote control should be boring. ⚠️&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things I'm working on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Approval bundling.&lt;/strong&gt; Right now each gated action is its own request. When the agent is migrating something, I get five approvals in a row. I want it to batch related requests into one card with one decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read-only sharing.&lt;/strong&gt; A teammate wants to watch the agent work on a shared repo without being able to steer it. That means separating the read and write auth properly instead of leaning on the reverse proxy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Longer term, I want the dashboard to show &lt;em&gt;why&lt;/em&gt; the agent is doing what it's doing, not just &lt;em&gt;what&lt;/em&gt;. A short "current reasoning" field in the heartbeat, written by the agent in one sentence, would go a long way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;If you're running any kind of long-lived AI coding agent and your intervention story is "SSH in and kill it", build the remote control. It's a weekend of work. The command queue is 40 lines, the approval gate is 30, the dashboard is an afternoon. You'll never go back to &lt;code&gt;tail -f&lt;/code&gt; on a phone.&lt;/p&gt;

&lt;p&gt;If this was useful, &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; 🚀 — I'm writing up the rest of this autonomous system piece by piece: the orchestrator, the self-healing agent, and the observability layer. And if you've built something similar, tell me in the comments what your "steer" equivalent looks like. I want to steal your ideas.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>showdev</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>How I Audited 1,400 npm Dependencies for License Risks With Claude Code</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:32:28 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-audited-1400-npm-dependencies-for-license-risks-with-claude-code-57bp</link>
      <guid>https://dev.to/yureki_lab/how-i-audited-1400-npm-dependencies-for-license-risks-with-claude-code-57bp</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;A customer's procurement team asked us for a license inventory of our product, and I discovered our npm lockfile contained 1,400+ dependencies nobody had ever audited. I used Claude Code to classify every license, dig into the sketchy ones, trace how the risky packages got into the tree, and wire a CI gate so this never silently regresses. Total time: about two days instead of the two weeks I'd budgeted. 💡&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Here's an email you never want to receive: &lt;em&gt;"Before we can renew, our legal team needs a complete inventory of third-party licenses in your product, including transitive dependencies."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Our &lt;code&gt;package.json&lt;/code&gt; listed 62 direct dependencies. Innocent enough. Then I ran:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;--all&lt;/span&gt; &lt;span class="nt"&gt;--parseable&lt;/span&gt; | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
&lt;span class="c"&gt;# 1417&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;1,417 packages.&lt;/strong&gt; Every single one of them ships with a license, and every single one of those licenses is a tiny legal contract we had implicitly agreed to. Nobody on the team — including me — had ever read more than a handful of them.&lt;/p&gt;

&lt;p&gt;The scary part isn't the well-behaved majority. Something like 90% of the npm ecosystem is MIT/ISC/Apache-2.0 and genuinely fine. The scary part is the tail:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Copyleft licenses&lt;/strong&gt; (GPL, AGPL) hiding three levels deep in the tree&lt;/li&gt;
&lt;li&gt;Packages whose &lt;code&gt;license&lt;/code&gt; field says one thing while the bundled &lt;code&gt;LICENSE&lt;/code&gt; file says another&lt;/li&gt;
&lt;li&gt;Packages with &lt;code&gt;"license": "SEE LICENSE IN LICENSE.txt"&lt;/code&gt; — which tells you exactly nothing&lt;/li&gt;
&lt;li&gt;Packages with no license metadata at all, which legally defaults to "all rights reserved"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Reading 1,400 license files by hand is exactly the kind of soul-crushing, judgment-requiring-but-barely, high-volume work that I now refuse to do myself. So I opened Claude Code (v2.x, on Node.js 22.x) and made it a compliance intern for two days.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Get the raw inventory mechanically
&lt;/h3&gt;

&lt;p&gt;First rule: don't make the AI do work a deterministic tool does better. The initial sweep came from &lt;code&gt;license-checker-rss&lt;/code&gt;, which reads the metadata for the whole installed tree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx license-checker-rss &lt;span class="nt"&gt;--json&lt;/span&gt; &lt;span class="nt"&gt;--production&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; licenses.json
jq &lt;span class="s1"&gt;'length'&lt;/span&gt; licenses.json
&lt;span class="c"&gt;# 1417&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I asked Claude Code to summarize the distribution before doing anything fancy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Read licenses.json. Group packages by license identifier and give me
a count per license, sorted descending. Flag anything that is not
MIT, ISC, BSD-2-Clause, BSD-3-Clause, Apache-2.0, or 0BSD.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The distribution came back looking like most JS projects:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bucket&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;MIT / ISC / BSD / Apache-2.0&lt;/td&gt;
&lt;td&gt;1,361&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MPL-2.0, LGPL (weak copyleft)&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPL family flagged&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;UNKNOWN&lt;/code&gt;, &lt;code&gt;SEE LICENSE IN&lt;/code&gt;, custom&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So 1,361 packages needed zero further attention, and 56 packages needed actual eyes. That's the whole game: &lt;strong&gt;use cheap tools to shrink the haystack, use the agent on what's left.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Let the agent read the weird ones
&lt;/h3&gt;

&lt;p&gt;The 33 "unknown/custom" packages are where metadata-only scanners give up — and where an agent that can read files shines. Every one of those packages has &lt;em&gt;something&lt;/em&gt; in &lt;code&gt;node_modules&lt;/code&gt;: a &lt;code&gt;LICENSE&lt;/code&gt; file, a license section in the README, a header comment in the source.&lt;/p&gt;

&lt;p&gt;I gave Claude Code a classification rubric up front, in the project's instruction file, so it wouldn't improvise legal opinions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## License classification rules&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Classify as exactly one of: PERMISSIVE / WEAK_COPYLEFT /
  STRONG_COPYLEFT / PROPRIETARY / CANNOT_DETERMINE
&lt;span class="p"&gt;-&lt;/span&gt; Quote the exact sentence from the license text that justifies
  the classification. No quote, no classification.
&lt;span class="p"&gt;-&lt;/span&gt; If the package.json license field and the LICENSE file disagree,
  report BOTH and classify by the LICENSE file.
&lt;span class="p"&gt;-&lt;/span&gt; Never guess. CANNOT_DETERMINE is a valid and welcome answer.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the sweep itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;For each package in unknowns.txt, read its files under node_modules/,
find the actual license text, and classify it per the rules in
CLAUDE.md. Output one JSON line per package: name, version,
classification, evidence quote, file path.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last rule — &lt;em&gt;quote the sentence or don't classify&lt;/em&gt; — is the single most important line. On my first attempt without it, the agent confidently labeled a package MIT because the README "looked standard." With the evidence-quote requirement, every classification came back verifiable, and I spot-checked them in minutes instead of re-reading everything.&lt;/p&gt;

&lt;p&gt;Results from the 33 unknowns: 26 were permissive licenses with lazy metadata, 4 were custom-but-clearly-permissive texts, 2 were &lt;code&gt;CANNOT_DETERMINE&lt;/code&gt; (dead packages with no license text anywhere — we replaced them), and 1 was a genuine surprise…&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: The AGPL in the basement 😱
&lt;/h3&gt;

&lt;p&gt;One of the four GPL-family flags turned out to be &lt;strong&gt;AGPL-3.0&lt;/strong&gt;, sitting four levels deep under a charting library. If you're not familiar: AGPL is the license where even &lt;em&gt;network use&lt;/em&gt; can trigger source-disclosure obligations. For a proprietary SaaS product, that's not a footnote — that's a drop-everything finding.&lt;/p&gt;

&lt;p&gt;The immediate question was: &lt;em&gt;how did this get here, and can it go away?&lt;/em&gt; This is where &lt;code&gt;npm why&lt;/code&gt; plus an agent that can read changelogs saved me hours:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm why aggressively-licensed-package
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Ask: which of our direct dependencies ultimately pulls this in,
is it actually imported at runtime or only used at build time,
and does a newer version of the parent drop this dependency?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Claude Code traced the chain, checked our bundle output to confirm the package actually shipped to production (it did — build-time-only would have been a much easier conversation), and then found that the &lt;em&gt;parent&lt;/em&gt; library had replaced this dependency two major versions ago, precisely because of the license. The fix was a scheduled upgrade we'd been putting off anyway. One &lt;code&gt;npm install&lt;/code&gt; later, the AGPL was gone.&lt;/p&gt;

&lt;p&gt;I want to be honest: an experienced engineer with &lt;code&gt;npm why&lt;/code&gt; and patience would have found the same answer. The agent didn't do anything superhuman. It just did 45 minutes of archaeology in 4, and it read the parent library's changelog so I didn't have to.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Make it impossible to regress
&lt;/h3&gt;

&lt;p&gt;An audit that runs once is a snapshot, not a control. The lockfile changes every week; next month's &lt;code&gt;npm install&lt;/code&gt; can quietly reintroduce the exact problem. So the last step was a CI gate — and this part I had Claude Code write end to end:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// scripts/check-licenses.mjs — runs in CI on every lockfile change&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;execSync&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:child_process&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ALLOWED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;MIT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ISC&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;0BSD&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;BSD-2-Clause&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;BSD-3-Clause&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Apache-2.0&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CC0-1.0&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Unlicense&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="c1"&gt;// Reviewed one-by-one; each entry links to the review note.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;EXCEPTIONS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
  &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;some-mpl-package@4.2.1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;MPL-2.0 ok: unmodified, dynamically linked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;]);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;execSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;npx license-checker-rss --json --production&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tree&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;violations&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;entries&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tree&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;filter&lt;/span&gt;&lt;span class="p"&gt;(([&lt;/span&gt;&lt;span class="nx"&gt;pkg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;info&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;license&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;info&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;licenses&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UNKNOWN&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;ALLOWED&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;license&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;EXCEPTIONS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pkg&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;violations&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;License gate failed for:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;pkg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;info&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;violations&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`  &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;pkg&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; → &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;info&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;licenses&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`License gate passed: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nb"&gt;Object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tree&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; packages checked ✅`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The design choice that matters: &lt;strong&gt;it's an allowlist, not a blocklist.&lt;/strong&gt; A blocklist of "bad" licenses silently passes anything new or weird. An allowlist fails closed — a never-before-seen license identifier stops CI and forces a human (or an agent with a rubric) to look at it once, add it to &lt;code&gt;ALLOWED&lt;/code&gt; or &lt;code&gt;EXCEPTIONS&lt;/code&gt; with a written reason, and move on.&lt;/p&gt;

&lt;p&gt;The whole flow now looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[lockfile change] --&amp;gt; B[license-checker in CI]
    B --&amp;gt; C{all in allowlist?}
    C -- yes --&amp;gt; D[merge ✅]
    C -- no --&amp;gt; E[CI fails]
    E --&amp;gt; F[agent-assisted review]
    F --&amp;gt; G[allowlist or replace]
    G --&amp;gt; D&lt;/code&gt;&lt;/pre&gt;



&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Shrink the haystack with dumb tools before you deploy the smart one.&lt;/strong&gt; Metadata scanning removed 96% of the work for free. Pointing an LLM at all 1,417 packages would have been slower, pricier, and noisier. The agent earns its keep on the ambiguous residue, not the bulk.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Demand evidence quotes, not conclusions.&lt;/strong&gt; "Classify this license" gets you plausible answers. "Classify it and quote the sentence that proves it" gets you &lt;em&gt;checkable&lt;/em&gt; answers. This one prompt rule turned spot-checking from re-doing the work into skimming it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;code&gt;CANNOT_DETERMINE&lt;/code&gt; is a feature.&lt;/strong&gt; Explicitly telling the agent that "I don't know" is a welcome answer is the difference between an audit and a hallucination generator. The two packages it refused to classify were exactly the two that deserved human escalation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The license field lies more often than you'd think.&lt;/strong&gt; Out of 33 manually-reviewed packages, several had metadata that disagreed with the shipped license text. If your compliance story is built purely on &lt;code&gt;package.json&lt;/code&gt; fields, it's built on the honor system.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;An audit without a CI gate is theater.&lt;/strong&gt; The one-time sweep satisfied procurement. The allowlist gate is what makes the answer stay true. Fail closed on anything unrecognized, and make every exception carry a written justification.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;⚠️ Obligatory disclaimer: I'm an engineer, not a lawyer, and neither is Claude. The agent did the reading, sorting, and evidence-gathering; the actual risk calls on the flagged packages went through humans (and for the AGPL one, actual counsel). Use AI to prepare the decision, not to make it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things on my list: generating a proper SBOM (CycloneDX format) from the same pipeline so the next procurement request is a one-command answer, and running the identical playbook against our Python service — &lt;code&gt;pip&lt;/code&gt; has its own flavor of license chaos, and I suspect its tail is even weirder than npm's.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;Dependency license auditing is the perfect AI-agent workload: massive volume, mostly mechanical, occasionally requiring real judgment — and an agent with file access plus a strict rubric handles the volume while surfacing exactly the cases that need your brain.&lt;/p&gt;

&lt;p&gt;If you've never looked at your own lockfile's license tail, run &lt;code&gt;npx license-checker-rss --summary&lt;/code&gt; today. It takes thirty seconds, and you might meet your own AGPL in the basement.&lt;/p&gt;

&lt;p&gt;Have you found something scary in your dependency tree? Tell me your worst license surprise in the comments — and &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; 🚀 for more write-ups on putting coding agents to work on the jobs nobody wants.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>opensource</category>
      <category>node</category>
    </item>
    <item>
      <title>How I Built a Slack Bot That Answers Codebase Questions With Claude</title>
      <dc:creator>yureki_lab</dc:creator>
      <pubDate>Wed, 16 Sep 2026 14:32:44 +0000</pubDate>
      <link>https://dev.to/yureki_lab/how-i-built-a-slack-bot-that-answers-codebase-questions-with-claude-apa</link>
      <guid>https://dev.to/yureki_lab/how-i-built-a-slack-bot-that-answers-codebase-questions-with-claude-apa</guid>
      <description>&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;I built a Slack bot that answers "how does X work?" and "where does Y live?" questions about our private codebase, using Claude with tool-calling retrieval instead of embedding-based RAG. It took about two weekends, it answers most questions in under 30 seconds, and the biggest engineering problems had nothing to do with AI — they were about trust, staleness, and permissions. Here's the full build log, including the parts that went wrong. 🚀&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Every team has that one person who knows where everything is. On my projects, that person is me — and I was tired of being a human &lt;code&gt;grep&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The questions were always the same shape:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Where do we validate webhook signatures?"&lt;/li&gt;
&lt;li&gt;"How does the retry logic work in the billing worker?"&lt;/li&gt;
&lt;li&gt;"Is there already a helper for parsing these date formats?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each one costs the asker 20 minutes of digging or costs me a context switch. Multiply that by a few questions a day and you're losing real engineering time to &lt;em&gt;archaeology&lt;/em&gt;, not building.&lt;/p&gt;

&lt;p&gt;The obvious answer in 2026 is "point an LLM at the codebase." But the obvious &lt;em&gt;implementation&lt;/em&gt; — chunk the repo, embed it, cosine-similarity your way to an answer — is where I started, and it's the first thing I threw away.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Solved It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The architecture that didn't work: embedding RAG
&lt;/h3&gt;

&lt;p&gt;My first version was textbook RAG: split every file into 500-token chunks, embed them, stuff the top-k matches into a prompt.&lt;/p&gt;

&lt;p&gt;It failed in a very specific way. Code questions are usually &lt;strong&gt;structural&lt;/strong&gt;, not &lt;strong&gt;topical&lt;/strong&gt;. When someone asks "how does retry logic work in the billing worker," the answer isn't in any single chunk. It's spread across a config file, a decorator definition, and the call site — three files that don't share much vocabulary. Embedding similarity finds you five chunks that all &lt;em&gt;mention&lt;/em&gt; the word "retry" (including two from tests and one from a changelog) and none of the ones that matter.&lt;/p&gt;

&lt;p&gt;I spent a week tuning chunk sizes and rerankers before admitting the approach was wrong for this job.&lt;/p&gt;

&lt;h3&gt;
  
  
  The architecture that worked: let the model drive the tools
&lt;/h3&gt;

&lt;p&gt;The fix was to stop pre-deciding what context the model needs and instead give it the same tools I would use: search and read. Claude decides what to look for, reads what it finds, and follows the trail — the same loop a human does, just faster.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    A[Slack mention] --&amp;gt; B[Bot server]
    B --&amp;gt; C{Claude agent loop}
    C --&amp;gt;|search_code| D[ripgrep over repo]
    C --&amp;gt;|read_file| E[File reader]
    D --&amp;gt; C
    E --&amp;gt; C
    C --&amp;gt;|final answer| F[Slack thread reply]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The whole thing is a Slack Bolt app (Python 3.13, &lt;code&gt;slack-bolt&lt;/code&gt; 1.21) plus the Anthropic SDK. Two tools, that's it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;TOOLS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;search_code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Search the codebase with a regex. Returns matching &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                       &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lines with file paths and line numbers.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input_schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pattern&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glob&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Optional file filter, e.g. &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;*.py&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pattern&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Read a file (or a line range) from the codebase.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input_schema&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;object&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;properties&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;string&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;start_line&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;integer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;end_line&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;integer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;path&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;search_code&lt;/code&gt; is just ripgrep behind a subprocess call, with output capped so a careless &lt;code&gt;.*&lt;/code&gt; doesn't blow up the context window:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;search_code&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--line-number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--max-count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--no-heading&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;cmd&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--glob&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cwd&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;REPO_ROOT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;6000&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;No matches.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent loop is the standard tool-use pattern: call the model, execute any tool calls, feed results back, repeat until it produces a text answer. I run it on &lt;code&gt;claude-sonnet-5&lt;/code&gt; — fast enough that a 6-hop investigation still comes back inside Slack-attention-span, and cheap enough that I don't think about the bill.&lt;/p&gt;

&lt;p&gt;A typical answer takes 3–6 tool calls. Watching the transcripts is genuinely fun: it greps for the obvious keyword, reads the hit, notices an import, follows the import, then answers with file paths and line numbers.&lt;/p&gt;

&lt;p&gt;For the numbers people: median answer latency is around 20 seconds end-to-end, the slowest multi-hop investigations land near 45, and a month of real usage cost less than a single takeout lunch in API fees. Compare that to the 20 minutes of human digging each question used to burn and the math stops being interesting — it's just obviously worth it. 💡&lt;/p&gt;

&lt;h3&gt;
  
  
  The three problems that actually mattered
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Hallucinated file paths destroy trust instantly. ⚠️&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Early on, the bot answered a question with a confident reference to &lt;code&gt;utils/validation.py&lt;/code&gt; — a file that did not exist. It had seen similar codebases in training and &lt;em&gt;invented&lt;/em&gt; a plausible path. One wrong answer like that and your teammates stop trusting every answer.&lt;/p&gt;

&lt;p&gt;The fix was mechanical, not prompt-based: before sending the reply, I extract every path-looking string from the answer and check it against the actual file tree. Any path that doesn't exist gets the whole answer flagged:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify_paths&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;repo_files&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;candidates&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[\w./-]+\.(?:py|ts|tsx|go|sql|yaml|toml)\b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;candidates&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;`&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;repo_files&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;verify_paths&lt;/code&gt; returns anything, the bot appends a visible warning and links the search query it ran, so the human can check. Since adding this, "confidently wrong path" incidents went from a few per week to zero reaching users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Stale checkouts give stale answers.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The bot answers from a local clone. A clone that's three days old will cheerfully describe code that was refactored away on Tuesday. My first "fix" was pulling on a timer; the real fix was pulling &lt;strong&gt;on demand&lt;/strong&gt;: every question triggers a &lt;code&gt;git fetch&lt;/code&gt; + fast-forward before the agent loop starts. It adds ~2 seconds and eliminated the entire class of "that function doesn't exist anymore" answers. The answer also includes the commit SHA it read from, which turned out to be the single most trust-building detail in the whole project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The bot must not see more than the asker can.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the one people skip. A codebase Q&amp;amp;A bot is effectively a &lt;em&gt;read amplifier&lt;/em&gt;: anyone who can ask it questions can extract anything in its clone. Two hard rules came out of this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The bot's clone excludes anything secret-shaped — &lt;code&gt;.env&lt;/code&gt; files never land in git anyway (right? 😅), but I also deny-list config directories and run the same secret scanner we use in CI against every tool result before it enters the model context.&lt;/li&gt;
&lt;li&gt;The bot only joins channels where everyone already has repo read access. No DMs. This one is organizational, not technical, and it's non-negotiable.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tool-calling retrieval beats embedding RAG for code Q&amp;amp;A.&lt;/strong&gt; Code questions are graph traversals, not similarity lookups. Give the model grep and read, and let it walk the graph. My answer quality jumped more from this one architectural change than from everything else combined.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Verify outputs mechanically, not with prompts.&lt;/strong&gt; "Do not invent file paths" in the system prompt reduced hallucinations. Checking paths against the file tree &lt;em&gt;eliminated&lt;/em&gt; them. Anything you can verify with code, verify with code.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Citations are a feature, not decoration.&lt;/strong&gt; Answers that end with &lt;code&gt;billing/worker.py:141&lt;/code&gt; (at commit &lt;code&gt;a3f9c21&lt;/code&gt;) get trusted and clicked. Answers without them get re-asked to a human. If your bot can't cite, it's a liability.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Freshness is a product requirement.&lt;/strong&gt; Nobody forgives "that code was deleted last week." Sync on demand, stamp the SHA, and staleness stops being a conversation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Scope the bot's eyes to the asker's eyes.&lt;/strong&gt; Decide what the bot can read and who can ask it &lt;em&gt;before&lt;/em&gt; the first deploy, not after the first awkward answer. Retrofitting access control onto a chatbot is miserable.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;Two things are on the list. First, &lt;strong&gt;cross-repo questions&lt;/strong&gt; — "which services call the notification API?" needs the bot to search several clones and merge results, which mostly means smarter tool routing. Second, &lt;strong&gt;letting it answer with diagrams&lt;/strong&gt;: it already understands structure well enough that generating a Mermaid diagram of a subsystem on request feels within reach.&lt;/p&gt;

&lt;p&gt;I'm deliberately &lt;em&gt;not&lt;/em&gt; giving it write access. An answer bot that occasionally opens "helpful" PRs is a different product with a very different blast radius — that's a post for another day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap-up
&lt;/h2&gt;

&lt;p&gt;Total damage: ~600 lines of Python, two weekends, and one discarded RAG pipeline. The bot now handles the majority of "where is / how does" questions in our Slack, and the ones it can't handle come to me with context attached instead of from zero.&lt;/p&gt;

&lt;p&gt;If you've been thinking about pointing Claude at your own codebase: skip the embedding pipeline, start with two tools and an agent loop, and spend your saved week on verification and access control. That's where the real product is. ✅&lt;/p&gt;

&lt;p&gt;If this was useful, &lt;strong&gt;follow me here on Dev.to&lt;/strong&gt; — I write weekly about building with AI coding agents, including the failures. And if you build your own version, I'd genuinely love to hear what broke first: drop it in the comments. 💬&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>python</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
