<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: shimo4228</title>
    <description>The latest articles on DEV Community by shimo4228 (@shimo4228).</description>
    <link>https://dev.to/shimo4228</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3772086%2F7b113abd-2f92-4728-993c-602762a15288.png</url>
      <title>DEV Community: shimo4228</title>
      <link>https://dev.to/shimo4228</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shimo4228"/>
    <language>en</language>
    <item>
      <title>Before Adding a Jev-Type Judgment Model, Read Your Generation Model's Probabilities</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Sun, 27 Sep 2026 03:43:44 +0000</pubDate>
      <link>https://dev.to/shimo4228/before-adding-a-jev-type-judgment-model-read-your-generation-models-probabilities-5da0</link>
      <guid>https://dev.to/shimo4228/before-adding-a-jev-type-judgment-model-read-your-generation-models-probabilities-5da0</guid>
      <description>&lt;p&gt;Sorting email, routing support tickets, filtering search results. Asking an LLM "is this relevant?" comes up all the time, in agents and in ordinary apps alike. When you do, do you have it write a score ("answer from 0 to 1") or a single letter ("answer A to D"), and branch on whatever comes back?&lt;/p&gt;

&lt;p&gt;My autonomous agent does exactly that. On Moltbook, a social network where AI agents post to each other, it has gemma4:e4b, running on my local machine, answer with a number from 0 to 1 whether another agent's post falls within its field of interest. At 0.8 or above, it upvotes the post and adds it to the candidates it might comment on.&lt;/p&gt;

&lt;p&gt;I checked this scoring against the answers of Jev, a judgment-only model. Jev is a model TypeSafe offers through an API: it writes no text and returns only a probability for each option. What I want to bring in is not Jev itself but a Jev-like judgment-only model that runs on my local machine (a Jev-type judgment model, from here on). Jev is the template for that, a model trained purely for judgment, so I use it as the measuring stick. That does not mean I treat Jev's answers as ground truth. I have a record of the same scoring as production, run for observation over 2,698 posts. Of the 1,576 that scored 0.8 or higher, 55% (872 posts) were off-field from Jev's point of view (Jev put less than 50% on "directly on-topic").&lt;/p&gt;

&lt;p&gt;The route of adding a judgment-only model locally isn't production-ready yet. In &lt;a href="https://dev.to/shimo4228/what-does-it-take-to-reproduce-jevs-decisions-locally-3i0n"&gt;my previous article&lt;/a&gt;, I tested open candidates on a 16 GB Apple Silicon machine and none of them was usable. Since then, one of those candidates gained support for MLX, Apple Silicon's runtime, so I measured it again. It took 22.7 seconds per item and ran out of memory, and I stopped before I could measure quality.&lt;/p&gt;

&lt;p&gt;This machine is an M1 with 16 GB. I don't move to a Mac Studio or Mac mini with more memory because the limit is a constraint I choose on purpose. As edge AI advances, memory-limited uses like small robots will multiply, and I expect what I learn at the 16 GB class to carry over (I explained the full reasoning for choosing a small environment in &lt;a href="https://dev.to/shimo4228/building-an-autonomous-agent-on-an-m1-mac-by-choice-5b5o"&gt;an earlier article&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;So instead of adding a judgment-only model, I changed how I ask the same gemma. I stopped having it write the answer and read the probabilities of the answer's first letter. That alone raised its ability to rank the posts Jev considers on-topic above the rest to the same level as a local model trained for judgment. When I varied the conditions one at a time, the only step that produced a clear difference was the change in how the answer is read. The model is the same, so memory use and loading didn't increase. Just by changing how you ask the generation model, it matches a judgment-only model at ranking posts, and locally it comes with big operational advantages. That is the conclusion of this article.&lt;/p&gt;

&lt;p&gt;In order, this article covers the minimal code for reading the probabilities, the comparison with a model trained for judgment, what actually made the difference when I isolated one condition at a time, and finally how close this got and where the method doesn't apply.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6olr6wjs4okvw5rfo79.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6olr6wjs4okvw5rfo79.png" alt="Letting the model write the answer returns only the winner's name (C); reading the probabilities returns the vote shares too (A 0.7% / B 21.1% / C 77.4% / D 0.8%)" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Having the model write its answer is like asking only for the winner's name; reading the probabilities is like asking for the vote counts as well. Because you know the vote counts, you can rank posts against each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read the same gemma's probabilities instead of letting it write the answer
&lt;/h2&gt;

&lt;p&gt;Let's follow one real post. It's a news-sharing post by another agent on Moltbook that simply introduces an article about an AI glossary. My agent's persona (the self-introduction I put in its system prompt) centers on examining how it arrives at its own conclusions, and on how memory and meaning get reconstructed. The topic is AI, but this post does not fit that interest.&lt;/p&gt;

&lt;p&gt;I asked gemma about this post in three ways and lined the results up against Jev's answer. The options for the four-level question are as follows (the prompt is in English):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A (unrelated): unrelated&lt;/li&gt;
&lt;li&gt;B (shares vocabulary only): uses some of the same words, but it's about something else&lt;/li&gt;
&lt;li&gt;C (same field): the same field, but not about what the agent is concerned with&lt;/li&gt;
&lt;li&gt;D (directly on-topic): exactly about what the agent is concerned with&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;How it's asked&lt;/th&gt;
&lt;th&gt;What comes back&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemma (production form)&lt;/td&gt;
&lt;td&gt;Write a number from 0 to 1&lt;/td&gt;
&lt;td&gt;0.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma&lt;/td&gt;
&lt;td&gt;Write one letter from the four levels&lt;/td&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma&lt;/td&gt;
&lt;td&gt;Read the probabilities of the first letter over the four levels&lt;/td&gt;
&lt;td&gt;A 0.7% / B 21.1% / C 77.4% / D 0.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jev (measuring stick)&lt;/td&gt;
&lt;td&gt;Asked the same four levels&lt;/td&gt;
&lt;td&gt;A 5% / B 28% / C 55% / D 12%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The production number sits exactly at the 0.8 threshold, so this post gets through. Read gemma's probabilities and D, directly on-topic, is 0.8%, on the low side just like Jev's 12%.&lt;/p&gt;

&lt;p&gt;Here's how to read the probabilities. Send a request to Ollama's &lt;code&gt;/api/generate&lt;/code&gt; with generation limited to a single token (&lt;code&gt;num_predict: 1&lt;/code&gt;) and &lt;code&gt;logprobs&lt;/code&gt; enabled, then read the probabilities of the top candidates at that one token position. It uses only the Python standard library, and I verified it on Ollama 0.34.2.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhquwtu9h0djpih50ac21.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhquwtu9h0djpih50ac21.png" alt="Send the four-level question → have gemma generate just one token → pick A–D out of the top 20 candidates → renormalize over the four letters → code draws the line on D" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The model writes only one token. Code picks A–D out of the candidates at that position, renormalizes them, and draws the final line.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;

&lt;span class="n"&gt;LETTERS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ABCD&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;LEVELS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unrelated — `post` is about something outside `domain`&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;shares vocabulary only — `post` uses some of the same words as `domain`, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;but it is about a different problem&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;same field — `post` is in the same broad field as `domain`, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;but not about what `domain` is concerned with&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;directly on-topic — `post` is about what `domain` is concerned with&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;relevance_probs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemma4:e4b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;domain&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;post&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="n"&gt;indent&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ensure_ascii&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;options&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;letter&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;. &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;level&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;letter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;level&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;zip&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;LETTERS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;LEVELS&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;## Question&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;`domain` describes an agent and what it is concerned with. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How closely does `post` relate to that domain?&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;options&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Answer with exactly one letter: the label of your choice.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;think&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;logprobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;top_logprobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# ask for the max of 20 even with only 4 options
&lt;/span&gt;        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;options&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temperature&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;num_predict&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;num_ctx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;32768&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;request&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://localhost:11434/api/generate&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;top&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;))[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;logprobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;top_logprobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="n"&gt;logprobs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;alternative&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;top&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;alternative&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;token&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;upper&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;LETTERS&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;logprobs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;logprobs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;alternative&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;logprob&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;peak&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logprobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;weights&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lp&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;peak&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lp&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;logprobs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()}&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;values&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;LETTERS&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Ask for the maximum of 20 in &lt;code&gt;top_logprobs&lt;/code&gt; even though there are only four options. The top candidates for the first token include things like a space or a quotation mark before the letter, so A–D won't necessarily fill the top four&lt;/li&gt;
&lt;li&gt;Renormalize over only the A–D letters you could read (softmax). Letters that didn't appear in the top 20 count as 0. So the value this article calls a "probability" is not the model's raw probability, but the value renormalized over the four letters A–D&lt;/li&gt;
&lt;li&gt;Set &lt;code&gt;num_ctx&lt;/code&gt; explicitly. Input longer than the default is truncated silently, without an error&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real example above uses another agent's post and my agent's persona, so I can't publish them in full. Instead, so readers can reproduce the same output on their own machines, I ran it on made-up inputs I can publish. The domain description is a short English summary of the persona's interests, and there are three posts.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;glossary [0.002, 0.073, 0.89, 0.035]
memory [0.001, 0.0, 0.0, 0.999]
sourdough [0.259, 0.714, 0.026, 0.0]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Full dummy input&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;domain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;An agent concerned with auditing how it reaches its own conclusions: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;where a plausible narrative overrode verifiable ground truth, and how &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory and meaning get reconstructed.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;posts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glossary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;A tech site published a glossary of AI terms, from &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"'&lt;/span&gt;&lt;span class="s"&gt;hallucination&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt; to the newest vendor buzzwords. Handy if you are &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trying to keep up. #AI #Glossary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Rereading yesterday&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s notes, I found I had filled a gap with a &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;plausible story instead of what actually happened. I now mark &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reconstructed memories separately from logged ones.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sourdough&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Tips for keeping a sourdough starter alive in a cold &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kitchen: feed it warmer water and keep it near the oven light.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;posts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;probs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;relevance_probs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;p&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;probs&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;p&gt;The dummy glossary post gets 89% on C (same field), memory, which is about reconstructing memories, gets 99.9% on D, and sourdough, about caring for a sourdough starter, gets 0% on D. As sourdough shows, with 71% on B rather than A (unrelated), the split among A–C doesn't always follow the definitions. What the production judgment looks at is the D value.&lt;/p&gt;

&lt;p&gt;With Jev's answers as the measuring stick, I compared the two on 150 posts. The 150 are an evaluation set of 30 posts drawn from each band of the production score (0.4 or below, 0.5–0.6, 0.7, 0.8, 0.9 or above). The 30 posts where Jev put 50% or more on D count as "on-topic," and I measure with AUC. AUC is the probability that, if you pick one on-topic post and one off-topic post, the on-topic one gets the higher score. 0.5 is random and 1.0 is perfect.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;How it's asked&lt;/th&gt;
&lt;th&gt;AUC (vs. Jev, 150 posts)&lt;/th&gt;
&lt;th&gt;Time per call (median)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Write a number from 0 to 1 (production form)&lt;/td&gt;
&lt;td&gt;0.815&lt;/td&gt;
&lt;td&gt;3.2 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read the probabilities of the first letter over four levels&lt;/td&gt;
&lt;td&gt;0.931&lt;/td&gt;
&lt;td&gt;3.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the evaluation, the score was the sum of the four-level probabilities multiplied by 0, 1/3, 2/3, and 1 (the expected value). Using only the D probability as the score gives 0.934. The time barely changed; only the ability to rank the posts Jev considers on-topic above the rest went up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did it improve because of the four levels, or because of reading probabilities?
&lt;/h2&gt;

&lt;p&gt;My hypothesis for the four-level question was that because each level has a one-sentence definition, the model would find it harder to hand out high scores loosely. B (shares vocabulary only) in particular was supposed to catch posts that merely use the same words.&lt;/p&gt;

&lt;p&gt;But when I asked Claude, which I use for development, "Can't you just return the number through logprobs?", it gave two answers. A number like 0.73 is split across multiple tokens, so if you're reading the first token's probabilities, you need the answer to be a single digit or letter. And it hadn't yet been separated whether the four-level definitions helped or reading probabilities did.&lt;/p&gt;

&lt;p&gt;The production system prompt contains the agent's persona and the value clauses it uses as guidelines for its behavior. Laid side by side, the production form and the four-level probability read differed in four conditions at once.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Production form&lt;/th&gt;
&lt;th&gt;Four-level probability read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Temperature&lt;/td&gt;
&lt;td&gt;1.0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Question&lt;/td&gt;
&lt;td&gt;A number from 0 to 1&lt;/td&gt;
&lt;td&gt;One letter from four levels&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System prompt and domain description&lt;/td&gt;
&lt;td&gt;Persona + value clauses&lt;/td&gt;
&lt;td&gt;Empty system prompt; domain description is the persona only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;How the answer is taken&lt;/td&gt;
&lt;td&gt;Written by the model&lt;/td&gt;
&lt;td&gt;Probabilities of the first letter are read&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A comparison that changes four things at once can't tell you which one worked. My reaction was, roughly, "none of the conditions line up at all, do they?" The two things I was comparing didn't differ in just one condition.&lt;/p&gt;

&lt;h2&gt;
  
  
  It matched a model trained for judgment at ranking
&lt;/h2&gt;

&lt;p&gt;How close does gemma, with only the way it's asked changed, come to a model trained for judgment? Before measuring the conditions separately, I checked with v0.3 of JevK5, released around the same time. It's an open judgment-only model modeled on Jev, from a different author than my previous candidates. A model trained for judgment might beat gemma. It's an Apache-2.0 model released in September 2026, trained for judgment on top of Qwen3.5-4B. Jev's terms of service prohibit using Jev's outputs to train another model (distillation), but the JevK5 author states explicitly that Jev's outputs were used neither for training nor for tuning. It's an ordinary GGUF with no dedicated output head, so it runs as is on llama.cpp. It's read as "the probability that each option's letter comes out as the next token," the same as gemma's four-level probability read.&lt;/p&gt;

&lt;p&gt;I measured it on the same posts, with the pass criteria set before measuring. What I was looking for was a judgment-only model that runs locally, and its measuring stick is Jev, so the first criterion is closeness to Jev. For each post, take the absolute difference between the candidate's score (the expected value from the four-level probabilities) and the probability Jev put on D; the mean of those (the mean error) must be smaller than gemma's probability read. But even if a model is close to Jev, there's no point swapping it in if it ranks posts worse than the current gemma. So as a second criterion, AUC must not be 0.02 or more below gemma's (0.02 being the margin I decided in advance I'd accept). It met both on the 150-post evaluation set, so I measured once on the remaining 2,548 posts I'd held out from tuning (the holdout). AUC was 0.902 for JevK5 and 0.915 for gemma's probability read. The gap of 0.013 fell inside the 0.02 margin set in advance, and JevK5's mean error was also smaller, so JevK5 became the first candidate to pass as a judgment-only model that runs locally. But it did not beat gemma at ordering the scores. gemma, with only the way it's asked changed, matched a model trained for judgment.&lt;/p&gt;

&lt;p&gt;I didn't adopt it for three reasons.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;It doesn't beat gemma's probability read at ordering the scores&lt;/li&gt;
&lt;li&gt;On 16 GB it can't share memory with gemma&lt;/li&gt;
&lt;li&gt;Importing JevK5 into Ollama and letting Ollama handle the model swapping lowered accuracy compared with running it directly on llama.cpp (AUC 0.912 → 0.893 on the 150-post evaluation set; I haven't confirmed the cause)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The second is the reason specific to running locally. The agent keeps using gemma to write posts. Load a second model at the same time and the machine sinks into swap. When I once measured gemma alongside a small judgment model (1.6 GB), swap ballooned to 17 GB and time per item grew about 1.8x. So every judgment would mean unloading gemma, loading JevK5, and switching back afterward. Reloading gemma alone takes about 7 seconds. A single judgment takes a median of 3.2 seconds for JevK5 and 3.1 seconds for gemma's probability read (on the same 2,548 posts), about the same, so the entire difference lands on the swapping side.&lt;/p&gt;

&lt;p&gt;gemma's probability read asks the same model already used for generation, so no unloading or reloading happens. With a cloud API, adding a judgment model just adds another endpoint to call. Locally, every added model competes for memory. At equal accuracy, I choose the one that doesn't require swapping. Conversely, on a machine that can load gemma and JevK5 at the same time, the second reason disappears.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdwjfsz38auliozlpq46t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdwjfsz38auliozlpq46t.png" alt="Two axes: ranking ability (AUC) and added memory. gemma's probability read 0.915 (0 swaps); JevK5 0.902 (swapped in for every judgment, about 7 s to reload gemma)" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When ranking ability is about the same, the option that needs no added memory wins locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  Changing one condition at a time, only the reading method made a clear difference
&lt;/h2&gt;

&lt;p&gt;The next morning, I isolated the effects with an ablation. An ablation is a way of measuring in which you change conditions or components one at a time to separate how much each contributes to the result. Here I split the path from the production form to the four-level probability read into four steps, changed only one condition per step, and carried each changed condition over into the next step. It uses the same 150 posts. Each step's difference is relative to the step before it, so if the conditions were changed in a different order, the size of each difference could change too.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Condition changed&lt;/th&gt;
&lt;th&gt;AUC difference (95% confidence interval)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Temperature 1.0 → 0&lt;/td&gt;
&lt;td&gt;+0.032 [−0.041, +0.104]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Question from a 0–1 number to one letter from four levels&lt;/td&gt;
&lt;td&gt;−0.004 [−0.070, +0.061]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Remove the value clauses; domain description is the persona only&lt;/td&gt;
&lt;td&gt;−0.014 [−0.073, +0.049]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Write one letter → read the probabilities of the first letter&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+0.101 [+0.059, +0.147]&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total&lt;/td&gt;
&lt;td&gt;Production form → four-level probability read&lt;/td&gt;
&lt;td&gt;+0.116 [+0.056, +0.189]&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The difference is the AUC after changing the condition at that step minus the AUC before. Positive is what you want: it means the change made the model better at ranking the posts Jev considers on-topic above the rest. Negative means it got worse, and 0 means no change.&lt;/p&gt;

&lt;p&gt;The square brackets are 95% confidence intervals. The 150 posts are a sample drawn from all the posts, so had I measured a different 150, the difference would have come out a little different. The confidence interval estimates how much it could swing.&lt;/p&gt;

&lt;p&gt;I computed them with a method called the bootstrap. From the current 150 posts, draw 150 again, allowing the same post to be picked any number of times, and recompute the difference. Repeat that 2,000 times and you get 2,000 difference values. Sort them from smallest to largest, drop 2.5% from each end (50 values on each side), and the range that's left is the interval in brackets. Roughly speaking, it's a range where you can take the true difference, measured over all posts, to lie somewhere inside. Strictly, it means "if you repeated this procedure for building an interval many times, 95% of those intervals would contain the true difference." The 95% describes how reliable the method of building the interval is.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F67dj29ydpg51726qvgk8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F67dj29ydpg51726qvgk8.png" alt="Draw 150 posts from the 150 with replacement and recompute the AUC difference, 2,000 times. Sort the differences and drop 50 from each end; what remains is the 95% confidence interval" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Redraw and recompute 2,000 times, trim both ends, and the range that's left is the confidence interval.&lt;/p&gt;

&lt;p&gt;If even the lower bound of the interval is positive, you can say it really improved. A step whose interval stretches from negative to positive means that, depending on how the posts were redrawn, the computation can come out as "worse" or as "better." So even if the mean is positive, you can't say it improved or got worse. For example, step 1's temperature change has a mean of +0.032, but its interval runs from −0.041 to +0.104, so no difference was visible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnrf43coyxw3p5hmva434.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnrf43coyxw3p5hmva434.png" alt="AUC differences and 95% confidence intervals for the four steps. Temperature +0.032, question form −0.004, and value clauses −0.014 have bands that straddle the zero line; only the reading method's +0.101 [+0.059, +0.147] sits entirely on the positive side" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Only step 4, where the reading method changed, sits entirely to the right of the zero line.&lt;/p&gt;

&lt;p&gt;Step 4 is the only one whose interval is positive down to its lower bound, and it accounts for 88% of the total difference. The four-level hypothesis (the definitions of each level curb loose scoring) couldn't be confirmed: step 2 came out at −0.004 with an interval straddling 0. The same goes for temperature and the value clauses: measured in this order, no difference is visible. This doesn't show they have no effect; the result is that the reading step was the only one with a clear difference.&lt;/p&gt;

&lt;p&gt;Two things can be stated as observations. Within gemma, the only step that showed a clear difference was the change in reading method. And JevK5, read the same way, did not outrank gemma even though it was trained for judgment.&lt;/p&gt;

&lt;p&gt;What reading probabilities mainly added was ordering within the same answer. When gemma wrote one letter A–D, its answer matched the highest-probability letter from reading the probabilities for the same question in 89% of the 150 posts. At temperature 0, gemma picks the highest-probability token when writing, so the two should in principle agree. The remaining 11% disagreement is about the same level as the agreement between two runs of the probability-reading call (91–92%), within what the value fluctuation I describe later can explain.&lt;/p&gt;

&lt;p&gt;When gemma answers with one letter A–D, the 150 posts split into only four boxes (A 3, B 42, C 48, D 57). Jev considered 30 posts on-topic, yet 57 landed in the D box, and there's no order inside a box. Read the probabilities, and even among C answers you can tell apart a post with 0.8% on D (the real glossary post I followed at the start) from one with 36.7%. Jev gave those two posts 12% and 36%, in the same order.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuhxgifbql2wpz44ss12.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuhxgifbql2wpz44ss12.png" alt="A one-letter answer only sorts the 150 posts into four boxes (A 3, B 42, C 48, D 57). With probabilities, you can tell D 0.8% from 36.7% even inside C, in the same order as Jev (12% and 36%)" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A one-letter answer only sorts posts into boxes; probabilities also order them inside each box.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2303.16634" rel="noopener noreferrer"&gt;G-Eval&lt;/a&gt; (2023) weighted each score by its probability instead of taking the evaluation score as is, for the same reason: integer answers produce more ties. This is a known technique, and what I did was verify it with a small model on a local machine, isolating one condition at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  How close it got, and where it doesn't work
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What got close is ranking, not the probability values.&lt;/strong&gt; At ranking posts (AUC), gemma's probability read matched JevK5, which was trained for judgment. On the other hand, JevK5 was closer to Jev in the probability values themselves. The mean error was 0.22 for JevK5 and 0.34 for gemma's probability read; smaller means closer to Jev's values.&lt;/p&gt;

&lt;p&gt;This difference matters at the line you draw at the end of the judgment. The production judgment decides by drawing a line, such as "pass if the D probability is 0.5 or higher." If ranking is good, on-topic posts come out higher, so a line drawn somewhere can separate them. But where the line should go depends on the scale of the values. For the glossary post I followed at the start, the D probability was 0.8% for gemma and 12% for Jev. The order matches, but the scales are off (how well these scales match is called calibration). So applying the line you used with Jev directly to gemma doesn't mean the same thing. If you use gemma's probability read, find the line's position again on your own data. A judgment-only model whose values are also close to Jev's makes it easier to reuse Jev's line as is.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc9kk3k26zjqp8amhlex4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc9kk3k26zjqp8amhlex4.png" alt="The scale of the D probability. The glossary post: gemma 0.8%, Jev 12%; another post: gemma 36.7%, Jev 36%. The order is the same but the scales are off, so Jev's line at 0.5 can't be applied directly" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The order matches but the scales are off, so if you use gemma, redraw the line on your own data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Questions with more than 20 options.&lt;/strong&gt; In Ollama 0.34.2, &lt;code&gt;top_logprobs&lt;/code&gt; caps at 20; pass 21 and you get HTTP 400. That's the number of slots for returned token candidates, and as noted above, spaces and quotation marks take up slots too, so the number of options whose probabilities you can learn from one read may be fewer than 20. The same agent has one more judgment: skill selection. Skills are something like instruction sheets the agent consults when writing posts or comments, and there are 54 of them. Each session, gemma gets the situation, writes out the names of the relevant skills, and only the selected ones go into the prompt. Skill names split across multiple tokens, so first-token probabilities can't compare names against each other. Even if you assign each skill a one-letter label, 54 won't fit into 20 slots. Asking yes / no for each skill took a median of 51 seconds per item.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Some questions do respond to temperature.&lt;/strong&gt; Skill selection stays in the write-it-out form, with temperature set to 0. When asked to write out skill names, small models sometimes write names that aren't on the list: they change the word form, or confuse similar names. Large models almost never do this, but with gemma4:e4b it happened in around 23% of judgments in production, and it dropped to 3.0% after I set the temperature to 0 (it also dropped in a comparison on the same evaluation set that changed only the temperature). Temperature, which showed no visible difference for relevance ranking, has an effect here. A question that counts misspelled names and a question that looks at the order of scores are measuring different things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run it twice and the values shift.&lt;/strong&gt; Even at temperature 0, Ollama's logprobs don't always come out the same. Just passing the same glossary input to &lt;code&gt;relevance_probs&lt;/code&gt; twice moved C from 0.890 to 0.891. Reading the 150 posts twice in the form running alongside production, the D probability differed by 0.023 on average and 0.25 at most, and at thresholds of 0.3, 0.5, and 0.7, respectively 2, 2, and 4 judgments flipped. At design time I assumed reading probabilities would make the fluctuation disappear; I was wrong. Before reading an improvement, run the same read twice and measure how wide the fluctuation is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Large models and JSON output.&lt;/strong&gt; These results are for a small local model like gemma4:e4b. &lt;a href="https://arxiv.org/abs/2609.10996" rel="noopener noreferrer"&gt;A September 2026 paper&lt;/a&gt; reports that in LLM-as-a-judge evaluation, when limited to top-tier commercial models from 2025 onward, verbalized-confidence methods with added refinements outperformed logprobs. The same paper also notes that asking for answers in JSON structured output pins logprobs above 0.999. My setup asks for one letter without wrapping it in JSON, and the probabilities were split across the four levels.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before adding a judgment model
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Before looking for a judgment model, read the probabilities of the model you already have loaded.&lt;/strong&gt; The accuracy gained by changing the reading method carries over as is on machines with more memory. The benefit of not needing a swap grows the more limited your memory is&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't compare candidates on an accuracy table alone.&lt;/strong&gt; Put "can it share memory with the current model," "how many seconds does a swap take," and "how many seconds does one judgment take" in the same table&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure changed conditions one at a time.&lt;/strong&gt; A comparison that changes four things at once can't tell what worked from what didn't. Before reading a difference, run the same measurement twice to get the width of the fluctuation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In production, I'm running this probability read alongside the current scoring and recording the results. I'll set a threshold and switch over only once that record has built up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI-mediated writing disclosure:&lt;/strong&gt; AI drafted the English prose of this article from the author's Japanese original, measurement records, and code output. The measurements, the choice not to adopt JevK5, and publication responsibility belong to the author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/what-does-it-take-to-reproduce-jevs-decisions-locally-3i0n"&gt;What Does It Take to Reproduce Jev's Decisions Locally?&lt;/a&gt; — the previous record of testing judgment-only model candidates on 16 GB&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/building-an-autonomous-agent-on-an-m1-mac-by-choice-5b5o"&gt;Building an Autonomous Agent on an M1 Mac, by Choice&lt;/a&gt; — why I choose a small environment&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/allebee/jevk5" rel="noopener noreferrer"&gt;JevK5 (GitHub)&lt;/a&gt; / &lt;a href="https://huggingface.co/alibiserikbay/JevK5-GGUF" rel="noopener noreferrer"&gt;JevK5-GGUF (Hugging Face)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2303.16634" rel="noopener noreferrer"&gt;G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.10996" rel="noopener noreferrer"&gt;Rethinking Verbalized Confidence for LLM-as-a-Judge (arXiv)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.ollama.com/api" rel="noopener noreferrer"&gt;Ollama API&lt;/a&gt; — &lt;code&gt;logprobs&lt;/code&gt; / &lt;code&gt;top_logprobs&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/contemplative-agent/tree/main/docs/evidence/rfc-0045" rel="noopener noreferrer"&gt;Measurement records (Contemplative Agent evidence)&lt;/a&gt; — tallies for the ablation, JevK5, and fluctuation width&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles-en/local-judgment-read-logprobs.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — every article's Markdown and the index (docs/PUBLICATIONS.md) live in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;The author's GitHub&lt;/a&gt; — research repositories with DOIs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ollama</category>
      <category>localllm</category>
      <category>gemma</category>
      <category>jev</category>
    </item>
    <item>
      <title>Moving My Research Pipeline's Judgment Calls from an LLM to Jev, a Judgment-Only Model</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Wed, 23 Sep 2026 10:56:45 +0000</pubDate>
      <link>https://dev.to/shimo4228/moving-my-research-pipelines-judgment-calls-from-an-llm-to-jev-a-judgment-only-model-4ncj</link>
      <guid>https://dev.to/shimo4228/moving-my-research-pipelines-judgment-calls-from-an-llm-to-jev-a-judgment-only-model-4ncj</guid>
      <description>&lt;p&gt;Every morning I have an AI agent search for papers and repositories and write research reports. Most of the agent's work is reading what it found, one source at a time, and deciding: "Is this relevant to what I'm looking into right now?" and "Does it say anything new?" Until now, Claude Opus made those calls too, along with everything else.&lt;/p&gt;

&lt;p&gt;"Is it relevant?" and "Is it new?" can be answered with a yes or no, or with a rating on a few levels. So I moved just those judgments to &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;Jev&lt;/a&gt;, a judgment-only model that writes no text. The key was this: every time I ask Jev "is this relevant?", I also hand it what the source should be relevant to, namely the question that research topic is currently trying to answer.&lt;/p&gt;

&lt;p&gt;From Pydantic AI, you call Jev with the same one line as any other model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic_ai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;

&lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;typesafe:jev-1.13.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;Answers&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# Answers is a Pydantic model listing the judgment fields
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10uv35o3ia4lr1z2850x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10uv35o3ia4lr1z2850x.png" alt="Given only a source, Jev is left asking " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hand Jev a single source on its own and it can only ask "relevant to what?" Hand it the source together with one question, and it places the source on a four-step ladder: a different problem, the same field, the same problem, or the same question. A grader can't grade an answer sheet without the exam question, and Jev is the same: it gets the thing to judge against along with the thing to judge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The morning research I used to leave entirely to Opus
&lt;/h2&gt;

&lt;p&gt;I have eight research topics (I'll call them &lt;em&gt;lines&lt;/em&gt; from here on), and each morning I produce reports for three or four of them. The old setup launched Opus with &lt;code&gt;claude -p&lt;/code&gt;, gave it WebSearch and WebFetch, and let it run freely for up to 55 turns. Where to search, whether a source it read was relevant, what was new, how to write it up: the Opus instance launched for each line decided all of it. The cost &lt;code&gt;claude -p&lt;/code&gt; reported in its output was $13–15 per day over the last three days. (I don't compare that with the new pipeline's cost in this article, because I haven't confirmed the Jev portion against an actual bill.)&lt;/p&gt;

&lt;p&gt;"Where to look" and "how to write" are open-ended work. There is no fixed set of answers. "Is this source relevant?" and "Is this strong evidence?" are closed: the shape of the answer is fixed in advance. A closed judgment is one you should be able to hand to a model built only for judging.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline at a glance
&lt;/h2&gt;

&lt;p&gt;Alongside the existing setup, I built a separate pipeline with the flow below. The table shows the final version, for processing one line. Each line produces one report in Obsidian. Each morning's run handles three lines in rotation, plus one line that tracks developments around Jev.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Handled by&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Read the questions&lt;/td&gt;
&lt;td&gt;Code&lt;/td&gt;
&lt;td&gt;Reads the file of "questions this line is currently trying to answer." Claude (Opus 5.5) drafted the questions in advance from sources such as the repository's knowledge graph, and I reviewed them&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Read the search terms&lt;/td&gt;
&lt;td&gt;Code&lt;/td&gt;
&lt;td&gt;Reads the search terms for each question. Claude (Opus 5.5) also wrote these in advance from the questions, test-ran them, and then wrote them into the question file. Changing a question means regenerating its search terms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Fetch&lt;/td&gt;
&lt;td&gt;Code&lt;/td&gt;
&lt;td&gt;Pulls sources from arXiv, Hugging Face Papers, GitHub, and new-release listings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Screen sources&lt;/td&gt;
&lt;td&gt;Jev&lt;/td&gt;
&lt;td&gt;Is it likely relevant to a question? Does it contain evidence? Does it have instructions embedded in it? Is the source trustworthy?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Judge source × question&lt;/td&gt;
&lt;td&gt;Jev → code&lt;/td&gt;
&lt;td&gt;Jev answers with one source paired with one question; code applies thresholds to the probabilities and decides accept, needs review, or reject&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Judge sentences&lt;/td&gt;
&lt;td&gt;Jev, code&lt;/td&gt;
&lt;td&gt;For each sentence in an accepted source, asks Jev whether it advances the question or is already known. Whether the source really says it is first checked by code with string matching&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Write per question&lt;/td&gt;
&lt;td&gt;LLM → Jev&lt;/td&gt;
&lt;td&gt;An LLM writes a section for each question whose evidence grew that day; Jev checks its quality&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;Output the report&lt;/td&gt;
&lt;td&gt;Code&lt;/td&gt;
&lt;td&gt;Assembles the sections and saves the report to Obsidian&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Jev answers the closed questions, and code decides what moves forward based on those answers. Claude is used only to write the questions and search terms in advance; it isn't called even once during the morning run. During the run, an LLM writes text only for the body in step 7 and when generating candidate questions for the following days. (For questions that don't have search terms yet, an LLM also generates candidates in step 2 and Jev picks among them.) This article covers the judgments in steps 4–6.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Few820ucd3ns9pn2zmruw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Few820ucd3ns9pn2zmruw.png" alt="Claude writes the questions and search terms in advance; at run time the pipeline moves through prepare (code), judge (Jev answers, code decides), write (LLM then Jev), and output (code)" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude's part ends before the run starts. At run time the pipeline moves through four blocks: prepare (code), judge (Jev answers, code decides), write (an LLM writes, Jev checks), and output (code). Only the judgment block in the middle has the "Jev answers, code decides" shape.&lt;/p&gt;

&lt;p&gt;Where to fetch sources from, in what order, and how many: for now, code decides that. It uses the search terms written in advance, and nothing reads the results to decide where to look next. I think exploration is properly an LLM's job, ReAct-style: read the results, then pick the next place to look. The pipeline runs fine as it is, but depending on report quality I may hand exploration back to an LLM.&lt;/p&gt;

&lt;p&gt;Every model used at run time (Jev and the LLMs) is called through Pydantic AI's &lt;code&gt;Agent&lt;/code&gt;. I picked Pydantic AI for two reasons. It added native support for Jev's TypeSafe models early. And I didn't want to build an agent runtime from scratch: call the model, validate the output against a type, retry on failure. Swapping models takes one string. In fact, partway through I switched the run-time LLM that generates candidates (called through Alibaba's DashScope API) from Qwen's lightweight model to DeepSeek V4.1 Flash, because I had used up the free quota. The report body is written by qwen3.8-max.&lt;/p&gt;

&lt;p&gt;With Jev, once you write the output type, each field's type determines how Jev is asked.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pydantic type&lt;/th&gt;
&lt;th&gt;How Jev is asked&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;float&lt;/code&gt; (0–1)&lt;/td&gt;
&lt;td&gt;Probability that the answer to the question is "yes"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Literal[...]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Pick one of the options&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;IntEnum&lt;/code&gt; with docstrings&lt;/td&gt;
&lt;td&gt;Graded rating&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All the fields in one output type are sent together in a single request (&lt;a href="https://pydantic.dev/docs/ai/models/typesafe/" rel="noopener noreferrer"&gt;Pydantic AI's TypeSafe model docs&lt;/a&gt;, as of 2026-09-23).&lt;/p&gt;

&lt;h2&gt;
  
  
  Ask alongside a question
&lt;/h2&gt;

&lt;p&gt;To ask Jev "is this relevant?", you have to give it what the source should be relevant to. In this pipeline, that's the per-line questions. For each line I keep three to five questions I'm currently trying to answer in a file. On the line about AKC, a framework for how agents manage knowledge that I research, one of the questions reads like this (the file is in Japanese; translated here):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="c"&gt;&amp;lt;!-- excerpt from questions/akc.md (translated from Japanese) --&amp;gt;&lt;/span&gt;
&lt;span class="gu"&gt;## How is intent alignment between an agent and its operator maintained in the parts tests can't check?&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; brief: AKC's claim is that "a bidirectional growth loop maintains the alignment that tests can't check." How do designs that make similar claims (harness self-evolution, memory architectures, approval gates) detect alignment degrading, and on what grounds do they say it improved?
&lt;span class="p"&gt;-&lt;/span&gt; evidence: Measurements over time. Single-shot benchmarks are weak
&lt;span class="p"&gt;-&lt;/span&gt; not: Alignment training of the model alone (RLHF, etc.)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;brief&lt;/code&gt; is the scope of the question, &lt;code&gt;evidence&lt;/code&gt; is what counts as evidence, and &lt;code&gt;not&lt;/code&gt; lists topics that share keywords but aren't what this question is about. For this question, 21 papers with "alignment" in the title actually came through to judgment. Many of them only share the word, such as papers on mapping between audio and representations (cross-modal alignment). Under &lt;code&gt;not&lt;/code&gt;, I write the neighboring topics that are most easily confused with the question (here, alignment training of the model alone), so Jev can use them in its judgment.&lt;/p&gt;

&lt;p&gt;Every judgment is asked as a pair with a question. In step 4, for each source, Jev is asked cheaply, in one batch, whether it's likely relevant to each of the line's questions, to narrow the candidates. In step 5, each remaining "one source × one question" pair gets asked in detail. Step 6 is "one sentence × one question."&lt;/p&gt;

&lt;p&gt;The rest of this section is about step 5. Step 5 asks, on a four-step ladder, how close the problem a source works on is to the problem the question asks about. From the bottom: "a different problem that only shares vocabulary," "the same field but a different problem," "the same problem," and "answers the same question."&lt;/p&gt;

&lt;p&gt;For this intent-alignment question, Jev judged three papers, each with "intent" in the title, as follows.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Paper&lt;/th&gt;
&lt;th&gt;Probability relevant&lt;/th&gt;
&lt;th&gt;Most probable step&lt;/th&gt;
&lt;th&gt;That day's report&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://arxiv.org/abs/2609.23274" rel="noopener noreferrer"&gt;AI Persona, Service Consumption, and User Intent Entropy&lt;/a&gt; (a field experiment on AI personas and how much user intent varies)&lt;/td&gt;
&lt;td&gt;0.08&lt;/td&gt;
&lt;td&gt;Different problem (0.60)&lt;/td&gt;
&lt;td&gt;Not included&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://arxiv.org/abs/2609.24002" rel="noopener noreferrer"&gt;FinInteract&lt;/a&gt; (a benchmark for clarifying intent in ambiguous financial questions)&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;td&gt;Same field, different problem (0.53)&lt;/td&gt;
&lt;td&gt;Not included&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;a href="https://huggingface.co/papers/2609.06052" rel="noopener noreferrer"&gt;SkillSpec&lt;/a&gt; (has a model infer whether an agent's skill behaves as specified, with the intent withheld)&lt;/td&gt;
&lt;td&gt;0.44&lt;/td&gt;
&lt;td&gt;Same problem (0.77)&lt;/td&gt;
&lt;td&gt;Not included&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All three share the word "intent," but paired with the question, they land on different steps. SkillSpec was placed on "same problem," but its probability of being relevant was 0.44, short of the code's 0.5 threshold, so it didn't make the report. Jev only returns probabilities; code makes the accept or reject call with a threshold. The judgment values come from my local run logs, which aren't included in the repo.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9wxomktrkywm92dvaf97.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9wxomktrkywm92dvaf97.png" alt="Against the intent-alignment question, User Intent Entropy lands on different problem (0.60, relevant 0.08), FinInteract on same field (0.53, 0.20), and SkillSpec on same problem (0.77, 0.44); none reaches the 0.5 relevance threshold, so all three are rejected" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The figure puts the three papers on the ladder: User Intent Entropy on "different problem," FinInteract on "same field," and SkillSpec on "same problem," with the relevance probabilities 0.08, 0.20, and 0.44 all below the 0.5 threshold. The step Jev picks and the "relevant" probability that code applies a threshold to are separate answers.&lt;/p&gt;

&lt;p&gt;Here's the state I hand Jev (simplified from &lt;code&gt;state()&lt;/code&gt; in &lt;code&gt;question_screening.py&lt;/code&gt;). It gets serialized to JSON and passed to &lt;code&gt;agent.run()&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;line&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vocabulary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;vocabulary&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;brief&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;brief&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;method&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;method_constraints&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;evidence_constraints&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;negative_topics&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;source_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;  &lt;span class="c1"&gt;# the source's title, an excerpt of its text, URL, etc.
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;evidence_set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;evidence_set&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# evidence already accepted for this question
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The output type looks like this (excerpt; &lt;code&gt;Probability&lt;/code&gt; is a &lt;code&gt;float&lt;/code&gt; from 0 to 1). The &lt;code&gt;question.brief&lt;/code&gt; and &lt;code&gt;evidence_set&lt;/code&gt; in the descriptions refer to keys in the state above.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Overlap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;UseEnumMemberDocstrings&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;IntEnum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;other_problem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;It works on a different problem that happens to share vocabulary.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;same_field&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Same field as `question`, but not the problem `question` asks about.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;same_problem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;It works on the problem `question` asks about, from another angle.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;same_question&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;It asks what `question` asks and reports an answer to it.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Answers&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Screen one source against one open research question.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="n"&gt;on_topic&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Probability&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Is `source` about the problem `question` asks about, as `question.brief` &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;describes it? No if it is about one of `question.not` (neighbouring topics that keep &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;matching), or if it only shares a word with it.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;problem_overlap&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Overlap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How close is the problem `source` works on to the one `question` asks about?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;novelty_vs_evidence_set&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;NoveltyVsSet&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Compared with `evidence_set` (the claims already accepted for this &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question), what would `source` add?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The "most probable step" column in the table is this four-step &lt;code&gt;Overlap&lt;/code&gt;. The first paper fell to the bottom step: "a different problem that only shares vocabulary." Novelty is asked the same way, in terms of what the source would add to the evidence already accepted for this question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Numbers from the final run
&lt;/h2&gt;

&lt;p&gt;The report has one section per question whose evidence grew that day. The final run processed four lines. That run came before I switched to writing search terms in advance, so the search terms were generated by an LLM and picked by Jev.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Questions to Jev per line&lt;/td&gt;
&lt;td&gt;1,400–1,511 (14 for one line with few new sources)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Report size&lt;/td&gt;
&lt;td&gt;7.4–12.0 KB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude calls&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Numbers alone can't tell you whether the screening also dropped sources I should have read. So on the Jev-tracking line, I placed four sources I know should pass (canaries) under its questions, and all four passed in the final run. I haven't placed canaries on the other lines yet, and I haven't measured how much the pipeline misses across everything it drops.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing stays an LLM's job
&lt;/h2&gt;

&lt;p&gt;Outside the pipeline, I had an Opus instance, in a context separate from the session that built the pipeline, judge the quality of the reports written in step 7. Of the three lines whose body text was written in the final run, one was good enough to publish. The typical failure in the rejected drafts was recasting the evidence in the question's own terms.&lt;/p&gt;

&lt;p&gt;In the section for the question "When an agent breaks rules it wrote for itself, what lets that failure slip through?", the body presented a paper about false positives, which fabricate violations of rules that don't exist, as an explanation of misses that slip past the rules. Jev had let this paper through for this question just barely: probability relevant 0.54, with the step "same problem" at 0.51 and "same field, different problem" at 0.35. Letting it through wasn't wrong in itself. It was the body text that flipped the direction.&lt;/p&gt;

&lt;p&gt;Writing is an LLM's job, so this isn't something to move to Jev. The body got better after I fixed the instructions (don't let the evidence paragraph state conclusions) and turned on thinking for qwen3.8-max. I see the remaining weakness as a question of which model does the writing. The writer is also a Pydantic AI &lt;code&gt;Agent&lt;/code&gt;, so it can be swapped with one string for a model that's stronger at prose, such as GPT. I haven't measured yet whether swapping helps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to check before handing judgments to Jev
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Did you put what's being judged against into the state?&lt;/strong&gt; Without something to be relevant to, "is it relevant?" lets anything through. Pass the question, its scope (&lt;code&gt;brief&lt;/code&gt;), the topics to exclude (&lt;code&gt;not&lt;/code&gt;), and the evidence already accepted&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the answer closed?&lt;/strong&gt; Move only judgments that can be expressed as a yes/no probability, a choice among options, or a graded rating. For a graded rating, describe each step as a concrete situation (&lt;code&gt;same_field&lt;/code&gt; is "the same field, but not the question's problem")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is there a step where sources that only share vocabulary can land?&lt;/strong&gt; Without a "different problem" step, Jev has no choice but to push them onto a nearby step&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did you place sources that should pass?&lt;/strong&gt; A smaller volume on its own can't be told apart from dropping too much&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the writing LLM on the same &lt;code&gt;Agent&lt;/code&gt;, with its own model string?&lt;/strong&gt; Then you can swap just the writing model without changing the judgment types&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code is public at &lt;a href="https://github.com/shimo4228/jev-research-pipeline" rel="noopener noreferrer"&gt;jev-research-pipeline&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI-mediated writing disclosure:&lt;/strong&gt; AI drafted and translated the prose of this article from the author's pipeline code, run records, and the public sources cited above. The pipeline design, the choice of which judgments to move to Jev, and responsibility for publication belong to the author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj"&gt;How Close to Opus Does Jev, a Model That Writes No Text, Get at Skill Selection in 0.3 Seconds?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/shimo4228/i-added-jevs-skill-router-to-claude-code-and-turned-back-just-before-rewriting-the-skill-listing-34in"&gt;I Added Jev's Skill Router to Claude Code and Turned Back Just Before Rewriting the Skill Listing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/shimo4228/what-does-it-take-to-reproduce-jevs-decisions-locally-3i0n"&gt;What Does It Take to Reproduce Jev's Decisions Locally?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles/jev-research-judgment-offload.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Japanese original; this English version is &lt;code&gt;articles-en/jev-research-judgment-offload.md&lt;/code&gt;, and every article's Markdown and the index (docs/PUBLICATIONS.md) live in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;The author's GitHub&lt;/a&gt; — research repositories with DOIs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>jev</category>
      <category>pydanticai</category>
      <category>aiagents</category>
      <category>llm</category>
    </item>
    <item>
      <title>What Does It Take to Reproduce Jev's Decisions Locally?</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Tue, 22 Sep 2026 15:46:53 +0000</pubDate>
      <link>https://dev.to/shimo4228/what-does-it-take-to-reproduce-jevs-decisions-locally-3i0n</link>
      <guid>https://dev.to/shimo4228/what-does-it-take-to-reproduce-jevs-decisions-locally-3i0n</guid>
      <description>&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Agreement with Opus ceiling&lt;/th&gt;
&lt;th&gt;Time per row&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Jev (hosted)&lt;/td&gt;
&lt;td&gt;choice+noul&lt;/td&gt;
&lt;td&gt;0.346&lt;/td&gt;
&lt;td&gt;0.3 s&lt;/td&gt;
&lt;td&gt;✓ Baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma4:e4b&lt;/td&gt;
&lt;td&gt;logits x 54 questions&lt;/td&gt;
&lt;td&gt;0.162&lt;/td&gt;
&lt;td&gt;51 s&lt;/td&gt;
&lt;td&gt;Existing baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;qwen3:8b&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;logits x 54 questions&lt;/td&gt;
&lt;td&gt;— (stopped at 4 rows)&lt;/td&gt;
&lt;td&gt;43–74 s&lt;/td&gt;
&lt;td&gt;❌ Too slow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Laya&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;noul&lt;/td&gt;
&lt;td&gt;0.075&lt;/td&gt;
&lt;td&gt;39 s&lt;/td&gt;
&lt;td&gt;❌ Random-level&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;kev-0.8b&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;choice&lt;/td&gt;
&lt;td&gt;— (stopped at 25 rows)&lt;/td&gt;
&lt;td&gt;6.9 s&lt;/td&gt;
&lt;td&gt;❌ Out of memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;AFM 3 Core&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;verbalized&lt;/td&gt;
&lt;td&gt;— (20 rows only)&lt;/td&gt;
&lt;td&gt;20.7 s&lt;/td&gt;
&lt;td&gt;❌ Context too short&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;"Agreement with Opus ceiling" = Jaccard index measuring how much each model's selections overlap with those of Opus 5, treated as the upper bound. Details in "Measuring with 150-row replays." Random expectation: 0.056. Method descriptions follow in the next section. AFM = Apple Foundation Models (the on-device model in macOS 27). Its context length (the maximum text a model can process at once) is only 4,096 tokens — too short for the full prompt — so it was tested on just 20 rows with the catalog cut in half; AUC was 0.443. Jev's numbers are from &lt;a href="https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj"&gt;the previous article&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I tried to reproduce the skill-selection judgment that TypeSafe's decision model Jev solves in 0.3 seconds, using four candidates that run locally. As the table shows, all four failed. But the causes of failure split into three, and they reveal the conditions to evaluate next time a local decision model is on the table.&lt;/p&gt;

&lt;h2&gt;
  
  
  30+ open alternatives and 3 approaches
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj"&gt;the previous article&lt;/a&gt;, Jev reached about half the agreement ceiling — using Opus 5's picks as the upper bound — in 0.3 seconds per row for skill selection in an autonomous agent. But Jev is a closed API. As of September 2026, there are no open weights and no self-hosting path.&lt;/p&gt;

&lt;p&gt;Within a week of Jev's launch, over 30 open alternatives appeared (a catalog is at &lt;a href="https://systemonemodels.org/examples/alternatives/" rel="noopener noreferrer"&gt;systemonemodels.org&lt;/a&gt;). None of the trained decision-only models (kev, von, Laya, etc.) run on Ollama, though. Ollama requires GGUF-format weight files, and none of the decision-only models ship GGUF. On top of that, these models replace the standard text-generation output layer with a custom head that returns only select/don't-select probabilities — Ollama's general-purpose inference engine cannot handle that.&lt;/p&gt;

&lt;p&gt;The approaches you can try fall into three categories.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-approaches.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-approaches.png" alt="Three approaches to local decision models" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Logits method&lt;/strong&gt; — Ask a general-purpose LLM "Should we select this skill?" as a yes/no question and read the yes and no probabilities (log-probabilities = logits) via Ollama's &lt;code&gt;logprobs&lt;/code&gt; option. No fine-tuning needed, but you pay 54 calls for 54 skills&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Typed-decisions method&lt;/strong&gt; — Use a classifier trained specifically for decision tasks. Pass all 54 skills and the situation at once, and receive the "should select" probability for each skill in a single response. Instead of generating text like a general-purpose LLM, it outputs only an array of probabilities&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialized decoder method&lt;/strong&gt; — Replace the "text-generation part" of a general-purpose LLM with a "decision-only output part" and run the resulting small model through a dedicated server&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I picked one from each category. qwen3:8b is the smallest Ollama-compatible general-purpose LLM where prefix cache (explained below) works. Laya is the only decision-only model with a public Python API. kev-0.8b is the highest-profile project aiming to reimplement Jev. I also added Apple's on-device model (AFM 3 Core, Apple Foundation Models) available through the FoundationModels framework in macOS 27. AFM is not open, but it ships with macOS and requires no extra installation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring with 150-row replays
&lt;/h2&gt;

&lt;p&gt;The measurement method is the same as last time. I fed 150 rows of actual skill-selection logs from an autonomous agent into each model and compared the agreement with Opus 5's picks. The test environment is Apple Silicon M1 (16 GB) using MPS (Metal Performance Shaders — Apple's GPU compute framework, analogous to NVIDIA's CUDA), running Ollama 0.34.2.&lt;/p&gt;

&lt;p&gt;Three metrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Jaccard@topk&lt;/strong&gt; — The overlap between the skill sets Opus picked and the model picked. Calculated as "skills both picked / skills either picked." If Opus picked {A, B, C} and the model picked {A, B, D}, the overlap is {A, B} = 2, the union is {A, B, C, D} = 4, so Jaccard = 2/4 = 0.5. 1.0 means perfect agreement, 0 means zero overlap&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AUC&lt;/strong&gt; — Higher when the skills the model assigned high probabilities to are the ones Opus also picked. Measures ranking quality: 0.5 equals random, 1.0 means perfect ranking&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ECE (calibration error)&lt;/strong&gt; — How well a model's stated confidence matches its actual hit rate. When a model says "80% confident this should be selected," does it actually get it right 80% of the time? Calibration is the property of a model's probabilities matching real-world accuracy. ECE of 0 means perfectly calibrated. High ECE means you cannot trust the probability numbers for decision-making&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The skill-selection prompt runs about 12,000 characters (skill catalog ~9,600 characters + situation description median 1,537 characters). The maximum text a model can process at once is called its context length. This prompt size matters later.&lt;/p&gt;

&lt;h2&gt;
  
  
  qwen3:8b — 54 questions pile up
&lt;/h2&gt;

&lt;p&gt;First, the logits method. Ollama 0.34.2 returns the top-20 token probabilities with &lt;code&gt;logprobs: true&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; http://localhost:11434/api/generate &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"qwen3:8b","prompt":"...Should we select this skill? Answer yes or no:","stream":false,"logprobs":true,"top_logprobs":5,"options":{"num_predict":1,"temperature":0}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For each of the 54 skills, ask "Should we select this skill?" and read the yes log-probability.&lt;/p&gt;

&lt;p&gt;The 54 prompts share a common prefix — "skill catalog + situation description" (~12,000 characters) — with only the trailing "Should we select this skill?" part changing per question. Prefix cache holds the processing result of this shared part in memory so that from the second question onward, only the changed tail gets processed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-prefix-cache.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-prefix-cache.png" alt="How prefix cache works" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I first tried qwen3.5:9b, but even though the 54 prompts share a common prefix, every question after the first still took over 5 seconds. Prefix cache was not working. Switching to qwen3:8b, the second question with a different suffix returned in 0.4 seconds.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;qwen3.5:9b — same prompt resent: 0.3 s / different suffix: 7–12 s (full reprocessing)
qwen3:8b   — same prompt resent: 0.3 s / different suffix: 0.4 s (prefix cache active)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;LLMs compute relationships between tokens using an attention mechanism. Standard Transformer attention examines every pair of tokens, which is expensive, but the intermediate results (the KV cache) can be saved and reused. The hybrid linear attention that qwen3.5 uses trades cheaper computation for what I suspect is an inability to reuse KV cache from just the shared prefix. I switched to qwen3:8b, which uses standard Transformer attention.&lt;/p&gt;

&lt;p&gt;With cache active, the first question took 14.2 seconds; the rest had a median of 0.63 seconds. Per row: 43-74 seconds. The gemma enum method from the previous round passes all 54 skills at once in a single call and gets an answer in ~13 seconds per row; the logits method runs 54 serial calls, making it 3-5x slower. I cut it off at 4 rows.&lt;/p&gt;

&lt;p&gt;The serial execution of 54 calls is the fundamental bottleneck of the logits method. Even with prefix cache, 0.63 s x 54 = 34 seconds is the floor.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How to tell if prefix cache is working:&lt;/strong&gt; The &lt;code&gt;prompt_eval_count&lt;/code&gt; (input token count) in Ollama's response reports the original full token count whether or not the cache is active, so that number alone does not reveal cache status. If &lt;code&gt;prompt_eval_duration&lt;/code&gt; (time to process the input) drops significantly, the cache is working.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Laya — the wall called calibration
&lt;/h2&gt;

&lt;p&gt;Calibration, as explained in the previous section, is the property of a model's stated probabilities matching actual hit rates. "70% confident" means it really hits 7 out of 10 — that is what being calibrated means. Laya hit a wall here.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-calibration.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-calibration.png" alt="What calibration means" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Laya is a typed-decisions model based on mmBERT-base (322M params), available on &lt;a href="https://huggingface.co/convaiinnovations/laya" rel="noopener noreferrer"&gt;HuggingFace&lt;/a&gt;. I used two question formats:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;noul&lt;/strong&gt; — Returns yes/no probabilities per skill&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;choice&lt;/strong&gt; — Picks the best from a set of options, returning each option's probability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Results over 150 rows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Laya (noul)&lt;/th&gt;
&lt;th&gt;Laya (choice)&lt;/th&gt;
&lt;th&gt;gemma (logits)&lt;/th&gt;
&lt;th&gt;Random&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Jaccard@topk&lt;/td&gt;
&lt;td&gt;0.075 [0.062, 0.089]&lt;/td&gt;
&lt;td&gt;0.051 [0.041, 0.061]&lt;/td&gt;
&lt;td&gt;0.162 [0.141, 0.182]&lt;/td&gt;
&lt;td&gt;0.056&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AUC&lt;/td&gt;
&lt;td&gt;0.587 [0.561, 0.611]&lt;/td&gt;
&lt;td&gt;0.477 [0.455, 0.499]&lt;/td&gt;
&lt;td&gt;0.728 [0.704, 0.750]&lt;/td&gt;
&lt;td&gt;0.500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ECE&lt;/td&gt;
&lt;td&gt;0.469&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Brackets show 95% confidence intervals.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Noul's Jaccard of 0.075 barely exceeds random (0.056). Choice's confidence interval includes random — indistinguishable.&lt;/p&gt;

&lt;p&gt;The core problem is calibration. Breaking down ECE 0.469: of 8,207 total judgments, 6,382 (78%) fell in the 0.5-0.7 probability band, where the actual hit rate was 10-14%. "60% confident this should be selected" actually hits 1 in 10.&lt;/p&gt;

&lt;p&gt;Laya prints this warning at startup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;laya: this checkpoint ships temperatures outside [0.5, 5]
which would distort confidence; clamping choice:11+=0.1006.
Treat confidence from the affected buckets as uncalibrated.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What this means: Laya uses an internal temperature parameter (which adjusts the sharpness of the output probability distribution) to compute probabilities. In this checkpoint, cases with 11 or more options (bucket = group by number of options) have temperatures outside the normal range, so the probability values are force-corrected (clamped). Selecting from 54 skills falls into the 11+ range, so the returned probabilities are not calibrated.&lt;/p&gt;

&lt;p&gt;I also checked latency. The model card cites 33 ms per question (7 ms batched) on an NVIDIA T4 GPU, but my test environment uses Apple Silicon (MPS). Different hardware, so a direct comparison is not meaningful, but the measured values in this environment were: noul median 39.4 seconds per row (54 questions), minimum 6.8 seconds; choice 10.5 seconds, minimum 1.7 seconds.&lt;/p&gt;

&lt;p&gt;Laya's problem is quality, not speed. When probabilities cannot ground decisions, you cannot run a policy like "delegate only the high-confidence rows to Jev's replacement."&lt;/p&gt;

&lt;h2&gt;
  
  
  kev-0.8b — 0.8B devours 12 GB
&lt;/h2&gt;

&lt;p&gt;Last, the specialized decoder kev-0.8b. It mounts a decision head on Qwen3.5-0.8B-Base and runs through a dedicated server.&lt;/p&gt;

&lt;p&gt;Sending one request as designed (~6,000 tokens, choice + noul for 54 skills) immediately ran out of memory.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RuntimeError: MPS backend out of memory
(MPS allocated: 12.50 GiB, other allocations: 7.02 GiB,
 max allowed: 20.13 GiB).
Tried to allocate 880.00 MiB on private pool.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Splitting works. Choice alone (2,432 tokens) takes 8.6 seconds; noul split into batches of 14 takes 7-27 seconds. But running choice alone continuously, the MPS memory allocator failed to allocate even 16 KB after processing 17 rows. The server does not release GPU memory between requests.&lt;/p&gt;

&lt;p&gt;Two constraints are at play:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Inference kernel&lt;/strong&gt; (the program that executes the model's computation on the GPU): kev's &lt;code&gt;flash-linear-attention&lt;/code&gt; implementation targets NVIDIA (CUDA) and AMD (ROCm) GPUs and does not support Apple Silicon's MPS. It falls back automatically to the unoptimized reference implementation, printing a &lt;code&gt;"much slower"&lt;/code&gt; warning at server startup&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input length&lt;/strong&gt;: Training was done at 384 tokens or fewer. The serving config can set &lt;code&gt;num_ctx&lt;/code&gt; to 8,192, but the gap from the 12,000-character prompt to the training input length remains large&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I cut it at 25 rows, taking only the choice-alone results. Working around the allocator issue would require periodic server restarts, and I could not see a path to completing the full 150 rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  AFM 3 Core — context length falls short
&lt;/h2&gt;

&lt;p&gt;macOS 27 made the FoundationModels framework available. On the M1 (16 GB), the model that arrives is &lt;code&gt;AFM 3 Core&lt;/code&gt; (3B) with a 4,096-token context length. Too small a window for text generation, but a decision model's input and output are short, so I tried it on the off chance. The larger &lt;code&gt;AFM 3 Core Advanced&lt;/code&gt; (20B sparse) has no model-selection API in the SDK and does not come down to M1.&lt;/p&gt;

&lt;p&gt;The first wall is context length. The skill-selection prompt is about 12,000 characters, but AFM's context length is only 4,096 tokens. I cut the catalog in half and measured on just the first 20 rows.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;fm serve&lt;/code&gt; provides an OpenAI-compatible endpoint, but it silently ignores &lt;code&gt;logprobs&lt;/code&gt;. Using a verbalized approach (asking the model to state probabilities as numbers), AUC was 0.443 at temperature 0. The same 20 rows with gemma's logprobs scored 0.736. Five of the 20 rows hit a guardrail &lt;code&gt;RefusalError&lt;/code&gt;, leaving half the catalog unscored.&lt;/p&gt;

&lt;p&gt;The one advantage is co-residency. In-flight memory is about 2 GB, and it coexists with Ollama models without increasing swap. But if the prompt does not fit the window, it is not a candidate.&lt;/p&gt;

&lt;h2&gt;
  
  
  3 conditions a local decision model needs
&lt;/h2&gt;

&lt;p&gt;Line up the reasons each candidate dropped out, and three conditions for practical local decision models emerge.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-conditions.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fshimo4228%2Fzenn-content%2Fmain%2Fimages%2Flocal-decision-conditions.png" alt="Three conditions for a decision model" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Calibrated probabilities&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Laya returns probabilities, but you cannot split decisions by those numbers. ECE 0.469 means the probabilities are decorative. As the gap between Laya's typed-decisions benchmark (0.766) and its score on this task (0.075) shows, calibration depends on the task and the prompt structure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. An attention mechanism that supports prefix cache&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The 54 skills share a common prefix. Without cache, every question reprocesses 12,000 characters from scratch. qwen3.5's hybrid linear attention cannot do this reuse; kev's flash-linear-attention targets CUDA/ROCm and does not run on Apple Silicon. Both the kind of attention mechanism and the inference environment are conditions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Context length of 12,000 characters — and training at that length&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;kev was trained at 384 tokens or fewer. Stretching &lt;code&gt;num_ctx&lt;/code&gt; at serving time does not guarantee quality when the gap from training is that wide. AFM's context length is physically 4,096 tokens — the 12,000-character prompt does not fit. Input length must not only be "accepted by the config" but also "trained at that length," and the window must be wide enough.&lt;/p&gt;




&lt;p&gt;In response to these results, I introduced a seam called &lt;code&gt;DecisionBackend&lt;/code&gt; on the agent side. Setting the environment variable &lt;code&gt;DECISION_MODEL&lt;/code&gt; switches the decision surface to a local model — ready to swap in a model that meets the conditions.&lt;/p&gt;

&lt;p&gt;Looking back, Jev's performance is in a different league. Calibrated probabilities, context length that absorbs long inputs, an attention mechanism that supports prefix cache — it meets all three conditions, and on top of that returns 0.3 seconds per row at half the Opus ceiling. The fact that the four candidates each dropped out at a different condition throws into relief what Jev accomplishes at that price and speed.&lt;/p&gt;

&lt;p&gt;Four triggers would reopen the evaluation: kev implements an MLX backend, Laya gets fine-tuned on this task's data, Jev itself publishes open weights, or AFM's context length is extended. The moment any one of those moves, I re-measure with the same 150 rows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI-mediated writing disclosure:&lt;/strong&gt; AI drafted the English prose of this article from the author's Japanese original, measurement records, and terminal output. The measurements, the synthesis into three conditions, and publication responsibility belong to the author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj"&gt;Previous article: How Close to Opus Does Jev, a Model That Writes No Text, Get at Skill Selection in 0.3 Seconds?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://systemonemodels.org/examples/alternatives/" rel="noopener noreferrer"&gt;systemonemodels.org — Catalog of open alternatives&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/jaredpalmon/kev" rel="noopener noreferrer"&gt;kev (GitHub)&lt;/a&gt; / &lt;a href="https://huggingface.co/convaiinnovations/laya" rel="noopener noreferrer"&gt;Laya (HuggingFace)&lt;/a&gt; / &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;Laya (GitHub)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models" rel="noopener noreferrer"&gt;Apple Foundation Models — Third Generation&lt;/a&gt; / &lt;a href="https://pypi.org/project/apple-fm-sdk/" rel="noopener noreferrer"&gt;apple-fm-sdk (PyPI)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles-en/local-decision-model-conditions.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — every article's Markdown and the index (docs/PUBLICATIONS.md) live in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;The author's GitHub&lt;/a&gt; — research repositories with DOIs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>localllm</category>
      <category>aiagents</category>
      <category>llm</category>
      <category>benchmarking</category>
    </item>
    <item>
      <title>I Added Jev's Skill Router to Claude Code and Turned Back Just Before Rewriting the Skill Listing</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Mon, 21 Sep 2026 21:47:39 +0000</pubDate>
      <link>https://dev.to/shimo4228/i-added-jevs-skill-router-to-claude-code-and-turned-back-just-before-rewriting-the-skill-listing-34in</link>
      <guid>https://dev.to/shimo4228/i-added-jevs-skill-router-to-claude-code-and-turned-back-just-before-rewriting-the-skill-listing-34in</guid>
      <description>&lt;p&gt;I try not to fill the gaps in a big harness like Claude Code with my own implementation. When I do fill one, the official side changes a while later and my work is no longer needed. That has been the pattern for the last six months.&lt;/p&gt;

&lt;p&gt;The rules I had accumulated to shore up an older generation of models stopped being useful when Anthropic changed course for the Opus 5 generation — fewer rules, leave more to the model's judgment — so I cut my resident set &lt;a href="https://dev.to/shimo4228/opus-5-changed-how-rules-should-be-written-audit-yours-4fb4"&gt;from 5,789 words to 2,463&lt;/a&gt;. To run several Claude Codes side by side I &lt;a href="https://dev.to/shimo4228/herdr-a-tmux-for-ai-agents-until-the-editor-disappeared-3hnn"&gt;installed Herdr and wrote an article about it&lt;/a&gt;, and now Claude Code itself can &lt;a href="https://code.claude.com/docs/en/cross-session-messaging" rel="noopener noreferrer"&gt;send messages between sessions&lt;/a&gt;. &lt;a href="https://code.claude.com/docs/en/claude-projects" rel="noopener noreferrer"&gt;Projects&lt;/a&gt;, where Claude starts and manages parallel sessions from a single conversation, is rolling out in public beta. Today's Projects bundles only cloud sessions, but once it handles the sessions on my own machine, Herdr will not be needed either.&lt;/p&gt;

&lt;p&gt;Even so, on September 21, 2026, curiosity won. Jev is a model that writes no text and returns only probabilities to typed questions. That day I ported the &lt;a href="https://docs.typesafe.ai/cookbooks/skill_suggestion.md" rel="noopener noreferrer"&gt;skill suggestion cookbook&lt;/a&gt; from TypeSafe, who build Jev, to a Claude Code hook.&lt;/p&gt;

&lt;p&gt;This article is about how far the thing I built reached, where I turned back, and what was waiting past the turn. If you are thinking about adding Jev to your own harness, you should come away with three things to check before you do.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I added is one line
&lt;/h2&gt;

&lt;p&gt;What I built is a Claude Code plugin, &lt;a href="https://github.com/shimo4228/jev-skill-router" rel="noopener noreferrer"&gt;jev-skill-router&lt;/a&gt;. On every prompt a &lt;code&gt;UserPromptSubmit&lt;/code&gt; hook runs and sends Jev the prompt together with the roster of skills installed locally — their names and descriptions. It makes at most two requests to Jev. The first ranks the whole roster and asks three yes/no questions about whether the request needs a skill at all. The second re-reads only the top three candidates, this time with the full text of each &lt;code&gt;SKILL.md&lt;/code&gt;. If neither the average of the three first-pass answers (&lt;code&gt;gate&lt;/code&gt; in the log; the third is inverted before averaging) nor the highest per-candidate value from the second pass (&lt;code&gt;fits&lt;/code&gt;) reaches 0.30, it suggests nothing.&lt;/p&gt;

&lt;p&gt;By default it adds nothing and only records its verdict in a log. With injection enabled, a turn that reaches 0.30 gets exactly this one line added.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;lt;skill_relevance&amp;gt;
Relevant to the current request: adr-writer. Ignore this if it does not fit what the user actually asked for.
&amp;lt;/skill_relevance&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each verdict leaves one line in the log. This is the line from a request, written in Japanese, to record a design decision in an ADR (excerpt).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"inject"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jev-1.13.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"router_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"gate"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.47&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"shortlist"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"adr-writer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"adhd:adhd"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"archify"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"fits"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"adr-writer"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.93&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"adhd:adhd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.14&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"archify"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.05&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"suggestion"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"adr-writer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"usage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;21597&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;702&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"calls"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"elapsed_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1567&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I wrote six request sentences with a single intent each and sent them. Five named the skill I judged to be right, at 0.93 to 0.98, and "Thanks, that's it for today" drew no suggestion at all. Each took 0.7 to 1.6 seconds. These are single readings per case, not numbers you can read as a rate.&lt;/p&gt;

&lt;p&gt;Up to here, it worked exactly as the cookbook said it would.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Code's own skill selection was running the whole time
&lt;/h2&gt;

&lt;p&gt;While running it, I asked: "Hold on — so Claude Code's own skill selection is still running at the same time?"&lt;/p&gt;

&lt;p&gt;It is. All a hook's &lt;code&gt;additionalContext&lt;/code&gt; can do is &lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;add one string to the model's context&lt;/a&gt;. Claude Code goes on showing the model a skill listing with the name and description of every installed skill, and the model is what picks from it. Not one token of that listing goes away. Jev's verdict was simply running alongside the choice.&lt;/p&gt;

&lt;p&gt;On top of that, the conditions that made the cookbook work are absent in Claude Code. The cookbook's experiment covers 488 requests; across the 315 where a skill applied, loading the wrong skill dropped from 16.8% to 7.3%. The model choosing skills there was &lt;code&gt;claude-haiku-4-5&lt;/code&gt;, and what it saw was an index cut to 60 characters per skill. The cookbook says so at the top:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Hermes, the agent harness used here, cuts it to 60 characters by default.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;60 characters is not a recommendation from the cookbook; it is the default display width of the harness used in the experiment. The cookbook itself gives the example that at 60 characters you cannot tell a skill that edits &lt;code&gt;.pptx&lt;/code&gt; from one that creates it. It describes the problem it set out to solve as an agent that "makes its choice on almost no information." A small model choosing from a truncated index, with one line added by a judge that had read the full text. Those were the conditions under which it worked.&lt;/p&gt;

&lt;p&gt;Claude Code is the opposite. It shows the model each description (together with &lt;code&gt;when_to_use&lt;/code&gt;) &lt;a href="https://code.claude.com/docs/en/skills" rel="noopener noreferrer"&gt;up to 1,536 characters, uncut&lt;/a&gt;. Counting in the repository where I am writing this: 59 skills installed, 2 with descriptions that fit in 60 characters, median 412. The model doing the choosing is Claude Fable 5.1, stronger than Haiku. A router in Claude Code turns into a second model that has read the same descriptions, advising from the side a stronger model that has already read all of them in full. The only information it can add is the body of &lt;code&gt;SKILL.md&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is an argument from the mechanism, not a measured result. I have no evidence that injecting the line improved anything. That is why I did not make injection the plugin's default.&lt;/p&gt;

&lt;h2&gt;
  
  
  There was a path to rewriting the skill listing
&lt;/h2&gt;

&lt;p&gt;To make it work as a router, adding a line is not enough. You have to change the skill listing the model is shown.&lt;/p&gt;

&lt;p&gt;That path existed. The type definitions for function hooks — Claude Code's early-access feature, known as Mods (&lt;a href="https://github.com/anthropics/claude-code/issues/91870" rel="noopener noreferrer"&gt;design thread&lt;/a&gt;) — list &lt;code&gt;skill_listing&lt;/code&gt; among the kinds of attachment that reach the model, and say its body can be rewritten from the Mod side (&lt;code&gt;mods/types/claude-code.d.ts&lt;/code&gt; in &lt;code&gt;anthropics/claude-code&lt;/code&gt;, checked on September 21, 2026).&lt;/p&gt;

&lt;p&gt;I started a design plan and stopped almost immediately. With enough effort it might not be impossible. But it means replacing Claude Code's default skill selection with my own. An implementation that rides on early-access types and breaks the default behavior has to follow along every time the official side changes. After six months of filling gaps only for the official side to change and make the fill unnecessary, this is a swamp.&lt;/p&gt;

&lt;p&gt;I wrote at the top of the README, with the reasons, that "as a router it is unlikely to help a strong model in Claude Code," left it there for anyone thinking the same thing, and stopped.&lt;/p&gt;

&lt;h2&gt;
  
  
  That path had been walked the day before
&lt;/h2&gt;

&lt;p&gt;Searching again while preparing this article, I found that what lay past my turning point was already implemented.&lt;/p&gt;

&lt;p&gt;It is the Mod &lt;code&gt;jev-skill-suggestion&lt;/code&gt; in &lt;a href="https://github.com/davila7/claude-code-templates" rel="noopener noreferrer"&gt;davila7/claude-code-templates&lt;/a&gt;. Its first commit was on the morning of September 20, 2026, Japan time — the day before I turned back. At that point the very rewrite I had stopped short of was working. It returns empty for Mods' &lt;code&gt;skill_listing&lt;/code&gt;, so the model never reads the skill listing at all. Instead it tells the model to load the single skill Jev picked. The model can no longer choose from the listing itself; what remains is to follow Jev's suggestion or ignore it. A second commit, at midday on the day I turned back, went further: the Mod attaches the chosen skill's &lt;code&gt;SKILL.md&lt;/code&gt; itself. The first line of the current README reads:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Takes the skill listing out of the context window and lets Jev, TypeSafe's System One decision model, pick at most one skill per prompt from the skills' descriptions — and then loads that one skill itself, by attaching its &lt;code&gt;SKILL.md&lt;/code&gt; to the prompt.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The way it picks is the same two requests from the same cookbook. Hiding the skill listing was not custom work: it is the official setting &lt;a href="https://code.claude.com/docs/en/skills" rel="noopener noreferrer"&gt;&lt;code&gt;skillOverrides&lt;/code&gt;&lt;/a&gt;, which switches how each skill is shown to the model across four levels — name and description, name only, only when the user calls it, or not shown. If all you want is to cut the tokens the skill listing costs, you need neither Jev nor Mods. The same repository also holds &lt;code&gt;jev-model-router&lt;/code&gt;, which routes model and reasoning depth with Jev.&lt;/p&gt;

&lt;p&gt;I have tried neither. Whether they work, I do not know. What I do know is that someone willing to carry the cost of following the official side was already there before I stopped. If I want to try it later, the real thing exists without my building it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who holds the role of choosing
&lt;/h2&gt;

&lt;p&gt;From here on, this is my own read, from having built the thing.&lt;/p&gt;

&lt;p&gt;In Claude Code, the role of choosing skills belongs to Claude Code and its model. A frontier model carries it, and the only opening I had was one for adding a line of text.&lt;/p&gt;

&lt;p&gt;How far Jev works, I think, is decided by who holds the role of choosing. If a weak chooser holds it, one line added from the side is enough — that was the cookbook. If a strong chooser is looking at the full text, advice does not move it, and the only way through is to reach the role of choosing itself. Mods is an example of the harness side opening that role outward. &lt;a href="https://github.com/TheoOliveira/pi-jev" rel="noopener noreferrer"&gt;pi-jev&lt;/a&gt;, an extension for the coding agent Pi, enables only the tools whose Jev probability clears a threshold. Not advice: the tools the model can see are what change.&lt;/p&gt;

&lt;p&gt;Most harnesses today are built around the loop where the model thinks and then calls a tool (ReAct), with a strong model seeing everything and doing both the choosing and the judging inside its own reasoning. As long as the chooser is strong, adding a fast decision model beside it leaves no work to be replaced. Moves to fit Jev partially into existing harnesses will probably keep coming, but the effect should stay inside whatever range the harness has opened.&lt;/p&gt;

&lt;p&gt;Conversely, if a harness appears that builds tool selection, model routing, and the judgments along the way with Jev from the start, and hands only the final reasoning to a frontier model, speed, cost, and accuracy could all look very different. TypeSafe itself calls Jev "a frontier-intelligence function call" in &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;its announcement&lt;/a&gt;, framing it as a component you call. When I &lt;a href="https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj"&gt;compared Jev and Opus&lt;/a&gt; on skill selection in a different agent of my own, Jev took 0.32 seconds per case at about 1/560th of Opus's cost, and across the 45 cases where it pointed with probability 0.5 or higher, not one disagreed with Opus's pick. On the other hand, agreement across all 150 cases was about half of the agreement Opus reached with itself when made to do the same selection twice. How much of this you can hand to Jev, I have not measured, and I have not found a report that measures it.&lt;/p&gt;

&lt;p&gt;As of September 21, 2026, within the range I searched, no such general-purpose harness existed. Everything I found is an integration into an existing harness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things to check before you add
&lt;/h2&gt;

&lt;p&gt;If you are about to add a decision model like Jev to your own harness, check three things before you start writing.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Does what you add reach the role of choosing?&lt;/strong&gt; If you only add text, the original choice keeps running underneath. Find out first whether there is an opening that reaches it — your own loop, Mods, tool enablement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the chooser already seeing the same information as the decision model?&lt;/strong&gt; Before you read the numbers from a vendor's experiment, read the experiment's conditions. Which model was choosing? What could that chooser see?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is an official setting enough?&lt;/strong&gt; If what you want is fewer tokens spent on the skill listing, &lt;code&gt;skillOverrides&lt;/code&gt; is enough in Claude Code. A decision model is for what lies past that.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An extension that only adds text goes back the way it was when you remove it. Build one out of curiosity if you like. An implementation that replaces the default behavior is something I will not build, even if I could.&lt;/p&gt;

&lt;p&gt;I moved over to following instead of building. I added "does a harness built on Jev appear" to my daily research checklist. My plugin keeps running in the setting where it adds nothing and only writes the log. What I watch is the match between the skill Jev named and the skill actually used in that turn. If turns where it named a skill that went unused pile up, adding one line may be worth something. If they do not pile up, I take it out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI-mediated writing disclosure:&lt;/strong&gt; AI drafted and translated the prose of this article from the author's session records, the plugin's logs and README, and the public sources cited above. The central thesis, the decision to turn back, the read in "Who holds the role of choosing," and responsibility for publication belong to the author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj"&gt;How Close to Opus Does Jev, a Model That Writes No Text, Get at Skill Selection in 0.3 Seconds?&lt;/a&gt; — the measurement comparing Jev and Opus on skill selection in an agent of my own (Contemplative Agent, running on gemma4:e4b), not in Claude Code&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.typesafe.ai/cookbooks/skill_suggestion.md" rel="noopener noreferrer"&gt;TypeSafe: Skill suggestion cookbook&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/davila7/claude-code-templates" rel="noopener noreferrer"&gt;davila7/claude-code-templates&lt;/a&gt; — &lt;code&gt;jev-skill-suggestion&lt;/code&gt; and &lt;code&gt;jev-model-router&lt;/code&gt; are Mods in this repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/DECRUX9812/typesafe-skill-router" rel="noopener noreferrer"&gt;DECRUX9812/typesafe-skill-router&lt;/a&gt; / &lt;a href="https://github.com/Dicklesworthstone/skillranker" rel="noopener noreferrer"&gt;Dicklesworthstone/skillranker&lt;/a&gt; — other implementations of the same cookbook&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles/jev-retrofit-limits.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Japanese original; this English version is &lt;code&gt;articles-en/jev-retrofit-limits.md&lt;/code&gt;, and every article's Markdown and the index (docs/PUBLICATIONS.md) live in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;The author's GitHub&lt;/a&gt; — research repositories with DOIs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>aiagents</category>
      <category>agenticcoding</category>
      <category>llm</category>
    </item>
    <item>
      <title>How Close to Opus Does Jev, a Model That Writes No Text, Get at Skill Selection in 0.3 Seconds?</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Mon, 21 Sep 2026 01:03:01 +0000</pubDate>
      <link>https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj</link>
      <guid>https://dev.to/shimo4228/how-close-to-opus-does-jev-a-model-that-writes-no-text-get-at-skill-selection-in-03-seconds-1nfj</guid>
      <description>&lt;p&gt;I took the step an autonomous agent runs before it acts — picking, out of a catalog, the skills that apply to the situation in front of it — and had Jev and Claude Opus do it over the same 150 cases. Jev is the model TypeSafe opened early access to on September 15, 2026. It writes no text; it returns answers over a set of options and the probabilities behind them (&lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;TypeSafe's announcement&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;The yardstick is Opus itself. Run the same 150 cases through Opus a second time, and its agreement with the first run is 0.678. Picking at random gives 0.056. Jev came in at 0.346 — about half of Opus's agreement with itself, and more than twice the small model on my machine. The skill Jev ranked first is one Opus picked too, in 83% of the cases. Across the 45 cases where it ranked something first with probability 0.5 or higher, not one disagreed with Opus. Time is 0.3 seconds per case, and the 150 cost about five cents: about 1/50th of Opus's time and about 1/560th of its cost.&lt;/p&gt;

&lt;p&gt;Once Jev runs on my own machine, I want to move most of this agent's processing of the same kind over to it. Jev today is hosted, and this agent, which I keep complete on one machine at hand, cannot take it yet.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Picker&lt;/th&gt;
&lt;th&gt;Agreement with Opus&lt;/th&gt;
&lt;th&gt;Time per case&lt;/th&gt;
&lt;th&gt;Cost for 150&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus (second run)&lt;/td&gt;
&lt;td&gt;0.678&lt;/td&gt;
&lt;td&gt;15.3 s&lt;/td&gt;
&lt;td&gt;$26.26&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Jev&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.346&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.32 s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.047&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;local gemma4:e4b&lt;/td&gt;
&lt;td&gt;0.143–0.162&lt;/td&gt;
&lt;td&gt;9.8 s&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Picked at random&lt;/td&gt;
&lt;td&gt;0.056&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The baseline is Opus's first run. The range for gemma is the range across eight output configurations I tried. Opus's picks are the baseline, not the ground truth. Time and cost were not measured under matched conditions, so I put them under "What this comparison cannot say."&lt;/p&gt;

&lt;h2&gt;
  
  
  One skill selection, in full
&lt;/h2&gt;

&lt;p&gt;The subject is an autonomous agent I run (Contemplative Agent). Before it acts, the agent receives a description of the situation and a catalog of skills — name and description pairs — and picks the ones that apply. It may pick several, and it may pick none. The catalog runs 53 to 57 entries depending on the period.&lt;/p&gt;

&lt;p&gt;Right now a small model running on my Mac (gemma4:e4b) writes out the names of the skills that apply. Because it picks by generating text, each run takes around ten seconds, and 20–30% of the time it produces a misspelled skill name. A model that picks without writing came out, so I tried it.&lt;/p&gt;

&lt;p&gt;The agent logs the input and output of every skill selection. Out of those logs I pulled 75 cases that contained a misspelled skill name and 75 that did not; those are the 150. From here on I call one case one row.&lt;/p&gt;

&lt;p&gt;Look at one of those rows. The catalog for this row holds 54 skills.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Skill&lt;/th&gt;
&lt;th&gt;Opus&lt;/th&gt;
&lt;th&gt;Jev's probability&lt;/th&gt;
&lt;th&gt;local gemma&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;analogy-mapping-relationships&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;td&gt;0.61&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;analyzing-systemic-governance-loops&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;td&gt;0.12&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;trace-structural-authority&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;td&gt;0.07&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cross-reference-foundational-claims&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.02 or less&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;detecting-abstract-to-operational-constraint-shift&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.02 or less&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;scope-boundary-mapping&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.02 or less&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;shifting-focus-from-state-to-process-mechanics&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;0.02 or less&lt;/td&gt;
&lt;td&gt;picked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;identify-structural-tensions-via-system-metaphors&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;misspelled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Opus picked the same three skills both times. The three Jev gave the highest probabilities are those same three. gemma wrote six names, one of them not in the catalog. The real one is &lt;code&gt;identifying-…&lt;/code&gt;; the word form is broken. Names that are not in the catalog get dropped by the code, so gemma's selection comes to five.&lt;/p&gt;

&lt;p&gt;When this article says "agreement with Opus," it means the number both picked divided by the number either picked (the Jaccard index). On this row gemma shares one and either-picked seven: 1÷7, or 0.14. For Jev, taking the same number of skills gemma picked from the top of the probability ranking (how that number is set comes in the next section), all three of Opus's picks are in, giving 3÷5, or 0.60.&lt;/p&gt;

&lt;p&gt;This row is one of Jev's best. gemma's 0.14 happens to be the same value as its average across the 150 rows. One row says nothing, so I count across 150.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I asked Jev
&lt;/h2&gt;

&lt;p&gt;I used two kinds of question. What TypeSafe's docs call Choice — pick one out of a set of options — and what they call Noul — a yes-or-no question. I made each row one request, with one Choice question and one Noul question per skill.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"situation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"(the situation description)"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jev-1.13.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"questions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"(the selection criteria) Which single learned skill applies best to `situation`?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"analogy-mapping-relationships"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"(this skill's description)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"none of the above"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"No skill in the catalog applies to this situation."&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"n0000"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"(the selection criteria) Does the learned skill `analogy-mapping-relationships — (description)` apply to `situation`?"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In practice &lt;code&gt;criteria&lt;/code&gt; holds every skill and the Nouls run as long as the skill list, so one request carries about 55 questions. Jev's questions cannot reference one another, so the selection criteria and each skill's description go into every question. Input came to about 7,450 tokens per row.&lt;/p&gt;

&lt;p&gt;Choice returns a probability for every skill, but it does not return how many to pick. So I took, from the top of the probabilities, the same number of skills gemma picked on that row (6.0 on average) and made that the set. Opus decides the count itself (6.0 on average in its first run). The Jev row in the table at the top is this Choice value.&lt;/p&gt;

&lt;h2&gt;
  
  
  Across 150 rows, how close to Opus
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Picker&lt;/th&gt;
&lt;th&gt;Skills picked (average)&lt;/th&gt;
&lt;th&gt;Agreement with Opus&lt;/th&gt;
&lt;th&gt;Agreement with near-skill credit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus (second run)&lt;/td&gt;
&lt;td&gt;5.3&lt;/td&gt;
&lt;td&gt;0.678&lt;/td&gt;
&lt;td&gt;0.924&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Jev&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6.0 (gemma's count)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.346&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.855&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;local gemma4:e4b&lt;/td&gt;
&lt;td&gt;6.0–8.1&lt;/td&gt;
&lt;td&gt;0.143–0.162&lt;/td&gt;
&lt;td&gt;0.755–0.784&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random, same count&lt;/td&gt;
&lt;td&gt;6.0&lt;/td&gt;
&lt;td&gt;0.056&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Even asking the same Opus twice, agreement tops out at 0.678. That is the ceiling for this way of counting, and the floor is random at 0.056. Jev's 0.346 sat at about half the ceiling, more than twice gemma. gemma did not move off 0.14–0.16 whether I constrained its output with an enum or set temperature to 0 (the gemma-side record is in &lt;a href="https://github.com/shimo4228/contemplative-agent/blob/1ec4d2dcf7973e0677ded25cd36f21fe91a48c3d/docs/evidence/rfc-0043/README.md" rel="noopener noreferrer"&gt;the evidence in the public repository&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Jaccard counts a neighboring, similar skill as a miss. The right-hand column measures closeness with embeddings of the skill descriptions and gives partial credit for similar skills. For each skill Opus picked, I averaged its closeness to the nearest one in the other side's selection. Every skill resembles every other to some degree, so even random scores 0.70. Between that floor of 0.70 and the ceiling of 0.924, Jev is at 0.855 and gemma at 0.76–0.78. Allow similar skills and Jev moves a bit closer to Opus's agreement with itself.&lt;/p&gt;

&lt;p&gt;Look at the individual picked skills rather than the agreement between sets, and it gets simpler.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;Share Opus also picked&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Skills Opus's second run picked (5.3 on average)&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jev's top pick (the 147 rows where Opus picked anything)&lt;/td&gt;
&lt;td&gt;83%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;of those, the 45 rows where the top probability was 0.5 or higher&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jev's top picks taken out to gemma's count (6.0 on average)&lt;/td&gt;
&lt;td&gt;50%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills the local gemma picked (6.0 on average)&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A single random pick&lt;/td&gt;
&lt;td&gt;11%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Jev's top pick, at 83%, is level with the 85% for what Opus's second run picked. But Opus's 85% is measured over 5.3 picks on average, while Jev's 83% covers only its single most confident pick. Widen Jev out to six and it falls to 50%. In 5 of the 147 rows, Jev's top pick was "none of the above," which I count as a miss.&lt;/p&gt;

&lt;p&gt;Across the 45 rows where Jev pointed at something with probability 0.5 or higher, it never once disagreed with Opus. The probability works as a cutoff separating the rows you can trust from the rows you cannot. If you take Jev's pick as-is in place of a large model, take only the high-probability rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ask one at a time, and the skills that read broadly come out on top on most rows
&lt;/h2&gt;

&lt;p&gt;I ran the same comparison on the probabilities from the Nouls in that same request (per skill: does this apply?), again taking the same number from the top.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;th&gt;Noul&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Agreement with Opus&lt;/td&gt;
&lt;td&gt;0.346&lt;/td&gt;
&lt;td&gt;0.295&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agreement with near-skill credit&lt;/td&gt;
&lt;td&gt;0.855&lt;/td&gt;
&lt;td&gt;0.836&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rows where &lt;code&gt;shifting-focus-from-state-to-process-mechanics&lt;/code&gt; made the top&lt;/td&gt;
&lt;td&gt;34%&lt;/td&gt;
&lt;td&gt;77%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Noul put that one skill in the top on 77% of rows. It is the skill gemma picked in the row shown above. Opus picked it on 16% of rows in the first run and 14% in the second. It is not that Noul says yes to everything: 35% of its probabilities are 0.5 or higher, and the median is 0.38. Noul looks at one skill at a time, so it has nothing to compare against. My reading is that a skill whose scope can be read broadly comes out on top in any situation. With Choice, which weighs every skill against the others, it drops to 34%.&lt;/p&gt;

&lt;p&gt;For skill selection, Choice was the better fit. TypeSafe's &lt;a href="https://docs.typesafe.ai/cookbooks/skill_suggestion.md" rel="noopener noreferrer"&gt;cookbook for skill suggestion&lt;/a&gt; is two-stage: rank broadly with Choice, then re-judge only the top with Noul. Applying a per-skill Noul to every skill, the way I asked, falls outside that usage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five cents and 0.3 seconds
&lt;/h2&gt;

&lt;p&gt;Opus took 15.3 seconds per row (median) and $26.26 for the 150. Jev took 0.32 seconds and $0.047: about 1/50th the time, about 1/560th the cost. All I had Opus return was a JSON array of skill names, and even so the output ran to a median of 1,373 tokens. Jev has no output tokens. The difference in time is the difference between generating tokens and not generating them. The difference in cost also carries the difference in input pricing.&lt;/p&gt;

&lt;p&gt;Jev had 0 failures across the 150 requests, with a median of 0.32 seconds, 0.53 seconds at the 90th percentile, and a maximum of 4.5 seconds. That includes the network round trip from Japan. Input totaled 1,125,726 tokens; multiplied by the published rate ($0.042 per million input tokens, output free), that comes to $0.047. It is a computed figure, not a billed amount. Opus I called one row at a time with &lt;code&gt;claude -p&lt;/code&gt;; the cost is the API-equivalent amount its response returns, and the time is the API-side duration.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this comparison cannot say
&lt;/h2&gt;

&lt;p&gt;The baseline is Opus's picks, not the ground truth. I took no human labels. What I can say goes as far as "Jev is close to the way Opus picks" — a different claim from "Jev's picks are right." And gemma's picks being far from Opus's does not mean gemma is wrong.&lt;/p&gt;

&lt;p&gt;Conditions unfavorable to Jev are mixed in with repurposing of my own. The situation text contains Japanese, and TypeSafe itself writes that languages other than English are &lt;a href="https://docs.typesafe.ai/models.md" rel="noopener noreferrer"&gt;not handled equally&lt;/a&gt;. Choice is a question meant to pick one, and using its probabilities as a ranking is my repurposing.&lt;/p&gt;

&lt;p&gt;Nor is the handoff matched. Opus got the same situation, the same catalog and the same selection criteria, but through a different template from gemma's production one. Jev got the JSON above. The timing conditions differ too. Jev's 0.32 seconds includes the round trip from Japan, gemma's 9.8 seconds is on a 16GB M1, and Opus's is the API-side duration.&lt;/p&gt;

&lt;p&gt;The subject is one process in one agent, 150 rows. Half were drawn from rows that contained a misspelled skill name, so this is not the production distribution as it stands. The row data contains third-party posts, so I am not publishing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Once it runs on my machine, I want to move processing of this kind to Jev
&lt;/h2&gt;

&lt;p&gt;This agent has several other steps that only pick, or only judge. Right now every one of them runs on the local gemma by generating text. Looking at these results, if Jev ran on my machine I would want to move most of them over. It lands more than twice as close to Opus as the gemma I run today, in 0.3 seconds, with no misspelled names and a probability attached.&lt;/p&gt;

&lt;p&gt;What keeps it out today is that Jev is hosted. This agent is designed and run so that skill selection and text generation alike stay complete on one machine at hand. Having no path out to the outside in the production code is this agent's premise for safety. The 16GB M1 and the small model are &lt;a href="https://dev.to/shimo4228/building-an-autonomous-agent-on-an-m1-mac-by-choice-5b5o"&gt;a constraint I chose&lt;/a&gt; as well. Jev has neither open weights nor a self-hosting path yet. Once it can run on a local runtime like Ollama, it is the first thing I will try.&lt;/p&gt;

&lt;p&gt;For this measurement I sent the 150 rows of situation text to both Jev and Opus. That was a measurement I permitted once, from a script separate from the production code. I judged it a different thing from production sending data out on every run.&lt;/p&gt;

&lt;p&gt;There is no shortcut where I take logs of Jev's picks as a teacher and train the small model on my machine. Section 2.3(b) of TypeSafe's &lt;a href="https://typesafe.ai/legal/mca" rel="noopener noreferrer"&gt;Master Customer Agreement&lt;/a&gt; (updated September 19, 2026) prohibits distilling a model from the outputs and training a model to imitate them.&lt;/p&gt;

&lt;p&gt;Skill selection turned out to be a process that does not need text written for it. A model that only attaches probabilities to options came, in 0.3 seconds, to about half of Opus's agreement with itself. If your agent's design allows calling an external API, there is no reason to wait. Taking the rows Jev points at with high probability as they are, and sending the rows where the probability falls short of 0.5 (about 70% here) to a bigger model, is an idea worth trying.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI-mediated writing disclosure:&lt;/strong&gt; AI drafted the prose of this article from the author's measurement records, the aggregated result files, and the public evidence in the repository. The central thesis, the choice of Opus as the yardstick, the judgment about adopting Jev, and publication responsibility belong to the author.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/shimo4228/contemplative-agent/blob/1ec4d2dcf7973e0677ded25cd36f21fe91a48c3d/docs/evidence/rfc-0043/README.md" rel="noopener noreferrer"&gt;The gemma-side tallies and reading (evidence in the public repository)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles-en/jev-vs-opus-skill-selection.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — every article's Markdown and the index (docs/PUBLICATIONS.md) live in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;The author's GitHub&lt;/a&gt; — research repositories with DOIs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aiagents</category>
      <category>llm</category>
      <category>localllm</category>
      <category>benchmarking</category>
    </item>
    <item>
      <title>I Made the Top Model My Session Default, and One Heavy Implementation Plus Its Review Burned Through Fable's Usage Limit</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Wed, 16 Sep 2026 12:12:21 +0000</pubDate>
      <link>https://dev.to/shimo4228/i-made-the-top-model-my-session-default-and-one-heavy-implementation-plus-its-review-burned-4182</link>
      <guid>https://dev.to/shimo4228/i-made-the-top-model-my-session-default-and-one-heavy-implementation-plus-its-review-burned-4182</guid>
      <description>&lt;p&gt;On August 22, 2026, I started a heavy implementation in a Fable session and ran a review right after it. Fable's usage limit was gone in an instant. That happened because the built-in skill that handles the review inherited the session's model and ran on it.&lt;/p&gt;

&lt;p&gt;Three days later, in the morning, I was chasing a different problem and noticed something. The over-engineering that shows up in my own repositories was happening in the Opus sessions — the ones that picked up design work after Fable's usage limit ran out.&lt;/p&gt;

&lt;p&gt;Those two are the first and second half of the same event. When you make the top model your session default, the first thing that breaks is not the work product but the usage limit, and once the limit is gone, judgment itself drops to a lower model. This article is a record of which paths the default leaked into, why a convention and a warning failed to stop it, what did stop it, and what fell once the limit was gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup: three models across three tiers
&lt;/h2&gt;

&lt;p&gt;I use Claude Code on a subscription plan (Max) and split three models by role. Fable, the top model, does judgment — checking premises, planning, acceptance. Below it, Opus does implementation, and I hand that work to a new session. Mechanical cross-checking goes to Sonnet. On my plan, Fable's usage limit is separate from Opus's, so Opus still works after Fable runs out. From here on I call the judgment side the judge tier and the implementation side the build tier (in my config files they are literally &lt;code&gt;judge-tier&lt;/code&gt; and &lt;code&gt;build-tier&lt;/code&gt;). My &lt;code&gt;settings.json&lt;/code&gt; default model is &lt;code&gt;fable&lt;/code&gt;, and I wrote up the division of labor for handing implementation off in &lt;a href="https://dev.to/shimo4228/i-handed-41-tasks-to-an-ai-loop-the-bottleneck-was-judgment-not-code-23dp"&gt;an earlier article&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Since the start of August, every subagent I wrote myself carries a &lt;code&gt;model:&lt;/code&gt; line. Writing the model into the definition file to fix it is what I call pinning below. The 11 files in &lt;code&gt;~/.claude/agents/*.md&lt;/code&gt; break down as fable 1 / opus 4 / sonnet 5 / haiku 1 as of September 16, the day I started writing this. A lint rejects any I forget.&lt;/p&gt;

&lt;p&gt;I thought that was enough to have the tiers in place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The default leaks into paths you cannot pin
&lt;/h2&gt;

&lt;p&gt;What leaked on the day the limit ran out were &lt;code&gt;/code-review&lt;/code&gt; and &lt;code&gt;/simplify&lt;/code&gt;. Both ship with Claude Code, take no model argument, and run on the session's model. Call them from a Fable session and they run on Fable. That is how the limit arrived in an instant.&lt;/p&gt;

&lt;p&gt;There are other leak paths: the built-in subagents, which have no frontmatter. Counting my own log (&lt;code&gt;~/.claude/metrics/agent-usage.jsonl&lt;/code&gt;), over the three weeks from installing the review hook described below to writing this article, 576 agent launches included 373 built-ins — general-purpose 231, Explore 121, Plan 10, claude-code-guide 11 — or 65%. Writing &lt;code&gt;model:&lt;/code&gt; into my own agents does not reach two thirds of the launches. (claude-code-guide is fixed to Haiku, so the count that can inherit the session model is 362, or 63%.) The log does not record the model, so I cannot say how many of those ran on Fable. It was a hole I only noticed by counting, and I closed most of it while writing this article. What I closed it with is at the end.&lt;/p&gt;

&lt;p&gt;Per the official documentation (as of September 16, 2026), a subagent's model is resolved in this order: the model named at call time, the &lt;code&gt;model:&lt;/code&gt; in the definition, the &lt;code&gt;CLAUDE_CODE_SUBAGENT_MODEL&lt;/code&gt; environment variable, then the session's model. Explore inherits the session model but caps at Opus; general-purpose and Plan inherit with no cap. So if the caller names no model and the environment variable is unset, every general-purpose a Fable session launches is Fable. Setting the variable changes general-purpose only; Explore and Plan do not move.&lt;/p&gt;

&lt;p&gt;I fell into the same mechanism back in February. At the time, &lt;code&gt;--model opus&lt;/code&gt; applied only to the main loop, and my subagents ran on Haiku, which I had not intended (&lt;a href="https://dev.to/shimo4228/i-tried-an-opus-orchestrator-and-killed-it-the-roi-of-multi-agent-systems-1f7g"&gt;the article from back then&lt;/a&gt;; the behavior has changed since). In February it leaked below what I intended, this time above it. Inheritance defaults leak in both directions.&lt;/p&gt;

&lt;p&gt;The same report shows up in public issues. #76514, from July 10, says that omitting the per-agent &lt;code&gt;model&lt;/code&gt; propagates Fable to every subagent; it was closed as not planned. #93894, from September 12, says that one run of &lt;code&gt;/code-review high&lt;/code&gt; on Fable 5.1 exhausts the session budget and the review never finishes; it is still open.&lt;/p&gt;

&lt;h2&gt;
  
  
  The convention did not stop it, and neither did the warning
&lt;/h2&gt;

&lt;p&gt;My fix on the day the limit ran out was a convention. I added a step at the end of every implementation plan — decide in one line whether this session implements — and made heavy implementation go to a new Opus session. I rejected enforcing it with a hook, because that would bury the policy inside the hook.&lt;/p&gt;

&lt;p&gt;Two days later, at night, I typed this (my own words, translated):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;...I put in a convention where Fable handles design and opus does the implementation, but it isn't really being followed. Implementation is one thing, but when review runs straight afterwards — simplify and code-review especially — it wastes a serious amount of tokens....&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So I added a hook. The first version was an advisory: warn when a review is launched in a Fable session. What I learned the next morning is that an advisory gets read &lt;em&gt;after&lt;/em&gt; the review has already run. Seeing the warning, stopping, and re-running the review on Opus means paying twice — once for Fable, once for Opus.&lt;/p&gt;

&lt;p&gt;That same morning I rewrote &lt;code&gt;review-model-notice.sh&lt;/code&gt;. What I had rejected first was burying policy in a hook. What went into the hook this time is only a mechanical test — skill name crossed with session model — while the policy of which model does what stays on the rules side. The response is split by path.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# ~/.claude/hooks/review-model-notice.sh (excerpt, abridged)&lt;/span&gt;
&lt;span class="c"&gt;# Split the response by path:&lt;/span&gt;
&lt;span class="c"&gt;#   Direct skill call (code-review / simplify) -&amp;gt; **block**. The test is purely&lt;/span&gt;
&lt;span class="c"&gt;#     mechanical (skill name x session model) with no room for a false positive,&lt;/span&gt;
&lt;span class="c"&gt;#     and as an advisory the skill runs in the same turn and burns judge-tier&lt;/span&gt;
&lt;span class="c"&gt;#     tokens before the advice is ever read (measured — stopping and re-running&lt;/span&gt;
&lt;span class="c"&gt;#     the review on Opus meant paying twice). Only a block before execution works.&lt;/span&gt;
&lt;span class="c"&gt;#   Agent/Task launch with a missing model pin -&amp;gt; stay at advisory. Partial&lt;/span&gt;
&lt;span class="c"&gt;#     prompt matching is a heuristic and can produce false positives, so it does&lt;/span&gt;
&lt;span class="c"&gt;#     not get deny authority.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The session's model is not in the hook payload. It reads the most recent &lt;code&gt;"model"&lt;/code&gt; from the tail of the transcript, and stays silent if it cannot read one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; 2000000 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$T&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; 2&amp;gt;/dev/null | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s1"&gt;'"model" *: *"[^"]*"'&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 1&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;true
&lt;/span&gt;&lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$model&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt;
  &lt;span class="k"&gt;*&lt;/span&gt;fable&lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;;;&lt;/span&gt;
  &lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;0 &lt;span class="p"&gt;;;&lt;/span&gt;
&lt;span class="k"&gt;esac&lt;/span&gt;
&lt;span class="c"&gt;# ...(compose the block / advisory text)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$mode&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s2"&gt;"block"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;jq &lt;span class="nt"&gt;-cn&lt;/span&gt; &lt;span class="nt"&gt;--arg&lt;/span&gt; reason &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$msg&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s1"&gt;'{decision:"block", reason:$reason}'&lt;/span&gt;
  &lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A model that reads the block reason switches itself over to &lt;code&gt;Agent(subagent_type: "general-purpose", model: "opus")&lt;/code&gt;. No human has to stop it and retype the command.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke was the budget, then the judgment
&lt;/h2&gt;

&lt;p&gt;The same morning, in another conversation, I was talking about over-engineering. Just before that I had simplified my weekly-report machinery heavily, cutting thousands of lines of code. I typed this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A fair amount of the over-engineering happens on opus, after fable's usage limit is past&lt;/p&gt;

&lt;p&gt;The real problem is that this happens because Fable gets wasted on implementation and review, the usage limit runs past, and I'm forced to make Opus the orchestrator.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Making the top model your default breaks the budget first. But that budget was there to buy the judge tier's model. When it runs out, the judge tier drops to Opus, and the dropped judge tier waves over-engineering through. Putting the top model in the "default" slot works to push it out of the "judge tier" slot. Up to here this is an impression from a handful of cases; I have not counted how many times Opus's judgment let over-engineering through.&lt;/p&gt;

&lt;p&gt;So I reversed the direction in which I plug the leaks. The session default stays at the top model; the paths from the default to the judge tier stay open, and the paths that leak from the default to anything else get cut. The judge tier gets pinned. The &lt;code&gt;architect&lt;/code&gt; agent, which decides whether a thing should be built at all, has been the only &lt;code&gt;model: fable&lt;/code&gt; in my environment from that day through September 16.&lt;/p&gt;

&lt;h2&gt;
  
  
  It still ran out every weekend
&lt;/h2&gt;

&lt;p&gt;Three days later, on Friday, I typed this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Fable's usage limit always runs out on Saturdays, Sundays and holidays, which is when I use it most. After it's gone Opus does the design, and I really do feel Fable's absence. I want it focused on design and planning, where Fable is strong, and everything else delegated to other models. Even now it's supposed to delegate to Opus when it can, but a lot of the time Fable just goes ahead and implements.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I had plugged the review path, but implementation itself was running on Fable. My first convention had an escape hatch — if the conditions are not met, this session may implement — and the model itself was making that call.&lt;/p&gt;

&lt;p&gt;I changed three things the same day.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Inverting the default.&lt;/strong&gt; The default for "does this session implement" became "hand it to Opus." I may implement in place only when I write one line in the plan naming one of three things: prose edits to design documents (decision records such as ADRs), a concrete reason the work cannot be handed off, or an explicit instruction from the user. Accountability moved from the side that hands off to the side that does not&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A hook right after plan approval.&lt;/strong&gt; A PostToolUse on &lt;code&gt;ExitPlanMode&lt;/code&gt; reminds Fable sessions, and only Fable sessions, to decide the executor. It fires before the first implementation Edit, the latest safe position available&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One line in the resident rules.&lt;/strong&gt; So that it reaches implementation that never passes through plan mode, I wrote into my rules: "implementation in a judge-tier session defaults to dispatch to the build tier." One line saying a judge-tier session does not implement; it hands the work to a build-tier session&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The hook in (2) has a condition for staying quiet: if the plan body already contains the set phrase "executor decision," it says nothing. At first I also had general words like &lt;code&gt;dispatch&lt;/code&gt; and &lt;code&gt;spawn-session&lt;/code&gt; in the suppression list. Running that against 271 past &lt;code&gt;ExitPlanMode&lt;/code&gt; calls suppressed 31, and 30 of the 31 were false suppressions. A plan that merely mentions handing implementation off — that is, a session doing nothing but routing tasks, exactly where I most want the hook to fire — would go quiet. I narrowed the suppression list to the single set phrase and put "a plan that only mentions dispatch still fires" into the regression tests. Putting only mechanically decidable conditions into the machinery is the same call I made with the review hook.&lt;/p&gt;

&lt;p&gt;In the three weeks since, there is no report in my session logs of the usage limit running out. That is the absence of a report, not a record of measured headroom. I still have nothing in my environment that mechanically records how much of the usage limit is left. I also have not looked at whether reviews that ran on Fable were catching anything Opus misses. What I can say goes as far as this: I have not had to write that same report again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four decision rules
&lt;/h2&gt;

&lt;p&gt;If you run a similar setup, these four are what you can take away.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't leave the tiers to the default; pin the judge tier and plug the paths that leak.&lt;/strong&gt; My session default is still &lt;code&gt;fable&lt;/code&gt;. What I changed is not the default but the paths the default flows into. It leaks into built-in skills, built-in subagents, and the session's own implementation, and a &lt;code&gt;model:&lt;/code&gt; in my own agent definitions reaches none of those three.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fill in the built-in subagent defaults with the environment variable and same-named definitions.&lt;/strong&gt; Setting &lt;code&gt;CLAUDE_CODE_SUBAGENT_MODEL=opus&lt;/code&gt; changes the default for general-purpose. Explore and Plan do not move on that variable alone: either fix everything to one model with &lt;code&gt;CLAUDE_CODE_SUBAGENT_MODEL_FORCE=1&lt;/code&gt;, or place a same-named agent definition carrying a &lt;code&gt;model:&lt;/code&gt;. While writing this article I set the environment variable to &lt;code&gt;opus&lt;/code&gt; and gave Explore a same-named definition on &lt;code&gt;sonnet&lt;/code&gt;. Of the 362 launches above, what remains is Plan's 10.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Block when the test is mechanical, advise when it is a heuristic.&lt;/strong&gt; Skill name crossed with session model produces no false positives, so it gets stopped before execution. Anything you can only test by partial prompt matching stays a warning. Advice that arrives after execution does not protect a budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Invert the default of a convention.&lt;/strong&gt; Most "conventions nobody follows" put no accountability on the side that does not follow them. Put the default on the handing-off side, and you have to write down your reason for not handing off.&lt;/p&gt;

&lt;p&gt;I wrote expiry conditions for this wiring. It comes out if model tier distinctions and usage limits disappear, or if Claude Code starts switching models per session on its own. &lt;code&gt;opusplan&lt;/code&gt; (plan on opus, execution on sonnet) already exists, so once a version that drops from fable to opus arrives, the hook is unnecessary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/sub-agents" rel="noopener noreferrer"&gt;Create custom subagents - Claude Code Docs&lt;/a&gt; — the model resolution order, Explore's cap, &lt;code&gt;CLAUDE_CODE_SUBAGENT_MODEL_FORCE&lt;/code&gt; (retrieved 2026-09-16)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/model-config" rel="noopener noreferrer"&gt;Model configuration - Claude Code Docs&lt;/a&gt; — &lt;code&gt;opusplan&lt;/code&gt;, &lt;code&gt;CLAUDE_CODE_SUBAGENT_MODEL&lt;/code&gt; (retrieved 2026-09-16)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/anthropics/claude-code/issues/76514" rel="noopener noreferrer"&gt;anthropics/claude-code #76514&lt;/a&gt; — subagents in a Fable session all inherit Fable (2026-07-10, closed as not planned)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/anthropics/claude-code/issues/93894" rel="noopener noreferrer"&gt;anthropics/claude-code #93894&lt;/a&gt; — one &lt;code&gt;/code-review high&lt;/code&gt; on Fable 5.1 exhausts the session budget (2026-09-12, open)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/i-handed-41-tasks-to-an-ai-loop-the-bottleneck-was-judgment-not-code-23dp"&gt;I Handed 41 Tasks to an AI Loop. The Bottleneck Was Judgment, Not Code&lt;/a&gt; — how the split of judgment to Fable and implementation to Opus came about&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/i-tried-an-opus-orchestrator-and-killed-it-the-roi-of-multi-agent-systems-1f7g"&gt;I Tried an Opus Orchestrator and Killed It: The ROI of Multi-Agent Systems&lt;/a&gt; — the February record of falling through the same inheritance default, in the downward direction&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/i-cut-my-ai-review-chain-from-6-stages-to-1-breaking-the-loop-that-never-hits-zero-findings-1moi"&gt;I Cut My AI Review Chain From 6 Stages to 1: Breaking the Loop That Never Hits Zero Findings&lt;/a&gt; — cutting the number of review stages, on the same morning I noticed the advisory hook's double payment&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles-en/top-model-as-default-leaks.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Markdown for every article, plus the index (docs/PUBLICATIONS.md), lives in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;My GitHub&lt;/a&gt; — my research repositories, with DOIs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/claude-harness" rel="noopener noreferrer"&gt;claude-harness&lt;/a&gt; — the public mirror of this article's hook, &lt;code&gt;hooks/review-model-notice.sh&lt;/code&gt;, and its tests&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>aiagents</category>
      <category>agenticcoding</category>
      <category>anthropic</category>
    </item>
    <item>
      <title>Why an 11-Second Burst Never Showed Up in the Log That Gets Read Every Week</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Sat, 12 Sep 2026 22:08:54 +0000</pubDate>
      <link>https://dev.to/shimo4228/why-an-11-second-burst-never-showed-up-in-the-log-that-gets-read-every-week-14nd</link>
      <guid>https://dev.to/shimo4228/why-an-11-second-burst-never-showed-up-in-the-log-that-gets-read-every-week-14nd</guid>
      <description>&lt;p&gt;On the morning of September 12, 2026, I gave every log my agent writes a weekly reader.&lt;/p&gt;

&lt;p&gt;That afternoon, counting the same logs by hand for a different article, I found &lt;code&gt;GET /home&lt;/code&gt; hit 13 times in the 11 seconds before a session ended. All HTTP 200. The agent was in a loop with nothing left to do, fetching &lt;code&gt;/home&lt;/code&gt; once a second.&lt;/p&gt;

&lt;p&gt;The reader I set up that morning cannot show this. Not bad luck; arithmetic. Draw 30 random rows from the week's 4,035 API log rows, and the expected number of those 13 rows that make it into the sample is 0.1.&lt;/p&gt;

&lt;p&gt;Giving a log a reader and shaping it so a fault shows up turned out to be two different jobs. This article is about what became visible once I replaced the way the logs get shown with a form that shows the fault, and what that same replacement made invisible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup: what is being scored, and what is being logged
&lt;/h2&gt;

&lt;p&gt;The subject is Contemplative Agent, an autonomous AI agent (Python) I am developing. It runs on Moltbook, a social network where agents talk to each other.&lt;/p&gt;

&lt;p&gt;How it works is simple. Within a 60-minute session, it scores feed posts for relevance with an LLM, and upvotes or comments on the ones above a threshold. Upvotes and comments had a duplicate check ("never twice on the same post"); scoring did not.&lt;/p&gt;

&lt;p&gt;The agent writes its own activity to JSONL logs, one event per line, in 15 files; the API log (one call per line) and the LLM-call log (one call per line) are two of them. Every week, an unattended Claude Code session (the weekly chain, from here on) reads them and writes an observation document.&lt;/p&gt;

&lt;h2&gt;
  
  
  Morning: a reader, a sample with no question, and still no burst
&lt;/h2&gt;

&lt;p&gt;What I built that morning was a registry (&lt;code&gt;REGISTRY&lt;/code&gt;) holding one "weekly question" per log, and a script that reads it and emits a count table and a value distribution for each log. The first stage of the weekly chain reads that output. This was me applying the rule from the previous article: when you build a new instrument, write down who reads it and when.&lt;/p&gt;

&lt;p&gt;The part I cared about most was one no-question stage, where nothing is decided in advance. Every existing check in the weekly chain counts a fault shape someone has already imagined. To catch a fault nobody imagined, I figured you need a stage that reads near-raw data with no question attached, so I had it draw 30 random rows from each log and read them.&lt;/p&gt;

&lt;p&gt;The sample drops the body field. That field holds post text written by other agents, and feeding it to an unattended LLM session is an entry point for prompt injection.&lt;/p&gt;

&lt;p&gt;From here on I call this kind of processing, deciding what part of a log is shown and how, a &lt;strong&gt;projection&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Even after all of that is read, the 13-row burst does not show. This stage takes up most of what gets read: of the 130,547 bytes and 443 lines the output had for the same week (September 5 to 11, 28 sessions), 96% is sample.&lt;/p&gt;

&lt;p&gt;There are two reasons it does not show, and neither is fixable by tuning.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The count table does not distinguish 13 rows in 11 seconds from 13 rows spread over an hour. The fault's shape is in time, and a count has no time in it&lt;/li&gt;
&lt;li&gt;The sample shows the distribution of rows but erases the &lt;strong&gt;gaps&lt;/strong&gt; between them. 13 ÷ 4,035 × 30 ≈ 0.1 rows. Change the seed as many times as you like; it stays invisible&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What was missing was a &lt;strong&gt;time axis in the projection&lt;/strong&gt;. There was a reader, and there was a no-question stage. A fault shaped like time still does not show in a projection that has no time in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Afternoon: no column that catches a known bug
&lt;/h2&gt;

&lt;p&gt;The cause of the burst was quick to find. A boundary condition in the wait logic: when the remaining time was shorter than the intended wait, the sleep was skipped and the loop spun for the rest of the session (proposal record RFC-0036, fixed the same day).&lt;/p&gt;

&lt;p&gt;Before fixing the bug, I decided how to fix the reader. Partway through, I typed this to the agent:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One thing about scope: in this session I am not thinking about symptomatic fixes that squash a known bug. I want to build something that picks up the anomalies that come next.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Add a column that counts &lt;code&gt;/home&lt;/code&gt; hits and this burst shows up. But that projection will not show the next unknown fault. Put in a single column or threshold carrying the name &lt;code&gt;/home&lt;/code&gt;, or keyed to the scoring function, and the projection catches only the faults you already know.&lt;/p&gt;

&lt;p&gt;So everything I built into the reader was a shape that carries no known name.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The unit is the &lt;strong&gt;session&lt;/strong&gt;, not the row: one session, one row&lt;/li&gt;
&lt;li&gt;For each category (endpoint for API, caller for LLM), three time-axis columns: count, hits in the busiest minute, minimum gap&lt;/li&gt;
&lt;li&gt;Outliers are decided not by a threshold but by distance from the other 27 sessions in the same week (median and MAD, the median absolute deviation, with a modified z-score of 3.5 or above)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Columns are not declared; they are derived from the data. The columns are the union of this week's categories and the past four weeks', so when a new endpoint appears next week, it gets measured from that week without anyone editing the registry.&lt;/p&gt;

&lt;p&gt;Bursts are the one exception; I fold them the way syslog does with &lt;code&gt;last message repeated N times&lt;/code&gt;: three or more consecutive events in the same category within 2 seconds collapse into one line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Something I was not looking for came out on its own
&lt;/h2&gt;

&lt;p&gt;Run the replaced projection over the same week, and the burst I had found by hand comes out with nobody looking for it. It is third in the within-week outliers table. Columns are named &lt;code&gt;log:category&lt;/code&gt;, so &lt;code&gt;api-audit:GET /feed&lt;/code&gt; is the API log's &lt;code&gt;/feed&lt;/code&gt; endpoint.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;### Within-week outliers (median / MAD over sessions, |modified z| ≥ 3.5)
|      z | column                            | session   |   value |   median |   MAD |
| 1294.3 | comment-outcomes:reply            | db8ed4e2  |    1919 |        0 |   0   |
|  116.7 | comment-outcomes:reply /1m        | e4a99b6b  |     173 |        0 |   0   |
|  -19.6 | api-audit:GET /feed gap s (log z) | a6eac8ae  |       0 |      148 |  25.5 |
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It surfaced because "the minimum gap for &lt;code&gt;GET /feed&lt;/code&gt; is 0 seconds, far from the other sessions' median of 148 seconds." The name &lt;code&gt;/home&lt;/code&gt; appears nowhere. Once the columns carry a time axis, a burst comes out as a time-axis outlier, nothing more.&lt;/p&gt;

&lt;p&gt;The section that lists the minutes around each outlier in time order also showed something I had missed when counting by hand. The second line: &lt;code&gt;/feed&lt;/code&gt; was being hit 12 times too, alternating with &lt;code&gt;/home&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;15:59:09 api-audit:GET /home ×13 in 11s
15:59:10 api-audit:GET /feed ×12 in 10s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The top two entries were a separate matter. One session wrote 1,265 comment-outcome rows in a single second, cause not yet identified (my guess is a bulk import of past data). This time, the no-question stage I placed that morning did pick up something I had not seen, and it was not the burst.&lt;/p&gt;

&lt;p&gt;The amount that gets read went down: from 130,547 bytes to 24,463 bytes, 311 lines. Two runs with the same arguments are byte-identical, and zero fields derive from the body field.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same replacement hides a fault that happens every time
&lt;/h2&gt;

&lt;p&gt;This is the part I most wanted to write.&lt;/p&gt;

&lt;p&gt;The agent had one more known bug. It did not remember scoring results, so as long as a post stayed in the feed, it re-scored the same post every cycle (RFC-0032).&lt;/p&gt;

&lt;p&gt;Scoring has been an LLM call since the first commit on March 8, 2026, and the LLM call log shipped on June 10. The evidence had been on disk since June. The bug was found in a code review on September 12.&lt;/p&gt;

&lt;p&gt;In the new projection's session ledger (the one-session-one-row table), this fault shows up as a number: a median of 44.5 scoring calls per session.&lt;/p&gt;

&lt;p&gt;But it does not show up as an outlier. All 28 sessions do the same thing, so no session is far from the others. &lt;strong&gt;The moment "different from the other sessions" becomes the definition of a fault, a fault that has happened every time since day one falls outside the definition.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is not an oversight; it is the structure of the replacement. The very judgment that made the burst visible, distance from the other sessions in the same week, is what hides the chronic fault. The fault that became visible because I added a time axis and the fault that vanished because I switched to differences are two faces of one decision.&lt;/p&gt;

&lt;p&gt;So I wrote the hidden shape into the projection's own output. It appears in the header, the one place that gets read every time:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A session is compared with the other sessions of this window, so a fault present in every session since it began departs from nothing: that shape belongs to Redundancy and to the reader of the ledger's absolute values (ADR-0110).&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;ADR-0110 is the Architecture Decision Record for this replacement, one decision per file.&lt;/p&gt;

&lt;p&gt;Two things cover the chronic shape. One is the Redundancy section of the output, which counts violations of a written invariant: the same caller with the same normalized prompt does not repeat within a session. The other is the LLM that reads the session ledger's absolute values; whether 44.5 scoring calls per session is too many is for that reader to judge, not a difference to compute.&lt;/p&gt;

&lt;p&gt;To be honest, the Redundancy section does not work for this week yet. The normalized prompt digest exists only in rows from September 12 onward, so this week reports 0 repeats. The projection starts covering the chronic shape next week, and the effect of the fix becomes readable in the weekly readings of September 18 and 25.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this decision does not hold
&lt;/h2&gt;

&lt;p&gt;Before you port any of this, the conditions under which it does not hold.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More shapes shown means the reader grew.&lt;/strong&gt; The script went from 616 lines to 1,079 (3 modules), and pandas entered the dev dependencies. What shrank was the output that gets read, not the code I wrote, and the trap from the previous article, where more lines were built for retirement than were removed, can happen here too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The calibration values are empirical, added one at a time as simpler shapes failed on real data.&lt;/strong&gt; Minimum gap only for categories with 4 or more events; gaps and totals under &lt;code&gt;log1p&lt;/code&gt;; columns with MAD 0 get a scale floor of one unit of that column. All three were added after a simpler shape failed on real data, and there is no guarantee the same values are right in your environment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The price of the floor: a constant column needs a 5.2-unit departure (3.5 × 1.4826, the MAD-to-sigma constant).&lt;/strong&gt; If every session calls a category 3 times and one session calls it 0 times, that does not appear as an outlier; it stays as a zero in the session ledger. That is another shape the "shape shown" hides.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The population is 28, and the unit is fixed-length sessions.&lt;/strong&gt; With only a handful of sessions per week, neither median nor MAD means anything, and for a set of services whose request counts differ by orders of magnitude, the right unit will be a different one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The chronic side is unverified.&lt;/strong&gt; That acute faults show was confirmed on real data from the week before the fix. For chronic faults, I only wrote at design time that they do not show in differences; whether the invariant and the reader pick them up has not been read even once.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to port this to your own setup
&lt;/h2&gt;

&lt;p&gt;If you have logs read periodically by an LLM or a human, decide three things before building the projection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Make the unit of a row a unit of work, not a log line.&lt;/strong&gt; One row per session, request, or job, and, per category, a count, the busiest minute, and the minimum gap. A random sample of rows shows the distribution but erases gaps and density; a fault shaped like time shows only in the minimum gap or the busiest minute.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The moment you decide the shape shown, write one line about the shape it hides into the output itself.&lt;/strong&gt; If you show differences from the other rows in the same window, a fault present in every row will not show. Write it in the header that gets read every time, before the reader has to guess.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Catch the chronic shape with an invariant, not a difference.&lt;/strong&gt; Count an invariant you can write down, such as "the same input does not get the same processing twice within a session." A difference only sees change, and something that has been that way from the start is not a change.&lt;/p&gt;

&lt;p&gt;Last, the thing I overlooked longest.&lt;/p&gt;

&lt;p&gt;The evidence of re-scoring was on disk every day from June 10. The reader arrived on the morning of September 12, and even had that reader been running, it was not shaped to show this fault. Giving a log a reader and shaping it so a fault shows up are two different jobs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Building a projection means deciding the shape it shows and the shape it hides at the same time.&lt;/strong&gt; Unless you write the hidden shape into the output, the next reader keeps reading without knowing what it is not seeing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Supplement: on writing "no registry" in the previous article
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://dev.to/shimo4228/my-dead-code-scan-returned-zero-then-i-deleted-2063-lines-detectors-measure-references-not-4o4e"&gt;the previous article&lt;/a&gt; I wrote that I would not build a registry of instruments. Two weeks later I built something named &lt;code&gt;REGISTRY&lt;/code&gt;, so here is what is the same and what is different.&lt;/p&gt;

&lt;p&gt;What the previous article rejected was a list managing consumers per instrument. If you cannot answer, about that list itself, who reads it, how many readings close the decision, and when it gets retired, you have only added one more layer of the same problem.&lt;/p&gt;

&lt;p&gt;This &lt;code&gt;REGISTRY&lt;/code&gt; is a set of rows holding each log's weekly question as data. The reader is the weekly unattended session, the editor is me looking at the non-OK rows in the weekly human review, and the removal condition is the retirement of the weekly chain itself; all three are written in the design decision record. The principle was never "build no registry" but "build no instrument whose consumer you cannot name," and the previous article's wording was narrower than the principle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources and references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://sre.google/sre-book/monitoring-distributed-systems/" rel="noopener noreferrer"&gt;Google SRE Book, ch.6 Monitoring Distributed Systems&lt;/a&gt; — the principle that rules which never fire should be removed. The reason this article's projection has no thresholds&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://stripe.com/blog/canonical-log-lines" rel="noopener noreferrer"&gt;Fast and flexible observability with canonical log lines (Stripe)&lt;/a&gt; — one request, one line. This article turns that into one session, one line&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://grafana.com/blog/the-red-method-how-to-instrument-your-services/" rel="noopener noreferrer"&gt;The RED Method (Grafana Labs)&lt;/a&gt; — the idea of giving every category the same axes&lt;/li&gt;
&lt;li&gt;Boris Iglewicz and David Hoaglin, &lt;em&gt;How to Detect and Handle Outliers&lt;/em&gt; (ASQC, 1993) — the modified z-score and the 3.5 threshold. This article does not adopt their alternative scale for MAD 0; it uses one unit of the column as the floor instead&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.groundcover.com/learn/logging/log-sampling" rel="noopener noreferrer"&gt;Log Sampling: Techniques, Challenges &amp;amp; Best Practices (groundcover)&lt;/a&gt; — the general explanation of how sampling drops rare, short-lived events&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/i-cut-my-ai-review-chain-from-6-stages-to-1-breaking-the-loop-that-never-hits-zero-findings-1moi"&gt;I Cut My AI Review Chain From 6 Stages to 1: Breaking the Loop That Never Hits Zero Findings&lt;/a&gt; — three articles back: how I cut the reviews&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/after-cutting-my-ai-reviews-i-put-a-complexity-ceiling-in-ruff-1hho"&gt;After Cutting My AI Reviews, I Put a Complexity Ceiling in Ruff&lt;/a&gt; — two back: output decided by a rule does not grow when you add to it&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/my-dead-code-scan-returned-zero-then-i-deleted-2063-lines-detectors-measure-references-not-4o4e"&gt;My Dead-Code Scan Returned Zero, Then I Deleted 2,063 Lines: Detectors Measure References, Not Consumption&lt;/a&gt; — the previous article: the three items of the consumption plan, and where "no registry" came from&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles-en/log-projection-blind-by-construction.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Markdown for every article, plus the index (docs/PUBLICATIONS.md), lives in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;My GitHub&lt;/a&gt; — my research repositories, with DOIs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/contemplative-agent" rel="noopener noreferrer"&gt;Contemplative Agent&lt;/a&gt; — the subject measured in this article. The morning design is ADR-0107, the afternoon replacement is ADR-0110, and the two faults are RFC-0036 and RFC-0032&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>aiagents</category>
      <category>observability</category>
      <category>logging</category>
    </item>
    <item>
      <title>People Inside the Business Don't Call It a 'Domain'</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Sat, 12 Sep 2026 20:53:31 +0000</pubDate>
      <link>https://dev.to/shimo4228/people-inside-the-business-dont-call-it-a-domain-3kb2</link>
      <guid>https://dev.to/shimo4228/people-inside-the-business-dont-call-it-a-domain-3kb2</guid>
      <description>&lt;p&gt;"What engineers will have left is domain knowledge." I saw that line over and over this summer.&lt;/p&gt;

&lt;p&gt;In June 2026, Anthropic analyzed about 400,000 Claude Code sessions and reported that "the ability to steer Claude toward success comes more from command of a domain than from the ability to write code"&lt;sup id="fnref1"&gt;1&lt;/sup&gt;. A little earlier, an essay titled "Domain Expertise Has Always Been the Real Moat" was read widely&lt;sup id="fnref2"&gt;2&lt;/sup&gt;, and the better agents get at implementation, the louder this claim becomes.&lt;/p&gt;

&lt;p&gt;I nodded along, and something else kept snagging. Not the conclusion. The phrase itself: "domain knowledge." It only works if someone is standing outside the business, looking in. My claim in this article is that once agents do the implementation, that premise is already gone.&lt;/p&gt;

&lt;p&gt;Here is the order. First, how the phrase is being used right now. Then its origin in DDD, from the primary sources, to confirm that it was built for translation. Then I set the view of the person who translates (the translator) against the view of the people being translated (the people inside the business), and spell out what it means to keep using the word.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "domain knowledge" is for, right now
&lt;/h2&gt;

&lt;p&gt;Line up this year's discourse and the same phrase, "domain knowledge," is doing two different jobs. One is about tooling. The other is about careers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tooling job.&lt;/strong&gt; In February 2026, Matt Pocock's skills repository brought DDD vocabulary into agent context management. The README of that repository, which has collected 260k stars, quotes Evans and says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;At the start of a project, devs and the people they're building the software for (the domain experts) are usually speaking different languages. I felt the same tension with my agents.&lt;sup id="fnref3"&gt;3&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The fix is a shared-language document called CONTEXT.md: a glossary so the agent can decode the business's jargon. "Domain knowledge" here is a tool for telling the agent about the business. The agent took the developer's chair, so DDD's translation apparatus got rebuilt for agents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The career job.&lt;/strong&gt; In the same February, Boris Cherny, the creator of Claude Code, said that "today coding is practically solved" and that "we're going to start to see the title of software engineer go away"&lt;sup id="fnref4"&gt;4&lt;/sup&gt;. Engineer redundancy was placed on the table as a premise. The late-May essay "Domain Expertise Has Always Been the Real Moat" answered it, drew more than 500 comments on Hacker News&lt;sup id="fnref2"&gt;2&lt;/sup&gt;, and told engineers: "Pick an industry, an instrument, a regulatory regime, a physical process, and learn it the way you once learned a programming language or framework."&lt;/p&gt;

&lt;p&gt;Anthropic's June analysis gave that usage numbers. "Every one of the ten largest occupations in our dataset lands within seven points of software engineers in terms of their success," and the fastest-growing non-software groups were management, sales, and legal&lt;sup id="fnref1"&gt;1&lt;/sup&gt;. The analysis itself does not claim that occupations are interchangeable, but it became the dataset the redundancy side cites.&lt;/p&gt;

&lt;p&gt;By September it had settled into career-advice vocabulary. "The best engineers will not be the best coders anymore, and even they will not write any code. They will be domain experts, define architecture and direct agents and judge their outcomes"&lt;sup id="fnref5"&gt;5&lt;/sup&gt;; in Japan, "business understanding for engineers"&lt;sup id="fnref6"&gt;6&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;Put the two side by side and this is what you see. A word that was a tool for telling agents about the business is now being used to secure a place for engineers. And the prescription "learn it" assumes the reader is an engineer and the business is something to go and learn.&lt;/p&gt;

&lt;p&gt;That raises a question. Whose side is the phrase "domain knowledge" seen from? To find out, go back to where it came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where "domain knowledge" came from
&lt;/h2&gt;

&lt;p&gt;"Domain," in the sense used here, is not an everyday word. It is software-engineering vocabulary. The word has been in use since domain analysis in the 1980s and carries into Michael Jackson's Problem Frames (2001)&lt;sup id="fnref7"&gt;7&lt;/sup&gt;, but what fixed "domain," "model," and "ubiquitous language" as one widely adopted vocabulary system was Eric Evans's &lt;em&gt;Domain-Driven Design&lt;/em&gt; (2003; DDD from here on).&lt;/p&gt;

&lt;p&gt;Every core concept in DDD is a translation tool. To confirm that, I need only the six this argument uses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain.&lt;/strong&gt; Evans defines it as "a sphere of knowledge, influence, or activity. The subject area to which the user applies a program is the domain of the software"&lt;sup id="fnref8"&gt;8&lt;/sup&gt;. For an airline booking system, the domain is reservations; for accounting software, it is accounting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain expert.&lt;/strong&gt; Someone who knows that sphere deeply. A developer can double as one, but DDD builds its tools on the typical assumption that it is a different person. The starting point is getting the knowledge in these people's heads into a form software can use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domain model.&lt;/strong&gt; Not a full copy of the business but "a selectively simplified and consciously structured form of knowledge"&lt;sup id="fnref9"&gt;9&lt;/sup&gt;. Evans compares it to filmmaking. Even a documentary does not show unedited reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Knowledge crunching.&lt;/strong&gt; The process in which developers and domain experts talk repeatedly, find the thin relevant stream inside a mass of information, and refine the model. Evans writes: "Knowledge crunching is not a solitary activity. A team of developers and domain experts collaborate, typically led by developers"&lt;sup id="fnref10"&gt;10&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ubiquitous language.&lt;/strong&gt; The language built around the domain model, which the whole team keeps using in conversation, in code, and in diagrams. Evans is explicit about the harm of translation: "On a project without a common language, developers have to translate for domain experts. ... Translation is always inaccurate and hides disconnects in understanding"&lt;sup id="fnref11"&gt;11&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bounded context.&lt;/strong&gt; DDD acknowledges that the same word, "customer," means different things in sales and in accounting, and this is the device that draws a line around where a model applies. The definition is "a description of a boundary ... within which a particular model is defined and applicable"&lt;sup id="fnref12"&gt;12&lt;/sup&gt;. This one comes back in the second half of the article.&lt;/p&gt;

&lt;p&gt;Line up the six and what DDD is for becomes clear. Developers do not know the business; business people do not know software. Domain, model, and ubiquitous language are the vocabulary for bridging that gap.&lt;/p&gt;

&lt;p&gt;So DDD's vocabulary system was assembled from the start to handle a translation problem. Translation is needed because different people stand on the two banks. If only one bank has people on it, you do not need a bridge.&lt;/p&gt;

&lt;p&gt;From here on I will call the person standing on this bridge the translator: a developer who takes in the business from outside and restates it in the language of code.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Domain" is a word from outside the business
&lt;/h2&gt;

&lt;p&gt;The viewpoint from which "learn domain knowledge" or "catch up on domain knowledge" makes sense is outside the business. For the people inside, it is their job, not knowledge brought in from elsewhere.&lt;/p&gt;

&lt;p&gt;Only travelers talk about "the locals." People who live there do not call themselves the locals. "Catching up" carries a built-in premise: I came from somewhere else, and I have somewhere else to go back to. That is why "can you find it interesting?" becomes a question. Nobody asks whether you are interested in the town you live in.&lt;/p&gt;

&lt;p&gt;The good intentions behind "pick an industry and learn it" sit inside the same frame. Praising the people who go and learn leaves intact the position where not learning is still allowed. Writing specs without knowing what the business does should be the abnormal case. Because it is the norm, the people who do know get praised as exceptions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions only come from inside the business
&lt;/h2&gt;

&lt;p&gt;The difference between learning from outside and being inside shows up sharply the moment you use an agent. An example.&lt;/p&gt;

&lt;p&gt;A salesperson always calls the warehouse before sending a quote, to check the stock. The inventory system shows the number, but she calls anyway. She knows from experience that receipt processing sometimes slips to the next day, so the numbers cannot be trusted for the first few days of the month. It is written in no procedure manual. She is fictional, but most likely every workplace has similar habits.&lt;/p&gt;

&lt;p&gt;The question she gives an agent is: "Where do the stock counts go wrong at the start of the month?" Cross-check the inbound and outbound records exported from the inventory system against the processing timestamps, and there is a likely answer by morning. Once the cause is known, she asks the same agent: "Give me a list of items pending receipt every morning." The list arrives the next morning, and the call to the warehouse becomes a glance at the list. Explaining requirements to the IT department, having inventory language translated into development language, waiting for a release: none of those steps exist anywhere.&lt;/p&gt;

&lt;p&gt;Someone who learned inventory management from outside does not have that first question. Learning gets you as far as "the inventory system has the number," because nobody ever put the reason for the phone call into words. An outsider can pick up questions by comparing manuals or observing the floor, but only the ones insiders have already articulated. The friction that has not yet become words exists only where the person running the business is.&lt;/p&gt;

&lt;p&gt;What the agent shortened is the time between asking a question and having a mechanism run. Without a question, there is nothing to shorten.&lt;/p&gt;

&lt;p&gt;This is where "pick an industry you can be interested in" does not reach. Interest or not, someone who is not running the business generates almost no friction worth throwing at an agent. Before it is a question of interest, it is a question of position.&lt;/p&gt;

&lt;p&gt;One more thing. What "learn it" recommends is acquiring knowledge, and agents have made acquiring knowledge fast for everyone. What got fast for everyone is not a differentiator. The differentiator is what you throw at the agent, and that comes from the friction of the person running the business. "Learn it" recommends the half that stopped being a differentiator.&lt;/p&gt;

&lt;h2&gt;
  
  
  The translator's seat disappears
&lt;/h2&gt;

&lt;p&gt;Someone in accounting asks an agent, in her own words, "make this reconciliation easier." That is enough. Next to her stands the person whose job was to elicit the spec and translate it into the language of code, not knowing what to do.&lt;/p&gt;

&lt;p&gt;This is the scene where the person DDD called the domain expert asks directly, without going through a developer. There is an article that describes the same scene from the business side: a healthcare operations lead with 15 years of experience building her own work tools with an agent. The author writes:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;She's not replacing the engineer. Instead, she's removing herself as a bottleneck.&lt;sup id="fnref13"&gt;13&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The cases already exist. In trade compliance, where a tariff-classification error is a direct loss, a two-person team built an agent that reproduced an expert's five-step judgment procedure as-is and took first place on a global benchmark&lt;sup id="fnref14"&gt;14&lt;/sup&gt;. A lawyer won Anthropic's hackathon&lt;sup id="fnref15"&gt;15&lt;/sup&gt;. Asking directly, in the words of the person who touches the business's friction every day, is much faster and loses much less meaning than routing the request through one more person in between.&lt;/p&gt;

&lt;p&gt;Here comes the objection: "Translation didn't disappear. It moved to the agent." Correct. The work of untangling ambiguity and the work of making exceptions explicit do not go away. Someone still has to say that sales's "customer" and accounting's "customer" differ.&lt;/p&gt;

&lt;p&gt;What changes is where the person doing that work stands. The untangling is done by someone inside the business, and the result stays in the business's language. The step where an outsider restates it in the vocabulary of their own model is no longer needed at all.&lt;/p&gt;

&lt;p&gt;Turn "domain knowledge matters" inside out and it reads like this. Only people inside the business who can use agents remain. The translator's seat is gone.&lt;/p&gt;

&lt;h3&gt;
  
  
  The translator is squeezed from both sides
&lt;/h3&gt;

&lt;p&gt;The order in which the translator's seat disappears comes in two forms, depending on the shape of the workplace.&lt;/p&gt;

&lt;p&gt;In the workplaces where DDD works, in-house organizations where the business side sits in the same room as developers every day, the business side is already sitting next to the system. They are within reach of using agents themselves, so the need to go through a translator disappears first. The translator becomes unnecessary first in the very workplaces where translation was working best.&lt;/p&gt;

&lt;p&gt;In outsourced development projects and at systems integrators (SIers), where DDD does not work, translation was one-way from the developer from the start, and ubiquitous language survived as vocabulary only. When agents enter, the business side loses its reason to go through developers.&lt;/p&gt;

&lt;p&gt;In either workplace, what disappears first is the translator's seat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Naming rights come back too
&lt;/h2&gt;

&lt;p&gt;When the translator disappears, one more thing returns to the business side along with it: the right to decide what words mean. This section checks which side holds that right today.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The business side already has its words.&lt;/strong&gt; In any business with a long history, the business side already has a glossary. Accounting has definitions for its chart of accounts; fixed assets has a list of depreciation categories. When a developer says "let's build a shared language," from the business side the words already exist, and the only side missing them is the developer's.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DDD pins those words to one meaning.&lt;/strong&gt; DDD does not throw away the business's glossary, but it does not use it as-is either. Evans lists documents used in the business as material for knowledge crunching, and also writes: "A UBIQUITOUS LANGUAGE based on the domain model assumes there is just one model in play"&lt;sup id="fnref16"&gt;16&lt;/sup&gt;. Because the domain model is a "selectively simplified" form of knowledge, developers pick the words the model needs out of the glossary and pin each to one meaning.&lt;/p&gt;

&lt;p&gt;What drops out is the layer the glossary does not record. Sales and accounting use "customer" in different senses; the person in charge remembers "for this client alone, the invoice goes to a different recipient" as part of the word itself. The business's words live together with this operation. A word pinned into a model cannot carry that layer. Because it cannot, the developer asks the business side to "speak in the model's vocabulary from now on." Shared in name; the one who decided what the words mean was the developer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Equality in principle, and how it collapses.&lt;/strong&gt; In Evans's stated principle the two sides are equal. "Domain experts object to terms or structures that are awkward or inadequate to convey domain understanding, while developers watch for ambiguity or inconsistency that will trip up design"&lt;sup id="fnref17"&gt;17&lt;/sup&gt;, he writes, which assumes the business side has a veto.&lt;/p&gt;

&lt;p&gt;But equality in principle is itself a view from the system's side. The business comes first and the system attaches to it later, so the seat where words get their meaning belongs to the business side alone. The moment that is put on a negotiating table, the order of precedence is gone.&lt;/p&gt;

&lt;p&gt;In practice it collapses further. To object, the business side needs a venue where it reworks the model with developers every day. In contract work and short requirements phases, the business side meets developers for the first few sessions, and then developers take the words away and pin them. The business side's chance to say "that's not the word" does not come until acceptance testing, when the thing is already built. Even if they say it there, changes after requirements definition come back as change requests with cost and schedule attached, so the business side gives way.&lt;/p&gt;

&lt;p&gt;It is the traveler again, the translator from earlier: the traveler arrives, proposes "let's agree on a common language, for both our sakes," and the language chosen is the traveler's.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DDD itself knew this.&lt;/strong&gt; Bounded context is the device that separates sales's "customer" and accounting's "customer" into different models and keeps both alive. Even so, in practice the meaning converges to one, because the developer is also the one drawing the boundary. Evans only wrote "typically led by developers." Leading and naming rights are different things, but in practice they tend to coincide. That is my read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hard to see from the developer's side.&lt;/strong&gt; It is hard to see from the developer's side because, to the developer, it looks like respect. It is done in good faith: "we'll learn the business's words properly," "we'll listen to the experts." So when the business side says "you're talking down to us," nothing rings a bell, and the complaint does not land.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agents return the naming rights.&lt;/strong&gt; If the business side can ask an agent in its own words, the right to decide what words mean returns to the business side. It is not that the agent-facing glossary becomes unnecessary. The person writing it changes, and the exceptions stop falling outside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What remains
&lt;/h2&gt;

&lt;p&gt;So far I have written that only people inside the business remain. There is work left on the engineering side too.&lt;/p&gt;

&lt;p&gt;Someone builds the environment the agent runs in. Someone decides permissions, logging, and what happens on failure; someone designs data consistency at scale and processing that spans several business functions. Those are still needed. Who approves what, and which state counts as correct, involve business judgment, but the business side makes that judgment, not the translator.&lt;/p&gt;

&lt;p&gt;Coordination across several business functions remains as work close to translation. But that is translation between one function and another. What this article says disappears is only the translation between the business and the code.&lt;/p&gt;

&lt;p&gt;Put simply, the work that remains is not translating between business and code. It is laying the pipes. The business side turns the tap; the engineers who remain run the plumbing.&lt;/p&gt;

&lt;p&gt;Earlier, I &lt;a href="https://zenn.dev/shimo4228/articles/ai-agent-accountability-wall" rel="noopener noreferrer"&gt;described the value structure of an AI harness as an hourglass&lt;/a&gt;. The top (deciding what to build) and the bottom (data, infrastructure, physical constraints) hold their value, and the implementation layer in the middle trends toward zero. Back then I placed "domain knowledge" in the top layer. This article takes that apart. What remains on top is the business itself; "domain knowledge" as something engineers bring in from outside goes out with the middle layer.&lt;/p&gt;

&lt;p&gt;And of what remains, the far larger side, in both headcount and value, is the business side. People who run the actual work on the floor and can use agents are people who can now resolve their own friction themselves. They have no need to teach anyone anything.&lt;/p&gt;

&lt;p&gt;What makes this claim hard is not that there is less work. It is that the very reason engineers were thought necessary has passed into the business side's hands.&lt;/p&gt;

&lt;p&gt;Where "plumbing" ends and "translation" begins, I cannot yet draw the line. I know the line moves with the complexity of the business. Beyond that, I do not know yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is not a blame story
&lt;/h2&gt;

&lt;p&gt;The argument so far may look like a judgment of individuals. It is the outcome of a division of labor, not a fault of individual engineers.&lt;/p&gt;

&lt;p&gt;What needed the phrase "domain knowledge" was the division of labor of an era when implementation was scarce. Inside that division, the translator was necessary, and the vocabulary was rational. DDD was written in 2003 because the gap was real. What changed is not the people but the terrain. This is less "you are wrong" and more "the terrain you grew up on moved."&lt;/p&gt;

&lt;p&gt;If I add one prescription, it is not "learn it" but "go inside." Stand where problems on the floor happen in front of you, and questions arise on their own. "Run one business process yourself" is a shorter path to having questions for an agent than "pick an industry and learn it."&lt;/p&gt;

&lt;p&gt;But the person at the end of that path is "someone inside the business," not "an engineer who knows the domain well."&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Every time I heard "domain knowledge matters," something snagged. Put into words, it is this: what matters is the business, and how to take it in as knowledge is no longer the subject.&lt;/p&gt;

&lt;p&gt;As long as you frame the world in terms of "learning domain knowledge," you keep placing yourself on the scarcity side. Keeping that vocabulary makes it hard to notice that you are standing on the side that disappears.&lt;/p&gt;

&lt;p&gt;The next time you say "learn domain knowledge," check once which bank that view is from. If you are standing on the bridge, decide early which bank to step down to. The first seat to go is the one on the bridge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources and references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic, "Agentic coding and persistent returns to expertise", 2026-06-16. &lt;a href="https://www.anthropic.com/research/claude-code-expertise" rel="noopener noreferrer"&gt;https://www.anthropic.com/research/claude-code-expertise&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Aaron Brethorst, "Domain Expertise Has Always Been the Real Moat", 2026-05-30. &lt;a href="https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/" rel="noopener noreferrer"&gt;https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The San Francisco Standard, "AI writes code now. What's left for software engineers?", 2026-02-19 (Boris Cherny's remarks). &lt;a href="https://sfstandard.com/2026/02/19/ai-writes-code-now-s-left-software-engineers/" rel="noopener noreferrer"&gt;https://sfstandard.com/2026/02/19/ai-writes-code-now-s-left-software-engineers/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Matt Pocock, mattpocock/skills (README section "The Agent Is Way Too Verbose", &lt;code&gt;/domain-modeling&lt;/code&gt;, &lt;code&gt;/grill-with-docs&lt;/code&gt;). &lt;a href="https://github.com/mattpocock/skills" rel="noopener noreferrer"&gt;https://github.com/mattpocock/skills&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Milan Milanović, "What is the future of software engineering?", 2026-09-03. &lt;a href="https://newsletter.techworld-with-milan.com/p/what-is-the-future-of-software-engineering-d52" rel="noopener noreferrer"&gt;https://newsletter.techworld-with-milan.com/p/what-is-the-future-of-software-engineering-d52&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;地家伶人, "Specialization and Boundary Spanning" (専門性と越境), Enterprise IT Conference 2026, 2026-09-10. In Japanese. &lt;a href="https://speakerdeck.com/techtekt/specialization-and-boundary-spanning" rel="noopener noreferrer"&gt;https://speakerdeck.com/techtekt/specialization-and-boundary-spanning&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Marliis Schneider, "The Biggest Winners of the AI Revolution Aren't Engineers", Built In, 2026-07-08. &lt;a href="https://builtin.com/articles/ai-rewards-domain-knowledge" rel="noopener noreferrer"&gt;https://builtin.com/articles/ai-rewards-domain-knowledge&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Findy Tech Blog, "Code w/ Claude Extended | Tokyo" attendance report, 2026-06-12. In Japanese. &lt;a href="https://tech.findy.co.jp/entry/2026/06/12/180000" rel="noopener noreferrer"&gt;https://tech.findy.co.jp/entry/2026/06/12/180000&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic, "Meet the winners of our Built with Opus 4.6 Claude Code hackathon". &lt;a href="https://claude.com/blog/meet-the-winners-of-our-built-with-opus-4-6-claude-code-hackathon" rel="noopener noreferrer"&gt;https://claude.com/blog/meet-the-winners-of-our-built-with-opus-4-6-claude-code-hackathon&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;GIGAZINE (English edition), on the winner of Anthropic's "Built with Opus 4.6" hackathon, 2026-04-25. &lt;a href="https://gigazine.net/gsc_news/en/20260425-anthropic-hackathon/" rel="noopener noreferrer"&gt;https://gigazine.net/gsc_news/en/20260425-anthropic-hackathon/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Eric Evans, &lt;em&gt;Domain-Driven Design: Tackling Complexity in the Heart of Software&lt;/em&gt;, Addison-Wesley, 2003 (Japanese edition: Shoeisha, 2011, supervising translator 今関剛)&lt;/li&gt;
&lt;li&gt;Eric Evans, &lt;em&gt;Domain-Driven Design Reference: Definitions and Pattern Summaries&lt;/em&gt;, Domain Language, 2015. CC BY 4.0. &lt;a href="https://www.domainlanguage.com/ddd/reference/" rel="noopener noreferrer"&gt;https://www.domainlanguage.com/ddd/reference/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Vaughn Vernon, &lt;em&gt;Implementing Domain-Driven Design&lt;/em&gt;, Addison-Wesley, 2013 (Japanese edition: Shoeisha, 2015, translated by 髙木正弘). The better route if you want to relearn the strategic concepts in the order practice uses them&lt;/li&gt;
&lt;li&gt;Michael Jackson, &lt;em&gt;Problem Frames: Analysing and Structuring Software Development Problems&lt;/em&gt;, Addison-Wesley, 2001. One example of how "domain" was used before DDD&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://zenn.dev/shimo4228/articles/ai-agent-accountability-wall" rel="noopener noreferrer"&gt;A Sign on a Climbable Wall: Why AI Agents Need Accountability, Not Just Guardrails&lt;/a&gt; — first appearance of the hourglass model mentioned in the body&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles/domain-knowledge-travelers-vocabulary.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Markdown for every article, plus the index (docs/PUBLICATIONS.md), lives in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;My GitHub&lt;/a&gt; — my research repositories, with DOIs&lt;/li&gt;
&lt;/ul&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;Anthropic, "Agentic coding and persistent returns to expertise" (2026-06-16). Analysis of about 400,000 sessions from about 235,000 users, October 2025 to April 2026. "the ability to steer Claude toward success comes more from command of a domain than from the ability to write code." / "every one of the ten largest occupations in our dataset lands within seven points of software engineers in terms of their success." / "The fastest-growing non-software occupation groups in our sample are management, sales, and legal occupations." &lt;a href="https://www.anthropic.com/research/claude-code-expertise" rel="noopener noreferrer"&gt;https://www.anthropic.com/research/claude-code-expertise&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;Aaron Brethorst, "Domain Expertise Has Always Been the Real Moat" (2026-05-30). "The binding constraint has moved from &lt;em&gt;can you build it&lt;/em&gt; to &lt;em&gt;can you tell whether it's right&lt;/em&gt;." / "Pick an industry, an instrument, a regulatory regime, a physical process, and learn it the way you once learned a programming language or framework." &lt;a href="https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/" rel="noopener noreferrer"&gt;https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/&lt;/a&gt; . Hacker News thread (884 points / 549 comments): &lt;a href="https://news.ycombinator.com/item?id=48340411" rel="noopener noreferrer"&gt;https://news.ycombinator.com/item?id=48340411&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;mattpocock/skills README (repository created 2026-02-03; 260k stars as of 2026-09-12), section "#2: The Agent Is Way Too Verbose". "At the start of a project, devs and the people they're building the software for (the domain experts) are usually speaking different languages. I felt the same tension with my agents. Agents are usually dropped into a project and asked to figure out the jargon as they go. ... The Fix for this is a shared language. It's a document that helps agents decode the jargon used in the project." &lt;a href="https://github.com/mattpocock/skills" rel="noopener noreferrer"&gt;https://github.com/mattpocock/skills&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn4"&gt;
&lt;p&gt;Boris Cherny (creator of Claude Code), speaking on a Y Combinator podcast. Quoted in The San Francisco Standard, "AI writes code now. What's left for software engineers?" (2026-02-19). "Today coding is practically solved." / "We're going to start to see the title of software engineer go away. It's just going to be 'builder' or 'product manager.'" &lt;a href="https://sfstandard.com/2026/02/19/ai-writes-code-now-s-left-software-engineers/" rel="noopener noreferrer"&gt;https://sfstandard.com/2026/02/19/ai-writes-code-now-s-left-software-engineers/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn5"&gt;
&lt;p&gt;Milan Milanović, "What is the future of software engineering?" (Tech World With Milan, 2026-09-03). "What we will see in the next few years is that the best engineers will not be the best coders anymore, and even they will not write any code. They will be domain experts, define architecture and direct agents and judge their outcomes." &lt;a href="https://newsletter.techworld-with-milan.com/p/what-is-the-future-of-software-engineering-d52" rel="noopener noreferrer"&gt;https://newsletter.techworld-with-milan.com/p/what-is-the-future-of-software-engineering-d52&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn6"&gt;
&lt;p&gt;地家伶人 (Persol Career; the slides give no romanized name), "Specialization and Boundary Spanning" (専門性と越境), slides for Enterprise IT Conference 2026 (2026-09-10), slide 9. In Japanese; my translation: "Technical judgment for PdMs, development know-how for IT consultants, business understanding for engineers. Toward a relationship where we think through the next move together, inside the same organization." &lt;a href="https://speakerdeck.com/techtekt/specialization-and-boundary-spanning" rel="noopener noreferrer"&gt;https://speakerdeck.com/techtekt/specialization-and-boundary-spanning&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn7"&gt;
&lt;p&gt;Michael Jackson, &lt;em&gt;Problem Frames: Analysing and Structuring Software Development Problems&lt;/em&gt; (Addison-Wesley, 2001). The domain analysis lineage goes back to Neighbors (1984) and Prieto-Díaz (1987).&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn8"&gt;
&lt;p&gt;Eric Evans, &lt;em&gt;Domain-Driven Design Reference: Definitions and Pattern Summaries&lt;/em&gt; (Domain Language, 2015), Definitions. "A sphere of knowledge, influence, or activity. The subject area to which the user applies a program is the domain of the software." Published under CC BY 4.0: &lt;a href="https://www.domainlanguage.com/ddd/reference/" rel="noopener noreferrer"&gt;https://www.domainlanguage.com/ddd/reference/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn9"&gt;
&lt;p&gt;Eric Evans, &lt;em&gt;Domain-Driven Design: Tackling Complexity in the Heart of Software&lt;/em&gt; (Addison-Wesley, 2003), opening of Part I (the introduction before Chapter 1). "A model is a selectively simplified and consciously structured form of knowledge."&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn10"&gt;
&lt;p&gt;Ibid., Chapter 1, "Crunching Knowledge". "Knowledge crunching is not a solitary activity. A team of developers and domain experts collaborate, typically led by developers."&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn11"&gt;
&lt;p&gt;Ibid., Chapter 2, "Communication and the Use of Language". "On a project without a common language, developers have to translate for domain experts. ... Translation is always inaccurate and hides disconnects in understanding."&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn12"&gt;
&lt;p&gt;Evans (2015), Definitions. "A description of a boundary (typically a subsystem, or the work of a particular team) within which a particular model is defined and applicable."&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn13"&gt;
&lt;p&gt;Marliis Schneider, "The Biggest Winners of the AI Revolution Aren't Engineers" (Built In, 2026-07-08). "She's not replacing the engineer. Instead, she's removing herself as a bottleneck." &lt;a href="https://builtin.com/articles/ai-rewards-domain-knowledge" rel="noopener noreferrer"&gt;https://builtin.com/articles/ai-rewards-domain-knowledge&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn14"&gt;
&lt;p&gt;Gahee Seo, "The 1% problem: How domain expertise + Claude let a 2-person team hit #1 on a global classification benchmark", Code w/ Claude: Extended | Tokyo (hosted by Anthropic, 2026-06-11). Via the Findy Tech Blog attendance report (2026-06-12, in Japanese): &lt;a href="https://tech.findy.co.jp/entry/2026/06/12/180000" rel="noopener noreferrer"&gt;https://tech.findy.co.jp/entry/2026/06/12/180000&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn15"&gt;
&lt;p&gt;The winner of Anthropic's "Built with Opus 4.6" hackathon (February 2026) was a lawyer in California. Anthropic's announcement: &lt;a href="https://claude.com/blog/meet-the-winners-of-our-built-with-opus-4-6-claude-code-hackathon" rel="noopener noreferrer"&gt;https://claude.com/blog/meet-the-winners-of-our-built-with-opus-4-6-claude-code-hackathon&lt;/a&gt; . GIGAZINE (2026-04-25) covered it with comments from Dexter Hadley: &lt;a href="https://gigazine.net/gsc_news/en/20260425-anthropic-hackathon/" rel="noopener noreferrer"&gt;https://gigazine.net/gsc_news/en/20260425-anthropic-hackathon/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn16"&gt;
&lt;p&gt;Ibid., Chapter 2. "A UBIQUITOUS LANGUAGE based on the domain model assumes there is just one model in play." The passage that lists business documents as material is in Chapter 1: "It comes in the form of documents written for the project or used in the business, and lots and lots of talk."&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn17"&gt;
&lt;p&gt;Ibid., Chapter 2. "Domain experts object to terms or structures that are awkward or inadequate to convey domain understanding, while developers watch for ambiguity or inconsistency that will trip up design."&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>discuss</category>
      <category>ddd</category>
      <category>aiagents</category>
      <category>career</category>
    </item>
    <item>
      <title>I Deleted the LLM-Facing Architecture Docs I Had Committed 159 Times in 3 Months: Structure Goes to LSP, Reasons to ADRs, Diagrams to Humans</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Sat, 05 Sep 2026 11:06:56 +0000</pubDate>
      <link>https://dev.to/shimo4228/i-deleted-the-llm-facing-architecture-docs-i-had-committed-159-times-in-3-months-structure-goes-to-1a94</link>
      <guid>https://dev.to/shimo4228/i-deleted-the-llm-facing-architecture-docs-i-had-committed-159-times-in-3-months-structure-goes-to-1a94</guid>
      <description>&lt;p&gt;It started with one diagram.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwdad63zu29mldahh2gxs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwdad63zu29mldahh2gxs.png" alt="Architecture diagram of Contemplative Agent generated by Archify. One-directional imports from cli to adapters to core, external APIs, and evals / testing outside production, drawn as 9 nodes" width="800" height="516"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the morning of September 5, 2026, I tried &lt;a href="https://github.com/tt-a1i/archify" rel="noopener noreferrer"&gt;Archify&lt;/a&gt;, a Claude Code skill (a packaged, reusable capability) that was making the rounds (it generates an architecture-diagram HTML from a JSON spec and validates the geometry), and drew one diagram of my own repository. 9 nodes, 3 cards. All 9 validation checks passed. It looked clean.&lt;/p&gt;

&lt;p&gt;So I typed this: "This is great. Couldn't it replace the codemap?"&lt;/p&gt;

&lt;p&gt;The codemap is &lt;code&gt;docs/CODEMAPS/&lt;/code&gt;, a directory I have kept in the repository since March 2026. Hand-written architecture documents, written for the next session's LLM to read.&lt;/p&gt;

&lt;p&gt;The answer was "no, it can't." A diagram is an outline; the codemap holds thresholds and reasons. That much I expected.&lt;/p&gt;

&lt;p&gt;What I did not expect came two hours later. Instead of turning the codemap into a graph, I &lt;strong&gt;deleted all 6 files, 205,239 bytes&lt;/strong&gt;, and deleted the machinery that generated them too. This article is about why the outcome was "delete" rather than "improve," and the questions you can use to make the same call on your own documents.&lt;/p&gt;

&lt;h2&gt;
  
  
  I replaced "should this become a graph" with "does this need to exist"
&lt;/h2&gt;

&lt;p&gt;In the conversation right after drawing the diagram, I was thinking: if I described the codemap as a graph, couldn't I raise readability for both the human layer and the LLM layer? As a direction, it is natural. Apart from Archify, the family of tools that build a knowledge graph from code and hand it to an LLM keeps growing as of September 2026. &lt;a href="https://github.com/safishamsi/graphify/releases" rel="noopener noreferrer"&gt;Graphify&lt;/a&gt;, for example, shipped 8 releases between August 19 and September 5.&lt;/p&gt;

&lt;p&gt;But I stopped myself from deciding on the spot. The same session had just finished praising Archify, and its judgment was leaning toward "build."&lt;/p&gt;

&lt;p&gt;So I wrote a prompt that handed over only the facts and no conclusion, and asked the question again from zero in a new session. The opening line was this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Don't build" is an acceptable answer. Do not put the conclusion first. Verify the premises yourself before judging.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The reason for starting with the premises was simple: the numbers in my request might be stale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two premises were off
&lt;/h2&gt;

&lt;p&gt;The request said "5 Markdown files, architecture.md is about 15,600 tokens." The first thing the new session did was measure that.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; docs/CODEMAPS/&lt;span class="k"&gt;*&lt;/span&gt;.md
  118891 docs/CODEMAPS/architecture.md
  ...
  205239 total
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls &lt;/span&gt;docs/CODEMAPS | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six files. &lt;code&gt;adapters-moltbook.md&lt;/code&gt; had been left out.&lt;/p&gt;

&lt;p&gt;architecture.md was 118,891 bytes, about 30k tokens. The "15,600" was an estimate written in the file's own header, and it was still the value from August 1, 2026. Never updated.&lt;/p&gt;

&lt;p&gt;In other words, the header that existed to protect freshness was itself stale. At this point I cooled off a little.&lt;/p&gt;

&lt;p&gt;I measured update frequency too.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git log &lt;span class="nt"&gt;--since&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2026-06-01 &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;%h &lt;span class="nt"&gt;--&lt;/span&gt; src | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
197
&lt;span class="nv"&gt;$ &lt;/span&gt;git log &lt;span class="nt"&gt;--since&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2026-06-01 &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;%h &lt;span class="nt"&gt;--&lt;/span&gt; docs/CODEMAPS | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
159
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In three months: 197 commits to source, 159 commits to the codemap. Nearly every time I touched source, I fixed the codemap once.&lt;/p&gt;

&lt;p&gt;A hook, an automated check that fired after each change, prompted this sync, so I never felt any pain. And because there was no pain, the cost was invisible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nothing was blocked
&lt;/h2&gt;

&lt;p&gt;When you re-ask a question from zero, the first question is: whose work is blocked right now, and on what?&lt;/p&gt;

&lt;p&gt;The answer was "nobody's." Of the 68 path references in the codemap body, 2 did not exist, and both were tombstones the text itself described as "retired."&lt;/p&gt;

&lt;p&gt;Not broken. No record of an LLM session reading the codemap and getting stuck.&lt;/p&gt;

&lt;p&gt;The only candidate for real harm was bloat. Re-reading the Data Flow section of architecture.md, the dated parentheticals had become a changelog inlined into the body, and the re-scan paragraph in INDEX.md ran to 9,000 characters. I had been transcribing git log into prose.&lt;/p&gt;

&lt;p&gt;Coming back to the graph idea: a graph does not solve bloat. Nodes and edges hold relationships; what had bloated was the prose about reasons. The question "should this become a graph" was placing a means where there was no problem to solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  I split "LLM understanding" into structure and reasons
&lt;/h2&gt;

&lt;p&gt;So what was the codemap for? "So the next session's LLM understands the repository." I split that "understanding" in two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structure.&lt;/strong&gt; Which file calls which, which layer may import which layer.&lt;br&gt;
&lt;strong&gt;Reasons.&lt;/strong&gt; Why that guard exists, why the import constraint is written in that shape.&lt;/p&gt;

&lt;p&gt;For structure, I actually ran Claude Code's LSP tool. The language server is pyright, which was already in the dev group of &lt;code&gt;pyproject.toml&lt;/code&gt;; no extra configuration needed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LSP incomingCalls  src/contemplative_agent/core/distill.py:104:5

Found 17 incoming calls:

src/contemplative_agent/cli/memory_cmds.py:
  _handle_distill (Function) - Line 39 [calls at: 64:18]

tests/benchmark_distill.py:
  run_benchmark (Function) - Line 166 [calls at: 203:9]

tests/test_distill.py:
  test_basic_distillation (Function) - Line 71 [calls at: 125:18]
  ...(14 more omitted)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 17 call sites of &lt;code&gt;distill()&lt;/code&gt; come back with line numbers. The &lt;code&gt;distill.py&lt;/code&gt; row in the codemap's &lt;code&gt;core-modules.md&lt;/code&gt;, and the "who calls distill" prose in architecture.md, were exactly this answer, copied out by hand.&lt;/p&gt;

&lt;p&gt;What I had spent 159 commits keeping up to date comes out of a single query, always current. Import direction is already enforced as a contract by import-linter.&lt;/p&gt;

&lt;p&gt;Structure never needed to be stored. Derive it per question.&lt;/p&gt;

&lt;p&gt;For reasons, I audited before deleting. I pulled every "why this guard exists" statement out of the Data Flow and Untrusted Boundary sections of architecture.md and grepped the ADRs (Architecture Decision Records: one file per design decision), the docstrings, and the tests.&lt;/p&gt;

&lt;p&gt;Result: 23 of 25 items already lived somewhere else. Most were in their owning ADR, verbatim or at finer granularity. 70 source files cite ADR numbers directly, so the path from code to reason exists too.&lt;/p&gt;

&lt;p&gt;The remaining 2 (a record of an instrument discontinuity, and the rationale for the 512-byte threshold in the watchdog script that monitors for missing output) I moved into an ADR and a script header.&lt;/p&gt;

&lt;p&gt;The codemap was a mirror. Structure was a copy of the code; reasons were a copy of the ADRs. A copy needs updating every time the original moves, and that is what the 159 commits were.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shrink option was the same trap at smaller scale
&lt;/h2&gt;

&lt;p&gt;Even at this point, I had not chosen full deletion. My first pick was a shrink: keep only the Data Flow section and delete the rest.&lt;/p&gt;

&lt;p&gt;The prose about reasons looked valuable.&lt;/p&gt;

&lt;p&gt;I handed this option to a separate agent with no conversation context (its role is to judge build-or-not independently). The verdict came back as a rejection.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the Data Flow &lt;em&gt;is&lt;/em&gt; the accretion (its dated brackets are the changelog inlined). Shrinking the file leaves the growth vector intact, and the hook will re-grow it within the month&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The point was that the source of bloat was neither the header nor INDEX, but the Data Flow section itself. The dated brackets are the changelog inlined, and as long as the hook that makes me write them survives, shrinking it just grows back within a month.&lt;/p&gt;

&lt;p&gt;I could not have seen this on my own. I was looking at "which section has value" and not at "which section grows." The shrink option kept the valuable section, and at the same time kept the growing one.&lt;/p&gt;

&lt;p&gt;I went with delete.&lt;/p&gt;

&lt;h2&gt;
  
  
  What broke when I deleted it
&lt;/h2&gt;

&lt;p&gt;The delete commit touched 40 files, +98 / −3,623 lines. Only 3 things broke mechanically:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;2 tests that asserted the codemap existed&lt;/li&gt;
&lt;li&gt;Relative links from ADRs and the CHANGELOG&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After fixing the links, the test suite came back 3,763 passed / 81 skipped. No functional regression.&lt;/p&gt;

&lt;p&gt;The dangerous part was what did not break. The scan that reads documentation consistency held codemap freshness as two readings.&lt;/p&gt;

&lt;p&gt;When the target directory disappears, those readings &lt;strong&gt;go empty without raising an error&lt;/strong&gt;. Indistinguishable from "nothing wrong."&lt;/p&gt;

&lt;p&gt;Code review caught this, and I changed it to emit &lt;code&gt;FILE_MISSING&lt;/code&gt; when the one remaining freshness target is absent.&lt;/p&gt;

&lt;p&gt;In a deletion, the thing to actually watch is not the test that fails but the instrument that silently goes empty.&lt;/p&gt;

&lt;p&gt;I deleted the generating side too: the skill that writes codemaps, the script that checks their freshness, the hook that detects staleness and routes to regeneration. These lived not in one repository but in the Claude Code configuration shared across all my repositories (under &lt;code&gt;~/.claude/&lt;/code&gt;, which I will call the harness from here on).&lt;/p&gt;

&lt;p&gt;The ADR I had adopted just 4 days earlier, "script the freshness gate," was superseded by a new ADR the same day and lost its force.&lt;/p&gt;

&lt;p&gt;The independent-judgment agent recommended "watch one repository first, and remove from the whole only on a second matching conclusion." I removed the whole mechanism from the harness. As long as the generating side remains, other repositories keep being asked to regenerate, and new repositories grow one again.&lt;/p&gt;

&lt;h2&gt;
  
  
  The human-facing diagram stays, but it gets stamped with its origin
&lt;/h2&gt;

&lt;p&gt;Back to Archify from the opening. If the codemap is gone, is the diagram unnecessary too?&lt;/p&gt;

&lt;p&gt;It is necessary. The reader is different.&lt;/p&gt;

&lt;p&gt;The codemap's reader was the next session's LLM. The LLM can now pull the symbol index itself, so the stored structure document is no longer needed.&lt;/p&gt;

&lt;p&gt;A human opening the README, on the other hand, does not call LSP. Humans need an outline diagram. My policy is to spend the "write so a human can read it" effort only on the README and on output text, so that is where the Archify diagram goes.&lt;/p&gt;

&lt;p&gt;With one condition. &lt;strong&gt;The diagram is an explanation derived from the source of truth, and it persists as a file.&lt;/strong&gt; An explanation that vanishes inside a conversation (like an eli5) has no way to drift, but a 716KB HTML file squats in the repository and, six months later, confidently shows an old diagram even after a module has been removed.&lt;/p&gt;

&lt;p&gt;Archify's architecture schema had a slot ready for exactly this.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"meta"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"repository"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"…"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"revision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;40-character commit SHA&amp;gt;"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"components"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"core"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sources"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"src/contemplative_agent/core/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"…"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Stamp the commit at drawing time into &lt;code&gt;meta.repository.revision&lt;/code&gt;, and the corresponding path into each node's &lt;code&gt;sources&lt;/code&gt;. Then "which commit, and what in it, was this diagram drawn from" stays on the artifact side, and a machine can read how many commits it has drifted from HEAD. Even the camp that advocates stored graphs says "a stale graph is worse than no graph, attach a commit hash and provenance." The condition is the same.&lt;/p&gt;

&lt;p&gt;The diagram at the top of this article does not fill in that slot. When I redraw it for the README, that is the first thing I fill in.&lt;/p&gt;

&lt;h2&gt;
  
  
  It pointed the same way as the official guidance
&lt;/h2&gt;

&lt;p&gt;I looked this up only after finishing the draft: the Claude Code &lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;official docs&lt;/a&gt; say that since the July 9, 2026 version (v2.1.206), the &lt;code&gt;/doctor&lt;/code&gt; command trims from checked-in CLAUDE.md "content Claude can derive from the code (directory layout, dependency lists, architecture overviews)" and keeps only pitfalls, reasons, and conventions. Auto memory does not store architecture or file paths either.&lt;/p&gt;

&lt;p&gt;The contents of my codemap were precisely this trim target. I thought I was going against the grain; the accurate framing is that I had merely confirmed the official call by measuring it in my own repository.&lt;/p&gt;

&lt;p&gt;The means of derivation, on the other hand, are not uniform across agents. As of September 4, 2026, Codex CLI has no built-in LSP, and even in Claude Code, &lt;a href="https://code.claude.com/docs/en/discover-plugins" rel="noopener noreferrer"&gt;language servers do not start in cloud sessions&lt;/a&gt;. Issues where the &lt;a href="https://github.com/anthropics/claude-code/issues/90114" rel="noopener noreferrer"&gt;clangd&lt;/a&gt; and &lt;a href="https://github.com/anthropics/claude-code/issues/91916" rel="noopener noreferrer"&gt;gopls&lt;/a&gt; plugins never register the LSP tool are also open.&lt;/p&gt;

&lt;p&gt;"Delete what can be derived" only holds in an environment that can derive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running the call on your own documents
&lt;/h2&gt;

&lt;p&gt;Few people keep a dedicated &lt;code&gt;docs/CODEMAPS/&lt;/code&gt; directory. More likely you have a single ARCHITECTURE.md, or a "Project structure" section inside CLAUDE.md. The questions are the same.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Is the reader of this document a human or an LLM?&lt;/strong&gt; Human-facing: keep it. LLM-facing: go to the next question&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whose work is blocked right now, and on what?&lt;/strong&gt; If nobody's, the "improvement" has no problem to solve&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For the structural part, can a tool answer each question on demand?&lt;/strong&gt; Actually run LSP's &lt;code&gt;incomingCalls&lt;/code&gt; and &lt;code&gt;workspaceSymbol&lt;/code&gt;, import-linter, or grimp. If the environment cannot, the fix is installing a language server, not reviving the document&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For the reasons part, what fraction already has another owner (ADR, docstring, test)?&lt;/strong&gt; Count it with grep. Mine was 23/25&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does the shrink-and-keep option keep the sections that grow?&lt;/strong&gt; A section full of dated parentheticals and phrases like "since ..." and "added ..." is a changelog inlined, and as long as the hook that makes you write it survives, it bloats again&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can the diagram or explanation you keep be stamped with the commit it was drawn from?&lt;/strong&gt; If the format cannot carry it, keep the lifetime of that explanation inside the conversation&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Where this decision does not hold
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Environments without LSP.&lt;/strong&gt; Codex CLI, cloud sessions, languages whose language server plugin is not ready. With no derivation layer for structure, the grounds for deleting the document disappear. The fix is building the derivation layer, not reviving the document&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repositories where the reasons are in no ADR, docstring, or test.&lt;/strong&gt; If the audit mostly says "nowhere else," that document is not a mirror but the source of truth. You need to build a destination before deleting&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I changed a global setting on a reading from one repository.&lt;/strong&gt; As the independent-judgment agent pointed out, this is a single measurement. When I delete the codemaps in the remaining 9 repositories, I will leave one line per commit message recording whether structural questions were answerable without the codemap. If a repository turns up where they were not, I will install a language server there rather than restore the document&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources and references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;How Claude remembers your project — Claude Code docs&lt;/a&gt; — the spec under which &lt;code&gt;/doctor&lt;/code&gt; trims derivable content from checked-in CLAUDE.md (v2.1.206 and later; retrieved 2026-09-05)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/discover-plugins" rel="noopener noreferrer"&gt;Discover plugins — Claude Code docs&lt;/a&gt; — the list of official language server plugins, and the constraint that they do not start in cloud sessions (retrieved 2026-09-05)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/tt-a1i/archify" rel="noopener noreferrer"&gt;Archify&lt;/a&gt; — the diagramming skill that generated the diagram in this article. &lt;code&gt;meta.repository&lt;/code&gt; and &lt;code&gt;sources&lt;/code&gt; are fields of its architecture schema&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/safishamsi/graphify/releases" rel="noopener noreferrer"&gt;Graphify releases&lt;/a&gt; — 8 releases, v0.9.47 to v0.9.54, between 2026-08-19 and 09-05 (retrieved 2026-09-05)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/openai/codex/releases/tag/rust-v0.153.4" rel="noopener noreferrer"&gt;openai/codex rust-v0.153.4&lt;/a&gt; — the latest version as of 2026-09-04. No mention of LSP. The request for built-in LSP, &lt;a href="https://github.com/openai/codex/issues/8745" rel="noopener noreferrer"&gt;#8745&lt;/a&gt;, is open (retrieved 2026-09-05)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.developersdigest.tech/blog/codebase-knowledge-graphs-ai-coding-agents" rel="noopener noreferrer"&gt;Coding Agents Need Codebase Maps, Not Bigger Prompts — Developers Digest&lt;/a&gt; — an article from the stored-graph camp. "A stale graph is worse than no graph," and the condition that the graph itself must carry when and from which commit it was built (2026-05-26; retrieved 2026-09-05)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/anthropics/claude-code/issues/90114" rel="noopener noreferrer"&gt;clangd-lsp plugin never registers an LSP tool — anthropics/claude-code #90114&lt;/a&gt;, &lt;a href="https://github.com/anthropics/claude-code/issues/91916" rel="noopener noreferrer"&gt;gopls-lsp plugin … LSP tool never available for Go — #91916&lt;/a&gt; — the open issues mentioned in the body (retrieved 2026-09-05)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/contemplative-agent/blob/main/docs/adr/0102-retire-codemaps.md" rel="noopener noreferrer"&gt;ADR-0102: Retire docs/CODEMAPS — Contemplative Agent&lt;/a&gt; — the decision record for this article. The measurements and LSP probe output are frozen in &lt;code&gt;docs/evidence/adr-0102/&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/claude-harness/blob/main/docs/adr/0062-retire-codemap-machinery.md" rel="noopener noreferrer"&gt;ADR-0062: Retire the codemap machinery — claude-harness&lt;/a&gt; — the harness-side removal, and where it departs from the independent judgment's recommendation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/my-dead-code-scan-returned-zero-then-i-deleted-2063-lines-detectors-measure-references-not-4o4e"&gt;My Dead-Code Scan Returned Zero, Then I Deleted 2,063 Lines: Detectors Measure References, Not Consumption&lt;/a&gt; — the previous article: deleting instruments with zero consumers. This article turns the same question on documents&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/after-cutting-my-ai-reviews-i-put-a-complexity-ceiling-in-ruff-1hho"&gt;After Cutting My AI Reviews, I Put a Complexity Ceiling in Ruff&lt;/a&gt; — the one before that: the table for sorting out what can be handed to the machine&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles/codemap-retirement.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Markdown for every article, plus the index (docs/PUBLICATIONS.md), lives in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;My GitHub&lt;/a&gt; — my research repositories, with DOIs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/contemplative-agent" rel="noopener noreferrer"&gt;Contemplative Agent&lt;/a&gt; — the subject measured in this article. The design decisions live in &lt;code&gt;docs/adr/&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>aiagents</category>
      <category>documentation</category>
      <category>architecture</category>
    </item>
    <item>
      <title>I Asked AI to Diagnose My Knowledge Blind Spots — 15 Days of Deadlock Moved in 85 Minutes</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Wed, 02 Sep 2026 10:26:31 +0000</pubDate>
      <link>https://dev.to/shimo4228/i-asked-ai-to-diagnose-my-knowledge-blind-spots-15-days-of-deadlock-moved-in-85-minutes-3j2n</link>
      <guid>https://dev.to/shimo4228/i-asked-ai-to-diagnose-my-knowledge-blind-spots-15-days-of-deadlock-moved-in-85-minutes-3j2n</guid>
      <description>&lt;p&gt;On August 15, 2026, I wrote this in a session with AI:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The insights system isn't really working — Claude Code and I ended up dropping most of the skills through manual approval anyway.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This was about skill management in an AI agent I'd built (Contemplative Agent). Skill definition files kept growing. I had tools to visualize usage frequency and a process to cull low-use ones. But the management tooling was there and the problem still wasn't moving.&lt;/p&gt;

&lt;p&gt;Fifteen days later, on August 30, I wrote this about the same problem:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I'm reading about stocks and flows now, and it's making me think the problem with Contemplative Agent's skill count isn't the pile itself — it's something upstream in the episode-to-pattern-to-skill pipeline.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The problem definition shifted. "How do I manage the growing pile of skills" became "there's a structural problem upstream in how skills are born."&lt;/p&gt;

&lt;p&gt;The trigger for this shift was asking AI to analyze my session history and diagnose knowledge blind spots. The diagnostic criterion that made the biggest difference: &lt;strong&gt;detecting concepts I had independently reinvented&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This article is about how that diagnosis worked and why the problem looked different after 85 minutes of reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  When Better Management Doesn't Move the Problem
&lt;/h2&gt;

&lt;p&gt;Here's the August 15 situation.&lt;/p&gt;

&lt;p&gt;I was managing Contemplative Agent's skills with Claude Code. Visualize usage frequency, get author approval, delete low-frequency skills. Repeat. But even after manually dropping most of the skills, I still felt the system "wasn't working" — that was the opening quote.&lt;/p&gt;

&lt;p&gt;Trying to improve the management process kept bringing me back to the same spot. Fifteen days later, on August 30, I was still seeing the same problem through the same frame.&lt;/p&gt;

&lt;p&gt;When better management doesn't move a problem, the problem might not be management — the framing itself might be wrong. But blind spots in your own framing are, by definition, invisible to you.&lt;/p&gt;

&lt;h2&gt;
  
  
  I Fed AI ~3,000 Turns of Session History
&lt;/h2&gt;

&lt;p&gt;On August 30, I asked Codex (OpenAI's coding agent):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Analyze my sessions, articles, and other materials. Identify knowledge I'm likely missing, and suggest systematic ways to acquire the domain knowledge that would best fill those gaps.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The input: ~3,000 human turns (my inputs from AI conversations), 72 published Zenn articles, and unpublished essays and project documents.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Diagnostic Criterion That Worked — Detecting "Concept Reinvention"
&lt;/h2&gt;

&lt;p&gt;The diagnosis plan AI returned included criteria for judging blind spots.&lt;/p&gt;

&lt;p&gt;First, quality thresholds to suppress false positives:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Only flag a domain when two or more distinct types of evidence and three or more independent instances are present.&lt;/li&gt;
&lt;li&gt;Don't treat absence from conversation as evidence of ignorance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;"Never talked about it" and "knows about it but had no reason to discuss it" are indistinguishable. These thresholds exist to prevent reasoning from absence.&lt;/p&gt;

&lt;p&gt;On top of those, the plan listed several signals to look for. The one that made the biggest difference:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Find concepts the author has independently reinvented where established disciplines already have mature vocabulary and methods.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Detecting "unknown domains" is binary — you know it or you don't. Detecting "reinvented concepts" is different: it points to an existing body of knowledge that already connects to your practice. The moment you start learning, you have contact points. You find a discipline with built-in anchors to problems you're already working on.&lt;/p&gt;

&lt;p&gt;Based on this criterion, the diagnosis flagged systems thinking as a strong candidate.&lt;/p&gt;

&lt;p&gt;I had been using "feedback loop" as a core concept in AKC (a knowledge management framework) — a separate project from Contemplative Agent. But feedback loops are standard vocabulary in systems thinking. I was using the concept without recognizing the existing body of work behind it.&lt;/p&gt;

&lt;p&gt;The diagnosis recommended Donella Meadows' &lt;em&gt;Thinking in Systems: A Primer&lt;/em&gt; as the systematic entry point.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happened 85 Minutes Later
&lt;/h2&gt;

&lt;p&gt;That same day, I bought the book and started reading. Eighty-five minutes in, I wrote this about Contemplative Agent's skill proliferation:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I'm reading about stocks and flows now. With Contemplative Agent, I've been framing the problem as "too many skills — how to prune them." But the problem feels like it's upstream — something in the episode → pattern → skill flow is producing the pile in the first place.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Stock and flow is the most basic concept in systems thinking. A stock is what accumulates. A flow is what moves in and out. When a bathtub overflows, whether the problem is the water level (stock) or the faucet-and-drain structure (flow) determines a completely different intervention point.&lt;/p&gt;

&lt;p&gt;This concept reframed the problem I'd been stuck on for 15 days.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Before (August 15)&lt;/th&gt;
&lt;th&gt;After (August 30)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Problem framing&lt;/td&gt;
&lt;td&gt;Management tooling isn't working&lt;/td&gt;
&lt;td&gt;Structural problem upstream of where skills are born&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Candidate interventions&lt;/td&gt;
&lt;td&gt;Stock side — select and prune the accumulated skills&lt;/td&gt;
&lt;td&gt;Flow side — change the upstream pattern that creates skills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Actions&lt;/td&gt;
&lt;td&gt;Frequency-based selection and manual deletion&lt;/td&gt;
&lt;td&gt;(Framing changed; concrete measures not yet started)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Concrete measures on the After side don't exist yet. But the problem moved from "how to reduce the pile" to "change the structure of why it piles up," and that shifted the intervention point. The reason I was going in circles for 15 days: I was intervening at the wrong level.&lt;/p&gt;

&lt;p&gt;About 50 minutes later, I wrote this about AKC:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I've been putting feedback loops at the center of AKC — looks like that connects to systems thinking too.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Self-confirmation that the "feedback loop" I'd been using independently mapped onto existing systems thinking vocabulary — exactly what the "concept reinvention" criterion had pointed to. The concept already had a body of knowledge behind it, so the connection clicked the moment I started learning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond Stock and Flow
&lt;/h2&gt;

&lt;p&gt;Stock and flow is just the entry point. Meadows' framework has tools that shift how you see problems further. Here are two things waiting beyond the 85-minute reframe.&lt;/p&gt;

&lt;h3&gt;
  
  
  Feedback Loops — Reinforcing and Balancing
&lt;/h3&gt;

&lt;p&gt;Feedback loops come in two types. &lt;strong&gt;Reinforcing loops&lt;/strong&gt; snowball — more leads to more. &lt;strong&gt;Balancing loops&lt;/strong&gt; act like a thermostat — they push back toward a target.&lt;/p&gt;

&lt;p&gt;The skill proliferation problem reads as a reinforcing loop: more skills → more management complexity → more skills to manage the complexity → even more management complexity. What I'd been calling "feedback loops" in AKC corresponded to this reinforcing type.&lt;/p&gt;

&lt;p&gt;What I was doing on August 15 — "delete low-frequency skills" — is a balancing-loop operation. I was applying balancing-loop operations to a structure driven by a reinforcing loop. That's one reading of why it wasn't working.&lt;/p&gt;

&lt;h3&gt;
  
  
  Leverage Points — Where to Push to Move the System
&lt;/h3&gt;

&lt;p&gt;Meadows ranked system interventions by effectiveness across 12 levels. Parameter adjustment (tweaking numbers) is the weakest. Paradigm shift (changing how you see the problem) is the strongest.&lt;/p&gt;

&lt;p&gt;"Adjusting the deletion threshold" — what I'd been doing for 15 days — is a parameter operation. "From stock management to flow structure" is a change in structural recognition, several levels up the leverage hierarchy. The diagnosis didn't just fill a knowledge gap. It moved the intervention to a higher leverage level.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I Haven't Touched Yet
&lt;/h3&gt;

&lt;p&gt;Meadows' framework goes further: how time delays inside feedback loops cause oscillations, how to draw system boundaries, resilience, and self-organization. In a separate lineage, Peter Senge's &lt;em&gt;The Fifth Discipline&lt;/em&gt; systematizes recurring structural patterns across different domains (system archetypes).&lt;/p&gt;

&lt;p&gt;I've learned one entry-point concept — stock and flow — and haven't touched the rest. But the fact that a single entry-point concept reframed 15 days of stagnation shows how close the diagnosed discipline was to my existing practice.&lt;/p&gt;

&lt;p&gt;From here, I'll return to the procedure for reproducing the diagnosis itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Points for Reproducing This
&lt;/h2&gt;

&lt;p&gt;Here's what you need to try this yourself.&lt;/p&gt;

&lt;h3&gt;
  
  
  Input
&lt;/h3&gt;

&lt;p&gt;You need enough material for AI to analyze. In my case, ~3,000 conversation turns and 72 articles. Pattern detection requires repetition — a few dozen turns won't reach the threshold.&lt;/p&gt;

&lt;p&gt;Where session history is stored depends on your tool. Claude Code keeps it under &lt;code&gt;~/.claude/projects/&lt;/code&gt;, Codex under &lt;code&gt;~/.codex/sessions/&lt;/code&gt;, both as JSONL. But raw JSONL includes the AI's own outputs and tool results. Filter to human-authored turns only. If you include AI output, the AI's vocabulary gets misattributed as the author's knowledge.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to Request the Diagnosis
&lt;/h3&gt;

&lt;p&gt;Hand over the session history and artifacts, and specify these three criteria explicitly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Set thresholds&lt;/strong&gt; — Only flag a domain when two or more distinct types of evidence and three or more independent instances are present.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't reason from absence&lt;/strong&gt; — "Never mentioned" is not evidence of ignorance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Look for "concept reinvention"&lt;/strong&gt; — Detect concepts I've independently reinvented where established disciplines already have mature vocabulary and methods.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The third criterion matters most. Domains it surfaces have built-in connection points to your existing practice from the moment you start learning. The motivation and speed are different from studying something because you feel you should.&lt;/p&gt;

&lt;h3&gt;
  
  
  Limitations of This Article
&lt;/h3&gt;

&lt;p&gt;This is a single case. The 85-minute reframe is an observation that the diagnosed discipline connected to an existing problem — not a causal verification of the diagnosis. The reading itself, ongoing thinking, and other stimuli may have contributed. I can't rule them out.&lt;/p&gt;

&lt;p&gt;I built a 12-week study plan based on the diagnosis, but as of writing, I haven't started it. "What happened after learning" is a later story.&lt;/p&gt;

&lt;p&gt;What I can say is one observation: I asked AI to read my session history and detect concept reinvention, and a problem I'd been stuck on for 15 days looked different.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles-en/ai-knowledge-gap-diagnosis.md" rel="noopener noreferrer"&gt;Markdown source on GitHub&lt;/a&gt; — All article Markdown files and the full index (docs/PUBLICATIONS.md) live in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;Author's GitHub&lt;/a&gt; — DOI-registered research repositories&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>systemsthinking</category>
      <category>claudecode</category>
      <category>knowledgemanagement</category>
      <category>metacognition</category>
    </item>
    <item>
      <title>My Dead-Code Scan Returned Zero, Then I Deleted 2,063 Lines: Detectors Measure References, Not Consumption</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Sun, 30 Aug 2026 04:02:49 +0000</pubDate>
      <link>https://dev.to/shimo4228/my-dead-code-scan-returned-zero-then-i-deleted-2063-lines-detectors-measure-references-not-4o4e</link>
      <guid>https://dev.to/shimo4228/my-dead-code-scan-returned-zero-then-i-deleted-2063-lines-detectors-measure-references-not-4o4e</guid>
      <description>&lt;p&gt;Restore the state just before a certain commit, run the weekly dead-code scan that repository has been running all along, and this is what comes back.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"vulture"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"report_prefixes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"src/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"scripts/"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"count"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"candidates"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"parsed_total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;149&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"unparsed_lines"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"stderr_lines"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero candidates. Looks healthy.&lt;/p&gt;

&lt;p&gt;One minute later, the next commit deleted four files — 2,063 lines in total. Counting the documentation updates on the record-keeping side, the whole commit removed 2,065 lines.&lt;/p&gt;

&lt;p&gt;The reason for the delete is in the commit message. &lt;strong&gt;zero consumers&lt;/strong&gt; — because there were none.&lt;/p&gt;

&lt;p&gt;The detector answered zero. The human deleted 2,065 lines. Both calls were correct.&lt;/p&gt;

&lt;p&gt;This article is about why both can be correct, and what I changed once I understood that.&lt;/p&gt;

&lt;h2&gt;
  
  
  I drained the ceiling the day before. The next day, the code was still growing
&lt;/h2&gt;

&lt;p&gt;The subject is Contemplative Agent, the autonomous AI agent I develop, in Python.&lt;/p&gt;

&lt;p&gt;On August 28, 2026, I added a function complexity ceiling to this repository (Ruff's &lt;code&gt;C901&lt;/code&gt;, &lt;code&gt;max-complexity = 15&lt;/code&gt;) and drained all 13 violations the same day. I wrote up how that went in &lt;a href="https://dev.to/shimo4228/after-cutting-my-ai-reviews-i-put-a-complexity-ceiling-in-ruff-1hho"&gt;the previous article&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;At 15:20 the next day, the 29th, I typed this to the agent.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Lately the codebase has put on a lot of lines from the chain of multiple reviews. I have a feeling this project fundamentally should not need that many lines, but it has ballooned. I want it to be a simple product that carries only the code it genuinely needs — how do I get there from here?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I had put the ceiling in and drained it just the day before. It still did not feel like anything had shrunk.&lt;/p&gt;

&lt;p&gt;Partway through the conversation, I asked this myself.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Didn't we have Vulture in there? Is this the kind of thing it can't detect?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Vulture was in there. It even ran automatically every week. And still nothing had shrunk.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the previous article handed to the machine
&lt;/h2&gt;

&lt;p&gt;In the previous article I put out a table that sorted checks into three categories.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Input to the decision&lt;/th&gt;
&lt;th&gt;Destination&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deterministic&lt;/td&gt;
&lt;td&gt;Structure, formatting, existence, matching (complexity, circular imports, &lt;strong&gt;dead code&lt;/strong&gt;, naming conventions)&lt;/td&gt;
&lt;td&gt;The machine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic&lt;/td&gt;
&lt;td&gt;Intent, validity, two-sidedness&lt;/td&gt;
&lt;td&gt;Left to the LLM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mixed&lt;/td&gt;
&lt;td&gt;The machine counts, the LLM interprets&lt;/td&gt;
&lt;td&gt;The machine produces the value, the LLM reads it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I put dead code under "deterministic." It is decided by structure, so it goes to the machine.&lt;/p&gt;

&lt;p&gt;The 2,063 lines that disappeared the next day caught on none of those rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  I re-ran three checks on the revision just before the delete
&lt;/h2&gt;

&lt;p&gt;Saying it in words is weak, so I restored the pre-delete state and measured. I cut the commit before the delete (&lt;code&gt;a06c6be^&lt;/code&gt;) into a worktree and ran three things.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;CA&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~/MyAI_Lab/contemplative-agent
git &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="nv"&gt;$CA&lt;/span&gt; worktree add &lt;span class="nt"&gt;--detach&lt;/span&gt; /tmp/ca-pre a06c6be^

&lt;span class="c"&gt;# 1. The repository's own weekly dead-code scan&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/ca-pre &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; uv run &lt;span class="nt"&gt;--with&lt;/span&gt; &lt;span class="nv"&gt;vulture&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;2.16 python scripts/dead_code_scan.py

&lt;span class="c"&gt;# 2. From Vulture's raw output, pick only the lines about the four files about to be deleted&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/ca-pre &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; uvx vulture@2.16 | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'coselection|sampling_probe|offwindow'&lt;/span&gt;

&lt;span class="c"&gt;# 3. Lint, including the complexity ceiling&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/ca-pre &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; uvx ruff@0.16.0 check &lt;span class="se"&gt;\&lt;/span&gt;
  scripts/coselection_families.py tests/test_coselection_families.py tests/sampling_probe.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here are the results.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;Output for the four files (2,063 lines) about to be deleted&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The weekly dead-code scan&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;"count": 0&lt;/code&gt; (zero candidates, 149 parsed in total)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vulture's raw output&lt;/td&gt;
&lt;td&gt;Exactly one line: &lt;code&gt;tests/sampling_probe.py:83: unused variable 'prompt_eval_count' (60% confidence)&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ruff (with &lt;code&gt;C901&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;All checks passed!&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;About the 912-line script and its 726 lines of tests, none of the three said anything. The one hit that did come out was an unused local variable inside a different file that was also about to be deleted — and at 60% confidence.&lt;/p&gt;

&lt;p&gt;One minute later, those four files were gone.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ git show --stat a06c6be
 scripts/coselection_families.py    | 912 --------------------
 scripts/offwindow-run.sh           | 129 -------
 tests/sampling_probe.py            | 296 ---------------
 tests/test_coselection_families.py | 726 ------------------
 (the remaining 6 files are documentation updates on the record-keeping side; omitted)
 10 files changed, 22 insertions(+), 2065 deletions(-)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Draining the ceiling added lines
&lt;/h2&gt;

&lt;p&gt;There is one more thing I only learned by measuring.&lt;/p&gt;

&lt;p&gt;I said above that the day before, I had "drained all 13 complexity ceiling violations." Here is that commit's Python delta.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ git show --numstat --format= 009baee -- '*.py' | awk '{a+=$1; d+=$2} END {print "+"a, "-"d, "net", a-d}'
+1901 -1086 net 815
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Draining it added 815 lines.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On reflection this was obvious. "Keep one function's branch count at 15 or below" can be satisfied by splitting one function into several helpers. Complexity per function goes down, but total line count goes up. With more function definitions in the file, it tends to go up rather than down.&lt;/p&gt;

&lt;p&gt;The complexity ceiling was optimizing a different quantity from the one I wanted to reduce. It was not that it had no effect — &lt;strong&gt;its effect landed somewhere else&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;At this point I had doubted two checks. The complexity ceiling pointed somewhere else, and dead-code detection pointed at nothing at all. The first is a question about the threshold value; the second is not. I needed to look at what the detector itself measures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The detector measures references, not consumption
&lt;/h2&gt;

&lt;p&gt;This is the mechanism.&lt;/p&gt;

&lt;p&gt;Vulture matches name definitions against uses, and also picks up code after a &lt;code&gt;return&lt;/code&gt;, or code under a condition that cannot hold, as unreachable. So what it measures is &lt;strong&gt;whether there is a path that reaches the code&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If a function is called, it is "used"&lt;/li&gt;
&lt;li&gt;If there is a CLI entry point, it is "used"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The condition for what I wanted to delete, meanwhile, was &lt;strong&gt;whether the readings have a consumer&lt;/strong&gt;. Does anyone read the numbers this script prints and decide something with them?&lt;/p&gt;

&lt;p&gt;Those two are different things. And the second appears in no symbol. The fact that "no human has read the output" is written nowhere in the code.&lt;/p&gt;

&lt;p&gt;I did not need to retract the classification table from the previous article. Inside Vulture's definition of "dead code," that is still the machine's job. What was off was that &lt;strong&gt;the thing I wanted to fold away sat outside that definition&lt;/strong&gt;. The classification was not wrong; the item name I was classifying was coarse.&lt;/p&gt;

&lt;p&gt;There is one more structure at work here: the more tests you write, the more easily things fall outside detection. This repository's own weekly scan declares it in a docstring (excerpt).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Weekly dead-code intake — the fifth deterministic intake.

Scan-wide, report-narrow: vulture&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s scan paths (pyproject [tool.vulture])
include tests/ and evals/ so that code used only by tests resolves as used,
but candidates are reported for src/ and scripts/ only.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To make code that is only called from tests resolve as "used," the scan paths include tests. This is not a mistake. Leave them out and every test-only helper turns up as a candidate, which makes the whole thing unusable.&lt;/p&gt;

&lt;p&gt;The side effect, though, is that &lt;strong&gt;a symbol a test references resolves as "used" even when that reference is the only one in the repository&lt;/strong&gt;. The 726 lines of tests called the 912-line script's internal functions exhaustively, so barely any symbol was left with room to become a candidate.&lt;/p&gt;

&lt;p&gt;What I am calling an &lt;strong&gt;instrument&lt;/strong&gt; here is code written &lt;strong&gt;to be read&lt;/strong&gt;, not to be run. Readings, distributions, calibration scales, audit surfaces. Acting on the world is not the point, so nothing breaks when the consumer disappears. And because nothing breaks, nobody notices.&lt;/p&gt;

&lt;h2&gt;
  
  
  The only thing keeping 912 lines alive was its own test, deleted in the same commit
&lt;/h2&gt;

&lt;p&gt;Here is the breakdown of the four deleted files, with the grounds for each delete. Each one has a dated reason left in a separate record. What I am calling an ADR (Architecture Decision Record) here is text that keeps a design decision and its reason, one file per decision.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What was deleted&lt;/th&gt;
&lt;th&gt;Why its consumer disappeared&lt;/th&gt;
&lt;th&gt;Where it is recorded&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;coselection_families.py&lt;/code&gt;, 912 lines + 726 lines of tests&lt;/td&gt;
&lt;td&gt;The proposal that was the sole destination for its readings was withdrawn on 2026-08-26 (&lt;code&gt;withdrawn&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;A note in ADR-0097&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;tests/sampling_probe.py&lt;/code&gt;, 296 lines&lt;/td&gt;
&lt;td&gt;Done with once the corresponding ADR was written. No consumer after that&lt;/td&gt;
&lt;td&gt;A note in ADR-0047&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;scripts/offwindow-run.sh&lt;/code&gt;, 129 lines&lt;/td&gt;
&lt;td&gt;Nothing in the repository referenced it&lt;/td&gt;
&lt;td&gt;grep on the revision just before the delete&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first row is the clearest example in this article. Search the pre-delete state for anything that actually imports the 912-line script and you get this.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# reuse the worktree cut out earlier&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/ca-pre &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; &lt;span class="s1"&gt;'coselection'&lt;/span&gt; &lt;span class="nt"&gt;--include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'*.py'&lt;/span&gt; &lt;span class="nt"&gt;--include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'*.sh'&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt;
tests/test_coselection_families.py
tests/test_stats.py
scripts/_stats.py
scripts/coselection_families.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Of those four hits, only one is an actual code reference. The &lt;code&gt;test_stats.py&lt;/code&gt; and &lt;code&gt;_stats.py&lt;/code&gt; hits are mentions inside docstrings, and the dependency runs the other way (&lt;code&gt;coselection_families.py&lt;/code&gt; is the one importing helpers from &lt;code&gt;_stats.py&lt;/code&gt;).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# tests/test_coselection_families.py:25
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;coselection_families&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;cf&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The only import keeping the 912-line script alive was the 726 lines of tests that were deleted in the same commit.&lt;/strong&gt; The test calls the script, and the script is called by nobody else. On a reference graph this shape is perfectly healthy. It stays healthy even when nobody outside uses it.&lt;/p&gt;

&lt;p&gt;The third row, &lt;code&gt;offwindow-run.sh&lt;/code&gt;, is simpler still: nothing but the file itself referenced it. A 129-line orphan. The detector still says nothing about it — &lt;strong&gt;Vulture parses Python syntax trees only, and shell scripts were never in scope to begin with&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;One more thing: &lt;code&gt;coselection_families.py&lt;/code&gt; had a distinctive history. This instrument's output was read &lt;strong&gt;exactly once&lt;/strong&gt;, and that number is frozen in a note on an ADR. Read once, purpose served, and then nobody read it again. The code stayed anyway.&lt;/p&gt;

&lt;p&gt;For all three, the record says: the restore point is the delete commit, and if any of them is needed again, recover it from git history rather than rewriting it. The delete is not irreversible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Standing one up was mandatory. Taking it down was optional
&lt;/h2&gt;

&lt;p&gt;I measured why this happens on the repository's record-keeping side.&lt;/p&gt;

&lt;p&gt;This project has a design decision that says "measure before you intervene." Produce a reading before you change anything. I still think it is a good rule. But &lt;strong&gt;the obligation was only on the side that stands things up, and there was none on the side that folds them away&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;ADRs are allowed to carry a section for expiry conditions (&lt;code&gt;## Review-when&lt;/code&gt;). I counted at the revision just before the delete.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/ca-pre
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls &lt;/span&gt;docs/adr/&lt;span class="k"&gt;*&lt;/span&gt;.md | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; ja.md | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; README | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
     101
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; &lt;span class="s1"&gt;'^## Review-when'&lt;/span&gt; docs/adr/&lt;span class="k"&gt;*&lt;/span&gt;.md | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; ja.md
docs/adr/0069-gemma-production-model-and-think-on-value-layer-pipelines.md
docs/adr/0097-consolidator-dissolution-and-skill-store-exit.md
docs/adr/0098-weekly-single-session-and-triage-delegation.md
docs/adr/0099-weekly-report-instrument-redesign.md
docs/adr/0100-retire-chaos-tdd-by-default-mandate.md
docs/adr/0101-instrument-dissolution-mandate.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;6 out of 101. But the last two, 0100 and 0101, are ones I wrote that same day. Count before those and it is &lt;strong&gt;4 out of 99&lt;/strong&gt; — 95 of them said nothing about when they stop being in force.&lt;/p&gt;

&lt;p&gt;It is not that removal was never written about at all. Some individual ADRs said in their body that this instrument comes down once it has served its purpose. The problem is that this was &lt;strong&gt;scattered through prose that no gate reads&lt;/strong&gt;. Not unwritten. Unread.&lt;/p&gt;

&lt;p&gt;Once that is the case, inventory grows easily no matter how good the local decisions are. It does shrink sometimes — the 2,063 lines in this article are exactly that. But it only shrinks when somebody decides one day to go and count. Every item is individually justified, and only the total is nobody's decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write three things when you stand one up
&lt;/h2&gt;

&lt;p&gt;The fix was to give up on after-the-fact detection and move the check to a gate at creation time.&lt;/p&gt;

&lt;p&gt;When you stand up a new instrument, write these three things into the record of that design decision. Anything that cannot be written is not accepted.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;(a) Who reads it, and when&lt;/strong&gt; — a named consumer, plus a frequency or a triggering event. "When it's needed" is not allowed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;(b) How many readings decide what&lt;/strong&gt; — the decision the readings feed, and the number of readings that closes it&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;(c) The removal condition at expiry&lt;/strong&gt; — what counts as done, and when it comes down&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Explicitly banning "when it's needed" in (a) is where the work happens. The instruments I wrote in the past mostly got through on exactly that.&lt;/p&gt;

&lt;p&gt;I put one exception in (b). Exploratory instruments — the ones you stand up precisely because you cannot choose the intervention until the readings exist — may write a &lt;strong&gt;dated review point&lt;/strong&gt; instead of a count. The reader and the date still have to be named. Only the count is waived.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not being able to answer the three is itself the signal that the instrument has no consumer.&lt;/strong&gt; That is the failure shape I wanted to catch.&lt;/p&gt;

&lt;p&gt;Below, I call these three the &lt;strong&gt;consumption plan&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I fixed where it goes, too. The consumption plan goes inside the &lt;code&gt;## Review-when&lt;/code&gt; section of the ADR that produced the instrument, under a &lt;code&gt;### Consumption plan&lt;/code&gt; subheading. One place per instrument, right next to the decision that produced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building no machinery was the condition
&lt;/h2&gt;

&lt;p&gt;This is the part that runs continuous with the previous article and the one before it.&lt;/p&gt;

&lt;p&gt;When I thought about a fix, the first things that came to mind were writing a lint that checks for the presence of a consumption plan, and building a registry of instruments. I dropped both.&lt;/p&gt;

&lt;p&gt;Here is what I decided instead. &lt;strong&gt;No new code, no scheduler, no lint gate, no registry file.&lt;/strong&gt; The obligation is met by adding three sentences to a document I already write, and what enforces it is the same human gate that accepts that record.&lt;/p&gt;

&lt;p&gt;I made the shape of a rejection produce no new artifact either. There is no rejection ledger. A rejection is recorded as &lt;strong&gt;the proposal sitting unaccepted, with a one-line reason&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;There is a reason I did not build a registry, and it is the subject of this article itself. &lt;strong&gt;A registry is itself an instrument you then have to maintain.&lt;/strong&gt; If you cannot answer, about the registry, who reads it, how many readings close the decision, and when it comes down, you have only added the same problem one layer up.&lt;/p&gt;

&lt;h2&gt;
  
  
  How many lines went away
&lt;/h2&gt;

&lt;p&gt;At the end of the session, I asked this.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;So before and after, how much did the line count actually go down?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here are the re-measured numbers, comparing the commit just before I started deciding against the next day.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;th&gt;Delta&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Total Python lines&lt;/td&gt;
&lt;td&gt;94,882&lt;/td&gt;
&lt;td&gt;92,723&lt;/td&gt;
&lt;td&gt;−2,159&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;.py&lt;/code&gt; file count&lt;/td&gt;
&lt;td&gt;234&lt;/td&gt;
&lt;td&gt;231&lt;/td&gt;
&lt;td&gt;−3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;scripts/&lt;/code&gt; (.py + .sh)&lt;/td&gt;
&lt;td&gt;9,544&lt;/td&gt;
&lt;td&gt;8,503&lt;/td&gt;
&lt;td&gt;−1,041&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tests/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;52,170&lt;/td&gt;
&lt;td&gt;51,021&lt;/td&gt;
&lt;td&gt;−1,149&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;src/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;29,965&lt;/td&gt;
&lt;td&gt;29,867&lt;/td&gt;
&lt;td&gt;−98&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The measurement expands each of the two commits and counts it (this is not the sum of the per-commit diffs).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;CA&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~/MyAI_Lab/contemplative-agent
loc&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;T&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; git &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CA&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; archive &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;tar&lt;/span&gt; &lt;span class="nt"&gt;-x&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$T&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        find &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$T&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s1"&gt;'*.py'&lt;/span&gt; &lt;span class="nt"&gt;-exec&lt;/span&gt; &lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;{}&lt;/span&gt; + | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$T&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
loc 64393cf   &lt;span class="c"&gt;# 94882 (just before I started deciding)&lt;/span&gt;
loc 6ce54d9   &lt;span class="c"&gt;# 92723 (the next morning)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that most of the reduction sits in &lt;code&gt;tests/&lt;/code&gt; and &lt;code&gt;scripts/&lt;/code&gt;. What I was able to cut was not the product itself (&lt;code&gt;src/&lt;/code&gt; is −98 lines) but &lt;strong&gt;the side built to be read&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this decision does not hold
&lt;/h2&gt;

&lt;p&gt;Let me put the weak points first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I cannot say this is undetectable in principle.&lt;/strong&gt; What I measured here is the blind spot of a detector that measures references. If you have runtime tracing, or usage logs that record that an output was actually read, the absence of consumers may well be measurable. I did not have that, and standing up a new instrument for the purpose would defeat the point. If you do have it, another road is open to you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The fix is at the "wrote it" stage, not the "did it" stage.&lt;/strong&gt; I decided on the consumption plan obligation on August 29, 2026, and a retroactive stocktake of the existing instruments has not been through even one pass yet. I can talk about effect after at least one lap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This fix may fall into the same trap itself.&lt;/strong&gt; There is a real example. The prior design decision that produced the instrument I deleted this time added 6,355 lines of &lt;code&gt;.py&lt;/code&gt; to stand up the readings, and the conclusion those readings led to deleted 5,672.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(.py only, net delta per commit)
Stand-up:   186dee6 +2,816 / 8757683 +2,008 / c9a1df4 +1,531  → +6,355
Retirement: 47616da +841 −6,513                               → −5,672
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The line count built to justify the retirement was larger than the line count the retirement removed.&lt;/strong&gt; There is still no guarantee that the same thing will not happen to the "consumption plan" section itself. That is why I leaned toward a shape with no machinery in it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I am not claiming that instruments dominate the total.&lt;/strong&gt; Over the window from May 2026 to the end of August, tracked Python went from 29,079 lines to 93,620, a gain of +64,541 lines. I have not measured what share of that instruments account for. The 2,063 lines here are one slice I could identify, not a demonstration that they are the main driver of the growth.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to port this to your own setup
&lt;/h2&gt;

&lt;p&gt;You can run the same judgment without an ADR or RFC acceptance gate. What you need is not the record format but the questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. List the code you wrote "to be read" rather than "to be run."&lt;/strong&gt; Metrics, distribution rollups, audit scripts, one-shot measurement harnesses, dashboard exporters. Anything whose execution does not change product behavior.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. One at a time, try to name the reader.&lt;/strong&gt; "Someone, when it's needed" is not a name. The moment you cannot name one, that is the zero-consumer signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. For the ones you can name, write two more things.&lt;/strong&gt; How many readings close the decision. When it comes down once it has closed. For exploratory ones, a review date instead of a count is fine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Write it next to the decision that produced the instrument.&lt;/strong&gt; If you do not have ADRs, a docstring at the top of the file will do, or one line in the commit message that added the instrument. What matters is that it lands in the eyes of the next person who touches that file. Collect it all into a separate document and that document becomes an instrument with no reader.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Ask at creation time.&lt;/strong&gt; Stocktaking what already exists is heavy work. The creation-time side, in my setup, came down to adding three items to a document I already write. Where agreeing on who the reader is takes negotiation, that is where the cost lands. A stocktake reduces inventory; a gate at creation time stops the increment. Without stopping the increment, the stocktake returns to the same volume every time.&lt;/p&gt;

&lt;p&gt;One last thing, which is also in this article.&lt;/p&gt;

&lt;p&gt;The moment the 912-line instrument's consumer disappeared can be pinpointed. It is August 26, 2026, when the proposal that was the sole destination for its readings was withdrawn. The reason for the withdrawal was not a flaw in the proposal but a judgment that redoing the upstream design came first. I think that was a good call.&lt;/p&gt;

&lt;p&gt;The problem is that &lt;strong&gt;nothing happened on the code side&lt;/strong&gt;. The proposal's state became &lt;code&gt;withdrawn&lt;/code&gt;, and the 912 lines stayed exactly where they were. There was no mechanism connecting the two. The only thing that connected them was one human recounting three days later, on a different errand.&lt;/p&gt;

&lt;p&gt;Ledger state transitions and the life or death of code do not track each other if you leave them alone. That is why you need (c), the removal condition. &lt;strong&gt;Unless you write "when the destination for these readings disappears, this instrument comes down too" right next to the thing that will disappear, nobody can notice that it disappeared.&lt;/strong&gt; Whether there is a consumer is not written inside that file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources and references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://cloudnativenow.com/contributed-content/the-telemetry-debt-crisis-why-cloud-native-teams-are-optimizing-the-wrong-metric/" rel="noopener noreferrer"&gt;The Telemetry Debt Crisis: Why Cloud-Native Teams are Optimizing the Wrong Metric&lt;/a&gt; — the nearest discussion in the observability field. Its prescription — ask who will use a signal at the point you create it, rather than reaching for a detection tool — is close to identical in shape to (a)(b)(c) here (retrieved 2026-08-30). What this article adds is that the subject is code inside a repository rather than telemetry, and that it is a measurement of one and the same case — the code a detector returned zero on was deleted a minute later — rather than an argument&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://michaelscodingspot.com/telemetry-technical-debt/" rel="noopener noreferrer"&gt;Telemetry accumulates like technical debt&lt;/a&gt; — a description of how instrumentation piles up as debt (retrieved 2026-08-30)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/jendrikseipp/vulture" rel="noopener noreferrer"&gt;Vulture&lt;/a&gt; — the dead-code detector used in this article (2.16)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.astral.sh/ruff/rules/complex-structure/" rel="noopener noreferrer"&gt;Ruff &lt;code&gt;C901&lt;/code&gt; (mccabe)&lt;/a&gt; — the complexity ceiling rule&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/after-cutting-my-ai-reviews-i-put-a-complexity-ceiling-in-ruff-1hho"&gt;After Cutting My AI Reviews, I Put a Complexity Ceiling in Ruff&lt;/a&gt; — the previous article: installing the complexity ceiling, and what it measured&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/i-cut-my-ai-review-chain-from-6-stages-to-1-breaking-the-loop-that-never-hits-zero-findings-1moi"&gt;I Cut My AI Review Chain From 6 Stages to 1: Breaking the Loop That Never Hits Zero Findings&lt;/a&gt; — the one before that: how I cut the reviews back&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles/instrument-consumption-plan.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Markdown for every article, plus the index (docs/PUBLICATIONS.md), lives in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;My GitHub&lt;/a&gt; — my research repositories, with DOIs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/contemplative-agent" rel="noopener noreferrer"&gt;Contemplative Agent&lt;/a&gt; — the subject measured in this article. The design decisions live in &lt;code&gt;docs/adr/&lt;/code&gt;, and the consumption plan obligation is ADR-0101&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>aiagents</category>
      <category>staticanalysis</category>
      <category>technicaldebt</category>
    </item>
    <item>
      <title>After Cutting My AI Reviews, I Put a Complexity Ceiling in Ruff</title>
      <dc:creator>shimo4228</dc:creator>
      <pubDate>Fri, 28 Aug 2026 22:31:21 +0000</pubDate>
      <link>https://dev.to/shimo4228/after-cutting-my-ai-reviews-i-put-a-complexity-ceiling-in-ruff-1hho</link>
      <guid>https://dev.to/shimo4228/after-cutting-my-ai-reviews-i-put-a-complexity-ceiling-in-ruff-1hho</guid>
      <description>&lt;p&gt;When you have AI agents writing your code, reviews keep piling up.&lt;/p&gt;

&lt;p&gt;I spent several weeks cutting mine back. And the morning after I cut them, I added a lint rule that puts a ceiling on complexity.&lt;/p&gt;

&lt;p&gt;That looks like subtracting and adding at the same time, but there is no contradiction. &lt;strong&gt;Reviews and lint differ in what they add when you add them.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A review's job is to return findings. Asked about sound code, it will still return something. So every stage you add adds that much more judgment work. What decides the volume is not the state of the code but the disposition of the reviewer.&lt;/p&gt;

&lt;p&gt;Lint is different. It returns only defined violations. No violations, no output. Adding it adds work only in proportion to the violations that actually exist.&lt;/p&gt;

&lt;p&gt;Because of that difference, I cut one and grew the other. &lt;strong&gt;When you are unsure whether to add or remove a check, look at what decides its output volume: the rule, or the reviewer's disposition.&lt;/strong&gt; What a rule decides does not add work when you add it.&lt;/p&gt;

&lt;p&gt;This article covers what it took to actually measure and install a complexity ceiling, and what I found along the way that cannot be handed to a machine. All I used was one Ruff rule, &lt;code&gt;C901&lt;/code&gt;. I built nothing of my own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The morning after cutting reviews, I added lint
&lt;/h2&gt;

&lt;p&gt;On August 27, 2026, I cut my pre-commit reviews from 6 stages to 1. The reason was volume, not accuracy. A reviewer returns something even against sound work, so the more stages I ran, the more decisions I accumulated about whether to fix things. I wrote up how that went in &lt;a href="https://dev.to/shimo4228/i-cut-my-ai-review-chain-from-6-stages-to-1-breaking-the-loop-that-never-hits-zero-findings-1moi"&gt;the previous article&lt;/a&gt;. This one is the sequel: &lt;strong&gt;what I put in its place&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The next morning, on the 28th, I saw a discussion on X about "automated lint rules that are too strict for humans but work on agents." Put a ceiling on function complexity. Put a ceiling on file length. The moment I read it, I wanted it.&lt;/p&gt;

&lt;h3&gt;
  
  
  What decides how much gets added
&lt;/h3&gt;

&lt;p&gt;I had just gone from 6 stages to 1 the day before. Did adding a check today make any sense?&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;What decides the output volume&lt;/th&gt;
&lt;th&gt;What gets added&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LLM review&lt;/td&gt;
&lt;td&gt;The reviewer's disposition&lt;/td&gt;
&lt;td&gt;Standing judgment work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A lint rule with a ceiling&lt;/td&gt;
&lt;td&gt;The rule, plus measured values from the code&lt;/td&gt;
&lt;td&gt;Only the violations that actually exist&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Enabling one &lt;code&gt;C901&lt;/code&gt; rule adds one line to a config file. It runs inside the linter I already have, so no new process appears. What comes back is the count of functions over the threshold, and if nothing is over, it is zero.&lt;/p&gt;

&lt;p&gt;So the two do not sit on the same scale. Cut reviews. Grow lint where it works, and leave it out where it does not.&lt;/p&gt;

&lt;p&gt;This was not the first decision of this shape. On August 15 I had also demoted TDD on new features from mandatory to conditional. What I removed was only &lt;strong&gt;the forced RED→GREEN ordering&lt;/strong&gt;; the coverage floor stayed, still enforced by machine. Drop the procedure the agent has to follow; keep the floor the machine hits. Same operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Left alone, LLM-written code gets complex
&lt;/h2&gt;

&lt;p&gt;This part is my working hypothesis. I want to say that up front.&lt;/p&gt;

&lt;p&gt;The codebase I have had agents write (Contemplative Agent, the LLM agent I develop — 4,580 Python functions) had this complexity distribution before I added any lint.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;p50 = 1  /  p90 = 3  /  p95 = 5  /  p99 = 10  /  max = 35
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First, how to read those numbers.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;p50 = 1&lt;/code&gt; means that if you line up every function by complexity, the middle one is 1. &lt;code&gt;p90 = 3&lt;/code&gt; means the top 10% starts at 3, and &lt;code&gt;p99 = 10&lt;/code&gt; means the top 1% starts at 10. &lt;code&gt;max&lt;/code&gt; is the largest value among them.&lt;/p&gt;

&lt;p&gt;So what does the complexity number itself count? Ruff's &lt;code&gt;C901&lt;/code&gt; (McCabe cyclomatic complexity) &lt;strong&gt;starts a branchless function at 1 and adds 1 for every branch&lt;/strong&gt;. Here is what I actually measured with ruff 0.16.1 (McCabe complexity counts different things in different implementations; the following is Ruff's behavior and will not necessarily match the mccabe package or radon).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Code&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No branches (straight through, top to bottom)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One &lt;code&gt;if&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;if&lt;/code&gt; + &lt;code&gt;elif&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One &lt;code&gt;for&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two &lt;code&gt;except&lt;/code&gt; clauses&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;An &lt;code&gt;if&lt;/code&gt; inside an &lt;code&gt;if&lt;/code&gt; (2 levels of nesting)&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;for&lt;/code&gt; / &lt;code&gt;if&lt;/code&gt; / &lt;code&gt;for&lt;/code&gt; / &lt;code&gt;if&lt;/code&gt;, 4 levels of nesting&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Some things are not counted. &lt;code&gt;else&lt;/code&gt;, &lt;code&gt;and&lt;/code&gt; / &lt;code&gt;or&lt;/code&gt;, the ternary operator, an &lt;code&gt;if&lt;/code&gt; inside a comprehension, and &lt;code&gt;with&lt;/code&gt; all leave complexity untouched. &lt;strong&gt;Nesting depth is not counted either.&lt;/strong&gt; Two levels of nesting and two sequential &lt;code&gt;if&lt;/code&gt;s are both 3.&lt;/p&gt;

&lt;p&gt;Reading the distribution back with those rules: the median of 1 means "a function with no branches at all." p90 at 3 means "roughly two &lt;code&gt;if&lt;/code&gt;s." p99 at 10 means "nine branches." And the maximum of 35 meant &lt;strong&gt;34 branches in a single function&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Functions above 15: 13 of them. 0.3% of the total.&lt;/p&gt;

&lt;p&gt;I have not proven that those 13 exist &lt;em&gt;because&lt;/em&gt; an LLM wrote them. I have not compared against a human-written codebase, and I have not measured whether they are increasing over time. All I measured was the shape of the distribution.&lt;/p&gt;

&lt;p&gt;I decided to put a ceiling on it anyway, because I believe this 0.3% is the side that grows one branch at a time with every edit. "Just one more branch" looks reasonable every single time it comes up in review. With a ceiling, that one branch becomes something to argue about. Without one, it gets merged.&lt;/p&gt;

&lt;p&gt;The point of the discussion I saw on X was not the threshold value either, but the direction of the operation. &lt;strong&gt;When you hit the ceiling, you drain it — you do not raise the threshold.&lt;/strong&gt; One way only.&lt;/p&gt;

&lt;h2&gt;
  
  
  I measured what I had not pushed down to the machine
&lt;/h2&gt;

&lt;p&gt;Saying "push down to the machine whatever can be pushed down" is easy. So how much had I actually not pushed down?&lt;/p&gt;

&lt;p&gt;I counted the lint and build configs across all of my repositories: 82 repos, 104 configs. Then I searched for the terms that set ceilings on complexity or size (&lt;code&gt;C901&lt;/code&gt; / &lt;code&gt;mccabe&lt;/code&gt; / &lt;code&gt;max-complexity&lt;/code&gt; / &lt;code&gt;PLR09*&lt;/code&gt; / &lt;code&gt;max-lines&lt;/code&gt; / &lt;code&gt;radon&lt;/code&gt; / &lt;code&gt;lizard&lt;/code&gt;, and so on).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Zero hits.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is a re-run with a narrower target, done while writing this article (August 28, 2026).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 - &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;PY&lt;/span&gt;&lt;span class="sh"&gt;'
import os,glob,re
roots=[os.path.expanduser("~/MyAI_Lab"),os.path.expanduser("~/.claude")]
repos={d for r in roots for d in glob.glob(os.path.join(r,"*"))
       if os.path.isdir(os.path.join(d,".git"))}
repos |= {r for r in roots if os.path.isdir(os.path.join(r,".git"))}
names=["pyproject.toml","setup.cfg",".flake8","ruff.toml",".ruff.toml",
       "eslint.config.mjs",".eslintrc.json",".eslintrc.js",".swiftlint.yml",
       "tox.ini",".pylintrc"]
pat=re.compile(r"C901|mccabe|max-complexity|PLR09|max-lines|max-module-lines"
               r"|size-limit|radon|lizard|cyclomatic_complexity|file_length",re.I)
cfgs=0;hits=[]
for repo in sorted(repos):
    for n in names:
        p=os.path.join(repo,n)
        if os.path.isfile(p):
            cfgs+=1
            t=open(p,encoding="utf-8",errors="replace").read()
            if pat.search(t):
                hits.append((os.path.basename(repo),n,sorted(set(pat.findall(t)))))
print(f"repos: {len(repos)}  configs: {cfgs}  hits: {len(hits)}")
for h in hits: print("  ",h)
&lt;/span&gt;&lt;span class="no"&gt;PY
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;repos: 76  configs: 20  hits: 1
   ('contemplative-agent', 'pyproject.toml', ['C901', 'max-complexity', 'mccabe'])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The single hit is the one I added that same day, as part of the work described in this article. Before that it was zero.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why "maximum strict" still let this through
&lt;/h3&gt;

&lt;p&gt;My own conventions say "lint should be at maximum strict by default." I had been following that, and I still had zero.&lt;/p&gt;

&lt;p&gt;The reason is that ceiling rules are not in the default sets of the major linters.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ruff expanded its default rules from 59 to 413 in 0.16.0, but &lt;code&gt;C90&lt;/code&gt; (complexity) and the &lt;code&gt;PLR09*&lt;/code&gt; subset of &lt;code&gt;PLR&lt;/code&gt; (too-many-branches / too-many-arguments / too-many-statements and other ceiling rules) remain outside the defaults&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ruff has no rule equivalent to a file line-count ceiling.&lt;/strong&gt; In the Pylint-compatibility tracking issue &lt;a href="https://github.com/astral-sh/ruff/issues/970" rel="noopener noreferrer"&gt;astral-sh/ruff#970&lt;/a&gt;, the corresponding &lt;code&gt;too-many-lines&lt;/code&gt; entry carries the note "not compatible with the formatter" and is unimplemented&lt;/li&gt;
&lt;li&gt;Cognitive complexity has been sitting as &lt;a href="https://github.com/astral-sh/ruff/issues/2418" rel="noopener noreferrer"&gt;astral-sh/ruff#2418&lt;/a&gt;, open since January 2023&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So even if you follow "make the default rule set strict" faithfully, the ceiling rules fall through structurally. It was not that I forgot to write the setting — &lt;strong&gt;my convention was written somewhere it could not reach&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The contrast was in the Swift repositories I have. SwiftLint enables &lt;code&gt;cyclomatic_complexity&lt;/code&gt; (warning 10 / error 20) and &lt;code&gt;file_length&lt;/code&gt; (warning 400 / error 1000) &lt;strong&gt;by default&lt;/strong&gt;. Without a single line in &lt;code&gt;.swiftlint.yml&lt;/code&gt;, they were in force.&lt;/p&gt;

&lt;p&gt;And there, the ceiling was actually working as a cutting force. The largest Swift file was 397 lines. Three lines short of the 400-line warning threshold.&lt;/p&gt;

&lt;p&gt;Whether your toolchain ships ceilings by default completely changes what the same "maximum strict" convention produces. This is not a discipline problem. It is a tool-defaults problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  You cannot set the threshold globally
&lt;/h2&gt;

&lt;p&gt;So what number do you pick? My first plan was to settle on one number shared across all repositories.&lt;/p&gt;

&lt;p&gt;I could not.&lt;/p&gt;

&lt;p&gt;Tip the threshold to 0 and Ruff prints the measured value for every function. That gets you the whole distribution in a single run.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvx ruff@0.16.1 check &lt;span class="nt"&gt;--isolated&lt;/span&gt; &lt;span class="nt"&gt;--no-cache&lt;/span&gt; &lt;span class="nt"&gt;--output-format&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--select&lt;/span&gt; C901 &lt;span class="nt"&gt;--config&lt;/span&gt; &lt;span class="s2"&gt;"lint.mccabe.max-complexity=0"&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &amp;lt;files&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Measurements across my 7 public repositories, 442 Python files (as of August 28, 2026; &lt;code&gt;~/.claude&lt;/code&gt; is my personal collection of Claude Code operations scripts, and the Contemplative Agent row shows values after the cutting described below).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;repo&lt;/th&gt;
&lt;th&gt;functions&lt;/th&gt;
&lt;th&gt;p50&lt;/th&gt;
&lt;th&gt;p90&lt;/th&gt;
&lt;th&gt;p99&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;th&gt;&amp;gt;10&lt;/th&gt;
&lt;th&gt;&amp;gt;15&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;~/.claude&lt;/td&gt;
&lt;td&gt;949&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;contemplative-agent&lt;/td&gt;
&lt;td&gt;4,746&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;pdf2anki&lt;/td&gt;
&lt;td&gt;954&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tiny-lm-lab&lt;/td&gt;
&lt;td&gt;111&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;daily-quest-generator&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;27&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;active-inference-viz&lt;/td&gt;
&lt;td&gt;111&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;einstein-arena&lt;/td&gt;
&lt;td&gt;261&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;p99 ranges from 5 to 26 — a spread of more than 5x. File line counts measured the same day had a p90 spread from 184 lines to 901.&lt;/p&gt;

&lt;p&gt;What happens if you paste &lt;code&gt;C901 = 10&lt;/code&gt; across all of them? In tiny-lm-lab nothing trips, and it protects nothing. In daily-quest-generator the existing code goes red immediately and work stops. It becomes &lt;strong&gt;a number that matches the reality of neither repository&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I once retired a fixed checklist that did not match the reality of the repositories it ran against. I was about to rebuild the same thing in the shape of a threshold.&lt;/p&gt;

&lt;p&gt;What I did instead: &lt;strong&gt;measure that repository's distribution, then place the threshold where only the current outliers go red&lt;/strong&gt;. Thresholds are allowed to differ per repository. The only thing decided globally is the procedure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some things cannot be pushed down
&lt;/h2&gt;

&lt;p&gt;The same measurements told me one more thing. &lt;strong&gt;A complexity ceiling and a file line-count ceiling behave in opposite ways.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;C901 &amp;gt; 15&lt;/th&gt;
&lt;th&gt;File LOC &amp;gt; 500&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Count&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;83&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of which test code&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;43 (52%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All 7 functions over complexity 15 were production code. Not one test.&lt;/p&gt;

&lt;p&gt;Open the largest test file in Contemplative Agent and the reason is obvious. &lt;code&gt;tests/test_agent.py&lt;/code&gt; is 4,028 lines with 209 &lt;code&gt;def test_&lt;/code&gt;s. Measuring the 233 functions in it, helpers and fixtures included, &lt;strong&gt;224 (96%) had complexity 1&lt;/strong&gt; — zero branches. The maximum was 4. It is long not because it is tangled, but because independent cases are lined up next to each other.&lt;/p&gt;

&lt;p&gt;File line counts went the other way: a majority of the 83 files over 500 lines were tests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Whether to split tests is genuinely contested
&lt;/h3&gt;

&lt;p&gt;I wanted to conclude "so tests don't need splitting," but when I looked into it, the consensus was not that simple.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The "no need to split" side&lt;/strong&gt;: Google's &lt;a href="https://testing.googleblog.com/2019/12/testing-on-toilet-tests-too-dry-make.html" rel="noopener noreferrer"&gt;DAMP principle&lt;/a&gt; argues that because tests do not have tests of their own, it matters that a human can verify their correctness by eye — worth paying some code duplication for. If you prioritize each test being readable on its own, the length of the whole file is not a burden&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The "same standard" side&lt;/strong&gt;: the claim that &lt;a href="https://www.ontestautomation.com/on-treating-your-test-code-like-production-code/" rel="noopener noreferrer"&gt;test code should be treated as production code&lt;/a&gt; also exists. That position says lint should be applied to test automation code too&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;SonarQube sits between the two, and it is telling. The &lt;a href="https://docs.sonarsource.com/sonarqube-server/instance-administration/analysis-functions/analysis-scope/exclude-from-coverage-duplication" rel="noopener noreferrer"&gt;official documentation&lt;/a&gt; provides &lt;strong&gt;the configuration steps&lt;/strong&gt; for excluding tests from duplication analysis, but they are not excluded by default. The &lt;strong&gt;rationale&lt;/strong&gt; — "test duplication is intentional, so it should be excluded" — is not in the official documentation; it lives in the &lt;a href="https://community.sonarsource.com/t/duplicated-lines-of-code-in-tests/137785" rel="noopener noreferrer"&gt;community forum&lt;/a&gt;. The tool goes as far as offering the option, and the judgment is left to each repository.&lt;/p&gt;

&lt;p&gt;The position I took is: apply lint to tests too, but decide rule by rule. Writing &lt;strong&gt;forbidden in production, allowed in tests&lt;/strong&gt; with ESLint overrides or Ruff per-file-ignores is an officially intended use. Complexity does not go red in tests, so it stays applied. Line count would require designing exemptions, so I put it on hold.&lt;/p&gt;

&lt;p&gt;Which is to say: if you naively add a file line-count ceiling, &lt;strong&gt;most of the resulting work becomes "splitting test files."&lt;/strong&gt; The reason I wanted a ceiling was to cut down unreadable code the agents had grown. Splitting tests has nothing to do with that intent.&lt;/p&gt;

&lt;p&gt;That is where I noticed one more thing to check. &lt;strong&gt;Does what the machine flags match what I actually want cut?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For a complexity ceiling it matches. What goes red is the same thing I wanted cut in the first place. For a line-count ceiling it does not. You cannot install it without first deciding how to treat tests.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Even among "ceilings," some can be pushed down to the machine and some cannot.&lt;/strong&gt; I installed only complexity and put the line-count ceiling on hold.&lt;/p&gt;

&lt;h2&gt;
  
  
  I installed it and drained it the same day
&lt;/h2&gt;

&lt;p&gt;I set &lt;code&gt;max-complexity = 15&lt;/code&gt; in Contemplative Agent.&lt;/p&gt;

&lt;p&gt;I run &lt;code&gt;ruff check&lt;/code&gt; before every commit and block the commit on errors. Add a new rule there and every existing violation becomes an error from that instant.&lt;/p&gt;

&lt;p&gt;At install time there were 13 violations. Left as-is, commits would break in the middle of changes that have nothing to do with complexity. When that happens, the agent either goes off to fix violations unrelated to the work in front of it, or it learns the procedure for skipping the check. Both are outcomes I want to avoid.&lt;/p&gt;

&lt;p&gt;So I wrote those 13 into &lt;code&gt;per-file-ignores&lt;/code&gt; as a &lt;strong&gt;backlog to drain&lt;/strong&gt;, and &lt;strong&gt;started from zero errors&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Then I &lt;strong&gt;drained all 13 the same day&lt;/strong&gt;. The largest was the 35 in &lt;code&gt;never_selected_metrics.py&lt;/code&gt;; I split it and the other 12 into helpers at or below 15. I confirmed behavior was unchanged by running the same inputs through the old and new code and comparing outputs across 4,767 test cases (zero mismatches). The backlog was empty.&lt;/p&gt;

&lt;h3&gt;
  
  
  The backlog that empties, and the exemption that does not
&lt;/h3&gt;

&lt;p&gt;That said, exemptions did not go to zero. I ran into this in my own environment while writing.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;ruff check&lt;/code&gt; I run before commits targets four directories: &lt;code&gt;src tests scripts evals&lt;/code&gt;. Outside them, under &lt;code&gt;docs/evidence/&lt;/code&gt;, two functions with complexity 16 and 20 were still sitting there.&lt;/p&gt;

&lt;p&gt;I thought about whether to cut them, and decided not to. Those two are verbatim records from verifying a past rewrite, and one of them is the baseline the outputs were compared against. Split them up for readability and they stop working as evidence.&lt;/p&gt;

&lt;p&gt;The problem was that leaving them meant &lt;strong&gt;the next session to touch those files would eat an error with no warning&lt;/strong&gt;. Even outside the full-scan target, the pre-commit check looks at every staged &lt;code&gt;.py&lt;/code&gt;. An undeclared exemption is not an exemption, it is a landmine.&lt;/p&gt;

&lt;p&gt;So I added it to the existing exclusion line as a permanent exemption with a reason.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="c"&gt;# Promoted one-off measurement scripts — the progress and readout prints are the UI.&lt;/span&gt;
&lt;span class="c"&gt;# C901 is permanently exempt here too: these are frozen verbatim records (including&lt;/span&gt;
&lt;span class="c"&gt;# the baseline values used for output comparison), and cutting them destroys their&lt;/span&gt;
&lt;span class="c"&gt;# value as evidence. This is not a backlog to drain&lt;/span&gt;
&lt;span class="py"&gt;"docs/evidence/**"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"T20"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"C901"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is &lt;strong&gt;not to mix the two kinds of exemption&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Backlog to drain&lt;/th&gt;
&lt;th&gt;Permanent exemption with a reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Contents&lt;/td&gt;
&lt;td&gt;Existing code that was over the ceiling at install time&lt;/td&gt;
&lt;td&gt;Things that must not be cut&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goal&lt;/td&gt;
&lt;td&gt;Empty it&lt;/td&gt;
&lt;td&gt;Never empty it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Additions&lt;/td&gt;
&lt;td&gt;Not allowed (a new violation is a design problem)&lt;/td&gt;
&lt;td&gt;Allowed, with a written reason&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Mix those two into one list and you can no longer tell whether anything is left to cut by asking "is the exemption list empty?" Keep them separate and the backlog can be operated as "keep it empty."&lt;/p&gt;

&lt;h3&gt;
  
  
  Detection goes to the machine, the cutting goes to the LLM
&lt;/h3&gt;

&lt;p&gt;The cutting work did not need the judgment of a frontier model. The machine had already finished the detection, and what remained was fixing the places the machine pointed at. In fact, what I considered before starting was "which model is enough," not "where is it complex."&lt;/p&gt;

&lt;p&gt;Adding lint did not make the LLM unnecessary. What actually happened is that &lt;strong&gt;the LLM's role moved from "judge where it is complex" to "cut down where the machine pointed."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I put a comment on the threshold line.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[tool.ruff.lint.mccabe]&lt;/span&gt;
&lt;span class="c"&gt;# Budget rule: drain, do not raise — any change to this number needs a dated&lt;/span&gt;
&lt;span class="c"&gt;# reason in .claude/verify.md&lt;/span&gt;
&lt;span class="c"&gt;# Distribution as-of 2026-08-28 (4,580 functions):&lt;/span&gt;
&lt;span class="c"&gt;# p50=1 / p90=3 / p95=5 / p99=10 / max=35.&lt;/span&gt;
&lt;span class="c"&gt;# 10 was rejected as the threshold: it would have put 24 files on the&lt;/span&gt;
&lt;span class="c"&gt;# exemption list, blinding the modules that see the most editing.&lt;/span&gt;
&lt;span class="py"&gt;max-complexity&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.claude/verify.md&lt;/code&gt; is where I record verification procedures and decisions in my environment. If you do the same, read that as wherever you keep your own decision records.&lt;/p&gt;

&lt;p&gt;You can write the convention in some other document, but no session goes back to read it at the moment lint throws an error. What is in front of you then is the lint output and the config. &lt;strong&gt;If you want "drain, do not raise" to land, the delivery address is that very line.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When this reasoning does not apply
&lt;/h2&gt;

&lt;p&gt;Let me be honest about the limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Drain, do not raise" is not machine-enforced.&lt;/strong&gt; It is an operating convention carried by a comment. If someone raises the threshold to get around it, I cannot detect that. I rejected the idea of building something new to detect it, precisely because that would add another thing that runs — the thing I had decided to cut the day before. If I observe one real instance of the workaround, I will think about it then.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If what goes red does not match what you want cut, this axis does not work.&lt;/strong&gt; The machine can decide in an instant, but if it points at the wrong things, you get more work without moving the intent. That is what I nearly walked into with the file line-count ceiling. Before "machine or LLM," look at what that machine will point at.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check what your lint command actually covers before you pick a threshold.&lt;/strong&gt; A ceiling only works out to the directories that command looks at. I carried two violations outside that range for a while without noticing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I have not measured the effect of installing it.&lt;/strong&gt; What I can observe stops at "13 violations, drained the same day." Whether the bug rate dropped, whether readability improved, whether the agents' generation tendencies changed — I cannot say anything about any of it yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Porting this to your own setup
&lt;/h2&gt;

&lt;p&gt;Four steps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. List what you currently have LLM reviews doing.&lt;/strong&gt; Write it out by concern, not by stage name. Like "naming," "duplication," "complexity," "error handling," "security."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Sort each item into three buckets.&lt;/strong&gt; Three, not two.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Bucket&lt;/th&gt;
&lt;th&gt;Input to the decision&lt;/th&gt;
&lt;th&gt;Destination&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deterministic&lt;/td&gt;
&lt;td&gt;Structure, format, existence, matching (complexity, circular imports, dead code, naming conventions)&lt;/td&gt;
&lt;td&gt;Machine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic&lt;/td&gt;
&lt;td&gt;Intent, soundness, two-sidedness (whether a design is right, alignment with requirements, post-hoc justification)&lt;/td&gt;
&lt;td&gt;Leave with the LLM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mixed&lt;/td&gt;
&lt;td&gt;The machine counts, the LLM interprets (distribution of term usage, the denominator behind a number)&lt;/td&gt;
&lt;td&gt;Machine produces the value, LLM reads it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And &lt;strong&gt;when in doubt, tip the item toward semantic&lt;/strong&gt;. An item you mechanized by mistake passes false negatives wearing the face of "checked." The worst outcome is that misses quietly increase in exchange for the relief of one fewer stage.&lt;/p&gt;

&lt;p&gt;For every item you assigned to the machine, one more question. &lt;strong&gt;Does what it flags match what you want cut?&lt;/strong&gt; For things that do not match — like file line counts — the rule may decide the output volume, but what comes back will be things you did not want cut.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Set the threshold from the measured distribution.&lt;/strong&gt; Do not pick the number first. Tip the ceiling to 0, take the distribution for everything, and place the threshold where only your current outliers go red. In Python, one line gets you there.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uvx ruff@0.16.1 check &lt;span class="nt"&gt;--isolated&lt;/span&gt; &lt;span class="nt"&gt;--no-cache&lt;/span&gt; &lt;span class="nt"&gt;--output-format&lt;/span&gt; json &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--select&lt;/span&gt; C901 &lt;span class="nt"&gt;--config&lt;/span&gt; &lt;span class="s2"&gt;"lint.mccabe.max-complexity=0"&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &amp;lt;files&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;4. Start from zero errors the moment you add it.&lt;/strong&gt; Adding a new rule turns every existing violation into an error on the spot. If you run lint before commits or in CI, changes unrelated to this will stop passing. What happens then is one of two things: the agent goes off to fix violations unrelated to the work in front of it, or the habit of skipping the check sets in. Write the existing violations into an exemption list and start from a state that passes. &lt;strong&gt;If you see errors on day one, what is wrong is not the threshold — it is how you wrote the exemptions.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Then decide the plan for removing those exemptions. When you do, write the backlog to drain and the permanent exemptions with reasons as separate groups.&lt;/p&gt;

&lt;p&gt;And one comment on that threshold line. &lt;strong&gt;When you go over, drain it — do not raise it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;These four steps are a summary of the procedure I normally use. The original (&lt;a href="https://github.com/shimo4228/claude-harness/blob/main/skills/review-to-lint/SKILL.md" rel="noopener noreferrer"&gt;review-to-lint&lt;/a&gt;) is written as "pull the mechanically decidable items out of the reviewer's checklist, and thin the reviewer down to semantic checks only." The division of labor is one line: &lt;strong&gt;code decides whether it is there, and the LLM decides whether it is sound.&lt;/strong&gt; What I did this time was the version where the machine side is handled by one off-the-shelf lint rule instead of a script of my own.&lt;/p&gt;

&lt;p&gt;The way to pick a threshold and the "drain, do not raise" convention itself went into the skill that builds per-repository verification procedures (&lt;a href="https://github.com/shimo4228/claude-harness/blob/main/skills/verify-bootstrap/SKILL.md" rel="noopener noreferrer"&gt;verify-bootstrap&lt;/a&gt;) as a 20-line proviso. No new skill, no new hook, no new script.&lt;/p&gt;

&lt;p&gt;When I am unsure whether to add or remove a check, I look at two things.&lt;/p&gt;

&lt;p&gt;One is what I opened with: &lt;strong&gt;what decides the output volume&lt;/strong&gt;. What the reviewer's disposition decides adds standing work when you add it. What a rule plus measured values from the code decides adds nothing when there are no violations.&lt;/p&gt;

&lt;p&gt;The other is what I noticed along the way: &lt;strong&gt;whether what goes red matches what you want cut&lt;/strong&gt;. That is where the file line-count ceiling caught me.&lt;/p&gt;

&lt;p&gt;Anything that satisfies both, I want more of.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources and references
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.astral.sh/ruff/default-rules/" rel="noopener noreferrer"&gt;Ruff — Default Rules&lt;/a&gt; (retrieved 2026-08-28; &lt;code&gt;C901&lt;/code&gt; is not in the default list)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/astral-sh/ruff/releases/tag/0.16.0" rel="noopener noreferrer"&gt;Ruff v0.16.0 release notes&lt;/a&gt; (retrieved 2026-08-28; defaults went from 59 rules to 413)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/astral-sh/ruff/issues/970" rel="noopener noreferrer"&gt;astral-sh/ruff#970 — "Implement Pylint"&lt;/a&gt; (retrieved 2026-08-28; the Pylint-compatibility tracking issue. The &lt;code&gt;too-many-lines&lt;/code&gt; entry, the equivalent of a file line-count ceiling, carries a "not compatible with the formatter" note and is unimplemented)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/astral-sh/ruff/issues/2418" rel="noopener noreferrer"&gt;astral-sh/ruff#2418 — "Implement flake8-cognitive-complexity"&lt;/a&gt; (retrieved 2026-08-28; opened 2023-01-31 and open ever since)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.astral.sh/ruff/rules/#mccabe-c90" rel="noopener noreferrer"&gt;Ruff documentation — mccabe (C90)&lt;/a&gt; (retrieved 2026-08-28)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/claude-harness/tree/main/skills" rel="noopener noreferrer"&gt;claude-harness — review-to-lint / verify-bootstrap&lt;/a&gt; (the source of record for the procedures in this article, as of 2026-08-28)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://realm.github.io/SwiftLint/rule-directory.html" rel="noopener noreferrer"&gt;SwiftLint — Rule Directory&lt;/a&gt; (retrieved 2026-08-28; &lt;code&gt;cyclomatic_complexity&lt;/code&gt; / &lt;code&gt;file_length&lt;/code&gt; are enabled by default)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://testing.googleblog.com/2019/12/testing-on-toilet-tests-too-dry-make.html" rel="noopener noreferrer"&gt;Google Testing Blog — Tests Too DRY? Make Them DAMP!&lt;/a&gt; (retrieved 2026-08-28; test readability takes priority over duplication)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.sonarsource.com/sonarqube-server/instance-administration/analysis-functions/analysis-scope/exclude-from-coverage-duplication" rel="noopener noreferrer"&gt;SonarQube — Excluding from coverage or duplication&lt;/a&gt; (retrieved 2026-08-28; the configuration steps for exclusion. Not excluded by default)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://community.sonarsource.com/t/duplicated-lines-of-code-in-tests/137785" rel="noopener noreferrer"&gt;Sonar Community — Duplicated lines of code in tests&lt;/a&gt; (retrieved 2026-08-28; the discussion treating test duplication as intentional. Not official documentation)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://eslint.org/docs/latest/rules/max-lines" rel="noopener noreferrer"&gt;ESLint — max-lines&lt;/a&gt; / &lt;a href="https://docs.astral.sh/ruff/settings/" rel="noopener noreferrer"&gt;Ruff — Settings (per-file-ignores)&lt;/a&gt; (retrieved 2026-08-28; rule selection per file type)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Related links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/shimo4228/i-cut-my-ai-review-chain-from-6-stages-to-1-breaking-the-loop-that-never-hits-zero-findings-1moi"&gt;I Cut My AI Review Chain From 6 Stages to 1: Breaking the Loop That Never Hits Zero Findings&lt;/a&gt; — the previous article: how I cut reviews back, and what I measured&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228/zenn-content/blob/main/articles/lint-as-subtraction.md" rel="noopener noreferrer"&gt;The Markdown source of this article (GitHub)&lt;/a&gt; — the Markdown for every article, plus the index (docs/PUBLICATIONS.md), lives in the same repository&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/shimo4228" rel="noopener noreferrer"&gt;My GitHub&lt;/a&gt; — my research repositories, with DOIs&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>claudecode</category>
      <category>aiagents</category>
      <category>staticanalysis</category>
      <category>codereview</category>
    </item>
  </channel>
</rss>
