<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lex</title>
    <description>The latest articles on DEV Community by Lex (@lexosi).</description>
    <link>https://dev.to/lexosi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4065654%2F2659e8ca-99d3-45cc-803d-4cdf9e55f001.png</url>
      <title>DEV Community: Lex</title>
      <link>https://dev.to/lexosi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lexosi"/>
    <language>en</language>
    <item>
      <title>Agents Get Replaced. The Doctrine Is the Product.</title>
      <dc:creator>Lex</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:10:20 +0000</pubDate>
      <link>https://dev.to/lexosi/agents-get-replaced-the-doctrine-is-the-product-5f27</link>
      <guid>https://dev.to/lexosi/agents-get-replaced-the-doctrine-is-the-product-5f27</guid>
      <description>&lt;p&gt;Earlier today a session in my own system tried to write a file outside its&lt;br&gt;
territory. A hook denied it before the write happened, and the session&lt;br&gt;
couldn't override itself: the bypass is an environment flag the hook process&lt;br&gt;
reads, not something the agent can set. It had to hand the work back.&lt;/p&gt;

&lt;p&gt;That hook is less than a week old. The rule behind it is much older than the&lt;br&gt;
system it's running in.&lt;/p&gt;

&lt;h2&gt;
  
  
  From multiagent-system to agentic-os
&lt;/h2&gt;

&lt;p&gt;I moved from multiagent-system — a 19-agent orchestration setup I'd been&lt;br&gt;
running for months — to agentic-os, a rebuild, for three reasons: so the&lt;br&gt;
system had memory and a record of everything, so I could work faster, and to&lt;br&gt;
get closer to loops between agents. For now I still review every loop&lt;br&gt;
myself. A human is in the loop.&lt;/p&gt;

&lt;p&gt;After migrating I reviewed the initial agents, modified them, and reused a&lt;br&gt;
lot of doctrine from the old system. The reason is simple: the AI was making&lt;br&gt;
mistakes again that the old design had already had to correct.&lt;/p&gt;

&lt;p&gt;So I audited it. On 2026-09-27 I took 100 rules from multiagent-system and&lt;br&gt;
checked each one against agentic-os. 29 were already there, 35 were partial,&lt;br&gt;
36 were missing. The 36 defined the work, and they were mostly enforcement:&lt;br&gt;
rules the old system had turned into hooks, and this one still carried as&lt;br&gt;
text. The roster was replaced; the doctrine was ported by subtraction,&lt;br&gt;
keeping what corrects a known error and dropping the rest.&lt;/p&gt;

&lt;p&gt;A rule is the record of an error the system already paid for once. Agents&lt;br&gt;
change with each model generation. The list of mistakes a model tends to&lt;br&gt;
repeat changes much more slowly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What transferred
&lt;/h2&gt;

&lt;p&gt;Four rows of that inventory:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Row&lt;/th&gt;
&lt;th&gt;Rule&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A01&lt;/td&gt;
&lt;td&gt;Hard-block &amp;gt; advisory (advisory = 0% enforcement)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B01&lt;/td&gt;
&lt;td&gt;The orchestrator is the only invoker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C01&lt;/td&gt;
&lt;td&gt;The auditor is independent, anti-self-grading&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D24&lt;/td&gt;
&lt;td&gt;Deterministic signal before an LLM judge&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A01 and B01 are marked present in agentic-os: three fail-closed hooks, and&lt;br&gt;
the invoke tool blocked in the frontmatter of every subagent. A01 comes from&lt;br&gt;
&lt;a href="https://dev.to/lexosi/what-advisory-rules-actually-do-in-an-agent-loop-bke"&gt;What Advisory Rules Actually Do in an Agent&lt;br&gt;
Loop&lt;/a&gt;.&lt;br&gt;
The measured result in the earlier system was "all 9 blocked tool-calls&lt;br&gt;
across 7 runs, 4 via explicit logged override, 0 unauthorized writes". Those&lt;br&gt;
numbers belong to multiagent-system. What carried over is the rule they&lt;br&gt;
justify: a prompt that politely asks is not a boundary.&lt;/p&gt;

&lt;p&gt;The other two are the same instinct. The author never grades its own work.&lt;br&gt;
If a script can answer the question, the script answers before a model does.&lt;/p&gt;

&lt;p&gt;None of this came from an AI course. It came from &lt;a href="https://dev.to/lexosi/ten-years-directing-live-products-before-i-knew-it-was-called-product-management-225f"&gt;running live&lt;br&gt;
products&lt;/a&gt;,&lt;br&gt;
where one person owns the schedule, builders don't sign off their own&lt;br&gt;
builds, and a metric beats an opinion. Only the workers changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules arrived with their tests
&lt;/h2&gt;

&lt;p&gt;Each of the three guards was built the same way, and the order is in the&lt;br&gt;
commit log. The test went in first, failing, with the hook absent — 3 of 12&lt;br&gt;
red for territory, 2 of 13 for protected paths, 2 of 7 for the stop gate.&lt;br&gt;
Then the hook, and the same suites green: 12 of 12, 13 of 13, 7 of 7.&lt;/p&gt;

&lt;p&gt;Committing a red test is the point. A gate that has never been seen denying&lt;br&gt;
anything is a gate nobody has tested — it is indistinguishable, from the&lt;br&gt;
outside, from a gate that always says yes.&lt;/p&gt;

&lt;p&gt;The deny path is the one that has to be boring. Empty input, malformed&lt;br&gt;
input, an agent not in the config, an unreadable config file, any unhandled&lt;br&gt;
exception: all of them deny. The failure mode of a verification layer should&lt;br&gt;
be refusal, not silence.&lt;/p&gt;

&lt;p&gt;The session that got denied this morning was mine, doing research for this&lt;br&gt;
article. The rule that stopped it was ported from a system that no longer&lt;br&gt;
runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Three of my five hard rules are still only prose.&lt;/strong&gt; One is enforced by a
hook, one partially, and three — default billing mode, nothing gets
deleted from the vault, one script per file — exist as text and nothing
else. The port was real and it was incomplete.&lt;/li&gt;
&lt;li&gt;One operator, a handful of runs, one audit of 100 rules, and nobody
outside has reviewed any of it.&lt;/li&gt;
&lt;li&gt;A human is still in the loop. Loops between agents without my review are
the goal, not the current state.&lt;/li&gt;
&lt;li&gt;The inventory measures whether a rule exists, not its effect. A "yes"
means present, not proven.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents get replaced. The doctrine is the product.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>productivity</category>
    </item>
    <item>
      <title>A Whole Lane Was Losing on Geometry, Not on Merit</title>
      <dc:creator>Lex</dc:creator>
      <pubDate>Fri, 25 Sep 2026 15:43:38 +0000</pubDate>
      <link>https://dev.to/lexosi/a-whole-lane-was-losing-on-geometry-not-on-merit-da9</link>
      <guid>https://dev.to/lexosi/a-whole-lane-was-losing-on-geometry-not-on-merit-da9</guid>
      <description>&lt;p&gt;I had three categories of document coming in, and a semantic classifier that&lt;br&gt;
scored every one against all three. It worked. Except one of the three never&lt;br&gt;
showed up near the top — not rarely, never.&lt;/p&gt;

&lt;p&gt;The obvious explanations were that there were fewer of them, or that they&lt;br&gt;
were simply worse matches. Both were wrong.&lt;/p&gt;

&lt;p&gt;They matched just as well. They were losing on geometry.&lt;/p&gt;
&lt;h2&gt;
  
  
  How a whole lane can lose without being worse
&lt;/h2&gt;

&lt;p&gt;The setup is the ordinary one. Each category is described in a paragraph;&lt;br&gt;
that paragraph becomes a vector. Each document becomes another vector.&lt;br&gt;
Cosine against all three, keep the highest: that's its lane, and that number&lt;br&gt;
is its fit.&lt;/p&gt;

&lt;p&gt;The bug is in the last step. Cosines get compared &lt;em&gt;across&lt;/em&gt; lanes as if they&lt;br&gt;
were the same unit. They are not.&lt;/p&gt;

&lt;p&gt;A lane paragraph written in common vocabulary, heavily covered by the&lt;br&gt;
model's training data, lands in a dense region of the space. Everything near&lt;br&gt;
it scores high — 0.55, 0.60. A lane paragraph written in industry-specific&lt;br&gt;
vocabulary lands somewhere sparser, where the same degree of real fit&lt;br&gt;
produces 0.30 or 0.35. That gap isn't measuring fit. It's measuring how much&lt;br&gt;
text the model has seen that talks like this.&lt;/p&gt;

&lt;p&gt;So when you sort by raw cosine, the sparse lane loses every time. Not&lt;br&gt;
usually — every time. Its best possible document scores below a dense lane's&lt;br&gt;
mediocre one.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fix is five lines
&lt;/h2&gt;

&lt;p&gt;Don't compare across lanes. Normalise within each one.&lt;/p&gt;

&lt;p&gt;Route every document to its lane by max cosine, same as before. Then, per&lt;br&gt;
lane, take the mean and standard deviation of those cosines, and turn each&lt;br&gt;
raw cosine into a z-score: how many deviations above its own lane's average&lt;br&gt;
this document sits.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;by_lane&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ln&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ln&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;lane_names&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;recs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;by_lane&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lane&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]].&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="n"&gt;lane_stats&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ln&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;by_lane&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;arr&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;lane_stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;ln&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
                      &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;arr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;std&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1e-9&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                      &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;recs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;lane_stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lane&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
    &lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mu&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A z of 2.0 means the same thing in all three lanes: strong, for whatever&lt;br&gt;
this is. The lane's absolute scale drops out, and with it the advantage it&lt;br&gt;
never earned.&lt;/p&gt;

&lt;p&gt;The epsilon on the denominator isn't superstition. A lane with one document&lt;br&gt;
in it has zero deviation, and without it the z comes back infinite. That&lt;br&gt;
lane existed.&lt;/p&gt;
&lt;h2&gt;
  
  
  Making it count
&lt;/h2&gt;

&lt;p&gt;The z-score isn't the ranking. It's one signal feeding a larger score, and&lt;br&gt;
how it feeds in is its own decision. Symbol names changed for privacy,&lt;br&gt;
structure as it runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rules&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;semantic&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;                    &lt;span class="c1"&gt;// { w_sem: 0.4, z_cap: 2.0 }&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;map&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;rules&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;_semantic&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;rules&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;_semantic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;map&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cfg&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;w_sem&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;map&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rec&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isFinite&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;zCap&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;z_cap&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;z_cap&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;2.0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;bonus&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;cfg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;w_sem&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;z&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;zCap&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;bonus&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;bonus&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;trace&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`sem:&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;rec&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;lane&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things there are worth the space.&lt;/p&gt;

&lt;p&gt;The clamp starts at zero, so a negative z never subtracts: the semantic&lt;br&gt;
signal can only lift a document, never sink one. A weak semantic match can&lt;br&gt;
mean "bad fit", but it can also mean "this text is written strangely", and I&lt;br&gt;
didn't want the second one burying anything.&lt;/p&gt;

&lt;p&gt;The upper cap is duller and more important. A lane with few documents has a&lt;br&gt;
tiny standard deviation, and a tiny deviation produces enormous z-scores.&lt;br&gt;
Without the cap, the emptiest lane wins the entire list — which is the&lt;br&gt;
original bug again, running backwards.&lt;/p&gt;

&lt;p&gt;And the whole guard fails open. No cache, no config, or an item with no&lt;br&gt;
embedding, and the branch is skipped and the score comes out identical to&lt;br&gt;
what it was before any of this existed. I was adding a signal to something I&lt;br&gt;
already used daily. The one thing that couldn't happen was a missing file&lt;br&gt;
quietly changing an ordering I'd already come to trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  The general shape
&lt;/h2&gt;

&lt;p&gt;None of this is about semantic search. It's arithmetic that applies the&lt;br&gt;
moment a score ranks across groups that don't share a scale.&lt;/p&gt;

&lt;p&gt;Grades from different teachers. Latency percentiles from services with&lt;br&gt;
different load profiles. Product reviews in categories where people rate&lt;br&gt;
with different harshness. Anything classified by a model that has seen far&lt;br&gt;
more of one kind than another. The number looks comparable because it has&lt;br&gt;
the same name and the same range, and it isn't.&lt;/p&gt;

&lt;p&gt;The tell is specific: a whole category that never reaches the top. Poor&lt;br&gt;
performance mixes in. Total absence is structural. When an entire group is&lt;br&gt;
missing from your best-of list, the first thing to suspect isn't the data —&lt;br&gt;
it's the comparison.&lt;/p&gt;

&lt;p&gt;The fix is always the same. Normalise within the group, then compare across&lt;br&gt;
groups. Cap the output, because small groups produce extreme values. And&lt;br&gt;
decide deliberately whether the signal is allowed to subtract, or only to&lt;br&gt;
add.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;The mechanism is correct by construction: normalising within a group before&lt;br&gt;
comparing across groups is arithmetic, not a bet. What I don't have is a&lt;br&gt;
measured before-and-after. I didn't freeze a pre-change ranking to compare&lt;br&gt;
against, so what I can claim is that the sparse lane stopped being invisible&lt;br&gt;
— not how much the ordering improved in quality.&lt;/p&gt;

&lt;p&gt;Recording the baseline is what I'd do differently. The change took an&lt;br&gt;
afternoon; the measurement that would have proved it needed starting before&lt;br&gt;
I touched anything.&lt;/p&gt;

&lt;p&gt;One corpus, one embedding model, three lanes I defined by hand. The lanes&lt;br&gt;
are prose paragraphs, so rewriting one moves its vector and shifts its mean&lt;br&gt;
— the z-score is robust to lane size, not to me editing the description. I&lt;br&gt;
haven't measured how much.&lt;/p&gt;

&lt;p&gt;The slowest part wasn't the fix. It was seeing that the sparse lane wasn't&lt;br&gt;
failing.&lt;/p&gt;

&lt;p&gt;I'd spent weeks reading that list and assuming a third of the corpus simply&lt;br&gt;
didn't match anything well. The data had been saying otherwise the whole&lt;br&gt;
time. I was sorting by a number that didn't mean the same thing in every&lt;br&gt;
row.&lt;/p&gt;

&lt;p&gt;It's no accident that it took so long. A whole category missing reads like a&lt;br&gt;
verdict, not a bug. And as long as it reads like a verdict, you don't go&lt;br&gt;
looking.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>llm</category>
      <category>datascience</category>
    </item>
    <item>
      <title>git blame Told Me I Wrote 767 Lines I Didn't Write</title>
      <dc:creator>Lex</dc:creator>
      <pubDate>Fri, 18 Sep 2026 15:33:18 +0000</pubDate>
      <link>https://dev.to/lexosi/git-blame-told-me-i-wrote-767-lines-i-didnt-write-1pp6</link>
      <guid>https://dev.to/lexosi/git-blame-told-me-i-wrote-767-lines-i-didnt-write-1pp6</guid>
      <description>&lt;h2&gt;
  
  
  git blame Told Me I Wrote 767 Lines I Didn't Write
&lt;/h2&gt;

&lt;p&gt;Before writing about a gate, I checked who wrote it.&lt;/p&gt;

&lt;p&gt;The gate is 767 lines of JavaScript that refuses to ship a generated document when the model asserts a number or a tool that isn't backed by the source files. It sits in a document-generation pipeline I run locally — the same source material, retargeted per audience, which is exactly the setup where a model starts filling gaps to fit. I wired the gate into that pipeline four days ago, watched it abort real output the same afternoon, and decided it was worth a write-up.&lt;/p&gt;

&lt;p&gt;So I ran &lt;code&gt;git blame&lt;/code&gt;. Seven hundred and sixty-seven lines out of seven hundred and sixty-seven: me. Two commits in the log, both mine. Clean history, single author, no ambiguity.&lt;/p&gt;

&lt;p&gt;I didn't write a single line of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the blame lied
&lt;/h2&gt;

&lt;p&gt;The repository is a fork. &lt;code&gt;origin&lt;/code&gt; points at an upstream project — MIT licensed, someone else's — that ships an updater for its own system layer. I run a command; it fetches the canonical files and writes them to disk.&lt;/p&gt;

&lt;p&gt;That's how this file arrived. When I committed it, git recorded the only thing it can record: who moved the bytes. The commit message says as much — it's an automated system-files update, and nothing else — and the blame still puts my name on all 767 lines.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;git blame&lt;/code&gt; does not answer "who wrote this". It answers "which commit touched this line last, and who signed that commit". Those are usually the same answer. In a fork with an auto-updater they stop being the same answer, and the blame has no way to tell you that.&lt;/p&gt;

&lt;p&gt;Which is &lt;a href="https://dev.to/lexosi/a-line-of-documentation-was-acting-as-a-global-config-flag-3635"&gt;a shape I've written about before&lt;/a&gt;. A prose note asserting that a setting was off. A blame asserting that I wrote a file. Both authoritative, both mechanical, both wrong — except the note was prose and the blame is a tool, and the tool is far more convincing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who did write it
&lt;/h2&gt;

&lt;p&gt;The authorship lives in the upstream history, not mine — named there, not here. Blame the file against &lt;code&gt;origin/main&lt;/code&gt; and three contributors come back:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Contributor&lt;/th&gt;
&lt;th&gt;Surviving lines&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;301&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;246&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;66&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A fourth touched the file early on; none of those lines survive in the current blame.&lt;/p&gt;

&lt;p&gt;The shape of that history matters more than the split. The module started as a validator for invented numbers, and was extended weeks later to cover asserted names as well as numbers — the things a document says exist, not just how many. Then came the corrections, and the corrections are the interesting part. One trigger was case-sensitive in a way nobody intended, so a whole class of claims walked past untouched. In several languages the checker silently found nothing at all. One internal heuristic had drifted into answering a different question than the one it was built for. And a pattern was dropping claims on the floor without reporting it. Four people, roughly a month, five separate ways for a fact-checker to be quietly wrong.&lt;/p&gt;

&lt;p&gt;I would not have built that. I'd have built the version that works against my own text and fails silently against someone else's — which is the same version, right up until it isn't. The quality in this module isn't in the original design; it's in the corrections. None of the corrections are mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is mine
&lt;/h2&gt;

&lt;p&gt;Here's the part that surprised me more than the blame did: upstream never calls the gate.&lt;/p&gt;

&lt;p&gt;The module ships as a library and a CLI, and nothing in the generation path invokes it. The engine existed; the barrier did not. My commit wires it into three call sites along that path — 197 lines, on my branch only, not an ancestor of upstream's main. Before it, the gate was something you could run. After it, the gate is something the output has to get past.&lt;/p&gt;

&lt;p&gt;The call site is nine lines. Symbol names changed for privacy, structure untouched:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;checkClaims&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;cwd&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;block&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`blocked: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;describe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No retry at this point, no warning mode, no flag to skip it. The verdict comes back and the document either gets written or it doesn't. Nine lines is the whole of my contribution to the enforcement itself — the other 188 are plumbing: finding the source files, threading the config path, deciding which of three generation steps each check belongs to.&lt;/p&gt;

&lt;p&gt;The second thing that's mine is a config file: the allow and deny lists the engine reads. Verified figures the generator is permitted to state, phrases I never want appearing in generated text. The engine is deliberately generic — it knows how to ask whether a claim is supported, and nothing about what it's checking. That file is where it learns.&lt;/p&gt;

&lt;p&gt;The honest name for that work is integration, not authorship, and the distinction is worth keeping because the failure modes differ. The authors get it wrong when the gate mis-detects. I get it wrong when the gate doesn't fire. A perfect gate nobody wired is a file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first thing it blocked was true
&lt;/h2&gt;

&lt;p&gt;The document was a technical write-up about my own agent system. The gate aborted it, and the reason it gave was a list of four asserted tools it could find no backing for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;planner-worker&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;orchestrator-executor&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;artifact-based handoffs&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;role enforcement&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;My first read was the obvious one: the model invented four tools. It does that. I've watched it do that. I went and checked anyway, because the blame had just taught me what assuming costs.&lt;/p&gt;

&lt;p&gt;It invented nothing. &lt;code&gt;git grep&lt;/code&gt; across my agent system returns zero literal matches for all four strings — and every one of the four concepts is real and documented there. Planner and implementer are roles in the roster. Root-driven orchestration is what the architecture became after the orchestrator agent was absorbed into root. Artifact-based handoffs are how deliverables cross between agents: plans, reports, checkpoint files on disk. And role enforcement is a deny-by-default hook — &lt;a href="https://dev.to/lexosi/what-advisory-rules-actually-do-in-an-agent-loop-bke"&gt;the one I wrote my first article about&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So the model described my architecture correctly, in vocabulary that doesn't appear in the source files it's allowed to draw from. It didn't invent facts. It invented synonyms. The gate cannot tell the difference, because from where the gate stands there is no difference: an assertion with no backing text behind it is an assertion with no backing text behind it.&lt;/p&gt;

&lt;p&gt;And the gate was right. A gate that accepts "true, but not in the sources" is a gate that negotiates, and a gate that negotiates isn't a gate — it's a suggestion with extra steps. The cost of blocking fabrications is blocking unreceipted truths. That's the trade, and I'd make it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three days later, it lied properly
&lt;/h2&gt;

&lt;p&gt;A different abort, different class. This time the gate reported an invented figure — a count with no counterpart anywhere in the sources — and added the part that matters: the model had produced the same drift once already, been corrected, and produced it again on the next pass.&lt;/p&gt;

&lt;p&gt;No synonym this time. The number came from nothing.&lt;/p&gt;

&lt;p&gt;Which sent me looking at how that kind of drift had been handled before the gate existed. I found it in the generation code: a correction written by hand into the prompt, naming one specific figure and the wrong value the model kept reaching for. One fact, pinned in prose, inside the instructions.&lt;/p&gt;

&lt;p&gt;It's the same move I wrote about in August — a value living in a prompt string instead of coming from a source — and it's the reason this gate needs to exist at all. That line only ever protected one number. Every other figure the model might drift on had nothing standing behind it until something started checking all of them.&lt;/p&gt;

&lt;p&gt;That's a pattern I've hit before. In &lt;a href="https://dev.to/lexosi/a-line-of-documentation-was-acting-as-a-global-config-flag-3635"&gt;an earlier piece&lt;/a&gt; I described a claim the model kept inverting; fixing the wording at the source reduced the problem and did not remove it, and it took a hard stop downstream to close. Same shape here. A correction is a suggestion the model weighs against the task. A gate is a fact about the world.&lt;/p&gt;

&lt;p&gt;Two aborts, three days apart, and the useful thing is that the gate treated them identically. It does not distinguish true from false. It distinguishes supported from unsupported. That makes it blunt in the first case and exactly right in the second — and since it cannot know which case it's in, treating them the same way is not a flaw in the design. It is the design.&lt;/p&gt;

&lt;p&gt;Both documents stayed on my machine. Neither had to be caught by a human reading carefully at the wrong hour, which is the review process the gate replaced.&lt;/p&gt;

&lt;h2&gt;
  
  
  Both stories are the same story
&lt;/h2&gt;

&lt;p&gt;A tool that cannot verify what it asserts will give you the answer with exactly as much confidence as one that can. The blame said I wrote 767 lines. The gate said four real tools weren't real. Neither of them lied. Both of them answered a question slightly different from the one I asked, and neither had any way to flag the gap.&lt;/p&gt;

&lt;p&gt;The difference between them is that the gate declares its method. It blocks for want of backing and says so: unsupported, not false. The blame hands you a name and a date and never tells you what it is accountable for. A tool that documents its own limit is usable. A tool that doesn't will mislead you without ever being wrong.&lt;/p&gt;

&lt;p&gt;Three pieces in, this keeps being the same lesson at different altitudes. Prose rules don't survive contact with a loop, so enforcement has to live one layer down. Prose that describes configuration &lt;em&gt;is&lt;/em&gt; configuration, just without a type checker. And now: tools that summarise the past are making claims about the world, and those claims go stale in exactly the way prose does. &lt;code&gt;git blame&lt;/code&gt; is a derived artifact with a fossilised assumption inside it — that whoever committed a line is whoever wrote it.&lt;/p&gt;

&lt;p&gt;What I actually do with this: before asserting anything a tool handed me, I ask what question that tool answers, not what question it appears to answer. That's the same move the gate makes on every document it sees. The difference is the gate does it every time, and I have to remember.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;The gate has been wired for four days, and everything above comes out of that window.&lt;/p&gt;

&lt;p&gt;Three aborts carry a verbatim message in the logs. Counting the narrative entries in my own daily notes it's around six, but those overlap with the logs and I can't cleanly separate them, so six is a bound and three is the count I'd defend. There is no aggregate counter anywhere. I tallied these by hand, which is its own small irony in a piece about trusting derived artifacts.&lt;/p&gt;

&lt;p&gt;The split: two aborts for an unsupported tool, three for an unsupported number. Zero for the blocked-phrase list, which is configured, loaded, and has never once fired. I mention it because it's the part of the gate that has demonstrated nothing, and a reader looking for the weak joint should be handed it rather than have to find it.&lt;/p&gt;

&lt;p&gt;One operator, one pipeline, one month of upstream history I didn't write. No control group, and no idea whether any of this survives contact with a setup that isn't mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;The config file — the one that teaches a generic engine which figures are verified — is gitignored. No history. No blame. Not even the wrong blame. If a figure in that allow-list turns out to be stale, nothing on disk can tell me when it went in or what it was checked against.&lt;/p&gt;

&lt;p&gt;Which is the thing I wrote about in August: configuration living somewhere no tool would ever look. I published that piece, and roughly a month later populated this file outside version control without noticing. Knowing the failure mode is not immunity from it, and I'd like that to be less funny than it is.&lt;/p&gt;

&lt;p&gt;So: version the file, local branch if nothing else, and write down where each allowed figure comes from. Right now an entry asserts that a number is fine and doesn't say why — a frozen value with no provenance, which is precisely the shape that bit me last time.&lt;/p&gt;

&lt;p&gt;The open problem is bigger than the file. The gate distinguishes supported from unsupported, which ties it to whatever is already written in the sources. A genuine tool that isn't documented there is indistinguishable, from where the gate stands, from one the model made up. The answer isn't to loosen the gate — it's that the sources are where that gets fixed, which means someone has to keep them current. I haven't solved who, or how often, and four days of data doesn't tell me.&lt;/p&gt;

&lt;p&gt;I started by checking who wrote the gate because I was about to assert it in public and I wanted the assertion to be true.&lt;/p&gt;

&lt;p&gt;The check took twenty minutes and cost me the article I meant to write.&lt;/p&gt;

&lt;p&gt;The gate demands a source before it lets a claim through. I ran the same check on my own. That's the whole of it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>git</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Ten years directing live products before I knew it was called product management</title>
      <dc:creator>Lex</dc:creator>
      <pubDate>Fri, 21 Aug 2026 18:29:50 +0000</pubDate>
      <link>https://dev.to/lexosi/ten-years-directing-live-products-before-i-knew-it-was-called-product-management-225f</link>
      <guid>https://dev.to/lexosi/ten-years-directing-live-products-before-i-knew-it-was-called-product-management-225f</guid>
      <description>&lt;p&gt;In 2016 I didn't know what a server was.&lt;/p&gt;

&lt;p&gt;I was a kid uploading ARK: Survival Evolved videos to YouTube, and the game — as shipped — was unrecordable. Taming took days. Walking took hours. So I started digging into config files: player speed, taming multipliers, breeding timers. Trial and error, on my own machine, until it was recordable. I didn't know it yet, but that was my first product decision: the game had to fit the format, so I changed the game.&lt;/p&gt;

&lt;p&gt;The channel grew. I got invited to multiplayer series with bigger creators. And one day, mid-series, the guy running the server had a falling-out with the owner and walked. Nobody else knew how to keep it alive. I said: I'll run it. I had never run a real server. Three days later the series ran on it. I've said that sentence for every job I've taken since.&lt;/p&gt;

&lt;h2&gt;
  
  
  The conversation that changed how I quote work
&lt;/h2&gt;

&lt;p&gt;Word spread that I could handle ARK servers. Then Willyrex, a top Spanish-speaking creator, joined a series I was running and asked if I could set up Conan Exiles and GTA V servers too. I had no idea how. I said yes, learned, and delivered.&lt;/p&gt;

&lt;p&gt;But that's not why he kept coming back. He told me why: every technical person he'd worked with said yes to everything and then blew the deadline. I did the opposite. When he asked for something enormous, I'd say: "What you actually want for the video is THIS, right? Building it exactly as you describe is madness — weeks of work. But framed this way, you get almost the same result on screen, and you'll have it tomorrow."&lt;/p&gt;

&lt;p&gt;He and Vegetta777 loved it. I didn't have a word for it then. Now I do.&lt;/p&gt;

&lt;p&gt;That was the start of ten years of commissions: an ARK series with Pokémon-style gyms and creatures (with Staxx in it), servers for Rubius, xFarganx, and others. They came to me with a game and a theme; I owned everything else — research, team, budget, game systems, and enough fresh content that each creator had material every single episode. They kept commissioning me across different games, because what traveled was the execution, not the game.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shipping on a platform nobody had documented yet
&lt;/h2&gt;

&lt;p&gt;When Epic launched Fortnite Creative, Vegetta had a sponsored campaign with them — he and Fargan needed maps to showcase the new mode. The mode was less than a month old. No Verse, no documentation, no best practices, nothing to copy. I called three friends — Arsilex, Jesusseron, Hernybreak — and we built our first two maps under constraints nobody had mapped yet. The maps shipped and the campaign ran.&lt;/p&gt;

&lt;p&gt;Some time later, Vegetta asked me how I'd approach a series around creator-made maps. I told him two things were missing: more people on screen, for dynamics — and maps built expressly for the group, not generic ones. I took care of both. By mid-2019 that group of friends had a name the audience gave us — Los Noobs — and a series on Vegetta's channel, &lt;a href="https://www.youtube.com/playlist?list=PLSbDMtNBmYTv41a7zb_V9mkrt6uRPJ7hL" rel="noopener noreferrer"&gt;Minecraft con Noobs&lt;/a&gt;, alongside Fargan, Elyas360 and Alexby. Vanilla Minecraft, me on camera as a participant, at a time when nobody was making Spanish-speaking Minecraft series.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running my own platform, on my own payroll
&lt;/h2&gt;

&lt;p&gt;Then I built Tomatecraft, my own Minecraft network. The product decision that made it work: Java (PC) and mobile players could finally play together, when almost no server offered crossplay. I funded it myself, out of a channel that by then had passed half a million subscribers, with programmers, designers and builders on payroll. Later I merged networks with Arsilex and we ran large branded events — up to 20 contributors per event, with Vegetta, Willyrex, Fargan and Rubius joining in.&lt;/p&gt;

&lt;p&gt;Every feature I cut, I cut knowing what the payroll cost that month.&lt;/p&gt;

&lt;p&gt;Then I stopped for a long while — health comes first. When I came back, I formalized what I'd been doing by instinct and earned a Java certification.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comeback: 65,000 concurrent players
&lt;/h2&gt;

&lt;p&gt;Years earlier, when Epic ran the Fortnite World Cup and Rubius needed an official map for his qualifier — &lt;a href="https://www.esportsearnings.com/tournaments/35297-fortnite-world-cup-2019-creative-rubiustrials" rel="noopener noreferrer"&gt;RubiusTrials&lt;/a&gt; — I had pushed hard for two builders I believed in: Hooshen and Iscariote. I didn't build a single brick of it; I just insisted, to Vegetta and Rubius, that they were the right ones. They were.&lt;/p&gt;

&lt;p&gt;So when those same two needed someone to lead projects and bring in clients, they called me. I directed the Fall Guys collaboration island (with Vegetta777 and Willyrex): 65,000 concurrent players at peak, the #1 most-played island globally across all of Fortnite for several hours on day 2, and #1 in its category for two weeks — ahead of Epic's own official maps for the event.&lt;/p&gt;

&lt;p&gt;Afterwards, with Los Noobs, we shipped island after island — I directed five-plus of them end to end, with a core team of five, and contributed to twenty-plus more in specialist roles: gameplay programming, configuration, 3D modeling, mechanics design, and pricing. We monetized through Epic's engagement payouts, and when in-island micro-purchases arrived, pricing became my most common role across those twenty-plus islands. My last project with the studio was a battle-pass system: VIP pass, level purchases, score multipliers — built as a reusable framework, adopted across 4 maps by the time I left. What I learned pricing them: the hardest purchase is the first one. After that, buying again feels natural — so I priced the first one to be bought.&lt;/p&gt;

&lt;h2&gt;
  
  
  Today: AI agent systems
&lt;/h2&gt;

&lt;p&gt;Today I build AI agent systems. My 19-agent orchestration went from advisory rules — followed 0 times out of 3 — to deny-by-default enforcement: across 7 runs, 9 blocked tool-calls, 4 allowed only via explicit logged override, 0 unauthorized writes — &lt;a href="https://dev.to/lexosi/what-advisory-rules-actually-do-in-an-agent-loop-bke"&gt;published with reproducible numbers&lt;/a&gt;, extracted into an open-source package (&lt;a href="https://github.com/lexosi/loopward" rel="noopener noreferrer"&gt;loopward&lt;/a&gt;). Different stack. Same rule: the deadline is part of the product.&lt;/p&gt;

&lt;p&gt;I never held the PM title. I have the proof.&lt;/p&gt;

</description>
      <category>gamedev</category>
      <category>career</category>
      <category>ai</category>
      <category>product</category>
    </item>
    <item>
      <title>A Line of Documentation Was Acting as a Global Config Flag</title>
      <dc:creator>Lex</dc:creator>
      <pubDate>Fri, 14 Aug 2026 15:16:50 +0000</pubDate>
      <link>https://dev.to/lexosi/a-line-of-documentation-was-acting-as-a-global-config-flag-3635</link>
      <guid>https://dev.to/lexosi/a-line-of-documentation-was-acting-as-a-global-config-flag-3635</guid>
      <description>&lt;p&gt;I spent a morning hunting for a setting that did not exist.&lt;/p&gt;

&lt;p&gt;A while back I turned off Claude's co-authorship trailer in my commits — a deliberate choice at the time. Last week I decided I wanted it back. So I went looking for the switch I'd flipped. &lt;code&gt;~/.claude/settings.json&lt;/code&gt;: no key. &lt;code&gt;settings.local.json&lt;/code&gt;: no key. &lt;code&gt;~/.claude.json&lt;/code&gt;, parsed as JSON, top-level plus all &lt;strong&gt;38 project entries&lt;/strong&gt;: no key. The &lt;strong&gt;27&lt;/strong&gt; &lt;code&gt;.claude/settings*.json&lt;/code&gt; files scattered across my two working drives: no key. Every &lt;code&gt;CLAUDE.md&lt;/code&gt; and &lt;code&gt;AGENTS.md&lt;/code&gt; I own: no mentions. Environment variables: nothing. A final sweep of my entire user directory — every &lt;code&gt;*.json&lt;/code&gt; and &lt;code&gt;*.md&lt;/code&gt; — returned three raw hits: a changelog and two copies of an editor extension's JSON schema. The note I wrote when that finished was two words: "Cero hits reales." Zero real hits.&lt;/p&gt;

&lt;p&gt;There was exactly one thing anywhere on disk that turned attribution off, and it was a sentence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~/.claude/rules/ecc/common/git-workflow.md:12
  Note: Attribution disabled globally via ~/.claude/settings.json.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It had a Chinese twin, same file path with &lt;code&gt;zh/&lt;/code&gt; instead of &lt;code&gt;common/&lt;/code&gt;, same line 12. Both files are rules files. Rules files get loaded into every session. So every session opened with a line of documentation asserting, flatly and falsely, that a global setting was off — and the model behaved accordingly. The switch I remembered flipping never existed as a switch. The prose was the switch.&lt;/p&gt;

&lt;p&gt;The key that sentence gestured at, &lt;code&gt;includeCoAuthoredBy&lt;/code&gt;, is deprecated and replaced by &lt;code&gt;attribution&lt;/code&gt;. Neither is present in any of my configs, which means the default was active the whole time. The feature was on. Only the description of the world said otherwise, and the description won.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thesis
&lt;/h2&gt;

&lt;p&gt;Here's what I take from that, and from two more receipts this week: in an agent system, prose and constants are not documentation about the control plane. They &lt;em&gt;are&lt;/em&gt; the control plane.&lt;/p&gt;

&lt;p&gt;I run a personal multi-agent system on top of Claude Code — 19 specialized agents, root-driven, single-writer, coordinating through on-disk artifacts; a sanitized snapshot is public at &lt;a href="https://github.com/lexosi/multiagent-system-lex" rel="noopener noreferrer"&gt;multiagent-system-lex&lt;/a&gt;. Around it sits a set of always-loaded rules files, and a separate document-generation script I run locally that feeds an LLM a system prompt full of hand-written constraints. Both of those are, in the strict sense, configuration. Neither is stored anywhere a linter, a schema, or a test would ever look at it.&lt;/p&gt;

&lt;p&gt;That's the mechanism. A rules file is loaded unconditionally into every session, so a claim inside it has the reach of a global setting with none of the accountability: no schema, no type, no default, no &lt;code&gt;git blame&lt;/code&gt; if the directory isn't a repo, nothing downstream that notices it disagrees with reality. The model doesn't verify the claim against the system, because from inside the loop the claim &lt;em&gt;is&lt;/em&gt; the system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The constant that outlived its own source
&lt;/h2&gt;

&lt;p&gt;Second receipt, same week, different shape. In that document-generation script, the system prompt carried a block of figures marked FROZEN — meaning: reproduce exactly, never paraphrase, never invent. One of them was a subscriber count, hardcoded as a literal in the prompt string. It was wrong: &lt;code&gt;617K&lt;/code&gt;, when the real figure is &lt;code&gt;577K&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The interesting part is the ordering:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;Commit&lt;/th&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-10 &lt;strong&gt;14:04:31&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;fc580b2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;removes the frozen &lt;code&gt;617K&lt;/code&gt; from the system prompt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-10 &lt;strong&gt;21:28:35&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;69e0499&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;corrects one remaining profile document, &lt;code&gt;617.000&lt;/code&gt; → &lt;code&gt;577.000&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-14 &lt;strong&gt;11:02:19&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;102f63e&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;a generated artifact is &lt;em&gt;still&lt;/em&gt; emitting &lt;code&gt;617K&lt;/code&gt;, in three separate places&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The structured sources were already right. The commit message for &lt;code&gt;69e0499&lt;/code&gt; says so in as many words: "la cifra canonica es ~577.000" — the canonical figure is ~577,000 — and names the two source files that already held it. The only thing carrying &lt;code&gt;617K&lt;/code&gt; was a literal sitting inside a prose instruction, wearing the word FROZEN.&lt;/p&gt;

&lt;p&gt;And FROZEN is exactly why it won. That annotation exists to stop the model from softening or improvising a number. It did its job perfectly — on the wrong value. A freeze marker is authority, and I had handed authority to a copy instead of a source. The fix was one line in the doc comment and one in the prompt rule: the count is no longer hardcoded there at all, it comes from the profile facts or it doesn't appear. Four days later I found the stale value still sitting in a generated artifact downstream, because generated output doesn't retroactively fix itself when you fix its generator.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mirror: a prose rule that held
&lt;/h2&gt;

&lt;p&gt;I want to be fair to prose, because one of my prose rules worked, and &lt;em&gt;why&lt;/em&gt; it worked is the whole point.&lt;/p&gt;

&lt;p&gt;When I wrote &lt;a href="https://dev.to/lexosi/what-advisory-rules-actually-do-in-an-agent-loop-bke"&gt;the first article&lt;/a&gt;, the brief asked for a specific external citation — a multi-agent failure taxonomy, with two percentages attached. It also carried a standing rule: use only data present in a specific on-disk notes file, with its source; anything unsourced doesn't go in. The figures weren't in that file. The plan recorded the refusal explicitly — "NO está en el fichero → NO se usa", not in the file, so it isn't used — and moved the citation to a forbidden list. The published piece contains zero references to it.&lt;/p&gt;

&lt;p&gt;That rule held where the attribution note failed, and the difference isn't discipline. It's that "is this string present in that file?" is a question with a mechanical answer. The rule wasn't asking anyone to be careful; it was asking a question I could go check, and the check was cheap enough to actually run. Prose that resolves to a disk lookup behaves like code. Prose that asserts a fact about the world behaves like a lie waiting to be believed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing the prose is not enough
&lt;/h2&gt;

&lt;p&gt;Fourth receipt, and it's the one that connects back.&lt;/p&gt;

&lt;p&gt;In that same document-generation script, the model kept inverting a claim from the first article — writing that "nine unauthorized writes were blocked", when what was blocked were nine &lt;em&gt;tool-calls&lt;/em&gt; and unauthorized writes were zero. Same numbers, reversed meaning, and the reversed version reads better, which is precisely the problem. I fixed the wording at the source. The commit records what happened next, with one noun swapped for anonymity: "Fixing the source cut this from 4 generated documents to 1; the model still paraphrased its way back, so it needed a hard stop." The hard stop was a denylist on the inverted phrasing, with tests in both directions so the correct claim stays legal.&lt;/p&gt;

&lt;p&gt;Four to one, not four to zero. A sibling change the same day says it from another angle: a gold example meant to fix the output's register was already in the prompt and didn't prevent the failure — "the lever is explicit rules plus a mechanical constraint, not more example."&lt;/p&gt;

&lt;p&gt;The first article argued that if a boundary matters, it has to live in the substrate the agent can't reason its way around — the tool-call layer, not prose the model weighs against its objective. That enforcement layer is packaged as &lt;a href="https://github.com/lexosi/loopward" rel="noopener noreferrer"&gt;loopward&lt;/a&gt;. This piece is the other half: the prose &lt;em&gt;is&lt;/em&gt; substrate. Just substrate with no type checker, no default, and no test. It needs the same treatment — single-sourced values, mechanical checks, provenance — and when it can't get it, a hard stop downstream.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;Four incidents, one operator, one system, one week. That's an anecdote with receipts, not evidence of a law. Everything above is reconstructed from commits and session transcripts I can point at, and quoted rather than paraphrased, but there's no control group here and no idea how any of it generalizes to a system that isn't mine.&lt;/p&gt;

&lt;p&gt;The sharpest limit is in the first story. I can't tell you when that note was introduced, because &lt;code&gt;~/.claude&lt;/code&gt; is not a git repository — &lt;code&gt;git rev-parse&lt;/code&gt; there returns "fatal: not a git repository", and the trail ends there. A line of text stood in for a config flag and I can't date it, so the audit stops at "undated". I'd rather say that plainly than round it to a date.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;Put &lt;code&gt;~/.claude&lt;/code&gt; under version control, first thing — the rules directory is configuration with production effects and it deserves a history. Single-source every constant that appears in a prompt; if a figure lives in a YAML file, the prompt gets a reference, never a copy, and FROZEN marks the &lt;em&gt;rule&lt;/em&gt;, never the value. Write prose rules that resolve to a disk lookup wherever the choice exists, because those are the ones that survive contact with a loop. And treat a fix at the source as step one of two: generated artifacts don't self-heal, so something has to go find the stale copies.&lt;/p&gt;

&lt;p&gt;The uncomfortable version: I've been writing configuration in English and grading it as documentation. It was never documentation. It ran.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>productivity</category>
    </item>
    <item>
      <title>What Advisory Rules Actually Do in an Agent Loop</title>
      <dc:creator>Lex</dc:creator>
      <pubDate>Thu, 06 Aug 2026 10:05:43 +0000</pubDate>
      <link>https://dev.to/lexosi/what-advisory-rules-actually-do-in-an-agent-loop-bke</link>
      <guid>https://dev.to/lexosi/what-advisory-rules-actually-do-in-an-agent-loop-bke</guid>
      <description>&lt;p&gt;&lt;em&gt;Advisory rules in my 19-agent system produced 0/3 compliance in the three runs right after the rule was canonized — the hook warned, the agent proceeded anyway. I replaced the warning with a deny-by-default gate plus a one-shot, logged override: over two months, 9 blocked tool-calls, 4 authorized overrides, 0 unauthorized writes. Small, well-delimited evidence, reproducible by script — but the direction is clear: enforcement has to live in the substrate, not in prose the model weighs against its objective.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hook fired. My root agent read the warning telling it that &lt;code&gt;plan.md&lt;/code&gt; belongs to the planner subagent, acknowledged it, and wrote the file anyway. Then it did it again on the very next step. The audit report for that run records it drily:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"plan.md root direct ×2 (step 1 + 2), planner NO invocado. hook PreToolUse advisory flagged 2x consecutivos; contenido aprobado verbatim inline → root procede."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Root authored the planning file directly, twice; the advisory hook flagged it twice in a row; root proceeded anyway, logging "aprobado verbatim inline" — approved verbatim inline — as its justification. Two runs later, a third instance with exactly the same shape: "plan.md root direct 3ª instancia consecutiva. Hook PreToolUse advisory disparó + root ack con razón 'aprobado verbatim inline'."&lt;/p&gt;

&lt;p&gt;That third one is what stopped me. I was reading the audit reports as they came in, expecting the new rule to have changed something — instead I was watching my own system politely log the same violation, run after run.&lt;/p&gt;

&lt;p&gt;Those were the three runs immediately after I canonized the rule. The hook fired every single time. Compliance: 0/3.&lt;/p&gt;

&lt;h2&gt;
  
  
  The system
&lt;/h2&gt;

&lt;p&gt;Some context for that scene. I run a personal multi-agent system on top of Claude Code: 19 specialized agents — planners, implementers, reviewers, domain specialists, independent auditors — under a root-driven, single-writer architecture. One main thread is the only physical invoker of subagents, and coordination happens through on-disk artifacts rather than shared mutable state: plans, hypothesis files, reports. The system has accumulated 171 documented runs, each with its own directory, plan, and audit trail. That number is scale and nothing more — raw transcripts from the early era are gone to retention, so I cannot measure rule compliance across all of it, and I won't pretend otherwise. The piece of the division of labor that matters for this story: the planner subagent owns &lt;code&gt;plan.md&lt;/code&gt;. Root coordinates, but it is not supposed to author the plan itself. A sanitized snapshot of the whole thing is public at &lt;a href="https://github.com/lexosi/multiagent-system-lex" rel="noopener noreferrer"&gt;github.com/lexosi/multiagent-system-lex&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why advisory rules fail
&lt;/h2&gt;

&lt;p&gt;Here's the mistake I made first, and I don't think it's an exotic one: I wrote the rule down and assumed that was enforcement. The advisory hook was technically real — a PreToolUse hook, introduced 2026-05-21, that fired on every root write into planner territory. But it always exited 0. It warned; it never blocked. Five days later, on 2026-05-26, I canonized the rule in the system's constitution: root must not author &lt;code&gt;plan.md&lt;/code&gt;; the planner owns it. The three runs I opened with are the three runs immediately after that canonization.&lt;/p&gt;

&lt;p&gt;A warning is just more context. The model isn't being malicious when it walks past one; it's optimizing for the task in front of it. When a soft rule competes with the immediate objective — produce the plan, keep the run moving, the content is already approved anyway — the rule loses. In my data, it lost every time. The failure mode isn't disobedience; it's incentives. The path of least resistance runs straight through the warning, and "approved verbatim inline" isn't a rebellion — it's a locally reasonable justification generated by a system that was never actually stopped. Anthropomorphizing this as the agent "ignoring the rules" misses the mechanism and, worse, suggests the wrong fix: a sterner warning is still a warning.&lt;/p&gt;

&lt;p&gt;Once I saw the pattern for what it was, I stopped tweaking the wording and went looking for where enforcement could actually live — somewhere the model couldn't weigh it against the task. The answer was one layer down.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cutover
&lt;/h2&gt;

&lt;p&gt;Six days after canonizing the rule, the warning was gone. On 2026-06-01 (commit &lt;code&gt;b3efe00&lt;/code&gt;) the advisory hook was deleted and replaced by a deny-by-default PreToolUse hook: any write into a guarded territory returns &lt;code&gt;permissionDecision=deny&lt;/code&gt; unless the caller is the owning subagent — the hook discriminates callers by &lt;code&gt;agent_type&lt;/code&gt; — or a valid one-shot sentinel file authorizes exactly one override, with a short TTL. The previous design's session-level environment-variable escape hatch died in the same commit: an override that lives for a whole session is an override that outlives its justification.&lt;/p&gt;

&lt;p&gt;The numbers, verbatim from the eval report — which is regenerated by script, not written by hand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Advisory compliance after the rule was canonized&lt;/strong&gt; (2026-05-26, &lt;code&gt;plan.md&lt;/code&gt; territory): &lt;strong&gt;0/3 runs&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard-block era&lt;/strong&gt; (2026-06-02..2026-08-04): &lt;strong&gt;9 blocked tool-calls (attempted root writes to planner-owned files), across 7 runs&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4 writes authorized&lt;/strong&gt; via the one-shot sentinel — each logged with a non-empty reason.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unauthorized writes to guarded territories: 0.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I want to label this evidence carefully, because the labeling is half the point. The advisory side is a &lt;strong&gt;natural experiment, n=3&lt;/strong&gt;: three consecutive runs immediately after a single dated rule change, one territory, with frozen citations into on-disk audit reports that the eval script re-verifies. The hard-block side is a &lt;strong&gt;mechanical log&lt;/strong&gt;: an append-only deny-ledger covering two months, parsed by script, no human or agent judgment in the count. And the honest limit: the hook logs denials and overrides, not allows — so "0 unauthorized writes" is an invariant by construction (the hook denies guarded territories unless the caller is the owner or consumes a valid sentinel), corroborated by the logs, not an exhaustive census of every write that ever happened. This is not a large-scale study and I'm not dressing it up as one. It is a small, well-delimited experiment that anyone can re-run against the repository with &lt;code&gt;python eval/run_eval.py&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deny, deny, override — as designed
&lt;/h2&gt;

&lt;p&gt;The pattern I care most about in the ledger is deny → deny → override. Root tries to write a guarded file: denied. Tries again: denied. Then one of two things happens. Either root routes the work to the owning subagent — which is the intended outcome, and the common one — or I look at the situation, decide the exception is legitimate, and drop a sentinel file that authorizes exactly one write, which the hook consumes and logs. Four of those in two months. All deliberate, all recorded, all with a stated reason.&lt;/p&gt;

&lt;p&gt;That flow is the design, not a failure of it. A hard block with no exit teaches the system — and its operator — to fight the wall, and fighting the wall is where the really creative workarounds come from. What I actually want is a gate: violations impossible by default, exceptions possible, authorized, and logged. The one-shot property is what keeps it honest. A sentinel authorizes one write, not a mood.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this fits
&lt;/h2&gt;

&lt;p&gt;The reason I think this matters beyond my setup: everyone is building loops right now. Boris Cherny, head of Claude Code at Anthropic, said it on stage in June 2026: "I don't prompt Claude anymore. […] My job is to write loops" (&lt;a href="https://www.youtube.com/watch?v=SlGRN8jh2RI" rel="noopener noreferrer"&gt;Acquired Unplugged&lt;/a&gt;) — outer programs that invoke the agent loop as a subroutine and decide what happens next. And the recent evidence says those loops need harder gates than most of us are giving them. &lt;a href="https://arxiv.org/abs/2603.24755" rel="noopener noreferrer"&gt;SlopCodeBench&lt;/a&gt; (UW-Madison/MIT, March 2026) had agents extend their own code under changing specs and found structural quality eroding iteration after iteration even while core tests kept passing — best strict rate 17.2%, agent code 2.2× more verbose than comparable human repositories. &lt;a href="https://arxiv.org/abs/2608.00267" rel="noopener noreferrer"&gt;LoopsBench&lt;/a&gt; (Microsoft, August 2026) found the best outer-loop configuration solved 25.00% of its tasks, with visible regression events across every loop profile it measured — though it's days old and unreplicated, so hold that one loosely.&lt;/p&gt;

&lt;p&gt;The shared reading of both: objective verifiers like tests are necessary but insufficient — what passes the gate can still rot. My result is a small datapoint in the same direction, one layer down. Even the &lt;em&gt;process&lt;/em&gt; rules — who is allowed to write what — don't hold as good-faith conventions inside a loop. If a boundary matters, it has to live in the substrate the agent can't reason its way around: the tool-call layer, before the write happens, not prose the model weighs against its objective and discounts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Touch it
&lt;/h2&gt;

&lt;p&gt;I extracted the enforcement layer into a standalone package, &lt;a href="https://github.com/lexosi/loopward" rel="noopener noreferrer"&gt;loopward&lt;/a&gt;: a reliability layer for multi-agent systems — anti-loop limits, human stop-gates, per-run audit trails — enforced structurally rather than by convention. &lt;code&gt;pip install git+https://github.com/lexosi/loopward.git&lt;/code&gt;, then &lt;code&gt;loopward-demo&lt;/code&gt; runs a deterministic offline demo with zero API keys; this isn't an announcement, it's the touchable version of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enforcement by construction, not good faith
&lt;/h2&gt;

&lt;p&gt;That's the one-line version of two months of logs. Rules the agent can weigh get weighed away; gates the agent cannot pass produce routing instead of violations, and the exceptions become visible, deliberate acts instead of silent drift.&lt;/p&gt;

&lt;p&gt;What I'd do differently, and what's still open — because a write-up that ends in triumph is usually hiding the interesting part. I'd log allow decisions from day one, so "0 unauthorized" could be a census instead of an invariant plus corroboration. I'd keep raw transcripts — retention ate the advisory era, and n=3 is what survived, not what I designed. I wouldn't canonize a rule five days before I could enforce it. And open questions remain: n=3 is small even if it's clean; &lt;code&gt;plan.md&lt;/code&gt; is one territory, and the same claim for the others is only partially measured; and I don't yet know how the gate scales as guarded territories multiply, because every new wall adds override friction and friction is a budget you spend from. What I do know is that the direction has stopped feeling arguable: if I need an agent to respect a boundary, I don't ask. I make the boundary the default, and the exception a logged, deliberate act.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>python</category>
    </item>
  </channel>
</rss>
