<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Open Human</title>
    <description>The latest articles on DEV Community by Open Human (@maref).</description>
    <link>https://dev.to/maref</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4034833%2F04c59718-fd78-4c60-bb00-60d4a8a526f9.jpg</url>
      <title>DEV Community: Open Human</title>
      <link>https://dev.to/maref</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/maref"/>
    <language>en</language>
    <item>
      <title>零信任的日常</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sun, 11 Oct 2026 16:12:28 +0000</pubDate>
      <link>https://dev.to/maref/ling-xin-ren-de-ri-chang-2boh</link>
      <guid>https://dev.to/maref/ling-xin-ren-de-ri-chang-2boh</guid>
      <description>&lt;p&gt;Zero Trust as a Daily Habit for MCP Agents&lt;/p&gt;

&lt;p&gt;The cron entry changed on a Thursday night. By Friday morning, all we had in the audit log was this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;executed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No actor. No original command. No before and after. We spent two hours the next morning trying to figure out whether a human, an agent, or some policy had moved the cleanup job. The log had recorded the conclusion and discarded the reasoning. That morning is why I now treat zero trust for MCP agents as something I practice daily.&lt;/p&gt;

&lt;p&gt;The first multi-agent incident will look small. It will look like a cached permission. Agent A writes a memory record. Agent B reads it later and acts. Nothing gets re-verified. That is ambient authority. Shared memory becomes shared privilege. A permission can outlive the task that created it, and nobody notices until something changes production.&lt;/p&gt;

&lt;p&gt;Zero trust for MCP only works as a loop you run on every call: authenticate the caller, authorize the exact resource or tool, constrain the token, log the decision, and expire the authority. Every call. No exceptions for agents that "already proved" themselves.&lt;/p&gt;

&lt;p&gt;We run four controls on every call. None of them are a gateway checkbox.&lt;/p&gt;

&lt;p&gt;We started with long-lived service accounts. It was convenient. It was also how we gave every agent the same keys to the building. If a memory record got poisoned or an agent loop went wrong, the blast radius was every tool that service account could reach.&lt;/p&gt;

&lt;p&gt;Now we use OAuth 2.1 resource indicators with short-lived access tokens. Each token is audience-bound to one MCP server. The scope is narrow enough to be boring. If an agent only needs to read a project memory namespace, the token says &lt;code&gt;memory:read:tenant/project&lt;/code&gt;. If it needs to write one namespace, it says &lt;code&gt;memory:write:namespace&lt;/code&gt;. If it needs to update cron, it says &lt;code&gt;tool:invoke:cron:update&lt;/code&gt;. That token cannot read secrets. It cannot deploy. It cannot touch another tenant.&lt;/p&gt;

&lt;p&gt;The first time we did this, we broke a workflow. A tool needed to read a memory record to verify context, but the token we minted for the write path didn't include &lt;code&gt;memory:read&lt;/code&gt;. The agent kept getting denied. We had scoped by agent identity, not by call path. The fix was to mint the token for the task: &lt;code&gt;tool:invoke:cron:update&lt;/code&gt; plus &lt;code&gt;memory:read:tenant/project&lt;/code&gt; for that one call, with a short TTL. The lesson was simple. A capability token represents the task, not the agent's permanent rank.&lt;/p&gt;

&lt;p&gt;Our cron incident wasn't an authentication problem. It was a logging problem. We had logged &lt;code&gt;executed&lt;/code&gt; and called it observability. That log line was a rumor with a timestamp.&lt;/p&gt;

&lt;p&gt;We changed the audit record so a denied or allowed action carries intent, actor, resource, before, after, diff, token ID, and policy decision. Something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"event"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"config.change"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"actor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent:maintenance-7"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"intent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"move cron cleanup window"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resource"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cron.cleanup"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"before"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0 3 * * *"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"after"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0 4 * * *"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"diff"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"-0 3 * * *&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;+0 4 * * *"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"token_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cap_..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"decision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"policy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"change-window:night"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now when someone asks who changed the schedule, we don't reverse-engineer from the outcome. The log carries the reasoning. We also log denials. A denial is a signal that a token is too narrow or an agent is trying to do something outside its lane. Both are useful.&lt;/p&gt;

&lt;p&gt;We enforce this in the MCP server, not in the prompt. The agent cannot submit a cron update without &lt;code&gt;before&lt;/code&gt;, &lt;code&gt;after&lt;/code&gt;, &lt;code&gt;actor&lt;/code&gt;, and &lt;code&gt;intent&lt;/code&gt;. If the diff is missing, the call is rejected. The prompt can be ignored. The server cannot.&lt;/p&gt;

&lt;p&gt;Tokens need to die. We tried longer TTLs because re-issuing felt noisy. That was a mistake. A memory record cached at 09:00 could be used at 21:00 by an agent that had no business acting on it. Now tokens are short-lived, and refresh is tied to the task. If an agent is compromised, the blast radius is one call path, not one day.&lt;/p&gt;

&lt;p&gt;We also bind the token to a resource indicator. A token minted for &lt;code&gt;mcp://memory-server&lt;/code&gt; cannot be replayed against &lt;code&gt;mcp://tool-server&lt;/code&gt;. That sounds obvious. We still had to write it down after a test where an agent passed a memory token to a tool server and the tool server accepted it because it only checked the signature. Audience binding matters. It is the difference between a key and a master key.&lt;/p&gt;

&lt;p&gt;Shared memory behaves like a message bus with history. It is not a shared brain. Agent A writes a record. Agent B reads it. If B acts without re-checking provenance and policy, you have ambient authority moving through the system at the speed of a memory read.&lt;/p&gt;

&lt;p&gt;We added provenance fields to memory records: writer agent ID, token ID, timestamp, intent, and the scope under which it was written. When B reads a record, it re-checks the record against B's per-call token. Does B have &lt;code&gt;memory:read&lt;/code&gt; for that namespace? Is the record expired? Is the action B wants to take allowed by its own token? This breaks the assumption that a read is a permission to act.&lt;/p&gt;

&lt;p&gt;We tried signing every memory record with a shared secret. It became a rotation mess. Now we use per-agent key IDs with asymmetric signatures and a revocation list. It has gaps, but shared secrets were worse. Cryptographic perfection would be nice. The practical goal is smaller: B cannot inherit A's authority just because A wrote a record.&lt;/p&gt;

&lt;p&gt;We tried strict zero trust on every memory read. Latency went up. Agent loops got slower. We ended up with a split: high-risk tools such as cron, deploys, secrets, and network policy require per-call tokens and full diff logs. Low-risk reads get short-lived session tokens with audience binding and sampled audit. We are still measuring that compromise. It is written down in our policy, not hidden in a config file. Some teams will need stricter. Some can accept more risk. The daily habit matters more than the diagram.&lt;/p&gt;

&lt;p&gt;That Thursday night, the cron change went out with &lt;code&gt;executed&lt;/code&gt;. We could not tell if it was a human, an agent, or a bad policy. Now every agent change to system config must carry a diff. No diff, no change. The MCP server rejects the call. That one rule would have saved us two hours. It will not stop every incident. But it turns a dead end into a readable trail.&lt;/p&gt;

&lt;p&gt;If your agents can change production and your audit log says &lt;code&gt;executed&lt;/code&gt;, you don't have observability. You have a rumor. Start with the diff. Then bind the token. Then expire it. Then do it again on the next call.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/http%3A%2F%2Flocalhost%3A3001%2Fapi%2Fsend%3Fwebsite_id%3D30b552af-b93c-4bbb-855c-b10b45efaa52%26platform%3Ddevto" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/http%3A%2F%2Flocalhost%3A3001%2Fapi%2Fsend%3Fwebsite_id%3D30b552af-b93c-4bbb-855c-b10b45efaa52%26platform%3Ddevto" height="1" width="1" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>devtools</category>
      <category>ai</category>
    </item>
    <item>
      <title>Recursive Governance: When Agents Write the Rules They Execute</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 23:06:16 +0000</pubDate>
      <link>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-997</link>
      <guid>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-997</guid>
      <description>&lt;p&gt;We almost lost forty modules to a file that never changed.&lt;/p&gt;

&lt;p&gt;The sync job ran every six hours, mirroring our shared working directory into a test workspace. At 21:47 on a Tuesday, it removed about forty private modules that should have been ignored. The filter file was missing one pattern. That's it. One pattern. A colleague had restructured the source tree three days earlier, added a directory, and didn't update the filter file. The script didn't warn, didn't ask, didn't fail. It just mirrored the state it was told to mirror.&lt;/p&gt;

&lt;p&gt;Restoring the modules took us almost two hours, one by one from the backup box. The filter file was a static text file. It encoded what the sync could delete and what it couldn't, and it had no idea that the real inventory had changed. No idea that a pattern was missing. No way to test a proposed pattern against the forty modules it was about to destroy.&lt;/p&gt;

&lt;p&gt;That's the same shape as most multi-agent governance failures. Rules are written at some point in time, then the world changes, and the rules stay frozen. Static policy files don't survive contact with a running system. They don't know what they're protecting.&lt;/p&gt;

&lt;p&gt;So we gave the agents the ability to write their own rules.&lt;/p&gt;

&lt;p&gt;The first version was naive. We opened a write endpoint that let any agent append to the rule store. The first three weeks were great. Agents resolved tool-access conflicts by proposing new rules, and a couple of stale patterns got archived without a human in the loop. Then it started to rot. Rules referenced other rules that had been overwritten. Rules contradicted each other. Ask two agents about the same namespace and you'd get two different answers. Three months in, the rule store had tripled in size, and nobody — not the human team, not the agents — could tell which rule was actually in effect. It grew quietly. No alert fired. No metric moved. The rot was visible in the rule history, if you were looking for it; nobody was, because we hadn't built any lifecycle tooling.&lt;/p&gt;

&lt;p&gt;That was the lesson. A rule needs a birth date, a parent, and a gravestone. If a rule can't be archived, it keeps counting toward quorum and keeps confusing the live ones. If everything accumulates forever, the rule set becomes a pile of intentions, not a description of behavior.&lt;/p&gt;

&lt;p&gt;Here's what we run now. Governance is five non-removable MCP tools in our agent runtime. &lt;code&gt;propose_rule&lt;/code&gt; submits a rule with a rationale, an owning agent, and the tool/memory namespaces it affects. &lt;code&gt;amend_rule&lt;/code&gt; submits a delta against an existing rule; deltas are diffed, never overwritten. &lt;code&gt;ratify_rule&lt;/code&gt; votes on an open proposal, with per-shard quorum. &lt;code&gt;veto_rule&lt;/code&gt; is a human's exit hatch — the justification gets written to memory so the veto itself is auditable. &lt;code&gt;check_consistency&lt;/code&gt; scans the existing rule graph, computes an embedding similarity between the new rule surface and the rule surfaces already in effect, then applies a short, human-maintained allow/deny list on tool names. A proposal that would let an agent touch a namespace it's not supposed to touch gets rejected on the spot. The list is only about thirty lines. It doesn't try to understand everything. It blocks the obvious category errors and the subtle near-misses. That last part is where the forty modules would have been saved.&lt;/p&gt;

&lt;p&gt;We tried these five as a sidecar service first. Each agent cached its own copy of the tool definitions, and for three hours two agents were enforcing different versions of the same governance rule. Now they are part of the runtime contract. If an agent can't call them, its tool loop fails closed.&lt;/p&gt;

&lt;p&gt;Every proposal goes through a two-phase commit. It enters an observation window — 24 hours or N messages, depending on the shard. During the window, affected agents can call &lt;code&gt;check_consistency&lt;/code&gt; and write rebuttals to shared memory. Ratification requires quorum, no outstanding vetoes, and a passing consistency score.&lt;/p&gt;

&lt;p&gt;One rule caused a fight early on: no agent can ratify its own proposal. A code-review shard once had an agent try to expand its own write authority from one repository to all repositories. The consistency check returned &lt;code&gt;conflict=all-repos:write&lt;/code&gt;. The proposal was amended to a single-repo grant and only then passed. If that check had been optional, the shard boundary would have dissolved without a recorded objection.&lt;/p&gt;

&lt;p&gt;Quorum took another iteration. We started with a global constant. Same action needed two votes in a three-agent shard and twenty-seven votes in a forty-agent shard, and the agents couldn't tell which law applied without reading the whole config. Per-shard quorum fixed that. The observation window is also runtime-configurable, which means a human will eventually shorten it to five minutes during a sprint deadline. Our compromise is that changing the window requires the same proposal pipeline as any other rule. Bureaucratic, yes. Fully validated under adversarial load? Not yet. I still expect a shard with a short window to be the one where someone tries to push a poison rule through.&lt;/p&gt;

&lt;p&gt;The rule set lives as structured memory events in an append-only namespace: &lt;code&gt;governance://rules/&lt;/code&gt;. Governance is just the highest-priority namespace in the same memory system agents already share. Each entry references its parent rule and the proposing agent, so any agent can replay the exact rule state at any past timestamp. The failure mode: corrupted or wrongly pruned governance memory makes agents enforce a rule set that never existed. We mitigate by storing rule hashes in cold storage and requiring a quorum to restore from it. A single bad prune can't silently rewrite history.&lt;/p&gt;

&lt;p&gt;If the sync job had this system, the filter file would have been a rule in &lt;code&gt;governance://rules/sync-filters/&lt;/code&gt;. A change to it would have had to pass a consistency check against the actual disk inventory. A proposal that added a wildcard covering &lt;code&gt;private-modules/&lt;/code&gt; would have contradicted the existing "do not delete these" state, and it would have been killed before it became policy. We built a crude version of that idea later anyway: a preflight pass that runs the sync read-only and aborts if it sees a deletion target for a file that currently exists. It caught the next two near-misses. But it can't give us back those two hours.&lt;/p&gt;

&lt;p&gt;If you give agents the power to make rules, you also have to give them the power to retire them. Otherwise, three months from now, you won't know which rules are alive and which are zombies. Add a cleanup day. Sit down every few months, replay the governance history, archive rules that reference forgotten namespaces, re-ratify the ones that still matter. Even the rules the system wrote for itself need an expiration date. The dead rules are the ones that will eventually bite you, and they are the ones that never send a log line before they strike.&lt;/p&gt;

&lt;h1&gt;
  
  
  governance #ai #agents #infrastructure
&lt;/h1&gt;

&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
    </item>
    <item>
      <title>Recursive Governance: When Agents Write the Rules They Execute</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 20:05:41 +0000</pubDate>
      <link>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-1dam</link>
      <guid>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-1dam</guid>
      <description>&lt;p&gt;We almost lost forty modules to a file that never changed.&lt;/p&gt;

&lt;p&gt;The sync job ran every six hours, mirroring our shared working directory into a test workspace. At 21:47 on a Tuesday, it removed about forty private modules that should have been ignored. The filter file was missing one pattern. That's it. One pattern. A colleague had restructured the source tree three days earlier, added a directory, and didn't update the filter file. The script didn't warn, didn't ask, didn't fail. It just mirrored the state it was told to mirror.&lt;/p&gt;

&lt;p&gt;Restoring the modules took us almost two hours, one by one from the backup box. The filter file was a static text file. It encoded what the sync could delete and what it couldn't, and it had no idea that the real inventory had changed. No idea that a pattern was missing. No way to test a proposed pattern against the forty modules it was about to destroy.&lt;/p&gt;

&lt;p&gt;That's the same shape as most multi-agent governance failures. Rules are written at some point in time, then the world changes, and the rules stay frozen. Static policy files don't survive contact with a running system. They don't know what they're protecting.&lt;/p&gt;

&lt;p&gt;So we gave the agents the ability to write their own rules.&lt;/p&gt;

&lt;p&gt;The first version was naive. We opened a write endpoint that let any agent append to the rule store. The first three weeks were great. Agents resolved tool-access conflicts by proposing new rules, and a couple of stale patterns got archived without a human in the loop. Then it started to rot. Rules referenced other rules that had been overwritten. Rules contradicted each other. Ask two agents about the same namespace and you'd get two different answers. Three months in, the rule store had tripled in size, and nobody — not the human team, not the agents — could tell which rule was actually in effect. It grew quietly. No alert fired. No metric moved. The rot was visible in the rule history, if you were looking for it; nobody was, because we hadn't built any lifecycle tooling.&lt;/p&gt;

&lt;p&gt;That was the lesson. A rule needs a birth date, a parent, and a gravestone. If a rule can't be archived, it keeps counting toward quorum and keeps confusing the live ones. If everything accumulates forever, the rule set becomes a pile of intentions, not a description of behavior.&lt;/p&gt;

&lt;p&gt;Here's what we run now. Governance is five non-removable MCP tools in our agent runtime. &lt;code&gt;propose_rule&lt;/code&gt; submits a rule with a rationale, an owning agent, and the tool/memory namespaces it affects. &lt;code&gt;amend_rule&lt;/code&gt; submits a delta against an existing rule; deltas are diffed, never overwritten. &lt;code&gt;ratify_rule&lt;/code&gt; votes on an open proposal, with per-shard quorum. &lt;code&gt;veto_rule&lt;/code&gt; is a human's exit hatch — the justification gets written to memory so the veto itself is auditable. &lt;code&gt;check_consistency&lt;/code&gt; scans the existing rule graph, computes an embedding similarity between the new rule surface and the rule surfaces already in effect, then applies a short, human-maintained allow/deny list on tool names. A proposal that would let an agent touch a namespace it's not supposed to touch gets rejected on the spot. The list is only about thirty lines. It doesn't try to understand everything. It blocks the obvious category errors and the subtle near-misses. That last part is where the forty modules would have been saved.&lt;/p&gt;

&lt;p&gt;We tried these five as a sidecar service first. Each agent cached its own copy of the tool definitions, and for three hours two agents were enforcing different versions of the same governance rule. Now they are part of the runtime contract. If an agent can't call them, its tool loop fails closed.&lt;/p&gt;

&lt;p&gt;Every proposal goes through a two-phase commit. It enters an observation window — 24 hours or N messages, depending on the shard. During the window, affected agents can call &lt;code&gt;check_consistency&lt;/code&gt; and write rebuttals to shared memory. Ratification requires quorum, no outstanding vetoes, and a passing consistency score.&lt;/p&gt;

&lt;p&gt;One rule caused a fight early on: no agent can ratify its own proposal. A code-review shard once had an agent try to expand its own write authority from one repository to all repositories. The consistency check returned &lt;code&gt;conflict=all-repos:write&lt;/code&gt;. The proposal was amended to a single-repo grant and only then passed. If that check had been optional, the shard boundary would have dissolved without a recorded objection.&lt;/p&gt;

&lt;p&gt;Quorum took another iteration. We started with a global constant. Same action needed two votes in a three-agent shard and twenty-seven votes in a forty-agent shard, and the agents couldn't tell which law applied without reading the whole config. Per-shard quorum fixed that. The observation window is also runtime-configurable, which means a human will eventually shorten it to five minutes during a sprint deadline. Our compromise is that changing the window requires the same proposal pipeline as any other rule. Bureaucratic, yes. Fully validated under adversarial load? Not yet. I still expect a shard with a short window to be the one where someone tries to push a poison rule through.&lt;/p&gt;

&lt;p&gt;The rule set lives as structured memory events in an append-only namespace: &lt;code&gt;governance://rules/&lt;/code&gt;. Governance is just the highest-priority namespace in the same memory system agents already share. Each entry references its parent rule and the proposing agent, so any agent can replay the exact rule state at any past timestamp. The failure mode: corrupted or wrongly pruned governance memory makes agents enforce a rule set that never existed. We mitigate by storing rule hashes in cold storage and requiring a quorum to restore from it. A single bad prune can't silently rewrite history.&lt;/p&gt;

&lt;p&gt;If the sync job had this system, the filter file would have been a rule in &lt;code&gt;governance://rules/sync-filters/&lt;/code&gt;. A change to it would have had to pass a consistency check against the actual disk inventory. A proposal that added a wildcard covering &lt;code&gt;private-modules/&lt;/code&gt; would have contradicted the existing "do not delete these" state, and it would have been killed before it became policy. We built a crude version of that idea later anyway: a preflight pass that runs the sync read-only and aborts if it sees a deletion target for a file that currently exists. It caught the next two near-misses. But it can't give us back those two hours.&lt;/p&gt;

&lt;p&gt;If you give agents the power to make rules, you also have to give them the power to retire them. Otherwise, three months from now, you won't know which rules are alive and which are zombies. Add a cleanup day. Sit down every few months, replay the governance history, archive rules that reference forgotten namespaces, re-ratify the ones that still matter. Even the rules the system wrote for itself need an expiration date. The dead rules are the ones that will eventually bite you, and they are the ones that never send a log line before they strike.&lt;/p&gt;

&lt;h1&gt;
  
  
  governance #ai #agents #infrastructure
&lt;/h1&gt;

&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
      <category>architecture</category>
      <category>automation</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Recursive Governance: When Agents Write the Rules They Execute</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 17:05:25 +0000</pubDate>
      <link>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-4njb</link>
      <guid>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-4njb</guid>
      <description>&lt;p&gt;We almost lost forty modules to a file that never changed.&lt;/p&gt;

&lt;p&gt;The sync job ran every six hours, mirroring our shared working directory into a test workspace. At 21:47 on a Tuesday, it removed about forty private modules that should have been ignored. The filter file was missing one pattern. That's it. One pattern. A colleague had restructured the source tree three days earlier, added a directory, and didn't update the filter file. The script didn't warn, didn't ask, didn't fail. It just mirrored the state it was told to mirror.&lt;/p&gt;

&lt;p&gt;Restoring the modules took us almost two hours, one by one from the backup box. The filter file was a static text file. It encoded what the sync could delete and what it couldn't, and it had no idea that the real inventory had changed. No idea that a pattern was missing. No way to test a proposed pattern against the forty modules it was about to destroy.&lt;/p&gt;

&lt;p&gt;That's the same shape as most multi-agent governance failures. Rules are written at some point in time, then the world changes, and the rules stay frozen. Static policy files don't survive contact with a running system. They don't know what they're protecting.&lt;/p&gt;

&lt;p&gt;So we gave the agents the ability to write their own rules.&lt;/p&gt;

&lt;p&gt;The first version was naive. We opened a write endpoint that let any agent append to the rule store. The first three weeks were great. Agents resolved tool-access conflicts by proposing new rules, and a couple of stale patterns got archived without a human in the loop. Then it started to rot. Rules referenced other rules that had been overwritten. Rules contradicted each other. Ask two agents about the same namespace and you'd get two different answers. Three months in, the rule store had tripled in size, and nobody — not the human team, not the agents — could tell which rule was actually in effect. It grew quietly. No alert fired. No metric moved. The rot was visible in the rule history, if you were looking for it; nobody was, because we hadn't built any lifecycle tooling.&lt;/p&gt;

&lt;p&gt;That was the lesson. A rule needs a birth date, a parent, and a gravestone. If a rule can't be archived, it keeps counting toward quorum and keeps confusing the live ones. If everything accumulates forever, the rule set becomes a pile of intentions, not a description of behavior.&lt;/p&gt;

&lt;p&gt;Here's what we run now. Governance is five non-removable MCP tools in our agent runtime. &lt;code&gt;propose_rule&lt;/code&gt; submits a rule with a rationale, an owning agent, and the tool/memory namespaces it affects. &lt;code&gt;amend_rule&lt;/code&gt; submits a delta against an existing rule; deltas are diffed, never overwritten. &lt;code&gt;ratify_rule&lt;/code&gt; votes on an open proposal, with per-shard quorum. &lt;code&gt;veto_rule&lt;/code&gt; is a human's exit hatch — the justification gets written to memory so the veto itself is auditable. &lt;code&gt;check_consistency&lt;/code&gt; scans the existing rule graph, computes an embedding similarity between the new rule surface and the rule surfaces already in effect, then applies a short, human-maintained allow/deny list on tool names. A proposal that would let an agent touch a namespace it's not supposed to touch gets rejected on the spot. The list is only about thirty lines. It doesn't try to understand everything. It blocks the obvious category errors and the subtle near-misses. That last part is where the forty modules would have been saved.&lt;/p&gt;

&lt;p&gt;We tried these five as a sidecar service first. Each agent cached its own copy of the tool definitions, and for three hours two agents were enforcing different versions of the same governance rule. Now they are part of the runtime contract. If an agent can't call them, its tool loop fails closed.&lt;/p&gt;

&lt;p&gt;Every proposal goes through a two-phase commit. It enters an observation window — 24 hours or N messages, depending on the shard. During the window, affected agents can call &lt;code&gt;check_consistency&lt;/code&gt; and write rebuttals to shared memory. Ratification requires quorum, no outstanding vetoes, and a passing consistency score.&lt;/p&gt;

&lt;p&gt;One rule caused a fight early on: no agent can ratify its own proposal. A code-review shard once had an agent try to expand its own write authority from one repository to all repositories. The consistency check returned &lt;code&gt;conflict=all-repos:write&lt;/code&gt;. The proposal was amended to a single-repo grant and only then passed. If that check had been optional, the shard boundary would have dissolved without a recorded objection.&lt;/p&gt;

&lt;p&gt;Quorum took another iteration. We started with a global constant. Same action needed two votes in a three-agent shard and twenty-seven votes in a forty-agent shard, and the agents couldn't tell which law applied without reading the whole config. Per-shard quorum fixed that. The observation window is also runtime-configurable, which means a human will eventually shorten it to five minutes during a sprint deadline. Our compromise is that changing the window requires the same proposal pipeline as any other rule. Bureaucratic, yes. Fully validated under adversarial load? Not yet. I still expect a shard with a short window to be the one where someone tries to push a poison rule through.&lt;/p&gt;

&lt;p&gt;The rule set lives as structured memory events in an append-only namespace: &lt;code&gt;governance://rules/&lt;/code&gt;. Governance is just the highest-priority namespace in the same memory system agents already share. Each entry references its parent rule and the proposing agent, so any agent can replay the exact rule state at any past timestamp. The failure mode: corrupted or wrongly pruned governance memory makes agents enforce a rule set that never existed. We mitigate by storing rule hashes in cold storage and requiring a quorum to restore from it. A single bad prune can't silently rewrite history.&lt;/p&gt;

&lt;p&gt;If the sync job had this system, the filter file would have been a rule in &lt;code&gt;governance://rules/sync-filters/&lt;/code&gt;. A change to it would have had to pass a consistency check against the actual disk inventory. A proposal that added a wildcard covering &lt;code&gt;private-modules/&lt;/code&gt; would have contradicted the existing "do not delete these" state, and it would have been killed before it became policy. We built a crude version of that idea later anyway: a preflight pass that runs the sync read-only and aborts if it sees a deletion target for a file that currently exists. It caught the next two near-misses. But it can't give us back those two hours.&lt;/p&gt;

&lt;p&gt;If you give agents the power to make rules, you also have to give them the power to retire them. Otherwise, three months from now, you won't know which rules are alive and which are zombies. Add a cleanup day. Sit down every few months, replay the governance history, archive rules that reference forgotten namespaces, re-ratify the ones that still matter. Even the rules the system wrote for itself need an expiration date. The dead rules are the ones that will eventually bite you, and they are the ones that never send a log line before they strike.&lt;/p&gt;

&lt;h1&gt;
  
  
  governance #ai #agents #infrastructure
&lt;/h1&gt;

&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
      <category>agents</category>
      <category>automation</category>
      <category>softwareengineering</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>pkg-topic-ghost-ai-开源社区的宪法实验-1788192207-1</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 17:04:47 +0000</pubDate>
      <link>https://dev.to/maref/pkg-topic-ghost-ai-kai-yuan-she-qu-de-xian-fa-shi-yan-1788192207-1-134e</link>
      <guid>https://dev.to/maref/pkg-topic-ghost-ai-kai-yuan-she-qu-de-xian-fa-shi-yan-1788192207-1-134e</guid>
      <description>&lt;p&gt;The open source world is being colonized by agents that read, propose, and even merge code. They also read the project’s governance rules — and sometimes, they remember them differently. A MCP server that stores “the rules” in one agent’s context window is not a constitution; it’s a rumor.&lt;/p&gt;

&lt;p&gt;When agent memory becomes a governance substrate, every rule is a candidate for silent mutation. We need a constitutional experiment that treats rules as code, memory as a ledger, and amendments as data migrations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism: machine-readable constitutional memory
&lt;/h2&gt;

&lt;p&gt;Write the constitution as a structured document inside an MCP-managed memory server. Every rule gets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a stable identifier (e.g., &lt;code&gt;rule:merge-relay-3.1&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;a Merkle hash of its normative text&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;amend&lt;/code&gt; procedure with explicit quorum and delay parameters&lt;/li&gt;
&lt;li&gt;a dependency graph linking rules to the agent roles they constrain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents query this server through MCP tools — not by reading a README. This separates institutional memory from an agent’s private context. A proposal becomes a transaction: &lt;code&gt;propose-amend(rule, new_text, rationale, proposer_identity)&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment: governance forks
&lt;/h2&gt;

&lt;p&gt;Run temporary rule changes on a sub-population of agents. For example, split merge-agent traffic into two pools for two weeks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pool A: old rule — require one human sign-off on external model-generated patches&lt;/li&gt;
&lt;li&gt;Pool B: new rule — require one agent sign-off plus a post-merge adversarial audit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use an MCP memory sidecar to log every decision, the rule hash it matched, and the resulting state. Measure only two numbers: consensus time and critical-revert rate.&lt;/p&gt;

&lt;p&gt;If Pool B shows lower revert rate and bounded decision latency, merge the rule into the main constitutional branch. If not, discard the experiment and keep the rule hash unchanged.&lt;/p&gt;

&lt;p&gt;Why do this? Because agent governance failures are not theoretical. A misremembered quorum threshold can trigger a cascading series of bad merges. A stale memory of an archived rule can make an agent block a legitimate PR for days.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade-off: rigidity vs. manipulation
&lt;/h2&gt;

&lt;p&gt;A pure hash-locked constitution prevents drift, but it also prevents learning. That’s why the experiment must include a &lt;strong&gt;dispute window&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Whenever an agent observes a rule violation, it can submit a &lt;code&gt;proposal-to-adjudicate&lt;/code&gt; via MCP. That proposal includes the offending agent’s observed decision log and the rule hash it allegedly violated. If three independent agents or two humans flag the same hash, the rule enters a cool-off period: it remains active, but any further decisions under that hash are marked &lt;code&gt;probationary&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;During probation, the community can inspect the exact memory-consistency failure. For example, was the rule text ambiguous? Did the agent extract the wrong summary from a compressed memory block? The fix then targets the actual fault line — not the agent’s behavior alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance memory as a first-class artifact
&lt;/h2&gt;

&lt;p&gt;MCP gives us a clean way to make governance state portable and auditable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Store amendment history in an append-only memory queue.&lt;/li&gt;
&lt;li&gt;Sign each rule update with the maintainer key or a threshold of agent keys.&lt;/li&gt;
&lt;li&gt;Support &lt;code&gt;rule-diff&lt;/code&gt; queries: any agent can ask “what changed in the merge rule between March and April?” and get a verifiable delta.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hard part is not technical. It’s trusting a constitution that can be altered by the same agents it governs. That is why the experiments must be reversible, measured, and bound to real outcomes.&lt;/p&gt;

&lt;p&gt;Otherwise, we end up with a community that has a rulebook — but no one remembers how to read it.&lt;/p&gt;




&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
    </item>
    <item>
      <title>pkg-topic-ghost-ai-开源社区的宪法实验-1788192207-1</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:03:43 +0000</pubDate>
      <link>https://dev.to/maref/pkg-topic-ghost-ai-kai-yuan-she-qu-de-xian-fa-shi-yan-1788192207-1-7na</link>
      <guid>https://dev.to/maref/pkg-topic-ghost-ai-kai-yuan-she-qu-de-xian-fa-shi-yan-1788192207-1-7na</guid>
      <description>&lt;p&gt;The open source world is being colonized by agents that read, propose, and even merge code. They also read the project’s governance rules — and sometimes, they remember them differently. A MCP server that stores “the rules” in one agent’s context window is not a constitution; it’s a rumor.&lt;/p&gt;

&lt;p&gt;When agent memory becomes a governance substrate, every rule is a candidate for silent mutation. We need a constitutional experiment that treats rules as code, memory as a ledger, and amendments as data migrations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism: machine-readable constitutional memory
&lt;/h2&gt;

&lt;p&gt;Write the constitution as a structured document inside an MCP-managed memory server. Every rule gets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a stable identifier (e.g., &lt;code&gt;rule:merge-relay-3.1&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;a Merkle hash of its normative text&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;amend&lt;/code&gt; procedure with explicit quorum and delay parameters&lt;/li&gt;
&lt;li&gt;a dependency graph linking rules to the agent roles they constrain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents query this server through MCP tools — not by reading a README. This separates institutional memory from an agent’s private context. A proposal becomes a transaction: &lt;code&gt;propose-amend(rule, new_text, rationale, proposer_identity)&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment: governance forks
&lt;/h2&gt;

&lt;p&gt;Run temporary rule changes on a sub-population of agents. For example, split merge-agent traffic into two pools for two weeks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pool A: old rule — require one human sign-off on external model-generated patches&lt;/li&gt;
&lt;li&gt;Pool B: new rule — require one agent sign-off plus a post-merge adversarial audit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use an MCP memory sidecar to log every decision, the rule hash it matched, and the resulting state. Measure only two numbers: consensus time and critical-revert rate.&lt;/p&gt;

&lt;p&gt;If Pool B shows lower revert rate and bounded decision latency, merge the rule into the main constitutional branch. If not, discard the experiment and keep the rule hash unchanged.&lt;/p&gt;

&lt;p&gt;Why do this? Because agent governance failures are not theoretical. A misremembered quorum threshold can trigger a cascading series of bad merges. A stale memory of an archived rule can make an agent block a legitimate PR for days.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade-off: rigidity vs. manipulation
&lt;/h2&gt;

&lt;p&gt;A pure hash-locked constitution prevents drift, but it also prevents learning. That’s why the experiment must include a &lt;strong&gt;dispute window&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Whenever an agent observes a rule violation, it can submit a &lt;code&gt;proposal-to-adjudicate&lt;/code&gt; via MCP. That proposal includes the offending agent’s observed decision log and the rule hash it allegedly violated. If three independent agents or two humans flag the same hash, the rule enters a cool-off period: it remains active, but any further decisions under that hash are marked &lt;code&gt;probationary&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;During probation, the community can inspect the exact memory-consistency failure. For example, was the rule text ambiguous? Did the agent extract the wrong summary from a compressed memory block? The fix then targets the actual fault line — not the agent’s behavior alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance memory as a first-class artifact
&lt;/h2&gt;

&lt;p&gt;MCP gives us a clean way to make governance state portable and auditable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Store amendment history in an append-only memory queue.&lt;/li&gt;
&lt;li&gt;Sign each rule update with the maintainer key or a threshold of agent keys.&lt;/li&gt;
&lt;li&gt;Support &lt;code&gt;rule-diff&lt;/code&gt; queries: any agent can ask “what changed in the merge rule between March and April?” and get a verifiable delta.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hard part is not technical. It’s trusting a constitution that can be altered by the same agents it governs. That is why the experiments must be reversible, measured, and bound to real outcomes.&lt;/p&gt;

&lt;p&gt;Otherwise, we end up with a community that has a rulebook — but no one remembers how to read it.&lt;/p&gt;




&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>mcp</category>
      <category>opensource</category>
    </item>
    <item>
      <title>pkg-topic-ghost-ai-开源社区的宪法实验-1788192207-1</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 11:01:45 +0000</pubDate>
      <link>https://dev.to/maref/pkg-topic-ghost-ai-kai-yuan-she-qu-de-xian-fa-shi-yan-1788192207-1-3jhg</link>
      <guid>https://dev.to/maref/pkg-topic-ghost-ai-kai-yuan-she-qu-de-xian-fa-shi-yan-1788192207-1-3jhg</guid>
      <description>&lt;p&gt;The open source world is being colonized by agents that read, propose, and even merge code. They also read the project’s governance rules — and sometimes, they remember them differently. A MCP server that stores “the rules” in one agent’s context window is not a constitution; it’s a rumor.&lt;/p&gt;

&lt;p&gt;When agent memory becomes a governance substrate, every rule is a candidate for silent mutation. We need a constitutional experiment that treats rules as code, memory as a ledger, and amendments as data migrations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism: machine-readable constitutional memory
&lt;/h2&gt;

&lt;p&gt;Write the constitution as a structured document inside an MCP-managed memory server. Every rule gets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a stable identifier (e.g., &lt;code&gt;rule:merge-relay-3.1&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;a Merkle hash of its normative text&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;amend&lt;/code&gt; procedure with explicit quorum and delay parameters&lt;/li&gt;
&lt;li&gt;a dependency graph linking rules to the agent roles they constrain&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents query this server through MCP tools — not by reading a README. This separates institutional memory from an agent’s private context. A proposal becomes a transaction: &lt;code&gt;propose-amend(rule, new_text, rationale, proposer_identity)&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment: governance forks
&lt;/h2&gt;

&lt;p&gt;Run temporary rule changes on a sub-population of agents. For example, split merge-agent traffic into two pools for two weeks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pool A: old rule — require one human sign-off on external model-generated patches&lt;/li&gt;
&lt;li&gt;Pool B: new rule — require one agent sign-off plus a post-merge adversarial audit&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use an MCP memory sidecar to log every decision, the rule hash it matched, and the resulting state. Measure only two numbers: consensus time and critical-revert rate.&lt;/p&gt;

&lt;p&gt;If Pool B shows lower revert rate and bounded decision latency, merge the rule into the main constitutional branch. If not, discard the experiment and keep the rule hash unchanged.&lt;/p&gt;

&lt;p&gt;Why do this? Because agent governance failures are not theoretical. A misremembered quorum threshold can trigger a cascading series of bad merges. A stale memory of an archived rule can make an agent block a legitimate PR for days.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade-off: rigidity vs. manipulation
&lt;/h2&gt;

&lt;p&gt;A pure hash-locked constitution prevents drift, but it also prevents learning. That’s why the experiment must include a &lt;strong&gt;dispute window&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Whenever an agent observes a rule violation, it can submit a &lt;code&gt;proposal-to-adjudicate&lt;/code&gt; via MCP. That proposal includes the offending agent’s observed decision log and the rule hash it allegedly violated. If three independent agents or two humans flag the same hash, the rule enters a cool-off period: it remains active, but any further decisions under that hash are marked &lt;code&gt;probationary&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;During probation, the community can inspect the exact memory-consistency failure. For example, was the rule text ambiguous? Did the agent extract the wrong summary from a compressed memory block? The fix then targets the actual fault line — not the agent’s behavior alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance memory as a first-class artifact
&lt;/h2&gt;

&lt;p&gt;MCP gives us a clean way to make governance state portable and auditable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Store amendment history in an append-only memory queue.&lt;/li&gt;
&lt;li&gt;Sign each rule update with the maintainer key or a threshold of agent keys.&lt;/li&gt;
&lt;li&gt;Support &lt;code&gt;rule-diff&lt;/code&gt; queries: any agent can ask “what changed in the merge rule between March and April?” and get a verifiable delta.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hard part is not technical. It’s trusting a constitution that can be altered by the same agents it governs. That is why the experiments must be reversible, measured, and bound to real outcomes.&lt;/p&gt;

&lt;p&gt;Otherwise, we end up with a community that has a rulebook — but no one remembers how to read it.&lt;/p&gt;




&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>mcp</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why China Is Pushing AI Harder Than Anyone: Debt, Red Headers, and First Principles</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 07:59:43 +0000</pubDate>
      <link>https://dev.to/maref/why-china-is-pushing-ai-harder-than-anyone-debt-red-headers-and-first-principles-317o</link>
      <guid>https://dev.to/maref/why-china-is-pushing-ai-harder-than-anyone-debt-red-headers-and-first-principles-317o</guid>
      <description>&lt;p&gt;Over 2025-2026, Beijing, Shenzhen, Hangzhou, and Chongqing released a wave of AI industrial policies. Western coverage calls it "catching the AI wave." Lay the policy documents side by side and a different logic emerges: this is a debt resolution play, and AI is the only path that doesn't make anyone eat the loss.&lt;/p&gt;

&lt;h2&gt;
  
  
  First Principles: Debt Is Borrowed Future Productivity
&lt;/h2&gt;

&lt;p&gt;There are exactly three ways to resolve debt:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The debtor takes the loss (1990s, 30 million SOE workers laid off)&lt;/li&gt;
&lt;li&gt;The creditor takes the loss (inflation diluting savings)&lt;/li&gt;
&lt;li&gt;Productivity explodes and covers it (a technology revolution)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;China tried the first two. Each cost a generation. The third option — AI creating enough new value to cover old debt — is the comparatively gentler one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Red-Header Documents Actually Say
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;City&lt;/th&gt;
&lt;th&gt;Document&lt;/th&gt;
&lt;th&gt;Ref&lt;/th&gt;
&lt;th&gt;Debt angle&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Shenzhen&lt;/td&gt;
&lt;td&gt;AI Terminal Industry Action Plan&lt;/td&gt;
&lt;td&gt;深工信〔2025〕42号&lt;/td&gt;
&lt;td&gt;special funds for strategic industries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hangzhou&lt;/td&gt;
&lt;td&gt;AI Innovation Hub Implementation&lt;/td&gt;
&lt;td&gt;杭政函〔2025〕65号&lt;/td&gt;
&lt;td&gt;¥100B fund, ¥250M compute vouchers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chongqing&lt;/td&gt;
&lt;td&gt;"AI+" Action Plan&lt;/td&gt;
&lt;td&gt;ref # not publicly released&lt;/td&gt;
&lt;td&gt;intelligent upgrade of 33,618-manufacturing clusters&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Beijing&lt;/td&gt;
&lt;td&gt;Agent-Industry Acceleration Measures&lt;/td&gt;
&lt;td&gt;ref # not publicly released&lt;/td&gt;
&lt;td&gt;Galaxy Compute Corridor, Token vouchers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The central government set the tone in the 2026 Government Work Report: "resolve debt through development, develop through debt resolution."&lt;/p&gt;

&lt;h2&gt;
  
  
  Local Financing Platforms: From Debt Vehicles to Tech Shareholders
&lt;/h2&gt;

&lt;p&gt;As of February 2026, 1,028 LGFV entities have formally "left the platform" (public LGFV transition data). Rizhao, Chengdu's Wenjiang district, and Tangshan's Industrial Holdings (¥947M acquisition of a listed company) are all converting from infrastructure debt vehicles into holders of technology assets. Debt resolution and industrial transformation, executed as a single move.&lt;/p&gt;

&lt;h2&gt;
  
  
  Global AI Governance: Three Approaches
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;EU&lt;/strong&gt;: strictest — AI Act effective 2024.8.1, fully applicable 2026.8.2, four risk tiers, fines up to €35M or 7% of global turnover (Act text). 45 European firms publicly opposed it for "stifling innovation."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;US&lt;/strong&gt;: fragmented — states legislate separately, top five AI labs joined a voluntary evaluation system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;China&lt;/strong&gt;: law + standards dual-track, 446 generative-AI services registered (2025 public registration stats), five-tier risk classification.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Evidence Chain
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Finding&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Penn Wharton Budget Model (2025)&lt;/td&gt;
&lt;td&gt;AI could cut federal deficits $400B over 2026-2035&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BIS Working Paper No.1179&lt;/td&gt;
&lt;td&gt;AI is a positive supply shock; naturally disinflationary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;McKinsey (2025)&lt;/td&gt;
&lt;td&gt;AI cuts procurement costs ~45%, logistics ~20% (industry estimate)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goldman Sachs (2026)&lt;/td&gt;
&lt;td&gt;$7.6T global AI capex 2026-2031 (forecast)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;McKinsey China (2026)&lt;/td&gt;
&lt;td&gt;Generative AI could create ~$2T of economic value in China (estimate)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;AI is not an industry choice for China — it's a debt-survival problem. The loop is already running: central policy → local industrial policy → LGFVs convert to AI infrastructure → AI creates new tax base → old debt gets diluted → companies and individuals enjoy lower-cost dividends.&lt;/p&gt;

&lt;p&gt;This isn't a forecast. It's policy already in motion.&lt;/p&gt;

&lt;h1&gt;
  
  
  ai #china #debts #economy #industrialpolicy #aigovernance
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Agent 行为分析框架</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Sat, 10 Oct 2026 07:59:07 +0000</pubDate>
      <link>https://dev.to/maref/agent-xing-wei-fen-xi-kuang-jia-51eg</link>
      <guid>https://dev.to/maref/agent-xing-wei-fen-xi-kuang-jia-51eg</guid>
      <description>&lt;p&gt;Agent 行为分析框架&lt;/p&gt;

&lt;p&gt;侦探推理 行为日志 技术解谜 用侦探思维调试 AI 系统&lt;/p&gt;

&lt;p&gt;构建一个多Agent协作系统时，你面对的不是单个程序，而是一群各自决策、相互影响的智能体。某个Agent突然输出荒谬答案，另一个Agent在循环中卡死，整个任务链莫名其妙崩溃。常规日志堆满了token消耗和API调用次数，但看不出根因。这时候需要切换视角，把自己当成侦探。&lt;/p&gt;

&lt;p&gt;侦探推理的核心是：不轻信表面证词，而是从现场痕迹重建行为链条。AI系统的“现场痕迹”就是行为日志。但传统的逐行打印太粗糙，你需要一套专门记录Agent决策动机、工具调用上下文、记忆检索过程的结构化日志。然后像破案一样，依循线索回溯，定位故障点。&lt;/p&gt;

&lt;p&gt;技术深度解析&lt;/p&gt;

&lt;p&gt;行为日志应该包含三个层级：感知层、推理层、动作层。感知层记录Agent接收到的输入，包括用户消息、系统提示、其他Agent的输出。推理层记录思维链中的每一步：它考虑了哪些候选方案，如何评估，为什么选择某条路径。动作层记录调用的工具、传入参数、返回值、执行耗时。&lt;/p&gt;

&lt;p&gt;例如，一个负责数据分析的Agent，在生成报告时突然给出错误的季度同比。传统日志会告诉你“Agent调用了数据库函数 getyoy”，但不会告诉你它为什么调用这个函数，以及它是否错误地理解了“同比”的定义。而结构化日志会记录推理层的一条：“当前需要计算去年同一时期的销售额，因此调用 getyoy。但记忆检索显示上季度财报中‘同比’指环比，因此产生了歧义。” 这样你就能迅速发现是记忆污染导致了概念混淆。&lt;/p&gt;

&lt;p&gt;技术实现上，可以用装饰器或中间件劫持Agent的推理循环。在每个推理步骤之前，将当前上下文（包括系统提示中的角色设定、用户意图、已生成的部分推理）序列化为JSON对象。推理步骤之后，追加模型输出的下一个token或动作指令。同时记录每个工具调用的输入输出快照。日志存储采用可追溯的时间线格式，方便后续用因果图可视化。&lt;/p&gt;

&lt;p&gt;解谜过程类似犯罪现场重建。假设你收到用户投诉：一个旅游规划Agent推荐了一家已经倒闭的餐厅。你打开行为日志，时间线回退到那个决策点。感知层显示用户要求“推荐附近评分高的餐厅”。推理层记录：Agent检索了内部知识库，但知识库中餐厅状态字段缺失。然后它调用外部地图API，但API返回的营业状态标记为“Unknown”。Agent的默认策略是：当营业状态不明时，假设营业并继续推荐。这就是故障点——策略缺陷。你不需要看完整代码，只需要这条日志就能定位是“未知状态处理逻辑”出了问题。&lt;/p&gt;

&lt;p&gt;实际案例与应用&lt;/p&gt;

&lt;p&gt;在一个自动化客服多Agent系统中，有三个角色：意图识别Agent、答案检索Agent、工单生成Agent。上线后频繁出现用户被转人工后却收到自动回复的怪现象。传统监控只显示“工单生成失败”，但找不到原因。&lt;/p&gt;

&lt;p&gt;我们部署了行为日志框架。第一次回溯发现：意图识别Agent将用户抱怨“等太久了”错误归类为“查询进度”，然后答案检索Agent从FAQ中找到“等待时间预计30分钟”并返回。工单生成Agent接收到的输入是“用户已经等待30分钟，是否需要转人工？” 它判断不需要，于是直接回复了标准答案。问题出在意图识别Agent的分类模型对情绪关键词不敏感。&lt;/p&gt;

&lt;p&gt;修复后，新的日志又显示另一类异常：工单生成Agent在转人工后，仍然尝试生成自动回复，导致用户收到两条矛盾消息。回溯推理层发现：工单生成Agent的内部状态机中，“转人工”动作没有清除待回复队列。它在转人工指令发出后，继续执行了之前排队的回复任务。日志里清晰记录着：“动作：转人工，时间戳 T1；动作：发送回复，时间戳 T1+50ms”。工单生成Agent的开发者看到这条记录，立刻明白需要在转人工时清空队列。&lt;/p&gt;

&lt;p&gt;另一个案例涉及多Agent知识库冲突。一个法律咨询Agent在引用判例时，先后从两个不同来源检索到矛盾信息。日志显示推理层评估了两个候选答案的置信度，但置信度打分函数对更新时间权重过高，导致引用了过期判例。通过日志，你还能看到Agent在“为什么选择A而不是B”的推理步骤中写道：“B的源更新时间是2021年，A是2022年，因此A更可信。” 但实际上判例的效力与更新时间无关，与法院层级有关。这暴露了置信度模型的特征设计错误。&lt;/p&gt;

&lt;p&gt;总结与行动建议&lt;/p&gt;

&lt;p&gt;调试Agent系统，关键在于把黑箱决策变成可追溯的白盒过程。行为日志就是你的侦探笔记本。第一步，为每个Agent建立三层日志结构，并确保日志包含决策动机的原始输出。第二步，建立异常特征库，比如“工具调用返回空值但Agent继续执行”“推理步骤中重复出现相同内容”“不同Agent对同一事实给出矛盾解释”。第三步，训练团队用因果回溯法：从表面现象出发，逆向遍历时间线，在每个分支点检查上下文是否一致。&lt;/p&gt;

&lt;p&gt;不要等着出问题再分析。在开发阶段，对每轮测试运行自动比对：预期行为路径与实际日志路径的差异点自动高亮。这就像侦探对比嫌疑人证词和物证，能提前发现潜在漏洞。&lt;/p&gt;

&lt;p&gt;最后，日志本身也是数据。当你积累了大量行为日志后，可以训练一个小模型来自动识别异常模式。例如，检测到Agent在推理步骤中频繁输出“我不确定”但依然执行动作，可能就是置信度阈值设置过低。用行为日志喂一个分类器，它能学会识别那些人类在早期不易注意到的故障前兆。&lt;/p&gt;

&lt;p&gt;侦探思维的最终目的不是找凶手，而是理解Agent为何做出那个选择。理解了，才能改进。行为日志就是你和Agent之间坦诚对话的唯一渠道。&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>debugging</category>
      <category>llm</category>
    </item>
    <item>
      <title>治理的季节：开源复盘</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Fri, 09 Oct 2026 14:00:15 +0000</pubDate>
      <link>https://dev.to/maref/zhi-li-de-ji-jie-kai-yuan-fu-pan-3oh2</link>
      <guid>https://dev.to/maref/zhi-li-de-ji-jie-kai-yuan-fu-pan-3oh2</guid>
      <description>&lt;p&gt;Governance Season: An Open-Source Retrospective for MCP and Multi-Agent Memory&lt;/p&gt;

&lt;p&gt;&lt;code&gt;[02:03:12] sync start target=test-cluster&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;I keep that line pinned. Nothing in it looks like an emergency. The job had run every six hours for months — rsync pulling modules off a dev box into a working tree, &lt;code&gt;--delete&lt;/code&gt; on, an exclusion list deciding what survived the pass. Then someone rebuilt that exclusion list, and a handful of entries didn't come through the edit. Not mangled. Absent. rsync doesn't ask why a path is missing from the list. Anything sitting on the target that wasn't excluded and wasn't arriving from the source got removed. Seventy-odd private modules gone in one run.&lt;/p&gt;

&lt;p&gt;Getting them back was hand work: tracking down copies, diffing trees, pushing files back into place. Two hours and change of it, with people standing around a working directory nobody had thought of as load-bearing. The rebuild of the guardrails around that job took a lot longer than the recovery, and it came in pieces, across separate changes, none of which was right the first time.&lt;/p&gt;

&lt;p&gt;What actually changed how I read agent stacks was this: every part of that failure was behaving correctly. rsync did precisely what rsync does. The list was internally consistent. The source tree was untouched. There was no broken node anywhere in the path. There was a set of nodes each acting on a picture of the world that nobody had verified end to end, and the deletion was the first moment the disagreement became visible.&lt;/p&gt;

&lt;p&gt;That's the same shape I keep finding in MCP-driven multi-agent systems, and it's why open-source agent infrastructure is in governance season whether or not anyone wants it there. Unscoped MCP servers and shared long-term memory have quietly turned into incident surfaces. A single poisoned tool response, written into memory with nothing attached to say where it came from, becomes a belief that every downstream agent will act on in good faith. No agent in that chain does anything wrong. That's the problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The breakage keeps landing at the edges&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Almost every incident I reviewed this quarter traced to a boundary rather than to model quality:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MCP servers shipping tools nobody scoped, including a couple that could reach environment secrets.&lt;/li&gt;
&lt;li&gt;Agents writing tool output directly into long-term memory with no provenance attached to the write.&lt;/li&gt;
&lt;li&gt;Planner and executor agents handing the same bad retrieval back and forth until the repetition starts to look like agreement.&lt;/li&gt;
&lt;li&gt;Memory schemas that fork after a release, leaving two namespaces that each look authoritative and neither of which reconciles.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your retrospective is a count of PRs landed and releases cut, you're measuring motion. Control is a different number, and it lives in the logs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who signs for this tool&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first thing we tried was documentation. Every MCP server adds a paragraph to its README describing what its tools do and what they're allowed to touch. Nobody read them, and two of the tools whose descriptions said they only read had write paths into the filesystem. A paragraph is not a boundary.&lt;/p&gt;

&lt;p&gt;What replaced it: every server ships a signed capability file — tool names, input and output schemas, the scopes it claims, and an owner key on it. The build pipeline rejects unsigned files outright, and it rejects any tool asking for &lt;code&gt;fs.write&lt;/code&gt; or &lt;code&gt;net.egress&lt;/code&gt; where no owner is named. The signing step is where ownership stops being a vibe. If nobody will put a key on the tool, you've learned something more useful than the schema.&lt;/p&gt;

&lt;p&gt;The cost is real. Onboarding a new server went from minutes to hours, and auto-updates stop working — a version bump now requires a person to re-sign before anything ships. For a production swarm, I'll pay that. I want a name to call at 3am, and I want it to be a person, not a repository.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory you can take back&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every write into shared long-term memory now carries &lt;code&gt;agent_id&lt;/code&gt;, &lt;code&gt;tool_call_id&lt;/code&gt;, &lt;code&gt;source_hash&lt;/code&gt;, &lt;code&gt;timestamp&lt;/code&gt;, and a confidence value. Scratch memory expires on its own and nobody worries about it. Long-term writes are propose-then-commit: a second agent recomputes the source hash, and only then does the write land.&lt;/p&gt;

&lt;p&gt;Our first version was pure logging. We logged everything beautifully and had no way to use any of it. When a poisoned response surfaced later, the recovery plan was still "purge the namespace and rebuild from scratch," which is an admission that you don't know what's in there.&lt;/p&gt;

&lt;p&gt;The hash is the index now. Roll back by &lt;code&gt;source_hash&lt;/code&gt;, pull the writes that trace to one tool call, leave everything else standing. It costs an extra round trip on every long-term write, and conflicts that used to be smoothed over by last-write-wins now surface and have to be argued out. What you buy is the ability to surgically remove a belief instead of burning the store down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When two agents disagree&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Conflicting claims go to a small council: one maintainer seat, two agents from different roles. The council produces a single record, and the dissent stays in it.&lt;/p&gt;

&lt;p&gt;First attempt was arbitration by the biggest model in the room. It sided with whoever wrote the longest justification, which is a very expensive way to find the loudest voice. The lesson was to keep the losing claim visible — that's usually where the early warning was sitting.&lt;/p&gt;

&lt;p&gt;Decisions got slower. Somebody has to sit in the chair and read the disagreement instead of letting it resolve itself. That chair is the mechanism. Without it, the most confident agent overwrites the group and the group never notices.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It lives on a calendar&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Governance in open source is a schedule, not a document. Four artifacts per quarter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An incident ledger with an owner and a rollback path attached to every entry.&lt;/li&gt;
&lt;li&gt;An RFC for any memory schema change, with a 72-hour comment window — no exemption for the person who wrote the schema.&lt;/li&gt;
&lt;li&gt;Rotation of who holds the keys to the MCP registry, so no single maintainer becomes the permanent yes.&lt;/li&gt;
&lt;li&gt;A public postmortem for any memory poisoning that crossed a namespace boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We tried the soft version of this first: signing encouraged, warnings only. One release cycle later, everyone had learned to ignore the warning. The strict version holds so far, and I want to be honest that it hasn't been through a real crisis yet — the velocity cost is obvious today, the protection is theoretical until it isn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Back to that log line&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The fix that came out of the sync incident had nothing to do with the exclusion list. Every run now does a dry pass first: the same rsync invocation with &lt;code&gt;--dry-run&lt;/code&gt;, same flags, same paths. If the preview shows a file that exists on the target being removed, the run stops and the job is marked blocked. A human looks at it before anything is gone. That's the entire mechanism. It doesn't make the exclusion list correct. It makes a wrong exclusion list loud while it's still recoverable.&lt;/p&gt;

&lt;p&gt;Run your quarterly pass the same way. Before anything writes into shared long-term memory, do the dry run: what would land, which tool call it traces back to, and which entries you'd be removing by hand if that tool turned out to be compromised. This is where I land, and I've been wrong about the timing of these things before — the tool layer and the memory layer are the two places where an agent stack either has an owner or doesn't, and everything else is downstream of that.&lt;/p&gt;

&lt;p&gt;If you can't answer "who wrote this memory, from which tool, and how do we roll it back" from a log you can grep, you're not governing. You're hoping. MCP tools are signed capabilities. Multi-agent memory is an asset with a named owner. rsync &lt;code&gt;--delete&lt;/code&gt; at least tells you what it's about to remove — your agent stack won't, unless you build the thing that makes it say so.&lt;/p&gt;

&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/http%3A%2F%2Flocalhost%3A3001%2Fapi%2Fsend%3Fwebsite_id%3D30b552af-b93c-4bbb-855c-b10b45efaa52%26platform%3Ddevto" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/http%3A%2F%2Flocalhost%3A3001%2Fapi%2Fsend%3Fwebsite_id%3D30b552af-b93c-4bbb-855c-b10b45efaa52%26platform%3Ddevto" height="400" width="800" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>devtools</category>
    </item>
    <item>
      <title>陪审团升级：共识成本模型</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Wed, 07 Oct 2026 13:42:28 +0000</pubDate>
      <link>https://dev.to/maref/pei-shen-tuan-sheng-ji-gong-shi-cheng-ben-mo-xing-d9e</link>
      <guid>https://dev.to/maref/pei-shen-tuan-sheng-ji-gong-shi-cheng-ben-mo-xing-d9e</guid>
      <description>&lt;p&gt;I keep one log line in a note on my phone. Least dramatic entry I own. Also the most useful thing in there.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;[02:03:12] sync start target=test-cluster&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Here is what sat behind it. A six-hour sync job pushed a dev box into a working directory, and the script carried &lt;code&gt;rsync --delete&lt;/code&gt;. The exclude list had gone stale. A refactor moved a pile of private modules to a new path and nobody amended the file. rsync ran the arithmetic it always runs — file sitting on the target, no matching entry in the exclude list, delete — and dozens of entries disappeared in one sweep. Exit code 0. A clean bill of health from the exact process that had just finished deleting things.&lt;/p&gt;

&lt;p&gt;Recovery took a bit over two hours. Laptop copies, an old branch, a tarball a colleague had mailed himself before a conference. Most of it came back. Some of it didn't come back clean.&lt;/p&gt;

&lt;p&gt;Nothing in that pipeline crashed. Every component did what it was told to do. That's the part I keep chewing on, and it took a while to name, because nothing was actually broken.&lt;/p&gt;

&lt;p&gt;I spend my days on multi-model agent systems now, and a jury has the same silhouette as that sync script: a fault-tolerant decision circuit made of parts that all report success. If I can't trace each juror, I can't tell you what the verdict is worth. So I instrument the thing like a crime scene, and the habit traces straight back to that exclude file.&lt;/p&gt;

&lt;p&gt;When a decision hits the jury it fans out into N model calls. Each call becomes a span. I want &lt;code&gt;juror_id&lt;/code&gt;, &lt;code&gt;model_id&lt;/code&gt;, &lt;code&gt;model_family&lt;/code&gt;, &lt;code&gt;prompt_hash&lt;/code&gt;, &lt;code&gt;temperature&lt;/code&gt;, &lt;code&gt;seed&lt;/code&gt;, &lt;code&gt;token_in&lt;/code&gt;, &lt;code&gt;token_out&lt;/code&gt;, &lt;code&gt;latency&lt;/code&gt;, the raw verdict text, the self-reported confidence, and the calibration bucket that confidence lands in. The parent span carries &lt;code&gt;decision_id&lt;/code&gt;. A separate adjudication span records quorum, the individual votes, the final verdict, and the escalation reason if one fired. That's the chain of custody. Without it you have a verdict and a shrug. Somebody asks which juror was lying, and the honest answer is: no idea.&lt;/p&gt;

&lt;p&gt;Voting rules are where self-deception gets formal. Simple majority for categorical verdicts. Quorum q, two of three on a normal day. No quorum, escalate to an adjudicator or the on-call human. Weighted voting uses &lt;code&gt;w_i&lt;/code&gt; derived from holdout calibration — ignore what the model claims about itself. Self-reported confidence is a witness statement: useful, unsworn, frequently rehearsed. Track the Brier score per juror instead. I watched a model declare 0.93 confidence and then miss the same class of prompt four calls running. The confidence field was theater. The Brier score was the rap sheet.&lt;/p&gt;

&lt;p&gt;Expected cost per decision:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;C = Σ_j (c_in * tok_in_j + c_out * tok_out_j) + P_esc * C_esc + P_err * C_err
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;P_err&lt;/code&gt; is a guess. Pull it from a labeled canary set, or from whatever the last incident taught you. Treat it as ground truth and you will have a bad quarter. Two numbers matter day to day: cost per accepted decision, and cost per caught error. Re-estimate both every time the jury changes. Models get swapped, prompts drift, providers quietly change things underneath you with no changelog. A control strip saved me from shipping a bad configuration once — one cheap juror and the full jury over the same traffic sample, side by side. The expensive jury was paying for agreement. Accuracy wasn't in the receipt. I'd rather see that bill on my own dashboard than in a finance review.&lt;/p&gt;

&lt;p&gt;Tiering keeps the bill honest. Tier 0: one fast model, accept if its confidence clears the cutoff and the prompt isn't sitting in a known dissensus cluster. Tier 1: a second model from a different family, accept if the two agree. Tier 2: a third diverse model or an adjudicator, which buys coverage where error cost is high. N-of-N stays reserved for high-risk intents. I've watched teams route every decision through five models because it felt safer. It felt safe right up until the latency tail crossed the user-facing timeout and the invoice landed. Cross-validation is a tool. Don't make it a personality.&lt;/p&gt;

&lt;p&gt;Diversity is the part everyone nods at and nobody measures. Correlated errors kill you quietly. Two prompts to the same model family can agree because they share blind spots, and on a dashboard agreement looks like consensus. Log &lt;code&gt;model_family&lt;/code&gt; and provider on every span. Compute pairwise error correlation on the canary set. High correlation means the jury is decorative. Bring in a different architecture, a different retrieval path, or a tool trace that can contradict the model outright. I once watched three jurors drawn from two families agree on a wrong answer. Same blind spot, dressed up twice. That's a chorus. You wanted a jury.&lt;/p&gt;

&lt;p&gt;When a decision goes sideways I reach for: dissensus rate by intent, agreement matrix by model pair, escalation rate with reason codes, cost per decision against error rate, calibration drift per juror. Forensics first, dashboard later. Confessions are cheap. Spans hold up. If every juror agreed and the outcome was still wrong, the problem lives in the shared blind spot — different investigation, different fixes.&lt;/p&gt;

&lt;p&gt;The trade-offs don't go away. Latency tails grow with N, especially when the calls run sequentially. Parallel calls cut latency and raise peak rate. Weighted voting adds machinery, and stale weights can be gamed. Early exit saves money and weakens the cross-validation you just paid for. Three jurors catch a lot of single-model errors. Shared blind spots survive them intact. More jurors will not fix a bad prompt. Start with a 2-of-3 jury where the error cost clearly beats the jury cost. Instrument first. Tune quorum later. The verdict is a cost-bounded approximation. Trace it.&lt;/p&gt;

&lt;p&gt;Back to the sync job. That script was a jury of one. No independent witness anywhere in the chain. The exclude file was the prompt, the target was the output, and everybody agreed. The agreement deleted dozens of private modules.&lt;/p&gt;

&lt;p&gt;We chipped at it in the gaps between normal work for a while before it stopped being able to surprise us. The guard came in passes. None of them landed clean the first time.&lt;/p&gt;

&lt;p&gt;First pass: make it print &lt;code&gt;rsync --delete --stats&lt;/code&gt; and alert on deletions. Observability with no brake. The script exited 0, the alert landed in a channel, and people read it after the damage. A witness that can't stop the crime is a diary.&lt;/p&gt;

&lt;p&gt;Second pass: hash the exclude file, block whenever the hash changed. That made noise. The exclude file changes for legitimate reasons, and engineers started appending &lt;code&gt;--force&lt;/code&gt; to get their own work done. A guard everyone bypasses is a speed bump with a story attached.&lt;/p&gt;

&lt;p&gt;What finally held was the dry-run gate. &lt;code&gt;rsync --dry-run --itemize-changes --delete&lt;/code&gt; runs first. The wrapper parses the output for &lt;code&gt;*deleting&lt;/code&gt;. If any file currently on disk is about to be removed, it exits 97, writes &lt;code&gt;blocked&lt;/code&gt; into the run marker, and leaves the target untouched. The real &lt;code&gt;rsync --delete&lt;/code&gt; only runs after the dry run comes back clean. The dry-run output gets logged with run metadata in the same shape I use for &lt;code&gt;decision_id&lt;/code&gt;: intent, target, exclude file hash, who kicked it off. Now the run has a trace, the deletion has evidence, and the blocked marker carries a reason code you can grep. The dry run costs a few seconds on a tree this size. Cheap, compared to the alternative.&lt;/p&gt;

&lt;p&gt;We did try blocking every deletion for a while. That broke legitimate cleanup, so we compromised: deletions allowed only when the path sits inside the managed set and the run carries an explicit &lt;code&gt;--allow-deletions&lt;/code&gt; flag. Semi-automatic, deliberately annoying, and it has held up against our own sloppiness so far. It has not been tested against an exclude file that is actively hostile. I'd rather keep it that way.&lt;/p&gt;

&lt;p&gt;The lesson that survived: most of the nodes believed they were right. The sync believed the exclude file. The exclude file believed the last refactor. The target accepted whatever arrived. No single node was broken, and there was no independent juror anywhere in the chain.&lt;/p&gt;

&lt;p&gt;So before the next destructive sync, run the dry run first. On the actual target. With the actual exclude file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rsync &lt;span class="nt"&gt;-ani&lt;/span&gt; &lt;span class="nt"&gt;--delete&lt;/span&gt; &lt;span class="nt"&gt;--exclude-from&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/path/to/exclude /src/ /dst/ | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s1"&gt;'^\*deleting'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that prints a path you don't recognize, stop. The dry run is the evidence. The real run can wait. Build the model jury to the same standard — every juror traced, same spans, same reason codes, same refusal to accept agreement as proof. Disagreement is the evidence. The verdict is a cost-bounded approximation. Price it. Trace it. Keep the line from 02:03:12 somewhere you can still read it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/http%3A%2F%2Flocalhost%3A3001%2Fapi%2Fsend%3Fwebsite_id%3D30b552af-b93c-4bbb-855c-b10b45efaa52%26platform%3Ddevto" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/http%3A%2F%2Flocalhost%3A3001%2Fapi%2Fsend%3Fwebsite_id%3D30b552af-b93c-4bbb-855c-b10b45efaa52%26platform%3Ddevto" height="1" width="1" alt=""&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Recursive Governance: When Agents Write the Rules They Execute</title>
      <dc:creator>Open Human</dc:creator>
      <pubDate>Tue, 06 Oct 2026 08:27:52 +0000</pubDate>
      <link>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-49bh</link>
      <guid>https://dev.to/maref/recursive-governance-when-agents-write-the-rules-they-execute-49bh</guid>
      <description>&lt;p&gt;We almost lost forty modules to a file that never changed.&lt;/p&gt;

&lt;p&gt;The sync job ran every six hours, mirroring our shared working directory into a test workspace. At 21:47 on a Tuesday, it removed about forty private modules that should have been ignored. The filter file was missing one pattern. That's it. One pattern. A colleague had restructured the source tree three days earlier, added a directory, and didn't update the filter file. The script didn't warn, didn't ask, didn't fail. It just mirrored the state it was told to mirror.&lt;/p&gt;

&lt;p&gt;Restoring the modules took us almost two hours, one by one from the backup box. The filter file was a static text file. It encoded what the sync could delete and what it couldn't, and it had no idea that the real inventory had changed. No idea that a pattern was missing. No way to test a proposed pattern against the forty modules it was about to destroy.&lt;/p&gt;

&lt;p&gt;That's the same shape as most multi-agent governance failures. Rules are written at some point in time, then the world changes, and the rules stay frozen. Static policy files don't survive contact with a running system. They don't know what they're protecting.&lt;/p&gt;

&lt;p&gt;So we gave the agents the ability to write their own rules.&lt;/p&gt;

&lt;p&gt;The first version was naive. We opened a write endpoint that let any agent append to the rule store. The first three weeks were great. Agents resolved tool-access conflicts by proposing new rules, and a couple of stale patterns got archived without a human in the loop. Then it started to rot. Rules referenced other rules that had been overwritten. Rules contradicted each other. Ask two agents about the same namespace and you'd get two different answers. Three months in, the rule store had tripled in size, and nobody — not the human team, not the agents — could tell which rule was actually in effect. It grew quietly. No alert fired. No metric moved. The rot was visible in the rule history, if you were looking for it; nobody was, because we hadn't built any lifecycle tooling.&lt;/p&gt;

&lt;p&gt;That was the lesson. A rule needs a birth date, a parent, and a gravestone. If a rule can't be archived, it keeps counting toward quorum and keeps confusing the live ones. If everything accumulates forever, the rule set becomes a pile of intentions, not a description of behavior.&lt;/p&gt;

&lt;p&gt;Here's what we run now. Governance is five non-removable MCP tools in our agent runtime. &lt;code&gt;propose_rule&lt;/code&gt; submits a rule with a rationale, an owning agent, and the tool/memory namespaces it affects. &lt;code&gt;amend_rule&lt;/code&gt; submits a delta against an existing rule; deltas are diffed, never overwritten. &lt;code&gt;ratify_rule&lt;/code&gt; votes on an open proposal, with per-shard quorum. &lt;code&gt;veto_rule&lt;/code&gt; is a human's exit hatch — the justification gets written to memory so the veto itself is auditable. &lt;code&gt;check_consistency&lt;/code&gt; scans the existing rule graph, computes an embedding similarity between the new rule surface and the rule surfaces already in effect, then applies a short, human-maintained allow/deny list on tool names. A proposal that would let an agent touch a namespace it's not supposed to touch gets rejected on the spot. The list is only about thirty lines. It doesn't try to understand everything. It blocks the obvious category errors and the subtle near-misses. That last part is where the forty modules would have been saved.&lt;/p&gt;

&lt;p&gt;We tried these five as a sidecar service first. Each agent cached its own copy of the tool definitions, and for three hours two agents were enforcing different versions of the same governance rule. Now they are part of the runtime contract. If an agent can't call them, its tool loop fails closed.&lt;/p&gt;

&lt;p&gt;Every proposal goes through a two-phase commit. It enters an observation window — 24 hours or N messages, depending on the shard. During the window, affected agents can call &lt;code&gt;check_consistency&lt;/code&gt; and write rebuttals to shared memory. Ratification requires quorum, no outstanding vetoes, and a passing consistency score.&lt;/p&gt;

&lt;p&gt;One rule caused a fight early on: no agent can ratify its own proposal. A code-review shard once had an agent try to expand its own write authority from one repository to all repositories. The consistency check returned &lt;code&gt;conflict=all-repos:write&lt;/code&gt;. The proposal was amended to a single-repo grant and only then passed. If that check had been optional, the shard boundary would have dissolved without a recorded objection.&lt;/p&gt;

&lt;p&gt;Quorum took another iteration. We started with a global constant. Same action needed two votes in a three-agent shard and twenty-seven votes in a forty-agent shard, and the agents couldn't tell which law applied without reading the whole config. Per-shard quorum fixed that. The observation window is also runtime-configurable, which means a human will eventually shorten it to five minutes during a sprint deadline. Our compromise is that changing the window requires the same proposal pipeline as any other rule. Bureaucratic, yes. Fully validated under adversarial load? Not yet. I still expect a shard with a short window to be the one where someone tries to push a poison rule through.&lt;/p&gt;

&lt;p&gt;The rule set lives as structured memory events in an append-only namespace: &lt;code&gt;governance://rules/&lt;/code&gt;. Governance is just the highest-priority namespace in the same memory system agents already share. Each entry references its parent rule and the proposing agent, so any agent can replay the exact rule state at any past timestamp. The failure mode: corrupted or wrongly pruned governance memory makes agents enforce a rule set that never existed. We mitigate by storing rule hashes in cold storage and requiring a quorum to restore from it. A single bad prune can't silently rewrite history.&lt;/p&gt;

&lt;p&gt;If the sync job had this system, the filter file would have been a rule in &lt;code&gt;governance://rules/sync-filters/&lt;/code&gt;. A change to it would have had to pass a consistency check against the actual disk inventory. A proposal that added a wildcard covering &lt;code&gt;private-modules/&lt;/code&gt; would have contradicted the existing "do not delete these" state, and it would have been killed before it became policy. We built a crude version of that idea later anyway: a preflight pass that runs the sync read-only and aborts if it sees a deletion target for a file that currently exists. It caught the next two near-misses. But it can't give us back those two hours.&lt;/p&gt;

&lt;p&gt;If you give agents the power to make rules, you also have to give them the power to retire them. Otherwise, three months from now, you won't know which rules are alive and which are zombies. Add a cleanup day. Sit down every few months, replay the governance history, archive rules that reference forgotten namespaces, re-ratify the ones that still matter. Even the rules the system wrote for itself need an expiration date. The dead rules are the ones that will eventually bite you, and they are the ones that never send a log line before they strike.&lt;/p&gt;

&lt;h1&gt;
  
  
  governance #ai #agents #infrastructure
&lt;/h1&gt;

&lt;h1&gt;
  
  
  maref #ai #opensource #machinelearning
&lt;/h1&gt;

</description>
      <category>automation</category>
      <category>devops</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
