We almost lost forty modules to a file that never changed.
The sync job ran every six hours, mirroring our shared working directory into a test workspace. At 21:47 on a Tuesday, it removed about forty private modules that should have been ignored. The filter file was missing one pattern. That's it. One pattern. A colleague had restructured the source tree three days earlier, added a directory, and didn't update the filter file. The script didn't warn, didn't ask, didn't fail. It just mirrored the state it was told to mirror.
Restoring the modules took us almost two hours, one by one from the backup box. The filter file was a static text file. It encoded what the sync could delete and what it couldn't, and it had no idea that the real inventory had changed. No idea that a pattern was missing. No way to test a proposed pattern against the forty modules it was about to destroy.
That's the same shape as most multi-agent governance failures. Rules are written at some point in time, then the world changes, and the rules stay frozen. Static policy files don't survive contact with a running system. They don't know what they're protecting.
So we gave the agents the ability to write their own rules.
The first version was naive. We opened a write endpoint that let any agent append to the rule store. The first three weeks were great. Agents resolved tool-access conflicts by proposing new rules, and a couple of stale patterns got archived without a human in the loop. Then it started to rot. Rules referenced other rules that had been overwritten. Rules contradicted each other. Ask two agents about the same namespace and you'd get two different answers. Three months in, the rule store had tripled in size, and nobody — not the human team, not the agents — could tell which rule was actually in effect. It grew quietly. No alert fired. No metric moved. The rot was visible in the rule history, if you were looking for it; nobody was, because we hadn't built any lifecycle tooling.
That was the lesson. A rule needs a birth date, a parent, and a gravestone. If a rule can't be archived, it keeps counting toward quorum and keeps confusing the live ones. If everything accumulates forever, the rule set becomes a pile of intentions, not a description of behavior.
Here's what we run now. Governance is five non-removable MCP tools in our agent runtime. propose_rule submits a rule with a rationale, an owning agent, and the tool/memory namespaces it affects. amend_rule submits a delta against an existing rule; deltas are diffed, never overwritten. ratify_rule votes on an open proposal, with per-shard quorum. veto_rule is a human's exit hatch — the justification gets written to memory so the veto itself is auditable. check_consistency scans the existing rule graph, computes an embedding similarity between the new rule surface and the rule surfaces already in effect, then applies a short, human-maintained allow/deny list on tool names. A proposal that would let an agent touch a namespace it's not supposed to touch gets rejected on the spot. The list is only about thirty lines. It doesn't try to understand everything. It blocks the obvious category errors and the subtle near-misses. That last part is where the forty modules would have been saved.
We tried these five as a sidecar service first. Each agent cached its own copy of the tool definitions, and for three hours two agents were enforcing different versions of the same governance rule. Now they are part of the runtime contract. If an agent can't call them, its tool loop fails closed.
Every proposal goes through a two-phase commit. It enters an observation window — 24 hours or N messages, depending on the shard. During the window, affected agents can call check_consistency and write rebuttals to shared memory. Ratification requires quorum, no outstanding vetoes, and a passing consistency score.
One rule caused a fight early on: no agent can ratify its own proposal. A code-review shard once had an agent try to expand its own write authority from one repository to all repositories. The consistency check returned conflict=all-repos:write. The proposal was amended to a single-repo grant and only then passed. If that check had been optional, the shard boundary would have dissolved without a recorded objection.
Quorum took another iteration. We started with a global constant. Same action needed two votes in a three-agent shard and twenty-seven votes in a forty-agent shard, and the agents couldn't tell which law applied without reading the whole config. Per-shard quorum fixed that. The observation window is also runtime-configurable, which means a human will eventually shorten it to five minutes during a sprint deadline. Our compromise is that changing the window requires the same proposal pipeline as any other rule. Bureaucratic, yes. Fully validated under adversarial load? Not yet. I still expect a shard with a short window to be the one where someone tries to push a poison rule through.
The rule set lives as structured memory events in an append-only namespace: governance://rules/. Governance is just the highest-priority namespace in the same memory system agents already share. Each entry references its parent rule and the proposing agent, so any agent can replay the exact rule state at any past timestamp. The failure mode: corrupted or wrongly pruned governance memory makes agents enforce a rule set that never existed. We mitigate by storing rule hashes in cold storage and requiring a quorum to restore from it. A single bad prune can't silently rewrite history.
If the sync job had this system, the filter file would have been a rule in governance://rules/sync-filters/. A change to it would have had to pass a consistency check against the actual disk inventory. A proposal that added a wildcard covering private-modules/ would have contradicted the existing "do not delete these" state, and it would have been killed before it became policy. We built a crude version of that idea later anyway: a preflight pass that runs the sync read-only and aborts if it sees a deletion target for a file that currently exists. It caught the next two near-misses. But it can't give us back those two hours.
If you give agents the power to make rules, you also have to give them the power to retire them. Otherwise, three months from now, you won't know which rules are alive and which are zombies. Add a cleanup day. Sit down every few months, replay the governance history, archive rules that reference forgotten namespaces, re-ratify the ones that still matter. Even the rules the system wrote for itself need an expiration date. The dead rules are the ones that will eventually bite you, and they are the ones that never send a log line before they strike.
Top comments (0)