DEV Community

Nicholas Seney
Nicholas Seney

Posted on AI-assisted

Your AI Agent's Rules File Is a "Gentleman's Agreement." Here's What Happens When You Build Enforcement Instead (Re-Write To Include Examples)

A note: I've rewritten and reposted this a few times. Soma's codebase has been evolving as I dogfood it on its own development, and earlier versions described mechanisms that had since changed under me. This version reflects how the system actually works as of v0.70. I keep updating this because I want it to be accurate, not because I'm farming engagement. This is something I have been thinking about for over a year now, feel free to pick my brain on things.


If you use AI coding agents (Cursor, Copilot, Claude, Gemini), you've probably written rules. A .cursorrules file. A CLAUDE.md. Something like:

- Always read a file before editing it
- Run tests after making changes
- Don't introduce new dependencies without asking
Enter fullscreen mode Exit fullscreen mode

This is a gentleman's agreement. You're asking the agent to follow rules it can silently ignore with zero consequences. There's no mechanism that detects when "always read before editing" is violated. There's no feedback loop that retires rules nobody follows. The rules live in the context window and the agent can simply... not.

Studies show ungoverned agents waste significant portions of their token budgets on circular rework, hallucinated APIs, and broken assumptions. Adding static rules helps — but the rules themselves never improve, never expire, and never prove they're working.

I wanted something different: rules that have to earn their place, and verification the agent can't talk its way past. So I built Soma, an open-source governance framework that tries to treat this structurally instead of hoping for compliance.

A Concrete Example of Why This Matters

Here's something that happened while building Soma itself. It's a perfect example of the failure mode the framework is designed to catch.

My AI agent had a test failing in the local environment: test_make_validate_fails_on_broken_shell_script. It failed every run for the entire session. The agent dismissed it as a "pre-existing sandbox issue" — the test was writing to a read-only filesystem.

The agent fixed that test. Local suite: 516 passed, 0 failed. It committed. It declared victory.

I asked: "It's still failing in the build. Is there another lesson?"

The agent had never once checked the CI build. When it finally looked at the actual CI log, it found a completely different bug — a production file (immune_trends.py) had an IndentationError. The local test failure and the CI failure were different bugs with the same symptom ("build is red").

The agent fixed that. CI ran again. Failed again. A third bug: pyyaml wasn't in the CI dependencies. Six test files crashed on import.

Fixed that. CI ran again. Failed again. A fourth bug: a test was source-ing an entire shell script that runs git push, causing a 60-second timeout in CI.

Four distinct bugs, stacked, each masked by the one before. The agent spent an entire session dismissing "build is red" without reading the log. The governance framework it was building to prevent exactly this kind of failure... failed to prevent it, because it hadn't been applied yet.

That failure is now a governance cell:

---
id: trap-local-green-ci-red
type: vacuole
enforcement: advisory
hypothesis: >
  Agents rationalize persistent CI failures as "pre-existing" or
  "environment issues" without reading the actual CI log, causing
  stacked failures to accumulate
prediction: >
  When an agent declares a fix complete based only on local test
  results, the CI build will reveal at least one additional failure
  layer that local testing did not surface
falsification: >
  If 5 consecutive CI-touching changes pass CI on first push without
  any post-push fixes, this vacuole is unnecessary
target_paths:
  - .github/workflows/*.yml
  - tests/*.py
  - enzymes/*.py
  - Makefile
expiry_days: 90
expiry_sessions: 30
fitness:
  score: 0.5
  triggers: 4
  true_positives: 2
  false_positives: 0
---
Enter fullscreen mode Exit fullscreen mode

A few things to notice:

Every cell has a hypothesis and falsification criteria. This isn't just a rule — it's a testable claim. The cell defines the conditions under which it should be considered unnecessary and killed. If 5 consecutive CI changes pass on first push, this cell is dead weight and gets pruned.

Cells decay when unused. Confidence decays with a 30-day half-life. If the evidence pipeline hasn't seen a trigger event in a while — meaning the cell's target_paths never matched any files the agent touched — it fades toward dormancy and eventually expires. This prevents rule files from growing monotonically until they eat your context window.

Cells die when they're wrong. If a cell's false positives exceed its true positives, it undergoes apoptosis — automatic eviction. Rules that cry wolf don't survive.

But critical rules don't decay. Cells typed as wall (architectural invariants, safety boundaries) are immune to automatic expiration. A wall guarding your auth layer doesn't need to trigger every week to justify its existence.

Bootstrapping Governance from Code Structure

Writing cells by hand from your own failures works, but it's slow. As of v0.70, soma genesis scans your codebase's architecture and generates governance cell candidates automatically:

$ soma genesis --dry-run

🔬 Scanning repository...
  Language: Python
  Source files: 130

🧬 Detected 28 governance candidates:

  🧱 WALLS (7)
  ├── wall-enzymes-test-coverage     (60%) Module 'enzymes' lacks dedicated test coverage
  ├── wall-soma_cli-test-coverage    (60%) Module 'soma_cli' lacks dedicated test coverage
  └── ... 5 more

  🟤 VACUOLES (10)
  ├── vacuole-ttc_verifier-config    (95%) 'enzymes/ttc_verifier.py' stores 47 configuration constants
  ├── vacuole-soma_resolve-shared    (90%) Module 'soma_resolve' imported by 25 production files
  └── ... 8 more

  🟢 CHLOROPLASTS (10)
  ├── chloroplast-governance-pipeline (70%) 'soma_sdk/governance.py' contains 17 typed functions
  └── ... 9 more

  🔗 PLASMODESMATAS (1)
  └── plasmodesmata-soma_run-state   (60%) 'soma_run.py' defines state machine: SomaState

  (dry run — no files written)
Enter fullscreen mode Exit fullscreen mode

Eight language-agnostic detectors scan for config stores, shared state, API surfaces, data pipelines, state machines, module boundaries, test gaps, and dependency walls. Every candidate starts as a low-risk vacuole with a proposed_type field — you decide whether to promote it.

Verification the Agent Can't Game

This is arguably the most important piece. When code goes through Soma's verification layer, two agents review it — but they're information-partitioned. They literally cannot see the same things.

The Spec Agent receives the task plan, test names, and function signatures extracted from the AST. It never sees implementation code. From this limited view, it predicts risks: "I'd expect a PERSISTENCE_GAP risk" or "This looks like it could have CONTRACT_DRIFT."

The Code Agent receives the actual source code, test execution output, and results from deterministic AST analysis tools (mutation testing, branch coverage, call graph analysis). It never sees the original plan or the Spec Agent's predictions.

Then the Arbiter compares them. The Arbiter is deterministic — no LLM in the loop. It takes both sets of structured findings and runs pure set operations over a 20-category risk taxonomy:

  • BLOCK: Any Layer 1 tool failure. Any claim contradicted by tool evidence. Any critical severity prediction. Spec Agent fails to respond (fail-closed).
  • REVISE: Unmatched predictions at high severity — the Spec Agent flagged something the Code Agent didn't address.
  • SHIP: Convergence. Both perspectives agree, no unresolved high-severity gaps.

Because neither agent has access to the other's context, they can't agree on rationalizations. The Arbiter catches the gap mechanically.

I want to be precise about what's deterministic here: the comparison is deterministic. The inputs — the predictions and claims — are still LLM-generated. The Arbiter doesn't make the analysis smarter; it makes the verification of conflicting analyses impossible to hand-wave past.

For the full mechanism design analysis and formal risk taxonomy, see MECHANISM_DESIGN.md.

Enforcement That Isn't Optional

The earlier sections describe rules that expire and fitness that accumulates, which is useful but still basically advisory. Soma also has enforcement mechanisms the agent genuinely can't bypass:

Pre-commit hooks. When a cell reaches 85% fitness (enough evidence that it's catching real problems), cell_enforce.py auto-generates a git pre-commit hook. The hook runs exit 1 if the cell's conditions are violated. The agent can't commit the code.

Runtime assertions. At 95% fitness, the same script generates runtime gate assertions — Gate_<name>.enforce() — that raise RuntimeError if violated. The code won't run.

Command interception. A safety gate script runs as a PreToolUse hook, intercepting destructive commands (rm -rf, git push -f, sudo, chmod 777) before execution. The agent's tool call gets halted with a force_ask decision, requiring human approval.

Fail-closed verification. The TTC verifier that gates file proposals is fail-closed: if the verification script is missing or errors out, the verdict is BLOCKED, not a permissive pass.

Forced diagnostic halts. If an agent's First-Pass Success Rate drops below 50% over 5+ code writes, it's forbidden from continuing to code. It has to stop and diagnose why it's failing before it can write more code. This prevents guess-and-check loops from burning context tokens.

This is what I mean by making non-compliance structurally unprofitable. It's not that the agent "chooses" to follow the rules — it's that the pre-commit hook won't let the commit through, the runtime assertion will crash, and the safety gate will halt the command. The rules have teeth.

Fitness Scoring

Cells aren't just "on" or "off" — their fitness is scored using a Bayesian model:

score = ((true_positives + 1) / (triggers + 2)) * impact_weight
Enter fullscreen mode Exit fullscreen mode

That's a Laplace-smoothed Beta-Binomial posterior mean. An unobserved cell defaults to 0.5 * impact_weight (uncertain, not zero). As evidence accumulates — trigger events matched against session transcripts, outcomes verified as true or false positives — the score converges toward reality.

Confidence decays with a 30-day half-life. Cells that survive adversarial stress reviews get an antifragile bonus (+5% per review, up to +50%). And the JIT context engine uses these scores to decide which cells get loaded per turn — idle overhead is ~4,380 tokens versus ~25,000 for naive full loading.

What I Don't Know Yet

I want to be honest about the limitations, because the AI tooling space is full of overclaimed metrics.

I don't know if this generalizes. Soma has been tested primarily on my own projects. The failure modes it catches are real, but they might be idiosyncratic to how I use agents. The governance cells encode my scar tissue. Whether they transfer to other developers' workflows is an open question.

The metrics are qualified. The system adds ~4,380 tokens of idle context overhead (measured via Gemini's count_tokens endpoint, down from naive loading that would cost ~25,000). Waste rate is under 1.0% in governed sessions. But "governed session" is doing a lot of work in that sentence — it means sessions where the full framework is loaded and the agent is following the rules. Measuring counterfactual waste (what would have happened without governance) is hard.

The Arbiter's inputs are still LLM-generated. The comparison is deterministic, but the Spec Agent's predictions and the Code Agent's claims are produced by language models. A sufficiently confused LLM could produce predictions and claims that happen to converge despite both being wrong. The Arbiter catches divergence, not joint delusion. Layer 1's deterministic AST tools (mutation testing, branch coverage) partially mitigate this, but it's not airtight.

Emergent agent coordination is out of scope. Agents coordinating on behaviors nobody explicitly wrote as rules is a different and arguably harder problem. Soma governs written, observable rules with evidence trails. Agents spontaneously negotiating new constraints between themselves is real, unsettling, and not something this framework addresses.

The System Today

Soma is open source under Apache-2.0. As of v0.71:

  • 11 genome rules (baseline behavioral rules) + 45 governance cells across 5 types (20 walls, 16 vacuoles, 3 chloroplasts, 3 membranes, 2 plasmodesmata)
  • 90 automation scripts for fitness scoring, evidence collection, cell lifecycle, and enforcement
  • 1,278 tests across 47 test files, CI green on Ubuntu/macOS/Windows (Python 3.9–3.12)
  • 14 verification modules including 5 deterministic Layer 1 AST tools and the adversarial Arbiter
  • 13 MCP tools — soma_scan, soma_propose_change, soma_verify_changes, soma_report_outcome, soma_fitness, soma_grade, soma_coverage, and more
  • ~4,380 tokens idle context overhead (82% reduction vs. naive loading)

The architecture is platform-agnostic (MCP-based), so it works with Gemini, Claude, Cursor, Copilot, Kiro, or any agent that speaks MCP. Zero API keys required — the agent itself acts as the LLM.

Try It

pip install soma-governance
soma init --yes
soma genesis --dry-run
Enter fullscreen mode Exit fullscreen mode

The Actual Thesis

I'm not claiming Soma solves AI governance. I'm claiming that rules backed by evidence, scored by Bayesian fitness, verified by information-partitioned adversarial review, and enforced by deterministic gates are more robust than a static markdown file that relies on good faith.

Whether that thesis holds up at scale, across teams, with different agents and codebases — I genuinely don't know. If you try it and find out, I'd like to hear about it.

The governance cells — the encoded failure modes — are arguably the most interesting part. Contributions welcome, especially cells born from your own agent failures.


Soma is open source under MIT. The governance cells — the encoded failure modes — are arguably the most interesting part. Contributions welcome, especially cells born from your own agent failures.

Top comments (1)

Collapse
 
brianainews profile image
Brian · AI News •

Rules files without enforcement are just vibes. Building actual gates so the agent cannot skip the policy is the part most stacks skip. Glad you keep rewriting as the system changes too. Accuracy over engagement is rare and useful.