DEV Community

Nicholas Seney
Nicholas Seney

Posted on AI-assisted

Rules That Can't Prove Themselves Die: Adaptive Governance for AI Coding Agents

Darwinian Governance for AI Coding Agents: How Prism AI Steering Reduces Wasted Computation by 10× Through Adaptive Rule Evolution

Nick Seney · September 2026

Prism AI Steering is an open-source adaptive governance framework that generates, measures, and evolves its own rules based on observed agent behavior. Over 17 governed sessions and 7,015 analyzed steps, it reduced agent waste from 56% to under 5% — while adding only ~3.4% idle context overhead.


1. Introduction

Large language model (LLM) coding agents — GitHub Copilot, Cursor, Amazon Q, Google's Gemini in Antigravity — are transforming software engineering. Yet a critical problem remains largely unaddressed: ungoverned agents waste extraordinary amounts of computation on circular rework, hallucinated APIs, and broken assumptions.

Static rule files (.cursorrules, CLAUDE.md, .github/copilot-instructions.md) emerged as the community's first attempt at governance. These files inject behavioral constraints into the agent's system prompt. They help. But they share a fundamental flaw: they never adapt. A rule that prevents hallucinated hotkeys in a game development project is dead weight in a REST API codebase. A rule that enforces virtual environment hygiene is irrelevant in a Go monorepo.

This paper introduces Prism AI Steering, an open-source framework that treats agent governance as an evolutionary system. Rules are not written once and forgotten — they are generated from observed failures, measured against empirical waste metrics, promoted or demoted based on effectiveness, and specialized per-repository through a biological cell metaphor. The result is a governance system that improves itself over time.

1.1 Contributions

  1. Empirical waste taxonomy: A 23-pattern classification of AI agent failure modes, derived from post-mortem analysis of 7,015 agent steps across 17 sessions.
  2. Darwinian rule lifecycle: A formal protocol where rules are born from observed failures, measured against First-Pass Success Rate (FPSR), and retired if they stop proving effective.
  3. Three-layer biological architecture: A Biome → Forest Floor → Cell hierarchy that separates global operational modes, composable analytical prongs, and repository-specific adaptive invariants.
  4. Adaptive review orchestration: A self-escalating code review protocol that autonomously selects review depth based on finding severity, achieving NASA IV&V-inspired adversarial falsification with bounded computational cost.
  5. Conditional loading optimization: A token economics strategy that reduces per-turn governance overhead by 82.6% (from 25,163 to 4,380 tokens) while preserving full coverage.

2. The Problem: Quantifying Agent Waste

2.1 Baseline Measurement

The foundational observation behind Prism came from a 1,339-step post-mortem on a real coding session (a Dwarf Fortress reinforcement learning agent). Rather than speculating about what rules might be useful, we analyzed what actually went wrong:

Failure Mode Steps Wasted % of Session
Hallucinated game hotkeys (agent invented controls without verifying) 180 13.4%
Virtual environment confusion (host Python vs. project venv) 110 8.2%
Infinite polling loop in tests (no timeout on hardware boundary) 103 7.7%
Method shadowing (action_masks() silently overridden) 50 3.7%
Stale references after rename (no global grep) 30 2.2%
Total waste 749 55.9%

A second independent session (TAB AI, 1,551 steps) confirmed many of the same patterns — hallucinated controls, venv confusion, zero subagent delegation — ruling out project-specific bias.

A third session (Dwarf Fortress, 1,579 steps) was the most catastrophic: 8 architectural pivots, 2 safety incidents (the agent clicked through a Steam dialog), zero subagent delegation across 17 compactions, and an estimated 8 million tokens consumed with zero working autonomous runs produced.

2.2 Cross-Session Pattern Analysis

Analyzing all sessions revealed which failure patterns are systemic versus project-specific:

Pattern Session A (1,579 steps) Session B (1,551 steps) Session C (856 steps)
Hallucinated controls ✅ 160 steps ✅ 180 steps —
Venv/path confusion ✅ ✅ —
Zero subagent delegation ✅ (0 subagents, 17 compactions) Improved Heavy use
Bulk media context pollution — ✅ (127 files in brain) —
Desktop automation safety breach ✅ (clicked through Steam) — —
Premature framework extraction ✅ (150 steps on unused SDK) — —

This cross-session analysis established that agent waste is not random — it follows predictable, classifiable patterns that can be targeted by specific governance rules.


3. Architecture

Prism's governance architecture follows a three-layer biological hierarchy, each operating at a different scope and granularity.

3.1 Layer 1: Biome (Global Operational Modes)

The Biome layer defines Review Modes — globally applicable protocols that determine the rigor of code review. These range from lightweight targeted fixes to full adversarial assurance:

Mode Dispatches Token Budget Use Case
🌱 Breeze 2 Flash ~3–4k Known defects, renames, typo fixes
🌬️ Gale 3–4 Flash ~4k Quick feature reviews
🔱 Trident 5–8 Flash ~8–12k Features, refactors
🌊 Maelstrom 7–12 Flash ~15–20k Architecture changes, security
⛈️ Tempest 8–12 Flash ~30–50k Catastrophic risk, infrastructure

The key design decision is that mode selection is not manual. An escalation sentinel, implemented as a zero-token bash script (scripts/escalation_sentinel.sh), classifies modified files by sensitivity (HIGH: scripts/infra/auth, MEDIUM: rules/code, LOW: docs/README) and recommends a minimum protocol level at the PreInvocation lifecycle hook.

3.2 Layer 2: Forest Floor (Composable Analytical Prongs)

Each Review Mode composes a subset of six Review Prongs — modular analytical passes that can be combined in different configurations:

Prong Metaphor Function Output
🍄 Spores Fungal spore dispersal Width-first heuristic survey; scouts for problems Problem inventory
🍄 Mycelium Underground fungal network Blast-radius mapping; 2-hop import chain analysis Dependency impact assessment
🌿 Roots Deep root system Root-cause analysis with concrete fix proposals Fix specifications
🌹 Thorns Defensive thorns Adversarial falsification (NASA IV&V-inspired) Attack vectors against proposed fixes
🪨 Bedrock Geological foundation Structural verification gate Binary SHIP/BLOCK verdict
🍂 Mulch Decomposition cycle Post-review learning extraction Taxonomy updates, governance proposals

This design draws on NASA's Independent Verification & Validation (IV&V) tripartite falsification methodology: independent reviewers with conflicting analytical incentives attempt to falsify each other's conclusions before a final structural gate.

Empirical Refutation Gate: Before accepting any critical finding from Spores or Thorns, the system requires 2-of-3 evidentiary criteria to be met. This eliminates approximately 80% of false positives (empirically observed across governed sessions), saving 15–30 steps per avoided phantom bug chase.

3.3 Layer 3: Cell (Repository-Local Adaptive Invariants)

The Cell layer is where Prism departs most significantly from static governance systems. Instead of shipping a fixed set of rules, Prism generates repository-specific governance cells through a process called Cytogenesis (the fifth stage of the Genesis onboarding skill).

Five cell types are generated:

Cell Biological Analog Function Example
Vacuoles Storage organelles Repository-specific traps and known pitfalls "This codebase uses SQLAlchemy 1.4 — do not use 2.0-style select() syntax"
Cell Walls Structural boundaries Architectural boundaries that must not be crossed "Never import from internal/ in api/ handlers"
Membranes Selective permeability Override rules for project-specific exceptions "Allow eval() in the template engine (contrary to global security rule)"
Chloroplasts Energy producers Persona definitions for specialized subagents "Use 'SRE Debugger' persona when analyzing Kubernetes manifests"
Plasmodesmata Cell-to-cell channels Cross-service connection mappings "The users-service event bus feeds notifications-service via Kafka topic user.events"

Cells are generated automatically by analyzing the repository's structure, dependencies, git history, and conventions. They are ephemeral — when the codebase changes, cells can be regenerated through a subsequent Genesis run that diffs against the previous report.


4. Rule Evolution: The Darwinian Lifecycle

4.1 Birth: Rules From Observed Failures

Every rule in Prism traces its origin to a specific, documented failure. The mapping is explicit:

Observed Failure (steps wasted) Rule Created
Hallucinated hotkeys (180 steps) UI Grounding Gate — verify via screenshots, never guess
Venv confusion (110 steps) Makefile Venv Guard — bind $(VENV)/bin/python explicitly
Infinite test polling (103 steps) Mock-First Mandate — hardware boundaries always mocked
Method shadowing (50 steps) Symbol Collision Guard — grep before defining in large files
Stale references after rename (30 steps) Refactoring Sweep — global grep confirms zero orphans

This is the core methodological claim: rules derived from observed waste outperform rules derived from assumed best practices, because they target the failure modes that actually cost tokens in practice.

4.2 Measurement: First-Pass Success Rate (FPSR)

The primary quality metric is First-Pass Success Rate (FPSR) — the percentage of code changes that pass review (Bedrock SHIP) on the first attempt without requiring rework.

  • Target: >80%
  • Halt threshold: <50% over N≥5 consecutive reviews
  • Calibration benchmark: Tencent's SiriusDeliver system reported 87.2% FPSR in 2026 industrial deployment

FPSR is tracked per-rule: if a rule's post-adoption waste rate equals or exceeds its pre-adoption rate, the rule is flagged INEFFECTIVE. If post-adoption waste drops to <30% of pre-adoption levels, the rule is flagged EFFECTIVE.

4.3 Death: Ineffective Rules Are Retired

Session analyzed → waste classified by 23-pattern taxonomy (v2.1) → metrics stored
    ↓
effectiveness.json updated with before/after waste rates
    ↓
IF post_adoption_rate >= pre_adoption_rate → rule flagged INEFFECTIVE
IF post_adoption_rate < 30% of pre → rule flagged EFFECTIVE
IF new pattern with no rule → flagged UNCOVERED, priority = wasted steps
Enter fullscreen mode Exit fullscreen mode

This continuous governance loop caught 7 additional findings that manual review missed, including the insight that copying files into Antigravity's brain scratch directory pollutes checkpoint metadata.

As of the latest audit: 2 rules have earned EFFECTIVE verdicts (providence §11, §8), and 3 rules were flagged INEFFECTIVE (requiring mechanical enforcement rather than prompt-based instruction).


5. Token Economics

A governance system that prevents waste but costs more than the waste it prevents is counterproductive. Prism's token economics are designed around conditional loading — a strategy that keeps the idle overhead minimal while preserving full coverage on demand.

5.1 Per-Turn Cost Structure

Component Naive Loading (all rules + skills) Optimized (conditional) Savings
Always-on rules (3) 3,743 3,743 0
Conditional rules (8) 3,579 248 (idle headers only) 93.1%
Skills (15) 17,841 387 (idle headers only) 97.8%
Total per turn 25,163 4,380 82.6%

The 4,380 tokens of idle overhead represents approximately 3.4% of a typical agent turn (estimated 128k context window with ~40k active content). This is the cost of governance — the price of preventing 56% waste.

5.2 Multi-Tier Cost Model

Prism operates a two-tier cost model:

  • Tier 1 (Parent Orchestrator): Carries full rules + conversation history. Expensive over long sessions but handles decision-making.
  • Tier 2 (Ephemeral Subagents): ~200 token pre-seeding via Context Pre-Seeding. No parent rules loaded, zero conversation history. Highly disposable.

This separation means that review prongs (Spores, Roots, Thorns, Bedrock) operate as lightweight Flash-tier subagents without inheriting the parent's full governance context — dramatically reducing the per-dispatch cost.

5.3 Automation Savings

Mechanism Cost Steps Prevented
Auto-Preflight (session-preflight) 0 tokens (bash) + ~881 if Flash probe dispatched 40–120 steps on venv/git drift
Context Pre-Seeding ~200 tokens per dispatch 2–3 exploratory cold-start steps per subagent
Governance Sweep ~500 tokens per run (local Python, zero LLM cost) Continuous compliance without model invocation
Escalation Sentinel 0 tokens (bash) Prevents under-reviewed high-risk changes

6. Empirical Results

6.1 Aggregate Telemetry (17 Sessions, 7,015 Steps)

Metric Value
Total steps analyzed 7,015
Sessions analyzed 17 (8 deep, 9 sweep)
Overall waste rate (governed) 18.8% (1,317 / 7,015 steps)
Best-case waste rate 1.1%
Waste patterns tracked 23 (taxonomy v2.1)
Rule sections monitored 13
Rules with EFFECTIVE verdict 2
Rules with INEFFECTIVE verdict 3

6.2 Top Waste Sources Under Governance

Even under governance, waste is not zero. The remaining waste concentrates in predictable categories:

Rank Pattern Steps Wasted % of Governed Waste
1 Desktop automation guessing 255 19%
2 Rework loops (no read-before-write) 230 18%
3 Scope inversion + micro-prototyping 222 17%
4 Environment blindness / venv drift 120+ ~9%

These residual waste sources represent the frontier of what prompt-based governance can achieve. The three INEFFECTIVE rules target patterns where mechanical enforcement (bash scripts, git hooks) outperforms prompt-based instruction — a finding that itself validates the Darwinian lifecycle: rules that can't prove themselves in prompt form are retired in favor of mechanical alternatives.

6.3 Maelstrom Dogfooding

The first operational Maelstrom review (Phase 6) caught 5 critical bugs that would have shipped broken under traditional single-pass review. The Thorns adversarial prong specifically identified:

  • 3 broken fixes that passed Roots analysis
  • 5 weakened fixes with incomplete edge case coverage
  • 0 of 4 Thorns attack vectors survived to Bedrock — all were addressed in the revision cycle

6.4 Comparison to Baseline

Metric Ungoverned Baseline Prism-Governed Improvement
Waste rate 25–56% <5% (best case 1.1%) 5–50× reduction
Subagent utilization 0 dispatches (worst case) Heavy delegation ∞ improvement
Context compactions 17 (worst case session) Minimal Dramatic reduction
Idle governance overhead 0 ~3.4% of context Acceptable cost

7. Related Work

7.1 Static Rule Files

The .cursorrules, CLAUDE.md, and .github/copilot-instructions.md formats represent the state of practice for agent governance. These are flat text files injected into the system prompt. They share three limitations that Prism addresses:

  1. No adaptation: Rules are written once and never measured. There is no mechanism to determine whether a rule is actually preventing waste.
  2. No repository specialization: The same rules apply to every project. A game development governance file is applied identically to a REST API codebase.
  3. No review protocol: Static files govern how the agent writes code but not how the agent reviews code. There is no escalation, no adversarial falsification, no structural verification gate.

7.2 Multi-Agent Code Review Systems

Recent work on multi-agent review systems (Tencent's SiriusDeliver [2026], MIT's multi-agent collaboration benchmarks, Coasty's concurrent agent framework) has established the viability of dispatching multiple LLM agents for code review. Prism's contribution is the composable prong architecture — rather than dispatching N identical reviewers, each prong has a distinct analytical persona and operates under an empirical refutation gate.

The Orthogonal Persona Mandate (drawing from MAR/ICML 2026) bans homogeneous reviewer fan-outs and mandates conflicting analytical incentives across prongs. This directly addresses the consensus bias failure mode observed in naive multi-agent systems where identical agents converge on the same false positive.

7.3 Structured Diagnostic Protocols

The Diagnose Before Repair schema (adapted from REFLEX/ICML 2026) requires a structured diagnostic artifact (failure_mode, root_cause, broken_invariant, fix_spec) before any fix code is written. This addresses the symptom-patching loop — a pattern where agents apply surface-level fixes without understanding the root cause, leading to cascading rework.


8. Limitations and Future Work

8.1 Current Limitations

  1. Self-reported metrics: Waste classification is performed by the same agent system being governed. Independent human annotation of waste patterns would strengthen the empirical claims.
  2. Platform coupling: While Prism supports multiple platforms (Gemini/Antigravity, AWS Kiro, GitHub Copilot), the deepest testing has been on Gemini. Cross-platform validation at scale is ongoing.
  3. Sample size: 17 governed sessions (7,015 steps) is sufficient for pattern identification but insufficient for statistical significance testing of individual rule effectiveness. The Phase 11 expanded dataset (66 sessions) addresses this partially but uses masked alternative models.
  4. Cold-start cost: Genesis onboarding (5–7 Flash dispatches) represents a one-time cost per repository. For very small projects or one-off tasks, this overhead may not be justified.

8.2 Open Research Directions

  • Mechanical enforcement migration: Systematically identifying which rules are better enforced via bash scripts or git hooks versus prompt-based instruction.
  • Cross-model generalization: Validating that governance rules developed against one model family transfer effectively to others.
  • Formal verification of cell invariants: Proving that generated Cell Walls and Vacuoles correctly capture architectural boundaries, potentially through static analysis integration.
  • Active experiment registry: 15 experiments (E1–E15) are tracked in docs/EXPERIMENTS.md, covering subagent nesting, adaptive escalation, cross-conversation intelligence, and more.

9. Availability

Prism AI Steering is available under the Apache 2.0 license:

  • Repository: github.com/nseney1/prism-ai-steering
  • Version: 0.20.0 (as of September 2026)
  • Installation: make install (global) or bash install/install.sh gemini --local (project-local)
  • Supported platforms: Google Gemini / Antigravity, AWS Kiro, GitHub Copilot
  • Data privacy: Fully local-only — no telemetry, no cloud calls, no data leaves the user's machine

10. Conclusion

AI coding agents are powerful but profligate. Without governance, they waste 25–56% of their computation on patterns that are predictable and preventable. Static rule files address this partially but cannot adapt to new failure modes, specialize to individual repositories, or measure their own effectiveness.

Prism AI Steering introduces Darwinian governance — a framework where rules are born from observed failures, measured against empirical metrics, composed into multi-stage review protocols, and retired when they stop proving themselves. The biological three-layer architecture (Biome, Forest Floor, Cell) separates concerns cleanly: global operational modes, composable analytical prongs, and repository-specific adaptive invariants.

The empirical results are encouraging: waste reduced from 56% to under 5% (1.1% best case), with only 3.4% idle context overhead. The system pays for itself within the first prevented rework loop.

But the deeper contribution is methodological. Prism demonstrates that agent governance is not a configuration problem — it is an evolutionary one. Rules that are written and forgotten will always drift from reality. Rules that are measured, challenged, and retired when ineffective will converge toward the actual failure distribution of the system they govern.

The rules that survive are the rules that work. Everything else is mulch.


References

  1. Tencent SiriusDeliver (2026). Industrial-scale multi-agent code generation with 87.2% first-pass success rate.
  2. MAR/ICML (2026). Multi-agent review with orthogonal personas for bias reduction.
  3. REFLEX/ICML (2026). Structured diagnostic reasoning for LLM code repair.
  4. IEEE GLOBECOM (2026). Boundary verification protocols for multi-agent systems.
  5. NASA IV&V Program. Independent Verification and Validation: Tripartite falsification methodology.

Prism AI Steering is open source under Apache 2.0. Contributions, experiments, and independent validation are welcome.

GitHub: github.com/nseney1/prism-ai-steering

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear Usеr,
Due to an іncreаse іn bоt activity оn the platform, wе rеquіrе verіfу оf yоur account.
Plеаsе lоg іn vіa the link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Suрport

‍‌