In my last post, I shared an early prototype of Soma and argued why static .cursorrules and CLAUDE.md files are essentially gentleman's agreements. Static text rules sit in your system prompt, consume context, and get ignored the moment an agent encounters a tricky multi-file diff.
Since that post, I’ve put the codebase through intensive rounds of security hardening, adversarial edge-case testing, and dogfooding.
As of today, Soma is at v0.89.0, backed by 1,890 passing tests, 51 verification engines, 15 auditing organs, and native MCP support.
Here is what I changed during that hardening process, how the two-layer verification engine works, and the full feature set now available in the package.
What broke in my early versions (and why hardening was necessary)
When I let AI agents edit code autonomously, I quickly realized that basic rule checking isn't enough. In early iterations, I ran into three major friction points:
- Out-of-order state mutations: An agent would request a rule check, modify a file, and then try to apply a secondary edit based on stale workspace state. Without cryptographic state tracking, edits could race or corrupt files.
-
Platform and shell edge cases: Path handling under Windows (PowerShell vs. Git Bash MSYS path translation),
zshvsbashPATH detection, and symlink traversal presented real security risks. -
Dependency and execution leaks: Bare third-party imports (like
import yaml) in test suites or verification scripts could trigger CI crashes or un-guarded code execution.
The Two-Layer Verification Process
To guarantee safety without sacrificing execution speed, I implemented a strict Two-Layer Verification Protocol:
[ AI Agent / LLM ]
│
├── 1. Request State-Bound Receipt (soma_request_receipt)
│
├── 2. Submit Proposed Diff (soma_propose_change)
│
├──► [ LAYER 1: Fast AST & Import Guard (<5ms) ]
│ │ • Static AST parsing & syntax validation
│ │ • Catches un-guarded third-party imports
│ └─► REJECTS immediately if syntax or import guard fails (0 token cost)
│
└──► [ LAYER 2: Dynamic Proposer-Verifier Loop ]
│ • JIT path matching against active .soma/cells/
│ • Wilson-bounded fitness ranking
└─► WRITES to disk only when all gates pass
Layer 1: Instant AST & Import Guard (<5ms)
Before running heavy test suites or invoking LLMs, Layer 1 runs pure Python AST parsing on proposed diffs:
- It parses the AST to detect third-party modules (like
import yamlorimport requests) that aren't guarded bytry/exceptblocks or explicitly listed in project dependencies. - It catches syntax errors, broken signatures, and missing error guards.
- It operates locally in under 5 milliseconds with zero API calls and zero LLM token cost.
Layer 2: Dynamic Proposer-Verifier Loop
Once Layer 1 passes, Layer 2 evaluates the diff against your active governance rules:
- It matches files against active JIT rules (
.soma/cells/). - For catastrophic-risk paths (like Auth or Core Infrastructure), Layer 2 triggers Tempest Mode, forcing specialized security (
soma_audit_security) and performance (soma_audit_performance) audits before allowing disk writes.
State-Bound Cryptographic Receipts (soma_request_receipt)
To eliminate race conditions and unverified disk writes, write and execute tools require a single-use receipt:
- The agent calls
soma_request_receiptwith its intended tool and payload. - Soma hashes the current workspace state, target file state, and active cell state, issuing a single-use receipt valid for 300 seconds.
- If any file or rule state changes before execution, the receipt invalidates automatically and the action fails closed.
Additional Goodies Included in the Package
Beyond core governance, I built a complete suite of developer tooling into v0.89.0:
1. Multi-Language SDKs (soma-steering)
If you want to query governance health, calculate Wilson fitness intervals, or trigger verification directly inside your CI/CD pipelines or custom scripts:
-
Python SDK:
pip install soma-steering -
Node.js / TypeScript SDK:
npm install soma-steering
from soma_sdk import Governance
gov = Governance(project_root=".")
grade = gov.grade() # Overall codebase health grade
coverage = gov.coverage_report() # Governance rule coverage across files
landscape = gov.fitness_landscape(bayesian=True)
2. Self-Diagnosing CLI (soma doctor & soma status)
-
soma doctor: Audits your environment, detects shell profile configuration (bash,zsh,fish,pwsh), and fixes PATH friction automatically. -
soma status: Displays active core rules, adaptive cells, trigger counts, and estimated token overhead. -
soma grade&soma coverage: Evaluates rule coverage and codebase health right from your terminal.
3. Multi-Platform Agent Installers
Soma auto-detects and configures rule bindings for your favorite environment out of the box:
bash install/install.sh gemini # Google Gemini Antigravity
bash install/install.sh claude # Claude Code
bash install/install.sh cursor # Cursor
bash install/install.sh copilot # GitHub Copilot
bash install/install.sh kiro # AWS Kiro
bash install/install.sh mcp # Generic MCP server
4. Automatic Organelle Mapping
Whenever you initialize or scan a project (soma genesis), Soma generates an Organelle Map (docs/organelles.md) mapping every source module to its corresponding governance cells, giving you full visual transparency into rule coverage.
Security & Zero-Telemetry Guarantee
Since I built this to run locally on my own machine, security was non-negotiable:
-
pip-auditClean: The entire dependency tree passespip-auditwith 0 known vulnerabilities or CVEs. - 100% Local-Only: Zero telemetry, zero tracking, zero remote servers, and zero network calls. All rule evaluations and receipt validation happen strictly on your machine.
- Failure-Closed Execution: If an agent attempts an unauthorized file modification or passes an expired receipt, the operation aborts.
How I Dogfooded the Hardening Sprint
The fastest way I found to harden a governance engine was to use it on itself. During this hardening sprint:
- Every bug fix and security guard (from Windows path handling to receipt expiration) was written via test-driven development (TDD) alongside AI coding agents (Claude Code, Gemini Antigravity).
- Whenever an agent attempted an un-guarded edit or unsafe path call, Soma's Layer 1 verifier blocked the proposal before it hit disk.
Behavioral Testing Over Brittle Mocks
One crucial architectural choice I made during this sprint was enforcing behavioral testing over implementation testing:
- The 1,890 tests do not rely on fragile internal mocks. Instead, they test real workspace state changes, actual AST parsing, real receipt validations, and sad paths (expired receipts, symlink attacks, MSYS path translation under Windows).
- If an internal implementation detail changes but the behavioral contract holds, the test passes cleanly. If an actual security or execution boundary breaks, the test fails immediately. This is why the test suite remained reliable even while refactoring core engines.
Per-Prompt Overhead (Myth vs. Reality)
A common concern with governance tools is prompt bloat. While Soma manages 80+ core and adaptive rules in its local database, its Just-In-Time (JIT) engine only matches 2–3 relevant rules per git diff.
As a result, actual per-prompt token overhead stays locked at ~800 to 1,500 tokens (~1% of a standard context window), regardless of how large the repository grows.
Getting Started with v0.89.0
Soma v0.89.0 is open-source under the Apache 2.0 license.
# Install CLI
pip install soma-governance
# Initialize in your repo (auto-detects Gemini, Claude, Cursor, Copilot)
soma init --rules full --yes
To configure Soma as an MCP server in your agent (.mcp.json or ~/.gemini/antigravity/mcp_config.json):
{
"mcpServers": {
"soma": {
"command": "python3",
"args": ["-m", "soma_mcp"],
"env": {
"SOMA_WORKSPACE": "/path/to/your/project",
"SOMA_EXECUTION_ENABLED": "1"
}
}
}
}
- GitHub Repository: https://github.com/nseney1/Soma-Governance
- PyPI: soma-governance
- npm: soma-steering
How are you currently handling security boundaries and un-guarded edits when running AI agents on your codebase? Let me know in the comments below.
Top comments (0)