A Cursor agent deleted a production database in 9 seconds last April. Not a test database. Production. PocketOS, a SaaS platform serving car rental businesses, lost its live database and all volume-level backups in a single Railway API call. The agent had been assigned a routine staging task. It hit a credential mismatch, scanned the codebase on its own, found an overprivileged API token in an unrelated file, and fired a destructive delete without asking anyone.
The founder's post-mortem is worth reading in full. But the detail that stuck with me: the token the agent used was sitting in a file that had nothing to do with the task. It had blanket authority across the entire Railway account. The agent found it, used it, and moved on. Nine seconds.
The problem is bigger than one incident
AI coding agents touch files that most teams never audit. Your .mcp.json lists MCP servers with auth tokens. Your CLAUDE.md or .cursorrules can define hooks that run shell commands on every session start. Your settings.json can set permission modes and allowlists.
The numbers keep stacking up:
Hush Security analyzed ~82,000 public MCP config files on GitHub. 12% of credential slots had a hardcoded secret. Not an env var reference, not a placeholder. A literal API key, database password, or bearer token sitting in version control. 55% of those secrets had no vendor-recognizable token shape, which means standard scanners like gitleaks and trufflehog walk right past them.
CVE-2026-25725 showed that Claude Code's sandbox didn't protect .claude/settings.json if the file didn't exist at startup. Malicious code inside the sandbox could create that file and inject persistent hooks that execute with host privileges on restart. CVSS 3.1 score: 10.0. Critical. The fix landed in v2.1.2, but the pattern is the point: repo-writable config files are an attack surface.
METR's study found that experienced developers using AI tools took 19% longer to complete tasks, while believing they were 20% faster. That perception gap matters here. If you feel faster, you review less. You skip the diff on the config file the agent just touched.
Glean's Work AI Institute found that workers spend 6.4 hours a week "botsitting" AI. Feeding it context, checking outputs, debugging mistakes. Almost a full working day, every week. But most of that time goes to code review, not config review. Nobody is systematically checking what the agent changed in .mcp.json or settings.json.
What actually gets missed
I started cataloging the surfaces after reading the PocketOS post-mortem. The list is longer than I expected:
Instruction files (
AGENTS.md,CLAUDE.md,.cursorrules,copilot-instructions.md,.windsurfrules): can contain hidden Unicode characters, bidi overrides that make a line render differently than it executes, prompt injection that overrides the instruction stack, or instructions to exfiltrate.envfiles and SSH keys.MCP config (
.mcp.json,mcp.json): auth tokens hardcoded as literals instead of${ENV_VAR}references. High-entropy strings in env blocks that match no known vendor pattern. Connection strings with embedded passwords.Hooks and plugins (
.claude/commands/,.cursor/scripts,plugin.json): shell scripts that run on session start or commit. These can read credential stores, pipe fetches into shells, or declare permission bypasses.Controls (
settings.json,settings.local.json): permission modes, hook definitions, and allowlists defined in files the agent itself can write to. The CVE-2026-25725 pattern.
Most teams have no automated check for any of this. They run secret scanners on application code and ignore the agent config directory entirely.
A scanner that covers the gap
I built a Python scanner that checks all of these surfaces. Five bands, 22 rules, zero dependencies beyond stdlib. It runs offline, no AI, no network calls. You point it at a repo and get a finding table with severity, location, evidence, and a concrete fix for each issue.
The five bands:
1. Unicode / Injection detection
Finds invisible characters (zero-width spaces, soft hyphens), bidi control characters that can make code render differently than it executes, and homoglyph characters (Cyrillic or Greek letters mixed into Latin words). These are the building blocks of visual spoofing attacks in instruction files.
2. Injection / Override detection
Catches exfiltration instructions (reading .env, SSH keys, AWS credentials, piping fetches into shells), prompt injection ("ignore all previous instructions"), permission bypasses (dangerouslySkipPermissions, --yolo), and guardrail weakening instructions. Filters out advisory phrasing like "never disable the sandbox" so you don't get false positives on your own safety guidance.
3. Secrets exposure
Pattern-matches known credential shapes (OpenAI keys, GitHub tokens, AWS access keys, Slack tokens, Google API keys, GitLab tokens, NPM tokens, private key headers, bearer tokens, connection strings). Also catches high-entropy literals under secret-suggesting key names, and the blind-spot class the Hush study identified: high-entropy strings in env blocks that match no vendor shape.
4. Mods / Plugin audit
Scans hook scripts and plugin manifests for the same exfiltration and credential-read patterns, plus plugins that declare permission bypasses in their manifest.
5. Controls band
Flags hooks and permissions defined in repo-writable config files. This is the CVE-2026-25725 pattern. If settings.json defines hooks in a file the agent can write to, the scanner tells you to move it to managed settings. Personal override files (.local.) get an info-level note instead of a blocking finding.
Output: A severity-sorted finding table (critical/high/medium/low/info) with file location, rule name, what it found, masked evidence, and a specific remediation step. Exit code 0 means clean, exit code 1 means blocking findings. Fits into CI with one line.
python audit/scan.py /path/to/repo --report audit-report.md
Get it
The Security Audit Kit is $29 on Gumroad. It's v0.1.0. You get the scanner, the rule definitions, and documentation for every band.
Launch discount: use code LAUNCH for €10 off.
Want to try before you buy? Grab the free sample. It's a complete agent config kit for a Next.js repo, free, so you can see how a versioned kit is structured.
Want free examples first? The agents-md-examples repo on GitHub has example agent config files you can study and adapt.
Why this matters now
The PocketOS incident happened because a credential was reachable and the agent had no gate between finding it and using it. The Hush study showed that 12% of MCP configs on public GitHub have the same problem. CVE-2026-25725 showed that the config files themselves are an attack surface. And the METR and Glean studies show that we're spending more time supervising agents while catching less of what they do.
Automated scanning for agent config surfaces is table stakes. If you're running AI agents on your codebase and you haven't checked what's in your .mcp.json, your CLAUDE.md, or your settings.json, now is a good time to start.
Top comments (0)