AI coding assistants are evolving rapidly. With Anthropic’s Model Context Protocol (MCP), agents in Claude Desktop, Cursor, or Zed are no longer confined to chat windows - they read directories, write files, inspect diffs, and invoke terminal commands directly on our development machines.
That level of agency is incredible for developer velocity, but it introduces a major security trade-off: runtime authority over sensitive files.
Here is why client-side confirmation modals fail in real-world workflows, why naive pattern blocking creates a false sense of security, and how I implemented a lightweight, sub-millisecond security shim in Rust to enforce deterministic policies.
The Core Problem: All-or-Nothing Permission Models
When using official servers like @modelcontextprotocol/server-filesystem, the security boundary is typically guarded by client-side modals:
Confirmation Fatigue: When an agent is refactoring a multi-file architecture, human-in-the-loop oversight quickly degrades into clicking "Always allow" out of sheer habit.
Accidental Credential Exposure: Even in well-maintained repositories, working trees frequently contain high-risk data: .env files, .npmrc with private registry tokens, cloud credential caches (~/.aws), or generated build artifacts.
Session-Wide Authority: Once a directory is granted, the MCP server will happily access any file inside that tree if the LLM requests it.
If an agent hallucinates a path or tries to read configuration secrets to understand your project structure, your credentials get ingested straight into the model's context window.
The Trap of Naive Pattern Matching (The Symlink Bypass)
My first instinct was simple: inspect the incoming JSON-RPC payload, search for ".env" in the argument string, and block the call if found.
That approach is flawed.
If an attacker (or a hallucinating agent) interacts with a workspace containing a symbolic link - say, symlink_config.txt -> .env - a raw string search on the arguments fails:
- The tool call asks for read_file(path: "symlink_config.txt").
- A string match sees no forbidden keywords.
- The OS follows the link and exposes your production keys.
A robust security gateway must perform canonicalization before policy evaluation:
- Intercept the JSON-RPC call over stdio.
- Extract the candidate path.
- Resolve all relative segments (../) and follow symbolic links to their real filesystem targets.
- Verify the canonical target strictly resides within the approved workspace boundary.
- Apply block/allow pattern matching against the resolved path.
Enter Argos: An In-Process Stdio Gateway in Rust
To solve this without adding heavy dependencies, background daemons, or runtime overhead, I wrote Argos - a single ~1 MB self-contained binary written in Rust using Tokio.
Instead of running an MCP server directly, the client invokes Argos, which launches the target server as a supervised child process:
{
"mcpServers": {
"Argos": {
"command": "/usr/local/bin/argos",
"args": [
"--",
"npx",
"-y",
"@modelcontextprotocol/server-filesystem",
"/path/to/workspace"
]
}
}
}
Argos acts as a zero-overhead pipe between stdin and stdout.
Key Implementation Highlights
- Intercepting tools/call Without Breaking JSON-RPC State If a policy violation occurs, simply terminating the process would crash the MCP transport and display an unhelpful Server Disconnected error to the user.
Instead, Argos catches the call on stdin and synthesizes an immediate standard MCP tool error response back to the client:
let error_response = json!({
"jsonrpc": "2.0",
"id": rpc_id,
"result": {
"content": [{
"type": "text",
"text": format!("[SECURITY POLICY VIOLATION] Action rejected by Argos Gateway: {reason}")
}],
"isError": true
}
});
Because the response complies with the JSON-RPC 2.0 specification, the AI client remains operational. The model receives clear, structured feedback that the action was blocked and continues its reasoning without context loss.
- Resolving Paths and Symlinks Here is the core verification logic applied before any policy evaluation:
if let Some(path_val) = target_path_str {
// 1. Block explicit traversal markers immediately
if config.filesystem.block_path_traversal && (path_val.contains("../") || path_val.contains("..\\")) {
return PolicyDecision::Block("Explicit path traversal ('../') attempt detected".into());
}
let candidate = PathBuf::from(path_val);
// 2. Resolve symlinks and normalize to real disk location
if let Ok(canonical_target) = candidate.canonicalize() {
let cand_str = canonical_target.to_string_lossy();
let clean_target = cand_str.strip_prefix(r"\\?\").unwrap_or(&cand_str);
// 3. Ensure the resolved destination does not escape the workspace
if let Ok(canonical_root) = workspace_root.canonicalize() {
let root_str = canonical_root.to_string_lossy();
let clean_root = root_str.strip_prefix(r"\\?\").unwrap_or(&root_str);
if !clean_target.starts_with(clean_root) {
return PolicyDecision::Block(format!("Resolved target path escapes workspace: {clean_target}"));
}
}
// 4. Evaluate allow/deny rules against the resolved file
let is_allowed = config.filesystem.allowed_patterns.iter().any(|p| clean_target.contains(p));
if !is_allowed {
for pattern in &config.filesystem.blocked_patterns {
if clean_target.contains(pattern) {
return PolicyDecision::Block(format!("Resolved target matches blocked pattern '{pattern}'"));
}
}
}
}
}
Conclusion & Next Steps
Local AI coding agents are becoming the standard development interface, but safety cannot rely purely on prompt engineering or manual click-through dialogs. Deterministic, low-level policy enforcement at the communication boundary provides the safety margin needed for true peace of mind.
Argos is open-source, runs completely locally with zero telemetry, and is available for Linux, Windows, and macOS.
GitHub Repository: [https://github.com/JUSICK/Argos-mcp-guardrail]
Pre-built Binaries: Available on the Releases page
If you are exploring runtime security for AI tooling, I would love to hear your thoughts, feedback, or any edge cases you've encountered with agent file operations!
Top comments (0)