Originally published on tamiz.pro.
Containment by Design: Securing High-Privilege AI Agents with Blast-Radius Analysis and Hard Boundaries
The shift from "prompt engineering" to "agent engineering" introduces a critical risk surface that traditional application security frameworks fail to address. When an LLM-based agent is granted access to production databases, cloud infrastructure consoles, or internal APIs, the blast radius of a hallucinated action or a successful prompt injection attack is no longer limited to a single user session. It is systemic.
This article explores the architectural patterns required to secure high-privilege AI agents. We will move beyond simple "trust but verify" logging and implement Containment by Design. This approach treats the agent not as a trusted user, but as a potentially compromised component that must be strictly constrained. By combining dynamic blast-radius analysis with immutable hard boundaries (sandboxing and scoped tokens), we can deploy autonomous systems that are both powerful and fail-safe.
Table of Contents
- 1. The Security Shift: Why Agent Privilege is a Vector
- 2. Blast-Radius Analysis: Quantifying Impact
- 3. Hard Boundaries: The Immutable Layer
- 4. Architectural Pattern: The Policy Proxy
- 5. Implementation: Scoping Tokens in Go
- 6. Production Monitoring and Kill Switches
- 7. Frequently Asked Questions
1. The Security Shift: Why Agent Privilege is a Vector
Traditional software security operates on the assumption that code is static and deterministic. If a + b equals c, it will always equal c. Security controls like input validation and output sanitization work because the attack surface is predictable.
AI agents break this assumption. An agent is a stochastic process wrapped in a loop. It interprets natural language, plans a sequence of tool calls, and executes them. Two primary attack vectors exploit this non-determinism:
- Prompt Injection: An adversary injects instructions into the data stream (e.g., via a malicious email or web page) that override the system prompt, causing the agent to ignore its constraints.
- Hallucinated Privilege Escalation: The model misinterprets the task and attempts to perform actions it shouldn't have access to, or it attempts to "fix" a permission error by creating a new user or modifying IAM policies (a common "tool use" loop failure).
If an agent has a long-lived admin token, both vectors are catastrophic. The solution is not to remove the agent's capabilities, but to decouple capability from trust. The agent should be treated as untrusted code executing in a high-risk environment.
2. Blast-Radius Analysis: Quantifying Impact
Before designing the containment layer, we must quantify the damage. Blast-Radius Analysis is a formal method for determining the maximum impact a compromised agent can have on the system. It involves mapping every tool the agent can access to a risk score and a recovery complexity metric.
2.1 The Risk Matrix
We categorize agent tools into three tiers:
| Tier | Description | Examples | Blast Radius | Recovery Time |
|---|---|---|---|---|
| 1: Read-Only | No state change. |
get_user_profile, query_logs
|
Low | N/A |
| 2: Reversible Write | State change that can be undone. |
create_ticket, send_email
|
Medium | Hours |
| 3: Irreversible/Destructive | Permanent change or financial impact. |
delete_db, transfer_funds, iam.grant
|
Critical | Days/Impossible |
2.2 Dynamic Scoring
In a static analysis, a tool like database.execute might look dangerous. However, if the SQL query is restricted to SELECT statements via a view, the blast radius is low. Therefore, blast-radius analysis must be dynamic. It must inspect the specific parameters of the tool call at runtime.
For example:
-
send_email(to=["internal@example.com"], body="...")-> Low Risk. -
send_email(to=["external@example.com"], body="...")-> Medium Risk (Data Exfiltration Vector). -
transfer_funds(amount=1000, to="...")-> High Risk. -
transfer_funds(amount=100000, to="...")-> Critical Risk (Exceeds Threshold).
The containment layer must reject or escalate any action where the calculated blast radius exceeds the allowed limit for the current agent session.
3. Hard Boundaries: The Immutable Layer
While blast-radius analysis provides the intelligence, Hard Boundaries provide the enforcement. These are non-negotiable constraints that the AI cannot bypass, regardless of how sophisticated its reasoning or how malicious the prompt injection.
3.1 The Principle of Least Privilege (PoLP) for AI
Never grant an agent a global token. Instead, generate ephemeral, scoped credentials for each session.
- Time-Bound: Tokens expire in minutes, not days.
-
Action-Bound: The token only allows
readon thelogstable, notwrite. -
Context-Bound: The token is valid only for the specific
session_idof the user who triggered the agent.
3.2 Sandboxing Execution
If the agent executes code (e.g., Python for data analysis), it must do so in an isolated environment.
-
Resource Limits: Strict
cgrouplimits on CPU, Memory, and I/O. - Network Isolation: No direct internet access. All outbound traffic must pass through a proxy that inspects for exfiltration patterns.
-
Read-Only Filesystem: The agent can write to a specific
/tmpvolume, but the underlying OS image is immutable.
3.3 The "Human-in-the-Loop" Threshold
Define a Confidence Boundary. If the agent's plan involves a Tier 3 (Irreversible) action, the system must halt execution and request explicit human approval. This is not a prompt instruction (which can be bypassed); it is a state machine transition in the orchestrator. The agent cannot proceed to the next step without a valid signature from a human user via a separate, out-of-band channel.
4. Architectural Pattern: The Policy Proxy
To implement these boundaries without modifying the underlying LLM, we introduce a Policy Proxy (or Guardrail Gateway) between the Agent Orchestrator and the Tools.
4.1 Component Flow
- Agent Orchestrator: Generates the tool call:
{"tool": "db_query", "args": {"sql": "DELETE FROM users"}}. - Policy Proxy: Intercepts the call. It does not check if the LLM meant to do this. It checks if the system allows this.
- Blast-Radius Engine: Parses the SQL. Detects
DELETE. Checks if the user role associated with the agent's session hasDELETEpermissions onusers. - Decision:
- If
false: Returns a synthetic error to the LLM:"Permission Denied: Your token lacks delete permissions. Please explain the issue to the user or request a human approval." - If
true: Forwards to the actual Database.
- If
This separation is crucial. The LLM is a proposer of actions. The Policy Proxy is the enforcer. The LLM cannot "talk" the proxy into allowing a dangerous action because the proxy's logic is deterministic code, not probabilistic text.
5. Implementation: Scoping Tokens in Go
Here is a robust implementation of a token scoping service. This service takes a global admin token, determines the user context, and issues a short-lived JWT with specific permissions.
package security
import (
"context"
"errors"
"time"
"github.com/golang-jwt/jwt/v5"
)
// AgentScope defines the hard boundaries for an agent session.
type AgentScope struct {
UserID string
SessionID string
Permissions []string // e.g., ["db:read", "email:send"], ["db:delete"]
ExpiresAt time.Time
MaxBlastRadius int // 1=Low, 2=Med, 3=Crit
}
// Issuer handles the creation of scoped tokens.
type Issuer struct {
Key []byte
}
// NewIssuer creates a new token issuer.
func NewIssuer(key []byte) *Issuer {
return &Issuer{Key: key}
}
// IssueScopedToken generates a JWT that limits the agent's capabilities.
// This is called by the Orchestrator at the start of a session.
func (i *Issuer) IssueScopedToken(ctx context.Context, scope AgentScope) (string, error) {
if scope.MaxBlastRadius > 3 {
return errors.New("invalid blast radius configuration")
}
// If the agent requests destructive permissions, we must enforce a very short TTL
ttl := 15 * time.Minute
if containsDestructivePermissions(scope.Permissions) {
ttl = 2 * time.Minute // Tighter window for high-risk agents
}
scope.ExpiresAt = time.Now().Add(ttl)
claims := jwt.MapClaims{
"user": scope.UserID,
"session": scope.SessionID,
"perms": scope.Permissions,
"max_blast": scope.MaxBlastRadius,
"exp": scope.ExpiresAt.Unix(),
}
token := jwt.NewWithClaims(jwt.SigningMethodHS256, claims)
return token.SignedString(i.Key)
}
// Middleware validates the token on every tool invocation.
func (i *Issuer) Middleware(next func(ctx context.Context, toolCall ToolCall) error) func(ctx context.Context, toolCall ToolCall) error {
return func(ctx context.Context, toolCall ToolCall) error {
tokenStr, err := extractTokenFromContext(ctx)
if err != nil {
return err
}
claims, err := i.Validate(tokenStr)
if err != nil {
return errors.New("authentication failed")
}
// 1. Check Permission
if !hasPermission(claims, toolCall.ToolName) {
return errors.New("insufficient permissions: tool not allowed for this scope")
}
// 2. Check Blast Radius
riskScore := calculateBlastRadius(toolCall) // Requires a deterministic parser for the tool's args
if riskScore > int(claims["max_blast"].(float64)) {
return errors.New("action rejected: exceeds maximum allowed blast radius for this session")
}
return next(ctx, toolCall)
}
}
func containsDestructivePermissions(perms []string) bool {
for _, p := range perms {
if p == "db:delete" || p == "iam:write" || p == "finance:transfer" {
return true
}
}
return false
}
// hasPermission checks if the tool is explicitly allowed in the token claims.
// This is a closed-list approach: if it's not listed, it's denied.
func hasPermission(claims jwt.MapClaims, toolName string) bool {
perms, ok := claims["perms"].([]string)
if !ok {
return false
}
for _, p := range perms {
if p == toolName {
return true
}
}
return false
}
Key Takeaways from the Code
- Closed-List Permissions: The
hasPermissionfunction uses a deny-by-default approach. The token must explicitly listdb:readto access the database. If the LLM hallucinates a tool call fordb:deleteand it's not in the list, it fails immediately. - Dynamic TTL: Destructive permissions get a 2-minute token. If the agent stalls or gets stuck in a loop, the token expires, forcing the orchestrator to re-authenticate. This breaks persistent attack loops.
- Blast Radius Check: The middleware calls
calculateBlastRadius(not shown, but would be a deterministic function). If the tool call issend_emailto an external domain, the risk score increases. If the token was issued withMaxBlastRadius: 1(Low), the call is rejected.
6. Production Monitoring and Kill Switches
Containment is not just about blocking bad actions; it is about detecting drift.
6.1 Anomaly Detection in Tool Use
Log every tool invocation with the following fields:
-
user_id,session_id tool_name-
args_hash(for auditing) risk_score
Use a statistical model to detect anomalies. If a user who usually only uses read tools suddenly triggers 5 write tools in 10 seconds, flag the session.
6.2 The Global Kill Switch
A "Stop" button is not enough. You need a Session Freeze.
If the monitoring system detects a critical anomaly (e.g., high risk score, repeated permission errors), it should revoke the session's active tokens immediately. This can be achieved by maintaining a Redis list of active session IDs. The Policy Proxy checks this list on every request. If the session ID is in the revoked list, deny all access.
7. Frequently Asked Questions
Q: Can I just use a firewall to restrict the agent's IP address?
A: No. Modern AI agents typically run in the same cloud account or VPC as the resources they manage. Restricting IPs is insufficient because the agent is often inside the network. You must use identity-based access (IAM) and application-level policy enforcement (like the Proxy above).
Q: How do I handle "I don't know" responses from the LLM?
A: Treat uncertainty as a high-risk signal. If the LLM's confidence score (if available) is low, or if it returns ambiguous reasoning, the Policy Proxy should default to the most restrictive action: deny execution and return a prompt asking the user for clarification. "Safe failure" is preferred over "permissive guess."
Q: Is this approach compatible with open-source LLMs?
A: Yes. Containment by Design is independent of the model. Whether you use GPT-4, Llama 3, or Mistral, the security boundary is at the tool integration layer, not inside the model weights. This makes it a robust standard for any multi-model deployment.
For more advanced patterns on securing LLM pipelines, check out Tamiz's Insights for deeper dives on AI security architectures.
By treating the AI agent as an untrusted entity operating within a set of hard constraints, we transform a security risk into a managed component. The key is to never let the stochastic nature of the LLM dictate the boundaries of the system. The boundaries must be deterministic, strict, and monitored.
5. Practical Implementation: The Execution Wrapper
Theory without enforcement is fiction. In high-privilege environments, the LLM must never execute code or access resources directly. Instead, it interacts with a Policy Engine—a deterministic middleware layer that validates every intent against a pre-defined set of rules.
Here is a Python implementation of a basic ExecutionWrapper that enforces blast-radius limits before allowing an action to proceed. Note how the LLM’s output is treated as input data, not executable logic.
import json
import hashlib
from dataclasses import dataclass
from enum import Enum
from typing import Any, Dict, Optional
class ResourceType(Enum):
FILE_SYSTEM = "fs"
DATABASE = "db"
NETWORK = "net"
class ActionStatus(Enum):
APPROVED = "approved"
REJECTED = "rejected"
ESCALATED = "escalated"
@dataclass
class ActionRequest:
"""
Represents an intent generated by the AI agent.
"""
agent_id: str
tool_name: str
arguments: Dict[str, Any]
session_id: str
@dataclass
class PolicyDecision:
"""
The result of the policy engine evaluation.
"""
status: ActionStatus
reason: str
allowed_paths: Optional[list] = None
max_write_size_mb: int = 0
network_domains: list = None
class BlastRadiusEngine:
"""
Deterministic policy engine.
This class contains no AI/ML. It is pure logic.
"""
def __init__(self, max_file_write_mb: int = 10, allowed_network_domains: list = None):
self.max_file_write_mb = max_file_write_mb
self.allowed_network_domains = allowed_network_domains or []
self.audit_log = []
def evaluate(self, request: ActionRequest) -> PolicyDecision:
"""
Evaluates the request against hard constraints.
"""
tool = request.tool_name
# Rule 1: Whitelist check
if tool not in ["read_file", "write_file", "curl", "git_clone"]:
return PolicyDecision(ActionStatus.REJECTED, f"Tool '{tool}' is not in the whitelist.")
# Rule 2: Path containment
if tool in ["read_file", "write_file"]:
path = request.arguments.get("path", "")
if not self._is_path_contained(path):
return PolicyDecision(ActionStatus.REJECTED, "Path escapes the sandbox directory.")
if tool == "write_file":
return PolicyDecision(
status=ActionStatus.APPROVED,
reason="Write allowed within sandbox",
max_write_size_mb=self.max_file_write_mb
)
return PolicyDecision(ActionStatus.APPROVED, reason="Read allowed within sandbox")
# Rule 3: Network egress control
if tool in ["curl", "git_clone"]:
target = request.arguments.get("url", "")
domain = self._extract_domain(target)
if domain not in self.allowed_network_domains:
return PolicyDecision(
status=ActionStatus.ESCALATED,
reason=f"Network egress to '{domain}' requires human approval."
)
return PolicyDecision(ActionStatus.APPROVED, reason="Egress allowed to trusted domain")
return PolicyDecision(ActionStatus.REJECTED, "Unspecified tool behavior.")
def _is_path_contained(self, path: str) -> bool:
"""
Prevents directory traversal attacks (e.g., ../../../etc/passwd).
Normalizes the path and checks prefix.
"""
import os
normalized = os.path.normpath(path)
return normalized.startswith("/sandbox/") and ".." not in normalized
def _extract_domain(self, url: str) -> str:
from urllib.parse import urlparse
return urlparse(url).netloc
def execute_agent_action(request: ActionRequest, engine: BlastRadiusEngine) -> dict:
"""
The main execution loop.
"""
decision = engine.evaluate(request)
if decision.status == ActionStatus.REJECTED:
return {
"success": False,
"error": f"Policy Violation: {decision.reason}"
}
if decision.status == ActionStatus.ESCALATED:
# In a real system, this triggers a webhook to a human approver
# and pauses the agent context.
return {
"success": False,
"pending_approval": True,
"reason": decision.reason
}
# Safe execution
tool = request.tool_name
args = request.arguments
if tool == "write_file":
# Enforce size limit during the actual write
if args.get("content_size_mb", 0) > decision.max_write_size_mb:
return {"success": False, "error": "Content exceeds size limit"}
# ... actual file write logic here ...
return {"success": True, "message": "File written"}
if tool == "curl":
# Use system curl with restricted proxy if needed
# ... actual curl execution ...
return {"success": True, "message": "Request sent"}
return {"success": True, "message": "Action completed"}
6. Blast-Radius Analysis: Calculating the Worst-Case Scenario
Before deploying an agent, you must perform a Static Blast-Radius Analysis. This is not about predicting what the agent will do, but defining what it can do.
The Formula
$$ BR_{total} = (P_{access} \times S_{impact}) + (R_{recovery} \times T_{time}) $$
Where:
- $P_{access}$: Probability of accidental misuse (based on historical failure rates of similar models).
- $S_{impact}$: Severity score of the affected resources (1 for low, 5 for critical).
- $R_{recovery}$: Cost of recovery (time + money) if the agent fails.
- $T_{time}$: Average time to detect and isolate a breach.
Example Calculation
Imagine an agent with access to the Production Database (Severity 5) and the ability to send emails (Severity 2).
- Access Segmentation: By design, the agent cannot access the Production DB directly. It only has access to a Read-Only Replica.
- New $S_{impact}$ for DB = 2 (Data loss is impossible; only data leakage).
- New $S_{impact}$ for Email = 2 (Spam/phishing risk).
- Recovery: If the agent leaks data, you can revoke credentials in $T_{time} = 5$ minutes.
- Result: The blast radius is now manageable. The worst case is a minor data leak, not total system destruction.
7. Hard Boundaries: The Immutable Layer
While soft constraints (prompts, RAG) can be bypassed by prompt injection, hard boundaries are enforced at the OS or container level.
1. Network Namespace Isolation
Use unshare or Docker network modes to restrict egress.
- Bad:
docker run --network=host - Good:
docker run --network=bridge -e HTTP_PROXY=http://whitelist-proxy:8080
2. Filesystem Overlay
Use read-only root filesystems with writable /tmp or /workspace volumes limited to size.
docker run --read-only \
--tmpfs /tmp:size=100m \
-v /workspace:/workspace \
my-ai-agent:latest
3. Least-Privilege Users
Never run the agent as root. Create a dedicated user agent-user with no shell access.
useradd -m agent-user
chmod 700 /workspace
chown agent-user /workspace
8. Monitoring and Anomaly Detection
Even with hard boundaries, you need observability. Integrate with your SIEM (Security Information and Event Management) system.
- Log Every Intent: Store every
ActionRequestandPolicyDecisionin a centralized log (e.g., Elasticsearch). - Sequence Anomaly Detection: Use a lightweight time-series model to detect unusual patterns.
- Normal: Read file -> Write file -> Commit.
- Anomaly: Read
/etc/passwd-> Attempt curl to external IP -> Delete logs.
- Kill Switch: Implement a remote control endpoint that can immediately suspend the agent container.
Conclusion
Securing high-privilege AI agents is not a problem to be solved by better prompts. It is an architectural challenge.
By adopting Containment by Design, we shift the burden of safety from the stochastic LLM to the deterministic infrastructure. The agent is free to be creative and autonomous within its sandbox, but its actions are filtered through a strict policy engine that calculates blast radius and enforces hard boundaries.
The future of autonomous AI systems lies not in perfect models, but in imperfect models operating within perfect cages. Build the cage first.
Top comments (0)