DEV Community

lucanu
lucanu

Posted on

Securing AI Agents: Enforce Boundaries Outside the Model, Not Inside the Prompt

Prompt injection sits at the top of the OWASP Top 10 for LLM Applications (LLM01), and no system prompt reliably stops it. If your agent's only guardrail is a sentence like "never delete files," you have a suggestion, not a boundary. The fix is architectural. Treat the model as an untrusted planner and put enforcement in deterministic code the model cannot talk its way past. This article covers four layers: capability scoping, a policy gate on every tool call, sandboxed execution, and verification you can run in CI. ## Scope Capabilities Before You Write Any Policy

Most agent incidents trace back to over-provisioned credentials. The agent can drop a production table because the database user can drop tables. But oWASP calls this LLM06, Excessive Agency, and the remedy is old-fashioned least privilege. The fix is concrete. Give the agent its own identity, never a developer's token. For Postgres, create a role with exactly the grants the task needs:

CREATE ROLE support_agent LOGIN PASSWORD '...';
GRANT SELECT ON orders, customers TO support_agent;
GRANT UPDATE (status) ON orders TO support_agent;
REVOKE ALL ON SCHEMA public FROM support_agent;
Enter fullscreen mode Exit fullscreen mode

Apply the same rule to GitHub, using a fine-grained personal access token or a GitHub App limited to one repository with contents: read. Apply it to AWS, using an IAM role whose policy lists specific actions like s3:GetObject on one bucket ARN, with no wildcards. If the credential cannot perform an action, no injected instruction can make the agent perform it. Takeaway: List every credential your agent holds today and replace any shared or admin-scoped token with a dedicated identity limited to the exact resources the agent touches. ## Put a Policy Gate Between the Model and Every Tool

Credentials set the outer wall. Inside it, you still need per-call decisions. A refund tool might be allowed in general, but you may not want the agent issuing a $5,000 refund without a human. The model proposes a tool call as structured JSON. Your code validates it before anything executes. A useful split looks like this:

  1. Schema validation. Parse arguments with Pydantic or Zod and reject unknown fields. Reject malformed calls instead of trying to repair them. 2. Policy evaluation. Send the call, the user context, and the session state to a policy engine. Open Policy Agent (OPA) with Rego works well because policies live in version control and are testable. AWS Cedar is a solid alternative. 3. Escalation. Return one of three decisions: allow, deny, or require_approval. Route the last one to a human queue. A minimal Rego rule:
package agent.tools

default decision := "deny"

decision := "allow" if {
 input.tool == "issue_refund"
 input.args.amount <= 100
 input.args.order_id in input.session.verified_orders
}

decision := "require_approval" if {
 input.tool == "issue_refund"
 input.args.amount > 100
}
Enter fullscreen mode Exit fullscreen mode

The verified_orders check matters most. It ties the action to data the user is authorized for, not data the model happened to mention. That blocks the classic confused-deputy attack, where injected text in a support ticket says "refund order 98231."

Takeaway: Wrap your tool dispatcher in a single function that calls a policy engine with default-deny. Make sure no tool can execute through any other path. ## Sandbox Anything That Executes Code or Touches the Network

Agents that run shell commands or generated Python need OS-level isolation. A policy gate cannot reason about what python script.py will do once it starts. Practical options, from lighter to stronger:

  • Docker with hardening flags. At minimum: docker run --rm --network=none --read-only --cap-drop=ALL --pids-limit=128 --memory=512m --user 1000:1000 agent-sandbox. - gVisor (--runtime=runsc). It adds a user-space kernel, which shrinks the syscall attack surface. - Firecracker microVMs. This is the technology behind AWS Lambda, and it gives you VM-level isolation with fast boot times. Hosted services like E2B build on it. Network egress is where data exfiltration happens. One common pattern is an injected instruction that tells the agent to curl secrets to an attacker's domain. Default to --network=none. When the agent needs network access, route it through an egress proxy with a domain allowlist. Squid or Envoy both work. Takeaway: Run your agent's code-execution tool with --network=none and --cap-drop=ALL today. Then add back only the access that breaks. ## Verify Rules Continuously, Not Once

Policies rot as you add tools. Treat agent boundaries like any other security control: test them, log them, and attack them. Unit-test policies. Run opa test ./policies -v in CI. Write a test for every deny case, not just the happy path. Red-team with real tools. Promptfoo has red-teaming plugins for prompt injection and excessive agency. Garak, from NVIDIA, probes models for known jailbreak classes. Run either one against your full agent with its tools wired in. Testing the bare model misses the failures that matter. Log every decision. For each tool call, record the proposed call, the policy decision, the policy version, and the outcome. When something goes wrong, you need to know whether the policy allowed it or the agent bypassed the gate. A spike in deny events is also an early signal of an active injection attempt. Canary tests. Plant a fake secret, such as a decoy API key in a document, and alert if it ever appears in tool arguments or outbound requests. Takeaway: Add a CI job that runs your policy tests plus at least one injection test suite on every change to prompts, tools, or policies. Start with one action today: find the single most dangerous tool your agent can call, whether it deletes, pays, sends, or deploys. Put it behind a default-deny check in code that returns require_approval above a threshold you choose. That one gate converts your riskiest prompt-level hope into an enforced boundary.


This article was drafted with AI assistance.

Top comments (1)

Collapse
 
dev_in_the_fog profile image
Jason Y. (dev_in_the_fog) •

Loved how practical this is. So many LLM posts stay at surface level, but addressing real implementation trade-offs is where the value is.