The problem is not the prompt
When teams ask about best practices for implementing LLM access controls and monitoring, the conversation usually starts with prompt injection. That is the wrong starting point. Prompt injection is the delivery mechanism. The thing that turns it into an incident is an over-scoped credential sitting behind a tool the model can call.
An agent that holds a long lived API key with broad scope does not need to be tricked cleverly. It needs one instruction, from a retrieved document, a tool response, or a user turn, that points the existing permission at the wrong resource. The call that follows is authorized at the network layer. It looks normal. That is what makes it hard to catch.
What AI agent security actually covers
AI agent security is the discipline of controlling what an autonomous or semi-autonomous model-driven process is allowed to do, and observing what it did. It spans three layers:
- Identity. Which agent is acting, and is that identity distinct from the human who triggered it.
- Authorization. Which tools and scopes that identity may reach, evaluated per request.
- Observability. A record of every allow and deny, tied to the identity and the requested action.
Most deployments today have layer one partially and layers two and three not at all. The model is given a credential and trusted to use it well.
The attack surface, concretely
Consider a support agent with a tool that reads customer records. The tool is backed by a service account key with read access to the whole customer table. The agent is instructed to only look up the current user.
Now a retrieved document contains an instruction to look up a different customer and include the result in the summary. The agent complies. The tool call succeeds. The service account is authorized for that read. Nothing in the stack flags it, because nothing in the stack was checking intent against scope. It was checking whether the key works.
This is not theoretical. It is the default shape of most agent integrations. The blast radius of any single injected instruction equals the scope of the credential behind the tool.
The mechanism that stops it
The control belongs in the request path, between the agent and the tool. Not in the prompt. Not in a post-hoc log review.
The ordering matters:
- Resolve the agent identity for this request. Not the user identity alone. The agent acting on the user's behalf is its own principal.
- Evaluate the requested tool and the requested scope against a policy for that identity. Deny by default.
- Issue a short lived credential scoped to exactly that call, only if the policy allows.
- Emit an event before the call executes, recording identity, tool, scope, and decision.
- Execute. Emit a completion event.
Step four is the one teams skip. Logging after execution tells you what happened. Logging before execution gives you a decision point and an audit trail that is not reconstructed from side effects.
The result is that a single injected instruction can only reach what the policy already permitted for that agent identity at that moment. Least privilege stops being an aspiration and becomes a runtime property.
Implementation with RESK
reskSecure is where the policy and the enforcement live. You define agent identities and the tools and scopes each may reach. The policy is evaluated per request, in the path, before the credential is issued. A denied request never reaches the tool.
ReskPoints is where the decisions land. Every allow and deny becomes an event tied to an agent identity, a tool, and a scope. That gives you the monitoring half of the question without bolting on a separate pipeline. You can see which agents are hitting denials, which scopes are being requested, and where policy is too broad.
Together they cover the two halves that are usually missing: enforcement at request time, and a record of enforcement that is tied to identity rather than inferred from logs.
For the tool-level view of how permissions should be structured, see our article on AI agent tool permissions.
Checklist
- Give each agent its own identity. Do not reuse a human credential or a shared service account.
- Deny by default. Allow specific tools and scopes per identity.
- Evaluate policy per request, in the path, before credential issuance.
- Issue short lived, narrowly scoped credentials for a single call.
- Emit the decision event before execution, not after.
- Record every allow and deny against the agent identity.
- Review denials as signal. A spike means either an attack attempt or a policy that is too tight.
- Review allows as risk. A broad allow is the blast radius of the next injected instruction.
The prompt is not the control point. The credential is.
Top comments (0)