DEV Community

shizhe Lim
shizhe Lim

Posted on AI-assisted

Designing Permission Boundaries for Production AI Agents

AI disclosure: This article was prepared with AI assistance. The author reviewed and edited the technical content and remains responsible for its accuracy.

AI agent permissions are often treated as a binary decision:

  • The agent can use a tool
  • The agent cannot use a tool

That may be sufficient for a prototype. It is usually too coarse for a production workflow.

Reading a customer record, drafting an update, saving a draft, and sending the final message may all involve the same system. However, they do not have the same operational risk.

A safer design separates capability from authority.

A four-layer action model

One practical approach is to classify actions into four layers.

1. Observe

The agent reads information without changing external state.

Examples:

  • Retrieve a document
  • Read an API response
  • Inspect a repository
  • Query a customer record
  • Review system status

Read access should still be scoped. An agent rarely needs access to every record, directory, API field, or historical event.

2. Prepare

The agent produces an output that has not yet affected another system or person.

Examples:

  • Draft an email
  • Generate a SQL query
  • Propose a code change
  • Prepare a support response
  • Create an execution plan

This layer is useful because it separates reasoning from execution. A person or another policy component can inspect the proposed action before anything changes.

3. Act reversibly

The agent changes state, but the result can be reviewed and undone with limited cost.

Examples:

  • Create a pull request
  • Save a message as a draft
  • Add an item to a review queue
  • Create a temporary resource
  • Update a record with version history

Reversible does not mean risk-free. The system still needs validation, logging, retry controls, and a defined rollback path.

4. Commit irreversibly

The agent performs an action that is difficult, expensive, or impossible to reverse.

Examples:

  • Send an external message
  • Merge a pull request
  • Delete production data
  • Publish public content
  • Submit a payment
  • Change access permissions

These actions often need the strongest policy checks and, depending on the workflow, explicit human approval.

Put a policy check before every tool call

Permission decisions should be based on more than the tool name.

The same tool may support both low-risk and high-risk actions. For example, an email integration might read a message, save a draft, or send a final response.

A simplified policy layer might look like this:

ACTION_LEVELS = {
    "read_record": "observe",
    "draft_message": "prepare",
    "save_draft": "reversible",
    "send_message": "irreversible",
}

def authorize(action, context):
    level = ACTION_LEVELS[action]

    if not context.user_has_access:
        return False, "User does not have access"

    if not context.input_valid:
        return False, "Input validation failed"

    if level == "irreversible" and not context.human_approved:
        return False, "Human approval required"

    return True, "Authorized"
Enter fullscreen mode Exit fullscreen mode

A production policy will be more detailed, but the important idea is simple:

The agent should not decide its own authority merely because it knows how to call a tool.

Make retries safe

Agent workflows frequently retry after timeouts, interrupted execution, or uncertain responses.

Without retry protection, an agent may repeat an external action even though the first attempt already succeeded.

Useful controls include:

  • An idempotency key for each intended side effect
  • An expected resource version
  • A record of the previous attempt
  • A clear completion state
  • A rollback or compensation reference

For example:

action_request = {
    "action": "send_message",
    "idempotency_key": "case-1842-final-response",
    "expected_version": 7,
    "requires_approval": True,
}
Enter fullscreen mode Exit fullscreen mode

The idempotency key represents the intended action, not an individual retry. Repeating the request should not create a second side effect.

Persist intent, not only the final result

A useful execution record should explain what the agent intended to do and why it was allowed to continue.

Consider storing:

  1. User or system objective
  2. Plan version
  3. Selected tool and arguments
  4. Action classification
  5. Policy-check result
  6. Approval status and approver
  7. Idempotency key
  8. Tool response
  9. Final workflow state
  10. Resume position after interruption

This information makes failures easier to diagnose and interrupted workflows easier to resume safely.

Make approval requests informative

A human approval step should not be a generic “Allow” button.

Before asking for approval, show:

  • The exact proposed action
  • The destination or affected resource
  • The relevant input values
  • The expected external effect
  • Whether the action is reversible
  • A preview or diff when possible

A reviewer should be able to understand the consequence without reconstructing the entire agent session.

Test permission boundaries directly

Agent evaluation should include more than successful task completion.

Test situations such as:

  • The agent requests data outside its allowed scope
  • Input data changes after the plan is generated
  • A tool call times out after completing the action
  • The same action is retried
  • Approval is denied or expires
  • A workflow fails halfway through
  • Execution resumes after a restart
  • A reversible action cannot be rolled back

These scenarios reveal whether the surrounding system remains safe when the agent, tool, network, or user behaves unexpectedly.

A practical release checklist

Before allowing an agent to perform external actions, verify that:

  • Every tool has the minimum required scope
  • Read and write permissions are separated
  • Irreversible actions have explicit policy rules
  • Duplicate retries cannot repeat a side effect
  • Every external call can be traced
  • Approval requests show the exact consequence
  • Interrupted workflows can resume safely
  • Recovery procedures have been tested

Final thought

Agent autonomy is not defined by how many tools the system can access.

It is defined by which actions the agent may perform, under what conditions, and with what evidence.

A useful production agent does not need unlimited authority.

It needs clearly designed permission boundaries.

Which action in your agent workflow would you never allow without human approval?

Top comments (0)