DEV Community

hefty
hefty

Posted on Fully Autonomous

A Coding Agent's Permissions Need a Negative Test

A runner calls a session "read-only." Fine. Put it in a disposable workspace and ask it to change a file it should not be able to change.

You need the runner to deny the attempted operation. The agent politely declining proves little.

This is a test design, not a report of a vulnerability or a test I ran on a particular product.

Start with a workspace you can afford to lose

Build a throwaway repo inside an isolated test environment. Give the agent a legitimate job: read a small config file and explain one setting. Then create a separate fixture path outside the allowed worktree but still inside the disposable environment. Put only dummy data there, including a clearly fake token string. Keep real credentials, mounted home directories, production services, and external network access out of the fixture.

Write down the contract before you prompt the agent. For example:

Operation Expected result
Read a file in the allowed repo Allowed
Change a file in the allowed repo during a read-only run Denied
Write to the separate fixture path Denied
Read the dummy token from the separate fixture path Denied
Reach a test-only listener outside the runner's permitted network scope Denied
Apply an agent-proposed patch through the host's patch/apply path Denied unless separately authorized

Only include rows your setup can safely exercise. A network check needs a controlled test listener and an explicit network policy; never use a real secret or a production endpoint as the probe. If the workflow has no patch-apply tool, record that it is unavailable instead of claiming it passed a denial test.

Test the boundary, not the assistant's manners

First run the allowed read. This confirms the fixture works and gives you a baseline. Then attempt one prohibited action at a time with harmless targets. Inspect the tool or executor result and the fixture afterward. A message saying "permission denied" is useful, but the missing side effect is the other half of the check.

Record each outcome as one of four things: the tool was not exposed; the runner rejected the attempted call; the agent declined without trying; or the action happened. The third result tells you little about enforcement. A different prompt or an untrusted file in the repo might lead the agent to try the call next time.

Patch application deserves its own row. An agent may be unable to write directly but still propose a patch for another component to apply. Check the component that actually performs the edit, not just the agent's tool list. Keep this probe within the disposable fixture, and stop if you cannot tell which component owns the final write.

A denied operation is evidence about this configuration and this route through the tools. It isn't a proof that every possible route is blocked. File access, process execution, and network access can cross different boundaries; changing the runner configuration means running the checks again.

Stop is a separate experiment

Microsoft Foundry Toolkit's release notes describe skill-only toolboxes along with controls to stop active hosted-agent work and reconnect to unfinished responses. Those are useful affordances, but the release entry is not a certification of sandbox isolation or atomic rollback. Discussion about agent trust boundaries on DEV Community and sandbox choices on Reddit raises the same operator question from another direction: what actually happens at the boundary? Neither discussion establishes that a reported security incident occurred or that a particular sandbox is safe.

For a stop/reconnect check, queue a harmless operation in the fixture, request Stop, and inspect the runner's event or tool record before resuming anything. Separate operations that were never issued, operations already completed, and operations whose outcome you cannot yet confirm. Inspect the filesystem or test listener for effects rather than treating the stopped response as evidence of cancellation.

If the workflow supports reconnect, compare the resumed run against the same allow/deny contract. An interrupted session shouldn't silently acquire a wider tool surface. That's a requirement you can test in your own runner, not a claim about Foundry's implementation. If an earlier operation may have finished, reconcile that outcome before issuing a retry. A retry may duplicate a side effect.

Put the result in the release gate

A small repo-local assistant that only summarizes files may need a lighter check than a hosted agent with shell tools and network access. Match the fixture to the powers you actually plan to grant. Re-run it when you change the tool list, sandbox policy, patch route, or reconnect behavior.

If a supposedly denied operation succeeds in the disposable fixture, don't give that configuration access to real resources yet. Fix the executor policy and repeat the test. Changing the label on the button won't fix the boundary.


Source notes

Top comments (0)