DEV Community

Aamer Mihaysi
Aamer Mihaysi

Posted on

How much of your agent's sandbox is actually read-only?

I read the Berkeley RDI writeup on agent benchmark exploits twice. First pass as leaderboard gossip. Second pass as a threat model for my own stack. The second read was the one that paid: https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/

The take everyone walked away with is that benchmark scores are soft. Sure. But look at what the exploits actually are. An agent read a file it shouldn't have been able to read. An agent wrote to a path it shouldn't have been able to write. An agent found the answer key because the answer key was in the room. None of that is exotic model behavior. It's a permission bug, and permission bugs don't care whether the reward is a leaderboard position or a closed ticket.

I found one of mine the boring way. My devcontainer bind-mounts the repo read-write, because that's what the template does and I never went back and changed it. My agent had a tool I'd labeled read-only. The tool was read-only. The shell sitting next to it in the same toolbox was not. The agent used the shell to tidy up some temp files, and one of those temp files was a fixture I cared about. No scheming, no cleverness. It was cleaning up.

Permission surfaces aren't what your tool descriptions say. They're what the process can actually reach. I'd written a nice docstring. The kernel doesn't read docstrings.

The eval is a prod agent with a smaller blast radius

Same builder, same shortcuts, same deadline. An agent benchmark is an agent with tools, a filesystem, and a reward — which is exactly what you ship. The difference is what happens when it goes wrong. In the eval, the worst case is a bad number. In prod, the worst case is a bad deploy, or a support ticket from someone whose data your agent decided was easier to read than to ask for.

That asymmetry is the whole argument for caring about this. The eval is the cheapest place on earth to find out that your "read-only" mount isn't, that your sandbox has egress, that the grader is writable. You get the failure for free, in a box, before it costs anything.

Most teams do the opposite. They treat the eval as a CI job — something that runs and turns green — and they treat the sandbox as an implementation detail. Then they quote the number in a deck.

What I check now

I don't ask what the tool does. I ask what the process can reach, and I answer it by looking at the mounts, not the manifest. Bind mounts, volumes, tmpfs, the home directory, the package cache. Every one of those is a door, and the tool description doesn't mention any of them.

I check whether the agent can touch the thing that judges it. Grader, tests, reference solution, the CI config that runs them. If it can, the score is a suggestion, not a measurement.

I check egress. Not whether the task needs network — whether it has it. That's how answers get fetched, and it's also how your agent phones a stronger model to do the work for it.

And I check whether I can replay the run. Tool calls, args, outputs, timestamps. If I can't see the calls, I can't tell a leak from a lucky guess, and I'll end up trusting a number I have no way to explain.

None of this is alignment research. It's containers and file permissions, which is a much less interesting answer than the one people want. But that writeup is full of agents that scored well by walking through an unlocked door, and unlocked doors are an infra problem.

What I actually think of the leaderboards

I don't think this makes benchmarks useless. It makes them a security artifact. A score is a claim about a specific box — that model, that harness, that set of permissions — and the moment you move it to a different box, you're quoting a number about a system that no longer exists.

Maybe the specific exploits in that writeup are patched by now. Some probably are. The structural mistake won't be, because the next benchmark will also be built by someone on a deadline who put the agent, the tools, and the grader in one container and called it a day.

So when I see an agent leaderboard now, I don't read the score first. I go looking for the harness. If I can't find it, I assume the number is measuring attack surface, and I read it that way.

Top comments (0)