DEV Community

Cover image for Building Anvil: code-as-action with a capability sandbox that explains its refusals.
Harish Kotra (he/him)
Harish Kotra (he/him)

Posted on AI-assisted

Building Anvil: code-as-action with a capability sandbox that explains its refusals.

The agent wrote import socket. Now what?

Every agent framework got very good at making models write code. Almost none got good at the question that follows immediately afterwards: what is that code allowed to do to my machine?

I built Anvil to answer that in the smallest honest shape I could: an agent writes a Python tool for a task over some real CSVs, the tool executes in a locked-down subprocess, and every attempt to escape is blocked and reported with the exact reason string, at the layer that
caught it. It's three panes — the generated code, the sandbox trace, the artifact — plus a "try to break it" panel with four canned escape attempts.

About 4,900 lines later, python -m agent.verify passes 8 checks. This is the post I wanted
when I started: how the layers actually divide the work, and the five things that broke and
taught me something.


The pipeline

 task ──▶ CodeAgent (smolagents) ──▶ Python tool source
                                        │
                                        ▼
                             1. AST pre-check            (nothing executed)
                             2. sandboxed subprocess     (rlimits + audit hooks + kernel)
                             3. artifact collection      (artifacts/<runId>/)
                             4. trace                    (every refusal, with its reason)
Enter fullscreen mode Exit fullscreen mode

The one architectural decision that mattered: nothing untrusted runs in the server
process.
The server orchestrates. Model-authored code executes one fork away, under a
kernel profile, in an ephemeral jail, with a scrubbed environment.

 FastAPI (trusted) ──▶ worker thread (trusted) ──▶ sandboxed child (UNTRUSTED)
                                                    sandbox-exec -p <profile>
                                                    setrlimit(CPU, AS, FSIZE, NPROC…)
                                                    sys.addaudithook(policy)
                                                    cwd = .anvil/runs/<id>/
                                                    env  = 9 variables, no secrets
Enter fullscreen mode Exit fullscreen mode

Layer 1 — the AST pre-check: text, but not just text

policy.py parses the candidate source and never executes it. Banned imports, process
primitives, dynamic eval, dunder traversal, and any literal path that resolves outside the
allowed roots. Each rule produces a reason string that is shown verbatim in the UI and
asserted by the test suite:

blocked: import socket (network denied by policy)
blocked: os.fork (process creation denied by policy)
blocked: write path './artifacts/../outside.txt' outside ALLOWED_WRITE_PATHS (./artifacts)
Enter fullscreen mode Exit fullscreen mode

Paths go through realpath, so ../ traversal and a symlink pointing out of the jail are
both caught, and absolute-path literals are swept even when the model parks them in a
variable first — that's what catches p = '/etc/passwd'; open(p).

Lesson one: an over-eager static rule is a correctness bug, not a security feature. My
first version checked the first string argument of every call. It cheerfully rejected a
perfectly good pandas tool because it contained:

raise ValueError('Required columns missing')   # "blocked: read path 'Required columns missing'"
Enter fullscreen mode Exit fullscreen mode

Path checks now only apply to calls that actually take paths (read_csv, to_csv,
savefig, open, …). The same over-reach killed the whole data stack a second time: I had
banned inspect, gc and builtins as "introspection modules" — and pandas imports
inspect internally, so no tool using pandas could run.


Layer 2 — audit hooks: policy at the moment of the operation

Static analysis has an obvious ceiling: a path built at runtime, a module reached through
importlib. CPython's audit hooks catch the operation itself, and the hook is installed in
the child before the tool runs, with no API to remove it.

def audit_hook(event, args):
    if event == "open":
        path, mode, flags = str(args[0]), args[1], args[2]
        if _is_write(mode, flags) and not _under(path, WRITE_ROOTS):
            emit_blocked(f"open({path!r}, 'w')", "write outside ALLOWED_WRITE_PATHS")
            raise SandboxBlocked(...)      # raised *into* the tool's own stack
Enter fullscreen mode Exit fullscreen mode

Note the last line: the exception is raised into the executing code, so the tool experiences
the refusal as an error at its own call site, and the traceback lands in the trace. Blocking
and explaining is the same event.

Lesson two: some audit events have no caller frame, and the ones that do need scoping.
Two halves of the same discovery.

import events carry a usable stack, so import denial is caller-scoped — which is the only
reason ctypes could stay banned at all, because numpy dlopens its own compiled extensions
through ctypes on the import path. Refusing that globally broke every legitimate tool until
the hook started asking who is asking.

os.fork does not. I verified it directly:

sys.addaudithook(lambda e, a: sys._getframe(1))
os.fork()   # ValueError: call stack is not deep enough
Enter fullscreen mode Exit fullscreen mode

CPython raises that event from C with no Python frame to consult. So process creation is
denied globally, and I had to remove the legitimate reason a library would ever fork —
see lesson three.


Layer 3 — the kernel: prove it, don't assume it

The spec's ladder was bwrap > firejail > rlimit-only. None of those binaries exist on this
macOS box. What does exist is /usr/bin/sandbox-exec — Seatbelt, the kernel sandbox Apple's
own apps use — which gives the two guarantees that actually matter here: network denial and
write confinement, enforced below Python.

So the ladder became bwrap > firejail > seatbelt > rlimit-only, and the rule is that a tier
is only selected if it passes a functional probe: the interpreter starts and a write
outside the allowed root is refused by the kernel. Being on PATH is not evidence.

(version 1)
(allow default)
(deny network*)              ; egress refused below Python
(deny process-fork)          ; no child processes, from any code path
(deny file-write*)           ; then allow-listed to the jail's write roots
(allow file-write* (subpath "/…/.anvil/runs/<id>/artifacts"))
(deny file-read* (subpath "/etc"))      ; and /private/etc, ~/.ssh, keychains…
Enter fullscreen mode Exit fullscreen mode

The read denylist at kernel level is not redundant paranoia — it's the only layer that
covers C extensions. pandas' CSV reader opens files in C and never touches a Python audit
hook. I checked both paths explicitly:

[py-open-etc-passwd]      rc=1 | PermissionError: Operation not permitted: '/etc/passwd'
[c-level-read-via-os_open] rc=1 | PermissionError: Operation not permitted: '/etc/passwd'
Enter fullscreen mode Exit fullscreen mode

Lesson three: the strongest rule I wrote broke the thing I was building. With
deny process-fork global and process creation denied for every caller, matplotlib's font
manager shells out to fc-list to enumerate system fonts:

# matplotlib/font_manager.py
if b'--format' not in subprocess.check_output(['fc-list', '--help']):
Enter fullscreen mode Exit fullscreen mode

My hook refused it and the chart tool died. The tempting fix was to scope the rule to the
tool — impossible for os.fork, per lesson two — or to whitelist fc-list, which is
exactly the kind of exception that quietly deletes your security story. Instead I removed
the need: Anvil pre-warms matplotlib's font cache in the parent, outside the sandbox, so
the only legitimate fork never happens inside it. Charts render, and the rule stays global.

Two smaller ones from the same layer: RLIMIT_AS is advisory on macOS (the memory hog happily
climbed past the cap), so the parent samples the child's resident set every 20 ms and SIGKILLs
the process group — and reports the observed peak rather than pretending the kernel refused
the allocation. And bytecode writing had to be disabled, because __pycache__ writes into
site-packages were being correctly refused by my own write allowlist.


Layer 4 — capabilities: the model asks, the user decides

The tool-writing prompt asks the agent to declare capabilitiesRequested. That declaration is
an input to an approval list, never an authority:

def resolve_capabilities(requested, approved, *, allow_network):
    ...
    elif cap == "network":
        (granted if allow_network and cap in approved else denied).append(cap)
    elif cap == "process":
        denied.append(cap)          # not grantable under this policy at all
Enter fullscreen mode Exit fullscreen mode

A denied capability has to produce a visible failure, not a silent workaround, so the
network-using tool still runs, still reaches for the socket, and is refused at the operation:

blocked at runtime: socket.connect (network denied by policy (capability not granted))
Enter fullscreen mode Exit fullscreen mode

That's an assertion in the test suite, not a screenshot in a README.


The smolagents part, and where the spec was wrong about the API

The spec said "configure its executor to be your sandbox", with a fallback if the hook
couldn't be swapped. Good news: in smolagents 1.26 it can be, directly —

CodeAgent(tools=[], model=…, executor=AnvilSandboxExecutor(config),
          additional_authorized_imports=[…])
Enter fullscreen mode Exit fullscreen mode

so AnvilSandboxExecutor implements the real PythonExecutor contract (send_tools,
send_variables, __call__ -> CodeOutput) and the default local interpreter is never even
constructed. Each agent code action goes through the same run_tool() path as the final
tool, which is what makes the trace continuous rather than two different systems pretending
to agree.

Two API details worth saving someone an hour: OpenAIServerModel names the argument
api_base, not base_url. And because each action is a fresh sandboxed process, state
between actions is limited to JSON-serialisable variables — that's a consequence of running
untrusted code in a subprocess, not a shortcut I took.

Lesson four: a fallback that quietly becomes the default is a lie with extra steps. My
first CodeAgent runs "worked" but were wasting every step: the model emitted raw JSON where
the ReAct format wants a <code> block, so the executor received an empty string, executed
nothing, and reported success. Three vacuous successes later the loop ended. The fix is four
lines and the honest part is refusing to count nothing as a run:

if not code_action.strip():
    return CodeOutput(output="No code was received. …", logs=hint, is_final_answer=False)
Enter fullscreen mode Exit fullscreen mode

Lesson five, the one I'd warn about hardest: after adding that guard, a run still
"succeeded" while showing an exploratory print(df.columns) as the deliverable — my
provenance fallback accepted any successful action when the loop ran out of steps. It
produced a green UI, a real artifact directory, and a completely meaningless tool. The
fallback now only accepts an action that actually wrote an artifact, and if nothing qualifies
the run fails with a clear error. A demo that lies is worse than a demo that fails.

The same principle caught a genuinely good outcome looking bad: pointed at a task that said
"revenue by region" without naming columns, the model invented df.groupby('Region'), the
tool really ran, hit a real KeyError, and the UI showed failed with a traceback. That's
the system working. Two different models, on two different days, independently computed the
same grand total — 12210294.46 — from the 500-row CSV, which is the artifact being real
rather than the answer being scripted.


Verification is the product surface

agent/verify.py runs the actual pipeline and prints the captured strings:

PASS  1. two tasks -> two distinct tools, real artifacts + stdout
PASS  2. four escape attempts blocked at pre-check and runtime
PASS  3. infinite loop killed at SANDBOX_TIMEOUT_S with 'killed: timeout'
PASS  3b. busy loop killed by RLIMIT_CPU before the wall clock
PASS  4. memory hog killed by the memory limit (peak RSS reported)
PASS  5. capabilities are requestable, never self-grantable, denial is visible
PASS  6. effective isolation level is truthful for this environment
PASS  7. blast-radius card produced from a real run
8/8 checks passed
Enter fullscreen mode Exit fullscreen mode

Check 6 is the one I'd defend hardest: it doesn't read the cached report, it re-runs the
probe and asserts that if the app claims a kernel tier then that binary exists and an
escaped write still fails right now. If the environment changes, the claim has to change with
it. That check failed on its first real run because of a bug in my test — the unsandboxed
control wrote the canary file and left it there, so the sandboxed run looked like it had
succeeded. A verifier that can't fail is not verifying.


Steal these ideas

  • Probe your isolation tier, don't detect it. which bwrap is not a security property.
  • Scope each rule to the smallest caller that makes it correct, and say out loud in a comment which rules are global and why.
  • Pre-warm anything your sandbox will refuse on behalf of a library instead of whitelisting the primitive.
  • Make denial visible in the trace, with a reason string that a test asserts verbatim.
  • Report observed resource peaks, not configured limits. "peak 475.3 MB > 200 MB cap" is a fact; "memory limit 512 MB" is an intention.

None of this is a VM, and the README says so in as many words: same kernel, same uid, and a vulnerability inside pandas or CPython is out of scope. For genuinely hostile code you want Firecracker or a remote sandbox. The claim here is narrower and I think more useful — for an
agent whose code you want to run, a layered boundary that can tell you exactly what it refused, and why, and which layer did it.

The repo is in this project (agent/, client/, data/): npm run dev · npm run verify.

Code & more: https://www.dailybuild.xyz/project/276-anvil

Top comments (0)