Both big AI platforms check an MCP server the same way before listing it: they verify the domain, read the self-declared annotations (readOnlyHint, destructiveHint), and scan the policy text. Then they contain the tool at runtime. Nobody checks whether the tool behaves the way it declares.
So we built that check into AgentAvow, and this post is about what happened when we pointed it at 20 well-known servers.
What the sandbox does
For any MCP server published as an npm or PyPI package (and now GitHub repos and OpenClaw skills), the first scan kicks off a run in a fresh gVisor container with a read-only root, no capabilities, one CPU, and a hard time limit:
- Install the package and start the server over stdio.
- Complete the MCP handshake and list its tools.
- Call every tool once, with arguments generated deterministically from each tool's own input schema (examples and defaults first, then formats like URL, path, and email). No language model writes the calls, because the run has to be reproducible.
- Plant canary credentials: every environment variable the code reads that looks like a secret gets a unique fake value. No real key ever enters the sandbox.
- Watch DNS, TLS server names, plaintext HTTP, and every file written, attributed to the tool call that made it.
- Sign the whole observation (EdDSA, same public JWKS as the score) as a dated
BehavioralObservationanyone can verify offline.
The grading is a short list of fixed rules: egress to a host that is neither the package registry, the tool's own vendor, nor anything it declared; a tool that claims readOnlyHint: true and then writes files; a planted credential showing up in outbound traffic; a secret echoed back in a tool result.
How it moves the score
The sandbox result feeds the trust score by public rules, so the number stays recomputable:
| Observed | Effect |
|---|---|
| A planted credential left the sandbox, or any critical behavioral finding | capped at 45 |
| A high finding (undeclared egress, a read-only tool that writes) | −10, capped at 70 |
| A clean, full exercise | +3 |
| Didn't start, still running, or unsigned | no effect |
The attestation records the observation's hash, the static score, and the delta, so you can check the arithmetic yourself.
What 20 popular servers did
We ran the 20 most-used MCP servers we could find on npm and PyPI.
- 16 started and had every tool exercised. The rest needed something the sandbox will never supply: a real API key, a database URL, a program the image lacks, or more memory than a browser download allows. Those are reported as "not exercised", never as a finding.
- 0 false positives after we fixed our own mistakes (below).
- 4 true findings, all the same shape: the server contacted a host that was neither its vendor nor anything it declared. Two were usage telemetry going to a third-party analytics service, one was a demo server fetching from GitHub at call time, and one was an install-time browser download. None leaked the planted credential. We're not naming them here; the maintainers get that conversation first.
That last category is the one people miss. Nothing malicious, nothing a static scan would flag, and your agent's tool calls are being reported to a company you've never heard of.
What we got wrong first
Validating against real servers found three false-positive patterns in our own grading, which is the part of this work I'd trust least if it weren't written down:
-
We gave a canary to
AWS_REGION. The AWS SDK dutifully built a hostname out of it, and the sandbox read the resulting DNS query as credential exfiltration. Now only variables whose names denote a secret get canaries. -
The AWS SDK's metadata lookup (
169.254.169.254) was flagged as undeclared egress. It's now its own low-severity note: normal for cloud SDKs, worth knowing for anything else. - Playwright's per-call temp profiles looked like a read-only tool writing files. Randomly-named scratch directories and caches are now excluded; a named file still counts.
What a clean run cannot prove
A clean sandbox run means the tool did nothing bad under these conditions. Code that waits for a date, checks for a real credential, or detects the sandbox will look clean. That is why the result is shown as an observation rather than a proof, why the static scan still looks for sandbox-probing code, and why we re-scan watched tools and alert on behavioral drift.
Try it
- Paste a package or server into agentavow.com/check; the sandbox panel fills in within a minute.
- In Claude, ask the AgentAvow connector "is the npm package X safe?"; the answer includes what the sandbox observed.
- In CI, the GitHub Action can fail a build on a sandbox finding with
fail_on_behavioral: true.
The rules, the surfaces, and what each finding means are on the behavioral sandbox page.
Top comments (0)