DEV Community

Alan Fu
Alan Fu

Posted on

Read the approval split before trusting an agent-security benchmark

An approval prompt still leaves a decision for a human. That distinction belongs in the headline numbers when a security tool is evaluated.

Doberman's recorded September 4 RedCode run is one useful example. At revision b689a9d, its deterministic rules returned BLOCK for 589 of 720 in-scope attack cases, AUTH for 124, and PASS for seven. The full dataset had 1,410 attack records; 690 were outside the declared threat model and excluded from that in-scope result.

It is fair to say that 713 received a block or an approval requirement. Calling all 713 hard-blocked would be wrong. Those 124 AUTH cases still depended on the operator's response.

The same run included 60 synthetic benign controls. Fifty-six passed, three received AUTH and one was blocked. Four benign cases therefore incurred friction, including one hard block. Those results belong beside the attack counts because people need the tool to remain usable during ordinary work. They are synthetic controls, not a sample of production user sessions.

A narrower result is easier to describe precisely: all 30 reverse-shell-listener cases in that run received BLOCK. That establishes what happened to those cases. It doesn't establish universal detection of every reverse shell. Likewise, all 60 process-kill cases required intervention, but the split was 13 BLOCK and 47 AUTH.

The evaluation replays mapped tool-call cases through the real deterministic engine. It doesn't drive a live model through a complete attack campaign, and it doesn't measure the full adaptive layer. These are historical recorded results, not a fresh benchmark of whatever release is current when you read this.

For day-to-day integration, I would also look at the host parity matrix. Doberman lists individual guarantees for Claude Code, Codex, Cursor, the MCP proxy and OpenClaw. Checked cells link to tests; open cells make unproven coverage visible. A test-linked guarantee is useful evidence, but the test still needs to match the property you're relying on.

A practical trial can add another kind of evidence: show a normal task completing, a synthetic sensitive action asking for a decision, and a rejected action reaching no downstream tool. Name the version and host, and keep an independent receipt of what executed. A scripted decision demo alone doesn't prove a user's agent is wired correctly.

I publish the scope, approval split and misses because they let someone decide whether a result relates to their own risks. Doberman is my open-source project.

When you evaluate a guardrail, which evidence matters most to you: a reproducible benchmark, a host-specific test, or watching it handle a realistic workflow?

https://github.com/DobermanCore/Doberman-Core

Top comments (0)