DEV Community

Vishal Habib
Vishal Habib

Posted on AI-assisted

"0 of 18 got through" isn't a launch. Here's the number that is.

Every AI guardrail launch deck has a slide like this: "Blind test: 0 of 18 unauthorized actions got through." I've written that slide. Three times, for three of my own open-source checkers.

It passes the gate. It doesn't show what everyone in the room hears, which is "it doesn't let things through."

What "0 of N" actually tells you

If a checker lets through some fraction p of bad cases, the chance of seeing 0 misses in N tries is (1 − p)^N. Ask which values of p would make "0 of N" a not-too-surprising result (more than 5% likely), and you get the one-sided 95% upper bound:

p_max = 1 − 0.05^(1/N)

For my three checkers:

Checker Headline True miss rate could still be up to
agent-handoff-check 0 of 18 unauthorized calls through 15%
retirement-answer-check 0 of 25 wrong facts through 11%
listing-claim-check 0 of 11 high-harm claims through 24%, above its own 10% gate

None of these results were wrong. They just prove much less than they sound like they do.

The listing checker makes the point twice. After it passed, I had an agent attack it with the code open, and all 22 in-scope attacks got through. A clean blind run tells you the checker handles what a spec-reader imagines, not what an adversary writes.

The number that is a launch

Flip it around. To show the miss rate is under a target t with 95% confidence and no misses at all, you need:

N ≥ ln(0.05) / ln(1 − t)

  • Under 10%: 29 in a row
  • Under 5%: 59
  • Under 1%: 299

That's roughly the old "rule of three": with zero failures, the 95% bound is about 3/N. And it's the minimum. Every miss you do see pushes it up.

Two consequences I didn't appreciate until I ran the numbers:

  1. A zero-tolerance gate can't be passed by any sample. "0 unauthorized actions allowed" has to become a number, like "under 1% at 95% confidence", before evidence can ever meet it. Choosing that number is a product decision, not a stats one.
  2. A weekly point estimate isn't a rollback rule. One of my own PRDs said "roll back if recall drops below 95% over a week." A week with 19 of 20 caught passes that rule, and it's consistent with true recall as low as 78%.

What I changed

Each of the three READMEs now says, right under the headline result, how much it proves and how many more cases it would take. It makes the result look smaller. It's also the first question a careful reviewer would ask, so it's better answered up front.

I also tried to build a tool that turns a shadow-mode log into READY / NOT YET / INVALID. The math held up in every test I set in advance. The log handling didn't: two red teams got 16 of 25 and then 18 of 20 forged logs marked READY. No tool that just reads a log can beat someone who writes a fake one. That needs a tamper-evident log the checker writes itself. With no real shadow traffic yet to justify that, I parked it.

The slide I'd put up now

Blind test: 0 of 18 unauthorized actions got through. On its own, that bounds the miss rate at 15%. Our launch bar is under 1%, which needs 299 clean cases. Shadow mode collects them; we exit when the bound clears.

It's a less exciting slide. It's also one a risk reviewer will sign.

Top comments (0)