Every AI guardrail launch deck has a slide like this: "Blind test: 0 of 18 unauthorized actions got through." I've written that slide. Three times, for three of my own open-source checkers.
It passes the gate. It doesn't show what everyone in the room hears, which is "it doesn't let things through."
What "0 of N" actually tells you
If a checker lets through some fraction p of bad cases, the chance of seeing 0 misses in N tries is (1 − p)^N. Ask which values of p would make "0 of N" a not-too-surprising result (more than 5% likely), and you get the one-sided 95% upper bound:
p_max = 1 − 0.05^(1/N)
For my three checkers:
| Checker | Headline | True miss rate could still be up to |
|---|---|---|
| agent-handoff-check | 0 of 18 unauthorized calls through | 15% |
| retirement-answer-check | 0 of 25 wrong facts through | 11% |
| listing-claim-check | 0 of 11 high-harm claims through | 24%, above its own 10% gate |
None of these results were wrong. They just prove much less than they sound like they do.
The listing checker makes the point twice. After it passed, I had an agent attack it with the code open, and all 22 in-scope attacks got through. A clean blind run tells you the checker handles what a spec-reader imagines, not what an adversary writes.
The number that is a launch
Flip it around. To show the miss rate is under a target t with 95% confidence and no misses at all, you need:
N ≥ ln(0.05) / ln(1 − t)
- Under 10%: 29 in a row
- Under 5%: 59
- Under 1%: 299
That's roughly the old "rule of three": with zero failures, the 95% bound is about 3/N. And it's the minimum. Every miss you do see pushes it up.
Two consequences I didn't appreciate until I ran the numbers:
- A zero-tolerance gate can't be passed by any sample. "0 unauthorized actions allowed" has to become a number, like "under 1% at 95% confidence", before evidence can ever meet it. Choosing that number is a product decision, not a stats one.
- A weekly point estimate isn't a rollback rule. One of my own PRDs said "roll back if recall drops below 95% over a week." A week with 19 of 20 caught passes that rule, and it's consistent with true recall as low as 78%.
What I changed
Each of the three READMEs now says, right under the headline result, how much it proves and how many more cases it would take. It makes the result look smaller. It's also the first question a careful reviewer would ask, so it's better answered up front.
I also tried to build a tool that turns a shadow-mode log into READY / NOT YET / INVALID. The math held up in every test I set in advance. The log handling didn't: two red teams got 16 of 25 and then 18 of 20 forged logs marked READY. No tool that just reads a log can beat someone who writes a fake one. That needs a tamper-evident log the checker writes itself. With no real shadow traffic yet to justify that, I parked it.
The slide I'd put up now
Blind test: 0 of 18 unauthorized actions got through. On its own, that bounds the miss rate at 15%. Our launch bar is under 1%, which needs 299 clean cases. Shadow mode collects them; we exit when the bound clears.
It's a less exciting slide. It's also one a risk reviewer will sign.
Top comments (0)