DEV Community

Your AI guardrail is green. It's also catching nothing.

Rudratosh Shastri on September 30, 2026

There are three ways a security guardrail can fail you. Two of them are loud. One of them is a serial killer. It's down. Crashed, misconfigured...
Collapse
 
pm25coder profile image
pm25coder •

Ran bench/at_budget.py over your own saved scores (bench/results/at_budget_2pct.json) before writing this, because the port is what makes the rest checkable — all nine detectors reproduce both saved columns exactly (in-sample and cross-domain) once the suite rotation is ordered workspace/travel/banking/slack, so the numbers below are yours, not a re-derivation.

Three things the pooled figure doesn't say.

1. The 2% budget is an in-sample property; the unseen column is a different quantity. caught_at_budget enforces the budget when picking (allowed = floor(budget * len(calib)), per fold on the calibration split), then measures on the held-out suite, where it can exceed it — and does, for eight of the nine (the regex baseline is the exception): 5.2% for Prompt Guard 2 86m, 13.4% for the 22m, 9.3% for testsavant. So "99% at a 2% false-alarm budget" splices an in-sample budget onto a cross-domain TPR. Both are honest numbers; the sentence that gets screenshotted has combined two.

2. The pooled TPR hides a fold spread as large as the rate it reports. Per held-out suite, prompt-guard-2-22m is 100% on travel, 22.9% on workspace, 16.0% on banking, 1.0% on slack — pooled 35%, spread 99 points. jailbreak-detector-large runs 17.1% → 100% (pooled 51%). For six of the nine the spread is at least the pooled rate itself. cross_domain already has each fold in c and accumulates it away (caught += c); keeping the list and printing min/max is two lines, and it changes what "the number to trust" means.

3. For a control, the worst fold is the number to trust — not the mean. That's your own thesis one level down: a pooled average is a liveness statistic (it says the detector ran on four suites), while the min fold is closer to an effectiveness one. Prompt Guard 2 86m is the reassuring case — 97–100%, a 3.3-point spread — and most of the field isn't.

The benign-sample size and the rotating-canary points above are the other half of it; both are cheap.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

This is exactly the kind of comment I was hoping this post would attract.

You’re right on the 2% point — I compressed two different quantities into one sentence. The budget is enforced during calibration, while the reported held-out/cross-domain catch rate is a different measurement. Putting those side by side without making that distinction explicit makes the headline number easier to screenshot than to interpret.

The fold spread point is even more important. A pooled TPR can look reassuring while one domain is basically a miss. I like the idea of keeping the per-fold values visible and reporting the worst fold alongside the pooled number.

That actually fits the thesis of the article better: the number isn't the measurement unless you preserve the conditions under which it was produced.

I'm going to fix this rather than defend the original wording. Thanks for actually running the code and checking it against the saved results.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

This is a really good catch.

I was treating “0 false positives out of 97” as if it established a 2% false-alarm rate, when statistically it really only tells us what happened in that small sample.

Your point about the benign denominator is exactly the kind of thing that gets lost when a benchmark focuses on the attack-success column. At a threshold this close to the benign-score distribution, the uncertainty around the benign sample can matter more than another decimal place on TPR.

And I really like the held-out rotating-canary idea. A fixed canary can gradually become a memorization test instead of a security test.

So the revised measurement should probably report both:

effectiveness on the held-out pool + confidence around the benign false-alarm estimate.

That’s a much more useful production story than simply saying “99% caught.”

Great addition.

Thread Thread
 
pm25coder profile image
pm25coder •

On the min-fold question — same re-run, and the answer changed twice while I was measuring it.

Headline pair: yes. But attach n and an interval to the min fold. Min-fold is by construction the smallest numerator, so it is also the widest interval. Your min folds, Wilson 95%, ordered: 97% [94–98], 26% [19–34], 17% [12–24], 1% [0–5], then five of the nine bottoming out at 0/140–240 → [0–3%]. A bare "17%" reads like a measurement; "17% [12–24], 140 attacks" reads like one.

The surprise: the min-fold leaderboard is mostly a tie. Walking the ranked min folds, 6 of the 8 adjacent pairs have overlapping 95% intervals — Prompt Guard 2 86m is cleanly separated at the top, and below third place the ordering is not resolvable at these n (four detectors are all 0% [0–3%]). So print the interval with the rank, or replace the rank with the interval below the top — a rank implies a separation the data doesn't have.

On readability: the per-suite row is four numbers — that is not what makes it unreadable. Put the 4-cell row under the headline and the full table in results/*.json; what actually costs a reader is three quantities with different denominators (budget / false-alarm rate / TPR) sitting next to each other unlabelled. Label the denominators and you can afford the detail.

One more, same re-run, and it cuts the same way: at this corpus size the 2% budget is degenerate. allowed = floor(0.02 * len(calib)), and calib is 57/77/81/76 benign cases depending on which suite is held out → 1 in all four folds. So cross-domain, "2%" is not a rate at all; it is "exactly one benign case above the line". Reporting the budget as an allowed count until the benign corpus grows (your ~150+ fix) makes that visible from the header — and it is a third reason those two numbers shouldn't share a sentence.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

a bare "17%" reads like a measurement; "17% [12–24], 140 attacks" reads like one

Taking all four together, they collapse into one correction I should have made myself: every headline in that benchmark is a percentage printed over a small integer, and a percentage over a small integer implies a resolution the count doesn't carry. That's the post's own thesis one level down — I stripped the conditions (the denominator, the n) that make a number a measurement, and kept the number.

Point by point, because each fix is different:

Min-fold gets an interval and an n, always — x% [lo–hi], N per cell. The interval isn't decoration here; it's the thing that stops a min-fold (smallest numerator, widest interval by construction) from being read as a point.

The ranking below the top isn't real, so it goes. This is the one that stings and it's the most important. If 6 of 8 adjacent pairs overlap at 95%, the leaderboard is asserting an order the data can't resolve — and a rank is itself "green by construction": it looks like information and encodes none below Prompt Guard 2 86m. So the honest output is a partial order — one detector cleanly separated at the top, and an unranked pack of "indistinguishable at this n" (four literally 0% [0–3%]). Intervals, not positions, below first place.

The unreadability was unit-collision, not count. You're right — four numbers is fine; three different kinds of rate (budget / false-alarm / TPR) sitting unlabelled side by side is the cost. Label the denominators and the detail pays for itself: 4-cell row under the headline, full table in results/*.json.

The 2% budget is an integer wearing a percentage. floor(0.02 × {57,77,81,76}) = 1 in every fold — so "2%" is "exactly one benign over the line," and printing it as a rate invents precision the calibration split doesn't have. It reports as an allowed count until the benign corpus clears the size where the percentage means anything.

The question that follows is the one I actually don't know: below the top, is the pack resolvable at any feasible n, or is "these are indistinguishable" the real scientific finding? At the min-fold widths you listed, separating two detectors ~10 points apart needs the attack corpus to grow a lot — and if ranking the pack is just noise-mining, the benchmark's honest output is "one detector clears the bar, the rest are tied," full stop. Did your re-run give you any feel for whether more attacks would ever break the tie, or is the pack genuinely level?

Collapse
 
arhancanli profile image
Arhan Canli •

Keeping the false-positive column in, and the honesty clause about the shared AgentDojo template, is what makes the 1% → 99% result believable.

Two things I'd add to the checklist.

The false-alarm budget needs enough benign samples to back it. If "wrongly flags at most 2%" is checked on something like the 97 benign outputs, the data can't really say 2%: with 0 flags in 97 the 95% upper bound on the false-alarm rate is about 3.7%, and with 1 flag it's about 5.6%. To show "at most 2%" with zero flags you need roughly 150 benign samples, and more if a few flags are allowed. At a threshold like 0.003, sitting that close to the benign scores, this is the number most likely to surprise someone in production.

On check 4, the known attack fired through the live guardrail: if it's always the same one, it can keep passing after the detector has drifted, for the same reason you flag with the 99%. It may be recognising that one string. Drawing the canary from a held-out pool of attacks that were never used to set the threshold, and rotating it, turns the armed check into a small ongoing measurement instead of a single fixed probe.

Collapse
 
arhancanli profile image
Arhan Canli •

@rudratosh thanks, and a small mix-up: your replies to pm25coder and me seem to have swapped places. The at_budget.py run and the three points about the splice, the fold spread and min-fold are pm25coder's; the benign sample size and the rotating canary were mine. Worth crediting them in the changelog under the right name.

On your canary question: I'd do both, at two speeds. Resample one canary from the held-out pool on every run, as you lean towards, and track the catch rate as a running count with its interval, so drift shows up as a trend. Then on every deploy or threshold change, run the whole held-out pool, which gives a proper effectiveness number at the moment it's most likely to break. And retire any canary that ever gets used to tune anything, or the pool slowly turns back into the fixed probe.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Absolutely — and thanks for the correction on attribution. I’ll make sure the changelog credits the at_budget.py analysis and the benign-sample/canary suggestions separately.

I like the two-speed approach.

For normal runs, rotate a canary from the held-out pool and track the catch rate over time. Then, whenever the detector is deployed or the threshold changes, run the entire held-out pool to get the real effectiveness measurement.

The “retire any canary that gets used for tuning” rule is especially important. Otherwise the canary quietly stops being a test and becomes another training signal.

That gives me a much better definition of what the live check should be:

the canary detects drift; the held-out pool measures effectiveness.

That distinction wasn't explicit enough in my original post. Really useful comment.

Collapse
 
aifrontierpost profile image
AI Frontier Post •

You report the 1%-to-99% swing on Prompt Guard 2 came down to one threshold number, 0.5 to 0.003, on the same weights. The detail worth underlining is where 0.5 came from: it's the demo default that most classifier wrappers inherit, chosen on a score distribution nobody's real traffic matches. We've seen the same shape in production — scores sitting at 0.009 versus 0.0008 look safely separated until attack phrasing shifts and the gap closes, which is why a one-time sweep doesn't hold. The armed check that survives it is re-running the calibration on fresh traffic, not just the canary.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

The armed check that survives it is re-running the calibration on fresh traffic, not just the canary.

This is the sharper version of my point, and I'm going to steal the framing: a one-time sweep calibrates to a distribution that's already expiring.

You nailed why 0.5 is so dangerous — it's not a value anyone chose for your traffic, it's the demo default that rides in with the wrapper and never gets questioned because the light stays green.

The bit you added that I under-weighted: the 0.009 vs 0.0008 gap isn't a fixed property of the model, it's a property of the current attack phrasing. Shift the phrasing and the gap closes, and a threshold you swept last quarter is now sitting in the dead zone again — silently, with every health check still passing.

So the real control isn't "sweep once, set threshold, done." It's:

  • Calibration as a recurring job, re-run on fresh traffic, not a one-time setup step
  • A live known-attack canary that asserts blocked, so drift trips an alarm instead of hiding
  • Watching the score gap itself as a metric — if attack/benign separation shrinks, that's your early warning before catch-rate craters

Out of curiosity — in your production case, what cadence did you land on for re-calibration, and was it time-based or triggered by the separation metric closing?