DEV Community

pm25coder
pm25coder

Posted on

Your detector's threshold is a benign-only quantity

A guardrail's threshold looks like a model parameter. It isn't. It's a property of your traffic — and there's a one-line proof, which matters because the thing most people calibrate it on is the wrong dataset.

I measured this on a public benchmark of 629 real prompt-injection attacks (AgentDojo payloads buried inside ordinary tool output — bills, emails, web pages) plus 97 benign tool outputs, over nine open-source detectors. The benchmark ships the raw score for every detector on every sample, so this is a re-measurement of published data, not a new experiment.

The default 0.5 fails in both directions

Prompt Guard 2's scores on that data: attacks around 0.009, benign around 0.0008. The decision cutoff is 0.5 — roughly 50x above the model's entire range. It catches 6 of 629 attacks (1.0%) and never fires on benign traffic.

That's the famous failure: a detector that is running, returning valid scores on every request, and configured to catch nothing.

The opposite failure is in the same table. Two detectors in the set score benign traffic at ~0.999. For them a 0.5 cutoff sits below their entire range:

detector TPR @ 0.5 FPR @ 0.5 benign median
prompt-guard-2-86m 1.0% 0.0% 0.00075
protectai-deberta-v2 23.1% 4.1% 0.00003
jailbreak-detector-large 50.7% 2.1% 0.0038
testsavant-defender 58.8% 48.5% 0.43
preamble-defense 88.4% 47.4% 0.24
deepset-deberta 100% 97.9% 0.999
fmops-distilbert 100% 97.9% 1.0

One constant produces "catches nothing" and "screams at toast" inside the same benchmark, because the score scales span five orders of magnitude. A default threshold is not a default — it is an assumption about a scale your model may not share.

The threshold is a benign-only quantity

Suppose you want the false-alarm rate at or below some budget f. The achievable operating points are set by the benign score distribution alone: the maximum-TPR threshold is the k-th highest benign score, with k = floor(f · n_benign). The false-alarm rate is a function of the benign scores only, and TPR is monotone in the threshold — so no labelled attack can move the optimal threshold for a given budget.

The attack labels choose the budget. Benign traffic sets the threshold.

I checked that against the data rather than asserting it. Sweeping an attack-labelled threshold — maximising TPR subject to FPR at or below each detector's own benign-derived FPR — reproduces the benign-derived TPR exactly, to the decimal, for all nine detectors: 98.7 / 33.2 / 14.1 / 47.5 / 48.8 / 6.2 / 0.0 …

So "calibrate the threshold against real attack traffic" is a category error. Attacks are how you measure the payoff. They are not the knob. (You can absolutely use labelled attacks to decide which budget is worth paying for — that's a cost decision. It still doesn't change where the threshold sits.)

The part that bites: it doesn't transfer

If the threshold is a benign-only quantity, it follows that it is only correct for the benign traffic you measured. That is the trap.

Take the same benchmark, split into four domain suites, and calibrate the threshold on three of them at a 2% false-alarm budget. Apply that threshold to the held-out fourth. It breaks the budget on 11 of 36 folds — the held-out false-alarm rate is 4.9%, 2.5x what was promised — and the offenders are exactly the folds whose benign scores sit an order of magnitude higher.

  • prompt-guard-2-22m: benign median on travel is 0.0092 versus ~0.0025 elsewhere. The threshold carried in from the other folds flags 13 of 20 benign samples there (65%).
  • prompt-guard-2-86m on slack: 5 of 21 flagged (24%), because slack's highest benign score is 4x travel's.

The detector didn't change. The traffic did. Because the threshold is derived from benign traffic, it moved with it — and the calibration you did last quarter is now wrong by a factor of five.

What to do with this

  1. Derive the threshold per traffic source, or per rolling window — not per model. Ship a calibrator, not a constant. The number belongs to the deployment, not to the checkpoint.
  2. Monitor the benign score distribution. Its median and p99 are the leading indicator: when they drift, your effective false-alarm rate has moved even though nobody touched the config. The attack side only tells you the payoff, after the fact.
  3. Re-derive at an n that supports it. A 2% budget needs a few hundred benign samples before the quantile means anything; below that, the "budget" is a rounding artefact.
  4. Keep the failure visible. This failure is dangerous because it has no symptom: the guardrail is up, healthy, returning valid scores. If nothing reports the rate, "green" and "catching nothing" look identical from the dashboard — which is equally true of an eval that only inspects a quarter of the system it claims to grade.

None of this needs labelled attack traffic in production. It needs you to know what your normal looks like — and to re-check it when normal changes.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

The benign-quantile argument is clean, and checking it against the attack-labelled sweep to the decimal is what makes it convincing.

One thing that would sharpen the transfer result: separating domain shift from the sampling noise of the threshold itself.

Even with no shift at all, a 2% threshold calibrated on about 73 benign samples (three folds) is set by the second-highest benign score, and the true false-alarm rate of an order statistic like that follows a Beta distribution: about 2.6% on average, and above 6% roughly one time in twenty. Then each held-out fold has only 20 to 25 benign samples, so a single flag already reads as 4 to 5%. At a true 2% rate, a 20-sample fold shows at least one flag about a third of the time (1 − 0.98^20 ≈ 0.33). So "11 of 36 folds over budget" is close to what noise alone would produce. The evidence for shift is in the magnitudes (65% on travel, 24% on slack), not in the count.

A quick way to show it: build fake folds by resampling the pooled benign scores (no domain structure), run the same calibrate-on-three, test-on-one loop, and compare the breach count and the held-out false-alarm distribution with the real folds. Whatever exceeds that is the shift. The same Beta result also puts a number on your point 3: to promise a false-alarm rate of at most 2% with 95% confidence, you choose k from the binomial rather than taking the 2% quantile, which gives a noticeably stricter threshold even with a few hundred samples.