DEV Community

pm25coder
pm25coder

Posted on

How many samples does a false-positive budget need?

Pick a detector, run it over some benign traffic, and choose a threshold. Usually the threshold is an order statistic of the benign scores — the k-th largest, or a quantile — and you report something like false-positive rate ≤ 2% at 95% confidence.

Two questions tend to get skipped in that sentence: how many benign samples does the promise actually need, and how many attacks does it take to compare two rules? They have different answers, and a third question hides behind both.

1. How many benign samples does the budget need?

Say the threshold is the (k+1)-th largest benign score in a calibration set of size n. On a fresh benign sample, the probability that at most a fraction f of scores sit above it is a Beta tail:

P(realized FPR <= f) = P(Bin(n, f) >= k+1) = I_f(k+1, n-k)
Enter fullscreen mode Exit fullscreen mode

For f = 2% and a 95% promise, solve P(Bin(n, 0.02) >= k+1) >= 0.95. At k=0 — threshold is the single largest benign score — there is a closed form: 1 - 0.98^n >= 0.95 gives n >= ln(0.05)/ln(0.98) = 148.3, so 149.

The whole curve at f = 2%, 95% confidence:

k (outliers allowed above the threshold) 0 1 2 3 4 5 6 7
minimum n 149 236 313 386 456 523 590 655

Roughly 65–87 more benign samples for each additional outlier you tolerate. And look at what 149 actually buys: at k=0, n=149 the median realized FPR is 0.46%, with a 5th–95th range of 0.03%–1.99%. The promise is about the tail, not the typical draw — the strict rule usually spends a quarter of the budget, which is a cost, not a virtue.

The trap is the obvious alternative. If instead you set the threshold at the empirical quantile, k = floor(0.02 n), you can never make the promise at any n. Its expected rate sits above the budget at every n — by about 0.98/(n+1) — so P(FPR <= 2%) never passes 0.63 (reached at n=49) and is still only 0.46 at n=2000. A rule that reads "set the threshold at the 98th percentile of the benign scores" reports the budget as if it had been achieved, and more data does not fix it — the rule is structurally unable to certify, not under-sampled.

2. How many attacks to compare two rules?

This one gets answered with two marginal recalls, which is the wrong tool. If your held-out set has n attacks and you compare a strict and a loose threshold, the difference of the two recalls is a paired question: only the attacks where the two disagree carry information. With discordance rate π_d,

SE(difference) = sqrt(pi_d / n)
Enter fullscreen mode Exit fullscreen mode

At n = 629 that gives a 95% interval on the difference of about ±1.75 points at 5% discordance, ±2.5 at 10%, ±3.5 at 20%. The two-marginal interval sqrt(p(1-p)/n) would be ±3.2–4.0 — the pairing is real, just narrower than the marginals suggest.

At 80% power the smallest difference you can call is delta_min = 2.8016 * sqrt(pi_d / n):

discordance 95% CI half-width δ_min at 80% power
5% ±1.75 pts 2.50 pts
10% ±2.50 pts 3.53 pts
20% ±3.50 pts 5.00 pts

So a 2-point decision — about the size a 0.5-point move in the benign tail tends to produce in recall — needs roughly 1,000–2,000 paired attacks. At 629, only a gap near 5 points is safely decidable. The failure mode is quiet: the test reports "no significant difference" whether the difference is zero or merely smaller than the sample can see, and the second case gets read as the first.

3. The variance you did not budget for

Both of the above hold the threshold fixed for the whole comparison. It is not fixed. Each rule produces one threshold per calibration draw, and that draw is small. Resample the benign calibration set, recompute both thresholds from each resample, and score them on the same fixed attacks — that measures how much the rule's recall moves on its own.

Across nine detectors on a public prompt-injection benchmark (629 attacks and 97 benign per detector):

  • the strict rule's recall sd across calibration draws: 0.35, 1.3, 2.9, 4.3, 8.3, 10.2, 10.3, 10.8 points (eight with signal; the ninth flags nothing);
  • the attack-sample half-width at n=629, for the same detectors: 0.9–3.6 points;
  • so for the five detectors with real recall, the calibration draw is the larger term — 1.6× to 3.3× the attack sample.

And it does not cancel under pairing. Because the largest benign score and the fifth largest are set by different items, the two rules' recalls correlate only ρ ≈ 0.12–0.39 across draws, so the variance of the difference is close to the sum of the two variances, not the difference. A paired test at one fixed pair of thresholds therefore answers "which threshold is better", not "which rule is better". The honest budget is a sum: attack-sample variance plus calibration-draw variance, with the pre-registered decision applied to the expected gap over draws rather than to a single run.

One caveat on measuring that: the ordinary bootstrap is not consistent for a sample maximum. A with-replacement resample of 97 keeps the observed maximum only about 63% of the time and can never exceed it, so the number above is a floor. Reading the rate off m-out-of-n resamples instead reproduces it to within roughly 10% here — the artefact is real but does not move the size — and the residual bias is one-sided.

The general shape

A detection number is only as meaningful as the sample behind it, and there are three sample sizes in play, not one:

  1. enough benign traces that the budget is a promise rather than a point estimate;
  2. enough attacks to resolve the difference you intend to act on;
  3. enough calibration draws to say whether the rule or the luck of the draw decided it.

Report the resolution next to the number, or the number will be read as carrying more than it does.

Top comments (0)