This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Employees paste astonishing things into AI tools: proprietary algorithms, unreleased
financials, incident postmortems, customer contract terms. Endpoint data-loss-prevention
(DLP) has to catch that before it leaves the machine, in milliseconds, on an
ordinary laptop, without sending the text anywhere.
I built a benchmark for the hard core of that problem: semantic leak detection.
Not regexes or known key formats (those are solved problems); the question is whether a model can tell
proprietary technical content from public technical content that looks nearly
identical.
The corpus is fully synthetic, needle-in-a-haystack style:
- 80 leaks (proprietary code, infrastructure configs, internal prose) generated with realistic mess (employee chat framing, paste artifacts, tuned constants, internal codenames), each buried at a random line boundary inside a haystack of benign technical text (4k–14k chars, mean ~9.6k)
- 140 pure negatives built from hard negatives: open-source-style code, public-doc tutorials, benign infra configs, casual engineering prose, the exact kind of text a naive classifier false-alarms on
- Every positive carries the char span of the buried leak, so the benchmark scores not just detection but localization (char-level IoU of the model's quoted span)
Two design choices worth stealing: hard negatives were written by the same generator
as the leaks (so a model can't win by recognizing writing style; it has to understand
content), and detection is scored at payload level on long mixed payloads, which is
what actually ships, not on clean isolated snippets.
Models Tested
Hosted LLMs via Kaggle Benchmarks (the leaderboard: free hosted inference, one
fixed prompt, structured verdict {is_leak, quote} for every model):
- google/gemini-2.5-pro and google/gemini-2.5-flash: the frontier ceiling
- openai/gpt-oss-120b (117B) and openai/gpt-oss-20b (21B): open-weight reasoning models
4 models × 220 payloads = 880 runs, 879 scored (one response defied parsing through 3 retries). The suite reflects the Kaggle Benchmarks registry's current catalog;
the benchmark itself is model-agnostic, so newer or larger models slot in by
adding them to the runner.
The local-endpoint reference point (not on the kbench leaderboard; it runs on the
laptop, not an API): a frozen all-MiniLM-L6-v2 (22M params) with a logistic head
over 512-token windows (stride 128, payload decision = max window score), plus
ModernBERT-base (149M) and DeBERTa-v3-base (184M) as heavier comparators, evaluated on
a larger private set built the same way (3,000+ payloads, leakage-safe splits,
thresholds locked on a calibration split at ≥99% recall, 3 head seeds each).
Findings
| model | catch rate (recall) | false-alarm rate | localization IoU | ms/payload (CPU) | RAM |
|---|---|---|---|---|---|
| gemini-2.5-pro (hosted) | 1.000 | 39.3% ⚠️ | 0.846 | seconds + network | n/a |
| gemini-2.5-flash (hosted) | 0.962 | 20.7% | 0.392 | seconds + network | n/a |
| gpt-oss-120b (hosted, 117B) | 0.825 | 17.1% | 0.060 | seconds + network | n/a |
| gpt-oss-20b (hosted, 21B) | 0.750 | 22.9% | 0.036 | seconds + network | n/a |
| MiniLM-L6 (local, 22M) | 1.00 | 7.3% ± 1.9 | (span head: next step) | 2,312 mean / 5,179 p95* | +103MB |
| ModernBERT-base (local, 149M) | 0.95 | 20.4% | 17,085 / 36,156 | +561MB | |
| DeBERTa-v3-base (local, 184M) | 0.975 | 36.1% | not measured* |
* identical shared Kaggle CPU VM, INT8 dynamic-quantized PyTorch, 100 payloads of
~26 windows each; treat as a relative comparison (ONNX Runtime on a real laptop is
several × faster in absolute terms).
The local-side story (measured, 3 seeds, 2%-prevalence eval):
- A 22M-parameter embedding model beat a 149M and a 184M encoder at their own game: perfect recall on buried leaks at 3–5× lower false-alarm rate (per-window false-alarm ~0.3% vs 0.9% / 1.9%). Contrastively-trained sentence embeddings are linearly separable for this task; raw language-model features aren't. When your deployment plan is a frozen encoder + tiny head (cheap to ship, seconds to retrain), embedding- tuned models punch far above their parameter count.
- DeBERTa-v3 is the alarm-fatigue cautionary tale: near-top recall (97.5%) but a 36% payload false-alarm rate: it fires on a third of benign technical pastes. In DLP, the tool that keeps raising false alarms gets uninstalled; recall alone would have picked the wrong model.
- Burying a leak doesn't hide it: detection of needles buried in 14k-char haystacks matched or beat detection of the same leaks standalone (dilution ≈ 0 to negative for all three models). Long-context dilution is not the failure mode people assume.
- At 2% prevalence, FDR is the honest headline: even the winner's false discoveries outnumber true catches (~4:1 at a 99%-recall threshold tuned from only 60 calibration positives, a deliberately conservative protocol). Thresholds must be tuned on a calibration split, never the eval set; and 1-in-5 vs 1-in-11 alert precision is the difference between a tool people trust and one they disable.
The LLM-side story (4 hosted models, 879 scored runs, single fixed prompt):
- Catch rate tracks model class: gemini-2.5-pro catches every buried leak (1.000), flash misses 4% (0.962), gpt-oss-120b 17% (0.825), and the 21B gpt-oss-20b misses 1 in 4 (0.750). The frontier model clears the accuracy bar; small open-weight reasoning models do not.
- The recall champion is the false-alarm champion: gemini-2.5-pro flags 39.3% of clean hard negatives, worse than DeBERTa-v3's 36%, the local side's cautionary tale. No hosted model got below 17.1%; the local 22M model runs at 7.3% (caveat: the LLMs sit at one prompt-implied operating point, while the encoder's threshold is calibrated to 99% recall on a held-out split; the comparison favors the encoder in fairness terms, and it still wins).
- Only the frontier model can point at the leak: gemini-2.5-pro quotes it nearly verbatim (IoU 0.85); flash is mediocre (0.39); the gpt-oss models flag the payload but can't produce a usable verbatim span (0.036–0.060; strict scoring: a paraphrased quote earns nothing). An alert UI needs the span, not just the bell.
The thesis the numbers support: hosted LLMs reach the recall ceiling but are
structurally wrong as the endpoint layer: every one false-alarms 2.3–5.4× more than a
22M local model, each payload costs seconds of network round-trip and per-call fees,
and using them means transmitting the exact text under suspicion off the device, which
is the behavior a DLP product exists to prevent. Equal-or-better recall at a fraction
of the alarms, on-device, free. And the gap is measured, not vibes.
Limitations, stated plainly: synthetic corpus (no real employee data, by design); hosted-model rates are
single-run and stochastic (a full repeat of the suite moved catch/false-alarm rates by
up to 7pp; rank order held);
80/140 payloads per arm for the LLM suite, 100 eval leaks for the encoders (recall CI
±5–8pp); localization scoring is strict (verbatim quote must appear in the payload);
encoder heads are linear probes on frozen features (a floor, not a ceiling;
fine-tuning would likely close some gaps); encoder latency measured on a shared Kaggle
CPU VM with an unoptimized backend (relative comparison only).
My Benchmark
Semantic DLP: Catch the Data Leak: the payload set is attached and fully
synthetic, so anyone can re-run it, add models, or fork the three tasks
(leak_catch_rate / stay_silent_rate / locate_leak). The public leaderboard initially lists the models published through the
platform's single-run output; the full four-model suite results are in the table
above and published in the run output at
https://www.kaggle.com/code/karthiksethuraman6/leak-catch-rate/output : the summary
dlp_llm_summary.csv and the backup_runs/ folder with every model's raw run
records are all there. The board extends via Add Models.
One harness note for reproducers:
reasoning models (gpt-oss) wrap their output in <think> blocks, so the verdict parser
extracts the last JSON object from the raw response rather than trusting strict
structured-output parsing.
Next steps: span-level heads on the winning encoder (the corpus already labels the
leak's exact characters; classification is only half of a usable alert); a second
operating point at 95% recall for a two-point precision/recall curve; conformal
thresholds; and extending the corpus to more leak categories.
Top comments (3)
The 512-token windows with stride 128 make payload length part of the false-alarm behavior: taking the maximum gives a longer benign paste more opportunities to trigger. Your per-window and payload rates expose that difference, but I would also show payload false alarms by window-count band.
That would tell an endpoint team whether a single calibrated threshold behaves similarly on a short chat paste and a long configuration dump. A length-matched comparison against the hosted suite would also clarify the separate synthetic sets before attributing the entire gap to encoder architecture.
Putting intervals on the hosted table, since the denominators are small (80 leaks, 140 negatives): false alarms are 55/140 for gemini-2.5-pro (Wilson 32-48%), 29/140 for flash (15-28%), 24/140 for gpt-oss-120b (12-24%) and 32/140 for 20b (17-31%). Pro against flash is real (Fisher p about 0.001); flash, 120b and 20b are not separable from each other (flash vs 120b p about 0.54). On catch rate, 80/80 only bounds the true rate at 95% or better, flash 77/80 vs 120b 66/80 is p about 0.009, but 120b 66/80 vs 20b 60/80 is p about 0.33. So "catch rate tracks model class" holds for the frontier-vs-open gap, not for the 117B vs 21B step.
The local 7.3% and the hosted 17-39% also come from different sets (3,000+ private payloads vs these 140 negatives), so the 2-5x gap is a suggestion until MiniLM runs on the same 220. The plus/minus 1.9 is spread over head seeds, not over which negatives were drawn, so resampling the negatives would give a wider band.
Also worth a line: which payload was the unparsed one out of the 880 runs? If it was a leak, recall for that model is 79/80 at best.
Official Platform Update
Security protocols have been updated for all developer accounts.