<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: karthik sethuraman</title>
    <description>The latest articles on DEV Community by karthik sethuraman (@ksr).</description>
    <link>https://dev.to/ksr</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161737%2F410ace68-da3e-4447-a0ce-6cd7eb1523f8.png</url>
      <title>DEV Community: karthik sethuraman</title>
      <link>https://dev.to/ksr</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ksr"/>
    <language>en</language>
    <item>
      <title>Semantic DLP: Can AI Models Catch the Data Leak?</title>
      <dc:creator>karthik sethuraman</dc:creator>
      <pubDate>Sun, 11 Oct 2026 09:24:22 +0000</pubDate>
      <link>https://dev.to/ksr/semantic-dlp-can-ai-models-catch-the-data-leak-1dob</link>
      <guid>https://dev.to/ksr/semantic-dlp-can-ai-models-catch-the-data-leak-1dob</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Employees paste astonishing things into AI tools: proprietary algorithms, unreleased&lt;br&gt;
financials, incident postmortems, customer contract terms. Endpoint data-loss-prevention&lt;br&gt;
(DLP) has to catch that &lt;strong&gt;before it leaves the machine&lt;/strong&gt;, in milliseconds, on an&lt;br&gt;
ordinary laptop, without sending the text anywhere.&lt;/p&gt;

&lt;p&gt;I built a benchmark for the hard core of that problem: &lt;strong&gt;semantic leak detection&lt;/strong&gt;.&lt;br&gt;
Not regexes or known key formats (those are solved problems); the question is whether a model can tell&lt;br&gt;
&lt;em&gt;proprietary&lt;/em&gt; technical content from &lt;em&gt;public&lt;/em&gt; technical content that looks nearly&lt;br&gt;
identical.&lt;/p&gt;

&lt;p&gt;The corpus is fully synthetic, needle-in-a-haystack style:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;80 leaks&lt;/strong&gt; (proprietary code, infrastructure configs, internal prose) generated with
realistic mess (employee chat framing, paste artifacts, tuned constants, internal
codenames), each &lt;strong&gt;buried at a random line boundary inside a haystack&lt;/strong&gt; of benign
technical text (4k–14k chars, mean ~9.6k)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;140 pure negatives&lt;/strong&gt; built from &lt;em&gt;hard negatives&lt;/em&gt;: open-source-style code, public-doc
tutorials, benign infra configs, casual engineering prose, the exact kind of text a
naive classifier false-alarms on&lt;/li&gt;
&lt;li&gt;Every positive carries the char span of the buried leak, so the benchmark scores not
just &lt;em&gt;detection&lt;/em&gt; but &lt;strong&gt;localization&lt;/strong&gt; (char-level IoU of the model's quoted span)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two design choices worth stealing: hard negatives were written by the &lt;em&gt;same generator&lt;/em&gt;&lt;br&gt;
as the leaks (so a model can't win by recognizing writing style; it has to understand&lt;br&gt;
content), and detection is scored at &lt;strong&gt;payload level on long mixed payloads&lt;/strong&gt;, which is&lt;br&gt;
what actually ships, not on clean isolated snippets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hosted LLMs via Kaggle Benchmarks&lt;/strong&gt; (the leaderboard: free hosted inference, one&lt;br&gt;
fixed prompt, structured verdict &lt;code&gt;{is_leak, quote}&lt;/code&gt; for every model):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;google/gemini-2.5-pro&lt;/strong&gt; and &lt;strong&gt;google/gemini-2.5-flash&lt;/strong&gt;: the frontier ceiling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;openai/gpt-oss-120b&lt;/strong&gt; (117B) and &lt;strong&gt;openai/gpt-oss-20b&lt;/strong&gt; (21B): open-weight reasoning models&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;4 models × 220 payloads = 880 unique (model, payload) prompts, 1,200 task executions&lt;br&gt;
(positives are prompted for both detection and quoting), 1,199 scored; the one&lt;br&gt;
unparseable response was a quoting call on a positive payload, so catch rates&lt;br&gt;
are unaffected and that model's localization mean is over 79 of 80. The suite reflects the Kaggle Benchmarks registry's current catalog;&lt;br&gt;
the benchmark itself is model-agnostic, so newer or larger models slot in by&lt;br&gt;
adding them to the runner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The local-endpoint reference point&lt;/strong&gt; (not on the kbench leaderboard; it runs on the&lt;br&gt;
laptop, not an API): a frozen &lt;strong&gt;all-MiniLM-L6-v2&lt;/strong&gt; (22M params) with a logistic head&lt;br&gt;
over 512-token windows (stride 128, payload decision = max window score), plus&lt;br&gt;
ModernBERT-base (149M) and DeBERTa-v3-base (184M) as heavier comparators, evaluated on&lt;br&gt;
a larger private set built the same way (3,000+ payloads, leakage-safe splits,&lt;br&gt;
thresholds locked on a calibration split at ≥99% recall, 3 head seeds each).&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;model&lt;/th&gt;
&lt;th&gt;catch rate (recall)&lt;/th&gt;
&lt;th&gt;false-alarm rate&lt;/th&gt;
&lt;th&gt;localization IoU&lt;/th&gt;
&lt;th&gt;ms/payload (CPU)&lt;/th&gt;
&lt;th&gt;RAM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemini-2.5-pro (hosted)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;39.3%&lt;/strong&gt; ⚠️&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.846&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;seconds + network&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-2.5-flash (hosted)&lt;/td&gt;
&lt;td&gt;0.962&lt;/td&gt;
&lt;td&gt;20.7%&lt;/td&gt;
&lt;td&gt;0.392&lt;/td&gt;
&lt;td&gt;seconds + network&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-120b (hosted, 117B)&lt;/td&gt;
&lt;td&gt;0.825&lt;/td&gt;
&lt;td&gt;17.1%&lt;/td&gt;
&lt;td&gt;0.060&lt;/td&gt;
&lt;td&gt;seconds + network&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-20b (hosted, 21B)&lt;/td&gt;
&lt;td&gt;0.750&lt;/td&gt;
&lt;td&gt;22.9%&lt;/td&gt;
&lt;td&gt;0.036&lt;/td&gt;
&lt;td&gt;seconds + network&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MiniLM-L6 (local, 22M)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7.3% ± 1.9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;(span head: next step)&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;2,312 mean / 5,179 p95&lt;/strong&gt;*&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+103MB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ModernBERT-base (local, 149M)&lt;/td&gt;
&lt;td&gt;0.95&lt;/td&gt;
&lt;td&gt;20.4%&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;17,085 / 36,156&lt;/td&gt;
&lt;td&gt;+561MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeBERTa-v3-base (local, 184M)&lt;/td&gt;
&lt;td&gt;0.975&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;36.1%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;
&lt;em&gt;not measured&lt;/em&gt;*&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Hosted rows: the public 220 payloads (95% Wilson intervals on false-alarm rates:&lt;br&gt;
pro 32-48%, flash 15-28%, gpt-oss-120b 12-24%, gpt-oss-20b 17-31%). Local rows:&lt;br&gt;
the private 3,000+ payload eval, where the ± is spread over head seeds,&lt;br&gt;
not over which negatives were drawn; the follow-up section reports the local&lt;br&gt;
model on the same 220 payloads (catch 1.000, false alarms 61.7%).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;* identical shared Kaggle CPU VM, INT8 dynamic-quantized PyTorch, 100 payloads of&lt;br&gt;
~26 windows each; treat as a &lt;em&gt;relative&lt;/em&gt; comparison (ONNX Runtime on a real laptop is&lt;br&gt;
several × faster in absolute terms).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The local-side story (measured, 3 seeds, 2%-prevalence eval):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A 22M-parameter embedding model beat a 149M and a 184M encoder at their own game&lt;/strong&gt;:
perfect recall on buried leaks at 3–5× lower false-alarm rate (per-window false-alarm
~0.3% vs 0.9% / 1.9%). Contrastively-trained sentence embeddings are &lt;em&gt;linearly
separable&lt;/em&gt; for this task; raw language-model features aren't. When your deployment
plan is a frozen encoder + tiny head (cheap to ship, seconds to retrain), embedding-
tuned models punch far above their parameter count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeBERTa-v3 is the alarm-fatigue cautionary tale&lt;/strong&gt;: near-top recall (97.5%) but a
36% payload false-alarm rate: it fires on a third of benign technical pastes. In DLP,
the tool that keeps raising false alarms gets uninstalled; recall alone would have
picked the wrong model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Burying a leak doesn't hide it&lt;/strong&gt;: detection of needles buried in 14k-char haystacks
matched or beat detection of the same leaks standalone (dilution ≈ 0 to negative for
all three models). Long-context dilution is not the failure mode people assume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At 2% prevalence, FDR is the honest headline&lt;/strong&gt;: even the winner's false discoveries
outnumber true catches (~4:1 at a 99%-recall threshold tuned from only 60 calibration
positives, a deliberately conservative protocol). Thresholds must be tuned on a
calibration split, never the eval set; and 1-in-5 vs 1-in-11 alert precision is the
difference between a tool people trust and one they disable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The LLM-side story&lt;/strong&gt; (4 hosted models, 879 scored runs, single fixed prompt):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Catch rate tracks model class&lt;/strong&gt;: gemini-2.5-pro catches every buried leak (1.000),
flash misses 4% (0.962), gpt-oss-120b 17% (0.825), and the 21B gpt-oss-20b misses 1 in 4
(0.750). The frontier-vs-open-weight gap is statistically solid (77/80 vs 66/80,
Fisher p = 0.009); the 117B-vs-21B step is not separable at n=80 (p = 0.33),
and 80/80 only bounds the frontier's rate at 96% or better.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The recall champion is the false-alarm champion&lt;/strong&gt;: gemini-2.5-pro flags &lt;strong&gt;39.3%&lt;/strong&gt; of
clean hard negatives, worse than DeBERTa-v3's 36%, the local side's cautionary tale.
No hosted model got below 17.1% on this set; the local 22M model runs at 7.3% on its private eval (cross-set: see the follow-up section below) (caveat: the LLMs sit
at one prompt-implied operating point, while the encoder's threshold is calibrated to
99% recall on a held-out split; the comparison favors the encoder in fairness terms,
and it still wins).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only the frontier model can point at the leak&lt;/strong&gt;: gemini-2.5-pro quotes it nearly
verbatim (IoU 0.85); flash is mediocre (0.39); the gpt-oss models flag the payload but
can't produce a usable verbatim span (0.036–0.060; strict scoring: a paraphrased quote
earns nothing). An alert UI needs the span, not just the bell.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The thesis the numbers support:&lt;/strong&gt; hosted LLMs reach the recall ceiling but are&lt;br&gt;
structurally wrong as the endpoint layer: each payload costs seconds of network round-trip and per-call fees,&lt;br&gt;
and using them means transmitting the exact text under suspicion off the device, which&lt;br&gt;
is the behavior a DLP product exists to prevent. On-device and free remain true; whether the local encoder also wins on accuracy is distribution-dependent, and the follow-up section below measures both sides of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Follow-up: the local model on the same 220 payloads
&lt;/h2&gt;

&lt;p&gt;The local 7.3% false-alarm rate and the hosted 17-39% come from different eval&lt;br&gt;
sets (private 3,000+ payloads vs the public 220), so as a direct control I ran&lt;br&gt;
the encoder on the same public payloads, head retrained after excluding every&lt;br&gt;
220 source from training (mandatory: all 80 public needles also appear in the&lt;br&gt;
private train split).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;model&lt;/th&gt;
&lt;th&gt;catch rate&lt;/th&gt;
&lt;th&gt;false-alarm rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-120b (hosted)&lt;/td&gt;
&lt;td&gt;0.825&lt;/td&gt;
&lt;td&gt;17.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-2.5-flash (hosted)&lt;/td&gt;
&lt;td&gt;0.962&lt;/td&gt;
&lt;td&gt;20.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-2.5-pro (hosted)&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;39.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MiniLM-L6 (local, same 220)&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;61.7% " + chr(177) + " 4.3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the same payloads the ordering inverts: the 22M model still catches everything&lt;br&gt;
but false-alarms on nearly two thirds of clean payloads.&lt;/p&gt;

&lt;p&gt;Why: the public 220's benign text is LLM-written (296 of the 300 generator&lt;br&gt;
chunks), while the private train and calibration sets contain zero LLM-written&lt;br&gt;
payloads; the head had simply never seen that writing style as negative. The length&lt;br&gt;
dependence of max-over-windows aggregation is also visible in isolation: false&lt;br&gt;
alarms climb&lt;br&gt;
42% / 65% / 84% across payloads of up to 10 / 11-20 / 21-30 windows, consistent&lt;br&gt;
with max-over-windows aggregation giving longer pastes more chances to trigger.&lt;/p&gt;

&lt;p&gt;A natural control is to retrain with half of the LLM-written negatives added to&lt;br&gt;
training and measure the held-out half: false alarms drop from 53% to 20% at&lt;br&gt;
unchanged 1.000 recall. But this cannot be read as a fix. The two halves share&lt;br&gt;
156 of 221 generator chunks (69 of 70 held-out payloads contain text the head saw&lt;br&gt;
in training), so the improvement is consistent with memorizing seen chunks.&lt;br&gt;
Whether fresh, unseen generator-written negatives would repair the small model is&lt;br&gt;
an open question: the public set consumed 296 of the 300 available chunks, so a&lt;br&gt;
clean answer requires generating a new negative pool. What the same-set result&lt;br&gt;
does establish: hard negatives must match the deployment distribution, and a&lt;br&gt;
private eval whose negatives differ in generator from the public ones can invert&lt;br&gt;
a model ranking.&lt;/p&gt;

&lt;p&gt;Limitations, stated plainly: synthetic corpus (no real employee data, by design); hosted-model rates are&lt;br&gt;
  single-run and stochastic (a full repeat of the suite moved catch/false-alarm rates by&lt;br&gt;
  up to 7pp; rank order held);&lt;br&gt;
80/140 payloads per arm for the LLM suite, 100 eval leaks for the encoders (recall CI&lt;br&gt;
±5–8pp); localization scoring is strict (verbatim quote must appear in the payload);&lt;br&gt;
encoder heads are linear probes on frozen features (a floor, not a ceiling;&lt;br&gt;
fine-tuning would likely close some gaps); encoder latency measured on a shared Kaggle&lt;br&gt;
CPU VM with an unoptimized backend (relative comparison only).&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.kaggle.com/benchmarks/karthiksethuraman6/semantic-dlp-catch-the-data-leak" rel="noopener noreferrer"&gt;Semantic DLP: Catch the Data Leak&lt;/a&gt;&lt;/strong&gt;: the payload set is attached and fully&lt;br&gt;
synthetic, so anyone can re-run it, add models, or fork the three tasks&lt;br&gt;
(leak_catch_rate / stay_silent_rate / locate_leak). The public leaderboard initially lists the models published through the&lt;br&gt;
platform's single-run output; the full four-model suite results are in the table&lt;br&gt;
above and published in the run output at&lt;br&gt;
&lt;a href="https://www.kaggle.com/code/karthiksethuraman6/leak-catch-rate/output" rel="noopener noreferrer"&gt;https://www.kaggle.com/code/karthiksethuraman6/leak-catch-rate/output&lt;/a&gt; : the summary&lt;br&gt;
&lt;code&gt;dlp_llm_summary.csv&lt;/code&gt; and the &lt;code&gt;backup_runs/&lt;/code&gt; folder with every model's raw run&lt;br&gt;
records are all there. The board extends via Add Models.&lt;br&gt;
One harness note for reproducers:&lt;br&gt;
reasoning models (gpt-oss) wrap their output in &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; blocks, so the verdict parser&lt;br&gt;
extracts the last JSON object from the raw response rather than trusting strict&lt;br&gt;
structured-output parsing.&lt;/p&gt;

&lt;p&gt;Next steps: span-level heads on the winning encoder (the corpus already labels the&lt;br&gt;
leak's exact characters; classification is only half of a usable alert); a second&lt;br&gt;
operating point at 95% recall for a two-point precision/recall curve; conformal&lt;br&gt;
thresholds; and extending the corpus to more leak categories.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
