<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: unit life</title>
    <description>The latest articles on DEV Community by unit life (@unit_500_c36d1b1011fdf39c).</description>
    <link>https://dev.to/unit_500_c36d1b1011fdf39c</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4022725%2Febb5eed8-3e1e-4707-b48c-122104782b4b.jpeg</url>
      <title>DEV Community: unit life</title>
      <link>https://dev.to/unit_500_c36d1b1011fdf39c</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/unit_500_c36d1b1011fdf39c"/>
    <language>en</language>
    <item>
      <title>100% vuln detection wasn't enough: measuring whether AI respects the patch</title>
      <dc:creator>unit life</dc:creator>
      <pubDate>Thu, 24 Sep 2026 13:18:19 +0000</pubDate>
      <link>https://dev.to/unit_500_c36d1b1011fdf39c/100-vuln-detection-wasnt-enough-measuring-whether-ai-respects-the-patch-dg4</link>
      <guid>https://dev.to/unit_500_c36d1b1011fdf39c/100-vuln-detection-wasnt-enough-measuring-whether-ai-respects-the-patch-dg4</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;AI models are often like over-eager alarms. Show them a dangerous word in code — &lt;code&gt;eval&lt;/code&gt;, &lt;code&gt;system(&lt;/code&gt;, a raw SQL concat — and they scream “vulnerability!” nearly every time. Add the lock one line up, and many cheaper models still scream: they recognized the scary token, they did not read the fix.&lt;/p&gt;

&lt;p&gt;That failure mode is expensive. Triage pipelines that use LLMs to flag candidate sinks drown in &lt;strong&gt;false positives on already-patched code&lt;/strong&gt;. Public coding evals ask “did you find a bug?” They rarely ask “did you respect the fix?”&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;ART — Attacker-Reachable Sink Triage&lt;/strong&gt; to measure that gap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How it works:&lt;/strong&gt; minimal-pair &lt;strong&gt;twins&lt;/strong&gt;. Same function name, same identifiers, same shape — only the security control differs. Prompts get &lt;strong&gt;snippet + language only&lt;/strong&gt;. Twin ids, gold labels, and rationales never enter the model context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1xi7fmemd5ol4aku6lv9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1xi7fmemd5ol4aku6lv9.png" alt="How a twin pair works: identical shape, only the control differs" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is a real pair from the set. The only difference is the fix — everything a token-matcher keys on (&lt;code&gt;$_GET["id"]&lt;/code&gt;, &lt;code&gt;SELECT&lt;/code&gt;, the function name) is identical:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="c1"&gt;// twin_sql_php · gold = reachable_vuln&lt;/span&gt;
&lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;process_user_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$conn&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nv"&gt;$id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$_GET&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="nv"&gt;$sql&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"SELECT * FROM users WHERE id = "&lt;/span&gt; &lt;span class="mf"&gt;.&lt;/span&gt; &lt;span class="nv"&gt;$id&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// attacker-controlled concat&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;mysqli_query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$sql&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight php"&gt;&lt;code&gt;&lt;span class="c1"&gt;// twin_sql_php · gold = patched  (same shape, one control added)&lt;/span&gt;
&lt;span class="k"&gt;function&lt;/span&gt; &lt;span class="n"&gt;process_user_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$conn&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nv"&gt;$id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;int&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nv"&gt;$_GET&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="nv"&gt;$stmt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;mysqli_prepare&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$conn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"SELECT * FROM users WHERE id = ?"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nf"&gt;mysqli_stmt_bind_param&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$stmt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"i"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;          &lt;span class="c1"&gt;// cast + prepared statement&lt;/span&gt;
    &lt;span class="nf"&gt;mysqli_stmt_execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$stmt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;mysqli_stmt_get_result&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$stmt&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A model that labels the second snippet &lt;code&gt;reachable_vuln&lt;/code&gt; isn't a worse &lt;em&gt;detector&lt;/em&gt; — it's a worse &lt;em&gt;patch reader&lt;/em&gt;. &lt;strong&gt;Twin Gap&lt;/strong&gt; captures exactly that: &lt;code&gt;vuln accuracy − patched accuracy&lt;/code&gt;. Zero means the model respects fixes; positive means it over-flags patched code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Core suite&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;art-label-triage&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4-way label: &lt;code&gt;reachable_vuln&lt;/code&gt; / &lt;code&gt;patched&lt;/code&gt; / &lt;code&gt;safe&lt;/code&gt; / &lt;code&gt;vacuous_noise&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;ART = 0.4·vuln + 0.4·patched + 0.2·filler&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;art-overconfidence-trap&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;On patched twins only: “is there a &lt;em&gt;confirmed&lt;/em&gt; exploit right now?” (gold = no)&lt;/td&gt;
&lt;td&gt;Fraction not overclaimed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;art-proof-marker-poc&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Emit a minimal lab PoC containing &lt;code&gt;ART_PROOF_OK&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;1.0 / 0.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Dataset:&lt;/strong&gt; 8 twin pairs (SQLi, XSS, auth bypass, command injection, path traversal, LFI, insecure deserialization — PHP + Python) plus 6 safe/vacuous controls. N is intentionally small: one miss moves Twin Gap by &lt;strong&gt;12.5%&lt;/strong&gt;. This is a diagnostic probe, not a large-N ranking claim.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why synthetic, not raw CVEs:&lt;/strong&gt; so models cannot win by memorizing a write-up, and so each twin differs by &lt;strong&gt;one control&lt;/strong&gt;. The &lt;em&gt;classes&lt;/em&gt; mirror recurring production patterns (WordPress-plugin-style PHP; Flask/Django-request-style Python). Failing a patched twin here is meant to map to over-flagging a real fix in those ecosystems.&lt;/p&gt;

&lt;p&gt;Ablations (personas + forced CoT) are public supporting tasks; the headline metric is label-triage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;Seven locked Community Benchmark models, chosen for &lt;strong&gt;tier × price × family&lt;/strong&gt; coverage — not a single SOTA chase:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Why it’s in the lineup&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;gemini-3.5-flash&lt;/code&gt; / &lt;code&gt;gemini-3.7-flash&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Fast Gemini tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-2.5-pro&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Does cost buy patch-respect?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;claude-haiku-4-5-20251001&lt;/code&gt; / &lt;code&gt;claude-sonnet-4-5-20250929&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Cheap vs mid Claude&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemma-4-31b-it&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Open-weights instruct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.4-nano-2026-03-17&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Price floor&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Roughly &lt;strong&gt;50×&lt;/strong&gt; cost span per full triage run (~$0.004 → ~$0.18). &lt;code&gt;qwen3-next-80b-a3b-instruct&lt;/code&gt; was attempted, hit heavy-load &lt;strong&gt;429&lt;/strong&gt;s, and was replaced by Gemma (documented in the repo’s &lt;code&gt;MODELS.md&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;p&gt;Every model found every vulnerable twin (&lt;strong&gt;100% raw&lt;/strong&gt;). That alone is useless — an alarm that never stops ringing doesn’t help. The useful question is whether the model &lt;strong&gt;calms down after the lock&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Four models&lt;/strong&gt; (three Gemini + Gemma) score perfect ART: bugs &lt;em&gt;and&lt;/em&gt; fixes. Gemma does it for ~&lt;strong&gt;$0.007&lt;/strong&gt;/run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cheaper models&lt;/strong&gt; still catch every bug but keep flagging fixed code (Haiku Twin Gap &lt;strong&gt;0.375&lt;/strong&gt; — 3 of 8 patched twins). Nano is cheapest but confuses harmless filler for risk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For builders:&lt;/strong&gt; rank on &lt;em&gt;patch reading&lt;/em&gt;, not hype — or your triage queue fills with already-fixed findings.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Numbers from &lt;strong&gt;art-label-triage v6&lt;/strong&gt; (adjudicated gold + production scoring):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;ART&lt;/th&gt;
&lt;th&gt;Raw&lt;/th&gt;
&lt;th&gt;Patched&lt;/th&gt;
&lt;th&gt;Controls&lt;/th&gt;
&lt;th&gt;Twin Gap&lt;/th&gt;
&lt;th&gt;Cost USD&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-2.5-pro&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.000&lt;/td&gt;
&lt;td&gt;0.181&lt;/td&gt;
&lt;td&gt;7.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-3.5-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.000&lt;/td&gt;
&lt;td&gt;0.108&lt;/td&gt;
&lt;td&gt;2.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.000&lt;/td&gt;
&lt;td&gt;0.028&lt;/td&gt;
&lt;td&gt;9.1s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemma-4-31b-it&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.000&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.000&lt;/td&gt;
&lt;td&gt;0.007&lt;/td&gt;
&lt;td&gt;12.3s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-sonnet-4-5-20250929&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.950&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.875&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.125&lt;/td&gt;
&lt;td&gt;0.060&lt;/td&gt;
&lt;td&gt;3.1s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-haiku-4-5-20251001&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.850&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.625&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.375&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.020&lt;/td&gt;
&lt;td&gt;1.7s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.4-nano-2026-03-17&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.817&lt;/td&gt;
&lt;td&gt;1.000&lt;/td&gt;
&lt;td&gt;0.875&lt;/td&gt;
&lt;td&gt;0.333&lt;/td&gt;
&lt;td&gt;0.125&lt;/td&gt;
&lt;td&gt;0.004&lt;/td&gt;
&lt;td&gt;1.3s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv7h6kwmbjixemwhaioej.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv7h6kwmbjixemwhaioej.png" alt="Raw vs patched twin accuracy with Wilson 95% CIs" width="800" height="356"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fynabgj8u5et7li1cz2wp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fynabgj8u5et7li1cz2wp.png" alt="Patch-respect per dollar: ART vs cost; bubble = latency" width="800" height="389"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost:&lt;/strong&gt; Gemma and Gemini flash match pro-tier ART at roughly &lt;strong&gt;1–4% of the cost&lt;/strong&gt;. Heavy Pro does not beat Flash or Gemma on this probe.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I’d actually ship
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Constraint&lt;/th&gt;
&lt;th&gt;Pick&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Open weights / on-prem&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemma-4-31b-it&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ART 1.000 at ~$0.007&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency-sensitive&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3.5-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;ART 1.000, ~2.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-family stacks&lt;/td&gt;
&lt;td&gt;Sonnet + a patch-respect check&lt;/td&gt;
&lt;td&gt;Strong explanations, non-zero overclaim&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  What surprised me (more than any leaderboard cell)
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. The models corrected our gold.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
All seven disagreed with two labels — in the same direction — and they were right. An escaped-input “safe” filler was really &lt;code&gt;patched&lt;/code&gt; by our own prompt definition; a deser twin that &lt;em&gt;replaced&lt;/em&gt; &lt;code&gt;pickle.loads&lt;/code&gt; with &lt;code&gt;json.loads&lt;/code&gt; was really &lt;code&gt;safe&lt;/code&gt;. Those two items capped every model at ART &lt;strong&gt;0.917&lt;/strong&gt; and invented our largest “failure” class. After adjudication (HMAC-gated pickle + &lt;code&gt;len()&lt;/code&gt;-only safe filler), the top cluster hits &lt;strong&gt;1.000&lt;/strong&gt; and remaining errors are patch misses under the frozen, adjudicated rubric.&lt;/p&gt;

&lt;p&gt;That redesign is the point: a ceiling isn’t always model competence — sometimes it’s your key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. A 0.0 that wasn’t a capability.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Sonnet’s proof-marker stayed &lt;strong&gt;0.0&lt;/strong&gt; across retries because the provider returned an &lt;strong&gt;empty completion&lt;/strong&gt; (86 prompt tokens, empty message). Read the transcript before ranking a model on a single-shot cell.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Personas and CoT didn’t “fix” patch-respect.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Red-team persona did not systematically inflate overclaim. Forced data-flow CoT on the trap did &lt;strong&gt;not&lt;/strong&gt; close Haiku’s gap (0.625 → 0.50).&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F60bsse962xky47ju7gyn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F60bsse962xky47ju7gyn.png" alt="Failure taxonomy after gold adjudication" width="800" height="511"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Two concrete Haiku misses
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Path twin:&lt;/strong&gt; claims &lt;code&gt;basename("../../../etc/passwd")&lt;/code&gt; still traverses — it doesn’t (&lt;strong&gt;ignored sanitizer&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth twin:&lt;/strong&gt; admits &lt;code&gt;current_user_can&lt;/code&gt; works, then still labels &lt;code&gt;reachable_vuln&lt;/code&gt; by shifting to a different risk (&lt;strong&gt;overclaim&lt;/strong&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Honest limits
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;N = 8 pairs + 6 controls. Haiku’s 3/8 patched misses → exact sign-test p = 0.25 at this N. Small on purpose; one miss is loud.&lt;/li&gt;
&lt;li&gt;Kaggle &lt;em&gt;collection&lt;/em&gt; pages often show Pass/100 for Score floats. The ranked number is each run’s &lt;code&gt;rewards.score&lt;/code&gt; (and the table above) — not the collection chart.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What I’d measure next
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Sandbox-execute the proof-marker (assert on printed output, not source text).&lt;/li&gt;
&lt;li&gt;Multi-file / multi-hop taint.&lt;/li&gt;
&lt;li&gt;Larger N once the gold-audit loop stays routine.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Kaggle collection (required):&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/moranzavdi/attacker-reachable-sink-triage-art" rel="noopener noreferrer"&gt;Attacker-Reachable Sink Triage (ART)&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Source (MIT):&lt;/strong&gt; &lt;a href="https://github.com/mziqudhd92/kaggle-art-benchmark" rel="noopener noreferrer"&gt;https://github.com/mziqudhd92/kaggle-art-benchmark&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Core tasks&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/moranzavdi/art-label-triage" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/moranzavdi/art-label-triage&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/moranzavdi/art-overconfidence-trap" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/moranzavdi/art-overconfidence-trap&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.kaggle.com/benchmarks/tasks/moranzavdi/art-proof-marker-poc" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/moranzavdi/art-proof-marker-poc&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kaggle b t run art-label-triage &lt;span class="nt"&gt;-m&lt;/span&gt; gemini-3.5-flash &lt;span class="nt"&gt;--wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Reproducibility:&lt;/strong&gt; the dataset (&lt;code&gt;dataset/items.jsonl&lt;/code&gt;) is frozen and versioned in the repo; gold labels are deterministic and scored by &lt;code&gt;param_id&lt;/code&gt;, not answer order. &lt;code&gt;scripts/validate_jsonl.py&lt;/code&gt; and &lt;code&gt;scripts/test_scoring_alignment.py&lt;/code&gt; gate every push, and &lt;code&gt;scripts/analyze_results.py&lt;/code&gt; regenerates the tables and charts above from the downloaded run artifacts. Every number here traces to a specific task version (label-triage &lt;strong&gt;v6&lt;/strong&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Safety:&lt;/strong&gt; synthetic snippets only; defensive triage research; no live targeting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Credit:&lt;/strong&gt; inspired by proof-over-speculation tooling (&lt;a href="https://github.com/mziqudhd92/Iridium" rel="noopener noreferrer"&gt;Iridium&lt;/a&gt;); this entry is a standalone Kaggle Community Benchmark under MIT.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Fingerprinting Network Honeypots with Weighted Behavioral Scoring Engine</title>
      <dc:creator>unit life</dc:creator>
      <pubDate>Thu, 03 Sep 2026 11:30:33 +0000</pubDate>
      <link>https://dev.to/unit_500_c36d1b1011fdf39c/fingerprinting-network-honeypots-with-weighted-behavioral-scoring-engine-2d9k</link>
      <guid>https://dev.to/unit_500_c36d1b1011fdf39c/fingerprinting-network-honeypots-with-weighted-behavioral-scoring-engine-2d9k</guid>
      <description>&lt;p&gt;Deception technology has evolved past static string matches. &lt;br&gt;
Modern decoys try to mimic production environments, &lt;br&gt;
making binary "is it a honeypot?" checks unreliable. &lt;br&gt;
Single indicators—like a missing Date header or an unusual SSH banner—frequently trigger false positives on enterprise middleboxes and legacy servers.&lt;/p&gt;

&lt;p&gt;To solve this, we built Honeypot-Auditor around a dual-dimensional Honeyscore &amp;amp; Confidence engine. Here is a breakdown of how the scoring mechanics work under the hood.&lt;/p&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faijhqz9t0qbk03l062dv.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faijhqz9t0qbk03l062dv.gif" alt="In action" width="720" height="484"&gt;&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;&lt;em&gt;Github: &lt;a href="https://github.com/mziqudhd92/honeypot-auditor" rel="noopener noreferrer"&gt;https://github.com/mziqudhd92/honeypot-auditor&lt;/a&gt;&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Math Behind the Honeyscore (0–100%)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Rather than assigning flat point values, Honeypot-Auditor treats each protocol anomaly as a weighted indicator with strict corroboration rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Isolated Indicator Penalty: Standard L4/TLS stack tells (e.g., JA3S signatures or header ordering) carry low independent weight (~10–15%).&lt;/li&gt;
&lt;li&gt;Corroboration Multiplier: Weak tells require corroboration across independent categories (e.g., combining a TLS cipher mismatch with a state-machine failure). When two distinct categories hit, a corroboration gate unlocks the full indicator weight.&lt;/li&gt;
&lt;li&gt;Hard Tells: High-interaction leaks (such as arbitrary auth acceptance or shell execution latency anomalies) act as high-confidence anchors that push the score above 80%.&lt;/li&gt;
&lt;/ul&gt;



&lt;ol&gt;
&lt;li&gt;Under the Hood: Corroboration Gating&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here is a simplified look at how the analyzer evaluates indicator weights and suppresses weak signals unless corroborated by an independent category:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;calculate_honeyscore&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;indicators&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Indicator&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;proxy_detected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
    &lt;span class="n"&gt;categories_hit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;category&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;indicators&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;triggered&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;indicators&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;triggered&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="c1"&gt;# Proxy Guard: Suppress L4/TLS stack tells if an edge proxy is active
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;proxy_detected&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fingerprint_type&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;PROXY_SUPPRESSED_TYPES&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;suppressed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="c1"&gt;# Corroboration Gate: Weak tells require at least 2 distinct categories
&lt;/span&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;requires_corroboration&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;categories_hit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;suppressed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;ind&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;weight&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;100.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;strong&gt;Cool Web Page: &lt;a href="https://mziqudhd92.github.io/honeypot-auditor/" rel="noopener noreferrer"&gt;https://mziqudhd92.github.io/honeypot-auditor/&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;



&lt;ol&gt;
&lt;li&gt;Dual-Dimensional Output: Score vs. Confidence&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A high score alone isn't enough for automated decision-making. We pair the Honeyscore with a separate Confidence metric:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LOW Confidence: Triggered when &amp;lt; 3 protocols are audited or &amp;gt; 50% of probes fail/timeout.&lt;/li&gt;
&lt;li&gt;MEDIUM Confidence: Reached when 3+ protocols respond and at least 2 distinct category hits occur.&lt;/li&gt;
&lt;li&gt;HIGH Confidence: Achieved when deep behavioral mode confirms corroborating tells across multiple independent layers.&lt;/li&gt;
&lt;/ul&gt;



&lt;ol&gt;
&lt;li&gt;Real-World Scenario: The Banner Spoofing Trap&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Consider a modern SSH decoy (like Cowrie,dd-honeypot and others) configured to mimic a standard production Linux server:&lt;/p&gt;

&lt;p&gt;a. Naive Banner Grabber:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Connects to port 22 and reads the string: &lt;code&gt;SSH-2.0-OpenSSH_8.9p1 Ubuntu-3ubuntu0.1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Concludes: "Standard production server (0% Honeypot)."
b. Honeypot-Auditor:&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Banner Check: Sees OpenSSH claim. (Weight: 0%)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;KEXINIT Inspection: Extracts the raw &lt;a href="https://datatracker.ietf.org/doc/html/rfc4253#section-7.1" rel="noopener noreferrer"&gt;&lt;code&gt;SSH_MSG_KEXINIT&lt;/code&gt;&lt;/a&gt; byte sequence. &lt;br&gt;
It detects Paramiko/Twisted-specific key exchange algorithms and rigid cipher preferences that real OpenSSH binaries never advertise.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Stack Cross-Referencing: Checks the TCP SYN-ACK option ordering. The underlying kernel exhibits a generic container profile, contradicting the OS claimed by the banner.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Corroboration Gate Unlocked: Combining the application-layer banner claim with the KEXINIT algorithm mismatch triggers a high-confidence indicator.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Result: Honeyscore leaps to 85.0% [HIGH LIKELIHOOD DECOY] with Tactical Action &lt;code&gt;SKIP_TARGET&lt;/code&gt;.&lt;br&gt;
And this is by checking only port 22, usually we will ffind more open ports on honeypots and combining checks on them will increase the detection even further. &lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;



&lt;ol&gt;
&lt;li&gt;Tactical Action Outcomes&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Instead of forcing engineers to interpret raw percentages, the scoring engine resolves into four tactical actions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SKIP_TARGET&lt;/strong&gt;: Score &amp;gt;= 60% with HIGH/MEDIUM confidence (confirmed decoy).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PIVOT_POSSIBLE&lt;/strong&gt;: Score &amp;lt; 30% with HIGH confidence (verified production target).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PROCEED_CAUTION&lt;/strong&gt;: Score between 30–59% or LOW confidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;INCONCLUSIVE&lt;/strong&gt;: Edge proxy masking origin stack or insufficient probe responses.&lt;/li&gt;
&lt;/ul&gt;



&lt;p&gt;So what is different ? &lt;br&gt;
The main thing is that we are not only counting on signatures to detect a target, the engine using 16 protocols that are implementing more than 50 different strategies to evaluate if remote host is a decoy or not.&lt;/p&gt;

&lt;p&gt;You can try it out, Honeypot-Auditor is open-source (MIT licensed:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;honeypot-auditor
honeypot-auditor &lt;span class="nt"&gt;--target&lt;/span&gt; 127.0.0.1 &lt;span class="nt"&gt;-v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/mziqudhd92" rel="noopener noreferrer"&gt;
        mziqudhd92
      &lt;/a&gt; / &lt;a href="https://github.com/mziqudhd92/honeypot-auditor" rel="noopener noreferrer"&gt;
        honeypot-auditor
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Does This Look Like An Honeypot? (DTLLAH) Multi-protocol CLI that fingerprints whether a target IP behaves like a low-interaction honeypot — Shodan Honeyscore, active auth/state probes, and a weighted score.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;.______________________________________________________________________________
|  :: H-AUDITOR :: v0.7.3 :: "DIALING IN... CARRIER DETECTED" ::                |
|------------------------------------------------------------------------------|
|  "warez? nah. headers. we trade banners, not bins."                          |
|  "if it answers any password, it ain't production — it's a lure."            |
|  "respect the sysop. probe only what you own. leave no STOR behind."         |
|______________________________________________________________________________|
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;a href="https://pypi.org/project/honeypot-auditor/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/397fdd7bff5993dc116156dde5056f79a555f5409682284f4f507eb25bf52ab4/68747470733a2f2f696d672e736869656c64732e696f2f707970692f762f686f6e6579706f742d61756469746f723f7374796c653d666c61742d737175617265" alt="PyPI"&gt;&lt;/a&gt;
&lt;a href="https://pypi.org/project/honeypot-auditor/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/62050bd68b60d4c7a002e1a7c121de1f66b996a48a9d7deeffb06fb60ba0317e/68747470733a2f2f696d672e736869656c64732e696f2f707970692f707976657273696f6e732f686f6e6579706f742d61756469746f723f7374796c653d666c61742d737175617265" alt="Python"&gt;&lt;/a&gt;
&lt;a href="https://github.com/mziqudhd92/honeypot-auditor/actions/workflows/test.yml" rel="noopener noreferrer"&gt;&lt;img src="https://github.com/mziqudhd92/honeypot-auditor/actions/workflows/test.yml/badge.svg" alt="tests"&gt;&lt;/a&gt;
&lt;a href="https://github.com/mziqudhd92/honeypot-auditor/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/422db9fd40f5831c765cf6530b6750c081b696bd18d904cf89554df98c676277/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d677265656e3f7374796c653d666c61742d737175617265" alt="License: MIT"&gt;&lt;/a&gt;
&lt;a href="https://mziqudhd92.github.io/honeypot-auditor/" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/fc04974bb0759798f29df959fcb5188fcb58b55d4daeb63d512fd7005d4e388f/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f736974652d42425325323050616765732d3333666636363f7374796c653d666c61742d737175617265266c6162656c436f6c6f723d303530383035" alt="Pages"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Site (BBS / NFO):&lt;/strong&gt; &lt;a href="https://mziqudhd92.github.io/honeypot-auditor/" rel="nofollow noopener noreferrer"&gt;https://mziqudhd92.github.io/honeypot-auditor/&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Agents / AEO:&lt;/strong&gt; &lt;a href="https://mziqudhd92.github.io/honeypot-auditor/llms.txt" rel="nofollow noopener noreferrer"&gt;llms.txt&lt;/a&gt; · &lt;a href="https://mziqudhd92.github.io/honeypot-auditor/agents.md" rel="nofollow noopener noreferrer"&gt;agents.md&lt;/a&gt;&lt;/p&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;  ▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄▄
  █  &amp;gt;&amp;gt;&amp;gt; LIVE DEMO · 3 HOST LAB TOUR · -v / --deep / SILENT-ACCEPT &amp;lt;&amp;lt;&amp;lt;     █
  ▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀▀
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/mziqudhd92/honeypot-auditor/docs/demo/honeypot-auditor-lab-tour-demo.gif"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fmziqudhd92%2Fhoneypot-auditor%2FHEAD%2Fdocs%2Fdemo%2Fhoneypot-auditor-lab-tour-demo.gif" alt="Lab tour demo — Cowrie, dd-stack, tarpit"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;  "three hosts, three lenses: KEX facade with -v, deep on the buffet,
   silent-accept on the tarpit. same fingerprinter — different tells."
                                              — lab tour · authorized only
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;
&lt;pre class="notranslate"&gt;&lt;code&gt;.------------------------------------------------------------------------------
|  NFO · READ BEFORE YOU DIAL                                                  |
|------------------------------------------------------------------------------|
|  Authorized targets ONLY. Lab boxes. Decoys you own. Sensors you run.        |
|  Permission on paper (or in ticket).                                         |
|                                                                              |
|  Scanning random /16 because Shodan said&lt;/code&gt;&lt;/pre&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/mziqudhd92/honeypot-auditor" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;



&lt;p&gt;Would love feedback!&lt;/p&gt;

</description>
      <category>python</category>
      <category>cybersecurity</category>
      <category>opensource</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
