<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bro Force</title>
    <description>The latest articles on DEV Community by Bro Force (@bro_force_af8dd8c3202a933).</description>
    <link>https://dev.to/bro_force_af8dd8c3202a933</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4103222%2F3a86ccc0-c372-4b0a-8096-c38009e175d9.jpg</url>
      <title>DEV Community: Bro Force</title>
      <link>https://dev.to/bro_force_af8dd8c3202a933</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bro_force_af8dd8c3202a933"/>
    <language>en</language>
    <item>
      <title>I Gave My Secret Scanner a Confidence Score — Then Found a Number I Didn't Want to Publish</title>
      <dc:creator>Bro Force</dc:creator>
      <pubDate>Mon, 31 Aug 2026 19:12:40 +0000</pubDate>
      <link>https://dev.to/bro_force_af8dd8c3202a933/forge-a-zero-dependency-secret-scanner-and-the-stdlib-potholes-i-hit-building-it-49k6</link>
      <guid>https://dev.to/bro_force_af8dd8c3202a933/forge-a-zero-dependency-secret-scanner-and-the-stdlib-potholes-i-hit-building-it-49k6</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; I built &lt;code&gt;forge&lt;/code&gt; for raptors.dev's Zero Dependency Hackathon&lt;br&gt;
(Track F) — a single-binary secret scanner that ranks its own findings by&lt;br&gt;
confidence instead of just flagging them, on top of a hand-rolled&lt;br&gt;
log-structured store. Nothing outside Python's standard library. This is&lt;br&gt;
the story of a 327-false-positive run I almost buried, a build hash that&lt;br&gt;
changed on a machine I hadn't touched, and a benchmark number I genuinely&lt;br&gt;
did not want to put in the README — and put in anyway.&lt;/p&gt;

&lt;p&gt;For raptors.dev's Zero Dependency Hackathon (Track F, Open/Wildcard), I&lt;br&gt;
built &lt;strong&gt;Forge&lt;/strong&gt;: one &lt;code&gt;forge&lt;/code&gt; binary that does secret scanning, file&lt;br&gt;
search, dedup, repo stats, a terminal dashboard, run-to-run diffing, a&lt;br&gt;
risk timeline, and per-finding confidence explanations. Underneath all&lt;br&gt;
of it sits a hand-rolled log-structured key-value store. Nothing outside&lt;br&gt;
Python's standard library.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5,502 lines in &lt;code&gt;src/forge/&lt;/code&gt;. 313 tests. 0 runtime dependencies. ~77 KB&lt;br&gt;
single-file build.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'm not going to walk the whole feature list — the README does that.&lt;br&gt;
This is the honest version: what I actually rebuilt, a real bug that&lt;br&gt;
testing against live repos actually caught, where the standard library&lt;br&gt;
fought back, which real package I think I made pointless, and the one&lt;br&gt;
build bug that cost me an entire afternoon and taught me more than&lt;br&gt;
anything else in the project.&lt;/p&gt;
&lt;h2&gt;
  
  
  The number I didn't want to publish
&lt;/h2&gt;

&lt;p&gt;Every one of the numbers above makes Forge look good, so here's the one&lt;br&gt;
that doesn't. I benchmarked &lt;code&gt;forge scan&lt;/code&gt; against &lt;code&gt;detect-secrets&lt;/code&gt; on two&lt;br&gt;
targets — a small tree (42 files, 291 KB) and a large one (480 files,&lt;br&gt;
5.1 MB) — 7 timed runs each, cold scans, no result cache, matching&lt;br&gt;
&lt;code&gt;detect-secrets&lt;/code&gt;' no-cache behavior:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Forge&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;detect-secrets&lt;/code&gt; 1.5.0&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Startup baseline (&lt;code&gt;--version&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.52s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.14s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small tree (42 files, 291 KB)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.79s&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.61s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large tree (480 files, 5.1 MB)&lt;/td&gt;
&lt;td&gt;4.66s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.81s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scan-only throughput, large tree&lt;/td&gt;
&lt;td&gt;~1.3 MB/s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~7 MB/s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Forge starts twice as fast and wins the small-tree case, because&lt;br&gt;
&lt;code&gt;detect-secrets&lt;/code&gt; pays import cost for &lt;code&gt;pyyaml&lt;/code&gt;, &lt;code&gt;requests&lt;/code&gt;, &lt;code&gt;urllib3&lt;/code&gt;,&lt;br&gt;
and its full plugin registry before it scans a single byte. On the large&lt;br&gt;
tree, that startup advantage stops mattering and Forge's &lt;strong&gt;scan loop&lt;br&gt;
runs 5–6× slower per byte&lt;/strong&gt;. Twelve regexes plus entropy tokenization on&lt;br&gt;
every line, no real thread parallelism because of the GIL, versus a&lt;br&gt;
library that's had years to get its throughput right. I checked whether&lt;br&gt;
the confidence scorer was the cause before writing this — it's under 4%&lt;br&gt;
of scan time, so ranking isn't the tax. Raw regex throughput is.&lt;/p&gt;

&lt;p&gt;I don't have a fix for this that doesn't compromise something else I&lt;br&gt;
actually care about (zero deps, zero install, ranked output), so instead&lt;br&gt;
of hiding it in a footnote I put it in the README next to the win. The&lt;br&gt;
trade is real: Forge gives up large-tree throughput for zero&lt;br&gt;
installation and confidence-ranked findings; &lt;code&gt;detect-secrets&lt;/code&gt; gives up&lt;br&gt;
both of those for raw speed. Different tools optimizing for different&lt;br&gt;
things, and a judge should get to see the actual trade instead of a&lt;br&gt;
cherry-picked number.&lt;/p&gt;
&lt;h2&gt;
  
  
  Architecture at a glance
&lt;/h2&gt;

&lt;p&gt;One CLI entrypoint, ten subcommands, everything fanning down into a&lt;br&gt;
single storage layer. This is the actual internal import graph, not an&lt;br&gt;
idealized version of it — grouped into three layers so the shape reads&lt;br&gt;
at a glance:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;graph TD
    subgraph L1["CLI"]
        CLI["cli.py — 10 subcommands"]
    end

    subgraph L2["Feature modules"]
        Scanner["scanner.py&amp;lt;br/&amp;gt;12 regexes + entropy"]
        Dedup["dedup.py"]
        Stats["stats.py&amp;lt;br/&amp;gt;hand-parsed .git objects"]
        Dashboard["dashboard.py&amp;lt;br/&amp;gt;curses / ANSI fallback"]
        Watcher["watcher.py&amp;lt;br/&amp;gt;polling watch"]
        Baseline["baseline.py&amp;lt;br/&amp;gt;new-vs-existing gate"]
        Sarif["sarif.py&amp;lt;br/&amp;gt;SARIF 2.1.0"]
        History["history.py&amp;lt;br/&amp;gt;diff + risk timeline"]
    end

    subgraph L2b["Scoring"]
        Confidence["confidence.py&amp;lt;br/&amp;gt;naive-Bayes log-odds scorer"]
        Entropy["entropy.py"]
    end

    subgraph L3["Persistence"]
        Cache["cache.py"]
        Storage[("storage.py — LogStore&amp;lt;br/&amp;gt;CRC32 · crash recovery")]
    end

    CLI --&amp;gt; Scanner &amp;amp; Dedup &amp;amp; Stats &amp;amp; Dashboard &amp;amp; Watcher &amp;amp; Baseline &amp;amp; Sarif &amp;amp; History

    Scanner --&amp;gt; Cache
    Scanner --&amp;gt; Confidence
    Dashboard -.-&amp;gt;|reuses| Scanner
    Watcher -.-&amp;gt;|reuses| Scanner
    Confidence --&amp;gt; Entropy

    Cache --&amp;gt; Storage
    History --&amp;gt; Storage

    style Storage fill:#2d2d2d,stroke:#f5a623,stroke-width:2px,color:#fff
    style Confidence fill:#2d2d2d,stroke:#f5a623,stroke-width:2px,color:#fff&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Solid arrows are "needs this to function"; dashed arrows labeled &lt;em&gt;reuses&lt;/em&gt;&lt;br&gt;
are "calls this subcommand's logic directly" (&lt;code&gt;dashboard&lt;/code&gt; and &lt;code&gt;watcher&lt;/code&gt;&lt;br&gt;
both call into the scanner rather than re-implementing scan logic). Two&lt;br&gt;
things worth noticing about the &lt;em&gt;shape&lt;/em&gt;: everything that needs to&lt;br&gt;
persist between runs — the cache and the risk/diff history — funnels&lt;br&gt;
through one storage engine, so &lt;code&gt;scan&lt;/code&gt;, &lt;code&gt;dedup&lt;/code&gt;, &lt;code&gt;stats&lt;/code&gt;, and &lt;code&gt;diff&lt;/code&gt; all&lt;br&gt;
get crash-safe persistence for free instead of each subcommand rolling&lt;br&gt;
its own file format. And confidence scoring gets its own layer, not&lt;br&gt;
tucked inside the scanner as an afterthought — ranking isn't&lt;br&gt;
post-processing, it's a first-class step between "regex matched" and&lt;br&gt;
"finding reported," which is why &lt;code&gt;forge explain&lt;/code&gt; can reconcile its&lt;br&gt;
per-term breakdown back to the exact score &lt;code&gt;forge scan&lt;/code&gt; printed.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I reimplemented
&lt;/h2&gt;

&lt;p&gt;The obvious piece is the scanner itself: twelve regexes for AWS keys,&lt;br&gt;
GitHub tokens, Stripe secret keys, GitLab PATs, JWTs, PEM blocks, plus a&lt;br&gt;
Shannon-entropy fallback for anything that doesn't match a known shape.&lt;br&gt;
That part is table stakes for this hackathon category — most stdlib&lt;br&gt;
secret scanners stop here.&lt;/p&gt;

&lt;p&gt;The part I actually care about is &lt;code&gt;confidence.py&lt;/code&gt;: every finding gets&lt;br&gt;
scored &lt;strong&gt;0.00–1.00&lt;/strong&gt; by a hand-tuned naive-Bayes classifier evaluated in&lt;br&gt;
log-odds space — the exact math behind a spam filter, done by hand&lt;br&gt;
instead of &lt;code&gt;scikit-learn.fit()&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;log_odds&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;logit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prior_for_this_rule&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;entropy_term&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;character_diversity_term&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;ordered_sequence_term&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;match_length_term&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;comment_marker_term&lt;/span&gt;
          &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;file_type_term&lt;/span&gt;
&lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sigmoid&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;log_odds&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seven signal families feed it — pattern specificity, Shannon entropy,&lt;br&gt;
character-class diversity, monotonic-run detection (catches a&lt;br&gt;
hand-written &lt;code&gt;ABCDEF...&lt;/code&gt; charset that would otherwise score as&lt;br&gt;
high-entropy &lt;em&gt;and&lt;/em&gt; high-diversity), match length against known&lt;br&gt;
fixed-width formats, inline &lt;code&gt;# nosec&lt;/code&gt;/&lt;code&gt;# noqa&lt;/code&gt; markers, and file-path&lt;br&gt;
context. A hard override kills obvious placeholders like &lt;code&gt;changeme&lt;/code&gt; or&lt;br&gt;
&lt;code&gt;&amp;lt;token&amp;gt;&lt;/code&gt; outright. Every weight is documented inline where it's added,&lt;br&gt;
specifically so a judge — or me, six months from now — can read the&lt;br&gt;
score instead of trusting it.&lt;/p&gt;

&lt;p&gt;Underneath that, &lt;code&gt;storage.py&lt;/code&gt; is an append-only, CRC32-checksummed,&lt;br&gt;
log-structured key-value store with crash recovery, backing the shared&lt;br&gt;
cache/history layer for every subcommand that needs to remember a&lt;br&gt;
previous run.&lt;/p&gt;
&lt;h2&gt;
  
  
  Testing it on real repos found a real bug, not a hypothetical one
&lt;/h2&gt;

&lt;p&gt;Talk is cheap for a secret scanner, so I shallow-cloned three public&lt;br&gt;
repos I didn't pick for a favorable result — &lt;code&gt;pallets/click&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;psf/requests&lt;/code&gt;, &lt;code&gt;expressjs/express&lt;/code&gt;, 596 files, ~10.4 MB — and read every&lt;br&gt;
single finding by hand. None of the three has a real committed secret,&lt;br&gt;
so the right answer for all three is "nothing," and the interesting&lt;br&gt;
question is how much noise Forge generates getting there.&lt;/p&gt;

&lt;p&gt;The first run of that validation found &lt;strong&gt;327 findings in &lt;code&gt;requests&lt;/code&gt;&lt;br&gt;
alone.&lt;/strong&gt; Not 2 — 327. 325 of them were base64 chunks inside one file:&lt;br&gt;
&lt;code&gt;ext/requests-logo.ai&lt;/code&gt;, a PDF-wrapped Adobe Illustrator asset. My binary&lt;br&gt;
check only sniffed the first 4 KB of a file for a NUL byte, and &lt;code&gt;.ai&lt;/code&gt;&lt;br&gt;
files carry a long plain-text header (PDF structure plus XMP metadata)&lt;br&gt;
that outruns 4 KB before the binary payload starts — so Forge read the&lt;br&gt;
file as text and dutifully entropy-flagged every base64 run in it.&lt;/p&gt;

&lt;p&gt;Fix: run the binary check over the full file content Forge already has&lt;br&gt;
in memory, not just a 4 KB prefix, so a NUL byte anywhere marks the file&lt;br&gt;
binary. After that, &lt;code&gt;requests-logo.ai&lt;/code&gt; contributes &lt;strong&gt;zero&lt;/strong&gt; findings, and&lt;br&gt;
the three-repo total drops to 14 — down to genuine low/medium noise (a&lt;br&gt;
hash in a doc, an &lt;code&gt;ABCD…wxyz&lt;/code&gt; alphabet in a docstring, a dozen&lt;br&gt;
ETag/hash literals in &lt;code&gt;express&lt;/code&gt;'s own test fixtures). Every one of those&lt;br&gt;
12 &lt;code&gt;express&lt;/code&gt; findings landed in confidence's medium band (0.41–0.49);&lt;br&gt;
&lt;strong&gt;zero reached high-confidence&lt;/strong&gt; — which is the actual claim &lt;code&gt;--allowlist&lt;/code&gt;&lt;br&gt;
and the pre-commit hook lean on: the ranking keeps ordinary noise out of&lt;br&gt;
the tier that would actually block a commit.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;repo&lt;/th&gt;
&lt;th&gt;files&lt;/th&gt;
&lt;th&gt;findings (after fix)&lt;/th&gt;
&lt;th&gt;high-confidence&lt;/th&gt;
&lt;th&gt;true positives&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;click&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;195&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;requests&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;159&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;express&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;242&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I'm including the 327-finding number, not just the fixed 14, because a&lt;br&gt;
write-up that only shows the clean result after the fact isn't proof of&lt;br&gt;
anything — the bug is the evidence the validation actually happened.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where the standard library actually made me suffer
&lt;/h2&gt;

&lt;p&gt;Two places, honestly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;curses&lt;/code&gt; doesn't exist on Windows.&lt;/strong&gt; Not "harder to use" — it's simply&lt;br&gt;
not part of the CPython standard distribution there. &lt;code&gt;forge dashboard&lt;/code&gt;&lt;br&gt;
needed a live-updating terminal view, and the natural stdlib answer&lt;br&gt;
(&lt;code&gt;curses&lt;/code&gt;) evaporates the moment a judge runs this on Windows, which —&lt;br&gt;
given this hackathon's toolchain — a lot of them will. The fallback is a&lt;br&gt;
plain-text view that redraws on an interval using raw ANSI escape codes.&lt;br&gt;
It works, but it's not the same widget, and I say so directly in the&lt;br&gt;
README instead of pretending the two code paths are equivalent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Determinism across platforms is not free.&lt;/strong&gt; &lt;code&gt;os.walk&lt;/code&gt; order is sorted&lt;br&gt;
on NTFS and effectively hash-order on ext4. &lt;code&gt;forge stats --json&lt;/code&gt; returned&lt;br&gt;
its &lt;code&gt;files_by_ext&lt;/code&gt;/&lt;code&gt;bytes_by_ext&lt;/code&gt;/&lt;code&gt;lines_by_ext&lt;/code&gt; maps straight from that&lt;br&gt;
walk, so the same repo produced byte-different JSON depending on the&lt;br&gt;
filesystem underneath it — invisible in the human-readable table (which&lt;br&gt;
was already sorted by count for display), silently wrong in the machine&lt;br&gt;
contract judges would actually diff. Fix was mechanical once I found it&lt;br&gt;
— sort the maps at construction — but "the stdlib gives you an iteration&lt;br&gt;
order, not a &lt;em&gt;guarantee&lt;/em&gt;" is a lesson that doesn't show up until you&lt;br&gt;
actually run the same code on two OSes.&lt;/p&gt;
&lt;h2&gt;
  
  
  The package I think I made unnecessary
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;detect-secrets&lt;/code&gt; (Yelp) — 1–2M downloads/month by PyPI's own tracker,&lt;br&gt;
and the closest real category match to what &lt;code&gt;forge scan&lt;/code&gt; does: regex&lt;br&gt;
plus entropy heuristics, a baseline file, a CI/pre-commit gate. I&lt;br&gt;
installed &lt;code&gt;detect-secrets==1.5.0&lt;/code&gt; into a throwaway venv and measured&lt;br&gt;
instead of guessing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Forge&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;detect-secrets&lt;/code&gt; 1.5.0&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Runtime deps&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2 direct, 6 total installed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Installed footprint&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0 bytes&lt;/strong&gt; (whole 10-subcommand CLI is ~77 KB)&lt;/td&gt;
&lt;td&gt;~4.3 MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ranks findings by confidence&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No — flags only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the actual pitch, not the dependency count.&lt;br&gt;
&lt;code&gt;detect-secrets&lt;/code&gt; treats every regex hit as equally worth a human's time.&lt;br&gt;
Forge ranks them, so a bare AWS key ID sitting in &lt;code&gt;config/production.env&lt;/code&gt;&lt;br&gt;
sorts to the top and a high-entropy string on a line ending in &lt;code&gt;# noqa&lt;/code&gt;&lt;br&gt;
inside &lt;code&gt;tests/fixtures/&lt;/code&gt; sorts to the bottom — the same signal every&lt;br&gt;
commercial scanner (GitGuardian, GitHub secret scanning) sells as a&lt;br&gt;
headline feature, done here in ~170 lines of &lt;code&gt;math&lt;/code&gt; and &lt;code&gt;re&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I'll say the honest gap too: &lt;code&gt;detect-secrets&lt;/code&gt; has a real plugin&lt;br&gt;
architecture and an interactive &lt;code&gt;audit&lt;/code&gt; workflow for annotating false&lt;br&gt;
positives over time that persists across a team. Forge's flat&lt;br&gt;
&lt;code&gt;--allowlist&lt;/code&gt; file is the lighter-weight version of that, not a clone.&lt;/p&gt;
&lt;h2&gt;
  
  
  The edge case that ate an afternoon
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Reproducible builds and &lt;code&gt;core.autocrlf&lt;/code&gt; do not get along.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Forge claims a byte-reproducible build — rebuild &lt;code&gt;dist/forge.pyz&lt;/code&gt; from&lt;br&gt;
source and it should hash identically every time, so a judge can verify&lt;br&gt;
the exact artifact they're running matches the repo. On my machine, it&lt;br&gt;
did. Then I checked out the repo fresh on a different Windows setup with&lt;br&gt;
&lt;code&gt;core.autocrlf=true&lt;/code&gt; (the Windows default) and the hash &lt;em&gt;changed&lt;/em&gt;. Same&lt;br&gt;
source, same commit, different SHA-256.&lt;/p&gt;

&lt;p&gt;Two separate bugs stacked on top of each other:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;scripts/build.py&lt;/code&gt; was reading source bytes verbatim into the zipapp.
A Windows autocrlf clone silently converts every &lt;code&gt;\n&lt;/code&gt; in the checkout
to &lt;code&gt;\r\n&lt;/code&gt; on the way out of git — so the "same" source file was
actually different bytes depending on which machine cloned it, before
my build script ever touched it.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;scripts/bundle_single_file.py&lt;/code&gt; was writing its output with
&lt;code&gt;write_text()&lt;/code&gt;, which on Windows happily re-introduces platform line
endings on the way &lt;em&gt;out&lt;/em&gt;, even if the input was clean.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Neither bug is visible from reading either script in isolation — both&lt;br&gt;
looked correct. It only showed up as "the hash doesn't match" after a&lt;br&gt;
clean clone on a machine I hadn't built on before, which is exactly the&lt;br&gt;
scenario a judge running the reproducibility check would hit.&lt;/p&gt;

&lt;p&gt;The fix, once diagnosed: normalize all source to &lt;code&gt;\n&lt;/code&gt; before archiving,&lt;br&gt;
switch the bundler to &lt;code&gt;write_bytes()&lt;/code&gt; with an LF-pinned bootstrap, and&lt;br&gt;
add a &lt;code&gt;.gitattributes&lt;/code&gt; pinning &lt;code&gt;* text=auto eol=lf&lt;/code&gt; (with &lt;code&gt;*.pyz&lt;/code&gt;&lt;br&gt;
declared binary so git doesn't touch the archive itself). I also found&lt;br&gt;
&lt;code&gt;zipfile&lt;/code&gt; stamps each entry's &lt;code&gt;create_system&lt;/code&gt; byte from the &lt;em&gt;build OS&lt;/em&gt;&lt;br&gt;
(0 on Windows, 3 elsewhere) — so even with identical &lt;code&gt;zlib&lt;/code&gt; output, a&lt;br&gt;
Windows-built and Linux-built &lt;code&gt;.pyz&lt;/code&gt; still differed on that one metadata&lt;br&gt;
byte. Pinned that too.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;tests/test_build.py&lt;/code&gt; now has explicit regression tests —&lt;br&gt;
&lt;code&gt;test_pyz_hash_is_independent_of_source_line_endings&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;test_single_file_hash_is_independent_of_source_line_endings&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;test_committed_source_and_artifacts_are_lf_only&lt;/code&gt; — so this doesn't come&lt;br&gt;
back quietly. And &lt;code&gt;judge_mode.py&lt;/code&gt; re-derives both artifact hashes as one&lt;br&gt;
of its 13 checks, so "reproducible" is something a judge runs in seconds&lt;br&gt;
instead of taking on faith.&lt;/p&gt;

&lt;p&gt;The lesson wasn't really about git or zipfile. It was that "reproducible&lt;br&gt;
build" is a claim about &lt;em&gt;bytes&lt;/em&gt;, and bytes have opinions about line&lt;br&gt;
endings and OS metadata that never show up until someone else's&lt;br&gt;
environment disagrees with yours.&lt;/p&gt;
&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;Knowing what I know now, I'd build this exact reproducible-build check&lt;br&gt;
&lt;em&gt;first&lt;/em&gt;, not near the end. &lt;code&gt;judge_mode.py&lt;/code&gt;'s hash verification came&lt;br&gt;
after the CRLF bug above had already shipped once — which meant the bug&lt;br&gt;
reached a committed artifact before anything caught it. If the&lt;br&gt;
byte-comparison check had existed from day one, that whole category of&lt;br&gt;
bug gets caught in CI on the first Windows run, not discovered by&lt;br&gt;
accident on a second machine days later. The fix itself — normalize line&lt;br&gt;
endings before archiving — took twenty minutes once diagnosed.&lt;br&gt;
Diagnosing it took the afternoon. Building the verification first would&lt;br&gt;
have collapsed those into the same twenty minutes.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where it landed
&lt;/h2&gt;

&lt;p&gt;313 tests pass. &lt;code&gt;dist/forge.pyz&lt;/code&gt; and &lt;code&gt;dist/forge_single.py&lt;/code&gt; are&lt;br&gt;
committed, byte-reproducible, and independently rebuildable from&lt;br&gt;
&lt;code&gt;src/forge/&lt;/code&gt;. &lt;code&gt;deps-proof.txt&lt;/code&gt; statically checks every import against&lt;br&gt;
the standard library, then re-checks it under an isolated interpreter&lt;br&gt;
with no site-packages on &lt;code&gt;sys.path&lt;/code&gt; at all. &lt;code&gt;STDLIB.md&lt;/code&gt; documents 21&lt;br&gt;
substitutions, each backed by an actual import in the codebase, not a&lt;br&gt;
hypothetical.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python scripts/judge_mode.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;runs all of the above — plus the confidence-score reconciliation, the&lt;br&gt;
risk-timeline formula, the SARIF 2.1.0 output, and a deliberate&lt;br&gt;
storage-corruption test — as one command, in under 15 seconds, and exits&lt;br&gt;
non-zero if any claim in this post doesn't hold up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forge is my submission to &lt;a href="https://dev.to/raptorsdev"&gt;Hackathon Raptors&lt;/a&gt;'&lt;br&gt;
Zero Dependency Hackathon — Track F (Open/Wildcard).&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing thought
&lt;/h2&gt;

&lt;p&gt;Anyone with an AI coding agent can produce a scanner that finds AWS keys&lt;br&gt;
with a regex — that part isn't the hard 20%. The hard part was building&lt;br&gt;
enough ways to catch myself being wrong: a &lt;code&gt;judge_mode.py&lt;/code&gt; that&lt;br&gt;
re-derives every claim in this post instead of asking a judge to trust&lt;br&gt;
it, a real-world validation that kept the 327-finding run instead of&lt;br&gt;
only publishing the fixed one, and a benchmark table honest enough to&lt;br&gt;
put a 5–6× loss next to a 2× win. The scanner is the artifact. The&lt;br&gt;
willingness to publish the run that didn't flatter it is the actual&lt;br&gt;
submission.&lt;/p&gt;

&lt;h3&gt;
  
  
  Project
&lt;/h3&gt;

&lt;p&gt;Built for Zero Dependency 2026, run by &lt;a href="https://dev.to/raptorsdev"&gt;Hackathon Raptors&lt;/a&gt;&lt;br&gt;
(&lt;a href="https://dev.to/partnerships_raptors"&gt;@partnerships_raptors&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/broforce6909-cmd/forge" rel="noopener noreferrer"&gt;GitHub Repository — broforce6909-cmd/forge&lt;/a&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  ZeroDependencyHack #Python #OpenSource #DeveloperTools #CLI
&lt;/h1&gt;

&lt;h1&gt;
  
  
  PythonStandardLibrary #BuildInPublic
&lt;/h1&gt;

</description>
      <category>python</category>
      <category>opensource</category>
      <category>security</category>
      <category>hackathon</category>
    </item>
  </channel>
</rss>
