A monorepo breaks the naive secret scan in two directions at once. The full-repository scan is slow because it walks every file in every service, and it is noisy because the test fixtures of twelve services all look like secrets to a pattern matcher. The fix is not a smarter scanner. It is pointing the same scanner at the right directories, and the tool's path flag exists exactly for this. You scan the deploy directory where the real credentials live, you scan the one service that just changed, and you skip the fixture farm entirely. The scan that used to take a while takes seconds, and the findings that used to be buried in noise are the only findings in the report.
The other direction is the whole-tree sweep, which is the recursive flag pointed at the projects directory itself, and it is the quarterly deep clean rather than the daily check. Leaks do not respect repository boundaries; the forgotten experiment from last year is a live credential the moment someone finds it, and a monorepo is where forgotten experiments live. So the practice is two rhythms: the targeted path scan on every change to a service that touches production config, and the full recursive sweep once a quarter over everything. The same command, the same rules, two cadences. The monorepo does not need a different security tool; it needs the same tool aimed at the right scope, which is a cheaper and more reliable fix than any amount of scanner cleverness.
The walkthrough assumes the tool is already one command away. dotguard is Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds. Run npx @wuchunjie/dotguard and the help output is the table of contents for everything below. The sections follow the order a first run actually goes, not the order the README presents it in, because the first run is where the questions happen and the README is written for the second run. Where a step can fail, the failure is named, because the failure is the part the walkthrough exists to cover, and the covered failure is the difference between the tutorial and the experiment. By the end you have run the full path once, and the once is what makes the next run the fast one.
The Rotation Checklist the Report Triggers
The finding without the rotation is a list, and the list is the part that ages badly. The rotation checklist is the four steps that turn the finding into the closed loop, and the steps are short enough to be a habit. Step one is the verify, the five-second check that the value is live and not a fixture. Step two is the rotate, the call to the provider that replaces the key and invalidates the old one. Step three is the update, the new value in the place the build reads, and the old value deleted from everywhere the working tree still holds it. Step four is the re-scan, the same command run again, and the re-scan is the proof, because the proof is the old value not in the report. The four steps take the length of a coffee, and the coffee is the part that makes the loop close on the day of the finding instead of the day of the reminder. The report is the trigger, and the trigger without the loop is the notification, and the notification is what gets ignored.
Monorepos, Home Directories, and the Recursive Flag
The scanner assumes your secrets live in one project, and then breaks that assumption on purpose. The path flag lets you point it at a subdirectory — scan just the deploy configs, or just one service in the tree. The recursive flag lets you do the opposite: point it at your projects directory and audit everything you have ever written, including the old side projects you forgot about. That last use is the one that pays for itself. Leaks do not respect project boundaries; the forgotten demo repository from a few years ago is a live credential the moment someone finds it. A scanner that can sweep an entire directory tree in seconds turns auditing all your code from a month-long project into a coffee break. The monorepo use is the inverse: one command covers every service, so the scan cost does not grow with the organization. Either direction, the same principle: the scan should be cheap enough to run often, and running it often is what catches the leak while it is still cheap to fix.
What Zero Dependencies Buys You in a Security Tool
A security tool has a special trust problem: you are asking it to read your most sensitive files. The natural question is what the tool itself trusts. For this one, the answer is nothing. No dependencies, no network calls, no telemetry, no update daemon. It reads files, matches patterns, and writes a report. That matters in the places where security tooling gets blocked: restricted CI runners, air-gapped builds, compliance environments that require an audit of every third-party package in the pipeline. A single-file scanner with zero dependencies is auditable in an afternoon by a security reviewer who would never approve a forty-package tree. In security, small is not a feature. Small is the product. The auditability is the trust model: you can read the whole thing, you can verify what it matches, and you can be confident that the thing reading your secrets is not also phoning home with them. That confidence is not a nice property of the design; it is the design.
Two Rhythms: The Daily Path and the Quarterly Tree
The scanner runs on two cadences, and the two cadences are the practice. The daily one is the targeted scan: the path pointed at the directory that changed, the service that touched the config, the deploy folder that holds the real credentials. The targeted scan is fast enough to run before the push, and the fast is what makes it the reflex. The quarterly one is the recursive sweep: the flag pointed at the whole projects directory, the tree that includes the forgotten experiment and the demo repository and the fork that nobody maintains. The sweep is the deep clean, and the deep clean is what catches the leak that is already old, because the old leak is the one the targeted scan never points at, and the never-pointed-at is where the leak waits. The same command, the same rules, two scopes, two rhythms. The monorepo does not need a different tool for the two cadences, it needs the same tool aimed at the two scopes, and the aiming is the practice, and the practice is what the scanner's flags exist to make cheap.
Scanning the Worktree, and the History Behind It
The scanner covers what is on disk: the environment files, the config files, the source in your working tree. That catches the obvious case — the file you just created and are about to commit. For the subtler case, the secret that is already in history, you combine it with the version control system: find the files that ever contained the pattern, then scan them. A committed secret does not stop being a secret when it is deleted from the current branch; it lives in every clone, every fork, and every mirror. The honest workflow is two steps: scan the present automatically, and audit the past with the same rules applied to the files that history touched. Detection is a habit, and habits are easier to keep when the tooling is small enough to run by reflex. The present scan is the reflex; the past audit is the quarterly deep clean, and both use the same rules, which is what makes the pair coherent instead of two unrelated chores.
The Rules: What the Scanner Actually Matches
The scanner works on two rule families, and knowing both is what makes the report readable. The first family is the known-shape rule: the cloud provider key prefixes, the git token formats, the webhook secret shapes, the database URL patterns. These are the rules with names, and they fire on the credentials the ecosystem has already standardized. The second family is the entropy heuristic: a string with the statistical fingerprint of a generated secret, assigned to a variable whose name says what it is. The heuristic is the net for the custom key, the one that follows no documented format, and it is the family that needs the five-second human check, because the heuristic is a candidate, not a verdict. The report marks which family fired, and the mark is what separates the rotate-today finding from the probably-a-fixture one. The two families together cover the cases that matter, and the report tells you which one caught the finding, which is the part that makes the triage fast instead of anxious.
Secrets Hide in Unlikely Places
The environment file is where people look, and it is where secrets are least well hidden. The real distribution is wider: a database URL inside a container compose file, an API key in a YAML config under a deploy directory, a webhook token in a JSON file checked in for one test, a token pasted into a README as a working example. The scanner covers source code, not just dotfiles, precisely because the pattern of where secrets actually live is messier than the pattern of where people think they live. If your security review only opens environment files, you are reviewing one out of five places. The scanner's job is to make the review exhaustive by making it automatic, and the exhaustive part is the whole value. The surprising findings — the ones in files nobody would call secret files — are the ones that justify the tool, because they are the ones no checklist would ever reach. A checklist asks about secrets; the scanner asks about every file, which is a different and better question.
The Cost of a Leaked Key, Quantified
The numbers on public cloud incidents are uncomfortable in a specific way: they are all larger than the annual budget of the team that leaked the key. Compromised cloud credentials have produced bills in the hundreds of thousands of dollars in a single weekend. Leaked repository tokens have been used to push malicious code into downstream packages that millions of installs then inherit. Leaked database passwords have emptied tables into the clear. The economics are one-directional: the cost of a scan is a line of workflow file and a few seconds of CPU; the cost of a miss is measured in six figures and a postmortem you will give in the morning with your team in the room. The asymmetry is the entire argument for running the scanner by default, on every push, in every repository, whether or not you think you have secrets in it. The repositories that are sure they have no secrets are the ones that have never checked, and the check is the only difference between those two statements.
Writing the Detection Rules Is a Window Into Developer Habits
The detection rule set is a fossil record of how developers actually handle secrets. The cloud provider rule catches the most leaks by volume, which says something about how much infrastructure runs on a single provider. The web-token rule catches a specific habit: developers pasting a full token into code to debug one endpoint, then forgetting it is there. The high-entropy rule exists because a surprising number of teams generate strong keys and store them with weak discipline — random value, obvious variable name, plaintext file. Building the rules taught me that secret leaks are not a knowledge problem; everyone knows the environment file should not be committed. They are a friction problem. The right thing is slow, the wrong thing is fast, and the scanner removes the friction from the right thing by making the check automatic. Every rule in the set is a documented instance of the friction, and the rule set is a map of where the friction lives in a codebase.
False Positives Are the Real Cost of a Scanner
Every secret scanner has to make a tradeoff: miss a real key, or flag a false one. The scanner leans toward flagging, and the price is occasional false positives — an encoded blob that looks like a key, a test fixture with a fake credential, a documentation example that uses a real-looking prefix. That is why the output includes the file, the line, and the matched rule: the cost of verifying a finding should be five seconds, not a forensic exercise. A scanner with zero false positives that also misses real keys is a liability, not a tool. Budget a minute per finding, verify, rotate if it is real, and note the known test values so the team stops re-checking them. The loop is the product. Over time the false-positive list becomes its own artifact — a record of the places in the codebase that look like secrets and are not, which is useful information in its own right, because it maps the codebase's sensitive-looking surfaces.
The Team-Wide Hook: Distributing the Habit
Personal hooks die with the person who set them up. The pattern that actually works is distribution: put the scan in a hook that lives in the repository itself, committed to version control, so every clone gets the behavior automatically. A new developer clones the repository, runs their first command, and the scanner is already there. Nobody has to be told, nobody has to remember, and the behavior cannot be forgotten because it is part of the repository. This is the difference between a tool and a practice: a tool is what one person runs, a practice is what the repository does to everyone. The scanner is small enough that the distributed version costs nothing to maintain — one file, one command, zero configuration — which is exactly the size at which a team habit becomes cheaper than the individual habit it replaces. The repository becomes the enforcer, and enforcement by repository is the only enforcement that survives turnover.
The takeaway
The walkthrough is done, and the done is the state where the next run is the fast one. dotguard is the tool that made the path the short one: Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds. The install line, for the next reader who is starting now, is npx @wuchunjie/dotguard, and the repository is https://github.com/wuchunjie00/dotguard. The steps above are the ones that work, and the work is the part that the version keeps working, because the keeps is the part the maintenance is. If a step failed in your run, the failure is the thing to report, because the report is what the next run reads, and the reads is what the fix becomes.
Top comments (0)