Security habits fail for a boring reason: they depend on remembering. This setup removes the remembering. It takes ten minutes, and after it the check happens before every push without any input from you, which is the entire design point. The first two minutes are the tool itself. Run it via npx, point it at the current directory, read what it finds. A clean repo gives you nothing, which is the best possible output from a scanner. A dirty one gives you a file, a line, and the pattern that matched, which turns verification into a five-second job instead of a forensic one.
The middle five minutes wire the habit into the motion you already make. A pre-push hook with a single line runs the scan on every push from your machine, so a leaked key dies in your terminal instead of on the remote. The final three minutes add the second layer: one line in the CI workflow, so a push that slips past your local hook still gets scanned before it reaches the default branch. Two layers, ten minutes, zero configuration. The habit stops living in your head and starts living in the repository, and the repository does not get tired, does not skip the check because the day is busy, and does not leave the company.
The walkthrough assumes the tool is already one command away. dotguard is Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds. Run npx @wuchunjie/dotguard and the help output is the table of contents for everything below. The sections follow the order a first run actually goes, not the order the README presents it in, because the first run is where the questions happen and the README is written for the second run. Where a step can fail, the failure is named, because the failure is the part the walkthrough exists to cover, and the covered failure is the difference between the tutorial and the experiment. By the end you have run the full path once, and the once is what makes the next run the fast one.
Monorepos, Home Directories, and the Recursive Flag
The scanner assumes your secrets live in one project, and then breaks that assumption on purpose. The path flag lets you point it at a subdirectory — scan just the deploy configs, or just one service in the tree. The recursive flag lets you do the opposite: point it at your projects directory and audit everything you have ever written, including the old side projects you forgot about. That last use is the one that pays for itself. Leaks do not respect project boundaries; the forgotten demo repository from a few years ago is a live credential the moment someone finds it. A scanner that can sweep an entire directory tree in seconds turns auditing all your code from a month-long project into a coffee break. The monorepo use is the inverse: one command covers every service, so the scan cost does not grow with the organization. Either direction, the same principle: the scan should be cheap enough to run often, and running it often is what catches the leak while it is still cheap to fix.
What a Clean Scan Does for the Team
The clean scan is the output nobody celebrates, and the no-celebration is the point, because the clean scan is the state the team should be in by default. The value of the default-clean is the attention: the team's security attention goes to the findings, and the findings are the exceptions, and the exceptions are what get the rotate-today treatment instead of the quarter-review treatment. The clean scan also does the onboarding: the new developer who runs the scan and gets nothing learns the habit without the incident, and the habit learned without the incident is the habit that holds, because the first experience is the quiet one instead of the alarming one. The quiet is the design goal, and the quiet is what the pre-push hook and the CI gate buy, because the gate catches the leak before the team sees it, and the team that never sees the leak is the team that never has the incident meeting. The clean scan is the boring output, and the boring output is the expensive one to keep.
The 8-Month Leak That Started This
The scanner was built after an audit of my own public repositories found three leaked secrets. The worst was a production database password that had been committed, in plaintext, for eight months. Nothing dramatic happened — no breach, no incident report — which is exactly the problem. Most leaks are silent. They sit in a file that was committed in month one, get copied into forks, become the default credential in a demo, and nobody notices until the rotation happens by accident. A scanner is not a luxury for teams with a security budget. It is the same category of tool as a spellchecker: cheap, automatic, and the only thing standing between a careless commit and a very bad quarter. The eight-month leak is not a story about a mistake; it is a story about the absence of a check, and the check is the entire product.
Rotation Beats Detection, Every Time
There is a comfortable misconception that finding a secret is the win. It is not. The win is the key that no longer works. A detected secret that stays valid is a secret that is still leaking; the only difference is that now you know about it, which is a strange kind of liability. The workflow that matters is: scan, confirm, rotate, and then scan again to confirm the rotation happened and the old value is gone from the working tree or accepted as a known dead value. Rotation is boring, it touches other teams, and it is the step everyone skips. A detection tool that makes rotation feel like a natural next step — by giving you the exact file and line, so the fix is a two-minute job — is doing more than pattern matching. The finding is a ticket, and the ticket's definition of done is the rotated key, not the closed alert. Designing the tool around that definition of done is what separates a scanner from a security workflow.
How the Scan Actually Works
The scanner looks for the patterns that real secrets actually take. Cloud provider keys start with a known prefix, Git tokens start with a known prefix, messaging platform tokens have a known shape, payment processor live keys have a known prefix, and web tokens start with a known base64 header. It also checks for high-entropy strings assigned to suspicious variable names — password, token, secret, key — because the values do not always follow the format, but variable names are a reliable signal. The scan covers environment files, config files, and source code, and it reports the file, the line, and the matched pattern so a human can verify in seconds rather than minutes. It is deliberately a detection tool, not a verdict tool: it finds candidates, a human confirms, and the key gets rotated either way. That division of labor is what keeps false positives from becoming noise fatigue.
What Zero Dependencies Buys You in a Security Tool
A security tool has a special trust problem: you are asking it to read your most sensitive files. The natural question is what the tool itself trusts. For this one, the answer is nothing. No dependencies, no network calls, no telemetry, no update daemon. It reads files, matches patterns, and writes a report. That matters in the places where security tooling gets blocked: restricted CI runners, air-gapped builds, compliance environments that require an audit of every third-party package in the pipeline. A single-file scanner with zero dependencies is auditable in an afternoon by a security reviewer who would never approve a forty-package tree. In security, small is not a feature. Small is the product. The auditability is the trust model: you can read the whole thing, you can verify what it matches, and you can be confident that the thing reading your secrets is not also phoning home with them. That confidence is not a nice property of the design; it is the design.
The Files It Reads, and the Ones It Skips
The scanner's file coverage is the difference between the scan and the audit, and the coverage list is worth knowing. It reads the environment files, the config files in the common formats, the compose files, the YAML under the deploy directories, the JSON configs, and the source code, because the source is where the debug paste lives. It skips the directories that are noise by definition: the dependency folders, the build output, the lockfiles, because the dependency folders are the supply chain and the supply chain is audited separately, and the build output is derived and the derived is not the source of truth. The skip list is the tuning that makes the scan fast on a monorepo, and the fast scan is what runs often enough to matter. The coverage list is in the documentation, and the documentation is the thing to read once, because the one read is what separates the person who trusts the scan from the person who wonders what it missed, and the wondering is what stops the trust.
Secrets Hide in Unlikely Places
The environment file is where people look, and it is where secrets are least well hidden. The real distribution is wider: a database URL inside a container compose file, an API key in a YAML config under a deploy directory, a webhook token in a JSON file checked in for one test, a token pasted into a README as a working example. The scanner covers source code, not just dotfiles, precisely because the pattern of where secrets actually live is messier than the pattern of where people think they live. If your security review only opens environment files, you are reviewing one out of five places. The scanner's job is to make the review exhaustive by making it automatic, and the exhaustive part is the whole value. The surprising findings — the ones in files nobody would call secret files — are the ones that justify the tool, because they are the ones no checklist would ever reach. A checklist asks about secrets; the scanner asks about every file, which is a different and better question.
The JSON Report That CI Can Actually Use
A security tool is only as good as its integration surface, and the scanner's is a JSON report. Run it with the report flag and you get a machine-readable list of every finding: file, line, rule, severity. Exit codes are meaningful, so a pipeline can fail the build on any finding without parsing colored terminal output. That design decision pays off in the boring way: the tool fits into CI or a local pre-push script with zero glue code. When the report is data instead of text, other tools can consume it — a dashboard, a chat alert, a compliance export. Security tooling that cannot be integrated is a demo, not infrastructure. The JSON report is the part of the design that turns a one-person tool into a team habit, and it costs nothing to use: the same command, one extra flag, and the output becomes something the rest of the system can act on.
One Line in GitHub Actions
The entire CI integration is one step: a run line that invokes the scanner via npx. No service container, no token to configure, no daemon. Every push gets a full scan, and the build fails if a secret is found, which means the secret never reaches the default branch. The beauty of the one-line integration is the maintenance cost: there is nothing to update, no version to pin, and no vendor to renew. When a security control costs one line of workflow file, the only question is why it is not already there. That is the bar every pre-merge security control should meet, and the scanner was designed to meet it on purpose. The one line is also the onboarding story: new contributors see the check in the workflow file, understand what it is doing, and never have to be told to run it. The pipeline is the policy, and the policy is one line long.
The Rules: What the Scanner Actually Matches
The scanner works on two rule families, and knowing both is what makes the report readable. The first family is the known-shape rule: the cloud provider key prefixes, the git token formats, the webhook secret shapes, the database URL patterns. These are the rules with names, and they fire on the credentials the ecosystem has already standardized. The second family is the entropy heuristic: a string with the statistical fingerprint of a generated secret, assigned to a variable whose name says what it is. The heuristic is the net for the custom key, the one that follows no documented format, and it is the family that needs the five-second human check, because the heuristic is a candidate, not a verdict. The report marks which family fired, and the mark is what separates the rotate-today finding from the probably-a-fixture one. The two families together cover the cases that matter, and the report tells you which one caught the finding, which is the part that makes the triage fast instead of anxious.
The takeaway
The walkthrough is done, and the done is the state where the next run is the fast one. dotguard is the tool that made the path the short one: Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds. The install line, for the next reader who is starting now, is npx @wuchunjie/dotguard, and the repository is https://github.com/wuchunjie00/dotguard. The steps above are the ones that work, and the work is the part that the version keeps working, because the keeps is the part the maintenance is. If a step failed in your run, the failure is the thing to report, because the report is what the next run reads, and the reads is what the fix becomes.
Top comments (0)