DEV Community

ke jia
ke jia

Posted on

dotguard vs detect-secrets: The Node Scanner Against the Python Ecosystem Favorite

The secret scanning world has a Python favorite and a Node shop's answer, and they solve the same problem with different assumptions about where the work happens. detect-secrets is the tool the Python ecosystem reached for first, and it is a good tool with a specific shape: it runs in the Python world, it integrates with the pre-commit framework that Python teams already use, and its baseline mechanism for known false positives is well worked out. dotguard is the answer for the JavaScript and TypeScript side of the house: it runs via npx with zero dependencies, it reads the same kinds of files, and it is built to be one line in a workflow file rather than a plugin in a hook system.
The honest comparison is about the rule set and the integration surface, not about which tool is better, because the better tool is the one that runs where your code lives. The rule sets overlap on the common cases, the cloud key prefixes and the git tokens, and they differ at the edges in the ways that matter per ecosystem. The integration story is where the shop type decides: a Python monorepo with pre-commit already wired in gets more value from the tool that speaks pre-commit, and a Node service with a GitHub Actions workflow gets more from the tool that is one run line. The section below is the head to head: the same set of test files through both tools, the findings side by side, and the verdict for each shop type, which is the part the benchmark tables never get to.

The comparison is set up the way a benchmark should be: the same input, the same conditions, and the verdict stated before the detail, because the detail is for the reader who wants to check the work and the verdict is for the reader who wants the answer. The tools in question are dotguard, which is Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds., and the alternative, which gets the same courtesy. The alternative is not the villain; the alternative is the option the reader might have, and the might-have is the part the comparison is for. The sections below are the head-to-head: the setup, the run, the findings, and the honest note about where each one loses, because the note is what makes the verdict the verdict instead of the ad.

False Positives Are the Real Cost of a Scanner

Every secret scanner has to make a tradeoff: miss a real key, or flag a false one. The scanner leans toward flagging, and the price is occasional false positives — an encoded blob that looks like a key, a test fixture with a fake credential, a documentation example that uses a real-looking prefix. That is why the output includes the file, the line, and the matched rule: the cost of verifying a finding should be five seconds, not a forensic exercise. A scanner with zero false positives that also misses real keys is a liability, not a tool. Budget a minute per finding, verify, rotate if it is real, and note the known test values so the team stops re-checking them. The loop is the product. Over time the false-positive list becomes its own artifact — a record of the places in the codebase that look like secrets and are not, which is useful information in its own right, because it maps the codebase's sensitive-looking surfaces.

The Cost of a Leaked Key, Quantified

The numbers on public cloud incidents are uncomfortable in a specific way: they are all larger than the annual budget of the team that leaked the key. Compromised cloud credentials have produced bills in the hundreds of thousands of dollars in a single weekend. Leaked repository tokens have been used to push malicious code into downstream packages that millions of installs then inherit. Leaked database passwords have emptied tables into the clear. The economics are one-directional: the cost of a scan is a line of workflow file and a few seconds of CPU; the cost of a miss is measured in six figures and a postmortem you will give in the morning with your team in the room. The asymmetry is the entire argument for running the scanner by default, on every push, in every repository, whether or not you think you have secrets in it. The repositories that are sure they have no secrets are the ones that have never checked, and the check is the only difference between those two statements.

The Team-Wide Hook: Distributing the Habit

Personal hooks die with the person who set them up. The pattern that actually works is distribution: put the scan in a hook that lives in the repository itself, committed to version control, so every clone gets the behavior automatically. A new developer clones the repository, runs their first command, and the scanner is already there. Nobody has to be told, nobody has to remember, and the behavior cannot be forgotten because it is part of the repository. This is the difference between a tool and a practice: a tool is what one person runs, a practice is what the repository does to everyone. The scanner is small enough that the distributed version costs nothing to maintain — one file, one command, zero configuration — which is exactly the size at which a team habit becomes cheaper than the individual habit it replaces. The repository becomes the enforcer, and enforcement by repository is the only enforcement that survives turnover.

The False Positive Conversation

Every scanner has false positives, and the honest relationship with them is the part that keeps the tool in the workflow. The fixtures are the main source: the test that asserts against a fake key, the README example that shows a token shape, the mock that returns a secret-looking string. The false positive is not a bug in the scanner, it is the scanner doing its job on data that looks like the job, and the cost of the false positive is the five-second check, which is the cost the tool is designed around. The known-false-positive list is the management: the findings you have already classified stay classified, and the classification is what the report's rule field supports, because the rule tells you the shape that fired, and the shape is what the fixture is. The conversation to have with the team is not about removing the false positives, it is about budgeting for them: the five seconds each, and the list that grows, and the list that makes the next triage faster. The scanner that has no false positives is the scanner that misses the custom key, and the custom key is the one that matters.

The Exit Code Is the Integration

The scanner's interface to the machine is the exit code, and the exit code is the part that makes it CI-ready without any glue. A clean scan exits zero, a scan with findings exits nonzero, and the nonzero is what the pipeline turns into a failure, and the failure is what blocks the push. The pre-push hook uses the same code: the hook runs the scan, and the push goes through only if the exit code is zero. The JSON report is the second interface, for the things that consume data instead of pass/fail: the dashboard that charts findings over time, the chat alert that posts the file and the line, the compliance export that needs the record. The two interfaces cover the two consumers, the pipeline and the person, and the coverage is the design. A tool that needs a wrapper script to be useful in CI is a tool that will not be in CI, because the wrapper is the step that does not get written, and the not-written is the check that does not run.

One Line in GitHub Actions

The entire CI integration is one step: a run line that invokes the scanner via npx. No service container, no token to configure, no daemon. Every push gets a full scan, and the build fails if a secret is found, which means the secret never reaches the default branch. The beauty of the one-line integration is the maintenance cost: there is nothing to update, no version to pin, and no vendor to renew. When a security control costs one line of workflow file, the only question is why it is not already there. That is the bar every pre-merge security control should meet, and the scanner was designed to meet it on purpose. The one line is also the onboarding story: new contributors see the check in the workflow file, understand what it is doing, and never have to be told to run it. The pipeline is the policy, and the policy is one line long.

The 8-Month Leak That Started This

The scanner was built after an audit of my own public repositories found three leaked secrets. The worst was a production database password that had been committed, in plaintext, for eight months. Nothing dramatic happened — no breach, no incident report — which is exactly the problem. Most leaks are silent. They sit in a file that was committed in month one, get copied into forks, become the default credential in a demo, and nobody notices until the rotation happens by accident. A scanner is not a luxury for teams with a security budget. It is the same category of tool as a spellchecker: cheap, automatic, and the only thing standing between a careless commit and a very bad quarter. The eight-month leak is not a story about a mistake; it is a story about the absence of a check, and the check is the entire product.

The JSON Report That CI Can Actually Use

A security tool is only as good as its integration surface, and the scanner's is a JSON report. Run it with the report flag and you get a machine-readable list of every finding: file, line, rule, severity. Exit codes are meaningful, so a pipeline can fail the build on any finding without parsing colored terminal output. That design decision pays off in the boring way: the tool fits into CI or a local pre-push script with zero glue code. When the report is data instead of text, other tools can consume it — a dashboard, a chat alert, a compliance export. Security tooling that cannot be integrated is a demo, not infrastructure. The JSON report is the part of the design that turns a one-person tool into a team habit, and it costs nothing to use: the same command, one extra flag, and the output becomes something the rest of the system can act on.

How the Scan Actually Works

The scanner looks for the patterns that real secrets actually take. Cloud provider keys start with a known prefix, Git tokens start with a known prefix, messaging platform tokens have a known shape, payment processor live keys have a known prefix, and web tokens start with a known base64 header. It also checks for high-entropy strings assigned to suspicious variable names — password, token, secret, key — because the values do not always follow the format, but variable names are a reliable signal. The scan covers environment files, config files, and source code, and it reports the file, the line, and the matched pattern so a human can verify in seconds rather than minutes. It is deliberately a detection tool, not a verdict tool: it finds candidates, a human confirms, and the key gets rotated either way. That division of labor is what keeps false positives from becoming noise fatigue.

Five Seconds to Verify a Finding

The verification is the step that turns the finding into a decision, and the five seconds is the design target, because the five seconds is what keeps the check happening. The report gives the file, the line, and the rule, and the verification is open the file, look at the line, and answer one question: is this value live, or is it a fixture? The live value is the one that the build reads, that the deploy uses, or that a request would present to a real service. The fixture is the one that the test asserts against, that the README shows as an example, that the mock returns. The answer changes the next step: live means rotate and re-scan, fixture means note it and move on. The five seconds works because the report is specific, and the specificity is the feature. A scanner that says suspicious content found is a scanner that takes an hour to triage, and the hour is what stops the habit. The five seconds is the habit.

Two Rhythms: The Daily Path and the Quarterly Tree

The scanner runs on two cadences, and the two cadences are the practice. The daily one is the targeted scan: the path pointed at the directory that changed, the service that touched the config, the deploy folder that holds the real credentials. The targeted scan is fast enough to run before the push, and the fast is what makes it the reflex. The quarterly one is the recursive sweep: the flag pointed at the whole projects directory, the tree that includes the forgotten experiment and the demo repository and the fork that nobody maintains. The sweep is the deep clean, and the deep clean is what catches the leak that is already old, because the old leak is the one the targeted scan never points at, and the never-pointed-at is where the leak waits. The same command, the same rules, two scopes, two rhythms. The monorepo does not need a different tool for the two cadences, it needs the same tool aimed at the two scopes, and the aiming is the practice, and the practice is what the scanner's flags exist to make cheap.

The takeaway

The verdict, restated: the choice is by the context, and the context is the part the reader has and the comparison does not. dotguard wins the contexts the sections showed, and the showed is the part the detail supports, because the supports is what the verdict is not floating on. Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds. The install is npx @wuchunjie/dotguard, and the repository is https://github.com/wuchunjie00/dotguard, for the reader who wants the tool after the comparison instead of before it, because the after is the part that decides with the data in hand, and the data-in-hand is the reader the comparison is written for.

Top comments (0)