DEV Community

ke jia
ke jia

Posted on

7 Places Secrets Hide in a Typical Node Monorepo (Besides .env)

The .env file gets all the attention in the secret talk, and the attention is deserved, and the distribution of real leaks does not match the attention, which is the whole point of this list. Seven places in a typical Node monorepo hold secrets that the .env review will never see, and each one below has the file type, the credential shape that hides there, and the reason the review process misses it. The scanner's job is to make the review exhaustive by making it automatic, and the exhaustive part is the value, because a checklist that asks about secrets only ever finds the secrets people expected to be there.
The compose file is the first, and it is the most common offender outside the .env, because the database URL and the service credentials live there as environment variables for the containers, and the file is reviewed as infrastructure, not as a secret store. The YAML under the deploy directory is second, because the CI reads it and the credential in it is the one that makes the deploy work, which means it is a live credential in a file nobody calls a secret file. The remaining five are in the section below, each with the same three fields, and the list ends with the scan that covers all seven at once, because the point of the list is not to memorize the places, it is to run the check that finds them regardless of whether you remembered.

The list is the format, and the format is the promise: each entry is one thing, one use case, and one honest note, and the three are the unit the list is made of. The entries are ordered by the weight they carry in actual use, not by the order they were discovered, because the discovered order is the story and the weight order is the tool. The entries come from dotguard, Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds., and the tool is the context for the list, because the list is the tool's shape made explicit, and the explicit is the part the feature page does not do, and the does-not-do is what the list is for. The section below is the list, and the list is the section.

What Zero Dependencies Buys You in a Security Tool

A security tool has a special trust problem: you are asking it to read your most sensitive files. The natural question is what the tool itself trusts. For this one, the answer is nothing. No dependencies, no network calls, no telemetry, no update daemon. It reads files, matches patterns, and writes a report. That matters in the places where security tooling gets blocked: restricted CI runners, air-gapped builds, compliance environments that require an audit of every third-party package in the pipeline. A single-file scanner with zero dependencies is auditable in an afternoon by a security reviewer who would never approve a forty-package tree. In security, small is not a feature. Small is the product. The auditability is the trust model: you can read the whole thing, you can verify what it matches, and you can be confident that the thing reading your secrets is not also phoning home with them. That confidence is not a nice property of the design; it is the design.

The Files It Reads, and the Ones It Skips

The scanner's file coverage is the difference between the scan and the audit, and the coverage list is worth knowing. It reads the environment files, the config files in the common formats, the compose files, the YAML under the deploy directories, the JSON configs, and the source code, because the source is where the debug paste lives. It skips the directories that are noise by definition: the dependency folders, the build output, the lockfiles, because the dependency folders are the supply chain and the supply chain is audited separately, and the build output is derived and the derived is not the source of truth. The skip list is the tuning that makes the scan fast on a monorepo, and the fast scan is what runs often enough to matter. The coverage list is in the documentation, and the documentation is the thing to read once, because the one read is what separates the person who trusts the scan from the person who wonders what it missed, and the wondering is what stops the trust.

Writing the Detection Rules Is a Window Into Developer Habits

The detection rule set is a fossil record of how developers actually handle secrets. The cloud provider rule catches the most leaks by volume, which says something about how much infrastructure runs on a single provider. The web-token rule catches a specific habit: developers pasting a full token into code to debug one endpoint, then forgetting it is there. The high-entropy rule exists because a surprising number of teams generate strong keys and store them with weak discipline — random value, obvious variable name, plaintext file. Building the rules taught me that secret leaks are not a knowledge problem; everyone knows the environment file should not be committed. They are a friction problem. The right thing is slow, the wrong thing is fast, and the scanner removes the friction from the right thing by making the check automatic. Every rule in the set is a documented instance of the friction, and the rule set is a map of where the friction lives in a codebase.

Secrets Hide in Unlikely Places

The environment file is where people look, and it is where secrets are least well hidden. The real distribution is wider: a database URL inside a container compose file, an API key in a YAML config under a deploy directory, a webhook token in a JSON file checked in for one test, a token pasted into a README as a working example. The scanner covers source code, not just dotfiles, precisely because the pattern of where secrets actually live is messier than the pattern of where people think they live. If your security review only opens environment files, you are reviewing one out of five places. The scanner's job is to make the review exhaustive by making it automatic, and the exhaustive part is the whole value. The surprising findings — the ones in files nobody would call secret files — are the ones that justify the tool, because they are the ones no checklist would ever reach. A checklist asks about secrets; the scanner asks about every file, which is a different and better question.

Scanning the Worktree, and the History Behind It

The scanner covers what is on disk: the environment files, the config files, the source in your working tree. That catches the obvious case — the file you just created and are about to commit. For the subtler case, the secret that is already in history, you combine it with the version control system: find the files that ever contained the pattern, then scan them. A committed secret does not stop being a secret when it is deleted from the current branch; it lives in every clone, every fork, and every mirror. The honest workflow is two steps: scan the present automatically, and audit the past with the same rules applied to the files that history touched. Detection is a habit, and habits are easier to keep when the tooling is small enough to run by reflex. The present scan is the reflex; the past audit is the quarterly deep clean, and both use the same rules, which is what makes the pair coherent instead of two unrelated chores.

Monorepos, Home Directories, and the Recursive Flag

The scanner assumes your secrets live in one project, and then breaks that assumption on purpose. The path flag lets you point it at a subdirectory — scan just the deploy configs, or just one service in the tree. The recursive flag lets you do the opposite: point it at your projects directory and audit everything you have ever written, including the old side projects you forgot about. That last use is the one that pays for itself. Leaks do not respect project boundaries; the forgotten demo repository from a few years ago is a live credential the moment someone finds it. A scanner that can sweep an entire directory tree in seconds turns auditing all your code from a month-long project into a coffee break. The monorepo use is the inverse: one command covers every service, so the scan cost does not grow with the organization. Either direction, the same principle: the scan should be cheap enough to run often, and running it often is what catches the leak while it is still cheap to fix.

Five Seconds to Verify a Finding

The verification is the step that turns the finding into a decision, and the five seconds is the design target, because the five seconds is what keeps the check happening. The report gives the file, the line, and the rule, and the verification is open the file, look at the line, and answer one question: is this value live, or is it a fixture? The live value is the one that the build reads, that the deploy uses, or that a request would present to a real service. The fixture is the one that the test asserts against, that the README shows as an example, that the mock returns. The answer changes the next step: live means rotate and re-scan, fixture means note it and move on. The five seconds works because the report is specific, and the specificity is the feature. A scanner that says suspicious content found is a scanner that takes an hour to triage, and the hour is what stops the habit. The five seconds is the habit.

The JSON Report That CI Can Actually Use

A security tool is only as good as its integration surface, and the scanner's is a JSON report. Run it with the report flag and you get a machine-readable list of every finding: file, line, rule, severity. Exit codes are meaningful, so a pipeline can fail the build on any finding without parsing colored terminal output. That design decision pays off in the boring way: the tool fits into CI or a local pre-push script with zero glue code. When the report is data instead of text, other tools can consume it — a dashboard, a chat alert, a compliance export. Security tooling that cannot be integrated is a demo, not infrastructure. The JSON report is the part of the design that turns a one-person tool into a team habit, and it costs nothing to use: the same command, one extra flag, and the output becomes something the rest of the system can act on.

Rotation Beats Detection, Every Time

There is a comfortable misconception that finding a secret is the win. It is not. The win is the key that no longer works. A detected secret that stays valid is a secret that is still leaking; the only difference is that now you know about it, which is a strange kind of liability. The workflow that matters is: scan, confirm, rotate, and then scan again to confirm the rotation happened and the old value is gone from the working tree or accepted as a known dead value. Rotation is boring, it touches other teams, and it is the step everyone skips. A detection tool that makes rotation feel like a natural next step — by giving you the exact file and line, so the fix is a two-minute job — is doing more than pattern matching. The finding is a ticket, and the ticket's definition of done is the rotated key, not the closed alert. Designing the tool around that definition of done is what separates a scanner from a security workflow.

What a Clean Scan Does for the Team

The clean scan is the output nobody celebrates, and the no-celebration is the point, because the clean scan is the state the team should be in by default. The value of the default-clean is the attention: the team's security attention goes to the findings, and the findings are the exceptions, and the exceptions are what get the rotate-today treatment instead of the quarter-review treatment. The clean scan also does the onboarding: the new developer who runs the scan and gets nothing learns the habit without the incident, and the habit learned without the incident is the habit that holds, because the first experience is the quiet one instead of the alarming one. The quiet is the design goal, and the quiet is what the pre-push hook and the CI gate buy, because the gate catches the leak before the team sees it, and the team that never sees the leak is the team that never has the incident meeting. The clean scan is the boring output, and the boring output is the expensive one to keep.

False Positives Are the Real Cost of a Scanner

Every secret scanner has to make a tradeoff: miss a real key, or flag a false one. The scanner leans toward flagging, and the price is occasional false positives — an encoded blob that looks like a key, a test fixture with a fake credential, a documentation example that uses a real-looking prefix. That is why the output includes the file, the line, and the matched rule: the cost of verifying a finding should be five seconds, not a forensic exercise. A scanner with zero false positives that also misses real keys is a liability, not a tool. Budget a minute per finding, verify, rotate if it is real, and note the known test values so the team stops re-checking them. The loop is the product. Over time the false-positive list becomes its own artifact — a record of the places in the codebase that look like secrets and are not, which is useful information in its own right, because it maps the codebase's sensitive-looking surfaces.

The takeaway

The list is the section, and the section is done, which means the entries are the content and the content is the part the reader scans. dotguard is the tool behind the list: Scan .env files, config files, and source code for exposed secrets. Zero dependencies, JSON reports, 1000+ files in seconds. The install line is npx @wuchunjie/dotguard, and the line is the part that makes the entries runnable, because the runnable is what the list entry promises and the promise is the part the reader checks. The repository is https://github.com/wuchunjie00/dotguard, and the repository is where the next entry goes, because the next-entry is the part the list that stops growing loses, and the loses is what the maintained list does not.

Top comments (0)