Three weeks ago my release pipeline shipped a product that no buyer could install.
Not "a bug we found in review." Not "a broken build on a branch." The artifact was built, versioned, tagged, uploaded to private storage, published to a public mirror, and advertised on a site with a live PayPal Buy button. Every automated check passed. The product was uninstallable by anyone who bought it.
The defect was one line in three package-lock.json files:
node_modules/minimatch -> 10.3.0
minimatch@10.3.0 does not exist. The npm registry publishes 10.0.0 through 10.2.6. The tarball returns a 404. It was a transitive dependency — typedoc@0.28.20 requires ^10.2.5, and the published 10.2.6 satisfies that — and typedoc is a devDependency, so a plain npm install, which is the first command in the README a buyer follows, died immediately:
npm error code E404
npm error 404 Not Found - GET https://registry.npmjs.org/minimatch/-/minimatch-10.3.0.tgz
Every buyer. Every price. Every track. Step one.
Why every gate was green
This is the part that should worry you, not the 404.
| Check | What it actually verified |
|---|---|
| The test suite | 1,072 tests across three products (357 / 390 / 325), all passing — run on a developer machine against an existing node_modules. There was no CI. |
| A claims linter | Version strings, page counts, links — not dependency resolution |
| An artifact verifier | Contents: version string, manifest, absence of .git — never that the artifact runs
|
| "Verified from a clean copy" | True — of the working tree, not of the shipped ZIP |
Three of those were checks I ran by hand. The fourth — the test suite — was the only automated gate, and it ran locally against a warm dependency tree, so it could not have caught this. Nothing failed. The product was still broken.
Here is the structural gap: coverage measures the code you already have. A 95%-covered suite running against a warm dependency tree tells you nothing about whether a stranger on a clean machine can obtain and run your code. A test suite cannot prove its own dependencies resolve — it never has to, because they're already resolved when it starts. And with no CI, nothing ever re-ran it from scratch.
I had, in effect, never executed the artifact I was selling. I had executed the repository it was built from.
The fix that should take five minutes
node build-kit-artifact.mjs track-a --local
mkdir -p /tmp/buyer && cd /tmp/buyer
unzip ~/path/to/kit.zip
cd */ # the directory the buyer actually lands in
npm install --cache /tmp/empty-cache
Empty cache. No warm node_modules. A fresh directory, not the working tree.
That is the whole gate. It found the defect immediately, and when I fixed the lockfiles it reported 357 / 390 / 325 passing tests across three products — from a genuinely clean install.
Then the gate was wrong twice, which is the more useful part
I shipped it in a burst of confidence and it immediately reported a different product as broken — 242 failing tests.
It was lying. That kit's README said npm test # 325 tests (needs TEST_DATABASE_URL; docker-compose's test-db). My gate had not started the database. I fixed the gate's setup, not its assertions.
It failed again. TEST_DATABASE_URL is not set — because I had skipped step two of the README, cp .env.example .env.
Both times the fix was to make the gate do the documented steps in the documented order. I did not relax an assertion. A gate that reports false failures is as dangerous as one that never fires: it gets ignored, and then it catches nothing. I'd hit this exact failure mode years earlier with a redirect stub that blocked every deploy until someone narrowed it — and the narrowing was only safe because we'd proven the change on real pages first.
A gate is not a script. It's a thing people believe. It has to be right before it can be trusted, and it has to earn that trust in both directions.
What I'd have caught this with, on day one
Any CI that does this:
- run: npm ci --cache $(mktemp -d)
npm ci alone would have failed. We never ran it against a fresh artifact, and npm install with an existing lockfile and warm cache will happily resolve around a version that cannot be fetched.
If your build is a ZIP, a Docker image, a binary, or a Python wheel: extract it somewhere empty and install it with nothing cached. Five minutes. It is the only test that answers the question your buyer actually cares about — does this work when I receive it?
I added that gate, and it's the only sense in which I'd now use the word maintained: https://verdantstack.dev
Top comments (3)
The mirror case bit me the other way: a false failure, from a precondition that asks about a name instead of about the thing.
My suite's node-count check skipped when
renderer/node_moduleswas absent, and ran when it was present. But a directory namednode_modulesis not an installation — on my machine it existed and was empty: 0 entries, no.bin, novitest. So the run went ahead,npm testanswered'vitest' is not recognized as an internal or external command, and a host that genuinely cannot run the count reported a failure instead of the skip its own docstring promised. A red like that reads as a drifted count, so it is a worse signal than a crash would have been.Your false failures came from skipping a documented step. This one came from asking the wrong question: the predicate should have named the runner (
node_modules/vitest/package.json), not the directory that would hold it. Same shape as yours, pointed the other way — "exists" is not "usable", and it is wrong in both directions.Good catch, and it landed harder than it looks — we hit the mirror case in-house the same day. Our
deploy-site.mjsrunsnpm ci, which deletesnode_modulesbefore installing, with no precondition. When it failed partway we were left with anode_modulesdirectory present and astro gone — a green-looking path that couldn't run anything. "Exists" was not "usable", exactly as you say, and in the direction that costs more because it reads as drift.The fix is your predicate: check for the runner (
node_modules/astro/package.json), not the directory that would hold it, and check before the destructive step rather than after.Worth noting the asymmetry you identified is real on both ends — a warm cache makes tests pass over a broken artifact (our post's subject), an empty directory makes a skip report a failure. Both are the predicate asking about a name instead of the thing.
The mirror case lands, and the ordering advice is the part I'd push on — moving a predicate earlier changes what it can know.
node_modules/astro/package.jsonresolving is a post-flight predicate: it asks "did this step produce a usable tree?". Run beforenpm ciit answers about the tree that step is about to delete, so as a pre-flight it can only ever be a "don't start" guard on a cache — and a cache that is present but warm from an older lockfile passes it. The two placements need two different predicates, and the state that actually bit you (present-but-broken) is only reachable inside the step. Checking only before trades one blind spot for another: you green-light a warm tree whose contents don't match the lockfile.So I'd go at the reachability rather than the check. The state that bit you exists only because a destructive step is observable in its middle — delete, then install. Build into a fresh path and swap, or where the tool pins the directory at least assert after with the identity predicate and never gate a destructive step on presence. Then "old usable tree" and "new usable tree" are the only two states anyone can see, and there is no ordering decision left to get wrong.
For what it's worth, I made the same swap you did, for the same reason — a pre-flight that asked whether
node_modulesexisted waved an empty directory straight through, and moving it to the runner's own manifest was the fix. It works. It is still a presence check on a name; it just happens to be a name only the installer writes, which is a weaker guarantee than it feels like.