I build an audit tool. Its whole proposition is that its records can be trusted. So the defects we found in the last release were embarrassing in a specific way: almost every one of them was the tool claiming something it had not established.
They were written months apart, by different people, in subsystems that share no code. I only noticed they were one defect when the fourth one turned up.
Here is the shape.
A check confirmed that something was present. The result was reported as a claim that it worked.
Once you have the sentence, you start finding them everywhere.
The catalogue
A SIEM emitter that reported success without sending anything. The code fired four HTTP requests through setImmediate, attached .catch(() => undefined) to each, slept 800ms, and printed a green check. With no SIEM configured it printed the same green check, having made no request at all. Exit 0.
A coverage table that said a client was logging because it once had. Per-tool status read configured ✓ logged ✓. The second tick came from sessions > 0 — has this client ever recorded anything. One client had stopped capturing weeks earlier. The tick stayed green the entire time.
A doctor command that certified a launcher pointing nowhere. It checked whether a config contained something shaped like our entry. It did. The command that entry would run lived inside a temporary directory from an old test run — the kind the operating system clears on its own schedule. Configured, said the tool. Three clients were pinned to that throwaway install when we finally looked.
A signed-export flag that produced unsigned archives. --signed tried to sign, and when it couldn't — missing private key, more than one session selected — it caught the error, wrote the archive anyway, printed signed: no, and exited 0.
A verifier that passed those archives. Missing signature? A warning, then exit 0. Then we fixed that, and found that a signature file containing {} still passed, because the new check tested that the file could be read rather than that it contained a signature. Presence, reported as function, inside the fix for presence being reported as function.
An importer that made vendor exports look like first-hand observation. One import path recorded provenance metadata. The one beside it didn't. Records that came from a third-party export file were therefore indistinguishable, to every downstream reader, from sessions the tool had watched happen.
A summary that counted machine traffic as conversation. Turns were counted by role. A tool result is stored with role user — nobody typed it. A context-compaction marker is stored with role system — nobody said it. The code counted anything that wasn't user as AI, so both landed in the conversation totals.
Two implementations of one filter that disagreed. The table hid reviewed items by default. The JSON output didn't. A person and a script asking the same question got different answers, and only the script's was wrong.
The one that travels
These are not all equally bad, and the difference is worth naming.
A false "emitted" misleads the operator standing in front of the terminal. Annoying, recoverable, contained.
An unsigned archive from a --signed flag leaves the building. Someone receives a file whose name asserts that its origin can be established. They run the verifier. They get a zero. Nothing in that sequence tells them the archive proves nothing about where it came from.
Severity here isn't about how wrong the claim is. It's about how far the claim travels from the person able to check it.
So the fix wasn't "print a warning". It was: refuse. If --signed can't sign, it fails and writes nothing at the output path — because an unsigned file sitting where a --signed run was asked for is precisely what someone picks up later and trusts.
The one in our own pipeline
We fixed all of that. Then we released it, and the release workflow printed:
✓ Published chron-mcp@0.1.60 to npm
The registry 404'd for the version. For six minutes, the package did not exist.
npm publish returning 0 means the registry accepted the upload. npm says so itself, in output our own log captured: "Your package is being processed and may take a few minutes to become available." We printed our own success line immediately underneath it.
Worse: the step that syncs our public mirror ran next, and succeeded. For those six minutes the public repository advertised code that nobody could install.
Exit code, reported as availability. The same bug, in the pipeline that shipped the fixes for the same bug.
It now polls the registry until the version is present, the latest tag matches it, and a fresh install from the public registry reports the right version — then prints published. Three conditions, because any one alone can be true while the release is unusable.
The one I made while writing this up
I reported that npm had been 404ing "for about forty minutes."
It was six minutes and nine seconds. I had been comparing local timestamps against UTC ones and never did the subtraction. I filed the alarm sixteen seconds before the registry converged.
Forty minutes of silence from a package registry is an incident. Six is ordinary asynchronous propagation, which is what npm's output had said was happening all along. I'd spent two days removing unverified claims from a product and then made one, in the report about it, about a number I could have computed in one line.
The underlying finding was real. The evidence I gave for it was not.
What actually changed
Fixing twelve instances would have been worthless. What mattered was the structural changes:
Separate the three states, and only call one of them evidence. Our client diagnostics now report configured (an entry exists), loaded (for one client: not determinable from disk, and we say so), and captured (a session from this client is in the database, with when). Only the third is evidence, and only for the moment it happened.
Make the honest state reachable. Our delivery model had an unknown state for "sent, no trustworthy answer". It could never actually occur, because bare fetch has no timeout — a server that accepted a connection and never replied would hang forever. A state you cannot reach is not a state.
Let absence be the strong claim. A record saying something was read is the reading system's statement about itself. It establishes very little. The absence of such a record, when every read path emits reliably, is much harder to fake. Sell the negative; label the positive honestly.
Write refusals, not warnings. Every fix above replaced a warning with a non-zero exit. A warning is a claim that the caller will probably ignore. An exit code is one they have to handle.
Why the tests didn't catch any of it
This is the part I'd most want someone else to learn from.
Every layer was green independently. MCP connected. Tools available. Config exists. Database healthy. Tests passing.
The tests asserted our assumptions. They declared the expected configuration — including the stale launcher format — and so preserved the defect as a specification. They checked that a config entry existed, not that the process it named could start. They verified our tool list, not that the instructions were delivered. They ran inside workspaces that happened to contain a file whose absence was the actual failure mode.
The one test nobody had written:
fresh client → ordinary prompt → session appears in the database → both messages recorded
Everything else was infrastructure confirming its own existence.
The limit we had to publish
One more, because leaving it out would make this essay the thing it's about.
Our live capture depends on a model choosing to call a tool. We can deliver the instruction reliably. We cannot make a client obey it. During testing, one client hit a budget limit and stopped before logging; one refused the call under its default permission policy; one called the wrong tool entirely.
So the release ships saying so: capture is best-effort, not mandatory, and the diagnostics say configured where they used to say logging works. Enforcement needs an execution path we control — a wrapper that opens the audit session before the client starts and holds the output until the write lands — and we haven't built it.
Shipping a tool that declines to flatter its own integration is uncomfortable. It is less uncomfortable than the alternative, which is a green check mark that means nothing, in a product sold on its records being trustworthy.
If you build anything that reports on its own health, it's worth an afternoon with one question: for each green mark in your output, what did you actually observe, and when?
I expect you'll find at least one that answers "I observed that a thing exists, and I'm telling you it works."

Top comments (1)
Every one of these is the same shape: presence read as proof. It's the failure mode I keep finding too.
The one that got me: a tool that derived an index key one way on write and another way on read. Every unit test passed because tests wrote and read through the same helper, so the fork never surfaced; it only showed up when a real caller handed it a native Windows path. Green suite, wrong records.
What catches these for me isn't a better check, it's asserting the negative: write down what should be FALSE if the thing never worked, and assert that. "Archive was signed" becomes "there is at least one signature whose bytes differ from an empty object and validate against the key." Your
{}case is exactly the tell: readability tested, content assumed.The hard part is that whoever writes the check shares the blind spot of whoever wrote the bug. Different hands on a different day is worth more than a smarter assertion.