A diagnostic tool has exactly one job: tell the truth about the machine in
front of it. So when I built Linux Doctor — a read-only checker that
reports what is wrong and prints the fix without ever running it — I did the
thing a doctor dreads. I audited it against reality: five distro families,
minimal container images, and machines that had been running for years.
It lied about twenty times. None of them were exotic. Every one is a class of
bug you can hit in any script that reads the output of system tools.
Here they are, grouped by what they teach.
1. A pipeline exits with the status of its last command
df -P /boot | tail -n 1 # a failed df looks successful
journalctl -p err | grep MCE # grep exits 1 when nothing matches
The boot check read /boot's usage through tail. If df failed, the exit
status still belonged to tail, which succeeded — so a failed probe looked
like a good one.
Worse was hardware. It used the exit status of journalctl … | grep … as its
readability gate. But grep exits 1 when there is nothing to match —
so "no hardware errors" (good news) was read as "I could not read the log", and
the check stayed silent on a healthy machine instead of saying "No hardware
errors logged". A benign EDAC banner made grep exit 0 on my own box, which
hid the bug for weeks.
Lesson: never gate readability on a pipeline's status. Take the raw output
and decide in code.
2. "Nothing found" and "the command is missing" are not the same
- Minimal Debian, Fedora and Ubuntu images do not ship
ip(iproute2). The probe's failure was read as an empty route table — a medium "No default network route" on a machine with a perfectly good route. - Minimal openSUSE has no
awk.zypper … | awkproduced nothing, so Tumbleweed reported "System is up to date" with thirteen lines of updates on screen. - The engine's own
run()only flagged a command as missing when the shell itself was absent — which never happens. A missing tool exits127through the shell. So every check that gated a "could not check" skip onmissingstayed silent, and a minimal image scored as if those checks had passed.
Lesson: distinguish "the tool answered" from "the tool isn't there."
Exit 127 is the tell.
3. Success messages contain the words you search for
-
pacman -Dkprints "No database errors have been found!" — a bareerrormatch flagged a clean Arch system as having broken packages. - The EDAC driver's startup line is "EDAC ie31200: No ECC support" — the check matched it as an ECC problem. It was the top item in the report and cost 9 points.
-
mce: CPU supports N MCE banksis printed once per CPU at boot on every Intel machine. A baremcematch read it as a machine-check exception.
Lesson: a keyword match is not a diagnostic. Require an actual event, and
explicitly reject the routine preamble lines.
4. Kernel interfaces are not namespaced
Inside a 256 MB container, free -b reported the host's 15 GB.
/proc/loadavg is the host's load average while nproc reports the
container's CPU count — and the ratio between them invented an overloaded
system out of an idle container. (All five test images reported it.) lsblk
and /proc/swaps are not namespaced either, so the container was told about
the host's disks and swap.
Lesson: detect the container and say why you are skipping. That is a
limitation of the environment, not a verdict about the host.
5. Sometimes the bug is you
A fuser-based lock probe ran next to Linux Doctor's own apt-get check. So
it found the tool itself holding the dpkg lock — and told the user to wait for,
or kill, a process that was the tool. It only happened when run as root, so
almost nobody would ever have seen it.
Lesson: when you inspect a global resource, exclude your own process group.
The wrong direction of wrong
The dangerous bug is not a false alarm. It is a false all-clear.
- A fresh image that never ran
apt updateanswers "0 upgraded" with exit 0 — which read as "up to date". That is not being up to date. The check now says it cannot tell and points atapt update. -
apk info -uis not a valid command at all (it exits1withunrecognized option 'u'), so Alpine said nothing. - Void had no update branch, so a machine with 54 pending updates was skipped and scored as current.
-
flatpak remote-ls --updatesprints a column table whose fields contain no/, which is exactly what the count looked for → "apps are up to date" with updates pending.
Lesson: prefer "unknown" to a confident lie. A check that says "I could
not determine this, and here is the package that would let me" is worth more
than one that silently scores a broken machine as healthy.
And sometimes it answered a different question than the one asked
-
processeswarned whenever one app used more than 20% of total RAM, ignoring what was actually free — so a 15 GB box with 9.4 GB free got a medium warning about a browser. - The same check listed a browser once per process (its memory is spread across many processes sharing one binary) and reported the largest single process as "the app".
-
timerscalleddnf-makecache.timera broken schedule. It is enabled on an immutable system and can never run — its start condition is unmet by design, andsystemctl statussays so. - The KDE lock screen logs
Authentication attempt too soonevery time you retype a wrong password quickly. A healthy desktop got "6 recognized errors" and lost 8 points. The same string fromsshdis worth seeing — only the screen locker's copy is noise.
Lesson: "the number is real" is not the same as "the number means what you
think it means."
How these are caught now
- A clean-image gate runs the engine inside Fedora, Debian, Ubuntu, Alpine and Arch containers, and fails when an unexplained high or medium finding appears.
- Recorded fixtures from real machines replay through the same pipeline, and every high/medium finding they produce needs a written reason.
- A severity rubric and a finding-code registry so severity cannot drift silently.
- No fix lands without a regression test that fails first. A wrong result is reproduced before it is changed.
- The whole list is public, in both directions:
docs/limitations.mdnames every false positive and false negative it has shipped, with the test that guards each one.
Why tell on yourself
The entire value of a diagnostic is trust — and trust is built by being
explicit about where you might be wrong, not by never being wrong. A tool that
admits "I could not check this; here is the package that fixes it" is far
more useful than one that quietly scores a minimal image as perfect.
Linux Doctor is read-only by construction: it prints the fix and never runs it.
One engine sits behind a CLI, a web dashboard and a desktop app — and the
honesty document ships with it.
→ https://github.com/7sh1d0w7x/linux-doctor
If you run it and it gets something wrong, that is the most useful bug report
the project can get — and it comes with a replayable fixture.
Top comments (0)