DEV Community

Cover image for Linux Doctor, heal thyself: 20 ways my Linux health checker was wrong
7sh1d0w7x
7sh1d0w7x

Posted on

Linux Doctor, heal thyself: 20 ways my Linux health checker was wrong

A diagnostic tool has exactly one job: tell the truth about the machine in
front of it. So when I built Linux Doctor — a read-only checker that
reports what is wrong and prints the fix without ever running it — I did the
thing a doctor dreads. I audited it against reality: five distro families,
minimal container images, and machines that had been running for years.

It lied about twenty times. None of them were exotic. Every one is a class of
bug you can hit in any script that reads the output of system tools.

Here they are, grouped by what they teach.


1. A pipeline exits with the status of its last command

df -P /boot | tail -n 1     # a failed df looks successful
journalctl -p err | grep MCE # grep exits 1 when nothing matches
Enter fullscreen mode Exit fullscreen mode

The boot check read /boot's usage through tail. If df failed, the exit
status still belonged to tail, which succeeded — so a failed probe looked
like a good one.

Worse was hardware. It used the exit status of journalctl … | grep … as its
readability gate. But grep exits 1 when there is nothing to match —
so "no hardware errors" (good news) was read as "I could not read the log", and
the check stayed silent on a healthy machine instead of saying "No hardware
errors logged". A benign EDAC banner made grep exit 0 on my own box, which
hid the bug for weeks.

Lesson: never gate readability on a pipeline's status. Take the raw output
and decide in code.

2. "Nothing found" and "the command is missing" are not the same

  • Minimal Debian, Fedora and Ubuntu images do not ship ip (iproute2). The probe's failure was read as an empty route table — a medium "No default network route" on a machine with a perfectly good route.
  • Minimal openSUSE has no awk. zypper … | awk produced nothing, so Tumbleweed reported "System is up to date" with thirteen lines of updates on screen.
  • The engine's own run() only flagged a command as missing when the shell itself was absent — which never happens. A missing tool exits 127 through the shell. So every check that gated a "could not check" skip on missing stayed silent, and a minimal image scored as if those checks had passed.

Lesson: distinguish "the tool answered" from "the tool isn't there."
Exit 127 is the tell.

3. Success messages contain the words you search for

  • pacman -Dk prints "No database errors have been found!" — a bare error match flagged a clean Arch system as having broken packages.
  • The EDAC driver's startup line is "EDAC ie31200: No ECC support" — the check matched it as an ECC problem. It was the top item in the report and cost 9 points.
  • mce: CPU supports N MCE banks is printed once per CPU at boot on every Intel machine. A bare mce match read it as a machine-check exception.

Lesson: a keyword match is not a diagnostic. Require an actual event, and
explicitly reject the routine preamble lines.

4. Kernel interfaces are not namespaced

Inside a 256 MB container, free -b reported the host's 15 GB.
/proc/loadavg is the host's load average while nproc reports the
container's CPU count — and the ratio between them invented an overloaded
system out of an idle container. (All five test images reported it.) lsblk
and /proc/swaps are not namespaced either, so the container was told about
the host's disks and swap.

Lesson: detect the container and say why you are skipping. That is a
limitation of the environment, not a verdict about the host.

5. Sometimes the bug is you

A fuser-based lock probe ran next to Linux Doctor's own apt-get check. So
it found the tool itself holding the dpkg lock — and told the user to wait for,
or kill, a process that was the tool. It only happened when run as root, so
almost nobody would ever have seen it.

Lesson: when you inspect a global resource, exclude your own process group.

The wrong direction of wrong

The dangerous bug is not a false alarm. It is a false all-clear.

  • A fresh image that never ran apt update answers "0 upgraded" with exit 0 — which read as "up to date". That is not being up to date. The check now says it cannot tell and points at apt update.
  • apk info -u is not a valid command at all (it exits 1 with unrecognized option 'u'), so Alpine said nothing.
  • Void had no update branch, so a machine with 54 pending updates was skipped and scored as current.
  • flatpak remote-ls --updates prints a column table whose fields contain no /, which is exactly what the count looked for → "apps are up to date" with updates pending.

Lesson: prefer "unknown" to a confident lie. A check that says "I could
not determine this, and here is the package that would let me" is worth more
than one that silently scores a broken machine as healthy.

And sometimes it answered a different question than the one asked

  • processes warned whenever one app used more than 20% of total RAM, ignoring what was actually free — so a 15 GB box with 9.4 GB free got a medium warning about a browser.
  • The same check listed a browser once per process (its memory is spread across many processes sharing one binary) and reported the largest single process as "the app".
  • timers called dnf-makecache.timer a broken schedule. It is enabled on an immutable system and can never run — its start condition is unmet by design, and systemctl status says so.
  • The KDE lock screen logs Authentication attempt too soon every time you retype a wrong password quickly. A healthy desktop got "6 recognized errors" and lost 8 points. The same string from sshd is worth seeing — only the screen locker's copy is noise.

Lesson: "the number is real" is not the same as "the number means what you
think it means."

How these are caught now

  • A clean-image gate runs the engine inside Fedora, Debian, Ubuntu, Alpine and Arch containers, and fails when an unexplained high or medium finding appears.
  • Recorded fixtures from real machines replay through the same pipeline, and every high/medium finding they produce needs a written reason.
  • A severity rubric and a finding-code registry so severity cannot drift silently.
  • No fix lands without a regression test that fails first. A wrong result is reproduced before it is changed.
  • The whole list is public, in both directions: docs/limitations.md names every false positive and false negative it has shipped, with the test that guards each one.

Why tell on yourself

The entire value of a diagnostic is trust — and trust is built by being
explicit about where you might be wrong, not by never being wrong. A tool that
admits "I could not check this; here is the package that fixes it" is far
more useful than one that quietly scores a minimal image as perfect.

Linux Doctor is read-only by construction: it prints the fix and never runs it.
One engine sits behind a CLI, a web dashboard and a desktop app — and the
honesty document ships with it.

→ https://github.com/7sh1d0w7x/linux-doctor

If you run it and it gets something wrong, that is the most useful bug report
the project can get — and it comes with a replayable fixture.

Top comments (0)