DEV Community

Cover image for Why DKMS says "Manual intervention is required!" (and how I fixed it upstream)
hopsayer
hopsayer

Posted on

Why DKMS says "Manual intervention is required!" (and how I fixed it upstream)

Or: how one line in someone else's Bash script saves a million human-hours.


Ever found an annoying problem in software you use daily, but couldn't do anything about it? That's usually where the story ends. This time it didn't.

How I met the message

I was updating my kernel on EndeavourOS (Arch-based). The usual routine: pacman -Syu, reboot, everything works. But after one update, this appeared in my terminal:

find: '/usr/lib/modules/6.12.3-arch1-1/': No such file or directory
nvidia/565.57.01: broken
Error! nvidia/565.57.01: Missing the module source directory or the symbolic link pointing to it.
Manual intervention is required!
Enter fullscreen mode Exit fullscreen mode

I read it. I understood it was about the proprietary nvidia driver. I understood manual intervention was required. I understood there was no module source directory or symlink. What I did not understand was where I was supposed to intervene.

So I googled. I found dozens of threads on the EndeavourOS, Manjaro, and Arch forums. People in the same situation: Manual intervention is required! — and then everyone guessing. Some wipe /var/lib/dkms/ entirely. Some reinstall nvidia-dkms. Some nuke everything and start over.

I eventually fixed my own system. But it bothered me that this is how it works. One error message, printed once, generates a constant stream of forum threads — that's a discoverability problem. A single error message, changed once, gets printed in the terminals of millions of users. A bad message means a million searches. A good one means a million saved hours. It's a quality-of-life improvement, and it reduces noise for everyone. Answer without asking.

And it's an easy fix. This is where the practical advantage of open source kicks in: users can actually influence the software instead of waiting helplessly for updates.

Why DKMS matters

DKMS (Dynamic Kernel Module Support) is fundamental infrastructure. It lets drivers (NVIDIA, VirtualBox, ZFS, WireGuard) rebuild automatically when the kernel updates, saving users from manually rebuilding modules and wiring them back into the boot process. Along with akmod, it's the kind of project that runs (or can run) on every major distro — Arch, Fedora, Debian — worldwide. Millions of systems. And I saw a chance to improve it.

I checked the project's activity, contribution policy, and the odds my patch would land. The project lives on GitHub under dkms-project, receives patches from Dell maintainers and the community, but it's not React or Kubernetes in terms of activity. Patches get accepted slowly, but they get accepted. I decided it was worth a try.

What was actually wrong

I searched for occurrences of Manual intervention is required! with rg across the cloned repo.

Turns out DKMS isn't missing the path — it knows it. It checks it in is_module_broken():

is_module_broken() {
    [[ $1 && $2 ]] || return 1
    [[ -d $dkms_tree/$1/$2 ]] || return 2
    [[ -L $dkms_tree/$1/$2/source && ! -d $dkms_tree/$1/$2/source ]] && return
    [[ ! -L $dkms_tree/$1/$2/source && -d $source_tree/$1-$2/ ]] && return
}
Enter fullscreen mode Exit fullscreen mode

The problem is that the path never makes it into the error message. Four places in the code (module_is_broken_and_die, do_status, run_match, autoinstall) print the same sentence:

Missing the source directory or the symbolic link pointing to it.
Manual intervention is required!
Enter fullscreen mode Exit fullscreen mode

The user reads it, then googles. What they could have seen instead:

Missing the source directory or the symbolic link pointing to it:
/var/lib/dkms/nvidia/565.57.01/source
Manual intervention is required!
Enter fullscreen mode Exit fullscreen mode

And immediately understood: here's the empty spot. Go look at it, figure out what happened.

"But find: already shows the path!"

That's what I thought too, when I went back to my log:

find: '/usr/lib/modules/6.12.3-arch1-1/': No such file or directory
nvidia/565.57.01: broken
Error! nvidia/565.57.01: Missing the module source directory or the symbolic link pointing to it.
Manual intervention is required!
Enter fullscreen mode Exit fullscreen mode

The path is right there. Why do we need another one?

Because these are two different paths in two different places.

  • find: '/usr/lib/modules/6.12.3-arch1-1/' is the path to the kernel directory. find scans it while cleaning up old modules and complains it's gone. It's side noise from the find utility, not a DKMS message.
  • /var/lib/dkms/nvidia/565.57.01/source is the path to the module source. This is what is_module_broken() actually checks, and this is what the error message doesn't print.

These paths don't overlap. /usr/lib/modules/6.12.3-arch1-1/ is about the kernel. /var/lib/dkms/... is about the module. If your module is broken, cleaning /usr/lib/modules/6.12.3-arch1-1/ is pointless — it's already gone, and it has nothing to do with the module.

Then there's context. find: appears once, buried in a wall of pacman output, unhighlighted, easy to miss. Manual intervention is required! appears everywhere DKMS encounters a broken module: in dkms status, dkms build, dkms install, dkms autoinstall. The user sees that one and googles.

Check it against a real scenario. After I saw find: '/usr/lib/modules/6.12.3-arch1-1/', I:

  1. Googled "dkms broken manual intervention".
  2. Tried dkms remove nvidia/565.57.01 --all — it didn't work.
  3. Saw nvidia/565.57.01: broken in dkms status.
  4. Couldn't tell whether to clean /var/lib/dkms/nvidia/565.57.01/ or /usr/lib/modules/6.12.3-arch1-1/.
  5. Found the answer via an EndeavourOS forum thread from 2023, where someone explained to look in /var/lib/dkms.

If the message had simply said:

Missing the source directory or the symbolic link pointing to it:
/var/lib/dkms/nvidia/565.57.01/source
Enter fullscreen mode Exit fullscreen mode

I would have understood in 10 seconds. No googling, no forums, no guessing.

Imagine your car's Check Engine light comes on. And someone has taped a sticky note under the hood that says "front-left wheel is loose." Technically the information exists. But it doesn't answer the question that triggered the light, and you have to find it. My patch is what happens when, instead of "Check Engine," the dashboard says "Check tire pressure on front-right."

That's why the improvement mattered. Not because the path was missing, but because it was in the wrong place, and about the wrong thing.

Having confirmed the problem was real and worth fixing, I moved to the next step: checking whether anyone had done it already.

Searching for prior work

Before writing code, I dug through the project's issues and PRs. Found:

  • PR #357 — added the broken status. But it did not add the path to the message.
  • Issue #94 — about improving diagnostics, but for a different scenario and a different stage.
  • Issue #463 — about return codes, tagged help wanted.

No open PR did exactly what I wanted. The niche was free. Time to write.

What I did

Three independent, layered commits:

  1. Added the path to all four call sites that print the message. Also updated 10 test blocks in run_test.sh.
  2. Added endeavouros to the distro case in the test script. One line.
  3. Added a two-line hint about what to do:
If this module version is no longer needed, you can remove the stale directory.
Otherwise, reinstall the package that provides its source.
Enter fullscreen mode Exit fullscreen mode

The third commit was the most debatable. Maintainers might say it's too verbose. So I:

  • kept it as a separate commit, so it can be reverted with one command,
  • left a note in the PR saying I'm not attached to it: "Happy to drop the hint and keep just the path if preferred",
  • deliberately avoided phrasing like "clear this folder", because that nudges users toward rm -rf in a system directory without understanding the consequences.

"Remove it if it's not needed, or reinstall if it is" is not an instruction — it's a fork in the road the user takes themselves. Google becomes meaningful again.

How I tested it

This took far more time and iterations than the actual patch. The DKMS test suite is run_test.sh, a Bash script that requires:

  • root,
  • kernel headers matching uname -r,
  • actually compiling modules through make,
  • a clean /var/lib/dkms (I had nvidia/580.178.04 sitting there — had to move it aside temporarily).

Along the way I hit:

  • unknown Linux distribution ID endeavouros — the tests didn't know my distro. Adding it to the case became commit #2. Tests pass with just that one line — the folder layout is identical to Arch.
  • /lib/modules vs /usr/lib/modules — on Arch-likes, /lib is a symlink to /usr/lib, and DKMS prints the canonical path. The tests expected /lib because they were written for Fedora/Debian. A temporary substitution in my local copy fixed it, but it wasn't needed in the PR.
  • Upstream build via make install uses MODDIR=/lib/modules, while the Arch package patches it to /usr/lib/modules. First test run failed on this, second passed.
  • Conflict with the system dkms — the test script calls /usr/bin/dkms, which was the system package. Temporarily removed it with a force flag — pacman -Rdd dkms.

The pipeline that ended up working:

sudo pacman -Rdd dkms
sudo mv /var/lib/dkms/nvidia /root/nvidia.dkms.bak
cd dkms
sudo make install
sudo ./run_test.sh
sudo make uninstall
sudo mv /root/nvidia.dkms.bak /var/lib/dkms/nvidia
sudo pacman -S dkms
Enter fullscreen mode Exit fullscreen mode

Final run:

*** All tests successful :)
Enter fullscreen mode Exit fullscreen mode

On EndeavourOS with kernel 7.2.7-zen1-1-zen. Not in CI, but on a live system where the nvidia driver is running in parallel.

The result

  • Issue #606 — describes the problem with a real example and links to the existing #357 and #94.
  • PR #607 — merged into main. Three commits, full diff, tests pass.
  • For users — instead of "Manual intervention is required!", they now see where to look and what to do.

From opening the issue to merge, it took less than a week. For DKMS, that's fast — PR #357, for comparison, sat for three months.

Honestly, I'd be satisfied just knowing I noticed the problem, didn't ignore it, proposed a solution, and did it politely. Opportunities like this feel like luck.

What I learned

1. "Bad UX" isn't only about graphical interfaces. CLI tools have it too. An error message is only useful if it's informative. If it doesn't help the user — even if it's technically correct — it's not worth much.

2. In open source, the justification matters as much as the code. A patch without context and research signals is as useful to a reviewer as an uninformative error is to a user. A patch with an issue, links to related discussions, Before/After, and test results is more respectful, more readable, and far more likely to be merged.

3. Atomic commits give flexibility to both sides. Maintainers get to silently accept only the commits they need. Time spent on coordination goes down. At a single commit, someone would have had to hand-edit the diff — either them or me. Layered PRs also read better, because they encode the author's intent in their structure. That's genuinely pleasant if you're the reviewer.

4. The environment eats more time than the code. The patch is 28 lines in dkms.in and 12 in the tests. The headers, symlinks, kernel versions, and distro detection took a day.

5. Small contributions to fundamental projects are worth it. DKMS is infrastructure. Contributing to it feels good just by itself. Fixing one line in an error message in a major project can matter more than 100 new features in a pet project.

If you want to make your first PR

The recipe that worked for me:

  1. Notice annoying behavior in software you use. Not a "typo fix" — a real UX rough edge or uninformativeness. Weigh the value: how many other people are doomed to hit it, and what does fixing it mean at the code and project level?
  2. Check that nobody has done it already. Search issues and PRs. Duplicated work helps neither you nor the reviewers. Gauge the scope of your change and aim for minimal.
  3. Describe the problem in an issue: where, when, what's inconvenient, actual vs expected behavior, why you think it matters. The issue is a document in itself: if you don't make the PR, you leave an idea for someone else to pick up, and that's also valued. If you do make the PR, it rests on your own issue, which is good documentation hygiene.
  4. Keep the patch small. One or two files, not ten.
  5. Split it into atomic commits. Even in a single PR. Yes, it takes longer, but it pays off.
  6. Run the tests on the final commit.
  7. Write the PR as a story. What was, what became, what problem it solves, whether tests pass, and note that the commits are layered.
  8. Be ready to roll back any part. "Happy to drop X if preferred" isn't weakness — it's strength.

PR #607 (merged): dkms-project/dkms#607
Issue #606: dkms-project/dkms#606


Top comments (0)