My homelab dashboard has a tile for each service. Most of those services sit at fixed addresses. A few don't: my 3D printer gets its address from DHCP, and when it moves, the tile points at nothing.
So on 12 July I wrote a small job to keep the tiles right. It runs every 15 minutes, works out the current address of each tracked device, and rewrites the dashboard's config only if something changed. It tracks three devices: the printer, Home Assistant and Uptime Kuma.
For three and a half weeks it did nothing at all, and its log said so every quarter of an hour in a way I read as good news.
Built to be safe
I made two design choices on purpose, and both come back into this story.
First, it is a plain Python script with no language model anywhere in it. An agent that "works out" an IP address can invent one, and a wrong address written into a dashboard is worse than a stale one.
Second, it fails safe. If it can't find a device, it leaves that device's tile alone. It never writes a guess. A device that is switched off shouldn't get its tile blanked, and the next run can pick it up again.
That second choice is the one that hid the problem.
How it found devices
My router doesn't serve forward DNS for local .lan names, so asking "what is the printer's address?" by name never worked. Reverse lookups did: DHCP leases were mirrored into the resolver, so asking "what name belongs to this address?" gave an answer.
That resolver rate-limits reverse lookups hard. Sweeping all 254 addresses at 64 workers turned up about 13 names. Even 8 workers was flaky. So the job had two paths:
- Confirm. Do a reverse lookup on each device's last-known address, three queries in all, and accept it if the name still matches.
- Sweep. Only for a device that has moved: ping the whole subnet twice, then do reverse lookups on live hosts only, four at a time, and stop as soon as it's found.
It worked. The log from the first days shows no changes (3 resolved, 0 miss).
The log line I kept reading as fine
On 16 August I actually read the log. Run after run looked like this:
MISS: Voron 3D Printer (...) not on <subnet> — leaving tile untouched
MISS: Home Assistant (...) not on <subnet> — leaving tile untouched
MISS: Uptime Kuma (...) not on <subnet> — leaving tile untouched
no changes (0 resolved, 3 miss)
"No changes" was true. It was also useless. Since the evening of 23 July, almost every run had found none of the three devices. It couldn't see anything, so it changed nothing, so it reported nothing to change. Nothing alerted, because the only thing it ever notified on was a change.
The log has 4,422 lines reading no changes (0 resolved, 3 miss). That isn't 4,422 runs, though. It's about half that, because of a second bug I found in the same pass: every line was written twice. The script printed to stderr and appended to its log file, and cron redirected stderr into that same file. Even the count of the failure was wrong.
Why it went blind
Reverse DNS had stopped working entirely. .lan names came back NXDOMAIN from the resolver the job uses, and also when I asked the router directly. Reverse lookups of addresses I knew were good came back empty. The router was otherwise healthy: it answered pings with no loss and resolved normal forward DNS. Bare hostnames still resolved, but only to IPv6 addresses on my VPN overlay, never the LAN IPv4 the tiles needed.
I never found out why the router stopped answering reverse lookups. I stopped depending on it instead.
The fix: ask something that knows
The rule I took from it: don't infer an address from DNS when an authoritative source exists. Two such sources were sitting there already.
Containers ask Proxmox. Two of the three devices are LXC containers, and Proxmox knows their addresses. The job now reads them from the nodes over SSH. Only running containers with an address on the right subnet count. One container name exists on both nodes. That is logged as ambiguous and refused, not guessed.
Physical boxes are found by MAC. The printer runs on a Raspberry Pi. Its MAC address doesn't change when its IP does, which is exactly the drift this job exists for. It checks the ARP table first. If the MAC isn't there, it pings the last-known address to refresh the table. Only then does it do a full ping-sweep.
Reverse DNS is still there, but only as the last fallback.
The first run after the change: 3 resolved, 0 miss. Then the embarrassing part. I checked the live dashboard config, and all three tiles had been correct the whole time. None of the devices had moved. The sync had been blind, not wrong, so the dashboard never showed a symptom.
Two more fixes in the same pass
The double-logging: the script now writes to the file and only echoes to stderr when it's attached to a terminal. It also echoes if the file write fails, so output can't disappear silently.
And timing: 38.0 seconds of a 38.1-second run were spent asking Proxmox for container addresses. That was two SSH calls, each running a helper that called pct three times per container. It came to around 3,600 pct invocations a day for addresses that almost never change. They're cached for an hour now. But the hour isn't the main guard: if a cached address stops answering, the job refreshes from Proxmox straight away. I tested that by poisoning the cache with a dead address. It noticed, refreshed and recovered the right one. The run dropped from 38 seconds to 0.34.
Did it earn its keep?
Since then the printer has moved twice. On 7 September and again on 10 September it came back on a new DHCP address. Both times the job found it by MAC and rewrote the tile, backing up the config first. That is the job it was built for, done twice, after three and a half weeks of finding nothing while saying "no changes".
"No change" and "couldn't look" are different answers
The log still has a problem, and I'd rather say so. Most runs today end no changes (2 resolved, 1 miss). The miss is the printer: its MAC isn't on the network much of the time. The script does keep misses in a separate list in its state file. But it still notifies only on a change, so a device it can't find is still a log line nobody reads.
What I'd do next, and what I'd suggest for any job shaped like this:
- Return three outcomes, not two. "Checked, unchanged", "checked, changed" and "couldn't check" are different results. Don't let the third collapse into the first.
- Alert on consecutive misses, not on each one. One miss is a printer that's switched off. Every device missing on every run for a day is a broken job.
-
Put the resolved count where you'll see it.
0 resolvedwas in every line. I just never looked.
This wasn't a one-off, either. This month my blog's publish queue ran empty after 18 September. Both publishers logged "queue empty — nothing to publish" and exited 0, and nothing alerted. A healthy, idle publisher looks just like a healthy, busy one unless you check how deep the queue is. There is now a check that alerts when the queue is empty.
And on 7 September, a canary that watches my paid model provider switched four agent profiles back to local models after its probe failed on every model. It turned out to be the probe's input, not the models. It only knew how to switch back after a credit problem, not after a failover, so it logged "nothing to watch" for 17 days. It now re-probes after a failover and switches back on its own.
All three were healthy processes reporting a true, reassuring sentence about a state they hadn't actually checked.
🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.
Top comments (0)