DEV Community

Geek Consulting
Geek Consulting

Posted on Originally published at geekconsulting.au on

The Someday Project, Part II: How an AI Agent Built the Homelab Dashboard I’d Given Up On

My media stack worked fine. Two Sonarrs, a Radarr, two Prowlarrs, qBittorrent, Plex, Jellyfin, Seerr (my request manager) — all humming along on three machines. The problem was that “working fine” lived in a dozen browser tabs and a mental map only I could read. Homarr — the self-hosted homelab dashboard — promised to collapse all of that into one pane of glass. I’d tried. Every time I’d wire up two integrations, hit the fiddly bits — an API key here, a reverse-proxy hostname there, a widget that refused to authenticate — lose the thread, and the tab would gather dust until I forgot the admin password.

The payoff never justified the faffing about. It’s the textbook someday project: obviously nice to have, never quite worth the afternoon.

What changed this time isn’t Homarr. It’s that I could hand the faffing about to an agent.

This is the follow-up to last week’s write-up, where Claude Code turned my Home Assistant instance into a proper Grafana data platform. Same collaborator, same house, same principle: I supply direction, the agent supplies execution. This time the target was the media stack and the monitoring layer that had been on my “someday” list for a year.

Three goals

  1. One control plane. Every *arr app, both media servers, downloads, requests, and Home Assistant behind a single dashboard I’d actually open — with the ability to start/stop containers from it.
  2. Real metrics, not green dots. Throughput, queue depth, container CPU/RAM, and internet speed flowing into the same VictoriaMetrics + Grafana I stood up yesterday, so the media stack lives next to the smart-home dashboards.
  3. Do it without me doing it. My job was to decide what and approve the shape. Everything tedious — the keys, the widgets, the Docker plumbing — was the agent’s.

The stack

Homarr for the control plane. I picked it over Homepage, Dashy and friends for one reason: it has the best machine interface of any dashboard I’ve seen — a full OpenAPI + tRPC API, and its own built-in MCP server. That last part matters more than it sounds. Once Homarr was running, Claude didn’t have to click a UI or hand-edit YAML; it could talk to the dashboard through the dashboard’s own agent interface — list boards, create integrations, add widgets, all as first-class tool calls. The dashboard configured itself, supervised.

For metrics I reused yesterday’s platform rather than building a new one:

  • Scraparr  — one container that scrapes every *arr instance (plus Kavita and Seerr) and exposes Prometheus metrics.
  • cAdvisor + node-exporter  — Docker container and host metrics.
  • vmagent → VictoriaMetrics → Grafana  — the same pipeline the smart-home data already flows through. Ten-year retention, LAN-only, behind Caddy.

Homarr itself sits behind Caddy on my Linux Docker host, reaching Docker through a socket-proxy  — a small container that exposes a scoped, read-mostly slice of the Docker socket (start/stop/restart, but no arbitrary exec) so the dashboard can manage containers without being handed root over the whole host. Direction from me: “front the socket, don’t bind-mount it raw.” Execution from the agent.

The build

Thirteen integrations, and a recurring lie

The first session was integrations: Sonarr (anime + TV), Radarr, both Prowlarrs, qBittorrent, Seerr, then Plex, Jellyfin, Home Assistant, Glances, Uptime Kuma, and Speedtest Tracker. Thirteen in total.

Almost immediately we hit a quirk worth its own lesson. Homarr’s MCP integration_create tool would return a long, alarming schema-validation error — a wall of red about invalid_union and “expected string, received undefined.” Every single time. And every single time, the integration had actually been created. The error was in how the tool serialised its success response, not in the action itself.

A human reads that error and assumes failure. The agent did something better: it ignored the error text and asked the system what was true  — called integration_all and looked for the new integration in the list. There it was, connection-tested and saved. From then on, every “failed” create was verified by a read, not a retry.

_ Lesson: _ A tool’s error envelope is not the source of truth. The running system is. When a write “fails,” read the state back before you believe it.

The same discipline caught a real failure too. Uptime Kuma’s integration genuinely wouldn’t connect — a clean 404, not the cosmetic one. With no credentials it defaults to reading a status page at slug default, which didn’t exist. Rather than guess, the agent read Uptime Kuma’s own SQLite database on disk, found a published status page with the slug main and 24 monitors, and pointed the integration at that. No trial and error — it looked.

The board nobody wants: one big wall of widgets

By the time all thirteen integrations had widgets, the single board was a mess — calendar, indexers, downloads, requests, streams, Docker, weather, all fighting for space. Exactly the state that made me abandon Homarr the last three times.

So we restructured into five boards: an overview home board, then media , downloads-indexers , infrastructure , and — the fun one — a home board for Home Assistant.

Here’s where the API depth paid off. Homarr’s MCP can add items to a board but not remove them, and it can’t create the collapsible “category” sections at all. The agent went and read Homarr’s source on GitHub, found the saveBoard tRPC route and its exact Zod schema, and reconstructed the whole board — sections, items, per-widget grid layouts — as a single validated payload. It didn’t guess the shape; it fetched the schema and matched it.

And it did the cautious thing without being told: instead of destructively rewriting my live board, it built the four new boards fresh and left the old one untouched as a backup for me to delete once I was happy. Judgment I’d expect from a careful colleague, not a script.

129 smart-home tiles, generated from the house itself

The home board is my favourite artifact of the whole project. Home Assistant knows my house as areas  — Bar, Deck, Garage, Kitchen, Living Room, Master Bedroom, Study, and so on. I wanted one Homarr category per area, each filled with toggle tiles for that room’s lights, switches and fans.

Doing that by hand across 1,261 entities is precisely the work that kills a someday project. The agent did it by asking Home Assistant to describe itself: it hit HA’s REST template API with a Jinja snippet using areas() and area_entities(), got back every area and its controllable entities as JSON, and then generated the board — 14 categories, ~129 tiles — via that same saveBoard route. Lights, switches and fans became clickable toggles; media players became read-only status tiles. Five tiles a row, laid out programmatically.

Then it read the board back and counted the tiles per category to confirm the whole thing landed. Verification as the last step, again.

Moving devices Home Assistant wouldn’t let me move from the UI

One area was a near-duplicate — my study existed twice under two slightly different names, a legacy split, with three stray entities stranded on the wrong copy. I asked the agent to merge them. It found the three (a Samsung Odyssey monitor and an air purifier, both from SmartThings), realised HA’s area assignment is only editable over the WebSocket API, not REST — and, finding no WebSocket library installed on the box, wrote a minimal one from the Python standard library: raw socket, the upgrade handshake, frame masking, just enough to authenticate and call config/device_registry/update. Moved the devices, verified the source area was now empty, rebuilt that board section.

_ Lesson: _ When the tool you need isn’t installed, sometimes the fastest path is to implement the ten lines of protocol you actually use. An agent doesn’t get bored writing a throwaway WebSocket client.

Where it got real: two migrations mid-project

The abandoned image

I’d started a Speedtest Tracker container a while back and asked the agent to just “add it to Homarr.” It checked, and pushed back: the image was henrywhitaker3/speedtest-tracker — the original author’s build, now archived — installed through CasaOS, which I’d stopped using when I moved to ZimaOS. That’s why I couldn’t find the API-token screen: it wasn’t in that old UI.

I told it to bin the lot. It listed every container first, deleted exactly the three I named — the dead Speedtest image and two orphaned CasaOS containers — and pointedly left the unrelated openspeedtest and other stale containers alone. Then it deployed the maintained image, lscr.io/linuxserver/speedtest-tracker (v1.14), as a clean stack.

Getting metrics out of it was a chain of small, verified problems:

  • The production image ships no Tinker and no speedtest:run command , so there was no obvious way to mint an API token or trigger a test. The agent bootstrapped the Laravel framework in a throwaway PHP script to mint a Sanctum token directly, and dispatched a test through the app’s own action class.
  • Homarr’s connection test hits /api/v1/results/latest, which 404s until at least one test has completed  — so the integration refused to save until a real speed test had run. The agent curled the endpoint, read the 404, understood why, ran a test, and tried again. It never assumed the token worked; it checked.
  • Speedtest Tracker has a built-in Prometheus exporter at /prometheus, but it’s gated behind a database setting and returns “no data available” until a test completes after the exporter is enabled (it keys off a cache entry). Enable, run one more test, then metrics flow.

The one-character trap

Wiring vmagent to scrape the new exporter produced a genuinely instructive bug. I edited the scrape config with sed -i, and the target stubbornly stayed “down” — reading the old config. The config file is bind-mounted into the container as a single file, and sed -i doesn’t edit in place; it writes a new file and renames it, swapping the inode. The container’s mount still pointed at the original inode. The fix was to rewrite the file with cat > (which truncates in place, same inode) and restart the agent.

_ Lesson: _ Never sed -i a bind-mounted config file. The rename swaps the inode and the container keeps reading the file you thought you replaced.

After that, one more networking wrinkle — vmagent scrapes its neighbours by container name, and the new container was on its own Docker network — so the agent attached it to the monitoring network and scraped it by name. Then, of course, it queried VictoriaMetrics for the exact metric names to confirm the 22 speedtest_tracker_* series were landing before declaring victory.

A few smaller traps, for the record

  • Tautulli (Plex history and stats) came up pre-seeded to skip its setup wizard — but the LinuxServer image appends its own default config on first boot, and autodiscovery quietly reset the Plex address to 127.0.0.1. The fix was to pin pms_url_manual and dedupe the config, killing the container rather than stopping it gracefully so it wouldn’t save the broken state on the way out.
  • Timezones in Speedtest Tracker are set by an app-wide DISPLAY_TIMEZONE env var, not a per-user preference — a five-minute rabbit hole that ended in one line of compose.
  • The host has broken IPv6 , which made Homarr’s weather widget hang on a location lookup. We worked around it with hard-coded coordinates and logged the host issue as deferred — it’ll likely evaporate in the Proxmox migration, and chasing it now wasn’t worth it.

What it looks like now

  • 13 integrations in Homarr, each connection-tested.
  • Five boards : an overview home page, media, downloads/indexers, infrastructure, and a Home Assistant control board with a category per room and ~129 live tiles.
  • Container management from the dashboard via the socket-proxy.
  • Grafana dashboards for the *arr apps (Scraparr), Docker/host (cAdvisor), and internet speed (Speedtest Tracker) — feeding the same VictoriaMetrics as the smart-home data.
  • Everything documented  — a repo, a knowledge-base project, and the agent’s own persistent memory of the gotchas, so next session starts warm.

If you want to replicate it

  1. Pick a dashboard with a real API. Homarr’s OpenAPI + tRPC + MCP is what made delegation possible; a YAML-config dashboard would have just moved the tedium.
  2. Reuse your metrics backend. If you already have Prometheus/VictoriaMetrics + Grafana, don’t stand up a second one — point Scraparr and the exporters at what you have.
  3. Front the Docker socket with a proxy. Never hand a dashboard the raw socket.
  4. Verify every write with a read. Integrations, tokens, metrics, board layouts — confirm against the running system, don’t trust the success message.
  5. Prefer maintained images, and check who maintains them before you build on top.
  6. Default dashboards to LAN-only and keep credentials out of your logs.
  7. Log the things you’re deliberately not fixing (like the host IPv6 quirk) so “done” is honest.

The honest part

I want to be clear about what actually unblocked this, because it isn’t a tool recommendation. I already knew what I wanted — I’ve known for a year. Homarr didn’t get easier. VictoriaMetrics didn’t get easier. What changed is that the work between “I want this” and “this exists” — the API keys, the schema-matching, the inode trap, the WebSocket client, the six different ways a token can be almost-but-not-quite valid — could be delegated to something that doesn’t find that work tedious and doesn’t lose the thread when the fourth rabbit hole opens up.

And the thing it did best wasn’t writing code. It was refusing to guess. Over and over, when a component misbehaved, it didn’t theorise — it asked the running system. It read the 404 instead of assuming the token was fine. It queried the database for the metric name instead of hoping. It read the board back to count the tiles. It checked which image was actually deprecated before deleting anything. That’s the difference between a script and a collaborator: a script assumes; this verified.

Not direction. Execution — with judgment.

The dashboard I’d abandoned three times took an afternoon of my attention and a lot of the agent’s. Now every arrow in my media pipeline is on one screen, every container is a click from a restart, and the whole thing trends in Grafana alongside the smart-home dashboards from Part I.

Build the boring glue once — or rather, get it built. Then go actually use the thing you kept meaning to finish.

Top comments (0)