TL;DR
I keep a task board for everything the house and my side projects need doing, and the AI agent I work with keeps it up to date as we go. I read the board but I don’t work it, the agent does, and reading it meant switching to a browser, so the work and the record of the work were always in different windows.
Now the board sits in a panel beside the conversation. It shows what’s in progress, what’s stuck and what’s waiting. When the agent picks up a task, the panel opens that task by itself and updates while it works. There’s a button to start a fresh session on any task, in the right folder.
It took an evening and the morning after. Most of it works and I use it. Two parts have only been tested against a pretend version of the board, and I say which further down. It’s on GitHub, and getting it there taught me something about what “scrubbed” means.
If you’re here for the code, it’s at the bottom.
The board in the other window
The tracker is Vikunja, self-hosted, and it’s the one piece of the homelab that’s about the work rather than the house. Every job gets a card. The agent creates them, moves them to Doing when it starts, writes what it built into the description when it finishes. The Done column is the history of the whole lab.
I look at it. I don’t touch it. The agent does that.
Which is how I want it, but it put the board in the wrong place. It’s a browser tab, and the work happens in Claude Code. When the agent said “moved to Doing”, seeing that meant leaving the conversation and finding the card. I’m the reader of this board, not its editor, and the reading was happening in a different window from the thing I was reading about.
Claude Code recently grew a way to fix that: mods. A mod is a small plugin of hooks that runs inside the session and can draw its own panel next to the conversation. I asked for one in a single sentence: let’s create a mod to see our tasks when we’re working on them.
I’ve written about building things this way before: a data platform, a dashboard, a car alert, a spa that heats on spare sunshine. Same principle every time: I supply direction, the agent supplies execution. This one tested the first half of that sentence harder than usual, because the thing being built was a user interface, and only one of us could see it.
One sentence, and a question about lanes
Before it wrote a line, the agent went and asked the tracker some questions, read-only. 32 projects. 179 open tasks. And one awkward fact: ask Vikunja for a list of tasks and every one comes back saying it’s in bucket zero. The lane a card sits in is only reported when you read a project’s board, one project at a time.
So it read one board to see what that cost. 1.2 megabytes , for a single project, because every card comes with its full description. Thirty-two of those a minute was not a design.
It tried the other door instead: would the task list accept a filter on a bucket it refuses to report? It would. One small call per project to learn which bucket is called Doing and which is Blocked, cached for half an hour, then one filtered query per lane.
That’s the pattern of the whole evening in miniature. It didn’t assume the sensible-looking route worked. It measured it, found it was absurd, and went looking for a better one before building on either.
The first version was up a few minutes later: a panel with two lists, and a section at the top for any task the session had touched.
Every button needed two clicks
I clicked a task and it opened in the browser, which was the exact thing I was trying to get away from. That got fixed quickly: click a task, the panel shows the task, a Back button returns to the list.
Then I told it the Back button was unresponsive. It took a few goes to work.
The agent had a theory, and it was a reasonable one. Back and the background refresh were both writing to the same stored value, so a press landing mid-refresh could be overwritten. It rewrote that, gave Back its own tiny value that nothing else touches, and said plainly that this was a fix for its best guess, not a confirmed cause.
It didn’t fix it. My next report was more precise: I had to click everything twice. Tasks, Back, all of it.
Here is the part I like. The agent could not look at the panel. It tried: it asked for permission to take a screenshot of the Claude app, and the app refused, on the grounds that an agent operating its own window could change its own permissions. Fair enough. So it had no eyes on the thing it was debugging, and one failed theory behind it.
It didn’t produce a second theory. It built an instrument. The mod got a temporary trace that wrote every press, every focus change and every redraw to a log file with a timestamp, and I was asked to click a task, click Back, and click another task.
The log said this, 35 seconds in:
- 35.27 s: the focus ring moves to Back. No press.
- 35.95 s: a press on Back.
The first click never reached the button. It handed the panel the keyboard and stopped there. The second click, 0.7 seconds later, was the press. And every time the view changed, the button that held the focus disappeared with the old view, so the panel dropped the keyboard and the next click was spent getting it back.
The fix was to move the focus onto the new view’s first button whenever the view swaps. In the same log, after it went in: fourteen presses in a row, each one landing on a single click.
The trace came out again afterwards and the log file was deleted. One extra click when you first move from the prompt into the panel is still there, and it’s the app’s to decide who holds the keyboard.
Lesson: When you can’t see the thing, don’t guess harder. Make it write down what it’s doing, and read that.
Building a screen blind
“Fix the UX, it’s pretty ugly so far.”
That’s a hard instruction to give something that can’t see. The agent’s answer was to stop trusting itself and build a test: mount the panel on a pretend desktop, feed it a pretend tracker with one project and a handful of tasks, press the buttons, and check what was drawn. It doesn’t show what anything looks like. It does prove the app will accept the layout, that a click opens a task, and that Back comes back.
It used the same trick to settle a question about colour. Rather than hope a colour name was valid, it built a throwaway mod that drew one box in nine different colour spellings and ran it. Seven were accepted, including one it had made up to see what happened. ansi:blue was refused. That told it exactly how far the checking went, and it deleted the throwaway.
Then the test missed something, and it missed it for an instructive reason.
I’d asked for a way to start a session on a task, with a folder picker, the folders all living on my D: drive. Shortly after that went in, I opened a task and the whole panel went blank.
The agent’s test still passed. So it stopped testing with pretend folders and fed the test my real folder list. It failed at once, with the app’s own words: a dropdown takes between 1 and 64 options. My D: drive has 80 folders. One dropdown over the limit and the app declines to draw anything, not just the dropdown.
The picker now offers the 58 most recently changed folders and a box to type any other name. The test keeps a 90-folder case in it permanently.
Lesson: A test built on tidy pretend data proves the code works on tidy pretend data. The bug was in how much stuff I actually own.
The first time it saw its own work
Eventually I did the obvious thing and pasted a screenshot into the chat.
It was the first time the agent had seen its own work, and it found two faults in it before I’d said what I wanted. Titles were being cut off with a third of the row still empty, because it had counted characters as if they were all the same width and the panel uses a proportional font. And the titles didn’t line up on the left, because each one was being centred in its own space.
I’d only asked for the priority badges to move to the end of the row.

Where it ended up. This is the built-in demo data, not my board; more on why further down.
After that we did the visual work the way you’d do it with a designer: one change at a time, and I said keep or drop.
- A highlight when the pointer is over a row. Kept.
- Priority and status as filled badges instead of coloured letters. Kept, then moved to the end of the row, then made all the same width, because “P1” is narrower than “P4” and a ragged column of badges is worse than none.
- A bordered card around each section. Kept.
- A summary bar across the top , showing the proportions of the three lanes. Dropped the moment I saw it. The sections already have counts.
That last one was the agent’s own suggestion, built properly, and it was redundant. It took me about four seconds of looking to know. The agent couldn’t have known at all.
Scrubbed, published, and still showing my board
The next morning I had it moved somewhere permanent and put on GitHub.
The agent did what it always does before something goes public. It pulled the house out of the code: the server’s address, my hostname, the D: drive and the home-improvement rule all became settings. Then it searched every file in the history for addresses, hostnames, user names, emails and anything shaped like a token. Nothing. It published, and opened a pull request containing the whole repository so a review bot could scan it for personal details as a second opinion. The bot came back with three minor code findings and nothing about privacy.
A while later I asked one idle question: is my “all but home improvement” filter in the committed code?
The filter wasn’t. But the agent went and looked rather than answering from memory, and came back with something else. The test it had written to check the panel used a pretend tracker, and it had populated that pretend tracker with three of my real project names , and real id numbers off my board. They’d been public since the first push.
Its search had been for the things that grant access or locate a house. Project names are neither, so the search wasn’t looking for them, and the review bot hadn’t blinked either. They’re not secrets. They are, though, a list of what I’m working on, sitting in a public repository because nobody asked the right question.
The names are invented ones now. The old ones are still in the history; I chose not to rewrite it.
The same morning produced the fix for the other half of that problem. I wanted screenshots in the README, and every screenshot I’d taken showed my real tasks. Blurring them would have looked rough and risked missing one. The agent built a demo mode instead: type /vikunja demo and the panel switches to a small invented tracker held in memory, with made-up projects, people and folders. A test proves it makes no network, disk or process calls while it’s on. Both screenshots in this piece are of that.

A task, in demo mode. On my own board that folder picker lists the real folders on my drive.
Lesson: A scrub finds what you told it to look for. Decide what counts as private before you search, and include the boring things: names, numbers, the shape of your to-do list.
Smaller traps, for the record
- The validator on my machine was too old. The command-line Claude Code on my PATH was 48 builds behind the one inside the desktop app, and rejected the mod’s format outright. The agent found the app’s own copy and validated with that.
- “5 of 6”. With nothing filtered, the Doing section read “5 of 6”. The sixth was a home-improvement task the panel hides by default. A count that makes you ask what’s missing, when you haven’t hidden anything yourself, is a bug. Hidden projects are now out of both numbers.
-
Vikunja has no custom fields. A task’s folder is stored as a label on the task,
folder:and then the path. It shows up as a chip in Vikunja itself, and anything else that reads the task can see it. - Updating a task replaces the whole task. Send Vikunja only a new description and it resets the rest, assignees included. The mod never updates a task. It only adds or removes that one label.
-
The mod can make a folder only by writing a file into it. So a folder created from the panel starts life with one empty
.gitkeep. -
Git Bash ate a slash. Checking that a fresh session loaded the mod, the agent ran the
/vikunjacommand from a shell that helpfully rewrote it into a Windows path. The fresh session received a file path, shrugged, and did something else entirely. Twice, before the shell was told to leave it alone. - A clear button for a toggle is silly. A Show/Hide switch displays its own state. It doesn’t need a Clear next to it.
- The wrong number. Every task has two: the one on its card, and an internal one that only appears in its web address. The agent had been naming tasks by the internal one, which I can’t find on the board. It now uses the number I can see, and that’s written down where every future session will read it.
- A link can’t have a keyboard shortcut. So “Open in Vikunja” became a button that asks the operating system to open the page. The app then drew its own little key badge beside it, which made the “(o)” in the label redundant within the minute.
- One review an hour. The review bot’s free allowance stalled two pull requests with every finding fixed and nothing left to approve them. It’s been replaced with a secret scanner and a review that runs on the Claude subscription I already pay for.
What it looks like now
- Three lanes in cards: Doing, Blocked, To-Do, grouped by project, with a priority badge and who it’s assigned to. To-Do stops at 80 rows and tells you how many more there are.
- Filters for who, which project, what priority and what’s due (overdue, this week), and a Show/Hide for Blocked.
- Find # : type the number on a card and it finds the task in any project or lane, finished ones included.
-
Keyboard shortcuts :
rto refresh,oto open the task in the browser,bto go back. - A demo mode with an invented board, for screenshots and for trying it without a server.
- Home-improvement projects hidden by default , because this panel sits next to code.
- The task the session is on opens by itself , and is re-read every 10 seconds. It only redraws when something changed, so it doesn’t fight your scrolling.
- A Start session button on every task, with its folder.
- About 1,200 lines in the main file, a test that runs the panel on two surfaces, and three settings so it isn’t welded to my house.
The honest part
The other pieces ended on what unblocked the thing: the tedious middle could be handed off, or the earlier projects already existed, or the brief was wrong and something that could see the house said so. This one’s different, because nothing was blocked. I asked for it in a sentence and had a working panel in minutes.
What I noticed instead was which of my contributions mattered.
I didn’t supply a design. I didn’t know the plugin API existed in any detail, and I couldn’t have told you a dropdown stops at 64. Every technical decision in this, the filtered query, the label instead of a field, the focus fix, was the agent’s, and each one was checked against the real system before it was built on.
What I supplied was eyes. “It opens in the browser.” “I have to click everything twice.” “It’s pretty ugly.” “Put the badges last.” “That bar is redundant.” “Is my filter in the committed code?” None of those is a specification. Every one of them is something only a person looking at the screen could say, and every one moved the thing further than the paragraph of requirements I might have written instead.
And the agent’s side of that bargain was knowing what it couldn’t see. It said so each time: I haven’t seen it draw. This is a fix for my best guess. I’m least sure about the borders. When it was refused a screenshot it didn’t pretend; it built a trace, then a test, then asked me for a picture. I’d take that over confidence.
Now the part that isn’t finished. Two things in the mod have only ever run against the pretend tracker: writing the folder label to the real one , and Start session , which hands off to the desktop app to offer a new session. The test says both do what they should. I’ve used the lanes, the filters, the search and the shortcuts on my real board. I haven’t pressed either of those two buttons in anger. And the new session probably starts in a fresh working copy of the folder, which will suit a git repo and may well fail on a plain one. I don’t know yet.
There’s also a small, pleasing loop in what this is. The panel exists so I can watch the agent work through the board. The agent built the panel. It was told how the panel looked by the one participant who could see it, and it now sits beside the conversation.
Build the tool that lets you watch the work. Then be the one who looks.
Under the hood: if you want to build one
What it runs on
- Claude Code 2.1.286, desktop app. Mods are plugins of function hooks: one TypeScript module that hooks session events and draws panels. The API is early access and moves between releases.
- Vikunja v2.4.0, self-hosted. Read over its REST API with a token from the environment.
- vikunja-mcp-ng v0.6.0: the MCP server the agent uses to work the board. The mod watches for its tool calls to know which task the session is on.
Decisions
- Read the tracker directly, not through the agent. The panel polls on a timer and costs no model calls.
-
Find lanes by bucket name, then filter tasks by bucket id. A task list reports bucket 0 for everything, but accepts
bucket_id in …as a filter. Reading whole project boards was 1.2 MB a project. -
The folder is a label, not a field. Vikunja has no custom fields;
folder: <path>is visible in its own UI and to every other client. - Never send a task update. It replaces the whole task. Add and remove a label instead.
- Write state only when it changed. A write redraws the panel; redrawing every 10 seconds fights scrolling.
- Site details are settings. API address and token from the environment; web address, folder root and the hidden project tree from the plugin’s options.
- A demo mode instead of blurred screenshots. An in-memory tracker with invented data, proven by test to make no outside calls.
- Search by the card’s number, not the internal id. The number is per project, so one number can list several tasks; the project name tells them apart.
- Rejected: a custom-drawn board with drag and drop. Possible, and a much bigger build than a list you can click.
Gotchas
- The first click on a panel only gives it the keyboard. And a view change drops it again unless you move the focus to an element in the new view.
- A dropdown takes 64 options at most. One more and the entire tree is refused; the panel goes blank.
- The desktop draws text in a proportional font. Character-count truncation runs short. Cut generously and let the box clip.
- Badges of equal width need a fixed-width box , with the label centred in it.
- Colour values aren’t checked against a palette. Any word passes validation; only the format is refused.
- The agent can’t screenshot the app it runs in. Budget for a trace, a mount test, and a person.
- Check which Claude Code your shell finds. An older CLI rejects the plugin outright.
- A mod makes directories only on the way to writing a file.
- A Link can’t carry a hotkey. Use a Button and open the page through the host. Don’t put the key in the label: the desktop draws its own badge.
- The focus ring is drawn outside its button , and the pane’s edge clips it. Leave a cell of padding on the top row.
- Scrub test data too. Fixtures are where real names hide.
The focus fix
The few lines that turned two clicks into one. PANE is the pane’s id; the key is whichever button the next view leads with.
// Denied while the prompt holds the keyboard, which is the person's to give.
const keepFocus = ($, key) =>
$.ui.focus({ requestId: PANE, key }).catch(() => {})
// opening a task
await update($, openId, () => task.id)
void keepFocus($, 'back')
// going back to the lists
await update($, openId, () => 0)
void keepFocus($, 'refresh')
Take it with you
The whole mod is one repository: angusmaul/claude-code-vikunja-tasks. MIT. It needs a Vikunja address and token in the environment, and takes three optional settings. The README says what it reads, the two things it writes, and how to run the validator and the test.
Used by hand on a real board, in the desktop app on Windows: the lanes, filters, search, task detail, auto-open, live refresh and keyboard shortcuts. Proven only against a fake tracker: saving a folder label, creating a folder, and Start session. Not run at all: the terminal, macOS, Linux.

Top comments (0)