DEV Community

yashraj jangir
yashraj jangir

Posted on AI-assisted

Focus Agent: A Local AI Focus Companion Built for a Friend

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

Focus Agent: A Local AI Focus Companion Built for a Friend

This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend, including the Best Use of Gemma category.

Project: Focus Agent on GitHub

Demo: Watch the captioned walkthrough

Focus Agent dashboard

The idea: help Raj return to his own priorities

Raj is a friend who is talented and ambitious, but can get pulled into interesting side quests and lose track of the task he meant to finish. A to-do list can remember the plan. A screen-time chart can count applications. Neither can connect the two in a way that gives him a useful, timely nudge.

So I built Focus Agent: a desktop companion that lets Raj state his priorities, observes which application is active, and checks whether that activity seems connected to the current task. If the relationship is unclear, a local Gemma model helps interpret a small, allowlisted context. If distraction remains likely and Raj ignores the check-in, the intervention becomes more visible. Raj can correct the classification, switch tasks, take a break, or pause monitoring.

The agent is meant to help someone act on their own intentions, not to judge their productivity. It explains its reasoning and leaves the decision with Raj.

What works today

  • Turn a natural-language plan into a short list of tasks and set the current priority.
  • Track application changes through a small KDE KWin script and local D-Bus bridge.
  • Save activity sessions, tasks, corrections, breaks, and intervention responses in local SQLite.
  • Use Gemma through Ollama for ambiguous activity, with strict JSON validation and deterministic policy checks.
  • Remember Raj's “still related” corrections for a task and application.
  • Escalate ignored drift prompts from a quiet notification to a full-window intervention.
  • Pause monitoring for 15 or 30 minutes, 1, 2, or 4 hours, or until Raj resumes it.
  • Keep planning usable with a deterministic fallback if Ollama or Gemma is unavailable.

Gemma matters to the core loop: application names alone do not tell us whether a window is useful for the task. Running the model locally means activity context does not need to go to a hosted inference service. The project is open source, so the collection and decision flow can be inspected.

Development environment

This prototype currently targets Linux with KDE Plasma 6 on Wayland. To run it, you need:

  • Node.js 22 or newer and npm;
  • KDE's KWin script and D-Bus support;
  • Ollama with a Gemma-family model (the README uses gemma3:1b as a small starting point).

After installing dependencies, the repo documents npm test and npm run dev, plus commands to install the KWin activity bridge. The model download is needed for local reasoning; once the model and dependencies are installed, inference does not need a hosted AI API.

This is a Linux prototype, not a cross-platform release. The activity bridge is KDE-specific, and I have not yet validated it on X11, Windows, or macOS.

What I saw while testing

In one short real desktop session, Firefox, VS Code, and ChatGPT were classified as related to the active task. I then used Chrome for about ten minutes; Focus Agent left it as unknown. That conservative result is useful: an unfamiliar browser should not automatically be called a distraction just because it is a browser. It also showed me a current limitation: without browser-domain or page context, application-level signals can be too coarse to tell what someone is doing.

I also exercised the reasoning path with gemma3:1b: an obvious game launcher produced a high-confidence distraction result, while an unknown browser stayed low-confidence/unknown. These are prototype checks, not a broad accuracy study. More testing with Raj and other users is still needed.

The demo includes representative preview data for the full intervention sequence. Waiting through real cooldowns made it impractical to capture every escalation level in a short walkthrough, so I used preview data for that part rather than implying the video is a continuous live session.

Setbacks that changed the design

One UI bug taught me that activity tracking and interface actions can interfere with each other. When I clicked Still related, Focus Agent itself became the active window before the correction landed, so the action could apply to the wrong activity. I changed the flow to preserve the intended activity target and filter Focus Agent's own window from collected activity.

Another challenge was deciding what an application name can really prove. My Chrome observation remained unknown, and that is preferable to inventing certainty. The bridge currently sees application-level context, not browser URLs, page contents, or documents. I kept that boundary explicit and made corrections available to the user.

I also started after missing the first few days of the challenge, so the scope had to stay focused: a working Linux prototype, a clear demo, and a privacy model I could explain rather than a premature promise of support for every desktop.

What comes next

I want to finish this as a dependable cross-platform project. The next work is to add activity adapters for other Linux desktop environments and X11, then build and test native activity integrations for Windows and macOS. I also want to test with Raj, improve the app-level ambiguity handling without collecting more personal data than necessary, and make packaging and first-run setup easier.

If you work on desktop integrations, local AI, privacy-respecting productivity tools, or testing across Linux setups, contributions and issue reports are welcome. Please open an issue or pull request on the GitHub repository; small, well-scoped contributions are especially helpful.

KDE Wayland release candidate

The Focus Agent v0.1.0-rc.1 GitHub release is available for KDE Plasma 6 on Wayland, x86_64. The AppImage bundles Electron and its Node.js runtime, so you do not need to install Node.js separately. Ollama with Gemma is still needed for local reasoning, and the KWin activity bridge is available as a separate download on the release page. The README includes setup steps. This is an unsigned release candidate built on Kali Linux Rolling; it has not been validated on other distributions.

Related reading from DEV

These community posts cover adjacent parts of the problem: running Gemma locally with Ollama and designing a local-first desktop AI workspace.

Agent Session summary

Focus Agent: From Friend's Focus Problem to KDE Wayland Release
You

I am building a project for Hacktoberfest Weekend Challenge: Build for a Friend. I have a PRD and want to discuss enhancements, boundaries, and improvements.

Agent

We framed this as a one-person build for the October 5 deadline, with a focused MVP and room to go further. The project should solve a concrete problem for a friend, respect privacy, and state prototype boundaries honestly.

You

The friend is Raj, a pseudonym. He is talented but gets distracted and forgets priorities. The prototype targets Linux with KDE Plasma; I do not yet know whether the session is Xorg or Wayland.

Agent

We chose KDE as the first prototype target and later used the available KDE Plasma Wayland environment for activity integration and packaging. Raj's chosen priority remains the reference point; the agent should use activity as evidence, not as a verdict.

You

My personal stack is JavaScript. Should we shift to Electron?

Agent

We aligned the desktop app with the JavaScript stack: Electron for the shell, React and TypeScript for the interface, a KDE KWin script and D-Bus bridge for Linux activity events, SQLite for local state, and Gemma through Ollama for local reasoning.

You

I want the dashboard to contain only essential statistics and core functions. Put the story and other material in an About tab, keep the logo, remove the native window frame, and add in-app minimize, maximize, and close controls.

Agent

The interface was reorganized around a minimal dashboard and an About view, with a custom frameless window and matching window controls.

You

Interventions should escalate after repeated ignores, eventually showing a full window with current work statistics. The user needs choices to delegate, dismiss, take a break, or pause monitoring for 15 minutes, 30 minutes, 1 hour, 2 hours, 4 hours, or until they resume it.

Agent

We implemented progressive intervention levels and user-controlled responses, including timed and indefinite monitoring pauses. The high level shows the current priority and relevant work statistics.

You

When clicking Still related for an activity, the active context switches to Focus Agent so the correction can affect the wrong activity. What can we do?

Agent

We adjusted activity targeting so the correction retains the intended session instead of relying on whichever window has focus after the click, and added filtering so Focus Agent does not classify its own window as user activity.

You

I completed the early test phases. Firefox, VS Code, and ChatGPT were related to the task. I then used Chrome for about ten minutes and it stayed unknown.

Agent

We recorded that as one observed session, not a general accuracy result. An unknown browser was treated conservatively because application-level tracking does not see the domain or page content. Separate Gemma checks showed high-confidence distraction for an obvious game launcher and low-confidence/unknown for an unfamiliar browser.

You

I am recording the demo. The first recording was imperfect, so we need to redo it. I am considering a voiceless captioned version first, then a later voiceover.

Agent

We prepared the demo from the project's video, created synchronized captions, and produced a captioned walkthrough. Because waiting through live cooldowns was impractical, the complete escalation sequence in the walkthrough uses representative preview data. The user uploaded the demo to Google Drive.

You

Rewrite the README with dashboard.png, caption-free screenshots from the demo, and the demo video link. There also seems to be a Mermaid syntax error.

Agent

The README was rewritten for judges with dashboard and walkthrough screenshots, the Drive demo link, feature and privacy explanations, environment requirements, test observations, known limitations, and an architecture diagram. The Mermaid graph was simplified to use ordinary directed arrows and standard SQLite node syntax.

You

Create a DEV submission focused on the idea, current capabilities, environment needs, observed results, setbacks, cross-platform goals, and an invitation to contribute.

Agent

A post was drafted with the required hf26challenge tag and links to the repository and demo. The user reviewed and published the post, then asked to add release information.

You

After I submit the article, build a release candidate for the current KDE and Wayland configuration and provide it on GitHub.

Agent

We added electron-builder configuration and release instructions, then built an x86_64 AppImage on Kali Linux Rolling under KDE Plasma 6 Wayland. The AppImage includes Electron's Node.js runtime; Ollama with Gemma and the separate KWin/D-Bus bridge remain requirements. A Debian target was dropped after its maintainer-email requirement could not be met. The release candidate is unsigned and has not been validated on other distributions.

Agent

The AppImage, KWin bridge archive, and SHA-256 manifest were published as the prerelease v0.1.0-rc.1 at https://github.com/yasshhhraj/focusagent/releases/tag/v0.1.0-rc.1. The source changes and screenshots were committed as 6f8db3e and pushed to main. The published DEV post was updated with the release link. The package build and AppImage runtime metadata check succeeded, but no interactive launch test was performed.

Tags: #hf26challenge #gemma #linux #productivity

Top comments (0)