Persistent Agents Are a Systems Problem
A persistent personal agent must retain context, operate across devices, reach applications, continue working when its interface is closed, and remain accountable for consequential actions.
That is the context for nanoMuse’s October 6, 2026 paper, written by Guangyi Liu, Yong Liu, and Jiangning Zhang. nanoMuse is presented not as another language model, but as an end-to-end software system in which the model is replaceable.
Unlike model-centered projects such as Meta Muse Glimmer-30B, nanoMuse supplies surrounding machinery: persistent identity, execution environments, routing, memory, interfaces, schedules, and approval boundaries.
The complete implementation is published in the official nanoMuse repository.
Architecture: One Account, Multiple Agents
nanoMuse describes one personal agent from the user’s perspective, but each device runs its own agent runtime. After sign-in, those runtimes meet through a relay and share a conversation.
That design avoids treating every phone or computer as a thin terminal for one remote virtual machine. A task can be entered on a phone and executed on a computer. Prefixing a message with a target such as @Mac routes it to that device, while approval requests can return to the device currently in the user’s hand.
The relay carries conversation text and coordinates devices, while files and screenshots remain where they were created. It can use the project’s community infrastructure or a user-controlled server.
This boundary is not an automatic privacy guarantee. Self-hosting controls message routing, but an externally hosted model provider may still receive model requests. Relay location, model location, and device execution are separate decisions.
Device-Local Execution
Phone and Desktop Runtimes
On Android, nanoMuse places an Alpine Linux environment under proot inside the APK, with a shell, browser, MCP support, and skills. The project credits OpenMinis as the foundation for this on-device agent.
On desktop, nanoMuse uses DeepSeek Harness with a Python runtime for screen-operation capabilities. The desktop application targets Windows, macOS, and Linux. The wider project also includes iPhone, iPad, and browser clients, although the repository’s comparison table says screen-based “hands” are not available on iOS.
These runtimes give the system more than a chat interface. The agent can use a Linux shell, browser, MCP servers, and application-specific skills. It can also schedule routines, check goals, and produce a feed while the visible application is closed.
The Relay Is Coordination, Not Execution
The relay joins devices and lets them request work from one another, but execution remains associated with a device runtime or configured server environment.
A disconnected device may be unable to receive a routed task, and a local shell is not remotely reachable without the relay path. Moving coordination to a personal server also does not make every dependency local.
Memory You Can Inspect
nanoMuse stores its identity, user knowledge, and wake schedule in Markdown files. Users can inspect and correct those files directly rather than relying entirely on an opaque memory store.
The paper nevertheless treats stronger memory provenance as roadmap work. Editable Markdown is not a completed system explaining when, why, and from where every remembered statement originated.
The repository contains a working system, but the paper identifies evaluation, provenance, and open screen-operation models as continuing work.
Giving the Agent “Hands”
Many applications do not expose a suitable API. nanoMuse addresses that gap with “hands”: screen-based operation for Android applications and desktop windows.
The desktop implementation builds on UI-TARS-desktop, while other acknowledged projects contribute phone-operation components and traces. Screen interaction lets the agent work through visible interfaces when no dedicated integration is available.
It also introduces uncertainty. Graphical interfaces change, and the project’s paper lists an evaluation suite for the hands as roadmap work. The available material does not establish that screen operation is uniformly reliable across applications or workflows. Developers should treat it as an execution mechanism requiring observation and recovery, not as equivalent to a stable API contract.
Sentinel and Consequential Actions
Every action is described as passing through a Sentinel. The repository says nanoMuse pauses before deleting, sending, or paying, with approvals scoped for one occurrence, the current chat, or always.
Passwords and verification codes remain for the user to type. Logins and CAPTCHAs can be handed back to the user, after which selecting “Done” resumes the task.
This approval model places a human checkpoint before actions the project classifies as difficult to undo. It does not eliminate risk. Broad “always” approvals trade repeated interruptions for a larger standing authorization, while shell access, browser control, and screen operation create a wider potential impact than chat alone.
These are project-described controls, not evidence of independent security verification. Anyone deploying the system should evaluate relay exposure, provider selection, device permissions, approval scope, and the consequences of an agent operating personal interfaces.
What Is Actually Open and Self-Hostable?
The repository says the phone app, desktop app, web console, and relay are included under GPL-3.0-or-later. Users can run their own relay, use their own model-provider keys, or configure a custom runtime for the web application.
That is substantial system-level openness, but it does not mean every possible deployment is fully local. The community relay is a hosted service with a model-use allowance, and many supported model choices are external providers. The model itself is deliberately interchangeable rather than supplied as one mandatory open component.
The paper also places an open model for screen operation, provenance-aware memory, and a hands evaluation suite on the roadmap. Those items should not be reported as finished merely because the surrounding applications and relay are available.
Why Developers Should Care
nanoMuse offers a concrete architecture for a problem often reduced to model selection. Persistent agents also require routing, device identity, local execution, memory formats, approvals, scheduling, and fallbacks for interfaces without APIs.
Developers can inspect where conversation text travels, where files remain, which device executes a task, what users can edit, and when approval is required.
nanoMuse does not settle the reliability or security questions raised by access to personal devices. It provides an open, self-hostable system in which those questions can be examined beyond the model layer.
Top comments (0)