DEV Community

Cover image for I Ran PI-Desktop's Plan, Goal and Subagent Modes on a 1.7B Local Model. Here's What Broke.
Ishank Choudhary
Ishank Choudhary

Posted on Originally published at Medium

I Ran PI-Desktop's Plan, Goal and Subagent Modes on a 1.7B Local Model. Here's What Broke.

Pi, the open-source coding agent engine, reached 1.0 on 1 October. Three days later, on 4 October, PI-Desktop shipped v0.16.1, the first release built on Pi 1.0.0. PI-Desktop is a desktop GUI for that engine: Electron on the front, a Rust "host core" underneath, and the Pi agent packages running in a Node sidecar.

One thing to clear up first. PI-Desktop is a community project by vastsa, not something from the Pi team. It's LGPL-3.0, has about 6.4k GitHub stars, and its README calls the 0.16.x line an "Early Preview". Several other projects use the "pi-desktop" name, so if you go looking, start from the repo and the official docs.

I didn't want to write a feature list. I wanted to know what actually happens when you install it on a plain Linux box and run the three things that set it apart from a terminal agent: Plan mode, Goal mode and subagents.

The honest caveat up front

There was no cloud API key on my test machine. So I ran everything against qwen3:1.7b through Ollama, on CPU only: Debian 13, 8 vCPUs, no GPU. That model managed about 250 tokens/s on prompt processing and about 8 tokens/s on generation.

That is a stress test, not a fair benchmark. A 1.7B model is the weakest setup the app would ever see. Some of what broke is the model's fault, and I'll say which parts. Other things broke regardless of the model, and those are the more useful findings.

The test project was lukeed/clsx v2.1.1: 21 files and 32 passing tests. Small enough for a tiny model to have a chance.

Install: fast and uneventful

Linux gets an AppImage, a .deb and an .rpm (glibc 2.35+, so Ubuntu 22.04+, Debian 12+ or Fedora 36+). I used the AppImage, extracted it instead of mounting it so libfuse2 wasn't needed:

wget https://github.com/vastsa/PI-Desktop/releases/download/v0.16.1/PI-Desktop-0.16.1-linux-x86_64.AppImage
chmod +x PI-Desktop-0.16.1-linux-x86_64.AppImage
./PI-Desktop-0.16.1-linux-x86_64.AppImage --appimage-extract
./squashfs-root/AppRun --no-sandbox
Enter fullscreen mode Exit fullscreen mode

What I measured:

  • Download: 164.8 MB, and the SHA-512 matched latest-linux.yml.
  • Extract: about 3 seconds, 414 MB on disk. No missing shared libraries according to ldd.
  • First launch: the window appeared in 1.56 s cold and 0.54 s warm.
  • Memory: about 630 MB RSS across the Electron, host-core and sidecar processes, before the agent did anything.
  • --no-sandbox was needed only because the extracted tree's chrome-sandbox isn't setuid. The installed packages shouldn't need it.

Inside the bundle you get Chromium 150, the Rust pi-desktop-host-core binary, the Pi sidecar, a bundled models.dev catalog of 8,344 models, and two built-in plugins (a browser and a file manager).

First launch: the

The first screen is a calm "What can I help you build?" with a Get started checklist. The composer has three chips that matter for this test: the mode (Agent / Plan / Goal), the permission (Ask every time / accept edits / Auto) and the model.

Hooking up a local model

There's no dedicated Ollama tile in 0.16.1. Local models go in through Settings → Agent → Models → Add provider → Custom endpoint, which takes "any OpenAI- or Anthropic-compatible URL".

Base URL:   http://127.0.0.1:11434/v1
API key:    ollama        (dummy)
API format: OpenAI Chat Completions
Enter fullscreen mode Exit fullscreen mode

Fetch list found the model and the provider showed "Connected · 1 model found". The provider list is long, too: subscription sign-ins (Claude Pro/Max, ChatGPT, GitHub Copilot, SuperGrok/X Premium and others) plus around 30 API-key providers.

A smoke test ("Reply with exactly: PI-Desktop smoke test OK") worked, but it took 1 minute 54 seconds, and it showed me the first real number:

  • The app's system prompt plus tool schemas came to about 4.1k tokens (Ollama cached 4,098 of 4,717 prompt tokens).
  • The context ring in the composer read 96% after one turn, because I'd set Ollama to a 16k context and the custom model's context window showed as "— —" in the UI.
  • The reasoning setting appended /think to my prompt (Qwen's reasoning switch), and the model echoed it back in its answer.

Lesson one: when you add a local model, open the model's Advanced settings and set the Context Window yourself. With 16k of context, the harness takes a quarter of it before you type anything.

Plan mode: it worked, but the plan was fiction

I opened the clsx folder as a project, set the chip to Plan with "Ask every time", and asked it to plan a test for nested arrays and falsy values.

Plan mode result: a plan card with Approve/Reject next to the saved plan file

The mechanics were solid:

  • The plan was saved as a real file, .pi/plan/add-nested-array-falsy-values-tests-for-clsx-20261007-1108.md (315 bytes).
  • A plan card appeared with Open plan, Reject and Approve (Ask), and the File Manager pane opened showing the markdown.
  • The chat got an automatic, sensible title.

The content was not. The plan said to add test files in /work/clsx/tests, a path that doesn't exist. The run log showed "3 issues · 4 tools": a search failed, a read failed, the first SubmitPlan failed and the second one succeeded. The app reported "Processed for 5m 19s", but wall-clock time from sending the prompt to seeing the card was about 17 minutes. The context meter sat at 95%.

That's mostly the model. A 1.7B model with a quarter of its context taken by instructions invents paths. But the workflow around it (a separate artifact, an explicit approval gate) is exactly what I'd want to put in front of a stronger model.

Goal mode: the permission chip changed on its own

Goal mode is meant to write a .pi/goal/*.md file with an outcome and acceptance criteria, wait for approval, then switch to Agent and work until it can report which criteria it verified.

I clicked the mode chip twice to get to Goal. Then I noticed the permission chip next to it had silently switched from "Ask every time" to "Auto". I hadn't touched it. For a mode whose whole job is to run autonomously toward a target, that's the one setting I'd expect it never to loosen on its own.

The run didn't go well:

  • The first model call took 4 min 26 s.
  • No goal file and no acceptance criteria were ever written.
  • The model went looking in /work/tests again.

Goal mode: even under Auto, a Grep on a path outside the workspace triggered a permission prompt

The safety net did work: even under Auto, a Grep on /work/tests stopped with "Accesses a path outside the session workspace", rated LOW RISK. I denied it. The model then tried to Edit /work/tests/new-test.js, which failed. I stopped it at about 12 minutes with nothing verified, and the chat had been auto-renamed in Chinese ("新增嵌套假值测试用例", roughly "add nested falsy value test cases").

The failed run is the model's fault. The permission flip and the Chinese title happen in the app, and they're worth reporting to the maintainers. Fairly, the docs do say Plan and Goal are "contract modes, not strict read-only security profiles". I'd still check that chip every time you switch modes.

Subagents: the one clean win

PI-Desktop ships five built-in subagents, each with a defined tool set:

  • Explorer: Read, Glob, Grep, Bash
  • Code reviewer: Read, Glob, Grep (read-only)
  • Test runner: Read, Glob, Grep, Bash
  • Fixer: adds Edit and Write
  • UI designer: adds BrowserPreview, Edit and Write

The five built-in subagents in Settings, each with its own tool set

You can add your own as markdown files in ~/.agents/subagents/ (up to 16). Delegation uses a Task tool that runs in the background and returns a delegation ID, with TaskWait, TaskList and TaskStop alongside it, and up to 10 at once per session. Subagents only work in Agent mode.

In Agent mode I asked: "Use the test-runner subagent to run npm test and report the result."

The test-runner subagent asks for approval before running npm test

This went well:

  • The UI drew a small delegation graph: Main agent → test-runner, "Coordinating 1 delegated task".
  • The permission prompt said clearly that it was "Asked by the test-runner subagent", showed the exact command (npm test) and labelled it HIGH RISK. Running a shell command is the right thing to flag.
  • After Allow once, the subagent finished in 2 min 44 s (3 steps). The whole turn took 5 min 55 s.
  • The result was exit 0, 32/32 passed, which matches what I got running npm test in a shell myself.

The subagent's own transcript opened in a side pane marked read-only ("Subagents are driven by the main agent"). And the chat was renamed again, to "npm测试成功" ("npm test succeeded").

The lesson: give a small model a narrow, well-defined job and the harness carries it. Ask it to plan open-ended work and it falls apart.

Two findings that have nothing to do with the model

1. "OS keychain" isn't quite true yet. The landing page and the README say credentials live in the OS keychain. I ran gnome-keyring with Secret Service available, saved the Ollama provider, and zero items went into the keyring. Instead, two files appeared:

~/.pi-desktop/secrets/<sha>.bin    48 B   mode 644
~/.pi-desktop/secrets/.machine-key 32 B   mode 600
Enter fullscreen mode Exit fullscreen mode

The project's own storage spec explains it: the shipped backend is an AES-256-GCM file store keyed by a machine key that "host-core generates once and keeps beside them", and the OS keychain backend is one "that neither host-core nor Electron main implements today, so a same-user process that can read the data directory can also decrypt the secrets." The spec is honest about this. The marketing isn't. If you put real cloud keys in it, treat the ~/.pi-desktop folder as sensitive.

2. The window only redraws when the mouse moves. On my Xvfb/xfwm4 desktop, timers and streaming output froze until I moved the cursor. That could be specific to my headless setup, but it made the long local runs feel even longer.

The rest, briefly

  • MCP Market: Settings → Agent → MCP has a Market tab that merges a built-in catalog with the official MCP registry. It listed Memory, Sequential Thinking, Filesystem, Playwright, Context7, Fetch, Git, Time, Serena and others, most marked Verified, each with its npx/uvx command. Some descriptions showed up in Chinese. As of v0.16.1, tools from user-added MCP servers need approval.
  • Session import: it can import sessions from Claude Code, Codex, OpenCode and Pi.
  • Scheduled tasks: hourly, daily or weekly, but they only run while the app is open.
  • No telemetry and no required account, according to the README.

What a bigger cloud model would likely change

I didn't test this, so treat it as an expectation, not a result. With a capable cloud model and a large context window, I'd expect:

  • The 4.1k-token system prompt to stop mattering. At 128k+ context it's a rounding error, not 25%.
  • Plan mode to produce real paths, since a stronger model would actually read the repo before planning.
  • Goal mode to write its acceptance criteria and run through them in minutes, not stall after 12.
  • Total time per turn to fall from minutes to seconds.

What I don't expect to change: the permission chip flipping to Auto, the keychain wording, the Chinese auto-titles and the redraw bug. Those are in the app.

Verdict

PI-Desktop v0.16.1 installs in seconds, starts in under two, and its Plan, Goal and subagent workflows are thought through: separate artifacts, approval gates, clear risk labels, a visible delegation graph. Subagents did their job even on a 1.7B model on CPU.

It's also clearly an early preview. Check the permission chip when you switch to Goal, don't trust the "OS keychain" line yet, and set your local model's context window by hand.

If you run it with a frontier model, I'd like to know whether Goal mode holds up. That's the test I want to see next.


Originally published on Medium.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to