DEV Community

Hari N for GetPullRequest

Posted on Originally published at getpullrequest.com on

Ship code from your phone with the AI coding agents you already pay for

Ship code from your phone with the AI coding agents you already pay for

We run AI coding agents like Cursor, Claude Code and Codex from our phones. The agents don't run in a vendor's cloud sandbox. They run on our own laptops, and each of us has also paired a cloud desktop as a second machine. The phone is how we give them work and check on them. The tool that does this is GetPullRequest (GPR), and we have been building GPR with GPR since May. Every branch a GPR task or session creates is named gpr/task-tsk_…, and by now that is most of our branch list.

TL;DR: GPR runs the AI coding agents you already pay for (Cursor, Claude Code, Codex and five more) on your own machine, using the Agent Client Protocol. Tasks can come from the app, GitHub, GitLab, Jira or Slack. You answer the agent's questions and merge its pull request from your phone. Your machine needs to be on.

This post covers how a task runs, why we keep the agent on your machine, and some of the strange problems we ran into while building it. If you want the story of why we built GPR in the first place, read the launch post.

Claude Code remote control, Codex, Cursor background agents: why not just use those?

It's a fair question. Claude Code has Remote Control, Codex works with the ChatGPT app, and Cursor has background agents (now called Cloud Agents) that you can start from Slack or their iOS app. If you only use one agent and just want to check on one session from your phone, those tools are good and you should use them.

The problem is that each one is a chat with a single vendor's agent. Claude Code's Remote Control, for example, only works with Anthropic's models and a claude.ai login. If you run Claude Code with an API key or through OpenRouter, Remote Control won't start. Our team was using four different agents, and nobody wanted to switch to just one. Our work also didn't start as a chat. It came in as GitHub issues, Jira tickets from the product team, and screenshots posted in Slack. What we wanted was an assistant for every developer that owns a task board, keeps track of the work on it, and gets it done with whichever agent fits.

How a task actually runs

You create a task on the board in the app. It can also come from a GitHub issue with the gpr:agent label, from a Jira project, or from an @gpr mention in Slack. The task goes to the workspace where you ran gpr setup.

The GPR daemon on that machine picks up the task and creates a git worktree for it under ~/.gpr/worktrees/. This keeps the agent away from the branch you are working on. The daemon then starts the agent you chose for that workspace, and the agent plans, implements, verifies, writes a commit summary and opens a pull request. GPR saves a checkpoint after each step, so if something interrupts the run, it continues from the last checkpoint instead of starting again.

If the agent has a question while it is planning, the task moves to "Needs your input" and you get a notification on your phone. You reply on the task and the agent continues. When the pull request is open, the task stays in "In review" until it is merged. You can also open the live session at any time and watch what the agent is doing.

None of this is complicated on paper. Most of our time went into making it work reliably on real laptops, real networks and with real agents.

Why AI coding agents should run on your machine: faster, and already paid for

Our repo is an Nx monorepo with a FastAPI backend, a Flutter app, a Go daemon, a generated Dart SDK and a couple of other products. Building it from a fresh checkout takes quite a few steps.

We tried hosted sandboxes first, using Daytona. The first problem was setup. We had to get the backend, the Flutter app, the database and the daemon running together inside a sandbox, and keep them working as the code changed. The second problem was cost , which was hard to justify for a startup when our laptops were sitting idle half the day. The biggest problem was testing. Every time we wanted to check an agent's change, someone had to check out the branch and build everything from scratch.

"We were paying to rebuild an environment in the cloud that was already fully built on the laptop in front of us."

When the agent runs on your machine, it runs the tests the same way you do. Your code, your agent's login and your API keys stay on your machine. The GPR backend only holds the board and the task queue, passes the session to your phone and sends notifications. It never has a copy of your repo.

The downside is that your machine has to be on. If your laptop goes to sleep, the task waits as "Waiting for machine" and continues when the laptop wakes up. That's why each of us also paired a cloud desktop.

Everything works with nothing connected

Our first plan was to get diffs and merge status from GitHub. We dropped it before writing any code. The agent's changes are on your machine, often not even committed yet, so GitHub doesn't know about them. GPR reads them with git on your machine instead. If you connect GitHub, it adds a bit extra, like whether the PR can be merged or is behind the base branch.

This became a rule for the whole product, and we call it the enrichment model. Coding, reviewing, merge checks, resolving conflicts and changing a task's base branch all work using the daemon and git on your machine. Integrations bring work in and handle things like opening the pull request on GitHub. They never replace information your machine already has.

Works with nothing connected Added when you connect a provider
Coding, plan/implement/verify Issues synced in from GitHub, GitLab or Jira
Diffs and merge checks from local git "Mergeable" and "behind" hints from GitHub
Resolving conflicts, changing base branch Opening the hosted PR
The board, queue, live session and @gpr mentions Tasks started from Slack

In the code, each provider lives in its own folder and plugs into a shared interface. We added GitHub first, then GitLab (which mostly copied the GitHub structure), then Jira and Slack. In August we found GitHub-specific code in places that were supposed to work for every provider, and we spent some time cleaning that up. It still isn't perfect.

This approach helped us in September. Atlassian removed an old Jira search endpoint, and our Jira calls started failing with 410 Gone. The agents kept working as normal. Only the syncing of new tickets stopped. We wrote the fix up as a Jira ticket, and GPR picked it up and opened the pull request.

Some weird issues we hit while building it

I want to share a few of these because building reliable software takes time, even when agents write a lot of the code. In most cases, finding the problem took much longer than fixing it.

The connection that was dead on one side.

A Windows machine had a connection that looked fine from the machine's side, but the relay could not reach it. The daemon thought it was online, so it never reconnected. Tasks sent to that machine just waited and nothing happened.

Split brain after a deploy.

When we deploy a new version of the backend, the old servers keep running for a short time before they shut down. The daemon's live connection stayed on one of those old servers, while its normal API calls went to the new ones. So one part of the backend thought the machine was connected, and another part thought it was not. We fixed both problems the same way: the backend now starts the reconnect itself, and the daemon checks its status regularly and reconnects if the two don't agree.

52 agents on one laptop.

In August one developer's machine had 52 agent processes running, each using about 300 MB of memory, a few minutes after the daemon started. Several Cursor agents had started at the same time and were all trying to use the same config file, so their startup failed. We retried, but the failed agents did not stop when we asked them to. We had also stopped tracking them, so nothing ever cleaned them up. Now we give an agent two seconds to stop, and if it is still running, we kill it.

Tasks marked done with no changes.

Sometimes an agent said it was done and wrote a good summary, but it had not changed any code. GPR still marked the task as successful. Now a task with no commits is marked as failed. In another case, a log line that should have explained why tasks were timing out printed nothing useful, because it used Go formatting in Python code. When we fixed the log, it showed us a second bug where one session could start two agents. Both stories are in AI can run the software factory. It can't turn the lights off yet.

Updates on Windows.

On some Windows machines gpr update did nothing, but the output looked the same as a successful update. Windows does not let you replace a program while it is running, but it does let you rename it. So now we rename the old program and put the new one in its place. Every change to the daemon has to work on macOS, Linux and Windows, and Windows caused most of these problems.

Duplicate Slack messages.

Slack sends us every @gpr mention twice, but only one of the two has the screenshot attached. We were reading the other one, so screenshots people attached to bug reports were lost.

You may have noticed that several of these problems came from the agents themselves, like the Cursor processes that would not stop. That happens because GPR runs coding agents from existing vendors instead of an agent we built ourselves. We chose that on purpose, early on, and the next section explains why.

Agent Client Protocol: why we never built our own coding agent

Almost every developer we talked to was already paying for a coding agent, whether that was a Cursor plan, a Claude subscription or ChatGPT for Codex. Asking them to pay for another agent just to try GPR would have been a barrier, and they would be paying twice for the same thing. That's also why I find it hard to justify paying for Devin on top of those plans.

What we didn't know was how to support all of these agents without building a separate integration for each one. Then I watched a video with Ben Brandt from Zed and people from JetBrains about the Agent Client Protocol (ACP). Without a shared protocol, every editor needs its own integration with every agent. With ACP, each editor and each agent only has to support the protocol once.

If Zed and JetBrains were building this standard for editors, a phone app could use it too.

Here is how GPR uses it. The daemon talks to the agent using ACP on your machine. The daemon also connects out to our relay over a WebSocket, and the relay sends the same messages on to your phone. The relay keeps a copy of the session, so if your phone goes offline it can catch up later, and it sends you a notification when the agent needs you. Because of this, the app doesn't have to read a terminal screen. It receives the plan, tool calls, diffs and the agent's slash commands as structured data, which is why one app works with every agent.

Cursor and Kiro support ACP directly, so they were the first agents we fully supported, in July. Claude Code and Codex need an adapter, and we added them in September. Today the daemon supports eight agents: Cursor, Claude Code, Codex, Kiro, OpenCode, Qwen, Cline and Copilot. Our ACP post explains the protocol in more detail, including how it differs from MCP.

We built GPR using Cursor, Claude Code, Kiro and OpenCode, and nobody on the team had to give up the agent they liked. 370 commits in our repo are co-authored by GPR. Each agent starts up, logs in and shuts down a little differently, and the daemon handles those differences. Switching agents is just a dropdown in the app.

What a phone is actually good for (the Claude Code mobile app question)

A lot of people search for a Claude Code mobile app. In practice, the things you need to do on a phone are fairly small: answer a question about the plan, approve a command, read a short diff, give feedback and merge. On Android you can also share a screenshot from your gallery directly into an agent session, and a lot of our own bug reports now start that way.

We spent a lot of time making these small things easy on a small screen. You can leave a comment on a diff by tapping or long-pressing the line, and the agent picks it up in the next round. Long agent sessions have a conversation minimap, so you can jump to the part you care about instead of scrolling through everything. The slash commands in the app are the ones your agent actually supports, and switching the model or the agent is a single dropdown. When you add reviewers to a pull request, you pick them by name and photo, because nobody remembers GitHub usernames on a phone. If the base branch has moved, you can update the pull request from the phone, and the agent fixes any conflicts in new commits so you can see exactly what it changed. And if your phone loses signal halfway through a session, it catches up from where it left off when it reconnects.

A phone is not the right place to review a big architectural change. When I need to compare big files side by side, still i prefer big screen. But this pushes to keep agent tasks small, and agents do better work on small tasks anyway. You can read more about how we review and merge from the phone. If you're wondering why a person still needs to approve the merge at all, the software factory post covers that in detail.

Still rough

The iOS app isn't in the App Store yet. Live preview of your running app on the phone is still being built. Diffs come from git on your machine, so if your laptop is asleep, you can't see the diff either. And when an agent's login expires, GPR can tell you that the agent failed to start, but someone still has to go to the machine and log in again.

Trying it

To try it, you need the Android app from Google Play (iOS app coming soon), a Mac, Linux or Windows machine, and one coding agent installed on it. GPR is free for projects, sessions, and the GitHub and Slack connections. You pay only for more workspaces or more tasks running at once.

curl -fsSL https://getpullrequest.com/install | bash
cd ~/code/your-repo
gpr setup # scan the QR code with the app
gpr discover # lists which agents are ready

Enter fullscreen mode Exit fullscreen mode

We'd suggest starting with a small task, something you've been putting off, so you can see how it works before giving it anything bigger.

Related posts

Top comments (0)