Coding agents changed what a working session looks like. You describe the work, and it happens without you. You can walk away for twenty minutes and come back to a finished change.
The tools around them haven't caught up. Almost all of them still assume you're sitting at the machine the agent runs on, watching it.
I often wasn't. Some days I was nowhere near my PC, and on others I was near it but couldn't do any real work on it. Long-running agents looked like the obvious answer: start the work, walk away, let it run. What I wanted was a way to hand work to an agent while my PC was off.
That raised two questions right away. If I'm away for hours, how do I keep the agent's output good enough to be worth keeping? And when I get back, how do I review what it did without digging through a terminal?
Nautilus is my answer to both. I built it over the five days of the holiday.
Ask on the phone, review on the PC
A phone is a good place to ask for work and a bad place to review it. A small screen is fine for a sentence of instructions and a glance at a live page. It's a poor place to read a 300-line diff.
So Nautilus splits the job.
On the phone, an installable web app lets you pick a project, send a prompt and watch the agent's turn stream in: every file it reads, every command it runs, every step of its reasoning. When it wants to run a command, it asks, and you approve or deny it from the phone. A preview tab shows the project's dev server live, so a UI change is something you can see, not just read about. Then you lock the screen, and it tells you when it's done.
At your desk, a desktop app shows you what changed as a real diff, file by file. Nothing reaches your files until you pull. If you changed the same file while you were away, you choose a side per file, and the last pull can be undone.
Your .env files never leave the PC. If the preview needs some of their values, you pick them in the desktop app. It keeps them in the OS keychain and sends them to the runner for the dev server alone. The agent runs in a sandbox and only ever sees their names.
The PC doesn't need to be on while the agent works. It only has to be reachable when you sync.
The constraint: zero dollars
I set one rule before writing anything. It had to run on infrastructure that costs nothing. No credit card, no domain, no Docker, no paid server.
That ruled out the easy version: rent a small server, point a domain at it, done. What was left was a free Lightning AI Studio, a CPU machine in the cloud with an SSH endpoint and ports you can expose over HTTPS.
Free comes with terms, and three of them shaped most of the design.
The machine restarts every four hours. Not gracefully, and not on your schedule. Processes die, installed dependencies disappear, and a file written just before the restart can come back older or cut short.
You don't control the front door. Traffic reaches the machine through the provider's proxy, and that proxy has its own ideas about cookies and headers.
And the PC is a laptop on a home network. It can't accept connections from the internet, and I didn't want it to.
Each of those could have been a reason to pay for something. Treated as requirements instead, they pushed the system to be more careful than it would have been on a server I trusted.
Planning before code
Before any code, I wrote down every feature end to end: what the phone shows, what the runner does, what the PC does, and what happens at each step when something fails halfway. The design I settled on then is the one that shipped.
Three machines, one job each. The runner, in the cloud, runs the agent, the project's dev server, the API and the phone app, all behind a single gateway. The phone steers. The PC keeps the real project and decides what gets into it.
Most of the five days went into one part of that plan: the failure cases.
Three problems that took the most thought
A machine that forgets everything every four hours
The easy mistake is to treat a restart as a hiccup: bring the processes back up and carry on. That works until a restart catches something mid-write and the runner starts serving a project that's quietly older than it claims to be.
So Nautilus treats every boot as a full recycle. The boot script reinstalls what disappeared and takes a lock so two boots can't race each other. Then the server checks its own state before it serves anything. Each project's files are compared with its last checkpoint, and a project that doesn't match is marked unhealthy and left alone. Refusing to show a project is better than showing the wrong version of it.
A turn the agent was in the middle of is never reported as finished. It comes back as interrupted, and the phone offers to retry it. The phone reconnects on its own and picks up the event stream from the last event it saw.
Keeping your Git out of it
The agent needs history: a checkpoint after every turn, three-way merges, undo. The obvious place for that is your project's own Git repository, and that's exactly where it shouldn't go. Your branches, hooks and remotes are yours.
So each side keeps a second, separate repository that tracks the same folder. The agent's history lives there. Sync moves only what's new, every merge is tried in full before a single file is written, and an apply that gets interrupted rolls back on the next start. Your .git directory is never read or written.
It's the part I'm proudest of, and it gets its own article next week.
A front door that works against you
The provider's proxy rewrites every cookie so the browser sends it cross-site, and marks every response as open to any origin. The app's own cookie settings stop meaning anything, and any page on the internet could send requests that carry your session.
You can't fix the proxy, so the app assumes it's there. Every write needs a matching origin and a custom header that a cross-site form can't set. The proxy's permissive headers are stripped on the way out. Admin routes don't exist on the public side at all. And one-time preview links only work on a button press, so a link-preview bot that fetches a shared URL can't use one up.
What I didn't expect
Two things were harder than I planned for.
The first was releasing for three operating systems. Building the desktop app is one command. Shipping it is not. Each platform needs its own system libraries, its own installer format and its own build machine. The app carries a small program that handles syncing on the PC, and that program has to be built as a standalone binary for each platform, so nobody needs Node installed. One Linux package format compressed that runtime slowly enough to stall the build, so I dropped it. And because the builds aren't signed, macOS and Windows both warn on first launch. That's now a paragraph in the README instead of a problem I could solve for free.
The second was thinking in services instead of one app. Most of what I'd built before was a Next.js app with its own server: one process, one thing to deploy, and if it's down, everything is. Nautilus is a gateway, an API, the phone app, the agent and each project's dev server, all separate processes on one machine, plus a desktop app and a sync agent on another. Sync alone is a conversation between three machines with a review step in the middle. The hard part was never any one service. It was the edges between them, and making sure one failing piece costs you that piece and not the rest. When a single project is unhealthy, the runner reports itself as degraded and keeps serving everything else.
Try it
Nautilus is open source under GPL-3.0. Setting up a runner takes one command, and the desktop app has installers for Linux, macOS and Windows. It's maintained on demand, so if something breaks, open an issue and I'll fix it.
Have you handed long-running work to an agent yet? What made you trust the result when you came back?
Nautilus is on GitHub: https://github.com/itamarhanan/nautilus


Top comments (0)