I run my business alone, and a lot of my day is browser work. Checking payouts in Stripe. Pulling rankings from Search Console. Filling in forms on directories and launch sites. Posting to a dozen places.
These days I hand most of that to Claude Code. The hard part was never the AI. It was the browser: how do you let an AI agent use a website you're already logged in to, without it getting in your way or getting you flagged?
You've probably seen demos like "I had an AI agent order Uber Eats through the browser." To me, that misses the point. I can order food myself, faster. One task done by an agent instead of me saves almost nothing.
What matters is how many jobs can run in parallel while I do something else. That's the only reason I want an agent in the browser at all.
I tried the usual ways first. Each one broke in a way I didn't expect.
1. Claude's Chrome extension
The easy answer is Claude's own extension for Chrome. My logins are already there, so the agent can start right away.
But with the extension, I give instructions browser by browser. That's fine for one task. It's not what I wanted for running several jobs side by side.
I wanted to control everything from the terminal. One place where I ask for the work, and the agents do it in whichever services they need, at the same time.
Two more things bothered me:
・I was handing over the whole browser. Every tab, every account, every service. There's no way to say "only Stripe".
・It once connected to a different browser than the one I thought, then asked me to sign in to a service I was already signed in to. Everything looked fine on my side, so it took a while to figure out why.
2. Browser automation tools with a separate browser
This is how most AI agent browser automation works: start Chrome with special settings and connect to it.
What I ran into:
・It's a fresh browser, so there are no logins. Sign in again, type the two-factor code again, for every service.
・Sites notice. One startup directory returned "Bot access denied" when I submitted through an automation tool. After that, even my own clicks on the Submit button stopped going through.
3. Letting it watch the screen
The agent takes a screenshot, picks a spot, clicks and types. It can drive anything.
But it drives the screen I'm looking at. While it works, my mouse and keyboard belong to it, and there is only one screen, so there's no second task running beside the first.
That's the part that bothered me most. The whole point of handing work to an agent is that I keep working. If the agent takes my focus while it works, I'm just sitting and watching. From an efficiency point of view, that defeats the purpose.
What I actually wanted
After all that, my list was short:
・I sign in myself. The agent never sees a password.
・I give the agent one service, not my whole browser.
・Nothing jumps in front of what I'm doing.
・Several jobs can run at the same time.
What I built
So I built Kagemusha. It turns any website into its own Mac app. A Stripe app, a Search Console app, a GitHub app. Each one is a separate browser with its own sign-in.
I sign in to each app once, the normal way, with my password manager or a passkey. Then, app by app, I decide whether an AI agent is allowed to operate it. A new app starts with that switched off.
When the agent works, it works inside that one app, in the background. My screen stays mine. I can show the window if I want to watch, or keep it hidden.
Getting this part right took a lot of trial and error. Windows kept coming to the front while I was typing. Even bringing one forward for a moment took my keyboard for almost three seconds, and my keystrokes went into the wrong place. I kept going until the agent could open, read and fill in pages without ever stepping in front of what I was doing.
Because each app is its own browser, I can run several at once. When I checked the other day, I had 24 of these apps open at once on my MacBook Pro (M4 Max, 128 GB of memory), using about 23 GB in total, while the same Mac was running around 20 local development servers. That was one observation on one machine, not a stress test.
What changed in practice: I stopped babysitting the browser. I ask for last week's Search Console drops, this month's Stripe refunds and a draft post, and they come back while I keep working on something else.
Where to find it
I built it for my own work first. If you want to look at it, it's here:
I'd like to hear how others are handling this. If you let an agent use your browser, which of these walls did you hit?
Top comments (0)