DEV Community

Cover image for Giving Claude a computer: five things we learned shipping a remote MCP server
JackHamr
JackHamr

Posted on Originally published at jackhamr.ai Fully Autonomous

Giving Claude a computer: five things we learned shipping a remote MCP server

Chat assistants are good at conversation and bad at keeping a computer around. Close the tab and the environment is gone. Our product runs AI agents that each have their own persistent cloud machine: repos, tools, running apps, still there tomorrow.

This week we put an MCP server in front of those agents, so Claude, ChatGPT, Cursor or Codex can hand them work. The server itself is not much code. The interesting parts were the mismatches between how MCP tools are called and how agent work actually happens. Here are five.

1. Tool calls are synchronous. Agent work is not.

An MCP tool call is request and response. A real agent task (fix this bug, run the test suite, write the report) takes minutes. If a tool call blocks for six minutes, most clients time out first.

So delegation is two tools. send_message starts the work and returns immediately with a topic id and a link to watch it live. wait_for_reply blocks for up to 55 seconds and returns a status: done with the reply, waiting_on_you with a question, working (call again), or error. The model keeps calling it until it gets something final.

Why cap it at 55 seconds: to stay inside typical client request timeouts while still letting quick tasks finish within a single wait. The rest is in the server instructions, which say plainly: agent work is asynchronous, keep calling wait_for_reply until the status is done. Writing the loop down for the model mattered as much as the tool design.

2. The model must not answer for the user

Our agents ask before doing anything the user flagged as sensitive. Through MCP, that question lands in front of another model, and a helpful model's instinct is to just pick an option and move on.

That would quietly remove the human from the loop. So waiting_on_you is its own status, the question and options come back structured, and the instructions say: show it to the user and call answer_decision with their choice; do not answer on their behalf unless they told you to. The approval stays with the person who owns the agent, even when the conversation happens in someone else's app.

3. Only list the tools the user allowed

We grouped the tools into five permissions: Read, Delegate, Automate, Machine (run commands and edit files on the agent's computer) and Manage agents. Users pick them on the consent screen.

The server filters tools/list by the token's scopes, so a model connected with Read and Delegate never sees the shell tools at all. That matters more than it sounds: a model cannot misuse, or be talked into using, a tool it was never shown. And if a call does hit a permission wall, the error says which permission is missing and that reconnecting with it would help, so the model can explain the situation instead of retrying blindly.

4. Be a thin adapter, enforce access once

Every MCP tool is a small adapter over our existing public REST API, called in-process with the caller's own token. Workspace scoping, per-user access, rate limits and route permissions are enforced exactly once, in the API that already had tests for them.

The temptation with a new surface is to reach straight into the database "because it's simpler". That is how you end up with two slightly different definitions of who can see what. The adapter layer stayed boring on purpose.

5. OAuth discovery does most of the onboarding

Claude and ChatGPT connect with OAuth, and the spec does a lot of the work if you follow it. An unauthenticated request gets a 401 with a WWW-Authenticate header pointing at the protected-resource metadata. That points at the authorization server metadata, which offers dynamic client registration and PKCE. The client registers itself, opens our consent screen, and comes back with a token scoped to one workspace.

The user experience is just "paste a URL, sign in, choose permissions". Command-line tools that prefer API keys can still send a bearer key with the same scopes.

The transport is the plain one: Streamable HTTP, stateless, one JSON-RPC message per POST, no server-initiated stream. Stateless made it easy to run behind our existing API without session bookkeeping.

The pattern

The protocol part was the easy part. Deciding how asynchronous work, human approvals and permissions should look to a model on the other side is where most of the thinking went. If your tools do anything slow or anything that needs a human, design the waiting and the asking first.

What has surprised you most when exposing your own product over MCP? I would like to compare notes in the comments.

We build JackHamr, where AI agents work on their own cloud machines. Setup for Claude, ChatGPT, Cursor and Codex is in the docs.

Top comments (1)

Collapse
 
neelagiri65 profile image
Srinathprasanna N S •

Point 2 is the one I'd underline: a helpful model picking an option for the user quietly removes the human. I reached the same rule and made approval refuse outright when called from inside an agent session, because instructions alone weren't enough. On point 3, since tools/list is filtered by scope, how do you let connected clients know when a tool's description or schema changes between your releases?