The protocol is the easy part. Nobody warns you about the rest.
I shipped a remote MCP server this month. Building the thing that speaks MCP took an afternoon — one route on a backend I already ran, and no new dependency. Everything after it took days, and almost none of it is in the protocol docs.
This is the whole path in order: the decisions that are cheap, the ones you only get to make once, and the point where most people should stop. Where a step cost me a full day, I've linked the write-up with the exact commands and error messages, so you can skip the part where I was confused.
First: are you sure you need a server?
The cheaper alternative is a skill — a file of instructions the user installs, which their agent reads and follows. No hosting, no security surface, and it carries far more nuance than a tool description ever will, because it arrives as prose in the agent's context rather than a one-paragraph summary.
The honest test is where the work has to happen. If everything your integration needs is already on the user's machine or the public web, and they can do it with their own credentials, write the skill.
You need a server when the work needs data or a service only you can run; when it has to act as the user, which means sign-in and permissions; when you need to change behaviour for everyone at once without asking anyone to update a file; or when you want to be findable in registries and directories.
The asymmetry to weigh: a skill is a file; a server is an authorization surface you own for as long as it's up. If a file would do, write the file.
They aren't exclusive, and the strongest setups use both with a clean division — the server owns the data and the operations, the skill owns the judgment.
Step 1: tool names are a one-way door
Directory scorecards reward tool names that read as a navigable tree — a dotted path like stories.search. My listing scored 98/100, and those were the missing two points.
Don't take them. OpenAI's function-name pattern is ^[a-zA-Z0-9_-]{1,64}$; dots are rejected outright, and Anthropic's is the same shape. A dotted tool name is refused by the exact clients the listing exists to reach. It's a criterion you should knowingly fail.
What you can do is pick a grouping convention up front and apply it without exception, so the structure stays legible with an underscore as the separator. Decide it at planning time and write it down — because the moment a registry has scanned you, a directory has listed you, and a client's app directory has accepted your submission, every published tool name is one you have to keep. Renaming later isn't a refactor, it's a migration across systems you don't control.
Step 2: build the server (genuinely the easy part)
An MCP server speaks JSON-RPC. A client asks it to introduce itself, asks for its tool list, then calls tools one at a time. For a server people reach over the internet, use the streamable HTTP transport — one endpoint handles everything.
Three things that are easy to skip and shouldn't be:
- Write tool descriptions for an agent, not a human. They're the entire basis on which a model decides whether to call you.
- Set the annotations honestly — read-only, destructive, open-world. Clients show these before approving a call.
-
Declare an
outputSchemaand returnstructuredContent, so consuming agents can rely on shape instead of parsing prose.
Then point the MCP Inspector at your own endpoint before any real client sees it. It's a test client you drive by hand, with no model in the loop — which matters, because when a real client doesn't call your tool you cannot tell from outside whether your description was vague, your schema was malformed, or the server errored. Those are three different bugs that look identical.
It also runs a schema portability check. Mine flagged that nullable fields written as one type with two entries are legal JSON Schema and silently mishandled by several clients — a defect none of my own tests were looking for, because my tests shared my assumptions.
Step 3: where it lives, and what it costs
An MCP server is a small, bursty HTTP service: idle most of the time, awake when somebody's agent calls. That fits serverless. The default instinct is a new repo, the official SDK and a container; I added a route to the backend I already ran. The entry point is a single branch in the existing request handler, and the endpoint is standard library plus the cloud SDK already in the runtime. Zero new dependencies.
Concrete numbers: marginal compute is pennies a month and free tiers swallow it. Listing in the official MCP Registry is free. Submitting a ChatGPT app is free. The only line item that isn't free is the Claude connector directory, which needs a Team organisation — about $40/month with a two-seat minimum.
The recurring cost isn't compute. It's that you now run an authorization server, so you own its security; that every client release can surface a new incompatibility; and that without request logging you'll have no idea who calls what.
Full walkthrough: An MCP server for the price of one more endpoint
Step 4: four problems that are all really sign-in
This is the part I'd most like to have known in advance. Identity doesn't arrive as one task. It arrives four times wearing a different mask, and each one reads as a small detour from what you were actually doing.
Acting as a person. The moment a tool writes anything, you aren't adding a password field — you're running an OAuth 2.1 authorization server: discovery documents, dynamic client registration, PKCE, a consent screen. Clients register themselves; there's nothing to pre-issue. This is the expensive one. → OAuth 2.1 on AWS Lambda, start to token
Deciding what needs no sign-in at all. Having built the gate, the easy mistake is putting everything behind it — including the list of what your server offers. Mine did, and a directory listed me with a name, a description, and no tools, because a registry mirror has nobody to sign in as. The fix was about fifteen lines, and the rule is: expose anonymously exactly what is already public elsewhere, and nothing else. → My MCP server hid its own tool list behind a login
Giving a stranger a way in. An app reviewer has to use your server. If your only sign-in is a social provider, there is nothing to hand them, and you'll build a sandboxed account under deadline. → App review wants test credentials
Having the right token yourself. Publishing under an organisation rather than your own name needs a token carrying a scope you probably aren't using, and the error won't tell you that.
Step 5: instrument it before anyone uses it
An MCP server is one URL. Every call arrives as POST /mcp, so the access log you get for free is a column of identical lines. Which tool ran, for whom, in what order — none of it is there.
One structured line per call, emitted in the dispatcher rather than in each handler, fixes it. Trim deliberately: log identifiers and lengths, never the content people send you.
I added this to a server almost nobody was using, which felt pointless. It immediately caught a client calling a method my server doesn't implement, getting an error, and silently falling back — five times over two days, with nothing visibly broken. None of that is visible in a generic access log or an APM dashboard: MCP-aware logging is something you have to write yourself, and this is the argument for writing it early. → Your MCP server is one URL
Step 6: you may already be done
Everything above was building. The server runs, it signs people in, it says what it can do, and you can see who calls it. That is a finished thing, and plenty of good servers stop exactly here.
What follows is distribution, and it's optional in a way none of the earlier steps were.
If your users already live in a terminal or an IDE, you're done now. Publish the endpoint and one install line in your README — claude mcp add --transport http <name> <url>, or the equivalent for their client — and they add it themselves in about ten seconds. No forms, no reviewer, no waiting on anybody. For an internal tool, a team server, or anything with a known audience, this isn't the lesser option. It's the right one.
Read on only if you want the other thing: to appear inside a client's app directory, where someone who has never heard of you can find you.
Step 7: the documents, and the one rule
Submitting means writing a pile of prose: a privacy policy, terms, a public documentation URL, a domain-verification file, tool-by-tool justifications, example prompts, and credentials a stranger can sign in with.
None of it is hard. All of it is slow. And one property matters more than any individual document: they must describe the server as it runs today.
I learned that by failing. My privacy policy was written when the connector was four read-only tools with no sign-in. By review time it was six tools behind OAuth, two of which write on your behalf. The policy still said "read-only tools" and "no account is required". The rejection was one sentence: the policy does not clearly disclose all data uses, and must reflect current tool inputs and outputs.
Entirely fair. I'd shipped four significant changes and updated none of the documents describing them. A reviewer reads your prose against your live endpoint — which is something you never do.
The second thing worth knowing: these documents are fetched at review time, not uploaded. A fix needs no resubmission of anything but the form; a drift needs no negligence, only time.
Step 8: being listed is not being found
Getting into the official registry feels like the finish line. It's a database entry.
Eight days after publishing there, searching my server's name in every browsable directory returned nothing. But "each one wants its own submission" turned out to be only half right, and the half that's right has a pattern underneath it.
Scanners want a submission; mirrors don't. A directory that connects to your live endpoint, or is human-reviewed, needs you to submit. A directory that reads the official registry needs nothing — the registry already holds what it displays. One tell that you're looking at a mirror: its listing shows your description and URL but no tools. It never connected. → Submitting an MCP server to the directories
Step 9: the only number that tells you whether any of this worked
Every step above has a satisfying finish line. That's exactly what makes them easy to mistake for the goal.
The real question isn't an engineering question, and no directory will answer it: has anyone who is not you ever connected?
Here's what that looked like for my server, from its own request log over 24 hours: 93 protocol calls — connect, list the tools, list the resources, repeat — and exactly one actual tool call, which was mine. Two distinct users, both of them me.
Connected and used are not the same number, and until you log it you cannot tell them apart.
What I'd do differently
- Decide tool names in the planning doc, not in the code.
- Add request logging in the first deploy, not the fifth.
- Re-read every public document against the live tool list before submitting, not when I wrote them.
- Ask "is a directory listing what I actually want" before doing the work that a listing requires.
Run this instead of re-typing it
Every write-up linked above carries a machine-readable "Reproduce this" contract — prerequisites, the inputs it needs from you, ordered steps, and a verification section your agent runs to prove it actually worked. You point your agent at it and it does the work in your project.
They're collected in build order as a path: Build and ship your own MCP server. It's honest about its own state — one step is currently marked a proposal rather than a result, because I've resubmitted after that rejection and the verdict isn't in yet.
If you're weighing whether an MCP server is the right move at all, or you're stuck on the OAuth step, I do consultations — details at worklore.dev.


Top comments (3)
We hit a close cousin of your nullable-field bug on a finance MCP server, and it came from the same place: tests that shared our assumptions. The suite called each tool handler directly and passed, while clients that validate results against the schemas they received from tools/list rejected responses carrying fields the schema didn't declare. The durable fix was to make the harness behave like a real client, in order: initialize, list tools, then validate every result against the outputSchema it just listed. That one change would also have caught your tool list hiding behind the login, since the harness can't get past step two. And your point about dotted names being a criterion to knowingly fail is a good one: scorecards reward the tree, but the clients are what you're building for.
Ran your harness against my server the same evening. It found a live one on the first pass.
get_story declared an outputSchema and was answering every call with isError and "Object of type Decimal is not JSON serializable" — a DynamoDB Decimal in an x-ray finding's line number reaching json.dumps. It had been shipping that way. None of my 79 tests saw it, because nothing called through tools/call. Exactly your diagnosis: the suite agreed with me. I had a coercion helper that handled the two numeric fields I'd thought about, and anything nested deeper went straight through.
Two notes from running it.
Your step two is a bigger deal than you made it. An endpoint that won't answer tools/list without a token can't be conformance-tested by anyone — and that's the same property that had a registry-mirroring directory listing my server with a description and zero tools. I'd found those two symptoms separately and never connected them. One cause.
Watch the fake. My first run of your harness passed before the fix, because my fake DynamoDB hands back plain ints. Seeding real Decimals is what made it fail. A fake that's more forgiving than production will happily bless broken code.
The harness is in the repo now, in your order: initialize, tools/list anonymously, then validate every result against the schema it was just handed. Thanks — that was a better bug report than most bug reports.
That's a great find, and the fake is the part I'd underline for anyone reading. Our stock-split bug had the same shape: the test fixtures agreed with the code's convention, the real data used the other one, and everything passed. What helped most was building fixtures from recorded real responses instead of hand-written ones, so the fake inherits production's types (your Decimals) rather than our assumptions about them. One small guard to add if the harness doesn't have it yet: assert that happy-path calls don't come back with isError. An error result usually carries no structuredContent, so a schema check can wave it through without anything having worked. And agreed on step two, one cause for both symptoms is a much better story than two unrelated bugs.