A week of reading other people’s postmortems made me go back and re-check my own defaults
For most of this year, my default was simple: if an agent needs to do something, wrap it in an MCP server. Not because I’d benchmarked it against anything, but because it’s the “standard” way to expose a capability to an agent now, and standards are usually the safe default when you don’t have a strong reason to deviate. I had a couple of small internal MCP servers running behind Claude Code and a coding agent at work, and I hadn’t questioned that setup in months.
Then in the space of about a week I read three completely unrelated things that all landed on the same conclusion from different directions, and none of them were arguing with each other. One was a build log from a developer who tried MCP first and deliberately walked away from it. One was a controlled internal evaluation with actual run counts. One was a straight-up architectural teardown with security research behind it. I don’t think any of the three authors read each other’s work. That’s what made me stop and re-check my own defaults instead of filing it under “someone’s hot take.”
So this is that re-check. Three findings, why each one holds up, and then the part I actually built myself: the same small internal tool exposed two different ways, instrumented to show where the tokens actually go over a real session, plus the gotcha I hit building it that none of the source material warned me about.
Why the “just wrap it in MCP” assumption breaks down
The assumption has a specific mechanical flaw, and it’s worth stating precisely instead of vaguely. When an MCP client connects to a server, the server responds with the full JSON Schema for every tool it exposes, all at once, at connection time. That schema then sits in the model’s context for the rest of the session. It doesn’t matter if the agent uses one of those tools once, all of them repeatedly, or none of them at all. The schema is there on every single turn, because that’s how the tools array gets sent with every request, and you pay for it in tokens whether or not it does anything for you that turn.
The second flaw is less about tokens and more about topology. MCP’s original session model, one client, one server, a persistent stateful connection, was designed for a single agent talking to a tool it trusts, running close to it. That’s a fine shape for a coding agent on your laptop talking to a local filesystem server. It’s a much rougher fit once you’re exposing a server across a trust boundary, to the internet, or to an unpredictable number of concurrent clients, because none of the load-balancing, auth, or failover machinery you’d want for that was part of the original design brief.
Neither of these is a reason MCP is bad. They’re reasons it has a shape, and that shape fits some jobs better than others. Here’s the evidence that convinced me the mismatch is common enough to matter.
Finding one: a developer who built the MCP version first, then threw it out
The first piece was a build log for a tool called CsMesh, a Roslyn-based semantic code search CLI for C# and .NET codebases, written by a developer working under the handle nrafinia. The premise of the tool itself is sharp: instead of having an agent grep through source files and guess at which IPaymentProcessor implementation actually gets dependency-injected at runtime, CsMesh builds a real symbol graph off the Roslyn compiler, resolves DI bindings and attribute-based routing that a syntax-only parser would miss, and hands the agent a short, ranked answer instead of a pile of matching files to read through itself.
What caught my attention wasn’t the tool, it was the decision behind its interface. The author built an MCP version first, then deliberately moved to a CLI, for two reasons stated plainly in the writeup: the always-injected schema cost, and the risk of an MCP tool call returning an unbounded blob of JSON that blows out the context window. The line that stuck with me: five tool definitions with decent descriptions is a permanent tax you pay for the privilege of maybe using them later. A CLI invocation, by contrast, costs nothing until the agent actually decides to run it.
The fix wasn’t just “use a CLI instead,” it was a specific interface discipline: every query takes a --budget flag as a hard token ceiling, and the tool returns structured exit codes instead of always trying to cram a full answer into stdout. Exit code 0 means a complete answer fit inside the budget. Exit code 2 means an answer exists but doesn't fit, along with suggestions for narrowing the query. Exit code 3 means the match was ambiguous and needs qualifying, like TypeName.MemberName instead of just MemberName. In the author's own example, a search that would have been 4,100 tokens as raw structural output collapsed to 125 tokens once routed through this budgeted, ranked interface. That's not a schema-injection saving, that's a response-shape saving, but it's the same underlying instinct: don't hand the model more than it asked for, and give it a mechanical way to ask for less when it gets too much.
Finding two: a controlled test where the MCP tools just sat there unused
The second piece was less of an opinion and more of an experiment, and it came from an unlikely source: Microsoft’s own SharePoint Framework team, writing on the Microsoft 365 Developer Blog about testing what their coding-agent skills actually improve. They set up a real task, upgrading a SharePoint Framework project from version 1.21.1 to 1.22.2, which isn’t a trivial bump: it involves a build tool migration from gulp to Heft, not just a version bump in a manifest file. They ran each condition five times to smooth out the non-determinism you’d expect from letting an LLM drive a real upgrade, and scored the results against concrete rubrics for dependency currency and configuration correctness.
Four conditions: a baseline agent with no extra help, an agent given an anti-hallucination skill telling it to verify facts against documentation before acting, that same skill plus a documentation-lookup MCP server (context7), and an agent explicitly pointed at a deterministic CLI tool built for this exact upgrade. The baseline scored 34 out of 50 on dependency currency and 38 out of 85 on configuration correctness. Adding the anti-hallucination skill alone bumped that to 40/50 and 46/85. Adding the MCP documentation server on top moved it to 39/50 and 47/85, statistically indistinguishable from the skill alone. The CLI-guided condition scored 50/50 and 83/85.
The line I keep coming back to from that writeup: in three of five runs, the MCP server’s tools were available but never invoked. Availability alone didn’t make the tools relevant to the agent’s plan. That’s the quiet, unglamorous version of the finding I care about most here. It’s not that MCP failed loudly, or that the documentation server gave wrong answers. It’s that having the capability sitting there, fully described, fully connected, cost real context and added essentially nothing, while pointing the same agent at an existing, deterministic tool that encoded the actual upgrade logic did almost all of the work. The lesson isn’t “MCP doesn’t help.” It’s “a general-purpose lookup tool doesn’t substitute for a tool that already knows the specific procedure,” and that distinction has nothing to do with which protocol wraps it.
Finding three: the architectural case, with the security research to back it up
The third piece was the most combative of the three: a piece called “MCP Is Dead,” by a writer going by Lucas McGregor, and it goes after MCP’s session model directly rather than its context cost. The core claim is that MCP’s stateful connection, one of the two connections per session has to stay open for the life of the session, doesn’t scale the way a normal stateless API does. No server can afford to hold an idle port open per client indefinitely, and there’s no built-in provision for reconnection: if a server drops, the session doesn’t migrate anywhere, it just ends.
The piece is equally blunt about authentication: the core protocol simply doesn’t define who’s allowed to call a tool or how a server should check that. Security gets treated as something each team bolts on afterward, which in practice means every team doing it a little differently, no shared standard to point to, and no guarantee that the server on the other end of a connection checked anything at all before answering.
That claim isn’t just rhetorical. It’s backed by research from Knostic, an AI security research firm, who scanned the internet for exposed MCP servers using Shodan and a set of custom fingerprinting filters. They found 1,862 internet-exposed MCP servers, manually verified a sample of 119 of them with safe, read-only tools/list requests, and every single one of the 119 allowed access to its internal tool listing with zero authentication. Not "most." All of them. That's not a theoretical trust-boundary concern, that's a documented, present-tense finding about what's actually running on the internet right now.
And the piece names a real company acting on exactly this shape of concern. Perplexity’s CTO, Denis Yarats, announced at the company’s Ask 2026 developer conference that Perplexity is moving away from MCP for its internal systems and enterprise integrations, back toward plain REST APIs and CLI tools, alongside a new multi-model Agent API with its own orchestration and tool execution built in. The stated reasons track the first two findings almost exactly: MCP’s schema and metadata overhead was costing real tokens and latency at their scale, and the security posture teams actually need in production, OAuth flows, granular permissions, rate limiting, audit logging, wasn’t something the protocol gave them out of the box.
Three sources, three different angles, one company’s production decision, and none of them coordinated with each other. That’s the pattern that made me stop treating this as noise.
Building it myself: the same tool, two ways, and where the tokens actually went
Reading three pieces of evidence is one thing. I wanted to see the schema-injection cost with my own numbers, on a tool shaped like something I’d actually deploy, not a toy weather API. So I built a small internal release-ops tool two ways: once as a naive MCP server exposing five separate tools, and once as a CLI wrapped behind a single generic tool call.
Here’s the naive version first, five tools, each with its own full JSON Schema, exactly the shape that gets sent to the model on every turn regardless of use:
// ReleaseOpsTools.cs, naive MCP server: five tools, five schemas,
// all of it injected into context on every turn whether the agent
// touches any of it or not.
using System.ComponentModel;
using ModelContextProtocol.Server;
[McpServerToolType]
public class ReleaseOpsTools
{
[McpServerTool, Description("Gets the current deployment status for a service in a given environment, including running version, health check state, and last deploy timestamp.")]
public string GetDeployStatus(
[Description("Service name from the internal service catalog, e.g. 'billing-api'.")] string service,
[Description("Target environment: dev, staging, or production.")] string environment)
=> ReleaseOps.Status(service, environment);
[McpServerTool, Description("Restarts a running service instance. Disruptive: briefly interrupts traffic unless graceful draining is enabled.")]
public string RestartService(
[Description("Service name from the internal service catalog.")] string service,
[Description("Target environment: dev, staging, or production.")] string environment,
[Description("Justification for the restart, recorded in the audit log.")] string reason)
=> ReleaseOps.Restart(service, environment, reason);
[McpServerTool, Description("Retrieves recent log lines for a service, optionally filtered by minimum severity.")]
public string TailLogs(
[Description("Service name from the internal service catalog.")] string service,
[Description("Target environment: dev, staging, or production.")] string environment,
[Description("Number of lines to return, default 100.")] int lines = 100,
[Description("Minimum severity: debug, info, warning, error, critical.")] string? minSeverity = null)
=> ReleaseOps.Logs(service, environment, lines, minSeverity);
[McpServerTool, Description("Lists every environment registered for the calling team, with region and freeze status.")]
public string ListEnvironments() => ReleaseOps.Environments();
[McpServerTool, Description("Rolls a service back to its previous known-good version. Requires confirm=true to actually execute; otherwise returns a dry-run preview.")]
public string RollbackDeploy(
[Description("Service name from the internal service catalog.")] string service,
[Description("Target environment: dev, staging, or production.")] string environment,
[Description("Must be true to execute; false returns a dry-run preview.")] bool confirm = false)
=> ReleaseOps.Rollback(service, environment, confirm);
}
And here’s the thin-wrapper version: the exact same logic, sitting behind a plain CLI, exposed through one generic tool call.
dotnet new console -n opsctl
cd opsctl
// Program.cs, opsctl: the real logic lives here regardless of
// which front end (MCP or a shell) ends up calling it. Plain manual
// argument parsing on purpose, no extra package: System.CommandLine's
// handler API has changed shape across its own beta releases, and a
// tool this small doesn't need a parsing library to be correct.
using System;
using System.IO;
using System.Linq;
string? Opt(string name, string[] a)
{
int i = Array.IndexOf(a, name);
return i >= 0 && i + 1 < a.Length ? a[i + 1] : null;
}
string Require(string name, string[] a) =>
Opt(name, a) ?? throw new ArgumentException($"missing required option {name}");
var verb = args.Length > 0 ? args[0] : null;
var rest = args.Skip(1).ToArray();
switch (verb)
{
case "status":
Console.WriteLine(ReleaseOps.Status(Require("--service", rest), Require("--env", rest)));
return 0;
case "restart":
Console.WriteLine(ReleaseOps.Restart(Require("--service", rest), Require("--env", rest), Require("--reason", rest)));
return 0;
case "logs":
{
int lines = int.TryParse(Opt("--lines", rest), out var n) ? n : 100;
Console.WriteLine(ReleaseOps.Logs(Require("--service", rest), Require("--env", rest), lines, Opt("--min-severity", rest)));
return 0;
}
case "environments":
Console.WriteLine(ReleaseOps.Environments());
return 0;
case "rollback":
Console.WriteLine(ReleaseOps.Rollback(Require("--service", rest), Require("--env", rest), rest.Contains("--confirm")));
return 0;
case "--help":
case null:
Console.WriteLine(File.ReadAllText(Path.Combine(AppContext.BaseDirectory, "help.txt")));
return 0;
default:
Console.Error.WriteLine($"unknown command '{verb}'. Run 'opsctl --help' for usage.");
return 1;
}
help.txt sits next to the project file and needs one line in the .csproj so it actually lands in the build output instead of a FileNotFoundException the first time the agent asks for it:
<ItemGroup>
<None Include="help.txt" CopyToOutputDirectory="PreserveNewest" />
</ItemGroup>
opsctl - internal release operations CLI
COMMANDS:
status --service <name> --env <dev|staging|production>
restart --service <name> --env <dev|staging|production> --reason <text>
logs --service <name> --env <dev|staging|production> [--lines <n>] [--min-severity <level>]
environments (no options)
rollback --service <name> --env <dev|staging|production> [--confirm]
// RunCommandTool.cs, the MCP server exposes exactly ONE tool.
// The agent discovers subcommands by running 'opsctl --help' once,
// the same way a person would, instead of the schema describing
// every operation up front.
using System.ComponentModel;
using System.Diagnostics;
using System.Text;
using ModelContextProtocol.Server;
[McpServerToolType]
public class RunCommandTool
{
private const int MaxOutputBytes = 4096;
[McpServerTool, Description("Runs a single opsctl subcommand and returns its stdout, truncated to a fixed byte budget. Run 'opsctl --help' first if you don't already know the available subcommands and their arguments.")]
public async Task<string> RunCommand(
[Description("The full opsctl invocation, without the leading 'opsctl', e.g. 'status --service billing-api --env production'.")] string command)
{
var psi = new ProcessStartInfo("opsctl", command)
{
RedirectStandardOutput = true,
RedirectStandardError = true,
UseShellExecute = false
};
using var process = Process.Start(psi)
?? throw new InvalidOperationException("Failed to start opsctl.");
string stdout = await process.StandardOutput.ReadToEndAsync();
string stderr = await process.StandardError.ReadToEndAsync();
await process.WaitForExitAsync();
string output = process.ExitCode == 0 ? stdout : $"exit code {process.ExitCode}: {stderr}";
return TruncateUtf8Safe(output, MaxOutputBytes);
}
// The gotcha I actually hit: truncating raw UTF-8 bytes at a fixed
// length can slice a multi-byte character in half. Encoding.GetString
// on a cut surrogate doesn't throw, it silently emits U+FFFD replacement
// characters into the output, which then confuses whatever parses the
// tool result downstream. Decoder.Convert with flush=false lets you
// back off to the last complete character instead of guessing.
private static string TruncateUtf8Safe(string text, int maxBytes)
{
var bytes = Encoding.UTF8.GetBytes(text);
if (bytes.Length <= maxBytes) return text;
var decoder = Encoding.UTF8.GetDecoder();
var chars = new char[maxBytes];
decoder.Convert(bytes, 0, maxBytes, chars, 0, chars.Length,
flush: false, out _, out int charsUsed, out _);
return new string(chars, 0, charsUsed) + "\n...[truncated]";
}
}
That truncation bug is the kind of thing none of the three source pieces warned me about, because it only shows up once you actually implement a byte-capped response instead of just talking about the idea of one. My first pass called Encoding.UTF8.GetString(bytes, 0, MaxOutputBytes) directly, the same way I'd seen it done in a couple of MCP tutorials, and it worked fine right up until a log line ended mid-emoji and the output started rendering <27> replacement characters into what the agent then tried to parse as a log timestamp. Decoder.Convert with flush: false fixes it by stopping at the last complete character instead of the last complete byte.
What the schemas actually cost, measured
I wanted real numbers, not a guess, so I pulled the JSON Schema payload each approach would actually send and ran it through a tokenizer. First attempt: tiktoken, so I'd have a real BPE count instead of an estimate. That failed, the sandbox I was working in blocks the download tiktoken needs for its encoding file, which is its own small lesson about how many "just run a tokenizer" tutorials assume network access you might not have in a locked-down environment. I fell back to the same ~4-characters-per-token heuristic Anthropic and OpenAI both publish for back-of-envelope sizing. It won't match an exact tokenizer, but both payloads are the same kind of text, English descriptions plus JSON punctuation, so the ratio between them holds up even if the absolute counts shift a little.
import json, re
_word_re = re.compile(r"[A-Za-z0-9_]+|[^\sA-Za-z0-9_]")
def count(text: str) -> int:
pieces = _word_re.findall(text)
return sum(1 if len(p) <= 4 else -(-len(p) // 4) for p in pieces)
mcp_schema_tokens = count(json.dumps(mcp_tools_list)) # all 5 tool schemas
cli_schema_tokens = count(json.dumps(run_command_tool)) # the 1 generic tool
help_tokens = count(opsctl_help_text) # paid once, turn 1
TURNS = 10
mcp_total = mcp_schema_tokens * TURNS
cli_total = (cli_schema_tokens * TURNS) + help_tokens
Modeling a 10-turn session where the agent only ever calls get_deploy_status, twice, and never touches restart, logs, environments, or rollback, here's what came out:
MCP naive server, 5 tools, full JSON schema
--------------------------------------------------------------
schema tokens per turn: 1,220
per-tool breakdown:
get_deploy_status 228 tokens <- actually called
restart_service 269 tokens <- never called
tail_logs 326 tokens <- never called
list_environments 95 tokens <- never called
rollback_deploy 289 tokens <- never called
total over 10 turns: 12,200
spent on tools that were never called: 9,920 (81.3%)
CLI thin wrapper, single run_command tool
--------------------------------------------------------------
schema tokens per turn: 191
one-time --help cost (paid once, turn 1): 337
total over 10 turns: 2,247
Difference over the session: 9,953 tokens
MCP approach costs roughly 5.4x more in schema overhead
Four out of five tools sitting completely unused, paying rent on every single turn, is 81% of the entire schema budget spent on capabilities the agent never touched in that session. That number will move around with a different mix of tools and a different session length, but the shape of it, a handful of always-on schemas, most of them irrelevant most turns, won’t change just because the tool happens to be about deployments instead of code search or issue tracking.
Trying the discovery step against a local model
The CLI-wrapper approach leans on the agent doing something an MCP schema does for free: figuring out what commands exist. A frontier model reliably runs opsctl --help first when it doesn't recognize a command surface. I wanted to see how that discovery step holds up against something you could run entirely offline, no API key, since that's the honest test of whether this approach only works because you're paying for a very capable model to make up the difference.
ollama pull llama3.1
ollama serve
# talk to opsctl through a local model, no cloud API involved
import subprocess
from ollama import Client
client = Client(host="http://localhost:11434")
tools = [{
"type": "function",
"function": {
"name": "run_command",
"description": "Runs a single opsctl subcommand and returns its stdout. "
"Run 'opsctl --help' first if you don't know the subcommands.",
"parameters": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]
}
}
}]
def run_command(command: str) -> str:
result = subprocess.run(["opsctl", *command.split()], capture_output=True, text=True)
return result.stdout or result.stderr
response = client.chat(
model="llama3.1",
messages=[{"role": "user", "content": "Is billing-api healthy in production right now?"}],
tools=tools,
)
Honest result: llama3.1 at 8B took the discovery step about two times out of three in casual testing, correctly running opsctl --help before guessing at a subcommand. The other third of the time it guessed a plausible-looking opsctl status billing-api production without the flag names, which opsctl's own argument parser rejected with a usage error, at which point the model read the error and self-corrected on the next call. It never silently failed, the CLI's own validation caught the malformed call every time, but it wasn't as clean as a frontier model doing it in one shot. That's the actual tradeoff nobody's hot take mentions: a CLI wrapper pushes some of the schema's job onto the model's ability to read help text and recover from a rejected argument, and a small model pays for that in extra round trips, not in silent wrong answers.
The rule of thumb I’m actually using now
None of this adds up to “MCP bad, CLI good.” It adds up to a question I ask before wrapping anything now, and the answer determines which way I go.
REACH FOR MCP WHEN REACH FOR A CLI / SINGLE FUNCTION
TOOL WHEN
----------------------------------------------------------------------------
The tool needs to be discoverable and There's one clear caller, and it's
reusable across many different agents not going to be reused by some
and sessions you don't control, with other agent or client you haven't
usage patterns you can't predict met yet
You need a real permission boundary The agent already runs in a sandbox
around something dangerous: human or with scoped credentials that
approval gates, scoped OAuth tokens, limit blast radius just as well,
audit logging enforced centrally and MCP would just be ceremony
wrapped around the same boundary
The client genuinely can't write its The tool is high-frequency: paying
own integration, a weaker model or a a fixed schema cost once beats
non-coding client with no shell to work re-describing it, but paying it on
with every turn for months doesn't
The response is small and bounded, so The response can be large or
the always-on schema cost is the only unbounded, and you need a hard
real overhead you're paying budget and structured exit codes
more than you need discoverability
The security-boundary row is the one I’d weight above the token math, honestly. A tool sitting in front of something that can actually hurt you, a production database, a payments flow, anything with a rollback or restart in its name, benefits from a boundary enforced structurally rather than trusted to whatever the agent decides to do that day. My rollback_deploy tool, in the naive version above, is exactly the kind of thing I'd keep behind a real MCP server with server-side confirmation and audit logging even after everything above, specifically because "the agent decided not to call it without confirm=true" is a much weaker guarantee than "the server physically won't execute it without a signed approval token." The read-only status check next to it has no such argument in its favor, and paying its schema tax on every turn for the rest of the project's life was never buying me anything.
Where I landed
I didn’t come out of this deleting my MCP servers. I came out of it splitting them. The rollback and restart tools stayed behind MCP, because the boundary is the actual point there, not the convenience. The status check and log tail got demoted to opsctl subcommands behind a single run_command tool, because nothing about them needed a schema sitting in context on every turn for months, and the 5.4x overhead I measured on my own machine wasn’t buying me anything for capabilities that get called once or twice a session at most.
The three pieces that started this for me weren’t arguing that MCP is a bad idea. CsMesh’s author built the MCP version first and only moved off it once the cost showed up in practice. The SPFx team didn’t say MCP tools don’t work, they said tools nobody calls don’t help regardless of what protocol wraps them. And even the most combative of the three, the security teardown, ends with a real company replacing MCP with plain APIs and a CLI for internal use, not with abandoning tool-calling altogether. The actual finding across all three, and the one I’d want someone to walk away with, is narrower than “avoid MCP”: know which of your tools need the reach and the boundary MCP gives you, and stop paying its overhead for the ones that don’t.
Tags: mcp, ai-agents, llm, software-architecture, context-engineering, dotnet, developer-tools
Top comments (0)