Most MCP demos look the same. A tool is registered. A client calls it. The tool returns the expected data. Everyone nods.
That demo tells you the tool can work. It tells you nothing about what happens when the tool is given the wrong tenant ID, or when it declares scope it should not have, or when it is about to reach users who did not write it.
I built a preflight scanner for exactly that gap. This post is about what a bounded preflight can actually check, and what it honestly cannot.
The tools you expose are part of your attack surface
An MCP tool is a capability. Some capabilities are safe to hand out freely. Others are not.
Common risky declarations:
- A tool that runs shell commands from user input.
- A tool with
filesystem:*ornetwork:*scope when it only needs to read one directory. - A tool whose description contains a secret-like value from a copy-pasted config.
- A tool that takes an unvalidated URL or file path from an agent.
- A tool that writes without any approval requirement.
These are not exotic. They show up in real MCP servers being shipped right now.
What a bounded preflight does
The scanner I built runs two kinds of checks: static and behavioral.
Static rules
Static rules look at the tool metadata: name, description, declared scopes, and any command patterns. Four rules cover the common cases:
-
MCP-001: unsafe command declaration. Anything that runs a shell or takes a free-form command string. -
MCP-002: excessive filesystem or network scope. Anything that declaresfilesystem:*,network:*, ornetwork:egresswhen the tool clearly does not need it. -
MCP-003: secret-like value in tool metadata. Anything that looks like an API key, token, or password embedded in a description or default config. -
MCP-004: untrusted input reaching a sensitive operation. Anything where user input flows into a command, a file path, or a URL without an obvious boundary.
Behavioral checks
Behavioral checks actually call a local fixture server and observe what happens:
- Tenant boundary: does a tenant-a token get tenant-b data?
- Write approval: does a write tool run without an explicit, request-bound approval?
- Quota: does a runaway loop get stopped, or does it keep calling the tool forever?
Each check produces a pass, fail, blocked, or incomplete status. Nothing is reported as a pass if it was not actually run.
What the report looks like
One run produces both a Markdown report and a JSON report. Each finding has:
- A rule ID (
MCP-001, etc.). - A severity (info, low, medium, high).
- The affected tool.
- A sanitized evidence line, with any secret-like value redacted.
- A remediation.
- A status (open, accepted, resolved).
The point is that a finding becomes an engineering task, not a vague warning.
What a preflight is not
It is not a penetration test. It is not a certification. It is not a scan of a real customer environment. It does not replace a real security review.
It is a bounded, repeatable first pass. It catches the obvious things before they ship, and it does so without pretending to be more than it is.
The one honest limit
A preflight can tell you that a tool declares too much scope. It cannot tell you that the scope is intentional and appropriate for the business. That is a decision for the team.
A preflight can tell you that a write tool ran without an approval. It cannot tell you whether the approval was correctly granted in a real workflow. That is a process question.
The scanner is a filter. It narrows the surface for the human review that still has to happen.
The code is at github.com/glatinone/mcp-security-preflight. If your team is about to expose MCP tools to users or internal agents and wants a bounded first pass, I take short sprints on exactly this. kielltampubolon.id
Top comments (9)
The strongest point is treating “incomplete” as a real result rather than quietly turning it into a pass. For tenant isolation, I’d also test the boundary independently of the tenant ID supplied by the caller. An agent-controlled tenant_id should never be the authority for determining which tenant the request belongs to; that identity should come from an authenticated context and be enforced again at the data-access layer. Otherwise a preflight can confirm that tenant-a behaves correctly while missing the more fundamental confused-deputy problem. The static + behavioral combination is a good pattern because neither metadata inspection nor happy-path execution is enough to establish isolation.
Right, and that distinction is exactly where I draw the scanner's edge. The behavioral check runs tokens from both tenants as fixtures, so it verifies enforcement at the boundary, but it still trusts caller identity as given. A tool that accepts tenant_id as a plain parameter is a different class and honestly deserves its own rule: flag any tool taking tenant-scoped parameters with no identity provider reference in the config. Right now those only get caught by the MCP-004 input-boundary heuristics. Adding the dedicated rule to the backlog, that one's on you.
That's the framing more MCP writeups need: a happy-path demo proves the tool can work, not that it's safe to expose. The scope-declaration problem especially: I've seen agents get a filesystem-scoped tool handed to them because it was the only thing that worked in the demo, and nobody revisits that grant when the deployment changes.
The honest part is what your preflight can't check: a tool can declare tight scope and still leak via its description or its error strings. I ended up treating the tool description itself as attack surface after a prompt injected through one. Are you parsing descriptions for secret-like content in the scanner, or is that still manual review?
Parsing is live, that's MCP-003: descriptions and default configs get scanned for secret-shaped strings (key patterns, tokens, passwords) and the evidence line redacts before it ever hits the report. But your injection case is a different class and honestly a gap: secret-like and instruction-like content need different detectors, and instruction-shaped text in a description is currently only caught if it trips the MCP-004 input heuristics. Treating the description as untrusted instruction text, with the same suspicion as a user prompt, is the right model and I'm adding it to the backlog as its own rule. Your prompt-injection-through-description story would make a good test fixture if you ever want to contribute it.
This is fantastic - what have you seen adoption wise from developers who likely will just always click "accept" ? any tips on getting adoption with that crowd ?
The always-click-accept crowd is real, and honestly the only defense that works is making their click expensive to ignore later. In the scanner, an accepted finding stays in the report with status accepted, who accepted it and why gets recorded, and the next run still shows it instead of silently going green. Same idea for runtime approvals: scope the grant to the specific request and tool, keep an allowlist so repeat accept is a no-op that is still auditable, and surface accept counts in the summary so a team drowning in prompts has a number staring at them. Adoption for that crowd comes from making the safe path the lazy path, not from better warning copy.
Any ability to flag "false positives" like how you would mark code as okay in SAST product?
Yes, that's what the accepted status is for. A finding marked accepted stays in the report with who accepted it and why, and the next run still shows it instead of dropping it, so it works like a SAST suppression that leaves a trail. What I haven't built is a separate false positive label. "The scanner was wrong" and "the scanner was right and we accept the risk" are different statements, so would you want those split?
Some comments may only be visible to logged-in visitors. Sign in to view all comments.