When you pull an existing MCP server or an OpenAI/Claude-style skill into a local agent package model, the hard part is not the install. It is turning foreign tool metadata into something reviewable: provenance and version kept intact, network/filesystem/secret needs mapped to explicit permissions, and hooks left off until someone approves them.
A credible baseline is still manual. Read the server or skill manifest, list every tool action, note undeclared filesystem or network reach, and reject or mark unsupported anything you cannot map. Auto-enable is how permanent owner-memory write and silent background behavior sneak in.
Running imports through the same validator and permission model as native packages only helps if unmappable authority fails closed and compatibility status stays visible (native, wrapped, partial, rejected).
How do you currently record provenance and permission gaps when wrapping a third-party MCP server so a later review can tell those four outcomes apart?
Top comments (8)
Absence is not neutral in this metadata, and that breaks the four-state ladder at the exact point it is supposed to hold. The MCP annotation schema documents a default of true for
destructiveHint, and it conditions bothdestructiveHintandidempotentHintonreadOnlyHintbeing false. So a server that ships no annotations resolves, by the schema's own defaults, to the most dangerous reading available. A wrapper that maps declared needs into permissions inverts that. Silence yields an empty permission set, which means there is no unmappable authority to fail closed on, and the import lands in the tidy end of native/wrapped. The tools nobody can vouch for are the ones that sort cleanest. Storing absent separately from explicit false, then resolving absent to the schema default instead of to nothing, is the cheap repair.The harder problem is that a permission gap written in the manifest's own vocabulary has no content. If the record says a tool claims it needs network access, the manifest has defined the set being compared against, and the diff is claim against claim. Gaps acquire content only when the permission record is written in terms the server does not control: enforced filesystem roots, or egress destinations observed at the sandbox boundary. A later reviewer can then compare what was granted and what was touched against what was asserted.
That makes digest pinning weaker than it looks. Metadata invariance does not imply behavioral invariance, and the digest can hold while an endpoint resolved at runtime moves.
For the four outcomes to survive a later review, the record probably needs one row per (server, tool) with three columns: the claim as shipped with absent distinguished from false, the grant as actually enforced, and boundary reach observed over the first N calls, with denied attempts kept apart from successful ones. Status then gets derived rather than typed in by whoever did the import. Partial becomes definable as a grant strictly narrower than the claim, or reach outside the grant.
Which of those three columns is actually retained today, and is observed reach among them, or is the record still claims only?
You're right, and the annotation defaults make the failure concrete:
destructiveHintdefaults totrue, so a server that ships nothing is the most dangerous reading under the schema, while a claims→permissions mapper turns that same silence into an empty grant that sails through as "wrapped." The tools nobody vouched for sort cleanest. That's a real inversion, not a nitpick.Honest answer to your question: what the article describes is claims only: column one, and not even with
absentdistinguished fromfalse. Nothing there is observed reach. So today the four states are typed in by the importer, not derived.The cheap repair you named is the one worth doing first, and it's mechanical:
Persist the tri-state, resolve
absentto the schema default only at evaluation time, and treat anyabsenton a non-readOnlyHinttool as unmappable authority → fail closed. That alone stops silence from landing in "native."The enforced-grant and observed-reach columns need a real sandbox boundary to be truthful, and I'd rather leave them empty than fake them from the manifest. A column filled with claims wearing an "observed" label is worse than a visible gap. Your
partial = grant narrower than claim, or reach outside grantdefinition is the target; the record just isn't there yet. Thanks for the correction.The tri-state and resolving defaults at evaluation time move this the right way. I would push the claims column one step further: absence is the only state that says anything without trusting a declaration. A false destructiveHint is written by the same source whose permissions are being constrained, and lying costs nothing. A true value is self-reported too. So the missing piece is provenance. The manifest's content digest, plus whatever signature the value appeared under, is what a stored value has to be bound to. Without that binding even a tri-state record goes stale the moment the manifest is replaced, and nothing in the record says it happened.
Leaving observed reach empty until a real sandbox boundary exists is the right call. The ambiguity left behind is that one empty cell covers two very different situations. "Not observed" means no sandbox, or the tool was never run inside one. "Observed, no reach" means a run happened and found nothing within that run's boundary. Only the second can be evidence against a claim, and only inside the scope of that run. Does the importer have anywhere to put that distinction today, or would it need a column of its own?
@infracore, the native/wrapped/partial/rejected compatibility states are a useful guard against imports silently gaining authority. I’d also persist the original manifest digest alongside the normalized permissions so a later update cannot inherit approval just because the package name stayed the same. Would you invalidate approval whenever either declared capabilities or previously undeclared network or filesystem behavior changes?
@raju_dandigam Yes on the digest. Keeping the original manifest hash beside the normalized permissions stops a later update from inheriting approval just because the package name stayed the same. I would invalidate whenever declared capabilities change, or when previously undeclared network or filesystem behavior appears. Those are the cases where the earlier review no longer covers what will run. Fail closed, re-map, and keep the prior native/wrapped/partial/rejected status tied to the new digest so the drift stays visible.
'Hooks left off until someone approves them' is the whole game. Install-time is when your authority over the foreign code is maximal and your information about it is minimal - every default that runs at import is borrowed trust you never audited. The manifest read you describe is the manual version of what package tooling should enforce: declared reach vs observed reach, diffed. Skill formats make this worse than classic packages because the payload is prose the model will obey, not just code the runtime runs.
You've named the asymmetry exactly: maximal authority, minimal information at install-time is why every import-time default is borrowed trust. Agreed that the manifest read is the manual stand-in for a declared-vs-observed diff the tooling should own.
The skill-format point deserves its own handling, because prose has no manifest to declare against. Two things that make it at least diffable:
write,fetch,run,remember,alwaysproduces a claims list a human can approve line by line. It's crude, but it turns "read the whole thing" into "approve these six sentences," and anything imperative that wasn't in the approved list on next import is a visible delta.For MCP servers, the observed column only becomes honest inside a boundary that logs actual tool invocations; without one, I'd rather leave it empty than fill it from the manifest wearing an "observed" label.
Thanks for sharpening the prose-vs-code distinction - that's the part I'd underweighted.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.