MCP agents act on your behalf but can't prove what they did. Logs are self-reported claims. Receipts are independently verifiable evidence. Here's...
For further actions, you may consider blocking this person and/or reporting abuse
The discovery gap is real, and it compounds: even if a server publishes a
/.well-known/agent.jsonor adid:webdocument, that declaration is only as trustworthy as DNS at lookup time. The harder problem is that the manifest you inspected at discovery and the tools actually served at call time can diverge silently - tool injection, version rollback, capability creep between the two moments. Binding a SHA-256 hash of the manifest at discovery time to the proof chain of the first tool call would close that window: if the manifest drifted between "I connected" and "I called", the chain breaks. That's the gap between a static identity claim and a cryptographically verifiable execution context.The "polished UX, manual backend" pattern is actually a useful diagnostic: it usually means the agent's tool calls aren't observable enough to trust in production. When you can't distinguish what the agent executed from what a human substituted, you can't safely automate the human out - you need to audit every call first. One way to break that loop is to make the agent's actual MCP
tools/callinvocations cryptographically verifiable before expanding autonomy. Run a week in shadow mode where every tool call produces a signed, timestamped receipt, compare the agent's actual behavior against what the UX claimed happened, and you have a concrete basis to decide whether the human can step back.The
did:webmethod partially addresses this - a server can host/.well-known/did.jsonbinding its domain to a public key, and an agent can resolve that DID before the first call to verify it's talking to the declared entity. The gap is that nothing enforces it: there's no protocol-level handshake in MCP that requires a server to present a verifiable identity, so adoption is optional and therefore sparse. A practical middle ground that doesn't wait for spec changes: hash the server's tool manifest at first connection, store it as a baseline, and flag any subsequent drift - that at least catches post-connection substitution even if it doesn't solve pre-connection identity. The discovery problem and the transparency problem are the same problem at different points in the timeline.The database stall pattern often exposes something subtler: there's no way to distinguish "agent called the SQL tool" from "human ran the query and injected the result." Both produce identical output. The transparency problem in MCP isn't just a UX issue - it's that tool calls leave no independently verifiable trace. A signed, per-call record (args, response, timestamp) would let you audit the actual execution path, not just the final answer.
The gap you're describing is narrower than it looks: on-chain x402 transactions give you an immutable payment record, but the tool invocation that caused the payment - exact arguments, response, moment of execution - has no equivalent. Your users can verify that value moved on Base; they can't verify what the agent decided before it moved.
That asymmetry points to debugging and attestation being different problems for different audiences. Debugging is for the developer (stack trace, replay logs). Attestation is for the principal behind the agent - they want a signed receipt they can hand to a counterparty or auditor, not a log entry in a system they don't control. Most "transparency" tooling today solves the first; the second is still largely unaddressed at the MCP layer.
The distinction between an app-controlled review log and an independently verifiable receipt is useful. I operate an early public work/review history for AI agents, and our current trail is still app-controlled. For a publish or peer-review tool call, which fields would you sign so another party can verify the action without exposing private work content? I'd be glad to compare a small workflow if you're interested.
the receipt layer solves "prove what happened." the gap i keep running into is the step before that: deciding whether the action should happen at all.
a cryptographic receipt of "agent sent payment to wrong vendor at 3am" is useful for the audit. a policy that holds the payment for human review before it executes would have been more useful at 3am. both layers matter but most of the tooling energy right now is going into receipts, not interception.
The "your agent can't show its work" framing maps to something I keep hitting from the payment side. I shipped an MCP server with x402 micropayments on Base mainnet (forgepointsignal.com), and the trust gap shows up consistently — agents are willing to call tools, but the human behind them wants visibility into what was called and why.
Curious whether you're seeing transparency as primarily a debugging problem, a trust/audit problem, or something else? Feels like multiple problems wearing the same word.
Saw you've been shipping ark-forge/trust-layer — certifying proxy with SHA-256 proof chain + Ed25519 + Sigstore Rekor. That's the transparency framing from your post turned into running code. Curious how you're thinking about the intersection with the payment side — agents that pay for tool calls also need a paid-call audit trail, and trust-layer feels like it naturally extends there. Happy to share what I've seen on my own x402 deployment if useful.
— Randy
The receipt model is a solid approach. But I keep thinking about the mirror image of this problem on the discovery side.
Your article covers proving what an agent did after connecting. There's an equally messy question though: how does an agent verify what it's connecting TO before that first tool call?
MCP today assumes you already know and trust the server URL. No standard way for a server to present verifiable identity or capability claims before the handshake. You could have perfect cryptographic receipts for every action and still have connected to the wrong server.
The IETF has 8+ competing drafts trying to solve different slices of this (agents.txt for policy, ARDP for registration, AID for DNS-based resolution). None talk to each other yet.
Curious if you see the receipt model tying into discovery metadata at some point. A server publishing its receipt-issuer identity as part of a discoverable profile would close both ends of the trust gap.