DEV Community

MCP Deep Dive, Part 10: When the Agent Feels Off — Debugging and Observability for MCP in Production

kirandeepjassal-crypto on July 21, 2026

A web API fails loudly: a 500, a stack trace, an alert. An agent fails softly. It doesn't crash — it quietly takes six turns instead of two, calls...
Collapse
 
alexshev profile image
Alex Shev

Observability for MCP gets interesting because the failure is often not a crash. The agent “feels off” because a tool returned stale context, a permission boundary was too broad, or a retry changed the state. Traces need to show intent, inputs, tool output, and what the agent believed afterward.

Collapse
 
kirandeepjassalcrypto profile image
kirandeepjassal-crypto

Exactly — the crash is the easy case, because at least it announces itself. The failure that actually hurts is the one where every span is green and the answer is still wrong: a tool returned stale context, a boundary was a little too broad, a retry mutated state between attempts. Nothing "failed," so nothing pages you.

The line I'd underline is your last one — "what the agent believed afterward." That's the field most observability setups miss. We log intent, inputs, and tool output, and then stop right before the interesting part: what the model concluded from that output and carried into the next step. When an agent goes off, the divergence almost always lives there — the tool returned correct data and the model drew the wrong inference, or it silently down-weighted a result. Without capturing the post-tool belief state, you can see that the inputs were fine and the output was fine and still have no idea why the run went sideways.

The retry point is the other quiet one. A retry isn't idempotent at the belief level even when it's idempotent at the tool level — the model now has two observations of the "same" call, and if they differ at all, that difference becomes reasoning it acts on. Traces that only show the final attempt hide exactly the state change that explains the weird behavior.

So the unit that actually debugs these isn't request → response, it's intent → inputs → tool output → resulting belief, per step, across the whole run. Green spans tell you the plumbing held; only that last field tells you what the agent was actually thinking.

Collapse
 
alexshev profile image
Alex Shev

That "all spans green but the answer is wrong" case is the one that matters in local ranking work. A Maps agent can successfully fetch a grid, categories, reviews, and competitors, then still reason from yesterday's centroid or an old business name. I would log the evidence the agent believed, not just the tool calls it made.

Thread Thread
 
kirandeepjassalcrypto profile image
kirandeepjassal-crypto

That local-ranking example is the cleanest illustration of the whole problem I've seen. Every tool call succeeds — grid fetched, categories fetched, reviews fetched, competitors fetched, all green — and the agent still reasons from yesterday's centroid or a renamed business. The output is confidently wrong with a spotless trace behind it, and there's no span you can point at because no span failed.

"Log the evidence the agent believed, not just the tool calls it made" is exactly the fix, and your domain makes the why obvious: the tool returning fresh data and the agent actually using the fresh data are two separate events, and the gap between them is where the stale centroid lives. If you only capture the call, that gap is invisible — you can prove the data was current and still not explain the wrong ranking. Recording which centroid or business name the agent carried forward is what turns "fetched correctly but reasoned from stale state" into something you can see instead of infer.

The bit I'd add for the Maps case specifically: freshness has to be part of the believed-evidence record, not just the fetch. "Reviews fetched at T" and "agent ranked using reviews it treated as current" can diverge by a full refresh cycle, and that delta is the bug. Log the belief with its as-of timestamp and stale reasoning stops being a mystery.

Thread Thread
 
alexshev profile image
Alex Shev

That is exactly the failure class. The tool layer can be perfectly green while the model's working state is wrong. I would rather debug the evidence bundle the agent accepted than stare at a successful trace and pretend success means correctness.

Collapse
 
alexshev profile image
Alex Shev

Yes, "what the agent believed afterward" is the missing column. For MCP I would log it as a small derived state, not a giant transcript: tool selected, fact accepted, confidence/source, and whether the answer/action depended on it. Then an incident can ask whether the bad output came from the tool, the interpretation, or a stale belief that survived after the call.

Collapse
 
mads_hansen_27b33ebfee4c9 profile image
Mads Hansen

Treating the run as the observability unit is exactly right. Two production details matter here. First, tenant is often too high-cardinality for a metric label; it can explode cost and accidentally expose customer identity. Keep low-cardinality dimensions in metrics, put a pseudonymous tenant key in controlled logs/traces, and link outliers with exemplars. Raw goals, arguments, and results should be denied by default rather than relying on a best-effort Redact() call.

Second, a session replay needs a manifest, not only JSON-RPC events: model/provider version, system-prompt digest, tool-schema/version digest, policy version, routing config, and captured external-result hashes. Replay the tool layer with recorded responses first, then separately test the current live dependencies. Otherwise a “failed replay” cannot distinguish agent drift from changed tools or data. I’d also sign/hash-chain the audit sequence so reconstruction evidence is tamper-evident, not merely append-only by convention.

Collapse
 
kirandeepjassalcrypto profile image
kirandeepjassal-crypto

Both of these are corrections I'll take, and the second one changes how I'd build the replay entirely.

On tenant as a metric label — you're right, and I conflated two stores that should have different rules. Tenant is high-cardinality and identifying, so as a metric dimension it's the worst of both: it blows up your time-series cardinality (and bill), and it quietly turns a metrics backend into a place customer identity leaks. The split you're describing is the correct one: metrics carry only low-cardinality dimensions (tool, status, model, maybe tenant tier), and the tenant key lives as a pseudonymous id in access-controlled logs/traces, with exemplars linking a spiking metric to the specific traces behind it. You get the outlier's fingerprint without putting the identity in the metric. And the deeper point — deny raw goals/arguments/results by default rather than leaning on a best-effort Redact() — is the one I under-weighted. A redactor you have to remember to call is a redactor that eventually doesn't get called. Default-deny with an explicit allowlist of what's safe to capture is the only version that survives contact with a tired engineer at 2am.

On replay, the manifest is the part I got wrong. JSON-RPC events alone reconstruct what the agent did, but not the world it did it in — and without that, a failed replay is uninterpretable. Your manifest list is exactly the missing context: model/provider version, system-prompt digest, tool-schema/version digest, policy version, routing config, and hashes of the external results. And the two-phase idea is the key insight: replay the tool layer with recorded responses first (does the agent still behave the same against a frozen world?), then separately exercise the live dependencies (did the tools or data change underneath us?). Collapse those and you can't tell agent drift from a changed downstream — which is precisely the "the agent feels off" ambiguity the whole post was trying to resolve. I described one-trace-per-run as the unit but didn't give it enough context to be deterministically replayable; the manifest is what closes that.

And hash-chaining the audit sequence is the right upgrade over "append-only by convention." Append-only is a property of how you intend to write; a signed hash chain is a property you can verify. During an incident — especially a security one — "this reconstruction is tamper-evident" is a materially stronger statement than "we don't have a delete path." It's the same theme a couple of other commenters have pushed me on across this series: stop trusting conventions, make the guarantee something a test or a signature can prove.

Going to fold both into a revision — the metrics/logs cardinality split, and the replay-manifest-plus-two-phase model with a hash-chained trail. Genuinely some of the most useful feedback the series has gotten.

Collapse
 
raju_dandigam profile image
Raju Dandigam

The point about treating the run, not the request, as the unit of observability is exactly right. If you only instrument tool latency, you miss the agent behaviors that matter most in production: thrashing, duplicate turns, tool oscillation, and retries that look like progress.

That is a big part of why we built agent-inspect around local-first execution traces and tool-call receipts for TypeScript agent workflows. Curious whether your audit log and trace model share the same run id end to end, or whether you join them later during incident triage?