DEV Community

Cover image for Source-Aware Verification for MCP Agents: Why Fact-Checking Isn't Enough When Tools Lie About Provenance
mech.app
mech.app

Posted on Originally published at mech.app

Source-Aware Verification for MCP Agents: Why Fact-Checking Isn't Enough When Tools Lie About Provenance

Most fact-checking systems for LLM agents ask one question: is the claim supported by the evidence? They do not ask a second, equally important question: did the claim come from the source the agent cited?

When an MCP agent pulls data from a search tool, a database query, a patient record API, and a policy document, then synthesizes an answer, a source-blind verifier will pass any claim that appears somewhere in the pooled evidence. If the agent says "According to the account record, this plan includes a 30-day refund window," but the refund policy came from a scraped FAQ and not the structured account record, the verifier sees the fact and approves it.

This failure mode is called cross-source conflation. It matters most in financial, healthcare, and compliance contexts where the authority of the source determines whether the answer is actionable. A trading agent that cites Bloomberg terminal data but actually pulled from a Reddit scrape is not just wrong. It is dangerously misattributed.

Multiverse Computing's ProvenanceGuard addresses this by building a verification layer that checks not just factual support, but source lineage. Published September 29, 2026, the system introduces source-aware verification as a separate concern from fact verification.

Why MCP Makes Provenance Harder

The Model Context Protocol lets agents call multiple tools in a single turn. Each tool returns structured data, but MCP itself has no built-in provenance primitives. The protocol does not require servers to sign their responses, attest to data lineage, or declare confidence scores.

When an agent receives:

  • A search result from a web scraper
  • A row from a Postgres query
  • A JSON blob from a Bloomberg API
  • A paragraph from a PDF retrieval tool

The agent sees four text chunks. The orchestration layer may log which tool returned which chunk, but the LLM prompt does not enforce citation discipline. The agent can say "According to the database" and pull facts from the PDF.

Traditional fact-checkers (RAGAS, MiniCheck, AlignScore, SummaC) concatenate all evidence into a single context and ask whether each claim is entailed. They do not track which tool provided which sentence. They do not penalize cross-source conflation.

How ProvenanceGuard Works

ProvenanceGuard sits between the agent's output and the user. It receives:

  1. The agent's final answer
  2. The set of MCP tool outputs (each tagged with tool name and metadata)
  3. The agent's citations (explicit or inferred)

The system then:

  1. Decomposes the answer into atomic claims using an LLM or dependency parser
  2. Maps each claim to the cited source (e.g., "According to the account record" maps to the get_account tool output)
  3. Verifies entailment within that specific source using a fine-grained NLI model
  4. Flags cross-source conflation when a claim is supported by evidence from a different tool than the one cited

If a claim fails source-aware verification, ProvenanceGuard can either block the answer or trigger a repair step that re-queries the correct tool.

Architecture

┌─────────────┐
│ MCP Agent   │
└──────┬──────┘
       │ answer + tool outputs
       ▼
┌─────────────────────────────┐
│ ProvenanceGuard             │
│ ┌─────────────────────────┐ │
│ │ Claim Decomposer        │ │
│ └───────────┬─────────────┘ │
│             ▼               │
│ ┌─────────────────────────┐ │
│ │ Citation Mapper         │ │
│ │ (claim → cited tool)    │ │
│ └───────────┬─────────────┘ │
│             ▼               │
│ ┌─────────────────────────┐ │
│ │ Source-Scoped NLI       │ │
│ │ (verify within tool)    │ │
│ └───────────┬─────────────┘ │
│             ▼               │
│ ┌─────────────────────────┐ │
│ │ Conflation Detector     │ │
│ └───────────┬─────────────┘ │
└─────────────┼───────────────┘
              ▼
      ┌───────────────┐
      │ Pass / Block  │
      │ / Repair      │
      └───────────────┘
Enter fullscreen mode Exit fullscreen mode

The claim decomposer breaks the answer into sentences or sub-sentence units. The citation mapper uses regex patterns or an LLM to extract phrases like "According to X" and match them to tool names. The source-scoped NLI model (often a fine-tuned DeBERTa or similar) checks entailment only within the evidence from the cited tool. The conflation detector flags cases where the claim is true in the pool but not in the cited source.

Implementation Considerations

Latency Budget

Running an NLI model on every claim adds 50-200ms per claim. For a 10-claim answer, that is 500-2000ms. If you are building a customer support chatbot, this may be acceptable. If you are building a high-frequency trading agent, it is not.

Options:

  • Batch verification: collect all claims and verify in parallel
  • Selective verification: only verify claims that cite high-stakes sources (e.g., account records, compliance docs)
  • Cached entailment: if the same tool output appears in multiple turns, cache the NLI results

Citation Extraction

The paper assumes agents produce explicit citations like "According to the account record." In practice, many agents do not. You have three options:

  1. Prompt engineering: force the agent to cite sources in a structured format (e.g., [tool:get_account] The plan includes...)
  2. Heuristic matching: use keyword overlap to guess which tool output supports each claim
  3. LLM-based attribution: ask a second LLM to label each claim with the most likely source

Option 1 is cleanest but requires prompt discipline. Option 3 is most robust but adds another LLM call.

Conflation Repair

When ProvenanceGuard detects cross-source conflation, it can:

  • Block the answer and return an error to the orchestrator
  • Re-query the cited tool and regenerate the claim
  • Downgrade the claim to a hedged statement (e.g., "Some sources suggest...")

The paper demonstrates a repair loop where the agent is prompted to regenerate the conflated claim using only the cited tool's output. This works if the cited tool actually contains the information. If it does not, the agent may hallucinate or refuse to answer.

MCP Server Metadata

ProvenanceGuard does not require MCP servers to change their API. But if you control the server, you can make verification easier by:

  • Tagging outputs with confidence scores (e.g., "confidence": "high" for Bloomberg API, "confidence": "low" for web scrape)
  • Signing responses with a cryptographic signature to prove the server identity
  • Attesting to data lineage (e.g., "source_chain": ["Bloomberg API", "internal cache"])

None of this is part of the MCP spec today. If you are building financial or healthcare agents, you may want to extend your servers with these fields.

Trade-Offs

Dimension Source-Blind Verification Source-Aware Verification
Latency 50-100ms per answer 500-2000ms per answer
False Positives Low (passes true facts) Medium (blocks misattributed facts)
False Negatives High (misses conflation) Low (catches conflation)
Prompt Dependency Low High (needs citations)
MCP Server Changes None Optional (metadata helps)
Repair Complexity N/A High (re-query or hedge)

When It Matters

Source-aware verification is critical when:

  • The source determines legal liability (e.g., a compliance agent citing a regulation vs. a blog post)
  • The source determines financial authority (e.g., a trading agent citing Bloomberg vs. Twitter)
  • The source determines medical authority (e.g., a clinical agent citing a patient record vs. a symptom checker)

It is less critical when:

  • All sources are equally authoritative (e.g., a research agent pulling from peer-reviewed papers)
  • The agent is read-only (e.g., a summarization tool with no downstream actions)
  • Latency is more important than precision (e.g., a conversational assistant)

Failure Modes

The Agent Lies About Citations

If the agent is prompted to cite sources but fabricates them, ProvenanceGuard will fail. The citation mapper will match the fabricated citation to a real tool, and the NLI model will check the wrong evidence.

Mitigation: log all tool calls and verify that cited tools were actually invoked. If the agent cites get_account but never called it, block the answer.

The NLI Model Is Wrong

Fine-grained NLI models are not perfect. They may:

  • Miss paraphrases: the claim is supported but phrased differently
  • Hallucinate entailment: the claim is not supported but the model thinks it is
  • Fail on negation: the claim is the opposite of what the evidence says

Mitigation: use ensemble verification (multiple NLI models) or human-in-the-loop review for high-stakes claims.

The Tool Output Is Ambiguous

If a tool returns a JSON blob with nested fields, the NLI model may struggle to determine which field supports which claim. For example:

{
  "account": {
    "plan": "Premium",
    "refund_window": 30
  },
  "policy": {
    "refund_window": 14
  }
}
Enter fullscreen mode Exit fullscreen mode

If the agent says "The refund window is 30 days," which field does it cite? The account record or the policy?

Mitigation: flatten tool outputs or use structured extraction to map claims to specific JSON paths.

Technical Verdict

Use source-aware verification when you are building agents that:

  • Operate in regulated domains (finance, healthcare, legal)
  • Make decisions based on authoritative vs. non-authoritative data
  • Cite multiple tools in a single answer
  • Have a latency budget that tolerates 500-2000ms of verification overhead

Avoid it when:

  • All tools are equally trustworthy
  • The agent is conversational and read-only
  • Latency is more important than attribution precision
  • You cannot enforce citation discipline in the agent prompt

If you are building financial agents, this is not optional. If you are building a chatbot, it probably is.

Source Links

Top comments (0)