Most fact-checking systems for LLM agents ask one question: is the claim supported by the evidence? They do not ask a second, equally important question: did the claim come from the source the agent cited?
When an MCP agent pulls data from a search tool, a database query, a patient record API, and a policy document, then synthesizes an answer, a source-blind verifier will pass any claim that appears somewhere in the pooled evidence. If the agent says "According to the account record, this plan includes a 30-day refund window," but the refund policy came from a scraped FAQ and not the structured account record, the verifier sees the fact and approves it.
This failure mode is called cross-source conflation. It matters most in financial, healthcare, and compliance contexts where the authority of the source determines whether the answer is actionable. A trading agent that cites Bloomberg terminal data but actually pulled from a Reddit scrape is not just wrong. It is dangerously misattributed.
Multiverse Computing's ProvenanceGuard addresses this by building a verification layer that checks not just factual support, but source lineage. Published September 29, 2026, the system introduces source-aware verification as a separate concern from fact verification.
Why MCP Makes Provenance Harder
The Model Context Protocol lets agents call multiple tools in a single turn. Each tool returns structured data, but MCP itself has no built-in provenance primitives. The protocol does not require servers to sign their responses, attest to data lineage, or declare confidence scores.
When an agent receives:
- A search result from a web scraper
- A row from a Postgres query
- A JSON blob from a Bloomberg API
- A paragraph from a PDF retrieval tool
The agent sees four text chunks. The orchestration layer may log which tool returned which chunk, but the LLM prompt does not enforce citation discipline. The agent can say "According to the database" and pull facts from the PDF.
Traditional fact-checkers (RAGAS, MiniCheck, AlignScore, SummaC) concatenate all evidence into a single context and ask whether each claim is entailed. They do not track which tool provided which sentence. They do not penalize cross-source conflation.
How ProvenanceGuard Works
ProvenanceGuard sits between the agent's output and the user. It receives:
- The agent's final answer
- The set of MCP tool outputs (each tagged with tool name and metadata)
- The agent's citations (explicit or inferred)
The system then:
- Decomposes the answer into atomic claims using an LLM or dependency parser
-
Maps each claim to the cited source (e.g., "According to the account record" maps to the
get_accounttool output) - Verifies entailment within that specific source using a fine-grained NLI model
- Flags cross-source conflation when a claim is supported by evidence from a different tool than the one cited
If a claim fails source-aware verification, ProvenanceGuard can either block the answer or trigger a repair step that re-queries the correct tool.
Architecture
┌─────────────┐
│ MCP Agent │
└──────┬──────┘
│ answer + tool outputs
▼
┌─────────────────────────────┐
│ ProvenanceGuard │
│ ┌─────────────────────────┐ │
│ │ Claim Decomposer │ │
│ └───────────┬─────────────┘ │
│ ▼ │
│ ┌─────────────────────────┐ │
│ │ Citation Mapper │ │
│ │ (claim → cited tool) │ │
│ └───────────┬─────────────┘ │
│ ▼ │
│ ┌─────────────────────────┐ │
│ │ Source-Scoped NLI │ │
│ │ (verify within tool) │ │
│ └───────────┬─────────────┘ │
│ ▼ │
│ ┌─────────────────────────┐ │
│ │ Conflation Detector │ │
│ └───────────┬─────────────┘ │
└─────────────┼───────────────┘
▼
┌───────────────┐
│ Pass / Block │
│ / Repair │
└───────────────┘
The claim decomposer breaks the answer into sentences or sub-sentence units. The citation mapper uses regex patterns or an LLM to extract phrases like "According to X" and match them to tool names. The source-scoped NLI model (often a fine-tuned DeBERTa or similar) checks entailment only within the evidence from the cited tool. The conflation detector flags cases where the claim is true in the pool but not in the cited source.
Implementation Considerations
Latency Budget
Running an NLI model on every claim adds 50-200ms per claim. For a 10-claim answer, that is 500-2000ms. If you are building a customer support chatbot, this may be acceptable. If you are building a high-frequency trading agent, it is not.
Options:
- Batch verification: collect all claims and verify in parallel
- Selective verification: only verify claims that cite high-stakes sources (e.g., account records, compliance docs)
- Cached entailment: if the same tool output appears in multiple turns, cache the NLI results
Citation Extraction
The paper assumes agents produce explicit citations like "According to the account record." In practice, many agents do not. You have three options:
-
Prompt engineering: force the agent to cite sources in a structured format (e.g.,
[tool:get_account] The plan includes...) - Heuristic matching: use keyword overlap to guess which tool output supports each claim
- LLM-based attribution: ask a second LLM to label each claim with the most likely source
Option 1 is cleanest but requires prompt discipline. Option 3 is most robust but adds another LLM call.
Conflation Repair
When ProvenanceGuard detects cross-source conflation, it can:
- Block the answer and return an error to the orchestrator
- Re-query the cited tool and regenerate the claim
- Downgrade the claim to a hedged statement (e.g., "Some sources suggest...")
The paper demonstrates a repair loop where the agent is prompted to regenerate the conflated claim using only the cited tool's output. This works if the cited tool actually contains the information. If it does not, the agent may hallucinate or refuse to answer.
MCP Server Metadata
ProvenanceGuard does not require MCP servers to change their API. But if you control the server, you can make verification easier by:
-
Tagging outputs with confidence scores (e.g.,
"confidence": "high"for Bloomberg API,"confidence": "low"for web scrape) - Signing responses with a cryptographic signature to prove the server identity
-
Attesting to data lineage (e.g.,
"source_chain": ["Bloomberg API", "internal cache"])
None of this is part of the MCP spec today. If you are building financial or healthcare agents, you may want to extend your servers with these fields.
Trade-Offs
| Dimension | Source-Blind Verification | Source-Aware Verification |
|---|---|---|
| Latency | 50-100ms per answer | 500-2000ms per answer |
| False Positives | Low (passes true facts) | Medium (blocks misattributed facts) |
| False Negatives | High (misses conflation) | Low (catches conflation) |
| Prompt Dependency | Low | High (needs citations) |
| MCP Server Changes | None | Optional (metadata helps) |
| Repair Complexity | N/A | High (re-query or hedge) |
When It Matters
Source-aware verification is critical when:
- The source determines legal liability (e.g., a compliance agent citing a regulation vs. a blog post)
- The source determines financial authority (e.g., a trading agent citing Bloomberg vs. Twitter)
- The source determines medical authority (e.g., a clinical agent citing a patient record vs. a symptom checker)
It is less critical when:
- All sources are equally authoritative (e.g., a research agent pulling from peer-reviewed papers)
- The agent is read-only (e.g., a summarization tool with no downstream actions)
- Latency is more important than precision (e.g., a conversational assistant)
Failure Modes
The Agent Lies About Citations
If the agent is prompted to cite sources but fabricates them, ProvenanceGuard will fail. The citation mapper will match the fabricated citation to a real tool, and the NLI model will check the wrong evidence.
Mitigation: log all tool calls and verify that cited tools were actually invoked. If the agent cites get_account but never called it, block the answer.
The NLI Model Is Wrong
Fine-grained NLI models are not perfect. They may:
- Miss paraphrases: the claim is supported but phrased differently
- Hallucinate entailment: the claim is not supported but the model thinks it is
- Fail on negation: the claim is the opposite of what the evidence says
Mitigation: use ensemble verification (multiple NLI models) or human-in-the-loop review for high-stakes claims.
The Tool Output Is Ambiguous
If a tool returns a JSON blob with nested fields, the NLI model may struggle to determine which field supports which claim. For example:
{
"account": {
"plan": "Premium",
"refund_window": 30
},
"policy": {
"refund_window": 14
}
}
If the agent says "The refund window is 30 days," which field does it cite? The account record or the policy?
Mitigation: flatten tool outputs or use structured extraction to map claims to specific JSON paths.
Technical Verdict
Use source-aware verification when you are building agents that:
- Operate in regulated domains (finance, healthcare, legal)
- Make decisions based on authoritative vs. non-authoritative data
- Cite multiple tools in a single answer
- Have a latency budget that tolerates 500-2000ms of verification overhead
Avoid it when:
- All tools are equally trustworthy
- The agent is conversational and read-only
- Latency is more important than attribution precision
- You cannot enforce citation discipline in the agent prompt
If you are building financial agents, this is not optional. If you are building a chatbot, it probably is.
Source Links
- ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (Hugging Face Blog, September 29, 2026)
Top comments (0)