An AI agent is asked, "Is this French supplier still active, and what is its legal name?" It calls a tool, gets some text back, and replies confidently. What the user doesn't get is where the answer came from, when it was checked, or which parts the model filled in by itself.
This tutorial is about a small pattern that closes that gap: instead of an answer, the tool returns an evidence envelope, and the agent is written to read it, cite it, and pass on what's missing instead of papering over it. The worked example is FACTRAIL MCP, an open-source (Apache-2.0) MCP server. We'll go through one real response line by line, including the field it could not resolve.
Note on the name: this is FACTRAIL MCP (factrail.online, github.com/baronsigma/factrail). An unrelated GitHub project uses the same name.
1. Why agents make up facts
An LLM predicts plausible text. For a question like "what is the legal name of company X?", a plausible answer and a correct one look the same from the inside. Tool calling helps, but only partly:
- A search or scraping tool returns text, and the model still has to work out which part is the fact.
- Most tools return a bare value (
"active"), with no source, no date, and no way of saying "I couldn't find this." - When a field is missing, the model tends to fill the gap rather than report it.
What the agent needs is a structured way to say: this part is established, by this source, at this time; this other part is unknown.
2. What an evidence envelope is
An evidence envelope is a response format in which every fact is tied to its evidence. In FACTRAIL MCP the envelope (schema 1.2) contains:
| Part | What it tells the agent |
|---|---|
status |
Overall outcome: supported, contradicted, insufficient_evidence, stale, or conflicting_sources. Not a boolean. |
facts[] |
Each value, with a support_level (for example authoritative, derived_provisional, caller_input) and the IDs of the evidence behind it. |
evidence[] |
Each source: publisher, URL, retrieval time, source status. |
coverage |
Which requested fields were resolved and which weren't, and why. |
conflicts[] |
Competing claims from different sources, kept visible. |
freshness |
When the evidence was retrieved and whether it's stale. |
receipt_id |
A content-addressed ID you can use to fetch the same envelope later and check it hasn't been altered. |
The design rule behind it: unknown is better than invented.
3. Connect in about a minute
The hosted endpoint is https://mcp.factrail.online/mcp. It uses Streamable HTTP, is free and read-only, and needs no account or API key. Each client IP can make up to 60 MCP requests per minute (full fair-use terms in section 7).
Quick check with curl (tested)
curl -s https://mcp.factrail.online/mcp \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}'
Real response (29/09/2026):
{"jsonrpc":"2.0","id":1,"result":{"capabilities":{"experimental":{},"tools":{"listChanged":false}},"protocolVersion":"2025-06-18","serverInfo":{"name":"factrail","version":"2.4.1"}}}
The server is stateless, so tools/list works straight away with the same headers and {"jsonrpc":"2.0","id":2,"method":"tools/list"}. It returns 7 tools. Four are the current ones:
| Tool | Title | Description starts with |
|---|---|---|
factrail_capabilities |
FACTRAIL Capabilities | "Use when you need to discover what FACTRAIL can verify or assess before calling another tool." |
factrail_verify |
Verify French Company | "Use when you need verified, source-backed facts about a French company from its SIREN or SIRET." |
factrail_assess |
Assess EU Import (Beta) | "Use when you need a source-backed assessment of importing a product into the EU…" |
factrail_get_receipt |
Get Evidence Receipt | "Use when you need to re-fetch or audit a previously returned FACTRAIL result by its receipt ID." |
The other three (verify_french_company, assess_import, analyze_company) are older compatibility tools. Their titles start with [Deprecated] and their descriptions point to the replacement. Everything is annotated readOnlyHint: true.
This matters for agents: the "Use when…" first sentence is what a model reads when it decides which tool to call, and the [Deprecated] marker steers it away from the legacy tools. The same metadata is published as a server card at https://mcp.factrail.online/.well-known/mcp/server-card.json, so directories and clients can read it without opening an MCP session.
Python, official MCP SDK (tested)
Tested with mcp 2.2.0 on Python 3.13:
python3 -m venv .venv && . .venv/bin/activate && pip install mcp
import asyncio, json
from mcp import ClientSession
from mcp.client.streamable_http import streamable_http_client
URL = "https://mcp.factrail.online/mcp"
async def main():
async with streamable_http_client(URL) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
caps = await session.call_tool("factrail_capabilities", {})
ver = await session.call_tool("factrail_verify", {
"subject_type": "company_fr",
"identifier": "552081317",
"fields": ["status", "legal_name", "legal_form",
"head_office", "naf_code", "creation_date"],
})
envelope = json.loads(ver.content[0].text) # also in ver.structuredContent
rec = await session.call_tool("factrail_get_receipt",
{"receipt_id": envelope["receipt_id"]})
print(json.dumps(envelope, indent=2))
asyncio.run(main())
(Older 1.x releases of the SDK call the helper streamablehttp_client, and it yields three values. Check your installed version.)
Claude Desktop through the mcp-remote bridge (bridge tested)
{
"mcpServers": {
"factrail": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://mcp.factrail.online/mcp"]
}
}
}
We tested the bridge itself (mcp-remote 0.14.3, driven over stdio): initialize and tools/list came back with all 7 tools. Startup took about 16 seconds, because the bridge first probes for OAuth, which this server doesn't use.
Claude Code, Cursor, VS Code (checked against official docs)
Claude Code:
claude mcp add --transport http factrail https://mcp.factrail.online/mcp
Cursor, in .cursor/mcp.json (project) or ~/.cursor/mcp.json (global):
{ "mcpServers": { "factrail": { "url": "https://mcp.factrail.online/mcp" } } }
VS Code, in .vscode/mcp.json:
{ "servers": { "factrail": { "type": "http", "url": "https://mcp.factrail.online/mcp" } } }
Browser-based clients
Browser-based MCP clients work too: the endpoint accepts requests from any Origin and answers CORS preflight requests. We checked this with a request carrying an Origin header and a preflight OPTIONS request, and both came back 200 with access-control-allow-origin: *.
4. The agent loop: capabilities, verify, handle the gaps, keep the receipt
Step 1: ask what's covered. factrail_capabilities returns the scope in machine-readable form. In our run it listed two capabilities:
-
company_fr(operationverify): sources "INSEE Sirene" and "BODACC", limitation "French companies only". -
import(operationassess):official_taric.statusis"not_installed". The tool itself is titled "Assess EU Import (Beta)".
An agent that reads this first knows not to use the tool for a German company, and not to treat any import result as official TARIC data.
Step 2: verify. We asked for six fields about SIREN 552081317, Électricité de France (EDF), a large public company.
Step 3: pass on the gaps, don't fill them. This is the part that belongs in your agent code or prompt. Here's a small helper that turns an envelope into text the model can quote, with an explicit "do not guess" line for every unresolved field:
def summarize_for_agent(envelope: dict) -> str:
"""Turn an EvidenceEnvelope into text an agent can pass on without filling gaps."""
sources = {e["id"]: e for e in envelope["evidence"]}
lines = [f"Overall: {envelope['status']} (coverage: {envelope['coverage']['level']})"]
for fact in envelope["facts"]:
src = ", ".join(
f"{sources[i]['publisher']} @ {sources[i]['retrieved_at']}" for i in fact["evidence_ids"]
)
lines.append(f"- {fact['field']}: {fact['value']} [{fact['support_level']}; {src}]")
reasons = envelope["coverage"].get("metadata", {}).get("unresolved_field_reasons", {})
for field in envelope["coverage"]["fields_unresolved"]:
lines.append(f"- {field}: UNKNOWN ({reasons.get(field, 'no reason given')}). Do not guess.")
if envelope["conflicts"]:
lines.append(f"- {len(envelope['conflicts'])} source conflict(s): report them, don't pick one silently.")
lines.append(f"Receipt: {envelope['receipt_id']}")
return "\n".join(lines)
Run against the envelope below, it prints:
Overall: insufficient_evidence (coverage: partial)
- status: active [authoritative; INSEE Sirene 3.11 @ 2026-09-29T05:18:42.087910Z]
- legal_name: ELECTRICITE DE FRANCE [authoritative; INSEE Sirene 3.11 @ 2026-09-29T05:18:42.087910Z]
- head_office: {'line_1': '22 AVENUE DE WAGRAM 22-30', 'line_2': None, 'locality': 'PARIS', 'postal_code': '75008', 'country': 'FR'} [authoritative; INSEE Sirene 3.11 @ 2026-09-29T05:18:42.249656Z]
- naf_code: 35.11Z [authoritative; INSEE Sirene 3.11 @ 2026-09-29T05:18:42.087910Z]
- creation_date: 1955-01-01 [authoritative; INSEE Sirene 3.11 @ 2026-09-29T05:18:42.087910Z]
- legal_form: UNKNOWN (unsupported_fact). Do not guess.
Receipt: fr_952a43c43fba95f2ceeef5040132dd3206c4fe6a7dd0c7b5e1a8140a32012cef
Step 4: keep the receipt. Store receipt_id next to whatever the agent decided. Any agent can re-fetch the same observation later with factrail_get_receipt.
5. Walkthrough of the real output
This envelope was captured on 29/09/2026 (05:18 UTC) from the hosted server, then at version 2.4.0. We re-ran the same call on 2.4.1 later that morning and got the same format and values; only generated_at changed. It's trimmed only where marked …:
{
"schema_version": "1.2",
"status": "insufficient_evidence",
"subject": {
"type": "company_fr",
"name": "ELECTRICITE DE FRANCE",
"identifiers": { "siren": "552081317", "siret": "55208131766522" }
},
"facts": [
{ "field": "status", "value": "active",
"evidence_ids": ["insee-siren"], "provenance_type": "sourced", "support_level": "authoritative", … },
{ "field": "legal_name", "value": "ELECTRICITE DE FRANCE",
"evidence_ids": ["insee-siren"], "support_level": "authoritative", … },
{ "field": "head_office",
"value": { "line_1": "22 AVENUE DE WAGRAM 22-30", "line_2": null,
"locality": "PARIS", "postal_code": "75008", "country": "FR" },
"evidence_ids": ["insee-siret"], "support_level": "authoritative", … },
{ "field": "naf_code", "value": "35.11Z",
"evidence_ids": ["insee-siren"], "support_level": "authoritative", … },
{ "field": "creation_date", "value": "1955-01-01",
"evidence_ids": ["insee-siren"], "support_level": "authoritative", … }
],
"evidence": [
{ "id": "insee-siren", "source_type": "government_registry",
"authority_class": "primary_official_registry", "publisher": "INSEE Sirene 3.11",
"url": "https://api.insee.fr/api-sirene/3.11/siren/552081317",
"retrieved_at": "2026-09-29T05:18:42.087910Z", "source_status": "available", … },
{ "id": "insee-siret", "publisher": "INSEE Sirene 3.11",
"url": "https://api.insee.fr/api-sirene/3.11/siret/55208131766522",
"retrieved_at": "2026-09-29T05:18:42.249656Z", "source_status": "available", … },
{ "id": "bodacc", "source_type": "government_bulletin",
"authority_class": "official_publication", "publisher": "DILA BODACC",
"retrieved_at": "2026-09-29T05:18:42.318612Z", "source_status": "available",
"metadata": { "source_error": null, "truncated": true }, … }
],
"conflicts": [],
"coverage": {
"level": "partial",
"fields_requested": ["status", "legal_name", "legal_form", "head_office", "naf_code", "creation_date"],
"fields_resolved": ["status", "legal_name", "head_office", "naf_code", "creation_date"],
"fields_unresolved": ["legal_form"],
"metadata": {
"unresolved_field_reasons": { "legal_form": "unsupported_fact" },
"authority_policies": [
{ "policy_id": "company_fr.status.current_registry_over_publication",
"precedence": ["primary_official_registry", "official_publication"],
"reason": "INSEE/Sirene reports current administrative status; BODACC notices are historical publication evidence.", … },
…
]
}
},
"freshness": {
"generated_at": "2026-09-29T05:18:42.961558Z",
"oldest_supporting_retrieved_at": "2026-09-29T05:18:42.087910Z",
"newest_supporting_retrieved_at": "2026-09-29T05:18:42.318612Z",
"stale": false,
"recheck_after": null
},
"receipt_id": "fr_952a43c43fba95f2ceeef5040132dd3206c4fe6a7dd0c7b5e1a8140a32012cef",
"state_fingerprint": "fs_6a25339254aa8ac976fa5987c691fa5b42bcfde8269457d850de135a04dd575e",
"generated_at": "2026-09-29T05:18:42.961558Z"
}
How to read it:
-
Five facts, each with a source. Status, legal name, head office, NAF code and creation date are all
authoritative, and each points to an evidence entry (insee-sirenorinsee-siret) with a URL and a retrieval time. -
One field it wouldn't invent. We asked for
legal_form. It's listed underfields_unresolvedwith the reasonunsupported_fact, and there's no value for it. Because one requested field is missing, the top-levelstatusisinsufficient_evidenceandcoverage.levelispartial, rather than a success that quietly skips a field. An agent can pass that gap on to the user as it is.-
The source hierarchy is explicit.
authority_policiesstate that for current status and legal name, the INSEE register takes precedence over BODACC publications. BODACC was queried (source_status: "available",truncated: true), but no fact in this envelope was taken from it, andconflictsis empty.
-
The source hierarchy is explicit.
-
Freshness. All the evidence was retrieved within about 0.25 seconds, and
staleisfalse.
6. Receipts: integrity, not truth
We then called factrail_get_receipt with the returned receipt_id. It returned the same envelope, with the same receipt ID and the same state_fingerprint, and no integrity error.
What a receipt gives you:
- The ID is a hash of the envelope's canonical content. When you fetch it, the server recomputes the hash and rejects content that has been altered.
- A second agent, or you next week, can check exactly what was observed and when.
What it does not give you: a digital signature, a blockchain record, or proof that the facts are true. It proves the envelope hasn't changed since it was produced. Whether INSEE was right is a separate question.
7. Limits, fair use and current coverage
As of 29/09/2026 (server 2.4.1, public beta):
-
French company verification (
company_fr, from Sirene and BODACC) works. It covers French companies only, and on its own it isn't a complete KYC/KYB compliance check. -
Import assessment (
factrail_assess,import) is beta. Curated tariff and VAT values stay marked curated or provisional. -
Official EU TARIC data is not installed (
not_installed), so there's no authoritative tariff coverage. - Not provided: general fact-checking, web search, global company data, sanctions or beneficial-ownership checks, signed receipts.
Fair use of the public endpoint:
FACTRAIL's public MCP endpoint (https://mcp.factrail.online/mcp) is free, read-only, and needs no account. To keep it fair for everyone, each client IP may make up to 60 MCP requests per minute, while traffic arriving through known MCP gateways such as Smithery shares a separate allowance of 600 requests per minute. Requests over the limit receive HTTP 429 with a
Retry-Afterheader, and clients that keep sending requests after being limited are paused for 5 minutes. If you need more capacity, please get in touch via https://factrail.online/ rather than working around the limits.
In practice: an agent that gets a 429 should wait for the number of seconds in Retry-After before retrying, instead of looping.
On Apify: if your agents run on Apify, the Actor spherical_distinction/factrail-evidence is a thin wrapper around this same endpoint, billed per event: $0.005 per successful verification and $0.03 per successful import assessment (Apify pricing checked on 29/09/2026). The direct MCP endpoint stays free.
Call factrail_capabilities first: it reports the live scope, which may be newer than this article.
8. Self-hosting
The server is a Python package. From the repository README:
git clone https://github.com/baronsigma/factrail && cd factrail
python3 -m venv .venv && . .venv/bin/activate
pip install -e '.[dev]'
cp .env.example .env
# Set INSEE_API_KEY in .env for live French company lookups
python3 -m factrail.mcp_http_server # POST http://localhost:8765/mcp
Notes from the README:
- You need your own INSEE API key (portail-api.insee.fr).
-
FACTRAIL_HOSTandFACTRAIL_PORTchange the bind address and port. Behind a public hostname, setFACTRAIL_ALLOWED_HOSTS; by default only localhost is accepted. -
FACTRAIL_ALLOWED_ORIGINScontrols browser access. When it's unset, requests that carry anOriginheader are rejected;*allows any Origin with CORS (this is what the hosted endpoint does). - Rate limits are configurable (
FACTRAIL_RATE_LIMIT_PER_MIN, default 60, and related variables). - Set
FACTRAIL_CACHE_PATHto a persistent path if you want receipts to survive restarts (the default is/tmp). - For stdio clients:
python3 -m factrail.mcp_server.
Takeaways
- Give your agent tools that return evidence, not answers: source, support level, freshness and coverage for every fact.
- Treat "unresolved" as a real result, and write your agent so it says so. An envelope that says
insufficient_evidenceis more useful than a guessed value. - Call a capabilities tool first, prefer tools whose descriptions say when to use them, and keep receipts for anything you might need to re-check.
Links: factrail.online · github.com/baronsigma/factrail · endpoint https://mcp.factrail.online/mcp
Top comments (1)
The part I'd add is a measurement of step 3, because the envelope only helps if the agent actually passes the gap on. On our FilingFacts benchmark (questions about SEC filings, some deliberately unanswerable), gpt-5.4-mini declined 93.3% of the unanswerable items closed-book, but only 76.7% once it had a tool that reads the filings. Having a source made it more willing to produce a number the data couldn't support. The run records are public: huggingface.co/datasets/arhancanli...
So I'd test the "Do not guess" line the same way: a set of requests where one field is always unresolved, and a count of how often the final answer still states a value for it. That number tells you whether the envelope's honesty survives the model.
A smaller design question: with 5 of 6 fields authoritative, the top-level
statusisinsufficient_evidence. An agent that branches onstatusalone will throw away five good facts. Is the intent that agents always readcoveragefirst, or would a per-field status at the top level be clearer?