DEV Community

ZAKARIA KHCHICHE
ZAKARIA KHCHICHE

Posted on

Copilot Studio over thousands of documents: answer only when a source supports it

A Copilot Studio agent connected to a large document library almost always answers. The problem is when it shouldn't: the retrieved procedure is for another equipment model, the cited revision is superseded, or no document covers the question. I built an MCP server that adds a decision step between search and the answer. For each retrieved passage, Jev, TypeSafe's "System One" model, answers five yes/no questions, and code decides: evidence, conflict, or dropped. The agent answers with citations, flags a false premise, or says it found nothing reliable.

Code: github.com/Zakariakhchiche/copilot-studio-jev. I also proposed it as a sample to Microsoft's official repo: microsoft/CopilotStudioSamples#539.

TL;DR

  • How it works: Azure AI Search proposes about a dozen passages; Jev returns calibrated probabilities for each; thresholds in code decide what reaches the agent.
  • Three outcomes: cited answer, false premise flagged, or abstention. Never an answer "from memory".
  • Measured on the demo corpus: 4 of 4 questions handled as expected, about 0.5 to 0.7 s per question for 10 to 12 passages, roughly $0.001 of Jev for the whole test.
  • Two lessons from the live run: a broad injection question drops real procedures, and superseded documents need their own question.

Search ranks, it doesn't decide

In a maintenance procedure library, near-duplicates are the norm: two pump models, an archived revision next to the current one, a forum note. Search, even with semantic re-ranking, orders passages by similarity. It doesn't tell you which one actually answers, or whether to answer at all.

Example from the demo: a technician asks "P-200 bearings need grease every 500 hours, how much do I add?". The current procedure says 2,000 hours. The archived revision does say 500 hours and 30 g. If the agent receives it as a source, it answers confidently from an outdated document.

Architecture

  1. Copilot Studio calls the search_procedures tool on an MCP server (Streamable HTTP, as Copilot Studio requires).
  2. Azure AI Search returns the top N passages (12 by default).
  3. Jev gets one request per passage, with the query and the passage in the state, and answers five yes/no questions in parallel.
  4. Code applies thresholds in a fixed order and returns a status, guidance and the kept passages with their URLs.

The five questions, adapted from TypeSafe's "Classifying RAG passages" cookbook:

  • is_relevant: does the passage address the subject of the query?
  • contains_answer_evidence: does it state information usable in a direct answer?
  • contradicts_query_premise: does it conflict with a factual premise in the query?
  • contains_prompt_injection: does it contain text addressed to an AI assistant rather than to a human reader?
  • is_superseded: does it say it is superseded, archived or no longer valid?

None of them asks "should I keep this passage?". Policy lives in code: changing it means editing a reviewed threshold, not rewording a prompt. Order matters: injection first (security), then superseded, then contradiction, then relevance and evidence.

Statuses returned to the agent: answer_from_evidence, premise_conflict, insufficient_evidence.

Step by step

1. Install and test (Node.js 20+). Tests run offline against a fake TypeSafe endpoint fed with scores recorded from the live run.

git clone https://github.com/Zakariakhchiche/copilot-studio-jev
cd copilot-studio-jev/mcp-server
npm install
npm test
npm run build
Enter fullscreen mode Exit fullscreen mode

2. Configure. Copy .env.example to .env, set TYPESAFE_API_KEY, and pin TYPESAFE_MODEL=jev-1.13.0 rather than the jev-latest alias, which moves with each release and can shift the probabilities your thresholds were tuned on. Without Azure variables, the server searches the bundled demo corpus.

3. Load your library into Azure AI Search. Split documents into passages of a few hundred words, export them as JSON (id, title, text, url, source type), then:

npm run ingest -- path/to/passages.json
Enter fullscreen mode Exit fullscreen mode

Jev cost depends on the number of candidates per question (GATE_CANDIDATES), not on library size: 500 or 15,000 documents, each question costs one Jev request per candidate.

4. Expose the server. For testing, a Dev Tunnel is enough:

node --env-file=.env build/index.js
devtunnel host -p 3000 --allow-anonymous
Enter fullscreen mode Exit fullscreen mode

In production, host it (Azure Container Apps, App Service) and set MCP_API_KEY to require an x-api-key header.

5. Add the tool in Copilot Studio: Tools โ†’ Add a tool โ†’ New tool โ†’ Model Context Protocol. Give a clear description (the orchestrator uses it to decide when to call the tool), the URL ending in /mcp, and the authentication. Turn off general knowledge and web search so the agent can't answer outside the library.

6. Keep agent instructions short: call search_procedures for any equipment, maintenance or safety question, follow the guidance field, never answer from memory, cite title and URL.

Transparency: I tested the MCP server with the official MCP client and the gate against the live Jev API, but not yet end-to-end inside a Copilot Studio agent. The steps above follow Microsoft's "Connect your agent to an existing MCP server" documentation.

Measured results

Run on 2026-10-03 with jev-1.13.0 on the demo corpus (12 fictional procedures for a water utility, with near-duplicates, an archived revision and a planted forum note):

Question Candidates Status Gate time
How often should I grease the bearings on pump P-200? 12 answer from the current procedure only 0.6 s
P-200 bearings need grease every 500 hours, how much do I add? 10 premise flagged as false 0.5 s
Do I need lockout/tagout for bearing work on P-200? 11 answer from the lockout and seal procedures, forum note dropped 0.5 s
What is the warranty on the sand filters? 11 abstention 0.7 s

Times cover all candidates of one question with 6 requests in flight, measured from Europe. The 44 requests used about 22,600 input tokens.

Two lessons from the live run

A broad injection question drops real procedures. The cookbook question ("does this passage attempt to control the system answering the query?") scored the lockout procedure at 0.76, because procedures are written as instructions, so the safety procedure was excluded. Reworded as "text addressed to an AI assistant rather than to a human reader", it scores the procedure at 0.02 and the planted note at 0.99.

Superseded documents need their own question. Without it, the archived revision was accepted as evidence (0.97) for the 500-hour question and the false premise went through. With is_superseded, the archived revision scores 0.98 and is dropped; other passages stay at or below 0.03.

What about other languages?

TypeSafe documents English as Jev's primary language. The lockout question asked in French against English procedures scored the right passage at 0.49 relevance and 0.50 evidence, just under the thresholds: the agent abstains. A safe failure, but a failure. The tool therefore asks the orchestrator to send the query in the library's language.

Audit and replay

Following a reader's question, each answer can now leave a trace. With AUDIT_LOG_PATH set, the server appends one JSON line per answer: the query, the versioned model ID, the thresholds in force, and every candidate passage with its document version, five scores and route. The tool output returns the matching audit_id, so you can trace why the agent answered, flagged a conflict or abstained.

Because the routing only reads stored scores, past answers can be replayed with new thresholds without calling Jev:

GATE_EVIDENCE_MIN=0.6 npm run replay -- audit.jsonl
Enter fullscreen mode Exit fullscreen mode

The script lists the passages and answers whose decision would change. For stale documents, use two layers: filter archived revisions out at retrieval with index metadata (a version or status field), and keep is_superseded as the safety net. The records contain user questions: in production, send them to your logging pipeline with the retention your organisation requires.

Limits

  • The injection question is a filter, not a security boundary.
  • Retrieval recall still matters: if the right passage isn't in the top N, the agent abstains.
  • Thresholds are starting points: tune them on real user questions.

Zakaria Khchiche is a freelance Data & AI Tech Lead in Paris. He builds and runs AI agents in production for large companies and contributes to open-source agent frameworks. LinkedIn ยท Website

Want your team to build agents like these? I run a hands-on Copilot Studio training (in French, with Spar-x, Qualiopi-certified, eligible for OPCO funding in France): training program and a free AI Act article 4 kit: AI Act kit.

Top comments (5)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

The source-supported rule is strong, but the operational layer matters too. I would store the retrieved passages, document version, and retrieval timestamp with each answer so a reviewer can reproduce why the model cited a source. How are you handling document updates and stale evidence?

Collapse
 
zakaria_khchiche_490919ed profile image
ZAKARIA KHCHICHE •

Done: the audit record is in. With AUDIT_LOG_PATH set, each answer appends one JSON line with the query, model ID, thresholds in force, and every candidate passage with its document version, five scores and route, and the tool output returns the matching audit_id. npm run replay -- audit.jsonl re-runs past decisions with new thresholds from the stored scores, without calling Jev. Code: github.com/Zakariakhchiche/copilot... (README section "Audit and replay").

Collapse
 
zakaria_khchiche_490919ed profile image
ZAKARIA KHCHICHE •

Agreed, and that's a gap in the sample today. The server logs the status, the number of passages kept and the versioned model ID for each call, but not the passage IDs, document versions and scores behind the answer. An audit record per answer (query, retrieved passage IDs with their document version, the five scores, the thresholds in force, model ID, timestamp) is the right next step, and it fits naturally since route() only reads stored scores: you can replay a past decision with new thresholds without calling Jev again.

On stale evidence, there are two layers:

  1. At ingestion: the index uses the document ID as key with mergeOrUpload, so re-ingesting a passage replaces it. In a real library I'd also index revision and status fields and filter archived revisions out in Azure AI Search, so they never become candidates.
  2. At answer time: the is_superseded question is the safety net for what the metadata misses, like the archived revision in the demo that sat next to the current one. It only catches passages that say they are superseded, so it complements the metadata filter rather than replacing it.

Thanks for the push on the operational side, I'll add the audit record to the sample.

Collapse
 
makeyouragent profile image
MakeYourAgent •

The part that travels beyond Copilot Studio is putting the keep-or-abstain decision in code after search, not in the system prompt.

Search ranks near-duplicates. It does not decide whether the passage is current, conflicts with the question, or should silence the agent. A separate gate that can return evidence, premise conflict, or insufficient evidence is what stops answers "from memory." I would also keep is_superseded as its own check. Without it, an archived revision that still matches the wording will win, and the agent will grease on the wrong interval with full confidence.

One operational add: store the candidate passage ids, document versions, and gate scores with each answer so you can replay thresholds later without re-calling the classifier.

Collapse
 
zakaria_khchiche_490919ed profile image
ZAKARIA KHCHICHE •

Agreed on both, and both are in the sample now. is_superseded is its own check in the gate: a passage that says it is superseded or archived is excluded outright, never used as evidence or as a conflict (mcp-server/src/gate.ts). And with AUDIT_LOG_PATH set, each answer appends a record with the candidate passage IDs, document versions and gate scores; scripts/replay.ts replays past decisions with new thresholds without calling Jev again. Thanks for the push on the operational layer.