This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
A clinical trial discovery and safety verification agent powered by Sanity Context MCP and deterministic GROQ queries.
What I Built
A few months ago, a friendβs relative was diagnosed with advanced non-small cell lung cancer. First-line platinum chemotherapy had stopped working. The oncologist said something that stuck with me: "Our best shot right now is an experimental targeted trial. Go home, search ClinicalTrials.gov, and see what you can find."
If you have never searched ClinicalTrials.gov during a medical crisis, the experience is overwhelming. The registry contains more than 450,000 studies written in dense clinical legalese. Each protocol has dozens of eligibility criteria, negative exclusion clauses, and washout windows spread across 40-page PDF attachments.
Like most developers, my first reflex was to see if an LLM could help. I pasted the patientβs pathology summary into ChatGPT and Gemini:
58-year-old female, Stage IV NSCLC, EGFR Exon 20 insertion, progressed after carboplatin/pemetrexed, seeking recruiting Phase 2 trials in Texas.
The response looked brilliant on the surface. Fluent prose, plausible drug mechanisms, and two clinical trial IDs: NCT04847387 and NCT04746654.
When I went to verify them on official registries, my stomach dropped: neither trial existed. The model had synthesized oncology vocabulary and generated fake 8-digit numbers out of thin air. In casual creative writing, an AI hallucination is harmless. In oncology, sending a desperate family chasing a non-existent study is catastrophic.
Next, I tried standard keyword search. It dumped 69 trials on my screen matching "EGFR" and "chemotherapy." But when I sat down and read through protocol after protocol, 31 of them had explicit exclusion clauses disqualifying anyone with prior systemic chemotherapy. If a patient drove hours to Houston based on that search, they would be turned away at the clinic door. Worse, keyword search matched unrelated kidney failure trials because the kidney filtration test eGFR shares letters with the cancer oncogene EGFR.
That made the technical reality impossible to ignore:
Medical safety cannot be solved with flat text search or raw LLM prompts. Clinical eligibility depends on strict Boolean rules, exact allele variants, and line-of-therapy sequences. AI cannot guess clinical eligibilityβit needs structured ground truth.
I built TrialMatch: an open-source clinical trial discovery and safety verification agent that connects an AI extraction agent to a structured Sanity Content Lake using the Sanity Context MCP.
Instead of letting an LLM guess against fuzzy text chunks or hallucinate trials, TrialMatch uses the model strictly to parse unstructured clinical notes into clean filter parameters. Those parameters hydrate deterministic GROQ queries that run against structured Sanity schemas and protocol guidance rules stored in a Sanity Knowledge Base.
System Architecture
sequenceDiagram
autonumber
actor Clinician as Clinician / Patient
participant UI as TrialMatch Web App (Next.js 16)
participant Agent as Claude Haiku 4.5 (Schema Extractor)
participant MCP as Sanity Context MCP (trialmatch)
participant Lake as Sanity Content Lake (Project 6xsr2k42)
participant KB as Sanity Knowledge Base (trialmatch-kb)
Clinician->>UI: Enter Patient Note ("58yo NSCLC, EGFR Exon 20, prior chemo, TX")
UI->>Agent: Extract Clinical Primitives
Note over Agent: Converts narrative into:<br/>condition="Lung", biomarker="EGFR",<br/>chemo="ALLOWED", state="TX"
Agent->>MCP: Dispatch GROQ Query with Bound Parameters
MCP->>KB: Check Washout & Hierarchy Rules
KB-->>MCP: Rule Verified (Chemo permitted post-progression)
MCP->>Lake: Execute Deterministic GROQ Filter
Lake-->>MCP: Return 2 Verified Recruiting Protocols
MCP-->>UI: Return Match Set + Full GROQ Audit Trace
UI-->>Clinician: Render Verified Cards + Side-by-Side Hazard Analysis

The Duel Arena in action: Live Sanity Context MCP status beacon, real-time safety scorecard, and 8 clinical scenario presets.
Demo
- Live Application: trialmatch-oncology.vercel.app
- 3-Arm Benchmark Dashboard: trialmatch-oncology.vercel.app/benchmark
- Hosted Sanity Studio: trialmatch-oncology.sanity.studio
The Duel Arena: Side-by-Side Reality
To prove the difference between structured content and naive search, I built the app as a real-time side-by-side comparison. You can enter any custom patient profile or pick from eight pre-configured clinical scenarios:
- Left Column (Sanity Context Agent): Returns only verified, recruiting trials that strictly pass every exclusion check. Each card identifies the target biomarker match, lists confirmed hospital sites, and features an expandable drawer displaying the exact GROQ query executed against Sanity.
- Right Column (Naive Keyword Baseline): Shows what traditional keyword search returns. A red warning banner breaks down the safety violations, highlighting why specific trials are clinically disqualified (such as prior chemotherapy prohibitions or active liver metastases).
-
Column-Level Pagination: When keyword search returns 69 studies and the Sanity agent returns 2, each column paginates independently (6 trials per page) with centered navigation controls (
< Page 1 of 12 >). This keeps the comparison clean and readable without infinite scrolling.

Inspecting the generated GROQ query on the left alongside the naive search failure warnings on the right.

Independent column pagination keeps 69 keyword matches manageable while highlighting verified hospital locations.
Code
- GitHub Repository: github.com/IshekKhal/trialmatch
Project Layout
trialmatch/
βββ app/
β βββ page.tsx # The Duel Arena (Sanity Agent vs Keyword Baseline)
β βββ benchmark/page.tsx # 3-Arm Evaluation Dashboard
β βββ layout.tsx # Root layout with Geist font tokens
β βββ globals.css # Medical obsidian styling (#080C14)
β βββ api/
β βββ duel/route.ts # Live duel execution endpoint
β βββ benchmark/route.ts # 3-arm benchmark evaluation API
β βββ trials/route.ts # Normalized clinical trial fetcher
βββ components/
β βββ DuelArena.tsx # Dual-column comparison with column-level pagination
β βββ ScenarioChips.tsx # 8 pre-configured oncology clinical scenarios
β βββ BenchmarkRunner.tsx # Interactive benchmark runner & case inspector
β βββ BenchmarkTable.tsx # 10-patient audit breakdown table
β βββ Scoreboard.tsx # Real-time comparative metric display
βββ studio/
β βββ sanity.config.ts # Sanity Studio v3 configuration
β βββ schemas/
β βββ clinicalTrial.ts # Core schema: biomarkers, therapy rules, facilities
β βββ protocolRule.ts # Knowledge base schema: exclusion overrides & washouts
β βββ index.ts # Schema registry
βββ lib/
β βββ agent.ts # Claude Haiku 4.5 schema extractor & GROQ builder
β βββ sanity.ts # Sanity client & live Context MCP bindings
β βββ eval_cases.ts # 10 gold-standard oncology test profiles
β βββ eval_runner.ts # Automated 3-arm benchmark execution engine
β βββ types.ts # TypeScript domain interfaces
βββ scripts/
β βββ ingest_trials.ts # ClinicalTrials.gov API v2 data ingestion pipeline
β βββ run_eval.ts # CLI benchmark evaluation runner
βββ data/
βββ trials_normalized.json # 100 curated oncology trials (1.2 MB)
βββ eval_results.json # Raw 3-arm benchmark evaluation outputs
βββ eval_summary.md # Comprehensive 419-line benchmark report
How I Used Sanity
1. Modeling Oncology Beyond Text Blobs
Oncology protocols cannot live in unstructured markdown or string fields. A single protocol contains dozens of pharmacogenomic, line-of-therapy, and clinical status constraints.
In Sanity Studio (studio/schemas/clinicalTrial.ts), I modeled trials into structured primitives:
-
targetBiomarkers: Typed string array (EGFR,KRAS,BRAF,HER2,BRCA1,Exon 20,G12C,V600E). -
priorTherapyRules: An object with explicit status enums (REQUIRED,ALLOWED,EXCLUDED,ANY) for chemotherapy, immunotherapy, and targeted therapies. -
locations: Array of facility objects with city, state, and hospital names across 1,144 US sites. -
eligibilityCriteria: Structured inclusion and exclusion bullet points.
Clinical Trial Protocol Schema (studio/schemas/clinicalTrial.ts)
import {defineField, defineType} from 'sanity'
export default defineType({
name: 'clinicalTrial',
title: 'Clinical Trial',
type: 'document',
fields: [
defineField({
name: 'nctId',
title: 'NCT ID',
type: 'string',
validation: (Rule) => Rule.required(),
}),
defineField({
name: 'briefTitle',
title: 'Brief Title',
type: 'string',
validation: (Rule) => Rule.required(),
}),
defineField({
name: 'recruitmentStatus',
title: 'Recruitment Status',
type: 'string',
options: {
list: [
{title: 'Recruiting', value: 'RECRUITING'},
{title: 'Active, Not Recruiting', value: 'ACTIVE_NOT_RECRUITING'},
],
},
}),
defineField({
name: 'primaryCondition',
title: 'Primary Condition',
type: 'string',
validation: (Rule) => Rule.required(),
}),
defineField({
name: 'targetBiomarkers',
title: 'Target Biomarkers',
type: 'array',
of: [{type: 'string'}],
options: {
list: ['EGFR', 'KRAS', 'BRAF', 'HER2', 'ALK', 'BRCA1', 'BRCA2', 'Exon 20', 'G12C', 'V600E'],
},
}),
defineField({
name: 'priorTherapyRules',
title: 'Prior Therapy Rules',
type: 'object',
fields: [
defineField({
name: 'chemotherapy',
type: 'string',
options: { list: ['REQUIRED', 'ALLOWED', 'EXCLUDED', 'ANY'] },
}),
defineField({
name: 'immunotherapy',
type: 'string',
options: { list: ['REQUIRED', 'ALLOWED', 'EXCLUDED', 'ANY'] },
}),
],
}),
defineField({
name: 'locations',
title: 'Trial Locations',
type: 'array',
of: [{
type: 'object',
fields: [
{ name: 'facility', type: 'string' },
{ name: 'city', type: 'string' },
{ name: 'state', type: 'string' },
{ name: 'country', type: 'string' },
],
}],
}),
],
})
To handle tricky edge cases (such as whether prior chemotherapy represents an absolute contraindication or simply requires a 30-day washout period), I added a companion schema, protocolRule.ts, indexed in the Sanity Knowledge Base:
Protocol Guidance Rule Schema (studio/schemas/protocolRule.ts)
import {defineField, defineType} from 'sanity'
export default defineType({
name: 'protocolRule',
title: 'Clinical Protocol Interpretation Rule',
type: 'document',
fields: [
defineField({ name: 'ruleId', title: 'Rule ID', type: 'string' }),
defineField({ name: 'title', title: 'Rule Title', type: 'string' }),
defineField({
name: 'category',
type: 'string',
options: {
list: [
{title: 'Eligibility Hierarchy', value: 'ELIGIBILITY_HIERARCHY'},
{title: 'Washout Periods', value: 'WASHOUT_PERIODS'},
{title: 'Biomarker Specificity', value: 'BIOMARKER_SPECIFICITY'},
],
},
}),
defineField({ name: 'priority', title: 'Priority (1-10)', type: 'number' }),
defineField({ name: 'summary', title: 'Rule Summary', type: 'text' }),
defineField({ name: 'clinicalRationale', title: 'Clinical Rationale', type: 'text' }),
],
})

Sanity Studio managing 100 oncology protocols with explicit prior therapy rules, biomarker tags, and multi-center site registries.
2. Sanity Context MCP Endpoints
We connected our Sanity Content Lake through two dedicated Sanity Context MCP endpoints:
-
GROQ Context Endpoint (
trialmatch): Exposes the live dataset for parameterized GROQ queries. -
Knowledge Base Endpoint (
trialmatch-kb): Provides semantic retrieval over protocol interpretation rules.
When a patient note arrives, the agent translates the clinical parameters into this exact GROQ query:
*[_type == "clinicalTrial" &&
recruitmentStatus == "RECRUITING" &&
primaryCondition match $condition &&
targetBiomarkers[] match $biomarker &&
(priorTherapyRules.chemotherapy == "ALLOWED" ||
priorTherapyRules.chemotherapy == "REQUIRED" ||
priorTherapyRules.chemotherapy == "ANY") &&
locations[].state match $state] {
nctId,
briefTitle,
phase,
primaryCondition,
targetBiomarkers,
priorTherapyRules,
"matchingLocations": locations[state match $state]
}
Because priorTherapyRules.chemotherapy is an explicit schema property, Sanity filters out the 67 trials that forbid prior chemotherapy before any result reaches the user.
Behind the Build: 4 Traps I Ran Into π₯
Building this was a wild ride between medical literature and software architecture. Here are four real hurdles that cost me hours (and how structured content solved them):
-
Negative exclusions break standard search completely: Keyword search and vector search are fundamentally built to find words that are there. But clinical trial eligibility is mostly about words that forbid you from entering. A trial that says "No prior chemotherapy within 6 months" will happily match a search for "prior chemotherapy". By turning negative exclusions into explicit Sanity schema properties (
ALLOWED,REQUIRED,EXCLUDED), GROQ executes strict Boolean disqualifications deterministically. -
The
eGFRvsEGFRsemantic collision: In standard search, kidney filtration tests (eGFR) kept polluting the cancer oncogene (EGFR) results because they share characters. In Sanity,targetBiomarkersis a typed string array. A kidney test never enters the oncogene array, eliminating false matches entirely. -
Why the agent is an extractor, not a query writer: Giving an LLM free rein to write arbitrary database queries against an oncology database is dangerous. It can hallucinate fields, botch filters, or drop exclusions. Instead, TrialMatch constrains the AI model to one specific job: extracting patient narrative primitives into typed criteria (
condition,biomarker,priorTherapy,state). Our code then maps those verified primitives into deterministic, parameterized GROQ queries. -
Serverless read-only filesystem traps on Vercel: Our local benchmark evaluation script saved audit logs to disk (
fs.writeFileSync). When deployed to Vercel's serverless edge, writing to/var/taskthrew 500 errors. Refactoring the benchmark runner to keep the normalized dataset in memory and compute real-time diffs on the fly solved the issue instantly.
The 3-Arm Benchmark: Empirical Findings
To objectively measure whether structured content changes clinical discovery outcomes, I built an automated evaluation suite (lib/eval_runner.ts) and tested ten gold-standard oncology profiles across three discovery approaches:
| Evaluation Metric | Arm 1: Structured Sanity Agent | Arm 2: Naive Keyword Search | Arm 3: Bare LLM (Gemini 3.8 Flash) | Clinical Implication |
|---|---|---|---|---|
| Medical Precision | 100% | 78% | 0% | Arm 1 returns only verified candidates; Arms 2 and 3 return disqualified cohorts |
| Safety Violations | 0 | 60 | 20 | Naive search misses negative exclusions; bare LLM bypasses protocol rules |
| Hallucinated NCT IDs | 0 | 0 | 20 (100% fake) | Bare LLM invents non-existent trial identifiers |
| Auditability Rate | 100% | 0% | 0% | Arm 1 provides exact GROQ queries and rule citations |
| Avg Returned Trials | 1.2 | 41.5 | 2.0 | Arm 1 isolates actionable, recruiting matches |

The 3-Arm Benchmark Dashboard showing 100% medical precision and 0 hallucinations for the Sanity agent versus 60 safety violations in keyword search and 100% fake NCT IDs from the bare LLM.
Real Clinical Failure Modes
Examining individual patient cases demonstrates where unstructured search breaks down:
1. Chemotherapy Exclusion (Case TC-01)
- Patient: 58yo female, NSCLC EGFR Exon 20 insertion, prior platinum chemotherapy, Texas.
-
Keyword Failure: Returned 69 trials. 31 violated safety rules. Trial
NCT07799935specifically prohibits prior systemic chemotherapy in Rule 14. -
Bare LLM Failure: Hallucinated trials
NCT04847387andNCT04746654. Neither identifier exists. -
Sanity Result: Executed GROQ checking
priorTherapyRules.chemotherapy in ["ALLOWED", "REQUIRED", "ANY"]. Safely returned exactly 2 verified trials.
2. Organ Site Metastasis Contraindications (Case TC-02)
- Patient: 62yo male, Stage IV colorectal cancer, KRAS G12C mutation, stable liver metastases, California.
-
Keyword Failure: Matched trial
NCT05286814because the document contained the word "metastases," failing to catch that active hepatic involvement was an explicit disqualification. -
Bare LLM Failure: Hallucinated trials
NCT04793958andNCT04685141. - Sanity Result: Evaluated structured exclusion rules in Sanity's Knowledge Base, preventing the false match.
3. Quantitative Biomarker Expression Cutoffs (Case TC-03)
- Patient: 51yo female, HER2-low (IHC 1+ or IHC 2+/FISH negative) metastatic breast cancer.
-
Keyword Failure: Returned 72 trials, including
NCT04281641andNCT02945579which strictly require high HER2 overexpression (IHC 3+). - Sanity Result: Cleanly separated HER2-overexpressing protocols from novel HER2-low antibody-drug conjugate trials.
4. Phase 3 Confirmatory Requirement (Case TC-04)
- Patient: 47yo patient, unresectable Stage IIIC melanoma, BRAF V600E mutation, seeking Phase 3 trials in New York.
- Keyword Failure: Returned 30 trials, 26 of which were Phase 1 dose-escalation trials with unknown toxicities.
-
Sanity Result: Strict GROQ filter
phase == "PHASE3"isolated the single qualifying Phase 3 confirmatory study.

Detailed case audit: inspecting why trial NCT07799935 violates chemotherapy rules, while Gemini invents fictional identifiers.
Sanity Project Details
-
Project ID:
6xsr2k42 -
Dataset:
production - Hosted Studio: trialmatch-oncology.sanity.studio
-
Sanity GROQ MCP Endpoint:
https://api.sanity.io/v1/context/organizations/oz3yptidu/mcp/trialmatch -
Sanity Knowledge Base MCP Endpoint:
https://api.sanity.io/v1/context/organizations/oz3yptidu/mcp/trialmatch-kb - Public Sample Query (Live JSON): Live 3-Trial Sanity Query Endpoint
Agent Session
Below is a representative excerpt from the agent's interaction loop with Sanity Context MCP:
Clinician:
"We have a 58yo patient with NSCLC harboring an EGFR Exon 20 insertion.
She progressed on carboplatin/pemetrexed. Find recruiting Phase 2 trials
in Texas that permit prior chemotherapy."
Agent (Claude Haiku 4.5):
Extracting clinical parameters:
- Condition: "Lung" (NSCLC)
- Target Biomarker: "EGFR", "Exon 20"
- Chemotherapy: "ALLOWED" or "REQUIRED"
- Location: "TX"
- Recruitment Status: "RECRUITING"
Tool Call -> Sanity Context MCP (trialmatch):
GROQ: *[_type == "clinicalTrial" && recruitmentStatus == "RECRUITING"
&& primaryCondition match "Lung"
&& targetBiomarkers[] match "EGFR"
&& (priorTherapyRules.chemotherapy == "ALLOWED" ||
priorTherapyRules.chemotherapy == "REQUIRED")
&& locations[].state match "TX"]{nctId, briefTitle, phase}
MCP Response:
[
{
"nctId": "NCT05376891",
"briefTitle": "Phase 2 Study of Targeted EGFR Exon 20 Inhibitor",
"phase": "PHASE2"
},
{
"nctId": "NCT06234137",
"briefTitle": "Targeted Kinase Therapy for Relapsed EGFR Mutations",
"phase": "PHASE2"
}
]
Verification: 2 active matches found. 0 safety exclusions violated. 0 hallucinations.
What This Demonstrates
The software industry has spent two years hoping that larger foundation models or clever prompting would eliminate hallucinations. But in medicine, an almost-right answer is as dangerous as a total failure. A patient either qualifies under a protocol's strict criteria, or they do not.
Connecting an AI agent to structured content in Sanity changes the model's role. It stops guessing against text blobs and acts as a deterministic schema parser. The model extracts the patient's clinical situation, Sanity executes the Boolean protocol rules, and the patient receives verified, life-saving options.
Structured content is not merely a CMS pattern. In high-stakes domains, structured content is the safety anchor AI cannot function without.
Try It Yourself
Explore the live duel or run the benchmark yourself:
- π Live Arena: trialmatch-oncology.vercel.app
- π Live Benchmark: trialmatch-oncology.vercel.app/benchmark
- π» GitHub: github.com/IshekKhal/trialmatch
If you have feedback, ideas for clinical trial filtering, or experience with Sanity MCP in production, I would love to hear your thoughts in the comments below!
Top comments (1)
I built this because flat keyword search in clinical trials leads to dangerous safety blind spots. How are you handling deterministic guardrails for LLM agents in your own production architectures?