This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
We asked our own website 87 questions that buyers had emailed us. It fully answered only 8 of them.
What I Built
What it is. Cross-Check is an agent that reads our website through a Sanity Knowledge Base and compares it with two things: the questions prospects actually ask us, and the answers our sales team actually gave them. Where the site is silent, unclear or says something different, it files a finding with a suggested fix that an editor reviews in Sanity Studio.
How it works.
- Index. A build step turns 73 pages and case studies into plain-text kbSource documents, and Sanity Context builds a Knowledge Base from them.
- Ask. For each of the 87 real buyer questions, an agent searches and reads the Knowledge Base through Context MCP.
- Judge. A judge model compares the agent's answer with the page text and with what our sales team actually told clients, then files a verdict (gap, ambiguity, contradiction or covered) with a suggested fix.
- Review. Every finding is a Sanity document. An editor opens it in Studio, sees what buyers ask, what the site says and what sales told them side by side, and marks it real or a false alarm.
The problem. We are a technology company, so the stack moves, our offer moves, and what clients need moves with it. The content does not keep up. Some of what is published is current, some of it is two offers old, and nothing marks which is which. That mattered less when a buyer skimmed a few pages and asked us the rest. Now an agent reads the whole site more carefully than any buyer ever did, and hands the buyer a summary and a recommendation on whether to work with us.
What makes it work. Every input is something a real person wrote.
- 87 buyer questions: what prospects actually emailed us, distilled from 209 questions across 76 email threads.
- 65 sales answers: anonymised records of what our sales team replied, covering rates, SLA tiers, payment terms and contract mechanics.
- 73 pages and case studies from pagepro.co.
Ask any model to guess what buyers want to know about a development agency and it will produce a plausible list in seconds. That list describes agencies in general, so it finds nothing. These 87 questions name the things that stalled our own deals, and the 65 answers are what we are on the hook for.
What came back.
| Verdict | Count |
|---|---|
| Gap: buyers ask and the site does not answer | 57 |
| Ambiguity: the site answers two ways | 13 |
| Contradiction: the site disagrees with sales or with itself | 9 |
| Covered | 8 |
Most gaps are contract terms and process, not price: SLA severity tiers, what gates a kickoff, what happens to unused hours. The contradictions are the expensive ones. Our site publishes one hourly rate; four separate offers quoted materially higher. It lists one post-launch support price; sales quoted a different one in three currencies.
The second pass. The audit covers what the site contains. It says nothing about whether that reaches a buyer, so every question the site answers at all goes three ways: a model with no tools, the same model with web search, and the agent reading the Knowledge Base. The first two are what a buyer gets today. The third is what the right answer was. We call this pass Answer drift: it measures how far what an assistant tells a buyer drifts from what our Knowledge Base says.
| Verdict | Count |
|---|---|
| Not communicated: the model had the site in reach and still missed | 27 |
| Search misled: right from memory, wrong after searching | 6 |
| Stale priors: wrong from memory, right once it searched | 6 |
| Aligned | 6 |
| World guesses for us: our content disagrees with itself, so the models fill the gap | 2 |
Of those 27: in 9 the model with web search answered wrong, in 18 it gave no real answer and told the buyer to contact us. In 22 of the 27 it cited pages it had read and still came away without the fact.
The question prospects ask more than any other is what a WordPress to Next.js migration costs. The model with web search read our pages and told the buyer to get in touch.
Demo
Code
Pagepro
/
cross-check
An agent that checks a website against what sales actually told buyers: gaps, ambiguities and contradictions, filed in Sanity Studio. DEV.to Sanity Challenge entry.
Cross-Check
An agent that compares what Pagepro's website says with what our sales team actually told prospects, and files gaps (buyers ask, the site doesn't answer), ambiguities (the site answers two ways) and contradictions (the site disagrees with sales, or with itself). Findings land in Sanity Studio, next to the evidence, for the editor who fixes the page. Entry for the DEV.to Sanity Challenge, Path One.
The Vercel deploy still carries the project's earlier working name, coverage-gap-probe.
How it works
Pagepro site (private dataset `pagepro`)
└─ scripts/stage1/build-kb-sources.mjs → 73 `kbSource` docs (one per page / case study: URL + text)
└─ Knowledge Base "Pagepro site" (Sanity Context, built with `sanity context`)
└─ Context MCP, Knowledge Base mode ── KB agent (lib/probe/kb-agent.ts)
│
buyerQuestion (87, from 209 prospect emails) ──────┤
salesAnswer (65, what we told clients) ────────────┤ judge (lib/probe/judge.ts)
page text search (gap check, independent of KB) ───┘ verdict, priority,…How I Used Sanity
What the Knowledge Base is built from. A copy of our pages and case studies lives in Sanity as page-builder documents, where a large share of each document is layout and plumbing rather than words a reader sees. Pointed straight at those, the Context build skipped the largest service pages, the ones that carry pricing and process. So a build step walks each document, keeps only the text, and writes one kbSource document holding the page's full URL and its body. 73 of them: 52 pages and 21 case studies. Seven utility pages are excluded by hand (careers, privacy, cookies, newsletter, thank-you and two forms) because no buyer question touches them. Answers can then cite the subpage they came from.
How the agent queries it. Through Context MCP in mode=knowledge_base, using knowledge_base_search to find candidate entries and knowledge_base_read to pull their full text. Knowledge Base search matches exact keywords with no stemming, so the agent is instructed to run at least three searches with different wordings before concluding anything, and to read every hit rather than trusting the outline. Without that instruction it searches once, gives up, and reports something as missing when it is not. The search is plain; the structure is what makes the verdict possible, because what the site says, what buyers ask and what sales answered are separate document types the agent compares.
Why the sales corpus stays outside the Knowledge Base. The Knowledge Base is the site's view of itself. The evidence against it has to come from somewhere the site cannot see, so the 65 sales answers sit in the same dataset but outside the index, retrieved separately. That also keeps them clear of the document cap.
One guard in the pipeline. Before any gap is filed, the tool searches the page text itself, independently of the Knowledge Base. A gap is only recorded when neither the Knowledge Base answer nor the raw page text answers the question. Without it, a keyword miss becomes a false gap, and the tool sends an editor off to write content that already exists.
How we know it reads the Knowledge Base. A model that already knows something about us can look like it read our content when it did not. So we plant a fact that exists nowhere else, ask about it, and remove it again: every new client receives a printed project handbook bound in green linen. An answer that repeats it came from the Knowledge Base. scripts/stage1/plant-fact.mjs plants and removes it, and it only ever touches our private copy of the content.
What the agent does with what it retrieves. A judge model compares the Knowledge Base answer, the page text excerpts and the matching sales records, and returns a verdict, a priority, a one line headline for the editor and a concrete fix. For Answer drift, both outside answers are graded against the Knowledge Base answer and the verdict itself is computed in code, as a pure, unit tested function, rather than written by a model.
Where it is written back. Findings are ordinary Sanity documents. That is why Studio renders them with no custom interface at all, and why the custom tabs only had to add the side by side view and the review action.
The honest limits. Knowledge Bases are in beta and cap at 150 documents, at organization level rather than per project, and larger documents get split during a build, so the real ceiling arrives sooner than it looks. Web search results do not reproduce, so every run is stored with its date and model. The sample is our own site, so the findings are ours and not a general claim about anyone else's content.
The review is a human one, not a test suite. We have gone through 14 findings so far and confirmed 13 as real, marking one as a false alarm where the site did answer the question after all.
Sanity Project Details
Project ID: jupwoq5l
Dataset: pagepro (private)
Knowledge Base: built with Sanity Context from the 73 kbSource documents
Context MCP: Knowledge Base mode, knowledge_base_search and knowledge_base_read
Models: Claude Sonnet 5 through the Vercel AI Gateway, with Anthropic's provider executed web_search tool for the middle column
Document types: kbSource (page text and URL), buyerQuestion (the corpus, with its pre-audit label and measured site coverage), salesAnswer (what we told clients, anonymised), finding (the audit output), comparison (an Answer drift row).
Top comments (0)