DEV Community

Nikema
Nikema

Posted on AI-assisted

Ask the Archive: a Q&A agent grounded in 150 stories of Black California history

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Ask California Black Stories is a chat agent that answers questions about Black history in California, and it only answers from a real archive: 150 researched stories, 436 claims, 45 thematic pillars, with every claim carrying at least one source URL.

I run California Black Stories, a content project on Black history in California. The research archive behind it is large and structured, but asking it a question meant digging through notes by hand. And asking a general-purpose chatbot gets you a confident answer stitched from who knows where. I wanted an interface where every answer traces back to a specific story and a real source, and where the agent says "that is not in the archive" instead of guessing.

So the agent has one hard rule: if the archive does not cover the question, it says so and stops. No general knowledge, no gap filling.

Demo

The landing page has sample questions you can click to see it in action. Ask about Biddy Mason or the Palace Hotel sit-in and you get a cited answer. Ask about the Montgomery bus boycott and it tells you that is outside the archive, because it is.

Try it live: https://californiablackstories.com
Demo video: https://www.youtube.com/watch?v=hYWFyAtU46I

Code

GitHub logo prophen / california-black-stories

A Q&A agent that answers questions about Black history in California, grounded in a 150-story Sanity archive with cited sources. Built with Sanity Context MCP for the Sanity Challenge, Path One.

Ask California Black Stories

A chat agent that answers questions about Black history in California, grounded in a real research archive. Every answer cites its sources. If the archive does not cover your question, the agent says so and stops.

Live: https://californiablackstories.com

Video walkthrough: https://www.youtube.com/watch?v=hYWFyAtU46I

Ask California Black Stories landing page

Built for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

The archive

California Black Stories is my content project on Black history in California. The Sanity dataset holds 150 story documents: 436 fact-check claims across 45 thematic pillars, with every claim carrying at least one source URL.

Content model: story, person, place, source. Stories carry structured claim objects; people and places are their own documents so questions about who and where resolve against structure, not keyword matching.

How it works

  1. Research notes (an Obsidian vault with three note formats) are parsed into normalized records (ingest/parse_notes.py…

How I Used Sanity

The content model. Four schema types: story, person, place, source. Stories carry structured claim objects, and every claim has at least one source URL. People and places are their own documents so questions about who and where resolve against structure, not keyword matching.

The ingest. My research lives in an Obsidian vault with three different note formats. I wrote a parser that normalizes all three into story records, converted them to NDJSON, and imported 150 documents into a production dataset with sanity dataset import.

Sanity Context. I built a Knowledge Base ("California Black Stories") on the query *[_type == "story"], exposed through a Context MCP endpoint. The KB distilled 33 entries from the 150 stories, each entry linked to the story it came from.

The agent. A Next.js app. The /ask chat UI streams from /api/ask, which uses AI SDK v7 with an MCP client pointed at the Sanity Context endpoint. The model reads KB entries as its only context.

Two grounding bugs I caught in testing:

  1. After refusing a question, the agent answered it anyway from general knowledge (my test: the Montgomery bus boycott). The system prompt rule was not strict enough. I tightened it and retested until the refusal held.

  2. Dead citation links. KB entries cite story titles but carry no URLs, so the model invented plausible-looking source links, including fake URLs on the site's own domain. A story lookup tool fixed it at first, but the model sometimes skipped the tool and faked a link anyway. So the route now injects every story's real source URLs from Sanity directly into the prompt, and the model copies from that list instead of calling a tool. I retested with the Palace Hotel sit-in and the citations now link to the real sources.

The prompt also instructs the agent to surface disagreements side by side when sources tell a story differently, instead of picking one.

An honest note on verification. I ran mechanical checks over the KB output: all 479 citation markers resolve to real story titles, no two stories cover the same person or event, and the corpus report confirms every claim has at least one source URL. That is the level I am claiming: structural and citation-level verification, not a line-by-line re-fact-check of all 436 claims.

Sanity Project Details

  • Project ID: bq2hoxdt
  • Dataset: production
  • 150 story documents; schema types story, person, place, source

Top comments (0)