This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
Market Evidence Desk is a public research desk for crypto investor protection questions. It links original publications to dated claims about proof of reserves, custody, deposit insurance, stablecoin redemption, and the gap between international recommendations and local implementation. An exchange may say customers can check inclusion in a reserve snapshot, while accounting and regulatory sources explain why that snapshot cannot establish full liabilities or solvency. SEC staff and a Commissioner also take different positions on what reserve reports demonstrate. Flattening those positions into a single confident answer would lose the point.
The site lets a visitor inspect published source records, follow each claim to its question and original URL, compare dated views, and ask a scoped AI agent. The published Sanity graph contains 101 research questions linked to 22 source records (21 distinct source URLs); the evidence atlas shows them in manageable batches. Source scope notes make clear whether a document is investor education, analysis, or a recommendation to regulators. The agent returns a draft with source links and the names of the Sanity Context tools it used. It declines live Bitcoin prices and buy/sell requests because the Sourcebook does not contain current market feeds. Research drafts remain subject to human review.
Demo
- Live Desk: market-evidence-desk.vercel.app
- Evidence Atlas: market-evidence-desk.vercel.app/web/atlas.html
- Compare Evidence: market-evidence-desk.vercel.app/web/compare.html
- Sourcebook: market-evidence-desk.vercel.app/web/sources.html
- AI Agent: market-evidence-desk.vercel.app/web/agent.html?case=dispute
Video Walkthrough
Make the claim. Show the receipt.
Watch the walkthrough on YouTube
Instructions for Judges
Judges can test the public app without signing in:
- Start on the Home page, select Show the disagreement, and press Ask the sources.
- Compare the cited answer with the original sources and the guided comparison matrix.
- On the Research desk, switch to the question about worldwide protections to see 2023 international recommendations linked alongside 2025 implementation reviews.
- Try the live-price boundary case as a negative test to verify the agent's refusal mechanism.
Code
GitHub repository, production branch · Architecture and setup · Scope and limitations
How I Used Sanity
The production dataset models source, marketEvent (the research question), and evidenceClaim as separate linked documents. Each claim points to a source and a question, with a stance such as supports, context, or conflicts, a date, and review state. The public pages query published records and show each position with its origin, date, and unresolved limits. This structure makes it possible to follow opposing positions and avoid treating a company's description of its own process as independent assurance.
Sanity Context points at both the dataset and the Market Evidence Desk Sourcebook Knowledge Base, whose selection is capped at 140 source and claim documents plus a separate SEC investor-alert page. The latest Sourcebook rebuild, on September 27, succeeded with 13 cited entries and zero open issues, and the proof-of-reserves entry keeps a source-scoped correction that separates Kraken's company description from PCAOB and SEC cautions. Import, build and public-graph counts measure different things, so the 140 cap is not a count of current published documents; the full graph stays available through the public pages and GROQ. The Python agent calls groq_query for the linked published graph and knowledge_base_read for relevant Sourcebook entries before answering, and the server withholds an answer if either retrieval was skipped. Credentials stay in server environment variables. The guided comparisons read Sanity records directly and are not AI-generated.
The hard case is the same-day April 4, 2025 stablecoin material. The SEC Division of Corporation Finance staff described an issuer practice around reserve reports; Commissioner Caroline A. Crenshaw challenged whether such reports can establish adequate backing. The desk shows both attributed positions and does not turn either into a Commission rule or a present-day solvency finding. It also separates FSB and IOSCO recommendations from the rules actually adopted by individual countries. Human review remains separate from synthesis.
Review status is shown honestly: 1 of the 112 claims carries an editor source check with the exact excerpt and locator, and the rest are labeled as awaiting review. Adding more checks is a Studio task, and the desk never presents an unchecked claim as verified.
Sanity Project Details
- Project ID:
cxjysvlq - Dataset:
production - Public Sanity dataset query
- Hosted Sanity Studio
- Sourcebook Knowledge Base ID:
kbu9WNgZ9ocF
What I Learned
Structured content matters most where sources disagree or cover different scopes. Dates, authors, stances, and source references let the interface show what each source can establish. Sanity Context gives the agent a way to read both records and the Sourcebook, while the human review state reminds readers that a generated explanation is a draft. The useful boundary is as visible as the answer: historical evidence cannot provide a current price or investment recommendation.
Top comments (0)