This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.
What I Built
A customer disputes an old fictional card fee. The statement was issued in July, but the charge posted in June. There are older and newer policies, a special-customer rule, and scoped clarifications. Which rule applied when the charge was posted, and when must support stop rather than promise a waiver?
PolicyTrace is an English-first, Chinese-supported support agent. A model routes the question, Sanity Context supplies policy records, deterministic checks handle effective dates, customer conditions, conflicts and authority, and a model confirms the citation set. Missing facts trigger a question. An unresolved conflict or waiver request produces an unsent human handoff. There is no refund or waiver execution tool.
Demo
Open the public evidence viewer and video. The viewer replays saved results; it does not invoke a live model. The current narrated video shows evidence cards and an application screenshot, not a continuous recording of the latest version.
The key scenario compares the same disputed premium-customer fee before and after a scoped clarification is reviewed. Before review, both applicable policies remain visible and the case goes to a human. After review, the clarification permits a sourced fee explanation, while a waiver request still goes to a human. Evidence includes an actual Sanity Context read and a same-code pending-versus-reviewed comparison. It does not establish a full unchanged-code cloud publication experiment.
Code
Download the reproducible v4 source ZIP. An independently extracted v4 package passed 55 included offline tests in an existing Python environment. It includes the FastAPI app, bilingual UI, Sanity adapter, policy verifier, model integration, synthetic cases, comparison scripts and traces. Live local operation requires the evaluator's own Sanity and model credentials; none are distributed.
How I Used Sanity
In the saved evaluation, Sanity Context MCP read six structured records in a Knowledge Base: policy versions and clarifications. Effective-date ranges, customer conditions, amounts, source IDs, revisions and authority information change the decision. An old posting date excludes a later policy; conflicting premium-customer rules trigger a handoff; a reviewed, scoped clarification changes what can be explained but never grants waiver authority. Saved traces include retrieval time and source fingerprint.
A server-side reviewed content hash controls clarification approval; a Knowledge Base entry cannot approve itself by naming a reviewer. Source failures, unknown conditions and mismatched model citations fail closed. The adapter accepts a bounded record format and requires review for arbitrary new pages. Customer facts are synthetic input, not identity verification.
Sanity Project Details
- Project ID:
7d711ssk - Context Knowledge Base ID:
kbuLoYcmxK7A - Content source: two uploaded structured Markdown files represented by six Context entries at the evaluated snapshot; the Knowledge Base can contain more entries after updates. Products, bills, customers and policies are fictional.
Evaluation and Limits
An actual Qwen3-8B → Sanity Context → policy checks → model citation-confirmation run passed seven developer-authored cases after a citation-prompt correction. The earlier first-case citation failure is retained. In ten predeclared fictional cases using one Sanity snapshot and the same Qwen3-8B model, PolicyTrace met the recorded decision, amount/null, valid nonduplicated citation and no-action checks in 10/10; a full-context direct-model baseline met them in 1/10 and a simple keyword Top-3 baseline in 0/10. Frozen cases and raw results are in the package.
These cases were written by the developer, not independently sampled or blinded. The keyword baseline is simple, not a strong RAG implementation. The system has task-specific date and authority checks, so the methods have different capabilities. This experiment cannot establish production accuracy, a statistical win rate or superiority over other entries. No real bank connection or measured business return exists.
Built with AI-assisted development and testing. The author reviewed the design, code, results and limitations. No Agent Session transcript is submitted.
Top comments (0)