DEV Community

Ling Zhou
Ling Zhou

Posted on

ClauseHound: a pet-insurance decision engine that refuses to guess

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

πŸ”— Demo: https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app
πŸ—„οΈ Sanity project ID: yijqzehr (dataset production)
🏷️ Tag: #sanitychallenge


What I Built

Your Bernese Mountain Dog needs hip surgery at age three. You file the claim, confident β€” and get denied. The 12-month orthopedic waiting period was in Section 4 all along. You just never found it before you bought the policy.

That's the moment ClauseHound is built for β€” except before it happens. ClauseHound is a pet-insurance decision engine for people shopping for a policy. Not a chatbot that answers trivia about pet insurance: it takes the four decisions shoppers actually agonize over and answers each one from the insurers' own policy text, with the exact clause cited:

  1. Insure vs. save β€” "Should I get pet insurance or just put $60/month in a savings account?" ClauseHound runs the deterministic math (premiums vs. vet bills over time, in integer cents, in code β€” the LLM never does arithmetic) and tells you plainly when saving wins. It says "save" when saving wins.
  2. Lifetime cost comparison β€” paste two quotes and it computes the true 10-year cost per plan: premiums + deductible + reimbursement math, side by side, exact to the cent.
  3. Claim-payment analysis β€” "My dog needs a $4,200 cruciate surgery in month 8 β€” what would each insurer actually pay?" Answered per-carrier, from the coverage clauses, exclusions, and waiting periods that govern that claim.
  4. Switching guidance β€” "My dog has a pre-existing skin condition; which insurers will still cover her if I switch?" The highest-stakes question in pet insurance, answered carrier by carrier, with the pre-existing-condition clauses cited.

One grounding contract runs through all four:

  • Verdict first, citations follow. Every answer opens with the decision, then shows its work.
  • No evidence, no answer. Ask about something the corpus doesn't cover and you get "I couldn't verify this in the policy documents" β€” not a guess, a refusal, with pointers to what is answerable.
  • Conflicts stay unresolved. When a coverage clause and an exclusion both match, both surface side by side, marked unresolved. The agent never picks a winner.
  • Owner-reported facts stay visibly separate from verified policy facts. What you told it and what the policy says are never blended.
  • Carrier-specific claims require policy citations. No citation, no claim.

Demo

No login needed. The landing page frames the four decisions; each job card (and the hero) opens the chat modal with its question already prefilled β€” you press Send, it never auto-sends.

Try these in order:

  1. Insure vs. save β€” press Send on the prefilled question. It asks three quick questions that change the answer (dog's age, monthly savings, breed/size) β€” nothing more. Watch the loading stages narrate the pipeline ("Pulling the relevant policy text…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…"), then the verdict: for a healthy young dog with no breed risks, it will tell you saving wins, with the math shown.
  2. The pre-existing switch β€” "My dog has a pre-existing skin condition. Which insurers will still cover her?" Five carriers, five cited answers, one table.
  3. The conflict β€” "Are cruciate ligament injuries covered in the first year?" Some carriers cover, some exclude as pre-existing within the waiting period. Both clauses surface, side by side, unresolved.

Landing page: the four decisions
The landing page frames the four decisions, not a chat box.

Chat modal with the insure-vs-save question prefilled
Each job card opens the modal with its question prefilled. Nothing auto-sends.

Answered state: verdict first, then the math
The answered state: "There's no single winner β€” it depends on when the emergency hits." Verdict first, then the month-by-month math, then the cited clauses.

Code

Next.js + TypeScript app. The interesting parts:

  • app/lib/mcp.ts β€” Sanity Context MCP client (JSON-RPC over HTTPS). The app never queries Sanity directly; every fact comes through the Context endpoint's tools (initial_context, groq_query, …).
  • app/lib/cost.ts β€” deterministic money math in integer cents: premiums, deductibles, reimbursement rates, annual limits, multi-year scenarios. The LLM explains the numbers; it never computes them. (Tested to the cent β€” cost.test.ts.)
  • app/lib/insureVsSave.ts, switching.ts, denial.ts, renewals.ts, appeals.ts β€” the four decision jobs as typed modules, each with its own test suite.
  • app/components/ChatModal.tsx + ChatApp.tsx β€” the modal chat: prefilled questions, streaming answers with loading stages, citation cards, contradiction panels, cost cards.

The repo is private (lymcho/clausehound, now merged to main); the test suite (1,653 tests) and the eval fixtures are part of the submission evidence below.

How I Used Sanity

Everything ClauseHound knows lives in Sanity as structured content β€” typed documents with relationships, modeled from the insurers' published US sample policies:

  • 5 insurers β†’ 5 policy documents β†’ 28 coverage clauses, 30 exclusion clauses, 18 waiting periods, 20 claim steps β€” 106 documents in the production dataset, verified live.

The schema is the trust strategy. Every clause carries a required sectionRef (e.g. Sec. 3.4) β€” the citation anchor β€” and a required policy reference back to its policyDocument, which references its insurer:

insurer β†’ policyDocument β†’ coverageClause / exclusionClause / waitingPeriod / claimStep
Enter fullscreen mode Exit fullscreen mode

The agent reads through a Sanity Context MCP endpoint backed by an Agent Context over that dataset. A GROQ content filter at the context level scopes every query to the four answerable types and only documents attached to a real policy:

_type in ["coverageClause", "exclusionClause", "waitingPeriod", "claimStep"]
&& defined(*[_type == "policyDocument" && _id == ^.policy._ref][0])
Enter fullscreen mode Exit fullscreen mode

That filter is the guardrail: the agent cannot wander into irrelevant content no matter what the user asks. And because coverage, exclusions, and waiting periods are distinct document types, they stay distinct in answers β€” a coverageClause and an exclusionClause matching the same question surface as two typed records with their sectionRefs, not two paragraphs to blend. That's what makes the conflict view possible.

The app narrates the pipeline at real phase boundaries β€” "Pulling the relevant policy text…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…" β€” so the grounding is visible while you wait.

The bet: the agent only works because the content was structured. Keyword search could find the words "hip dysplasia"; it couldn't keep five carriers' answers from bleeding into each other, keep a coverage clause distinct from the exclusion that overrides it, or compute a 10-year cost from typed premium/deductible/reimbursement fields. The structure is the product.

What the live runs showed

A 19-fixture held-out eval (eval/fixtures.json, frozen β€” never edited to make the app pass). Each fixture posts a real user question to the deployed /api/chat and checks the response structurally: HTTP 200, citations resolve to real clause IDs, no dangling markers, carrier scope correct, refusals refuse, contradictions surface as unresolved. The two cost fixtures re-derive every dollar figure with an independent port of the integer-cent math β€” exact match required, no LLM judging.

16/19 passed on the production build, 2026-10-01 (merged to main as 9957cb1; the merge tree is byte-identical to the evaluated commit, so the results stand for production). Live answers stream in 18–95s; the deterministic cost fixtures verify in under a second.

The three failures, because they matter more than the score:

  • inv-10: correctly refused an unanswerable question (cloning coverage) but emitted 5 citations alongside the refusal. A refusal should be clean β€” citations on a refusal are a leak.
  • inv-11: asked about acupuncture coverage, returned 0 citations. The corpus is thin on alternative-therapy clauses; it should have said "couldn't verify" with pointers instead.
  • cfl-01: the expected unresolved cruciate-waiting-period contradiction didn't surface in the eval run β€” but did in a manual re-run minutes later (Lemonade 6 months vs. Nationwide 12-month exclusion). Flaky across runs rather than absent, which is arguably worse, and the highest-priority fix: the conflict view is the feature.

Known limitations: five carriers' published US sample policies β€” not every plan variant, rider, or state endorsement. Fixtures are authored by me, so they measure agreement with my labels, not independent accuracy. Hard multi-carrier questions can take ~2 minutes.

Sanity Project Details

  • Project ID: yijqzehr, dataset production (106 documents, queried live 2026-10-01)
  • Document types: insurer, policyDocument, coverageClause, exclusionClause, waitingPeriod, claimStep
  • Reads via the Sanity Context MCP endpoint (SANITY_AGENT_CONTEXT_URL) over an Agent Context on the production dataset; the GROQ filter above scopes all retrieval to answerable, policy-attached documents.

No login needed to try ClauseHound. Live answers are streamed; deterministic cost comparisons run without the model.


Demo only β€” research aid, not professional insurance advice. Verify with your insurer before buying.

Top comments (0)