<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ling Zhou</title>
    <description>The latest articles on DEV Community by Ling Zhou (@ling_zhou_78cf1e40f3e605d).</description>
    <link>https://dev.to/ling_zhou_78cf1e40f3e605d</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3892870%2F08f96e6a-56f6-465e-b634-9b4b4898086a.png</url>
      <title>DEV Community: Ling Zhou</title>
      <link>https://dev.to/ling_zhou_78cf1e40f3e605d</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ling_zhou_78cf1e40f3e605d"/>
    <language>en</language>
    <item>
      <title>ClauseHound: a pet-insurance decision engine that refuses to guess</title>
      <dc:creator>Ling Zhou</dc:creator>
      <pubDate>Thu, 01 Oct 2026 19:35:27 +0000</pubDate>
      <link>https://dev.to/ling_zhou_78cf1e40f3e605d/clausehound-a-pet-insurance-decision-engine-that-refuses-to-guess-14f</link>
      <guid>https://dev.to/ling_zhou_78cf1e40f3e605d/clausehound-a-pet-insurance-decision-engine-that-refuses-to-guess-14f</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge&lt;/a&gt;, Path One: Ship an Agent That Queries Real Content.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;🔗 &lt;strong&gt;Demo:&lt;/strong&gt; &lt;a href="https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app" rel="noopener noreferrer"&gt;https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app&lt;/a&gt;&lt;br&gt;
🗄️ &lt;strong&gt;Sanity project ID:&lt;/strong&gt; &lt;code&gt;yijqzehr&lt;/code&gt; (dataset &lt;code&gt;production&lt;/code&gt;)&lt;br&gt;
🏷️ &lt;strong&gt;Tag:&lt;/strong&gt; &lt;code&gt;#sanitychallenge&lt;/code&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Your Bernese Mountain Dog needs hip surgery at age three. You file the claim, confident — and get denied. The 12-month orthopedic waiting period was in Section 4 all along. You just never found it &lt;em&gt;before&lt;/em&gt; you bought the policy.&lt;/p&gt;

&lt;p&gt;That's the moment ClauseHound is built for — except &lt;em&gt;before&lt;/em&gt; it happens. &lt;strong&gt;ClauseHound is a pet-insurance decision engine for people shopping for a policy.&lt;/strong&gt; Not a chatbot that answers trivia about pet insurance: it takes the four decisions shoppers actually agonize over and answers each one from the insurers' own policy text, with the exact clause cited:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Insure vs. save&lt;/strong&gt; — "Should I get pet insurance or just put $60/month in a savings account?" ClauseHound runs the deterministic math (premiums vs. vet bills over time, in integer cents, in code — the LLM never does arithmetic) and tells you plainly when saving wins. It says "save" when saving wins.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lifetime cost comparison&lt;/strong&gt; — paste two quotes and it computes the true 10-year cost per plan: premiums + deductible + reimbursement math, side by side, exact to the cent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claim-payment analysis&lt;/strong&gt; — "My dog needs a $4,200 cruciate surgery in month 8 — what would each insurer actually pay?" Answered per-carrier, from the coverage clauses, exclusions, and waiting periods that govern that claim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Switching guidance&lt;/strong&gt; — "My dog has a pre-existing skin condition; which insurers will still cover her if I switch?" The highest-stakes question in pet insurance, answered carrier by carrier, with the pre-existing-condition clauses cited.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One grounding contract runs through all four:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Verdict first, citations follow.&lt;/strong&gt; Every answer opens with the decision, then shows its work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No evidence, no answer.&lt;/strong&gt; Ask about something the corpus doesn't cover and you get &lt;em&gt;"I couldn't verify this in the policy documents"&lt;/em&gt; — not a guess, a refusal, with pointers to what &lt;em&gt;is&lt;/em&gt; answerable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conflicts stay unresolved.&lt;/strong&gt; When a coverage clause and an exclusion both match, both surface side by side, marked &lt;code&gt;unresolved&lt;/code&gt;. The agent never picks a winner.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Owner-reported facts stay visibly separate from verified policy facts.&lt;/strong&gt; What you told it and what the policy says are never blended.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Carrier-specific claims require policy citations.&lt;/strong&gt; No citation, no claim.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;No login needed. The landing page frames the four decisions; each job card (and the hero) opens the chat modal with its question already prefilled — you press Send, it never auto-sends.&lt;/p&gt;

&lt;p&gt;Try these in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Insure vs. save&lt;/strong&gt; — press Send on the prefilled question. It asks three quick questions that change the answer (dog's age, monthly savings, breed/size) — nothing more. Watch the loading stages narrate the pipeline ("Pulling the relevant policy text…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…"), then the verdict: for a healthy young dog with no breed risks, it will tell you saving wins, with the math shown.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The pre-existing switch&lt;/strong&gt; — "My dog has a pre-existing skin condition. Which insurers will still cover her?" Five carriers, five cited answers, one table.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The conflict&lt;/strong&gt; — "Are cruciate ligament injuries covered in the first year?" Some carriers cover, some exclude as pre-existing within the waiting period. Both clauses surface, side by side, &lt;code&gt;unresolved&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkrf5f5t4vcfsu3qe3rug.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkrf5f5t4vcfsu3qe3rug.png" alt="Landing page: the four decisions"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The landing page frames the four decisions, not a chat box.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsm9h5v4pmfqdaifai0ks.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsm9h5v4pmfqdaifai0ks.png" alt="Chat modal with the insure-vs-save question prefilled"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Each job card opens the modal with its question prefilled. Nothing auto-sends.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp11uvbxf4015c5e1iv25.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp11uvbxf4015c5e1iv25.png" alt="Answered state: verdict first, then the math"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The answered state: "There's no single winner — it depends on when the emergency hits." Verdict first, then the month-by-month math, then the cited clauses.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;Next.js + TypeScript app. The interesting parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;app/lib/mcp.ts&lt;/code&gt; — Sanity Context MCP client (JSON-RPC over HTTPS). The app never queries Sanity directly; every fact comes through the Context endpoint's tools (&lt;code&gt;initial_context&lt;/code&gt;, &lt;code&gt;groq_query&lt;/code&gt;, …).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;app/lib/cost.ts&lt;/code&gt; — deterministic money math in integer cents: premiums, deductibles, reimbursement rates, annual limits, multi-year scenarios. The LLM explains the numbers; it never computes them. (Tested to the cent — &lt;code&gt;cost.test.ts&lt;/code&gt;.)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;app/lib/insureVsSave.ts&lt;/code&gt;, &lt;code&gt;switching.ts&lt;/code&gt;, &lt;code&gt;denial.ts&lt;/code&gt;, &lt;code&gt;renewals.ts&lt;/code&gt;, &lt;code&gt;appeals.ts&lt;/code&gt; — the four decision jobs as typed modules, each with its own test suite.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;app/components/ChatModal.tsx&lt;/code&gt; + &lt;code&gt;ChatApp.tsx&lt;/code&gt; — the modal chat: prefilled questions, streaming answers with loading stages, citation cards, contradiction panels, cost cards.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The repo is private (&lt;code&gt;lymcho/clausehound&lt;/code&gt;, now merged to &lt;code&gt;main&lt;/code&gt;); the test suite (1,653 tests) and the eval fixtures are part of the submission evidence below.&lt;/p&gt;
&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;Everything ClauseHound knows lives in Sanity as structured content — typed documents with relationships, modeled from the insurers' published US sample policies:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;5 insurers&lt;/strong&gt; → &lt;strong&gt;5 policy documents&lt;/strong&gt; → &lt;strong&gt;28 coverage clauses&lt;/strong&gt;, &lt;strong&gt;30 exclusion clauses&lt;/strong&gt;, &lt;strong&gt;18 waiting periods&lt;/strong&gt;, &lt;strong&gt;20 claim steps&lt;/strong&gt; — 106 documents in the &lt;code&gt;production&lt;/code&gt; dataset, verified live.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The schema is the trust strategy. Every clause carries a required &lt;code&gt;sectionRef&lt;/code&gt; (e.g. &lt;code&gt;Sec. 3.4&lt;/code&gt;) — the citation anchor — and a required &lt;code&gt;policy&lt;/code&gt; reference back to its &lt;code&gt;policyDocument&lt;/code&gt;, which references its &lt;code&gt;insurer&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;insurer → policyDocument → coverageClause / exclusionClause / waitingPeriod / claimStep
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent reads through a &lt;strong&gt;Sanity Context MCP endpoint&lt;/strong&gt; backed by an Agent Context over that dataset. A GROQ content filter at the context level scopes every query to the four answerable types and only documents attached to a real policy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;_type in ["coverageClause", "exclusionClause", "waitingPeriod", "claimStep"]
&amp;amp;&amp;amp; defined(*[_type == "policyDocument" &amp;amp;&amp;amp; _id == ^.policy._ref][0])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That filter is the guardrail: the agent &lt;em&gt;cannot&lt;/em&gt; wander into irrelevant content no matter what the user asks. And because coverage, exclusions, and waiting periods are distinct document types, they stay distinct in answers — a &lt;code&gt;coverageClause&lt;/code&gt; and an &lt;code&gt;exclusionClause&lt;/code&gt; matching the same question surface as two typed records with their &lt;code&gt;sectionRef&lt;/code&gt;s, not two paragraphs to blend. That's what makes the conflict view possible.&lt;/p&gt;

&lt;p&gt;The app narrates the pipeline at real phase boundaries — "Pulling the relevant policy text…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…" — so the grounding is visible while you wait.&lt;/p&gt;

&lt;p&gt;The bet: &lt;strong&gt;the agent only works because the content was structured.&lt;/strong&gt; Keyword search could find the words "hip dysplasia"; it couldn't keep five carriers' answers from bleeding into each other, keep a coverage clause distinct from the exclusion that overrides it, or compute a 10-year cost from typed premium/deductible/reimbursement fields. The structure &lt;em&gt;is&lt;/em&gt; the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the live runs showed
&lt;/h2&gt;

&lt;p&gt;A 19-fixture held-out eval (&lt;code&gt;eval/fixtures.json&lt;/code&gt;, frozen — never edited to make the app pass). Each fixture posts a real user question to the deployed &lt;code&gt;/api/chat&lt;/code&gt; and checks the response structurally: HTTP 200, citations resolve to real clause IDs, no dangling markers, carrier scope correct, refusals refuse, contradictions surface as &lt;code&gt;unresolved&lt;/code&gt;. The two cost fixtures re-derive every dollar figure with an independent port of the integer-cent math — exact match required, no LLM judging.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;16/19 passed&lt;/strong&gt; on the production build, 2026-10-01 (merged to &lt;code&gt;main&lt;/code&gt; as &lt;code&gt;9957cb1&lt;/code&gt;; the merge tree is byte-identical to the evaluated commit, so the results stand for production). Live answers stream in 18–95s; the deterministic cost fixtures verify in under a second.&lt;/p&gt;

&lt;p&gt;The three failures, because they matter more than the score:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;inv-10:&lt;/strong&gt; correctly refused an unanswerable question (cloning coverage) but emitted 5 citations alongside the refusal. A refusal should be clean — citations on a refusal are a leak.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;inv-11:&lt;/strong&gt; asked about acupuncture coverage, returned 0 citations. The corpus is thin on alternative-therapy clauses; it should have said "couldn't verify" with pointers instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;cfl-01:&lt;/strong&gt; the expected &lt;code&gt;unresolved&lt;/code&gt; cruciate-waiting-period contradiction didn't surface in the eval run — but did in a manual re-run minutes later (Lemonade 6 months vs. Nationwide 12-month exclusion). Flaky across runs rather than absent, which is arguably worse, and the highest-priority fix: the conflict view is the feature.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Known limitations: five carriers' published US sample policies — not every plan variant, rider, or state endorsement. Fixtures are authored by me, so they measure agreement with my labels, not independent accuracy. Hard multi-carrier questions can take ~2 minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Project ID: &lt;code&gt;yijqzehr&lt;/code&gt;, dataset &lt;code&gt;production&lt;/code&gt; (106 documents, queried live 2026-10-01)&lt;/li&gt;
&lt;li&gt;Document types: &lt;code&gt;insurer&lt;/code&gt;, &lt;code&gt;policyDocument&lt;/code&gt;, &lt;code&gt;coverageClause&lt;/code&gt;, &lt;code&gt;exclusionClause&lt;/code&gt;, &lt;code&gt;waitingPeriod&lt;/code&gt;, &lt;code&gt;claimStep&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Reads via the Sanity Context MCP endpoint (&lt;code&gt;SANITY_AGENT_CONTEXT_URL&lt;/code&gt;) over an Agent Context on the production dataset; the GROQ filter above scopes all retrieval to answerable, policy-attached documents.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No login needed to try ClauseHound. Live answers are streamed; deterministic cost comparisons run without the model.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Demo only — research aid, not professional insurance advice. Verify with your insurer before buying.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>sanitychallenge</category>
    </item>
  </channel>
</rss>
