DEV Community

Cover image for Retrieval Isn't Enough — I Built a Claim Relationship Resolver
Emmimal P Alexander
Emmimal P Alexander

Posted on

Retrieval Isn't Enough — I Built a Claim Relationship Resolver

Sanity Challenge Path One Submission

Claim Relationship Resolver

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

What I Built

I built a Claim Relationship Resolver, a deterministic, pure-Python agent using only the standard library and no LLM API. It answers questions from structured claims stored in Sanity, but instead of simply retrieving a matching document, it determines how the retrieved claims relate to each other before producing an answer. The agent takes a structured question, retrieves the relevant claims through Sanity Context MCP, applies the relationship rules, and returns a typed result rather than generating an answer from retrieved text.

The idea started from one distinction: two claims can look different without actually disagreeing. A pricing limit of 500 for new accounts and 100 for legacy accounts isn't necessarily a contradiction, it can just be a scope difference. But two claims with the same applicable scope that disagree, with no relationship resolving them, represent a genuine conflict. And a claim explicitly superseded by another claim shouldn't remain current simply because it's still sitting in the knowledge base.

A retrieval system can return those claims. The missing layer is deciding what relationship they have.

My resolver classifies each question into one of five outcomes:

  • SUPPORTED — one applicable claim supports the answer.
  • CONTEXTUAL — different valid answers apply to different scopes.
  • SUPERSEDED — an explicit supersedes relationship makes an older claim stale.
  • CONFLICTING — applicable claims for the same scope disagree and no relationship resolves the difference.
  • INSUFFICIENT — no eligible claim exists for the question.

The rules are frozen and documented before evaluation. There's no "newest wins" heuristic and no source-authority ranking. A claim becomes stale only through an explicit supersedes relationship.

To evaluate the system, I built:

  • A 404-question held-out benchmark generated from a fixed seed, with ground truth determined independently by how each test case was constructed, not by running the resolver.
  • A blind 36-question audit, where I manually labeled the cases from the frozen rules before inspecting the generated answers: 36/36 agreement.
  • Three comparison baselines: a retrieval-only BM25 baseline, newest claim wins, and newest claim wins within the matching scope.
  • Two real-content datasets: 36 real TDS Contributor Portal articles and the 20 most recent pages from my EmiTechLogic sitemap, run through the same resolver and baselines with no synthetic benchmark data involved.

Held-out results

On the 380 headline questions in the held-out benchmark:

System Correct Confidently wrong Missed conflicts
Claim Relationship Resolver 380/380 0/380 0/50
B+scope 276/380 38/380 50/50
B_newest 176/380 180/380 50/50
A_retrieval 125/380 255/380 50/50

The 404 questions come from 240 claim clusters, so they aren't independent observations. I report the cluster count explicitly rather than treating the 404 questions as 404 unrelated experiments. The remaining 24 questions are Tier-2 cases outside the headline accuracy calculation, and the resolver correctly returns unsupported_case for all 24 rather than guessing (that's the 380 + 24 = 404).

The resolver produced 0 false conflicts among the 330 non-conflict headline questions.

Real-Content Demo

The benchmark tests the relationship rules systematically. The real-content demo shows what happens when the same resolver is pointed at actual content.

The real-content build contains 149 documents: 57 scopes and 92 claims, drawn from 36 real TDS articles and 20 recent EmiTechLogic blog pages.

For a specific TDS article where no eligible claim exists, the resolver returns INSUFFICIENT rather than borrowing a value from another article. The baselines return values like submitted or draft instead.

The clearest example is a portfolio-wide question. Ask "what's the status of my whole TDS portfolio," and the resolver reports:

CONTEXTUAL across 34 scopes

rather than collapsing the portfolio into one status. The newest-value baseline answers submitted, and the retrieval baseline answers published, each picking a single value that's only true for one article.

The EmiTechLogic sitemap produces the same pattern: a blog-wide query spans 20 scopes, and the baselines collapse those multiple scoped claims into a single answer.

The real-content dataset intentionally contains no SUPERSEDED or CONFLICTING cases. Those relationship types are evaluated in the controlled held-out benchmark, not here.

Demo

https://youtu.be/35ab7M0CmBI

The demo is a silent walkthrough: the Sanity Studio data, terminal output, and benchmark results show the complete flow.

  1. Sanity Studio — a real claim document: subject, attribute, value, scope reference, and an explicit supersedes relationship.
  2. A live run against my EmiTechLogic blog data. Ask about the whole blog and the resolver reports CONTEXTUAL across 20 scopes. Ask about a page with no eligible claim and it returns INSUFFICIENT. Meanwhile the baselines collapse the scoped claims into a single answer regardless.
  3. The held-out benchmark result, run live against Sanity: 380/380 for the resolver, zero false conflicts and zero missed conflicts, against 276/380, 176/380, and 125/380 for the three baselines.

There's also a static, no-setup version of this: demo.html in the repo (below) is a self-contained page with the same real examples, viewable by opening the file directly in any browser, no server or credentials needed.

Sanity isn't failing to retrieve the information here. It provides the structured claims; the resolver decides how those claims relate before an answer comes out.

Code

Repository: https://github.com/Emmimal/claim-relationship-resolver

The repository contains:

  • resolver.py — the frozen resolver logic
  • RESOLVER_RULES.md — the frozen rules and amendments log, including two bugs I found and fixed during live testing
  • heldout/ — the 404-question benchmark and blind audit materials
  • real_content/ — the real datasets and demonstration script
  • demo.html — a static, self-contained demo page with real examples, no server needed
  • compare.py — runs the resolver and baselines against local, mock, or live Sanity data and checks local/live parity

compare.py --source local and --source mock reproduce the exact same results with no Sanity credentials needed, so the benchmark is checkable without any setup. --source live re-runs everything against the actual Sanity Context MCP endpoint.

How I Used Sanity

I used Sanity Context in GROQ mode rather than Knowledge Base mode, because the data consists of structured facts: subject, attribute, value, scope, version, effective date, and explicit supersedes relationships.

That structure is central to the project. It's what lets the agent reason about relationships between claims instead of treating retrieved text as isolated answers.

Why I used GROQ mode

I used a Sanity Context MCP endpoint backed directly by my structured dataset and queried it through the groq_query tool.

That choice follows the shape of the data. My claims aren't primarily prose documents that need semantic distillation. They're typed records with fields like subject, attribute, value, scope, version, effectiveFrom, and supersedes, and the resolver needs those fields intact because they're part of the relationship rules.

The challenge supports pointing an agent at a full dataset through Context MCP as an alternative to the Knowledge Base beta's document limit. I chose the dataset route because this resolver needs exact matches on subject and attribute, followed by deterministic relationship logic. I query the Context MCP groq_query tool directly rather than using semantic similarity for retrieval.

There's no LLM anywhere in the retrieval or resolution loop. Sanity provides the structured content through Context MCP; the resolver determines how the returned claims relate.

This also keeps the retrieval boundary explicit: Sanity answers "which structured claims match this subject and attribute?" and the Python resolver answers "what relationship do those claims have?"

I also discovered and fixed a subject-isolation bug during live testing: an early query could retrieve unrelated claims when multiple subjects shared the dataset. The fix and how I verified it are documented in RESOLVER_RULES.md.

Schema

There are two document types:

  • scope — hierarchical scopes connected through a parent reference, such as enterprise under all.
  • claim — a structured claim containing its subject, attribute, value, scope, version, effective date, and an optional supersedes array containing {claim, inScope} references.

The schema was deployed from a Sanity Studio project using sanity schema deploy.

Retrieval

The Python client talks to the Sanity Context MCP endpoint over JSON-RPC and calls the groq_query tool, using two query shapes: fetch claims for the requested subject and attribute, and fetch the scope hierarchy needed to determine containment.

There's no LLM involved in this retrieval or resolution loop. The query inputs come directly from the structured question fields.

Resolution

After retrieval, the resolver applies the frozen rules in sequence:

  1. Temporal validity — is the claim effective on or before the question's as-of date?
  2. Scope containment — does the claim apply to the requested scope, and does a narrower applicable scope take precedence over a broader one?
  3. Explicit supersession — does an explicit relationship replace an older claim, and does that relationship apply to the requested scope?
  4. Relationship classification — do the remaining claims produce a supported answer, contextual answers, a genuine conflict, or insufficient evidence?

Missing metadata is never silently guessed. A missing scope isn't assumed to mean global scope, and a missing effective date doesn't automatically make a claim eligible.

Sanity Project for Submission

  • Project ID: ouvupih6
  • Project Name: Claim Relationship Resolver
  • Dataset: production
  • Organization ID: o4oimb7tz
  • Dataset URL: https://ouvupih6.apicdn.sanity.io/v2024-01-01/data/query/production?query=*[_type=="claim"]
  • Studio URL: https://claim-relationship-resolver-emitechlogic.sanity.studio/

Top comments (0)