DEV Community

ANIRUDDHA  ADAK
ANIRUDDHA ADAK Subscriber

Posted on

I built an adjudication desk for data that contradicts itself (Both Sides)

Sanity Challenge Path One Submission

When your sources disagree, most systems pick one silently and are confidently wrong. I built Both Sides — an adjudication desk that makes you look at both sides first.

Live app: https://both-sides-eta.vercel.app
Source: https://github.com/aniruddhaadak80/both-sides
Agent endpoint: https://both-sides-eta.vercel.app/api/mcp
Sanity project: 4npxmu4m · dataset production

The landing page showing a real contradiction with every factor itemised

The problem, with a real example

I went looking for a knowledge base that would embarrass itself, and it took about thirty seconds.

Kyoto — Q34600 — has 35 different population claims published in Wikidata. Not 35 duplicates. 35 different numbers, attached to different references, retrieved on different dates. Osaka carries four population figures and three different founding dates (1889, roughly 1500, roughly 500) as equally plausible statements about the same city.

This is not an edge case. It is the normal condition of structured knowledge that many people contribute to over many years.

Any system that reads one of those values and answers confidently is making a choice it never told you about. An LLM with a keyword search over that data will confidently return whichever number it saw first.

Both Sides takes the opposite position: you cannot adjudicate a contradiction you have not seen in full.

What it does

Load an entity. Every competing claim for a property appears side by side, with its upstream rank, its reference URLs, its retrieval date, and the sentence that produced each score contribution. Then a deterministic engine rules on which claim governs — and you record the ruling, which gets sealed into a hash chain.

The two-podium adjudication desk, with the leading claim marked

flowchart TB
  classDef live fill:#22d3ee,color:#05202b,stroke:#0b7285
  classDef engine fill:#a78bfa,color:#1b1033,stroke:#6d4fd0
  classDef verified fill:#34d399,color:#04231a,stroke:#12805c
  classDef infra fill:#94a3b8,color:#111827,stroke:#4b5563

  A[Competing claims]:::live
  B[Normalise rank refs date]:::engine
  C[Five weighted factors]:::engine
  D[Contribution per factor]:::engine
  E{Tie on score}:::engine
  F[Break on refs then id]:::engine
  G[Leader and margin]:::verified
  H[Verdict band]:::verified
  I[Recommendation]:::verified
  J[Store seal with ruling]:::infra

  A --> B --> C --> D --> E
  E -->|yes| F --> G
  E -->|no| G
  G --> H --> I --> J

The five factors, and the two rules it will not bend

Factor Weight What it measures
editorialRank 0.28 preferred / normal / deprecated upstream
evidenceDepth 0.22 how many independent references back it
corroboration 0.20 does the second source quote a matching figure
recency 0.16 reference age, decaying over 12 years
specificity 0.14 precise measurement versus a vague band

Rule one: a deprecated claim cannot govern. This came directly out of a failing test. I had a deprecated claim carrying nine references beating an active claim with none — technically "more evidence", but upstream editors had already said they reject it. So deprecated is now a structural bar, not a low weight: it sorts last and reports why. No amount of piling on references overrides an explicit editorial rejection.

Rule two: exact ties break on reference count, then lexicographically on claim id. Two runs on identical input must produce identical output, or the hash chain means nothing.

Sanity behind it

The content layer is the Sanity Content Lake, read with GROQ, with the schema defined in TypeScript in a standalone Studio.

  • entity — the QID, labels, and the Wikipedia and OSM references
  • dispute — one property, referencing its entity, with claims[]
  • claim — value, numeric, rank, reference count, reference URLs, retrieved date, precision
  • sanity.agentContext — the scoped MCP configuration for agent access

The schema carries the idea that matters: a dispute is a document type, not a flag. You do not mark an entity as "disputed". You publish a dispute document with at least two claim objects that disagree. A single source of truth cannot express that, which is exactly why keyword search keeps returning confident nonsense.

Every entity and disputed property in the corpus

flowchart LR
  classDef live fill:#22d3ee,color:#05202b,stroke:#0b7285
  classDef engine fill:#a78bfa,color:#1b1033,stroke:#6d4fd0
  classDef verified fill:#34d399,color:#04231a,stroke:#12805c
  classDef external fill:#fbbf24,color:#2a1d00,stroke:#b07908
  classDef risk fill:#fb7185,color:#2b0710,stroke:#c2334d

  A[Content Lake GROQ]:::live
  B[Session imports]:::live
  C{Any content?}:::verified
  D[status live]:::verified
  E[Sealed dated snapshot]:::risk
  F[status fallback plus notice]:::risk
  G[Upstream APIs]:::external

  A --> C
  B --> C
  C -->|yes| D
  C -->|no| E --> F
  G -.->|import reads live| B

What is actually verified

I did not want to hand-wave this, so the repository ships a verifier that drives real HTTP against the deployment.

  • 42 unit tests — engine determinism, boundary cases, empty and malformed input, corroboration scaling, canonical JSON, the seal chain, and a test that the tool manifest cannot drift from tools/list
  • 76 end-to-end HTTP checks — health, corpus, engine, validation, create, read-back, update, agent mutation, idempotent retry, integrity replay, export, share, teardown, and the rendered GitHub link
  • GitHub Actions green — typecheck, lint, tests, production build, and the full 76-check journey against a real Postgres service
  • Production store confirmed as neon-postgres, not an in-memory map
  • A security audit that fails on a leak — it greps tracked files, then fetches all seven deployed pages and every client JavaScript chunk looking for the database URL, Neon keys, Sanity tokens or OIDC tokens
  • Content-Security-Policy and eight other headers, asserted by scripts/check-headers.mjs against a real production build
  • Zero console errors in a Playwright pass across desktop and mobile

Recording a ruling, with the seal reported back

git clone https://github.com/aniruddhaadak80/both-sides.git
cd both-sides/web && npm install && npm run dev
Enter fullscreen mode Exit fullscreen mode

No API keys. No account. With no environment variables at all it runs on an embedded PGlite database, so the first paint never breaks.

An agent that mutates through the same service layer

Nine MCP tools over JSON-RPC 2.0 at /api/mcp, published in public/mcp.json. The mutating ones call the same repository functions the interface uses, so the audit chain cannot be bypassed by going through the agent.

The agent console issuing a real tools/list call

sequenceDiagram
  participant A as Agent
  participant M as MCP route
  participant R as repository.ts
  participant E as engine.ts
  participant P as Postgres

  A->>M: tools/call record_ruling
  M->>R: createRuling with idempotencyKey
  R->>E: adjudicate the dispute
  E-->>R: ranked factors and verdict
  R->>P: insert ruling and audit event
  R->>P: insert idempotency key
  R-->>M: ruling with SHA-384 seal
  M-->>A: content JSON

Retrying with the same idempotencyKey returns the original ruling with idempotentReplay: true instead of creating a duplicate. Every tool is scoped to the calling anonymous session, so an agent cannot read anyone else's rulings.

Integrity you can replay without this app

flowchart LR
  classDef verified fill:#34d399,color:#04231a,stroke:#12805c
  classDef engine fill:#a78bfa,color:#1b1033,stroke:#6d4fd0
  classDef risk fill:#fb7185,color:#2b0710,stroke:#c2334d
  classDef infra fill:#94a3b8,color:#111827,stroke:#4b5563

  A[Genesis seal]:::infra
  B[Event n]:::engine
  C[Canonical JSON]:::engine
  D[seal n SHA-384]:::verified
  E{Replay matches}:::verified
  F[Report first broken link]:::risk
  G[Tombstone retained]:::infra

  A --> B --> C --> D --> E
  E -->|yes| D
  E -->|no| F
  D --> G
seal_0 = SHA-384("both-sides/genesis/v1:" + scopeId)
seal_n = SHA-384( UTF-8(seal_{n-1}) || canonicalJson(event_n) )
Enter fullscreen mode Exit fullscreen mode

Canonical JSON sorts keys recursively and drops undefined, so equal payloads always hash equally. Deleting a ruling leaves a tombstone, so replay still succeeds after a removal — the verifier asserts exactly that.

The bugs I hit, because they are the interesting part

PGlite exec silently dropped my bind parameters. exec() takes no params argument, so my adapter passed $1…$19 into a call that discarded them. Every insert failed with there is no parameter $1. Anything parameterised now goes through query().

Each route bundle opened its own database. Route handlers get separate module registries, so every bundle built its own PGlite instance. When the directory lock was contended my code fell back to a fresh in-memory database — so the tables created by /api/health were invisible to /api/rulings. Fixed with a process-wide singleton, and I removed the silent fallback that hid it.

A deprecated claim won on evidence. Covered above. My first instinct was to lower its weight; that was the wrong fix, because a weight can always be outvoted.

The CI journey failed and the reason was an assumption. My adapter sent every postgres:// URL through Neon's serverless HTTP driver, which cannot reach a normal Postgres server. The scheme is not the transport. It now selects on DATABASE_DRIVER and host detection, with a real TCP driver alongside the HTTP one.

My own landing page lied about a number. The stats strip said "10 agent tools" while the endpoint served nine, because the count was a hardcoded literal. There is now one manifest in src/lib/agent-tools.ts that the page, the console and the MCP route all read, plus a test that fails if tools/list drifts from it.

Corroboration was quietly broken. The factor compared the claim against every number in the source text, so a stray "8" in "area 64 km2, founded 8 AD" counted as a rival figure and pushed corroboration to zero for every large value. Figures are now filtered to a comparable magnitude first, and the evidence line shows only real competitors.

The embedded database did not work in a production build. It worked in dev, so I nearly shipped it. PGlite ships WebAssembly assets that must be read from node_modules at runtime, and bundling produced a build whose embedded adapter could not start, which would have broken every self-hoster with no database. Caught only because scripts/check-headers.mjs boots next start rather than next dev.

Where it stands, honestly

Three things are not what I would ship as finished:

  1. The Content Lake has no published disputes yet. Seeding needs an Editor token, so today /api/corpus serves the sealed snapshot and labels itself fallback. The import path is genuinely live: importing an entity reads Wikidata, Wikipedia and OSM at request time.
  2. Vercel's free tier capped at 100 deployments per day, so the most recent commits are merged and CI-green but not yet on the live alias.
  3. POST /api/rulings returned a 500 in two of three browser passes while 24 sequential and 36 concurrent writes all succeeded against the same deployment. I could not reproduce it, and I could not ship the logging that would diagnose it because of the deploy cap. I added a bootstrap retry and a real error envelope rather than pretending it is fixed.

Try it

Open https://both-sides-eta.vercel.app/desk, search Kyoto, import it, and open the population dispute. You will get 35 competing claims ranked with every factor exposed. Rule on one, export the dossier, then replay the chain.

The interface on a 390px viewport

Source and issues: https://github.com/aniruddhaadak80/both-sides

MIT licensed. Data courtesy of Wikidata (CC0 1.0), Wikipedia (CC BY-SA 4.0) and OpenStreetMap (ODbL). Both Sides ranks claims and records a human decision: it does not certify that the chosen value is correct.

#sanitychallenge

Top comments (0)