When your sources disagree, most systems pick one silently and are confidently wrong. I built Both Sides — an adjudication desk that makes you look at both sides first.
Live app: https://both-sides-eta.vercel.app
Source: https://github.com/aniruddhaadak80/both-sides
Agent endpoint: https://both-sides-eta.vercel.app/api/mcp
Sanity project: 4npxmu4m · dataset production
The problem, with a real example
I went looking for a knowledge base that would embarrass itself, and it took about thirty seconds.
Kyoto — Q34600 — has 35 different population claims published in Wikidata. Not 35 duplicates. 35 different numbers, attached to different references, retrieved on different dates. Osaka carries four population figures and three different founding dates (1889, roughly 1500, roughly 500) as equally plausible statements about the same city.
This is not an edge case. It is the normal condition of structured knowledge that many people contribute to over many years.
Any system that reads one of those values and answers confidently is making a choice it never told you about. An LLM with a keyword search over that data will confidently return whichever number it saw first.
Both Sides takes the opposite position: you cannot adjudicate a contradiction you have not seen in full.
What it does
Load an entity. Every competing claim for a property appears side by side, with its upstream rank, its reference URLs, its retrieval date, and the sentence that produced each score contribution. Then a deterministic engine rules on which claim governs — and you record the ruling, which gets sealed into a hash chain.
flowchart TB
classDef live fill:#22d3ee,color:#05202b,stroke:#0b7285
classDef engine fill:#a78bfa,color:#1b1033,stroke:#6d4fd0
classDef verified fill:#34d399,color:#04231a,stroke:#12805c
classDef infra fill:#94a3b8,color:#111827,stroke:#4b5563
A[Competing claims]:::live
B[Normalise rank refs date]:::engine
C[Five weighted factors]:::engine
D[Contribution per factor]:::engine
E{Tie on score}:::engine
F[Break on refs then id]:::engine
G[Leader and margin]:::verified
H[Verdict band]:::verified
I[Recommendation]:::verified
J[Store seal with ruling]:::infra
A --> B --> C --> D --> E
E -->|yes| F --> G
E -->|no| G
G --> H --> I --> J
The five factors, and the two rules it will not bend
| Factor | Weight | What it measures |
|---|---|---|
editorialRank |
0.28 | preferred / normal / deprecated upstream |
evidenceDepth |
0.22 | how many independent references back it |
corroboration |
0.20 | does the second source quote a matching figure |
recency |
0.16 | reference age, decaying over 12 years |
specificity |
0.14 | precise measurement versus a vague band |
Rule one: a deprecated claim cannot govern. This came directly out of a failing test. I had a deprecated claim carrying nine references beating an active claim with none — technically "more evidence", but upstream editors had already said they reject it. So deprecated is now a structural bar, not a low weight: it sorts last and reports why. No amount of piling on references overrides an explicit editorial rejection.
Rule two: exact ties break on reference count, then lexicographically on claim id. Two runs on identical input must produce identical output, or the hash chain means nothing.
Sanity behind it
The content layer is the Sanity Content Lake, read with GROQ, with the schema defined in TypeScript in a standalone Studio.
-
entity— the QID, labels, and the Wikipedia and OSM references -
dispute— one property, referencing its entity, withclaims[] -
claim— value, numeric, rank, reference count, reference URLs, retrieved date, precision -
sanity.agentContext— the scoped MCP configuration for agent access
The schema carries the idea that matters: a dispute is a document type, not a flag. You do not mark an entity as "disputed". You publish a dispute document with at least two claim objects that disagree. A single source of truth cannot express that, which is exactly why keyword search keeps returning confident nonsense.
flowchart LR
classDef live fill:#22d3ee,color:#05202b,stroke:#0b7285
classDef engine fill:#a78bfa,color:#1b1033,stroke:#6d4fd0
classDef verified fill:#34d399,color:#04231a,stroke:#12805c
classDef external fill:#fbbf24,color:#2a1d00,stroke:#b07908
classDef risk fill:#fb7185,color:#2b0710,stroke:#c2334d
A[Content Lake GROQ]:::live
B[Session imports]:::live
C{Any content?}:::verified
D[status live]:::verified
E[Sealed dated snapshot]:::risk
F[status fallback plus notice]:::risk
G[Upstream APIs]:::external
A --> C
B --> C
C -->|yes| D
C -->|no| E --> F
G -.->|import reads live| B
What is actually verified
I did not want to hand-wave this, so the repository ships a verifier that drives real HTTP against the deployment.
-
42 unit tests — engine determinism, boundary cases, empty and malformed input, corroboration scaling, canonical JSON, the seal chain, and a test that the tool manifest cannot drift from
tools/list - 76 end-to-end HTTP checks — health, corpus, engine, validation, create, read-back, update, agent mutation, idempotent retry, integrity replay, export, share, teardown, and the rendered GitHub link
- GitHub Actions green — typecheck, lint, tests, production build, and the full 76-check journey against a real Postgres service
-
Production store confirmed as
neon-postgres, not an in-memory map - A security audit that fails on a leak — it greps tracked files, then fetches all seven deployed pages and every client JavaScript chunk looking for the database URL, Neon keys, Sanity tokens or OIDC tokens
-
Content-Security-Policy and eight other headers, asserted by
scripts/check-headers.mjsagainst a real production build - Zero console errors in a Playwright pass across desktop and mobile
git clone https://github.com/aniruddhaadak80/both-sides.git
cd both-sides/web && npm install && npm run dev
No API keys. No account. With no environment variables at all it runs on an embedded PGlite database, so the first paint never breaks.
An agent that mutates through the same service layer
Nine MCP tools over JSON-RPC 2.0 at /api/mcp, published in public/mcp.json. The mutating ones call the same repository functions the interface uses, so the audit chain cannot be bypassed by going through the agent.
sequenceDiagram
participant A as Agent
participant M as MCP route
participant R as repository.ts
participant E as engine.ts
participant P as Postgres
A->>M: tools/call record_ruling
M->>R: createRuling with idempotencyKey
R->>E: adjudicate the dispute
E-->>R: ranked factors and verdict
R->>P: insert ruling and audit event
R->>P: insert idempotency key
R-->>M: ruling with SHA-384 seal
M-->>A: content JSON
Retrying with the same idempotencyKey returns the original ruling with idempotentReplay: true instead of creating a duplicate. Every tool is scoped to the calling anonymous session, so an agent cannot read anyone else's rulings.
Integrity you can replay without this app
flowchart LR
classDef verified fill:#34d399,color:#04231a,stroke:#12805c
classDef engine fill:#a78bfa,color:#1b1033,stroke:#6d4fd0
classDef risk fill:#fb7185,color:#2b0710,stroke:#c2334d
classDef infra fill:#94a3b8,color:#111827,stroke:#4b5563
A[Genesis seal]:::infra
B[Event n]:::engine
C[Canonical JSON]:::engine
D[seal n SHA-384]:::verified
E{Replay matches}:::verified
F[Report first broken link]:::risk
G[Tombstone retained]:::infra
A --> B --> C --> D --> E
E -->|yes| D
E -->|no| F
D --> G
seal_0 = SHA-384("both-sides/genesis/v1:" + scopeId)
seal_n = SHA-384( UTF-8(seal_{n-1}) || canonicalJson(event_n) )
Canonical JSON sorts keys recursively and drops undefined, so equal payloads always hash equally. Deleting a ruling leaves a tombstone, so replay still succeeds after a removal — the verifier asserts exactly that.
The bugs I hit, because they are the interesting part
PGlite exec silently dropped my bind parameters. exec() takes no params argument, so my adapter passed $1…$19 into a call that discarded them. Every insert failed with there is no parameter $1. Anything parameterised now goes through query().
Each route bundle opened its own database. Route handlers get separate module registries, so every bundle built its own PGlite instance. When the directory lock was contended my code fell back to a fresh in-memory database — so the tables created by /api/health were invisible to /api/rulings. Fixed with a process-wide singleton, and I removed the silent fallback that hid it.
A deprecated claim won on evidence. Covered above. My first instinct was to lower its weight; that was the wrong fix, because a weight can always be outvoted.
The CI journey failed and the reason was an assumption. My adapter sent every postgres:// URL through Neon's serverless HTTP driver, which cannot reach a normal Postgres server. The scheme is not the transport. It now selects on DATABASE_DRIVER and host detection, with a real TCP driver alongside the HTTP one.
My own landing page lied about a number. The stats strip said "10 agent tools" while the endpoint served nine, because the count was a hardcoded literal. There is now one manifest in src/lib/agent-tools.ts that the page, the console and the MCP route all read, plus a test that fails if tools/list drifts from it.
Corroboration was quietly broken. The factor compared the claim against every number in the source text, so a stray "8" in "area 64 km2, founded 8 AD" counted as a rival figure and pushed corroboration to zero for every large value. Figures are now filtered to a comparable magnitude first, and the evidence line shows only real competitors.
The embedded database did not work in a production build. It worked in dev, so I nearly shipped it. PGlite ships WebAssembly assets that must be read from node_modules at runtime, and bundling produced a build whose embedded adapter could not start, which would have broken every self-hoster with no database. Caught only because scripts/check-headers.mjs boots next start rather than next dev.
Where it stands, honestly
Three things are not what I would ship as finished:
-
The Content Lake has no published disputes yet. Seeding needs an Editor token, so today
/api/corpusserves the sealed snapshot and labels itselffallback. The import path is genuinely live: importing an entity reads Wikidata, Wikipedia and OSM at request time. - Vercel's free tier capped at 100 deployments per day, so the most recent commits are merged and CI-green but not yet on the live alias.
-
POST /api/rulingsreturned a 500 in two of three browser passes while 24 sequential and 36 concurrent writes all succeeded against the same deployment. I could not reproduce it, and I could not ship the logging that would diagnose it because of the deploy cap. I added a bootstrap retry and a real error envelope rather than pretending it is fixed.
Try it
Open https://both-sides-eta.vercel.app/desk, search Kyoto, import it, and open the population dispute. You will get 35 competing claims ranked with every factor exposed. Rule on one, export the dossier, then replay the chain.
Source and issues: https://github.com/aniruddhaadak80/both-sides
MIT licensed. Data courtesy of Wikidata (CC0 1.0), Wikipedia (CC BY-SA 4.0) and OpenStreetMap (ODbL). Both Sides ranks claims and records a human decision: it does not certify that the chosen value is correct.
#sanitychallenge






Top comments (0)