This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
An agent that answers "what computer do I need to run this AI model?" — including "don't buy, rent" and
"you already have enough". It reads a Sanity dataset of criteria, not a catalog: laws (formulas for weights
memory, KV cache, speed ceiling, own-vs-rent breakeven), machine-checkable rules, 11 solution paths
(keep what you have · upgrade · used · one GPU · multi-GPU · unified memory · MoE experts in system RAM ·
CPU-only small model · smaller model · cloud API · rented GPU), run reports and reference hardware.
New models appear every week, so anything not in the base is looked up on the web by the base's own checklist.
- Sanity project:
onwa0wvs, public datasetv2(200 documents), Studio: https://hw-for-ai-lab-v2.sanity.studio - Agent connection: Sanity Context MCP on the full dataset with embeddings enabled
(
/context/mcp/onwa0wvs/v2); all numbers below are measured through this connection - Second entry: a Knowledge Base (
kbW7wsbtJQkl, built from the core 138 documents of the same dataset, refreshed weekly) behind its own Context MCP endpoint in KB mode — created and connected, not measured separately -
Demo — try it with your own agent: https://hw-advisor.helgardorlm.tech (MCP
https://hw-advisor.helgardorlm.tech/mcp, no login) · Replay of all 27 measured runs: https://helgard-orlm.github.io/sanity-ai-hardware-advisor/ - Repo: https://github.com/helgard-orlm/sanity-ai-hardware-advisor
Demo
Live — bring your own agent (no login, no keys): https://hw-advisor.helgardorlm.tech · MCP endpoint https://hw-advisor.helgardorlm.tech/mcp
Replay of all 27 measured runs: https://helgard-orlm.github.io/sanity-ai-hardware-advisor/
Code
https://github.com/helgard-orlm/sanity-ai-hardware-advisor (the public proxy with check_answer is in public_mcp/)
Sanity Project Details
- Project ID
onwa0wvs, public datasetv2— 200 documents, 15 types (law, rule, solutionPath, aiModel, gpu, cpu, runReport, offer, cloudOffer, situationTemplate, …) - Query it directly: https://onwa0wvs.api.sanity.io/v2025-01-01/data/query/v2?query=*[_type=="law"]
- Studio: https://hw-for-ai-lab-v2.sanity.studio
Try it with your own agent (no keys, no login)
The demo is a public MCP endpoint, not a hosted chatbot: you bring the agent, the base brings the knowledge and the checking.
Add https://hw-advisor.helgardorlm.tech/mcp to ChatGPT (developer mode), Claude, Claude Code, Codex, Cursor or VS Code.
There are one-click buttons for Cursor and VS Code on the page. Or paste one line into any agent that can run commands:
"Read https://hw-advisor.helgardorlm.tech/setup.md and follow it."
Behind it is a ~230-line stdlib proxy. It passes the four Sanity Context MCP tools through unchanged and keeps the token
server-side. It puts today's date and the advisor's work order in front of initial_context. It adds one tool, check_answer,
which is the same validator the app uses: it recomputes every calculation with the law.formula stored in Sanity.
In the first outside test (Codex, its own subscription), check_answer rejected the first draft with 4 errors. The second draft passed:
5 calculations recomputed, 11/11 solution paths walked. A test from ChatGPT (developer mode) also went through to a passing check. Agents that can only browse get the same tools as plain HTTPS links
(/api/context, /api/query?q=…, /api/check), listed in /llms.txt.
Known limit: the check covers numbers, completeness, sources and identity. It does not cover whether a verdict makes sense. In one test an agent said yes to "MoE experts in RAM" for a dense model, and the check passed.
How I Used Sanity
Sanity Context MCP tools the agent calls: initial_context, groq_query, schema_explorer, array_field_reader (plus our check_answer). What it does with the content it retrieves — the structure is used, not just searched:
-
Paths are data. The agent must give a yes/no/maybe with a number for every
solutionPathdocument. Before this, the same model answered a 321B-MoE question with "4×H200" and never mentioned the one-GPU + large-RAM path. -
Rules are the search checklist. For any model/card/CPU that is not in the base, the app requires the
fields that the base's
ruledocuments test for that type (e.g.cpu.instructionSets,aiModel.variants.fileGb). -
Laws are recomputed by code. The agent writes every calculation with its inputs into a hidden check block;
the app evaluates the
law.formulafrom Sanity and sends the answer back if the number is wrong. - The person's situation is stored, not remembered: country, budget, existing PC survive a change of mind.
The same code checks the base itself: every law variable and rule path must be a real schema field
(coverage check), and laws get property tests. That caught the filler model writing an own-vs-rent formula that
added electricity instead of subtracting it and billed rent for 24 h/day — 12.3 vs 110 months on the same inputs.
A real trace (GLM-5.3-Flash, not in the base) — open it in the replay
- GROQ over the base → no
aiModeldocument → 8 web searches by the base's own checklist: the model card andconfig.json(layers, KV heads, head dim), file sizes, and the GPU fields the base's rules test (slots, bandwidth, compute capability). - Laws from Sanity: weights
320 × 4 / 8 = 160 GB(NVFP4), KV for 8k context 22.5 GiB (flagged as rough for hybrid attention). -
Validator round 1 sent the answer back: the speed ceiling was written as 19.91 tok/s — code evaluated the
law.formulafrom Sanity with the agent's own inputs and got 199.1 (a 10× slip); plus three entities missing fields that the base's rules check (e.g.variants.fileGb,slots). - Round 2: no errors. All 11 solution paths get a verdict — multi-GPU = yes, MoE-in-RAM / unified memory = maybe, cloud = no (the person ruled it out) — and the build is 3 × RTX PRO 6000 (288 GB) with the numbers above.
Other catches by the same recomputation in the measured run: KV cache 0.949 → 6.75 GiB (A4), monthly electricity cost 0.675 → 2.025 (A5).
Measured, honestly
Blind grader (a different model, sees only the dialogue and a frozen checklist), 9 scenarios × 3 runs,
78 checklist items, same base state for all arms:
| version | score | median time |
|---|---|---|
| v2: base + instructions | 47/78 | 62 s |
| v3: + solution paths, validator | 53/78 | 104 s |
| v3.1 via Sanity Context MCP: + rules-as-checklist, recomputed laws, dated prices | 59/78 | 181 s |
What this does not show: that structured retrieval beats plain text for reading. The same 200 documents
pasted as text notes scored the same as v3 (53/78) — at this size the model reads either. Where structure
mattered in our runs is checking: fewer first answers rejected by the validator with structured queries
(8/27 vs 14/27), and the gains of v3.1 are exactly in the scenarios where code recomputes Sanity's laws
(KV for 8 users 7/9 → 9/9, own vs rent 5/12 → 7/12, "don't buy" 4/6 → 6/6, v2 → v3.1). The price: answers take ~3× longer.
An earlier full run of v3.1 scored 62/78 with two bugs in my own checker (nested fields, European number format) that sent
correct answers back for rework; after fixing them the clean run scored 59/78 — with N=3 per scenario, 59 vs 62 is noise.
How it was built
Architect method: a risk catalog (16 classes of agent failure seen in earlier sessions) → acceptance scenarios
written before changes → build → blind grading. The base was filled by a model (GPT Luna) from an empty schema and
two prompts; I never wrote data by hand. An independent judge model watched the session and corrected the method
several times (frozen checklist, snapshot baseline, text control).
Known limits
- Answers are slow (minutes) because the validator can send them back up to twice.
- Ratings N=3 per scenario; differences of 1–2 points are noise.
- "unknown after search" is allowed and can be abused for vague entities.
- In the measured run 12 of 27 answers came out in Russian to English questions (an account language setting leaked in); the app now checks the answer's language in code and sends it back (3/3 English on a re-test).
Top comments (0)