This is a submission for the Sanity Challenge, Path One: ship an agent that queries real content.
What I Built
Museums put glue, foam, paint and plastic in display cases next to objects. Some of these materials give off fumes that corrode metal or leave deposits. Conservators screen them with the Oddy test: strips of silver, copper and lead are sealed with the material for 28 days at 60 °C, then inspected for corrosion. Each strip is rated P (permanent use), T (temporary) or U (unsuitable).
A rating reads like a property of the material. It isn't. It describes one sample, tested by one lab, with one procedure.
- In a 2024 study, seven institutions ran their own versions of the test on the same ten materials. They disagreed on several of them.
- A structural adhesive, reported in a 2025 review as Terostat-9220, passed an Oddy test. Its emitted compound, TMP-ol, was later linked to white crystalline deposits on objects in display cases at several institutions.
- In a 2022 study of 60 commercial cellulose ethers, Klucel G bought in 2017 passed on every strip. Samples bought in 2021 from two suppliers corroded lead.
ArtifactGuard answers one question: for this material, this object and this display, what does the published evidence actually support? You pick the three, or type a question like "Is Terostat-9220 okay next to lead in a sealed case, not touching?". It returns one of six answers, each with the reasons, the ratings behind it and links to the sources:
| Answer | Meaning |
|---|---|
SUPPORTED |
The evidence shows no conflict. Not a safety guarantee. |
CONDITIONAL |
Usable only under the conditions the evidence states. |
CONFLICTED |
Sources disagree. |
INSUFFICIENT |
Not enough evidence for this combination. |
RETEST |
The evidence is old. |
AVOID |
The evidence shows a problem. |
It never says "safe" and never gives a percentage. There is no language model in the verdict: fixed rules read records in Sanity, so the same inputs give the same answer. If you leave something out or a name is ambiguous, it asks instead of guessing.
Limits, stated up front:
- The benchmark compares against a simple first-match keyword lookup that I wrote. It is not a production RAG system. I also wrote the expected answers, so the benchmark guards behaviour; it doesn't prove accuracy.
- Two products are synthetic test fixtures and are labelled as such.
- The data covers what seven published sources report. It is not a materials database, and it doesn't replace a conservator.
Demo
https://artifactguard.onrender.com (no login)
Three quick ways in:
-
The PASS wasn't enough: a flat lookup says
SUPPORTED; ArtifactGuard saysCONDITIONAL. - Same product, different batch: Klucel G from 2021 against lead.
- Labs disagree: the eight ratings behind an Ethafoam result.
Every result shows the Oddy ratings behind it, Knowledge Base guidance under the deductions it explains, and a trace of the Sanity calls. /compare runs 25 scenarios and /method explains the pipeline.
The site runs on a free host that sleeps when idle, so the first visit after a quiet spell can take a minute.
Code
https://github.com/sagarchive/artifactguard
npm run qa # lint, seed integrity, tests, benchmark, red-team checks, fuzz, HTTP smoke test
Plain Node 22, no runtime dependencies. The Sanity Studio schema is in studio/. The README has a plain-language summary, setup, the data provenance and the threat model.
How I Used Sanity
What Sanity does and doesn't do here. Sanity holds the facts as linked records and serves them live, so the agent can follow a test to its protocol, its product batch and its source. Those links make the warning "this pass missed a later problem" possible, which a flat table couldn't express. Sanity does not decide the verdict. Fixed rules do, and the Knowledge Base only explains.
Structured content. 10 document types (product, product revision, Oddy test, test protocol, emission observation, compound, interaction, artifact material, evidence claim, source), 363 documents and 848 native Sanity references. 141 of the 144 Oddy tests are real ratings from the three studies. The verdict follows the references: revision → Oddy tests → protocols, revision → emission → compound → interaction, and evidence claims that challenge what a passing test shows.
- Revisions carry supplier, purchase date and batch, which is how two Klucel G batches get different answers.
- Artifact materials carry the coupon metal they correspond to, so a result for lead reads the lead strip from every test instead of a blended rating.
- Protocols carry the vessel, volumes and coupon preparation each institution used, so the answer can say why labs disagree.
- The engine contains no product ids. Remove the challenged claim from the dataset and the "a PASS isn't enough" deduction disappears; a test checks this.
Context MCP, two endpoints.
- A dataset endpoint in GROQ mode, filtered to those document types. The agent calls
groq_queryfor the evidence graph. - A Knowledge Base endpoint. After the rules run, the agent reads the outline (
initial_context), searches each topic its own deductions raise withknowledge_base_search, and reads the top hit withknowledge_base_read. The flagship case reads two entries, not all eight. The Sanity calls are shown in the result's trace.
Knowledge Base. Eight entries, far under the 150-document beta budget: what an Oddy test does and doesn't cover, variation between laboratories, enclosure effects, the Terostat case, object vulnerability, flexible PVC, how to word evidence, and source hierarchy. The matching passage appears under the deduction it explains. It can't change the verdict.
Sanity's builder writes the entries, and my first build added claims my files never made. It called Terostat-9220 "also formulated as" Terostat MS 937, when the sources describe two formulations. Sanity's own conflict check caught that one. I fixed the rest with a standing instruction that limits entries to statements in the files, and wrote npm run verify:kb, which flags built sentences whose words aren't in my files. A lesson from it: an instruction is tied to source ids, so replacing a file silently detaches it, and the next build invented a ranking of source types.
Why both. The dataset holds facts with relationships. The Knowledge Base holds guidance that doesn't fit a schema. The rules only need the dataset to decide.
Checking the data against the papers. The dataset is generated from the papers' tables, not typed. Table 7 of the 2024 study is published as an image, so I transcribed it and checked it against the paper's own Table 8: all 63 per-coupon counts and 21 totals match. Every emission statement is stored with a quote I verified is in the paper text. Checking turned up problems in my first draft, which I fixed:
- My Oddy record for Terostat had an exact test date and per-metal results. The paper only says the adhesive passed a test run by museum staff. I deleted both and changed the schema, which had marked per-metal results as required.
- A second paper I'd cited for the Terostat claims never mentions Terostat.
- The Smithsonian study names the adhesive Terostat MS 937 and says only that formulation was found in their cases, while the review says Terostat-9220. The dataset records the disagreement and shows it with the result.
- Two of my rules were hardcoded to specific products. They now come from the data.
- The papers contradict themselves in places, for example a table rating four materials P for silver while the text says they tarnished it.
research/sources.mdlists these, and the dataset leaves out anything that depends on them.
What I'd change. I didn't use schema_explorer. With more time I'd let the agent discover document types with it instead of using a fixed GROQ projection.
Sanity Project Details
Project ID: 3yfxc6mg (dataset production, public)
Public dataset: https://3yfxc6mg.api.sanity.io/v2025-02-19/data/query/production?query=count(*)





Top comments (0)