This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
I'm Italian, and part of my work is checking the Italian copy of software websites. The same mistakes keep shipping: "Fidato da" for "Trusted by", "Utensileria" (a hardware store) as the menu label for "Tools", "5 Mld+ Altoparlanti nativi" (native loudspeakers) for "5B+ native speakers". A general model catches some of these and misses others, and when it does catch one it cannot tell you which house rule it applied.
So I built a reviewer agent. You give it a string, where it appears and, if you have it, the English source. It reads the team's Italian style guide from a Sanity Knowledge Base, answers ok or error, proposes the corrected text, and lists the Knowledge Base entries it relied on. The style guide is structured content in Sanity: 15 rules and 19 glossary terms, each rule carrying real strings that shipped on live product pages next to their native correction.
I also measured it, three times, and I report every run. On 12 held-out string pairs the Knowledge Base never saw, the latest grounded agent got 9 pairs fully right and the same model without the Knowledge Base got 9. The Knowledge Base started four pairs behind and is now level; it has not pulled ahead. The details are below.
Demo
Three real reviews in the local web UI, grounded in the Knowledge Base. Each took 12 to 17 seconds on the free Gemini tier; the GIF cuts the wait, and every card shows the real time.
The same three reviews as a video: https://vimeo.com/1231260014
A review looks like this (real output from the Sep 30 test run, gemini-3.1-flash-lite, input Fidato da oltre 100.000 imprese):
{
"verdict": "error",
"reason": "The phrase 'Fidato da' is a literal calque of 'Trusted by' and is considered unnatural in Italian marketing copy.",
"evidence": [
"Fidato da"
],
"corrected": "Scelto da oltre 100.000 imprese",
"rules": [
"Calques, False Friends & Idiom Errors/\"Trusted by\" calque (major)"
]
}
Code
https://github.com/exochard/italian-copy-reviewer
TypeScript on Node 22 with the Vercel AI SDK; a Sanity Studio schema; a small Python script that turns the rules into Sanity documents. npm test runs without network.
How I Used Sanity
The content model. Two document types in the Studio schema:
-
l10nRule: title, slug, category (calque, agreement, number format…), severity, the rule in plain words, preferred forms, forms to avoid, and an array of real shipped examples (product type, where the text appears, English source, shipped Italian, native correction). -
glossaryTerm: the English term, the preferred Italian, the forms to avoid, a note.
The examples are the point. "Don't translate literally" is useless to a model; "a software company shipped Fidato da oltre 9,500 marchi e organizzazioni globali and a native reviewer changed it to Scelto da oltre 9.500 marchi e organizzazioni in tutto il mondo" is not.
The Knowledge Base. Two sources: the production dataset above, and Microsoft's public Italian Localization Style Guide (a PDF). They sometimes disagree. When it built the Knowledge Base, Sanity Context flagged three such conflicts as Issues instead of merging them silently; I read each one and resolved it in the dashboard before the final runs.
The agent. The AI SDK connects to the Context MCP endpoint in knowledge-base mode. The system prompt tells the model to call initial_context first and then knowledge_base_read on the entries that could apply. The output is a typed object (verdict, reason, the wrong fragments quoted from the text, corrected text, cited entry paths), so every answer says which rules it used.
The first measured run taught me something. With the Knowledge Base the agent flagged correct text more often than the model alone, because it applied entries whose trigger was not there: it flagged "AI" in a sentence with no "AI" in it, and asked for "oltre" on a number that had no "+". So every flag now has to quote the wrong words, and code drops a flag whose quote is not in the text. The prompt also says an entry applies only when the text contains what the entry is about. I tuned this on the training pairs only.
That was not enough. The second test run showed flags whose quote was the whole string: the agent read the capitalisation entry and called Automazione batch con un clic and RECENSIONI DEGLI UTENTI "Title Case". So a flag that relies only on the Title Case, "+" or AI entry now survives only when the text holds that entry's trigger: a capitalised word in mid-sentence (ALL CAPS never counts), a plus sign, or "AI" and "IA" mixed in one string, which is exactly what the AI entry says. Re-applied offline to the saved training outputs, this took the grounded agent from 12/17 to 16/17 without losing a caught bug. It stays off the model-alone arm, where it would have dropped a real bug.
One fix came from reading a test failure, and I'd rather say so. On USER REVIEW both arms read the text between the prompt's <<< and >>> markers as an empty placeholder. The prompt now says the text between the markers is the literal string.
In 24 of 24 grounded test runs the agent read at least one entry before answering.
Keeping the test honest. The benchmark has 29 pairs. The train half feeds the Knowledge Base; the test half never enters it, and a unit test asserts the two halves are disjoint. Each pair is scored only when the agent flags the shipped string and leaves the native fix alone, so a reviewer that flags everything scores zero.
Every test run, same 12 pairs, gemini-3.1-flash-lite, pairs fully right:
| Test run | Changed before it (tuned on training pairs) | Model alone | + Sanity Knowledge Base |
|---|---|---|---|
| Sep 28 | first version | 7/12 | 4/12 |
| Sep 29 | quote check; "entry applies only when triggered" prompt; two entries reworded | 9/12 | 5/12 |
| Sep 30 | per-rule trigger check; literal-text prompt | 9/12 | 9/12 |
The latest run in full:
| Bugs caught | Fixes left alone | Pairs fully right | |
|---|---|---|---|
| gemini-3.1-flash-lite alone | 11/12 | 10/12 | 9/12 |
| gemini-3.1-flash-lite + Sanity Knowledge Base | 10/12 | 11/12 | 9/12 |
The model alone moves between runs too: both arms share the prompt, and the model is not fully repeatable even at temperature 0. Twelve pairs is a small test, so one pair is eight points.
What still fails in the Sep 30 run: both arms passed Transcrizione, which should be Trascrizione, so neither checks spelling letter by letter. With the Knowledge Base the agent called the correct Ottieni informazioni utili e pratiche a calque of "Get", and let an English string, Larger file per task, through as correct. The model alone failed on two correct strings instead: it carried the "+" of the English source into oltre 100 siti and called it redundant, and it objected to the preposition in per attività. What the Knowledge Base does well is the explanation: every flag names the entry it came from, so a reviewer can see why it fired and fix the entry instead of the prompt.
Sanity Project Details
Project ID: 3ezm3tm1 (dataset production). The schema is in studio/schemaTypes/ and the seed documents in content/seed.ndjson.
How this was made
I wrote the benchmark strings and their corrections from my own QA audits; they predate the challenge, and the brand names are masked. The code and the Knowledge Base content were written during the challenge with an AI coding assistant (Claude Code): I set the spec and the acceptance checks, it drafted the code, and I reviewed, ran and tested every part. The numbers above come from npm run eval runs (results/eval-*.json) and are reported as measured, including the ones that went against the Knowledge Base. src/rescore.ts re-applies the flag checks to a saved run without calling any model.

Top comments (0)