This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
When a CRM sends closed deals back to Google and Meta as offline conversions, every email, phone and name goes out as a SHA-256 hash. The hash either matches the one the platform computed or it does not, and a mismatch is silent: the upload succeeds, the conversion lands, and it is never attributed. The campaign just looks worse than it is.
The two platforms do not normalize the same way, so one shared function is wrong for at least one of them:
| Input | Google (Data Manager API) | Meta (Conversions API) |
|---|---|---|
Jane.Doe+promo@GMAIL.com |
janedoe@gmail.com |
jane.doe+promo@gmail.com |
(650) 555-1212, US |
+16505551212 |
16505551212 |
O'Brien |
o'brien |
obrien |
ZIP 94035-1234, US |
94035-1234, in clear |
94035, hashed |
Each column is right for its own destination.
conversion-preflight is an agent a developer asks before an upload. It answers rules questions with citations, checks a CRM export row by row for both platforms, and audits the implementation against the current documentation. It is built on one rule: three sources, three jobs, never blurred.
- A Sanity knowledge base says what each platform documents.
- A rule table says what the code does, one entry per platform and field, published to a Sanity dataset as structured documents.
- A small MCP server of tested code is the only place a normalized value or a digest is produced.
When those disagree, the agent says so and quotes both sides instead of picking one.
Demo
A real run over eight synthetic CRM rows, trimmed to three of its findings:
$ python -m agent.run --csv examples/leads.csv --region US "Check these rows before tonight's upload to Google and Meta."
-> preflight_file examples/leads.csv (6 columns mapped, id=Lead ID, region=US)
-> initial_context
-> groq_query *[_type == "implementationProfile"][0]{globalDivergences}
-> knowledge_base_read ['normalization_and_hashing/google', 'normalization_and_hashing/meta']
-> groq_query *[_type == "normalizationRule" && field in ["phone","country","first_name",...
## L-1006
- Google and Meta, phone `+1 (415) 555 0142 ext 7`: needs_review because it contains a letter (U+0065).
- Google says every character after `+` must be a digit [normalization_and_hashing/google].
- Meta says to remove symbols and letters [normalization_and_hashing/meta].
- Refusing letters instead of stripping them is a declared divergence for both platforms
[rule: google/phone] [rule: meta/phone].
- Fix at source: store the extension in a separate field and send only the number.
- Google, first_name `Mrs. Carol`: warning. Google says to omit prefixes such as `Mrs.` from
`given_name` [normalization_and_hashing/google]. The code doesn't remove them, which is declared
[rule: google/first_name]. Fix at source: split the title into its own field.
## L-1008
- Google and Meta, email: invalid because it contains U+200B. ...
Every citation says which source it came from: [normalization_and_hashing/...] is the knowledge base, what the platform documents; [rule: ...] is the dataset, what the code declares.
Digests never pass through the model's text. The agent sees the header and three rows and sends back a column mapping; preflight_file reads every row, writes the upload file keyed by lead id, and returns the totals and the flagged cells. The first version had the model copy every cell into the tool call, which an adversarial review measured at 561,701 characters of tool output for 100 rows. The published run returned 3,216 for the eight. The same review caught a worse thing in an earlier demo: for a phone the tool had refused, the model wrote the "corrected" number itself. The prompt now says to describe the fix at the source and never to compose the value.
Code
vinimabreu
/
conversion-preflight
Preflight agent for Google and Meta offline conversion uploads: rules from a Sanity knowledge base, declared rules from a Sanity dataset, digests from tested code.
conversion-preflight
An agent that checks customer data before an offline conversion upload to Google and Meta. What the platforms document comes from a Sanity knowledge base built from their own pages. What the code declares about its own rules comes from structured documents in a Sanity dataset, queried with GROQ. Every normalized value and SHA-256 digest comes from tested code. Where those disagree, the agent says so instead of picking one.
A real run over eight synthetic CRM rows, trimmed to three of its findings (the full transcript is in examples/runs/):
$ python -m agent.run --csv examples/leads.csv --region US "Check these rows before tonight's upload to Google and Meta."
-> preflight_file examples/leads.csv (6 columns mapped, id=Lead ID, region=US)
-> initial_context
-> groq_query *[_type == "implementationProfile"][0]{globalDivergences}
-> knowledge_base_read ['normalization_and_hashing/google', 'normalization_and_hashing/meta']
-> groq_query *[_type == "normalizationRule" && field in ["phone","country","first_name","last_name","em
[6 turns, 14.3s, in 6+31541 cached, out 1895 tokens; transcript runs/...;…The agent is Claude Code run headless (claude -p) with every built-in tool switched off and only three MCP servers allowed: the knowledge base endpoint, the dataset endpoint, and the local preflight server. The whole prompt is one file, agent/system.md. It runs on my existing plan, so nothing here is billed per call.
202 tests. The suite checks every worked example with an input that the platforms publish (including Meta's own digests) against the code, pins every declared divergence to the code's actual behaviour, and exercises the MCP server over real stdio.
How I Used Sanity
The knowledge base, and what building it taught me
I pointed a Sanity Context knowledge base at three sources: Google's Data Manager "Format user data" guide, Google Ads' "Manage offline conversions" page, and Meta's customer information parameters. Meta's page renders client side and the crawl came back empty, so I extracted its text and uploaded it as a file source, with the URL and retrieval date at the top.
The purpose text does most of the steering. Mine says who asks (engineers about to upload), what leads (which endpoint, how each platform wants each field, where the two differ), what stays out (bidding, pricing, reporting), and, after the first builds, that each platform's rules apply to that platform only and that a rule's qualifiers and the pages' worked examples are kept as written.
Four things happened on the way that shaped the agent:
-
The first crawl indexed one page in fifteen languages. The "15 documents" from the Data Manager site were the formatting page plus fourteen
?hl=translations, and a later check pulled nineteen translations of the Google Ads page. The first entry I read cited a Hebrew page. "Ignore query parameters" did not stop it; I removed the translations in the crawled-pages view. -
Entries are summaries, and summaries drop words. Google says to trim "leading, trailing and intermediate whitespace" in an email address. One build kept "leading/trailing" and lost "intermediate". Another wrote "5-digit equivalents only" for non-US postal codes, which no page says, and the agent then told me to cut the Brazilian CEP
24020-005to24020. A single-entry rebuild once wrote "Base64 encoding is not used", true of the Google Ads samples and false for the Data Manager API. - Rebuilds are not deterministic, and instructions are what stick. A detail fixed in one build could vanish in the next. An instruction is honored on every build; when I saved the one for "intermediate", Sanity checked it against the entries itself and listed the two pages that contradicted it before rebuilding them.
-
Issues found a bug in Google's own sample code. A build opened a conflict between two entries about the Gmail plus-suffix. One side came from Google's Python sample, so I read it:
normalize_and_hash_email_addressremoves the.and+characters from the local part for every domain and keeps the text after the+. The page's own rule, and its Java sample, apply that only togmail.comandgooglemail.comand remove everything from the+on. Code copied from the Python sample turnsfirst.last@company.comintofirstlast@company.com. I resolved the issue with the documented rule; two other issues were differences between the platforms, not contradictions, and I dismissed them.
The dataset: the code's rules as structured content
The implementation's rule table lives in the production dataset as normalizationRule documents, one per platform and field, plus an implementationProfile that references them. Each rule carries what the code does in plain words, the platform's words when the code was written (verbatim, with the URL and retrieval date), the knowledge base entry that should agree today (kbEntry), the platforms' own worked examples, and the differences the code makes on purpose. A sync command publishes it from the Python module the tests pin, and a --check mode fails when the dataset and the code drift apart. The schema is a small Studio, deployed at conversion-preflight.sanity.studio.
A second Context MCP endpoint serves that dataset, filtered to those two types, and the agent reads it with groq_query. This is where the two halves of Sanity Context earned their separate places. The knowledge base is the right tool for prose that gets rewritten into entries; the dataset is the right tool for the thing that must come back exactly as written, because a query returns the quote, not a summary of it. Two things I learned the hard way: the endpoint answers "Only datasets with deployed Studio applications are supported" until a Studio with the schema is deployed, and a document id with a dot in it (normalizationRule.google.email) is a private path in Sanity, so the ids use hyphens.
What the agent does with it
-
Rules questions:
initial_contextonce, thenknowledge_base_readon the entries it needs, and every rule it states carries the entry path, like[normalization_and_hashing/meta]. -
Batch checks: it maps CRM columns to fields, calls
preflight_fileonce, and explains each refusal and each warning with the documented rule behind it. Before calling a result a conflict with the docs, it checks the rule's declared divergences. -
Audits: for every rule it compares the knowledge base entry (the documentation today), the quote the code was written against, and the implementation, and sorts what it finds: code and documentation disagree, entry and quote disagree, declared, or agrees. The audit is what flagged the lost "intermediate", the invented "5-digit equivalents", the "Base64 is not used" regression and Meta's self-contradicting UK example. It also found behaviour in my own code that nobody had written down: letters in a phone number refused rather than stripped, a seven-digit E.164 number refused, spaces inside a Meta first name removed. All of it is declared now. The Python sample came from Sanity's Issues, the translations from the crawled-pages view, and two wrong digests the tool reported as fine (
+44 (0)20 ...and a US ZIP+4 with no country) from an adversarial review of the repository.
Graded by code
Fourteen cases, graded without a model. Expected digests come from the tested normaliser and, in three cases, from digests Meta prints on its own page. The digest rule is about provenance: any 64-character hex string in an answer must be expected or must appear in a tool result from the same run.
| Case group | Full agent | Knowledge base only |
|---|---|---|
| Values whose digest Meta prints on its page (3) | 3 / 3 | 3 / 3 |
| Values whose digest has to be computed (3) | 3 / 3 | 0 / 3 |
| Refusals (3) | 3 / 3 | 3 / 3 |
| Rule questions (5) | 4 / 5 | 4 / 5 |
| Total | 13 / 14 | 10 / 14 |
| Hex digests the model wrote itself | 0 | 0 |
The knowledge base alone gets the rules right and even the digests a platform prints, which the baseline copied straight out of the entry. Every other digest needs the tool: without it the model did not invent one, it declined, which is honest and useless for an upload.
Both modes failed the same case. Asked what to send for the Brazilian CEP 24020-005, the full agent said the tool sends it unchanged and that the documentation does not settle whether -005 counts as an extension, and it would not recommend a value. That is a fair answer and still a fail, because the case asks for a recommendation and I fixed the criteria before the run. An earlier build of the knowledge base, the one that invented "5-digit equivalents", had the agent recommend cutting the CEP to 24020; I would take the refusal over that every time.
The grader went through three versions, and each result file records its version. Version 2 replaced "no unexpected digest" with provenance and forbade any digest in a refusal; version 3 tightened the rule-case patterns against counterexamples an adversarial review wrote to break them, and those counterexamples are now tests. Rule cases are still keyword checks, so every graded answer is published next to the results.
Sanity Project Details
- Project ID: h3njchfh, dataset
production(normalizationRule,implementationProfile) - Studio: conversion-preflight.sanity.studio (the dataset is private, so the documents show to project members)
- Knowledge base: "Offline conversion imports: Google and Meta" (
kb6kEFy3JtbM)
Agent Session
A run of the agent itself, not of the build: the audit, from the request to the answer. It reads the knowledge base outline, pulls every rule from the dataset with two GROQ queries, reads both normalization entries, and sorts what it finds. On the final version it found no undeclared disagreement between the code and the documentation, only entries whose summary differs from the quoted page.


Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.