DEV Community

Cover image for One email, two correct hashes: an agent that checks conversion uploads against the docs and the code
Vinicius Pereira
Vinicius Pereira

Posted on

One email, two correct hashes: an agent that checks conversion uploads against the docs and the code

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

When a CRM sends closed deals back to Google and Meta as offline conversions, every email, phone and name goes out as a SHA-256 hash. The hash either matches the one the platform computed or it does not, and a mismatch is silent: the upload succeeds, the conversion lands, and it is never attributed. The campaign just looks worse than it is.

The two platforms do not normalize the same way, so one shared function is wrong for at least one of them:

Input Google (Data Manager API) Meta (Conversions API)
Jane.Doe+promo@GMAIL.com janedoe@gmail.com jane.doe+promo@gmail.com
(650) 555-1212, US +16505551212 16505551212
O'Brien o'brien obrien
ZIP 94035-1234, US 94035-1234, in clear 94035, hashed

Each column is right for its own destination.

conversion-preflight is an agent a developer asks before an upload. It answers rules questions with citations, checks a CRM export row by row for both platforms, and audits the implementation against the current documentation. It is built on one rule: three sources, three jobs, never blurred.

  • A Sanity knowledge base says what each platform documents.
  • A rule table says what the code does, one entry per platform and field, published to a Sanity dataset as structured documents.
  • A small MCP server of tested code is the only place a normalized value or a digest is produced.

When those disagree, the agent says so and quotes both sides instead of picking one.

Demo

Demo: one email with two correct hashes, the three sources, a CRM export checked, the audit, and the evaluation

A real run over eight synthetic CRM rows, trimmed to three of its findings:

$ python -m agent.run --csv examples/leads.csv --region US "Check these rows before tonight's upload to Google and Meta."
  -> preflight_file examples/leads.csv (6 columns mapped, id=Lead ID, region=US)
  -> initial_context
  -> groq_query *[_type == "implementationProfile"][0]{globalDivergences}
  -> knowledge_base_read ['normalization_and_hashing/google', 'normalization_and_hashing/meta']
  -> groq_query *[_type == "normalizationRule" && field in ["phone","country","first_name",...

## L-1006
- Google and Meta, phone `+1 (415) 555 0142 ext 7`: needs_review because it contains a letter (U+0065).
  - Google says every character after `+` must be a digit [normalization_and_hashing/google].
  - Meta says to remove symbols and letters [normalization_and_hashing/meta].
  - Refusing letters instead of stripping them is a declared divergence for both platforms
    [rule: google/phone] [rule: meta/phone].
  - Fix at source: store the extension in a separate field and send only the number.
- Google, first_name `Mrs. Carol`: warning. Google says to omit prefixes such as `Mrs.` from
  `given_name` [normalization_and_hashing/google]. The code doesn't remove them, which is declared
  [rule: google/first_name]. Fix at source: split the title into its own field.
## L-1008
- Google and Meta, email: invalid because it contains U+200B. ...
Enter fullscreen mode Exit fullscreen mode

Every citation says which source it came from: [normalization_and_hashing/...] is the knowledge base, what the platform documents; [rule: ...] is the dataset, what the code declares.

Digests never pass through the model's text. The agent sees the header and three rows and sends back a column mapping; preflight_file reads every row, writes the upload file keyed by lead id, and returns the totals and the flagged cells. The first version had the model copy every cell into the tool call, which an adversarial review measured at 561,701 characters of tool output for 100 rows. The published run returned 3,216 for the eight. The same review caught a worse thing in an earlier demo: for a phone the tool had refused, the model wrote the "corrected" number itself. The prompt now says to describe the fix at the source and never to compose the value.

Code

GitHub logo vinimabreu / conversion-preflight

Preflight agent for Google and Meta offline conversion uploads: rules from a Sanity knowledge base, declared rules from a Sanity dataset, digests from tested code.

conversion-preflight

An agent that checks customer data before an offline conversion upload to Google and Meta. What the platforms document comes from a Sanity knowledge base built from their own pages. What the code declares about its own rules comes from structured documents in a Sanity dataset, queried with GROQ. Every normalized value and SHA-256 digest comes from tested code. Where those disagree, the agent says so instead of picking one.

One email with two correct hashes, the three sources, a CRM export checked, the audit, and the evaluation

A real run over eight synthetic CRM rows, trimmed to three of its findings (the full transcript is in examples/runs/):

$ python -m agent.run --csv examples/leads.csv --region US "Check these rows before tonight's upload to Google and Meta."
  -> preflight_file examples/leads.csv (6 columns mapped, id=Lead ID, region=US)
  -> initial_context
  -> groq_query *[_type == "implementationProfile"][0]{globalDivergences}
  -> knowledge_base_read ['normalization_and_hashing/google', 'normalization_and_hashing/meta']
  -> groq_query *[_type == "normalizationRule" && field in ["phone","country","first_name","last_name","em
[6 turns, 14.3s, in 6+31541 cached, out 1895 tokens; transcript runs/...;
…

The agent is Claude Code run headless (claude -p) with every built-in tool switched off and only three MCP servers allowed: the knowledge base endpoint, the dataset endpoint, and the local preflight server. The whole prompt is one file, agent/system.md. It runs on my existing plan, so nothing here is billed per call.

202 tests. The suite checks every worked example with an input that the platforms publish (including Meta's own digests) against the code, pins every declared divergence to the code's actual behaviour, and exercises the MCP server over real stdio.

How I Used Sanity

The knowledge base, and what building it taught me

I pointed a Sanity Context knowledge base at three sources: Google's Data Manager "Format user data" guide, Google Ads' "Manage offline conversions" page, and Meta's customer information parameters. Meta's page renders client side and the crawl came back empty, so I extracted its text and uploaded it as a file source, with the URL and retrieval date at the top.

The purpose text does most of the steering. Mine says who asks (engineers about to upload), what leads (which endpoint, how each platform wants each field, where the two differ), what stays out (bidding, pricing, reporting), and, after the first builds, that each platform's rules apply to that platform only and that a rule's qualifiers and the pages' worked examples are kept as written.

Four things happened on the way that shaped the agent:

  1. The first crawl indexed one page in fifteen languages. The "15 documents" from the Data Manager site were the formatting page plus fourteen ?hl= translations, and a later check pulled nineteen translations of the Google Ads page. The first entry I read cited a Hebrew page. "Ignore query parameters" did not stop it; I removed the translations in the crawled-pages view.
  2. Entries are summaries, and summaries drop words. Google says to trim "leading, trailing and intermediate whitespace" in an email address. One build kept "leading/trailing" and lost "intermediate". Another wrote "5-digit equivalents only" for non-US postal codes, which no page says, and the agent then told me to cut the Brazilian CEP 24020-005 to 24020. A single-entry rebuild once wrote "Base64 encoding is not used", true of the Google Ads samples and false for the Data Manager API.
  3. Rebuilds are not deterministic, and instructions are what stick. A detail fixed in one build could vanish in the next. An instruction is honored on every build; when I saved the one for "intermediate", Sanity checked it against the entries itself and listed the two pages that contradicted it before rebuilding them.
  4. Issues found a bug in Google's own sample code. A build opened a conflict between two entries about the Gmail plus-suffix. One side came from Google's Python sample, so I read it: normalize_and_hash_email_address removes the . and + characters from the local part for every domain and keeps the text after the +. The page's own rule, and its Java sample, apply that only to gmail.com and googlemail.com and remove everything from the + on. Code copied from the Python sample turns first.last@company.com into firstlast@company.com. I resolved the issue with the documented rule; two other issues were differences between the platforms, not contradictions, and I dismissed them.

The dataset: the code's rules as structured content

The implementation's rule table lives in the production dataset as normalizationRule documents, one per platform and field, plus an implementationProfile that references them. Each rule carries what the code does in plain words, the platform's words when the code was written (verbatim, with the URL and retrieval date), the knowledge base entry that should agree today (kbEntry), the platforms' own worked examples, and the differences the code makes on purpose. A sync command publishes it from the Python module the tests pin, and a --check mode fails when the dataset and the code drift apart. The schema is a small Studio, deployed at conversion-preflight.sanity.studio.

A second Context MCP endpoint serves that dataset, filtered to those two types, and the agent reads it with groq_query. This is where the two halves of Sanity Context earned their separate places. The knowledge base is the right tool for prose that gets rewritten into entries; the dataset is the right tool for the thing that must come back exactly as written, because a query returns the quote, not a summary of it. Two things I learned the hard way: the endpoint answers "Only datasets with deployed Studio applications are supported" until a Studio with the schema is deployed, and a document id with a dot in it (normalizationRule.google.email) is a private path in Sanity, so the ids use hyphens.

What the agent does with it

  • Rules questions: initial_context once, then knowledge_base_read on the entries it needs, and every rule it states carries the entry path, like [normalization_and_hashing/meta].
  • Batch checks: it maps CRM columns to fields, calls preflight_file once, and explains each refusal and each warning with the documented rule behind it. Before calling a result a conflict with the docs, it checks the rule's declared divergences.
  • Audits: for every rule it compares the knowledge base entry (the documentation today), the quote the code was written against, and the implementation, and sorts what it finds: code and documentation disagree, entry and quote disagree, declared, or agrees. The audit is what flagged the lost "intermediate", the invented "5-digit equivalents", the "Base64 is not used" regression and Meta's self-contradicting UK example. It also found behaviour in my own code that nobody had written down: letters in a phone number refused rather than stripped, a seven-digit E.164 number refused, spaces inside a Meta first name removed. All of it is declared now. The Python sample came from Sanity's Issues, the translations from the crawled-pages view, and two wrong digests the tool reported as fine (+44 (0)20 ... and a US ZIP+4 with no country) from an adversarial review of the repository.

Graded by code

Fourteen cases, graded without a model. Expected digests come from the tested normaliser and, in three cases, from digests Meta prints on its own page. The digest rule is about provenance: any 64-character hex string in an answer must be expected or must appear in a tool result from the same run.

Case group Full agent Knowledge base only
Values whose digest Meta prints on its page (3) 3 / 3 3 / 3
Values whose digest has to be computed (3) 3 / 3 0 / 3
Refusals (3) 3 / 3 3 / 3
Rule questions (5) 4 / 5 4 / 5
Total 13 / 14 10 / 14
Hex digests the model wrote itself 0 0

The knowledge base alone gets the rules right and even the digests a platform prints, which the baseline copied straight out of the entry. Every other digest needs the tool: without it the model did not invent one, it declined, which is honest and useless for an upload.

Both modes failed the same case. Asked what to send for the Brazilian CEP 24020-005, the full agent said the tool sends it unchanged and that the documentation does not settle whether -005 counts as an extension, and it would not recommend a value. That is a fair answer and still a fail, because the case asks for a recommendation and I fixed the criteria before the run. An earlier build of the knowledge base, the one that invented "5-digit equivalents", had the agent recommend cutting the CEP to 24020; I would take the refusal over that every time.

The grader went through three versions, and each result file records its version. Version 2 replaced "no unexpected digest" with provenance and forbade any digest in a refusal; version 3 tightened the rule-case patterns against counterexamples an adversarial review wrote to break them, and those counterexamples are now tests. Rule cases are still keyword checks, so every graded answer is published next to the results.

Sanity Project Details

  • Project ID: h3njchfh, dataset production (normalizationRule, implementationProfile)
  • Studio: conversion-preflight.sanity.studio (the dataset is private, so the documents show to project members)
  • Knowledge base: "Offline conversion imports: Google and Meta" (kb6kEFy3JtbM)

Agent Session

A run of the agent itself, not of the build: the audit, from the request to the answer. It reads the knowledge base outline, pulls every rule from the dataset with two GROQ queries, reads both normalization entries, and sorts what it finds. On the final version it found no undeclared disagreement between the code and the documentation, only entries whose summary differs from the quoted page.

conversion-preflight: auditing the code against Google and Meta docs with Sanity Context claude-sonnet-5-5
You

Audit the implementation against the documentation, for both platforms.

Agent
Agent
Agent
Agent
Agent

Code and documentation disagree, undeclared

None found. I compared all 14 implemented rules against both knowledge base entries and their docQuote text.

Entry and quote disagree

Each of these needs someone to read the source page. From here I can't tell whether the page changed or the entry lost or added a detail.

  • Meta ct (city): [normalization_and_hashing/meta] says "Lowercase, no punctuation, no special characters, no spaces. Special characters must be UTF-8 encoded." docQuote starts with "Hashing required. Using Roman alphabet a-z characters is recommended." The entry drops the Roman-alphabet recommendation for city. Source: https://developers.facebook.com/docs/marketing-api/conversions-api/parameters/customer-information-parameters [rule: meta/city]
  • Meta fn/ln (names): [normalization_and_hashing/meta] adds "(do not transliterate)". docQuote says only "If using special characters, the text must be encoded in UTF-8 format." Same source page. [rule: meta/first_name] [rule: meta/last_name]
  • Google administrative_area (state): [normalization_and_hashing/google] says "address_line and city are used only for Google Analytics connections". docQuote says the state field is also "Used only for Google Analytics". The entry doesn't list state. Source: https://developers.google.com/data-manager/api/devguides/concepts/formatting [rule: google/state]
  • Google source attribution: the entry cites the Google Ads API pages as sources [1][2] for the rules. The Data Manager formatting page, which docQuote uses, is listed only as source 3, and no rule is tied to it. Check which page each rule comes from.

Declared divergences

  • Global: invisible characters and non-ASCII whitespace (U+00A0, U+200B, U+3000 and the like) are refused in every field.
  • Global: a phone number without a country code is completed from the record's country or default_region, for US, CA, GB, AU, BR and DE only.
  • Google phone: letters go to needs_review instead of being stripped. Fewer than 8 digits after + is refused. [rule: google/phone]
  • Google first_name/last_name: prefixes and suffixes ("Mrs.", "Jr.") are not removed, so the source system has to split them out. [rule: google/first_name] [rule: google/last_name]
  • Meta email: whitespace inside the address is removed, not only trimmed. [rule: meta/email]
  • Meta phone: letters are refused instead of removed. Fewer than 8 digits is refused. "Leading zeros" is applied only to the 00 prefix and to a trunk 0 for GB, AU and DE. [rule: meta/phone]
  • Meta fn/ln: inner spaces and symbols (category S) are removed. [rule: meta/first_name] [rule: meta/last_name]
  • Meta st: a full US state name is not abbreviated. [rule: meta/state]
  • Meta zp: with no region, a value shaped like a ZIP+4 is needs_review. For the UK the code follows the m11ae example, not the prose. [rule: meta/postal_code]
  • Meta country: names are refused, so the "United States" → us example is deliberately not reproduced. [rule: meta/country]

Agrees

Google email, country (region_code) and postal_code. Meta city, apart from the entry-vs-quote note above.

Not implemented

  • Google city: [normalization_and_hashing/google] says lowercase, remove symbol characters, trim whitespace, do not hash. The tool returns not_implemented. [rule: google/city]
  • Google state (administrative_area): [normalization_and_hashing/google] says lowercase, use the two-letter abbreviation or the full name, remove symbol characters, trim whitespace, do not hash. The tool returns not_implemented. [rule: google/state]

Not in the rule table

The dataset has no rule for these fields, and the tool's field list doesn't include them either. The knowledge base documents them:
- Google address_line: lowercase, remove symbol characters, trim whitespace, hash with SHA-256. [normalization_and_hashing/google]
- Meta db: YYYYMMDD format. [normalization_and_hashing/meta]
- Meta ge: f or m. [normalization_and_hashing/meta]
- Meta external_id: hashing is recommended, not required. [normalization_and_hashing/meta]

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.