DEV Community

Cover image for Drum Bun: an agent that knows Romania rewrote its vignette this week
Butnar Daniel
Butnar Daniel

Posted on

Drum Bun: an agent that knows Romania rewrote its vignette this week

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Every December, hundreds of thousands of cars drive home between Germany or Austria and Romania. The questions before the drive are always the same: which vignettes, for which days, bought where, and what else gets you fined on the way? The answers are spread over four toll operators, four languages and a lot of outdated blog posts.

This year they changed under everyone's feet:

  • Romania replaced its rovinietă on 1 October 2026: new seller (TollRo), prices in lei by Euro class, and a car whose class cannot be shown pays the Euro 0 rate. A Senate bill could still postpone it.
  • Hungary added an M1 regional vignette in 2026 that makes a yearly crossing less than half the price of the national one, if you know that the Pest and Komárom-Esztergom county vignettes leave a gap on the M1 that only Fejér covers.
  • Austria sells only digital vignettes for validity from 1 December 2026, raised the on-the-spot penalty to €200, and its 2027 prices are not published yet, so a Christmas trip cannot be priced exactly.

Drum Bun ("have a good road", what Romanians say before a journey) tells you what your car needs for your trip:

  • the cheapest set of products per country for your dates and vehicle, with the price valid on each travel day (and a flag when it is not published yet);
  • the rules that apply on those dates (winter tyres, emission zones at the destination, border checks);
  • the outdated claims you probably read, each linked to the fact that corrects it;
  • an agent you can ask in Romanian, German, Hungarian or English, which shows every Knowledge Base entry and query it used.

Demo

Live: https://drum-bun-agent.vercel.app (no login; the planner works without the model, the chat is rate-limited)

The Brașov to Munich plan: a strip map of the route through Romania, Hungary, Austria and Germany, the warnings before you go, and what to buy in each country with prices for each travel day
The route as a strip map drawn from the data: tolled sections red with a yellow core, free sections hollow, borders dashed.

The agent checks two claims from a blog; the expanded trace shows a Knowledge Base search and the two entries it opened
Every answer shows its steps: planner run, Knowledge Base entries opened, GROQ queries.

Code

GitHub logo danielbutnar / drum-bun

Road-trip agent on Sanity Context: vignettes, tolls and rules between Romania and Germany/Austria, priced for your dates

Drum Bun

What your car needs on the road between Romania and Germany or Austria: vignettes and tolls for your dates (the cheapest combination), winter-tyre and emission-zone rules, and which claims you read online are out of date. Every answer is traced to an official source.

Live: https://drum-bun-agent.vercel.app · Demo video (2:52): https://youtu.be/X2qOuYowQro · Sanity project pd5e7gez, dataset production (public) · Built for the DEV Sanity Challenge, Path One.

"Drum bun!" is what Romanians say before a journey, and what the sign says when you leave a town.

Why this needs structure

"What does Brașov → Munich on 20 December, back 3 January, cost in a Euro 5 diesel?" has no page that answers it. The answer is computed from:

  • the route: which road sections, in which Hungarian counties, which are tolled (the Salzburg Nord → Walserberg stretch is free, the M1 between Bicske and Szárliget…

How I Used Sanity

1. Structured content that computes

"Brașov to Munich, 20 December, back 3 January, Euro 5 diesel" has no page that answers it. So the content is modelled as the pieces the answer is computed from:

  • route → ordered legs → roadSection references. A section has tolled, coveredBy[] (the products that make it legal, any one is enough), counties[] and exemptVehicles[].
  • tollProduct with validity (days, months, calendarYear with overlap windows) and dated prices[]: each price has validFrom/validTo and an optional emission band (euroMin, euroMax, electric, appliesWhenUnknown).
  • rule with a yearly season window, conditional + condition, effectiveFrom/To and severity; zone with dieselMinEuro and status.
  • claim: a statement found online, its verdict (outdated, wrong, misleading), the source it was seen on, and correctedBy[] references to the facts that are true now.
  • source with trust (official, club, press, blog, forum), language and checkedAt. Every fact references at least one.

A pure, unit-tested planner walks that structure: it picks the price valid on each date, runs a small dynamic program for the cheapest cover (two 1-day vignettes beat one 2-month; a 2026 Austrian annual vignette still covers 3 January 2027), solves an exact set cover over Hungarian county vignettes and the M1 regional one, and applies Austria's 18-day rule for online purchases. The agent calls it as a tool. The model never computes a price.

2. A Knowledge Base from official pages in four languages

The Knowledge Base "Drum Bun road rules" is built from the dataset (a GROQ query that flattens each document into readable fields) plus 24 web pages: the toll operators, ministries and cities in Romanian, Hungarian, German and English, and the blogs drivers actually read. The first build filed 14 issues; after all fixes and a full rebuild the Knowledge Base has 22 entries from 103 dataset documents and 24 web pages, 21 decided issues and 9 standing instructions. I resolved them with @sanity/client's context API instead of clicking through the Dashboard, so every decision is in the repo with its reason (decisions.md):

  • Two of the four conflicts came from my own vocabulary: the build read my enum value carTrailer as "a trailer" and concluded trailers must carry warning triangles. One standing instruction fixed the vocabulary for every future build. Of the other two, one was true for different dates (Austria's substitute toll was €120 before 2026, €200 since), so an instruction now makes entries state the date; the other was a Hungarian motorcycle price that is right once you know motorcycles buy the car's annual vignette.
  • It invented Hungarian purchase points (post offices, the automobile club) that no source names. An instruction limits entries to the channels the sources name.
  • It missed one: an entry quoted the Romanian Interior Ministry page's winter-tyre fine, computed with an old penalty-point value. I wrote an instruction scoped to that page; the background contradiction check then filed an issue against the stale entry by itself, and applying it fixed the number. That loop, human instruction → automatic check → issue → rebuild, is the part I would not want to build myself.

A Knowledge Base conflict on the How it knows page: the side that was not kept, the side that was kept, and why
One of the conflicts the build filed, with the decision and the reason.

The public How it knows page shows every issue with both sides, what was kept, the standing instructions and the outline, from a snapshot the script writes into the dataset.

3. Sanity Context MCP, two endpoints

One endpoint serves one mode, so there are two: drum-bun-kb (Knowledge Base mode: outline + knowledge_base_read/knowledge_base_search) and drum-bun-rules (GROQ mode over the dataset, embeddings enabled). The agent (AI SDK 7, AI Gateway) fetches both initial contexts once, inlines them, and gets the planner as a third tool. Conversations are saved with Sanity Context Insights; a daily cron classifies them.

4. Does structure beat search?

Sixteen questions in Romanian, German and English, each with facts an answer must contain and outdated facts it must not (questions). The baseline gets the same Sanity documents and the same model, but retrieves with an any-word keyword search (what a site search box does) instead of the planner, the Knowledge Base and GROQ.

correct
Drum Bun agent (planner + Knowledge Base + GROQ via Sanity Context) 15 / 16
Keyword search over the same content 11 / 16

The one question the agent missed: someone typed the wrong plate on an Austrian 10-day digital vignette that has already started. The agent said the right thing, that it can no longer be changed and that a wrong plate counts as no vignette (€200 on the spot), but it left out that annual vignettes are the exception (re-registration for €18). The test requires that detail, so it counts as a miss. I kept the test as written instead of tuning the prompt until it passes; keyword search missed it too.

Search did fine on single facts: the dataset states them plainly, which is itself a point for structured content. It failed exactly where the answer has to be computed: it found the Munich zone rules but no vignette prices for the Christmas trip, found the rovinietă but could not pick the Euro-class price, named 2 of the 6 Hungarian counties and missed the M1 regional vignette, and found the surcharge rules but not the amount.

Honest notes: the first run scored the agent 6 / 16. It asked "which city?" for general questions, answered an English question in Romanian, returned empty answers when reasoning tokens ate the output budget, and hit the free model tier's 5-requests-a-minute limit. The fixes were a tool for single-product prices, language detection in code, a bigger output budget and a paced runner. After the final run I widened two patterns that rejected correct answers ("1, 10, 30, 60 de zile", a non-breaking hyphen in "60‑minute") and re-scored both systems; every run is in evals/runs.

What did not go smoothly

  • A dataset source query with select() made the ingest fail with a generic server error; plain projections work.
  • A website source crawls everything under the URL's path: the Munich city page pulled in 306 pages of a 150-source budget. Single deep URLs import one page.
  • GROQ mode needs a deployed Studio, not only a deployed schema.
  • The AI Gateway's free tier does not include Claude models and allows 5 requests a minute, so the live agent runs on gpt-5-mini, the example questions replay recorded answers (labelled as such, with an "ask it live" button), and a busy model says so instead of failing silently. The planner makes the model choice matter less.
  • The first eval run lost to keyword search (above). Measuring early is what made the agent good.

Sanity Project Details

Agent Session

The whole build ran as one Claude Code session over two days: the research into four countries' toll rules, the schema, the planner and its tests, the Knowledge Base decisions, the eval run that lost to keyword search, and the video. Steps that touched my unrelated private notes are removed.

Building Drum Bun with Claude Code claude-opus-5-5
You

I found a new challange and i need you to max the probabillity for winning. Here are the details. Please research, plan then execute:

<pasted_content id="49e8">
Sanity Challenge
Signed upView Entries
Build an AI agent on structured content, or vibe-code an app with Sanity behind it.
Challenge Status:Live
We're excited to announce our newest challenge with Sanity!
The Sanity Challenge runs September 18 to October 4. Build an AI agent on structured content, or vibe-code an app with Sanity behind it. $2,500 in prizes.
New to Sanity? Sanity is the AI Content Operating System. Your content lives in the Content Lake as JSON documents, with schemas you define in TypeScript and query with GROQ. You can use Sanity to build anything that you can imagine with structured content as your starting point.
Read on to learn more.
Key Dates

  • Contest start: September 18, 2026
  • Submissions due: October 04, 2026
  • Winners announced: October 22, 2026

Badge Rewards
Sanity Challenge Winner Badge
Sanity Challenge Completion Badge
Find Out More
Ask questions and share your ideas on the Sanity Challenge Launch Post.
View Launch Post
Sponsored by Sanity
Sanity is the AI Content Operating System. Your content lives in the Content Lake as JSON documents, with schemas you define in TypeScript and query with GROQ.
Learn More →
Challenge Prompts
Path One: Ship an Agent That Queries Real Content
Build an agent, then point it at a Sanity Context MCP endpoint backed by a Knowledge Base. Any agent framework, hosted anywhere.
Build anything that needs an answer it can't afford to get wrong. A board game companion that knows the errata contradicts the rulebook. A better interface to your favorite open-source docs. A eurorack planner that knows what actually fits in your case. An award-travel agent that untangles which transfer partner story is current. A car repair agent? A camera gear-head compendium? The sky is the limit.
Point Sanity Context at a website, a set of files, or your own Sanity content, and it distills a navigable Knowledge Base your agent reads through MCP. Every entry stays linked to the source it came from. When two sources contradict each other, both claims surface side by side with their sources, and the decision you make carries across future builds. It all lives in your Sanity Dashboard.
The strongest submissions will show an agent that only works because the content was structured. If a keyword search would have gotten you the same answer, aim higher.
Note: Knowledge Bases are in beta and currently index up to 150 documents. Design for that budget, or point your agent at your full dataset through a Context MCP endpoint with embeddings enabled. Both count for Path One.
[Submission Template](https://dev.to/new?prefill=---%0Atitle%3A%20%0Apublished%3A%20%0Atags%3A%20devchallenge%2C%20sanitychallenge%2C%20sanity%2C%20ai%0A---%0A%0A%2AThis%20is%20a%20submission%20for%20the%20%5BSanity%20Challenge%2C%20Path%20One%3A%20Ship%20an%20Agent%20That%20Queries%20Real%20Content%5D%28https%3A%2F%2Fdev.to%2Fchallenges%2Fsanity-2026-09-16%29%2A%0A%0A%23%23%20What%20I%20Built%0A%3C%21--%20Provide%20an%20overview%20of%20your%20project%20and%20what%20problem%20it%20solves%20or%20experience%20it%20creates.%20--%3E%0A%0A%23%23%20Demo%0A%3C%21--%20Embed%20a%20video%20walkthrough%20or%20share%20a%20link%20to%20your%20deployed%20project.%20--%3E%0A%0A%23%23%20Code%0A%3C%21--%20Embed%20or%20share%20a%20link%20to%20your%20repository.%20--%3E%0A%0A%23%23%20How%20I%20Used%20Sanity%0A%3C%21--%20Explain%20how%20your%20agent%20reads%20and%20queries%20content%20stored%20in%20Sanity.%20What%20did%20you%20point%20Sanity%20Context%20at%20to%20build%20your%20Knowledge%20Base%2C%20which%20Sanity%20Context%20tools%20did%20you%20use%2C%20and%20what%20did%20the%20agent%20actually%20do%20with%20the%20content%20it%20retrieved%3F%20--%3E%0A%0A%23%23%20Sanity%20Project%20Details%0A%3C%21--%20Required%3A%20share%20your%20Sanity%20project%20ID%20or%20a%20link%20to%20a%20public%20dataset%20URL%20so%20the%20Sanity%20team%20can%20see%20how%20you%20modeled%20and%20used%20your%20structured%20content.%20--%3E%0A%0A%23%23%20Agent%20Session%0A%3C%21--%20Optional%20but%20encouraged%3A%20upload%20your%20transcript%20at%20https%3A%2F%2Fdev.to%2Fagent_sessions%2Fnew%2C%20curate%20or%20slice%20the%20parts%20worth%20showing%2C%20and%20embed%20it%20right%20here.%20Supported%20tools%20include%20Claude%20Code%2C%20Gemini%20CLI%2C%20Codex%2C%20GitHub%20Copilot%20CLI%2C%20and%20Pi.%20--%3E%0A%3C%21--%20Sessions%20are%20unlisted%20by%20default%2C%20so%20hit%20Make%20Public%20before%20you%20publish%20or%20judges%20won%27t%20be%20able%20to%20open%20it.%20Check%20your%20transcript%20for%20keys%20and%20sensitive%20data%20first.%20--%3E%0A%0A%3C%21--%20Don%27t%20forget
[… 6246 characters trimmed]

Agent
Agent
Agent
Agent
Agent

Three research tracks are still running. When they report back I'll put together a plan: which path or paths to enter, the concept, a build schedule up to the Oct 4 deadline, and which steps only you can do, such as creating a Sanity account.

You

Research report from a helper agent:

Sanity Context and Knowledge Bases: research for the DEV.to Sanity Challenge, Path One (checked 2026-09-29)

Challenge page: https://dev.to/challenges/sanity-2026-09-16
- Submissions close Oct 4, 2026, 11:59 PM PDT. Winners are announced Oct 22.
- You must give the Sanity project ID or a public dataset URL, and use the tag #sanitychallenge. If the demo needs a login, give test credentials.
- The announcement says a Knowledge Base (KB) or an embeddings-enabled full-dataset endpoint "Both count for Path One".
- Sanity's own example ideas: a board-game companion where the errata contradicts the rulebook, a better interface to open-source docs, a eurorack planner, an award-travel agent.
- Support is in the #mcp-server channel on Discord.

1. What Sanity Context is and how a KB is built

Docs: https://www.sanity.io/docs/ai/sanity-context

Sanity Context is a hosted MCP server. It is read-only and does not run the agent loop for you. It has two modes:
- GROQ mode: queries the live dataset.
- Knowledge Base mode: serves an index built ahead of time. The build reads the sources, resolves conflicts, and writes entries the agent reads directly.

Enabling it. An org admin turns on Context on the Labs page (https://www.sanity.io/manage/org/labs). KBs are an "opt-in beta".

Creating a KB in the Dashboard (https://www.sanity.io/docs/ai/sanity-context-create-knowledge-base):
1. Dashboard → Context → New knowledge base.
2. Fill in Title and Purpose. The purpose is one or two sentences about who the KB serves, and it steers the outline.
3. Create knowledge base → Add source (Dataset / Website / Files).
4. Build entries, then wait for Entries up to date.
5. Review Entries (a tree with summaries) and Issues.

Creating a KB from the CLI (https://www.sanity.io/docs/cli-reference/cli-context). I did not find a minimum CLI version.
```sh
sanity context create --organization org-abc123 --title "Support docs" [--description ...]
sanity context imports create kb-abc123 --url https://example.com/docs

other flags: --file <path> --content-type <mime> --text <value> --title

--sanity-project <id> --sanity-dataset <name> --query <GROQ>

sanity context build kb-abc123 --watch # also --cancel
sanity context jobs get kb-abc123 job-def456 --watch
sanity context refresh kb-abc123
sanity context update kb-abc123 --refresh-enabled --refresh-frequency ...
``
Other subcommands:
list,get --json,delete --yes, andimports list/get/delete/download`. The CLI has no command for Issues.

Source types (https://www.sanity.io/docs/ai/sanity-context-source-types). One KB can mix all three kinds and needs at least one source.
- Dataset: a full GROQ query such as *[_type == "article"], optionally with a projection. A bare filter is rejected. Up to 5,000 documents per source, published documents only. Connecting one needs the Administrator or Developer role on the source project plus unrestricted read access.
- Website: "A crawl starting from a URL. Crawls respect robots.txt." The docs recommend the most specific URL you can. Crawl depth, page cap and JS rendering are not documented.
- Files:
- Size limits: PDF 500 MB, DOCX/PPTX 100 MB, XLSX 50 MB, HTML/images/AsciiDoc 25 MB. Markdown, text and code have no stated limit.
- ZIP/TAR archives are expanded. Images under 50 KB skip text extraction.
- Total upload is capped at 5 GiB. Uploaded files "never re-sync".

The 150-document limit appears only in the challenge text. The docs only say that plans cap the number of KBs and the number of sources per KB ("enforced when you build"). How documents are counted (per crawled page, per file, per dataset document) is not found / unclear.

Build time is not documented; the docs only say "Larger builds take longer".

Refreshes and rebuilds (https://www.sanity.io/docs/ai/sanity-context-maintain-knowledge-base):
- Automatic refresh can be Weekly, Monthly or Off.
- A refresh is incremental. It recrawls, compares with the last build, and files issues; only affected entries change.
- A full rebuild reruns the whole pipeline and replaces the outline, so "anything the previous build decided is reconsidered".
- You need a rebuild after changing the purpose, when a "Rebuild required" callout appears, or when the KB has only uploaded files.
- Version history lets you pick an earlier outline and use Restore this version.

Contradictions (https://www.sanity.io/docs/ai/sanity-context-resolve-issues):
- A conflict shows the two claims side by side with the source of each. Sanity's example is a help center saying 30-day returns while a product page says 45.
- You choose Keep the current entry or Accept the incoming claim, then Resolve issue.
- Other issues are structural proposals (add, remove, split or merge entries). You can Apply, Edit manually or Dismiss them.
- Issues that touch a fact an agent is likely to state are marked Critical.

What "the decision carries across future builds" means concretely: the resolution is "Saved as an instruction that shapes your content on the next rebuild".
- The instruction lasts until the source material changes or someone reopens the issue.
- If new source material no longer supports it, it is "archived automatically, with a reason".
- You can also write standing Instructions yourself: a rule in plain language, anchored to chosen sources.
- The docs also advise fixing conflicting content in the dataset itself, because "The next build inherits the correction."

Issues are resolved only in the Dashboard. No MCP tool, CLI command or API for reading Issues is documented. One entry ("Still True") has a list_conflicts tool, but it looks like the author's own wrapper and is unverified.

What "navigable" looks like. The outline has entry paths with one-line summaries and [core]/[peripheral] tags. Example from https://www.sanity.io/docs/ai/sanity-context-knowledge-bases:

products/latex/gloves [core]
Glove grades, sizes, and what each is rated for
topics: Grades, Sizing, Ratings
shipping/import-routes
Customs paperwork and lead times by region
related: support/returns

Each entry is a Markdown document with citations back to its sources. Entries cannot be edited by hand; builds rewrite them.

2. The Context MCP endpoint

Reference pages: https://www.sanity.io/docs/ai/sanity-context-mcp and https://www.sanity.io/docs/ai/sanity-context-configure-mcp

URL: https://api.sanity.io/v1/context/organizations/:organizationId/mcp/:mcpEndpointName

Creating an endpoint. Endpoints are created in the Context app with these fields:
- title (up to 100 characters)
- name (lowercase letters, numbers and hyphens, up to 64 characters, immutable)
- sources (1–100 entries), for example {"type":"knowledge-base","id":"kb…"} or {"type":"dataset","id":"PROJECT.DATASET"}
- instructions (up to 10,000 characters)
- groqFilter (applies in GROQ mode only)

No HTTP API or CLI is documented for creating endpoints.

One mode per endpoint. If an endpoint has both a dataset source and KB sources, "the dataset source wins and the Knowledge Base sources are ignored". The official starter therefore runs separate GROQ and KB endpoints.

URL parameters:
- mode=groq|knowledge_base
- knowledgeBases=kb…,kb…
- tools= (allowlist)
- embeddings=true|false (omit to auto-detect)
- instructions, groqFilter, perspective, workspace

Transport. The docs never name it. The OpenAI Agents SDK guide uses MCPServerStreamableHttp (with a 30 s timeout recommended), and the AI SDK guide uses type: 'http'. So it is streamable HTTP, inferred from those guides.

Auth: Authorization: Bearer <token>.
- Use an organization API token with Context Viewer permissions, created under Manage → API → Tokens (https://www.sanity.io/manage/org/api/tokens).
- The underlying grant is sanity.knowledge-base.read.
- Project tokens get 403 contextGrantRequired.
- Other errors: 401 (bad token), 422 invalidGroqFilter, JSON-RPC -32005 (KB mode with no KBs configured), -32004 (schema not deployed).

Tools (https://www.sanity.io/docs/ai/sanity-context-mcp-tools):

Mode Tool Parameters Returns
KB initial_context none the outline of each KB the endpoint serves
KB knowledge_base_read knowledgeBase (kb id, required), paths (string[], 1–20) full content of those entries
GROQ initial_context none compressed schema plus query instructions
GROQ schema_explorer type, optional path detail for one schema type
GROQ groq_query query {meta:{executedQuery, perspective, resultCount, returnedCount, warnings, hint}, result}
GROQ array_field_reader mode (range/filter/continue/outline), documentId, field, plus mode options cropped array items

Initial context over HTTP. You can fetch it once and put it in the system prompt instead of calling the tool:
sh
curl https://api.sanity.io/v1/context/organizations/ORG/mcp/NAME/initial-context -H "Authorization: Bearer $SANITY_ORGANIZATION_TOKEN"

It returns Markdown; ?heading_offset controls heading levels. After inlining it, drop the initial_context tool.

KB endpoint versus full dataset with embeddings:
- A KB endpoint serves pre-reconciled, cited entries through the outline and knowledge_base_read. Text search does not apply.
- A dataset endpoint runs live GROQ with BM25 keyword search. Keyword search has no fuzzy matching or stemming. Semantic and hybrid ranking uses text::semanticSimilarity(), which works only inside score().
- GROQ mode needs sanity schema deploy from Studio 5.1.0 or later.

Enabling embeddings (https://www.sanity.io/docs/content-lake/dataset-embeddings):
```sh
sanity datasets embeddings enable <name> --wait

or: sanity datasets create <name> --embeddings --embeddings-projection='{ title, summary, category }'

sanity datasets embeddings status <name> # updating | ready | error
``
- Embeddings are available on all plans, and generating them costs nothing.
- Each **query** that uses them counts against the monthly semantic-search quota.
- A document is embedded in at most 10 chunks.
- The endpoint uses semantic search only when status is
ready` and the project has AI usage credits left.

3. Plans and costs

  • Pricing page (https://www.sanity.io/pricing): "Agent Context" and "Semantic search" are listed as included on Free, Growth and Enterprise. Knowledge Bases are not on the pricing page.
  • Free plan: 250k API requests a month, 1M CDN requests, 10k documents, 2 datasets (public only), 500 embeddings queries a month, 1,000 AI credits.
  • Growth plan: $15 per seat; 1k embeddings queries a month, then $1.50 per 1k.
  • Beta gating: Labs opt-in by the org admin, not a waitlist. The docs say "Features and limits may change"; Enterprise customers can ask for higher limits.
  • KB caps per plan: not published. Whether the free plan can hold a KB, and how many, is unclear. Other challenge entrants appear to have built KBs, but their plans are unknown.
  • Rate limits for Context MCP: not documented. A third-party article (robotostudio.com/blog/sanity-knowledge-bases) says tool calls are "priced like ordinary Sanity API calls" with no per-token retrieval fee.
  • For a public demo judges will use: the 500 semantic queries a month on Free is the tightest limit. A KB-mode endpoint avoids it. Other entrants added rate limits and same-origin guards to their public routes.

4. Connecting an agent

(a) Vercel AI SDK (https://www.sanity.io/docs/ai/sanity-context-patterns). Sanity's public-assistant route, condensed:
ts
import {createMCPClient} from '@​ai-sdk/mcp'
const mcp = await createMCPClient({transport:{type:'http', url: MCP_URL,
headers:{Authorization:`Bearer ${process.env.SANITY_ORGANIZATION_TOKEN}`}}})
const tools = await mcp.tools()
const result = streamText({model: anthropic('claude-sonnet-4-5'), system: SYSTEM_PROMPT,
messages: convertToModelMessages(messages), tools,
stopWhen: ({steps}) => steps.length >= 8, onFinish: () => mcp.close()})
return result.toUIMessageStreamResponse()

Version notes:
- The quick start pins npm install @​ai-sdk/mcp@^1, matched to the ai major version.
- @​sanity/context v2.1.0 (Sept 28) adds support for AI SDK v7. On v7, telemetry moves to telemetry:{integrations:[…]} with no isEnabled.
- Current AI SDK docs close the client in onEnd rather than onFinish, so check the callback name for your version.
- The Vercel AI SDK guide (https://www.sanity.io/docs/ai/sanity-context-vercel-ai-sdk) shows the pattern of fetching /initial-context and removing the tool: const {initial_context: _, ...tools} = await mcpClient.tools().
- There are also official guides for the OpenAI Agents SDK and LangChain.

(b) Anthropic Messages API MCP connector (https://platform.claude.com/docs/en/agents-and-tools/mcp-connector):
- Beta header mcp-client-2025-11-20. The docs also show mcp-client-2026-09-15 for pinned tool lists.
- Only tool calls are supported, and the server must be reachable over Streamable HTTP or SSE.
- It is not available on Bedrock or Vertex.
ts
await anthropic.beta.messages.create({ model, max_tokens, messages, betas:["mcp-client-2025-11-20"],
mcp_servers:[{type:"url", url: MCP_URL, name:"sanity", authorization_token: ORG_TOKEN}],
tools:[{type:"mcp_toolset", mcp_server_name:"sanity"}] })

Sanity does not document this route. It should work because the connector sends a Bearer token, but I have not verified it.

Claude Agent SDK (verified in its docs):
ts
options:{ mcpServers:{ sanity:{ type:"http", url: MCP_URL,
headers:{Authorization:`Bearer ${process.env.SANITY_ORGANIZATION_TOKEN}`} } },
allowedTools:["mcp__sanity__*"] }

(c) Claude Code:
sh
claude mcp add --transport http sanity-context <URL> --header "Authorization: Bearer <token>"

Claude.ai and Desktop custom connectors do not accept a Bearer header (anthropics/claude-ai-mcp issue #112). A bridge such as mcp-remote would be needed; I did not verify that.

Official examples:
- https://github.com/sanity-labs/starters/tree/main/knowledge-base (merged Sept 23) is the best reference. It is a Next.js help center with Claude and two chats, /chat (public) and /internal. It uses four endpoints: beacon-catalog (GROQ), beacon-support-kb (a KB built from a dataset query plus 10 PDFs), beacon-ops and beacon-ops-kb. The KBs are created in the UI and conflicts are resolved in the Issues queue.
- https://github.com/sanity-io/context has the @​sanity/context package, examples/ecommerce (a Next.js shopping assistant with a classify-conversations function), and three skills: create-agent-with-sanity-context, dial-your-context and shape-your-agent. Install them with npx skills add sanity-io/context --all.

Token safety: keep the token server-side only. The docs say "never embed it in client code". In KB mode there is no per-item filtering (groqFilter is ignored). Since knowledgeBases= is a URL parameter, keep the endpoint URL server-side too.

Insights (a bonus worth showing). Import sanityInsightsIntegration from @​sanity/context/ai-sdk, or call client.context.conversations.save({threadId, messages, metadata:{mcpEndpoints:[…]}}). A scheduled Sanity Function running classifyConversations() then scores each conversation with successScore (1–10), sentiment and contentGaps. Results show on a dashboard in the Context app. Docs: https://www.sanity.io/docs/ai/sanity-context-insights

5. The general Sanity MCP server

Docs: https://www.sanity.io/docs/ai/mcp-server

  • URL: https://mcp.sanity.io, over HTTP with OAuth. Add it with claude mcp add Sanity -t http https://mcp.sanity.io --scope user.
  • Schema and deploy: get_schema, list_workspace_schemas, deploy_schema, deploy_studio
  • Documents: create_documents, patch_documents (up to 25 per call), query_documents (64 KiB response limit), get_document, publish_documents, unpublish_documents, create_version, discard_drafts
  • Projects and datasets: create_project, create_dataset, update_dataset, list_*, plus CORS tools
  • Assets: upload tools, plus generate_image and transform_image, which use AI credits
  • Releases: create_release, list_releases
  • Docs and search: search_docs, read_docs, semantic_search, run_sanity_cli
  • No tools for managing Context or KBs were found.

6. How Sanity markets it (a hint at what judges value)

Sources: https://www.sanity.io/blog/everything-we-shipped-so-you-can-trust-agents-with-your-content, https://www.sanity.io/everything-2026-recap, https://www.sanity.io/context

  • Their framing: "One set of facts your whole organization can rely on" and "Contradictions surfaced (and hopefully fixed) instead of shipped".
  • KBs are described as an "agentic map": the agent reads the outline, then drills into entries, instead of retrieving top-k chunks.
  • Use cases they highlight:
    • support and ticket deflection
    • product Q&A (their example: "Does this speaker work with eARC TVs?")
    • shopping with filters ("hiking boots under $200")
    • internal research
    • style-guide checks
    • multi-brand operations (MJH Life Sciences, 60 brands)
  • Customer quotes come from Braze and PUMA.
  • The docs' guidance for the system prompt: answer only from retrieved content, and say so when nothing is found.

Hints for the entry:
- Make the KB's citations and conflict resolution visible in the UI.
- Structure content in Sanity rather than dumping it in.
- Consider pairing a KB endpoint with a GROQ endpoint.
- Add Insights.
- Show the 30-versus-45 style contradiction being resolved and the change carrying into the next build.

Not found / unclear

  • How the 150-document limit is counted.
  • Website crawl depth and page cap.
  • Typical build time.
  • Per-plan caps on KBs and sources, and whether the free plan can use KBs.
  • Rate limits on Context MCP.
  • Any API or MCP access to Issues.
  • The minimum CLI version for sanity context.
  • Whether the Anthropic MCP connector works with Context: not tested.
Agent

The Path One (Sanity Context) research is back. Still waiting on the other two tracks: the rules and competition scan, and the App SDK and Workflows research.

You

Research report from a helper agent:

Sanity research for the DEV.to Sanity Challenge, Path Two (checked 2026-09-29)

Versions come from npm view today: sanity 6.16.0, @​sanity/cli 8.13.0, @​sanity/sdk-react / @​sanity/sdk 3.6.0, @​sanity/client 8.9.0, next-sanity 13.3.4, @​sanity/ui 4.2.7, @​sanity/blueprints 0.27.0, @​sanity/functions 1.8.0, @​sanity/runtime-cli 17.15.0, and all @​sanity/workflow-* packages at 0.35.0 (published 2026-09-23).


1. App SDK

Docs: https://www.sanity.io/docs/app-sdk/sdk-introduction, /sdk-quickstart, /sdk-react-hooks, /sdk-authentication, /sdk-deployment (all under https://www.sanity.io/docs/app-sdk/).

  • Scaffold: npx sanity@latest init --template app-quickstart, or --template app-sanity-ui if you want the Sanity UI starter. It needs Node ≥22.12 and React 19 (peer react ^19.2.0).
  • Structure: three files matter:
    • sanity.cli.ts holds defineCliConfig({app: {organizationId, entry: './src/App.tsx'}}). After the first deploy it also holds app.id.
    • src/App.tsx wraps your components in <SanityApp config={[{projectId, dataset}]} fallback={...}>.
    • src/ExampleComponent.tsx is a sample component.
  • Dev: npm run dev serves on port 3333 and opens through the Dashboard at https://sanity.io/@<org-id>?dev=http://localhost:3333. Studio also defaults to 3333, so start the app with npm run dev -- --port 3334.
  • Hooks: I read these from the 3.6.0 type definitions:
    • Documents: useDocuments, usePaginatedDocuments, useDocument, useDocumentProjection, useDocumentPreview, useQuery
    • Editing and actions: useEditDocument, useApplyDocumentActions (with createDocument, publishDocument, unpublishDocument, discardDocument, deleteDocument), useCreateDocument
    • Events and status: useDocumentEvent, useDocumentSyncStatus, useDocumentPermissions
    • Presence: usePresence, usePresenceForDocument, useReportPresence
    • Users and orgs: useCurrentUser, useUsers, useUser, useProjects, useDatasets, useOrganization
    • Releases: useActiveReleases, useApplyReleaseActions
    • Comments: useComments, useCommentThreads, useCommentActions
    • Agent Actions: useAgentGenerate, useAgentTransform, useAgentTranslate, useAgentPrompt, useAgentPatch
    • Other: useClient, useNavigate, useToast tsx const {data, hasMore, loadMore} = useDocuments({documentType:'article', filter:'status == $s', params:{s:'review'}}) const {data} = useDocumentProjection({...handle, projection:`{title,"author":author->name}`}) const setTitle = useEditDocument({...handle, path:'title'}) const apply = useApplyDocumentActions(); apply([createDocument(h,{title:'x'}), publishDocument(h)])
  • Real time: documents are live by default. useDocument combined with useEditDocument gives local-first optimistic edits. Edits go to the draft unless the handle has liveEdit: true. Several actions passed to one apply run as one transaction. useDocumentEvent reports changes made by others as 'remote-patches'.
  • Auth (important for judges): the docs say the flow "presumes a Sanity user is already authenticated within the host Dashboard". Logged-out users are sent to sanity.io/login. The docs describe no anonymous or public mode. An SDK app also works inside a Studio, taking the Studio's auth.
    • Judges without a Sanity login cannot open the app. They would have to be invited to the organization.
    • Whether the free Viewer role can open custom apps: not found / unclear.
  • Deploy: npx sanity deploy puts the app in the organization Dashboard. It needs the org admin or Developer role. The build is capped at 2 GB. npx sanity undeploy removes it.
    • A public hosting URL format is not documented.
    • Hosting on Vercel, or embedding the app in a public Next.js site: not found / unclear, and effectively unsupported because of the auth model.
  • UI kit: @​sanity/ui with ThemeProvider and buildTheme() from @​sanity/ui/theme. The docs recommend styled-components@npm:@​sanity/css-in-js for React 19. Tailwind also has a guide.
  • Free plan: the pricing page has no row for App SDK. Sanity says Dashboard is available to all users. Explicit plan limits for custom apps: not found.

2. Workflows — a real product, in early access

Docs: https://www.sanity.io/docs/workflows plus /getting-started, /studio-plugin, /app-sdk, /effects-and-runtimes, /sanity-functions, /mcp, /limits, /prerelease, /cookbook.

It is a named product, not just a cookbook pattern: "Workflows is in early access, built in public." The packages are 0.x, so a minor version can break the API.

  • Packages:
    • @​sanity/workflow-engine (with /define) and @​sanity/workflow-cli, which provides the sanity-workflows binary
    • -studio-plugin, -sdk (App SDK adapter), -react, -studio, -components, -diagram
    • -mcp, -blueprint, -engine-test (a test bench)
  • Definition: a sanity.workflow.ts config file plus the workflow definition. ts import {defineWorkflowConfig, defineWorkflow, defineStage, defineActivity, defineAction, defineTransition, defineField, defineEffect} from '@​sanity/workflow-engine/define' export default defineWorkflowConfig({deployments:[{name:'dev', tag:'dev', expectedMinReaderModel:10, workflowResource:{type:'dataset', id:'PROJECT.DATASET'}, definitions:[articleReview]}]}) // stages[] → activities[] → actions[{name, status:'done', when?, ops?}], transitions[{to}] defineAction({name:'drafted', when:"$effectStatus['ai-draft'] == 'done'", status:'done'})
    • A transition fires on $allActivitiesDone by default.
    • The document being processed is a subject field.
    • Instances are stored as sanity.workflow.instance documents.
  • CLI:
    • npx sanity-workflows deploy (add --check or --dry-run to validate first)
    • npx sanity-workflows start <wf> --field subject='{...}'
    • npx sanity-workflows fire-action <id> --activity write --action submit
    • npx sanity-workflows show <id>
  • Agent moves the draft:
    • An effect handler calls Agent Actions: content.agent.action.generate({schemaId, documentId, instruction, target}).
    • An agent can also connect through the MCP server (npx -y @​sanity/workflow-mcp with an org-level SANITY_AUTH_TOKEN). Its tools that change state are workflows_start, workflows_fire_action and workflows_deploy_definition.
  • Human approves: in code, engine.fireAction({instanceId, activity:'verify', action:'approve'}). In an App SDK app, useWorkflowSession({engine, instanceId}) gives session.fireAction(activity, action).
    • @​sanity/workflow-sdk also exports useDocumentWorkflows, useWorkflowInstances and useWorkflowEngine.
    • The docs warn not to edit instance documents with useEditDocument.
  • Studio plugin: workflowStudioPlugin({tag}) plus structureTool({defaultDocumentNode: workflowDefaultDocumentNode()}). It adds a form strip, a Workflows view on each document, a Workflows tool, stage badges and a document-action lock.
    • Requires sanity ≥6.15, React ≥19.2.7, styled-components ≥6.4.2 and @​sanity/sdk ≥3.1.
    • Guards are advisory only: "Content Lake enforcement" has not shipped.
  • External API calls: you declare defineEffect({name, bindings, outputs, retry}). The engine only queues effects; your runtime runs the handler, e.g. fetch(...) with ctx.setProgress. Runtimes that can do this:
    • a document Function that calls engine.drainEffects
    • your own server
    • an App SDK client

Time-based conditions need a runtime to call engine.tick().
- Plans and quotas: Workflows has no pricing row. Access conditions: not found / unclear; the packages are public on npm. Hard caps from /limits:
- 100 hops per cascade
- subworkflow depth 6
- 5-minute claim lease
- 24-hour idempotency window

Definition sharing is on by default during early access.
- Free-plan risk: the Functions integration generates a heartbeat that asks for once-per-minute scheduling. Free-plan scheduled Functions run at most daily. Effect draining by document Function should still work; a minute-level tick would need Vercel cron or your own server. This is my inference and not verified.


3. Functions and Blueprints

Docs: https://www.sanity.io/docs/functions/function-quickstart, /functions-introduction, /functions-local-testing, https://www.sanity.io/docs/blueprints/blueprint-config

npx sanity@latest blueprints init . --type ts --stack-name production --project-id <id>
npx sanity@latest functions add --name log-event --type document-create --type document-update
npx sanity blueprints deploy | npx sanity functions logs <name> | npx sanity blueprints destroy
npx sanity functions env add KEY value
defineBlueprint({resources:[defineDocumentFunction({name:'log-event',
  event:{on:['create','update'], filter:'_type == "post"', projection:'{_id,title}'}})]})
export const handler = documentEventHandler(async ({context, event}) => {
  const client = createClient({...context.clientOptions, apiVersion:'2026-02-27'}) })
  • Triggers: on takes create, update or delete. The publish event is deprecated.
  • Other function types: defineScheduledFunction({event:{expression:'0 0 * * *'}}) (organization-scoped, needs a robot token), Media Library, sync-tag and PubSub.
  • Local testing:
    • npx sanity functions dev opens a playground on :8080.
    • npx sanity functions test <name> accepts --document-id, --file, --data, --event, --document-id-before/after and --with-user-token.
    • context.local is true during local runs.
  • Runtime limits:
    • Node 24.x
    • timeout 10 s by default, configurable 1–900 s
    • memory 1 GB by default, up to 10 GB
    • package size 200 MB
    • at most 16 chained invocations
    • rate limit 200 invocations per document per 30 s, 4,000 per project per 30 s
  • Free plan: 500K invocations, 20K GB-seconds, 5 scheduled functions, and a daily minimum cadence (Growth is hourly, Enterprise minutely).

4. Agent Actions

Docs: https://www.sanity.io/docs/agent-actions/introduction, /operations, /generate-quickstart, https://www.sanity.io/docs/platform-management/how-ai-credits-work

  • Requirements: @​sanity/client ≥7.4 (Generate, Transform and Translate work from 7.1) with apiVersion: 'vX' and a token.
  • Schema: every action needs a schemaId, e.g. _.schemas.default. npx sanity@latest schemas deploy uploads the schema and schemas list shows the id; a Studio sanity deploy also uploads it. ts await client.agent.action.generate({schemaId, documentId, instruction:'Summarise $d', instructionParams:{d:{type:'field', path:'body'}}, target:{path:'summary'}}) client.agent.action.transform({schemaId, documentId, targetDocument:{operation:'edit', _id}, instruction}) client.agent.action.translate({schemaId, documentId, targetDocument:{operation:'create'}, fromLanguage:{id:'en-US'}, toLanguage:{id:'de-DE'}}) client.agent.action.patch({schemaId, targetDocument:{operation:'edit', _id}, target:{path:'title', operation:'set', value:'x'}})
  • prompt: exists (it needs client 7.4+ and has the useAgentPrompt hook), but I could not confirm its parameter shape (not found).
  • Options: async: true and noWrite are supported. Actions never write to a published document unless forcePublishedWrite: true. File fields are not supported, and hidden/readOnly fields are skipped.
  • Cost: each Agent Action request costs 1 AI credit ($0.05). The pricing page lists 1,000 AI credits a month on Free; a third-party site says 100, so check in Manage. When the cap is reached, AI features pause until the next month.

5. Studio customisation (with free-plan status)

  • Custom document actions: document.actions: (prev, ctx) => [...prev, MyAction], where an action returns {label, onHandle, dialog} and can use useDocumentOperation(id, type).patch.execute(...). See https://www.sanity.io/docs/studio/document-actions.
  • Badges: document.badges, returning {label, title, color: 'primary'|'success'|'warning'|'danger'}. You can package actions, badges and schema in definePlugin.
  • Custom inputs and Structure Builder: use components: {input} on a field and structureTool({structure}). I did not re-check their docs in this session; these are standard and free.
  • Presentation tool and Visual Editing: Free (the pricing page lists them). Setup with next-sanity:
    • defineLive({client, serverToken, browserToken}) returns {sanityFetch, SanityLive}, imported from next-sanity/live.
    • Render <SanityLive/> and <VisualEditing/> (from next-sanity/visual-editing) in the root layout.
    • Add a defineEnableDraftMode route and configure presentationTool({previewUrl: {previewMode: {enable: '/api/draft-mode/enable'}}}).
    • The pricing page lists "Real-time database and live previews" on Free.
  • Comments, Tasks, Scheduled drafts, private datasets: not on Free.
  • Content Releases and Media Library: the pricing page shows "not included" on both Free and Growth. Releases are an add-on there and part of Enterprise.

6. Free plan (https://www.sanity.io/pricing)

  • Price: $0
  • Seats: 20
  • Roles: 2 (Administrator, Viewer)
  • Datasets: 2, public only
  • Documents: 10k
  • Assets: 100 GB; bandwidth: 100 GB a month
  • API CDN requests: 1M a month; API requests: 250k a month
  • Functions: 500K invocations, 20K GB-seconds, 5 scheduled functions at daily cadence
  • AI credits: 1,000 a month
  • Draft review window: 3 days
  • Overages: no pay-as-you-go on Free, only an upgrade. Growth is $15 per seat per month with a trial.

7. Windows gotchas and the fastest path

Gotchas:
- Cannot find module @​rollup/rollup-win32-x64-msvc after a fresh npm create sanity. This is npm's optional-dependency bug; delete node_modules and package-lock.json, then reinstall (https://github.com/sanity-io/sanity/issues/12725).
- sanity init --output-path inside an existing Next.js project crashed with ENOENT on Windows 11 with Node 24.18 (https://github.com/sanity-io/cli/issues/1647, now closed). Scaffold the Studio in a sibling folder to be safe.
- An old Windows-only circular-import bug in sanity 4.20 is not relevant to 6.x.
- Node 24 fits: Functions run on Node 24, sanity needs ≥22.12, and Workflows needs ≥20.
- The docs' multi-line \ commands and JSON --field subject='{...}' arguments will quote badly in PowerShell; run them in Git Bash. This is my inference.
- @​sanity/sdk 3.1.0 needed an @​sanity/mutate@0.18.2 override. Whether 3.6.0 still needs it: unclear.
- SDK apps have known dev connection issues in Safari.

Fastest path (my own suggested sequence, pieced together from the docs; not tested):
1. Scaffold the site and Studio: npm create sanity@latest -- --template sanity-io/sanity-template-nextjs-clean. This gives Next 16, a Studio, Presentation, the Live Content API and Visual Editing.
2. In the Studio folder run npx sanity@latest deploy. That gives a hosted Studio at *.sanity.studio and uploads the schema, which Agent Actions need for their schemaId.
3. Deploy the Next.js site to Vercel with SANITY_API_READ_TOKEN. This is the public URL judges can open.
4. npx sanity@latest init --template app-sanity-ui, then npm run dev -- --port 3334, then npx sanity deploy. The app appears in the Dashboard, members only.
5. npx sanity blueprints init → functions add → blueprints deploy for the document-event Function that calls Agent Actions or an external API.
6. Optionally add Workflows: install the packages, write sanity.workflow.ts, run npx sanity-workflows deploy, add the Studio plugin, and use useWorkflowSession in the SDK app.

Implication for judging: judges cannot open the App SDK app or the Studio without being invited to your Sanity org. For those parts, plan on screenshots or video in the writeup, or invite the judges.

Nothing was written to the project. Package tarballs were unpacked for type inspection in the scratchpad at ~\AppData\Local\Temp\claude\C--Users-danie-code-danielbutnar\1d3dd969-55b7-437f-822b-82660f6f0d7c\scratchpad.

Agent

The App SDK and Workflows research is back too. Only the competition scan is still running.

You

Research report from a helper agent:

Sanity Challenge research (checked 2026-09-29)

Romania is eligible. The main risk is the rule that development must start inside the entry period. Path One is crowded with "agent with receipts / contradiction" entries, so the idea itself won't stand out. What does stand out is using Sanity's features deeply and giving judges a way to test the app.

1. Rules

Sources: challenge page https://dev.to/challenges/sanity-2026-09-16, launch post https://dev.to/devteam/join-the-sanity-challenge-2500-in-prizes-for-five-winners-514m, contest rules https://dev.to/page/sanity-challenge-v26-09-16-contest-rules, general rules https://dev.to/page/official-hackathon-rules

  • Prizes: 3 winners in Path One and 2 in Path Two. Each gets $500, DEV++ and a winner badge. Every valid entry gets a completion badge.
  • Deadline: "October 4, 2026 at 11:59 PM PDT". That is 5 Oct at 06:59 UTC, or 09:59 in Brașov. Winners are announced Oct 22. DEV has delayed several recent announcements (Notion, Algolia, Gemma, Copilot CLI).
  • Eligibility: 18+. The excluded countries are Afghanistan, Belarus, CAR, Cuba, Equatorial Guinea, Iran, Iraq, Kosovo, Libya, Myanmar, North Korea, Russia, South Sudan, Sudan, Syria, Tanzania, Venezuela and Yemen. Romania and the EU are not excluded. Sponsor employees and their families are excluded. Entries that "violate an Entrant's employer's policies" are ineligible.
  • Pre-existing work: "development of your Entry was started during, and not prior to, the Entry Period" (general rules). The FAQ allows building on open source if the changes are "significant enough", with credit.
  • Teams and entries: teams of up to 4, and one post per team. "Only one submission per path is allowed". You may enter both paths with separate posts.
  • AI: "Use of AI is allowed as long as all other rules are followed."
  • Required in the post: the Sanity project ID or a public dataset URL (without it the entry "may be considered incomplete"), the tag #sanitychallenge, and test credentials if login is needed. Posts must be in English to win prizes.
  • Judging.
    • Path One: meaningful use of Sanity Context and structured content, technical implementation and code quality, use of Knowledge Bases, usability.
    • Path Two: quality and honesty of the build writeup, functionality, schema thoughtfulness, creativity.
    • Tiebreak: most positive reactions.
  • Staff guidance in the posts:
    • Path One: "If a keyword search would have gotten you the same answer, aim higher."
    • Knowledge Bases are beta and "index up to 150 documents". Pointing the agent at the full dataset through a Context MCP endpoint with embeddings enabled also counts.
    • Path Two: "A rough app with an honest writeup beats a polished one with three sentences."
  • Staff comments: DEV staff (heyitsjem) only replied with emoji. There were no clarifications on free plan, credits, deployment or Claude Code. Not found.
  • Practical limits other entrants hit:
    • One commenter (mpmike) hit an organisation-level cap: "3536 of 150 used… Upgrade the plan".
    • One Context MCP endpoint serves only one mode (GROQ or Knowledge Base), so several entrants run two endpoints (ClauseWatch, Detection Debt, Will It Focus).
    • Embeddings must be switched on per dataset with sanity datasets embeddings enable (Lux Stay).
    • App SDK apps require a Sanity login, so judges can't open them. VARdict moved its public side to Next.js, and Contradiction Triage created an Editor-role judge account.
    • Workflows is early access (0.35), its Studio plugin needs Studio 6, and there were peer-dependency conflicts.
    • Scheduled Functions on the free plan run at most daily.
    • Sanity's docs say an org admin enables Knowledge Bases under Labs settings (https://www.sanity.io/docs/ai/sanity-context-knowledge-bases). Trials or credits: not found.

2. Competition

The tag lists 140 posts. After removing the launch post, spam, empty templates and duplicates, there are about 128 posts from 104 authors: 71 in Path One and 57 in Path Two.
- Path One: 45 have a live URL, 19 a video, 61 mention a Knowledge Base, 9 embed an agent session.
- Path Two: 38 live, 13 video, 12 sessions. Real App SDK apps: at least 7.
- The full per-entry table (all columns you asked for, plus a heuristic quality score) is in ~\AppData\Local\Temp\claude\C--Users-danie-code-danielbutnar\1d3dd969-55b7-437f-822b-82660f6f0d7c\scratchpad\sanity\sanitychallenge_entries.tsv.

Strongest entries (reactions/comments, then my assessment):

Entry Path R/C Concept Assets Quality
FinePrint 1 27/4 Checks a hackathon entry against a rules Knowledge Base (website sources + dataset) with a typed checker and a visible trace live without login, 65 s video, KB, session, test suite high
INKSHIFT (same author) 2 22/1 Photo of a crossed-out paper plan moves a live schedule, bookings kept, version "time machine" live, 2 videos, App SDK + Workflows, session, "what is still unverified" section high
Counts My Receipts 1 34/9 Separates recorded claims from findable receipts over 120 docs live, KB + GROQ high
TSB Oracle 1 12/22 Car-repair bulletins; checks every quote; contradictions go to a Studio-approved decision live, KB high
Poisoned pages 1 5/1 Security test: Knowledge Base vs raw mode against 7 planted documents live, 2 min video, session high
Faux Pas Atlas 2 5/2 Etiquette claims with no verdict field Workflows, App SDK Dispute Board, 3 Functions, Agent Actions, session high
Bug Graveyard, VARdict, Cryptid Field Office, Contradiction Triage 2 ~0 App SDK + Workflows (+ Functions / Agent Actions) videos, sessions high
Oniria 2 21/4 Walkable 3D library of DEV articles no live link, video "TBA" medium-high
Detection Debt, ClauseWatch, Will It Focus, ZéroJour, Chain of Custody 1 low Knowledge Base compared against a graph / keyword baseline evals high

Saturated ideas:
- Agents that "refuse to answer without receipts / citations" (30+).
- Errata and version drift: Magic ×2, Next.js ×3, pygame, Blender, Raspberry Pi HATs, camera, MCP spec.
- Surfacing contradictions. The launch post itself suggested it, so it is expected, not a differentiator.
- Incident response ×3, cryptids ×3–4, hackathon-rules agents ×2, "Museum of Almost" ×2, legal / tax / mortgage 5+, hotel / travel ×2.
- A measured baseline ("keyword search fails") is now standard among the top entries.

What the top entries share:
- The structured field decides and the LLM only explains. Quotes are checked in code.
- Two Context endpoints and an honest account of what the Knowledge Base did and did not catch.
- A deployed app with no login, a short video, an agent session, and a limitations section.

Gaps:
- No Path One entry uses the App SDK, and only about 5 use Workflows, for a human-review loop around the agent.
- No Romanian content at all, and German appears only twice. Nobody does a multilingual Knowledge Base where the contradictions are between language versions.
- Few entries use Knowledge Bases built from website or PDF sources on a real public-interest corpus.
- No entry uses Media Library or Canvas. Presentation / visual editing appears in 3 entries, Content Releases in about 7.
- A public App SDK demo is missing because of the login wall. A judge account plus a video fills it.

3. What wins on DEV

Posts checked:
- Notion MCP: https://dev.to/devteam/congrats-to-the-notion-mcp-challenge-winners-28ab
- Algolia Agent Studio: https://dev.to/devteam/congrats-to-the-algolia-agent-studio-challenge-winners-3ocn
- Hermes Agent: https://dev.to/devteam/congrats-to-the-hermes-agent-challenge-winners-3on0
- Agentic Postgres: https://dev.to/devteam/congrats-to-the-winners-of-the-agentic-postgres-challenge-with-tiger-data-51m6

Winning posts opened: NoteRunway, Verdict, Refusal Engine, Relay and the 25-agent Polish Parliament.

  • Structure: the template headings (What I Built / Demo / Code / How I Used X), then one deep section on why the sponsor tech is needed.
    • NoteRunway: "direct API access for speed on bulk reads, and Notion MCP as a safety layer for all destructive writes".
    • Hermes judges praised use "not as a wrapper around a chat model".
  • Length: 500–1,700 words (2–7 minute read). No winner was an essay. Most Sanity entrants write 2,000–3,500 words.
  • Cover image: 4 of 5 had one.
  • Live URL: 4 of 5.
  • Video: 4 of 5. Lengths were 1:49, about 2 min, 5:30 and 25 min, so length did not decide.
  • Diagrams: rare.
  • Honesty sections: few. They matter for Sanity Path Two because honesty is a judging criterion.
  • Reactions: winners had 5–47. Verdict won with 6, so reactions are only a tiebreaker.
  • Precedent: Algolia winner "Verdict — When Policies Collide" is a contradiction resolver. Real-world, high-stakes problems won repeatedly.
  • Earlier Sanity challenge: not found (checked devteam posts back to Aug 2025 and a web search).

4. Agent Sessions

Sources: https://dev.to/devteam/share-embed-and-curate-agent-sessions-on-dev-beta-5bj6 and Forem source (app/views/agent_sessions/new.html.erb, app/javascript/agentSessionParsers/*, app/liquid_tags/agent_session_tag.rb, app/models/agent_session.rb).

  • Format: .jsonl or .json. For Claude Code, the uploader names ~/.claude/projects/<project>/*.jsonl and detects the tool automatically. Only user and assistant records are kept, with tool results attached to their tool calls.
  • Size: parsing and curation now run in the browser, so the upload limit was removed. The saved curated data is capped at MAX_CURATED_DATA_SIZE = 10.megabytes. Parsing stops after 50,000 records.
  • Curation: you click messages to keep them (shift-click selects a range) and can create several named slices. A saved session is unlisted until you press Make Public. Unpublished sessions can only be embedded by their owner.
  • Embed syntax: {% agent_session SLUG_OR_ID %}, a message range {% agent_session slug 358..398 %}, or a named slice {% agent_session slug slice_name %}.
  • Redaction: about 81 regex patterns covering AWS, GitHub, Stripe, OpenAI, Anthropic, database URLs, private keys, JWT, Bearer tokens, generic secret=/token= values, home directories, emails and IPs. The page warns that "we cannot guarantee all sensitive information has been caught". There is no Sanity-specific token pattern, so check Sanity tokens by hand.
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

The Sanity CLI works. Now your first step: I'm starting the Sanity sign-in, which opens a browser window.

Agent
Agent
Agent

Sign-in worked. Next I'm creating the Sanity project from the CLI.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

Sanity is set up: organization "Daniel Butnar" (ob2cyckj9) and project "Drum Bun" (pd5e7gez) with a public dataset. Next I'm checking whether Knowledge Bases still need to be switched on.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

The Studio (Sanity's content editor) and its content model are installed. Now I'm typechecking the schema and deploying it to Sanity.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

Writing tests for the planner: cheapest vignette choice, year-end price changes, Austria's 18-day online rule, and emission zones.

Agent
Agent
Agent
Agent
Agent

All 15 planner tests pass. Now I'm writing the project README with the plan and schedule, then making the first commit.

Agent
Agent

The first commit is in. Next is the look of the site: your rules say you pick from 2–3 design directions before I build the UI, so I'm preparing those.

Agent
You

Research report from a helper agent:

I wrote both research files; they end with "Unverified" lists of gaps.

Summary (about 150 words). Romania changes its road charging on 1 Oct 2026, two days from now. The legal basis becomes Legea 226/2023, and car rovinieta prices come from Ordinul MTI 888/2026, published 25 Sep 2026. Prices move from euro amounts converted each month to fixed lei amounts by Euro class. A 12-month rovinieta costs 228 lei (electric), 254 (Euro VI), 292 (Euro IV–V) or 330 (Euro 0–III or unknown class). Until 30 Sep 23:59 it is sold on erovinieta.ro at 50 EUR = 262.92 lei; from 1 Oct only on the CNAIR TollRo portal (portal.etoll.ro). Durations are 1, 10, 30 and 60 days and 12 months; there are no 7- or 90-day rovinietas any more. The fine becomes 2× to 4× the 12-month price. A Senate bill registered 21 Sep would postpone everything to 1 Apr 2027, so the agent must re-check. Germany: no car toll in 2026 (the minister ruled one out in June 2026). All 35 active Umweltzonen require the green sticker; driving without it costs 100 € and no points. Munich bans only Euro 4 diesels, Stuttgart still bans Euro 5 in its inner zone, and since 1 Oct 2024 only tyres with the alpine (3PMSF) symbol count as winter tyres.

Most surprising facts:
- Unknown Euro class: a car whose Euro class cannot be shown is charged the Euro 0 rate, 330 lei a year.
- Danube bridge tolls: the minister announced that the Fetești–Cernavodă (19 lei) and Vadu Oii tolls end on 1 Oct 2026. I found only news reports of this, no CNAIR or ministry text.
- Rovinieta scope: from 1 Oct the law excludes national-road sections inside municipalities.
- M+S tyres: they are enough in Romania but not in Germany, so the same car can be legal in one country and fined in the other.
- Official pages with wrong fines: berlin.de and the federal environment ministry still say the Umweltzone fine is 80 €; the law says 100 €.
- Romanian winter-tyre fine: it is now 1,946–4,325 lei, and the police keep the registration certificate.
- Border checks: Germany and Austria both run border checks until 15 Mar 2027, so drivers need ID or a passport.

Outdated-claims pages: 6 for Romania and 6 for Germany (three of the German ones are official sites).

Things I could not confirm (all listed under "Unverified"): whether motorcycles, trailers and campers need a rovinieta; the exact fine range from 1 Oct; the TÜV SÜD sticker price; and the Munich council's 2023 decision to suspend the Euro 5 ban.

Files are in ~\code\danielbutnar\drum-bun\research:
- romania.md
- germany.md

Agent
Agent
Agent
Agent
Agent
Agent
Agent

The schema now supports Euro-class prices and pending law changes. Next, the planner has to pick the right price band for your car.

Agent
Agent
Agent
Agent
You

Research report from a helper agent:

I wrote ~\code\danielbutnar\drum-bun\research\austria.md, in the same layout as germany.md: Facts (every bullet ends with its source), a Prices table, Outdated claims, and Unverified. Most facts come from the ASFINAG toll regulations valid from 1 Jan 2026, which I read in both German and English. The rest come from the 2026 and 2025 price ordinances on RIS, BStMG and KFG texts on RIS, asfinag.at, oesterreich.gv.at and bmimi.gv.at.

Summary: Car prices for 2026 are €106.80 (annual), €32.00 (2 months), €12.80 (10 days) and €9.60 (1 day). In 2025 they were €103.80, €31.10, €12.40 and €9.30. Motorcycles pay €42.70, €12.80, €5.10 and €3.80 in 2026. New prices have always started on 1 December and have been announced between late September and mid-November. No official 2027 price existed on 29 Sep 2026, so I did not guess one.

The 18-day wait applies only to annual and 2-month vignettes bought online (shop or app). 1-day and 10-day vignettes are valid immediately, and so is anything bought at a petrol station, toll station or vending machine. Trailers need no vignette.

On the Nickelsdorf → Vienna → Salzburg → Walserberg, Vienna → Graz and Vienna → Suben routes, the vignette is the only toll. Salzburg Nord → Walserberg needs no vignette. I found 7 low-trust outdated pages (RO, DE, EN) and 2 official pages with stale or misleading wording.

Most surprising facts:
- Substitute toll went up: since 1 Jan 2026, driving without a valid vignette costs €200 on the spot for a car (was €120) and €100 for a motorcycle (was €65). If unpaid, the fine is €300–€3,000.
- Stickers end on 1 Dec 2026: under the toll law, any vignette starting from 1 Dec 2026 is digital only.
- Wrong plate counts as no vignette: a typo can be fixed free before validity starts. After that, only annual vignettes can be changed, for €18.
- Campers: what counts is the technical maximum mass (field F.1 on the registration). A camper downrated on paper to 3.5 t keeps the vignette until 31 Jan 2029 only if it was downrated before 1 Dec 2023.
- Two official pages disagree with the law:
- The English toll regulations can be read as exempting motorcycles from tunnel tolls. The binding German text only exempts sidecars.
- The transport ministry's winter-tyre page says fines go up to €5,000 when others are endangered. The law (§ 134 KFG) allows up to €10,000.
- Styria removed its IG-L 100 km/h limits on 18 Apr 2025. Since 26 May 2026 the A2 at Graz has an ordinary 100 km/h limit that also applies to electric cars.
- Motorcyclists must carry a first-aid kit.

Still open, listed under Unverified: 2027 prices, the exact on-the-spot winter-tyre fine (€35 per the ministry vs about €100 per ÖAMTC), trailer speed limits, and whether foreign electric cars are exempt from IG-L limits.

Agent
Agent

The Knowledge Base exists (kbv1SRpT2A3t), so that Labs step is either done or wasn't needed. Reading the Austria research now to build the seed data.

Agent
Agent
Agent
Agent

Next is the seed script: every price, rule, zone, source and outdated claim from the research as structured Sanity documents. Hungary's data slots in when its research finishes.

Agent
Agent
You

Research report from a helper agent:

I finished the Hungary research and wrote it to ~\code\danielbutnar\drum-bun\research\hungary.md. It has the sections you asked for: Facts, Prices (2025 and 2026), Corridors (with annual cost of county-vignette combinations), Outdated claims (10 pages) and Unverified (12 items).

Summary (about 150 words). Hungary's e-vignette has four categories. D1 covers cars, cars with a trailer or caravan (no extra vignette needed), and campers up to 3.5 t if the registration shows M1 and at most 7 seats. D1M is motorcycles, U is only the trailer of a D2 vehicle, and D2 is everything else. The 2026 D1 prices are daily 5,550, 10-day 6,900, monthly 11,170, annual 61,760 and county 7,190 HUF. The fine is 27,790 HUF if paid within 60 days, and 95,730 HUF after that. You must buy before entering, but drivers who enter by mistake have 60 minutes to buy. A wrong plate can be corrected within 60 days of the fine notice, and only if at most 3 characters are wrong; the fee is 1,470 HUF. There is no winter-tyre obligation in Hungary. Headlights are required outside towns, and the alcohol limit is zero. The 2027 prices are not official yet: media calculate about 62,560 for the annual, using the 1.3% August inflation. Collection agencies chase foreign plates, including one in Oradea for Romanian plates.

Most surprising facts:
- New in 2026: an "M1 regional" vignette for 15,000 HUF covers all tolled roads in the four M1 counties. With it, Nădlac to Hegyeshalom costs 29,380 HUF a year, against 61,760 for the national annual vignette.
- A trap for county-vignette users: the Pest and Komárom-Esztergom vignettes do not cover the M1 between Bicske and Szárliget. That stretch needs the Fejér vignette.
- Borș I/Ártánd to Kisújszállás is all free: main roads 42 and 4, then the M4 section to Törökszentmiklós, which has been toll-free since 2025-12-22.
- The 60-minute rule is still official, although your brief listed it as a possibly outdated claim. The limit: if you refuse to buy at a roadside check, you lose it.
- One daily vignette does not cover a night trip across midnight. It is valid only for one calendar day.
- Index.hu (2026-09-08) wrongly says the fine deadline is 75 days. The decree says 60.
- "Vármegye" becomes "megye" from 2026-10-01, so the county product will likely appear under both names.

Things you should know:
- The official site cannot be read with WebFetch. nemzetiutdij.hu is a JavaScript app. I read it through its public JSON endpoints, which are listed at the top of the file so your crawler can use them.
- The decree text comes from a copy. njt.hu refused connections, so the legal text is from the Wolters Kluwer copy on net.jogtar.hu. The county junction limits come from the decree's Annex 1 there, and they match the official 2026 maps.
- A few supporting facts come from secondary sources. These are the Borș II opening date, the M4 construction timeline and the ministry's statement on the M1 regional vignette for 2027. They are marked as secondary in the file.
- Outdated claims 1 and 10 are special cases. Claim 1 (Index.hu) is a wrong statement rather than an old one. Claim 10 is an official Hungarian agency page about Romanian winter tyres, left for the Romania research to check.
- An 11th outdated page was dropped. The 2023 Romanian news articles that call the 60-minute rule new are not wrong, so they are not in the table.
- Your brief's assumptions, corrected: the 10-day product ("heti") was never renamed, and foreigners can buy county vignettes.

Agent
Agent
Agent
Agent
Agent

The Germany data is written (no car toll, the 5 destination low-emission zones, the Munich and Stuttgart diesel bans, winter-tyre rules, outdated claims). Now Hungary, the most complex country.

Agent
Agent

Hungary's data is written. The planner now needs a real set cover: the cheapest combination of county vignettes and the M1 regional vignette that covers every section, the Fejér gap included.

Agent
Agent
Agent
Agent

21 tests pass, including the Fejér-gap case. Now the places, routes and the seed runner.

Agent
Agent
Agent
Agent

The dry run is clean: every reference resolves. Writing the dataset to Sanity now.

Agent
Agent

Sources and countries reference each other, so they have to be written in one transaction. Switching to a single commit.

Agent
Agent

224 documents are live in Sanity. Next I'm running the planner against the real data, with a Brașov → Munich Christmas trip in a Euro 5 diesel.

Agent
Agent
Agent
Agent

Found a bug: the Romanian price always came out as the electric-car band (20 lei), because the query never fetched the band field. Fixing it.

Agent
Agent

Fixed: a Euro 5 diesel now pays the Euro IV–V price (26 lei a day), and a car without a stated class pays the Euro 0 rate (29 lei), as Order 888/2026 says. Now tightening the warnings.

Agent
Agent

Warnings now show only what depends on your dates: winter rules in season, unpublished prices, the Munich sticker and Romania's pending Senate bill. Committing, then building the Knowledge Base.

Agent
Agent
Agent
Agent
Agent

The query returns 103 readable documents. Binding it to the Knowledge Base.

Agent
Agent
Agent

The dataset import failed with a server-side error on Sanity's end. I'm checking the website import, then testing whether a simpler query works.

Agent
Agent
Agent
Agent

Same failure twice, so it's likely our query rather than a hiccup. I'm testing a minimal query to narrow down the cause.

Agent
Agent
Agent

Only one dataset source is allowed per Knowledge Base, so the failed one has to go first.

Agent
Agent

The minimal query works, so something in the big projection trips Sanity's ingester. I'm narrowing it down by removing the unusual parts: select() and the nested .fact lookup.

Agent
Agent

The simplified query works. Now I'm adding back every field except the two suspects, rewritten in plain GROQ.

Agent
Agent
Agent
Agent
Agent

Down to 122 sources, under the 150 limit. Starting the Knowledge Base build in the background.

Agent
Agent

While the Knowledge Base builds, I'm switching on dataset embeddings (free, used for semantic ranking) and writing the configuration for the two Context MCP endpoints.

Agent
Agent
Agent

Embeddings are on. Next I'm checking which Claude models Vercel's AI Gateway offers and what they cost, since the free $5 has to last through judging.

Agent
Agent
Agent

The Knowledge Base is still building. Meanwhile I'm checking the exact AI SDK v7 API in the installed package, since v7 renamed things.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

Writing the agent now: it connects to both Context MCP endpoints, adds the deterministic planner as a tool, and streams answers with sources.

Agent
Agent
Agent
Agent
Agent

The agent typechecks. Next come the two API routes, both rate-limited and same-origin only:

  • /api/plan runs the planner alone. It's free and instant.
  • /api/chat runs the agent.
Agent
Agent

Both routes typecheck. Checking on the Knowledge Base build.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

There's an undocumented but typed client.context.mcpEndpoints API in @​sanity/client. If it works, you won't have to click through the Dashboard to create the endpoints. Reading its signature now.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

The planner now returns the road sections the map needs. Next is the strip map itself.

Agent
Agent
Agent
Agent

The components typecheck (the editor warnings are stale). Now the page itself: hero, trip form, and the chat panel with its visible trace.

Agent
Agent
Agent
Agent
Agent

The UI typechecks. Starting the dev server to look at it.

Agent
Agent
Agent
Agent
Agent
Agent

The website is running locally. Taking a first look.

Agent

It looks right: the expanded "Drum bun!" title, the form, and €59.36 for the example trip with the estimate note. Scrolling to the strip map.

Agent

The browser pane's screenshot is garbled, a known quirk on this machine. I'll use a small Playwright script for full-page screenshots at 1440 and 375 px.

Agent
Agent
Agent
Agent
Agent

The Knowledge Base build succeeded. Next I'm pulling its Issues (the contradictions it found) through the client API. First, checking how the client addresses a Knowledge Base.

Agent
Agent
Agent
Agent

19 well-organized entries, including a dedicated fact-check entry per country. Opening the winter-tyre fact-check entry to see how it cites sources.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

I have the source IDs for Romania and Austria. Getting the Hungarian ones, then applying the decisions.

Agent
Agent

Applying the decisions now: conflicts first, then the standing instructions.

Agent
Agent
Agent

Four instructions are saved. I typed the last one's second source ID from memory, so I'm verifying it exists.

Agent
Agent

The ID was wrong (the real one is c3edf50e-c40c-…), so I'm correcting the instruction's scope.

Agent
Agent
Agent
Agent
Agent

Scope fixed. Applying the gap fixes (road names the entries should mention, plus the Hungarian purchase-channel issue):

Agent
Agent
Agent
Agent

Now the snapshot command. It writes each decision's reasoning next to the issue, so the page can say why.

Agent
Agent

The apply job succeeded. Checking whether the winter-tyre entry now carries the right Romanian fine, and whether the rebuild filed new issues.

Agent
Agent

The background check against my instruction worked: it filed a new issue saying the fact-check entry still quotes the old 603–1,340 lei and proposed the exact fix. Applying it:

Agent
Agent
Agent

Now the "How it knows" page. It shows the pipeline from sources to agent, every contradiction with both sides and the decision, the standing instructions, and the Knowledge Base outline.

Agent
Agent

Taking the Knowledge Base snapshot, then screenshotting the new page.

Agent
Agent

One entry field (tldr) comes back as an object, not text. Checking its shape:

Agent
Agent
Agent

The Hungarian surcharge (two tiers of one penalty) and the other ranges are fine as they are. Re-seeding, then an incremental Knowledge Base refresh so it picks up the corrected penalty.

Agent
Agent
Agent

The claims now mix countries and lead with what you actually buy. Checking whether you've created the MCP endpoint yet:

Agent
Agent

The CLI can create the project, point it at the web folder and deploy. Deploying makes the site public, which is part of the plan you approved (live demo on Vercel). Linking the project:

Agent
Agent

The project exists (scope daniel-51ed/drum-bun). Now pointing its root at web and setting the framework:

Agent
Agent
Agent

Git Bash rewrote the /v9/... path, a known quirk here. Running it from PowerShell instead:

Agent
Agent

The root directory is set. Adding the two endpoint URLs as environment variables (they're not secret, but the code keeps them server-side), then deploying:

Agent
Agent
Agent

The build failed on the snapshot's old tldr objects: I fixed the script but hadn't re-run it. Re-running the snapshot, and making the page tolerate either shape:

Agent
Agent

The snapshot is refreshed. The refresh finished too, and the web sources are now 23, still under the budget at 126 in total. Redeploying:

Agent
Agent

The site is live at https://drum-bun-khaki.vercel.app. (drum-bun.vercel.app was taken; I'll try for a cleaner name later.) Checking the live "How it knows" page:

Agent
Agent
Agent
Agent
Agent

The fine-range instruction is saved, and the rebuild also re-checks two related Hungarian entries. While that runs, I'm checking the home page at phone width.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

The evaluation is ready: 16 questions and a runner. It needs the agent live, so it waits for your endpoint step. Checking now that the keyword-search baseline query works against Sanity:

Agent
Agent
Agent

score() rejects coalesce() inside a match. Rewriting it the documented way, one predicate per field:

Agent
Agent

The runner file is fixed; my scratch test didn't receive the path variable. Rewriting the test directly:

Agent
Agent
Agent

The keyword search now returns sensible hits, but my edit to the runner's call site didn't apply. Fixing it by hand:

Agent
Agent
Agent

The runner is fixed. Checking whether the endpoint exists yet:

Agent
Agent
Agent
Agent

The counts match the README. Committing, then checking what token Sanity's conversation Insights needs, so you only have to create tokens once.

Agent
Agent
Agent
Agent

The integration plugs into AI SDK 7's telemetry option. Wiring it in with the chat's thread ID, plus a note under the chat input that questions are stored.

Agent
Agent

Insights is wired in. Next, the daily job that classifies saved conversations (success score, sentiment, content gaps), as a Vercel cron route on Haiku to save budget.

Agent
Agent

Generating a random secret for the cron route and storing it straight in Vercel. It's never printed and I never see it.

Agent

Here's where things stand.

Built and live (checked this session):
- Site: https://drum-bun-khaki.vercel.app. It plans all 20 routes (4 Romanian start cities × Vienna, Munich, Stuttgart, Frankfurt, Berlin) with no login and no AI needed. It shows the cheapest vignette combination for your dates, Romania's new Euro-class prices from 1 Oct, winter and emission-zone rules, outdated claims you may have read, and every source. The layout checked clean at 1440 px and 375 px.
- Sanity: 224 documents built from fresh research in RO/HU/DE/EN, with 85 sources and 26 outdated claims.
- Knowledge Base: 19 entries from 126 sources. I resolved its 15 issues through an API that isn't documented but exists in @​sanity/client, and added 9 standing instructions. It even caught mistakes in my own data model. The "How it knows" page shows all of it.
- Tests: 23 planner unit tests pass.
- Ready but not run yet: the chat agent (two MCP endpoints, Insights logging, daily classification) and the 16-question comparison against keyword search. Neither has run: both need the endpoints below.

Blocked on you. The agent can't run until the MCP endpoints exist, and there's still none in your organization.

Your step 3:
1. Open https://www.sanity.io/@​ob2cyckj9 and click Context in the left sidebar.
2. Create a new MCP endpoint:
- Title: Drum Bun road rules (Knowledge Base)
- Name: drum-bun-kb
- Source: only the Knowledge Base Drum Bun road rules
- Instructions: paste this:

Road rules for private cars, campers and motorcycles driving between Romania, Hungary, Austria and Germany. Entries cite their sources. Official sources (issuer, ministry, law) outrank automobile clubs, press, blogs and forums; when an official web page and the law disagree, the law wins. Entries about "claims seen online" describe statements drivers read that are outdated or wrong: never present them as current, say what is true now and why. Prices and dates in entries are only valid for the periods they name; for a price on a specific travel date, prefer the structured planner.

3. Save, then reply done (or send a screenshot if the form looks different).

After that come the second endpoint, one access token, a public GitHub repo, and publishing the post. I'll give you those one at a time.

Agent

Step 3 is done and set up correctly: drum-bun-kb, with only the Knowledge Base as its source.

Your step 4: create the second MCP endpoint, in the same place (Context → new MCP endpoint):
- Title: Drum Bun structured rules (GROQ)
- Name: drum-bun-rules
- Source: the dataset production of project Drum Bun (pd5e7gez). Only the dataset, no Knowledge Base.
- GROQ filter (if the form has this field):

_type in ["country", "place", "route", "roadSection", "tollProduct", "rule", "zone", "claim", "source", "exchangeRate"]

- Instructions:

Structured road rules for Romania (RO), Hungary (HU), Austria (AT) and Germany (DE).
- tollProduct.prices[] are dated: the price for a travel date is the one with validFrom <= date <= validTo. Several prices on the same dates differ by band (Romania from 2026-10-01 prices by Euro class; band.appliesWhenUnknown marks the rate charged when the class is unknown). If no price covers the date, say the price is not published yet and give the latest one as an estimate.
- tollProduct.validity.unit: "days" = N calendar days, start day included; "months" = to the same calendar day; "calendarYear" = annual with overlap windows (yearStartsPrevious, yearEndsNext).
- tollProduct.activation.onlineDelayDays > 0 means an online purchase by a private customer is only valid after that many days (Austria: 18 days for annual and 2-month vignettes).
- roadSection.coveredBy lists the products that make a section legal; any one is enough. tolled == false means free. exemptVehicles owe nothing on that section.
- rule.season is a yearly MM-DD window; rule.conditional means it only applies under rule.condition.
- zone.dieselMinEuro is the lowest Euro class a diesel may have; status other than "active" means not in force.
- claim documents are statements found online. verdict "outdated", "wrong" or "misleading" means they are not true now; correctedBy points at the facts that are.
- source.trust ranks sources: official > club > press > blog > forum.
- Never state a price, fine or date that is not in a retrieved document.

Save, then reply done. I'll check the setup from here.

You

<pasted_content id="49e8">
Not ready. Some things need fixing.
Studio
No Studio application found for this project/dataset
Schema
Studio check failed — no schema descriptor available
Semantic search
Ready
</pasted_content id="49e8">

Agent

That readiness check says GROQ mode needs a deployed Studio: Sanity reads the schema from the Studio application. We only uploaded the schema, which isn't enough. Deploying the Studio now (it lands at drum-bun.sanity.studio and only project members can sign in).

Agent
Agent
Agent

The Studio is deployed at https://drum-bun.sanity.studio, and the deploy also uploaded its schema descriptor, which fixes both red items.

Your step: in the endpoint form, re-run the check. If the form has no re-check button, close it and start the drum-bun-rules endpoint again with the same values. Studio and Schema should now show as ready. Save, then reply done, or paste the check again if something is still red.

Agent

The guard objected because I piped that command's output through a filter. Running it plainly, since its output holds no values:

Agent

All variables are in place for every environment. Restarting the local dev server so it loads them, then one real test question.

Agent
Agent
Agent

The stream hides the real error behind a generic message. Adding server-side logging and a readable error in development:

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

The free tier does allow some strong models: openai/gpt-5.1 and gpt-5-mini both answered. Probing a few newer ones for the best allowed balance of quality and cost:

Agent
Agent
Agent
Agent
Agent
Agent
Agent

Still tuning the model and citations. Gemini just returned an empty answer for a simple question, which is a reliability red flag. Looking at the raw response:

Agent
Agent
Agent
Agent
Agent

The evaluation's baseline needs model access too, and that key lives only inside the dev server. So the baseline runs as a development-only route, and the runner calls both over HTTP.

Agent
Agent

The per-IP chat limit would stop the 16-question run partway, so it now applies only outside development. Then running the full evaluation in the background:

Agent
Agent
Agent

The evaluation is running (about five minutes). Meanwhile I'm checking that nothing secret is tracked in git before the repo goes public.

Agent
Agent

The evaluation output is buffered until it finishes. Meanwhile I'm making the cover image in the same strip-map style.

Agent
Agent
Agent
Agent

The cover works: the same strip-map route, with the one key fact per country. Saving it to the repo and checking the evaluation:

Agent
Agent

The repo is public at https://github.com/danielbutnar/drum-bun with everything committed so far. Checking the evaluation:

Agent
Agent
Agent
Agent
Agent
Agent

The code typechecks. Pulling the new Insights token into local env, adding a language-detection test, and pacing the evaluation to stay under 5 requests a minute:

Agent
Agent
Agent

The English question with "Brașov" in it was detected as Romanian, the exact failure the check is meant to catch. The fix: ignore place names, and count English words too.

Agent
Agent
Agent
Agent

My scripted edit lost the regex backslashes (\s became s). Rewriting the function directly:

Agent
Agent
Agent
Agent

All 22 tests pass. Checking the editor's two type complaints against the real compiler:

Agent
Agent

Both were stale; the compiler is clean. Committed and pushed. Starting the full paced evaluation in the background (about 20 minutes, to respect the rate limit):

Agent
Agent
Agent
Agent

Now the chat side: an example click replays its recording instantly, with a label and an Ask live button.

Agent
Agent

Replay is wired in and typechecks. Checking the evaluation's progress:

Agent
Agent

Staying free. Updating the build notes while the evaluation finishes:

Agent
Agent

It started at about 21:37 and should finish within minutes. Meanwhile I'm checking that yesterday's Knowledge Base fixes reached the outline summaries, then refreshing the public snapshot.

Agent
Agent
Agent
Agent
Agent

After re-scoring both systems: agent 15/16, keyword search 11/16. Checking why keyword search failed its four, so the post can say precisely where structure matters:

Agent
Agent
Agent

While that records, I'm confirming the evaluation run files are committed (they were git-ignored before):

Agent
Agent

The run files are committed, and the apply job has finished. While the recording runs, I'm looking at the video pipeline from the Lovable challenge to reuse it for the demo video.

Agent

The pipeline is reusable: narration → local voice synthesis → a paced browser recording with captions → ffmpeg build. I'll adapt it for Drum Bun tomorrow. Checking on the example recording:

Agent

Status, about 22:05:

  • Evaluation: agent 15/16, keyword search 11/16 on the same Sanity content, after widening two patterns and re-scoring both systems equally. The agent wins where the answer must be computed: the Christmas trip, the Euro-class price, the Hungarian county cover, the surcharge amount. The earlier bad run (6/16) and its causes are kept honestly in evals/.
  • Model: gpt-5-mini on the free tier, with recorded example answers and a clear "busy" message for when the 5-requests-a-minute limit hits.
  • Knowledge Base: it keeps catching stale text by itself after my standing instructions, and I apply the fixes.

Running now: recording the four example answers (about 5 minutes). Then I deploy the new agent to the live site, run one full Knowledge Base rebuild so the outline summaries pick up every fix, and refresh the "How it knows" snapshot.

Tomorrow (Wed):
- Morning: the demo video, with narration, filmed from the live site.
- Then: finalize the DEV post (numbers, screenshots, cover) and run a full QA pass.

Nothing is needed from you tonight. The remaining steps only you can do are all at the end: uploading the agent session, then publishing the post.

Agent
Agent

Deployed. Testing one live question against production:

Agent
Agent

The safety guard blocks sending data from the shell to an outside site, and I won't route around it. Testing the way a judge would instead: in the browser, on the live site.

Agent

The browser pane isn't drawing while it's in the background, a known quirk here. I'll run the same judge-style UI test with Playwright: click an example, then type a live question.

Agent
Agent
Agent
Agent
Agent

The live site now runs the new agent, and a UI test on production passed:
- Recorded example: shows instantly.
- Live question: the Berlin fine, answered in 15 s with the correct €100 and a note that pages quoting €80 are outdated.
- Console errors: none.

Running now: a full Knowledge Base rebuild (about 15–20 minutes). When it finishes I'll review any new issues it files and refresh the public snapshot.

Tomorrow:
- The demo video with narration, filmed from the live site.
- Final post text with the evaluation table and screenshots.
- A full QA pass on the site.

Nothing is needed from you until the final publishing steps.

Agent

The rebuild is in its first stage and will take a while. Meanwhile I'm starting on the demo video, reusing the Lovable pipeline. Reading the rest of it first:

Agent
Agent

The voice models are in the Lovable session's scratchpad, so I'll use them from there. Checking whether the build and mix scripts are generic enough to reuse as they are:

Agent
Agent

ffmpeg is available through uv. Generating the 15 voice lines with the same voice as the Lovable video (am_michael):

Agent
Agent
Agent

137 seconds of narration, so the video will run about 2:30. Now the filming script. Reading the rest of the old one's helpers and its ending, so the build step gets the exact files it expects:

Agent
Agent

Everything the build step needs is clear. Writing the Drum Bun filming script: same helpers (captions, cursor, paced by the voice lines), new scenes on the live site.

Agent

The filming script is written. A dry run first: headless, no recording, only checks that every step finds its element on the live site.

Agent
Agent

All 14 scenes find their elements on the live site. Before filming for real, the Knowledge Base rebuild has to finish so the "How it knows" page shows the final state. Checking it:

Agent

The rebuild is in its review stage (7 of 22 entries), roughly 10 more minutes. Committing the video scripts meanwhile; the generated audio stays out of git.

Agent

While the rebuild finishes, a small polish: the live address drum-bun-khaki.vercel.app looks auto-generated. Trying a cleaner alias:

Agent
Agent
Agent

drum-bun-agent.vercel.app works. A plain alias would stay on this one deployment, so I'm adding it as a project domain; then every production deploy moves it automatically:

Agent
Agent

The domain is verified and attached to the project. Updating every reference in the repo and in memory:

Agent

The URL replacement is still running in the background. Meanwhile the Knowledge Base rebuild finished; reading its result:

Agent
Agent

The full rebuild reorganized the outline into a clearer tree of 22 entries, e.g. tolls/austria/penalties and emission_zones/germany/munich. Checking the two stale phrases across all entries:

Agent
Agent

Clean: none of the stale phrases remain anywhere in the 22 entries. Taking the final snapshot for the "How it knows" page, and checking the URL-replacement job:

Agent

The replacement never ran: rg without a path reads from stdin when there's no terminal, so it waited forever. Stopping that job and editing the four files directly:

Agent

The snapshot is refreshed (22 entries, 21 issues decided, 9 instructions) and pushed. The memory note has two statements that are no longer true, so I'm correcting them:

Agent

While that records, I'm generating the background music pad (synthesized, nothing to license) and loading the live "How it knows" page so it picks up the new snapshot:

Agent
Agent
Agent

The live page still shows the old snapshot. A few minutes of caching is normal; I'll check again after the recording and redeploy if needed. Checking the recording:

Agent
Agent
Agent

The new recordings use the new Knowledge Base paths (for example tolls/germany/toll_free_routes). Committing and deploying, which also rebuilds the "How it knows" page with the fresh snapshot:

Agent
Agent
Agent

The live page now shows the final Knowledge Base (22 entries). Before filming, a quick dry run to confirm every scene still finds its element:

Agent
Agent

Every scene checks out. Filming the real take now: a Chrome window will open on your screen for about 2.5 minutes. Please don't click into it.

Agent
Agent

The take is recorded (750 frames; the screencast only captures when something changes). Building the video with the voice and caption band, then mixing in the music:

Agent
Agent

A 2:42 video with voice, captions and music (26 MB). Checking it visually: a contact sheet of 12 frames spread across the video.

Agent
Agent
Agent
Agent
Agent

After the title card hides, the page is already scrolled down to the plan instead of showing the home page. Something scrolls on load. Measuring it:

Agent
Agent

The page doesn't scroll by itself (it stays at the top after load), so it happens during the take. Reproducing the opening of the take with scroll logging:

Agent
Agent

The page never scrolls under the card by itself, so this is likely a timing issue in the build, not the page. Comparing the frame timestamps with the voice markers:

Agent
Agent
Agent
Agent
Agent
Agent

The first cut of the demo video is in your chat above: 2:41 in 1080p, with voice, captions and background music, filmed from the live site. Please watch it and tell me anything you'd change.

Done tonight (all checked):
- Live site: https://drum-bun-agent.vercel.app (new, cleaner address). The planner, the recorded example answers, a live chat answer and the "How it knows" page all work. The code is on GitHub at https://github.com/danielbutnar/drum-bun.
- Evaluation: the agent got 15 of 16 test questions right, keyword search over the same content 11 of 16. The first run went 6 against 9, and the post says so honestly; that failure is how the agent got good.
- Knowledge Base: the full rebuild produced no new conflicts, so all earlier decisions carried through. It has 22 entries, and a script confirmed none of the old errors remain.
- Post draft: about 1,400 words, with the evaluation table and the honest notes. Still missing: the video link, two screenshots and the agent session.

Tomorrow (Wed), in order:
1. Your feedback on the video.
2. I add screenshots and the cover to the post; they're hosted on GitHub, so no manual uploads.
3. You upload the video to YouTube (I'll guide you step by step).
4. I prepare the Claude Code session transcript without the business/HQ parts; you upload it on DEV and make it public.
5. A full QA pass on the site.
6. You publish the post, by Sunday evening at the latest. The hard deadline is Monday 5 Oct, 09:59 Brașov time.

Nothing else is needed from you tonight.

You

shouldn t I have an expination for the succes rate witch is not perfect, or should I keep it like this?

Agent

Yes, add an explanation. A short, concrete one makes the number more believable than a perfect score would. Judges see a lot of "100% accurate" claims, and a 16/16 on a test I wrote myself would look tuned. The post already explains the bad first run, but not the one question the final run missed, so I'm adding that now:

Agent
Agent

The post now says this, right after the results table:

The one question the agent missed: someone typed the wrong plate on an Austrian 10-day digital vignette that has already started. The agent said the right thing, that it can no longer be changed and that a wrong plate counts as no vignette (€200 on the spot), but it left out that annual vignettes are the exception (re-registration for €18). The test requires that detail, so it counts as a miss. I kept the test as written instead of tuning the prompt until it passes; keyword search missed it too.

The same explanation is in evals/RESULTS.md in the repo.

Why I'd leave it at 15/16 rather than tweak the agent until it passes:
- Credibility: changing the prompt after seeing a test's answer is tuning to the test. Judges who read the eval section, and the Sanity team will, trust an explained miss more than a perfect score.
- The miss isn't dangerous: the answer was correct for the driver's situation. It only lacked a side detail.
- It matches the rest of the post: the post already shows the bad first run (6/16) and the two scoring patterns I widened for both systems.

The video just says "fifteen of sixteen", which is fine: the explanation belongs in the post, where people read the details.

You

the video is good, see if you can find improvments tho.
Then continue

Agent

Starting by checking the repo state: several files show as changed on disk since my last edits.

Agent
Agent

148 seconds of speech now, so the video lands around 2:50. That's fine for a demo this dense. Updating the scenes:

Agent
Agent

The dry run passes. Filming take 3, building, and checking the changed scenes:

Agent
Agent
Agent

Next is the post's screenshots, but first a full QA pass on the live site, since the screenshots should show the final state. Loading the QA skill:

Agent
Agent
Agent
Agent

The public folder disappeared when I removed the template files, so the cover copy failed. Recreating it, and adding a Content-Security-Policy header for production:

Agent
Agent

security.txt wasn't written either (it came after the failed copy). Writing it, then the CSP header:

Agent
Agent
Agent
Agent
Agent

Typecheck passes, but lint caught a React-rules issue in the strip map (a variable mutated inside map). Rewriting that with cumulative sums:

Agent
Agent

Lint, typecheck and all 22 tests are clean. The web package was missing typecheck and test scripts, so the repo-wide check couldn't run them. Adding them and running the full check:

Agent
Agent

pnpm check passes end to end: typecheck, lint, then tests (the steps chain, so reaching the tests means the first two passed). Committing, deploying and re-running site-check:

Agent
Agent
Agent

The retry deployed fine. Re-running site-check:

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

QA fixes are committed. Deploying them, then taking the two screenshots for the post:

Agent
Agent

Deployed. Taking the post screenshots at high resolution from the live site: the plan with its strip map, the chat trace showing Knowledge Base reads, and one Knowledge Base decision.

Agent
Agent
Agent
Agent
Agent
Agent
Agent

Now wiring the images into the post. They'll load from the public GitHub repo, so there's nothing for you to upload:

Agent

Top comments (0)