<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kiter</title>
    <description>The latest articles on DEV Community by Kiter (@kiter).</description>
    <link>https://dev.to/kiter</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1441774%2Fa8b26748-7bce-4038-aed3-75737911a526.jpeg</url>
      <title>DEV Community: Kiter</title>
      <link>https://dev.to/kiter</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kiter"/>
    <language>en</language>
    <item>
      <title>I shipped an agent that asks Sanity Context what it is allowed to say</title>
      <dc:creator>Kiter</dc:creator>
      <pubDate>Sat, 03 Oct 2026 18:25:04 +0000</pubDate>
      <link>https://dev.to/kiter/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say-1jol</link>
      <guid>https://dev.to/kiter/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say-1jol</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path One: Ship an Agent That Queries Real Content&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agent's MCP endpoint:&lt;/strong&gt; &lt;code&gt;https://api.sanity.io/v1/context/organizations/ovihgdwkx/mcp/ninety&lt;/code&gt;&lt;br&gt;
&lt;strong&gt;Live app:&lt;/strong&gt; &lt;a href="https://ninety-europe.vercel.app" rel="noopener noreferrer"&gt;https://ninety-europe.vercel.app&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Context evidence page:&lt;/strong&gt; &lt;a href="https://ninety-europe.vercel.app/context" rel="noopener noreferrer"&gt;https://ninety-europe.vercel.app/context&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/PhiBao/ninety" rel="noopener noreferrer"&gt;https://github.com/PhiBao/ninety&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Sanity project:&lt;/strong&gt; &lt;code&gt;jvgi63fz&lt;/code&gt; · dataset &lt;code&gt;production&lt;/code&gt; · Studio: &lt;a href="https://beyond-vibe.sanity.studio" rel="noopener noreferrer"&gt;https://beyond-vibe.sanity.studio&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  The rule the whole agent is built around
&lt;/h2&gt;

&lt;p&gt;Most travel assistants will confidently tell you how many days you have left. They will also be wrong sometimes, and you cannot tell which times.&lt;/p&gt;

&lt;p&gt;So the agent in this submission is &lt;strong&gt;architecturally forbidden from stating a number.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// The engine. No model, no network, no env vars.&lt;/span&gt;
&lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;snapshot&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;itinerary&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;asOf&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="c1"&gt;// → {used: 73, attribution: {date: '2026-10-26', blame: [...]}, unresolvedDays: [...]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model may retrieve, classify, explain and cite. Every figure it reports comes from that engine. The split is not stylistic — it is the reason I did not use a language model at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the agent actually is
&lt;/h2&gt;

&lt;p&gt;Type &lt;strong&gt;"three weeks on Tenerife, then an 8 hour airport layover where I never cleared immigration"&lt;/strong&gt; and it works out what that is, then what it costs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdnzb12khatux27yif2y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdnzb12khatux27yif2y.png" alt="The agent" width="800" height="625"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It resolves to three separate reads: &lt;strong&gt;Bulgaria by train — 10 days counted&lt;/strong&gt;, from the date-banded rule that opened its land crossings on 31 December 2024. An &lt;strong&gt;airport transit with no place named — no days counted&lt;/strong&gt;, at 100%. And in between, a leg it is only about 50% sure about, so it refuses that one.&lt;/p&gt;

&lt;p&gt;That last behaviour is the whole point. In the middle of that sentence, &lt;em&gt;"flew to Paris 20 to 25 February"&lt;/em&gt; is a fragment, and the agent says so rather than guessing. A chat completion would have answered it fluently and been indistinguishable from a correct answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  The pipeline
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;traveller's words
     │
     ├─ date extraction            deterministic, in code, unit-tested
     │
     ├─ Sanity Context (MCP)  ──▶  the candidate set. Retrieved, never invented.
     │                              groq_query over territory + presenceRule
     │
     ├─ TypeSafe System One   ──▶  typed classification + calibrated confidence
     │                              {choice, probabilities, confidence}
     │
     └─ the engine            ──▶  every number in the answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three properties fall out of that, and each one was bought by a bug I hit while building it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. It cannot invent a place.&lt;/strong&gt; The territory list is retrieved from the corpus at request time. If a place is not there, there is no option to select, so the output is &lt;code&gt;"No place named"&lt;/code&gt;. Not a guess at the nearest match — the set does not contain one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. It cannot do arithmetic.&lt;/strong&gt; Classification goes through TypeSafe's System One models, which return a value from a set I defined plus the full probability distribution. They are not text generators; there is no channel through which &lt;code&gt;90&lt;/code&gt; could travel. The whole web app has &lt;strong&gt;five runtime dependencies and no model-provider SDK at all&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. It can say it does not know.&lt;/strong&gt; Below 0.62 confidence the read is reported as uncertain rather than answered.&lt;/p&gt;

&lt;p&gt;One consequence worth stating, because it is a bug I shipped and then caught: the classifier also answers &lt;em&gt;"would this person be present somewhere that counts against the allowance?"&lt;/em&gt;, and it does not know that Bulgaria's land crossings were internal in February 2025. So it says no for Sofia-by-train. The engine, which does know, charges ten days. Rather than show two contradicting labels, &lt;strong&gt;the "N days counted" figure is read back out of the ledger the engine actually produced&lt;/strong&gt; — the classifier never gets to have an opinion about charging.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sanity Context as the retrieval layer
&lt;/h2&gt;

&lt;p&gt;The agent's knowledge is a hosted, read-only MCP endpoint. Not a hardcoded prompt, not a vector index built at deploy time — a Sanity-managed endpoint that serves the dataset's schema and answers GROQ on the wire.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;What the agent uses it for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;initial_context&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Schema overview: types, fields, relationships, document counts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;groq_query&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Structured retrieval with projections — the candidate set, every request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;schema_explorer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Field-level detail when the overview is not enough&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;array_field_reader&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Reading &lt;code&gt;accessBands[]&lt;/code&gt; and &lt;code&gt;competingClaims[]&lt;/code&gt; without pulling whole documents into context&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  This is live, not a description
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;/context&lt;/code&gt; page calls the endpoint server-side and prints what came back, including its failure modes. Everything on it was fetched when you loaded the page.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2icvgw71syz6jl6tfvc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2icvgw71syz6jl6tfvc.png" alt="Sanity Context" width="800" height="688"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Reproduce it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.sanity.io/v1/context/organizations/ovihgdwkx/mcp/ninety &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$SANITY_CONTEXT_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Accept: application/json, text/event-stream'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"jsonrpc":"2.0","id":1,"method":"tools/list"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And a live &lt;code&gt;groq_query&lt;/code&gt; through Context, asking about the fact that makes the same fortnight cost fourteen days in Split in 2023 and nothing in 2022:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"tools/call"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"params"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"groq_query"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
   &lt;/span&gt;&lt;span class="s2"&gt;"*[_type==&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;territory&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;&amp;amp;&amp;amp;code==&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;HR&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;][0]{name,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;bands&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:accessBands[]{window{from,to},counted,modes}}"&lt;/span&gt;&lt;span class="p"&gt;}}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Croatia"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"bands"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"counted"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"window"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"from"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2000-01-01"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"to"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2023-01-01"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"counted"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"window"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"from"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2023-01-01"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"to"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;}}]}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the product thesis in one payload. A keyword index returns the paragraph about Croatia joining in 2023. This returns the &lt;em&gt;bands&lt;/em&gt;, which is what you need in order to answer a question about a specific day.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the dataset has to be shaped this way
&lt;/h2&gt;

&lt;p&gt;An agent's answer quality is bounded by the structure it can query. Four decisions do the work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Membership is a date-banded array, not a boolean.&lt;/strong&gt; &lt;code&gt;isSchengen: true&lt;/code&gt; is a lie for any country that joined later, and it is the reason most assistants get Croatia and Bulgaria wrong. &lt;code&gt;accessBands[]&lt;/code&gt; with &lt;code&gt;{from, to, counted, modes}&lt;/code&gt; makes &lt;em&gt;"was this in the area on this day, and by which route"&lt;/em&gt; answerable at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Carve-outs are territories.&lt;/strong&gt; The Canary Islands, Madeira, Åland, Svalbard, the French overseas departments, Ireland, Cyprus and Iceland are each their own &lt;code&gt;territory&lt;/code&gt; document with their own bands. The agent can therefore say &lt;em&gt;where&lt;/em&gt; a day was spent, which is also what makes attribution possible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ambiguity is a value, not an exception.&lt;/strong&gt; A &lt;code&gt;presenceRule&lt;/code&gt; can be &lt;code&gt;counted&lt;/code&gt;, &lt;code&gt;not_counted&lt;/code&gt;, or &lt;code&gt;disputed&lt;/code&gt;. When it is &lt;code&gt;disputed&lt;/code&gt; the engine refuses to classify a day of that kind, and the agent is required to present both readings and name the disagreement. It cannot resolve it, because there is nothing in the corpus to resolve it with.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every rule resolves to a source.&lt;/strong&gt; Each band, presence rule and regime carries &lt;code&gt;sourceRef[]&lt;/code&gt; to a real document with a publisher, a URL and a retrieval date. The agent's citations are checkable rather than plausible — and so is the &lt;em&gt;absence&lt;/em&gt; of one, which is what tells the agent to stop.&lt;/p&gt;




&lt;h2&gt;
  
  
  Does the structured content actually earn its keep?
&lt;/h2&gt;

&lt;p&gt;The challenge's own test: &lt;em&gt;if a keyword search would have gotten you the same answer, aim higher.&lt;/em&gt; So I measured it.&lt;/p&gt;

&lt;p&gt;Two systems, the same corpus, 23 adversarial histories with hand-computed expectations written before the runner existed, so the runner cannot grade itself. &lt;strong&gt;Neither arm uses a language model&lt;/strong&gt; — the point is to isolate the contribution of the content structure, and this way anyone can rerun it with &lt;code&gt;pnpm eval&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Structured (GROQ + engine)&lt;/th&gt;
&lt;th&gt;Keyword search (BM25, same rules)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Correct answer&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23/23&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1/23&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Produced a day count&lt;/td&gt;
&lt;td&gt;21/23&lt;/td&gt;
&lt;td&gt;0/23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Named the exact breach date&lt;/td&gt;
&lt;td&gt;1/1&lt;/td&gt;
&lt;td&gt;0/1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refused instead of guessing&lt;/td&gt;
&lt;td&gt;2/2&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Full results with the reasoning for each case: &lt;a href="https://github.com/PhiBao/ninety/blob/main/web/eval/RESULTS.md" rel="noopener noreferrer"&gt;&lt;code&gt;web/eval/RESULTS.md&lt;/code&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The keyword arm is not handicapped. It is given the same territory records, the same presence rules, the same permit exemptions and the same precedents, flattened into prose, and it scores BM25 over them. It frequently retrieves the right paragraph.&lt;/p&gt;

&lt;p&gt;It cannot produce a verdict, and the reason is worth stating plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The corpus holds the rules, and never held the traveller.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;"I have eleven trips logged this year and one booked — am I still legal?" is not a retrieval problem. It is a join between the rule and eleven date ranges, two of which fall in carve-outs, one of which nobody has ruled on, evaluated day by day across a rolling window. Retrieval finds the paragraphs. It cannot join them to a calendar.&lt;/p&gt;




&lt;h2&gt;
  
  
  Knowledge Bases
&lt;/h2&gt;

&lt;p&gt;Being straight about this one: Ninety does &lt;strong&gt;not&lt;/strong&gt; use a Sanity Knowledge Base, and the rubric asks for one.&lt;/p&gt;

&lt;p&gt;What it uses is the dataset served through Context, which covers the structured half of the problem — the rules — but not the prose half. The EU guidance documents the corpus cites are currently referenced as sources, not ingested as searchable text. So an agent question like &lt;em&gt;"why is a same-day transit not a stay?"&lt;/em&gt; can be answered from the structured rule and its citation, but cannot yet quote the underlying paragraph.&lt;/p&gt;

&lt;p&gt;Ingesting those documents as a Knowledge Base is the next step, and it is the step that would let the agent say &lt;em&gt;"the border-crossing page says X, the visa-policy page says Y, and that is why the day is disputed"&lt;/em&gt; instead of citing two titles and stopping. I would rather name the gap than imply it is closed.&lt;/p&gt;




&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/PhiBao/ninety" rel="noopener noreferrer"&gt;https://github.com/PhiBao/ninety&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;infra/sanity.blueprint.ts        infrastructure: CORS origin + the guard function
infra/functions/                 a Sanity Function that enforces the invariant
web/src/lib/agent/describe.ts    free text → dated stays, deterministically
web/src/lib/agent/typesafe.ts    typed client for the System One primitives
web/src/lib/agent/resolve.ts     retrieve candidates, classify, hand off
web/src/lib/sanity/context.ts    JSON-RPC client for the Context MCP endpoint
web/src/app/context/page.tsx     the evidence page, rendered server-side
web/src/lib/engine/              pure TypeScript. No model, no network.
web/scripts/eval/                the 23-case evaluation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A note on why &lt;code&gt;context.ts&lt;/code&gt; speaks raw JSON-RPC instead of going through an agent SDK: the diagnostics should show what an agent actually receives on the wire, not what a library chooses to show it. Every function degrades to a typed &lt;code&gt;unavailable&lt;/code&gt; result instead of throwing, so the evidence page reports a missing credential plainly rather than pretending the integration exists.&lt;/p&gt;

&lt;p&gt;There are 56 unit tests. The ones that matter most here are about the free-text layer, because that is where the model is closest to the answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;that &lt;code&gt;"1 to 21 February 2026"&lt;/code&gt; is one stay and not two&lt;/li&gt;
&lt;li&gt;that &lt;code&gt;"30 to 31 February 2026"&lt;/code&gt; is rejected rather than rolled into March&lt;/li&gt;
&lt;li&gt;that &lt;em&gt;"flew to Paris"&lt;/em&gt; and &lt;em&gt;"took the train into Sofia"&lt;/em&gt; in the same sentence each get their own place, their own mode, and their own classification&lt;/li&gt;
&lt;li&gt;that &lt;em&gt;"never cleared immigration"&lt;/em&gt; is an airside transit, not a cleared one — the one distinction that decides whether the day is charged&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  What I got wrong, since it is relevant here
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Context did not work until the Studio was deployed.&lt;/strong&gt; For most of this build the endpoint answered &lt;code&gt;Only datasets with deployed Studio applications are supported&lt;/code&gt;. It was never an authorisation problem in the end — I had assumed that, and burned time on it. Deploying the Studio fixed it in one command.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Asking N questions about an array of N stays does not bind question &lt;em&gt;i&lt;/em&gt; to stay &lt;em&gt;i&lt;/em&gt;.&lt;/strong&gt; I batched all the classifications into one request to save latency. The state was an array of stays and each question said "this description", so an 8-hour airport layover in the second stay got classified as the Canary Islands because the first stay mentioned Tenerife. One request per stay now. Cheaper in latency than it sounds, because each request is small.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;\s&lt;/code&gt; in a template literal collapses to &lt;code&gt;s&lt;/code&gt;.&lt;/strong&gt; My date-stripping regexes compiled into "match a literal s" and silently removed nothing, so the classifier was being handed &lt;em&gt;"layover on 3 June 2026"&lt;/em&gt; with the date still in it. A regex that silently matches nothing is worse than one that throws.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two evaluation cases were passing for the wrong reason.&lt;/strong&gt; I had written a territory code as &lt;code&gt;es-canary&lt;/code&gt;, which is a document &lt;em&gt;key&lt;/em&gt;; the code is &lt;code&gt;XCI&lt;/code&gt;. The engine's response to an unknown territory is a warning and zero charged days — so the case passed because the Canary days had been dropped rather than classified. The number was right and the reasoning was nonsense. &lt;code&gt;pnpm check:codes&lt;/code&gt; now fails the build if any case contains a dropped stay.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Sanity Function that guards the invariant failed in three ways that all looked like success.&lt;/strong&gt; The agent's correctness rests on a ruling being recorded as a precedent, so I added a document function to enforce that wherever the dispute is adjudicated — including in the Studio, bypassing the app. It wrote the precedent fine, then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;client.patch()&lt;/code&gt; is lazy in &lt;code&gt;@sanity/client&lt;/code&gt; v8. The handler awaited it, logged "attached precedent to dispute", and changed nothing.&lt;/li&gt;
&lt;li&gt;A transaction id derived from the document id looks like idempotency. Sanity remembers transaction ids permanently, so the second adjudication returned &lt;code&gt;transactionAlreadyExistsError&lt;/code&gt; and the function failed forever after, silently.&lt;/li&gt;
&lt;li&gt;The guard read a denormalised &lt;code&gt;presenceKind&lt;/code&gt; that the document does not have, so it declined to act on a dispute whose &lt;code&gt;subjectKind&lt;/code&gt; was plainly &lt;code&gt;presence_kind&lt;/code&gt;. It now resolves the kind through the reference.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third one generalises. A guard that refuses because it looked in the wrong place is worse than no guard: it turns a loud failure into a quiet one. Every check in this repo looks at what the system &lt;em&gt;did&lt;/em&gt;, not at whether it reported success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I had Iceland in the Schengen Area.&lt;/strong&gt; It is EEA, not Schengen, and my corpus said otherwise because I had used a shared helper that stamped the founding date onto every state. Reykjavík days were being charged against 90/180 — a mistake a lot of tools make, made silently, in the one file I was treating as ground truth.&lt;/p&gt;




&lt;p&gt;Not legal advice. It covers a handful of passport classes and refuses outside that coverage, which is the only behaviour that makes it safe to trust. 🤖&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Ninety: I vibe-coded a Schengen day accountant that admits when it doesn't know</title>
      <dc:creator>Kiter</dc:creator>
      <pubDate>Sat, 03 Oct 2026 16:05:56 +0000</pubDate>
      <link>https://dev.to/kiter/ninety-i-vibe-coded-a-schengen-day-accountant-that-admits-when-it-doesnt-know-gja</link>
      <guid>https://dev.to/kiter/ninety-i-vibe-coded-a-schengen-day-accountant-that-admits-when-it-doesnt-know-gja</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path Two: Vibe-Code Something Strange&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live:&lt;/strong&gt; &lt;a href="https://ninety-europe.vercel.app" rel="noopener noreferrer"&gt;https://ninety-europe.vercel.app&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Code:&lt;/strong&gt; &lt;a href="https://github.com/PhiBao/ninety" rel="noopener noreferrer"&gt;https://github.com/PhiBao/ninety&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Sanity project:&lt;/strong&gt; &lt;code&gt;jvgi63fz&lt;/code&gt; · dataset &lt;code&gt;production&lt;/code&gt; · Studio: &lt;a href="https://beyond-vibe.sanity.studio" rel="noopener noreferrer"&gt;https://beyond-vibe.sanity.studio&lt;/a&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ninety&lt;/strong&gt; knows the exact day you went over.&lt;/p&gt;

&lt;p&gt;Every Schengen short-stay calculator tells you &lt;em&gt;how many&lt;/em&gt; days you have used. None of them tell you &lt;em&gt;which day&lt;/em&gt; you crossed the line, &lt;em&gt;why&lt;/em&gt; that day counted, or &lt;em&gt;what to change&lt;/em&gt;. Ninety does all three.&lt;/p&gt;

&lt;p&gt;It is a day accountant for the 90-in-any-180 rule. You give it your trips; it resolves the rules that apply on &lt;strong&gt;each individual day&lt;/strong&gt;, counts a rolling window, and tells you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;73 of 90 days used&lt;/strong&gt; — and 17 remain&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On 26 October the window reaches 91 days.&lt;/strong&gt; It is &lt;em&gt;Barcelona — booked&lt;/em&gt; that puts you over.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leave on the 26th, not the 27th.&lt;/strong&gt; Two days off the end of that trip fixes it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then it shows its work: every calendar day as a mark on a strip, colour-coded by whether it was charged, and clicking any day gives you the rule that decided it and the authority behind it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fadniipt07qg1hoah1gtk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fadniipt07qg1hoah1gtk.png" alt="The verdict" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  The three things it does that a calculator cannot
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. It names the day, not just the number.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjknh4zejpotzztjr1a17.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjknh4zejpotzztjr1a17.png" alt="The day you went over" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Attribution is a real algorithm, not a caption. On the breach day the engine finds the rolling window, works out which days pushed the count past the limit, and maps them back to the trips that contained them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. It classifies instead of adding.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Days are not all alike, and this is where people actually get hurt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Three weeks in the &lt;strong&gt;Canary Islands&lt;/strong&gt; — Spanish territory, outside the area. &lt;strong&gt;0 days.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A month in &lt;strong&gt;Dublin&lt;/strong&gt; — an EU member that opted out of the acquis. &lt;strong&gt;0 days.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A week in &lt;strong&gt;Reykjavík&lt;/strong&gt; — Iceland is in the EEA, which is &lt;em&gt;not&lt;/em&gt; the Schengen Area. &lt;strong&gt;0 days.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;airside&lt;/strong&gt; airport transit — never crossed a border. &lt;strong&gt;0 days.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;The same flight into &lt;strong&gt;Sofia by air&lt;/strong&gt; in February 2025 — &lt;strong&gt;0 days&lt;/strong&gt; — but &lt;strong&gt;overland the same week: 10 days.&lt;/strong&gt; Bulgaria opened its land and sea crossings on 31 December 2024 and kept air borders external until 31 March 2025.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;pending&lt;/strong&gt; residence application — not a permit. &lt;strong&gt;31 days&lt;/strong&gt;, exactly as if you had nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;3. It refuses, and it says so.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Give it a passport it has never heard of and it does not guess:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Ninety declined to answer&lt;br&gt;
Passport ZZ is not covered by this dataset, so Ninety will not guess.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the only behaviour that makes it safe to trust at all. A confident wrong number here costs someone money.&lt;/p&gt;
&lt;h3&gt;
  
  
  And then there is the desk
&lt;/h3&gt;

&lt;p&gt;When the sources genuinely conflict, Ninety refuses to classify the day and reports a range. This one is real: &lt;em&gt;does a layover that clears border control count as a day in the area?&lt;/em&gt; The border-crossing guidance says admission at a border is entry into the territory. The visa policy is framed around &lt;em&gt;staying&lt;/em&gt;, and a same-day transit is not a stay.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff9lbn02dq2ehc4pqh46j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff9lbn02dq2ehc4pqh46j.png" alt="The desk" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The second demo sits at &lt;strong&gt;exactly 90 of 90&lt;/strong&gt; with a landed transit booked for three days' time. Ninety will not pick a side, so the answer is a range and the headline stays honest.&lt;/p&gt;

&lt;p&gt;A person can then rule. That ruling is written as &lt;strong&gt;precedent&lt;/strong&gt; — typed, dated, attributed, and scoped to that presence kind from the date of the ruling — and every later calculation reads it. Rule it "counts" and the demo flips from &lt;em&gt;inside the limit&lt;/em&gt; to &lt;em&gt;crosses the limit on 6 October&lt;/em&gt;, because at 90 of 90 a single charged day is the whole question.&lt;/p&gt;

&lt;p&gt;Here it is flipping on camera — &lt;em&gt;inside the limit&lt;/em&gt; becomes &lt;em&gt;crosses the limit on 6 October&lt;/em&gt; the moment a person rules that the day counts:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj3vmf86qcupdmlwy8bhf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj3vmf86qcupdmlwy8bhf.png" alt="After the ruling" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The whole run, 48 seconds:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://raw.githubusercontent.com/PhiBao/ninety/main/docs/ninety-demo.mp4" rel="noopener noreferrer"&gt;https://raw.githubusercontent.com/PhiBao/ninety/main/docs/ninety-demo.mp4&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is a scripted recording against the deployed app, and the dataset is reset afterwards, so you will find the desk open.&lt;/p&gt;
&lt;h3&gt;
  
  
  And then there is the agent
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdnzb12khatux27yif2y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvdnzb12khatux27yif2y.png" alt="The agent" width="800" height="625"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Counting your own dates is easy. Deciding whether an airport transit counts is not — which is exactly the judgement the structured form was asking people to make before it would help them. So you can also just write it down:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Took the train into Sofia 1 to 11 February 2025, then flew to Paris 20 to 25 February 2025, then an 8 hour airport layover on 3 June 2026 where I never cleared immigration.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and the agent resolves it against the corpus, shows its confidence, and refuses the leg it is not sure about.&lt;/p&gt;

&lt;p&gt;Three rules, and the third is the one that matters:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It can only choose from a set it retrieved.&lt;/strong&gt; The territory list comes from &lt;code&gt;groq_query&lt;/code&gt; against Sanity Context. If a place is not in the corpus there is no option to select, and the output is "No place named" — which is what the layover in that example gets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It cannot do arithmetic.&lt;/strong&gt; Classification goes through TypeSafe's System One models, which return &lt;code&gt;{choice, probabilities, confidence}&lt;/code&gt; — a value from a set I defined, plus how sure it is. There is no channel through which a day count could travel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It says when it does not know.&lt;/strong&gt; Below 0.62 confidence the leg is marked uncertain rather than answered. In the screenshot above, &lt;em&gt;"flew to Paris 20 to 25 February"&lt;/em&gt; lands at 50% and is refused. I would not have got that from a chat completion, and in this domain it matters more than fluency.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Note the dependency count: the agent needs no chat-model SDK. The whole web app has five runtime dependencies and none of them is a model provider.&lt;/p&gt;


&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/PhiBao/ninety" rel="noopener noreferrer"&gt;https://github.com/PhiBao/ninety&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Monorepo. &lt;code&gt;studio/&lt;/code&gt; is a standalone Sanity Studio — 11 document types, 7 object types; &lt;code&gt;web/&lt;/code&gt; is Next.js 16.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;infra/      Sanity Blueprints — the CORS origin and a document function.
studio/     Sanity Studio. 11 document types, 7 object types.
web/
  src/lib/engine/     Pure TypeScript. No Sanity client, no model, no env vars.
  src/lib/agent/      Retrieve, classify, then hand off to the engine.
  src/lib/data/       The authored corpus, shared by the tests and the seed.
  src/lib/sanity/     GROQ, the snapshot loader, the Sanity Context client.
  scripts/            Seed, verification, evaluation, demo recording.
eval/RESULTS.md       The evaluation output, committed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The engine has no model in it
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;src/lib/engine/evaluate.ts&lt;/code&gt; makes &lt;strong&gt;no network call and imports no model&lt;/strong&gt;. It takes a snapshot of the rules and an itinerary and returns a verdict. Same inputs, same answer, every time — and you can check it by hand.&lt;/p&gt;

&lt;p&gt;There are 56 tests. They cover the things that are genuinely easy to get wrong: that the day of arrival counts and the day of departure does not, that a same-day visit is one day, that 1–29 February 2024 costs 28 days and the same calendar span in 2026 costs 27, that a pending application is not a permit, that Iceland is EEA but not Schengen, and that an air arrival into Sofia in February 2025 is not the same journey as a train.&lt;/p&gt;




&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;The prompt asks which AI-native IDE I used, which prompts worked, which did not, where the model got stuck, and how I course-corrected. I used &lt;strong&gt;OpenCode&lt;/strong&gt; with a large model. I built this in about a day, which I want to be honest about up front: it was fast because I made three decisions early that a slower build would have discovered too late.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I decided before writing code
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. The engine before the interface.&lt;/strong&gt; I wrote the date arithmetic and the verifier before a single component. Every bug after that was in plumbing, not in thinking, because the thinking was already pinned down by tests.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. No model in the verdict.&lt;/strong&gt; The agent may retrieve, classify and explain, but it may never produce a number. This is the whole technical thesis and it cost nothing to adopt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. The evaluation before the polish.&lt;/strong&gt; I built the case set in the first third. It immediately started catching me.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where it got stuck, honestly
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;The dataset looked empty and I nearly believed it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pnpm seed&lt;/code&gt; reported success, but every query returned zero documents. Not an error — a &lt;code&gt;200 OK&lt;/code&gt; with an empty result. That is the signature of a private dataset and it looks exactly like a broken schema. I burned time on the schema before noticing that an unauthenticated read returned &lt;code&gt;0&lt;/code&gt; while an authenticated read returned &lt;code&gt;36&lt;/code&gt;. The fix was a server-side read token, and I added &lt;code&gt;pnpm verify:data&lt;/code&gt; so that failure mode can never be silent again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GROQ is whitespace-sensitive, in a way that reads like a syntax error.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My snapshot query kept returning empty. The cause:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"allowances": *[_type == "allowance"]
{ "id": _id, ... }
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A newline between &lt;code&gt;]&lt;/code&gt; and &lt;code&gt;{&lt;/code&gt; makes GROQ read the braces as a block. The fix was to write every projection as &lt;code&gt;[...]{...}&lt;/code&gt; with no gap, and to leave a comment at the top of the file saying why so the next person does not "tidy" it back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;@sanity/client&lt;/code&gt; v8 silently no-oped my dispute update.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The adjudication wrote a precedent and returned &lt;code&gt;200&lt;/code&gt;. The dispute stayed open. A partial &lt;code&gt;patch&lt;/code&gt; on a document with nested typed objects did nothing, without throwing. I only found it because I checked the dataset afterwards instead of trusting the status code. Both the write and the reset now use a full document replace, and failures surface as a &lt;code&gt;503&lt;/code&gt; with the underlying message.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The module-level cache had no TTL.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My snapshot loader memoised forever, which is fine in a script and wrong on Vercel: a warm lambda served stale rules, so the demo appeared to ignore its own adjudication until the process happened to recycle. Thirty seconds fixed it. I would have shipped that bug if I had not recorded the demo against the &lt;em&gt;deployed&lt;/em&gt; site rather than locally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The evaluation caught my own arithmetic, six times.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I hand-wrote 23 expected answers so the runner could not grade itself. Six were wrong:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Five cases asked about a February 2025 rule "as of October 2026", so the rolling window had aged the trip out and the correct answer was &lt;code&gt;0&lt;/code&gt;. The engine was right; my expectation was wrong. I added a per-case &lt;code&gt;asOf&lt;/code&gt;, because a rule that applied in February 2025 has to be evaluated against February 2025.&lt;/li&gt;
&lt;li&gt;One case: I expected 29 days for 1–29 February 2024. The answer is 28, because the departure day is not charged.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The engine was right every single time. That is the argument for writing it before you write the prose about it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The two bugs I am most glad I found
&lt;/h3&gt;

&lt;p&gt;Both were invisible by construction, which is why they are worth writing down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A test that passed for the wrong reason.&lt;/strong&gt; My evaluation had a case for the Canary Islands, written against the territory code &lt;code&gt;es-canary&lt;/code&gt;. &lt;code&gt;es-canary&lt;/code&gt; is a &lt;em&gt;document key&lt;/em&gt;; the code is &lt;code&gt;XCI&lt;/code&gt;. The engine's response to an unknown territory is a warning and zero charged days — so the case passed, because the Canary days had been silently dropped rather than classified as a carve-out. The number was right. The reasoning was nonsense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A real-world error, in my own corpus.&lt;/strong&gt; I had modelled Iceland as a full member, using a shared helper that stamps the Schengen founding date onto every state. Iceland is in the EEA but has never joined the Schengen Area, so Reykjavík days were being charged against 90/180. That is a mistake a lot of tools make, and mine made it silently, in the one file I had been treating as ground truth.&lt;/p&gt;

&lt;p&gt;Neither was findable by reading the code. Both came from writing a check that looks at &lt;em&gt;what the engine actually did&lt;/em&gt; rather than at the number it returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pnpm check:codes      &lt;span class="c"&gt;# every territory code the product uses resolves&lt;/span&gt;
pnpm check:fallback   &lt;span class="c"&gt;# the bundled corpus and Sanity produce identical verdicts&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;check:codes&lt;/code&gt; also fails the build if any evaluation case contains a dropped stay. A case that passes because its input was ignored is worse than a case that fails.&lt;/p&gt;

&lt;h3&gt;
  
  
  The prompts that mattered
&lt;/h3&gt;

&lt;p&gt;The ones that worked were &lt;strong&gt;constraints, not descriptions&lt;/strong&gt;. "Do not let the model state a number" produced a better result than "make an agent". "Refuse visibly rather than guessing" produced visible refusals everywhere. "Every classification resolves to at least one source, or the day is undecided" produced a schema where that is structurally true.&lt;/p&gt;

&lt;p&gt;The one that did not work: I asked for a single-file schema at first. It produced a &lt;code&gt;territory.isSchengen: boolean&lt;/code&gt;, which is exactly the modelling error the product exists to catch. It took me restating the requirement as &lt;em&gt;"the same ten-day trip into Sofia counts overland and does not count by air in February 2025"&lt;/em&gt; before the schema stopped lying.&lt;/p&gt;




&lt;h2&gt;
  
  
  Thoughtfulness of the Schema
&lt;/h2&gt;

&lt;p&gt;Three decisions carry the weight. A fourth is the reason the whole thing works.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Membership is a band, not a flag.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;bg&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Bulgaria&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;code&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;BG&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;state&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;accessBands&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="na"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2000-01-01&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2024-12-31&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;counted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="na"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2024-12-31&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;to&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2025-03-31&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;counted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
   &lt;span class="na"&gt;modes&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;land&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sea&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;...},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="na"&gt;from&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;2025-03-31&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;counted&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...},&lt;/span&gt;
&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My first pass used a shared &lt;code&gt;FULL('2000-01-01')&lt;/code&gt; helper for every state, which is a lie for anyone who joined later — and, as the Iceland bug above shows, for anyone who never joined at all. Fixing it meant going and dating each accession properly: the 1995 founding group, the Nordic members, the 2004 and 2007 enlargements, Switzerland in 2008, Croatia in 2023, Bulgaria and Romania in 2024–25, and the four territories that are in the EEA, the EU, or neither but not in the area.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Carve-outs are first-class territories.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Canary Islands, Madeira, Åland, Svalbard, the French overseas departments, and Ireland, Cyprus and Iceland as external are each their own &lt;code&gt;territory&lt;/code&gt; document with their own bands — not a boolean on the parent state. That way the ledger can say &lt;em&gt;"Canary Islands (Spain) — outside the Schengen area despite being Spanish territory"&lt;/em&gt;, and name the exact place a day was spent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Ambiguity is data, not an exception.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;presenceRule&lt;/code&gt; can be marked &lt;code&gt;disputed&lt;/code&gt;. When it is, the engine will not classify a day of that kind, reports the count as a range, and routes the day to an adjudication. The ruling becomes a &lt;code&gt;precedent&lt;/code&gt; scoped to that presence kind and effective from the date of the ruling — which is why the September layover in the first demo correctly stays unresolved after a ruling made in October. Days already counted are not revisited.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The rules and the traveller are different kinds of thing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;allowance&lt;/code&gt;, &lt;code&gt;territory&lt;/code&gt;, &lt;code&gt;nationalityClass&lt;/code&gt;, &lt;code&gt;visaRegime&lt;/code&gt;, &lt;code&gt;permitExemption&lt;/code&gt;, &lt;code&gt;presenceRule&lt;/code&gt; and &lt;code&gt;source&lt;/code&gt; are reference data — effectively one document each, seeded under deterministic IDs. &lt;code&gt;itinerary&lt;/code&gt; and &lt;code&gt;trip&lt;/code&gt; are a person's record. &lt;code&gt;dispute&lt;/code&gt; and &lt;code&gt;precedent&lt;/code&gt; are decisions about the first group, made in light of the second. Keeping those three layers apart is what let me seed, test, snapshot and reset each independently — and what lets the app fall back to a bundled copy of the corpus and still produce byte-identical verdicts when Sanity is unreachable.&lt;/p&gt;




&lt;h2&gt;
  
  
  The infrastructure is declared too
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;infra/sanity.blueprint.ts&lt;/code&gt; holds the CORS origin the app reads through and a document function, and &lt;code&gt;blueprints plan&lt;/code&gt; shows the diff before anything is applied.&lt;/p&gt;

&lt;p&gt;The function is the interesting part. Ninety's claim is that a ruling is not a message — it is a typed, dated, scoped precedent that every later calculation reads. The desk writes that precedent. But a dispute can also be adjudicated by editing the document in the Studio, and that path can leave a dispute marked &lt;code&gt;adjudicated&lt;/code&gt; with nothing behind it.&lt;/p&gt;

&lt;p&gt;Which is exactly the bug this build shipped. A partial patch wrote the precedent and left the dispute open, and the product looked correct until someone checked the dataset. So the invariant now lives in the data layer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;any dispute marked &lt;code&gt;adjudicated&lt;/code&gt; has a &lt;code&gt;precedent&lt;/code&gt; behind it, whoever adjudicated it&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Verified by hand, twice: a Studio-style patch that leaves the broken state is repaired within about thirteen seconds, the precedent carries the presence rule's own two sources so its citations still resolve, and a second ruling supersedes the first rather than overwriting it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Three ways to fail while building a guard
&lt;/h3&gt;

&lt;p&gt;Every one of these logged success and did nothing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;client.patch()&lt;/code&gt; is lazy in &lt;code&gt;@sanity/client&lt;/code&gt; v8.&lt;/strong&gt; The handler awaited it, logged &lt;em&gt;"attached precedent to dispute"&lt;/em&gt;, and changed nothing. Writes now go through &lt;code&gt;mutate()&lt;/code&gt;, which takes an array and returns a transaction result — either the write happened or the call threw.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A transaction id derived from the document id looks like idempotency and is a trap.&lt;/strong&gt; Sanity remembers transaction ids permanently, so the second adjudication of the same dispute came back &lt;code&gt;transactionAlreadyExistsError&lt;/code&gt; and the function failed forever after, silently. Convergence now comes from &lt;code&gt;createOrReplace&lt;/code&gt; against a date-keyed precedent id, which is naturally idempotent; the transaction id only has to be unique.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The guard read a field the document does not have.&lt;/strong&gt; It looked for a denormalised &lt;code&gt;presenceKind&lt;/code&gt;, found nothing, and declined to act — on a dispute whose &lt;code&gt;subjectKind&lt;/code&gt; was very plainly &lt;code&gt;presence_kind&lt;/code&gt;. It now resolves the kind through the reference.&lt;/p&gt;

&lt;p&gt;That last one is the general lesson, and it is why I think a guard that refuses because it looked in the wrong place is worse than no guard at all: it converts a loud failure into a quiet one. Every check in this repo earns its place by looking at &lt;strong&gt;what the system actually did&lt;/strong&gt; rather than at whether it reported success.&lt;/p&gt;




&lt;h2&gt;
  
  
  Did the structured content actually matter?
&lt;/h2&gt;

&lt;p&gt;The challenge says: &lt;em&gt;"If a keyword search would have gotten you the same answer, aim higher."&lt;/em&gt; So I measured it.&lt;/p&gt;

&lt;p&gt;Two systems, the same corpus, 23 adversarial histories with hand-computed expectations. Neither arm uses a language model, so anyone can rerun it with &lt;code&gt;pnpm eval&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measure&lt;/th&gt;
&lt;th&gt;Structured (GROQ + engine)&lt;/th&gt;
&lt;th&gt;Keyword search (BM25, same rules)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Correct answer&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23/23&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1/23&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Produced a day count&lt;/td&gt;
&lt;td&gt;21/23&lt;/td&gt;
&lt;td&gt;0/23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Named the exact breach date&lt;/td&gt;
&lt;td&gt;1/1&lt;/td&gt;
&lt;td&gt;0/1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refused instead of guessing&lt;/td&gt;
&lt;td&gt;2/2&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Full results, including why each case is hard: &lt;a href="https://github.com/PhiBao/ninety/blob/main/web/eval/RESULTS.md" rel="noopener noreferrer"&gt;&lt;code&gt;web/eval/RESULTS.md&lt;/code&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The keyword arm is not handicapped by a weak index. It sees the same territory records, the same presence rules, the same permit exemptions, flattened into prose, and scores BM25 over them. It returns the right &lt;em&gt;sentences&lt;/em&gt; most of the time. It just cannot produce a number, because:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The corpus holds the rules, and never held the traveller.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Asking "am I still legal" is not a retrieval problem. It is a join between the rule and eleven date ranges, two of which fall in carve-outs and one of which nobody has ruled on, evaluated day by day across a rolling window. Retrieval finds the paragraphs. It cannot join them to a calendar.&lt;/p&gt;




&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Project ID:&lt;/strong&gt; &lt;code&gt;jvgi63fz&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dataset:&lt;/strong&gt; &lt;code&gt;production&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contents:&lt;/strong&gt; 5 sources, 1 allowance, 36 territories with date-banded access, 3 nationality classes, 3 visa regimes, 5 permit exemptions, 6 presence rules, 2 demo itineraries with 17 trips, 1 open dispute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproduce it:&lt;/strong&gt; &lt;code&gt;pnpm --dir web seed &amp;amp;&amp;amp; pnpm --dir web verify:data&lt;/code&gt;
&lt;strong&gt;Infrastructure:&lt;/strong&gt; &lt;code&gt;pnpm --dir infra bp:plan &amp;amp;&amp;amp; pnpm --dir infra bp:deploy&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every rule in the dataset carries a &lt;code&gt;sourceRef&lt;/code&gt; pointing at a real document with a publisher, a URL and a retrieval date. Clicking any day in the ledger shows its authorities. Nothing in Ninety's reasoning is unattributed.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this is not
&lt;/h2&gt;

&lt;p&gt;Not legal advice, and not a complete immigration reference. It covers the Schengen short-stay allowance for a handful of passport classes and refuses outside that coverage. The rule corpus is deliberately narrow and dated rather than broad and approximate, because a wrong answer in this domain is not an inconvenience.&lt;/p&gt;

&lt;p&gt;If I had another day I would widen the territory set, add the UK 180-day visitor rule and the US admission-parity rule on the same generic engine, ingest the EU source documents as a Knowledge Base so the agent can quote the paragraph rather than the title, and build a real Sanity App SDK app for the desk so adjudication lives inside the Dashboard instead of a separate page.&lt;/p&gt;




&lt;p&gt;Thanks to the Sanity team for the challenge, and to the DEV community for the review. 🤖&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
