<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Syed Jawad</title>
    <description>The latest articles on DEV Community by Syed Jawad (@syed_jawad_ead7feefa5b789).</description>
    <link>https://dev.to/syed_jawad_ead7feefa5b789</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3816599%2F459316a2-57f3-4da4-9fb7-1d297fee3143.png</url>
      <title>DEV Community: Syed Jawad</title>
      <link>https://dev.to/syed_jawad_ead7feefa5b789</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/syed_jawad_ead7feefa5b789"/>
    <language>en</language>
    <item>
      <title>I Built a Roman Urdu Stock Book for My Friend's Grocery Store</title>
      <dc:creator>Syed Jawad</dc:creator>
      <pubDate>Sun, 04 Oct 2026 16:38:24 +0000</pubDate>
      <link>https://dev.to/syed_jawad_ead7feefa5b789/i-built-a-roman-urdu-stock-book-for-my-friends-grocery-store-3bba</link>
      <guid>https://dev.to/syed_jawad_ead7feefa5b789/i-built-a-roman-urdu-stock-book-for-my-friends-grocery-store-3bba</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Hamza and I grew up together. Today he runs Hamza Mart, a grocery store, and his stock lives in a paper register and in his memory, written and thought in a mix of Urdu, Roman Urdu, and English. I wanted to turn those same notes into balances he could check.&lt;/p&gt;

&lt;p&gt;Stock Diary turns “20 carton basmati aaye, 5 packet haldi nikle” (“20 cartons of basmati came in, 5 packets of turmeric went out”) into two proposed stock movements. Each shows the product, quantity, unit, direction, and the balance before and after. He can correct the proposal before confirming it. Nothing is saved until he confirms.&lt;/p&gt;

&lt;p&gt;That confirmation matters. A plausible sentence is not a reliable stock entry if the unit is wrong. If Hamza writes “5 bori laal mirch” but chilli is counted in packets, Stock Diary stops the entry until he picks the right unit. It also stops an outgoing movement of five when only three are left. The model helps read the sentence; plain Python checks the numbers and posts the confirmed movement to a SQLite ledger.&lt;/p&gt;

&lt;p&gt;He can also ask “kis cheez ka stock kam hai?” (“What is running low?”) or “kal kya mangwana hai?” (“What should I order tomorrow?”). Gemma only works out which question he is asking; Python answers from the ledger. The reorder list gives the minimum needed to clear each low-stock alert, not a forecast. Manual entries and new products cover the everyday work, with a full Urdu interface (right to left) and a layout that fits a phone.&lt;/p&gt;

&lt;p&gt;I built this from what I know of how Hamza works, not from a session with him. He has not used it yet; I'm showing it to him this week. All stock figures in the public demo are sample data, not Hamza Mart's real stock.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Try the &lt;a href="https://stock-diary-lscy.onrender.com" rel="noopener noreferrer"&gt;live Stock Diary demo&lt;/a&gt;: enter the sentence above, check the proposed balances, then confirm and look at the updated stock. Try “5 bori laal mirch aayi” to see the unit warning. Render's free plan can take about a minute to wake after inactivity. The video shows hosted and local Gemma running the same app code.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/s6pBkxYO72M" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Each browser gets a separate sandbox with a reset button. AI requests are limited to five per minute and 30 per day per visitor, with extra per-IP and global daily limits. These keep usage low; they don't guarantee the demo stays inside Cloudflare's free allowance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/syedjawad11/stock-diary" rel="noopener noreferrer"&gt;repository&lt;/a&gt; has the app, tests, and setup instructions. The demo catalogue and entries are synthetic. You can also run Gemma locally with Ollama, without cloud keys.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/syedjawad11" rel="noopener noreferrer"&gt;
        syedjawad11
      &lt;/a&gt; / &lt;a href="https://github.com/syedjawad11/stock-diary" rel="noopener noreferrer"&gt;
        stock-diary
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Bilingual (Urdu/English) stock diary for a small wholesaler, powered by Gemma. DEV Hacktoberfest Weekend Challenge 2026.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Stock Diary&lt;/h1&gt;
&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Try it live&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Watch the &lt;a href="https://youtu.be/s6pBkxYO72M" rel="nofollow noopener noreferrer"&gt;2-minute demo video&lt;/a&gt; (live on Cloudflare and fully local with Ollama), or try the &lt;a href="https://stock-diary-lscy.onrender.com" rel="nofollow noopener noreferrer"&gt;live demo&lt;/a&gt;, where each visitor gets their own sandbox seeded with sample stock. It runs on Render's free plan, so the first load after idle can take about a minute. The model is Gemma 4 (26B A4B) on Cloudflare Workers AI, with limits of 5 model calls per minute and 30 per day per visitor.&lt;/p&gt;
&lt;p&gt;A bilingual (Urdu, Roman Urdu, English) stock book for Hamza Mart, a small
grocery store run by Jawad's friend Hamza. Type a line like "20 carton basmati aaye, 5 laal
mirch nikle"; Gemma parses it into proposed stock movements, you confirm, and
plain Python keeps the ledger, running balances, and low-stock flags. Ask
"kis cheez ka stock kam hai?" and get a plain-language answer, filled in from
the real numbers — Gemma only…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/syedjawad11/stock-diary" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;I had one Sunday inside the challenge window. My first idea was an invoice matcher, but I dropped it before writing code. It depended on reading bill images, and it was further from a problem I knew Hamza had. A stock book gave me a smaller and more useful starting point: one line of messy language, a review step, and a balance that must be correct.&lt;/p&gt;

&lt;p&gt;My rule for the architecture was: &lt;strong&gt;Gemma reads, Python counts.&lt;/strong&gt; Google's Gemma 4 is an open-weight model and the only AI in the running app. It turns mixed-language text into proposed product movements against the shop's catalogue. For questions, it selects an intent such as &lt;code&gt;low_stock&lt;/code&gt;, &lt;code&gt;restock&lt;/code&gt;, &lt;code&gt;balance&lt;/code&gt;, or &lt;code&gt;today&lt;/code&gt;. Python fills fixed bilingual answer templates with ledger values. Gemma never writes a number to the ledger. Quantities are integers, and the ledger arithmetic is plain Python with pytest tests.&lt;/p&gt;

&lt;p&gt;Both model routes go through &lt;code&gt;app/model_adapter.py&lt;/code&gt; and are selected with &lt;code&gt;MODEL_BACKEND&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Backend&lt;/th&gt;
&lt;th&gt;Model and place it runs&lt;/th&gt;
&lt;th&gt;Result on my 12 test sentences&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hosted demo&lt;/td&gt;
&lt;td&gt;Gemma 4 26B A4B on Cloudflare Workers AI&lt;/td&gt;
&lt;td&gt;12/12, about 1–2 seconds each&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local&lt;/td&gt;
&lt;td&gt;Gemma 4 e4b through Ollama on my 16 GB MacBook&lt;/td&gt;
&lt;td&gt;12/12, about 4–9 seconds each; first call about 13 seconds while it loads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The 12 sentences covered Roman Urdu, Urdu script, English, and mixed input. One deliberately said “chawal 10 aaye” when the catalogue contained two kinds of rice; a correct response had to leave the product unresolved instead of guessing. This is a useful check for the language boundary, but it is not a measure of accuracy across a working shop.&lt;/p&gt;

&lt;p&gt;I hit a specific problem with the hosted model: multi-line entries came back empty. Cloudflare's Gemma 4 had thinking mode enabled by default, and its reasoning used the available token budget before it returned the structured answer. Setting this in the request fixed the empty output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"chat_template_kwargs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"enable_thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After that change, the hosted model completed all 12 test cases in roughly one to two seconds each. A parse used about nine Cloudflare neurons in this check; the free allowance is 10,000 neurons per day. That is why the public demo has visitor and global limits.&lt;/p&gt;

&lt;p&gt;The rest of the stack is FastAPI, Pydantic, SQLite, uvicorn, httpx, pytest, and a vanilla JavaScript frontend. Render serves the backend and frontend; I triggered deploys through the Render API. There are 42 pytest tests, and GitHub Actions runs them on Python 3.9 and 3.12 on every push.&lt;/p&gt;

&lt;p&gt;The checks caught more than parsing mistakes. A review of the code found that editing a quantity could bypass the unit warning, that an insufficient-stock error could hide a unit mismatch, and that resetting a sandbox during a parse could return a server error. I fixed the frontend unit-warning logic and added API regression tests for the mismatch response and for a reset during a parse. A browser click-through at phone and desktop widths caught two frontend issues, which I also fixed. The demo video was recorded with Playwright, and I generated its narration with ElevenLabs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;A small grocery store can't plan around a per-call model bill. Open weights give Stock Diary a local option: I ran Gemma 4 through Ollama on my 16 GB MacBook with no cloud keys at all. That shows a working alternative to paying a provider for every entry. It doesn't yet tell me what hardware Hamza would need, and that's part of the handover.&lt;/p&gt;

&lt;p&gt;His stock is his business information. The public demo sends requests to Cloudflare so anyone can try it in a browser, but in a local setup every entry stays on the machine running the app. Both routes go through the same model adapter and the same stock workflow, so switching &lt;code&gt;MODEL_BACKEND&lt;/code&gt; changes the inference provider without touching the ledger. If a host changes its prices or stops serving the model, the local path still works.&lt;/p&gt;

&lt;p&gt;The code is MIT-licensed, so another shop can adapt the catalogue, the product aliases and the interface for its own stock.&lt;/p&gt;

&lt;p&gt;Open weights also let me test the language Hamza actually uses: Roman Urdu mixed with English and Urdu script. A rigid stock form would make him translate his own notes into the software's categories before recording them. Here the model proposes the interpretation, but the review screen and Python checks keep that flexibility from silently changing a balance.&lt;/p&gt;

&lt;p&gt;There are limits I want to be clear about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Hamza has not used Stock Diary yet. There has been no testing with his real stock.&lt;/li&gt;
&lt;li&gt;Model accuracy was checked on 12 sentences, not at shop scale.&lt;/li&gt;
&lt;li&gt;The free Render service has cold starts. Demo call quotas can reset if the server restarts.&lt;/li&gt;
&lt;li&gt;There is no voice input, bill photo input, or login. The public version uses a sandbox for each visitor.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;I used Claude Code as the build orchestrator for the plan, architecture, shared &lt;code&gt;API.md&lt;/code&gt; contract, integration, tests, and deploys. OpenAI Codex ran GPT models as workers on defined files for the ledger, restock intent, and frontend, then as a reviewer. Three workers built those parts in parallel against the written API contract. Claude Code and Codex were build tools only; the running app calls Gemma.&lt;/p&gt;

&lt;p&gt;I did not record a DevRelay session. The repository keeps the &lt;a href="https://github.com/syedjawad11/stock-diary/tree/main/.agent" rel="noopener noreferrer"&gt;task briefs and review notes&lt;/a&gt; in &lt;code&gt;.agent/tasks/&lt;/code&gt; and &lt;code&gt;.agent/reviews/&lt;/code&gt;, including the review that found the three bugs before deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma:&lt;/strong&gt; Gemma is the only runtime AI and runs both locally through Ollama and on Cloudflare Workers AI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Render:&lt;/strong&gt; Render hosts the FastAPI backend and frontend that call Gemma for the public demo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of GitHub Copilot:&lt;/strong&gt; GitHub Actions runs all 42 tests on Python 3.9 and 3.12 on every push, so a ledger regression shows up on the commit that caused it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of ElevenLabs:&lt;/strong&gt; ElevenLabs generated the narration for the demo video.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This week I'm handing Stock Diary to Hamza to see how it fits a real day at Hamza Mart, and I'll add what he says to the comments. If he wants it, voice input in Urdu is the next feature I would try. If you have built a tool for a shopkeeper, how did you handle notes that switch between scripts and product names?&lt;/p&gt;

</description>
      <category>hf26challenge</category>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>ai</category>
    </item>
    <item>
      <title>Museum of Unfinished Futures: an AI clerk proposes, a human curator decides</title>
      <dc:creator>Syed Jawad</dc:creator>
      <pubDate>Sat, 03 Oct 2026 20:58:06 +0000</pubDate>
      <link>https://dev.to/syed_jawad_ead7feefa5b789/museum-of-unfinished-futures-an-ai-clerk-proposes-a-human-curator-decides-2dki</link>
      <guid>https://dev.to/syed_jawad_ead7feefa5b789/museum-of-unfinished-futures-an-ai-clerk-proposes-a-human-curator-decides-2dki</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge&lt;/a&gt;, Path Two: Vibe-Code Something Strange&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;The Museum of Unfinished Futures is a small museum of made-up inventions from futures that never happened.&lt;/p&gt;

&lt;p&gt;There is a vending machine that sells extra Mondays. There is an umbrella that remembers every storm, and a telephone for calling roads not taken. A toaster prints notes from your future self. A switchboard reconnects conversations that ended too soon, at the exact word where they stopped. A kettle brews the weather of past visits.&lt;/p&gt;

&lt;p&gt;Each one stands in a case with a blueprint drawing, an accession note (where the museum "got it"), a plaque and a choice. The switchboard's accession note reads: &lt;em&gt;"Removed intact from the night exchange at Counterfactual Relay Station Four."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Seven exhibits hang in three wings, each wing lit in its own colour:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Civic Time Expansion Era&lt;/strong&gt; (amber): the vending machine and the toaster.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Counterfactual Communications Boom&lt;/strong&gt; (cyan): the telephone and the switchboard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Domestic Weather Memory Era&lt;/strong&gt; (violet): the umbrella, the kettle and the doormat, the first exhibit the Acquisitions Clerk drafted (more on that below).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A visit goes like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;You stand at an exhibit&lt;/strong&gt; and read its plaque.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You choose.&lt;/strong&gt; The switchboard offers three choices: answer the line that is still lit, pull every cord at once, or plug a cord into the blank jack. The other exhibits offer two.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You reach an ending.&lt;/strong&gt; Each choice has its own ending, with short consequence tags beneath it, such as &lt;code&gt;sentence-finished&lt;/code&gt; or &lt;code&gt;apology-not-required&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A door opens.&lt;/strong&gt; Every ending says "Continue to →" and sends you to a different exhibit, sometimes in another wing, &lt;em&gt;because of what you chose&lt;/em&gt;. No ending leads back to its own exhibit. The museum loops on purpose and has no exit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You take a ticket.&lt;/strong&gt; "Print your ticket" turns the endings you reached into a short, personal "unfinished future". The whole ticket lives in the page address (&lt;code&gt;/your-future?trace=…&lt;/code&gt;). There are no accounts and nothing is stored. The same link always shows the same ticket, so you can share it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One real ticket, from a walk past the toaster and the kettle, begins: &lt;em&gt;"Your walk is over. What follows is only what the rooms noticed."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There is no chatbot and no AI for visitors. All of the strangeness is written content stored in Sanity, and the content model decides where a visitor goes next.&lt;/p&gt;

&lt;p&gt;Behind the scenes, new exhibits can arrive through the &lt;strong&gt;Acquisitions Clerk&lt;/strong&gt;, an AI agent that works inside Sanity's official Workflows next to a human curator. A curator gives it a one-line brief and a wing. The Clerk drafts an exhibit with two endings, starts an &lt;em&gt;Exhibit review&lt;/em&gt; run and submits it. The curator reads it in Studio. If they send it back with a note, the Clerk reads the note, revises and resubmits. The Clerk can't approve, put on display or publish; only a person can. It runs on Sanity's free monthly AI credits or on a model running locally through Ollama, so it costs nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live site, no login needed:&lt;/strong&gt; &lt;a href="https://museum-of-unfinished-futures.netlify.app" rel="noopener noreferrer"&gt;https://museum-of-unfinished-futures.netlify.app&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Try it in two minutes
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Enter the hall.&lt;/strong&gt; Three wings, seven cases. The newest one, &lt;em&gt;The Doormat That Knows Who Is Coming&lt;/em&gt; in the violet wing, was drafted by the Acquisitions Clerk and approved by a human curator.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open the doormat&lt;/strong&gt; and choose: wipe your feet and let the house know you, or step over the threshold without being known.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Follow the door.&lt;/strong&gt; The ending tells you where to go next, and "Continue to →" takes you there, often into another wing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make one more choice&lt;/strong&gt;, at the switchboard if you can find it. It is the only exhibit with three.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Print your ticket.&lt;/strong&gt; The endings you reached become a short "unfinished future". Copy the address and open it in another browser: the same ticket comes back, because the whole ticket lives in the link.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fewgkifjuvxcvrpqz4otn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fewgkifjuvxcvrpqz4otn.png" alt="The hall: three wings and seven cases, the doormat in the violet wing"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5381x72mulj21h3uk8et.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5381x72mulj21h3uk8et.png" alt="The switchboard after a choice: the ending, its tags and the door to the umbrella"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F04kqvnowkfd3wk3jz52z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F04kqvnowkfd3wk3jz52z.png" alt="A visitor's ticket that quotes the Clerk's exhibit"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Behind the glass, in Studio:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc53skn3plz5xmlvoy5gh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc53skn3plz5xmlvoy5gh.png" alt="Studio's Exhibit review card after the curator's note: Drafting, round 2, with the reason on the card"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fachqgmatu6b8rt12uxvr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fachqgmatu6b8rt12uxvr.png" alt="Workflow history: the Clerk's moves under its robot id, the curator's under their own name"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2bjfgfi8shaay22nfi6k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2bjfgfi8shaay22nfi6k.png" alt="The Clerk's exhibit live on the site"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More screenshots (phone, wing pages, plates at full size) are in &lt;a href="https://github.com/syedjawad11/museum-of-unfinished-futures/tree/main/evidence" rel="noopener noreferrer"&gt;&lt;code&gt;evidence/&lt;/code&gt;&lt;/a&gt; in the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/syedjawad11" rel="noopener noreferrer"&gt;
        syedjawad11
      &lt;/a&gt; / &lt;a href="https://github.com/syedjawad11/museum-of-unfinished-futures" rel="noopener noreferrer"&gt;
        museum-of-unfinished-futures
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A small museum of fictional inventions from futures that never happened. Built with Next.js and Sanity for the Sanity Challenge on DEV.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Museum of Unfinished Futures&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;Next.js visitor experience and embedded Sanity Studio for a fictional museum of futures that never happened.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What Exists&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;A responsive home gallery at &lt;code&gt;/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A visitor route at &lt;code&gt;/exhibits/[slug]&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A typed Sanity Content Lake repository backed by project &lt;code&gt;wa27n68e&lt;/code&gt;, dataset &lt;code&gt;production_1&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A domain resolver that maps a selected artifact choice to its linked outcome.&lt;/li&gt;
&lt;li&gt;Sanity schemas for artifacts, eras, and outcomes.&lt;/li&gt;
&lt;li&gt;An embedded Studio at &lt;code&gt;/studio&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;An honest empty-gallery state for datasets with no published artifacts.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The local fixture &lt;strong&gt;The Vending Machine That Sells Extra Mondays&lt;/strong&gt; remains test/demo recovery data only. Public visitor routes do not fall back to it.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Local Commands&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Node 26 is pinned in &lt;code&gt;.nvmrc&lt;/code&gt; (the build's &lt;code&gt;--no-experimental-webstorage&lt;/code&gt; flag is rejected by Node 20). &lt;code&gt;fnm&lt;/code&gt;/&lt;code&gt;nvm&lt;/code&gt; pick it up automatically. &lt;code&gt;npm run typecheck&lt;/code&gt; relies on route types that &lt;code&gt;next build&lt;/code&gt; generates, so run it after a build…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/syedjawad11/museum-of-unfinished-futures" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h3&gt;
  
  
  Running it locally
&lt;/h3&gt;

&lt;p&gt;The app reads published content from a public Sanity dataset, so you don't need a token.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install
&lt;/span&gt;npm run dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open &lt;code&gt;http://localhost:3000&lt;/code&gt;. The embedded Studio is at &lt;code&gt;/studio&lt;/code&gt;. It asks you to sign in to Sanity, and your origin must be allowed in the project's CORS settings.&lt;/p&gt;

&lt;p&gt;The checks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm run &lt;span class="nb"&gt;test&lt;/span&gt;:unit
npm run typecheck
npm run lint
npm run build
npm run sanity:check
npx sanity schemas validate
npx playwright &lt;span class="nb"&gt;install &lt;/span&gt;chromium   &lt;span class="c"&gt;# once per machine&lt;/span&gt;
npm run &lt;span class="nb"&gt;test&lt;/span&gt;:e2e                  &lt;span class="c"&gt;# builds, serves on 127.0.0.1:3100, runs the browser tests&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Environment variables
&lt;/h3&gt;

&lt;p&gt;The public defaults are checked in. Copy &lt;code&gt;.env.example&lt;/code&gt; to &lt;code&gt;.env.local&lt;/code&gt; only if you want to point at a different project. Names only:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;NEXT_PUBLIC_SANITY_PROJECT_ID&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;NEXT_PUBLIC_SANITY_DATASET&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;NEXT_PUBLIC_SANITY_API_VERSION&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;No passwords, tokens or verification codes belong in these, or anywhere in the repository.&lt;/p&gt;

&lt;h3&gt;
  
  
  Credits
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fonts:&lt;/strong&gt; Fraunces (© The Fraunces Project Authors) and IBM Plex Mono, both under the SIL Open Font License 1.1. The licence files sit next to the fonts in &lt;code&gt;src/fonts/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blueprint plates:&lt;/strong&gt; six original SVG drawings, one per curator-made exhibit, made for this project. The Clerk's doormat shows a "Plate pending" card for now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exhibit text:&lt;/strong&gt; original fiction.&lt;/li&gt;
&lt;li&gt;Built with Next.js 16 and Sanity (Studio, Content Lake, and the Sanity Workflows early-access packages, version 0.35.0).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;This is the honest version. Everything here comes from the project's build log, task packets and saved evidence, all of which are in the repository.&lt;/p&gt;

&lt;h3&gt;
  
  
  Who did the work
&lt;/h3&gt;

&lt;p&gt;One founder, directing AI agents, in two phases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase one (September 20–21): Hermes directing OpenAI Codex.&lt;/strong&gt; An earlier setup, orchestrated by Hermes, used OpenAI Codex &lt;code&gt;gpt-5.5&lt;/code&gt; to build the first version and &lt;code&gt;gpt-5.6-sol&lt;/code&gt; as a fresh-context, read-only reviewer. It produced the first three exhibits, the live Sanity connection, the custom curator review flow, and a private draft deploy on Netlify's free plan.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase two (from the evening of September 21): Claude Code as orchestrator.&lt;/strong&gt; Claude Code (Claude Opus 5.5) took over. It did not write product code. For each job it wrote a &lt;strong&gt;task packet&lt;/strong&gt;: the objective, which files the worker may and may not touch, what to read first, the exact commands that decide pass or fail, and a time limit. It handed the packet to a worker, inspected the result, and &lt;strong&gt;re-ran the checks itself&lt;/strong&gt;. A worker's summary was never treated as proof. The workers were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI Codex &lt;code&gt;gpt-5.5&lt;/code&gt;&lt;/strong&gt; for logic, scripts, schema and tests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Sonnet 5&lt;/strong&gt; agents for the interface, the blueprint plates and the browser tests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Claude Opus writer&lt;/strong&gt; for the newer fiction, the endings' tags and the ticket's language.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-family reviewers.&lt;/strong&gt; Codex &lt;code&gt;gpt-5.6-sol&lt;/code&gt; reviewed Claude's work, and a Claude reviewer (Sonnet or Opus) reviewed Codex's. The family that built something never reviewed it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The final sprint (October 1–4) on a new machine.&lt;/strong&gt; Codex wasn't set up on the founder's new MacBook, so the orchestrator built the Acquisitions Clerk itself, under the same rules: tests first, a deliberate break for every new test, and every check re-run before a commit.&lt;/p&gt;

&lt;p&gt;Deploys, spending and writes to the live dataset waited for the founder's approval. No money was spent: the Netlify account stayed on the free plan with 0 credits used and no payment method. The founder ruled out paid AI keys entirely, so the Clerk writes with Sanity's free monthly AI credits or with a local model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prompts that worked
&lt;/h3&gt;

&lt;p&gt;The packets were the prompts. Two patterns earned their place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Tell the worker what already exists.&lt;/strong&gt; Before writing the wings packet (T-008), the orchestrator read the code and found that the era queries it needed had shipped, with tests, two tasks earlier. So the packet said:&lt;/p&gt;

&lt;blockquote&gt;
&lt;h3&gt;
  
  
  What already exists — reuse it, do NOT rebuild it
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;sanityExhibitRepository.listEras()&lt;/code&gt; and &lt;code&gt;.getEraBySlug(slug)&lt;/code&gt; in
&lt;code&gt;src/content/sanity-repository.ts&lt;/code&gt; — the GROQ already resolves each era with
its &lt;code&gt;exhibits[]&lt;/code&gt; (full artifact projection, including &lt;code&gt;image&lt;/code&gt; and &lt;code&gt;choices&lt;/code&gt;),
ordered by &lt;code&gt;title asc&lt;/code&gt;. T-003 shipped these; they are tested.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same packet looked ahead: "&lt;strong&gt;nothing may assume a wing has exactly one exhibit&lt;/strong&gt;. A wing with zero exhibits must render an honest empty line, not crash." When three more exhibits arrived, the wing pages needed no change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Say what the tests must prove, and make the worker show a test can fail.&lt;/strong&gt; From the Sanity Workflows packet (T-013a):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Tests first: using the engine's in-memory test bench (testing.md), write tests that drive: the happy path drafting→…→on-display; request-changes with a reason returns to drafting and the reason is stored; request-changes WITHOUT a reason is refused; approve/put-on-display are not available from drafting (no skipping); publishing permitted only in &lt;code&gt;approved&lt;/code&gt; if the engine exposes that verdict. Run RED first (definition stub), then GREEN. Mutation proof: remove the reason requirement, show the refusal test fails, restore.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That last sentence became standard. The story below explains why.&lt;/p&gt;

&lt;h3&gt;
  
  
  What went wrong, and what we changed
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;The sandbox couldn't install anything.&lt;/strong&gt; Codex's builder sandbox has no network. On the first run it couldn't even install the test runner (&lt;code&gt;npm error code ENOTFOUND … registry.npmjs.org&lt;/code&gt;). The supervising session installed it with normal network access. The first pinned Vitest version (3.2.4) showed a critical security advisory, so it went to 5.0.1. From then on, installs happened outside the worker sandboxes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Turbopack → Webpack.&lt;/strong&gt; The first production build failed under the default Turbopack bundler, because the sandbox wouldn't let a helper process open a port. The Next.js docs list &lt;code&gt;next build --webpack&lt;/code&gt; as supported, so the build moved to Webpack. That hit &lt;code&gt;Could not parse output from TypeScript's --showConfig&lt;/code&gt;, which a documented setting (&lt;code&gt;experimental.useTypeScriptCli: false&lt;/code&gt;) fixed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A server-side import broke the Studio build.&lt;/strong&gt; After the embedded Studio was added, the production build failed. A server component imported &lt;code&gt;sanity.config.ts&lt;/code&gt;, so Webpack picked the server build of a library Sanity depends on (SWR), and that build has no default export. The error trace named &lt;code&gt;sanity.config.ts&lt;/code&gt; via the Studio page. Loading the Studio and its config through a client component fixed the root cause.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The publish button skipped validation.&lt;/strong&gt; The custom publish action checked that a curator had approved the exact draft, but it no longer waited for Sanity's schema validation. A fresh-context reviewer flagged this as a blocking data-integrity issue. Publishing is now blocked while validation is running, when validation is out of date for the current draft, or when there are errors. A second reviewer passed the fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Smaller ones.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A review found that malformed published exhibits would be quietly dropped, so a broken dataset looked empty. It now fails loudly.&lt;/li&gt;
&lt;li&gt;Lint runs timed out because ESLint was crawling Netlify's generated &lt;code&gt;.netlify/&lt;/code&gt; folder. One ignore rule fixed it.&lt;/li&gt;
&lt;li&gt;A site-wide loading page made a missing exhibit return HTTP 200 instead of 404. The Next.js 16 docs explain why: a root &lt;code&gt;loading.tsx&lt;/code&gt; starts streaming before &lt;code&gt;notFound()&lt;/code&gt; runs. We dropped the loading page and kept the real 404.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Tests that couldn't fail, twice in one day, from two model families.&lt;/strong&gt; In the wings task, a Codex reviewer found that a new browser test, written by a Claude agent, checked a wing's subtitle instead of its summary. It would have passed with every summary missing. The same day, a Claude reviewer found this in a Codex-written test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;([].&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;formatConsequenceTag&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;([])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;.map&lt;/code&gt; never calls its function on an empty array, so the test passes even if the function is deleted. A green rerun can't prove a repair like that, so we deliberately broke one wing summary in the test data, watched the test fail, and restored it. Since then, every packet that includes tests asks the worker to prove that at least one assertion can fail. The rule kept catching things: the first Workflows attempt was rejected because its tests only checked the definition's shape, and a deliberate break in a publish script showed that one of its safety checks was dead code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A test run that passed without running.&lt;/strong&gt; One end-to-end run reported success while printing &lt;code&gt;Error: http://127.0.0.1:3100 is already used&lt;/code&gt;. A worker's screenshot server had outlived its task and kept the port, so the suite stopped before running a single test. Our rule now: never accept an exit code without also finding the literal "N passed" line. Separately, undoing that deliberate test breakage with &lt;code&gt;git checkout --&lt;/code&gt; also wiped a worker's uncommitted edits to the same file. We recovered them from a backup, and the rule now is to back up first and never use &lt;code&gt;git checkout --&lt;/code&gt; on uncommitted work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Your trace carries carrying…"&lt;/strong&gt; Every check was green, and a browser test even asserted the ticket text. But the live ticket read &lt;em&gt;"Your trace carries carrying a day that was never yours to keep…"&lt;/em&gt;. The sentence began with a verb, while every tag phrase was written to follow an implied "you". Only rendering a real ticket and reading it caught the problem. It's fixed, and a unit test now pins the full sentence for a real two-ending walk. The lesson we wrote down: &lt;em&gt;green gates prove the code runs, not that the prose is English. Read the page.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Publishing the new exhibits stopped halfway.&lt;/strong&gt; The first live run of the three-exhibit publish script published seven endings, uploaded the toaster's drawing and created its draft. Then submitting that draft for review failed. Each review record held a &lt;strong&gt;strong&lt;/strong&gt; reference to its exhibit, and a brand-new exhibit has no published document for a strong reference to point at. Our review flow had only ever been used on existing exhibits, so it could never have reviewed a new one, from a script or from Studio. Codex had answered "yes, the flow supports brand-new artifacts" by reading the code, and the reviewer had accepted that. Nobody had checked it against Sanity's rules for references. Nothing visitors could see broke. The fix was to make the reference weak.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then the resume failed too.&lt;/strong&gt; Instead of deleting live documents, the script gained a resume mode. Its dry run reported things that weren't true. That was the second failure on one task, and we never make a third identical attempt, so the task moved to another model family. A Claude Sonnet builder found two bugs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Results were matched to documents by list position rather than by id.&lt;/li&gt;
&lt;li&gt;Documents were compared with &lt;code&gt;JSON.stringify&lt;/code&gt;, which cares about key order, and Sanity returns keys in its own order.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Codex &lt;code&gt;gpt-5.6-sol&lt;/code&gt; reviewed that fix and returned &lt;strong&gt;fail&lt;/strong&gt;. A draft edited in Studio between the checks and the approval could have gone live unchecked, and resume mode skipped some safety checks. Those were fixed and re-reviewed, and the resume run finished cleanly. The lesson: &lt;em&gt;a dry run proves the plan, not the write.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An instruction that wasn't in the repository.&lt;/strong&gt; Mid-build, a message appeared in agent contexts claiming "bypass permissions mode is active" and telling agents to edit files through raw shell commands. The orchestrator first told the founder the text was in &lt;code&gt;AGENTS.md&lt;/code&gt; without opening the file. It wasn't; a search of the whole workspace found nothing. Later, two workers each spotted the same message, refused it because it conflicted with their packets' file rules, and said so in their reports. The rule now: read an instruction-file finding off disk before acting on it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Official Sanity Workflows, and why the custom gate stays
&lt;/h3&gt;

&lt;p&gt;Sanity Workflows was in early access, and we gave it a time-boxed spike (T-013). Codex wrote an &lt;code&gt;exhibit-review&lt;/code&gt; definition with four stages: drafting, curatorial-review, approved and on-display. Requesting changes requires a reason, and publishing is held during drafting and review. Six tests run it on the real engine in memory. Two deliberate breaks, removing the reason requirement and adding a publish hold to &lt;code&gt;approved&lt;/code&gt;, each made a test fail, as they should. The definition is deployed, and a live demo run walked every stage (details under Sanity Project Details).&lt;/p&gt;

&lt;p&gt;We did &lt;strong&gt;not&lt;/strong&gt; retire the custom flow. The spike's write-up (&lt;code&gt;docs/workflows-spike.md&lt;/code&gt;) lists three things that don't carry over:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Revision pinning.&lt;/strong&gt; The custom flow approves one exact draft revision and refuses to publish if the draft has changed since. The Workflows definition language has no built-in way to do that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validation-gated publishing.&lt;/strong&gt; The custom action waits for Sanity validation. The definition doesn't model that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guards are advisory in early access.&lt;/strong&gt; The Studio plugin and engine honour them, but Content Lake doesn't yet enforce them against direct writes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So we run both. The custom &lt;code&gt;artifactReview&lt;/code&gt; flow stays the hard publish gate, and Workflows coordinates the curator's stages in Studio.&lt;/p&gt;

&lt;h3&gt;
  
  
  An agent in the workflow: the Acquisitions Clerk
&lt;/h3&gt;

&lt;p&gt;An outside review of an earlier draft of this entry said what was missing: an agent and a person working through the same workflow. The organizers' own phrase for it is an agent moving a draft forward and a person approving it. So we built the Clerk (&lt;code&gt;src/agents/acquisitions-clerk/&lt;/code&gt;, &lt;code&gt;scripts/acquisitions-clerk.ts&lt;/code&gt;, &lt;code&gt;docs/acquisitions-clerk.md&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Three findings from Sanity's docs shaped it before any code was written:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The engine takes the actor from the token.&lt;/strong&gt; There is no parameter that says "this was the agent". If the Clerk borrowed the founder's login, the history would say the founder did it. So the Clerk runs on its own Editor robot token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Workflow role checks are advisory.&lt;/strong&gt; Sanity says so plainly. We still pinned &lt;code&gt;request-changes&lt;/code&gt;, &lt;code&gt;approve&lt;/code&gt; and &lt;code&gt;put-on-display&lt;/code&gt; to &lt;code&gt;roles: ["administrator"]&lt;/code&gt;, and the Clerk's own code refuses anything but &lt;code&gt;submit&lt;/code&gt;, with a test for each. But we don't claim the agent is &lt;em&gt;unable&lt;/em&gt; to approve; we claim it doesn't, and show where it's stopped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent Actions Generate can't fill reference fields&lt;/strong&gt; without an embeddings index that is deprecated with no replacement, and an exhibit is mostly references. So the Clerk uses Agent Actions &lt;strong&gt;Prompt&lt;/strong&gt;, which returns JSON and writes nothing. The Clerk checks that JSON itself and writes the documents.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Clerk's draft has to pass the same limits as Studio's schema, and more: every consequence tag must already have a sentence on the visitor's ticket, and an ending may only lead to a published exhibit. A rejected answer goes back to the model once, with the reasons. If it fails again, nothing is written. The endings stay drafts until the curator publishes, so no unreviewed text is ever public.&lt;/p&gt;

&lt;p&gt;Each Workflows move is mirrored onto the custom gate, which pins the exact revision. If anyone edits the draft after approval, publishing refuses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first live run (2 October).&lt;/strong&gt; The brief was one line: "A doormat that knows who is coming", for the Domestic Weather Memory wing. The Clerk wrote &lt;em&gt;The Doormat That Knows Who Is Coming&lt;/em&gt; with Sanity's Agent Actions on the free monthly AI credits. It passed every check on the first answer, and the Clerk submitted it. The curator asked for changes in Studio: &lt;em&gt;"The two choices are too plain. Make them feel like a decision about being known, for example wiping your feet or stepping over the threshold, and make the second ending as specific and sensory as the first."&lt;/em&gt; The Clerk read that note from the workflow and turned "Step onto the mat" / "Walk around the mat" into &lt;strong&gt;"Wipe your feet and let the house know you"&lt;/strong&gt; / &lt;strong&gt;"Step over the threshold without being known"&lt;/strong&gt;. It rewrote the second ending around crowded coat hooks, dim entry lamps and rain beading on your sleeves, then resubmitted. The curator approved and published, and the exhibit appeared in the hall within a minute. Its first ending leads on to the Memory Umbrella, and a visitor's ticket now quotes it. The whole run used three AI credits.&lt;/p&gt;

&lt;p&gt;Workflow history shows the Clerk's moves under its own robot id (&lt;code&gt;g-BHx7IW47nZRW&lt;/code&gt;) and the curator's under the curator's own name. It does not add a separate "agent" badge; the engine credits whoever holds the token, which is why the Clerk has its own. The proof is in &lt;code&gt;evidence/T-018/clerk-*-executed-*.json&lt;/code&gt; and &lt;code&gt;evidence/T-018/screens/&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One small blemish we left: the revised choices kept their original keys, so the address bar still reads &lt;code&gt;?choice=step-onto-the-mat&lt;/code&gt;. Renaming them would have meant another review round for a cosmetic change.&lt;/p&gt;

&lt;h3&gt;
  
  
  Outside Studio: the Curator's Desk
&lt;/h3&gt;

&lt;p&gt;To go beyond the Studio, the curator also gets a small App SDK app, the &lt;strong&gt;Curator's Desk&lt;/strong&gt; (&lt;code&gt;apps/curators-desk/&lt;/code&gt;). It runs in the Sanity Dashboard. One screen lists every exhibit, including drafts nobody has published, with what a curator checks first: are all the endings there, is the plate there with its alt text, where is it in review? Selecting an exhibit opens its live workflow run. The buttons come from the workflow engine's own evaluation for the person signed in, and every move also updates the revision-pinned gate.&lt;/p&gt;

&lt;p&gt;One surprise: inside this repository the Sanity CLI kept building the Studio instead of the app, because it looks for a Studio config in parent folders before it looks for an app. A small script stages the app outside the repository and runs the CLI there.&lt;/p&gt;

&lt;p&gt;The Desk is built and tested (seven unit tests, an app build, and a logged-out load that hands over to Sanity's sign-in), but we didn't capture it signed in for this post.&lt;/p&gt;

&lt;h3&gt;
  
  
  Test results
&lt;/h3&gt;

&lt;p&gt;The latest numbers the orchestrator re-ran itself:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unit tests: 266/266 passed&lt;/strong&gt; (Vitest), including ten behavioural tests of the Workflows definition, 49 for the Clerk and 7 for the Curator's Desk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser tests: 77 passed&lt;/strong&gt; (Playwright, against a production build).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Typecheck and lint:&lt;/strong&gt; clean.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;npx sanity-workflows deploy --check&lt;/code&gt;:&lt;/strong&gt; passed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;npx sanity documents validate&lt;/code&gt;:&lt;/strong&gt; 32/32 documents valid after the Clerk's exhibit went live.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;npx sanity schemas validate&lt;/code&gt;:&lt;/strong&gt; 0 errors, 0 warnings.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What was cut, and what isn't finished
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A public "In conservation" strip&lt;/strong&gt; listing exhibits under review was cut. Review records, drafts and workflow runs are private in the dataset (an anonymous query counts 0 of each), and the site deliberately reads without a token. Showing them would mean putting a read token on the web server or publishing titles nobody had reviewed yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Seven exhibits, not more.&lt;/strong&gt; We added three by hand instead of five, to leave time for the Workflows spike, and the Clerk added the seventh through review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The review flow covers exhibits only.&lt;/strong&gt; Endings were published by separate guarded scripts. Each one runs dry by default, checks every document before writing anything, and writes everything in one transaction that aborts if any document changed in between.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The custom publish guard is a Studio guard.&lt;/strong&gt; There is a brief gap between its final check and the publish, and anyone with enough API permissions can bypass it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every ticket shares one link-preview picture.&lt;/strong&gt; Only the title and description change, because of a limit in Next.js 16.3.5 that we traced into the framework's source. The preview cards also can't show the plates, because the image renderer can't draw them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A wing's floor plan&lt;/strong&gt; doesn't yet light up the room you're in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dependency advisories.&lt;/strong&gt; The Sanity command-line and Studio tooling carries 15 known advisories (12 moderate, 3 high). The only automatic fix is an incompatible Sanity downgrade, so we didn't apply it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Project ID:&lt;/strong&gt; &lt;code&gt;wa27n68e&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dataset:&lt;/strong&gt; &lt;code&gt;production_1&lt;/code&gt; (public read; the site reads it with no token)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API version:&lt;/strong&gt; &lt;code&gt;2026-09-20&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A public query you can open in a browser. It lists the three wings and their accent colours:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://wa27n68e.apicdn.sanity.io/v2026-09-20/data/query/production_1?query=*%5B_type%20%3D%3D%20%22era%22%5D%7Btitle%2C%20accentColor%7D
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is &lt;code&gt;*[_type == "era"]{title, accentColor}&lt;/code&gt;, URL-encoded.&lt;/p&gt;

&lt;h3&gt;
  
  
  The content model
&lt;/h3&gt;

&lt;p&gt;Four document types:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                ┌───────────────────────────┐
                │ era (a wing)              │
                │ title, slug, summary,     │
                │ accentColor  (#rrggbb)    │
                └───────────────────────────┘
                   ▲ era (required)      ▲ era (required)
                   │                     │
┌──────────────────┴───────┐        ┌────┴──────────────────────┐
│ artifact (an exhibit)    │ choices│ outcome (an ending)       │
│ title, slug,             │ 2–4,   │ title, body,              │
│ accessionNote, summary,  │───────▶│ consequenceTags [1–4]     │
│ artifactLabel,           │ each to│                           │
│ visualDescription,       │ its own│                           │
│ image + alt (the plate), │ ending │                           │
│ choices[ label, outcome ]│        │                           │
└──────────────────────────┘        └───────────────────────────┘
     ▲             ▲                         │
     │             └──── leadsTo (optional) ─┘
     │ artifact (weak reference)
┌────┴──────────────────────────────────────────┐
│ artifactReview (the custom review gate)       │
│ state, submittedRevision, approvedRevision,   │
│ changeRequestReason, timestamps               │
└───────────────────────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why each connection exists:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;artifact.era&lt;/code&gt; and &lt;code&gt;outcome.era&lt;/code&gt;&lt;/strong&gt; are required references. The first groups exhibits into wings and tints their plates. The second tells the ticket which wing's closing line to use, directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;artifact.choices&lt;/code&gt;&lt;/strong&gt; holds two to four &lt;code&gt;{ label, outcome }&lt;/code&gt; pairs. A custom rule rejects two choices pointing at the same ending. Adding a choice in Studio changes the exhibit page with no code change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;outcome.consequenceTags&lt;/code&gt;&lt;/strong&gt; holds one to four unique lowercase tags matching &lt;code&gt;^[a-z0-9-]{2,32}$&lt;/code&gt;. They appear under an ending, and the ticket turns them into sentences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;outcome.leadsTo&lt;/code&gt;&lt;/strong&gt; is the "Continue to →" door. It is a strong reference, so a door can't point at an exhibit that doesn't exist. That also forced a publishing order for new exhibits: first the endings without doors, then the exhibits, then all the doors in one guarded transaction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;artifact.image&lt;/code&gt;&lt;/strong&gt; is the blueprint plate, with required alt text. The site checks the SVG for unsafe markup and draws it inline so it can take the wing's colour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;artifactReview.artifact&lt;/code&gt;&lt;/strong&gt; is a &lt;strong&gt;weak&lt;/strong&gt; reference. The halted publish run above is why.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Real queries from the app
&lt;/h3&gt;

&lt;p&gt;This query, from &lt;code&gt;src/content/sanity-repository.ts&lt;/code&gt;, returns each wing with its exhibits, found by reverse reference. &lt;code&gt;artifactProjection&lt;/code&gt; is the full exhibit shape: choices, their endings, and the endings' doors.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;*[_type == "era" &amp;amp;&amp;amp; defined(slug.current)] | order(title asc) {
  title,
  "slug": slug.current,
  summary,
  accentColor,
  "exhibits": *[_type == "artifact" &amp;amp;&amp;amp; references(^._id) &amp;amp;&amp;amp; defined(slug.current)] | order(title asc) ${artifactProjection}
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This one, from the same file, powers the ticket. The ids come from the page address, which can't be trusted. So they are checked against &lt;code&gt;^[a-z0-9-]{3,64}$&lt;/code&gt;, capped at 12, and passed in as the &lt;code&gt;$ids&lt;/code&gt; parameter, never pasted into the query text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;*[_type == "outcome" &amp;amp;&amp;amp; _id in $ids] {
  _id,
  title,
  body,
  consequenceTags,
  leadsTo-&amp;gt;{
    title,
    "slug": slug.current
  },
  era-&amp;gt;{
    "slug": slug.current
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The custom review gate
&lt;/h3&gt;

&lt;p&gt;The custom gate is a set of document actions on &lt;code&gt;artifact&lt;/code&gt; in Sanity Studio. A curator can:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Submit the current draft revision for review.&lt;/li&gt;
&lt;li&gt;Request changes, with a required reason.&lt;/li&gt;
&lt;li&gt;Resubmit, but only after the draft has actually changed.&lt;/li&gt;
&lt;li&gt;Approve the submitted revision.&lt;/li&gt;
&lt;li&gt;Publish, but only if the draft's current &lt;code&gt;_rev&lt;/code&gt; still matches the approved one &lt;em&gt;and&lt;/em&gt; Sanity's validation has finished with no errors.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Review records hold only public-safe fields: no names, emails or private notes. When the drawings moved into Sanity, and later when the three new exhibits went live, the scripts used this same submit → approve → guarded publish path instead of going around it.&lt;/p&gt;

&lt;h3&gt;
  
  
  The official Workflows definition
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;workflows/exhibit-review.ts&lt;/code&gt;, deployed to &lt;code&gt;production_1&lt;/code&gt; as &lt;code&gt;production.exhibit-review.v1&lt;/code&gt;. Version 2 (&lt;code&gt;production.exhibit-review.v2&lt;/code&gt;, deployed 2 October) adds the curator-only roles and a &lt;code&gt;submittedBy&lt;/code&gt; field for the Clerk; the Clerk's first run used it, and the earlier demo run stays on v1.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;drafting ──submit──▶ curatorial-review ──approve──▶ approved ──put-on-display──▶ on-display
    ▲                        │
    └──request-changes───────┘   (reason required, 1–500 characters)

Publishing is held in drafting and curatorial-review.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the live demo run on the vending-machine exhibit (&lt;code&gt;evidence/T-013/live-demo.txt&lt;/code&gt;), the engine refused a change request with no reason, then one with an empty reason:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;✖ fire-action error:
  Action "request-changes" on activity "review" rejected: invalid params
    - reason: required but missing

✖ fire-action error:
  Action "request-changes" on activity "review" rejected: invalid params
    - reason: length must be greater than or equal to 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then it accepted &lt;em&gt;"Plaque says 'Tuesday' where the machine sells Mondays; fix the date line."&lt;/em&gt; and sent the exhibit back to drafting. The exhibit then went back through review, was approved, and was put on display.&lt;/p&gt;

&lt;p&gt;Workflow documents have dotted ids, and Sanity keeps those private even in a public dataset. An anonymous query after the demo returned &lt;code&gt;[]&lt;/code&gt;, and the exhibit itself was not modified.&lt;/p&gt;

</description>
      <category>sanitychallenge</category>
      <category>devchallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
