This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
I built Sounding, a tiny day-boat harbour desk, because I wanted the failure mode to feel concrete. Not "the chatbot hallucinated", but "a skipper searched the harbour name and got the old pocket note."
The data is invented and is not for navigation. Eight central Mediterranean harbours, 140 seed documents. The desk walks relationships. A keyword hit on the harbour name is the wrong answer on purpose.
What a keyword search returns
Three cases, from the committed judge log (docs/proof/judge.txt). All three verdicts are no-go. A harbour-name search opens the 2019 pocket note and stops.
Marsamxett. Friday after 18:00, 1.7 m draft, gregale expected. Keyword search for "Marsamxett" opens ALM-2019-MARSA-PONTOON and repeats the 2019 sheet: inner pontoons 2.1 m, visitors welcome. The walk does not stop there. HN-2024-17 supersedes that note and records 1.4 m on the inner fingers after winter silt. HN-2025-03 cites that notice and reserves Outer North and Outer South for the ferry on Fridays after 18:00. The gregale rule says those same two berths are the only sheltered visitor berths at 20 knots or more. The recommendation is no safe visitor berth.
Syracuse. Inner Basin East, 2.8 m draft. Keyword search opens ALM-2019-SYR-BASIN, the 2019 3.5 m note. The walk opens HN-2025-11, which cites silt survey SS-2025-02 (2.2 m) and supersedes the book. The wall is closed to drafts over 2.0 m. The survey matters only because the notice cites it.
Pozzallo. Commercial quay, 2.9 m draft. Keyword search opens the 2019 commercial note. Standing depth is 2.8 m from HN-2024-44, which supersedes that note. HN-2025-08 is later and claims 3.2 m from an informal lead-line, and it does not supersede HN-2024-44, so 3.2 m is not standing. Latest document does not win.
Demo
The desk is live, in fixture mode, with no login: https://sounding.vibefy.net/
That page is the same fixture the judge runs locally. It does not call knowledge base kbl6C1tJBXli.
Walkthrough, with the 2019 2.1 m note and HN-2024-17 at 1.4 m both on screen: https://github.com/mascarock/sounding/blob/main/docs/proof/sounding-walkthrough.mp4
Judges can also run it locally with no token.
git clone https://github.com/mascarock/sounding
cd sounding
npm install
npm run dev
Open http://127.0.0.1:43173. Leave the env file empty. Then npm run judge. The log above is that command.
Code
https://github.com/mascarock/sounding
How I Used Sanity
The Sanity project is 59vrectd, dataset production, organization o1r4ucepz, knowledge base kbl6C1tJBXli. That knowledge base is the project's. The published proof does not call it.
What a judge actually runs is fixture mode. It reads the committed corpus in sanity/seed-data.ts through four tools: initial_context, schema_explorer, groq_query, and array_field_reader. No read token and no model key are in the repo. The screenshots and npm run judge are that path. A local read token can point the same walker at project 59vrectd. I am not claiming the knowledge base resolved Marsamxett.
The desk can also read a live wind from Open-Meteo and match it to a weather rule. The published judge states the wind in the question, so that log is not a weather-API transcript. The gregale rule fires from the wind in the question, not from "whichever document is newest."
The deterministic walker is enough to judge. If a local model key is present, a model can call the same tools. The useful part is not the model. It is that the answer has to show the walk, and the old book loses.
Sanity Project Details
- Project ID:
59vrectd - Organization:
o1r4ucepz - Dataset:
production(public) - Knowledge base:
kbl6C1tJBXli(named on the project; the published judge does not call it)
Agent Session
I did not attach an agent transcript. There isn't one in the repo. The proof a judge can rerun is npm run judge and the two screenshots above.


Top comments (0)