This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
On 1 November 2026, Canada stops changing its clocks — one province at a time. British Columbia, Alberta and the Northwest Territories already went permanent this year; Manitoba follows on 31 October. IANA shipped five tzdata releases in 2026 to keep up, and for several of these changes it deliberately models the switch on 1 November at 02:00 instead of the legal date (a documented "temporary hack"). Morocco quietly went back to GMT on 20 September.
Meanwhile, the Node.js you're running ships whatever tz data it was built with. On my machine that's 2025b: 11 of 597 zones are wrong in 2026–27. Even the Vercel runtime this project is deployed on (tzdata 2026c) still puts Winnipeg an hour off. Ask a search engine whether Ukraine abolished daylight saving time and you'll get "yes" — the bill passed in 2024 and was never signed.
So I built an agent that answers what the clock and the calendar legally say at a place and moment — and on whose authority.
Clockwrit reads the law, the IANA time zone database and your runtime's own clock, and tells you which one is right — with the decree, the date and the source quote.
- Ask. "A weekly call is set for Mondays 10:00 in New York. What time is it in Winnipeg on 2 November 2026, and why?" The answer comes back as cards, not prose: the local time, and four clocks side by side — the law, IANA 2026e, this server, and your browser, read live — plus every instrument the answer rests on, with its authority tier (§ primary law, † reference data, ‡ community) and its legal status.
- Pending laws as Content Releases. The US Sunshine Protection Act, Ukraine's bill 4201, Florida's conditional permanent-DST law, the EU proposal and three more are modelled as Sanity Content Releases. Pick one and the agent answers as if it had passed — the same tools, reading the dataset through the release perspective.
- A ruling desk. When the Knowledge Base finds two sources that disagree, the agent drafts a ruling, a person signs it, and the decision becomes a standing instruction for every future rebuild. (More below — this is the part I'm proudest of.)
- A schedule auditor, a drift report, an eval page, an MCP server, a CLI and a Sanity Dashboard app.
Demo
Live: https://clockwrit.vercel.app — no login needed.
A five-minute path, if you're reviewing:
- Home — the page reads your browser's tz data for 2 November and grades it against the law (landing).
- Ask — try a suggested question, then switch "Answer under" to Ukraine bill 4201 and ask about Kyiv next July.
-
Ruling desk — draft and approve a ruling. Reviewer passcode:
every-hour-has-a-citation. - Audit — a weekly meeting across five cities, every occurrence checked.
- Eval — every question, every answer, every verdict.
Code
RaYYeR220
/
clockwrit
What time is it, legally? An agent on Sanity Context that reads the law, IANA tzdata and your runtime's clock, and tells you which one is right.
Clockwrit
What time is it — legally? An agent that reads the law, IANA tzdata and your runtime’s clock, and tells you which one is right, with the decree, the date and the source quote. Built on Sanity Context: a typed dataset for the facts, a Knowledge Base for what the sources say, and a ruling loop for when they disagree.
Live: https://clockwrit.vercel.app · Video (3 min): https://youtu.be/pSU875tMuzw · Ask: /ask · Ruling desk: /desk · Eval: /eval · MCP: https://clockwrit.vercel.app/api/mcp · Reviewing? Start with JUDGES.md.
Why this exists
In 2026 Canada stopped changing its clocks one province at a time: British Columbia, Alberta, the Northwest Territories, then Manitoba from 31 October. IANA shipped five tzdata releases (2026a–e) to keep up, and for some of them deliberately models the change on 1 November at 02:00 — a documented “temporary hack” — while the law takes effect on another…
A pnpm monorepo: a deterministic legal-time engine (packages/core, 30 tests), the Next.js app and agent (web), the Studio and schema (studio), the App SDK Dashboard app (desk), the dataset and the 90-source corpus (knowledge), the evaluation (eval) and the CLI (packages/cli). MIT.
How I Used Sanity
1. A typed dataset where every fact points at its law
The schema is the product. A ruleSegment is one zone's clock regime over one span of time — the standard offset, the DST rule written the way the law (and zic) writes it (Sun>=8 2:00 wall), validFrom/validTo as UTC instants — and a required basis[] that references the instrument it rests on. An instrument is anything that makes a claim: a statute, a decree, a bill, an IANA release, a Microsoft notice, a news article, a holiday library. Each has an authority tier (primary / secondary / community), a legal status (in force, enacted-not-effective, conditional, pending bill, proposed, vetoed…), dates, and verbatim quotes in their original language with translations.
The calendar side has holiday (with dateCertainty: fixed by law, announced, or merely calculated — Eid moves on a moon sighting), weekendRegime (the UAE moved its weekend in 2022), workdayOverride (China's adjusted working Saturdays) and holidaySuspension (Ukraine's martial law suspends holiday days off). Each jurisdiction records which years have a complete holiday list, so the agent knows when it's allowed to say "working day".
845 documents, 179 instruments (83 of them primary law), 47 zones, 58 jurisdictions. Every one of the 114 clock regimes is checked against IANA tzdata 2026e — daily and hourly around every transition, 2020–2028 — before it's loaded.
2. Content Releases as pending legislation
A release is a set of document versions that publish together. That's exactly what a bill is. Seven releases hold the clock regimes as they would read if each law took effect, and the agent's tools read the dataset with perspective: [releaseId]. Asking "what if Ukraine's bill is signed?" isn't a prompt trick — it's the same deterministic computation over a different perspective.
3. Two Sanity Context MCP endpoints
-
clockwrit-catalog— GROQ mode over the dataset. The agent usesgroq_queryandschema_explorerto find instruments, statuses and dates. The endpoint's instructions explain what authority and status mean. -
clockwrit-sources— Knowledge Base mode. The agent usesknowledge_base_search/knowledge_base_readfor what the sources actually say.
Both endpoints' /initial-context is inlined into the system prompt so the agent doesn't spend a tool call on orientation. Every Ask answer shows an evidence trail of the Context tool calls it made. Conversations are saved to Context Insights through @sanity/context's AI SDK integration.
The agent never states an offset or a working day from memory. Three deterministic tools — legal_time, working_day, audit_schedule — read the typed records and compute. Code decides; the model explains.
4. A Knowledge Base built from 90 real sources — in two passes
I imported the corpus in two phases, each source with a header naming its authority tier:
- Phase A, the law first: statutes, proclamations, gazettes, government announcements, IANA NEWS entries.
- Phase B, then the internet: the news headline that overstated a bill, the Microsoft notice with a different date, the Wikipedia line its own citation doesn't support, the holiday library that's a day off for Pakistan's Eid.
5. The ruling loop: Knowledge Base issues → instructions
When the build finds a fact the sources disagree on, it files a conflict issue with both sides and the exact quoted lines. The one that matters most: Manitoba's proclamation says 31 October; IANA 2026e says 1 November 02:00.
On the ruling desk the agent drafts a ruling — which side, why (authority, status, dates), and a cross-check against the typed records — through a forced tool call. It cannot apply it. Applying is a code path behind a gate: the reviewer passcode, or — from the Dashboard app — the reviewer's own Sanity token, accepted only if a dry-run edit of that ruling succeeds with their permissions and they administer the organisation that owns the Knowledge Base. Approval resolves the issue through the Knowledge Base API, which mints an instruction that every later build follows, and writes the ruling back to the dataset.
6. The App SDK desk
A Dashboard app lists rulings live (useDocuments, useDocumentProjection), lets an editor approve with their own identity (useAuthToken), and shows the pending laws (useActiveReleases).
Why this only works because the content is structured
I wrote 40 questions from the verified research before the agent existed, each with an answer key and a trap reason. Ten were used while building; the 30 held-out ones were run once against three arms, graded by a model from a different family:
| Correct | Traps correct | |
|---|---|---|
| Clockwrit | 30 / 30 | 14 / 14 |
| Same model, no tools | 22 / 30 | 8 / 14 |
| Keyword search (BM25) over the same 90 sources + same model | 16 / 30 | 6 / 14 |
The model alone is confidently wrong exactly where the world changed this year: it says Casablanca is UTC+1 (it isn't, since 20 September), that 01:30 on 20 September happened once there (twice), that Pakistan's Eid al-Adha was the 28th. Keyword search made it worse than no retrieval at all — 8 abstentions — because the passages that match the words rarely carry the effective date, the legal status or the instant that decides the answer. Every answer and verdict is on the eval page.
What surprised me (and what I'm not claiming)
- The Knowledge Base is smarter than I planned for. My first test fed it a news headline ("Rada abolished the clock change") and the bill card ("never signed"). It didn't raise a conflict — it wrote "adopted but not in force", which is correct. Conflicts only appear when claims truly can't both be true. That's why the second import pass matters.
- It doesn't rate authority on its own. Conflict sides came back with no authority tier, and in one test it rated a community source as primary. Clockwrit takes authority from the typed records instead.
- A full rebuild restructures every entry. Decisions survive (issues keep a stable fingerprint, instructions persist) — the entry paths don't.
- The legal date and tzdata disagree on purpose. For Manitoba and the NWT the dataset records the legal effective instant; the UTC offset is identical either way, only the DST flag and abbreviation differ.
- Not claimed: global coverage (47 zones and 58 jurisdictions where something changed or is contested; elsewhere it says so), discovery of unknown facts (the eval measures whether structure gets the agent to the right answer), legal advice. The full ledger is docs/CLAIMS.md.
Sanity Project Details
-
Project ID:
c9x90tjo— datasetproduction(public) - Public dataset query: every instrument with its authority and status
- Studio: https://clockwrit.sanity.studio
-
Context MCP endpoints:
clockwrit-catalog(dataset, GROQ mode),clockwrit-sources(Knowledge Base mode)
Use it from your own agent:
claude mcp add --transport http clockwrit https://clockwrit.vercel.app/api/mcp
Agent Session
No session transcript is attached to this entry.





Top comments (0)