This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
"Why build this? You can just look it up." That was my first reaction too, so I tested it. I took 22 real contests and asked an AI model on my laptop when each one closes, in four cities. That's 88 questions, and every one asks for the date on a final answer line.
Here's what came back:
- No lookup tools: 0 of 88 right. It still put a wrong date on the answer line 44 times. In 9 of those, the reply said up front that it couldn't look the contest up or was going from memory. Its prompt told it to say that.
- The contest record from Sanity: now it really could look it up, and it got 70 of 88. All 18 wrong answers read the right cutoff from the record, and 17 of them also wrote the right UTC time. Every one of them slipped when turning it into local time.
- The record plus deterministic time tools: 87 of 88, with no wrong date. The one miss ran out of steps.
Okay, so the lookup isn't the hard part. The clock is. Here's one question both ways. When does Build With AI: Basics close, in Los Angeles time? Both times, the model pulled the same Sanity record. The rules say 5:00 pm ET on Oct 26.
- Sanity data, model does the time math itself: 1:00 PM on Oct 26. Wrong. It had Pacific's summer and winter offsets backwards: "PT (Los Angeles) is UTC-7 (or UTC-8 during DST)".
- Sanity data plus the time tool: 2:00 PM on Oct 26. Correct.
Same facts, one hour off. That's why the agent is told to do every conversion with a tool.
What I Built
Real Cutoff is a Sanity dataset of 25 real contests, a board, and an agent that reads the dataset through Sanity Context. I checked each contest against its own official rules on 2026-09-26. It answers the four questions I had before entering anything:
- When does it really close, in my time zone? The board shows every cutoff in your zone. The agent is told to convert with a tool, never in its head.
- Does it actually pay cash? Instructables' Halloween Contest lists $5,100 in prizes. All of it is gift cards.
- What rights do I give up? Two contests make you hand rights over. huntr's terms assign all IP rights in your contribution to Palo Alto Networks. The OpenAI Safety Bug Bounty has you assign your testing results to Bugcrowd. A third, Hack-Nation, publishes no IP terms at all.
- When sources disagree, which cutoff is safe? The earliest one any official source gives. 7 of the 25 contests state different closing times in different official sources. This challenge is one of them.
I went looking for contests to enter and kept running into pages that contradict themselves:
-
This challenge's own page. Its Live/Ended status is coded to switch at
Sun Oct 04 2026 23:59:59 GMT+0000, which is 4:59 PM in Las Vegas. The written rules say 11:59 PM PDT, about seven hours later. The rules govern, but Real Cutoff plans around the earlier time. I re-checked the page's code on Sep 28. - An AI horror film contest. The rules say "October 9th, 2026, at 11:59 PM PST". Pacific clocks are on PDT until Nov 1. And the entry portal's own configuration closes entries on October 3 at 10:00 PM Los Angeles time.
- itch.io game jams. The raw timestamps are UTC, so "Oct 24 03:59" is really Oct 23, 8:59 PM in Las Vegas.
- "Oct. 2, 2026, midnight UTC". That's the start of Oct 2. In West Coast time it's Oct 1, 5:00 PM.
A keyword search can't answer "when does this close for me". You need every source's claim stored as data and a rule for which one is safe. Then you still have to get the time-zone math right.
The proof table
The same 88 questions go to the same model three ways. They cover 22 fixed-deadline contests, each in Los Angeles, New York, London and Tokyo time. The other 3 contests are rolling bounties with no fixed deadline, so they aren't in it. The model is Qwen3.5-9B (Q5_K_M), running on an M4 MacBook Air through llama.cpp. One prompt version, one dataset version and a 32k context the whole way through. The answer key is each contest's safeCutoffUtc, converted with luxon and cross-checked against Intl in a unit test. The model gets that same field from Sanity. So the test is whether it can read and convert that field, not whether it can choose between sources (see Limits).
| What the model had | Correct | Wrong date or time |
|---|---|---|
| No lookup tools (question + system prompt, no deadlines) | 0/88 | 44 |
| + Sanity Context (GROQ + Knowledge Base) | 70/88 | 18 |
+ Sanity Context + time tools (/chat's setup, capped at 10 steps) |
87/88 | 0 |
With time tools it got 22/22 in Los Angeles, New York and Tokyo, and 21/22 in London. It called a time tool on all 88 questions. I had an earlier run (v1) that scored 0, 69 and 88. But it used an older prompt and a 16k context, and 8 questions overflowed that context, so I reran all three from scratch. The table is the rerun (v2). Every miss is broken down below.
Every miss, and how it was scored
contestCards tool was attached in every config. It only echoes whatever the model passes it. With no data, the model called it on 18 questions with made-up records./chat allows 24 steps. I haven't re-run this question there.
It runs on a laptop. In Local mode, the chat's default, the model runs on my Mac through llama.cpp, so your question and the answer stay here. What does leave goes to Sanity's API: the two Context connections, their initial context, each lookup the model writes (it can include words from your question), and the page's live-update stream, which carries no messages. The time tools run in the app. Under every answer there's a "What left this machine" panel that lists the lookups and notes the connection.
Asked when this challenge closes, the live chat in Local mode answered Sunday, Oct 4, 2026, 4:59 PM PDT. That's right, and it came with both claims, the recheck line and the panel. The answer is recorded, unedited, and you can read it with its tool calls on the hosted /chat. It took 4 min 36 s on this laptop.
Demo
The chat needs my laptop, because that's where the model runs. You don't need it to check the work:
-
The board: https://real-cutoff.vercel.app. Every safe cutoff is shown in your own time zone. Your browser converts it with the same
convertInstantfunction the agent's time tool calls, so no model is involved. - This challenge's own page, both claims side by side: https://real-cutoff.vercel.app/contests/sanity-challenge-dev-x-sanity
-
The proof table, read live from the
evalRundocuments in Sanity: https://real-cutoff.vercel.app/proof -
Real chat answers, recorded from my laptop and unedited: https://real-cutoff.vercel.app/chat. The hosted copy has no model, so it replays four answers with their tool calls and the "What left this machine" panel. Each one is labelled with the date and how long it took. Every attempt, kept or not, is listed in
app/data/replays/NOTES.md. - Live GROQ against the public dataset, no login:
- Studio: https://real-cutoff.sanity.studio (needs a Sanity login with access to this project)
Code
https://github.com/holdrfoldr/real-cutoff
Built on Sanity's knowledge-base starter (MIT).
Where to look:
-
app/lib/time-math.tsand its tests: the five time functions the agent's tools wrap. -
app/lib/agent-core.ts: one agent for Local and Cloud. It connects both Context endpoints and merges their tools with the time tools. -
context/: the Knowledge Base purpose, its source query, and both endpoints' instructions. They're checked in so they can be diffed. -
app/scripts/eval/: the proof-table harness and its strict scorer. -
studio/structure.tsandfunctions/set-reverify-by/: the desk lists and the Sanity Function.
How I Used Sanity
The content model does the work. Each contest stores what each source says about the deadline, not just one deadline:
-
deadline.governingis the source that controls the cutoff. -
deadline.otherClaims[]holds every other official source that states a different instant. - Each
deadlineClaimkeeps the zone label as printed (PST,CUT, …) apart from the real IANA zone and the UTC instant. Mislabels stay visible. -
deadline.safeCutoffUtcis the earliest instant any official source gives. - Prize tiers carry a
form(cash|gift-card|credits|crypto| …). So "which pay cash" is a filter, not a vibe. - Rolling bounties are marked
cashTotalMeaning: "max-single-award". The board labels them as a per-award maximum, and the agent is told never to rank them against real prize pools. -
verifiedAtandreverifyByput a date on every record. A Sanity Function fills inreverifyBy(14 days afterverifiedAt) when a record is saved without one.
The Studio desk has the lists an editor actually needs: Closing soon, Sources disagree, Pays gift cards, not cash, Rolling programs, Needs reverify, Closed.
Two Context endpoints, because one can't hold both. A Context MCP endpoint serves one source type. Per Sanity's docs, if you attach a dataset and a Knowledge Base to the same endpoint, the dataset wins and the Knowledge Base is silently ignored. So the agent connects to two:
-
real-cutoff-groq(GROQ mode, filter_type == "contest"), withgroq_query,schema_explorerandarray_field_reader. It handles anything you'd filter, sort or compare: deadlines, cash totals, eligibility. -
real-cutoff-kb(Knowledge Base mode), withknowledge_base_read. It handles the prose questions: why a deadline is a trap, what rights you give up, how two contests compare.
Routing is instructions, not code. One model gets both tool sets plus the time tools, with each endpoint's initial context in its system prompt. That prompt says structured facts come from GROQ and explanations come from the Knowledge Base. The Knowledge Base endpoint's own instructions say the same. The "What left this machine" panel shows which one the model called. Every eval question asks for a deadline. In the full config the model called groq_query on all 88 questions and knowledge_base_read on 2.
The Knowledge Base is built from the same records. Its dataset source is a GROQ projection that turns each contest into a title and a prose body, so no number gets retyped:
// excerpt: context/knowledge-bases/real-cutoff.md
"deadline": select(
deadline.rolling => deadline.note,
"Governing: " + deadline.governing.sourceLabel + " says \"" + deadline.governing.statedText +
"\" (" + deadline.governing.utc + " UTC). " +
"Safe cutoff (earliest of every source): " + deadline.safeCutoffUtc + ". " +
select(defined(deadline.trap) => "Trap: " + deadline.trap, "")
),
Then I gave it the real pages. As of 2026-09-30, the Sanity Context app shows the Knowledge Base, "Real Cutoff - contest rules, rights and deadline traps", with 8 sources. 7 are URL sources: the official pages of the 7 contests whose sources disagree. The eighth is the Real Cutoff / production dataset (25 documents).
While it was building, it filed an Issue on its own, tagged Conflict and Critical:
The AI-Use Policies entry states the Curious Refuge AI Horror Film Contest deadline is October 9th, 2026 at 11:59 PM PST, but the Countdown vs Rules Discrepancies entry reports that the Easypromos portal configuration shows an earlier deadline of October 3, 2026 at 10:00 PM Los Angeles time, a difference of roughly 6 days.
That's the horror contest from earlier, the same contradiction I'd found by hand. I resolved it by choosing the portal's October 3, 10:00 PM entry, and Context saved that as an Instruction: "Saved as an instruction that shapes your content on the next rebuild." There's also one manual Instruction on the same contest. The pages stay attached, so a refresh can flag it again.
The agent is told never to do time math in its head. Five deterministic tools built on luxon (resolveWallClock, convertInstant, checkZoneLabel, earliestCutoff, timeUntil) have 29 unit tests covering every trap above. The system prompt requires a tool call for any conversion. Nothing in code forces it, so the eval counts it. With time tools, the model called convertInstant on all 88 questions. It skipped the other rules more often. It called earliestCutoff on 13 of the 28 questions whose sources disagree, and checkZoneLabel on 13 of the 32 whose rules print a zone label like PST.
Limits
- This is a snapshot of 25 contests as of 2026-09-26. Terms change, so every record has a
reverifyBydate. - The eval tests looking up and converting a deadline. The safe cutoff is precomputed in
safeCutoffUtc, and the question says to use the earliest one, so choosing it isn't tested. - 88 questions, one scored answer each, no fixed seed. Between the two runs the Sanity-only total moved by one, but 25 of its 88 answers flipped (13 became right, 12 became wrong). The runs also differed in prompt and context size, so that isn't all randomness. Still, honestly, treat any single total here as plus or minus a few.
- The model doesn't reliably add the recheck line: at least 15 of 88 full-mode replies had it. Contest pages print the line themselves.
- The eval found a real bug. In the first run, at the server's old 16k context, 8 of the 176 Sanity-backed questions (2 Sanity-only, 6 with time tools) ran out of room. The live chat sends the same prompt and allows more steps, so it would likely have hit the same wall. The server now runs at 32k, and every score in the table comes from a 32k run.
- The local model is slow. It took a median of 71 s per eval question with all tools on, on an M4 MacBook Air (24 GB) busy with other jobs, and about 3 to 4 minutes for the live chat's first answer. A Cloud mode (Claude) is in the code but not evaluated here.
Sanity Project Details
- Project ID:
bfxn45k2, datasetproduction(public) - All 25 contests with every deadline claim (no login)
- The eval results are Sanity content too: one
evalRundocument per config, which/proofreads (no login) - Studio: https://real-cutoff.sanity.studio (needs a project login)
One honest note on how this got made: I built Real Cutoff with Claude Code as my pair, and it helped me write this post too. The numbers come from the repo, the public dataset and the Context app, and the eval ones are live on /proof if you want to check them yourself.


Top comments (0)