This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
In Nigeria you can score well in JAMB's UTME, get offered admission, and still be turned away at clearance. Clearance is when the university checks your O'level results against the course's rules. The mistakes are small and specific:
- a Physics credit that came from a second sitting when the course allows only one
- no Further Maths credit for UNILAG Computer Science
- Mathematics counted as one of your UTME subjects for Law
One wrong subject costs a whole year.
Clearance Desk is an agent for applicants (and the parents and teachers helping them). It checks your UTME subjects, UTME score and O'level sittings against the published 2026/2027 requirements of UNILAG, UI, OAU, UNN and LASU, before you apply.
- For each course it tells you Eligible, At risk or Not eligible, and explains why in plain English, linking every source.
- Then you can ask the Knowledge Base what to do next: deadlines, screening windows, awaited results. Answers come only from JAMB and university notices.
The core design rule: the model finds and explains; deterministic code judges. Claude never decides a verdict. It finds the right rules through Sanity Context, a small TypeScript evaluator checks them, and Claude explains the result using the Knowledge Base.
Demo
Live: https://clearancedesk.vercel.app. No login, and it works on a phone. Tap one of the three sample candidates at the top:
- You'll watch each Sanity Context step land live.
- The verdict appears as soon as the code decides it, usually within 10–15 seconds.
- Then the explanation arrives.
- Finally, try one of the suggested questions under the result.
In the two-minute narrated video:
- Chioma (UTME 301) checks Medicine at UNILAG and gets Not eligible: UNILAG allows one sitting, and her Physics credit is from a second one.
- With the same results, UNN Nursing (two sittings allowed) gives Eligible.
- "Show courses I qualify for" checks 10 courses at once.
- She asks the Knowledge Base for her upload deadline, which turns out to have been extended.
- A look at how the Knowledge Base was built, and the trace.
- The eval.
Code
Clearance Desk
An AI agent that tells Nigerian university applicants whether their UTME subjects, UTME score and O'level results meet the published requirements for a specific course at a specific university — before they apply, so they don't get admitted and then rejected at clearance.
Built for the DEV × Sanity Challenge, Path One: "Ship an Agent That Queries Real Content".
Live: https://clearancedesk.vercel.app (no login). The About page shows the architecture and live data coverage.
Core principle: the model finds and explains; deterministic code judges.
🚧 Work in progress. See
BUILD_SPEC.mdfor the plan andBUILD_LOG.mdfor the build journal.
Layout
| Path | What |
|---|---|
studio/ |
Sanity Studio + schema (sources, subjects, institutions, programmes, requirements) |
data/ |
sources.yaml, catalog.yaml, generated seed.ndjson
|
scripts/ |
NDJSON builder, MCP endpoint checks, eval runner |
web/ |
Next.js app: form, /api/check agent route, deterministic eligibility evaluator |
Data (Phase 2)
34 admission requirements for the 2026/2027 session across 5…
The repo has:
- the Studio schema
- the source-traced data pipeline
- the evaluator, with 49 unit tests
- the agent and the Next.js app
- the eval
- an honest build log of every wrong turn
How I Used Sanity
1. Admission rules as structured content
Requirements live in a Sanity dataset as data a program can check, not as prose. There are five document types: source, subject, institution, programme and requirement. A requirement holds one programme's rules for one session:
| Field | Why it's data, not text |
|---|---|
utmeCompulsory, utmeChoices ({pick: 1, from: [subject refs]}) |
"English, Maths, Physics + Chemistry or Biology" becomes a set problem the code can solve, without guessing from a sentence |
olevelCompulsory with a minGrade per subject, plus olevelChoices
|
Further Maths at UNILAG CS is just one more compulsory subject |
olevelMinCredits, olevelMinCreditsCombined, olevelMaxSittings, olevelAcceptedExams
|
"Five credits at one sitting, or six at two" (UI) can only be checked if it's modelled |
utmeMinScore (nullable) |
UI publishes no 2026/27 minimum, so null says "unknown", not 0 |
citations[] (a source reference plus a locator such as "p. 23, COMPUTER SCIENCE row") |
Every rule traces to a page you can open |
verificationStatus + conflictNote
|
When JAMB's brochure and the university disagree, I store the stricter rule and explain both |
Subjects are referenced by _id, with aliases, so "Use of English" vs "English Language" can't break a match.
The dataset has 34 requirements across 5 universities. 13 are verified field by field against their sources, and 21 are marked conflicting because official sources really do disagree, which is the whole problem. All of it is built from 57 saved sources (51 official): JAMB's IBASS brochure API, JAMB's brochure PDFs, and each university's 2026 notices.
2. Two Sanity Context endpoints, and why there are two
A Context MCP endpoint serves one kind of source: if you attach a dataset and a Knowledge Base together, the dataset wins and the KB is silently ignored. So the agent connects to two endpoints and prefixes their tools:
| Endpoint | Source | Tools the agent uses |
|---|---|---|
clearance-rules |
dataset cynv9mfk.production, with a GROQ filter to the 5 types |
rules_groq_query, rules_schema_explorer
|
clearance-policy |
the Knowledge Base |
policy_knowledge_base_read, policy_knowledge_base_search
|
Following Sanity's own pattern, both endpoints' initial_context is fetched over HTTP and put into the system prompt. The agent starts out knowing the schema and the KB outline without spending a tool call.
Each endpoint also has Instructions. For the rules endpoint, the instructions say to:
- match subjects by
_id, never by name - treat choice groups as "pick N"
- always return
verificationStatus,conflictNoteand citations - never decide eligibility itself, and instead pass requirement
_ids to the evaluator
3. What the agent actually does
One loop (Vercel AI SDK 6 + Claude Sonnet 5.5):
-
rules_groq_queryfinds the requirement documents. In check mode that's the programme's requirement. In explore mode the model writes GROQ likecount(utmeCompulsory[@._ref in [...your UTME subjects]]) == count(utmeCompulsory). -
evaluate_eligibility, a local tool, loads those documents and runs the evaluator against your results, which the server holds so the model can never retype them. It checks:- UTME subjects, using bipartite matching for choice groups
- the UTME score
- every combination of your sittings up to the course's limit
- credits, accepted exams and awaited results
-
policy_knowledge_base_readreads the KB entries behind each failed or uncertain check, in one call. -
submit_verdictis a tool with noexecute, so calling it ends the loop. The model's explanation is merged with the evaluator's verdicts.
The route streams the loop as it runs: each Sanity Context step as it finishes, then the verdict as soon as evaluate_eligibility returns, before the explanation is written. On a phone you watch the rules query and the checks land, and the decided verdict shows up in about half the total time.
Some guarantees are enforced in code, not just in the prompt:
- a policy note is dropped unless its KB path was actually read in that run
- check mode can't evaluate a different course
- if the model never finishes, you still get the exact verdicts
Every answer ships with its trace:
4. The Knowledge Base
The rules say what; the Knowledge Base says why it matters and what to do: cut-off marks, Post-UTME screening, awaiting-result windows, upload deadlines, sitting rules. I built it from 32 sources: 26 official (JAMB, plus all five universities' notices and requirement PDFs) and 6 blogs. The blogs are in deliberately, so Context would surface where they disagree with official sources.
Context found 7 conflicts. I resolved 6 and dismissed 1 as a false conflict. Each resolution became a standing instruction. For example:
- LASU's 195: a blog called it a "cut-off mark"; LASU's own notices say "a minimum of 195 marks". Official wording won.
- UNILAG sittings: UNILAG requires five O'level credits at one sitting only. This is the rule behind the demo's "Not eligible".
- UNILAG's O'level upload deadline: the extension notice (Monday, 24 August 2026) beats the original date.
- OAU English: an OAU page says "a pass at O-Level … in English Language". I kept the stricter reading, a full credit, because a pass would get a candidate rejected if OAU means a credit pass.
- UNILAG's lowest merit cut-off: Education Economics (49.65), not Meteorology as an entry claimed.
I wrote 4 instructions by hand:
- Official JAMB sources outrank blogs.
- A university's own published requirement is ground truth over summaries.
- A rule for what an entry must do when JAMB's brochure and a university's own requirement disagree.
- Always name the admission session.
The Knowledge Base answers questions directly too. Under every result there's an "Ask about the admission policy" box with suggested questions for that school, such as "What is the deadline to upload my O'level result for UNILAG?". A second agent answers them:
- It reads KB entries through
clearance-policyand must cite the paths it read. Citations it didn't read are dropped in code. - If the KB doesn't cover the question, it says so ("answered: false") instead of guessing.
- It gets today's date, so it can say when a deadline has already passed.
Building this exposed a real Knowledge Base problem. Its post_utme_screening entry still gave UNILAG's original upload deadline (14 August), even though I had resolved that conflict in favour of the extension (24 August). I fixed it at the source, with Context's "Rewrite this part" on that paragraph. That creates a standing rule every future build honours: "UNILAG's 2026/2027 O'level upload deadline … is Monday, 24 August 2026 … extended from the original Friday, 14 August 2026." The entry rebuilt with the corrected line, and every other fact on the page survived. The follow-up agent also keeps a general safeguard: when entries give different dates, the later notice wins.
5. Would keyword search get the same answer? The eval
The organisers asked exactly this, so I measured it. I wrote 15 cases full of traps:
- UNILAG Computer Science's Further Maths credit
- one sitting vs two
- UI's "6 credits at two sittings"
- NABTEB results
- a UTME score of 196 against minimums of 195 and 200
- a course outside the data
An independent agent that could only read the original source files set the expected verdict for each case, with quoted evidence. It couldn't see my dataset or code. Each case then ran through three systems:
- Clearance Desk on production.
-
The same model with Knowledge Base search. It gets
knowledge_base_search/knowledge_base_readon the same Knowledge Base, plus its outline: keyword search over the exact same content. - The same model with no tools.
Both baselines got more thinking time than the agent. Here is run 4; run 3 had the same Clearance Desk score.
| Clearance Desk | Same model + KB search | Same model, no tools | |
|---|---|---|---|
| Correct | 13 / 15 (13 in run 3 too) | 6 / 15 (4 in run 3) | 6 / 15 (8 in run 3) |
| Told a candidate who fails a published rule "Eligible" | 0 (0) | 4 (2) | 3 (4) |
Keyword search doesn't save the model. With the KB it searched, read the right entries, even quoted "five credits at one sitting", and still told Chioma (two sittings, UNILAG Medicine) she was eligible. It also missed UNILAG CS's Further Maths and UI's six-credit rule. Reading a rule isn't applying it. Structured rules plus code that applies them is the difference.
The eval also caught three of my own bugs, all fixed and logged:
- A data bug: LASU's sources only ever say "SSCE (or equivalent)", but I had encoded that as WAEC/NECO, so a NABTEB candidate was wrongly rejected. Run 1 scored 12/15 because of it.
- A rate-limiter bug: it counted its own refusals, so retrying kept extending the lockout.
- A parser bug in the eval harness itself.
Clearance Desk's two remaining misses are deliberate:
- It marks every requirement with conflicting official sources "At risk", even when the candidate meets both versions.
- It says "no data" for a course outside its data instead of guessing.
Full results, every answer, and earlier runs
Sanity Project Details
-
Project ID:
cynv9mfk· Dataset:production(public) -
Public dataset query (every requirement):
*[_type=="requirement"]{_id,session,verificationStatus} - Live coverage counts are on the About page.
Agent Session
The whole thing was built with Claude Code, phase by phase. That covered:
- the schema and the source-traced data pipeline
- setting up the Knowledge Base and both Context endpoints, through the browser
- the evaluator and the agent
- the UI, the deploy and the eval
Every wrong turn, and how it was fixed, is in the build log. The best moments are when the eval caught my own LASU data bug, and when the follow-up questions exposed a stale date in the Knowledge Base.
Limitations
- Coverage: 5 universities, 34 programmes, one session (2026/2027), UTME entry only. Direct Entry, Post-UTME scores, aggregate scores and catchment quotas aren't calculated, and meeting the minimum never guarantees admission.
- Conflicts: 21 of 34 requirements have official sources that disagree. Clearance Desk stores the stricter rule and says "At risk", which can be over-cautious, as eval case C13 shows.
- Unpublished minimums: UI publishes no 2026/27 UTME minimum, so a UI course can never come out plain "Eligible".
- Unchecked conditions: age, first choice and upload deadlines can't be checked from results. They're listed as "check these yourself". Cambridge O'Level isn't modelled.
- Explanations: the verdict comes from code, but the explanation comes from a model. It occasionally adds generic advice no source states, which the eval and the build log both note.
- The Knowledge Base is in beta (up to 150 documents). This one uses 32 sources.









Top comments (0)