This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
What I Built
Mot is a friendly career assistant built on O*NET 31.0, the U.S. Department of Labor's database of about 1,000 occupations. For every job, O*NET records the tasks, the skills and knowledge it needs (each rated for importance on a 1–5 scale and level on a 0–7 scale), the typical education, the work styles, and an interest profile.
Ask a general-purpose chatbot "what does it take to become a data scientist?" and you get a confident, plausible answer that nobody can check. Career advice is a place where a wrong answer costs real time and money, so I wanted an agent that only says what the data says, and links to the source every time.
Mot has three agents, each on its own page:
-
Career Explorer (
/explore): chat about careers. Describe what you like ("working outdoors with animals") or name a job, and Mot finds matching occupations, explains them, compares two jobs as a career move, and suggests your next question. You can bookmark occupations and compare 2–4 of them side by side. -
Mock Interview Coach (
/interview): pick a job, a number of questions, and a focus (behavioral, skills, or mixed). Mot interviews you one question at a time and grades each answer against O*NET's own examples of what each skill level looks like, then writes a final report. -
Interest Quiz (
/quiz): a two-round version of the O*NET Interest Profiler. Your ratings become a RIASEC profile (Realistic, Investigative, Artistic, Social, Enterprising, Conventional) that is matched against every occupation.
Why this only works because the content is structured
The challenge asks for "an agent that only works because the content was structured." Here's where that's literally true:
-
Interview grading uses Level Scale Anchors. O*NET doesn't just say Programming matters for Software Developers. It says the job needs Programming at level 4.x on a 0–7 scale, and it publishes concrete anchors for levels 2, 4, and 6 (Programming at level 2 is "Write a program to sort objects in a database"). The coach's
getInterviewBrieftool joins an occupation's ratings with the matchingonetLevelScaleAnchordocuments, so the model pitches each question at the required level and quotes the closest anchor when it scores the answer. A keyword search over job descriptions can't give you "the level-4 example for this skill on this scale." -
Career moves are computed, not guessed.
compareOccupationsdiffs the rating arrays of two occupations and returns the skill, knowledge, and ability gaps, the shared strengths, the Job Zone change, and the new technologies. The model turns those gaps into next steps instead of inventing them. -
Quiz scoring is deterministic, and the model never sees your ratings. The rating card is a client-side tool. The server reads the ratings from the UI messages and scores them in code: each RIASEC type gets
1 + 6 × mean, the same 1–7 scale O*NET uses for occupations, and the match percent is 70% the Pearson correlation with each occupation's interest profile and 30% how much you liked its strongest Specific Interest Areas. The model narrates the result. It can't make one up. -
Every occupation link carries its O*NET-SOC code. The system prompt requires links like
[Data Scientists (15-2051.00)](https://www.onetonline.org/...), and the UI parses that format to add a bookmark toggle and a "Practice mock interview" button next to each occupation.
Demo
Live app: https://career-exploration-two.vercel.app/ (there's a quick Cloudflare Turnstile human check before the agents start, to protect the Gemini quota)
Things to try:
- On
/explore: "I like helping people and solving puzzles, but I don't want a 4-year degree." Then: "How would I move from Medical Assistant to Registered Nurse?" - On
/interview: "Data Scientist", 3 questions, skills focus. Give one vague answer and one specific one, and compare the feedback cards. - On
/quiz: finish both rounds, then ask Mot about one of your matches.
Code
The project is four GitHub repos: three apps plus a root workspace that pins them as submodules and holds the shared Sanity Function.
- Workspace (Sanity Function, Blueprint, docs): https://github.com/sea2709/career-exploration
- Agent service (Node, AI SDK
ToolLoopAgent, Gemini): https://github.com/sea2709/career-exploration-agent - Web app (Astro 7 + React 19): https://github.com/sea2709/career-exploration-web
- Sanity Studio (schemas, O*NET importer, Knowledge Base and workflow setup scripts): https://github.com/sea2709/career-exploration-studio
Browser ──▶ web (Astro) ──/api/*──▶ agent (Node) ──▶ Gemini
│
├─ GROQ ─────────▶ Sanity dataset (O*NET, public CDN)
├─ Context MCP ──▶ career-explorer endpoint (full dataset, embeddings)
├─ Context MCP ──▶ interview-coaching endpoint (Knowledge Base)
└─ Insights ─────▶ org Context store
▲
studio (Sanity Studio) ── schemas, O*NET import, coaching guides │
functions/ ── weekly classify-conversations job ─────────────────┘
How I Used Sanity
1. The O*NET dataset as structured content
I imported O*NET 31.0 into the production dataset with a custom importer in the Studio repo (pnpm import:onet). The modeling rule was simple: reference tables and occupations become documents, and per-occupation rows are embedded as arrays on the occupation. One onetOccupation document holds its tasks, job titles, software skills, work styles, ratings, related occupations, and interest profile, with references to its Job Zone and to the content model elements its ratings describe.
| Document type | Count |
|---|---|
onetOccupation |
1,016 |
onetLevelScaleAnchor |
483 |
onetInterest (RIASEC types and Specific Interest Areas) |
47 |
coachingGuide |
20 |
careerQuiz |
8 |
Plus reference types: onetContentModelElement, onetScale, onetRatingCategory, and onetJobZone.
2. Career Explorer: a Context MCP endpoint over the full dataset
The explorer connects to a dataset-backed Sanity Context MCP endpoint (career-explorer) over the whole O*NET dataset, with dataset embeddings enabled and scoped by a projection to onetOccupation (title, description, alternate titles, tasks).
-
initial_context: fetched once per endpoint, cached for 5 minutes, and injected into the system prompt. The agent then drops theinitial_contexttool, which saves a step on every conversation. -
groq_queryandschema_explorer: for the long tail. Questions like "which Job Zone 2 occupations score highest on Social?" don't fit a fixed tool, so the model writes GROQ against a schema it already knows from the initial context. -
Local GROQ tools first: for the common questions I wrote four fast tools (
searchOccupations,getOccupationProfile,compareOccupations,getRelatedOccupations) that query the public CDN directly and return compact, pre-shaped results. The prompt tells the model to prefer them and fall back togroq_query.
searchOccupations is hybrid. It runs a keyword ranking (match on title, alternate titles, and description) and a semantic ranking (text::semanticSimilarity()) in parallel, then merges the two lists with reciprocal rank fusion. I first tried blending both into one query, but semantic scores sit in a narrow band (about 6.8 to 8.3) while keyword scores range from 5 to 50, so boosting the semantic term barely changed the order. Merging by rank position worked. Results from the real tool:
| Query | Keyword only | Hybrid |
|---|---|---|
| working outdoors with animals | nothing | Farming supervisors, Farmworkers, Animal Trainers, Animal Caretakers |
| nurse | Nurse Practitioners first, Registered Nurses 6th | Registered Nurses first |
| AI specialist | Animal Breeders 2nd ("AI" as in artificial insemination) | Animal Breeders gone from the top 6 |
| software developer | Software Developers first | Software Developers first |
The last row is why I kept the keyword ranking: semantic search alone didn't have Software Developers in its top 8. If the semantic query fails (embeddings off, or the monthly quota used up), the tool logs a warning and returns keyword results only.
3. Mock Interview Coach: a Knowledge Base built from counselor-written guides
The interview coach needs two kinds of knowledge: what the job requires (O*NET, exact) and how to run a good interview (judgment). I kept them apart on purpose.
-
What the job requires comes from the O*NET brief (
getInterviewBrief), which is always the source of truth. -
How to interview comes from a Sanity Context Knowledge Base, "Interview coaching guidance". Career counselors write
coachingGuidedocuments in the Studio: answer structure, behavioral and skills questions, work styles, how to grade answers, how to give feedback, special situations (career changers, entry-level candidates, senior roles, employment gaps), and practice ideas. The Knowledge Base imports the published guides only, through a GROQ filter, and distills them into a navigable outline of entries. The 16 starter guides became 12 entries because the build merged the overlapping ones.
The coach connects to a second, Knowledge Base-backed MCP endpoint (interview-coaching). Its outline goes into the system prompt, and the coach reads entries with the endpoint's knowledge_base_read tool at specific moments:
- at the start, after the brief: the entries on the requested focus and on rating answers, plus the entry for the role's situation when its Job Zone is 1–2 or 4–5 (at most three reads);
- before writing its first score and the final report: the feedback guidance;
- whenever the candidate mentions a career change, a gap, or little experience: the matching entry, before replying.
Two rules keep the coach honest. If the guidance and the brief seem to disagree, the brief wins. And the coach applies the guidance without ever mentioning a knowledge base to the candidate. If the endpoint is unset or unreachable, the coach logs it and runs on the O*NET brief alone.
On contradictions: when I published a new guide about which kinds of experience count as evidence for entry-level candidates, the Knowledge Base refresh filed an issue saying it overlapped the existing "Coaching Entry-Level Candidates" entry and proposed folding it in. I found that issue through the API rather than the dashboard, because the dashboard's Issues page didn't list issues from a refresh. So I changed my pnpm kb:coaching script to print every open issue with its kind, severity, entry path, finding, and suggested fix.
Because a bad guide would go straight into the coach, guides go through a Sanity Workflow before publishing: Draft → Counselor review → Approved. The counselor-review tasks block the guide from moving on until they're done ("check that the Job Zones match the advice", "check the guide agrees with the rating scale"). A Retired off-ramp unpublishes a guide, which takes it out of the Knowledge Base at the next refresh.
4. Learning from real conversations
When the Insights variables are set, the explorer saves each conversation to the organization's Context store through Conversation Insights. A scheduled Sanity Function (classify-conversations, deployed with a Blueprint) runs every Monday at 06:00 Central and classifies up to 500 transcripts with Gemini, so I can see which kinds of questions people ask and where the agent struggles.
Sanity Project Details
-
Project ID:
rhq335ze -
Dataset:
production(public, readable through the CDN without a token) -
Schemas:
studio/schemaTypes
Try a query in the browser. This one returns Data Scientists with its Job Zone and RIASEC profile:
Agent Session
This session is how I turned on dataset embeddings for the Career Explorer. I asked "how to turn on embeddings for onetOccupation?" The agent found the current dataset embeddings docs (after first landing on the deprecated Embeddings Index API) and confirmed embeddings were off for production. It then designed a projection that embeds only onetOccupation: the title as occupation, the description, the alternate titles as alsoCalled, and the tasks. Without a projection, Sanity embeds the whole document, and each occupation's hundreds of numeric ratings would push the meaningful text out of the 10-chunk limit.
I built this in Cursor, which DEV's Agent Sessions don't support yet, so I converted the transcript to the Claude Code format with a small script and DEV labels it that way. Cursor doesn't save tool outputs in its transcripts, so the tool calls appear without results.
The build story, including what went wrong, is in my Path Two post.
Data: O*NET 31.0 Database by the U.S. Department of Labor, Employment and Training Administration (USDOL/ETA), used under the CC BY 4.0 license. O*NET® is a trademark of USDOL/ETA. USDOL/ETA has not approved, endorsed, or tested this app.
Top comments (0)