DEV Community

Dang Tran
Dang Tran

Posted on

Mot: a career coach that grades your interview answers against O*NET, not vibes

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Mot is a friendly career assistant built on O*NET 31.0, the U.S. Department of Labor's database of about 1,000 occupations. For every job, O*NET records the tasks, the skills and knowledge it needs (each rated for importance on a 1–5 scale and level on a 0–7 scale), the typical education, the work styles, and an interest profile.

Ask a general-purpose chatbot "what does it take to become a data scientist?" and you get a confident, plausible answer that nobody can check. Career advice is a place where a wrong answer costs real time and money, so I wanted an agent that only says what the data says, and links to the source every time.

Mot has three agents, each on its own page:

  • Career Explorer (/explore): chat about careers. Describe what you like ("working outdoors with animals") or name a job, and Mot finds matching occupations, explains them, compares two jobs as a career move, and suggests your next question. You can bookmark occupations and compare 2–4 of them side by side.
  • Mock Interview Coach (/interview): pick a job, a number of questions, and a focus (behavioral, skills, or mixed). Mot interviews you one question at a time and grades each answer against O*NET's own examples of what each skill level looks like, then writes a final report.
  • Interest Quiz (/quiz): a two-round version of the O*NET Interest Profiler. Your ratings become a RIASEC profile (Realistic, Investigative, Artistic, Social, Enterprising, Conventional) that is matched against every occupation.

Why this only works because the content is structured

The challenge asks for "an agent that only works because the content was structured." Here's where that's literally true:

  1. Interview grading uses Level Scale Anchors. O*NET doesn't just say Programming matters for Software Developers. It says the job needs Programming at level 4.x on a 0–7 scale, and it publishes concrete anchors for levels 2, 4, and 6 (Programming at level 2 is "Write a program to sort objects in a database"). The coach's getInterviewBrief tool joins an occupation's ratings with the matching onetLevelScaleAnchor documents, so the model pitches each question at the required level and quotes the closest anchor when it scores the answer. A keyword search over job descriptions can't give you "the level-4 example for this skill on this scale."
  2. Career moves are computed, not guessed. compareOccupations diffs the rating arrays of two occupations and returns the skill, knowledge, and ability gaps, the shared strengths, the Job Zone change, and the new technologies. The model turns those gaps into next steps instead of inventing them.
  3. Quiz scoring is deterministic, and the model never sees your ratings. The rating card is a client-side tool. The server reads the ratings from the UI messages and scores them in code: each RIASEC type gets 1 + 6 × mean, the same 1–7 scale O*NET uses for occupations, and the match percent is 70% the Pearson correlation with each occupation's interest profile and 30% how much you liked its strongest Specific Interest Areas. The model narrates the result. It can't make one up.
  4. Every occupation link carries its O*NET-SOC code. The system prompt requires links like [Data Scientists (15-2051.00)](https://www.onetonline.org/...), and the UI parses that format to add a bookmark toggle and a "Practice mock interview" button next to each occupation.

Demo

Live app: https://career-exploration-two.vercel.app/ (there's a quick Cloudflare Turnstile human check before the agents start, to protect the Gemini quota)

Things to try:

  • On /explore: "I like helping people and solving puzzles, but I don't want a 4-year degree." Then: "How would I move from Medical Assistant to Registered Nurse?"
  • On /interview: "Data Scientist", 3 questions, skills focus. Give one vague answer and one specific one, and compare the feedback cards.
  • On /quiz: finish both rounds, then ask Mot about one of your matches.

Code

The project is four GitHub repos: three apps plus a root workspace that pins them as submodules and holds the shared Sanity Function.

Browser ──▶ web (Astro) ──/api/*──▶ agent (Node) ──▶ Gemini
                                        │
                                        ├─ GROQ ─────────▶ Sanity dataset (O*NET, public CDN)
                                        ├─ Context MCP ──▶ career-explorer endpoint (full dataset, embeddings)
                                        ├─ Context MCP ──▶ interview-coaching endpoint (Knowledge Base)
                                        └─ Insights ─────▶ org Context store
                                                                ▲
studio (Sanity Studio) ── schemas, O*NET import, coaching guides │
functions/ ── weekly classify-conversations job ─────────────────┘
Enter fullscreen mode Exit fullscreen mode

How I Used Sanity

1. The O*NET dataset as structured content

I imported O*NET 31.0 into the production dataset with a custom importer in the Studio repo (pnpm import:onet). The modeling rule was simple: reference tables and occupations become documents, and per-occupation rows are embedded as arrays on the occupation. One onetOccupation document holds its tasks, job titles, software skills, work styles, ratings, related occupations, and interest profile, with references to its Job Zone and to the content model elements its ratings describe.

Document type Count
onetOccupation 1,016
onetLevelScaleAnchor 483
onetInterest (RIASEC types and Specific Interest Areas) 47
coachingGuide 20
careerQuiz 8

Plus reference types: onetContentModelElement, onetScale, onetRatingCategory, and onetJobZone.

2. Career Explorer: a Context MCP endpoint over the full dataset

The explorer connects to a dataset-backed Sanity Context MCP endpoint (career-explorer) over the whole O*NET dataset, with dataset embeddings enabled and scoped by a projection to onetOccupation (title, description, alternate titles, tasks).

  • initial_context: fetched once per endpoint, cached for 5 minutes, and injected into the system prompt. The agent then drops the initial_context tool, which saves a step on every conversation.
  • groq_query and schema_explorer: for the long tail. Questions like "which Job Zone 2 occupations score highest on Social?" don't fit a fixed tool, so the model writes GROQ against a schema it already knows from the initial context.
  • Local GROQ tools first: for the common questions I wrote four fast tools (searchOccupations, getOccupationProfile, compareOccupations, getRelatedOccupations) that query the public CDN directly and return compact, pre-shaped results. The prompt tells the model to prefer them and fall back to groq_query.

searchOccupations is hybrid. It runs a keyword ranking (match on title, alternate titles, and description) and a semantic ranking (text::semanticSimilarity()) in parallel, then merges the two lists with reciprocal rank fusion. I first tried blending both into one query, but semantic scores sit in a narrow band (about 6.8 to 8.3) while keyword scores range from 5 to 50, so boosting the semantic term barely changed the order. Merging by rank position worked. Results from the real tool:

Query Keyword only Hybrid
working outdoors with animals nothing Farming supervisors, Farmworkers, Animal Trainers, Animal Caretakers
nurse Nurse Practitioners first, Registered Nurses 6th Registered Nurses first
AI specialist Animal Breeders 2nd ("AI" as in artificial insemination) Animal Breeders gone from the top 6
software developer Software Developers first Software Developers first

The last row is why I kept the keyword ranking: semantic search alone didn't have Software Developers in its top 8. If the semantic query fails (embeddings off, or the monthly quota used up), the tool logs a warning and returns keyword results only.

3. Mock Interview Coach: a Knowledge Base built from counselor-written guides

The interview coach needs two kinds of knowledge: what the job requires (O*NET, exact) and how to run a good interview (judgment). I kept them apart on purpose.

  • What the job requires comes from the O*NET brief (getInterviewBrief), which is always the source of truth.
  • How to interview comes from a Sanity Context Knowledge Base, "Interview coaching guidance". Career counselors write coachingGuide documents in the Studio: answer structure, behavioral and skills questions, work styles, how to grade answers, how to give feedback, special situations (career changers, entry-level candidates, senior roles, employment gaps), and practice ideas. The Knowledge Base imports the published guides only, through a GROQ filter, and distills them into a navigable outline of entries. The 16 starter guides became 12 entries because the build merged the overlapping ones.

The coach connects to a second, Knowledge Base-backed MCP endpoint (interview-coaching). Its outline goes into the system prompt, and the coach reads entries with the endpoint's knowledge_base_read tool at specific moments:

  • at the start, after the brief: the entries on the requested focus and on rating answers, plus the entry for the role's situation when its Job Zone is 1–2 or 4–5 (at most three reads);
  • before writing its first score and the final report: the feedback guidance;
  • whenever the candidate mentions a career change, a gap, or little experience: the matching entry, before replying.

Two rules keep the coach honest. If the guidance and the brief seem to disagree, the brief wins. And the coach applies the guidance without ever mentioning a knowledge base to the candidate. If the endpoint is unset or unreachable, the coach logs it and runs on the O*NET brief alone.

On contradictions: when I published a new guide about which kinds of experience count as evidence for entry-level candidates, the Knowledge Base refresh filed an issue saying it overlapped the existing "Coaching Entry-Level Candidates" entry and proposed folding it in. I found that issue through the API rather than the dashboard, because the dashboard's Issues page didn't list issues from a refresh. So I changed my pnpm kb:coaching script to print every open issue with its kind, severity, entry path, finding, and suggested fix.

Because a bad guide would go straight into the coach, guides go through a Sanity Workflow before publishing: Draft → Counselor review → Approved. The counselor-review tasks block the guide from moving on until they're done ("check that the Job Zones match the advice", "check the guide agrees with the rating scale"). A Retired off-ramp unpublishes a guide, which takes it out of the Knowledge Base at the next refresh.

4. Learning from real conversations

When the Insights variables are set, the explorer saves each conversation to the organization's Context store through Conversation Insights. A scheduled Sanity Function (classify-conversations, deployed with a Blueprint) runs every Monday at 06:00 Central and classifies up to 500 transcripts with Gemini, so I can see which kinds of questions people ask and where the agent struggles.

Sanity Project Details

  • Project ID: rhq335ze
  • Dataset: production (public, readable through the CDN without a token)
  • Schemas: studio/schemaTypes

Try a query in the browser. This one returns Data Scientists with its Job Zone and RIASEC profile:

https://rhq335ze.apicdn.sanity.io/v2025-02-19/data/query/production?query=%2A%5B_type%3D%3D%22onetOccupation%22%20%26%26%20onetsocCode%3D%3D%2215-2051.00%22%5D%5B0%5D%7Btitle%2C%20onetsocCode%2C%20%22jobZone%22%3A%20jobZone-%3E%7BjobZone%2C%20name%7D%2C%20%22riasec%22%3A%20interestProfile.riasec%7D

Agent Session

This session is how I turned on dataset embeddings for the Career Explorer. I asked "how to turn on embeddings for onetOccupation?" The agent found the current dataset embeddings docs (after first landing on the deprecated Embeddings Index API) and confirmed embeddings were off for production. It then designed a projection that embeds only onetOccupation: the title as occupation, the description, the alternate titles as alsoCalled, and the tasks. Without a projection, Sanity embeds the whole document, and each occupation's hundreds of numeric ratings would push the meaningful text out of the 10-chunk limit.

Sanity Challenge - Part One - Enable Embeddings cursor-agent
You

how to turn on embeddings for onetOccupation?

Agent

I'll check the vendored Sanity Context skills and the docs for how embeddings are enabled.

I need accurate info on enabling embeddings for Sanity Context, so I'll check the vendored agent skills from sanity-io/context to see if they document this.

Agent
Agent
Agent
Agent

That page is the deprecated Embeddings Index API. I'll find the current dataset embeddings docs.

I need the dataset embeddings documentation, so I'll fetch the llms.txt file and search it for embeddings-related content.

Agent
Agent
Agent
Agent
Agent

Since embeddings are dataset-level but can be scoped to onetOccupation via a type-specific projection, I need to check the onetOccupation field names in the schema. I should also consider whether coachingGuide docs could be included too, and verify whether the Sanity CLI version in studio supports the datasets embeddings command before checking current status.

Since embeddings are set at the dataset level via a projection, I'll check which onetOccupation fields are worth embedding and see if the Studio CLI supports that command.

Agent
Agent

Embeddings are currently off for production, and the Studio's CLI supports enabling them. I'll check the field names inside the task and job title arrays so the projection uses the right paths.

I'm worried about projection limits since expanding references won't help, so I need to be selective about fields—title, description, job titles, and tasks—since occupations can have dozens of tasks and job titles that could blow up document size.

Agent

I'll check whether the agent's own search tool would also need to change once embeddings are on.

Without a projection, embedding the entire document with all its arrays would blow past the chunk limit, so I need to trim it down to title, description, task names, and maybe alternate job titles — keeping it lean enough to fit comfortably within the chunk budget.

Field names like "occupation", "alsoCalled", and "tasks" give useful semantic context, and projections should work fine on array fields like tasks[].task since they're just paths within the document.

For next steps: enable the command, check status, then update the career-explorer endpoint instructions to remove the "semantic search not enabled" rule and add proper guidance — though I need to figure out whether that's editable via code or only through the Context app dashboard. I should also check whether any local agent tools like a search function need updating to leverage semantic scoring.

Agent
Agent

Embeddings are a setting on the whole production dataset, turned on with one Sanity CLI command. A projection limits embedding to onetOccupation documents. They're off right now, and the Studio's CLI already supports the command. I haven't run it, since it changes the live dataset. Based on the Dataset Embeddings docs:

1. Enable embeddings with a projection. Run this from studio/:

npx sanity datasets embeddings enable production --wait --projection '{
  _type == "onetOccupation" => {
    "occupation": title,
    description,
    "alsoCalled": jobTitles[].jobTitle,
    "tasks": tasks[].task
  }
}'

Why this projection:

  • Only occupations get embedded. Document types the projection doesn't list are skipped, so coaching guides, job zones, and other reference documents stay out. The 939 real occupations are what free-text questions need to match.
  • Without a projection, Sanity embeds the whole document. Each occupation carries hundreds of ratings[] items, and each task has nested ratings. Sanity embeds at most 10 chunks per document and drops the rest, so the numeric noise would push out the text that matters. Software Developers comes to about 5,000 characters with this projection, which fits comfortably.
  • Field names give the model context. The docs say field names are used as semantic context, which is why the projection renames title to "occupation" and the job titles to "alsoCalled".
  • Projections can't follow references. That's fine here, because Job Zone and ratings are references and wouldn't help semantic matching anyway.

2. Wait for it to finish. With --wait, the command returns once embedding is done. You can also check with:

npx sanity datasets embeddings status production

You want Status: ready.

3. Update the career-explorer endpoint instructions. The endpoint's Instructions field (in the Context app under MCP endpoints) currently tells the agent semantic search returns an error. Replace that rule with something like:

Semantic search is enabled for onetOccupation only (title, description, alternate titles, tasks). Use score(text::semanticSimilarity("...")) with order(_score desc) for descriptive questions ("I like working outdoors with animals"). Keep match or exact onetsocCode lookups for specific job titles and codes. Semantic search only ranks, so still apply the "All Other" and military filters.

4. Optionally, update the agent's search tool. searchOccupations in agent/src/onet/data.ts uses keyword match only, and its tool description says "Every word must match". That's the tool the agent reaches for first, so without a change it will keep doing keyword search. You could:

  • add a semantic term to its score() (hybrid search, keeping the title boost), or
  • add a separate tool for descriptive interest searches.

Then test whether text::semanticSimilarity() works over the public CDN client in agent/src/sanity.ts, since that client sends no token.

Some things to know first:

  • Cost: Generating embeddings is free on all plans, but each query using text::semanticSimilarity() counts against the organization's monthly semantic search quota.
  • Writes may slow down: The docs warn that datasets with embeddings can have slower writes. That would mainly affect the O*NET importer (npm run import:onet).
  • Disabling is destructive: It deletes the embeddings, and turning them back on recomputes everything.
  • The coaching Knowledge Base isn't affected: It's a separate index, and this doesn't change it.

If you want, I can run step 1 and then make the step 4 change to searchOccupations.

I should explain the steps rather than enabling the setting myself, since this is a production dataset change the user needs to approve. I also want to flag that semantic queries count against quota on the public CDN dataset, that it's unclear whether semanticSimilarity even works there without a token, and that re-importing O*NET data carries write-speed considerations worth testing first.

I built this in Cursor, which DEV's Agent Sessions don't support yet, so I converted the transcript to the Claude Code format with a small script and DEV labels it that way. Cursor doesn't save tool outputs in its transcripts, so the tool calls appear without results.

The build story, including what went wrong, is in my Path Two post.

Data: O*NET 31.0 Database by the U.S. Department of Labor, Employment and Training Administration (USDOL/ETA), used under the CC BY 4.0 license. O*NET® is a trademark of USDOL/ETA. USDOL/ETA has not approved, endorsed, or tested this app.

Top comments (0)