DEV Community

Ashiku Sheban
Ashiku Sheban

Posted on

Building the parts of an AI tutor that don't fit in a demo

title: Building the parts of an AI tutor that don't fit in a demo
published: false
tags: ai, javascript, webdev, discuss

A few weeks ago I wrote about a debugging tutor that refuses to give students the answer. It asks questions. It requires a hypothesis before every hint. It points at the general area of a bug and lets the student figure out the rest.

That post ended with a list of features I'd add if I kept going. This is what happened when I did.

Four things shipped. Three were on the original list. One — auth — turned out to be a prerequisite I hadn't planned for.

Post-mortem turns

After a student fixes their bug, the tutor asks one more question: "In your own words, why did the bug happen?"

Their answer gets scored — via an LLM call when the key is set, or a keyword-matching heuristic otherwise — and stored against the bug pattern. The weak-spots dashboard now shows not just "you hit off-by-one five times" but "you correctly explained the cause four out of five times."

That second number is the honest one. It's the difference between "the hints helped" and "they stumbled into a fix." Without it, every pass looks the same. With it, an instructor can see which exercises students are actually learning from and which ones they're getting through by luck.

The heuristic scorer isn't good. It reads the student's explanation for words like "because" and "one too many" and marks it partial if it sees them. The LLM scorer — same prompt, same response shape — does the real work. I built the heuristic because I wanted the pipeline to function without an API key, not because I thought it would be accurate.

Cross-student fingerprinting

The same pattern data, aggregated across everyone. If 75% of a class hits an off-by-one, that's not a student problem — it's a teaching signal.

The instructor dashboard surfaces it as a callout:

"off-by-one" appears in 75% of students (6 of 8)
4 students hit null-undefined on exercise ex-1
Exercise "ex-1" has a 20% completion rate across 5 sessions

Those are actual output from my test data. Three sentences, and an instructor knows which exercise to walk through tomorrow.

The interesting thing about this feature is that it's invisible from any single-student view. Students only see their own patterns. Instructors only see the aggregate. Nobody can see both, which means the class-level insight genuinely can't be reconstructed from what a student sees. That's what makes it worth building.

The implementation is three SQL aggregations over tables that already existed. No schema changes. That's the payoff of building the persistence layer properly the first time — the second feature is a query away.

Struggle timer

Lock hint level 1 for N minutes of independent attempts before it's even offered. Per-exercise config; 0 minutes disables it.

The response when the timer is active is a soft refusal with a countdown, not an error. The server returns 200 with { requiresStruggle: true, remainingSeconds: 143, ... }, and the UI renders a yellow card:

PRODUCTIVE STRUGGLE
Give it a bit more time. You'll get more out of this if you try on your own first.
2:23 of focused time

The frame matters more than the timer. This feature could easily feel punitive — "you can't have help yet" — and I spent more time on the copy than on the code. The version that worked reads like a coach, not a bouncer.

Two design details worth calling out:

You can't game it by waiting. The timer requires at least one prior code submission. If a student opens the exercise, walks away for ten minutes, and comes back, the timer has expired but they still haven't engaged. The gate checks both.

The timer never disables the button.** The client always lets the student click "Ask for a hint." The server decides. If you refresh mid-countdown, you see the same countdown — the state is on the server, not in the UI. This is a small thing but it's the right default for anything involving timing.

Auth — the thing I didn't plan for

The original app was one page with four tabs and a shared x-admin-key gate on the admin endpoints. That was fine for a demo. For actual instructors to use it, they need individual logins, and their view needs to be scoped to their own students.

So I built:

  • users and auth_sessions tables
  • bcrypt-hashed passwords, 12 rounds
  • Random 32-byte session tokens with 30-day expiry
  • HTTP-only cookies, SameSite=Lax, Secure in production
  • Register / login / logout / me endpoints
  • Instructor auto-seeded on first boot from an env var

Then I gated every existing route. The debug tutor endpoints require authentication; the admin endpoints require the instructor role specifically.

The most important change: the server never trusts the client's studentId anymore. It derives it from the session cookie. The old version had a text field in the top bar where a student typed their ID — which meant a malicious student could read another student's history by changing the field. That was a real hole, and it took adding auth to see how bad it was.

Instructors can impersonate a specific student via ?asStudentId= for testing, but only instructors, and only through the middleware.

The old x-admin-key system is gone. That was a shared secret in a .env file, which is the "our first customer is a friend who we trust" tier of security. The moment there's more than one instructor, it stops being viable.

What this bought: two dashboards from one app. Students see the Tutor and Weak Spots tabs. Instructors see those plus Health and Class. Same code, different views, driven by the role on the session.

What's still on the list

  • Cohorts. The class fingerprint view currently shows every student in the system. Real classes have boundaries. A cohorts table with an instructor field, and every aggregation filtered by it. This is the piece that makes multi-instructor deployment meaningful.
  • Exercise authoring. The struggle timer config is a hardcoded TypeScript map. It wants to be an exercises table with a real admin UI, when there's a reason to build one.
  • Deployment. Still localhost:3001. A VPS, Caddy, and a domain is a weekend project.
  • The LLM key. The fallback path works, but the LLM-generated hints — the ones that respond to the student's specific hypothesis and code — are the reason this exists. That's the next thing to turn on.

The state of things

The repo is at github.com/janabi54/codeteach-debug-tutor. Around 3,500 lines of TypeScript and vanilla JS. 23 regression tests. SQLite persistence across everything. Working end to end, login included.

None of it is clever. Most of it is the boring kind of feature that a demo doesn't need: users, sessions, migrations, aggregations. That's the work that turns a prototype into something an instructor might actually use, and it's the work that isn't fun to write about.

Top comments (0)