DEV Community

Cover image for CrossCheck, a study tool I built for a friend that runs on an open model
Manish Yadav
Manish Yadav

Posted on

CrossCheck, a study tool I built for a friend that runs on an open model

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

What I built

CrossCheck is a study tool. You paste your notes, it makes flashcards, and then it gives you an exam in two parts.

In Part 1 you solve things on your own and lock your answers. Only then Part 2 opens (so you can't copy steps from it). In Part 2 you check someone else's answer. About 3 out of 4 of those answers have one small mistake and the rest are fully correct, and you have to say which is which, why, and how sure you are.

It runs on my laptop with an open model. The notes never leave the computer.

Who it is for

It is for my best friend. We grew up together, so he is like a brother to me. He is in class 12 and this year he has his board exams, which are very important exams in India.

He used the "Add your own notes" feature to make a test on limits (it is his strong topic), took the test, and was really amazed by the app.

A normal quiz only gives a score like 5/10. It does not tell you if you can't do it, or if you can do it but trust things too easily. My friend told me he can do Part 1 type questions by following the steps his teachers taught, but most of his doubts are in theory. So I made the exam in two parts, to see both.

How it works

  1. You paste notes. The model splits them into small concepts.
  2. You revise with flashcards. They are closed during the exam.
  3. Part 1, solve it yourself and lock it. Part 2, check a worked answer or a statement.
  4. The app shows the real answer, why the wrong ones are wrong, and what to do next.

What to do next is not decided by the model. It is a small table of rules in plain Python:

Part 1 Part 2 Next step
good good go on, next exam has trickier mistakes
good weak explain why it works, then 2 more checking items
weak good 3 new practice problems with hints
weak weak go back to the flashcards
not enough items give 2 more items, no verdict yet

"Not enough evidence yet" is a real answer. With only 3 items, getting 2 right does not prove anything, so the app says so. I used a simple Beta estimate for this. For example 3 of 3 gives 0.87 (good), 2 of 3 gives 0.52 (not sure), and 1 of 3 gives 0.18 (weak).

How I built it: do not trust the model

A small open model makes mistakes. So I made sure it is never the judge of the facts.

  • Maths items are made by code, including the wrong solutions. The answer is checked a second way (numerical derivative or numerical integral). The model is not involved.
  • Theory items come from the student's own sentences. A correct item is their own sentence, unchanged. A wrong item is their sentence with one small edit, and I keep the original as the truth. Code checks the sentence is really in the notes, and checks the edit is small. Then a second model call checks that the edit really makes it false. If anything fails, that item is thrown away.
  • The model never writes code that I run.
  • The answers are never sent to the browser before you submit, and Part 2 is not sent at all until Part 1 is locked. There is a test for both.

The model does the language jobs: splitting notes, writing flashcards, picking sentences, proposing edits, and reading the student's written reasons (0, 1 or 2 points).

Some real numbers from my own testing (the app shows them at /api/stats): the model tried 51 edits to make a sentence false, and only 20 passed my checks. 31 were thrown away because the change was too big, or the second check said the sentence was not really false. Also, 13 of the 65 sentences the model picked were not actually in the notes, so code removed them. So a small open model does get things wrong, and these checks mattered.

Why open innovation matters

  • Privacy. A friend's notes and their exam mistakes stay on their own computer. Nothing goes to a server I don't control.
  • It works without internet. The model runs locally with Ollama. If the model is off, the calculus part still works fully, and the physics demo uses a small saved set. The app tells you which one it is using.
  • I can swap the model. It is one setting (LLM_MODEL), and any server with the OpenAI chat API works.
  • It costs nothing to run. My friend can retry 20 exams and nobody pays per question.

I did not test a closed model against it, so I can't say the open one is better. What I can say is that a small open model gets things wrong sometimes, and building around that is what made the design good.

Demo

The home page, with the model running (gemma3:4b) and a different quote every time you refresh:

CrossCheck home page showing the concepts for physics and calculus, with the local model gemma3:4b running

My friend pasted his own notes on limits, and it split them into concepts:

The Limits notes split into concepts like Limit Existence, Indeterminate Forms and L'Hopital's Rule, with the box to add your own notes below

Flashcards made from the notes. They are closed during the exam:

Flashcards for the Indeterminate Forms concept, each card with a rule, why, definition or watch-out label

Part 1, solve it yourself. Part 2 stays hidden until you lock this:

Part 1 of the exam with two written questions and a Lock Part 1 button

Part 2, check someone else's work. Say correct or not, why, and how sure you are:

Part 2 of the exam showing one statement to judge as correct or not correct, with boxes for the reason and a fix

The result. It says "not enough evidence yet" when there are too few items, and it points out when you were confident but wrong:

Results page with a confident but wrong warning, a not enough evidence yet message, and scores for Part 1 accuracy, mistakes caught, false alarms and wrong when sure

After the test, each answer is shown with the reference answer and a short reason for the score:

Item by item review where each written answer is shown with its score out of 2 and the reference answer

The profile page tracks every concept across exams, so you can see where you are shaky:

My profile page with scores across 4 exams and a table of each concept's state, Part 1 and Part 2 result

What my friend said

He really loved the idea. I don't even remember how many times he thanked me for this. He liked Part 2 the most, because most of his doubts are in theory, and he is already able to do Part 1 by following the steps his teachers taught.

What is not done

A small model sometimes makes a weak question or a flashcard with broken maths symbols, which you can see in my screenshots. The tests use a fake model, so they check my rules and code checks, not how good any one real model is.

Code

https://github.com/manish-yadav-iitg/crosscheck

ollama pull gemma3:4b
pip install -r requirements.txt
python -m uvicorn app.main:app --port 8000
Enter fullscreen mode Exit fullscreen mode

Top comments (0)