This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I built
CrossCheck is a study tool. You paste your notes, it makes flashcards, and then it gives you an exam in two parts.
In Part 1 you solve things on your own and lock your answers. Only then Part 2 opens (so you can't copy steps from it). In Part 2 you check someone else's answer. About 3 out of 4 of those answers have one small mistake and the rest are fully correct, and you have to say which is which, why, and how sure you are.
It runs on my laptop with an open model. The notes never leave the computer.
Who it is for
It is for my best friend. We grew up together, so he is like a brother to me. He is in class 12 and this year he has his board exams, which are very important exams in India.
He used the "Add your own notes" feature to make a test on limits (it is his strong topic), took the test, and was really amazed by the app.
A normal quiz only gives a score like 5/10. It does not tell you if you can't do it, or if you can do it but trust things too easily. My friend told me he can do Part 1 type questions by following the steps his teachers taught, but most of his doubts are in theory. So I made the exam in two parts, to see both.
How it works
- You paste notes. The model splits them into small concepts.
- You revise with flashcards. They are closed during the exam.
- Part 1, solve it yourself and lock it. Part 2, check a worked answer or a statement.
- The app shows the real answer, why the wrong ones are wrong, and what to do next.
What to do next is not decided by the model. It is a small table of rules in plain Python:
| Part 1 | Part 2 | Next step |
|---|---|---|
| good | good | go on, next exam has trickier mistakes |
| good | weak | explain why it works, then 2 more checking items |
| weak | good | 3 new practice problems with hints |
| weak | weak | go back to the flashcards |
| not enough items | give 2 more items, no verdict yet |
"Not enough evidence yet" is a real answer. With only 3 items, getting 2 right does not prove anything, so the app says so. I used a simple Beta estimate for this. For example 3 of 3 gives 0.87 (good), 2 of 3 gives 0.52 (not sure), and 1 of 3 gives 0.18 (weak).
How I built it: do not trust the model
A small open model makes mistakes. So I made sure it is never the judge of the facts.
- Maths items are made by code, including the wrong solutions. The answer is checked a second way (numerical derivative or numerical integral). The model is not involved.
- Theory items come from the student's own sentences. A correct item is their own sentence, unchanged. A wrong item is their sentence with one small edit, and I keep the original as the truth. Code checks the sentence is really in the notes, and checks the edit is small. Then a second model call checks that the edit really makes it false. If anything fails, that item is thrown away.
- The model never writes code that I run.
- The answers are never sent to the browser before you submit, and Part 2 is not sent at all until Part 1 is locked. There is a test for both.
The model does the language jobs: splitting notes, writing flashcards, picking sentences, proposing edits, and reading the student's written reasons (0, 1 or 2 points).
Some real numbers from my own testing (the app shows them at /api/stats): the model tried 51 edits to make a sentence false, and only 20 passed my checks. 31 were thrown away because the change was too big, or the second check said the sentence was not really false. Also, 13 of the 65 sentences the model picked were not actually in the notes, so code removed them. So a small open model does get things wrong, and these checks mattered.
Why open innovation matters
- Privacy. A friend's notes and their exam mistakes stay on their own computer. Nothing goes to a server I don't control.
- It works without internet. The model runs locally with Ollama. If the model is off, the calculus part still works fully, and the physics demo uses a small saved set. The app tells you which one it is using.
-
I can swap the model. It is one setting (
LLM_MODEL), and any server with the OpenAI chat API works. - It costs nothing to run. My friend can retry 20 exams and nobody pays per question.
I did not test a closed model against it, so I can't say the open one is better. What I can say is that a small open model gets things wrong sometimes, and building around that is what made the design good.
Demo
The home page, with the model running (gemma3:4b) and a different quote every time you refresh:
My friend pasted his own notes on limits, and it split them into concepts:
Flashcards made from the notes. They are closed during the exam:
Part 1, solve it yourself. Part 2 stays hidden until you lock this:
Part 2, check someone else's work. Say correct or not, why, and how sure you are:
The result. It says "not enough evidence yet" when there are too few items, and it points out when you were confident but wrong:
After the test, each answer is shown with the reference answer and a short reason for the score:
The profile page tracks every concept across exams, so you can see where you are shaky:
What my friend said
He really loved the idea. I don't even remember how many times he thanked me for this. He liked Part 2 the most, because most of his doubts are in theory, and he is already able to do Part 1 by following the steps his teachers taught.
What is not done
A small model sometimes makes a weak question or a flashcard with broken maths symbols, which you can see in my screenshots. The tests use a fake model, so they check my rules and code checks, not how good any one real model is.
Code
https://github.com/manish-yadav-iitg/crosscheck
ollama pull gemma3:4b
pip install -r requirements.txt
python -m uvicorn app.main:app --port 8000








Top comments (0)