DEV Community

Jarvis Wick
Jarvis Wick

Posted on

Groundtruth: a revision partner for my cousin that only answers from his notes

Hacktoberfest: Maintainer Spotlight

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My cousin asked me to be his revision helper. He was working through college coursework and wanted someone to quiz him on it. I could say yes to the first session and not to the six after it, so I built the part that does not need me in the room.

He had a specific problem before this one. He re-reads his notes until they feel familiar, then walks into an exam believing he knows the material. When he asked a chatbot to quiz him instead, it explained things that were never in his syllabus, produced citations that looked completely real and told him he was right when he was not. That is not a bug in that tool. It has every fact on the internet and no idea which lines of his notes are the ones that matter.

Groundtruth reads his notes and does three things:

  1. writes questions that can be answered only from that material
  2. grades his free-text answers against it
  3. shows the exact line that settles every verdict

When his notes do not cover something, it says so instead of guessing.

He is using it now. I did not manage to get a quotable reaction out of him, which I choose to read as a good sign.

Demo

The video is the bundled cell biology sample, so you can watch the refusal happen without supplying your own notes. Ask it about mitosis and it will tell you the notes never covered mitosis, because they do not.

Code

GitHub logo Ja4V8s28Ck / groundtruth

A revision partner

Groundtruth

A revision partner that only answers from a student's own notes and shows the exact line it used for every verdict. Demo link.

Why I built it

My cousin asked me to be his revision helper. He was working through college coursework and wanted someone to quiz him on it. I could say yes to the first session and not to the six after it, so I built the part that does not need me in the room.

He had a specific problem before this. He re-read his notes until they felt familiar, then walked into an exam believing he knew the material. When he asked a chatbot to quiz him instead, it explained things that were never in his syllabus, invented citations that looked real and told him he was right when he was not. That is not a bug in that tool. It has every fact…

https://github.com/Ja4V8s28Ck/groundtruth

Python and Streamlit, about 1,500 lines. You need Ollama running and two models:

ollama pull qwen3:4b           # writes and grades, 2.5 GB
ollama pull nomic-embed-text   # finds passages, 274 MB
Enter fullscreen mode Exit fullscreen mode

Then pip install -r requirements.txt and streamlit run app.py. Click "Use the sample notes" and you have a working session in about 20 seconds. No API key, no account, nothing to sign up for.

How I Built It

The model is not bolted onto a search box. Take it out and there is no product left, because there is no keyword path to writing a good revision question.

Citations you can check

This is the part I care about most. Most AI study tools tell you whether you were right. This one shows why it thinks so, in the words of your own notes.

Every chunk keeps the line range it came from. When the model grades an answer it has to return a verbatim quote and the chunk id. Back in Python that quote is checked against the real text with a sliding window match, which tolerates the model paraphrasing slightly and then returns the exact span from the file.

If the model returns correct and cannot produce a quote that exists in the notes, the verdict is downgraded rather than shown. You can open the file and check every claim the app made.

Three guards, because a 4B model is generous

qwen3:4b is small enough to run on a laptop and loose enough to grade things correct that are not. Three checks in the store catch that:

  • Quote verification. A correct verdict with no matching quote in the notes is never displayed.
  • Term coverage. If the retrieved passage covers almost none of the terms in the answer, the verdict is downgraded. This is what makes the refusal falsifiable. The app cannot be wrong by inventing and it cannot be wrong by being vague.
  • No spoilers. The hint shown for a wrong answer is checked so it does not hand over the answer itself.

Term coverage came out of a failing test. A vague answer was being graded incomplete, but that word implies the notes covered it and for a vague answer that is not knowable. So I made the claim falsifiable instead. If the notes barely mention what you wrote, the app is not permitted to assert that the notes covered it.

The model thinks before it answers and you cannot stop it

This was the surprise of the build. My machine has an RTX 4060 and Ollama reports the model running at 100% GPU, yet a call whose entire real answer was one word took 4.6 seconds and generated 166 tokens:

wall 4.64s   generated 166 token done=stop
reply: 'Hmm, the user just asked me to reply with the single word "ready"...'
Enter fullscreen mode Exit fullscreen mode

qwen3 is a reasoning model and Ollama ignores think: false for it. So did enable_thinking through the chat template kwargs and through the options bag. So did asking it in the system prompt not to reason.

Attempt Tokens generated
think: false (the default we send) 166
chat_template_kwargs.enable_thinking=false 133
options.enable_thinking=false 162
System prompt "answer directly, do not reason" 146

Roughly 150 tokens of preamble on every single call and it is 80 to 90 percent of the latency. Nothing suppresses it on this model.

What did help was noticing the calls were independent. Each passage is asked about separately and the calls share no state, so they now run in a thread pool:

build_session(count=8)   before   16.81s
build_session(count=8)   after    12.18s
Enter fullscreen mode Exit fullscreen mode

That is only 28 percent and the reason is worth being plain about. The GPU is saturated on generation, not on request overhead, so overlapping the calls hides the setup cost and nothing else. 35 tokens per second is simply what a 4B model generates on that chip. The real fix is a model that does not think and I did not have time to swap one and re-verify the guards before submitting.

A bug that only appeared once I added an embedding model

Worth writing down, because it is the kind of failure that hides.

My indexer looped over the chunks and embedded them one at a time, but it returned early as soon as the vector array existed. Chunk 0 got a real vector and every other row stayed zero. Semantic search filters zero scores out, so it cheerfully returned nothing but chunk 0 while the header still said the search was semantic. Grading was citing the first section of the notes and nothing else.

It never fired while I had no embedding model installed, which is exactly the window in which I was testing everything else. The moment I pulled nomic-embed-text it was live. The fix was one batched call for the whole document:

rows with a real vector: 10 of 10
probe "sodium potassium pump"  -> chunks [3, 1, 6]
probe "glycolysis"             -> chunks [8, 3, 2]
Enter fullscreen mode Exit fullscreen mode

Why Does Open Innovation Matter?

His notes are not mine to upload. Course material, a lecturer's slides and the record of what a student is weak at are not the builder's to hand to a third-party server. It runs over localhost, so the notes never leave the machine. There is no server to leak them to because there is no server.

It has to work with no network. Revision happens at 7am in a quiet room the night before an exam and depending on a cloud API means depending on someone else's uptime and someone else's billing.

There is a cost argument as well and it is the weakest of the four. A small model on your own hardware is free and renting an equivalent one costs a few cents a session. Free is not the interesting part. It's about keeping our data to ourselves.

Top comments (0)