This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Saathi is a study companion built around one person's own notes. You add your notes. It writes quiz questions from them, asks "how sure are you?" before you answer, works out why a wrong answer was wrong, brings each question back just before you would forget it, and at the end of the week writes a short letter about how things went, in the way you like to be spoken to.
I built it for [[Rahul Kumar]], [[He is my classmate]]. [[He was often struggles to evaluate him during exam time ]]
A chatbot answers from the whole internet. [[Rahul Kumar]] does not need the internet's version of chapter 4. They need to find out what they do not understand about their notes. So Saathi works only from their notes:
- Grounded questions only. Every generated question must pass a check (four distinct options, and the right answer has to appear in the source passage) before it is saved.
- Confidence matters. A wrong answer you felt sure about is the most valuable signal, so it hurts the card's schedule the most and is called out in the letter.
- Mistakes get a reason. The model names the likely misunderstanding behind that specific wrong choice, and the tags are counted over the week.
- The letter only uses true numbers. It is written by the model but from facts computed in SQL.
Demo
[[Live link: https://saathi-1vua.onrender.com. username - admin , password - Hello@123]]
Code
[[
ypunit757-prog
/
Saathi
An offline study companion, built for one person
Saathi
An offline study companion, built for one person.
Saathi turns a friend's own notes into quiz questions, asks how sure they are before they answer, works out why a wrong answer was wrong, schedules each question for the moment they are about to forget it, and writes a short weekly letter in their own style of language. It runs an open-weight model (Gemma by default) through Ollama on a laptop. Nothing is sent anywhere.
Built for the DEV Hacktoberfest Weekend Challenge, "Build for a Friend" (Oct 2 to 5, 2026). All code in this repo was started inside that window.
Why it is shaped this way
- Their notes, not the internet. Questions are written only from the notes you add, and every generated question must pass a grounding check (the right answer has to appear in the source passage) before it is saved. Small models make things up; this…
How I Built It
The model. Gemma, an open-weight model from Google. Saathi has three ways to reach it, chosen by one environment variable:
-
ollama: Gemma (gemma3:4b, or any Gemma you have pulled) on the laptop through Ollama. Nothing leaves the machine. -
cloud: hosted Gemma 4 (gemma-4-31b-it) through the Gemini API. This is what the live demo uses. -
mock: a deterministic stand-in so the app and its 45 tests run with no model at all.
Saathi talks to the model over plain HTTP, so there is no SDK between my code and the model.
Stack. Python and Flask, one SQLite file, plain HTML/CSS/JS with no CDN and no framework. Hosted on Render as a web service with a persistent disk, started with gunicorn. A password (SAATHI_PASSWORD) protects the public URL.
Four ideas that make it feel like a tutor, not a quiz generator
- Grounded questions. The model returns JSON and I validate it:
ctx = set(rag.tokenize(context))
ans = rag.tokenize(options[idx])
if not ans or not any(t in ctx for t in ans):
return None # the model invented it; drop the question
- Confidence changes the schedule. Standard SM-2 only knows right or wrong. Saathi also asks whether you were sure. Wrong-and-sure scores 0, wrong-and-unsure 1, and right-but-guessing gets a shorter gap than right-and-sure.
- Every mistake gets a reason. A second request asks for the misunderstanding behind that wrong choice, so the right/wrong feedback stays instant.
- The letter only uses true numbers. The prompt forbids inventing anything outside the facts object.
Retrieval. "Ask my notes" uses BM25 in pure Python (plus embeddings when a local embedding model exists). If the notes do not hold the answer, Saathi says so. Optionally, the learner can tick "Also search the web" and web results are labelled as such.
Monitoring the agents. Each task is a named agent (Saathi Quiz Writer, Misconception Diagnoser, Notes Assistant, Letter Writer) traced with Sentry's agent tracing: every model call records model, tokens and latency. Prompts and notes are not sent to Sentry unless I turn that on.
Voice. The weekly letter can be read aloud with ElevenLabs.
Comparing models on the real job. bench/benchmark.py measures how often each model returns valid JSON and how many questions survive validation:
| model | json_ok | questions asked | kept after validation | sec/chunk |
|---|---|---|---|---|
| [[gemma3:1b]] | [[ ]] | [[ ]] | [[ ]] | [[ ]] |
| [[gemma3:4b]] | [[ ]] | [[ ]] | [[ ]] | [[ ]] |
[[Run python bench/benchmark.py gemma3:1b gemma3:4b yourself and paste the real numbers. Delete this table if you do not run it.]]
Why Does Open Innovation Matter?
- Their study history is private. It is a record of everything someone does not understand yet. Because Gemma is open-weight, Saathi can run on Rahul kumar's own laptop and nothing has to leave it. The hosted demo is a deliberate trade-off, and it is the same code with one setting changed.
- I could shape it. I could swap models when a small one was too slow, validate and discard its output, and rewrite the prompts so Saathi talks the way Rahul kumar talks .
- It works when the internet does not. A local model does not care.
- Cost. Spaced repetition means thousands of small questions. Per-token pricing turns "study a bit more" into a bill; a local open model does not.
- Not locked in. The hosted backend can also be any OpenAI-compatible endpoint, so moving from Google to a GPU server I run myself is a settings change, not a rewrite.

Top comments (0)