DEV Community

Cover image for I Built My Friend a Study Buddy That Reads Her Notes and Never Uploads Them
Livansh Malhotra
Livansh Malhotra

Posted on

I Built My Friend a Study Buddy That Reads Her Notes and Never Uploads Them

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend


What I Built

I built this for Kanishka, who is preparing for Engineering 1st sem exams. She has a problem every serious student knows: hundreds of pages of lecture slides, scanned handwritten notes and typed summaries, scattered across folders, none of it searchable. Before every revision session she spends the first hour just finding things.

The obvious fix is to upload everything to a chatbot. She won't, and I agree with her. Her notes hold her draft answers, her mnemonics and her margin scribbles about what she doesn't understand. That is a record of how she thinks, and she shouldn't have to hand it to a server she has never heard of.

StudyBuddy is a local-first study assistant:

  • Drop a PDF into a folder, and about a minute later it's searchable.
  • Ask a question in plain English and get a short, cited answer pointing at the exact page of her own notes.
  • Take quizzes (MCQ) generated from those same notes.
  • See her weak spots, with topics ranked by how often she gets them wrong.

It runs on a laptop. After the one-time model download it works with the Wi-Fi off, including on her commute.


Demo

The recording shows:

  • A new PDF dropped into /uploads, with its status flipping from pending to ready
  • "What is a pointer?" answered with two cited bullets instead of a page of text
  • A repeated question answered instantly from cache and counted as an LLM call skipped
  • A quiz, a wrong answer, and that topic moving up the weak-topics list

Her real notes never leave her machine.


Code

GitHub Repository: StudyBuddy


How I Built It

drop file → parse → chunk → embed → pgvector
question  → cache → hybrid search → rerank → gate → (LLM only if needed)
quiz      → cached questions → code-graded MCQs → weak-topic ranking
Enter fullscreen mode Exit fullscreen mode

The principle: the LLM is the last resort. A small local model is slow and sometimes unreliable, so every step that can work without it does. Ingestion has no LLM. Search has no LLM. Repeat questions hit a semantic cache. MCQs are graded in plain code. The model only runs when a question genuinely needs synthesis.

The open stack

Layer Tool Why
LLM Gemma via Ollama Answers, quiz generation, free-text grading
Embeddings BGE-M3 Multilingual, self-hosted
Reranker bge-reranker-v2-m3 Open-weight, CPU-friendly, calibrated scores
Database PostgreSQL + pgvector + full-text search Hybrid retrieval in one place
Durable ingestion Temporal A 600-page scan that fails at page 400 resumes instead of restarting
Backend / UI FastAPI, React + Vite No lock-in
Tracing [Sentry, if your DSN is live] Every request tagged llm_used

Drop-a-file ingestion

A watcher monitors the uploads folder. Each file is hashed, so duplicates are skipped and an edited file is re-ingested alone. Nothing is retrained, because adding a document to a RAG system is just an index update.

The bug that taught me the most

My first version answered "what is a pointer?" with this:

C++ - Quick Notes Page 1 of 2 C++ Syntax, memory, OOP, STL, templates and modern C++ 1. Program Structure & Basics #include using namespace std; int main() { int x = 5; ... 2. Pointers, References & Memory * Pointer stores an address... 3. Classes & OOP class Animal { protected: ...

The answer was in there, buried in a wall of unrelated text. Four problems were stacking up:

  1. The gate returned whole chunks. A "high-confidence" hit was shown verbatim.
  2. Chunks were page-sized, not idea-sized. A hard 400-token split cut across topics, so Pointers shared a chunk with OOP.
  3. Flat text extraction destroyed structure. Headings and bullets collapsed into one run-on string.
  4. Rank-based scores aren't relevance scores. Reciprocal rank fusion only knows ordering, so a weak match could still look confident.

The fix was small-to-big retrieval:

  • Parse PDFs into markdown so headings and bullets survive.
  • Split by heading, then index at bullet level, with each bullet prefixed by its heading.
  • Add a local cross-encoder reranker whose scores are calibrated enough to set real thresholds.
  • Rebuild the gate on three outcomes: NOT_FOUND, EXTRACT (show the best 1-3 bullets, no LLM) or SYNTHESIZE (Gemma over a tiny, clean context).
  • Never show a parent section as the answer. It only appears in a collapsed "Source" panel.

The same question now returns:

• Pointer stores an address: int* p = &x; and *p dereferences it. [C++ Quick Notes, p.1]
• Prefer smart pointers over raw new/delete to avoid leaks. [C++ Quick Notes, p.1]

Did it actually get better?

I wrote [N] question and expected-answer pairs from her real notes and ran them before and after.

Before After
Avg. characters shown [ ] [ ]
Answer contains the fact [ ]% [ ]%
Contains unrelated text [ ]% [ ]%
Answered without the LLM [ ]% [ ]%
p50 latency (no LLM / LLM) [ ]s / [ ]s [ ]s / [ ]s

[One honest sentence on any number that disappointed you.]


Why Does Open Innovation Matter?

Her notes stay hers. With a closed API, every page of her notes is a request to someone else's server. With open-weight models and a local database, the whole system runs on a laptop with the network unplugged. That is not a feature you can bolt onto a closed model, however good its privacy policy is.

Zero cost per question. In exam season she may ask thousands of questions. Free local inference means she never rations curiosity because of a bill.

I could fix the retrieval because I owned it. The pointer-dump bug lived in chunking, scoring and the gate. Every layer was readable code I could change. With a black-box RAG API, the best I could have done is tweak a prompt and hope.

Swappable models. The model is one line in .env. I started on gemma:2b and moved to gemma3:4b for grounded answering. [Say what actually changed.]

Where a closed model would have been better: a frontier model writes more fluent explanations and handles messy questions more gracefully than a 2B or 4B model on a laptop CPU. Mine is also slower, about [ ]s per synthesized answer. I accepted that because most queries never reach the model, and because the alternative was her notes leaving the device.



The Hand-over

Her first search was "What is a pointer?". She thought she knew almost every topix well but had scored 1/3 on the quiz. After ten minutes she said: "this will help me for sure, but its ofcourse not your brain", (actually it is my own idea).


Prize Categories

  1. Best Use of Gemma: Uses local Gemma 2b for answer synthesis, quiz generation, and grading.
  2. Best Use of TabPFN: Uses TabPFN tabular model to predict student weak topics from quiz attempt history.
  3. Best Use of Temporal: Durable ingestion pipeline (IngestWorkflow) handling background parsing, chunking, and embedding.
  4. Best Use of Tiger Data: PostgreSQL pgvector hybrid search combining vector embeddings (BGE-M3) + BM25 keyword search with RRF.
  5. Best Use of Sentry Agent Tracing: Sentry SDK integration monitoring RAG pipeline latency ($p50$/$p95$) and confidence gate metrics.

Top comments (0)