This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built an Offline Study Buddy for Lecture Slides for my batchmate studying Computer Engineering at the University of Peradeniya.
Like many engineering students, my friend relies heavily on lecture slides for revision. However, studying from dense slides on a laptop with patchy Wi-Fi and limited mobile data is frustrating. Lecture slides are designed for classroom presentationβnot for fast exam revision or self-assessment. They needed a fast way to query complex lecture topics and test their knowledge without relying on a cloud connection.
The Offline Study Buddy solves both problems right on a laptop with zero internet required:
π‘ Explain Mode: Ask any question about a course or topic. The system retrieves relevant slide excerpts and generates a clear, concise explanation grounded entirely in the lecture slides, complete with explicit slide citations (e.g., [Slide 5]).
π Quiz Mode: Pick a specific topic or slide range (e.g., Slides 1β15). The system generates multiple-choice questions (MCQs) directly from the slide content, scores your choices interactively, and explains the correct answers.
Everything runs 100% locally on a laptopβwith no internet connection, no user accounts, and zero API costs.
Demo
Code
π Offline Study Buddy for Lecture Slides
An offline, local-first RAG (Retrieval-Augmented Generation) application designed for Computer Engineering undergraduates at the University of Peradeniya. It transforms dense, raw lecture slides (PDFs & PPTXs) into an interactive study & revision assistant.
Everything runs 100% locally on a laptopβno internet connection, no external API keys, no cloud data transmission, and zero cost.
π Key Features
-
π‘ Explain Mode (Grounded Q&A with Slide Citations)
- Ask any question about a course or specific topic.
- Searches local ChromaDB vector store for relevant slide excerpts.
- Generates clear, student-friendly revision summaries.
-
Grounded & Cited: Every fact explicitly references its source slide (e.g.,
[Slide 5]or[Operating_Systems.pdf, Slide 12]). - If the uploaded slides do not cover the question, the system clearly informs you rather than hallucinating.
-
π Quiz Mode (Interactive MCQs with Strict JSON Parsing)
- Select a topic or slide range (e.g., Slides 1β15).
- β¦
How I Built It
Lecture slides (PDF/PPTX)
β
βΌ
Text extraction (PyMuPDF / python-pptx), keeping file + slide number
β
βΌ
Chunking + local embeddings (nomic-embed-text via Ollama)
β
βΌ
ChromaDB (local vector store)
β
ββββββ΄ββββββββββββββββββββββββββββββ
βΌ βΌ
Explain mode Quiz mode
(retrieve slides β (slides β MCQs as strict JSON β
simple explanation scoring + short explanation)
with slide cites) β
β β
βββββββββββββββββ¬ββββββββββββββ
βΌ
Ollama (local open-weight LLM) β Gradio UI
π οΈ Tech Stack
Model Runtime: Ollama running open-weight local models (Gemma 2 for grounded explanations & Qwen 2.5 for strict JSON quiz generation)
Embeddings: nomic-embed-text via Ollama (with fallback embedding support)
Vector Store: ChromaDB (persistent local folder)Document Extraction: PyMuPDF (fitz) for PDF slides & python-pptx for PowerPoint presentations
UI Framework: Gradio (with custom high-contrast CSS and responsive theme layout)
π¬ Model Selection & Evaluation
I didn't assume one open-weight model would fit all tasks. I built a built-in benchmarking tool (benchmark.py) to test candidate models on identical test questions extracted from real computer engineering slides.
Takeaway: Gemma 2 won for clear, concise explanations and citation adherence, while Qwen 2.5 excelled at reliably producing valid JSON arrays for quizzes without syntax errors.
π― Keeping Answers Grounded
To ensure the assistant never hallucinates outside course materials, the model only receives retrieved slide chunks and is enforced by a strict system prompt:
1.Strict Context Boundary: Base answers only on the provided slide excerpts.
2.Explicit Citations: Every fact must cite its source slide (e.g., [Slide 7]).
3.Graceful Fallback: If the uploaded slides do not contain the answer, the model explicitly states: "The provided lecture slides do not contain information on this topic."
π Robust Quiz Generation
Generating structured quizzes requires the local LLM to output valid JSON (question, options, answer_idx, explanation, slide_citation).
To handle occasional markdown wrapping or minor syntax quirks from local models, I built a custom JSON repair engine (repair_and_parse_json in ollama_client.py) that extracts JSON arrays inside markdown blocks, cleans trailing commas, and falls back to regex extraction if needed.
Why Does Open Innovation Matter?
100% Offline Capability: Students in areas with patchy connectivity and expensive mobile data can study without interruption. Cloud APIs would fail them when they need it most during exam prep.
Data Privacy & Security: Sensitive course materials, lecture slides, and study logs never leave the student's laptop.
Zero Operating Cost: No per-token charges or subscription feesβmy friend can run unlimited quizzes and explanations.
Model Choice & Control: Being able to benchmark multiple open-weight models side-by-side enabled me to select the exact model that fit the laptop's RAM and task requirements.
Custom-Tailored Behavior: Prompts and MCQ schemas are tuned specifically for engineering lecture revision rather than generic chat.

Top comments (0)