This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Every AI study tool wants your lecture PDFs. Uploading copyrighted course material to a stranger's inference endpoint is a trade-off, and the friend I built this for wasn't willing to make it.
So I built Gist, a study partner that never makes that trade. It runs entirely on the laptop. PDFs are parsed, embedded, and searched locally, the LLM runs locally in Ollama, and the app works with the Wi-Fi switched off. There is no API key in the codebase because there is no API involvement.
The core of Gist isn't another generic chatbot wrapper. It is an intelligent past question paper analyser. You drop in 3+ years of past university exam PDFs, and Gist reverse-engineers the exam pattern:
- Which questions and variations repeat.
- What each topic is actually worth in total marks weightage.
- How frequently each concept appears across semesters.
- How your personal quiz drills cross-reference against that weightage.
Demo
Product page: https://gist-landing.onrender.com
Demo Video: https://youtu.be/QMNTq5FTj8k
disclaimer: i'm not a good speaker
Code
dev-Aarish
/
gist
Turn your lecture notes and past papers into an offline exam training ground. Grounded RAG with page citations, historical mark weightages, high-yield weak spot detection, and timed mock exams. Runs 100% locally on open-weight LLMs via Ollama — zero cloud, zero API keys.
Gist
Gist is an offline, privacy-first AI study partner powered by local open-weight large language models (LLMs) via Ollama, FastAPI, and React. It combines page-grounded document Q&A with an exam intelligence engine that analyzes past question papers, detects topic marks weightages, prioritizes high-yield weak spots, and conducts timed mock exams ("Grill Me" mode).
Architecture Overview
graph TD
subgraph UserInterface ["Frontend Client (React 18 + Vite - Dark Theme)"]
UI["Web Interface: Chat, Past-Paper Workbench, Priority Matrix, Quizzes, Timed Mock Exams"]
end
subgraph BackendGateway ["Backend API (FastAPI)"]
API["REST & SSE Gateway (Port 8000)"]
Ingest["Document Ingestion Engine (PyMuPDF)"]
Analyzer["Past-Paper Analyzer Engine"]
RAG["Grounded RAG Pipeline"]
Quiz["Adaptive Quiz & Instant Bank Engine"]
Exam["Mock Exam Simulator ('Grill Me' Engine)"]
end
subgraph StorageLayer ["Local Persistent Storage"]
Chroma["ChromaDB: Vector Embeddings"]
SQLite["SQLite (tracker.db): Papers, Question Bank, Mastery & Attempt History"]
Uploads["PDF Storage (data/uploads)"]
end
subgraph LocalLLM ["Local Inference Engine (Ollama)"]
EmbedModel["Embedding Models: nomic-embed-text / bge-m3"]
ChatModel["Dynamic Open-Weight Models: Gemma…The parsing rules are the part I'd walk through first. Marks appear in university papers as [10], (5 marks), 15M, or a bare 4 + 6 + 2 = 12 breakdown, and getting this wrong means every downstream percentage is wrong:
# backend/analyzer.py:134
m_single = re.search(
r'(?:\[|\()?\b(\d{1,2})\s*(?:marks?|mark|m|pts?)\b(?:\]|\))?|\[(\d{1,2})\]',
line,
re.I,
)
Mark weightage is the input to the entire priority model, so it has to come out of the PDF exactly, not approximately.
How I Built It
Everything runs on the user's machine. No network calls leave the host, which makes the privacy claim checkable rather than a policy promise.
Pipeline
The stack is FastAPI, PyMuPDF, ChromaDB, SQLite, and Ollama on the backend, React on the front. About 4,700 lines of Python now, up from 2,000 when the quiz was the main feature.
Resources I took help of:
- Antigravity(CLI)
- Opencode
- Freebuff
- Skills(frontend-design, ui-ux-pro-max, emil-design-eng, impeccable, documentation-writer)
Why the Analyser is Rules-First, LLM-Second
My first prototype sent the entire exam paper to an LLM asking for structured JSON. It was sluggish, expensive on local compute, and hallucinated marks.
The current architecture is a three-stage hybrid pipeline:
-
Deterministic Regex Parse: Splits questions on structural boundaries (
Q1(a),2.,(i)), groups MCQ sub-options under their parents, filters boilerplate university headers, and extracts marks per line. Runs in milliseconds with 100% fidelity. -
Keyword Topic Inference: A domain-specific ontology maps question vocabulary to canonical syllabus topics (
bcnf,3nf,attribute closure→ Normalization;precedence graph,2pl,dirty read→ Transactions). Using deterministic mapping prevents small 7B models from inventing arbitrary topic names. -
Micro-LLM Pass for Edge Cases: Only unclassified questions falling through to the generic category are sent to Ollama, batched at ≤ 12 snippets with
num_predict=200. Full LLM extraction is reserved strictly as a fallback for unstructured, legacy scanned layouts.
A canonical topic dictionary folds all variations back into fixed buckets so percentages aggregate cleanly without fragmentation.
Turning Weightage into a Ranked Study Plan
Two mathematical formulas drive the recommendation engine:
- Exam Importance measures how heavily a topic features across years, with frequency acting as a credibility multiplier on total marks share:
exam_importance = marks_pct * (0.6 + 0.4 * (frequency_pct / 100))
(Multiplication ensures a topic worth 15% across all past papers ranks significantly higher than a topic appearing in only one isolated year).
- Composite Priority scales exam weight against your personal weakness:
priority = exam_importance * ((100 - accuracy) / 100) * 1.5
Why Open Innovation Matters
This project exists solely because open weights and local runtimes are now viable on standard consumer laptops. A closed, cloud-only model would have compromised the project in three fundamental ways:
-
Verifiable Privacy: No keys to leak, no remote database to breach, and no data harvesting. Anyone can audit the 4,700 lines of Python and verify that zero packets leave
localhost. - Zero-Cost Feedback Loops: Prompt tuning and evaluation loops take seconds. Tweaking extraction prompts and re-running against local weights is free, reproducible, and immune to upstream provider deprecations or rate limits.
- Hybrid Architecture Viability: A deterministic rules-first pipeline that saves LLM calls only makes architectural sense when local execution is unmetered.
My Agent Session
Here is the interactive session trace detailing the design iterations behind the local-first architecture, the 3-stage hybrid parsing engine, and the priority formula:
Prize Categories
-
Best Use of Gemma: Gist is built on Google's open-weight Gemma 2 (
gemma2:9band lightweightgemma2:2b), running locally via Ollama to power the micro-LLM topic classification pass, grounded RAG answering, and automated mock exam grading. -
Best Use of Render: The interactive product landing page and live web preview are deployed on Render at https://gist-landing.onrender.com (configured via
render.yaml). - Best Use of Entire: The agent session transcript behind Gist's local-first architecture and hybrid parser design is embedded directly in this write-up.





Top comments (0)