DEV Community

Cover image for I Reverse-Engineered 3 Years of Exam Papers on localhost (With Wi-Fi Off)
Aritro Bag
Aritro Bag

Posted on

I Reverse-Engineered 3 Years of Exam Papers on localhost (With Wi-Fi Off)

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Every AI study tool wants your lecture PDFs. Uploading copyrighted course material to a stranger's inference endpoint is a trade-off, and the friend I built this for wasn't willing to make it.

So I built Gist, a study partner that never makes that trade. It runs entirely on the laptop. PDFs are parsed, embedded, and searched locally, the LLM runs locally in Ollama, and the app works with the Wi-Fi switched off. There is no API key in the codebase because there is no API involvement.

The core of Gist isn't another generic chatbot wrapper. It is an intelligent past question paper analyser. You drop in 3+ years of past university exam PDFs, and Gist reverse-engineers the exam pattern:

  1. Which questions and variations repeat.
  2. What each topic is actually worth in total marks weightage.
  3. How frequently each concept appears across semesters.
  4. How your personal quiz drills cross-reference against that weightage.

The Past Paper Analyser Workbench and Priority Matrix showing topic weightage breakdowns, past questions, and high-yield weak spots

Grounded Q&A with direct PDF page citations and Timed Mock Exam drill

A snapshot from the landing page

Demo

Product page: https://gist-landing.onrender.com

Demo Video: https://youtu.be/QMNTq5FTj8k
disclaimer: i'm not a good speaker

Code

GitHub logo dev-Aarish / gist

Turn your lecture notes and past papers into an offline exam training ground. Grounded RAG with page citations, historical mark weightages, high-yield weak spot detection, and timed mock exams. Runs 100% locally on open-weight LLMs via Ollama — zero cloud, zero API keys.

Gist

Gist is an offline, privacy-first AI study partner powered by local open-weight large language models (LLMs) via Ollama, FastAPI, and React. It combines page-grounded document Q&A with an exam intelligence engine that analyzes past question papers, detects topic marks weightages, prioritizes high-yield weak spots, and conducts timed mock exams ("Grill Me" mode).


Architecture Overview

graph TD
    subgraph UserInterface ["Frontend Client (React 18 + Vite - Dark Theme)"]
        UI["Web Interface: Chat, Past-Paper Workbench, Priority Matrix, Quizzes, Timed Mock Exams"]
    end
    subgraph BackendGateway ["Backend API (FastAPI)"]
        API["REST & SSE Gateway (Port 8000)"]
        Ingest["Document Ingestion Engine (PyMuPDF)"]
        Analyzer["Past-Paper Analyzer Engine"]
        RAG["Grounded RAG Pipeline"]
        Quiz["Adaptive Quiz & Instant Bank Engine"]
        Exam["Mock Exam Simulator ('Grill Me' Engine)"]
    end
    subgraph StorageLayer ["Local Persistent Storage"]
        Chroma["ChromaDB: Vector Embeddings"]
        SQLite["SQLite (tracker.db): Papers, Question Bank, Mastery & Attempt History"]
        Uploads["PDF Storage (data/uploads)"]
    end

    subgraph LocalLLM ["Local Inference Engine (Ollama)"]
        EmbedModel["Embedding Models: nomic-embed-text / bge-m3"]
        ChatModel["Dynamic Open-Weight Models: Gemma
…

The parsing rules are the part I'd walk through first. Marks appear in university papers as [10], (5 marks), 15M, or a bare 4 + 6 + 2 = 12 breakdown, and getting this wrong means every downstream percentage is wrong:

# backend/analyzer.py:134
m_single = re.search(
    r'(?:\[|\()?\b(\d{1,2})\s*(?:marks?|mark|m|pts?)\b(?:\]|\))?|\[(\d{1,2})\]',
    line,
    re.I,
)
Enter fullscreen mode Exit fullscreen mode

Mark weightage is the input to the entire priority model, so it has to come out of the PDF exactly, not approximately.

How I Built It

Everything runs on the user's machine. No network calls leave the host, which makes the privacy claim checkable rather than a policy promise.

Architecture diagram showing every component inside a single box representing the user's machine: a React 18 and Vite client calls a local FastAPI server over localhost, which dispatches to four engines (PyMuPDF ingestion, the past-paper analyser, grounded RAG, and a mock exam grader). These read and write ChromaDB for vector search, SQLite at data/tracker.db for papers and mastery, and the original PDFs in data/uploads. All four engines call a local Ollama runtime on port 11434. No connection leaves the box.

Pipeline

Flowchart of the past-paper analyser. A PDF passes through a regex metadata scan for year, total marks, and title, then a decision on whether question boundaries were found by regex. If yes, questions are split on patterns like Q1(a), 2., and (i) with marks extracted per line, topics are inferred by keyword so terms like bcnf and 3nf map to Normalization, and a second decision sends only still-unclassified questions to a micro-LLM pass capped at 12 questions and 200 tokens. If no boundaries were found, full LLM extraction is the fallback. Both paths fold results onto a canonical topic map, write questions, marks, and topics to SQLite, and feed a priority matrix computed as marks percentage times weakness, producing a ranked revision list. A dotted background thread also indexes the paper into ChromaDB.

The stack is FastAPI, PyMuPDF, ChromaDB, SQLite, and Ollama on the backend, React on the front. About 4,700 lines of Python now, up from 2,000 when the quiz was the main feature.

Resources I took help of:

  • Antigravity(CLI)
  • Opencode
  • Freebuff
  • Skills(frontend-design, ui-ux-pro-max, emil-design-eng, impeccable, documentation-writer)

Why the Analyser is Rules-First, LLM-Second

My first prototype sent the entire exam paper to an LLM asking for structured JSON. It was sluggish, expensive on local compute, and hallucinated marks.

The current architecture is a three-stage hybrid pipeline:

  1. Deterministic Regex Parse: Splits questions on structural boundaries (Q1(a), 2., (i)), groups MCQ sub-options under their parents, filters boilerplate university headers, and extracts marks per line. Runs in milliseconds with 100% fidelity.
  2. Keyword Topic Inference: A domain-specific ontology maps question vocabulary to canonical syllabus topics (bcnf, 3nf, attribute closure → Normalization; precedence graph, 2pl, dirty read → Transactions). Using deterministic mapping prevents small 7B models from inventing arbitrary topic names.
  3. Micro-LLM Pass for Edge Cases: Only unclassified questions falling through to the generic category are sent to Ollama, batched at ≤ 12 snippets with num_predict=200. Full LLM extraction is reserved strictly as a fallback for unstructured, legacy scanned layouts.

A canonical topic dictionary folds all variations back into fixed buckets so percentages aggregate cleanly without fragmentation.

Turning Weightage into a Ranked Study Plan

Two mathematical formulas drive the recommendation engine:

  1. Exam Importance measures how heavily a topic features across years, with frequency acting as a credibility multiplier on total marks share:
exam_importance = marks_pct * (0.6 + 0.4 * (frequency_pct / 100))
Enter fullscreen mode Exit fullscreen mode

(Multiplication ensures a topic worth 15% across all past papers ranks significantly higher than a topic appearing in only one isolated year).

  1. Composite Priority scales exam weight against your personal weakness:
priority = exam_importance * ((100 - accuracy) / 100) * 1.5
Enter fullscreen mode Exit fullscreen mode

Why Open Innovation Matters

This project exists solely because open weights and local runtimes are now viable on standard consumer laptops. A closed, cloud-only model would have compromised the project in three fundamental ways:

  1. Verifiable Privacy: No keys to leak, no remote database to breach, and no data harvesting. Anyone can audit the 4,700 lines of Python and verify that zero packets leave localhost.
  2. Zero-Cost Feedback Loops: Prompt tuning and evaluation loops take seconds. Tweaking extraction prompts and re-running against local weights is free, reproducible, and immune to upstream provider deprecations or rate limits.
  3. Hybrid Architecture Viability: A deterministic rules-first pipeline that saves LLM calls only makes architectural sense when local execution is unmetered.

My Agent Session

Here is the interactive session trace detailing the design iterations behind the local-first architecture, the 3-stage hybrid parsing engine, and the priority formula:

Building Gist: Local-First Offline Exam Intelligence & 3-Stage Hybrid Parser
You

How can we build Gist as a 100% local, offline exam intelligence engine so that no university course PDFs ever leave the student's laptop?

Agent

We design an entirely local architecture powered by Ollama for local open-weight LLMs (Gemma 2 / Llama 3.2), FastAPI for the backend gateway, PyMuPDF for deterministic extraction, ChromaDB for local vector search, and SQLite for persistent tracking. Zero API keys, zero network telemetry.

You

The LLM is hallucinating marks and taking too long when parsing entire exam PDFs at once. How do we make question and marks extraction deterministic and fast?

Agent

We replace full-paper LLM generation with a 3-Stage Hybrid Parser:
1. Deterministic Regex Parsing: Extract question boundaries (e.g. Q1(a), 2.) and mark tags ([10], (5 marks), 15M) in milliseconds with 100% fidelity.
2. Keyword Ontology Inference: Map question keywords directly to syllabus topics (e.g. 'bcnf', '3nf' -> Normalization).
3. Micro-LLM Pass: Only send leftover, unclassified questions to Ollama in small batches with low token limits.

You

How do we turn the extracted topic mark weightage and student quiz accuracy into an actionable revision priority matrix?

Agent

We compute Exam Importance using question frequency as an amplifier over marks share, then multiply by student weakness:
- exam_importance = marks_pct * (0.6 + 0.4 * (frequency_pct / 100))
- priority = exam_importance * ((100 - accuracy) / 100) * 1.5
Untested topics are penalized with an assumed 75% weakness rather than 0% accuracy to prioritize high-risk unknowns safely.

Prize Categories

  • Best Use of Gemma: Gist is built on Google's open-weight Gemma 2 (gemma2:9b and lightweight gemma2:2b), running locally via Ollama to power the micro-LLM topic classification pass, grounded RAG answering, and automated mock exam grading.
  • Best Use of Render: The interactive product landing page and live web preview are deployed on Render at https://gist-landing.onrender.com (configured via render.yaml).
  • Best Use of Entire: The agent session transcript behind Gist's local-first architecture and hybrid parser design is embedded directly in this write-up.

Top comments (0)