This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
StudyLens — A Local-First AI Study Assistant
Built for: My younger sister, who's drowning in lecture PDFs and needs a study tool that works offline, respects her privacy, and doesn't require a $20/month subscription.
What I Built
StudyLens turns passive reading into active learning — completely on your machine. No API keys, no cloud, no accounts, no telemetry.
Upload a PDF/TXT → Get:
- 🧠 Understand — Ask questions, get grounded answers with citations from your material
- 📚 Revise — One-click structured revision sheet (concepts, key terms, relationships, "remember" facts, common confusions)
- 🎯 Quiz — 5 or 10 mixed MCQ/short-answer questions, auto-scored with weak-area identification
- ⏱️ Exam Mode — "I have 30 minutes" → prioritized study plan (Must Know / Good to Know / Skim) with time estimates
Demo
ollama pull gemma3:1b
ollama serve
pip install -r requirements.txt
streamlit run app.py
Opens at http://localhost:8501 — works on any machine that runs Python + Ollama (tested on macOS, Linux, Windows).
Code
github.com/ShashwatSinha03/HF_BuildForAFriend
study-lens/
├── app.py # Streamlit UI + orchestration
├── llm.py # Ollama client + prompts + robust JSON parsing
├── document.py # PDF/TXT extraction + keyword retrieval
├── requirements.txt
└── README.md
How I Built It
Stack: Python • Streamlit • Ollama • Gemma 3 1B • PyPDF • Requests
Architecture decisions that matter:
| Constraint | Solution |
|---|---|
| No cloud/API keys | Local Ollama + Gemma 3 1B (815MB) |
| Small model, unreliable JSON | Multi-strategy parser (standard → trailing-comma fix → duplicate-key removal → regex fallback) |
| No embeddings/vector DB | Lightweight keyword-overlap retrieval (TF-IDF-inspired, ~30% of doc sent to LLM) |
| Limited compute | Single LLM call per action; batch short-answer eval; conservative token limits |
| Privacy-first | Zero network calls except localhost:11434; session-only state |
Sprint-based development (6 sprints over 3 days):
- Foundation — upload, extraction, session state, Ollama health checks
- Explain — grounded Q&A with retrieval
- Revise — structured revision sheets
- Quiz — MCQ + short answer, batch evaluation, weak areas
- Exam Mode — time-bounded prioritization
- Polish — robust parsing, error UX, README, .gitignore
Why Open Innovation Matters
This project exists because of open weights and local inference.
If I'd used a closed API:
- ❌ Cost — $20/month per user minimum
- ❌ Privacy — lecture notes sent to third parties
- ❌ Offline — no studying on planes, in libraries, or with spotty campus WiFi
- ❌ Customization — can't tweak prompts, retrieval, or model behavior
- ❌ Longevity — API deprecation, price hikes, or service shutdown kills the tool
With Gemma 3 1B + Ollama:
- ✅ Free forever — runs on hardware she already owns
- ✅ Private — her notes never leave her machine
- ✅ Offline-first — works anywhere
- ✅ Hackable — every prompt, retrieval strategy, and UI decision is in plain Python
- ✅ Sustainable — no vendor lock-in; model runs as long as Python runs
The "small model" constraints forced better engineering: lightweight retrieval, strict prompting, defensive parsing, batch evaluation. Those are better patterns that scale up — not compromises.
The Friend Test
"Wait, it's actually free? And my PDFs don't get uploaded anywhere? Damn, broski"
Yep, that's my sister, after using it for her midterms prep
That's the whole point. Open innovation doesn't just lower barriers — it changes what's possible to build for the people you care about.
Top comments (0)