DEV Community

Shashwat Sinha
Shashwat Sinha

Posted on

StudyLens for my tensed sister juggling with her academics

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

StudyLens — A Local-First AI Study Assistant

Built for: My younger sister, who's drowning in lecture PDFs and needs a study tool that works offline, respects her privacy, and doesn't require a $20/month subscription.


What I Built

StudyLens turns passive reading into active learning — completely on your machine. No API keys, no cloud, no accounts, no telemetry.

Upload a PDF/TXT → Get:

  • 🧠 Understand — Ask questions, get grounded answers with citations from your material
  • 📚 Revise — One-click structured revision sheet (concepts, key terms, relationships, "remember" facts, common confusions)
  • 🎯 Quiz — 5 or 10 mixed MCQ/short-answer questions, auto-scored with weak-area identification
  • ⏱️ Exam Mode — "I have 30 minutes" → prioritized study plan (Must Know / Good to Know / Skim) with time estimates

Demo

ollama pull gemma3:1b
ollama serve
pip install -r requirements.txt
streamlit run app.py
Enter fullscreen mode Exit fullscreen mode

Opens at http://localhost:8501 — works on any machine that runs Python + Ollama (tested on macOS, Linux, Windows).


Code

github.com/ShashwatSinha03/HF_BuildForAFriend

study-lens/
├── app.py          # Streamlit UI + orchestration
├── llm.py          # Ollama client + prompts + robust JSON parsing
├── document.py     # PDF/TXT extraction + keyword retrieval
├── requirements.txt
└── README.md
Enter fullscreen mode Exit fullscreen mode

How I Built It

Stack: Python • Streamlit • Ollama • Gemma 3 1B • PyPDF • Requests

Architecture decisions that matter:

Constraint Solution
No cloud/API keys Local Ollama + Gemma 3 1B (815MB)
Small model, unreliable JSON Multi-strategy parser (standard → trailing-comma fix → duplicate-key removal → regex fallback)
No embeddings/vector DB Lightweight keyword-overlap retrieval (TF-IDF-inspired, ~30% of doc sent to LLM)
Limited compute Single LLM call per action; batch short-answer eval; conservative token limits
Privacy-first Zero network calls except localhost:11434; session-only state

Sprint-based development (6 sprints over 3 days):

  1. Foundation — upload, extraction, session state, Ollama health checks
  2. Explain — grounded Q&A with retrieval
  3. Revise — structured revision sheets
  4. Quiz — MCQ + short answer, batch evaluation, weak areas
  5. Exam Mode — time-bounded prioritization
  6. Polish — robust parsing, error UX, README, .gitignore

Why Open Innovation Matters

This project exists because of open weights and local inference.

If I'd used a closed API:

  • ❌ Cost — $20/month per user minimum
  • ❌ Privacy — lecture notes sent to third parties
  • ❌ Offline — no studying on planes, in libraries, or with spotty campus WiFi
  • ❌ Customization — can't tweak prompts, retrieval, or model behavior
  • ❌ Longevity — API deprecation, price hikes, or service shutdown kills the tool

With Gemma 3 1B + Ollama:

  • ✅ Free forever — runs on hardware she already owns
  • ✅ Private — her notes never leave her machine
  • ✅ Offline-first — works anywhere
  • ✅ Hackable — every prompt, retrieval strategy, and UI decision is in plain Python
  • ✅ Sustainable — no vendor lock-in; model runs as long as Python runs

The "small model" constraints forced better engineering: lightweight retrieval, strict prompting, defensive parsing, batch evaluation. Those are better patterns that scale up — not compromises.


The Friend Test

"Wait, it's actually free? And my PDFs don't get uploaded anywhere? Damn, broski"

Yep, that's my sister, after using it for her midterms prep

That's the whole point. Open innovation doesn't just lower barriers — it changes what's possible to build for the people you care about.

Top comments (0)