What We Built
LectureLens (Mozhi) turns a lecture recording or document into a study resource. You upload audio, a PDF, a Word file, a text file or an image, choose the language, and get a summary, key terms and a quiz. Then you can ask questions, and every answer comes only from your material and shows where it came from: a timestamp for audio, a page or section for documents. If the lecture doesn't cover the question, it says so instead of guessing.
Everything runs locally on a laptop with open-source tools, with no API keys and no cloud.
Code
https://github.com/bavishnu11/hacktoberfest-hack-day-coimbatore-x-init-club-and-idea-club
How It Works
-
Read:
faster-whispertranscribes audio with timestamps. PDFs, Word files, text files and images are read into page- or section-labelled pieces. - Index: the text is split into chunks, embedded with a multilingual model and stored in ChromaDB.
- Ask: a question is embedded, the closest chunks are retrieved, and Gemma (running in Ollama) answers using only those chunks.
- Cite: the sources under each answer come from the retrieval step, not from the model, so they can't be invented.
Tech Stack
| Part | Tool |
|---|---|
| Speech to text | faster-whisper |
| LLM | Gemma via Ollama |
| Embeddings | sentence-transformers (multilingual-e5-small) |
| Vector store | ChromaDB |
| UI | Gradio |
How I Used Open Source and AI
[This runs fully open-source AI models locally—using faster-whisper for timestamped speech-to-text transcription and Gemma via Ollama for grounded Q&A, summaries, and quizzes.]
Challenges and What I Learned
[Managing local performance constraints when running compute-heavy open-source AI models (faster-whisper and Gemma via Ollama) and ensuring strict RAG grounding. Practical experience in building modular end-to-end RAG pipelines with vector databases (ChromaDB), optimizing local multilingual embeddings, and structuring prompt constraints for accurate JSON quiz generation and local LLM orchestration.]
Team
| Member | Contribution |
|---|---|
| Sanjay Vijay | faster-whisper transcription, timestamps, language setting, transcript cleanup |
| Sujithbabu S S | Chunking, embeddings, ChromaDB, search function with timestamp metadata |
| Vishal P | Ollama setup, summary, key terms, quiz, Q&A prompt, multilingual answers |
| S J Bavishnu | Gradio app, repo and Git workflow, README, demo video, deployment, pitch |
Top comments (0)