This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Darvin is a terminal + FastAPI RAG chat where you give it the path to a PDF (or upload via API), and it answers questions only from that PDF, with citations.
Who I built it for: My friend is preparing for semester exams and has 10+ bulky PDFs (notes, previous papers, manuals). Ctrl+F wasn't cutting it, and uploading everything to ChatGPT felt wrong for his personal notes. He needed a patient practice partner that would quiz him and answer strictly from his material without hallucinating.
So I built Darvin for him: drop in his PDFs, ask in plain English, get grounded answers + follow-ups with chat memory.
What he said after trying it: "Bro this actually cites the page number, now I don't have to scroll 200 pages."
Demo
- Test video: (youtube)[https://www.youtube.com/watch?v=r6MwwP4BkHU]
- GitHub: https://github.com/Rishavj90/darvin
Demo flow:
-
python test.py->prompt > - Ingest PDF via
POST /upload(checks.pdf, parses with PyMuPDF, chunks, stores) - Ask via
POST /askor terminal loop -> answer + sources + contexts + updated history
Code
https://github.com/Rishavj90/darvin
Stack: Python, FastAPI, LangChain, langchain-groq, ChromaDB, sentence-transformers, PyMuPDF
Key files:
-
backend/api.py-/uploadand/askroutes -
backend/ingestion/parse.py, chunk.py, store.py- PDF -> pages -> chunks -> Chroma -
backend/query/result.py- top-5 retrieval + grounded prompt +ChatGroq -
backend/query/memory.py- rewrites follow-up into standalone question using last 5 turns -
test.py- terminal chat loop
How I Built It
Open-source AI at the core - 2 layers:
-
Open-weight model:
llama-3.1-8b-instant(Meta Llama 3.1 8B weights). I run inference through Groq for speed, but the brain itself is open weights - I can swapGROQ_MODELin.envto any other Llama / Mistral / Gemma endpoint without changing code. -
Open-source harness + retrieval:
LangChain + langchain-text-splitters + ChromaDB + sentence-transformers. No black-box vector API. Embeddings are local (EMBEDDING_MODEL), stored locally (CHROMA_PATH,COLLECTION_NAME).
Pipeline:
PDF -> PyMuPDF parse -> LangChain splitter chunk_pages -> sentence-transformers encode -> Chroma store
Query -> encode -> Chroma top-5 -> build context with [Source, Page] -> Llama-3.1 prompt (strict: "ONLY use context, say 'I cannot find...' otherwise, cite after each statement") -> answer
Follow-up -> memory.get_standalone_question() rewrites using history -> ask()
.example.env:
EMBEDDING_MODEL=
CHROMA_PATH=
COLLECTION_NAME=
GROQ_API_KEY=
GROQ_MODEL=llama-3.1-8b-instant
Why terminal first? My friend lives in terminal, and I had ~1 weekend. FastAPI is already there so I can add a Streamlit / MERN UI later (I already ship MERN on Render).
Why Does Open Innovation Matter?
This project only works because the pieces are open:
- Swappable brain: Closed API would lock the prompt + behavior. With Llama open weights, if Groq pricing/limits change, I move the same prompt to Ollama locally, Together, or self-hosted vLLM. My friend can even run a smaller quantized Llama fully offline in hostel with no internet.
- Inspectable retrieval: LangChain splitters + Chroma + sentence-transformers are all open. When answers were bad, I could actually debug chunk size, embedding model, and top-k=5. Closed RAG-in-a-box hides that.
-
Zero cost to gift: No per-seat SaaS. He clones the repo, sets 5 env vars, runs
pip install -r requirements.txt. Data stays in his./chromaDBand./tmp_files, not on someone else's server. - No hallucination by design: The system prompt forces grounding + citations. Because I control the harness, I can enforce "do not use outside knowledge" - something you can't guarantee with a generic chatbot.
Closed would have been faster for a demo, but open made it giftable, debuggable, and portable.
My Agent Session
Built mostly by hand + terminal testing. No DevRelay session to embed.
Prize Categories
None - not using a partner tech this weekend. Entering for overall + completion badge.
Top comments (0)