Building a RAG App That Knows When to Say "I Don't Know"
Large language models are impressive, but they answer confidently even when they have no idea. For my internship milestone at Valentius Kryptix, I set out to fix that by building a Retrieval-Augmented Generation (RAG) app: an assistant that answers only from a document I give it. RAG works by looking up the relevant text first and then asking the model to answer from it, which keeps answers grounded and checkable.
What I built
A small web app where you can chat with a set of machine learning study notes stored as a PDF. Ask "What test RMSE did the random forest reach?" and it replies from the notes and shows the source passage. Ask something unrelated, like the capital of France, and it politely says it could not find that in the documentation. You can try it here: https://rag-study-assistant-fro0.onrender.com/
How it works
- Chunk: the PDF text is split into pieces of about 800 characters with a small overlap, so ideas are not cut in half.
- Embed: each chunk is turned into a vector with a small open-source embedding model that runs locally, so no paid API is needed.
- Index: the vectors are stored in FAISS for fast similarity search.
- Retrieve: the user's question is embedded the same way, and the four closest chunks are fetched.
- Generate: those chunks are sent to an LLM through Groq's API, with strict instructions to answer only from the context.
The backend is FastAPI, the interface is a single HTML page, and the whole thing is deployed on Render.
Three takeaways
Retrieval quality decides everything. Chunk size, overlap and a similarity threshold determine whether the app answers correctly or admits it does not know. A chunk that is too small loses context, and one that is too large buries the answer. Tuning the threshold is what made off-topic questions get refused.
Prompts are part of the product. A clear system prompt, two few-shot examples and structured JSON output (an answer, an "answerable" flag and source numbers) made responses consistent and easy to cite.
Shipping teaches the most. Keeping API keys in environment variables, showing errors clearly in the interface, and debugging problems that only appeared on the server taught me more than any tutorial.
What's next
I want to try chunking by headings, support several documents at once, and measure answer quality with a small set of test questions.
Read the full article: 3 Powerful Lessons from Building My First RAG-Powered AI App
Top comments (0)