This project is a Retrieval-Augmented Generation (RAG) web application built to demonstrate how a large language model can be made to answer questions reliably from a specific set of documents, rather than relying solely on its general training knowledge. The system was developed as part of an internship task requiring the integration of four core AI engineering skills: LLM API usage, prompt engineering, retrieval pipeline construction, and full-stack deployment, combined into a single, working, publicly accessible product.
At a technical level, the application follows the standard RAG architecture. Source documents are first split into overlapping text chunks to preserve semantic specificity. Each chunk is converted into a numerical vector representation using Google's Gemini embedding model. When a user submits a question, that question is embedded using the same model, and cosine similarity is computed against every stored chunk to identify the most semantically relevant sections of the source material. These top-matching chunks are then inserted into a carefully constrained prompt and passed to Gemini's chat model, which generates a response using only the retrieved context, explicitly declining to answer when the necessary information isn't present in the documents.
The application is built with Next.js and deployed on Vercel, using entirely free-tier infrastructure: no paid API subscriptions, no dedicated backend server, and no external vector database. Retrieval is handled through a lightweight in-memory similarity search, which is sufficient for small-to-medium document sets and avoids unnecessary infrastructure overhead for a project of this scope.
A key design decision was transparency and verifiability. Rather than presenting a black-box chatbot, the interface explicitly displays which documents currently ground the assistant's answers, and includes a live upload feature: any user can add their own .txt file at runtime and immediately query it, without needing to redeploy the application or access its codebase. This makes the system's retrieval behaviour independently testable, a user can confirm, in real time, that answers are actually derived from the supplied content rather than the model's pre-existing knowledge.
Several practical lessons emerged during development. Retrieval quality, driven by chunking strategy and the number of chunks returned, had a greater impact on answer accuracy than adjustments to the prompt itself. Effective prompt engineering proved to be less about clever phrasing and more about imposing strict behavioural constraints, particularly instructing the model to acknowledge uncertainty rather than fabricate an answer. Meeting production, readiness standards also required attention beyond the core AI logic: graceful error handling for failed API calls, clear loading states during retrieval and generation, and secure handling of API credentials outside of version control.
The resulting application demonstrates that a functional, trustworthy retrieval-grounded AI system can be built and deployed end-to-end using freely available models and infrastructure, without requiring a dedicated backend, paid API access, or a managed vector database, making the approach accessible for small-scale, resource-constrained projects.
For further actions, you may consider blocking this person and/or reporting abuse
Top comments (0)