DEV Community

Bedirhan Kızılcık
Bedirhan Kızılcık

Posted on

Part 1-3: From Raw Health Text to a Working Retrieval Pipeline

I'm a first-year undergraduate student in Artificial Intelligence Engineering, and this summer I'm building a small RAG (Retrieval-Augmented Generation) project from scratch to learn how these systems actually work under the hood — not just by calling an API, but by building the pipeline piece by piece. This is a log of the first three parts of that journey.

Part 1: Understanding the Building Blocks

Before writing any real project code, I spent time understanding the core ideas behind RAG:

  • Calling an LLM API: sending a prompt programmatically and getting a response back. I used the Gemini API for this. llm_test.py
  • Embeddings: the idea that text can be converted into numerical vectors, where semantically similar sentences end up close to each other in vector space. I tested this with a handful of sentences using sentence-transformers and computed cosine similarity between them — it was satisfying to see related sentences actually cluster together numerically. embedding_test.py
  • Why not just dump everything into the LLM?: context windows are limited, it's expensive, and irrelevant information can actually hurt answer quality rather than help it.
  • Vector databases: conceptually, tools like Chroma or FAISS solve the problem of searching through large numbers of embeddings quickly to find the most relevant ones. No heavy coding in Part 1 — mostly small test scripts to confirm I understood each piece before combining them.

Part 2: Collecting and Preparing Real Data

With the concepts in place, Part 2 was about getting real data ready for retrieval:

  • Data collection: I gathered short health topic descriptions (e.g. diabetes, epilepsy) and saved them as structured JSON records, each with a source and text field. Keeping the source attached to each record matters — it means the system can eventually point back to where an answer came from. data/health_data.json
  • Chunking: I wrote a script to split each record's text into smaller, paragraph-level pieces while keeping the original source attached to every chunk. Since I intentionally kept the raw text short for this first test run, most entries ended up as a single chunk each — which is fine for now, since the goal was to validate the pipeline, not build a large dataset yet.
  • Why chunk at all?: embedding models represent short, focused pieces of text more accurately than long documents. Splitting text into meaningful chunks is what makes retrieval actually useful later.

Part 3: Setting Up the Vector Store and First Retrieval

With chunked, source-tagged data ready, Part 3 turned that data into something actually searchable:

  • Vector store setup: installed and configured Chroma, then embedded every chunk from Part 2 and stored it alongside its text and source metadata. embed_and_store.py
  • First retrieval test: wrote a query script, asked a sample question, and retrieved the most similar chunks from the vector store. The results were reasonable given how small and short the dataset still is — a good early sign that the pipeline itself works end to end. retrieval_test.py
  • A note on limitations: since the dataset is intentionally tiny at this stage, retrieval quality is limited. That's expected, and it'll improve as the dataset grows in later parts.

What's Next

At this point I have a full (if small-scale) pipeline: raw text → structured data → chunks → embeddings → retrieval. The next step is connecting retrieval to actual answer generation — feeding retrieved chunks into the LLM as context so it can answer health questions grounded in real source material.

👉 Code for this project: health-rag-assistant


Follow along as I build this project part by part — code on GitHub, progress here on Dev.to.

Top comments (3)

Collapse
 
david_crystal_67d49f98f74 profile image
David Crystal

This is a great approach, especially for a first year. You do not only apply the RAG process but you also understand the rationale of every layer, which is something many struggle with when working with such pipelines
I appreciate your focus on validation of each step, with emphasis on the importance of chunking, embeddings and source tracking, which is key to building a solid pipeline, which will be further upgraded when scaling the data set and improving retrieval performances.
I am looking forward to seeing how you will be integrating retrieval within the generation process, and how you will be designing your context and prompts.
I have experience with building such pipelines and architectures, and I am happy to share the knowledge and discuss potential collaborations if you are open to it.

Collapse
 
bedirhankizilcik profile image
Bedirhan Kızılcık • Edited

Thanks a lot, this really means a lot coming from someone with
hands-on experience in this space!

Understanding why each layer exists (not just making it work) has
been my main focus so far, so I'm glad that came across.

For the retrieval-generation part, my current plan is to start simple, pass the retrieved chunks into the prompt with their sources attached,
then iterate based on where the answers go wrong. I'm sure I'll run
into some interesting edge cases once I actually connect the two.

Would definitely be interested in hearing more about your experience
and exploring a collaboration.

Collapse
 
david_crystal_67d49f98f74 profile image
David Crystal

Appreciate your comment - your approach with iteration over a simple solution and using the failure cases for improvement is exactly the one that allows to build such systems correctly. Retrieval -> prompt -> assessment -> iteration indeed seems to be the right track, as it allows to identify the problems with context relevance and ranking quickly enough.
I also like how you’re attaching documents for easy inspection before responding - this will help a lot downstream when an incorrect answer slips through, and will contribute to trustworthiness and transparency as well.
Would you like me to share some thoughts and observations about my own experience with similar pipelines, or is there something specific you’re looking for? There’s quite a bit to say about the interesting cases that arise when combining retrieval and generation (in particular, about context length limitations and contradictions between documents). Let me know if you’d like to discuss collaboration as well!