DEV Community

Cover image for Build Production-Grade RAG Pipeline with Metadata Filtering & Incremental Updates
Tidiane Stano
Tidiane Stano

Posted on

Build Production-Grade RAG Pipeline with Metadata Filtering & Incremental Updates

This tutorial continues the hands-on RAG learning series. By this stage, we move beyond simple vector search demos and build a complete enterprise-ready retrieval augmented generation workflow, addressing core production challenges like source tracing, document categorization, filtering rules and incremental knowledge base refresh.

1. The Core Pain Points of Basic RAG

A naive RAG implementation can retrieve relevant text chunks, but it suffers from several practical flaws when deployed for real business documents. Common unresolved questions include:

  1. Which original document does this retrieved chunk come from?
  2. How to distinguish information from different document categories (HR policy, office address, salary rules)?
  3. Can we restrict retrieval to only one specific document type for a user query?
  4. Do we need to re-embed every file in the knowledge base every time new content is added?

Metadata filtering is the key technique to solve these problems. Together with chunk splitting, recall tuning, precision control, Top-K selection, reranking and hybrid search, it forms the full optimization stack for industrial RAG systems.

2. Document Loader: Raw File to Text Chunks

Before embedding, PDF, Word, TXT and other file formats must go through document loading and text extraction. Document loaders parse binary files and convert them into plain text that embedding models can process.

After extraction, long text is split into multiple chunks. We attach structured metadata tags to every chunk. Metadata stores categorical attributes of the source content, such as document category, source file name, creation date, department and reference ID.

Sample metadata categories defined in this project:

  • category: salary for compensation related documents
  • category: work_time for attendance and working hour rules
  • category: vacation for leave policy
  • category: office for company location and office information
  • category: reference for supplementary reference materials

Once chunks and metadata are ready, we feed them into Chroma vector database. The embedding model converts chunk text into vector embeddings, then both embeddings, raw text and metadata are persisted inside the vector collection.

When stored in Chroma, each record contains three core fields:

  1. Vector embedding for semantic similarity search
  2. Original text chunk
  3. Structured metadata dictionary with category and source information

3. Metadata Filter: Narrow Down Search Scope Before Vector Retrieval

Metadata filters run before vector search. They use structured attributes to filter out irrelevant records, reducing the candidate pool before expensive similarity calculation.

Take the query “Where is our company located?” as an example. The question clearly belongs to office category. We add a metadata filter condition {"category": "office"} to the retrieval function.

The full retrieval pipeline with filter works as follows:

  1. Accept user question
  2. Generate query embedding
  3. Apply metadata filter to keep only chunks with matching category
  4. Run vector similarity search within the filtered subset
  5. Return top matching chunks
  6. Apply rerank model to reorder results
  7. Feed final context to LLM for answer generation

Without metadata filtering, the vector database scans all chunks from salary, vacation, attendance and office documents. With filtering enabled, only records tagged office enter the vector search phase. This cuts computation cost, reduces noise and prevents cross-category content leakage.

Why Metadata Filter Must Execute Before Vector Search

This ordering rule is critical. The correct sequence is:
Metadata Filter → Vector Search → Rerank → LLM

The wrong approach executes vector search first and applies filtering afterwards. Suppose your vector database contains 10,000 total chunks, and only 100 belong to the target category.

  • Bad workflow: compute vector similarity for all 10,000 entries, then filter, wasting massive compute resources
  • Good workflow: filter down to 100 matching records first, run vector search only on this small subset

Metadata filtering works exactly like WHERE clauses in SQL databases. Structured filters prune candidates, and vector search handles semantic matching within the remaining subset. Combining structured filtering and semantic retrieval is the foundation of reliable enterprise RAG.

4. Full Pipeline: Metadata Filter + Vector Search + Rerank

We assemble all components into one complete retrieval chain. The end-to-end steps are:

  1. Receive user query
  2. Encode the user query into embedding vector
  3. Run metadata filter to select qualified chunks
  4. Perform vector similarity search on filtered candidates
  5. Extract raw text of retrieved chunks
  6. Invoke reranking model to reorder retrieved results by relevance
  7. Assemble final context prompt
  8. Send prompt to LLM
  9. Return final answer to user

A common pitfall during development: modifying metadata schema does not automatically update existing records in Chroma. Metadata is saved together with embeddings during ingestion. Changing metadata definitions later will not rewrite old entries already stored in vector DB.

When schema changes, the simplest way in development phase is deleting the entire Chroma collection and rebuilding it from scratch.

  1. Drop existing collection
  2. Re-run document loading, chunking and embedding scripts
  3. Insert all chunks with updated metadata
  4. Validate collection metadata schema
  5. Run test queries to verify filtering effect

After rebuild, test the query about company address again. The metadata filter isolates office-related chunks, vector search finds semantically close content, rerank optimizes the sequence, and LLM generates accurate final answer.

5. Key Engineering Insight: RAG is a Multi-Step Pipeline

Many beginners misunderstand RAG as just vector search plus LLM. In production systems, RAG is a chained pipeline with discrete stages, and each stage can introduce errors.

Every stage needs independent inspection:

  1. Document loader: is text extracted correctly?
  2. Chunk splitter: is chunk segmentation reasonable?
  3. Embedding model: does embedding capture semantic meaning properly?
  4. Metadata filter: are filter conditions correctly applied?
  5. Vector search: does it retrieve semantically relevant chunks?
  6. Rerank model: does reranking improve ordering quality?
  7. LLM inference: does the model use provided context faithfully?

If the final answer is wrong, developers must debug each stage sequentially instead of only tuning the LLM prompt. This mindset separates hobby prototype building from real AI system engineering.

6. Incremental Update: Avoid Rebuilding Entire Knowledge Base

Rebuilding the whole vector library every time new documents arrive is not feasible for enterprise systems. Imagine a knowledge base holding 10,000 chunks. If a new PDF is added, naive full rebuild reprocesses all 10,000 old items plus the new content. The computation overhead grows linearly with dataset size.

Incremental update solves this issue. The workflow only processes newly added or modified files:

  1. Detect newly uploaded or changed documents
  2. Run loader, chunking and embedding only for these new files
  3. Insert new chunks and metadata into existing Chroma collection
  4. Keep all historical vector records untouched

This approach saves enormous compute and storage cost. In our local demo environment, we drop and rebuild collection for easier metadata validation. Production systems strictly adopt incremental ingestion and never perform full rebuild for routine updates.

7. Final Production RAG Architecture

After all optimizations, the mature RAG pipeline can be split into two major phases: Knowledge Base Build and Online Inference.

Phase A: Knowledge Base Build (Offline)

  1. Watch document storage folder for new or modified files
  2. Trigger document loader for updated files
  3. Parse raw files into plain text
  4. Split text into fixed-length chunks
  5. Generate metadata tags for each chunk
  6. Create embedding vectors for each chunk
  7. Insert chunks, embeddings and metadata into Chroma vector database

Phase B: Online Inference (User Query Time)

  1. Receive end user question
  2. Encode user query into embedding
  3. Apply metadata filter to narrow candidate chunks
  4. Execute vector similarity search
  5. Run reranking model to reorder retrieved results
  6. Compile retrieved chunks into context window
  7. Build prompt with system instruction, context and user question
  8. Call LLM to generate grounded answer
  9. Return answer and source references to end user

This architecture is close to commercial enterprise knowledge assistant systems. The whole pipeline decouples offline data processing and online query serving. For multi-model service orchestration, 4sapi as an API gateway can unify access control, request routing and logging for embedding, rerank and LLM endpoints.

8. Core Takeaways from Day 10

  1. RAG ≠ Vector Search
    Vector search is only one retrieval step. Production RAG is a full pipeline covering document parsing, chunking, metadata filtering, vector search, reranking and LLM generation. Each stage can affect final output quality.

  2. Metadata is indispensable for enterprise knowledge
    Vector search handles semantic similarity, metadata filter manages structured attributes. Combining the two enables precise, scoped retrieval. Users can query within a specified department, document type or time range. Metadata also supports source citation, so users know exactly which original document the answer comes from.

  3. Vector DB stores more than embeddings
    A usable knowledge base stores triplets: embedding vector, raw text chunk and structured metadata. Metadata is what makes retrieval controllable and auditable. Without metadata, you cannot implement filtering, source tracing or permission control.

  4. RAG is a pipeline, not a single prompt
    Do not treat RAG as simply feeding retrieved text into LLM. Every module must be observable and testable independently. Debugging RAG requires checking each component step by step, instead of only adjusting LLM prompts.

  5. The biggest gap between demo and production is data update
    Demo code often rebuilds the whole vector library every run. Real-world systems rely on incremental updates. Only changed documents are reprocessed and inserted into vector storage. This is mandatory when knowledge bases scale to tens of thousands of files.

9. Summary

We have transformed a simple RAG prototype into a practical enterprise pipeline. Metadata filtering greatly improves retrieval precision and reduces unnecessary vector computation. The ordered workflow Metadata Filter → Vector Search → Rerank → LLM becomes our standard blueprint. Incremental update eliminates full knowledge rebuild and makes long-term maintenance affordable.

The next step will introduce more advanced techniques, including hybrid search combining keyword search and vector retrieval, together with more advanced chunking strategies. This pipeline lays solid groundwork for building private knowledge assistants, internal document Q&A bots and enterprise AI search services.

International access: [https://4sapi.com](https://4sapi.com)
Domestic access: [https://4sapi.org](https://4sapi.org)

Top comments (0)