<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Noor Yasser</title>
    <description>The latest articles on DEV Community by Noor Yasser (@nooryasserx).</description>
    <link>https://dev.to/nooryasserx</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4059533%2Fbd025ca9-efbf-490b-a591-314719d65e38.png</url>
      <title>DEV Community: Noor Yasser</title>
      <link>https://dev.to/nooryasserx</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nooryasserx"/>
    <language>en</language>
    <item>
      <title>Production RAG Is an Evidence Pipeline, Not a Vector-Database Demo</title>
      <dc:creator>Noor Yasser</dc:creator>
      <pubDate>Fri, 09 Oct 2026 14:26:52 +0000</pubDate>
      <link>https://dev.to/nooryasserx/production-rag-is-an-evidence-pipeline-not-a-vector-database-demo-325o</link>
      <guid>https://dev.to/nooryasserx/production-rag-is-an-evidence-pipeline-not-a-vector-database-demo-325o</guid>
      <description>&lt;p&gt;A production RAG system is not “put documents in a vector database and ask an LLM.” It is an evidence pipeline with two versioned paths: an offline path that parses, chunks, embeds, and indexes authorized content; and an online path that retrieves, reranks, assembles bounded context, and asks the model to answer only from that evidence.&lt;/p&gt;

&lt;p&gt;This distinction matters because most real failures happen before generation. The parser drops a table, the chunk loses its heading, the tenant filter is missing, or a deleted source survives in the index. A fluent answer can hide all of those problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  First decide whether RAG is the right tool
&lt;/h2&gt;

&lt;p&gt;Use RAG when answers depend on large, changing, unstructured knowledge and provenance matters: policies, manuals, contracts, support histories, or technical documentation.&lt;/p&gt;

&lt;p&gt;Use something simpler when the problem is simpler:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SQL&lt;/strong&gt; for exact structured facts and transactional state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lexical search&lt;/strong&gt; for known identifiers, error codes, names, or quoted phrases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A deterministic API&lt;/strong&gt; for permissions, account state, and actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tuning&lt;/strong&gt; when the missing capability is behavior or style, not knowledge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full-context prompting&lt;/strong&gt; when a small, stable corpus fits safely inside the model context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The router that makes this choice is part of the product. Evaluate it like every other component.&lt;/p&gt;

&lt;h2&gt;
  
  
  Split indexing from query-time work
&lt;/h2&gt;

&lt;p&gt;The offline path should be idempotent and resume-safe:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Store the original document and its version.&lt;/li&gt;
&lt;li&gt;Extract structured text and preserve headings, pages, and source locations.&lt;/li&gt;
&lt;li&gt;Create chunks using document structure.&lt;/li&gt;
&lt;li&gt;Generate embeddings with an explicit model and pipeline version.&lt;/li&gt;
&lt;li&gt;Write searchable points and payload indexes.&lt;/li&gt;
&lt;li&gt;Mark the document searchable only after every required artifact is durable.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The online path owns the latency budget:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Authenticate the requester and derive tenant and ACL scope.&lt;/li&gt;
&lt;li&gt;Classify or rewrite the query only when justified.&lt;/li&gt;
&lt;li&gt;Run dense and lexical retrieval branches.&lt;/li&gt;
&lt;li&gt;Fuse and rerank a bounded candidate set.&lt;/li&gt;
&lt;li&gt;Hydrate canonical text through an authorized repository.&lt;/li&gt;
&lt;li&gt;Build cited context under a deterministic token budget.&lt;/li&gt;
&lt;li&gt;Generate an answer, validate citations, and abstain when evidence is insufficient.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Join both paths with stable identifiers such as tenant_id, document_id, document_version, chunk_id, source location, checksum, embedding model, and index version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunk by meaning and structure
&lt;/h2&gt;

&lt;p&gt;A single global chunk size is rarely a production strategy. Prefer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;headings and paragraphs for prose;&lt;/li&gt;
&lt;li&gt;function or class boundaries for code;&lt;/li&gt;
&lt;li&gt;rows plus headers for tables;&lt;/li&gt;
&lt;li&gt;clauses for contracts;&lt;/li&gt;
&lt;li&gt;turns for conversations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A chunk should represent one retrieval idea while still carrying enough context to answer a useful sub-question. Large overlap is an expensive substitute for good boundaries because it multiplies embedding cost and returns near-duplicates.&lt;/p&gt;

&lt;p&gt;Parent-child retrieval is useful when narrow chunks search well but the answer needs surrounding explanation: index the child, then hydrate its parent section or immediate neighbors after retrieval.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat embeddings as a versioned contract
&lt;/h2&gt;

&lt;p&gt;An embedding belongs to a specific model, input mode, preprocessing pipeline, and output dimension. Changing any of those changes the vector space.&lt;/p&gt;

&lt;p&gt;Store the model, dimensions, and pipeline version with every point. Build a new named vector or collection for migrations, backfill it, compare quality and latency, switch reads gradually, and keep a rollback window.&lt;/p&gt;

&lt;p&gt;Never mark a document indexed until all intended chunks have the expected embedding version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Qdrant around retrieval and isolation
&lt;/h2&gt;

&lt;p&gt;A Qdrant collection should follow the embedding contract, not the number of business entities. Put filter and provenance fields next to the vector: tenant, document and version IDs, language, content type, ACL labels, timestamps, and source location.&lt;/p&gt;

&lt;p&gt;Create payload indexes for fields used in filters. A field being present in payload does not automatically make filtering efficient.&lt;/p&gt;

&lt;p&gt;For multitenancy, never accept a tenant filter directly from the public request. Derive it from authenticated scope and inject it into every retrieval branch, including dense, sparse, recommendation, and scroll endpoints.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieve broadly, rerank narrowly
&lt;/h2&gt;

&lt;p&gt;Dense retrieval catches semantic similarity. Lexical retrieval catches identifiers and exact terms. Hybrid retrieval runs both and fuses ranks to improve candidate coverage.&lt;/p&gt;

&lt;p&gt;Treat fusion as candidate generation, not final context. Rerank only a bounded candidate set because rerankers are more expensive. Skip reranking when the initial candidates already meet the quality target.&lt;/p&gt;

&lt;p&gt;Query rewriting can resolve pronouns or expand abbreviations, but preserve the original query and evaluate the rewrite independently. A generated rewrite can drift from user intent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build context as evidence
&lt;/h2&gt;

&lt;p&gt;Context construction should be deterministic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;remove duplicate and near-duplicate chunks;&lt;/li&gt;
&lt;li&gt;prefer the newest authorized version;&lt;/li&gt;
&lt;li&gt;merge adjacent fragments only when they share a source;&lt;/li&gt;
&lt;li&gt;cap evidence per document;&lt;/li&gt;
&lt;li&gt;reserve tokens for instructions, user input, and the answer;&lt;/li&gt;
&lt;li&gt;attach stable citation IDs to document versions and locations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Treat retrieved documents as untrusted data. They may contain prompt-injection strings. Retrieved text must never override system rules or request tools.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;candidates&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;dense&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;sparse&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ranked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;rerank&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;deduplicate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;candidates&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;hydrated&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;repository&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loadAuthorized&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;ranked&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chunkId&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;buildContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hydrated&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;maxTokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;maxPerDocument&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;includeSource&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;documentId&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;version&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;page&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;section&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;generateGroundedAnswer&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;requireCitations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those numbers are example budgets, not universal recommendations. Measure them against your corpus, model, and latency objective.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluate every gate
&lt;/h2&gt;

&lt;p&gt;A final-answer score is not enough. Measure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retrieval recall@k;&lt;/li&gt;
&lt;li&gt;ranking quality with MRR or nDCG;&lt;/li&gt;
&lt;li&gt;tenant and ACL isolation;&lt;/li&gt;
&lt;li&gt;reranker lift;&lt;/li&gt;
&lt;li&gt;evidence precision;&lt;/li&gt;
&lt;li&gt;citation validity;&lt;/li&gt;
&lt;li&gt;supported-claim rate;&lt;/li&gt;
&lt;li&gt;abstention quality;&lt;/li&gt;
&lt;li&gt;p50 and p95 latency;&lt;/li&gt;
&lt;li&gt;token, embedding, and reranking cost;&lt;/li&gt;
&lt;li&gt;queue age and index freshness.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Run ablations on the same evaluation set: lexical only, dense only, hybrid, hybrid plus reranking, neighbor expansion, and contextualized chunks. Change one component at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production checklist
&lt;/h2&gt;

&lt;p&gt;Before launch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;define which questions go to RAG, SQL, search, or tools;&lt;/li&gt;
&lt;li&gt;version the parser, chunker, embeddings, and prompts;&lt;/li&gt;
&lt;li&gt;index only complete document versions;&lt;/li&gt;
&lt;li&gt;derive authorization filters from trusted identity;&lt;/li&gt;
&lt;li&gt;create payload indexes for real filter fields;&lt;/li&gt;
&lt;li&gt;enforce a deterministic context budget;&lt;/li&gt;
&lt;li&gt;validate every citation;&lt;/li&gt;
&lt;li&gt;reconcile deletions against vector points;&lt;/li&gt;
&lt;li&gt;trace model and component versions;&lt;/li&gt;
&lt;li&gt;maintain permission-aware evaluation data;&lt;/li&gt;
&lt;li&gt;test cross-tenant retrieval and prompt injection.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical goal is not a clever demo. It is a system that fails closed, degrades predictably, and can explain which evidence produced every material claim.&lt;/p&gt;

&lt;p&gt;For the full bilingual guide, architecture diagram, and official references, read the canonical article: &lt;a href="https://nooryasser.com/articles/production-rag-pipeline-chunking-qdrant-reranking/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=personal_brand" rel="noopener noreferrer"&gt;Production RAG: from chunks to grounded LLM answers&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Prepared by Noor Yasser.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>rag</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
