A production RAG system is not “put documents in a vector database and ask an LLM.” It is an evidence pipeline with two versioned paths: an offline path that parses, chunks, embeds, and indexes authorized content; and an online path that retrieves, reranks, assembles bounded context, and asks the model to answer only from that evidence.
This distinction matters because most real failures happen before generation. The parser drops a table, the chunk loses its heading, the tenant filter is missing, or a deleted source survives in the index. A fluent answer can hide all of those problems.
First decide whether RAG is the right tool
Use RAG when answers depend on large, changing, unstructured knowledge and provenance matters: policies, manuals, contracts, support histories, or technical documentation.
Use something simpler when the problem is simpler:
- SQL for exact structured facts and transactional state.
- Lexical search for known identifiers, error codes, names, or quoted phrases.
- A deterministic API for permissions, account state, and actions.
- Fine-tuning when the missing capability is behavior or style, not knowledge.
- Full-context prompting when a small, stable corpus fits safely inside the model context.
The router that makes this choice is part of the product. Evaluate it like every other component.
Split indexing from query-time work
The offline path should be idempotent and resume-safe:
- Store the original document and its version.
- Extract structured text and preserve headings, pages, and source locations.
- Create chunks using document structure.
- Generate embeddings with an explicit model and pipeline version.
- Write searchable points and payload indexes.
- Mark the document searchable only after every required artifact is durable.
The online path owns the latency budget:
- Authenticate the requester and derive tenant and ACL scope.
- Classify or rewrite the query only when justified.
- Run dense and lexical retrieval branches.
- Fuse and rerank a bounded candidate set.
- Hydrate canonical text through an authorized repository.
- Build cited context under a deterministic token budget.
- Generate an answer, validate citations, and abstain when evidence is insufficient.
Join both paths with stable identifiers such as tenant_id, document_id, document_version, chunk_id, source location, checksum, embedding model, and index version.
Chunk by meaning and structure
A single global chunk size is rarely a production strategy. Prefer:
- headings and paragraphs for prose;
- function or class boundaries for code;
- rows plus headers for tables;
- clauses for contracts;
- turns for conversations.
A chunk should represent one retrieval idea while still carrying enough context to answer a useful sub-question. Large overlap is an expensive substitute for good boundaries because it multiplies embedding cost and returns near-duplicates.
Parent-child retrieval is useful when narrow chunks search well but the answer needs surrounding explanation: index the child, then hydrate its parent section or immediate neighbors after retrieval.
Treat embeddings as a versioned contract
An embedding belongs to a specific model, input mode, preprocessing pipeline, and output dimension. Changing any of those changes the vector space.
Store the model, dimensions, and pipeline version with every point. Build a new named vector or collection for migrations, backfill it, compare quality and latency, switch reads gradually, and keep a rollback window.
Never mark a document indexed until all intended chunks have the expected embedding version.
Model Qdrant around retrieval and isolation
A Qdrant collection should follow the embedding contract, not the number of business entities. Put filter and provenance fields next to the vector: tenant, document and version IDs, language, content type, ACL labels, timestamps, and source location.
Create payload indexes for fields used in filters. A field being present in payload does not automatically make filtering efficient.
For multitenancy, never accept a tenant filter directly from the public request. Derive it from authenticated scope and inject it into every retrieval branch, including dense, sparse, recommendation, and scroll endpoints.
Retrieve broadly, rerank narrowly
Dense retrieval catches semantic similarity. Lexical retrieval catches identifiers and exact terms. Hybrid retrieval runs both and fuses ranks to improve candidate coverage.
Treat fusion as candidate generation, not final context. Rerank only a bounded candidate set because rerankers are more expensive. Skip reranking when the initial candidates already meet the quality target.
Query rewriting can resolve pronouns or expand abbreviations, but preserve the original query and evaluate the rewrite independently. A generated rewrite can drift from user intent.
Build context as evidence
Context construction should be deterministic:
- remove duplicate and near-duplicate chunks;
- prefer the newest authorized version;
- merge adjacent fragments only when they share a source;
- cap evidence per document;
- reserve tokens for instructions, user input, and the answer;
- attach stable citation IDs to document versions and locations.
Treat retrieved documents as untrusted data. They may contain prompt-injection strings. Retrieved text must never override system rules or request tools.
const candidates = await retrieve({ query, scope, dense: 40, sparse: 40 });
const ranked = await rerank(query, deduplicate(candidates));
const hydrated = await repository.loadAuthorized(
scope,
ranked.map((item) => item.chunkId),
);
const context = buildContext(hydrated, {
maxTokens: 7_000,
maxPerDocument: 3,
includeSource: ['documentId', 'version', 'page', 'section'],
});
return generateGroundedAnswer({
query,
context,
requireCitations: true,
});
Those numbers are example budgets, not universal recommendations. Measure them against your corpus, model, and latency objective.
Evaluate every gate
A final-answer score is not enough. Measure:
- retrieval recall@k;
- ranking quality with MRR or nDCG;
- tenant and ACL isolation;
- reranker lift;
- evidence precision;
- citation validity;
- supported-claim rate;
- abstention quality;
- p50 and p95 latency;
- token, embedding, and reranking cost;
- queue age and index freshness.
Run ablations on the same evaluation set: lexical only, dense only, hybrid, hybrid plus reranking, neighbor expansion, and contextualized chunks. Change one component at a time.
Production checklist
Before launch:
- define which questions go to RAG, SQL, search, or tools;
- version the parser, chunker, embeddings, and prompts;
- index only complete document versions;
- derive authorization filters from trusted identity;
- create payload indexes for real filter fields;
- enforce a deterministic context budget;
- validate every citation;
- reconcile deletions against vector points;
- trace model and component versions;
- maintain permission-aware evaluation data;
- test cross-tenant retrieval and prompt injection.
The practical goal is not a clever demo. It is a system that fails closed, degrades predictably, and can explain which evidence produced every material claim.
For the full bilingual guide, architecture diagram, and official references, read the canonical article: Production RAG: from chunks to grounded LLM answers.
Prepared by Noor Yasser.
Top comments (5)
hello
Some comments may only be visible to logged-in visitors. Sign in to view all comments.