DEV Community

Tony Henein
Tony Henein

Posted on Originally published at acaciatechgroup.com

How to Build a Robust RAG System: Lessons from a Live Production Architecture

Most RAG demos are animations; we built a live system to surface the hard truths of production AI. This article shares the lessons learned, focusing on how retrieval quality depends on structural chunking and explicit refusal gates.

How a production RAG pipeline builds grounded answers

This figure illustrates how a production RAG pipeline orchestrates components like Document Store, Retrieval, and LLM to build grounded answers.

Most RAG demos are animations. We wanted one you could actually break.

Our How RAG Works walkthrough explains retrieval-augmented generation step by step: chunk, embed, retrieve, generate. It is a good explainer. It is also a simulation, and a simulation can't show you the parts of RAG that only appear when real questions hit real content.

So we built the real thing. Below the walkthrough is a live panel: ask a data-architecture question, and it answers from the 50+ articles and glossary entries we've published, shows exactly what it retrieved and how closely each passage matched, and cites a source for every claim. When our content doesn't cover the question, it says so.

This article walks through how it's built, layer by layer, and the lessons that changed the RAG architecture design along the way. If you're planning a retrieval-augmented generation implementation over your own documents, the lessons are the useful part. (New to RAG? Start with RAG Explained: Giving Language Models a Library Card.)

The RAG architecture at a glance

Like every RAG system, it runs in two phases that happen at completely different times.

Indexing runs once, offline. We split our published articles and glossary into chunks, turn each chunk into an embedding (a vector that captures its meaning), and store the vectors in a vector index. We re-run it after publishing new content.

Querying runs on every question. The question passes a bot check, is embedded with the same model, matched against the index, filtered through RAG guardrails, and — only if it survives — answered by a language model that may use nothing but the retrieved passages.

Component

What we used

Why

Embeddings

BGE base (open source, 768 dimensions)

Strong retrieval quality, runs on managed GPUs

Vector index

Cloudflare Vectorize (~960 chunks)

Same platform as the site; no extra service to run

Answer model

Llama 3.3 70B (open-weight)

Good at the "is this answerable?" judgment, supports structured JSON output

Runtime

A Cloudflare Pages Function

Deploys with the website; no separate backend

Bot protection

Turnstile + a rate-limit rule

Stops scripts before any AI cost

Everything runs on open-source or open-weight models. That was deliberate: it keeps the system portable, and it's a fair test of whether you need a proprietary API to do RAG well. For a vector search strategy of this shape, you don't.

Indexing: chunking is where quality is decided

It's tempting to treat chunking as plumbing. It isn't. Retrieval can only ever return the chunks you created, so their boundaries decide what an answer can be built from.

Three decisions mattered most:

  • Split by section, not just by size. We split each article at its headings first, then into roughly 800-character chunks with about 150 characters of overlap, breaking at a paragraph, then a sentence, then a space. A chunk never crosses a heading, so every chunk knows which section it came from — which makes its citation meaningful.

  • Leave the noise out. We excluded reference lists and series-navigation sections. Indexed, they let a question match a list of links instead of the argument it was looking for.

  • Embed with context, store without it. Each chunk is embedded with its article title and section heading prefixed, so a passage that never names its subject ("it also cuts cost…") still retrieves for questions about that subject. The stored text stays clean for display.

Two operational details saved us later. Every chunk gets a deterministic ID derived from its URL and position, so re-indexing overwrites in place instead of duplicating. After each run, we delete any vector the current content no longer produces—a shortened or retired article— with a guard that refuses to run cleanup on an empty corpus.

Querying: the same model, or nothing works

The question must be embedded with the same model — and the same settings — used for the documents. Otherwise, query and documents live in different mathematical spaces, and similarity scores are quietly meaningless. No errors. Answers get worse.

We nearly hit a subtle version of this. The embedding API offers two "pooling" modes, and its default isn't the one the BGE models were trained with. Choosing one mode at indexing time and inheriting the default at query time would have produced exactly that silent failure. The fix was to pin the setting in one shared constant that both the indexer and the query path import.

Retrieval then fetches the 20 closest chunks and keeps only the single best chunk per article. Early tests returned five passages from the same post, which looks broken to a reader and starves the answer of breadth. The top four distinct articles become the numbered sources the model can use.

The hardest part: deciding what it may not say

Here is the finding that reshaped the enterprise AI architecture design. Retrieval always returns something. There is no natural "no results" — just less relevant results. So a system that answers whenever it retrieves will answer everything, confidently.

The obvious fix is a minimum similarity score. We measured it on the real index, and it doesn't work on its own:

Question

Best match score

Should it answer?

Should we use dbt or SQLMesh?

0.879

Yes

What does Acacia charge for a migration?

0.668

No — we've never published pricing

How do we keep an LLM from making things up about our data?

0.641

Yes

Our warehouse spend keeps climbing every quarter

0.597

Yes

How do I fix a leaking kitchen faucet?

~0.43

No

The pricing question outscores two questions our content genuinely answers. Any threshold strict enough to refuse it would also refuse real questions asked casually.

Then testing surfaced a worse failure. Asked whether we had SAP experience, the model found an article about SAP and answered "yes." It wasn't inventing from nothing — it was over-reading a real source. That's the failure mode that pulls RAG systems out of production, because it looks exactly like a grounded answer.

So refusal became three layers of RAG guardrails, each enforced in code rather than left to a prompt:

  1. An "ask us" gate, before any AI call. Questions about us — pricing, staffing, timelines, our own experience, clients or certifications — go to a person. Content can only ever answer those wrongly, and "talk to us" is the outcome we want anyway.

  2. A relevance floor, before generation. Clearly unrelated questions (the faucet scored 0.43) are refused without calling the model at all, so they cost nothing.

  3. A structured answerability check, after. The model returns JSON with its answer, its citations, and an explicit "answerable" flag. Code then verifies that every citation points to a source that was actually retrieved; an answer that cites nothing real is refused. The refusal wording is ours, not the model's.

One prompt-design detail mattered more than expected. With the "answerable" field first in the output schema, the model decided "no" before reading closely, and refused a cost question while looking at a source titled "Cost Risk: No Guardrails = 2–3× Your Expected Spend." Moving the answer field first — draft, then judge — fixed it.

If you take one idea from this article: grounding is necessary, but not sufficient. A trustworthy RAG architecture needs explicit rules about what it may claim. This is the same principle behind AI governance that actually works, applied at the level of a single answer.

Show the Retrieval, not just the answer.

The live panel shows the retrieved passages and their similarity scores before the answer, with weak matches dimmed. Anyone can wire up a chatbot; showing what was retrieved, and why, is what lets a reader judge whether the answer deserves trust. For a technical evaluator, it's the most interesting part of the system.

It also makes failures legible. When the answer is thin, you can usually see why in the Retrieval: the right article scored fourth, or the question needed a page that isn't indexed yet. That's the signal for the next improvements, such as hybrid search to catch exact terms that embeddings blur, and reranking to reorder candidates with a more precise model.

Cost, abuse, and keeping it running

A public endpoint backed by a 70-billion-parameter model is a free GPU for whoever finds it, so the controls came before launch:

  • Bot check per question. Every question carries a fresh, single-use token that is verified server-side before any AI call. Scripts are stopped at no cost.

  • Rate limit. A network-level rule caps each visitor at a handful of requests every few seconds.

  • Hard caps. Question length, request size, sources per answer, and answer length are all bounded.

  • Exact cost tracking. The inference API reports its own billing unit with every answer, and we log each request with it, so our monitoring shows real spend against the daily free allowance rather than an estimate. A typical answer costs about a tenth of a cent; refusals that skip the model cost nothing.

  • A kill switch. One setting turns the live panel off; the page falls back to the walkthrough, and nothing else changes.

The request log keeps the questions people ask (so we can see what's useful and tune refusals) and a one-way, salted code instead of any IP address, and deletes everything after 90 days.

A checklist for your own retrieval-augmented generation implementation

  • Chunk by structure first (headings, sections), then by size — and exclude navigation and reference noise.

  • Pin the embedding model and its settings in one place that indexing and querying both use.

  • Deduplicate retrieval by source document before generation.

  • Don't rely on a similarity threshold alone. Measure in- and out-of-scope questions on your real index first.

  • Route questions about your own organization — pricing, commitments, experience — to people, before any model sees them.

  • Ask for structured output with an explicit answerability flag, and validate citations in code.

  • Show the retrieved sources to users. It builds trust and makes failures diagnosable.

  • Add bot protection, rate limits, and cost tracking before the endpoint goes public.

RAG reduces hallucination; it doesn't eliminate it. What turns a demo into a production-grade RAG architecture is the layer around the model: what it may retrieve, what it may claim, and what it must refuse.

References

  1. What Matters in Production RAG

  2. Six Lessons Learned Building RAG Systems in Production

  3. RAG System Architecture: A Production Implementation Guide

  4. The hardest part of building a production-grade RAG system isn’t retrieval.

  5. The Ultimate Guide to Chunking Strategies for RAG Applications with Databricks

  6. Building RAG Systems: From Zero to Hero

  7. Building Robust RAG: An Architecture for Production-Ready Retrieval-Augmented Generation

  8. Production-Grade RAG: Architecture, Trade-offs, and Hard-Won Lessons

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The SAP-experience failure is a useful example of a source being relevant without supporting the proposed claim. Your citation-ID check protects against invented sources, but a real retrieved ID can still accompany an unsupported statement, so I would evaluate those two failure modes separately.

The single-best-chunk-per-article rule is another boundary worth testing. A question comparing two sections of the same article may need both passages even when document diversity is helpful overall. A small multi-section question set could show whether an article quota with a limited second passage preserves breadth without losing evidence needed for a complete answer.