RAG is one of those things where the tutorial is 30 lines and the production system is a
vector-database procurement process. This post is about the middle ground: a working retrieval
pipeline you can drop into a project in one command, what it actually installs, and the exact
point where you'd want to replace it.
The command
cd my-nextjs-app
npx chimerai add rag
rag depends on the chat module, so if ai-chat isn't installed the CLI adds it first. What
you get is a self-contained Python AI service (FastAPI + LiteLLM) under services/ai/ with the
RAG module switched on:
services/ai/
├── config.py # pydantic settings, env loading
├── provider_client.py # LiteLLM — multi-provider routing
├── main.py # FastAPI entry (generated from manifest)
├── services/
│ ├── rag_service.py # ingest → retrieve → answer pipeline
│ ├── vector_store.py # FAISS index + persistence
│ └── embedding_service.py # text embeddings via LiteLLM
├── routes/rag_routes.py # /api/rag/upload, /query, /stats
└── data/ # FAISS index on disk
The Next.js side gets proxy routes that forward to AI_SERVICE_URL (default
http://localhost:8002), so your frontend calls /api/rag/... and never talks to Python
directly. chimerai dev starts both.
The pipeline, in the order it runs
Ingest — you hand it raw text, it chunks and embeds. Chunking is
RecursiveCharacterTextSplitter with defaults you can see rather than guess:
self.text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
length_function=len,
separators=["\n\n", "\n", ". ", " ", ""],
)
chunk_size=1000 is in characters, not tokens (length_function=len) — a detail that
matters if you're used to token-based splitters; 1000 chars is roughly 250 tokens of English.
The 200-char overlap is there so a fact split across a boundary still shows up whole in at
least one chunk. Each chunk keeps a pointer back to its source:
for j, chunk in enumerate(chunks):
all_chunks.append(chunk)
all_metadatas.append({
**meta,
"_chunk_index": j,
"_chunk_total": len(chunks),
"_source_doc_index": i,
})
Those _chunk_* fields are what let you show "source: page 3" in the UI instead of a bare
similarity score.
Store — a faiss.IndexFlatL2 on 1536-dim vectors (OpenAI text-embedding-ada-002). Flat
L2 means exact brute-force search: no approximation, so recall is perfect and the index just
scans everything. That's fine to tens of thousands of vectors and gets slow past that — the
code comment says as much:
# Use IndexFlatL2 for exact search (can be changed to IndexIVFFlat for large datasets)
self.index = faiss.IndexFlatL2(self.dimension)
Retrieve + answer — the RAG chat is a plain "stuff context into the system prompt" flow:
relevant_docs = await vector_store.similarity_search(query, k=k)
context_parts = []
for i, doc in enumerate(relevant_docs, 1):
context_parts.append(f"[Document {i}]\n{doc['text']}")
context = "\n\n".join(context_parts) if context_parts else "No relevant documents found."
system_message = f"""You are a helpful assistant. Use the following context to answer \
the user's question. If the context doesn't contain relevant information, say so and \
provide a general answer.
Context:
{context}"""
The [Document i] labels exist so the model can cite, and the "if the context doesn't contain
the answer, say so" line is the single most important prompt sentence in the whole file —
without it the model answers from parametric memory and your eval numbers are fiction.
Using it
Add documents:
POST http://localhost:8002/api/rag/documents
{ "documents": ["FastAPI is a modern web framework for Python."],
"metadatas": [{ "source": "docs", "page": 1 }] }
Ask a grounded question:
POST http://localhost:8002/api/rag/chat
{ "query": "Tell me about FastAPI", "model": "gpt-3.5-turbo", "k": 3 }
The response is a normal OpenAI-shaped completion plus a rag_metadata block that lists the
retrieved chunks and their scores — so you can render citations and debug "why did it answer
that" without a separate search call:
"rag_metadata": {
"retrieved_documents": 3,
"documents": [{ "text": "FastAPI is a modern web framework...",
"score": 0.123, "metadata": { "source": "docs" } }]
}
There's also /api/rag/search (just retrieval, no LLM), /api/rag/stats, and
/api/rag/clear.
The FAISS import is guarded — and that's the point
On Python 3.13 / some Windows setups, faiss-cpu or numpy won't install cleanly. Rather than
crash the whole service at import time, the module degrades:
try:
import faiss
import numpy as np
FAISS_AVAILABLE = True
except Exception as e:
FAISS_AVAILABLE = False
print(f"Warning: FAISS/Numpy not available: {e}")
Every RAG entry point calls _check_availability() first and raises a clear
"FAISS vector store is not available" instead of an ImportError three frames deep. Chat,
guardrails, and tools keep working; only retrieval is off. Small design choice, but it's the
difference between "my RAG doesn't work" and "my whole app won't boot."
Persistence
The index is saved to data/faiss_index and metadata to a matching .pkl — loaded on startup,
written after ingest. It's a single-process, on-disk store with no locking. Which is exactly the
right scope for a starter and exactly the wrong scope for multi-instance:
When you outgrow it
Honest boundaries, so you know what you're buying:
-
Flat index → fine to ~tens of thousands of chunks, then switch to
IndexIVFFlat/HNSW or a real vector DB (pgvector, Qdrant, Weaviate). - One global index → no per-tenant / per-user namespaces. Multi-tenant retrieval needs metadata filtering you'd add yourself, or a DB that does it.
- Local disk, single writer → no horizontal scaling; two instances = two divergent indexes.
-
kis a fixed int → no hybrid (BM25 + dense) search, no re-ranking, no MMR diversification. Those are the usual next levers when retrieval quality plateaus.
None of that is a criticism of a scaffold — it's the list of what a "RAG in one command" is
not, so you replace the right thing at the right time instead of rewriting from scratch.
Try it
chimerai create my-rag-app --sqlite --yes
cd my-rag-app
chimerai add rag
chimerai dev
--sqlite skips Docker for the app DB; the AI service still runs as its own Python process.
Repo: github.com/armbur19-collab/chimerai-kickstart
Top comments (0)