DEV Community

Armin Burger
Armin Burger

Posted on

A RAG system in one command — what chimerai add rag actually wires up

RAG is one of those things where the tutorial is 30 lines and the production system is a
vector-database procurement process. This post is about the middle ground: a working retrieval
pipeline you can drop into a project in one command, what it actually installs, and the exact
point where you'd want to replace it.

The command

cd my-nextjs-app
npx chimerai add rag
Enter fullscreen mode Exit fullscreen mode

rag depends on the chat module, so if ai-chat isn't installed the CLI adds it first. What
you get is a self-contained Python AI service (FastAPI + LiteLLM) under services/ai/ with the
RAG module switched on:

services/ai/
├── config.py                    # pydantic settings, env loading
├── provider_client.py           # LiteLLM — multi-provider routing
├── main.py                      # FastAPI entry (generated from manifest)
├── services/
│   ├── rag_service.py           # ingest → retrieve → answer pipeline
│   ├── vector_store.py          # FAISS index + persistence
│   └── embedding_service.py     # text embeddings via LiteLLM
├── routes/rag_routes.py         # /api/rag/upload, /query, /stats
└── data/                        # FAISS index on disk
Enter fullscreen mode Exit fullscreen mode

The Next.js side gets proxy routes that forward to AI_SERVICE_URL (default
http://localhost:8002), so your frontend calls /api/rag/... and never talks to Python
directly. chimerai dev starts both.

The pipeline, in the order it runs

Ingest — you hand it raw text, it chunks and embeds. Chunking is
RecursiveCharacterTextSplitter with defaults you can see rather than guess:

self.text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200,
    length_function=len,
    separators=["\n\n", "\n", ". ", " ", ""],
)
Enter fullscreen mode Exit fullscreen mode

chunk_size=1000 is in characters, not tokens (length_function=len) — a detail that
matters if you're used to token-based splitters; 1000 chars is roughly 250 tokens of English.
The 200-char overlap is there so a fact split across a boundary still shows up whole in at
least one chunk. Each chunk keeps a pointer back to its source:

for j, chunk in enumerate(chunks):
    all_chunks.append(chunk)
    all_metadatas.append({
        **meta,
        "_chunk_index": j,
        "_chunk_total": len(chunks),
        "_source_doc_index": i,
    })
Enter fullscreen mode Exit fullscreen mode

Those _chunk_* fields are what let you show "source: page 3" in the UI instead of a bare
similarity score.

Store — a faiss.IndexFlatL2 on 1536-dim vectors (OpenAI text-embedding-ada-002). Flat
L2 means exact brute-force search: no approximation, so recall is perfect and the index just
scans everything. That's fine to tens of thousands of vectors and gets slow past that — the
code comment says as much:

# Use IndexFlatL2 for exact search (can be changed to IndexIVFFlat for large datasets)
self.index = faiss.IndexFlatL2(self.dimension)
Enter fullscreen mode Exit fullscreen mode

Retrieve + answer — the RAG chat is a plain "stuff context into the system prompt" flow:

relevant_docs = await vector_store.similarity_search(query, k=k)

context_parts = []
for i, doc in enumerate(relevant_docs, 1):
    context_parts.append(f"[Document {i}]\n{doc['text']}")
context = "\n\n".join(context_parts) if context_parts else "No relevant documents found."

system_message = f"""You are a helpful assistant. Use the following context to answer \
the user's question. If the context doesn't contain relevant information, say so and \
provide a general answer.

Context:
{context}"""
Enter fullscreen mode Exit fullscreen mode

The [Document i] labels exist so the model can cite, and the "if the context doesn't contain
the answer, say so" line is the single most important prompt sentence in the whole file —
without it the model answers from parametric memory and your eval numbers are fiction.

Using it

Add documents:

POST http://localhost:8002/api/rag/documents
{ "documents": ["FastAPI is a modern web framework for Python."],
  "metadatas": [{ "source": "docs", "page": 1 }] }
Enter fullscreen mode Exit fullscreen mode

Ask a grounded question:

POST http://localhost:8002/api/rag/chat
{ "query": "Tell me about FastAPI", "model": "gpt-3.5-turbo", "k": 3 }
Enter fullscreen mode Exit fullscreen mode

The response is a normal OpenAI-shaped completion plus a rag_metadata block that lists the
retrieved chunks and their scores — so you can render citations and debug "why did it answer
that" without a separate search call:

"rag_metadata": {
  "retrieved_documents": 3,
  "documents": [{ "text": "FastAPI is a modern web framework...",
                  "score": 0.123, "metadata": { "source": "docs" } }]
}
Enter fullscreen mode Exit fullscreen mode

There's also /api/rag/search (just retrieval, no LLM), /api/rag/stats, and
/api/rag/clear.

The FAISS import is guarded — and that's the point

On Python 3.13 / some Windows setups, faiss-cpu or numpy won't install cleanly. Rather than
crash the whole service at import time, the module degrades:

try:
    import faiss
    import numpy as np
    FAISS_AVAILABLE = True
except Exception as e:
    FAISS_AVAILABLE = False
    print(f"Warning: FAISS/Numpy not available: {e}")
Enter fullscreen mode Exit fullscreen mode

Every RAG entry point calls _check_availability() first and raises a clear
"FAISS vector store is not available" instead of an ImportError three frames deep. Chat,
guardrails, and tools keep working; only retrieval is off. Small design choice, but it's the
difference between "my RAG doesn't work" and "my whole app won't boot."

Persistence

The index is saved to data/faiss_index and metadata to a matching .pkl — loaded on startup,
written after ingest. It's a single-process, on-disk store with no locking. Which is exactly the
right scope for a starter and exactly the wrong scope for multi-instance:

When you outgrow it

Honest boundaries, so you know what you're buying:

  • Flat index → fine to ~tens of thousands of chunks, then switch to IndexIVFFlat/HNSW or a real vector DB (pgvector, Qdrant, Weaviate).
  • One global index → no per-tenant / per-user namespaces. Multi-tenant retrieval needs metadata filtering you'd add yourself, or a DB that does it.
  • Local disk, single writer → no horizontal scaling; two instances = two divergent indexes.
  • k is a fixed int → no hybrid (BM25 + dense) search, no re-ranking, no MMR diversification. Those are the usual next levers when retrieval quality plateaus.

None of that is a criticism of a scaffold — it's the list of what a "RAG in one command" is
not, so you replace the right thing at the right time instead of rewriting from scratch.

Try it

chimerai create my-rag-app --sqlite --yes
cd my-rag-app
chimerai add rag
chimerai dev
Enter fullscreen mode Exit fullscreen mode

--sqlite skips Docker for the app DB; the AI service still runs as its own Python process.

Repo: github.com/armbur19-collab/chimerai-kickstart

Top comments (0)