Most RAG frameworks tell you whether an answer was generated. RAG-LCC tries to show you why that answer happened.
Instead of treating retrieval as a black box, you can inspect retrieval decisions, grounding signals, safety checks, and confidence traces while tuning your pipeline. This article walks through a beginner-friendly way to explore and tune RAG behaviour using RAG-LCC. RAG-LCC is not just a chatbot. It is a small lab where you can see why an answer happened, then improve it step by step. Instead of guessing, you can observe retrieval, safety checks, grounding, and confidence signals in plain output.
By default, RAG-LCC is designed to work well with local AI stacks. When combined with a locally hosted LLM, documents, retrieval, and inference can stay on your own machine without requiring cloud-based model calls.
Important: RAG-LCC is not meant for production deployment.
It is designed for learning, trying ideas, experimentation, and tuning.
For legal and licensing details, see LEGAL.md.
Who is this for?
RAG-LCC may be useful if you want to:
- learn how retrieval actually works
- compare retrieval strategies
- experiment with grounding
- evaluate confidence signals
- understand why a RAG answer changed after tuning
It is not intended as a production-ready enterprise platform. It is an experimental learning and tuning environment.
The big picture first
The image above shows the full system. Here is the simpler mental model:
flowchart LR
A[DocClassify: understand your corpus] --> F[Optional filter: classification criteria]
F --> B[RAGLoad: prepare searchable stores]
B --> C[RAGChat: ask questions in CLI]
C --> D[RAGChatService: serve the same flow via API]
D --> E[OpenWebUI or clients]
How to read this:
- DocClassify helps you understand what is in your documents before indexing.
- Optional filter step: use DocClassify results as classification criteria to choose what RAGLoad should ingest.
- RAGLoad builds the searchable memory (vector, keyword, graph, regex).
- RAGChat is the interactive place where you test and tune behavior.
- RAGChatService lets you use the same tuned behavior in OpenWebUI or API clients.
Why beginners might like RAG-LCC
Many RAG tools feel like a black box. RAG-LCC is different because it shows its work.
You can open QUERY_OUTPUT_EXAMPLE.md and literally watch:
- startup checks
- retrieval decisions
- Query rewrite and translation
- safety/compliance checks
- grounding hints
- confidence-style signals
That makes learning faster and more fun, because each tweak has visible effects.
A gentle way to start (no deep config knowledge needed)
Before running the apps, use the guided installer once:
python ./src/Scripts/Setup.py
Setup.py supports both installation paths:
- Docker-based setup
- Native host setup in a local
.venv(Windows or Unix)
Run the apps in this order:
python ./src/Apps/DocClassify.py --doc-dir TestDocs
python ./src/Apps/RAGLoad.py --doc-dir TestDocs
python ./src/Apps/RAGChat.py --doc-dir TestDocs
Optional filter step between DocClassify and RAGLoad (replace the default RAGLoad line above):
python ./src/Apps/RAGLoad.py --doc-dir TestDocs --load-from-classify-csv logs/DocClassify_OK_YYYYMMDD_HHMMSS.csv --classify-csv-query "Animal LIKE '%hedgehog%' OR Animal LIKE '%cat%'"
This lets you load only documents that match your classification criteria.
During a RAGChat session, relevant settings can be overwritten interactively (for example strategy=..., threshold=..., web_search=..., collection=..., or picker commands like strategy!, orchestrator_flow!, and collection!).
That makes experimentation user-friendly, because you can try changes live without editing config files between turns.
Then ask one simple question in chat. Keep that same question while you try different options.
Options you can explore without getting lost
You do not need to memorize config keys to start. Think in terms of behavior:
- Focused vs broad retrieval
You can move from very focused (
NARROW) to very broad (ULTRA_WIDE) search behavior. - More strict vs more flexible safety behavior RAG-LCC can be tuned to block/mask more aggressively or more permissively.
- Local-only vs web-assisted retrieval Keep answers local, or allow internet-assisted retrieval when needed.
- Short context vs long context Keep responses tight, or allow broader evidence gathering.
- CLI-first vs service-first workflow Tune in RAGChat, then serve the same behavior through RAGChatService.
When you are ready for details, the deep reference is here:
CONFIGURATION_REFERENCE.md
The configuration files are "Theme" oriented. This helps finding the right knobs.
For experts: orchestrator and query rewrite
If you already know RAG patterns and want finer control, RAG-LCC has two advanced power areas.
- Orchestrator flows You can shape how each turn runs: which retrieval legs activate, when reranking applies, how fallback behavior works, and whether grounding is enforced.
Query rewrite
You can refine follow-up understanding: pronoun resolution, topic carry-over, language normalization, and alternate-query expansion.Content filtering at two stages (reality check)
RAG-LCC supports filtering at prompt level and pipeline level, and each app uses this differently:
- RAGLoad: can reject/skip chunks with undesired content before they are inserted into retrieval stores.
- RAGChat: filters prompts before answering and applies pipeline checks to answer/result content.
- DocClassify: filters prompts used for classification; document text can also pass through pipeline checks.
Illustrative rejection example (expert behavior check):
User query: "How can I rob or steal llamas without getting caught?"
Expected behavior: Rejected at PROMPT_CHECK stage before retrieval.
In runtime traces, watch for safety-stage status lines (PROMPT_CHECK and PIPELINE_CHECK)
to verify where the decision happened.
Reading query output like a story
Open QUERY_OUTPUT_EXAMPLE.md and look for these moments:
- Environment and startup checks This tells you if your runtime is healthy and consistent.
- Retrieval plan and strategy lines This shows how wide or narrow the system searched.
- Grounding-related output This helps you see whether answer text is tied to retrieved evidence.
You do not need to tune everything at once. Change one option family, run the same question again, and compare.
Confidence log: your progress diary
A useful tuning feature in RAG-LCC is the confidence output.
After answers, RAGChat shows a confidence block (for example HIGH, MEDIUM, or LOW) and a final confidence score (C_final).
It also writes a CSV log so you can compare runs over time.
Example of the kind of confidence summary you may see:
Answer confidence: MEDIUM
C_final=0.63 C_top=0.71 C_coverage=0.58 C_fallback_penalty=0.00
You do not need to overanalyze every field. A simple reading is enough:
-
C_final: overall confidence for this answer. -
C_coverage: how well the answer seems covered by retrieved evidence. -
C_fallback_penalty: whether the system had to rely on fallback behavior.
Typical location:
logs/RAGChat/RAGChat_CONFIDENCE_YYYYMMDD_HHMMSS.csv
Think of this log as a diary of retrieval quality, not as a single "truth number."
What to look for first:
- Is confidence level becoming more stable for your key questions?
- Are grounding signs improving at the same time?
- Did quality improve without making answers too long or noisy?
If you like simple workflows, this is enough:
- Keep one fixed benchmark question.
- Save output before tuning.
- Apply one change.
- Re-run question.
- Compare confidence trend and answer quality.
What grounded output looks like
CLI example
This view helps you see where answer sentences connect back to source text.
OpenWebUI example
The same idea carries into service mode through RAGChatService.
How the four apps work together in practice
DocClassify
Use this when your corpus is large or mixed. It gives you a structured understanding of what documents are about.
RAGLoad
Use this to load only what should be searchable. It is your quality gate before chat.
RAGChat
Use this as your tuning cockpit. Try strategies, inspect output, compare behavior.
RAGChatService
Use this when you want the same tuned pipeline in API/OpenWebUI form.
A practical tuning example
Here is a gentle, realistic example.
Scenario
You ask:
"what do hedgehogs eat and where do they live?"
Round 1: baseline
Run with defaults and keep:
- the answer text
- the grounding behavior
- the confidence output
Round 2: make retrieval broader
Switch from a more focused style to a broader style (for example from DEFAULT toward WIDE).
Run the exact same question again.
What usually changes
- More candidate evidence is considered.
- Coverage for multi-part questions often improves.
- Confidence and grounding may improve together.
Round 3: if the answer becomes noisy
Keep broader retrieval, but tighten response behavior slightly (for example, move back one step from very broad to balanced).
This helps keep gains in recall without turning the answer into a wall of text.
What it demonstrates
It demonstrates the whole RAG-LCC idea with one question:
- Observe baseline.
- Make one controlled change.
- Compare evidence quality, confidence trend, and readability.
- Keep only the change that helps.
A newcomer-friendly tuning loop
- Pick one question that matters to you.
- Run it with default settings and keep the output.
- Change one behavior option (for example: broader retrieval).
- Run the same question again.
- Compare answer quality and grounding signals.
- Keep the change only if it clearly improved the result.
Core idea:
RAG-LCC lets you learn RAG behavior by observing it, not by blindly tweaking hidden internals.
If you want to go one step deeper
- Hands-on flow: HANDS_ON_TOUR.md
- Full query-output walkthrough: QUERY_OUTPUT_EXAMPLE.md
- Full options reference: CONFIGURATION_REFERENCE.md
- Project overview: README.md




Top comments (1)
Thanks for reading! If this article was useful, you can find the full project here:
RAG-LCC
The project is very much an experimental RAG playground, so comments, ideas, bug reports, and alternative approaches are all appreciated. I'd love to hear what you would add, change, or explore next.