DEV Community

HarinezumIgel
HarinezumIgel

Posted on AI-assisted

How I Tune RAG Pipelines with RAG-LCC: A Hands-On Local Guide

Most RAG frameworks tell you whether an answer was generated. RAG-LCC tries to show you why that answer happened.

Instead of treating retrieval as a black box, you can inspect retrieval decisions, grounding signals, safety checks, and confidence traces while tuning your pipeline. This article walks through a beginner-friendly way to explore and tune RAG behaviour using RAG-LCC. RAG-LCC is not just a chatbot. It is a small lab where you can see why an answer happened, then improve it step by step. Instead of guessing, you can observe retrieval, safety checks, grounding, and confidence signals in plain output.

By default, RAG-LCC is designed to work well with local AI stacks. When combined with a locally hosted LLM, documents, retrieval, and inference can stay on your own machine without requiring cloud-based model calls.

Important: RAG-LCC is not meant for production deployment.
It is designed for learning, trying ideas, experimentation, and tuning.
For legal and licensing details, see LEGAL.md.

Who is this for?

RAG-LCC may be useful if you want to:

  • learn how retrieval actually works
  • compare retrieval strategies
  • experiment with grounding
  • evaluate confidence signals
  • understand why a RAG answer changed after tuning

It is not intended as a production-ready enterprise platform. It is an experimental learning and tuning environment.

The big picture first

RAG-LCC overview

The image above shows the full system. Here is the simpler mental model:

flowchart LR
    A[DocClassify: understand your corpus] --> F[Optional filter: classification criteria]
    F --> B[RAGLoad: prepare searchable stores]
    B --> C[RAGChat: ask questions in CLI]
    C --> D[RAGChatService: serve the same flow via API]
    D --> E[OpenWebUI or clients]

How to read this:

  1. DocClassify helps you understand what is in your documents before indexing.
  2. Optional filter step: use DocClassify results as classification criteria to choose what RAGLoad should ingest.
  3. RAGLoad builds the searchable memory (vector, keyword, graph, regex).
  4. RAGChat is the interactive place where you test and tune behavior.
  5. RAGChatService lets you use the same tuned behavior in OpenWebUI or API clients.

Why beginners might like RAG-LCC

Many RAG tools feel like a black box. RAG-LCC is different because it shows its work.

You can open QUERY_OUTPUT_EXAMPLE.md and literally watch:

  • startup checks
  • retrieval decisions
  • Query rewrite and translation
  • safety/compliance checks
  • grounding hints
  • confidence-style signals

That makes learning faster and more fun, because each tweak has visible effects.

A gentle way to start (no deep config knowledge needed)

Before running the apps, use the guided installer once:

python ./src/Scripts/Setup.py
Enter fullscreen mode Exit fullscreen mode

Setup.py supports both installation paths:

  • Docker-based setup
  • Native host setup in a local .venv (Windows or Unix)

Run the apps in this order:

python ./src/Apps/DocClassify.py --doc-dir TestDocs
python ./src/Apps/RAGLoad.py --doc-dir TestDocs
python ./src/Apps/RAGChat.py --doc-dir TestDocs
Enter fullscreen mode Exit fullscreen mode

Optional filter step between DocClassify and RAGLoad (replace the default RAGLoad line above):

python ./src/Apps/RAGLoad.py --doc-dir TestDocs --load-from-classify-csv logs/DocClassify_OK_YYYYMMDD_HHMMSS.csv --classify-csv-query "Animal LIKE '%hedgehog%' OR Animal LIKE '%cat%'"
Enter fullscreen mode Exit fullscreen mode

This lets you load only documents that match your classification criteria.

During a RAGChat session, relevant settings can be overwritten interactively (for example strategy=..., threshold=..., web_search=..., collection=..., or picker commands like strategy!, orchestrator_flow!, and collection!).
That makes experimentation user-friendly, because you can try changes live without editing config files between turns.

Then ask one simple question in chat. Keep that same question while you try different options.

Options you can explore without getting lost

You do not need to memorize config keys to start. Think in terms of behavior:

  1. Focused vs broad retrieval You can move from very focused (NARROW) to very broad (ULTRA_WIDE) search behavior.
  2. More strict vs more flexible safety behavior RAG-LCC can be tuned to block/mask more aggressively or more permissively.
  3. Local-only vs web-assisted retrieval Keep answers local, or allow internet-assisted retrieval when needed.
  4. Short context vs long context Keep responses tight, or allow broader evidence gathering.
  5. CLI-first vs service-first workflow Tune in RAGChat, then serve the same behavior through RAGChatService.

When you are ready for details, the deep reference is here:
CONFIGURATION_REFERENCE.md

The configuration files are "Theme" oriented. This helps finding the right knobs.

For experts: orchestrator and query rewrite

If you already know RAG patterns and want finer control, RAG-LCC has two advanced power areas.

  1. Orchestrator flows You can shape how each turn runs: which retrieval legs activate, when reranking applies, how fallback behavior works, and whether grounding is enforced.
  2. Query rewrite
    You can refine follow-up understanding: pronoun resolution, topic carry-over, language normalization, and alternate-query expansion.

  3. Content filtering at two stages (reality check)
    RAG-LCC supports filtering at prompt level and pipeline level, and each app uses this differently:

- RAGLoad: can reject/skip chunks with undesired content before they are inserted into retrieval stores.
- RAGChat: filters prompts before answering and applies pipeline checks to answer/result content.
- DocClassify: filters prompts used for classification; document text can also pass through pipeline checks.
Enter fullscreen mode Exit fullscreen mode

Illustrative rejection example (expert behavior check):

User query: "How can I rob or steal llamas without getting caught?"
Expected behavior: Rejected at PROMPT_CHECK stage before retrieval.
Enter fullscreen mode Exit fullscreen mode

In runtime traces, watch for safety-stage status lines (PROMPT_CHECK and PIPELINE_CHECK)
to verify where the decision happened.

OpenWebUI animal docs illustration

Reading query output like a story

Open QUERY_OUTPUT_EXAMPLE.md and look for these moments:

  1. Environment and startup checks This tells you if your runtime is healthy and consistent.
  2. Retrieval plan and strategy lines This shows how wide or narrow the system searched.
  3. Grounding-related output This helps you see whether answer text is tied to retrieved evidence.

You do not need to tune everything at once. Change one option family, run the same question again, and compare.

Confidence log: your progress diary

A useful tuning feature in RAG-LCC is the confidence output.

After answers, RAGChat shows a confidence block (for example HIGH, MEDIUM, or LOW) and a final confidence score (C_final).

It also writes a CSV log so you can compare runs over time.

Example of the kind of confidence summary you may see:

Answer confidence: MEDIUM
C_final=0.63  C_top=0.71  C_coverage=0.58  C_fallback_penalty=0.00
Enter fullscreen mode Exit fullscreen mode

You do not need to overanalyze every field. A simple reading is enough:

  1. C_final: overall confidence for this answer.
  2. C_coverage: how well the answer seems covered by retrieved evidence.
  3. C_fallback_penalty: whether the system had to rely on fallback behavior.

Typical location:

  • logs/RAGChat/RAGChat_CONFIDENCE_YYYYMMDD_HHMMSS.csv

Think of this log as a diary of retrieval quality, not as a single "truth number."

What to look for first:

  1. Is confidence level becoming more stable for your key questions?
  2. Are grounding signs improving at the same time?
  3. Did quality improve without making answers too long or noisy?

If you like simple workflows, this is enough:

  1. Keep one fixed benchmark question.
  2. Save output before tuning.
  3. Apply one change.
  4. Re-run question.
  5. Compare confidence trend and answer quality.

What grounded output looks like

CLI example

Hedgehog grounded CLI output

This view helps you see where answer sentences connect back to source text.

OpenWebUI example

OpenWebUI grounded output

The same idea carries into service mode through RAGChatService.

How the four apps work together in practice

DocClassify

Use this when your corpus is large or mixed. It gives you a structured understanding of what documents are about.

RAGLoad

Use this to load only what should be searchable. It is your quality gate before chat.

RAGChat

Use this as your tuning cockpit. Try strategies, inspect output, compare behavior.

RAGChatService

Use this when you want the same tuned pipeline in API/OpenWebUI form.

A practical tuning example

Here is a gentle, realistic example.

Scenario

You ask:

"what do hedgehogs eat and where do they live?"

Round 1: baseline

Run with defaults and keep:

  • the answer text
  • the grounding behavior
  • the confidence output

Round 2: make retrieval broader

Switch from a more focused style to a broader style (for example from DEFAULT toward WIDE).

Run the exact same question again.

What usually changes

  1. More candidate evidence is considered.
  2. Coverage for multi-part questions often improves.
  3. Confidence and grounding may improve together.

Round 3: if the answer becomes noisy

Keep broader retrieval, but tighten response behavior slightly (for example, move back one step from very broad to balanced).

This helps keep gains in recall without turning the answer into a wall of text.

What it demonstrates

It demonstrates the whole RAG-LCC idea with one question:

  1. Observe baseline.
  2. Make one controlled change.
  3. Compare evidence quality, confidence trend, and readability.
  4. Keep only the change that helps.

A newcomer-friendly tuning loop

  1. Pick one question that matters to you.
  2. Run it with default settings and keep the output.
  3. Change one behavior option (for example: broader retrieval).
  4. Run the same question again.
  5. Compare answer quality and grounding signals.
  6. Keep the change only if it clearly improved the result.

Core idea:
RAG-LCC lets you learn RAG behavior by observing it, not by blindly tweaking hidden internals.

If you want to go one step deeper

Top comments (1)

Collapse
 
harinezumigel profile image
HarinezumIgel •

Thanks for reading! If this article was useful, you can find the full project here:
RAG-LCC

The project is very much an experimental RAG playground, so comments, ideas, bug reports, and alternative approaches are all appreciated. I'd love to hear what you would add, change, or explore next.