DEV Community

Cover image for How I Built a Local RAG System That Catches Documentation Lies
Angellina Joyce Paul
Angellina Joyce Paul

Posted on

How I Built a Local RAG System That Catches Documentation Lies

I spent two weeks building a tool that solves a problem I kept hitting: documentation that lies.

You join a project. The README says the API runs on port 8080 with JWT auth. You spin it up. Port 9000. Session cookies. You spend 20 minutes wondering if you're on the wrong branch. You're not. The README is just stale.

I built TraceFind to catch this automatically.

It indexes your codebase and documentation, then answers technical questions with exact file and line citations. More importantly, it detects when sources contradict each other — like a README that says one thing and the actual code doing another.

And it runs 100% locally, using a local LLM with zero API costs.

Here's how it works, what broke along the way, and what I learned building it.


What TraceFind Does

Three things:

  1. Indexes a codebase. Point it at any repo. It walks the files, parses code with AST awareness, and builds a searchable index of code and documentation.

  2. Answers developer questions. Ask it natural-language questions like "How is authentication handled?" and it responds with a synthesized answer, citing exact files and line numbers.

  3. Detects contradictions. When the documentation says X and the code does Y, TraceFind reports both sources with a confidence score.

The whole pipeline — parsing, indexing, retrieval, and generation — runs on your machine. No cloud, no API keys, no data leaving your laptop.

GitHub → github.com/AngellinaJoycePaul/TraceFind


The Unfinished Project

I didn't originally build this alone. TraceFind started as a submission for Recursion, a hackathon run by VIT.

We didn't finish it.

The event assigned us a room in the basement of a campus building. The WiFi was weak to the point of being unusable. The venue had no reliable power supply, so our laptops kept dying. Even our own mobile hotspots — which we tried to fall back on — couldn't get a signal from down there.

We stayed for a while, tried to work around it, and eventually left. The project went into a GitHub repo and sat there for weeks, essentially useless.

Then one afternoon I was scrolling through my repositories and saw it sitting there. Half-built. Working in pieces, but never assembled end-to-end. And I thought: why leave it hanging?

I had power at home. I had WiFi. I had time. The hackathon constraints that killed it were gone.

So I picked it back up. Indexed my own codebase. Fixed the cross-encoder scores. Wrote the contradiction detector. Built the VS Code extension. Shipped the whole pipeline.

The version you're reading about in this post has nothing to do with a hackathon. No team, no judges, no deadline. Just a project I didn't want to leave half-finished.

That's the version I'm proud of.


The Architecture

TraceFind has five stages. Each one solves a specific problem that generic RAG systems get wrong on code.

[ Codebase + Docs ]
        │
        ▼
[ 1. AST-Aware Chunking ]      ← Tree-sitter parses code into functions/classes
        │
        ▼
[ 2. Dual Indexing ]           ← BM25 (lexical) + ChromaDB (semantic)
        │
        ▼
[ 3. Hybrid Retrieval ]        ← RRF fusion + Cross-Encoder re-ranking
        │
        ▼
[ 4. Grounded Synthesis ]      ← Local Llama 3.1 8B via Ollama
        │
        ▼
[ 5. Contradiction Detection ] ← Cross-source conflict analysis
Enter fullscreen mode Exit fullscreen mode

1. AST-Aware Chunking (Not Naive Line Splitting)

Most RAG tutorials chunk text by fixed size — "every 500 tokens." That's fine for prose. It's terrible for code.

Chunk a Python file every 500 tokens and you'll slice a function in half. Now the retrieval system hands the LLM a fragment that starts mid-if statement and ends before the return. The model has to guess what the function does.

TraceFind uses Tree-sitter — the same parser GitHub uses for syntax highlighting — to extract exact function, class, and method boundaries. Every chunk is a complete, syntactically valid unit.

It supports Python, TypeScript, JavaScript, Java, and Go out of the box.

2. Dual Indexing (BM25 + Semantic Vectors)

If you only index code semantically, you miss exact identifier matches.

Ask "Where is validateSessionToken defined?" and a pure vector search will return chunks that are about session validation — but maybe not the actual function named validateSessionToken. Embedding models compress meaning; they blur exact names.

If you only index lexically (BM25), you miss semantic queries. "How does authentication work?" won't match a function named check_user_permissions unless the query happens to share tokens.

TraceFind indexes into both:

  • BM25 (rank-bm25) for exact keyword and symbol matching
  • ChromaDB with all-MiniLM-L6-v2 embeddings for semantic similarity

Then it fuses both result sets.

3. The Part That Makes It Work On Code: A Code-Aware Tokenizer

This is the detail that separates TraceFind from most RAG projects.

Standard BM25 tokenizers split on whitespace and punctuation. Feed them getUserProfile and you get one token: getuserprofile. A query for "user profile" returns nothing.

TraceFind's tokenizer splits identifiers intelligently:

"getUserProfile"   →  ["getuserprofile", "get", "user", "profile"]
"user_auth_token"  →  ["user_auth_token", "user", "auth", "token"]
"HTTPServer.start" →  ["httpserver.start", "http", "server", "start"]
Enter fullscreen mode Exit fullscreen mode

It handles camelCase, PascalCase, snake_case, and dot paths — and keeps the original compound identifier too, so exact matches still work when someone searches for the full name.

This one function is why TraceFind's retrieval returns useful results where vanilla BM25 wouldn't.

4. Hybrid Retrieval: Reciprocal Rank Fusion

BM25 gives you a ranked list. Vector search gives you a ranked list. They use completely different scoring systems — one returns unbounded scores, the other returns distances. You can't just add them together.

TraceFind uses Reciprocal Rank Fusion (RRF) — a standard technique from information retrieval literature:

score(doc) = Σ weight / (k + rank(doc))
Enter fullscreen mode Exit fullscreen mode

Where k is a constant (60 in TraceFind's config) and weight is per-source. This fuses rankings without needing score normalization.

Optionally, a Cross-Encoder re-ranker (ms-marco-MiniLM-L-6-v2) further refines the top candidates. It scores each (query, chunk) pair directly — slower than the initial retrieval, but much more accurate for the final top-k.

5. Local Synthesis with Llama 3.1

The top chunks go to a local Llama 3.1 8B model running through Ollama. The system prompt instructs the model to ground every claim in the provided snippets, cite specific files and line numbers, report contradictions when sources disagree, and never invent details.

No OpenAI. No Anthropic. No API keys. Everything on-device.


The Contradiction Detection

This is the feature I'm most proud of, and the one that actually works.

Here's a real example from a repo I indexed. I asked TraceFind:

"How is authentication handled in this codebase?"

The answer came back with two sections. First, the direct answer:

"Authentication is handled using Session cookie authentication, as implemented in backend/main.py:182-200."

Then, an orange contradiction alert:

⚠️  Contradiction Detected
Type:          Configuration & Authentication Mismatch
Source A:      README.md:228-251 (Port 8080 / JWT)
Source B:      server.py:12-18   (Port 9000 / Cookie)
Confidence:    92%
Enter fullscreen mode Exit fullscreen mode

The README claims JWT authentication on port 8080. The code uses session cookies on port 9000. Neither is wrong on its own — but together they contradict.

TraceFind found both, cited both, and told me exactly which lines to check.

For a new developer joining that project, this is the difference between an hour of confusion and a 5-second answer.


The Bugs I Fixed (And What They Taught Me)

Two bugs in this project were subtle enough that I want to share them.

Bug 1: Cross-Encoder Scores That Made No Sense

The cross-encoder re-ranker outputs raw logits — unbounded numbers that can be positive or negative. For a while, my top results had scores like 1.61, -5.27, -6.59, -7.07. Negative scores appearing below positive ones looked fine, but the numbers were meaningless — you can't compare 1.61 to -5.27 on any meaningful scale.

The fix was a sigmoid transform:

prob = 1.0 / (1.0 + math.exp(-raw_score))
Enter fullscreen mode Exit fullscreen mode

Now every score is monotonic in (0, 1), and ranking is preserved.

What I learned: always sanitize model outputs before displaying them. A score that "sort of works" is not the same as a score that means something.

Bug 2: False-Positive Contradictions

My contradiction detector initially fired on any mention of the word "contradiction" in the LLM output. That meant when the model said "No contradictions found between these sources" — a clean result — the detector saw the word and reported a fake contradiction.

The fix had two parts:

  1. Negation guard — skip if the LLM output contains phrases like "no contradiction", "does not conflict", "no discrepancy"
  2. Source requirement — only report a contradiction if the LLM cited at least two distinct sources with [file:line] references

Now the detector is conservative. It misses a few subtle contradictions, but it never reports fake ones.

What I learned: with LLM-based features, false positives are far more damaging than false negatives. A user who sees one fake contradiction will stop trusting the whole system.


What's Next

TraceFind is functional but not finished. Honest roadmap:

  • Hugging Face Space demo — a public Gradio interface so people can try it without installing anything
  • More languages — Tree-sitter supports dozens; I've wired up 5
  • VS Code Marketplace publish — right now the extension only runs locally during development
  • Tests — I have zero automated tests, which is the biggest gap in the project

Try It

TraceFind is MIT-licensed and free.

git clone https://github.com/AngellinaJoycePaul/TraceFind
cd TraceFind
pip install -r requirements.txt
ollama pull llama3.1:8b
uvicorn backend.main:app --reload
Enter fullscreen mode Exit fullscreen mode

Then open http://localhost:8000/docs and try the /ask endpoint on any repo you have lying around.

If you find a contradiction in your own codebase, I'd love to hear about it.

GitHub → github.com/AngellinaJoycePaul/TraceFind


Thanks for reading. I'm a second-year IT student building projects around local AI and developer tooling. If you have feedback or want to talk about RAG, find me on GitHub.

Top comments (0)