<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Angellina Joyce Paul</title>
    <description>The latest articles on DEV Community by Angellina Joyce Paul (@angellinajoycepaul).</description>
    <link>https://dev.to/angellinajoycepaul</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4168852%2Fd9d4d1f5-a4aa-467b-a41d-5855ad9e855d.jpg</url>
      <title>DEV Community: Angellina Joyce Paul</title>
      <link>https://dev.to/angellinajoycepaul</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/angellinajoycepaul"/>
    <language>en</language>
    <item>
      <title>How I Built a Local RAG System That Catches Documentation Lies</title>
      <dc:creator>Angellina Joyce Paul</dc:creator>
      <pubDate>Wed, 07 Oct 2026 12:34:50 +0000</pubDate>
      <link>https://dev.to/angellinajoycepaul/how-i-built-a-local-rag-system-that-catches-documentation-lies-46lo</link>
      <guid>https://dev.to/angellinajoycepaul/how-i-built-a-local-rag-system-that-catches-documentation-lies-46lo</guid>
      <description>&lt;p&gt;I spent two weeks building a tool that solves a problem I kept hitting: documentation that lies.&lt;/p&gt;

&lt;p&gt;You join a project. The README says the API runs on port 8080 with JWT auth. You spin it up. Port 9000. Session cookies. You spend 20 minutes wondering if you're on the wrong branch. You're not. The README is just stale.&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;TraceFind&lt;/strong&gt; to catch this automatically.&lt;/p&gt;

&lt;p&gt;It indexes your codebase and documentation, then answers technical questions with exact file and line citations. More importantly, it detects when sources contradict each other — like a README that says one thing and the actual code doing another.&lt;/p&gt;

&lt;p&gt;And it runs &lt;strong&gt;100% locally&lt;/strong&gt;, using a local LLM with zero API costs.&lt;/p&gt;

&lt;p&gt;Here's how it works, what broke along the way, and what I learned building it.&lt;/p&gt;




&lt;h2&gt;
  
  
  What TraceFind Does
&lt;/h2&gt;

&lt;p&gt;Three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Indexes a codebase.&lt;/strong&gt; Point it at any repo. It walks the files, parses code with AST awareness, and builds a searchable index of code and documentation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Answers developer questions.&lt;/strong&gt; Ask it natural-language questions like "How is authentication handled?" and it responds with a synthesized answer, citing exact files and line numbers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Detects contradictions.&lt;/strong&gt; When the documentation says X and the code does Y, TraceFind reports both sources with a confidence score.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The whole pipeline — parsing, indexing, retrieval, and generation — runs on your machine. No cloud, no API keys, no data leaving your laptop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/AngellinaJoycePaul/TraceFind" rel="noopener noreferrer"&gt;GitHub → github.com/AngellinaJoycePaul/TraceFind&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Unfinished Project
&lt;/h2&gt;

&lt;p&gt;I didn't originally build this alone. TraceFind started as a submission for Recursion, a hackathon run by VIT.&lt;/p&gt;

&lt;p&gt;We didn't finish it.&lt;/p&gt;

&lt;p&gt;The event assigned us a room in the basement of a campus building. The WiFi was weak to the point of being unusable. The venue had no reliable power supply, so our laptops kept dying. Even our own mobile hotspots — which we tried to fall back on — couldn't get a signal from down there.&lt;/p&gt;

&lt;p&gt;We stayed for a while, tried to work around it, and eventually left. The project went into a GitHub repo and sat there for weeks, essentially useless.&lt;/p&gt;

&lt;p&gt;Then one afternoon I was scrolling through my repositories and saw it sitting there. Half-built. Working in pieces, but never assembled end-to-end. And I thought: why leave it hanging?&lt;/p&gt;

&lt;p&gt;I had power at home. I had WiFi. I had time. The hackathon constraints that killed it were gone.&lt;/p&gt;

&lt;p&gt;So I picked it back up. Indexed my own codebase. Fixed the cross-encoder scores. Wrote the contradiction detector. Built the VS Code extension. Shipped the whole pipeline.&lt;/p&gt;

&lt;p&gt;The version you're reading about in this post has nothing to do with a hackathon. No team, no judges, no deadline. Just a project I didn't want to leave half-finished.&lt;/p&gt;

&lt;p&gt;That's the version I'm proud of.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Architecture
&lt;/h2&gt;

&lt;p&gt;TraceFind has five stages. Each one solves a specific problem that generic RAG systems get wrong on code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ Codebase + Docs ]
        │
        ▼
[ 1. AST-Aware Chunking ]      ← Tree-sitter parses code into functions/classes
        │
        ▼
[ 2. Dual Indexing ]           ← BM25 (lexical) + ChromaDB (semantic)
        │
        ▼
[ 3. Hybrid Retrieval ]        ← RRF fusion + Cross-Encoder re-ranking
        │
        ▼
[ 4. Grounded Synthesis ]      ← Local Llama 3.1 8B via Ollama
        │
        ▼
[ 5. Contradiction Detection ] ← Cross-source conflict analysis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. AST-Aware Chunking (Not Naive Line Splitting)
&lt;/h3&gt;

&lt;p&gt;Most RAG tutorials chunk text by fixed size — "every 500 tokens." That's fine for prose. It's terrible for code.&lt;/p&gt;

&lt;p&gt;Chunk a Python file every 500 tokens and you'll slice a function in half. Now the retrieval system hands the LLM a fragment that starts mid-&lt;code&gt;if&lt;/code&gt; statement and ends before the &lt;code&gt;return&lt;/code&gt;. The model has to guess what the function does.&lt;/p&gt;

&lt;p&gt;TraceFind uses &lt;strong&gt;Tree-sitter&lt;/strong&gt; — the same parser GitHub uses for syntax highlighting — to extract exact function, class, and method boundaries. Every chunk is a complete, syntactically valid unit.&lt;/p&gt;

&lt;p&gt;It supports Python, TypeScript, JavaScript, Java, and Go out of the box.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Dual Indexing (BM25 + Semantic Vectors)
&lt;/h3&gt;

&lt;p&gt;If you only index code semantically, you miss exact identifier matches.&lt;/p&gt;

&lt;p&gt;Ask "Where is &lt;code&gt;validateSessionToken&lt;/code&gt; defined?" and a pure vector search will return chunks that are &lt;em&gt;about&lt;/em&gt; session validation — but maybe not the actual function named &lt;code&gt;validateSessionToken&lt;/code&gt;. Embedding models compress meaning; they blur exact names.&lt;/p&gt;

&lt;p&gt;If you only index lexically (BM25), you miss semantic queries. "How does authentication work?" won't match a function named &lt;code&gt;check_user_permissions&lt;/code&gt; unless the query happens to share tokens.&lt;/p&gt;

&lt;p&gt;TraceFind indexes into &lt;strong&gt;both&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;BM25&lt;/strong&gt; (&lt;code&gt;rank-bm25&lt;/code&gt;) for exact keyword and symbol matching&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ChromaDB&lt;/strong&gt; with &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; embeddings for semantic similarity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then it fuses both result sets.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Part That Makes It Work On Code: A Code-Aware Tokenizer
&lt;/h3&gt;

&lt;p&gt;This is the detail that separates TraceFind from most RAG projects.&lt;/p&gt;

&lt;p&gt;Standard BM25 tokenizers split on whitespace and punctuation. Feed them &lt;code&gt;getUserProfile&lt;/code&gt; and you get one token: &lt;code&gt;getuserprofile&lt;/code&gt;. A query for "user profile" returns nothing.&lt;/p&gt;

&lt;p&gt;TraceFind's tokenizer splits identifiers intelligently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="s2"&gt;"getUserProfile"&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="err"&gt;→&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"getuserprofile"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"get"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"profile"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="s2"&gt;"user_auth_token"&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="err"&gt;→&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"user_auth_token"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"auth"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"token"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="s2"&gt;"HTTPServer.start"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;→&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"httpserver.start"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"server"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"start"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It handles camelCase, PascalCase, snake_case, and dot paths — and keeps the original compound identifier too, so exact matches still work when someone searches for the full name.&lt;/p&gt;

&lt;p&gt;This one function is why TraceFind's retrieval returns useful results where vanilla BM25 wouldn't.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Hybrid Retrieval: Reciprocal Rank Fusion
&lt;/h3&gt;

&lt;p&gt;BM25 gives you a ranked list. Vector search gives you a ranked list. They use completely different scoring systems — one returns unbounded scores, the other returns distances. You can't just add them together.&lt;/p&gt;

&lt;p&gt;TraceFind uses &lt;strong&gt;Reciprocal Rank Fusion (RRF)&lt;/strong&gt; — a standard technique from information retrieval literature:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;score(doc) = Σ weight / (k + rank(doc))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Where &lt;code&gt;k&lt;/code&gt; is a constant (60 in TraceFind's config) and &lt;code&gt;weight&lt;/code&gt; is per-source. This fuses rankings without needing score normalization.&lt;/p&gt;

&lt;p&gt;Optionally, a &lt;strong&gt;Cross-Encoder re-ranker&lt;/strong&gt; (&lt;code&gt;ms-marco-MiniLM-L-6-v2&lt;/code&gt;) further refines the top candidates. It scores each (query, chunk) pair directly — slower than the initial retrieval, but much more accurate for the final top-k.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Local Synthesis with Llama 3.1
&lt;/h3&gt;

&lt;p&gt;The top chunks go to a local Llama 3.1 8B model running through &lt;strong&gt;Ollama&lt;/strong&gt;. The system prompt instructs the model to ground every claim in the provided snippets, cite specific files and line numbers, report contradictions when sources disagree, and never invent details.&lt;/p&gt;

&lt;p&gt;No OpenAI. No Anthropic. No API keys. Everything on-device.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Contradiction Detection
&lt;/h2&gt;

&lt;p&gt;This is the feature I'm most proud of, and the one that actually works.&lt;/p&gt;

&lt;p&gt;Here's a real example from a repo I indexed. I asked TraceFind:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How is authentication handled in this codebase?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The answer came back with two sections. First, the direct answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Authentication is handled using Session cookie authentication, as implemented in &lt;code&gt;backend/main.py:182-200&lt;/code&gt;."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then, an orange &lt;strong&gt;contradiction alert&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;⚠️  Contradiction Detected
Type:          Configuration &amp;amp; Authentication Mismatch
Source A:      README.md:228-251 (Port 8080 / JWT)
Source B:      server.py:12-18   (Port 9000 / Cookie)
Confidence:    92%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The README claims JWT authentication on port 8080. The code uses session cookies on port 9000. Neither is wrong on its own — but together they contradict.&lt;/p&gt;

&lt;p&gt;TraceFind found both, cited both, and told me exactly which lines to check.&lt;/p&gt;

&lt;p&gt;For a new developer joining that project, this is the difference between an hour of confusion and a 5-second answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bugs I Fixed (And What They Taught Me)
&lt;/h2&gt;

&lt;p&gt;Two bugs in this project were subtle enough that I want to share them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bug 1: Cross-Encoder Scores That Made No Sense
&lt;/h3&gt;

&lt;p&gt;The cross-encoder re-ranker outputs raw logits — unbounded numbers that can be positive or negative. For a while, my top results had scores like &lt;code&gt;1.61, -5.27, -6.59, -7.07&lt;/code&gt;. Negative scores appearing below positive ones looked fine, but the numbers were meaningless — you can't compare &lt;code&gt;1.61&lt;/code&gt; to &lt;code&gt;-5.27&lt;/code&gt; on any meaningful scale.&lt;/p&gt;

&lt;p&gt;The fix was a sigmoid transform:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prob&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;raw_score&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now every score is monotonic in &lt;code&gt;(0, 1)&lt;/code&gt;, and ranking is preserved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I learned:&lt;/strong&gt; always sanitize model outputs before displaying them. A score that "sort of works" is not the same as a score that means something.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bug 2: False-Positive Contradictions
&lt;/h3&gt;

&lt;p&gt;My contradiction detector initially fired on any mention of the word "contradiction" in the LLM output. That meant when the model said &lt;em&gt;"No contradictions found between these sources"&lt;/em&gt; — a clean result — the detector saw the word and reported a fake contradiction.&lt;/p&gt;

&lt;p&gt;The fix had two parts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Negation guard&lt;/strong&gt; — skip if the LLM output contains phrases like "no contradiction", "does not conflict", "no discrepancy"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source requirement&lt;/strong&gt; — only report a contradiction if the LLM cited at least two distinct sources with &lt;code&gt;[file:line]&lt;/code&gt; references&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Now the detector is conservative. It misses a few subtle contradictions, but it never reports fake ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I learned:&lt;/strong&gt; with LLM-based features, false positives are far more damaging than false negatives. A user who sees one fake contradiction will stop trusting the whole system.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;TraceFind is functional but not finished. Honest roadmap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hugging Face Space demo&lt;/strong&gt; — a public Gradio interface so people can try it without installing anything&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;More languages&lt;/strong&gt; — Tree-sitter supports dozens; I've wired up 5&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;VS Code Marketplace publish&lt;/strong&gt; — right now the extension only runs locally during development&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tests&lt;/strong&gt; — I have zero automated tests, which is the biggest gap in the project&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Try It
&lt;/h2&gt;

&lt;p&gt;TraceFind is MIT-licensed and free.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/AngellinaJoycePaul/TraceFind
&lt;span class="nb"&gt;cd &lt;/span&gt;TraceFind
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
ollama pull llama3.1:8b
uvicorn backend.main:app &lt;span class="nt"&gt;--reload&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then open &lt;code&gt;http://localhost:8000/docs&lt;/code&gt; and try the &lt;code&gt;/ask&lt;/code&gt; endpoint on any repo you have lying around.&lt;/p&gt;

&lt;p&gt;If you find a contradiction in your own codebase, I'd love to hear about it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/AngellinaJoycePaul/TraceFind" rel="noopener noreferrer"&gt;&lt;strong&gt;GitHub → github.com/AngellinaJoycePaul/TraceFind&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Thanks for reading. I'm a second-year IT student building projects around local AI and developer tooling. If you have feedback or want to talk about RAG, find me on &lt;a href="https://github.com/AngellinaJoycePaul" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>llm</category>
      <category>python</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
