<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harshdip Saha</title>
    <description>The latest articles on DEV Community by Harshdip Saha (@harshdipsaha).</description>
    <link>https://dev.to/harshdipsaha</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156738%2Fc9621109-065c-4f5e-b4e0-8a50772951b8.png</url>
      <title>DEV Community: Harshdip Saha</title>
      <link>https://dev.to/harshdipsaha</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harshdipsaha"/>
    <language>en</language>
    <item>
      <title>My RAG pipeline answered 1. The real answer was 20.</title>
      <dc:creator>Harshdip Saha</dc:creator>
      <pubDate>Fri, 02 Oct 2026 08:50:48 +0000</pubDate>
      <link>https://dev.to/harshdipsaha/my-rag-pipeline-answered-1-the-real-answer-was-20-2bbo</link>
      <guid>https://dev.to/harshdipsaha/my-rag-pipeline-answered-1-the-real-answer-was-20-2bbo</guid>
      <description>&lt;p&gt;I asked three pipelines one question from the TigerGraph Agentic GraphRAG Hackathon benchmark: &lt;em&gt;how many athletics events at the 2004 Summer Olympics had more than 41 competitors?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Plain RAG retrieved 4 chunks and answered 1. GraphRAG expanded one hop, saw 12 chunks, and answered 2. The real answer is &lt;strong&gt;20&lt;/strong&gt;, out of &lt;strong&gt;43&lt;/strong&gt; matching events in the graph.&lt;/p&gt;

&lt;p&gt;Both pipelines sounded equally confident. Neither one had any way to know that plain RAG was looking at 4 of the 43 matching events (9.3%). Picture an analyst pasting that "2" into a report because it sounded sure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 9% problem
&lt;/h2&gt;

&lt;p&gt;Similarity search returns what looks relevant, but a counting question needs everything that is, and no value of &lt;code&gt;k&lt;/code&gt; fixes that, because the right &lt;code&gt;k&lt;/code&gt; depends on the answer you haven't computed yet.&lt;/p&gt;

&lt;p&gt;On the 100 public benchmark questions, RAG and GraphRAG scored &lt;strong&gt;0 out of 31&lt;/strong&gt; on aggregation and superlative questions.&lt;/p&gt;

&lt;p&gt;Also, as a user I want to know whether the answer I got is complete, and a confident tone tells me nothing about that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Investigation Certificates
&lt;/h2&gt;

&lt;p&gt;TigerGraph already knows how many events match &lt;code&gt;sport = athletics AND games = 2004 Summer&lt;/code&gt;. It is a &lt;code&gt;COUNT&lt;/code&gt;. I call that number the structural bound (the size of the complete answer set, straight from the graph).&lt;/p&gt;

&lt;p&gt;So every answer my pipeline emits ships with an &lt;strong&gt;Investigation Certificate&lt;/strong&gt;: a JSON record that puts the &lt;strong&gt;structural bound&lt;/strong&gt; next to the evidence the agent actually inspected.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"qid"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pub-045"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"completeness_class"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"exhaustive"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"structural_bound"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;43&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence_set_size"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;43&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"completeness_check"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pass"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"stop_reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"structural_bound_met"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"latency_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;227&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;43 equals 43 (0 tokens here, since this question never calls the LLM), so the evidence set matches the graph's own count. Also, you can check that yourself without trusting the model. If the numbers don't match, the certificate says so instead of hiding it. Run the same question through plain RAG and its certificate would read evidence 4 against bound 43, a visible fail instead of a confident "1".&lt;/p&gt;

&lt;h2&gt;
  
  
  How the routing works
&lt;/h2&gt;

&lt;p&gt;A router sorts each question into one of 3 completeness classes and picks the cheapest tool that can satisfy it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;existential&lt;/strong&gt; (lookup, multi-hop): exact match or venue+date traversal. The certificate proves the entity resolved uniquely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;chained&lt;/strong&gt; (temporal): a &lt;code&gt;PREV&lt;/code&gt;/&lt;code&gt;NEXT&lt;/code&gt; hop chain. The certificate proves the hops were actually walked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;exhaustive&lt;/strong&gt; (aggregation, superlative): a GSQL structural scan returning the &lt;code&gt;COUNT&lt;/code&gt; plus every match. The certificate proves evidence set equals the graph's own count.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The LLM is called only to break a tie or recover a missing field. It never counts.&lt;/p&gt;

&lt;p&gt;Besides the 5 regex templates, I added TypeSafe AI's Jev System One as a fallback. It is a non-autoregressive decision model, so it classifies intent and picks between tied candidates at &lt;strong&gt;0 output tokens&lt;/strong&gt;. On 5 un-templated questions the regex router sent all 5 to unverified RAG, and Jev routed all 5 correctly (&lt;strong&gt;5/5&lt;/strong&gt;). That small test (5 queries) is all the evidence I have for the classifier, so I'd call it promising and unproven.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;Same 100 public questions, same live TigerGraph graph, same live Groq model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;RAG&lt;/th&gt;
&lt;th&gt;GraphRAG&lt;/th&gt;
&lt;th&gt;Agentic&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy&lt;/td&gt;
&lt;td&gt;25%&lt;/td&gt;
&lt;td&gt;43%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;99%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per answer&lt;/td&gt;
&lt;td&gt;1,389&lt;/td&gt;
&lt;td&gt;2,203&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;18&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency per answer&lt;/td&gt;
&lt;td&gt;9.3s&lt;/td&gt;
&lt;td&gt;15.5s&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.41s&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Thus the agentic pipeline uses 122x fewer tokens (2,203 vs 18) and runs 37.8x faster than GraphRAG (15.5s vs 0.41s), at 2.3x the accuracy. The cost was always in stuffing chunks into a context window, and the graph queries are nearly free.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one miss
&lt;/h2&gt;

&lt;p&gt;I got 99 out of 100. The miss is &lt;code&gt;pub-060&lt;/code&gt;, a date+venue tie among 37 candidate events at ExCeL. Jev picked one with 0.90 confidence and picked the wrong one. The LLM guessed wrong on this question in my earlier run too, so Jev did not fix it.&lt;/p&gt;

&lt;p&gt;The certificate for it reports &lt;code&gt;pass_with_llm_recovery&lt;/code&gt;, not a clean &lt;code&gt;pass&lt;/code&gt;. That is deliberate. The certificate measures whether the evidence was complete, the way an auditor signs off on the books and leaves the strategy to someone else. &lt;code&gt;pub-060&lt;/code&gt; is a &lt;strong&gt;complete but wrong&lt;/strong&gt; case, where evidence and answer come apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Code: &lt;a href="https://github.com/HarshdipSaha/tigergraph-agentic-graphrag" rel="noopener noreferrer"&gt;https://github.com/HarshdipSaha/tigergraph-agentic-graphrag&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Live dashboard (side-by-side comparison): &lt;a href="https://investigation-certificates.streamlit.app/" rel="noopener noreferrer"&gt;https://investigation-certificates.streamlit.app/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Demo video: &lt;a href="https://www.loom.com/share/6c4edb59ed714f1694d1bc65cca2d934" rel="noopener noreferrer"&gt;https://www.loom.com/share/6c4edb59ed714f1694d1bc65cca2d934&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Built for the TigerGraph Agentic GraphRAG Hackathon with TigerGraph Savanna, TypeSafe AI Jev, Groq and Streamlit.&lt;/p&gt;

&lt;p&gt;Overall, a certificate vouches for the evidence, and answer correctness stays a separate claim, so a system should say which one it is making. If you have a RAG pipeline that answers counting questions, I'm curious what your structural bound would be. :)&lt;/p&gt;

</description>
      <category>graphdatabase</category>
      <category>ai</category>
      <category>rag</category>
      <category>python</category>
    </item>
  </channel>
</rss>
