<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ruchita Nimkar</title>
    <description>The latest articles on DEV Community by Ruchita Nimkar (@ruchita_nimkar_fb6fcaab1f).</description>
    <link>https://dev.to/ruchita_nimkar_fb6fcaab1f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162014%2Fd467f508-3a1d-4e34-bf22-a482cadafdfb.jpg</url>
      <title>DEV Community: Ruchita Nimkar</title>
      <link>https://dev.to/ruchita_nimkar_fb6fcaab1f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ruchita_nimkar_fb6fcaab1f"/>
    <language>en</language>
    <item>
      <title>Beyond the First Answer: When Retrieval Becomes Investigation</title>
      <dc:creator>Ruchita Nimkar</dc:creator>
      <pubDate>Sun, 04 Oct 2026 17:21:15 +0000</pubDate>
      <link>https://dev.to/ruchita_nimkar_fb6fcaab1f/beyond-the-first-answer-when-retrieval-becomes-investigation-1hn4</link>
      <guid>https://dev.to/ruchita_nimkar_fb6fcaab1f/beyond-the-first-answer-when-retrieval-becomes-investigation-1hn4</guid>
      <description>&lt;h2&gt;
  
  
  What happens when a question looks simple, but answering it correctly requires more than retrieving a few documents?
&lt;/h2&gt;

&lt;p&gt;For my &lt;strong&gt;TigerGraph Hackathon&lt;/strong&gt; project, I explored this question by implementing and comparing &lt;strong&gt;RAG, GraphRAG, and Agentic GraphRAG&lt;/strong&gt; on the same benchmark questions.&lt;/p&gt;

&lt;p&gt;The goal was not simply to prove that &lt;strong&gt;"agents are better."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead, I wanted to understand something more practical:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When does a question actually need an investigation instead of a retrieval?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That question became the central idea behind my project:&lt;/p&gt;

&lt;h1&gt;
  
  
  &lt;strong&gt;Beyond the First Answer&lt;/strong&gt;
&lt;/h1&gt;




&lt;h2&gt;
  
  
  The Problem: Retrieval Is Not Always Enough
&lt;/h2&gt;

&lt;p&gt;A traditional &lt;strong&gt;RAG (Retrieval-Augmented Generation)&lt;/strong&gt; pipeline is very good at answering questions where the required information exists directly inside a few relevant documents.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User Question
      ↓
Similarity Search
      ↓
Relevant Documents
      ↓
     LLM
      ↓
   Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For many questions, this is completely sufficient.&lt;/p&gt;

&lt;p&gt;But consider a question such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;At first glance, this sounds like a normal retrieval question.&lt;/p&gt;

&lt;p&gt;But it isn't just asking the system to &lt;strong&gt;find information&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The system needs to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Find the relevant Olympic event documents.&lt;/li&gt;
&lt;li&gt;Identify the biathlon events.&lt;/li&gt;
&lt;li&gt;Extract the competitor counts.&lt;/li&gt;
&lt;li&gt;Count the qualifying events.&lt;/li&gt;
&lt;li&gt;Produce the final answer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The important distinction is that the answer requires a &lt;strong&gt;sequence of operations&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is where I started thinking about retrieval differently.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Retrieval finds evidence. Investigation decides what to do with that evidence.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  My Agentic GraphRAG Architecture
&lt;/h1&gt;

&lt;p&gt;My implementation uses a routing and execution architecture built around several components.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         USER QUESTION
                              ↓
                       QUESTION ROUTER
                              ↓
                       AGENTIC AGENT
                              ↓
                       ┌───────────────┐
                       │ AUTO → PLANNED│
                       └───────┬───────┘
                               ↓
                            PLANNER
                               ↓
                        PLAN VALIDATOR
                          ↓         ↓
                       VALID     INVALID
                          ↓         ↓
                       EXECUTOR   REACTIVE
                          ↓         ↓
                          └────┬────┘
                               ↓
                         TOOL REGISTRY
                               ↓
              ┌────────────────┼────────────────┐
              ↓                ↓                ↓
       Similarity /       Graph /         Deterministic
       Hybrid Search     Structural       Aggregation
                         Retrieval
              └────────────────┼────────────────┘
                               ↓
                          TOOL RESULTS
                               ↓
                         FINAL ANSWER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The flow starts with the &lt;strong&gt;Question Router&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The router classifies the question into categories such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;lookup&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;aggregation&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;temporal&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;superlative&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;multi-hop&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The question then reaches the &lt;strong&gt;Agentic Agent&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In my current configuration, &lt;strong&gt;Auto resolves to Planned mode&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Planner&lt;/strong&gt; creates a structured plan, and the &lt;strong&gt;Plan Validator&lt;/strong&gt; checks whether that plan is valid.&lt;/p&gt;

&lt;p&gt;If the plan is valid, it is passed to the &lt;strong&gt;Executor&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If the generated plan is invalid, the system can fall back to &lt;strong&gt;Reactive mode&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Important clarification:&lt;/strong&gt; Invalid means the &lt;strong&gt;generated plan is invalid&lt;/strong&gt; — not that the user's question is invalid.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  What Does the Tool Registry Do?
&lt;/h1&gt;

&lt;p&gt;The &lt;strong&gt;Tool Registry&lt;/strong&gt; provides a common dispatch layer between the agentic execution logic and the actual tools.&lt;/p&gt;

&lt;p&gt;This allows the system to access capabilities such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Similarity search&lt;/li&gt;
&lt;li&gt;Hybrid search&lt;/li&gt;
&lt;li&gt;Graph / structural retrieval&lt;/li&gt;
&lt;li&gt;Contextual retrieval&lt;/li&gt;
&lt;li&gt;Document retrieval&lt;/li&gt;
&lt;li&gt;Deterministic aggregation&lt;/li&gt;
&lt;li&gt;Other registered graph tools&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of allowing the agent to directly execute arbitrary operations, tool calls pass through a &lt;strong&gt;common registry&lt;/strong&gt;, where the appropriate tool can be validated and dispatched.&lt;/p&gt;

&lt;p&gt;This becomes particularly important when the system is performing &lt;strong&gt;multi-step investigation&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Failure That Changed My Design
&lt;/h1&gt;

&lt;p&gt;The most interesting part of the project was &lt;strong&gt;not when everything worked&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It was when something &lt;strong&gt;looked correct but wasn't&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I tested the question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The correct answer is:&lt;/p&gt;

&lt;h1&gt;
  
  
  &lt;strong&gt;5&lt;/strong&gt;
&lt;/h1&gt;

&lt;p&gt;Hybrid retrieval was actually doing its job.&lt;/p&gt;

&lt;p&gt;It successfully found the relevant event documents.&lt;/p&gt;

&lt;p&gt;So I initially thought the retrieval pipeline was fine.&lt;/p&gt;

&lt;p&gt;But the final answer was inconsistent.&lt;/p&gt;

&lt;p&gt;In one run, the LLM returned:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;4&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In another run, it identified the correct five qualifying events but concluded:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;6&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That was an important observation.&lt;/p&gt;

&lt;p&gt;The evidence was there.&lt;/p&gt;

&lt;p&gt;The problem was the &lt;strong&gt;computation and interpretation of that evidence&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This changed the way I designed the agentic pipeline.&lt;/p&gt;

&lt;p&gt;Instead of relying entirely on the LLM to perform numerical reasoning over retrieved context, I introduced &lt;strong&gt;deterministic aggregation tools&lt;/strong&gt; into the execution layer.&lt;/p&gt;

&lt;p&gt;The idea was simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Let the retrieval system find the evidence, but use deterministic operations when the task is fundamentally computational.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  Comparing RAG, GraphRAG, and Agentic GraphRAG
&lt;/h1&gt;

&lt;p&gt;I evaluated all three approaches using the &lt;strong&gt;same benchmark setup&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Public-100 Evaluation
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pipeline&lt;/th&gt;
&lt;th&gt;Correct&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Normal RAG&lt;/td&gt;
&lt;td&gt;41 / 100&lt;/td&gt;
&lt;td&gt;41%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GraphRAG&lt;/td&gt;
&lt;td&gt;55 / 100&lt;/td&gt;
&lt;td&gt;55%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agentic GraphRAG&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;72 / 100&lt;/td&gt;
&lt;td&gt;72%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The results were encouraging.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic GraphRAG achieved 72% accuracy&lt;/strong&gt;, compared with &lt;strong&gt;55% for GraphRAG&lt;/strong&gt; and &lt;strong&gt;41% for Normal RAG&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This suggests that adaptive investigation can help when questions require more than a single retrieval operation.&lt;/p&gt;

&lt;p&gt;But there is an important second side to the story.&lt;/p&gt;




&lt;h1&gt;
  
  
  Agents Are Not Free
&lt;/h1&gt;

&lt;p&gt;Agentic reasoning comes with a cost.&lt;/p&gt;

&lt;p&gt;In my &lt;strong&gt;Hidden-50 evaluation&lt;/strong&gt;, the average cost per question was approximately:&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost and Latency Comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pipeline&lt;/th&gt;
&lt;th&gt;Avg. Tokens / Question&lt;/th&gt;
&lt;th&gt;Avg. Time / Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Normal RAG&lt;/td&gt;
&lt;td&gt;13,496.44&lt;/td&gt;
&lt;td&gt;22.96 sec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GraphRAG&lt;/td&gt;
&lt;td&gt;1,218.77&lt;/td&gt;
&lt;td&gt;8.38 sec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agentic GraphRAG&lt;/td&gt;
&lt;td&gt;60,737.54&lt;/td&gt;
&lt;td&gt;100.53 sec&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The difference is significant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic GraphRAG used substantially more tokens and took considerably longer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This creates an important engineering trade-off:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;More Investigation
       ↓
More Tool Calls
       ↓
More Reasoning
       ↓
Higher Accuracy
       +
Higher Cost / Latency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Therefore, simply replacing every RAG pipeline with an agentic system would not necessarily be a good engineering decision.&lt;/p&gt;




&lt;h1&gt;
  
  
  What I Learned
&lt;/h1&gt;

&lt;p&gt;So I don't think the conclusion should be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Agentic GraphRAG is always better."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That would miss one of the most important lessons from the experiment.&lt;/p&gt;

&lt;p&gt;The more useful conclusion is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;h2&gt;
  
  
  &lt;strong&gt;Use the simplest retrieval strategy that can reliably answer the question. Escalate to investigation when the question actually requires it.&lt;/strong&gt;
&lt;/h2&gt;
&lt;/blockquote&gt;

&lt;p&gt;This leads to a more practical architecture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    USER QUESTION
                          ↓
                  Can retrieval
                  answer it reliably?
                     ↙       ↘
                   YES        NO
                    ↓          ↓
                RETRIEVE   INVESTIGATE
                    ↓          ↓
                 ANSWER    PLAN → TOOLS
                                  ↓
                             VALIDATE
                                  ↓
                              ANSWER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is therefore not to make every question &lt;strong&gt;agentic&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The goal is to make the system &lt;strong&gt;adaptive&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Core Idea
&lt;/h1&gt;

&lt;p&gt;My project explores a simple principle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Not every question needs an agent. But some questions need more than retrieval.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A lookup question may only need document retrieval.&lt;/p&gt;

&lt;p&gt;A structural question may benefit from a graph.&lt;/p&gt;

&lt;p&gt;A multi-hop or aggregation question may require planning, multiple tool calls, and deterministic computation.&lt;/p&gt;

&lt;p&gt;The challenge is deciding &lt;strong&gt;when to escalate&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is the problem I explored with &lt;strong&gt;Beyond the First Answer&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Takeaway
&lt;/h1&gt;

&lt;p&gt;The biggest lesson from this project was not that one architecture wins every benchmark.&lt;/p&gt;

&lt;p&gt;It was that &lt;strong&gt;retrieval and reasoning solve different parts of the problem&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RAG&lt;/strong&gt; is efficient when the answer is directly present in relevant documents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GraphRAG&lt;/strong&gt; adds structural relationships that can improve retrieval and context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic GraphRAG&lt;/strong&gt; becomes valuable when answering the question requires an actual investigation across multiple steps.&lt;/p&gt;

&lt;p&gt;But that additional intelligence comes with a cost in &lt;strong&gt;tokens and latency&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So my design philosophy became:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Retrieve first. Investigate when necessary. Compute deterministically when possible.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is what I mean by:&lt;/p&gt;

&lt;h1&gt;
  
  
  &lt;strong&gt;Beyond the First Answer&lt;/strong&gt;
&lt;/h1&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
    </item>
  </channel>
</rss>
