<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Anant Kumar</title>
    <description>The latest articles on DEV Community by Anant Kumar (@anant_kumar_bf65d0d3994d3).</description>
    <link>https://dev.to/anant_kumar_bf65d0d3994d3</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152237%2Fd7f6f603-93a0-40f4-b8fa-cf651b71a4a6.png</url>
      <title>DEV Community: Anant Kumar</title>
      <link>https://dev.to/anant_kumar_bf65d0d3994d3</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/anant_kumar_bf65d0d3994d3"/>
    <language>en</language>
    <item>
      <title># When does an AI agent actually earn its cost? I measured it on 100 questions</title>
      <dc:creator>Anant Kumar</dc:creator>
      <pubDate>Wed, 30 Sep 2026 12:07:58 +0000</pubDate>
      <link>https://dev.to/anant_kumar_bf65d0d3994d3/-when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions-4j5l</link>
      <guid>https://dev.to/anant_kumar_bf65d0d3994d3/-when-does-an-ai-agent-actually-earn-its-cost-i-measured-it-on-100-questions-4j5l</guid>
      <description>&lt;p&gt;&lt;em&gt;Built for the TigerGraph Agentic GraphRAG Hackathon. Repo and dashboard linked at the end.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;RAG retrieves text. GraphRAG adds structure. Agentic GraphRAG lets a model plan&lt;br&gt;
its own investigation. Everyone assumes the third one is better — but by how&lt;br&gt;
much, and on which questions? That's the question this hackathon asked, so I&lt;br&gt;
built six pipelines over one corpus and measured them against the same 100&lt;br&gt;
questions, with the same generation model throughout.&lt;/p&gt;

&lt;p&gt;The short version: &lt;strong&gt;exact match went from 67% to 99%, a 32-point jump.&lt;/strong&gt; But&lt;br&gt;
the interesting part isn't that number. It's that an agent &lt;em&gt;by itself&lt;/em&gt; only got&lt;br&gt;
me 3 of those 32 points.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;2,951 Wikipedia articles in TigerGraph Savanna. 100 evaluation questions across&lt;br&gt;
five types: lookup, temporal, multi-hop, superlative and aggregation. Every&lt;br&gt;
pipeline uses &lt;code&gt;gemini-3.1-flash-lite&lt;/code&gt; — a cheap model, deliberately — with local&lt;br&gt;
BGE embeddings and TigerGraph's native vector index. No pipeline gets a better&lt;br&gt;
model than any other.&lt;/p&gt;

&lt;p&gt;Scoring is exact match against the gold answer, computed with no model in the&lt;br&gt;
loop. I also ran an LLM judge, and I'll come back to why I stopped trusting it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure that taught me the most
&lt;/h2&gt;

&lt;p&gt;Plain RAG scored 67%. GraphRAG — entity linking plus one-hop traversal — also&lt;br&gt;
scored 67%. Identical. That was my first surprise.&lt;/p&gt;

&lt;p&gt;Breaking it down by question type showed why. On aggregation questions —&lt;br&gt;
&lt;em&gt;"how many cycling events had more than 30 competitors?"&lt;/em&gt; — RAG got &lt;strong&gt;1 out of&lt;br&gt;
21&lt;/strong&gt;. GraphRAG got &lt;strong&gt;0&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The reason is structural, and no amount of prompt engineering touches it.&lt;br&gt;
Answering that question requires &lt;em&gt;every&lt;/em&gt; matching document; for one question,&lt;br&gt;
43 of them. Retrieval fetches the top five and counts those. The model then&lt;br&gt;
confidently reports a number that is simply the size of what it was shown.&lt;/p&gt;

&lt;p&gt;My LLM-extracted entity graph didn't help either. It had 12,000 entities with&lt;br&gt;
free-text relationship labels — but "competitors: 43" was never a property you&lt;br&gt;
could filter or count on. I had modelled the prose and not the facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually fixed it
&lt;/h2&gt;

&lt;p&gt;Every one of those articles opens with an infobox: event, games, venue, date,&lt;br&gt;
competitors, nations, gold medallist. So I built a second layer in the graph&lt;br&gt;
directly from those boxes — &lt;code&gt;OlympicEvent&lt;/code&gt; linked to its &lt;code&gt;Games&lt;/code&gt;, &lt;code&gt;Sport&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;Venue&lt;/code&gt;, plus a &lt;code&gt;PREV_GAMES&lt;/code&gt; edge so "the Olympics before 2016" is one hop.&lt;/p&gt;

&lt;p&gt;That ingestion makes &lt;strong&gt;zero LLM calls&lt;/strong&gt;. It's a parser. And it turns counting&lt;br&gt;
from a retrieval problem into a graph traversal: aggregation went from 0/21 to&lt;br&gt;
&lt;strong&gt;21/21&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The agent, and the ablation that matters
&lt;/h2&gt;

&lt;p&gt;My agent plans each step and chooses between five tools: a structured graph&lt;br&gt;
query, time-scoped fact lookup, entity linking, one-hop traversal, and vector&lt;br&gt;
search over chunks. It states what the evidence does and doesn't establish &lt;em&gt;in&lt;br&gt;
the same structured call&lt;/em&gt; as the routing decision, so self-evaluation costs no&lt;br&gt;
extra call. It answers when satisfied, or changes strategy.&lt;/p&gt;

&lt;p&gt;Here's the comparison that I think is the real result:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pipeline&lt;/th&gt;
&lt;th&gt;Exact match&lt;/th&gt;
&lt;th&gt;vs RAG&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RAG&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GraphRAG&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;td&gt;+0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent over text/entity tools only&lt;/td&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+3&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same agent + structured graph tools&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+32&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The agentic architecture alone bought 3 points. The agent &lt;em&gt;with the right&lt;br&gt;
retrieval surface&lt;/em&gt; bought 32. If I had only built the agent, I'd have concluded&lt;br&gt;
agentic RAG barely helps. If I had only built the graph layer, I'd have missed&lt;br&gt;
that the agent is what handles everything the layer doesn't model — ask it who&lt;br&gt;
directed a film and it switches to vector search on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result I didn't expect
&lt;/h2&gt;

&lt;p&gt;If the agent picks the structured query on 100 out of 100 questions, does it&lt;br&gt;
need to &lt;em&gt;reason&lt;/em&gt; at all? I tested it: everything the planner decides already&lt;br&gt;
exists as a row in the graph — 47 sports, 20 Games, 316 venues, 475 event names.&lt;br&gt;
So planning isn't generation, it's &lt;strong&gt;selection&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I replaced the generative planner with two typed selection calls (TypeSafe's&lt;br&gt;
System One), kept the same GSQL and the same templated answer, and got:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Same 99%. 1.7 seconds per question instead of 13. Zero generation-model calls.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That last number matters more than the speed. My free tier is 500 calls a day —&lt;br&gt;
one benchmark run — and every measurement I took that week died on that wall,&lt;br&gt;
three times in one day. A path with no generation calls can't hit a rate limit,&lt;br&gt;
which means a live demo can't fail halfway through.&lt;/p&gt;

&lt;p&gt;So the honest answer to the hackathon's question: &lt;strong&gt;agents earn their cost on&lt;br&gt;
open-ended questions, and they're overkill where a typed query settles it.&lt;/strong&gt; The&lt;br&gt;
agent still ships, because you don't know in advance which kind of question is&lt;br&gt;
coming.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things that went wrong
&lt;/h2&gt;

&lt;p&gt;A submission that only lists wins isn't worth much, so:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My LLM judge was unreliable.&lt;/strong&gt; It gave 4 or 5 out of 5 to &lt;strong&gt;14 answers that&lt;br&gt;
are wrong&lt;/strong&gt; — mostly fluent refusals like "the corpus does not contain this."&lt;br&gt;
That's why exact match is my headline metric. I built a verification pass that&lt;br&gt;
asks a narrower question — does this evidence support this claim? — and it&lt;br&gt;
flagged all 14.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A "fix" of mine cost 17 points.&lt;/strong&gt; I changed which answer field wins, and exact&lt;br&gt;
match dropped from 99 to 82. The agent couldn't see what its tools had computed,&lt;br&gt;
so it recounted a truncated evidence list and answered &lt;code&gt;12&lt;/code&gt; — the size of the&lt;br&gt;
cap — to seven counting questions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One fix made things worse before better.&lt;/strong&gt; My first repair made a question&lt;br&gt;
fail three times out of three instead of one. A regex only matched year and&lt;br&gt;
season when adjacent, so "the Summer Olympics held immediately before 2016"&lt;br&gt;
produced no season at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A stale artifact nearly shipped.&lt;/strong&gt; My hidden-set answers were generated two&lt;br&gt;
minutes before a bug fix. Five counts were wrong. I only caught it by answering&lt;br&gt;
the same questions with a second pipeline and scoring both against a&lt;br&gt;
deterministic oracle.&lt;/p&gt;

&lt;p&gt;Every one of those is now pinned by a regression test, because each was&lt;br&gt;
invisible in normal use — the system kept answering, it just answered wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;RAG&lt;/th&gt;
&lt;th&gt;GraphRAG&lt;/th&gt;
&lt;th&gt;Agentic (ablation)&lt;/th&gt;
&lt;th&gt;Agentic (full)&lt;/th&gt;
&lt;th&gt;Selection planner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Exact match&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;td&gt;70%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;99%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregation&lt;/td&gt;
&lt;td&gt;1/21&lt;/td&gt;
&lt;td&gt;0/21&lt;/td&gt;
&lt;td&gt;3/21&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;21/21&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;21/21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens/question&lt;/td&gt;
&lt;td&gt;3,586&lt;/td&gt;
&lt;td&gt;3,952&lt;/td&gt;
&lt;td&gt;6,065&lt;/td&gt;
&lt;td&gt;3,412&lt;/td&gt;
&lt;td&gt;2,267&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Seconds&lt;/td&gt;
&lt;td&gt;6.8&lt;/td&gt;
&lt;td&gt;10.2&lt;/td&gt;
&lt;td&gt;20.2&lt;/td&gt;
&lt;td&gt;13.2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1.7&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The agent reaches 99% on &lt;strong&gt;fewer tokens than plain RAG&lt;/strong&gt;. A cheap model with the&lt;br&gt;
right architecture beats an expensive one with a lazy pipeline.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/antcybersec/graph_rag" rel="noopener noreferrer"&gt;https://github.com/antcybersec/graph_rag&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Live dashboard:&lt;/strong&gt; &lt;a href="https://graphrag-c.streamlit.app/" rel="noopener noreferrer"&gt;https://graphrag-c.streamlit.app/&lt;/a&gt;&lt;br&gt;
Built on TigerGraph Savanna with GSQL and its native vector index.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
    </item>
  </channel>
</rss>
