DEV Community

Utkarsh varshney
Utkarsh varshney

Posted on

When do AI agents actually matter? Benchmarking RAG vs GraphRAG vs Agentic GraphRAG on TigerGraph

Everyone's building agents right now. But a better question than "can an agent do this?" is "when does an agent actually beat something simpler?" For the TigerGraph Agentic GraphRAG Hackathon, I built three question-answering pipelines side by side to find out.

The setup

The dataset was ~2,900 Wikipedia articles about Olympic events, plus 100 evaluation questions with answers and 50 hidden ones. The questions came in five types: simple lookups, multi-hop questions ("who won gold at this venue on this date?"), temporal ones ("…at the Olympics held immediately before 2016"), aggregations ("how many biathlon events had more than 73 competitors?"), and superlatives ("which event had the most competitors?").

One early discovery shaped everything: about 760 of the documents were distractors — articles about films and companies mixed into the corpus. A plain text-retrieval system can get pulled toward them. A graph built only from Olympic event records never sees them.

The core idea: the LLM plans, the graph computes

I parsed each event's Wikipedia infobox into a structured Event vertex covering sport, year, season, venue, date, competitor count, nations and medallists, and loaded 2,187 of them into TigerGraph Savanna. Then I wrote GSQL query endpoints for filtering, counting and lookups.

The key design decision was that the model never computes answers. A question becomes a structured plan, and TigerGraph executes it. Counting 11 biathlon events and filtering by competitor count is a database job. An LLM guessing at it from a few retrieved paragraphs is how you get wrong answers.

Three pipelines

  • RAG: retrieve the top 5 documents by similarity and read the answer out of them.
  • GraphRAG: plan, then run one graph query and return the result.
  • Agentic GraphRAG: an orchestrator loop. It plans, queries the graph, judges whether the evidence is sufficient, self-corrects if it's thin (relaxing filters, re-matching events), verifies against a second source, and stops when it's confident.

Results (100 questions)

Pipeline Accuracy Est. tokens/query
RAG 18% 1,573
GraphRAG 92% 247
Agentic GraphRAG 100% 1,295

RAG scored 0% on aggregation and superlative questions. You simply can't count 40 events from 5 retrieved documents. It also often picked the wrong Olympic year for "before 2016" questions, because "2012" and "1992" look equally similar to a text retriever.

GraphRAG took the big leap to 92% at roughly a fifth of RAG's token cost.

Where the agent earned its keep

The final 8 points came from multi-hop questions. Take "who won gold at Beijing National Stadium on 16 August 2008". A venue hosts dozens of events, so GraphRAG grabbed the first match. The agent noticed several candidates, treated that as a gap, and disambiguated using the exact date in each event's record, with a similarity-search tiebreak for true ties. That one investigative behaviour took multi-hop from 71% to 100%.

That's the real lesson. Agents aren't better everywhere. For most of these questions, a well-designed graph query was enough and was far cheaper. The agentic loop paid off exactly where the evidence was ambiguous and needed checking.

Lessons from deploying on TigerGraph Savanna

  • On Savanna 4.x, tokens come from /gsql/v1/tokens, not the older /restpp/requesttoken.
  • Turn on Auto Resume, or a suspended workspace returns HTTP 500 to every API call.
  • Installed GSQL queries called over REST need every parameter supplied, so use no-op defaults.

Code: https://github.com/repulsortechnologies-alt/agentic-graphrag

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The Beijing venue-and-date example suggests a useful ablation: give the one-query GraphRAG baseline the same candidate-disambiguation rule, then compare it with the agent again. Returning the first match is a specific baseline limitation, so the final eight points may partly measure that choice rather than a requirement for iterative planning.

I would report candidate cardinality after each constraint: venue, date, event category, then any tie-break. Cases still ambiguous after all available fields are applied should remain ambiguous in the output. A similarity tiebreak can choose a plausible event, but it cannot create a missing fact that makes the choice uniquely supported.