Agentic GraphRAG Hackathon by TigerGraph, Round 1. Code: https://github.com/anirudh12032008/agentic-graphrag-tigergraph
Ask a RAG system "How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?" and it will give you a confident number. It will also be wrong. The answer depends on 8 to 43 documents, and a top-8 vector search can't see all of them.
We built three pipelines on the same model and the same corpus to find out where graphs help, where agents help, and where neither is needed. The agentic pipeline scored 99% on 100 public questions. Plain RAG scored 48%.
The setup
The corpus is 2,951 English Wikipedia articles about Olympic events from 1987 to 2023, plus distractor pages. The 100 public questions come from five templates:
- lookup: a single fact from one document
- aggregation: count events over a threshold across a whole sport at a Games
- superlative: which event had the most competitors
- temporal: who won the same event at the Games immediately before 2016
- multi_hop: who won the event at venue X on date Y
All three pipelines use the same LLM (claude-sonnet-5), the same embeddings, and the same TigerGraph Savanna instance. An LLM judge grades answers against gold labels, after a deterministic pre-check. Judge tokens are never counted against a pipeline.
The graph
Event pages start with an infobox. We parse it deterministically, with no LLM extraction, so the graph is exact, free to build and reproducible. The TigerGraph schema has Event, Sport, Games, Venue, Athlete, Country, Document and Chunk vertices. Edges include IN_SPORT, AT_GAMES, HELD_AT, WON, PREV_EDITION and HAS_CHUNK. Chunks carry an embedding, so vector search runs inside the same database as the graph.
The three pipelines
- RAG: vector top-8 chunks, then one LLM call.
- GraphRAG: vector top-6, up to 3 seed events, then a fixed graph expansion (medalists, previous edition, sibling events), then one LLM call.
- Agentic GraphRAG: an LLM orchestrator picks tools one step at a time. Tools include entity linking, GSQL aggregation, graph traversal, venue and date lookup, vector search and document reading. An evidence evaluator checks every final answer.
The evaluator rejects an answer if it has no citations, if it cites a document no tool returned, or if it lists unresolved gaps with low confidence. The loop is capped at 8 steps. A circuit breaker stops the run after 2 consecutive tool errors, so a backend outage doesn't burn tokens.
Results (100 public questions)
| pipeline | accuracy | avg tokens | avg latency (s) | tokens per correct answer |
|---|---|---|---|---|
| RAG | 48% | 6,466 | 7.1 | 13,471 |
| GraphRAG | 69% | 3,815 | 5.7 | 5,529 |
| Agentic GraphRAG | 99% | 8,001 | 8.1 | 8,082 |
The agent uses 1.24x the tokens of RAG for 2.06x the accuracy. Per correct answer it is cheaper than RAG by 40%. GraphRAG is the cheapest per correct answer, but it tops out at 69%.
Where the agent matters
| question type | RAG | GraphRAG | Agentic |
|---|---|---|---|
| lookup (19) | 100% | 100% | 100% |
| aggregation (21) | 0% | 67% | 100% |
| superlative (10) | 40% | 80% | 90% |
| temporal (22) | 36% | 86% | 100% |
| multi_hop (28) | 57% | 21% | 100% |
Lookups don't need an agent. Every pipeline gets them right, and RAG does it in the fewest steps. Pay for the agent only when the question needs it.
Aggregation is where RAG breaks. It scored 0 of 21. The evidence is spread over too many documents for k=8 chunks, and no prompt fixes that. The agent runs one GSQL count and gets an exact answer.
GraphRAG can be worse than RAG. On multi-hop questions GraphRAG scored 21% against RAG's 57%. Its fixed expansion seeds from the wrong events, and then every fact it adds is confidently about the wrong thing. The agent first links the venue, then filters by date, then reads the winner.
Two things the agent did that we didn't script
It flagged real ambiguity. "The Olympic Aquatic Centre on August 14, 2004" matches three finals (Thorpe, Phelps and Klochkova). The venue tool returns every match with an ambiguous: true flag, and the agent reports all candidates with citations instead of guessing.
Its one miss was a tie in the data. In 2008 fencing, men's épée and women's foil both had 41 competitors. The agent reported the tie, and the gold label picked one. We counted that as a miss.
What the agent costs
It makes 2.6 LLM calls per question on average, and a typical trace has about 5 entries across orchestrator, evaluator and tool steps. On the 50 hidden questions it averaged 6,809 tokens and 7.2 seconds. The cost is real but small compared to the accuracy gain on the question types where it matters.
RAG sometimes couldn't finish at all. On a few aggregation questions it used up a 4,096-token output budget reasoning over 8 chunks and produced no answer. We recorded those as failures and did not hide them.
What we'd do next
The schema already links editions with PREV_EDITION and NEXT_EDITION. For Round 2 we plan a ClaimVersion vertex per sourced fact with a SUPERSEDES edge. That would let the agent notice when documents disagree and pick the latest or most authoritative claim, with provenance.
Takeaways
- Use a graph when the answer depends on counting, comparing or hopping across many documents. Top-k retrieval can't do that however good the embeddings are.
- Use an agent when the right retrieval depends on what you've found so far. A fixed pipeline expands from the wrong place and can't recover.
- Skip both for single-document lookups.
- Give the agent an evaluator. Requiring citations from tool results is what stops it from answering from memory.
Everything is open source: the graph schema, the GSQL queries, the three pipelines, the benchmark harness and the Streamlit dashboard. Clone the repo and run make ingest && make bench-public to reproduce the table above.
Corpus text is derived from English Wikipedia (CC BY-SA 4.0).
Top comments (0)