What happens when a question looks simple, but answering it correctly requires more than retrieving a few documents?
For my TigerGraph Hackathon project, I explored this question by implementing and comparing RAG, GraphRAG, and Agentic GraphRAG on the same benchmark questions.
The goal was not simply to prove that "agents are better."
Instead, I wanted to understand something more practical:
When does a question actually need an investigation instead of a retrieval?
That question became the central idea behind my project:
Beyond the First Answer
The Problem: Retrieval Is Not Always Enough
A traditional RAG (Retrieval-Augmented Generation) pipeline is very good at answering questions where the required information exists directly inside a few relevant documents.
User Question
↓
Similarity Search
↓
Relevant Documents
↓
LLM
↓
Answer
For many questions, this is completely sufficient.
But consider a question such as:
How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?
At first glance, this sounds like a normal retrieval question.
But it isn't just asking the system to find information.
The system needs to:
- Find the relevant Olympic event documents.
- Identify the biathlon events.
- Extract the competitor counts.
- Count the qualifying events.
- Produce the final answer.
The important distinction is that the answer requires a sequence of operations.
That is where I started thinking about retrieval differently.
Retrieval finds evidence. Investigation decides what to do with that evidence.
My Agentic GraphRAG Architecture
My implementation uses a routing and execution architecture built around several components.
USER QUESTION
↓
QUESTION ROUTER
↓
AGENTIC AGENT
↓
┌───────────────┐
│ AUTO → PLANNED│
└───────┬───────┘
↓
PLANNER
↓
PLAN VALIDATOR
↓ ↓
VALID INVALID
↓ ↓
EXECUTOR REACTIVE
↓ ↓
└────┬────┘
↓
TOOL REGISTRY
↓
┌────────────────┼────────────────┐
↓ ↓ ↓
Similarity / Graph / Deterministic
Hybrid Search Structural Aggregation
Retrieval
└────────────────┼────────────────┘
↓
TOOL RESULTS
↓
FINAL ANSWER
The flow starts with the Question Router.
The router classifies the question into categories such as:
- lookup
- aggregation
- temporal
- superlative
- multi-hop
The question then reaches the Agentic Agent.
In my current configuration, Auto resolves to Planned mode.
The Planner creates a structured plan, and the Plan Validator checks whether that plan is valid.
If the plan is valid, it is passed to the Executor.
If the generated plan is invalid, the system can fall back to Reactive mode.
Important clarification: Invalid means the generated plan is invalid — not that the user's question is invalid.
What Does the Tool Registry Do?
The Tool Registry provides a common dispatch layer between the agentic execution logic and the actual tools.
This allows the system to access capabilities such as:
- Similarity search
- Hybrid search
- Graph / structural retrieval
- Contextual retrieval
- Document retrieval
- Deterministic aggregation
- Other registered graph tools
Instead of allowing the agent to directly execute arbitrary operations, tool calls pass through a common registry, where the appropriate tool can be validated and dispatched.
This becomes particularly important when the system is performing multi-step investigation.
The Failure That Changed My Design
The most interesting part of the project was not when everything worked.
It was when something looked correct but wasn't.
I tested the question:
How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?
The correct answer is:
5
Hybrid retrieval was actually doing its job.
It successfully found the relevant event documents.
So I initially thought the retrieval pipeline was fine.
But the final answer was inconsistent.
In one run, the LLM returned:
4
In another run, it identified the correct five qualifying events but concluded:
6
That was an important observation.
The evidence was there.
The problem was the computation and interpretation of that evidence.
This changed the way I designed the agentic pipeline.
Instead of relying entirely on the LLM to perform numerical reasoning over retrieved context, I introduced deterministic aggregation tools into the execution layer.
The idea was simple:
Let the retrieval system find the evidence, but use deterministic operations when the task is fundamentally computational.
Comparing RAG, GraphRAG, and Agentic GraphRAG
I evaluated all three approaches using the same benchmark setup.
Public-100 Evaluation
| Pipeline | Correct | Accuracy |
|---|---|---|
| Normal RAG | 41 / 100 | 41% |
| GraphRAG | 55 / 100 | 55% |
| Agentic GraphRAG | 72 / 100 | 72% |
The results were encouraging.
Agentic GraphRAG achieved 72% accuracy, compared with 55% for GraphRAG and 41% for Normal RAG.
This suggests that adaptive investigation can help when questions require more than a single retrieval operation.
But there is an important second side to the story.
Agents Are Not Free
Agentic reasoning comes with a cost.
In my Hidden-50 evaluation, the average cost per question was approximately:
Cost and Latency Comparison
| Pipeline | Avg. Tokens / Question | Avg. Time / Question |
|---|---|---|
| Normal RAG | 13,496.44 | 22.96 sec |
| GraphRAG | 1,218.77 | 8.38 sec |
| Agentic GraphRAG | 60,737.54 | 100.53 sec |
The difference is significant.
Agentic GraphRAG used substantially more tokens and took considerably longer.
This creates an important engineering trade-off:
More Investigation
↓
More Tool Calls
↓
More Reasoning
↓
Higher Accuracy
+
Higher Cost / Latency
Therefore, simply replacing every RAG pipeline with an agentic system would not necessarily be a good engineering decision.
What I Learned
So I don't think the conclusion should be:
"Agentic GraphRAG is always better."
That would miss one of the most important lessons from the experiment.
The more useful conclusion is:
Use the simplest retrieval strategy that can reliably answer the question. Escalate to investigation when the question actually requires it.
This leads to a more practical architecture:
USER QUESTION
↓
Can retrieval
answer it reliably?
↙ ↘
YES NO
↓ ↓
RETRIEVE INVESTIGATE
↓ ↓
ANSWER PLAN → TOOLS
↓
VALIDATE
↓
ANSWER
The goal is therefore not to make every question agentic.
The goal is to make the system adaptive.
The Core Idea
My project explores a simple principle:
Not every question needs an agent. But some questions need more than retrieval.
A lookup question may only need document retrieval.
A structural question may benefit from a graph.
A multi-hop or aggregation question may require planning, multiple tool calls, and deterministic computation.
The challenge is deciding when to escalate.
That is the problem I explored with Beyond the First Answer.
Final Takeaway
The biggest lesson from this project was not that one architecture wins every benchmark.
It was that retrieval and reasoning solve different parts of the problem.
RAG is efficient when the answer is directly present in relevant documents.
GraphRAG adds structural relationships that can improve retrieval and context.
Agentic GraphRAG becomes valuable when answering the question requires an actual investigation across multiple steps.
But that additional intelligence comes with a cost in tokens and latency.
So my design philosophy became:
Retrieve first. Investigate when necessary. Compute deterministically when possible.
That is what I mean by:
Top comments (0)