DEV Community

Cover image for My RAG pipeline answered 1. The real answer was 20.
Harshdip Saha
Harshdip Saha

Posted on

My RAG pipeline answered 1. The real answer was 20.

I asked three pipelines one question from the TigerGraph Agentic GraphRAG Hackathon benchmark: how many athletics events at the 2004 Summer Olympics had more than 41 competitors?

Plain RAG retrieved 4 chunks and answered 1. GraphRAG expanded one hop, saw 12 chunks, and answered 2. The real answer is 20, out of 43 matching events in the graph.

Both pipelines sounded equally confident. Neither one had any way to know that plain RAG was looking at 4 of the 43 matching events (9.3%). Picture an analyst pasting that "2" into a report because it sounded sure.

The 9% problem

Similarity search returns what looks relevant, but a counting question needs everything that is, and no value of k fixes that, because the right k depends on the answer you haven't computed yet.

On the 100 public benchmark questions, RAG and GraphRAG scored 0 out of 31 on aggregation and superlative questions.

Also, as a user I want to know whether the answer I got is complete, and a confident tone tells me nothing about that.

Investigation Certificates

TigerGraph already knows how many events match sport = athletics AND games = 2004 Summer. It is a COUNT. I call that number the structural bound (the size of the complete answer set, straight from the graph).

So every answer my pipeline emits ships with an Investigation Certificate: a JSON record that puts the structural bound next to the evidence the agent actually inspected.

{
  "qid": "pub-045",
  "completeness_class": "exhaustive",
  "structural_bound": 43,
  "evidence_set_size": 43,
  "completeness_check": "pass",
  "stop_reason": "structural_bound_met",
  "tokens": { "total": 0 },
  "latency_ms": 227
}
Enter fullscreen mode Exit fullscreen mode

43 equals 43 (0 tokens here, since this question never calls the LLM), so the evidence set matches the graph's own count. Also, you can check that yourself without trusting the model. If the numbers don't match, the certificate says so instead of hiding it. Run the same question through plain RAG and its certificate would read evidence 4 against bound 43, a visible fail instead of a confident "1".

How the routing works

A router sorts each question into one of 3 completeness classes and picks the cheapest tool that can satisfy it:

  • existential (lookup, multi-hop): exact match or venue+date traversal. The certificate proves the entity resolved uniquely.
  • chained (temporal): a PREV/NEXT hop chain. The certificate proves the hops were actually walked.
  • exhaustive (aggregation, superlative): a GSQL structural scan returning the COUNT plus every match. The certificate proves evidence set equals the graph's own count.

The LLM is called only to break a tie or recover a missing field. It never counts.

Besides the 5 regex templates, I added TypeSafe AI's Jev System One as a fallback. It is a non-autoregressive decision model, so it classifies intent and picks between tied candidates at 0 output tokens. On 5 un-templated questions the regex router sent all 5 to unverified RAG, and Jev routed all 5 correctly (5/5). That small test (5 queries) is all the evidence I have for the classifier, so I'd call it promising and unproven.

The numbers

Same 100 public questions, same live TigerGraph graph, same live Groq model:

RAG GraphRAG Agentic
Accuracy 25% 43% 99%
Tokens per answer 1,389 2,203 18
Latency per answer 9.3s 15.5s 0.41s

Thus the agentic pipeline uses 122x fewer tokens (2,203 vs 18) and runs 37.8x faster than GraphRAG (15.5s vs 0.41s), at 2.3x the accuracy. The cost was always in stuffing chunks into a context window, and the graph queries are nearly free.

The one miss

I got 99 out of 100. The miss is pub-060, a date+venue tie among 37 candidate events at ExCeL. Jev picked one with 0.90 confidence and picked the wrong one. The LLM guessed wrong on this question in my earlier run too, so Jev did not fix it.

The certificate for it reports pass_with_llm_recovery, not a clean pass. That is deliberate. The certificate measures whether the evidence was complete, the way an auditor signs off on the books and leaves the strategy to someone else. pub-060 is a complete but wrong case, where evidence and answer come apart.

Try it

Built for the TigerGraph Agentic GraphRAG Hackathon with TigerGraph Savanna, TypeSafe AI Jev, Groq and Streamlit.

Overall, a certificate vouches for the evidence, and answer correctness stays a separate claim, so a system should say which one it is making. If you have a RAG pipeline that answers counting questions, I'm curious what your structural bound would be. :)

Top comments (0)