DEV Community

Cover image for Graph RAG: where it actually breaks
Tae Kim
Tae Kim

Posted on Originally published at hannune.ai

Graph RAG: where it actually breaks

Neo4j with a working schema: two days. Cypher traversal for the relationships I needed: another day or two, once I knew what I was querying. The graph structure, once committed, stayed mostly stable.

That's not where things kept breaking.

The routing problem

When someone asks "what companies are exposed to semiconductor export restrictions?" the system needs to figure out whether to do graph traversal, vector search, or some combination. Getting this wrong doesn't produce an obvious error. It returns something that cites real sources and sounds confident. It just misses context that was sitting in the graph the whole time.

I had one case that stayed with me. A query about a Korean chipmaker's supplier network kept returning news results. The answer was plausible — recent articles about the company, relevant details. It took longer than I'd like to admit to notice that the supplier relationships were already in the graph. The system just wasn't routing that query there. Once I found the gap, the fix was two lines of code. Finding it was the work.

The routing logic I ended up with uses the query's dependency structure as a signal:

def route_query(query: str, graph_searcher, vector_searcher) -> list:
    relational_signals = ["supplier", "subsidiary", "controlled by", 
                          "exposed to", "connected to", "supply chain"]

    needs_graph = any(s in query.lower() for s in relational_signals)

    if needs_graph:
        entity_ids = graph_searcher.find_related(query)
        return vector_searcher.search_scoped(query, entity_ids)
    else:
        return vector_searcher.search(query)
Enter fullscreen mode Exit fullscreen mode

The signal list is hand-tuned and the scoping logic has edge cases. But it reduced the miss rate more than anything else I tried.

Entity quality doesn't stay clean

The second place things break is entity quality, and this one is more insidious because it's gradual.

The initial entity resolution pass establishes a set of merged entities. After that, the source data keeps changing. A company I had cleanly resolved was acquired six months into the project. Post-acquisition news started using a new parent company name alongside the old subsidiary name. The graph started accumulating both as separate nodes for the same underlying entity.

Graph traversal from the old node gave one picture; traversal from the new node gave a different one. Neither looked wrong in isolation.

MATCH (e:Entity {name: "Old Name"})-[:SUPPLIES]->(s:Entity)
RETURN s.name, s.country

-- If the same entity now also exists as "New Name" post-acquisition,
-- those relationships are on a different node entirely
Enter fullscreen mode Exit fullscreen mode

I treat entity quality as something that degrades continuously now, not a state you establish once. Running a periodic audit of merged pairs helps — checking whether the alias-to-entity mapping still makes sense for recently active entities.

from datetime import date, timedelta

def stale_merges(registry: dict, days_threshold: int = 90) -> list:
    cutoff = date.today() - timedelta(days=days_threshold)
    stale = []
    for entity_id, meta in registry["entities"].items():
        last_verified = meta.get("last_verified")
        if last_verified:
            verified_date = date.fromisoformat(last_verified)
            if verified_date < cutoff and meta.get("alias_count", 0) > 3:
                stale.append(entity_id)
    return stale
Enter fullscreen mode Exit fullscreen mode

The honest version: figuring out which merged pairs have actually gone stale is still mostly manual review. The detection is automatable. What to do with a flagged pair isn't something the code decides.

I build er-api, a multilingual entity resolution service for Korean, Japanese, Chinese, and English corporate data. More at hannune.ai.

Top comments (5)

Collapse
 
hannune profile image
Tae Kim •

The successor link point is exactly what I would have missed without seeing a real case: writing "valid_to" closes the validity window, but it leaves the graph structurally silent about what filled the gap, and you only discover the cost when something tries to walk the chain backward. The rename trigger gap matches what I saw too, as a company rebranded during an acquisition and word overlap gave me nothing because the new name shared zero tokens with the old one, so the registered key was the only path. What does your registered key schema look like, a side table of known equivalences you populate manually, or does some upstream process feed it?

Collapse
 
gde03 profile image
Giulio D'Erme • • Edited

Tae, this matches what I have seen, especially your point that a misrouted query still returns a confident answer that cites real sources. That failure never shows up as an error. Two of my results bear on your two failure modes.

Routing. I stopped letting the graph compete for the top of the list. The graph cannot move the first eight direct results. It may promote at most two graph-linked candidates into positions 9 and 10, and only candidates that plain retrieval had already found. On the full LoCoMo set (1,536 questions), the share of questions with complete evidence in the top 10 was:

64.39% with plain retrieval
57.03% when the graph took five of the ten slots
66.67% with the 8 + 2 split (45 rescues, 17 regressions)

A second full run was also positive (64.26% to 66.21%) but missed the two point gate I had registered in advance. So I read it as a small, repeatable gain. What mattered more is that a routing mistake became cheap, because the most the graph can do is change the last two slots. In a separate routing layer, which decides the store and embedder for a query, routing on words in the query cost me recall. Enabling features based on what the stored data actually contains worked better.

Edges, not entities, were my bottleneck. On conversation transcripts my extractor produced almost 10,000 entities and zero relations, so I had a set of nodes rather than a graph. On authored notes it found several thousand relations, and every one was a generic references edge.

Two questions:

How do you get typed edges like supplier or subsidiary out of unstructured text, and what precision do you measure on them? I suspect extraction precision limits everything after it.
In the acquisition case, do you merge the two nodes, or keep both and add a dated edge ("subsidiary of Y from this date")? My semantic graph had no relation type for supersession, so a traversal could never follow "this replaced that". I can't tell whether merging or dating is the better fix.

PS the link for er-api doesn't work

Collapse
 
hannune profile image
Tae Kim •

Typed edges: I run a separate relation classifier after the NER pass, feeding each entity pair plus their surrounding context window through it, and on Korean financial filings the supplier type lands around 71% precision, which is honestly low enough that I gate everything behind a confidence threshold before it touches the graph. Statewave's answer on merge vs date matches what I landed on too, and the successor link is the piece that costs you if you skip it. The trigger is where it gets ugly: name-overlap gives you nothing during a rename, so I depend on an explicit acquisition event from a structured filing, and any deal that never hit that path just stacked up two nodes until the next audit ran.

Collapse
 
gde03 profile image
Giulio D'Erme •

Thanks, Tae. Before replying I tried your approach on my own notes, so I can answer with numbers instead of guesses.

It works. With no trigger, a strong model behind a confidence gate reached only about 0.42 precision, because only about 3.5% of my candidate pairs were real. Adding a structural trigger first (two notes written the same day, or sharing words in their titles) cut the candidates by twelve times, while keeping most real pairs. With the gate at 0.90, precision on the actual candidate set is about 0.80, and it finds 24 of my 58 known pairs; most of the other links it proposed were real supersessions I had never recorded. One caveat: on a second, quite different set of notes, the same trigger and threshold reached only about 0.50, so in my case they need setting per corpus.

On merge or date, I agree with you: taking the direction from an authored timestamp was right on 45 of 47 pairs, while asking gpt-4o-mini to judge it from the text flipped 12 of the 30 true pairs it had found.

And on how the link is used, the difference was large. Showing a small reader perfect successor links as notes raised its accuracy by only 8 points, and it still gave the stale answer about a fifth of the time. Removing the superseded note before the reader saw anything raised it by 23 points and almost eliminated stale answers. With the links from my actual pipeline, the gain was 9 points.

Two questions, if you have a minute:

Is your 71% measured before or after the confidence gate, and roughly what share of candidates survives it?
Does your threshold hold across filing types, or do you set it per source?

Collapse
 
statewave profile image
Statewave •

On the merge-or-date question, because we picked one of the two and got half of it wrong.

We keep both and date them. When a newer fact supersedes an older one, the old row stays, its validity window closes at the new one's start, and retrieval stops returning it while the audit trail still shows it. Merging would have destroyed the thing we need most often later: what the system believed before, and when it stopped believing it.

The half we got wrong: we closed the old row without recording what replaced it. The state ends up correct, but a traversal still cannot follow "this replaced that" — the successor link only exists as a read-time guess in an admin view. It's open work for us now. If you're choosing today, write that link at the moment you make the decision. It costs almost nothing then and is expensive to reconstruct later.

Whichever you pick, the harder half is the trigger. Ours fires on word overlap or an explicitly registered key, and a rename is exactly where overlap gives you nothing: "Old Name" and "New Name" share no tokens. Only the key catches that case, which matches Tae's acquisition example — the graph accumulated both nodes because nothing said they were the same thing.