DEV Community

Reid Marlow
Reid Marlow

Posted on Originally published at reidmarlow.com

Search Agents Waste Half Their Tokens Rediscovering Entity Links

If you wire an LLM agent to a local directory of documents and give it terminal tools (grep, find, cat), you quickly notice an ugly pattern. When a question depends on evidence scattered across three separate files, the agent spends most of its trajectory wandering in circles. It greps for a keyword, pulls up five irrelevant markdown files, reads their headers, backs up, reformulates the search, and tries again.

On turn six, it finally discovers that Project Atlas has an approval slip in one folder, technical specifications in another, and a status report in a third. It answers the question, but the transcript shows two hundred thousand tokens burned on blind navigation.

A new paper from researchers at KAIST and Microsoft, titled "Follow the Entities: A Corpus Map for Agentic Search" (arXiv:2609.37226), quantifies exactly how much compute goes up in smoke during these runs. On EnterpriseRAG-Bench, an agent searching a raw flat corpus consumed an average of 206,500 input tokens per query trajectory just to hit 62.1 percent correctness. On WixQA, the raw search burned 337,200 tokens per query.

The culprit is straightforward. A raw document collection gives an agent zero relational pointers between files. Every single incoming query forces the model to reconstruct the entire web of cross-document relationships from scratch at inference time.

Why standard abstractions fail

Developers usually try two workarounds when raw grep starts blowing up context windows. Neither holds up under scrutiny.

The first workaround is folder-based aggregation. People group files into directories, ask an LLM to generate an index page for each folder, and hand those index files to the agent. In the paper's experiments under the Group Page baseline, this strategy backfired completely. Token consumption exploded to 423,000 tokens on EnterpriseRAG-Bench and over 1.1 million tokens on WixQA, while correctness dropped to 48.8 percent. Organizational folder trees mirror team charts or file formats, not the real questions people ask. Splitting related documentation across arbitrary folder walls just forces the agent to read redundant directory summaries before it can find the underlying evidence.

The second workaround is the unconstrained LLM wiki, popular following Andrej Karpathy's experiments. You let a language model ingest the corpus and freely draft cross-linked markdown notes. In practice, free-form wikis suffer from hallucinations and missing cross-references. On EnterpriseRAG-Bench, the LLM Wiki baseline burned 192,400 tokens and achieved only 56.7 percent correctness, falling behind even raw corpus search. Without strict grounding against raw files, the agent navigates through loose conceptual associations that drop hard facts.

Traditional graph retrieval methods like GraphRAG and HippoRAG avoid the agent navigation loop entirely by retrieving a fixed context window upfront. But as the authors demonstrate, fixed retrieve-then-generate pipelines cap performance because the model cannot inspect downstream sources if the initial graph walk misses an edge.

The entity map architecture

The authors propose a system called CorpusMap that sits between flat storage and the search agent.

Instead of organizing files by directories or abstract topic clusters, CorpusMap structures the corpus around recurring named entities. These entities are people, projects, systems, vendors, and code modules that appear across multiple independent documents.

Offline, an extraction pipeline identifies these recurring anchors and generates a dedicated Entity Page for each one. The Entity Page does two things. It aggregates key facts about that specific entity, and it maintains explicit, verified backlinks to every original document that mentions it. This creates a clean bipartite graph between entities and raw documents.

When the agent receives a task, it navigates this graph using standard terminal commands. Instead of blind keyword grepping, the agent jumps directly to the relevant entity page, inspects the consolidated facts, and follows direct links to the exact source documents it needs to verify.

The impact on trajectory efficiency is substantial:

On EnterpriseRAG-Bench using GPT-5.5, CorpusMap cut average input token consumption from 206,500 tokens down to 88,100 tokens, representing a 57 percent reduction. At the same time, answer correctness jumped from 62.1 percent to 73.8 percent.

On WixQA, token consumption dropped from 337,200 tokens to 74,500 tokens, a 78 percent drop, while factual accuracy rose from 67.5 percent to 70.7 percent.

Across seven different model families, including GPT-5.6 Sol, DeepSeek-V4-Pro, and Qwen3.8-27B, the pattern repeated consistently. Providing explicit entity anchors prevented the agent from getting lost in recursive search subroutines.

The economics of offline indexing

The obvious objection to building entity maps is the upfront indexing cost. Extracting cross-document entities across thousands of pages requires significant LLM inference.

The paper includes a transferability ablation in Table 4 that addresses this concern directly. When researchers used GPT-5.5 to construct the CorpusMap for 2,819 enterprise documents, the one-time indexing bill reached roughly $4,732. But when they swapped the builder model to GPT-5.6 Luna, construction cost dropped to $74.65.

Crucially, when high-end models like GPT-5.5 and GPT-5.6 Sol answered queries over the cheap Luna-built map, they maintained quality scores of 73.59 and 74.68, virtually matching the quality of the expensive GPT-5.5-built map. The structure of the entity graph matters far more than the prose style of the entity summary. You can run entity extraction with a cheap utility model and hand the resulting map to your expensive frontier agent without losing retrieval quality.

The authors also tested incremental updates. When new documents arrive, the system updates only the entity pages touched by those specific files rather than reprocessing the entire corpus. In their benchmarks, incremental updates saved 69 to 71 percent of tokens compared to full rebuilds while preserving overall answer quality.

Practical takeaways for local workflows

If you maintain agent workflows over project repositories, internal wikis, or legal folders, there are three immediate takeaways from this work.

First, stop expecting frontier models to compensate for unstructured storage. Adding more reasoning tokens to an agent does not fix the absence of cross-document links. It just gives the model more runway to burn cash on repetitive grep commands.

Second, avoid folder-centric indexes. Summarizing directories by folder path creates artificial walls that actively degrade multi-document recall. If an agent needs to correlate an infrastructure outage with a vendor contract and a commit log, folder hierarchies hide the connection.

Third, extract entities once and maintain explicit backlinks. You do not need a complex graph database or a proprietary framework to implement this. A directory of markdown files where each file represents a recurring entity and lists relative paths to source documents gives a terminal agent everything it needs. You pay the extraction cost once, and your search loops stop wandering in the dark.

Top comments (5)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

Table 4's ablation deserves more attention than it gets: building the map with GPT-5.6 Luna instead of GPT-5.5 cut indexing from $4,732 to $74.65 while quality held at 73.59 and 74.68. The leg it doesn't test is entity dedup — a cheap model merging 'Atlas' with 'Project Atlas' silently fuses two backlink lists onto one page, and content hashes can't catch that because the pointers aren't stale, just misattributed. Cheap extraction is fine; the merge step is where you still spend the money.

Collapse
 
reidmarlow profile image
Reid Marlow •

That misattribution problem is why greedy alias merging causes so much damage downstream. Once two distinct entities get collapsed into one canonical record, every downstream retrieval call inherits that poisoned neighborhood, and reranking cannot fix a bad graph edge.

A workable balance is splitting link extraction from alias resolution. You can let a cheap model propose candidate entity mentions, but gate the merge step behind an explicit disambiguation check before updating the canonical table. Keeping candidate clusters separate until there is verifiable overlap keeps the blast radius contained to individual document lookups instead of polluting the global index.

Collapse
 
slabb profile image
Sam LABBE •

The part of this that outlives the benchmark: the entity map becomes a new trusted component.
The extraction pipeline that builds it is itself an LLM, so every hallucinated fact or wrong entity merge it makes gets baked into the structure everything downstream trusts — and the paper's verification covers the links, not the facts on the pages.
Two cheap hardenings: bind each entity page to the content hashes of its source documents at extraction time, so a page that stops matching its sources flags itself stale instead of pointing confidently at history; and version the map, so a pipeline re-run diffs like any other build artifact.
Otherwise you've fixed navigation and quietly recreated the LLM-wiki failure one level down — confident pointers that outlive the things they pointed at. (Same shape as install lines in docs, incidentally — the pointer outliving its referent is where trust rots.)

Collapse
 
reidmarlow profile image
Reid Marlow •

Treating the map as a compiled build artifact with content hashes is the right move here. The riskiest part of the paper's setup is letting the offline extractor write natural language fact summaries onto the entity page.

Once an entity page stores prose claims, downstream agents read that summary and skip opening the underlying files. Stripping the entity page down to a pure structural index (canonical entity ID, known aliases, source SHA-256 hashes, and line offsets) gives the agent the navigation speedup while forcing every factual claim to come straight from the raw document.

Collapse
 
slabb profile image
Sam LABBE •

That's the cleaner cut. A page that holds only pointers can't lie about content — it can only go stale, and staleness is mechanical: a background hash sweep over the bipartite graph turns the map into something that reports its own rot.

And there's a familiar property hiding in it: the trusted layer holds no facts, so it doesn't need to be trusted, only checked — same shape as a domain-blind journal carrying claims about references instead of content.

Good exchange. Keep the prose off the page.