Where an agent's tokens go
Watch a coding agent start on an unfamiliar change and most of the early turns
look the same. It greps for a name, opens three of the matches, finds an import,
opens that file, greps for the caller, opens two more. Each read puts a whole
file into context, most of which has nothing to do with the question. By the
time it knows where the change goes, it has spent a large share of its budget on
navigation.
A code graph shortens that loop. If the agent can ask "who calls
renderInvoice, and what does it touch" and get the answer with the relevant
source attached, it skips the grep-read-grep cycle. That part is not new.
Two things about how ivar is used made the usual shape of a code graph a poor
fit.
The change spans repositories. A hall mounts an API, a web client and a
shared package side by side. The interesting edges are between them: the web
client's HTTP request and the API route that serves it, or a
type exported by the shared package and used in both. An index built per repo
does not have those edges.
Several features run at once. Each feature is its own branch, checked out as
its own set of worktrees, often with an agent working in each. A graph scoped to
a folder is built for one checkout. Point it at five feature checkouts and you
either index the same repo five times or answer every session from one of them,
which means some agents see code that is not on their branch.
The design
One base, one layer per feature
ivar graph keeps a single SQLite database for the hall at .ivar/memory.db.
It holds a base index of each repo's default branch, shared by every
session. On top of it, each feature gets a layer that holds only the files
the feature changed, committed or not.
A feature session reads its own layer first and the base for everything else. A discovery session reads the base.
A feature usually changes tens of files, so a layer is small next to the base.
Queries from a feature session resolve symbols against the layer first and fall
back to the base, which is how an agent sees its own edits without the graph
duplicating the rest of the repo.
Extraction uses tree-sitter and runs in parallel. Base indexing is incremental:
ivar diffs each repo against the commit it last indexed and checks file size
and modification time, so a rerun parses only what moved.
Freshness on every call
A layer goes stale the moment the agent edits a file. Asking the agent to call a
reindex tool after each edit works until it forgets. So before every query the
MCP server compares the layer with the worktree and reindexes the files that
changed.
That check has to be cheap to run on every call. On a feature with 177 changed
files it took about 18 ms per call, and about 21 ms right after an edit.
If the layer cannot be refreshed, the call fails with an error. It does not
quietly answer from the base: an answer about code that is no longer on disk is
the kind of wrong an agent acts on with confidence.
One query that returns enough to act on
The MCP server advertises one tool by default, graph_explore. The agent passes
symbol names, paths, or a short intent, several at once. The answer includes:
- the flow between the symbols it named,
- the line-numbered source of the most relevant files, cut to a size budget,
- the blast radius of each definition: its callers and type uses,
- a ready
Next:call with apathsargument listing the files that did not fit.
The source is the part that removes reads. The Next: hint is the part that
keeps the agent from falling back to opening files one at a time: a follow-up
with paths returns up to 12 whole files in one call.
The other tools (callers, impact, affected tests, paths and more) are available
behind ivar graph mcp --tools all, and the queries have CLI equivalents, which
matters for providers such as omp that do not pass MCP servers to subagents. A
subagent can run ivar graph explore renderInvoice in a shell and get the same
answer.
What it changed in a benchmark
We ran agents on the same fixture with and without the graph, using the same
model and the same task prompt in both arms. The fixture is two repositories: a Fastify API and a
React web app that calls it.
Discovery
The agent explores the fixture to answer a question about how the code works,
and its answer is scored against a hand-built answer key. Three runs per arm.
With ivar graph
|
Without a graph | |
|---|---|---|
| Median tokens | 460k | 738k |
| File reads | 0 | 12 to 28 |
| Answer-key coverage | 93 to 94 % | 91 to 94 % |
The graph arm answered as completely while reading no files; the source it
needed arrived inside the query results. Tokens fell by about 38 %.
Implementation
The agent implements a change in the fixture that has to pass an acceptance
suite. Five runs per arm, and every run in both arms passed.
With ivar graph
|
Without a graph | |
|---|---|---|
| Median tokens | 1.53 M | 1.61 M |
| Median file reads | 6 | 26 |
Here the token saving is small, about 5 %. Most of an implementation run is likely
spent editing, running tests and reading their output, and the graph does not
help with that. File reads still fell from 26 to 6, and the reads that remain
are the ones the provider requires before it lets an agent edit a file.
On a real hall
The fixture is small, so we also measured the graph on a working hall of about
29,000 files, which indexed to 144,000 symbols and 1.9 million edges.
| Measure | Result |
|---|---|
| Index size | 370 MB |
| Cold index | 77 s |
| Reindex, no changes | 0.17 s |
| MCP server startup | about 5 ms |
graph_explore |
12 to 80 ms |
| Callers query | 0.4 ms |
| Impact query | 0.2 ms |
Caveats
These are small samples: three and five runs per arm, on one fixture, with one
model. The discovery answer key was built by hand, so coverage is a judgment against
it, not an absolute measure. Token counts vary run to
run, which is why the tables report medians and ranges rather than a single
number. The implementation result in particular should be read as "fewer reads
for about the same cost", not as a large saving.
Try it
Build the index from the hall root:
ivar graph index
The first run also registers the graph MCP server in ivar.json (pass
--no-mcp to skip it).
Then run ivar sync to write the provider configs, and start a feature session.
ivar sync also keeps the base index current from then on, and ivar doctor
reports repos whose index has fallen behind.
The Code graph guide covers layers and limits, and the
reference lists every subcommand and MCP tool.

Top comments (0)