DEV Community

Cover image for Where a code graph pays off: big repositories, small models
Rodion Kazennov
Rodion Kazennov

Posted on Originally published at allkeep.org

Where a code graph pays off: big repositories, small models

My last benchmark said a code graph did not make my coding agent cheaper. It was too small to say more than that: 24 questions, nothing bigger than 25k graph nodes, no Opus. So I ran it again at a size where the answer could change, and wrote the plan down first. The hypotheses, repositories, budget and analysis are in PLAN-v2.md, committed before any run.

The short version: a graph pays when the model would otherwise keep searching. Haiku on kubernetes and vscode is that case. Sonnet saves a little there and opus nothing, and on a small repository the graph costs sonnet more than plain grep.

The setup

The agent is GitHub Copilot CLI 1.0.90, run headless, one question per session, with two arms:

  • grep: Copilot's view, grep and glob tools and nothing else;
  • graph: the same three tools, plus cbm-lean (my thin wrapper around codebase-memory-mcp 0.9.0) and a routing rule at the top of the prompt: the graph for callers, callees and call chains, grep for exact strings.

Four public repositories, each pinned to a commit:

repository language graph nodes
full-stack-fastapi-template Python, TypeScript 1,601
dub TypeScript 25,023
kubernetes Go 146,421
vscode TypeScript 218,883

Each repository has 16 questions. Eight are structural: who calls a function, what it calls, the call path from a CLI command or an HTTP handler down to a named function, and which functions break if a signature changes. Eight are exact: a config value, where an environment variable is read, the line an error is raised on, the line a type is defined on. I checked every answer key by hand at the pinned commit, and an entry counts only as a whole word, so init_db does not pass for init.

Haiku 4.5 answered all 64 questions three times in each arm, and sonnet 5 twice. Opus 5.5 costs 15 premium requests a prompt in Copilot, so it answered twelve structural questions from kubernetes and vscode once. That is 664 runs and about 744 premium requests.

Cost is tokenEquiv, as in the first note: uncached input, plus 1.25 × cache writes, plus 0.1 × cache reads, summed over the model calls of a run.

Results

The graph's cost as a share of grep's: for each question the graph arm's median over repetitions divided by the grep arm's, then the median over questions, with a 95% bootstrap interval. Below 1 the graph is cheaper.

model repository size structural exact
haiku 4.5 small (1.6k nodes) 0.88 (0.65–1.29) 0.92 (0.72–1.15)
haiku 4.5 medium (25k) 0.61 (0.11–1.50) 0.96 (0.75–1.37)
haiku 4.5 large (146k, 219k) 0.29 (0.11–0.70) 1.01 (0.94–1.28)
sonnet 5 small 1.51 (1.45–1.70) 1.42 (1.20–1.71)
sonnet 5 medium 0.94 (0.57–2.00) 1.14 (1.05–1.30)
sonnet 5 large 0.74 (0.62–0.92) 1.13 (1.01–1.29)
opus 5.5 large 0.97 (0.72–1.10) –

Correct answers on the structural questions, graph against grep:

model small medium large
haiku 4.5 13/24 against 19/24 20/24 against 22/24 38/47 against 33/48
sonnet 5 13/16 against 15/16 16/16 against 16/16 32/32 against 28/32
opus 5.5 – – 12/12 against 12/12

Cost and accuracy together, as tokenEquiv per correct answer on the structural questions:

Tokens per correct answer on structural questions, with the code graph and with grep only, on the FastAPI template and on kubernetes and vscode

What holds up:

  • On kubernetes and vscode, haiku with the graph spent 0.29 of what it spent with grep on structural questions and was cheaper on 15 of 16 of them. It was also right more often, so per correct answer the gap is 47k against 209k tokenEquiv. This is the only result that survives the Holm correction across the 13 tests (p = 0.006).
  • Sonnet on the same questions: 0.74, cheaper on 12 of 16, and 32 of 32 correct against 28 of 32. Opus: 0.97, cheaper on 7 of 12, and every answer right in both arms.
  • On the small repository the graph never paid. Haiku spent about the same and got structural questions wrong more often. Sonnet spent 42 to 51% more.
  • On exact questions the graph made no measurable difference for haiku, and sonnet paid 13 to 42% more for it.
  • The graph's tool schemas and the routing rule add a fixed 1.56k prompt tokens to every model call for haiku and 2.07k for sonnet and opus, at every repository size.

Why the small model gains most

Count the model calls. With grep, haiku needed a median of 22 calls per structural answer on kubernetes and 20 on vscode, and one impact question took 114. With the graph it needed 5 and 4. Every call sends the conversation again, so a long search costs far more than its call count suggests: haiku's most expensive grep answer came to 471k tokenEquiv.

Sonnet's searches with grep were shorter, 8 calls against 4 with the graph. Opus took 4 or 5 calls either way, so the graph had nothing to cut, and its fixed overhead ate what little it saved.

My guess, which I did not test directly: a graph replaces search turns, so it pays when the model would otherwise take many of them, which means a weaker model on a larger repository. A strong model already searches efficiently with grep, and on a small repository nobody needs many turns.

Copilot bills premium requests per prompt, whatever the number of calls, so inside Copilot the graph saves no money. It saves time: haiku's structural answers on the large repositories took a median of 71 s with the graph and 157 s with grep. On the small repository the graph arm was slower.

Where each arm went wrong

  • A field called lines. Asked on which line HorizontalController is defined, the graph arm answered horizontal.go:43 in all five runs, haiku and sonnet alike. The struct starts on line 93 and is 43 lines long. The snippet that codebase-memory-mcp returns carries "start_line": 93, "end_line": 135 and "lines": 43, and the models read the last one as the line number. It happened again with LemonSqueezyClient (client.ts:129 for a class on line 59 that is 129 lines long), and haiku did the same with CursorsController. My first note blamed "line numbers from the graph" for that client.ts:129; this is where it came from. Sonnet with grep got all three right. cbm-lean now passes the field on as line_count.
  • Holes in the graph, trusted. codebase-memory-mcp 0.9.0 records no call edge for a bare call to a function imported with from x import y. Asked what recover_password calls, the graph arm answered "only get_user_by_email" in all five runs, haiku and sonnet alike, straight from the graph and without a single grep. The function also calls generate_password_reset_token, generate_reset_password_email and send_email, and the grep arm found all four every time.
  • A missing dash. Copilot's grep tool shows line numbers only when it gets "-n": true, and its view tool returns a file without them. Sonnet passed -n in 508 of its 571 grep calls and opus in all 53. Haiku passed n, without the dash, 920 times and -n 22 times out of 1,949. The tool ignores the unknown key, so haiku got no numbers, counted lines in the viewed file and missed by a few: for an error raised on line 112 it answered 105, 107, 108 and 113. All but one of haiku's 79 wrong answers to exact questions are a wrong line number.
  • Long searches that still fail. On the call chain from kubeadm token create down to encodeTokenSecretData, haiku with grep took 18 to 22 calls and got the path wrong all three times. It was wrong with the graph too. Sonnet was right in both arms.
  • Haiku tried to call bash three times in the grep arm. Copilot refused, since the tool does not exist there, so no run used a tool outside its arm.

What changed since the first note

The first note said not to expect token savings from a graph with Claude models on repositories up to 25k nodes. That still holds at those sizes. It missed what happens above them: on kubernetes and vscode the smaller models gain, and the smaller the model, the more.

What I would tell a team wiring a graph into an agent

  • Measure on your largest repository with your cheapest model first. That is where a graph earns its place.
  • With a strong model on a small or medium repository, leave it out. You pay for its schemas on every call and get nothing back.
  • Keep routing exact questions to grep. The graph never won them.
  • Do not let the model take line numbers or "this is the complete list" from graph output on trust. Drop or rename fields like lines before they reach the model.

Limits

  • One agent harness (Copilot CLI) and one graph (codebase-memory-mcp 0.9.0 behind my wrapper). Claude Code, another graph or a newer version could give other numbers.
  • Eight questions per cell on the small and medium repositories. With 13 tests under the Holm correction, a cell of eight cannot reach significance even if the graph wins every question, so only the large-repository haiku result is significant.
  • Opus answered twelve questions, once each.
  • tokenEquiv leaves output tokens out. Counting them would widen haiku's gap on the large repositories: its grep answers there produced a median of 7.0k output tokens against 0.9k with the graph.
  • One haiku run failed before Copilot opened a session and is left out. Three sonnet runs failed to fetch Copilot's model list before any model call and were run again. Both are recorded in the plan.

Reproduce

Everything is in code-graph-vs-grep: the plan written before the runs, the questions with the evidence for every answer key, the driver, all 664 runs with their answers, and the analysis.

node bench/run-copilot.mjs --model claude-haiku-4.5 --reps 3 --out results/v2/claude-haiku-4.5.json
python3 bench/stats-v2.py results/v2/*.json
Enter fullscreen mode Exit fullscreen mode

Top comments (0)