DEV Community

Ilia Alshanetsky
Ilia Alshanetsky

Posted on Originally published at ilia.ws

CodeSage: Code Intelligence for AI Agents, and Why I Rebuilt Its Models

Until this week, CodeSage ran its two models exactly as they ship on Hugging Face: Jina's code embedder and the ms-marco MiniLM reranker. A full index of php-src took 1,782 seconds, just under half an hour, and I got tired of waiting.

CodeSage 0.38 does the same full index in 373 seconds. Nothing was retrained. The weights are the same weights; what changed is the inference graph around them, a few things in the chunker, and one SQLite query that had been quietly eating most of the run. The two rebuilt models are now public on Hugging Face, along with the scripts that produce them byte-for-byte.

What CodeSage gives an agent

CodeSage is a code intelligence engine for AI coding agents: a structural graph (symbols, references, dependencies) plus semantic search (embedding retrieval with cross-encoder reranking), in one Rust binary that works as a CLI or an MCP server. An agent asks it questions grep can't answer, like "what depends on this class" or "where does session handling happen", and gets file-and-line answers from an index instead of guesses from a context window.

Parsing is tree-sitter, covering PHP, Python, C, C++, Java, Rust, JavaScript, TypeScript and Go. Each project gets one SQLite database at .codesage/index.db, with sqlite-vec for KNN and FTS5 for literal matches. The MCP side is a stdio shim that starts or reuses a per-user daemon over a Unix socket, so concurrent agent sessions share one project cache, one model pool and one CUDA context. There is no HTTP listener.

The MCP server exposes 22 tools. The ones agents lean on most:

  • find_symbol, find_references, trace_call_path: where something is defined, who calls it, and the shortest call chain between two symbols.
  • impact_analysis, assess_risk: what a change to this file can break, scored from the graph plus git history.
  • search: natural-language retrieval over ~50-line chunks, reranked.
  • from_trace: maps a pasted stack trace or ASan report (Python, PHP with Xdebug, Rust, Java, Go, Node, gdb) onto indexed symbols.
  • edit_check: diffs a proposed declaration against HEAD before the agent writes the file.

The graph is only as useful as its recall. In August, impact_analysis reported that AxiosError in axios had 0 dependents. It has 23. File-path imports like ./core/AxiosError.js never resolved; only Rust-style :: paths did. Fixing that, and counting Foo.staticMethod() and x instanceof Foo as usage, took mean dependent recall on axios from 0.37 to 0.66 at precision 0.85. An agent told "0 dependents" by a tool it trusts goes ahead and changes the class.

The index has to keep up with the code

An index that lags the code feeds the agent stale answers with full confidence, so CodeSage re-indexes from git hooks on every commit, checkout and merge. That puts the indexer's fixed cost on the hot path dozens of times a day, which is why a lot of the work across 46 releases since April has gone into cost rather than features.

The per-commit path got there first. In 0.23.0, a no-change incremental run stopped loading the ONNX model and mapping a CUDA context before checking whether any file had changed:

No-change incremental codesage index Time Peak memory
0.22.0 2.7-4.5 s 1.16 GB
0.23.0 0.08 s 150 MB

The file watcher re-embeds only chunks whose text changed: an 8-save burst marks 90 chunks dirty and embeds 1.

That leaves the full index, which nobody runs on every commit but everybody runs eventually: first onboarding, a model change, a chunker change. 0.38 itself forces one, because the embedder now reads up to 1,024 tokens per chunk instead of 512. When the full run takes half an hour on a large repo, people postpone it, and a postponed re-index is a stale index.

Same weights, rebuilt graphs

The embedder was slow because of how it was exported, not because of what it computes. The upstream Jina v2 ONNX file is a decomposed opset-11 export. Its attention uses ALiBi plus a QK LayerNorm, a combination ONNX Runtime's transformer optimizer doesn't recognize, so nothing fuses. Every layer materializes the full [batch, 12, seq, seq] score tensor several times over, and because the graph outputs per-token hidden states, every batch copies a [batch, seq, 768] tensor from the GPU back to the host just so the caller can average it.

IA0x00/jina-embeddings-v2-base-code-codesage carries two graphs:

  • onnx/model.onnx (fp32, every execution provider): masked mean pooling and L2 normalization folded into the graph, so the output is the 768-dimension sentence embedding and nothing else crosses back to the host.
  • onnx/model_cuda_fp16.onnx (CUDA only): the same pooling, plus attention fused by hand into com.microsoft MultiHeadAttention, with ALiBi and the key-padding mask passed as attention_bias. Weights are fp16; the output stays fp32.

One fp16 detail: the attention mask filled padded positions with -FLT_MAX, which overflows fp16. It's -10000.0 now, which is still negative enough to zero those positions after softmax.

Measured on an RTX 4080 Laptop GPU with ONNX Runtime 1.24.4, 2,048 real chunks per corpus of up to 512 tokens each, compared against upstream fp32 with pooling done on the host:

Corpus Upstream fp32 Pooled fp32 Fused fp16 Min cosine (fp16) Top-10 neighbour overlap
php-src (C) 59.7 chunks/s 89 249-257 0.99999 0.994
Laravel app (PHP) 67.2 113 301-317 0.999996 0.996
React app (TS/JS) 69.8 109 296-317 0.999995 0.997

A minimum cosine of 0.99999 across 2,048 chunks means the vectors are, for retrieval purposes, the same vectors. Those throughput figures are at the 512-token cap; CodeSage now runs Jina at 1,024, because real C in php-src runs a median of 611 tokens per 1,500-byte chunk and the old cap truncated more than half of them. So the production per-chunk rate is lower than this table, and the chunks carry more of the code.

The reranker needed far less. IA0x00/ms-marco-MiniLM-L6-v2-codesage is a standard BERT, so ONNX Runtime's stock optimizer fuses it (embedding LayerNorm, six Attention nodes, skip LayerNorm, bias GELU) and the work was mostly converting to fp16. Scoring 50 (query, chunk) candidates went from a median 108.0 ms per query to 20.6 ms, with the same top result on 40 of 40 test queries, top-10 overlap of 0.9975, and a maximum logit change of 0.024.

Both repositories are Apache-2.0, the upstream license, and the weights are unchanged apart from precision. CodeSage pins each one by Hugging Face commit and by the sha256 of every graph and tokenizer it loads, and scripts/derive-jina-onnx.py and scripts/derive-reranker-onnx.py in the CodeSage repo download the pinned upstream files, verify them, and write the published files byte-for-byte. You don't have to trust my upload; you can rebuild it. Your config.toml keeps the upstream model names, and the pin table maps them to the derived repos, so there is nothing to change on upgrade.

The biggest single win wasn't a model

Most of the half hour was SQLite. When a file changes, its old chunks are deleted first, and file_path lived in an auxiliary column of the sqlite-vec table and an UNINDEXED column of the FTS5 table. Neither can be searched by index, so DELETE ... WHERE file_path = ? scanned both tables for every file: about one second per file at php-src's 54,000 rows, roughly 20 of the 28 minutes a full index took when I measured it.

The fix is a small sidecar table, <chunk_table>_paths(id, file_path), kept in step on every insert and delete and rebuilt if it goes missing or diverges. Deletes now resolve the path to rowids and remove by rowid: 3.2 ms per file.

Two chunker changes ship in the same release. Chunks that are mostly digits and commas (lookup tables, timezone and Unicode data) are no longer embedded, since no natural-language query is looking for them, and batches are length-sorted and byte-bounded so a batch of short chunks isn't padded out to the length of its longest member.

Before and after, all of it together:

codesage index --full Before 0.38
php-src 1,782 s 373 s
private mobile backend 469 s 80 s
private frontend app 295 s 60 s

That is the combined effect of the fused fp16 graph, the 1,024-token cap, the dropped data chunks, the batching, and the delete fix. I don't have a clean per-change split beyond the delete scan's own estimate, so I'm not going to invent one. Retrieval held: recall@10 is unchanged on three private eval corpora and rose from 0.765 to 0.790 on a php-src session corpus.

Scores that say what they measured

Fast answers are worthless if the agent can't tell a measurement from a default, so several releases since August went into making CodeSage say what it didn't check. impact_analysis reports a lower bound, because dynamic calls and unresolved imports can hide real dependents. assess_risk returns unscored for a file with no indexed git history instead of a low number. find_references marks a count as a floor when the symbol name is ambiguous. An empty search result carries indexed-file counts per language, so "not found" and "never indexed" stop looking identical.

The same rule applies to my own benchmarks. In August I withdrew a published retrieval number (recall@10 0.932, NDCG@10 0.788 over 602 queries) after finding the harness had indexed whole repos instead of each benchmark's root and scored crashed queries as 0.000. The replacement table in the README runs 663 queries across 33 repos in 9 languages, reports per-language NDCG@10 from 0.7455 (C) to 0.9434 (JavaScript), and says it was measured on a modified 0.27.0 build. It hasn't been re-run on 0.38 yet.

Where it falls short

These are the limits I know about, all of them in the README:

  • Nine languages. On the public benchmark, 47% of queries are in languages CodeSage doesn't parse (Ruby, Kotlin, Swift, Scala) and get no recall at all.
  • Agents don't reach for it on their own. With CodeSage and Grep both available in Claude Code, agents picked CodeSage on code-identifier queries 1.1% of the time over 30 days of sessions, and 0 of 10 on a controlled harness (measured in April, not re-measured since). That is why the codesage-tools Claude Code plugin serves a codesage brief inline before an Edit or Write in an onboarded project instead of waiting to be asked. It stays silent when there is nothing worth saying.
  • Every speed number in this post is CUDA. CPU and Apple's CoreML get the pooled fp32 graph but not the fused fp16 one, and I haven't measured the CPU side.
  • No prebuilt binaries and no documented Windows support. You build it with cargo.
  • Read-only. It doesn't edit code and doesn't query across repositories.

Try it

cargo build --release -p codesage --features cuda   # drop --features cuda for CPU / CoreML
cd /path/to/your/project
codesage init
codesage index
codesage install-hooks
claude mcp add --scope user codesage -- codesage mcp
Enter fullscreen mode Exit fullscreen mode

Codex users run codesage install codex --global instead of the last line. If you're on Claude Code, whetstone is the other half: CodeSage tells the agent what the code is, whetstone covers how to investigate, review and ship. The first index downloads the models: about 370 MB on CUDA, about 736 MB for the fp32 graphs on CPU or CoreML.

Nothing about CodeSage 0.38 required a better model. It required looking at what the existing one did between input and output, and at what the database did on every file delete. Most of that half hour was overhead around the model, not the model itself.

github.com/iliaal/codesage

Top comments (0)