DEV Community

cos white
cos white

Posted on

Surgical Subgraphs: How We Cut Coding-Agent Token Costs by 95%

Every coding agent has the same expensive habit: when it needs context, it grabs too much of it.
A full-file dump here, a 50-chunk RAG retrieval there — and suddenly a simple "trace this API call"
task burns 12,000+ tokens before the model writes a single line of code. At Claude 3.5 Sonnet pricing
($3.00 / 1M input tokens, Sept 2026), that's $38.09 per 1,000 tasks. For a team running agents all day,
the context bill becomes the real bill.

We built LKIO (open-source, Apache 2.0) to fix exactly this: a local repository-intelligence engine
that hands agents the minimal sufficient subgraph instead of a pile of text.

Why the obvious fixes fail

We benchmarked three approaches on a real 4,899-file business codebase
(Vue 3 frontend + Spring Boot microservices + an enterprise dashboard — not a toy repo):

Naive full-file dump Chunk RAG (top-k=10) LKIO surgical subgraph
Avg task tokens 12,698 4,266 545
P95 tokens 24,012 5,000 590
Cost / 1k tasks $38.09 $12.80 $1.64
Cross-stack link recall (n=12) — 0/12 12/12
Hop precision, zero spurious hops (72 hops) — — 72/72

Chunk RAG looks reasonable until you ask a question that requires a join:
"trace this Vue form submission through the API client, the REST route, the Spring controller,
the service, the DTO, down to the DB table." Vector search doesn't do joins — it scored 0/12
on our 12 end-to-end cross-stack links (Wilson 95% CI: [75.8%, 100.0%]).
LKIO got 12/12 with zero spurious hops, because it doesn't retrieve text. It queries a graph.

The core idea: the repo is a database, not a document pile

LKIO continuously parses your codebase with Tree-sitter into fine-grained symbols
(classes, methods, interfaces, Vue SFC blocks) and the relations between them
(calls, imports, DTO field lineage, REST route mappings). That lives in a
copy-on-write in-memory snapshot — readers never block writers, so an agent can query
mid-index with zero stalls. A sliding-window retention policy (last 50 snapshots)
keeps memory bounded: +11.62MB over 1,000 update cycles, no slow leak.

When an agent asks a question, LKIO doesn't return chunks. It runs a bounded,
cycle-safe BFS from the anchor symbols and returns the smallest subgraph that answers
the question — typically a few hundred tokens. Signal-to-noise ratio goes from ~3%
to ~27%.

It plugs into any agent via MCP (9 read-only tools, stdio transport):
repo_overview, list_entities, search_knowledge, analyze_impact,
evaluate_decision, get_timeline, and more. Read-only by design —
the MCP layer physically cannot mutate your code.

Trust, not just tokens

Two things matter more than the token math for production use:

Calibrated confidence. On 120 real decision samples, temperature scaling cut
expected calibration error from 0.1850 to 0.0469 (−74.6%, Brier 0.0583).
When LKIO says 0.90 confidence, empirical accuracy is ~91%.
Confidence you can actually use as a probability.

Governance that doesn't slow you down. The rule is simple: confidence ≠ permission.
Even at 0.999 confidence, changes to payment cores or auth middleware require human sign-off.
We adversarially tested 8 attack scenarios (regression smuggling, NaN confidences,
empty evidence chains…) — 8/8 blocked. Then we tested the thing teams actually fear:
32 everyday benign changes (renames, extract-function, comment edits, CSS tweaks, i18n fixes) —
32/32 passed, zero false blocks (false-block rate bounded ≤10.7% at 95% confidence).
A gate developers hate is a gate developers disable.

It runs on a laptop

Local-first was a design goal, not an afterthought (measured on Intel Core Ultra 9 275HX,
32GB DDR5, Windows 11, Python 3.12.10):

  • Cold start: 1,000 files in 10.61s (94.3 files/s, 128.9MB peak RSS)
  • Incremental: file save → queryable in 56.4ms (incl. 50ms fs-debounce)
  • Query latency: symbol lookup ~2μs, depth-2 impact analysis P50 0.121ms
  • Embedding: all-MiniLM-L6-v2 (384-dim, ~90MB) — no GPU required

Honest limitations

This is a young project (started September 2026, small team).
Our benchmarks are self-reported on our own codebase — the full methodology,
Wilson confidence intervals, and machine-readable results are in
docs/benchmarks/production_acceptance_rigorous_report.md, and we'd love
independent reproduction.

The last mile: dogfooding

All figures above come from rigorous synthetic benchmarks on a real codebase.
But between "admitted to production" and "trusted" there is always one last mile:
a real dev team committing at high frequency for a month.

We hold the invariant Implementation Complete ≠ Benchmark Validated ≠ Production Gate Passed,
and we will not dress benchmark scores up as production proof. We are currently running
a 2-week dogfooding program with 1–2 engineers, tracking daily false-block experience
and real agent token spend. The first field report will be appended here when it lands.

Try it

git clone https://github.com/xxmasy/lkio.git && cd lkio
uv sync && uv run lkio verify --all   # 11 verification gates, ~2s, all green
Enter fullscreen mode Exit fullscreen mode

Star it, break it, file issues. Contact: 812866097@qq.com

Roadmap: SQLite-backed zero-dependency mode, more languages, and a hosted option
for teams that don't want to self-host.

GitHub: https://github.com/xxmasy/lkio

Top comments (0)