Every coding agent has the same expensive habit: when it needs context, it grabs too much of it.
A full-file dump here, a 50-chunk RAG retrieval there — and suddenly a simple "trace this API call"
task burns 12,000+ tokens before the model writes a single line of code. At Claude 3.5 Sonnet pricing
($3.00 / 1M input tokens, Sept 2026), that's $38.09 per 1,000 tasks. For a team running agents all day,
the context bill becomes the real bill.
We built LKIO (open-source, Apache 2.0) to fix exactly this: a local repository-intelligence engine
that hands agents the minimal sufficient subgraph instead of a pile of text.
Why the obvious fixes fail
We benchmarked three approaches on a real 4,899-file business codebase
(Vue 3 frontend + Spring Boot microservices + an enterprise dashboard — not a toy repo):
| Naive full-file dump | Chunk RAG (top-k=10) | LKIO surgical subgraph | |
|---|---|---|---|
| Avg task tokens | 12,698 | 4,266 | 545 |
| P95 tokens | 24,012 | 5,000 | 590 |
| Cost / 1k tasks | $38.09 | $12.80 | $1.64 |
| Cross-stack link recall (n=12) | — | 0/12 | 12/12 |
| Hop precision, zero spurious hops (72 hops) | — | — | 72/72 |
Chunk RAG looks reasonable until you ask a question that requires a join:
"trace this Vue form submission through the API client, the REST route, the Spring controller,
the service, the DTO, down to the DB table." Vector search doesn't do joins — it scored 0/12
on our 12 end-to-end cross-stack links (Wilson 95% CI: [75.8%, 100.0%]).
LKIO got 12/12 with zero spurious hops, because it doesn't retrieve text. It queries a graph.
The core idea: the repo is a database, not a document pile
LKIO continuously parses your codebase with Tree-sitter into fine-grained symbols
(classes, methods, interfaces, Vue SFC blocks) and the relations between them
(calls, imports, DTO field lineage, REST route mappings). That lives in a
copy-on-write in-memory snapshot — readers never block writers, so an agent can query
mid-index with zero stalls. A sliding-window retention policy (last 50 snapshots)
keeps memory bounded: +11.62MB over 1,000 update cycles, no slow leak.
When an agent asks a question, LKIO doesn't return chunks. It runs a bounded,
cycle-safe BFS from the anchor symbols and returns the smallest subgraph that answers
the question — typically a few hundred tokens. Signal-to-noise ratio goes from ~3%
to ~27%.
It plugs into any agent via MCP (9 read-only tools, stdio transport):
repo_overview, list_entities, search_knowledge, analyze_impact,
evaluate_decision, get_timeline, and more. Read-only by design —
the MCP layer physically cannot mutate your code.
Trust, not just tokens
Two things matter more than the token math for production use:
Calibrated confidence. On 120 real decision samples, temperature scaling cut
expected calibration error from 0.1850 to 0.0469 (−74.6%, Brier 0.0583).
When LKIO says 0.90 confidence, empirical accuracy is ~91%.
Confidence you can actually use as a probability.
Governance that doesn't slow you down. The rule is simple: confidence ≠ permission.
Even at 0.999 confidence, changes to payment cores or auth middleware require human sign-off.
We adversarially tested 8 attack scenarios (regression smuggling, NaN confidences,
empty evidence chains…) — 8/8 blocked. Then we tested the thing teams actually fear:
32 everyday benign changes (renames, extract-function, comment edits, CSS tweaks, i18n fixes) —
32/32 passed, zero false blocks (false-block rate bounded ≤10.7% at 95% confidence).
A gate developers hate is a gate developers disable.
It runs on a laptop
Local-first was a design goal, not an afterthought (measured on Intel Core Ultra 9 275HX,
32GB DDR5, Windows 11, Python 3.12.10):
- Cold start: 1,000 files in 10.61s (94.3 files/s, 128.9MB peak RSS)
- Incremental: file save → queryable in 56.4ms (incl. 50ms fs-debounce)
- Query latency: symbol lookup ~2μs, depth-2 impact analysis P50 0.121ms
- Embedding: all-MiniLM-L6-v2 (384-dim, ~90MB) — no GPU required
Honest limitations
This is a young project (started September 2026, small team).
Our benchmarks are self-reported on our own codebase — the full methodology,
Wilson confidence intervals, and machine-readable results are in
docs/benchmarks/production_acceptance_rigorous_report.md, and we'd love
independent reproduction.
The last mile: dogfooding
All figures above come from rigorous synthetic benchmarks on a real codebase.
But between "admitted to production" and "trusted" there is always one last mile:
a real dev team committing at high frequency for a month.
We hold the invariant Implementation Complete ≠ Benchmark Validated ≠ Production Gate Passed,
and we will not dress benchmark scores up as production proof. We are currently running
a 2-week dogfooding program with 1–2 engineers, tracking daily false-block experience
and real agent token spend. The first field report will be appended here when it lands.
Try it
git clone https://github.com/xxmasy/lkio.git && cd lkio
uv sync && uv run lkio verify --all # 11 verification gates, ~2s, all green
Star it, break it, file issues. Contact: 812866097@qq.com
Roadmap: SQLite-backed zero-dependency mode, more languages, and a hosted option
for teams that don't want to self-host.
GitHub: https://github.com/xxmasy/lkio
Top comments (0)