Code: Megapixel99/lexindex
lexindex is a code completion engine with no model file, no network and no dependencies: Witten-Bell interpolated n-grams over the repository on disk, blended at a fixed 0.5 with a cache over the buffer being edited. The mechanism is published research and the package does not claim it. CACHECA shipped this architecture inside Eclipse for Java eleven years ago, and the implementation here was validated against the reference family first: 15 of 15 identical top-5 lists across the full range of the blend parameter. The counting is not the story. The story is what it takes to say honestly whether a completion engine is worth installing. Completion accuracy is large everywhere; what varies is whether the tool, rather than something considerably simpler, earned it. This package grew out of the same research line as the completer whose eval never measured its own headline feature and the cache that overtakes a transformer.
npm i -D lexindex
npx lexindex ./src --stats # index, and print the recital rate
node node_modules/lexindex/tools/measure.mjs ./src # the harness behind every number here
The claim is one table, and it is conditional on one number. The recital rate is how often a four-token context in a held-out file was already somewhere in the index, and it ran from 13.5% to 72.9% across the nine JavaScript and TypeScript corpora measured. Against the word list your editor already gives you for free, the advantage holds from 35.9% recital up (0.583 against 0.194, z=3.74). The only null observed is at 13.5%. Against ten lines of frequency counting over real identifiers, the threshold is much higher. That comparison is significant at 61.7% and 71.2% recital, and null at 46.5%, 45.3% and 37.9%. Two baselines give two thresholds, and collapsing them into one band was a mistake an earlier version of the table made. So there are recital ranges where this tool is expected to lose to a ten-line frequency table, and the README says so in its own summary table. npx lexindex ./src --stats prints which band your repository is in rather than letting you discover it. The scoring metric matters as much as the split. The number to read is ident+1char, accuracy after one typed character of an identifier, because aggregate top-1 is mostly punctuation and flatters every engine; identifiers are only 25% to 40% of positions.
What I want to document most is how easy this benchmark is to fool. During development it inflated its own headline through three separate doors, all recorded in the README. A vendored assets/vendor/ holding 344 third-party libraries took 0.567 to 0.809 on the same repository. A fetched research corpus, 1,249 files of 1,266, reached through the .ts and .tsx extensions rather than any directory name, took 54.7 to 80.9. A .claude/worktrees/ directory held 14 whole duplicate copies of the repository. The same concern then arrived a fourth way, through the argument list itself: ./src ./src/, two spellings of one directory nobody would look at twice, double-counted every file, taking the package's own recital from 51.0% to 74.0% and its identifier accuracy from 34.6% to 83.3%, a different verdict entirely. Files are now keyed on their real path so it cannot happen, and both the CLI and the harness say when paths overlapped. Vendored bundles and duplicated checkouts repeat themselves enormously, and repetition is exactly what this tool measures. A completion accuracy without its corpus definition cannot be read.
The multi-language table is where that lesson bit after publication. Seven corpora across six languages all beat both baselines, and one row was published wrong. The falken/service row first read 213 files at 64.2% recital and 78.5% accuracy, from a corpus checked by hand for generated code, and more than half of it (140 files of 266) turned out to be FlatBuffers output the check did not look for. The corrected row is 101 files, 55.3% recital, 72.3% accuracy. The verdict survived the correction and the numbers did not, and both versions are in the README. The tool now scans file heads for generation markers itself, and counts rather than excludes. Across 2,288 files in six languages it flagged 141: the 140 FlatBuffers files, and its own source, which contains the marker strings it searches for.
Getting completions out of it is not left as an exercise. npx lexindex-lsp starts a language server that speaks completion and nothing else, meant to run beside your type-aware server rather than instead of it (the README carries the Neovim, Helix and Emacs stanzas), and lexindex/codemirror and lexindex/monaco ship adapters for the browser editors, over a DocumentSet holding whatever documents the page has open. The engine is also a plain library: buildIndex("./src") and completer.complete(textBeforeCursor). The lexer is one regular expression, so the mechanism is language-agnostic; the default indexes the JavaScript family, and --lang python (or go, rust, java, and ten more) widens it, bringing that language's build directories along as exclusions.
The harness ships in the package, and its design position is that a null is a result. The README shows a run on a small corpus where re-ranking the editor's own candidate list comes back z=0.26, NULL, printed with the exclusions that make the comparison fair. Positions offering a single candidate are dropped, because every ordering is right there. Positions where the truth was never offered are reported as coverage rather than scored, since a re-ranker returns a permutation and cannot invent candidates. When too few positions survive the exclusions, measure exits 2 rather than printing a number. The same code refuses whenever its scorer was never observed producing both a hit and a miss, because a clean result from an instrument that could not fail is worth nothing.
The absences are results too. Seven features were implemented, measured and rejected in the research this derives from, and the README keeps the table so nobody re-adds one without re-running the measurement. Confidence gating lost on both axes three separate times. Recency decay on the cache is a clean negative. Whole-line generation was right about 1 time in 10 mid-line, and a bundled pretrained corpus at 57 times the size of a real repository was worth +0.000 while the buffer cache was warm. Two caveats travel with that last figure, and the README insists on them: the same measurement gives +0.016 once the buffer empties (a poor trade either way), and it was taken in a configuration that included a small transformer this package does not have, so it is evidence about the trade rather than a number measured on this architecture. It is still the load-bearing product decision, and the reason there is nothing to download: the only corpus that pays is the one already on your disk, and indexing one project to complete another cost 0.167 to 0.197 top-1, so shipping a foreign index would be worse than shipping nothing.
Top comments (0)