When a coding agent needs something from a file that is too big for its context window, it usually does one of two things. It reads the file in slices and hopes the answer is in one of them, or someone has already built a RAG index with chunks and embeddings. Both lose information at the seams.
Matryoshka takes a third route. It is an open-source TypeScript project by Dmitri Sotnikov (yogthos, the author of the Luminus Clojure framework). It keeps the whole document on the server and gives the agent a small query language to work on it. Results stay server-side, and the agent gets back short "handles" instead of the raw text. The README says it is based on the Recursive Language Models paper by Alex L. Zhang, Tim Kraska and Omar Khattab of MIT CSAIL. Matryoshka is an independent implementation of that idea, not the authors' own code.
When I checked on 4 October 2026, the repo had 149 stars, 20 forks and no open issues. It is Apache-2.0 licensed, and the latest npm release was matryoshka-rlm 0.2.40 from 17 May 2026. It's a small project, so I wanted to see how it holds up on real files.
One constraint shaped the whole test: there was no LLM API key on my machine and no local Ollama. Matryoshka ships an MCP server, lattice-mcp, that needs no LLM at all. So I played the agent myself and wrote every query by hand. This tests the engine and the tool interface. It says nothing about how well a model would drive them.
How it works, in one paragraph
The model doesn't write JavaScript or Python. It writes commands in Nucleus, a small S-expression language: (grep "ERROR"), (filter RESULTS ...), (count RESULTS), (sum RESULTS). The Lattice engine parses and type-checks each command and runs it against the document. Results go into an in-memory SQLite store, and the client gets a stub like $grep_error: Array(1000) [preview...]. It only pulls real lines with lattice_expand when it needs them. The README claims this gives "97%+ token savings", and "80%+ token savings compared to reading files directly" for the MCP server. Those are the project's numbers, and I didn't benchmark them.
Setup
The box was Linux x86_64 with 8 vCPUs, about 15 GB of RAM, Node 20.19.2 and pnpm 10.33.4. The README suggests pnpm add -g matryoshka-rlm. I installed it locally in a small Node/TS project instead, next to the official MCP SDK:
pnpm add matryoshka-rlm@0.2.40 @modelcontextprotocol/sdk tsx typescript @types/node
It finished in 2.5 seconds and took 220 MB of node_modules. pnpm 10 skipped the native build scripts for the tree-sitter grammars. I left them unbuilt, and symbol listing still worked.
Gotcha 1: don't run plain npx lattice-mcp. From a folder where Matryoshka isn't installed, npx lattice-mcp downloaded an unrelated npm package called lattice-mcp (1.6.2), which asked for LATTICE_API_URL and LATTICE_API_TOKEN. Use the binary that comes with matryoshka-rlm (node_modules/.bin/lattice-mcp, or npx -p matryoshka-rlm lattice-mcp). The real one prints lattice-mcp v0.2.40.
The client is a few lines of TypeScript with the MCP SDK's stdio transport:
import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";
const transport = new StdioClientTransport({
command: "/workspace/matryoshka-test/client/node_modules/.bin/lattice-mcp",
cwd: "/workspace/matryoshka-test", // lattice_load only accepts files under this folder
});
const client = new Client({ name: "ishank-test", version: "0.0.1" });
await client.connect(transport);
await client.callTool({ name: "lattice_load", arguments: { filePath: "docs/war-and-peace.txt" } });
const r = await client.callTool({ name: "lattice_query", arguments: { command: '(grep "Natásha")' } });
It connected in 409 ms and reported the server as lattice 0.2.40, with 12 tools: lattice_load, lattice_query, lattice_expand, lattice_close, lattice_status, lattice_bindings, lattice_reset, lattice_memo, lattice_memo_delete, lattice_help, lattice_llm_respond and lattice_llm_batch_respond.
The test documents
- War and Peace from Project Gutenberg (Maude translation): 3.36 MB, 66,041 lines, 566,333 words, about 770,000 tokens (769,825 with tiktoken's o200k_base). That is far more than most models' context windows.
- A real Apache access log from Elastic's examples repo: 10,000 lines (2.37 MB) of traffic from May 2015.
- Matryoshka's own source files, for the code features.
Before asking Matryoshka anything, I worked out the true answers separately with grep and a short Python script.
War and Peace: facts, counts and quotes
Loading the novel took about 0.75–0.8 seconds. Every query after that took between 1 and 430 ms.
| What I asked | Nucleus command | Matryoshka | Ground truth |
|---|---|---|---|
| Chapter headings |
(grep "^CHAPTER [IVXLC]+$") → (count RESULTS)
|
365 | 365 ✅ |
| Book headings |
(grep "^BOOK [A-Z]+") → count |
17 | 15 ⌠|
| Mentions of Natásha |
(grep "Natásha") → count |
1,213 | 1,213 matches ✅ (on 1,197 lines) |
| Pierre lines that also mention Natásha | (filter RESULTS (lambda x (match x "Natásha" 0))) |
44 | 44 ✅ |
| Where is "Genoa"? |
(grep "Genoa") → expand |
lines 840, 1573, 9153 | same ✅ |
| Quote the opening |
(lines 840 846) → expand |
exact 7-line "Well, Prince, so Genoa and Lucca…" | ✅ |
| Mentions of Borodinó |
(grep "Borodinó") → count |
108 | 108 ✅ |
Gotcha 2: grep is case-insensitive and counts matches, not lines. The 17 books threw me until I expanded the handle. Two of the hits were ordinary sentences starting with a lowercase "book" ("book from the high desk."). The source confirms that grep runs with the gmi flags. It also returns one item per match, so 1,213 is the number of times "Natásha" appears, not the 1,197 lines she appears on. Once I knew that, every count matched grep -oi exactly. A tighter pattern, (grep "^BOOK [A-Z]+: "), returned the correct 15.
The ranked searches showed where the "no RAG" approach has its limits:
-
(bm25 "battle of Borodino" 5)returned five lines that just said "battle." The translation spells it Borodinó. With the accent,(bm25 "Borodinó battle" 5)returned five lines that all contain "battle of Borodinó". -
(fuzzy_search "Natasha Rostova" 3)found nothing.(fuzzy_search "Natásha Rostóva" 3)found her straight away. -
(semantic "wounded at Austerlitz looking at the sky" 5)didn't return the famous lines where Prince Andrew lies wounded under "the lofty sky". Its closest hit was "been looking at." on line 15862, seven lines before. Plain(grep "lofty sky")found the passage at line 15869. The README describessemanticas TF-IDF cosine similarity, not embeddings, so this isn't surprising.
Gotcha 3: spell it the way the document does. Lexical ranking does not fold accents, so with translated or non-English text an agent should check the document's spelling before trusting a ranked search.
Gotcha 4: lattice_expand needs a named handle. RESULTS works inside queries, but lattice_expand with RESULTS returns "Invalid handle: RESULTS". lattice_bindings lists the real names, such as $grep_genoa. Patterns full of quotes get generic names ($grep, $grep_2), so lattice_bindings is worth calling often.
Two smaller things worked as documented. Loading /etc/hostname was refused with "Path outside working directory is not allowed." And a memo I saved with lattice_memo while the novel was loaded could still be expanded after I switched to the log file.
How much text actually reached my "agent"? My first session made 27 tool calls against the 3.36 MB novel, and the tool responses added up to 5,625 characters. That's my own count, not a benchmark. A real agent wouldn't read the whole book either, but it shows the design: the document stays on the server and only stubs and the lines you ask for come back.
The access log: counting and summing
| What I asked | Matryoshka | Ground truth |
|---|---|---|
404 responses, (grep "\" 404 ") → count |
213 | 213 ✅ |
| Googlebot requests | 543 | 543 ✅ |
| 200 responses | 9,126 | 9,126 ✅ |
Bytes served with 200, (sum RESULTS) on the grep result |
1,067,060.464 | ⌠meaningless |
Same, after (map RESULTS (lambda x (match x "\" 200 (\\d+)" 1)))
|
2,735,455,845 | 2,735,455,845 ✅ |
Gotcha 5: sum on raw lines adds up the wrong number. When you call sum on grep results, it takes the first number on each line. In an access log, that's the start of the client IP (83.149…), so you get a confident, meaningless total. Pull out the field first with map and match, and sum returns the exact byte count.
I also tried the sub-LLM path without an LLM. My client doesn't support MCP sampling, so (llm_batch RESULTS (lambda x (llm_query "Is this request from a bot or a human? {item}" (item x) (one_of "bot" "human")))) over the three 500-error lines returned a single [LLM_BATCH_REQUEST id=… count=3] message. It held three ready-made prompts, each with its log line and an "answer with exactly one" instruction. I typed the answers myself (["bot","bot","human"]) and sent them back with lattice_llm_batch_respond. Matryoshka stored them as a new handle, and filtering on "bot" counted 2. So the protocol works with any MCP client. To be clear, no model judged those lines. I did.
Code files: symbols yes, graph no
On Matryoshka's own lc-solver.ts (2,658 lines), (list_symbols "function") returned 20 functions, the same count grep gives for top-level function declarations. Each came with its line range and signature. (get_symbol_body "evaluate") returned the function's source, but inline and 64,311 characters long. Unlike search results, it isn't returned as a handle, so a big function goes straight into the agent's context.
The knowledge graph didn't work for me. (callers ...) and (god_nodes ...) returned "No symbol graph available" on all four source files I tried. Each time the server's stderr showed Graph.addEdge: source & target are the same (...), with names like evaluate, init and run. The graph library rejects self-loops, so a recursive function or a same-name call seems to stop the graph from being built. That's on 0.2.40 with these four files. I can't say how widespread it is.
The REPL and the full RLM CLI
lattice-repl docs/war-and-peace.txt is useful for trying commands out: (grep "Austerlitz") then (count RESULTS) gave 51, matching grep -oi. Piping :load and a query in together failed, because the load hadn't finished when the query arrived. Passing the file as an argument worked.
The rlm CLI is the part where the LLM writes the Nucleus itself. With no config, it defaults to Ollama on localhost, retried three times, and stopped after about 9 seconds with [aborted: llm 3 consecutive LLM call failures: fetch failed]. That's expected. It needs a model.
What I didn't test
- A real agent choosing Nucleus commands. That is the main promise, and I wrote every command by hand.
-
rlm/rlm-mcpwith a model, recursiverlm_query/rlm_batch, MCP sampling, compaction and resource limits. -
lattice-setup(Claude Code wiring), the HTTP and pipe adapters, program synthesis,fuse/dampen/rerank, and multi-document loading. - Any of the README's token-savings figures or the paper's benchmark results.
Takeaways
- The engine was exact on literal tasks. Every count, lookup and quote matched ground truth once I understood the semantics. Most of the work happens on the server, and only stubs come back.
-
The semantics matter more than the syntax. Case-insensitive, per-match
grepand first-numbersumare reasonable defaults, but an agent that doesn't know them will report 17 books and a nonsense byte total with full confidence. - The ranked searches are lexical. BM25, fuzzy and "semantic" (TF-IDF) depend on exact spelling. Don't expect embedding-style recall.
- The code graph is the weak spot in 0.2.40. Symbol listing worked, but the call graph failed on every file I tried.
- It's worth a try if your agent works over big logs, transcripts or books and you want answers you can check, not chunk retrieval. Next I want to put a real model in the loop and see whether it avoids these gotchas on its own.
If you've wired lattice-mcp into Claude Code or another agent, I'd like to hear whether your model figured out the case-insensitive grep by itself.
Originally published on Medium.
Top comments (0)