DEV Community

Gonza Roman
Gonza Roman

Posted on

Long-term memory for coding agents that costs the same whether it's 50 notes or 50,000

This is my first post; I hope you like it.

Every new session with Claude Code or opencode started from scratch. I had to explain all over again how my machine was configured, what decisions we’d already made and why, and the changes I’d made the day before. I tried keeping a large notes file and pasting it in, but it just kept growing; it cost me tokens on every turn, and most of it had nothing to do with the issue at hand.

So I created the IHMT (Infinite Hierarchical Memory Tree): a local MCP server that provides programming agents with long-term memory, stored as a tree of plain-text files on disk.

Infinite Hierarchical Memory Tree

How It Works in Practice

Yesterday I worked with Claude Code on a small app of mine and added a new mode to it. Today I opened a new session, with no context, and asked, “What about the mafia mode?” The agent searched its memory and responded with the details: what that mode does, which files implement it, and a change I’d requested later that same day.

It also caught a mistake. One note said the feature had shipped as version 1.3. A later note from the same day said I had decided to fold everything into 1.2 instead. The search surfaced both, and the agent told me about the correction instead of repeating the old version number.

That is the whole idea: tell your agent something once, and it remembers it in every future session, in any project.

How it works

The memory is a folder of plain files:

  • Leaves are .txt files with your text exactly as saved.
  • Branches are JSON files that summarize the nodes below them.
  • The trunk, root.json, is the index every search starts from.

A search doesn't read everything. It walks down from the trunk to the few leaves that match, so the cost grows with the depth of the tree, not with how much you have stored. In my measurements, 200 notes meant opening 9 branches, and 5,000 notes meant 14. A typical answer is 200 to 900 tokens and takes about 13 ms.

(...Instead of embedding everything into one flat index and scanning it, IHMT organizes knowledge into a tree: raw text leaves at the bottom, recursive JSON summaries above them, and a single root.json trunk at the top. A query walks that tree — root → branch → branch → leaf — so the number of files opened grows with the depth of the tree (≈ beam × log_B(n)), not with the amount stored...)
number of files opened grows with the depth of the tree

A few things I cared about:

  • Time. When a fact changes ("I moved to Valencia"), the old one is kept as history. If it comes up later, it is flagged OUTDATED so the agent doesn't act on it.
  • Ambiguity. If a question matches several memories equally well ("Luis", but which one?), it returns AMBIGUOUS and the agent asks you instead of guessing.
  • One memory for all your agents. Claude Code, Codex and opencode can share the same memory: what one saves, the others find.
  • Local and readable. No cloud, no database, no account. Every memory is a text file you can open, back up or put under git.

There are also a few tools for code projects: a project map, finding a single method, and re-reading a file as a diff. The code splitter never cuts a class or method in half (Python goes through ast, C-family languages through a brace scanner that skips comments and strings). And there is a small browser interface (python3 gui.py, standard library only) to explore the tree and see why a search returned what it did.

Where it doesn't help

I'd rather say this up front:

  • With 10 notes it saves nothing. They fit in the context anyway. The value shows up when your memory is bigger than what you would want to paste every session.
  • I ran an A/B test with 22 headless Claude Code sessions. With IHMT available, Claude didn't use the project tools at all: it preferred grep and sed, which are already cheap. Forcing it to use them made sessions 42 to 71 % more expensive. So the project tools are optional, and the real value is memory between sessions.
  • Having the server enabled but unused costs about 460 tokens per conversation.
  • Bulk-loading thousands of documents is slow (5,000 in about 107 s). Saving one note at a time, which is what the agent does, is instant.

usage optimization

Trying it

You need Python 3.10+ and git. Paste this into your agent:

Install the IHMT memory MCP server for me from https://github.com/gonzaroman/IHMT-MEMORY — follow the instructions in its INSTALL.md.

The agent checks the requirements, downloads the code, registers the MCP server, adds the usage instructions to its instruction file (CLAUDE.md, AGENTS.md...) and checks that everything works. It is tested on macOS and Linux with Claude Code, Codex and opencode. Windows instructions are included but not tested yet.

It's MIT licensed: https://github.com/gonzaroman/IHMT-MEMORY

I'd love feedback, especially from people whose setup is different from mine. If something breaks or doesn't make sense, open an issue or tell me in the comments.

Top comments (0)