DEV Community

Avinash Jetwani
Avinash Jetwani

Posted on AI-assisted

I tested whether Claude Code follows decisions from earlier chats

I got tired of re-explaining my project to Claude Code every session.

We'd decide something in one chat. "We use Postgres." "Don't touch the public API." "That migration approach failed, and here's why." A few sessions later, I'd be typing it again.

CLAUDE.md fixes this if you keep it updated. I didn't.

So I built jevmem.

What it does

  • You decide something in a Claude Code chat.
  • After each message, Jev (TypeSafe AI's decision model) checks if it's worth keeping: a decision, a rule, a bug, an approach that failed.
  • If it is, jevmem writes one line to JEVMEM.md in your repo.
  • Next session, the lines that matter for your prompt go back to Claude.
  • Change your mind, and the old line is crossed out.

Your team gets the same file through git.

Does it work? I measured it.

24 tasks in 3 small projects, 3 runs each. Real Claude Code sessions with claude-sonnet-5. Each task has a right answer that depends on a project decision saved earlier.

Three setups: no project memory, jevmem, and the same lines written by hand into CLAUDE.md.

No memory jevmem CLAUDE.md
Followed the project's decision 28/72 66/72 67/72
Tried a change the project forbids 10/18 0/18 not run
Repeated an approach that had already failed 3/15 0/15 0/15

So jevmem is about as good as a hand-written CLAUDE.md. The difference is that you don't write it.

Where CLAUDE.md won: a convention nothing in the prompt points to, a rule for every user-facing string. CLAUDE.md got it 3 of 3 times, jevmem 0 of 3. Rules every task must follow still belong in CLAUDE.md.

What to know

  • It needs a TypeSafe API key.
  • Message text goes to TypeSafe to be scored. Common secrets are scrubbed first. No telemetry.
  • Automatic in Claude Code. Codex and Cursor use the same file over MCP.
  • These are my tasks and my projects. The scripts and results are in the repo, so you can check them.

Try it

  1. Claude app → Plugins → Discover → jevmem → Add
  2. npm i -g jevmem
  3. jevmem key, then jevmem enable in your project

Open source, MIT. Built on TypeSafe AI's Jev.

Repo: https://github.com/Avinash-jetwani/jevmem · Full results: https://avinash-jetwani.github.io/jevmem/results/

I'd love to hear where it saves the wrong thing.

Written with help from Claude. The numbers are from my own test runs, and the results files are in the repo.

Top comments (1)

Collapse
 
muhannad_salkini_1036ea1b profile image
Muhannad Salkini •

Nice to see actual numbers instead of vibes, and thanks for publishing the scripts.

The one CLAUDE.md won (a convention nothing in the prompt points to) looks like a retrieval problem more than a memory one. If entries are retrieved by similarity to the prompt, a rule like "every user-facing string goes through i18n" never matches "add a settings page". What's worked better for us is attaching the paths each decision applies to, and pulling decisions by the files the agent is about to touch, not just by the prompt. The settings page touches src/ui/**, so the i18n rule comes along even though the prompt never mentions strings.

Two things I'd be curious about in a follow-up run:

  • Superseded decisions. When you change your mind, does the crossed-out line ever leak back in? An agent turning an old rule into a "don't do X" constraint is a failure mode I've seen a lot.
  • Decisions made outside the chat. A lot of a team's real decisions live in PR review comments ("we don't do it this way because…"). That's a different source from what the agent hears in a session, and probably where team-level value is.