DEV Community

Statewave
Statewave

Posted on Originally published at statewave.ai

Four of six open-source AI memory tools hold features back from the free build

Four of six open-source AI memory tools hold features back from the free build
An Apache-2.0 badge does not tell you whether the feature you are switching for is in the free build. I counted 14 capabilities on Mem0's own Platform vs Open Source page that self-hosted users do not get, then went through five alternatives that come up in the same conversation, one vendor page at a time. Four of the six hold something back.

We build Statewave, one of the six, so weigh that entry accordingly. Every claim below links to the vendor's own page, our benchmark fork is public, and the section on our own tool ends with the things it does badly.

Mem0 is the default for good reasons

Worth saying before the criticism, because most posts in this genre skip it. Mem0 takes three lines of code to add, ships Python and JavaScript SDKs, and offers a hosted API. It raised $24M in October 2025 and reported API calls going from 35 million in Q1 to 186 million in Q3 2025. CrewAI, Flowise and Langflow use it natively, AWS picked it as the exclusive memory provider for its Agent SDK, and the repo is past 65,000 stars. For personalizing a chat assistant, that is hard to beat.

Trouble starts when you self-host, need an audit trail, or run several agents against one memory.

Where the self-hosted build stops

The OSS build is a different product. Mem0's own comparison page lists graph memory, memory decay, temporal reasoning, Dream background consolidation, webhooks, memory export, batch updates and deletes, feedback signals, summaries, custom categories, app-level scoping, orgs and projects, the project-wide event log and the web dashboard as Platform-only. Counted up, that is 14 capabilities, as of September 2026. Mem0's page also notes that the OSS SDK returns a not-supported error if you pass it the decay or temporal parameters.

Graph memory is the clearest case. Moving the open-source SDK to the v3 pipeline deleted roughly 4,000 lines of external graph store drivers covering Neo4j, Memgraph, Kuzu, Apache AGE and Neptune. Its migration guide tells OSS users who need graph memory to move to the Platform. Mem0's README carries its own caveat: benchmark scores reflect the managed platform, "which includes proprietary optimizations not available in the open-source SDK."

Graph memory and real retrieval volume start at $249 a month. Mem0's plans as of September 2026: Hobby free at 1,000 retrievals a month, Starter $19 at 5,000, Pro $249 at 50,000. Graph memory and Dream consolidation begin at Pro. Retrieval limits bite well before storage limits do. A support agent handling 300 conversations a day with three lookups per conversation is around 27,000 retrievals a month, more than five times the Starter allowance.

Audit logs, on-prem and SSO are Enterprise-only. Self-hosting does not fill that gap. Its OSS edition keeps a history for each individual memory and no project-wide log of adds, searches and deletes. It scopes memory by user_id, agent_id and run_id, with no app_id for separating tenants.

You cannot reproduce what the agent saw. Search returns whichever memories rank highest at the moment you query, and the OSS build does not store the exact set a given response was built from. When a customer asks why the agent told them their refund was approved, you are reading logs instead of looking up a record.

The edition line, across six tools

Here is what each vendor's own documentation says, as of September 2026.

  • Mem0: 14 Platform-only capabilities by our count, graph memory removed from OSS.
  • Zep: stopped maintaining Community Edition and moved its open-source work to Graphiti, which is a library rather than a memory service.
  • Cognee: running the whole memory layer on Postgres ships as a demo feature. The production-ready version is a licensed product.
  • Supermemory: the local build runs as one process on one machine with a single API key. Proprietary extraction models, multi-member orgs and connectors are Enterprise.
  • Letta and LangMem ship their memory engines in full.
  • Statewave has no paid edition. One question settles most of this: is the feature I am leaving Mem0 for included in the free build today, under the same license? Check the vendor's own edition or pricing page rather than a listicle. This one included.
Tool Memory model You run Audit in the free build
Mem0 OSS (Apache-2.0) Extracted facts + vector search Vector store, LLM, embedder Per-memory history only
Statewave (Apache-2.0) Typed memories compiled from events Postgres + pgvector Source links, signed receipts, policies
Graphiti (Apache-2.0) Temporal knowledge graph Neo4j, FalkorDB or Neptune Historical graph edges
Letta (Apache-2.0) Agent-edited memory blocks Letta App Server Git-tracked memory
Cognee (Apache-2.0) Graph + vector over docs and code Pluggable stores Inspectable evidence
LangMem (MIT) Memory tools + background manager LangGraph store None built in
Supermemory local (MIT) Memory + RAG engine One local process Server logs

The five alternatives, and the trade-off each one carries

Graphiti builds a knowledge graph that records when each relationship was true, so an agent can answer how a fact changed. In the Zep paper it beat MemGPT on DMR, 94.8% to 93.4%, and improved LongMemEval accuracy by up to 18.5% over baseline implementations. You pay for that in operations. Neo4j, FalkorDB or Neptune to run, Kuzu support deprecated, Zep's full memory service cloud-only now, and no receipts or access policies. Good fit for compliance and clinical agents, where "who owned this account before March?" is a real question.

Letta has the agent edit its own memory blocks, with context tracked in git and scheduled consolidation in the background. Its original repo is a landing page now, with active source in letta-code. One honest data point they published themselves: Letta agents on gpt-4o-mini scored 74.0% on LoCoMo just by storing conversation history in files, which Letta read as a sign that current memory benchmarks may not mean much. It is a framework you adopt rather than a layer you drop into an existing agent.

Cognee turns text into entities and relationships and code into a graph of symbols, behind four operations: remember, recall, improve, forget. Plugins for Claude Code and Codex. Watch the edition line here, since running the whole memory layer on a single Postgres database is a demo feature in the open-source build.

LangMem is MIT and gives agents tools to save and search memory mid-conversation, plus a background manager that consolidates what they learn, natively against LangGraph's store. Two cautions: the default in-memory store loses everything on restart, so production needs a Postgres-backed store, and the last PyPI release was 0.0.30 in October 2025.

Supermemory local puts memory and RAG behind the same API as the hosted platform, reports first place on LongMemEval and LoCoMo, and publishes MemoryBench so others can compare. Its local build is one process, one machine, one API key. Best for personal assistants and air-gapped experiments on a single box.

Our own entry, including where it loses

Statewave is a memory runtime rather than a vector store. Raw agent events are saved as immutable episodes, a compiler turns them into typed memories carrying a confidence score, a validity window and the IDs of the episodes they came from, and at answer time it builds a ranked context bundle that fits a token budget you set. Given the same subject, task and budget over the same memory state, it returns the same bytes.

We forked Mem0's own benchmark suite and changed only the memory backend, keeping Mem0's judge unchanged with gpt-4o as answerer and judge:

Benchmark Statewave Mem0 cloud Mem0 OSS
LoCoMo (n=1,540) 0.905 0.899 0.866
LongMemEval (n=30) 0.967 0.933 0.833

Read those carefully, because I do.

Each is a single run rather than an average. Our LoCoMo margin over Mem0 cloud is 0.006, and a rerun can land either side of it, so nobody should treat it as a win. LongMemEval ran 30 questions, which makes it directional at best.

Two things cut against us specifically. We applied three fixes to Mem0's client code and every one of them raised Mem0's scores. And Mem0 OSS retrieves at most 20 memories per query by default, against roughly 200 for Statewave and Mem0 cloud, so part of that gap is retrieval budget rather than ranking quality. Which makes the cloud column the fairer comparison of the two.

Mem0 also publishes higher figures for its newer algorithm, 92.5 on LoCoMo and 94.4 on LongMemEval, from a different model stack than our fork runs. The full methodology is public and the fork diffs against upstream.

Where it loses:

  • Self-hosted only, with no managed cloud.
  • No first-class entity graph. If your questions are graph questions, Graphiti is the better tool.
  • One Postgres database. Multiple API replicas are verified, cross-region clustering is not available.
  • Tenant isolation is enforced in the application, without Postgres row-level security.
  • Bundles are denser than a plain fact store, so each answer costs more tokens. Best for teams whose memory has to stay in their own infrastructure and be explainable to an auditor or a customer.

Four questions that settle most of these evaluations

  1. Is the feature you are switching for in the free build today, under the same license? Edition pages get revised, and features move across the line in both directions. Check the date on whatever you are reading.
  2. Can you prove what the agent saw? If an auditor or a customer might ask, you need links back to source events and a per-call record, and you need them on the retrieval path.
  3. How many systems do you want to patch? A graph database, a vector database and an LLM is three.
  4. Can you clone and rerun the benchmark? One you can rerun beats a chart you cannot. Pick by requirement rather than star count. Fastest real test is to replay a week of your own conversations through two backends and diff the context bundles side by side. Whatever the leaderboards say, that diff is the thing you will actually be shipping.

Statewave's core runtime is on GitHub under Apache-2.0 if you want to read the compiler and the ranking code rather than take the table's word for it.

Top comments (0)