A Question Nobody Wants to Answer
First, an uncomfortable question: how many layers of encryption stand between your agent memory store and your company database?
The company database has TLS, at-rest encryption, audit trails, compliance certifications. A memory engine, meanwhile, holds something even more sensitive — the verbatim record of every customer conversation, the full context behind every business decision, every employee's preferences and habits. It is not a log. It is the company's second brain.
Now look at how the mainstream open-source memory engines handle self-hosting:
- mem0 (48K stars, the de facto category leader): the open-source edition serves plaintext HTTP by default, and the storage layer is unencrypted. Want encryption? The official advice is "plug in a vector database that supports encryption yourself." Transport encryption is barely mentioned in the docs.
- Zep: the community edition is deprecated; self-hosting means assembling Graphiti plus an external graph database yourself, with transport security entirely on you.
- MemOS: an academic project with beautiful papers, but almost zero public disclosure on security mechanisms — not a criticism; research projects simply have different priorities.
One detail from this year's 16-dimension memory framework benchmark stings: across eight projects, only one offered native at-rest encryption. Transport encryption fares worse — almost every self-hosted quickstart begins with http://localhost:PORT, and that http URL gets copied straight into production.
We're not saying these projects are bad — we've publicly expressed respect for mem0's usability. What we're saying is: as memory engines move from personal toys to shared enterprise infrastructure, the "plaintext by default" setting is going to cause an incident sooner or later.
So in L2.5, we shipped TLS. And we did it our way.
L2.5 TLS: Two Env Vars, Four Protocol Stacks Encrypted Together
NylonME's TLS design goal fits in one sentence: turning it on should not require a manual.
NYLON_TLS_CERT=server.crt
NYLON_TLS_KEY=server.key
Set both variables, and gRPC, HTTP, the web console, and MCP — all four protocol stacks — switch to TLS. Don't set them, and everything stays plaintext, exactly as before. Same philosophy as the auth release: the security capability is in place, but no existing user is forced to change config on a Friday night.
The real craft hides in three unglamorous decisions.
Decision 1: Half-config must crash, and crash loudly
Only NYLON_TLS_CERT set, no key — what should happen?
Many systems answer: "ignore it, start plaintext" — a warning line in the log, and your ops team discovers three months later that this link was never encrypted.
Our answer: exit code 2, refuse to start. A half-set security config means the operator intended encryption but something went wrong. Silently downgrading them to plaintext at that point is a betrayal. Fail-fast is not harshness; it is the first virtue of a security feature: better down than pretending to protect you.
Decision 2: A failed handshake must not kill the accept loop
There's a classic self-inflicted wound in TLS servers: one bad client sends a malformed handshake, the failure panics or propagates and takes down the accept loop — the whole service is down. A security feature becoming the entry point of a denial-of-service attack is exactly backwards.
In our implementation, a failed handshake only costs that one connection; the accept loop keeps going. The encryption layer must be tougher than the plaintext layer — otherwise attackers will thank you for making the DoS trigger more convenient.
Decision 3: The worst enemy of a security feature is "you thought it was on"
This is the most valuable bug in L2.5, valuable for the universal lesson it teaches.
Our MCP remote bridge supports https connections to remote engines. Code written, https:// parameter present, tests passing. But during L2.5 integration we found: the bridge silently fell back to plaintext in https scenarios. Users see a successful connection, everything working as usual — except the data is running naked on the network, and nobody knows.
This is the most dangerous way for a security feature to break: not failing with an error, but failing while looking perfectly fine. It was caught not by monitoring alerts but because our e2e tests use rcgen to mint real certificates and perform real handshakes — 34 test cases all green, plus four rounds of live smoke tests. If the tests had only mocked up to "config parsed correctly," this bug would have lived until some user's security audit found it.
The lesson deserves its own line: security feature tests must run until cryptography actually happened. Config parsed does not mean encryption happened; connection succeeded does not mean a handshake was performed.
The client side is ready too: Python SDK 0.2.4 supports https and tls_ca parameters, with the MCP bridge and CLI passing them through. For the client, it's still a one-variable change.
Scores, Since We're Here: Dual Benchmarks, Full Splits, Both 80+
Security is the protagonist of this post, but while we're at it, here's a month of evaluation progress — because the word "enterprise" needs receipts before it needs "secure":
| Benchmark | Split | Evidence recall@10 | End-to-end J |
|---|---|---|---|
| LoCoMo | official full (10 sessions, 1,536 questions) | 85.9% | 82.9% |
| LongMemEval-S | official full (500 questions) | 97.8% | 83.2% |
Three "fulls" deserve a pause: two benchmarks, both official full splits, both above 80. LoCoMo's 82.9% exceeds every published system under the same comparison protocol (caveats fully disclosed in the paper); LongMemEval-S's 500-question full split ran on a 350K-node graph, compared question-by-question against the small-store run with retrieval state bit-for-bit identical — 20x scale-up, zero precision loss.
Two more things we're quietly proud of, neither of which has anything to do with scores:
- We caught a "false regression" mid-process: one question type dropped 4.3 points, and we nearly bisected the codebase for an engine regression. Per-question log forensics showed the answering model itself had grown more cautious over those nine days — same evidence, confident in September, refusing to answer in October. We now have a rule: cross-week, few-question fluctuations are never trusted; every conclusion must come from same-period paired comparisons.
- Batch evaluation was bitten twice by infrastructure pollution (API quota contention, host memory pressure); each time we confirmed it with per-question paired diagnosis, discarded the run, and re-ran. The paper discloses even these pollutions and the exclusion methodology — because we believe clean data only earns trust when you show exactly how the dirty data was handled.
One Thing We Haven't Done — Saying It First
TLS protects data in transit. The memory files on disk (WAL, snapshots, vector index) are still plaintext today — at-rest encryption is on our roadmap, but not done, and we're not going to pretend otherwise.
This doesn't change today's enterprise deployment decision (most enterprises rely on disk/filesystem-layer encryption such as LUKS or BitLocker for databases anyway), but we believe a memory engine should offer a native option. When it ships, it will meet the same delivery standard: off by default, zero behavior change, half-config fail-fast, real crypto with real tests.
Final Words
We've long held one belief: memory engines will walk the same road databases did — first a personal toy, then a team tool, and finally enterprise infrastructure. Databases took forty years on that road; memory engines may only need four.
But one road must not be retraced: the database industry only made encryption the default after countless security incidents. A memory engine stores things more sensitive than a database — it doesn't get to learn that lesson at the same price.
L2.5 is just the first class we made up. Audit, multi-tenancy, and backup drills were submitted earlier (Post 10); at-rest encryption is on the way. If you're building enterprise-grade agents too, come read our code — it's a two-env-var change, and your memory store can stop running naked tonight.
NylonME: a Rust single-binary memory engine, Apache-2.0, two minutes to plug into Claude Code / Codex / VSCode / Qoder via MCP. Find nylon-memory/NylonME on GitHub, or start from our Quick Start.
All benchmark numbers in this post are reproducible from the repo's docs/LOCOMO_BENCHMARK.md; the paper (full ablations and eight negative results included) is ready, with arXiv submission in progress.





Top comments (0)