The Problem: Why Vector Databases Aren't "Memory"
Over the past year, as developer workflows shifted toward autonomous coding agents like Claude Code, Cursor, and custom agent swarms, one glaring limitation became impossible to ignore:
AI agents have total amnesia.
Every time a session ends or a subagent is spawned, the model's working context is wiped clean. When you ask your agent to implement a new feature tomorrow, it has no recollection of:
- Architectural tradeoffs you spent two hours debating yesterday.
- Database constraints and schema conventions you established last week.
- Bug workarounds that took all morning to isolate.
To solve this, most teams default to the standard "RAG memory" blueprint:
- Capture conversation turns.
- Generate embeddings with OpenAI or Jina.
- Store chunks in a vector database (Pinecone, Chroma, pgvector).
- Run cosine similarity search on the next turn.
While this demo looks neat on a 3-slide pitch deck, it breaks down completely in real-world software engineering.
1. Vector Search is Blind to Time
Consider this real scenario:
- January 2026: You tell your agent: "We are using MySQL 8 for our auth database."
- August 2026: You tell your agent: "We migrated the auth database to PostgreSQL 18 with Row-Level Security."
Both sentences have virtually identical semantic embeddings when queried with: "What database do we use for auth?". A standard nearest-neighbor vector search will return both facts with near-equal similarity scores. The LLM gets confused, combines them, or hallucinates an incorrect hybrid setup.
2. Lack of Lexical and Negative Precision
Dense vector similarity is inherently probabilistic. It struggles with:
- Exact symbol names and identifiers (
ORA-00942,DiskANN,River). - Specific version numbers (
PostgreSQL 18vs16). - Negative rules ("Never use package X under any circumstances").
3. Context Window Bloat and Token Inefficiency
Because vector search lacks precision, developers compensate by increasing the top_k candidate count or shoving massive 3,000-line .cursorrules files into the prompt.
If your agent runs 100 queries a day, dumping 6,000 tokens of raw history on every turn wastes 600,000 tokens daily — inflating your LLM bill by hundreds of dollars a month without actually solving the amnesia problem.
The Solution: Designing Nexusyn
To solve agent memory at its structural roots, we engineered Nexusyn (https://nexusyn.ai). Rather than a simple wrapper around an embedding model, Nexusyn is a dedicated data plane designed for agentic long-term memory.
Here is the architectural breakdown of how it works.
┌─────────────────────────┐
│ AI Agent / IDE │
│ (Claude Code, Cursor, …)│
└────────────┬────────────┘
HTTP │ MCP (/v1/mcp)
▼
┌────────────────────────────────────────────────────────────────────────┐
│ NEXUSYN ENGINE │
│ │
│ POST /v1/ingest POST /v1/query │
│ │ │ │
│ ▼ ▼ │
│ [ Deduplication ] [ Sub-query Gen ] │
│ │ │ │
│ [ Chunker ] [ Hybrid Retrieval ] │
│ │ ├── Vector (DiskANN) │
│ ▼ ├── BM25 Full-Text │
│ [ River Queue (Async) ] ├── Entity Graph │
│ ├── Batch Embeddings └── Date Anchor Boost │
│ ├── Entity & Relation Extraction │ │
│ └── Wiki / Profile Compilation ▼ │
│ [ Reranker (Cross-Enc) ]│
│ │ │
│ ▼ │
│ [ Grounded Synthesizer ]│
└────────────────────────────────────────────────────────────────────────┘
1. High-Throughput Vector Indexing with DiskANN
We chose Go 1.24 and PostgreSQL 18 with pgvectorscale.
Traditional HNSW (Hierarchical Navigable Small World) vector indexes consume massive amounts of RAM because the entire graph must reside in memory. In contrast, DiskANN stores compressed vector representations in memory while keeping the primary graph index on SSD/NVMe.
This achieves:
- 3.4x lower memory consumption than HNSW at scale.
- Single-transaction ACID consistency: relational tables, temporal timestamps, and vector embeddings live inside the same database engine with Row-Level Security (RLS).
2. Serialized Hybrid Search with Reciprocal Rank Fusion (RRF)
When an agent queries Nexusyn, the engine executes hybrid search in a single read-only transaction:
- Dense Vector Retrieval: Cosine distance search over DiskANN indexes.
-
PostgreSQL BM25 Search: Lexical matching over
tsvectorwith language stemming. - Reciprocal Rank Fusion (RRF): The results are merged using the formula: $$RRF_Score(d) = \sum_{m \in M} \frac{1}{k + r_m(d)}$$ (where $k=60$).
- Cross-Encoder Reranking: The top 64 fused candidates pass through a hosted cross-encoder reranker, scoring the exact semantic relevance between the query and each chunk.
3. Bi-Temporal Fact Versioning
Every memory record in Nexusyn carries two distinct time axes:
- Transaction Time: When the fact was recorded in the database.
-
Valid Time (
valid_from/valid_to): When the fact was true in the real world.
When an agent records a new decision that contradicts or updates a previous one, the engine marks the previous fact's valid_to = now(). The historical audit trail remains intact, but active queries will never recall the obsolete state as current reality.
4. Asynchronous Knowledge Graph
Memory is not flat text; it is an interconnected web of entities, decisions, and constraints. An asynchronous background queue (powered by Go's River queue over Postgres) extracts named entities and relationships into a bi-temporal knowledge graph.
Through the get_related tool, agents can traverse multi-hop connections (e.g. connecting an API rate limit back to the security review that mandated it).
Native Model Context Protocol (MCP) Integration
Rather than forcing developers to write bespoke SDK wrappers, Nexusyn implements the official Model Context Protocol (MCP) specification over Streamable HTTP (/v1/mcp).
It exposes 7 native tools directly to any MCP-compliant client:
-
add_memory: Saves facts, decisions, preferences, and lessons. -
search_memory: Executes hybrid retrieval and returns grounded answers with cited sources. -
get_guideline: Deterministic retrieval of organizational standards. -
get_related: Explores 1-2 hops across the knowledge graph. -
graph_overview: Structural topology, central nodes, and concept connectors. -
update_memory: In-place memory corrections. -
delete_memory: Bi-temporal soft deletion.
Connecting in 10 Seconds
In Claude Code:
claude mcp add nexusyn --transport http \
"https://api.nexusyn.ai/v1/mcp" \
--header "Authorization: Bearer <YOUR_API_KEY>"
Or via the Smithery CLI:
npx -y @smithery/cli install @nexusyn/nexusyn --client claude
In Cursor (.cursor/mcp.json):
{
"mcpServers": {
"nexusyn": {
"url": "https://api.nexusyn.ai/v1/mcp",
"headers": {
"Authorization": "Bearer <YOUR_API_KEY>"
}
}
}
}
Benchmark Results: LongMemEval-S & LoCoMo
To validate that this architecture actually moves the needle, we evaluated Nexusyn against the public LongMemEval-S benchmark (350 rigorous multi-session recall and temporal update tasks) and LoCoMo (1,986 multi-turn reasoning questions):
| Benchmark | Nexusyn Score | Metric Focus |
|---|---|---|
| LongMemEval-S | 81.1% (284/350) | Multi-session recall, temporal update tracking, preference extraction |
| LoCoMo | 74.3% (1,476/1,986) | Complex conversational reasoning and contradiction resolution |
Conclusion & Getting Started
Giving AI agents long-term memory is not a prompt-engineering trick — it requires dedicated data-plane infrastructure that understands time, entities, and hybrid retrieval.
You can explore the open-source engine, deploy it self-hosted, or use the hosted cloud instance:
- Cloud SaaS (Free Tier with 2,000 memories): https://nexusyn.ai
- Open Source Core (Apache 2.0): https://github.com/nexusyn/engine
- Smithery Registry: https://smithery.ai/servers/@nexusyn/nexusyn
- Glama Registry: https://glama.ai/mcp/servers/nexusyn/engine
Top comments (0)