Most of my coding-agent sessions start the same way. The agent rediscovers that the repo uses pnpm, gets told again not to touch the legacy payment file, and trips over the test that only passes in UTC. None of that is in the code. It lives in someone's head, or in a previous session that is gone.
Hindsight by Vectorize is built for that gap. It is an open-source agent memory server (MIT) that agents write to and read from across sessions. When I checked on 3 October 2026 its GitHub repo had 44,886 stars, sat second on GitHub's weekly trending page with 16,183 stars that week, and its latest release was v0.10.2 from 29 September.
I wanted to see what it does in practice, so I installed it on a Linux box and called it from a small Node/TypeScript project. One constraint shaped the whole test: I didn't have an LLM API key on that machine. Hindsight documents a mode for exactly that, so this is a test of its no-LLM floor, not of everything it can do. I've marked clearly which parts I ran and which parts I only read about.
The setup
The box was Linux x86_64 with 8 vCPUs, about 15 GB of RAM and Node 20. There was no Docker, so I used the bare-metal path from the README (pip install hindsight-api), through uv:
uv venv .venv-hs --python 3.12
uv pip install hindsight-api==0.10.2
It resolved 227 packages and finished in 11.6 seconds, including the PyTorch download. The catch is size: the virtual environment came to 6.3 GB, mostly torch and CUDA wheels. Plan disk space accordingly.
The configuration docs describe HINDSIGHT_API_LLM_PROVIDER=none as "chunk storage + semantic search only, no API key needed." In this mode retain stores your text as chunks without LLM fact extraction, recall still works, reflect returns HTTP 400, and observations are switched off.
export HINDSIGHT_API_LLM_PROVIDER=none
hindsight-api
The startup banner showed v0.10.2 with an embedded PostgreSQL (pg0, started automatically), LLM none / none, local embeddings (BAAI/bge-small-en-v1.5), a local reranker (cross-encoder/ms-marco-MiniLM-L-6-v2), and MCP enabled at /mcp.
- The first start took about 38 seconds to reach "Uvicorn running", including model downloads from Hugging Face.
- A warm restart took about 19 seconds to a healthy
/health. - The worker process used roughly 1.35 GB of resident memory.
Session 1: one agent writes down what it learned
For the Node/TypeScript side I installed the official client (version 0.10.2) and ran the scripts with tsx:
npm install @vectorize-io/hindsight-client typescript tsx @types/node
The scenario was a fictional shop-api repo. The first script plays "Agent A" at the end of a session. It creates a bank, which is Hindsight's isolated memory store, and retains five project facts that aren't obvious from the code.
// src/agentA.ts
import { HindsightClient } from '@vectorize-io/hindsight-client';
const BANK = 'shop-api-repo';
const client = new HindsightClient({ baseUrl: process.env.HINDSIGHT_URL ?? 'http://localhost:8888' });
const learnings = [
{ content: 'The shop-api repo uses pnpm, not npm. Run `pnpm test` with Vitest; Jest was removed in March.', context: 'tooling' },
{ content: 'All money values are stored as integer cents in Postgres (column type bigint). Never use floats for prices.', context: 'data model convention' },
{ content: 'Route handlers live in src/routes/*.ts and must validate request bodies with zod schemas from src/schemas.', context: 'code convention' },
{ content: 'The flaky test checkout.spec.ts fails when TZ is not UTC; CI sets TZ=UTC, so set it locally too.', context: 'debugging note' },
{ content: 'Do not touch src/legacy/paypal.ts; it is scheduled for deletion and has no tests.', context: 'team decision' },
];
async function main() {
await client.createBank(BANK, {
reflectMission: 'Long-term memory for coding agents working on the shop-api TypeScript repo.',
});
for (const l of learnings) {
const t0 = performance.now();
await client.retain(BANK, l.content, { context: l.context, tags: ['repo:shop-api'] });
console.log(`retain: ${(performance.now() - t0).toFixed(0)} ms`);
}
}
main();
The five retains took 184, 42, 46, 56 and 41 ms. Every response reported zero LLM tokens, which is expected with no model. listMemories showed five memories, all with fact_type: "world". In chunks mode that means the text is stored as written: no extracted entities and no restructuring.
Session 2: a fresh process asks what it should know
The second script is a separate process with a new client and no shared state apart from the bank name. It asks four questions an agent would plausibly ask on day one, then tries reflect.
// src/agentB.ts
import { HindsightClient } from '@vectorize-io/hindsight-client';
const BANK = 'shop-api-repo';
const client = new HindsightClient({ baseUrl: process.env.HINDSIGHT_URL ?? 'http://localhost:8888' });
const questions = [
'How do I run the tests in this repo?',
'How should I store a product price?',
'Why does checkout.spec.ts fail on my machine?',
'Can I refactor the PayPal integration?',
];
async function main() {
for (const q of questions) {
const t0 = performance.now();
const r: any = await client.recall(BANK, q, { tags: ['repo:shop-api'], maxTokens: 512 });
console.log(`Q: ${q} (${(performance.now() - t0).toFixed(0)} ms)`);
for (const m of r.results.slice(0, 2)) console.log(' ->', m.text);
}
try {
await client.reflect(BANK, 'What should a new agent know before editing this repo?');
} catch (e: any) {
console.log('reflect error:', e.message);
}
}
main();
Here is what came back:
| Query | Latency | Top result |
|---|---|---|
| How do I run the tests in this repo? | 227 ms | the pnpm / Vitest note ✅ |
| How should I store a product price? | 150 ms | the integer-cents note ✅ |
| Why does checkout.spec.ts fail on my machine? | 62 ms | the TZ=UTC note ✅ |
| Can I refactor the PayPal integration? | 68 ms | the legacy/paypal.ts note ✅ |
The right memory came first on all four questions. Two caveats keep that in proportion. The bank had five entries, so every query returned all five, and what I was really seeing was ranking, not filtering. The second result was sometimes unrelated: the price question had the TZ note in second place. That's fine for a toy, but on a real bank you would want the maxTokens budget or the minScores floors the client exposes.
reflect failed exactly as documented:
Reflect requires an LLM provider. Current provider is set to 'none'.
Set HINDSIGHT_API_LLM_PROVIDER to a real provider (e.g., openai, anthropic, gemini).
Next I stopped the server, started it again and re-ran Agent B. The top results didn't change (201 ms and 91 ms on the first two queries), so the memories live on disk in the embedded Postgres, not just in the process.
A third client over MCP
Every Hindsight server exposes a Model Context Protocol endpoint per bank at /mcp/{bank_id}/. I talked to it with plain curl. You send an initialize request, read back the mcp-session-id header, and send that header with each later call:
curl -s -D hdr.txt -X POST http://localhost:8888/mcp/shop-api-repo/ \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1"}}}'
SID=$(grep -i mcp-session-id hdr.txt | awk '{print $2}' | tr -d '\r')
curl -s -X POST http://localhost:8888/mcp/shop-api-repo/ \
-H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' \
-H "mcp-session-id: $SID" \
-d '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"recall","arguments":{"query":"which package manager and test runner?"}}}'
The server identified itself as hindsight-mcp-server 0.10.2, and tools/list returned 36 tools. They cover retain and recall, reflect, mental models, directives, memory and document management, operations, tags, bank settings and knowledge-page CRUD. The recall call returned the pnpm/Vitest memory that Agent A had written through the TypeScript SDK. The SDK and MCP both work against the same bank, so a TypeScript service and any MCP-capable agent can share one memory.
What the coding-agent installer actually wires up
Coding agents are the headline use case. Vectorize ships @vectorize-io/hindsight-coding-agents, and the version I tried (0.8.0) lists 20 harnesses in its installer: Claude Code, Codex, Cursor CLI, GitHub Copilot CLI, opencode, Grok Build, Cline CLI, Devin CLI, Antigravity CLI, Qwen Code, Kimi Code, DeepSeek Harness and more.
I did not run a real coding-agent session. The only agent the installer detected on the box was a Cursor CLI install that other tools share, and I didn't want to change its config. I installed the integration into a throwaway home directory instead and inspected what it writes:
HOME=./fakehome npx -y @vectorize-io/hindsight-coding-agents install cursor-cli \
--server self-hosted --api-url http://localhost:8888
It reported hooks merged into ~/.cursor/hooks.json, an MCP server added to ~/.cursor/mcp.json, and a skill installed under ~/.cursor/skills/hindsight-coding-agent. Here are the hooks it wrote (paths shortened):
{
"version": 1,
"hooks": {
"sessionStart": [{ "command": "node \".../cursor-sessionstart-hook.js\"", "timeout": 30 }],
"beforeSubmitPrompt": [{ "command": "node \".../cursor-hook.js\"" }],
"stop": [{ "command": "node \".../cursor-stop-hook.js\"", "timeout": 30 }]
}
}
The MCP entry runs a local stdio server, node ~/.hindsight/coding-agents/dist/mcp-server.js ${workspaceFolder}. I started that server myself with my Node project as the workspace and listed its tools. It has eight: hindsight_sync_status, hindsight_diagnose, hindsight_search_knowledge_pages, hindsight_list_knowledge_pages, hindsight_read_knowledge_page, hindsight_reflect, hindsight_capture_initiative and hindsight_ingest_document.
hindsight_diagnose resolved the project to a bank called coding-agent::node-agents on my local server. hindsight_sync_status reported synced: false, no git log and zero knowledge pages. That fits the docs, which say ingestion of git history and past sessions starts from the session-start hook, and I never fired that hook with a real agent. So I can show you the wiring, but not how it behaves inside an agent.
What I could not test
All of the following needs an LLM, so I can only describe it from Vectorize's docs:
- Fact extraction on retain. With a model, retain pulls out facts, entities, relationships and dates instead of storing raw chunks.
- Observations. Related facts are consolidated in the background into deduplicated beliefs that keep their supporting evidence.
- Reflect. A slower, reasoning pass over the bank, as opposed to a lookup.
- Mental models and knowledge pages. Standing answers and wiki-like pages that the bank rewrites as it learns. The coding-agent integration builds on these.
-
The LLM wrapper (
hindsight-litellm), which recalls and retains automatically around each model call. - Benchmarks. The README says Hindsight is state of the art on LongMemEval and that Virginia Tech's Sanghani Center and The Washington Post independently reproduced its results. I didn't benchmark anything, and five memories prove nothing about accuracy at scale.
Should you try it?
Here is what I can say from what I ran. The install is quick but large. The server starts with no external services and no API key. The TypeScript client is simple. Memories persist across restarts. One bank serves both SDK calls and MCP tool calls, and on my small, deliberately easy set of questions recall put the right note first every time.
Without a model, though, Hindsight is a well-packaged local vector-plus-keyword store with an MCP front end. The "learning" in its tagline (extraction, observations, reflect, knowledge pages) is the part that needs an LLM, and that is the part I haven't verified. If you have a key, or a local model through Ollama or LM Studio (both are supported providers), that is where I'd point the next test. If you don't, the none mode is still an honest, quick way to see whether the API and the bank model fit your agent setup before you spend tokens.
Have you put Hindsight or another memory layer behind your coding agents? I'd like to hear in the comments what actually stuck between sessions.
Originally published on Medium.
Top comments (0)