DEV Community

Awareness
Awareness

Posted on Originally published at github.com

Your AI Coding Agent Forgets Everything. I Fixed It With a Local MCP Memory Server

Your AI Coding Agent Forgets Everything. I Fixed It With a Local MCP Memory Server

Every new chat window, you re-explain your own codebase.

"We use PostgreSQL."
"No, the auth middleware is in src/lib/auth, not src/api."
"Yes, we already decided against Redux — that was three sessions ago."

It is the single biggest tax on AI-assisted development, and it has a boring name: agents have no memory. Context windows grew, but sessions still start from zero.

So I built Awareness Local — a local-first MCP memory server that gives your coding agent persistent memory across sessions. No cloud account. Works offline. Your memory is stored as plain Markdown.

Start in one command

npx @awareness.market/setup
Enter fullscreen mode Exit fullscreen mode

The installer detects your IDE (13+ supported: Cursor, Claude Code, Windsurf, Cline, Copilot, Codex CLI, Zed, JetBrains Junie, ChatGPT Desktop, …), wires up the MCP server, and starts a local daemon on 127.0.0.1:37800.

From then on your agent has five tools available:

Tool What it does
awareness_init Loads session context — recent knowledge, open tasks, project rules
awareness_recall Searches memory with progressive disclosure
awareness_record Saves decisions, changes, and insights (with knowledge extraction)
awareness_lookup Fast lookup of tasks, knowledge cards, risks, session history
awareness_get_agent_prompt Agent-specific prompts for multi-agent setups

The design decision that matters: it's just files

Memory lives in .awareness/, as Markdown you can read, edit, diff, and commit:

.awareness/
├── memories/2026-03-22_decided-to-use-postgresql.md
├── knowledge/decisions/postgresql-over-mysql.md
├── tasks/open/implement-rate-limiting.md
└── index.db   # SQLite FTS5 index, rebuilt automatically
Enter fullscreen mode Exit fullscreen mode

This is the whole point of local-first. Your agent's memory is not locked in a vendor's vector store. It is in your repo, in your git history, subject to your review. If the daemon dies, you have lost nothing — you have a folder of Markdown.

Retrieval is hybrid: SQLite FTS5 keyword search + local embeddings, fused with Reciprocal Rank Fusion. No LLM calls on the retrieval path, so recall costs you nothing per query and stays fully offline.

How good is the recall, actually?

I evaluated it on LongMemEval (ICLR 2025) — 500 human-curated questions, ~115k tokens of history per question:

Recall@1    80.2%
Recall@3    92.8%
Recall@5    96.0%   ◀ primary
Recall@10   98.6%
Enter fullscreen mode Exit fullscreen mode

Method: Hybrid RRF (BM25 + vector), multilingual-e5-small, zero LLM calls on the retrieval path, run on an Apple M1 / 8GB RAM in ~35 minutes total.

The ablation is the interesting part — neither method alone gets there:

Vector-only:   92.6%
BM25-only:     91.4%
Hybrid RRF:    95.6%   ← +3% over either single method
Enter fullscreen mode Exit fullscreen mode

Retrieval code and scripts are in the repo, so you can re-run it yourself rather than trust a table in a blog post.

Progressive disclosure: the token trick

The naive version of "give the agent memory" is to dump everything into context and watch your token budget evaporate. Awareness does a two-phase recall instead:

Phase 1  awareness_recall(query, detail="summary")
         → ~80 tokens per hit: title + summary + score
         → the agent picks what is actually relevant

Phase 2  awareness_recall(detail="full", ids=[...])
         → full content, only for the selected items
Enter fullscreen mode Exit fullscreen mode

The agent decides what to load. You pay for relevance, not for volume.

What it changes day to day

Before: every session re-opens the same three questions.

After: the agent opens the session already knowing that you migrated MySQL → PostgreSQL, that you chose it for JSON support, and that two TODOs were left open.

That difference compounds. Long migrations, multi-week refactors, and team handoffs are where it stops being a convenience and starts being the reason the work is tractable at all.

Try it

npx @awareness.market/setup
Enter fullscreen mode Exit fullscreen mode

Python and TypeScript SDKs (wrap_openai() / wrap_anthropic() interceptors) and an OpenClaw plugin are in the same ecosystem if you want memory inside your application, not just your IDE.

If re-explaining your codebase to your agent is getting old, a star helps more developers find it: ⭐ https://github.com/everest-an/Awareness-Market

Top comments (0)