Memory Vault: Teaching AI That Yesterday's Truth Might Be Wrong
AI can remember a fact. The harder problem is knowing when that fact stopped being true.
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
I built Memory Vault, an open-source AI memory layer for people who constantly work across projects, decisions, deadlines, and changing technical choices.
I built it around a problem a friend working across projects kept running into: AI can remember previous decisions, but it doesn't reliably know when those decisions have changed.
Consider this conversation:
Monday:
"We're building this in React."
Wednesday:
"We switched to Next.js."
Friday:
"Next.js was rejected. We're back to React."
A traditional retrieval system can find all three statements.
The harder question is:
Which one is true now?
Memory Vault treats memories as structured, time-aware records instead of a flat collection of text chunks.
A memory can contain:
- content
- type
- entity
- scope
- confidence
- timestamp
- validity period
- source
- status
- relationship to previous memories
So the history becomes:
React
β
SUPERSEDED
β
Next.js
β
SUPERSEDED
β
React β CURRENT
If I ask:
What framework are we currently using?
Memory Vault can answer:
React. You previously switched to Next.js, but that decision was later reversed.
And, importantly, it can show why.
This focuses on one specific problem in AI memory: handling changing and contradictory information over time. Research and community discussions around AI memory repeatedly surface stale memories, contradictory facts, poor provenance, and weak user control as major problems.
Demo
Live Demo: https://memory-vault-ruddy.vercel.app
The interface is deliberately simple:
- Ask the memory a question.
- Retrieve relevant memories.
- Resolve their temporal state.
- Generate an answer from the selected evidence.
- Inspect the evidence behind the answer.
The goal isn't another chatbot with a fancy interface.
The goal is an auditable memory layer.
Code
GitHub: https://github.com/vallabhatech/memory-vault
The project is open source and built around a structured memory model rather than treating memory as ordinary chat history.
How I Built It
The architecture is roughly:
Conversation
β
βΌ
βββββββββββββββββ
β Memory β
β Classificationβ
βββββββββ¬ββββββββ
β
βΌ
βββββββββββββββββ
β Memory Store β
β β
β fact β
β entity β
β scope β
β confidence β
β timestamp β
β status β
β supersedes β
βββββββββ¬ββββββββ
β
βΌ
βββββββββββββββββββ
β Retrieval + β
β Ranking β
ββββββββββ¬βββββββββ
β
βΌ
βββββββββββββββββββ
β Conflict / β
β Temporal Logic β
ββββββββββ¬βββββββββ
β
βΌ
βββββββββββββββ
β Gemma β
β Answer β
ββββββββ¬βββββββ
β
βΌ
Answer + Evidence
The application is built with:
- Next.js
- TypeScript
- PostgreSQL / Supabase
- Gemma
- Tailwind CSS
Temporal Memory
The important part is that memories aren't simply overwritten.
Suppose the system initially stores:
{
"content": "The project uses React",
"status": "active"
}
Later:
"We switched the project to Next.js."
The new memory can reference the old one:
{
"content": "The project uses Next.js",
"status": "active",
"supersedes": "previous-memory-id"
}
The previous memory becomes:
{
"status": "superseded",
"valid_until": "..."
}
This preserves history instead of pretending the old information never existed.
That distinction is the entire point.
Gemma
Gemma is used for two important parts of the system.
1. Memory classification
When a new memory arrives, Gemma helps classify its relationship to existing candidate memories:
adds_information
duplicate
contradiction
supersedes
unrelated
2. Answer generation
Gemma receives the selected memory evidence instead of unrestricted access to the entire memory store.
The answer generation layer is designed to:
- use the supplied evidence
- treat active memories as current
- treat superseded memories as historical
- acknowledge uncertainty
- return the memory evidence supporting the answer
The model therefore acts as a reasoning layer over structured memory rather than being the memory itself.
Evaluation
I didn't want to stop at:
"It seems to work."
So I created 40 synthetic scenarios covering:
- technology decisions
- programming languages
- databases
- frameworks
- deadlines
- project requirements
- user preferences
- locations
- team decisions
- reversed decisions
I compared a simple retrieval baseline against the Memory Vault temporal-resolution pipeline.
The results were:
| Metric | Baseline | Memory Vault |
|---|---|---|
| Current-fact accuracy | 65% | 100% |
| Conflict-resolution accuracy | β | 100% |
| Stale-memory error rate | 35% | 0% |
| Precision@3 | 58.3% | 55.4% |
The most interesting result wasn't the overall retrieval score.
It was the stale-memory errors.
The baseline selected an outdated value in 14/40 scenarios.
Memory Vault selected the expected current value in 40/40 scenarios in this synthetic temporal-state evaluation.
There is an important caveat: these measurements evaluate the deterministic memory/retrieval layer. They are not a claim that Gemma produces perfect answers, and this evaluation does not measure production-scale LLM quality.
The evaluation is specifically designed to test temporal state transitions, which is the problem Memory Vault is trying to address.
Why Does Open Innovation Matter?
Memory is personal infrastructure.
If an assistant is going to remember decisions, preferences, projects, conversations, and potentially sensitive information, I don't want the memory layer to be inseparable from one proprietary model provider.
That's why using an open-weight model matters here.
Gemma gives the system a model layer that can potentially be swapped, deployed through different infrastructure, run locally, or adapted without redesigning the underlying memory model.
The architecture becomes:
Memory
β
Open-weight model
β
Structured reasoning
The memory system should belong to the user.
Not the model provider.
Open innovation also makes it easier to inspect the architecture, experiment with different models, and build the memory layer independently of a single closed API.
My Agent Session
LINK
Gemma
Memory Vault uses Gemma for memory classification and evidence-based answer generation.
MongoDB Atlas
Remove this category unless MongoDB Atlas was actually used in the submitted implementation.
What I Learned
The biggest lesson from building this was:
Memory is not storage.
A database can store everything.
That doesn't mean an AI knows what matters.
A useful memory system needs to reason about:
- what information belongs together
- what changed
- what became obsolete
- what contradicts something else
- where a fact came from
- whether the current evidence is sufficient
- when it should admit uncertainty
That's a much more interesting engineering problem than simply putting embeddings into a vector database.
Limitations
Memory Vault is deliberately an MVP.
It does not solve long-term AI memory.
Current limitations include:
- authentication
- production security
- large-scale memory stores
- advanced semantic/vector retrieval
- real-world continual learning
- privacy-preserving encrypted retrieval
- broader evaluation datasets
The current evaluation is synthetic and specifically designed to test temporal state transitions.
So my claim isn't:
"I solved AI memory."
It's:
I built and evaluated one approach to temporal, auditable personal memory.
What's Next?
The next version I'd like to explore:
- pgvector-based semantic retrieval
- better memory consolidation
- user-controlled memory editing and forgetting
- scoped memories such as personal/work/project
- encrypted local memory
- larger real-world evaluation datasets
- model swapping between different open-weight models
- feedback-driven memory correction
The long-term idea is simple:
An AI shouldn't merely remember what you said.
It should understand that what you said yesterday may no longer be true today β and it should be able to show you why.
Built with: Next.js Β· TypeScript Β· PostgreSQL/Supabase Β· Gemma Β· Tailwind CSS
Live Demo: https://memory-vault-ruddy.vercel.app
GitHub: https://github.com/vallabhatech/memory-vault
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
Top comments (0)