Every "AI with memory" demo starts the same way: embed the user's notes, store the vectors, retrieve the top-k nearest chunks, paste them into the prompt. It works well enough in a demo. Then real use arrives and the cracks show.
This post is about where vector search stops being enough for personal memory, and what has to sit next to it. It's a design note, not a benchmark: we're not claiming numbers, just describing failure modes that show up when the thing you're remembering is a person's life.
What vectors are good at
Embeddings capture semantic similarity. If someone saved "dentist said to book a cleaning in March" and later asks "what did I need to do about my teeth?", nearest-neighbour search finds it without any shared keywords. That's genuinely useful, and for fuzzy recall it's hard to beat.
Where it breaks for personal memory
1. Similarity is not relevance. "What's Sara's number?" and a note that says "Sara's birthday dinner went great" are semantically close. The one you need is a structured fact (a phone number), not the nearest paragraph.
2. Time doesn't exist in embedding space. "I moved to a new flat" saved in January and "I'm moving next month" saved in September both match a query about where you live. Which one is current? Vectors can't tell you; you need timestamps, supersession rules, or both.
3. Facts get updated, not just added. A dentist appointment moves from Tuesday to Thursday. A vector store happily returns both. Personal memory needs the concept of "this replaces that".
4. Entities are not chunks. People, places and dates recur across many notes. If "Sara" is only ever a token inside chunks, you can't ask "everything connected to Sara" without hoping the retriever gets lucky. Extracting entities into their own records makes that a lookup, not a gamble. (We wrote about the extraction side here: how AI extracts people, places and dates from notes.)
5. Forgetting is a feature. If everything is retained forever with equal weight, retrieval gets noisier as the store grows. Deciding what to keep, summarise or let go is a design problem, not a storage problem.
What to put beside the vector index
A pragmatic stack for personal memory tends to look like this:
- Raw capture — keep the original note, message or voice transcript as the source of truth.
- Structured layer — extracted entities, dates and relationships stored as real fields, so exact questions get exact answers.
- Vector index — for fuzzy, "what was that thing about..." recall.
- Recency and supersession logic — so updated facts beat stale ones.
- A retrieval step that routes — exact lookup first when the question is exact, semantic search when it's vague, and a combination when it's both.
Whether the structured layer is a plain relational table or a graph is a separate decision, and the answer depends on how connected your data actually is. We compare the two in knowledge graph vs plain notes: when structure helps, and the retrieval basics in retrieval-augmented generation for personal notes, explained.
Short-term vs long-term memory
One more distinction that gets lost: what the assistant needs right now in a conversation (short-term) is different from what it should durably know about you (long-term). Treating both as "stuff in a vector DB" hides that. See how AI assistants remember across short-term and long-term memory for the longer version.
Takeaway
Vector search is a retrieval tool, not a memory system. For personal AI, the hard parts are knowing what's current, what's connected, and what to drop. If you're building in this space, start with the structure and add embeddings where fuzziness actually helps.
We're working on exactly this problem at Brinn, a personal AI that handles reminders, lists and memory across WhatsApp, email, web and desktop. If you're building something similar, we'd like to hear what's worked for you in the comments.
Top comments (0)