Your self-hosted AI assistant keeps forgetting you. Here is how we measured a fix.
Every guide to running your own AI assistant covers the easy half: install Ollama, point a web UI at it, add an API key for a bigger model when you need one. None of them warn you about the hard half, which shows up around week three. The assistant starts forgetting things you told it in week one. Not in a dramatic way. It just re-asks your preferences, loses a project detail you corrected twice, and confidently invents an answer that used to be on a sticky note inside its own context window.
We hit this while building ScallopBot, an open-source self-hosted assistant (we work on it, and this post links to our site once at the end). Rather than guess at a fix, we tried to measure the problem and the fix against a published benchmark. Here is what we learned, in the order you will probably need it.
First, name the failure you actually have
"Bad memory" is three different bugs wearing one coat:
- Storage failure. The fact was never written down anywhere. When the conversation ended or got compacted into a summary, it was gone. If your assistant only remembers what fits in the context window, this is you.
- Retrieval failure. The fact was stored, but the search could not find it. You asked about "the container thing" and nothing came back for "Docker configuration" because the retrieval was pure keyword, or pure vector, and the phrasing did not line up.
- Fabrication failure. Retrieval found nothing relevant, and instead of saying so, the model generated a plausible answer from weak matches. This is the worst one, because it erodes trust in the two cases where memory was working fine.
Which fix you want depends on which bug you have. Dump the raw transcripts and grep them. If the fact is in the transcript but the assistant still forgot it, you have a retrieval problem, not a storage problem.
A memory layout that survives restarts
Keep the whole memory store in one file. We use SQLite: memories, relations between memories, and session transcripts all live in one database. This sounds like a trivially boring choice and it is, which is the point. Backups are one file copy. Other processes (skills, an MCP server) can open the same store concurrently in WAL mode. You can inspect it with the sqlite3 CLI when the assistant does something strange, which beats any observability dashboard for debugging memory.
Two details matter more than the schema:
- Store dates inside the memory, not beside it. "We moved the staging server last month" is useless to retrieval unless the date is embedded with the fact. Later, questions like "what changed since the review?" have something to match against.
- Store what you rejected. When a newer fact contradicts an older one, keeping both and marking the old one superseded costs you a little retrieval noise. A wrong overwrite loses information permanently. Asymmetric mistakes deserve asymmetric caution.
Retrieval: two searches and a gate
Single-method retrieval is where most home-built assistants plateau. Hybrid retrieval fixes a surprising amount:
- BM25 keyword search catches the exact name, ticket number, or identifier that embeddings blur away.
- Dense vector search catches the paraphrase that keyword search misses entirely.
- Merge the two result sets, then optionally rerank with an LLM before anything enters the context window.
The underrated piece is the gate at the end. Return results only above a relevance threshold; if nothing clears the bar, inject nothing and let the model say it does not know. That one check is the difference between an assistant with occasional gaps and an assistant that lies to you about having gaps.
Maintenance: memory needs a night shift
A memory store that only grows gets noisier to search every month. Ours runs on three clocks: a light decay tick every minute, a deeper pass roughly hourly that summarizes idle sessions and audits what has been forgotten, and a nightly consolidation cycle that merges near-duplicates, links related memories, and prunes what stopped being useful. The nightly pass is where the real cleanup happens, and it costs almost nothing because it runs while you sleep. If you build your own, start with the nightly job: dedupe on high embedding similarity, and reinforce instead of duplicating when the user repeats a fact.
Does any of this actually help? Numbers.
We ran this architecture against LoCoMo, a public benchmark of long-conversation memory (1,049 QA items across 138 sessions), using identical models and embeddings on both sides. The hybrid assistant with reranking and the score gate scored F1 0.48 against 0.38 for a file-based baseline. The more telling split: on adversarial questions with no correct answer stored, the gated system scored 0.97 against 0.77. Most of the gap is the refusal to confabulate, not smarter recall.
Two honest caveats. These numbers describe one benchmark and one model pair, not a universal league table; treat them as "this direction works" rather than "this wins." And the whole cognitive pipeline is cheap: our estimate from the per-operation cost table is about 5 to 10 cents a day of model spend at 100 messages a day, on top of a $5 to 8 monthly VPS, or nothing extra if you run it on hardware you already have.
If you want to poke at the implementation, read the memory architecture notes and the cost breakdown at scallopbot.com, or clone it and run your own numbers. The project is MIT licensed, and the memory store is one SQLite file you can open with your eyes open. That, more than any F1 score, is the actual point of self-hosting this stuff.
Top comments (1)
Keeping superseded rows rather than running hard deletes saves you from bad overwrites, provided the query drops them before ranking. In my own SQLite setup, letting vector or BM25 search pull candidates before filtering superseded rows in Python ended up filling the top-k buffer with stale revisions of the same config. Putting a partial index on non-superseded records keeps the candidate pool clean before the reranker or score gate touches it.