Every agent-memory demo I tried worked the same way:
conversation → chunks → embeddings → top-k → prompt
It answered confidently. Sometimes it answered confidently about things that never happened.
That is the failure mode nobody demos. A false memory is worse than no memory: an agent that can't recall will ask again; an agent that misremembers will act on fiction — re-applying a reverted decision, citing a superseded API, or "remembering" a preference the user never stated.
We kept hitting this while building real agent workflows, so we stopped treating "memory quality" as a single retrieval number and started treating memory as a lifecycle: what was stored, from what source, with what authority, what it safely replaced, and what happens when the evidence isn't good enough.
The single biggest behavior change that fell out of this: refusal.
When our memory layer can't support an answer from its evidence, it doesn't guess — it declines and tells you exactly why: which checks failed, what was missing, what would be needed. The first time you see it in a real workflow, it feels wrong. Then you realize every confident hallucination it replaced was a bug you would have shipped.
Three things this forced us to build properly:
Evidence attached to every recall. Not a score — the source, the authority (user statement vs. observed fact vs. assistant proposal), and whether it has been superseded. If you can't show why a memory was used, you can't trust the decision built on it.
Deletion that actually deletes. Removing a memory has to remove the vectors, the candidates, and the delayed reindex jobs — not just the primary record. We issue a signed deletion receipt so you can prove it.
Memory that evolves instead of overwriting. When a preference or decision changes, the current state is recorded and the earlier state kept as history. Nothing is silently overwritten, so "what did we decide, and when did it change" is always answerable.
None of this shows up in a retrieval benchmark. It shows up at 2am when your agent does something you can't explain — and now you can.
We run this as SPM-Polaris at spmos.ai — provider-neutral, MCP-native, with a free tier. If you're building agents that need to remember things correctly, not just plausibly, it's worth an afternoon.
Docs: https://docs.spmos.ai
Top comments (0)