Give an agent memory and it will cheerfully remember the wrong thing.
A support rep says "just refund it" once in a frustrated chat. A flat vector store happily indexes that line, and three weeks later the agent is quoting it as company policy.
The fix isn't less memory. It's putting a human between "the AI noticed something" and "this is now shared knowledge." The pattern is simple enough to build in an afternoon.
The loop
Observe → Detect → Review → Promote → Record
Observe: log what agents see in conversations.
Detect: only flag a pattern that recurs across enough distinct people and sessions.
Review: a human sees the exact text that would be saved.
Promote: approved text becomes shared memory.
Record: every decision goes in an audit trail.
Step 1: Detect recurring patterns, not one-offs
python
from collections import defaultdict
observations = defaultdict(lambda: {"users": set(), "sessions": set()})
def observe(pattern: str, user_id: str, session_id: str):
observations[pattern]["users"].add(user_id)
observations[pattern]["sessions"].add(session_id)
def candidates(min_users=3, min_sessions=5):
return [p for p, o in observations.items()
if len(o["users"]) >= min_users and len(o["sessions"]) >= min_sessions]
One frustrated rep can't poison the memory. It takes a pattern across several people.
Step 2: The review queue
python
import time, uuid
queue, memory, audit_log = [], {}, []
def propose(layer_path: str, text: str):
item = {"id": str(uuid.uuid4()), "layer": layer_path, "text": text, "status": "pending"}
queue.append(item)
return item
def review(item_id: str, reviewer: str, decision: str, edited_text: str = None):
item = next(i for i in queue if i["id"] == item_id)
if edited_text:
item["text"] = edited_text # reviewer can edit before approving
item["status"] = decision # "approved" | "rejected"
if decision == "approved":
memory.setdefault(item["layer"], []).append(item["text"])
audit_log.append({"ts": time.time(), "item": item_id, "reviewer": reviewer,
"decision": decision, "text_shown": item["text"]})
The reviewer sees the exact text that will be written, never a paraphrase. A summary like "agents learned a refund rule" hides the dangerous detail.
Step 3: Scope retrieval to a layer path
Flat indexes leak. A one-branch exception ("the Mumbai office waives this fee") shouldn't answer a Berlin customer's question.
python
def retrieve(layer_path: str):
"""Return memory from this layer and its ancestors only."""
parts = layer_path.split("/")
results = []
for i in range(1, len(parts) + 1):
results += memory.get("/".join(parts[:i]), [])
return results
propose("acme/emea/berlin", "Waive late fees under €20.")
after approval, only acme, acme/emea and acme/emea/berlin queries can see it
Step 4: Make demotion possible
Approval isn't forever. Add a demote(item_id) that removes text from memory and logs it. Knowledge that was right in January can be wrong in June.
Gotchas
Reviewer fatigue. If the queue is 400 items deep, people rubber-stamp. Raise the recurrence thresholds until the queue stays small and meaningful.
Static vs adaptive. Uploaded documents can be indexed immediately. Only learned patterns need the queue.
Don't trust the detector. It surfaces candidates, and the human decides.
Wrap-up
An agent's memory should be treated like production config: changes are proposed, reviewed, versioned and reversible. What's the strangest thing your agent "learned" that you had to delete? Tell me in the comments.
Top comments (0)