DEV Community

Sohail Shaik
Sohail Shaik

Posted on

What Retain and Recall Actually Look Like in a Working Agent

What Retain and Recall Actually Look Like in a Working Agent
I've read a dozen posts about "agent memory" that describe it in the abstract — store some embeddings, retrieve them later, profit. None of them showed me what it actually looks like when it fires. So when we wired Hindsight into an agent that watches a community's conversations over time, I paid close attention to exactly that: what happens on screen, in the data, the moment memory goes from a concept to a running system.
What the system does
The project is called Community Time Machine: an agent that sits on top of a stream of community messages — the kind you'd see in a Discord server for an open-source project — and helps moderators turn recurring questions into durable knowledge instead of answering each message as if it arrived in isolation. The loop the system runs is: Remember → Reuse → Intervene → Measure → Learn.
Two parts of that loop are where the mechanism actually earns its keep:
Recurring question detection and FAQ generation. Ask the system about a problem that's come up before, and instead of a generic answer, it recognizes the pattern, pulls the actual historical evidence, and drafts a community-grounded answer.
Intervention tracking. Once a moderator publishes that answer, the agent remembers the publication itself as an event. Ask it later whether the intervention worked, and it can compare activity before and after — not from a hunch, but from what it retained at the time.
Neither of these needs a human to manually point the agent back at old messages. That's the part that only makes sense because of Hindsight's agent memory — retain and recall aren't a feature bolted onto the agent, they're the mechanism the whole thing depends on. We didn't call Hindsight directly from the agents, either — everything goes through a thin MemoryRetriever wrapper, so the investigation and FAQ logic never needed to know anything about how memory was actually stored underneath.
The core technical story: what retain and recall actually do
It's easy to say "the agent remembers things." It's more useful to show what that means at the moment it happens.
Remembering is the write path. Messages and events — a product update, an FAQ being published, a release — get written into memory the same way, tagged with enough context to be found again later by meaning, not just by keyword:
async def _load_messages(self) -> None:
path = self.data_dir / "messages.json"

with open(path, "r", encoding="utf-8") as file:
    messages = json.load(file)

for message in messages:
    await self.client.remember(
        memory_id=message["id"],
        content=message["content"],
        memory_type="message",
        timestamp=message["timestamp"],
        metadata={
            "channel": message["channel"],
            "author": message["author"],
            "type": message.get("type", "message"),
        },
    )
Enter fullscreen mode Exit fullscreen mode

Events get written through the same call, just with a different memory_type. To the memory layer, a product update and a user's message are both just memories with a timestamp — which is what lets the system later connect an update to a spike in questions three weeks apart.
Recalling is the read path, and it's where the interesting behavior shows up. Every retrieval in the system — whether it's the FAQ agent gathering evidence or the investigation agent reconstructing a timeline — goes through the same method:
def search(
self,
query: str,
limit: int = 10,
) -> List[Dict[str, Any]]:
"""
Retrieve memories relevant to a query.
"""
return self.client.search(
query=query,
limit=limit,
)
MemoryRetriever sits between the agents and the memory store on purpose. Neither the FAQ agent nor the investigation agent know or care how memories are stored — they just call .search() and get back whatever's relevant, whether that's a message from an hour ago or an event from a month ago.
Turning memory into an answer, honestly
Here's where it gets more interesting than "recall some text and paste it into a prompt." The FAQ agent doesn't just retrieve evidence — it checks whether there's enough of it before it's willing to say anything at all:
async def generate(self, question: str, limit: int = 10):
memories = await self.retriever.search(query=question, limit=limit)

if not memories:
    return {
        "status": "INSUFFICIENT_EVIDENCE",
        "question": question,
        "answer": "",
        "source_ids": [],
        "reason": (
            "No relevant community memories were found. "
            "An FAQ draft cannot be generated safely."
        ),
    }

source_ids = [memory["id"] for memory in memories if memory.get("id")]
evidence = "\n".join(
    f"[{memory['id']}] {memory.get('content', '')}"
    for memory in memories
)
# ... evidence is passed to the LLM, which returns a draft answer ...

return {
    "status": "DRAFT",
    "question": question,
    "answer": data["answer"],
    "source_ids": source_ids,
    "notes": data["notes"],
}
Enter fullscreen mode Exit fullscreen mode

INSUFFICIENT_EVIDENCE isn't a fallback message — it's a distinct status the rest of the system has to handle. "I don't know yet" is a first-class outcome here, not something papered over by the model's confidence. Every FAQ that does get generated comes back as a DRAFT, never a PUBLISHED answer. A human has to approve it first.
The investigation agent — the part that reconstructs why a problem happened — has an even stricter constraint baked directly into its system prompt:
Important rules:

  1. Do not invent facts.
  2. Do not claim causation from temporal association alone.
  3. Distinguish observations from interpretations.
  4. Cite evidence using the provided memory IDs.
  5. If evidence is weak or conflicting, explicitly say so.
  6. Suggested actions must be recommendations only.
  7. Do not take actions yourself. That second rule is the one I keep coming back to. A drop in repeat questions after an FAQ goes live is evidence, not proof the FAQ caused it — and the system is required to say so rather than assert it. Every piece of evidence is timestamped and sorted before it reaches the model, and the model has to separate interpretation from raw patterns in its own output. If the evidence is weak, it has to say that explicitly in uncertainties instead of filling the gap with a confident guess. What this looks like in practice Here's the actual flow, end to end, as it runs in the app. A recurring question gets flagged from the community feed — the system recognizes that a cluster of messages are all circling the same unresolved problem, and surfaces it with a historical signal count: � From there, selecting "Add FAQ" retrieves the relevant memory and drafts a community-grounded answer. Note the status in the top right — this is still a draft, not published: � Once a moderator reviews it, approval is a single state change, and the state is visible in the UI, not just in a database column no one looks at: � The part that convinced me this wasn't just retrieval theater was the intervention record. After publishing, the system shows exactly which memories the answer was built from — real Hindsight memory IDs, not a summary of "some messages": � That last screen is the one I'd point a skeptical engineer to first. It's the difference between an agent that claims to use your community's history and one that shows its work. Lessons learned Memory is the differentiator, not a feature. Strip out remember and search, and what's left is a chatbot that answers whatever's in front of it. The moment you add persistent memory, the same architecture becomes something that can turn a recurring problem into permanent knowledge. Don't let the model do arithmetic — or make claims — it doesn't need to. Anything that can be computed or verified deterministically (like whether evidence exists at all) should be checked before the model ever gets a chance to guess. Make "I don't have enough information" a real code path, not an afterthought. INSUFFICIENT_EVIDENCE is treated the same as any other status in this system. That one design choice does more for trustworthiness than any amount of prompt tuning. Correlation isn't causation, and the agent should say so — in the prompt, not just in the marketing. The rule against overclaiming isn't something we talked ourselves into after the fact. It's a literal constraint in the system prompt: don't claim causation from temporal association alone. Show the memory IDs, not just the summary. The most convincing screen in the whole app isn't the polished FAQ — it's the one that lists which specific memories the answer came from. If you're building something with Hindsight, surface that evidence somewhere the user can actually see it. Where this goes next The version we built is deliberately narrow — synthetic community data, one end-to-end workflow (recurring question → FAQ → publish → measure), and no live platform integration yet. That scope was intentional; it kept the focus on one question: does persistent memory produce a more trustworthy, more useful response than a system with no memory of what's already been discussed? The interesting part was never the chat interface on top. It was what remembering and recalling actually looked like once they were running, and what the system refused to say when it didn't have enough to go on. Docs for anyone who wants to see how this works under the hood: Hindsight documentation.




Top comments (0)