After researching along the DSH and TencentDB Agent Memory line, my understanding of 'agent memory' has changed.
Initially, I mostly cared about whether things could be connected: whether conversations could be retained, and whether a new session could still find earlier content. Later, as I kept using Hermes and studied projects such as Mem0 and Hindsight, the question gradually became: how should knowledge bases, conversation history, and long-term memory divide responsibilities?
They can all store information and potentially feed it to the model. But in an agent that keeps working over time, what is stored, when it is retrieved, who uses it, and how stale conclusions are corrected directly affect the next task.
Earlier, Part 05 discussed how a project continues. Part 06 investigated the memory retrieval path. This article moves one level up and puts those studies together into my current understanding of a memory system.
First, separate where the content in front of the model comes from
I now look at four things.
Current context is the information the model actually receives for this inference. The current question, recent conversation, tool results, and selected memories may all be inside it. No matter how long the context window is, someone still has to decide what goes into it.
Conversation history keeps 'what was said and done at the time.' It is useful for retracing the process, but it also contains temporary guesses, failed attempts, and abandoned plans. Keeping the full history does not guarantee that the next session will retrieve a useful conclusion.
Knowledge base mainly stores material that can be checked against the original text: docs, tutorials, code, and study notes. It helps the agent answer 'what is the basis for this?'
Long-term memory is more concerned with preferences, agreements, experience, and related facts formed over continued interaction. It helps the agent understand 'how is this task related to the past?'
These types of information can use similar retrieval techniques, and their boundaries overlap. I care more about what responsibility each one has in a task: the knowledge base keeps sources, history keeps process, long-term memory keeps information worth reusing, and the current context finally brings relevant pieces to the model.
For example, official documentation that says how to call an interface is knowledge-based evidence. An environment that once failed because of a specific parameter, and how it was later resolved, is conditional experience. If the parameter documentation is updated, that experience may need to be rechecked. Using both together is more reliable than only remembering 'this configuration works.'
Hermes: observing memory in an agent that does real work
Hermes made it easier to observe these layers because it is first an agent that handles tasks and calls tools.
Its built-in memory uses USER.md and MEMORY.md: the former stores user preferences, while the latter stores environment facts, agreements, and learning points. It can also search past conversations through session_search. Key records and full history therefore have different retrieval entrances. Hermes memory docs
Built-in memory is not an infinitely growing warehouse. These two files have capacity constraints. At the start of a session, they enter the system prompt as a fixed snapshot; updates written during the session enter that part of the prompt at the next session start.
This design showed me one division of labor: a small amount of frequently used information is carried into the session, while details are searched when needed. Larger knowledge material can live in a separate knowledge base and be retrieved by tools according to the task.
As for connecting an external memory service, I still need to see how the client calls it: when it writes, when it recalls, and how the returned content enters the current context. A service name in the configuration only means an integration setting exists. To judge the actual effect, I need to trace one write and the next retrieval.
Mem0: making memory a pluggable capability
When studying Mem0, I focused on a different position: it treats memory as a capability layer provided to the application or agent above it.
The official example shows a straightforward path: first search related memories by the current question and user scope, add the results to the model input, complete the conversation, then hand new content to the memory layer for processing. It provides an SDK and service interfaces, and distinguishes between the open-source library, self-hosted deployment, and managed service. Mem0 project
Along this line, an agent such as Hermes can be responsible for understanding the task and using tools, while the external memory layer handles retention and recall, and the knowledge base continues to store raw material. This is my understanding of the responsibility split; the actual combination still depends on the real versions and adaptation.
What really needs to be thought through is the write rule.
For example, should 'the user explicitly changed a long-term preference' and 'the model temporarily guessed what the user might like' be stored in the same way? After a successful operation, should we record only 'it worked,' or also keep the applicable environment and verification result?
If a guess is already turned into a fact at write time, then no matter how accurate later retrieval is, it only retrieves the wrong content more reliably. So when I look at a memory layer, I look at both writes and recalls, not only at which database it uses.
Hindsight: after saving experience, how can it be organized further
Hindsight is another line I looked at while studying Hermes-related work.
It splits core operations into retain, recall, and reflect: organize information at write time, retrieve it when needed, then analyze based on existing memories. The official material also distinguishes external facts, the agent's experiences, and observations summarized from multiple memories. Hindsight project
What interests me is the extra layer of 'organizing experience.' After several similar problems, can the system extract reusable insights from scattered records instead of re-reading raw history every time?
But summarization also introduces new judgments. Three failures in the same environment does not prove the environment was the cause; one method working does not prove it suits every situation.
So for me, organized experience still needs to retain its evidence and applicable conditions. Facts, inferences, and suggestions should be distinguishable. When new evidence appears, there should also be a way to correct old conclusions.
This also clarifies the relationship with the knowledge base: source texts and records provide evidence, the memory system helps organize experience, and the agent uses them in the current task.
DSH and TencentDB Agent Memory: after integration, keep looking at layers and retrieval
On the DSH and TencentDB Agent Memory line, I did some integration and debugging. The most direct lesson was already written in Part 06: a normal chat response does not prove that the proactive retrieval tool is working too.
Later, after reading source code and organizing notes, I started paying attention to how the system divides work internally.
In the MemoryCore implementation I studied, L0 keeps raw conversation, L1 extracts atomic memories, L2 organizes scene information, and L3 summarizes user profiles. Different layers let raw records and organized information take on separate jobs: tracing, retrieval, and context organization. MemoryCore notes
At the same time, memory and knowledge material have different service and index paths. MemoryCore manages memory, while MemoryKnowledge handles knowledge content such as wikis and code graphs. DSH, as the client that executes tasks, needs to access them through the corresponding entry points.
This made one point clearer: finishing a knowledge index is only one segment. After that, there are still asset registration, visibility scope, tool discovery, request endpoints, and getting results into the current context.
From a memory-system perspective, these parts decide whether information that already exists can be used in the right task. When debugging, if I only check whether content exists in the database, it is easy to miss problems on the consumer side.
How I would organize a memory system now
Putting these projects together, I tend to draw the responsibilities first, then choose components.
Figure 1 | The read/write division I currently understand: retrieve information by task, then organize reusable memory from the results.
Step one, keep raw material that can be checked. Documents, code, and history records have their own sources and versions. Later conclusions should be able to return to their sources.
Step two, decide what deserves to become long-term memory. Stable preferences, explicit decisions, and repeatedly verified experience are more suitable for long-term retrieval than whole chat logs. Unverified inferences can be kept, but whoever uses them later should be able to recognize their nature.
Step three, organize the current context by task. Common agreements can be brought in ahead of time. For a specific question, retrieve knowledge and related experience, then check the source, time, and applicable scope. Not everything retrieved should be handed to the model.
Step four, let task results correct the memory in return. When a new decision replaces an old one, history can remain, but the next retrieval should be able to recognize what is currently valid. Failed experience should also record its conditions, so a local conclusion does not become a permanent ban.
One more thread runs through all of this: memory needs ownership. Whether content formed by one user, role, or project can be read by another scope needs explicit rules. Using the same storage does not mean all content should be shared.
This does not mean stacking Mem0, Hindsight, and TencentDB Agent Memory together. They have overlapping responsibilities. If multiple components automatically distill the same conversation without a consistent update rule, they may leave behind conflicting conclusions. First decide clearly who owns which segment, then verify whether the combination solves the actual problem.
How I will judge whether it works
In the next stage, I want to check memory effects around concrete behavior.
After switching to a different session, can it retrieve preferences and agreements that were already explicit? When asked about a technical conclusion, can it locate the related original text? After new information overturns an old conclusion, can it adopt the updated judgment? When switching to another user or project, can it keep the visibility boundary?
These checks correspond to persistence, evidence, updating, and ownership. As for which solution fits my way of working better, I still need to keep verifying in the same tasks. Project introductions and individual performance numbers alone are not enough to decide.
By now, I increasingly care about one complete path: how information is kept, how it is organized, how it is found, and how it gets corrected in later tasks.
When that path is clear, the knowledge base, long-term memory, and agent can each play their role. An agent's continuity can then move from 'remembering some words' toward 'continuing to work with experience.'
I'm Aiclaw, recording my learning process around AI tools, agent integration, and engineering practice.
More notes: Aiclaw's blog


Top comments (0)