By M Gnan Ujwal
Software Engineer | Continuum Project
I focused on the production design of Continuum's memory architecture — deciding what should be retained, separating private customer memory from shared knowledge, and designing the Hindsight integration behind a testable service boundary.
I expected the difficult part of giving a support agent memory to be connecting an LLM to a memory system. It wasn't. The harder problem was deciding what information should survive a conversation, who should be allowed to benefit from it, and where that information should live.
That became the central engineering problem behind Continuum, a customer-support agent for a fictional SaaS product called Nimbus Sync. My role focused on the production-design side of the system: memory boundaries, the Hindsight integration, data responsibilities, and the abstractions that keep those pieces testable and replaceable.
Continuum uses Hindsight for agent memory in two different ways. Each customer has a private memory bank containing their environment, previous troubleshooting attempts, unresolved issues, and relevant interaction context. Separately, the system maintains a shared known-issues bank containing distilled technical knowledge about recurring problems across customers.
The important part is not that we stored more information. It is that we gave different information different lifetimes and different scopes.
Memory is not a transcript
My first design constraint was simple: the support transcript and the agent's memory should not be the same thing.
A transcript answers:
What exactly happened?
Memory answers:
What should the agent remember because it may matter later?
Continuum therefore keeps the literal customer, ticket, and message records in SQLite. The database is the system of record for the conversation that the UI renders.
Hindsight has a different job. It stores semantic knowledge that can be recalled when it is useful: previous troubleshooting attempts, resolutions, customer context, and recurring technical issues.
That separation is deliberate. If I treated every chat message as memory, the memory bank would become noisy. If I treated the memory store as the transcript, I would lose the clean, deterministic record of what actually happened.
The architecture is essentially:
Customer message
|
+---------------------> SQLite
| raw transcript
|
v
Hindsight
distilled knowledge
This distinction ended up being one of the most reusable lessons from the project: storage and memory are different responsibilities, even when both contain text.
The more interesting problem: private versus shared knowledge
A support agent needs to remember things about a customer, but a support organization also needs to remember things about the product.
Those are not the same memory.
Continuum gives every customer a bank such as:
customer::{id}
and keeps cross-customer technical knowledge in:
product::known-issues
The customer bank is allowed to contain information such as:
- the customer's environment and plan
- troubleshooting steps they already tried
- previous resolutions
- relevant interaction context
The shared bank has a stricter rule: it contains the technical shape of recurring problems and their resolutions, but not the customer's personal details.
That boundary matters.
Imagine three customers report the same sync failure. The useful organizational memory is:
Recurring issue in sync-engine.
Root cause: sleep-wake-stall.
Resolution: reconnect the session.
It is not:
Marcus Chen reported this from his laptop on Tuesday.
The first is reusable product knowledge. The second belongs to a customer's private history.
This is one reason I think "give the agent memory" is too vague as a design requirement. The real questions are: memory of what, scoped to whom, and reusable by whom?
One reflect() call closes the learning loop
The most important part of the Hindsight integration is the closed loop around reflect().
For each memory-enabled message, Continuum does four things:
1. Retain the customer's message
2. Recall relevant shared known issues
3. Reflect using customer memory + issue context
4. Write the outcome back into memory
The fourth step is what makes the system accumulate knowledge.
The reflect() call returns not just the response text but structured signals. The response schema includes fields such as:
{
"reply": "...",
"sentiment": "frustrated",
"resolved": False,
"root_cause_tag": "sync-timeout-vpn",
"is_known_issue": True,
"module": "sync-engine",
}
The application can then turn those signals into memory.
A resolution gets retained in the customer's bank. If the outcome is identified as a known issue, the technical signature is also retained in the shared bank.
That means the loop looks like this:
customer message
|
v
recall shared known issues
|
v
reflect(customer memory + issue context)
|
+----> response to customer
|
+----> sentiment / resolution / root cause
|
v
retain outcome
|
+------+------+
| |
customer bank shared bank
This is the part I would describe as learning in Continuum. The underlying model is not being retrained after every ticket. Instead, the system changes the context available to future reasoning by retaining useful outcomes.
The next ticket therefore starts with more knowledge than the previous one.
I kept Hindsight behind one boundary
I did not want Hindsight calls scattered throughout the application.
The rest of the application talks to:
MemoryService
and MemoryService talks to the Hindsight SDK.
The adapter exposes operations that make sense to Continuum:
memory.retain_customer_message(...)
memory.recall_known_issues(...)
memory.reflect_reply(...)
memory.retain_resolution(...)
memory.retain_known_issue(...)
This is more useful than exposing generic SDK calls everywhere.
The rest of the application does not need to know how banks, recall parameters, structured outputs, or SDK objects are represented. That knowledge stays in one place.
It also makes testing practical. The memory layer depends on a small HindsightLike protocol rather than requiring the concrete SDK everywhere:
class HindsightLike(Protocol):
def create_bank(self, bank_id: str, **kwargs): ...
def retain(self, bank_id: str, content: Any, **kwargs): ...
def recall(self, bank_id: str, query: str, **kwargs): ...
def reflect(self, bank_id: str, query: str, **kwargs): ...
Our tests provide a fake implementation of that surface.
The result is a useful boundary:
AgentService
|
v
MemoryService
|
v
Hindsight SDK
If the memory provider changes, or the SDK surface changes, I have one integration boundary to revisit rather than an application-wide rewrite.
For developers evaluating agent-memory infrastructure, the Hindsight GitHub repository and Hindsight documentation are useful starting points. The broader idea is also captured well in Vectorize's overview of agent memory.
SQLite still has a job
It would have been tempting to put everything into Hindsight and call the problem solved.
I think that would have made the architecture worse.
SQLite stores:
customers
tickets
messages
The repositories expose operations such as:
ticket_repo.add_message(...)
ticket_repo.list_messages(...)
ticket_repo.set_status(...)
Hindsight stores the semantic memory that the agent needs for future reasoning.
This is a classic separation of concerns:
SQLite
= what happened
Hindsight
= what the agent should remember
That distinction also makes debugging easier. If an agent gives an unexpected answer, I can inspect the literal transcript independently from the semantic memory that was retrieved.
Repository pattern keeps SQL out of the agent
I used repositories for the SQLite side rather than letting the orchestration layer contain SQL.
There are separate CustomerRepository and TicketRepository classes. They receive a database object and expose typed application-level operations.
That means AgentService can focus on the support workflow:
self.tickets.add_message(...)
reply = strategy.respond(...)
self.tickets.set_status(...)
instead of mixing ticket logic with SQL statements.
It sounds like a small architectural decision, but it pays off when the system grows. The service layer describes business behavior; the repository layer describes persistence.
Strategy made the memory comparison honest
Continuum also has a memory-off baseline.
Rather than scattering conditionals throughout the code, I used a common AgentStrategy interface with two implementations:
AgentStrategy
|
+-- WithMemoryStrategy
|
+-- NoMemoryStrategy
The memory-enabled strategy performs the retain → recall → reflect → write-back loop.
The no-memory strategy makes a stateless LLM call and retains nothing.
That distinction matters because the UI can send the exact same customer message through both paths. The comparison is therefore structural rather than a prompt trick.
It also makes the code easier to reason about: the question "what happens when memory is off?" has a concrete implementation instead of being hidden behind branches inside the main agent logic.
Dependency injection made the boundaries testable
FastAPI's dependency injection is what ties the pieces together without hard-coding them into routes.
The route can request an AgentService, while the service receives its ticket repository, memory service, and baseline LLM.
Conceptually:
FastAPI route
|
v
AgentService
| | |
v v v
Repo Memory LLM
For tests, those dependencies can be substituted.
That matters especially for Hindsight. The test suite uses a FakeHindsight implementation, so the memory behavior can be exercised without requiring a live API key or network connection.
What was actually difficult
The difficult part was not writing the API endpoint that sends a message.
It was deciding what should happen after the message.
If the customer says something that sounds important, should I retain the entire message? Which parts are durable facts? Is an issue specific to this customer or systemic? If it is systemic, what can be safely shared? How do I avoid turning the shared bank into a dump of unrelated ticket details?
Those questions pushed the architecture toward explicit missions, directives, structured output, tags, separate banks, and a dedicated memory service.
We also had to make the memory behavior observable. The application exposes metadata such as memories used, directives applied, and known-issue matches. That makes it possible to inspect why a response had the context it did instead of treating memory as an invisible black box.
What I would reuse in another agent
Three lessons stand out.
1. Design memory boundaries before choosing storage
Don't start with "where can I store the chat?"
Start with "what should the agent remember, for how long, and who can use it?"
That naturally leads to different memory scopes and retention rules.
2. Keep semantic memory separate from the system of record
A memory system is optimized for helping future reasoning. A database transcript is optimized for accurately recording what happened.
Trying to make one system perform both jobs can make both less useful.
3. Put the memory provider behind an application-level interface
Agent code should ask for things like "recall known issues" and "retain this resolution," not manipulate provider-specific primitives everywhere.
That makes the integration easier to test, replace, and evolve.
4. Make the learning signal structured
Returning a response alone makes it difficult to build a reliable write-back loop.
Returning the response plus explicit fields such as resolved, root_cause_tag, and is_known_issue gives the application something concrete to retain and act on.
5. Treat privacy boundaries as architecture
Private customer memory and shared organizational memory should not be separated only by a prompt instruction. They should be separated by the data model itself.
In Continuum, that separation is represented directly by different Hindsight banks.
The takeaway
The biggest lesson I took from building Continuum is that agent memory is not primarily a storage feature.
It is a design decision about what survives a conversation.
Once I treated memory that way, the rest of the architecture became clearer: SQLite records the conversation, Hindsight stores distilled knowledge, customer banks preserve private context, the shared bank captures systemic issues, MemoryService isolates the provider, repositories isolate persistence, and strategies make memory-on versus memory-off behavior explicit.
The result is a support agent that does more than remember previous text. It accumulates operational knowledge while keeping that knowledge scoped to the people and problems it belongs to.
That is the part of agent memory I think is worth designing carefully: not how much an agent can remember, but whether it remembers the right things.
Top comments (0)