DEV Community

Vemula Srinath
Vemula Srinath

Posted on

I Gave an AI Support Agent Two Memories with Hindsight

The first time an AI support agent tells a customer to restart an application, that can be reasonable. The tenth time it gives the same advice to the same customer, it starts to feel broken.

That was the problem I wanted to solve with Continuum, a support agent for the fictional SaaS product Nimbus Sync. The interesting part wasn't making another chatbot that could answer support questions. It was deciding what the agent should remember—and at what level.

I ended up giving it two kinds of memory with Hindsight: one for the individual customer, and another for problems that appear across customers.

That distinction changed the way I thought about support memory.

Continuum two-level memory architecture

One conversation can contain two kinds of knowledge

A support conversation can tell us something about a customer:

  • what environment they use
  • what they have already tried
  • what happened in previous tickets
  • how they prefer to communicate
  • whether a previous workaround actually solved the problem

But the same conversation can also tell us something about the product.

If several customers report similar symptoms, the useful information is no longer only about one customer. It may be evidence of a recurring or systemic issue.

Those are different memories.

Customer memory answers:

What do I know about this customer?

Shared memory answers:

What do I know about this problem across customers?

Continuum keeps those concerns separate using Hindsight.

For the customer-specific side, the agent maintains a private memory bank containing the customer's relevant history and environment.

For the shared side, it maintains a common bank containing knowledge about recurring issues and patterns across the customer base.

The separation matters because I don't want a customer's private context to become everybody's context. At the same time, I don't want every new customer to start from zero when the support system has already seen the same failure repeatedly.

The support loop

Memory isn't bolted onto the side of the agent. It is part of the response loop.

The actual WithMemoryStrategy makes that explicit:

def respond(self, customer_id: str, ticket_id: str, message: str) -> AgentReply:
    self.memory.retain_customer_message(customer_id, ticket_id, message)

    known = self.memory.recall_known_issues(message)

    outcome: ReflectOutcome = self.memory.reflect_reply(
        customer_id, message, known
    )

    self.memory.retain_resolution(customer_id, ticket_id, outcome)
    if outcome.is_known_issue:
        self.memory.retain_known_issue(outcome, ticket_id)
Enter fullscreen mode Exit fullscreen mode

There is a useful sequence here:

new message
    ↓
retain customer context
    ↓
recall known issues
    ↓
reflect with memory + directives
    ↓
return response + structured signals
    ↓
retain the outcome
    ↓
if it is a known issue, update shared memory
Enter fullscreen mode Exit fullscreen mode

That last step is the part I care about most. The agent isn't just reading memory. It writes the outcome of the interaction back into memory.

Why one memory layer wasn't enough

Imagine a customer says:

"My files aren't syncing again. I already restarted the application."

A stateless support agent has very little to work with. It might respond with the same troubleshooting step:

"Try restarting the application."

But customer-specific memory can tell the agent that this person has already tried that.

Now consider a different customer reporting:

"My sync queue has been stuck since yesterday."

Suppose the support system has seen similar reports from several other customers. Shared memory can contain the emerging pattern:

Product: Nimbus Sync
Module: sync engine
Symptom: sync queue stops processing
Observed across: multiple customers
Known resolution/workaround: retained from previous tickets
Enter fullscreen mode Exit fullscreen mode

The second customer does not need to know anything about the first customer. The useful knowledge is about the problem itself.

This is the difference between remembering a customer and remembering a problem.

What Hindsight is actually doing

I put Hindsight behind a MemoryService adapter so the rest of the application doesn't need to talk directly to the Hindsight SDK.

The important operation is reflect():

response = self.client.reflect(
    bank_id=self.settings.customer_bank_id(customer_id),
    query=message,
    budget="mid",
    context="\n".join(context_lines),
    response_schema=AGENT_REFLECT_SCHEMA,
    apply_all_directives=True,
)
Enter fullscreen mode Exit fullscreen mode

The call is doing more than generating prose.

The response schema asks for structured information alongside the customer-facing reply. In the current implementation, that schema includes sentiment, resolved, root_cause_tag, is_known_issue, and module.

AGENT_REFLECT_SCHEMA = {
    "type": "object",
    "properties": {
        "reply": {"type": "string"},
        "sentiment": {
            "type": "string",
            "enum": ["positive", "neutral", "frustrated", "angry"],
        },
        "resolved": {"type": "boolean"},
        "root_cause_tag": {"type": ["string", "null"]},
        "is_known_issue": {"type": "boolean"},
        "module": {"type": ["string", "null"]},
    },
    "required": ["reply", "sentiment", "resolved", "is_known_issue"],
}
Enter fullscreen mode Exit fullscreen mode

The response goes back to the customer.

The structured fields go back into the application and memory system.

That makes the LLM call part of a memory pipeline rather than treating it as a simple text-generation endpoint.

A small naming detail that matters

The design brief for this work describes the systemic signal as is_systemic. In the current repository, I implemented the same boundary as is_known_issue.

I prefer being explicit about that distinction because the code is the source of truth.

is_known_issue is defined as:

true if this matches a systemic issue affecting multiple customers, not a one-off.

So the conceptual flow is still:

one customer's ticket
        ↓
does this match a broader recurring problem?
        ↓
       yes
        ↓
shared / known-issue memory
Enter fullscreen mode Exit fullscreen mode

But the actual field name in the implementation is is_known_issue, not is_systemic.

Keeping private and shared knowledge separate

The two banks have different missions.

The customer bank is instructed to remember things such as the customer's environment, previous troubleshooting attempts, resolutions, and sentiment.

The shared bank has a different boundary: it should retain the technical shape of recurring issues and their resolutions, without storing an individual customer's personal details.

That distinction is encoded directly in the Hindsight configuration:

SHARED_BANK_MISSION = (
    "You are the collective memory of every support ticket ever resolved for "
    "{product}, across all customers. Remember which bugs and issues are "
    "recurring, which modules they affect, and what fix or workaround "
    "resolved them. Do not store any individual customer's personal details "
    "here — only the technical shape of the issue and its resolution."
)
Enter fullscreen mode Exit fullscreen mode

I like this because the privacy boundary isn't merely an assumption in application code. It is part of the memory bank's mission.

A concrete before and after

Without memory, two similar tickets can look almost identical to the agent:

Ticket 1:
"My files aren't syncing."

Ticket 2:
"My files aren't syncing."
Enter fullscreen mode Exit fullscreen mode

The model has to rediscover the context each time.

With the two memory levels, the second interaction can have additional information available:

Customer memory:
- This customer has already tried restarting.
- They previously reinstalled the application.

Shared memory:
- Other customers have reported similar sync failures.
- A recurring issue has already been associated with the sync module.
Enter fullscreen mode Exit fullscreen mode

Now the agent has two different reasons to change its response.

Customer memory prevents it from repeating the customer's own failed troubleshooting steps.

Shared memory helps it recognize that the issue may not be isolated.

That is a much more useful form of continuity than simply replaying the last few chat messages.

The closed learning loop

The repository makes the distinction between memory and no memory structural.

There are two strategies:

class NoMemoryStrategy(AgentStrategy):
    mode = "off"

    def respond(self, customer_id: str, ticket_id: str, message: str) -> AgentReply:
        text = self.baseline.reply(message)
        return AgentReply(
            text=text,
            memory_mode=self.mode,
            meta={"note": "stateless — no memory used"},
        )
Enter fullscreen mode Exit fullscreen mode

The memory-backed strategy follows the retain → recall → reflect → retain loop.

That gives me a useful comparison: the exact same support input can be handled with memory on or with a stateless baseline.

I don't have to claim that memory is better because it sounds better. The application can show what context was used, how many memories and directives were involved, and which known issues matched.

The hard part: deciding what deserves to be remembered

The most interesting engineering problem here isn't storing information. Storage is easy.

The harder question is deciding what information is worth retaining.

If every sentence becomes a permanent memory, the memory layer becomes noisy.

If nothing is retained, the agent remains stateless.

The useful middle ground is to retain distilled information: outcomes, recurring issues, customer context, and other facts that can change how a future support interaction should be handled.

That is also why the shared bank shouldn't simply be a copy of every customer conversation.

It should contain knowledge that has crossed the boundary from:

"This happened to one customer."

to:

"We've seen this problem across customers."

The known-issue signal is part of that boundary.

One limitation I would not hide

Identifying a systemic issue is not the same thing as proving one.

A single unusual ticket can resemble an existing problem. Conversely, several customers can describe the same symptom while having different root causes.

That means is_known_issue should be treated as a structured signal used by the system, not as an infallible diagnosis.

The memory layer can accumulate evidence, but the quality of that memory still depends on the quality of the agent's interpretation and the policies around retention.

That is one reason I like keeping memory operations behind a dedicated service boundary. It gives me a clear place to improve retrieval, retention, and validation rules without spreading memory logic throughout the application.

Continuum support agent architecture showing customer and shared memory with Hindsight

What I learned

1. Memory needs scope

"Remember everything" isn't a useful memory strategy.

Some information belongs to one customer. Some belongs to the organization. Making that distinction explicit makes the system easier to reason about.

2. Retrieval and retention are equally important

Remembering something later is only useful if the right thing was retained in the first place.

The memory loop therefore has two sides:

retain → recall → use → retain again
Enter fullscreen mode Exit fullscreen mode

Ignoring either side weakens the system.

3. Structured output makes memory actionable

A response string is difficult for application code to reason about.

Signals such as resolved, root_cause_tag, and is_known_issue create explicit hooks between model output and application behavior.

4. Shared memory changes the unit of learning

Customer memory is about continuity for one person.

Shared memory lets the support system accumulate knowledge about recurring product behavior.

That changes the unit of memory from "this conversation" to "this problem pattern."

5. More memory isn't automatically better

The goal isn't to give an agent the largest possible history.

The goal is to give it the right context for the current decision.

That's the part I'm most interested in continuing to explore with Continuum: not how much an agent can remember, but how deliberately it can decide what belongs in memory—and who should benefit from that memory.

If you're interested in the memory layer itself, start with the Hindsight GitHub repository, the Hindsight documentation, and Vectorize's explanation of agent memory.

Top comments (0)