DEV Community

Likitha
Likitha

Posted on

The Problem With Making an Agent Remember Everything

Imagine you are a sales representative preparing for a call with the CFO of a company you have been pursuing for six months.
You know you have spoken to them before. You remember there was a pricing objection. You remember another stakeholder liked the product. You vaguely remember a competitor being mentioned. But before the call, you still have to open the CRM, search through old notes, read meeting summaries, and look through emails to reconstruct what actually happened.
The frustrating part is that the information is not missing.
It is just scattered.
A long sales cycle can produce hundreds of individual interactions. A customer can object to pricing in January, bring up integration in March, add a new decision-maker in May, and change their budget situation in June. By the next call, the salesperson is not simply looking for information. They are trying to reconstruct how the deal changed.
That is the problem I wanted to solve with my Deal Intelligence Agent.
I wanted a salesperson to be able to ask something as simple as:
"Brief me for my next call with the CFO."
and get an answer based on the history of that specific deal.
My first instinct was to solve this by giving the model more context.
That turned out to be the wrong abstraction.
The obvious solution: give the model the whole history
If the model needs to know what happened six months ago, why not just give it six months of history?
I could take CRM notes, meeting summaries, emails, and previous conversations and put them into the prompt. The LLM would then have everything it needed to answer the question.
At first, this sounds reasonable.
But a longer context creates another problem: not everything in the history is relevant to the current question.
Consider a deal with this timeline:
January
CFO: "Your pricing is too high."

March
CTO: "The product looks good, but integration is a concern."

May
Customer receives additional funding.

June
CFO: "Budget is no longer the main issue."

July
IT team: "We need more information about integration."
All five events are useful historical facts.
But if the salesperson is preparing for the July call, they do not need five equally weighted facts. They need to understand that the original pricing objection has changed, the CTO is already positive, and integration is now the active concern.
Giving the model more context does not automatically give it better context.
I was treating the problem as a context-window problem when it was really a memory problem.
I changed the architecture instead of making the prompt bigger
Instead of continuously feeding the LLM more historical information, I moved the long-term state outside the model.
I use Hindsight as the long-term memory layer.
The basic flow is:
Customer interaction
|
v
retain()
|
v
Persistent memory
|
| later
v
recall()
|
v
Relevant deal history
|
+------ Current query
|
v
GPT-OSS-120B
|
v
Deal-specific answer
The LLM remains responsible for reasoning about the current question.
Hindsight provides the persistent memory that allows information from previous interactions to become available later.
That separation is important.
I do not need the model itself to permanently remember every conversation. I need the application to preserve the experience and retrieve the relevant parts when a future question requires them.
The memory layer therefore becomes the bridge between two otherwise separate interactions.
The deal becomes the memory boundary
A sales conversation only makes sense in context.
The same sentence can mean very different things depending on the customer, stakeholder, and stage of the deal.
That is why the application is organized around individual deals. A salesperson selects the deal they are working on and asks questions about it.
When a new interaction or call outcome is retained, it is associated with that deal.
Conceptually, the operation looks like this:
client.retain(
content=call_outcome,
context=f"deal:{deal_id}"
)
The important part is not the number of lines of code.
It is the fact that the interaction becomes part of persistent deal-specific memory instead of remaining only in the current conversation.
Later, a query such as:
"Brief me for the next call with the CFO"
can trigger a recall operation for that deal.
The model then receives the current question together with the relevant memories.
The same question with and without memory
I wanted to make the difference between the two approaches visible, so the application includes a Use Memory toggle.
With memory disabled, the query is sent directly to the LLM.
With memory enabled, the application first recalls relevant memories and then gives those memories to the LLM along with the query.
The question stays the same.
For example:
"Brief me for my next call with the CFO."
Without memory, the model can only provide general sales advice:
Prepare to discuss pricing, ROI, implementation costs, and likely objections.
That answer is not wrong.
It is simply generic.
Now suppose the memory layer retrieves:
CFO rejected the previous 15% discount.

CTO supports the product.

Competitor X is offering a lower upfront price.

CFO requested a three-year cost comparison.

The previous pricing discussion did not move the deal forward.
The model can now produce a much more deal-specific briefing:
The CFO previously rejected a 15% discount, so repeating the same pricing approach may not address the current concern. The last unresolved request was a three-year cost comparison. The CTO is already supportive, while Finance remains concerned about implementation cost. Competitor X is also offering a lower upfront price.
The model did not change.
The query did not change.
The difference was the memory available to the model.
That became the main design principle behind the system.
Remembering everything is not the goal
This is where the problem became more interesting.
It is easy to confuse storing information with giving an agent useful memory.
Suppose a deal has accumulated 200 interactions.
I could retrieve all 200 and pass them to the LLM. Technically, the agent would have more information.
Practically, I would have recreated the same problem I started with.
The salesperson wanted a briefing, not a transcript archive.
The agent therefore needs selective recall.
For a question about the CFO, useful memories might include previous CFO objections, pricing discussions, financial concerns, recent requests, competitors, and outcomes of previous approaches.
A technical conversation from six months ago might be completely irrelevant.
This is where the memory layer matters.
Hindsight gives me a place to retain the history without forcing the entire history into every prompt. When a question arrives, the application can recall memories relevant to that situation and use those memories as context for the LLM.
The goal is not maximum context.
The goal is useful context.
The harder problem: old information can still be true
The most important limitation I encountered in thinking about this architecture is that historical information does not become false just because it becomes old.
That makes memory surprisingly difficult.
Suppose the customer says:
January:
"Your product is too expensive."

April:
"Budget is no longer the issue."

June:
"Integration is our biggest concern."
If I delete the January statement, I lose useful history.
If I retrieve only the January statement because it matches a query about objections, I can give the model misleading context.
If I retrieve all three statements and treat them equally, I make the model responsible for deciding which one represents the current state.
This is the dead end I wanted to avoid: building a system that technically remembers everything but repeatedly brings stale information back into the conversation.
That means persistent memory does not eliminate the context problem. It changes where the problem lives.
The engineering challenge moves from:
"How do I fit the entire history into the prompt?"
to:
"How do I retrieve the right history for this question?"
That is a much better problem to have, but it is still a problem.
It also means I should not describe the system as simply "learning what the customer wants."
The system is accumulating evidence over time. The LLM still has to reason about that evidence, including conflicting or outdated information.
That distinction matters.
The memory also needs to grow after every call
A useful long-term memory cannot stop at the initial CRM history.
The application includes a Log Call Outcome feature so that a salesperson can add what happened after a conversation.
For example:
"CFO accepted the three-year comparison but requested detailed implementation pricing."
That outcome is retained against the current deal.
The next time the salesperson asks for a briefing, the new information can become part of the available memory.
The loop becomes:
Past interactions
|
v
recall()
|
v
Pre-call briefing
|
v
Customer call
|
v
Call outcome
|
v
retain()
|
v
Updated deal memory
|
+------------------+
|
v
Next briefing
This is not model training.
I am not changing the weights of GPT-OSS-120B after every customer conversation.
The model stays the reasoning engine. The external memory changes.
That gives the agent a way to accumulate experience without retraining the underlying model every time a new event occurs.
I wanted the memory to be inspectable
There is another problem with agent memory: if the final answer is wrong, it is difficult to know why.
The application therefore shows the memories that were retrieved for an answer.
That gives me a useful debugging boundary.
Final answer is wrong
|
v
Were the retrieved memories wrong?
|
+--+--+
| |
Yes No
| |
v v
Memory/ LLM
retrieval reasoning
problem problem
If the agent retrieves an irrelevant old objection, changing the LLM prompt is unlikely to fix the underlying problem.
If the right memories were retrieved but the model drew the wrong conclusion, the problem is somewhere else.
For an agent with long-term memory, being able to inspect the retrieved evidence is almost as important as generating the final answer.
What I learned
The first lesson was that more context is not automatically better context.
A larger prompt can contain more information while making it harder to identify what actually matters.
The second was that memory and context are different things.
Context is what the model sees for the current request. Memory is information that persists beyond the current request and can be retrieved later.
The third was that memory quality matters as much as memory quantity.
A system that retrieves the wrong historical information can make a strong LLM produce a convincing but poorly grounded answer.
The fourth was that an agent can accumulate experience without continuously retraining its model.
For this use case, the important state is the history of the relationship, not a new set of model weights after every conversation.
The final lesson was the one that changed the architecture:
I did not actually want my agent to remember everything.
I wanted it to remember enough to understand what changed.
That is why I stopped feeding the agent more and more context and gave it a persistent memory layer instead.

Top comments (0)