A sales agent can have access to months of customer history and still know surprisingly little about the deal.
Imagine a salesperson preparing for a call with the CFO of a company they have been pursuing for six months. The CRM contains meeting notes, previous objections, stakeholder comments, pricing discussions, competitor mentions, and call summaries.
The information is there.
But the salesperson still has to piece together the story before every important call.
What changed? Which objection was resolved? Who is actually supporting the deal? What did the CFO ask for last time? Did the previous pricing approach work?
My first instinct was to make the agent see more of that history.
That was the wrong solution.
The problem was not a lack of context. It was a lack of memory.
The problem with giving an agent more context
Suppose the agent is asked:
"Brief me for my next call with the CFO."
The obvious approach is to collect everything known about the deal and put it into the prompt.
For a small deal, this works.
For a six-month deal, the context starts looking more like this:
January:
CFO: "The pricing is too high."
March:
CTO: "The product looks good, but integration is a concern."
April:
Competitor X is introduced.
May:
Customer receives additional budget.
June:
CFO: "Budget is no longer the main issue."
July:
IT team: "We need more information about integration."
Nothing here is useless.
But the information has different importance depending on when the question is asked.
If the next call is about integration, the July IT discussion matters more than the January pricing objection.
If the CFO has already said budget is no longer the main issue, bringing up the original pricing objection without that later update can give the agent the wrong picture of the deal.
The model now has another job besides answering the question.
It has to reconstruct the state of the deal from a pile of historical information.
That was the point where I stopped thinking about the problem as "How do I give the model more context?"
I started thinking:
"How do I make the agent remember what happened?"
Context and memory solve different problems
The distinction is simple.
Context is information the model receives for the current request.
Memory is information the application keeps across requests and can retrieve when it becomes relevant.
I wanted the system to work more like this:
Customer interaction
|
v
retain()
|
v
Persistent deal memory
|
| later
v
recall()
|
v
Relevant deal history
|
+---- Current question
|
v
LLM
|
v
Deal-specific answer
Instead of repeatedly sending the entire deal history to the model, the application stores the history and recalls relevant information when a new question arrives.
For the memory layer, I used Hindsight Cloud through hindsight-client.
The LLM remains responsible for reasoning.
Hindsight provides the persistent memory.
That separation became the core of my architecture.
I built the memory around the deal
A sales conversation does not exist in isolation.
A pricing objection from one customer should not become context for another customer. A concern raised by one stakeholder should not automatically become a concern for every stakeholder in the organization.
So the application is organized around individual deals.
A salesperson selects a deal and asks a question about it.
When an interaction happens, the relevant outcome can be retained against that deal:
client.retain(
content=call_outcome,
context=f"deal:{deal_id}"
)
Later, when the salesperson asks a question, the application recalls information for that same deal:
memories = client.recall(
query=user_query,
context=f"deal:{deal_id}"
)
The recalled memories are then used as context for the LLM.
This is a small architectural change, but it changes how the agent behaves.
The model no longer needs to be handed the entire history every time.
It gets the current question and the historical information that was retrieved for that question.
The same model, with and without memory
I wanted the difference to be easy to see, so I added a Use Memory toggle to the application.
The user can ask exactly the same question with memory enabled or disabled.
For example:
"Brief me for my next call with the CFO."
Without memory, the model can still answer.
It might say:
Prepare to discuss pricing, ROI, implementation costs, and likely objections.
That is reasonable advice.
It is also advice that could apply to almost any sales call.
Now suppose the memory layer retrieves:
CFO rejected the previous 15% discount.
CTO supports the product.
Competitor X is offering a lower upfront price.
CFO requested a three-year cost comparison.
The previous pricing discussion did not move the deal forward.
The model can now produce something much more specific:
The CFO previously rejected a 15% discount, so repeating the same pricing approach may not address the current concern. The last unresolved request was a three-year cost comparison. The CTO is already supportive, while Competitor X is competing on lower upfront cost.
The interesting part is that I did not need a different model to get this behavior.
The question stayed the same.
The model stayed the same.
The additional capability came from the information retrieved before reasoning.
That was the clearest demonstration of why memory mattered.
I was not training the model
There is an important distinction here.
The agent is not learning by changing the weights of GPT-OSS-120B.
I used openai/gpt-oss-120b through Groq's API with the OpenAI Python client.
When a new call happens, I do not fine-tune or retrain the model.
Instead, I add the new information to the external memory layer.
The model provides reasoning.
The memory provides accumulated experience.
That means a new customer interaction can affect future answers without changing the underlying model.
Conceptually:
GPT-OSS-120B
Reasoning layer
|
v
"What should I say now?"
^
|
Relevant memory
^
|
Hindsight
^
|
Past outcomes
The model is still stateless in the traditional sense.
The application is what gives the overall agent continuity.
The interesting part is what not to remember
Once I started using persistent memory, another problem appeared.
It is tempting to think that more memory must be better memory.
It is not.
Suppose a deal has 200 recorded interactions.
I could store all 200.
But if every query retrieves everything, I have effectively rebuilt the original context problem.
The agent would technically remember the deal.
It would just be buried under the deal.
A useful memory system needs selective recall.
For example:
January
Pricing is the main objection.
April
Customer receives additional budget.
June
Pricing is no longer blocking the deal.
July
Integration becomes the main concern.
If the salesperson asks:
"What is currently blocking this deal?"
The January pricing objection is useful historical evidence.
But it should not automatically be treated as the current blocker.
This is where memory becomes more interesting than simply storing chat history.
The agent needs to reason over change.
The important question is no longer just:
"What did the customer say?"
It becomes:
"What did the customer say, when did they say it, and what happened afterward?"
The limitation: correct memories can still produce the wrong answer
This was the main limitation I had to account for.
A memory can be completely accurate and still be wrong for the current situation.
Consider:
January:
"Your product is too expensive."
April:
"Budget is no longer an issue."
June:
"Integration is our biggest concern."
All three statements are true.
If the agent retrieves only the January statement when preparing for the June call, the LLM can produce a convincing answer based on outdated information.
The LLM did not necessarily hallucinate.
The retrieved context was incomplete.
That changed how I evaluated the memory layer.
I could not simply ask:
"Did the agent remember something?"
I needed to ask:
"Did the agent remember the right thing for this question?"
This is an important difference.
Memory does not eliminate the context problem.
It moves the context problem into retrieval.
Instead of trying to fit an ever-growing history into the prompt, I now have to build a system that can find the useful parts of that history.
That is a better architecture for this use case, but it is still an engineering problem.
The agent gets better history after every call
A memory system is only useful if it keeps changing with the relationship.
That is why the application includes a Log Call Outcome feature.
After a call, the salesperson can record something like:
"CFO accepted the three-year comparison
but requested detailed implementation pricing."
That outcome can be retained as new memory for the deal.
The next briefing can then use it.
The overall loop becomes:
Past interactions
|
v
recall()
|
v
Pre-call briefing
|
v
Customer call
|
v
Log outcome
|
v
retain()
|
v
Updated deal memory
|
+------------------+
|
v
Next briefing
This was an important shift in how I thought about the agent.
A normal chat response ends when the answer is generated.
Here, the outcome of one interaction can become input to a future interaction.
The application accumulates experience without retraining the model.
I made the retrieved memory visible
There is also a practical debugging problem.
If the final answer is bad, I need to know why.
Was the wrong information retrieved?
Was the right information retrieved but misunderstood?
Was nothing useful retrieved at all?
So the application shows the memories used for the response.
That creates a simple debugging boundary:
Bad answer
|
v
What was retrieved?
|
+----------------------+
| |
v v
Wrong / stale memory Relevant memory
| |
v v
Retrieval problem Reasoning problem
If the system retrieves an irrelevant old objection, I know to investigate memory retrieval.
If the right memories are present but the model reaches the wrong conclusion, I can investigate the reasoning prompt or generation step.
For an agent with long-term memory, inspectability matters because the memory is part of the reasoning pipeline.
The architecture ended up being fairly small
The final system uses Python and Streamlit for the application, Hindsight for persistent memory, and GPT-OSS-120B through Groq for generation.
The overall flow is:
┌─────────────────┐
│ Streamlit │
│ UI │
└────────┬────────┘
|
v
┌─────────────────┐
│ Select Deal │
└────────┬────────┘
|
┌─────────┴─────────┐
| |
v v
Hindsight Query
recall() |
| |
└─────────┬─────────┘
v
┌─────────────────┐
│ GPT-OSS-120B │
│ via Groq │
└────────┬────────┘
|
v
Deal briefing
|
v
Call outcome
|
v
Hindsight
retain()
The architecture is not complicated.
The important decision was deciding what each component should be responsible for.
The LLM reasons about the current situation.
Hindsight preserves the history.
The application decides when to retrieve that history and how to put it in front of the model.
What I learned
The first lesson was that more context does not automatically mean more understanding.
A model can have access to a large amount of information and still struggle to identify what matters now.
The second was that memory is not just a bigger context window.
Context is what the model sees for one request. Memory is information that can persist beyond that request and be retrieved later.
The third was that memory quality depends on retrieval.
A system that retrieves stale or irrelevant information can make a strong LLM produce a very convincing answer for the wrong situation.
The fourth was that external memory can give an agent continuity without retraining the underlying model after every interaction.
And the biggest lesson was this:
I started by asking how much more context I could give my sales agent.
The better question was what it actually needed to remember.
Once I treated memory as a separate part of the system, the architecture became much clearer.
The LLM did not need to remember the entire customer relationship.
It needed the right pieces of that relationship at the right time.
Top comments (0)