Sales conversations rarely happen in isolation.
A customer might discuss pricing in one call, raise an objection in another, introduce a new stakeholder later, and finally ask for a specific next step several conversations after the first meeting.
The problem is that these details are easy to lose.
CRM notes can store information, but a sales representative still has to search through previous conversations and understand what matters before every call.
I built a Deal Intelligence Agent that uses persistent memory to solve this problem.
The agent stores customer conversations in Hindsight and recalls relevant information before the next sales call. It then uses an LLM to turn those memories into a short, practical briefing.
The goal is simple:
Instead of asking a sales rep to remember every conversation, let the AI remember them.
The problem
Imagine a sales representative talking to a company called Zenith Corp.
During the first call, the customer says:
- They are interested in the product.
- They like the automation features.
- Pricing is their main concern.
- They want to know whether a discount is possible.
- Meera needs to approve the purchase.
- They currently have around 50 users.
A normal AI assistant that only receives the current question does not automatically know any of this.
Before the next call, the sales rep might have to search through CRM notes or previous call transcripts.
That creates two problems:
- Important context can be missed.
- The rep spends time researching instead of preparing for the conversation.
I wanted the agent to remember this information automatically.
The idea
The architecture is straightforward:
Customer call notes
↓
Hindsight persistent memory
↓
Recall relevant customer context
↓
Groq LLM
↓
Next-call sales briefing
Using Hindsight
I used the Hindsight Python client to connect the application to persistent memory.
The memory bank is configured like this:
BANK = "sales-deals"
def get_memory():
return Hindsight(
base_url=os.getenv("HINDSIGHT_API_URL"),
api_key=os.getenv("HINDSIGHT_API_KEY"),
)
When a call is saved, the application uses aretain():
def hindsight_retain(content, context):
async def run():
memory = get_memory()
return await memory.aretain(
bank_id=BANK,
content=content,
context=context,
)
return asyncio.run(run())
The actual call notes are stored together with useful context:
hindsight_retain(
content=f"{call_label} with {client_name}: {notes}",
context=f"Sales call notes for {client_name}",
)
This gives the memory system information about both the conversation and what that information represents.
Recalling previous conversations
When the sales rep wants to prepare for another call, the application sends a recall query to Hindsight.
def hindsight_recall(query):
async def run():
memory = get_memory()
return await memory.arecall(
bank_id=BANK,
query=query,
)
return asyncio.run(run())
The query focuses on the information that is useful for sales preparation:
result = hindsight_recall(
query=(
f"Everything about {client_name}: "
f"objections, competitors, pricing, stakeholders, "
f"promises made, previous conversations, next steps, "
f"and customer requests"
)
)
The returned memories are then converted into text that the LLM can use.
if result and result.results:
facts = "\n".join(
f"- {memory.text}" for memory in result.results
)
else:
facts = f"No previous memories were found for {client_name}."
Turning memory into a sales briefing
The recalled information is not shown to the sales representative as a long list of raw memories.
Instead, it is given to the LLM with a specific instruction.
briefing_prompt = f"""
I have a call with {client_name} coming up.
Create a short, practical sales briefing.
Use the previous call memories below.
Previous call memories:
{facts}
The briefing should include:
- What we know about the customer
- Important objections
- Pricing discussions
- Stakeholders involved
- Competitors if mentioned
- Promises or commitments made
- What the customer wants next
- Recommended talking points
- Suggested next step
Do not invent facts that are not present in the memories.
"""
That last instruction is important.
The agent should use the customer's history rather than inventing information that was never discussed.
Before and after
Without persistent memory, the assistant receives a question such as:
I have a call with Zenith Corp coming up. Brief me and suggest how to handle it.
The response has very little customer-specific context to work with.
With Hindsight memory, the same request can be accompanied by the previous conversation:
- Zenith is interested in the product.
- Automation features were well received.
- Pricing is an objection.
- The customer asked about a discount.
- Meera is involved in the purchasing decision.
- The company has around 50 users.
The resulting briefing can therefore focus on the actual situation instead of producing a generic sales checklist.
For example, the representative can be reminded to clarify the pricing concern, prepare the discount discussion, and make sure the purchasing stakeholder is included in the next step.
This is the main difference the project is designed to demonstrate:
The LLM generates the language, but persistent memory provides the context.
The complete application flow
The application is built with Python and Streamlit.
The user first enters the customer or company name and then chooses between two actions.
Log a call
The sales representative enters the call label and notes.
The application sends those notes to Hindsight.
Brief me for the next call
The application recalls relevant memories for that customer.
Those memories are passed to the LLM.
The LLM produces a concise briefing containing customer information, objections, pricing discussions, stakeholders, commitments, talking points, and suggested next steps.
The application also displays the raw memories used to create the briefing so the user can inspect the source context.
Technology stack
The project uses:
- Python
- Streamlit
- Hindsight for persistent memory
- Groq for LLM generation
Streamlit provides the simple user interface.
Hindsight handles persistent memory across conversations.
Groq provides the language model used to transform recalled information into a practical sales briefing.
The important part of the architecture is the separation between memory and generation.
The LLM does not need to permanently remember every customer conversation itself. Instead, relevant information can be retrieved from persistent memory when it is needed.
What I learned
The biggest lesson from building this was that an AI assistant becomes much more useful when it can access the right context at the right time.
A language model can generate a good response from a single prompt, but a sales assistant needs continuity.
The customer conversation from last week can affect the conversation today.
An objection raised in an earlier meeting can become the most important topic in the next meeting.
A stakeholder mentioned several calls ago can become critical when the deal reaches the purchasing stage.
Persistent memory makes these connections possible.
Another useful lesson was that retrieval quality matters.
Simply storing information is not enough. The application also needs to ask memory for the information that is relevant to the current task. That is why the recall query focuses on objections, pricing, stakeholders, competitors, commitments, previous conversations, and next steps.
A limitation
This project is a focused prototype rather than a complete CRM replacement.
The quality of the briefing depends on the quality of the information stored in memory and the relevance of the recalled results.
If the original call notes are incomplete, the resulting briefing may also be incomplete.
There is also a practical need to decide what customer information should be stored, how it should be organized, and how long it should remain available.
For a production system, these areas would need additional attention.
What's next
There are several directions I would explore next.
The agent could automatically process call transcripts instead of requiring manual notes.
It could track the progress of an entire deal and identify unresolved customer concerns.
It could also compare current conversations with previous successful deals and surface patterns that may help the sales representative.
Another useful feature would be automatic follow-up tracking, where commitments from previous conversations become reminders for the next interaction.
The underlying idea remains the same:
Remember the conversation, retrieve the relevant context, and help the representative act on it.
Conclusion
Sales teams already have a large amount of customer information.
The challenge is making that information useful at the moment it matters.
A Deal Intelligence Agent with persistent memory can bridge that gap by remembering previous conversations and turning them into a focused briefing before the next call.
Hindsight provides the persistent memory layer, while the LLM turns recalled context into a practical response.
The result is a simple pattern for building more context-aware AI agents:
Remember → Recall → Reason → Act
That pattern can extend beyond sales.
Any application where the past matters to the next interaction can benefit from persistent memory.
Top comments (0)