Recall, Respond, Retain: The Architecture Behind Our Memory-Powered Support Agent
python #flask #ai #agents #llm
What happens when you give an AI support agent long-term memory?
That was the question behind our hackathon project, SupportMind.
I'm [Your Name], and our team built SupportMind to explore a simple architecture:
Recall → Respond → Retain
Instead of treating every support message as a completely new conversation, the system retrieves relevant information from previous customer interactions before generating an answer.
The problem with stateless support
Consider two customers sending exactly the same message:
“My internet keeps dropping.”
For a new customer, general troubleshooting makes sense.
But what if another customer reported the same issue last week, tried restarting the router multiple times, and eventually solved the problem by updating the firmware?
Giving both customers exactly the same response ignores useful information.
We wanted our agent to understand that difference.
Step 1: Identify the customer
Each request contains a customer ID.
customer_id = request.json.get("customer_id")
message = request.json.get("message")
The customer ID becomes the key used to access that customer's memory.
This creates isolated histories rather than one giant shared memory.
Conceptually:
Customer A → Memory Bank A
Customer B → Memory Bank B
Customer C → Memory Bank C
This separation is especially important because information belonging to one customer should not accidentally influence another customer's support conversation.
Step 2: Recall
Before asking the LLM to answer, SupportMind searches for relevant memories.
recall_result = hindsight.recall(
bank_id=customer_id,
query=message
)
past_memories = [
result.text for result in recall_result.results
]
The important idea here is relevance.
We don't necessarily need every interaction a customer has ever had.
We need the memories that are useful for the current problem.
If the customer asks about WiFi, previous WiFi problems may matter much more than an unrelated billing question.
Step 3: Add memory to the prompt
The recalled information becomes additional context.
if past_memories:
memory_context = (
"Past history with this customer:\n"
+ "\n".join(past_memories)
)
else:
memory_context = (
"This is a new customer with no past history."
)
The language model now has two sources of information:
CURRENT MESSAGE
+
RELEVANT MEMORY
↓
LLM
↓
CONTEXT-AWARE RESPONSE
This is where the difference between a new and returning customer becomes visible.
Step 4: Generate the response
The memory context is included in the system message:
response = groq_client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[
{
"role": "system",
"content":
f"You are a helpful support agent.\n{memory_context}"
},
{
"role": "user",
"content": message
}
]
)
The LLM itself doesn't need to permanently remember every customer.
The application retrieves the appropriate memory and provides it when needed.
That separation was one of the most interesting parts of the project for us.
Step 5: Retain
After generating the response, we save the new interaction.
hindsight.retain(
bank_id=customer_id,
content=(
f"Customer said: {message}. "
f"Agent replied: {reply}"
)
)
Now the system has additional information available for the customer's next visit.
That creates a loop:
┌───────────────┐
│ Customer asks │
└───────┬───────┘
↓
┌───────────────┐
│ Recall memory │
└───────┬───────┘
↓
┌───────────────┐
│ Generate reply│
└───────┬───────┘
↓
┌───────────────┐
│ Retain result │
└───────┬───────┘
│
└────→ Future conversations
Reflecting instead of reading everything
We also experimented with another useful idea: customer briefings.
A human support representative doesn't want to read dozens of previous messages before answering a ticket.
Instead, SupportMind can ask the memory system to reflect on the customer's history and produce a short briefing.
hindsight.reflect(
bank_id=customer_id,
query="""
Write a briefing for a support agent about this
customer: past issues, what fixed them, and how
best to help next. Use 3 short bullet points.
"""
)
This turns long-term history into immediately useful context.
Making memory visible
One of our favorite parts of the prototype is the memory panel.
Whenever the AI responds, the interface can show which memories were retrieved.
That makes debugging much easier.
Instead of wondering:
“Why did the AI recommend this?”
we can inspect the memory context that influenced the response.
What we learned
The biggest lesson was that AI applications don't always need larger prompts.
They need better context selection.
Sending an entire customer history to the model can become inefficient and noisy as the history grows.
Retrieving relevant information first gives the model a much more focused context.
Where we can take it next
Our prototype currently demonstrates the memory layer.
A production version could connect the same architecture to:
- CRM systems
- Ticket databases
- Subscription systems
- Order history
- Authentication
- Human support dashboards
- Controlled actions such as refunds or plan changes
The hackathon gave us an opportunity to experiment with a simple idea:
AI support shouldn't only know how to answer. It should know what happened before.
That's the idea behind SupportMind.
Top comments (0)