A support agent can give the right answer and still make a customer repeat themselves. I built this system around a simple idea: use long-term memory to carry useful context between conversations, but never let memory outrank the account data that determines what is actually true.
The Context Problem
The application is a customer support console with chat, requests, orders, feedback, and account settings. A customer message enters through the chat interface, the system gathers the relevant account context, and an assistant produces a reply. Attachments are stored with the conversation, and each exchange is recorded so the next interaction has a history to build on.
That flow sounds ordinary until the customer comes back a week later. The previous conversation may contain a preference, a troubleshooting step that already failed, or the fact that an issue is still open. Sending the entire transcript to a language model is a crude way to recover those details. It increases prompt size, includes irrelevant turns, and makes it harder to distinguish what the customer said from what the system knows about their account.
I treated context as several different things with different jobs:
• The customer profile carries stable preferences and identifying details.
• Orders and requests provide current account facts.
• Recent chat history preserves the immediate conversational thread.
• Hindsight provides durable, queryable memories from earlier interactions.
• Historical support tickets offer possible patterns, not facts about this customer.
The engineering decision that shaped the whole system was to keep these sources separate until the moment a reply is composed. That makes the context richer without pretending every piece of retrieved text has the same authority.
One Memory Bank Per Customer
I use Hindsight on GitHub for the long-term memory layer. Each customer gets a distinct bank, and recall is performed with the current message as the query:
def _bank(customer_id) -> str:
return f"customer-{customer_id}"
**def recall_memories(customer_id, query: str, limit: int = 6) -> list[str]:
try:
result = _hindsight().recall(bank_id=_bank(customer_id), query=query)
items = getattr(result, "results", result) or []
return [str(getattr(item, "text", item)) for item in items][:limit]
except Exception:
return []
**
There are two important choices here. First, the customer identifier determines the memory boundary. Second, retrieval is relevant to the current question; the agent does not indiscriminately replay everything ever stored. The result limit keeps the amount of long-term context bounded. The deployed application resolves this identifier from authenticated server-side state, never from a value supplied by the chat client. A memory system is only as useful as its isolation guarantees.
Hindsight is not a replacement for the order service or request database. It is a way to retrieve conversational knowledge that is awkward to represent as a fixed profile field. A customer's preference for concise updates, the fact that they already tried a reset, or a recurring explanation they have given can matter in a future exchange. An order's current delivery status cannot safely be inferred from a remembered conversation; it has to come from the authoritative order record.
That division is why I wanted a memory system designed for agent memory rather than treating a vector store as a transcript bucket. The Hindsight documentation describes the memory interface this application uses. The broader distinction between an agent's working context and its longer-lived memory is also covered in Vectorize's overview of agent memory.
Memory Is Retrieved, Not Dumped
The reply path assembles a compact view of the customer. It retrieves memories with the new message, looks up relevant historical cases, and includes only the most recent turns from the active conversation:
memories = recall_memories(customer["customer_id"], user_msg)
cases = similar_tickets(user_msg)
messages = [{"role": "system", "content": SYSTEM_PROMPT + "\n\n" + context}]
messages += history[-8:]
messages.append({"role": "user", "content": user_msg})
return _llm(messages)
This is not a claim that eight turns or six memories is universally correct. They are explicit bounds that make the context policy inspectable. Recent turns help the model understand pronouns and follow-up questions. Hindsight can contribute a detail from a previous session even when it is no longer in that short window. The current message gives retrieval a concrete relevance signal.
I also kept similar-ticket search separate from customer memory. The application uses TF-IDF to retrieve a few related support cases from a ticket dataset. Those cases can suggest a useful next step, but they are not evidence about the customer in front of us. A similar invoice dispute does not prove that this customer's invoice was duplicated.
That distinction is enforced in the system instructions, not left to an informal convention:
- Use ONLY the customer facts, orders, requests and memories given below.
- Never invent order numbers, refunds, dates or policies.
- "Similar past cases" are general hints from historical tickets, not facts about this customer.
- Use them for suggested next steps only. I like this boundary because it is simple enough to review. Personal memory can shape how the agent responds; account records establish what happened. Historical cases can inform what to try next; they cannot fill gaps in a customer's record. That separation is more useful than asking a model to be generally careful and hoping it interprets every context source correctly. A Billing Conversation, End to End Consider a customer who returns and writes, “I’m seeing the double charge again. Can you check?” The current message can retrieve earlier context related to billing. The profile contributes the preferred contact method. The request system can confirm whether a billing ticket is open. Recent turns keep the follow-up understandable, while similar historical tickets may suggest what evidence the billing team typically needs The reply should be grounded in the actual account: “I can see your billing request is still open. I’ll follow up by email. If you have the latest invoice, attach it here so the billing team can compare the charges.” If memory says the customer already uploaded that invoice yesterday, the agent can avoid asking for it again. If the ticket record says it was resolved, that current record must take precedence over an older memory that says it was open. The scenario illustrates why long-term memory should not be the system of record. It is a retrieval layer for continuity. The account systems remain responsible for mutable facts such as order state, ticket status, and amounts. That also gives the user a sensible correction path: update the record or clarify the conversation, rather than trying to teach the model which version of reality to trust. Keeping Memory Off the Critical Path After a response, the system stores both sides of the exchange so future recall has the customer's request and the agent's answer. Retention can take longer than generating the immediate response, so it should not hold the chat open: content = f"Customer said: {customer_msg}\nSupport agent replied: {agent_reply}"
def _store():
try:
_hindsight().retain(
bank_id=_bank(customer_id),
content=content,
context="customer support chat",
)
except Exception:
pass
threading.Thread(target=_store, daemon=True).start()
The useful design boundary is that a memory write is not part of the synchronous reply contract. In the deployed system, durable background workers handle retention with retries and failure visibility; a fire-and-forget thread alone would not be a delivery guarantee. The customer still gets an answer if memory retention is temporarily unavailable, and operators can observe that degradation.
Recall follows the same principle. If Hindsight is unavailable, the retrieval function returns no memories and the rest of the context can still support a response. The app also has a rule-based fallback when the model path fails. These are deliberate availability choices: memory improves continuity, but it should not become a single point of failure for basic support. A production service pairs those fallbacks with operational signals rather than silently ignoring persistent failures.
What I Learned
First, separate memory from truth. Use long-term memory for continuity and preferences, and query business systems for facts that change. Putting both in one prompt does not make them equally reliable.
Second, scope memory before optimizing retrieval. A relevance score is useless if the search is allowed to cross customer boundaries. Identity and authorization have to be part of the memory lookup contract.
Third, bound every context source. A short recent window and a capped memory result set make cost and behavior easier to reason about. The right limits need measurement against real conversations, but an explicit limit is a better starting point than an unbounded transcript.
Fourth, make optional context actually optional. Memory lookup and retention can fail without taking down the reply path. That requires deliberate fallback behavior and operational signals; silently swallowing errors is not enough for a production service.
Finally, retrieved examples need labels. Historical tickets can help suggest a troubleshooting step, but calling them “similar past cases” and explicitly limiting their role prevents them from quietly turning into customer-specific evidence.



The goal was not to make an agent remember everything. It was to make it remember a small amount of relevant context, for the right customer, while staying honest about what it knows. That is the difference between a chat transcript with a search box and a support system that can pick up a conversation without losing track of its sources.

** THE ARCHITECTURE OF PROJECT**
CUSTOMER
│
▼
┌─────────────────┐
│ Streamlit UI │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Customer Agent │
│ Controller │
└────────┬────────┘
│
┌─────────┴─────────┐
▼ ▼
┌─────────────┐ ┌──────────────┐
│ LLM │ │ HINDSIGHT │
│ Groq │◄────►│ Long-term │
│ │ │ memory │
└─────────────┘ └──────┬───────┘
│
▼
Customer experience
history
│
▼
Feedback / outcome
│
└──────────►
Hindsight
the hindsight github repository:
https://github.com/manasapenchala/ai-agent-project.git
the documentation for hindsight:
https://hindsight.vectorize.io/
the agent memory page on vectorize:
https://vectorize.io/what-is-agent-memory
Top comments (0)