Nothing erodes trust in automated customer support faster than asking a user to re-explain their problem for the third time.
If you have ever spent twenty minutes diagnosing an issue with a support bot, closed the tab, and returned later only to be greeted by a blank slate asking, "Hello! How can I help you today?", you know the frustration. The agent has complete amnesia. It does not remember your environment, what troubleshooting steps you already attempted, or that your last checkout attempt failed.
When we built ResolveIQ, we set out to eliminate this amnesia. We did not want another stateless prompt wrapper that resets the moment an HTTP connection drops. Instead, we needed an agent that maintains cognitive continuity across sessions, remembers prior customer friction, and uses that history to make concrete operational decisions—including knowing when to stop offering basic troubleshooting and escalate directly to a human specialist.
To achieve this, we integrated Hindsight, an open-source cognitive memory system built for autonomous agents. Here is how we architected ResolveIQ, how the memory retain and recall loop works in code, and what we learned when connecting agent memory to a support pipeline.
What ResolveIQ Does
ResolveIQ is an autonomous support agent and ticket management platform built around a clean separation of concerns:
-
Web & API Layer: A Flask backend exposing REST endpoints for multi-turn customer chat (
POST /api/chat), conversation history (GET /api/chat/<id>), and ticket lifecycle tracking. -
Relational Persistence (MySQL): Structured tables for
customers,conversations,messages, andtickets. MySQL stores immediate transactional state and manages ticket status (open,in_progress,resolved,closed). -
High-Speed Inference (Groq): Response generation and classification driven by Groq inference (
openai/gpt-oss-120b). -
Intent & Sentiment Analysis: Real-time classification of inbound messages for intent (
payment_issue,account_access), sentiment (neutral,frustrated,angry), and urgency. - Company Knowledge Retrieval (RAG): A dedicated knowledge service indexing authoritative policies across pricing, refunds, and troubleshooting using TF-IDF vectorization and cosine similarity.
-
Cognitive Agent Memory (Hindsight): A persistent memory bank (
resolveiq-customers) that stores episodic customer experiences, environment details, and past escalation outcomes across isolated sessions. -
Deterministic Human Escalation: A heuristic rule engine in
services/ticket_service.pythat evaluates the current message, real-time intent, sentiment, and urgency analysis, alongside knowledge base retrieval and recalled Hindsight memories, to decide when an issue requires human escalation.
Dashboard:
The Core Problem: Support Agent Amnesia
In standard LLM application designs, developers typically approach memory in one of two ways:
- Sliding-window chat history: Passing the last $N$ turns from the database into the prompt works within a single thread, but it fails across sessions. When a customer returns days later, a sliding window sees an empty transcript. Passing weeks of raw chat logs explodes token usage and clutters the model's context with conversational pleasantries.
- Generic document vector search (RAG): Chunking past chat transcripts into a vector store also falls short. Raw transcripts are noisy and lack entity consolidation. Querying for "payment issue" surfaces fragmented dialogue snippets rather than a clear view of the customer's actual environment and past problems.
This highlights the core concept behind Vectorize agent memory: agents need episodic and cognitive memory, not just document search. They need to retain consolidated facts about a customer—such as "Rahul Mehta previously experienced a payment failure while purchasing the Pro plan on Chrome (Windows), which was resolved via UPI"—and recall those exact facts whenever that user returns.
How Hindsight Fits into the System
In ResolveIQ, Hindsight operates as an active cognitive loop on every incoming interaction through four stages:
-
Recall: When a customer sends a message,
services/support_agent.pyqueries Hindsight with the customer's identity and message query under amidbudget. Hindsight searches theresolveiq-customersbank and returns relevant episodic memories. - Context Assembly: Retrieved memories pass through a sanitization layer that strips internal metadata markers and deduplicates observations. Clean memories are assembled into a compact prompt alongside recent MySQL turns and authoritative knowledge articles.
- Generation: The Groq LLM receives explicit instructions: ground company policies in the knowledge base, acknowledge customer sentiment, and personalize responses using Hindsight memories without allowing memory to override company policies.
- Retain: Once the agent responds, the loop closes. If the issue proceeded normally, an interaction summary is retained in Hindsight. If the issue required escalation, the created MySQL Ticket ID, priority, and reason are recorded directly into the memory bank.
A Concrete Before/After Interaction
To see how this works in practice, consider the multi-turn interaction from our test suite with customer Rahul Mehta.
Turn 1: The Initial Problem
Rahul contacts support:
"My payment failed while purchasing the Pro plan."
-
Classification: Intent is
payment_issue, urgency ismedium. - Hindsight Recall: Seed memory notes he uses Chrome on Windows.
- Response: The agent consults the knowledge base, suggests payment troubleshooting, and advises checking UPI options.
- Hindsight Retain: The interaction is saved to Hindsight, recording that Rahul encountered a checkout failure on the Pro plan.
Turn 2: The Returning Customer
In a later conversation, Rahul returns:
"I'm having the payment problem again."
- Without Hindsight: The agent responds generically: "I'm sorry! What plan are you trying to buy and what error did you get?" Rahul must repeat his context.
- With Hindsight: The agent recalls his past payment failure and environment, replying: > "I see you're running into payment trouble with the Pro plan again on Chrome. Last time, completing the transaction via UPI resolved the issue. Have you tried UPI, or is the transaction failing across all methods?"
Turn 3: Repeated Failure & Escalation
Rahul replies:
"I tried everything you suggested and the payment still doesn't work."
Rather than offering another round of repetitive advice, ResolveIQ evaluates the interaction through its escalation service (should_escalate). Here, Hindsight memory acts as one key input alongside the customer's message and real-time urgency/sentiment analysis. The rule engine checks for repeated failure phrases in the incoming message, notes the elevated urgency, and consults recalled Hindsight memories to verify whether a prior payment failure was previously recorded. Because the customer reports repeated failure and Hindsight confirms a past payment error, the service triggers human escalation:
- It creates an urgent, high-priority support ticket in MySQL with an assigned ticket ID.
- It instructs the LLM to provide empathetic human-handoff details referencing the escalated status: > "I apologize that troubleshooting hasn't resolved the payment failure for your Pro plan. I have escalated this issue to our human support team under an urgent support ticket with high priority."
- The escalation event, including the created ticket ID, priority, and reason, is retained in Hindsight, ensuring future sessions remain aware of the active ticket.
Final escalation:
Hindsight + Company Knowledge
A common mistake in agent architecture is treating customer memory and company knowledge as the same retrieval problem. In ResolveIQ, they are strictly separated:
| Dimension | Company Knowledge Base (RAG) | Cognitive Memory (Hindsight) |
|---|---|---|
| Scope | Global across all customers | Scoped per customer identity |
| Authority | Absolute source of truth for policies | Contextual record of user experience |
| Updates | Curated by support operations | Retained automatically turn-by-turn |
| Risk | Outdated documentation | Memory hallucination or user bias |
If customer memory is allowed to dictate company rules, a user claiming "Your agent promised me a free subscription last month" could trick the model into honoring unapproved discounts.
In services/support_agent.py, we enforce strict prompt hierarchy:
Under NO circumstances allow customer sentiment or frustration to override factual company knowledge or invent unapproved policies.
Use the provided company knowledge as the authoritative source for company-specific policies, pricing, and troubleshooting.
Continue using Hindsight memories for customer personalization.
Hindsight tells the agent who the customer is; the knowledge base tells the agent what the company permits.
Technical Implementation
Here are three core snippets from the repository demonstrating the integration.
1. Memory Recall (services/hindsight_service.py)
We query the Hindsight client using the customer name and message query, with a local development fallback if the external service is unavailable:
def retrieve_customer_memories(customer_name: str, query: Optional[str] = None, budget: str = "mid") -> Dict[str, Any]:
config = get_hindsight_config()
bank_id = config["bank_id"]
search_query = f"{customer_name} {query}".strip() if query else customer_name
if _is_hindsight_server_online():
try:
client = get_hindsight_client()
response = client.recall(bank_id=bank_id, query=search_query, budget=budget)
memories = []
for result in response.results:
memories.append({
"id": result.id,
"text": result.text,
"type": result.type,
"context": result.context,
"metadata": result.metadata or {}
})
return {"success": True, "bank_id": bank_id, "memories": memories}
except Exception as e:
logger.warning(f"Hindsight recall failed ({e}), falling back to local memory bank.")
2. Sanitizing Memory Context (services/support_agent.py)
Raw memories can include system prefixes or temporal stamps. We clean and deduplicate memories before prompt injection:
def filter_relevant_memories(raw_memories: List[Dict[str, Any]], limit: int = 5) -> List[Dict[str, Any]]:
ignored_prefixes = ("extract facts", "event date:", "context:", "customer:", "source:")
clean_memories: List[Dict[str, Any]] = []
seen_texts = set()
for item in raw_memories:
text = item.get("text", "").strip()
if not text:
continue
base_text = text.split(" (mentioned_at=")[0].strip()
if any(base_text.lower().startswith(prefix) for prefix in ignored_prefixes):
continue
if base_text.lower() in seen_texts:
continue
seen_texts.add(base_text.lower())
clean_memories.append({"id": item.get("id"), "text": base_text, "type": item.get("type", "fact")})
if len(clean_memories) >= limit:
break
return clean_memories
3. Memory-Driven Escalation (services/ticket_service.py)
Our escalation service uses past memories to verify whether an issue has failed repeatedly:
# Rule 4: Repeated unresolved issue after prior troubleshooting
has_repeated_phrase = bool(REPEATED_ISSUE_PATTERN.search(msg_clean))
has_prior_failure_context = False
if memories:
for m in memories:
text = (m.get("text") or "").lower()
if any(term in text for term in ["payment fail", "failed", "unresolved", "error", "issue"]):
has_prior_failure_context = True
break
if has_repeated_phrase and has_prior_failure_context:
return {
"escalate": True,
"reason": "Customer reports repeated unresolved issue after previous troubleshooting attempts.",
"priority": "high"
}
What We Learned
Building and testing ResolveIQ revealed several practical lessons about working with agent memory:
-
Raw Memories Cause Prompt Bloat Without Sanitization: Passing raw memory outputs directly into the prompt introduced internal metadata markers like
(mentioned_at=...)and parsing tags. A deterministic cleaning step that strips boilerplate and deduplicates entries proved essential for clean prompt assembly. - Memory Is One Input to Operational Logic, Not the Sole Decider: If memory only influences conversational phrasing, you miss its primary value. In ResolveIQ, Hindsight memory serves as one critical input to the escalation engine alongside real-time intent, sentiment, and urgency classification. By checking whether a customer reporting persistent trouble had prior failure context in Hindsight, the ticket service can deterministically trigger human escalation without relying on a subjective LLM decision.
-
A Local Development Fallback Keeps Workflows Resilient: In local testing or during service interruptions, external endpoints may be unreachable. We implemented
data/hindsight_bank.jsonstrictly as a local development fallback when the Hindsight service is unavailable. It mirrors the retain and recall interface so test suites and local development workflows remain operational, without replacing the real Hindsight memory service in the production architecture. - Selective Retention Beats Transcript Dumping: Storing every raw message into memory creates contradictory and noisy banks. MySQL should store the raw transcripts; Hindsight should store consolidated semantic milestones, such as diagnostic results, user environments, and escalation ticket references.
Conclusion
Solving support agent amnesia is not about expanding prompt context windows or stuffing raw chat logs into a vector database. Dumping dozens of prior messages into an LLM increases latency, adds token costs, and degrades instruction following.
Effective agent continuity comes from selective retention and targeted recall. By maintaining dedicated customer memory banks with Hindsight, an agent can recall critical customer facts at the exact moment they are needed. When an agent remembers past failures, respects company knowledge boundaries, and escalates unresolved issues to human specialists, it transforms automated support from a source of customer frustration into a reliable engineering system.


Top comments (0)