



Memory-Enabled Customer Support Agent — Dev-Ready Specification
Project Overview 1.1 Problem Current customer-support systems are commonly stateless at the interaction level. A returning customer may have to explain the same issue again, repeat troubleshooting steps, and wait while a new agent re-diagnoses a problem that was already solved.
The core problem is not simply lack of stored data. The problem is retrieving the right customer-specific history at the moment the agent is about to respond.
The system described in this document adds per-customer agent memory so the support agent can recall:
what the customer tried,
what failed,
what worked,
what problem is currently open or resolved,
stable facts about the customer's environment.
The supplied design uses Hindsight as the production memory backend and keeps a local in-process/file-backed implementation for development and testing.
Goals Primary goals Store structured memories for each customer. Keep memory strictly isolated by customer_id. Recall relevant memories before generating every reply. Use previous successful fixes before previously failed troubleshooting steps. Extract new memories automatically from customer messages. Support both local development storage and Hindsight production storage. Make the storage layer replaceable through a common protocol. Keep the complete recall/decision metadata auditable. Allow the full system to run without a network connection in local/test mode. Avoid feeding the current message back into its own recall results. Non-goals Building a general global FAQ/knowledge base. Mixing memories between customers. Replacing the support agent's reasoning engine with the memory system. Requiring a human to manually create every memory. Making the LLM the storage layer.
Core Concepts Each customer has three main memory types.
3.1 ATTEMPT
A troubleshooting action that was attempted.
Possible outcomes:
worked
failed
in_progress
Example:
Cleared application cache → failed
Update application → worked
3.2 INCIDENT
The problem being reported.
An incident can remain open until the problem is resolved.
Example:
App crashes on large PDF upload
3.3 PROFILE
A stable customer/environment fact.
Example:
Uses Acme PDF Suite
Functional Requirements FR-01 — Customer isolation Every memory must belong to exactly one customer_id.
A retrieval for customer A must never return customer B's memories.
FR-02 — Memory extraction
The system must extract structured memories automatically from customer messages.
Example input:
My Acme PDF Suite crashes on large PDF upload.
I restarted it but that did not help.
Cleared the cache, still crashing.
Updated the app and it fixed it.
Expected memories:
INCIDENT App crashes on large PDF upload
ATTEMPT Restart application failed
ATTEMPT Cleared application cache failed
ATTEMPT Update application worked
PROFILE Uses Acme PDF Suite
FR-03 — Recall before reply
The system must retrieve relevant historical memories before the reasoning engine generates the response.
Required order:
Customer message
↓
Parse current context
↓
Recall existing customer memories
↓
Generate response using recalled memories
↓
Extract memories from current message
↓
Store new memories
The current message must not be stored before recall.
FR-04 — Known-fix handling
If a previous attempt worked, the response should surface that successful action before repeating known failures.
FR-05 — Known-failure handling
Previously failed actions should be available so the agent can skip unnecessary repeated troubleshooting.
FR-06 — Returning-customer detection
The response metadata must indicate whether the customer has relevant prior history.
FR-07 — Auditable response metadata
Each generated support response should expose:
whether the customer is returning,
known fixes,
known failures,
memories used,
relevance score,
reason for retrieval.
FR-08 — Deterministic fallback
If the LLM/reasoning service is unavailable, the system must still return a deterministic/template response instead of failing completely.
System Architecture
┌─────────────────────┐
│ Customer Message │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Message / Context │
│ Parser │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Memory Recall │
│ customer_id + query │
└──────────┬──────────┘
│
┌───────────────┴───────────────┐
│ │
▼ ▼
┌──────────────────┐ ┌──────────────────┐
│ InMemoryStore │ │ HindsightStore │
│ Development/Test │ │ Production │
└──────────────────┘ └──────────────────┘
│ │
└───────────────┬───────────────┘
▼
┌─────────────────────┐
│ Reasoning / Reply │
│ Engine │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ SupportReply │
│ + audit metadata │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Memory Extraction │
│ from current msg │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Store New Memories │
└─────────────────────┘
MemoryStore Interface
The agent must depend only on a storage protocol, not on a concrete backend.
Python contract:
from typing import Protocol
class MemoryStore(Protocol):
def add(self, memory: Memory) -> Memory:
...
def update(self, memory: Memory) -> Memory:
...
def all(self, customer_id: str) -> list[Memory]:
...
def forget(self, customer_id: str) -> int:
...
def search(
self,
customer_id: str,
query: str,
limit: int = 5
) -> list[tuple[Memory, float]]:
...
This gives the application one stable interface for:
development,
testing,
production,
future storage implementations.
Data Model 7.1 Memory Recommended logical structure:
class Memory:
id: str
customer_id: str
type: str
outcome: str | None
text: str
product: str | None
incident_id: str | None
created_at: str
updated_at: str
Required fields
Field Purpose
id Unique memory identifier
customer_id Customer isolation key
type ATTEMPT, INCIDENT, or PROFILE
text Human-readable memory
outcome Result for attempts
product Product/environment context
incident_id Associates memory with an incident
created_at Creation time
updated_at Last modification time
Memory Extraction The extraction layer converts natural-language support messages into structured records.
Example
Input:
My Acme PDF Suite crashes on large PDF upload.
I restarted it and cleared the cache but neither helped.
Updating the app fixed it.
Output:
INCIDENT:
App crashes on large PDF upload
ATTEMPTS:
Restart application -> failed
Cleared application cache -> failed
Update application -> worked
PROFILE:
Uses Acme PDF Suite
Extraction rules
Rule 1 — Explicit success
Phrases such as:
fixed it
worked
solved it
resolved the issue
should mark the associated attempt as:
worked
Rule 2 — Explicit failure
Phrases such as:
did not help
still crashing
didn't work
failed
should mark the associated attempt as:
failed
Rule 3 — Multiple failures
For:
I restarted it and cleared the cache but neither helped.
both previous attempts should be marked:
Restart application -> failed
Cleared application cache -> failed
Rule 4 — Incident lifecycle
An incident can begin as:
open
and become:
resolved
when the message indicates that a fix resolved it.
Rule 5 — Profile extraction
Stable environmental facts should be stored separately from troubleshooting attempts.
Example:
Uses Acme PDF Suite
becomes a PROFILE memory.
Recall Pipeline
The recall loop must execute in this exact logical order.
Receive message
Identify customer
Parse product/symptoms/context
Search existing memories
Rank/filter relevant memories
Build reasoning context
Generate response
Extract new memories
Persist new memories
Critical implementation rule
Recall must happen before ingest.
Do not do:
store(current_message)
recall(current_message)
because the current message can then appear as if it were historical memory.
Correct:
recall(current_message)
generate_reply()
extract_memory(current_message)
store_memory()
Retrieval 10.1 Local implementation The local store can use lexical overlap.
Conceptually:
query_tokens = tokenize(query)
memory_tokens = tokenize(memory.text)
score = overlap(query_tokens, memory_tokens)
The local approach is intended for:
development,
demos,
tests.
10.2 Hindsight implementation
The production store sends recall requests to Hindsight.
Logical request:
POST /recall
Payload:
{
"bank": "support",
"customer_id": "CUST-1042",
"query": "application crashing when uploading PDF",
"limit": 5
}
A memory search result should return:
{
"memory": {
"customer_id": "CUST-1042",
"type": "ATTEMPT",
"outcome": "worked",
"text": "Update application"
},
"score": 0.71
}
Hindsight provides semantic retrieval through embeddings, which can identify related wording even when exact words differ.
Example:
Current message:
"The software freezes"
Stored memory:
"Application crashes"
Semantic retrieval can associate these more effectively than simple token overlap.
HindsightStore The production implementation should translate the common MemoryStore interface into HTTP requests to Hindsight.
Conceptual implementation:
class HindsightStore:
def add(self, memory: Memory) -> Memory:
response = self._client.post(
f"{self.base_url}/memories",
json={
"bank": self.bank,
"customer_id": memory.customer_id,
"memory": memory.to_dict(),
},
)
response.raise_for_status()
return memory
def search(
self,
customer_id: str,
query: str,
limit: int = 5,
):
response = self._client.post(
f"{self.base_url}/recall",
json={
"bank": self.bank,
"customer_id": customer_id,
"query": query,
"limit": limit,
},
)
response.raise_for_status()
return [
(
Memory.from_dict(item["memory"]),
float(item.get("score", 0.0)),
)
for item in response.json().get("results", [])
]
The HTTP client should be injected into the store so the wire contract can be tested without a real Hindsight server.
InMemoryStore The development implementation should store memories in a Python dictionary and may optionally persist them to a local JSON file.
Example structure:
{
"CUST-1042": [
Memory(...),
Memory(...),
Memory(...)
]
}
Responsibilities:
add
update
all
forget
search
No external service should be required for local tests.
Backend Selection The application should choose the backend through one factory.
def build_store(path=None):
if os.getenv("HINDSIGHT_URL"):
return HindsightStore()
return InMemoryStore(path)
Development
HINDSIGHT_URL not set
↓
InMemoryStore
Production
HINDSIGHT_URL set
↓
HindsightStore
No storage-specific conditional logic should be scattered through the agent.
Environment Configuration Recommended production configuration:
export HINDSIGHT_URL=https://your-hindsight-instance.example.com
export HINDSIGHT_BANK=support-prod
Variables
Variable Purpose
HINDSIGHT_URL Hindsight service URL
HINDSIGHT_BANK Memory namespace/product/environment
MEMORY_LIMIT Maximum memories returned per recall
STORE_PATH Local JSON path for development
HINDSIGHT_BANK should be used to keep product/environment memory namespaces isolated.
API/Application Flow A support message should conceptually enter through:
POST /support/message
Request:
{
"customer_id": "CUST-1042",
"message": "The application is crashing again when I upload a PDF."
}
Internal processing:
customer_id
↓
parse message
↓
store.search(customer_id, query)
↓
construct reasoning context
↓
reasoning engine
↓
SupportReply
↓
extract memories
↓
store.add(...)
Example response:
{
"reply": "Welcome back. From your previous sessions, updating the application worked. Restart and cache clearing already failed, so I will skip those steps.",
"metadata": {
"is_returning_customer": true,
"known_fixes": [
"Update application"
],
"known_failures": [
"Restart application",
"Cleared application cache"
],
"used_memories": [
{
"outcome": "worked",
"text": "Update application",
"score": 0.71,
"reason": "same product, matched on app, crash, upload"
}
]
}
}
SupportReply Model Recommended structure:
class SupportReply:
reply: str
is_returning_customer: bool
known_fixes: list[str]
known_failures: list[str]
used_memories: list[dict]
Each used_memories item should contain:
{
"outcome": "worked",
"text": "Update application",
"score": 0.71,
"reason": "same product, matched on app, crash, upload"
}
This creates an audit trail showing:
What was recalled?
Why was it relevant?
What score did it receive?
What did the agent know before replying?
Reasoning Engine The reasoning engine receives:
Current customer message
+
Relevant customer memories
+
Known fixes
+
Known failures
+
Current product/context
It should use this information to generate a response.
Example
Memory
worked:
Update application
failed:
Restart application
Cleared application cache
Current message
The application is crashing again when I upload a PDF.
Expected behavior
The response should:
Recognize the customer as returning.
Surface the previously successful update.
Avoid repeating known failed actions.
Offer escalation context if the known fix no longer works.
Deterministic Fallback The application must not depend completely on the LLM.
A ScriptedEngine should exist for:
local testing,
demos,
API/model outages,
deterministic test assertions.
Example:
class ScriptedEngine:
def reply(self, message, memories):
known_fixes = [
m.text
for m in memories
if m.outcome == "worked"
]
known_failures = [
m.text
for m in memories
if m.outcome == "failed"
]
if known_fixes:
return (
"From your previous sessions, "
f"{known_fixes[0]} worked previously."
)
return (
"I will review the available troubleshooting "
"information and continue from there."
)
The exact response wording can be improved later without changing the memory architecture.
Example End-to-End Scenario First session Customer:
My Acme PDF Suite crashes whenever I upload a large PDF.
System stores:
INCIDENT:
App crashes on large PDF upload
Customer:
I restarted it but that did not help.
System stores:
ATTEMPT:
Restart application
Outcome:
failed
Customer:
Cleared the cache, still crashing.
System stores:
ATTEMPT:
Cleared application cache
Outcome:
failed
Customer:
Updated the app and it fixed it.
System stores:
ATTEMPT:
Update application
Outcome:
worked
The incident can now be marked resolved.
Returning session
Customer:
The application is crashing again when I upload a PDF.
Recall returns:
Update application worked
Restart application failed
Cleared application cache failed
Agent response should lead with:
Update application
and avoid repeating:
Restart application
Cleared application cache
If the previous fix no longer works, the response can use the accumulated context to move toward escalation.
One-Shot Extraction Customers may provide their history in a single message:
My Acme PDF Suite crashes on large PDF upload.
I restarted it and cleared the cache but neither helped.
Updating the app fixed it.
The extractor must produce:
Restart application -> failed
Cleared application cache -> failed
Update application -> worked
The extraction logic must correctly propagate:
"neither helped"
to both preceding troubleshooting actions.
Customer Isolation Customer memory must be scoped at every layer.
Correct:
search(
customer_id="CUST-1042",
query="PDF crash"
)
Incorrect:
search(
query="PDF crash"
)
The second form risks returning another customer's experience.
Two customers can have identical symptoms but different successful fixes.
Therefore:
customer_id
↓
memory namespace
↓
retrieval
↓
reasoning context
must remain isolated.
Testing Strategy 22.1 Unit tests Memory model Test:
creation,
serialization,
deserialization,
required fields,
outcome values.
InMemoryStore
Test:
add,
update,
all,
search,
forget,
customer isolation.
HindsightStore
Test:
POST /memories,
POST /recall,
payload structure,
bank handling,
customer ID handling,
response parsing,
HTTP failures.
Use an injected httpx.Client/transport mock.
Critical Integration Tests Test 1 — Returning customer gets contextual response Given:
CUST-1042
Update application -> worked
Restart application -> failed
When:
CUST-1042 reports the same crash
Then:
is_returning_customer == True
and the response context contains:
Update application
while recognizing:
Restart application
as a known failure.
Test 2 — Memory isolation
Given:
CUST-1042:
Update application -> worked
CUST-2077:
Reinstall application -> worked
When:
CUST-1042 recalls PDF crash
Then:
Reinstall application
from CUST-2077 must not appear in CUST-1042's recall.
Test 3 — Recall occurs before ingest
Given a new message:
Updating the app fixed it.
The current message must not appear in its own historical recall.
Verify the sequence:
recall()
reply()
extract()
store()
Test 4 — Multiple failure extraction
Input:
I restarted it and cleared the cache but neither helped.
Expected:
Restart application -> failed
Cleared application cache -> failed
Test 5 — Successful fix extraction
Input:
Updating the app fixed it.
Expected:
Update application -> worked
Test 6 — Deterministic fallback
Simulate unavailable reasoning/LLM service.
Expected:
SupportReply is still produced.
Test 7 — Hindsight wire contract
Use a mocked HTTP transport and verify exact requests/responses without a running Hindsight server.
Recommended Project Structure
support-memory-agent/
│
├── app/
│ ├── init.py
│ ├── api.py
│ ├── agent.py
│ ├── models.py
│ ├── extraction.py
│ ├── recall.py
│ ├── factory.py
│ │
│ ├── engines/
│ │ ├── init.py
│ │ ├── llm_engine.py
│ │ └── scripted_engine.py
│ │
│ └── stores/
│ ├── init.py
│ ├── protocol.py
│ ├── memory_store.py
│ └── hindsight_store.py
│
├── tests/
│ ├── test_models.py
│ ├── test_extraction.py
│ ├── test_recall.py
│ ├── test_memory_store.py
│ ├── test_hindsight_store.py
│ ├── test_isolation.py
│ └── test_agent.py
│
├── demo.py
├── requirements.txt
├── .env.example
├── README.md
└── data/
└── memories.json
Dependencies
The supplied design specifically uses httpx for the Hindsight HTTP client.
Example:
httpx
Additional dependencies should only be introduced where they are required by the selected reasoning/LLM implementation.
The local development path should remain as dependency-light as practical.
Production Deployment Local development HINDSIGHT_URL not set ↓ InMemoryStore ↓ local JSON persistence
Production
HINDSIGHT_URL set
↓
HindsightStore
↓
Hindsight service
The application code should remain unchanged between these modes.
Observability Each support interaction should make it possible to inspect:
customer_id
product
current message
retrieved memories
retrieval scores
retrieval reasons
known fixes
known failures
generated reply
new memories extracted
Avoid logging sensitive customer content unnecessarily.
The purpose of observability is to make AI behavior explainable:
"The agent said X because it retrieved memories A, B and C."
Security and Data Isolation Requirements Never return another customer's memories. Validate customer_id at the API boundary. Keep Hindsight credentials out of source code. Use environment variables/secrets for deployment configuration. Use HTTPS for production Hindsight communication. Do not expose raw internal storage credentials in SupportReply. Apply appropriate access controls to support-memory data. Provide a customer-memory deletion path through forget(customer_id). Avoid unnecessary logging of full customer messages. Keep product/environment namespaces separated with HINDSIGHT_BANK.
Performance Considerations The memory layer should optimize for the point immediately before response generation.
Target flow:
message
↓
context extraction
↓
fast recall
↓
reasoning
↓
reply
The system should limit recall results to a configurable number such as:
limit = 5
to avoid unnecessarily increasing reasoning context.
The local lexical implementation is sufficient for small histories and testing. Hindsight semantic retrieval is intended for larger/production histories.
Future Extensions 30.1 Semantic recall Move production retrieval from lexical overlap to Hindsight's embedding-based semantic recall.
Example:
"software freezes"
↓
recalls
"application crashes"
30.2 Cross-customer pattern detection
The current memory system intentionally keeps customer histories separate.
A future analytics layer could aggregate anonymized outcome patterns across customers to identify common successful troubleshooting actions without mixing individual customer memories into another customer's recall.
30.3 Proactive escalation
If the same symptom has repeatedly failed across sessions, the system can use stored outcomes to flag a case for escalation.
Example rule:
same symptom
+
2 or more failed sessions
↓
escalation candidate
This should be implemented as a separate decision layer rather than changing the basic memory contract.
30.4 Multi-product memory isolation
Use:
HINDSIGHT_BANK
to separate products/environments.
Example:
support-prod
support-staging
product-a
product-b
30.5 Voice and asynchronous channels
Because the extraction layer consumes text, future channels can convert their content into text before passing it through the same pipeline:
Voice
↓
Transcription
↓
Memory Agent
Email
↓
Message text
↓
Memory Agent
Async Chat
↓
Message text
↓
Memory Agent
The memory architecture remains the same.
Acceptance Criteria The implementation is ready when all of the following are true:
[ ] Customer memories are stored by customer_id.
[ ] ATTEMPT, INCIDENT, and PROFILE memories are supported.
[ ] Attempt outcomes include worked/failed/in-progress.
[ ] Relevant memories are recalled before reply generation.
[ ] Current messages are not recalled as their own history.
[ ] Known successful fixes are available to the reasoning engine.
[ ] Known failed attempts can be skipped.
[ ] New memories are extracted automatically.
[ ] Multiple failures in one sentence are correctly associated.
[ ] Resolved incidents can be represented.
[ ] InMemoryStore works without external services.
[ ] HindsightStore implements the same storage protocol.
[ ] Hindsight configuration is environment-based.
[ ] Customer isolation is tested.
[ ] Hindsight wire format is tested using an injected/mock HTTP client.
[ ] Deterministic fallback behavior exists.
[ ] SupportReply exposes recall/audit metadata.
[ ] Memory deletion is available through the storage interface.
[ ] Production configuration does not require code changes.
Definition of Done
A developer can consider the first implementation complete when:
A new customer can send a support message.
The system extracts and stores structured memories.
The customer returns later with the same problem.
The system recalls only that customer's relevant history.
The reasoning layer receives the recalled memories before generating the reply.
Known successful fixes are surfaced.
Known failed steps are avoided.
The current message is stored only after recall/reply.
The same application works with InMemoryStore and HindsightStore.
Automated tests verify memory extraction, ordering, isolation, storage, and fallback behavior.
Reference Interaction
First visit
Customer:
My Acme PDF Suite crashes whenever I upload a large PDF.
Agent:
I'll help troubleshoot the PDF upload issue.
Customer reports previous attempts
Customer:
I restarted it but that did not help.
Memory:
ATTEMPT
Restart application
failed
Customer:
Cleared the cache, still crashing.
Memory:
ATTEMPT
Cleared application cache
failed
Customer:
Updated the app and it fixed it.
Memory:
ATTEMPT
Update application
worked
Returning customer
Customer:
The application is crashing again when I upload a PDF.
Recall
worked:
Update application
failed:
Restart application
Cleared application cache
Agent context
Returning customer: yes
Known fix: Update application
Known failures:
Restart application
Cleared application cache
Response direction
Use the known fix first.
Do not repeat known failed steps.
If the known fix no longer works, preserve the history and escalate with context.
Architecture Principle The central design principle is:
Recall first, respond second, learn third.
In implementation terms:
memories = store.search(customer_id, current_message)
reply = engine.reply(
message=current_message,
memories=memories,
)
new_memories = extractor.extract(current_message)
for memory in new_memories:
store.add(memory)
This ordering is a core architectural requirement, not merely an implementation detail.
Source Alignment This specification is derived from the supplied project notes:
The problem statement identifies stateless support, repeated troubleshooting, lost context, and the need for retrieval immediately before the agent replies.
The memory design defines ATTEMPT, INCIDENT, and PROFILE records and requires per-customer storage.
The Hindsight architecture defines the shared MemoryStore protocol, InMemoryStore/HindsightStore split, HTTP operations, bank isolation, and environment-based backend switching.
The interaction example defines the returning-customer flow, known-fix/known-failure behavior, recall metadata, and one-shot outcome extraction.
The future-scope notes define semantic recall, anonymized cross-customer pattern detection, proactive escalation, multi-product isolation, and voice/async support.
References
Hindsight GitHub: https://github.com/vectorize-io/hindsight
Hindsight Documentation: https://hindsight.vectorize.io/
Vectorize — What Is Agent Memory: https://vectorize.io/what-is-agent-memory
HTTPX: https://www.python-httpx.org
Top comments (0)