On September 14, Salesforce published a customer service case study on the U.S. Transportation Security Administration (TSA). TSA's AI Agent, Ace, now handles approximately 100,000 tourist inquiries per month, with roughly 96% resolved without real human agents. For the remaining cases, human agents access the full prior conversation history in a unified workspace before continuing the interaction.
This pattern shows up constantly in everyday customer support: When a customer returns, they expect the system to remember what they already said.
Consider a headphone exchange in progress. A user requested a replacement through live chat a few days earlier — provided a shipping address, and specified "just text me updates." Today, they follow up by email with three words: "Did it ship?"
Answering this requires retrieving the correct order and support ticket, then checking current status. The previously confirmed address and notification preference also need to carry forward — the user shouldn't have to restate them.
This scenario surfaces 3 distinct categories of information: Facts the user has explicitly provided, reusable operational knowledge from handling similar cases, and policy maintained centrally by the business.
For example, "this user prefers SMS updates" is a personal preference, scoped to that individual. "Verify order and policy before processing an exchange" is a generalizable procedure, applicable across users. These operate at different scopes and require separate handling.
Using consumer after-sales support as a working example, this post walks through Python code demonstrating how to combine all three information types into a complete customer-service memory system.
Note: The order, ticketing, logistics, and notification tools referenced in this post are simulated business tools built for demonstration purposes only — MemOS does not ship with these systems built in. Policy text shown here is likewise illustrative; in a production integration, replace it with the business's actual systems and official policy documentation.
Starting with a headphone exchange: 3 scenarios below
We'll begin with a single headphone exchange request. The demo runs in 3 phases:
DAY 1 · Live Chat
customer_001 submits a headphone exchange request. The support Agent looks up the order, checks policy, creates a ticket, and records the delivery time, shipping address, and SMS notification preference.
DAY 4 · Email Ticket
customer_001 follows up on the exchange status using a new conversation_id. This step verifies whether user facts and preferences persist across sessions.
DAY 7 · Live Chat
customer_002 runs into a similar static-noise issue with their headphones. The support Agent tries to reuse the Skill formed earlier, verifying that user data stays isolated while general Skills can still be reused across users.
What we're really testing isn't whether the Agent can store chat history. It's 2 things:
"Can it pick up where it left off with the user it should remember", and "Can experience it has already formed be handed to the next user and reused?"
Start by separating the three types of information
A single after-sales task typically leaves behind 3 completely different kinds of information.
Take DAY 1, when the user submits a headphone exchange request:
- "What the order number is?"
- "What's wrong with the headphones?"
- "Whether a ticket has been created?"
- "What the shipping address is?"
- "Which channel the user wants progress updates on?"
These information answers: "What has already happened with this user?"
Meanwhile, as the support Agent handles the request, it may also form a reusable workflow:
Look up order → Verify policy → Create ticket → Update fulfillment info → Notify via the user's confirmed channel
This information answers: "How is this kind of issue usually handled?"
A third type comes from the business itself:
- What are the return and exchange conditions?
- How long does the warranty last?
- Under what circumstances is an exchange allowed?
These information answers: “What rules apply to this case?”
So this solution works with three types of context:
| INFORMATION TYPE | WHAT IT STORES | WHAT IT SOLVES | OWNERSHIP |
|---|---|---|---|
| User Memory | Stores orders, defects, ticket status, shipping address, and notification preferences. | Answers "where is this customer in the process?" | Belongs only to the current user_id. |
| Agent Skill | Stores general workflows that the support Agent distills from complete task tracks. | Answers “what steps does this kind of task usually require?” | Belongs to a stable agent_id. |
| Policy Knowledge Base | Official rules for returns, exchanges, warranty, logistics, invoicing, and more. | Verify “how this should be handled according to the rules?” | Maintained centrally by the business. |
When the boundaries between these three blur together, two typical failures follow:
User A's address gets exposed to User B, or one user's one-off request gets mistaken for a general workflow that applies to everyone.
So what matters isn't storing every historical message. It's putting each piece of information in the right place.
Solution overview
The approach comes down to two writes, one combined recall.
After each task, the same task record is written twice: once to user memory and once to Agent Skill, with allow_memory_view specifying a different memory type for each. When the user comes back, a single /search/memory call retrieves the current user’s memory, Agent Skills, and policy knowledge together. The current session history is still maintained by the support application itself.
Specifically:
- User memory is scoped by
user_id. - Agent Skill is reused across users via
agent_id. - The policy knowledge base joins the same search through
knowledgebase_ids.
User request
│
├─ Agent reads the current session history from the business application
│
├─ One combined recall: agent_id + knowledgebase_ids
│ filter selects the current user's memory (related_id) and Agent Skill (SkillMemory)
│ → user facts, prefs, Skills, policies
│
├─ Merge the context and run business tools: orders, tickets, notifications, etc.
│
├─ LLM generates the final reply based on tool results
│
├─ Write 1: user_id + agent_id → facts, preferences
│
└─ Write 2: agent_id → request to generate or update a Skill
Before the integration
Before getting hands-on, get your API access and project configuration in place:
- A MemOS API Key, along with access to the policy knowledge base for the project that key belongs to.
- An OpenAI or OpenAI-compatible LLM endpoint.
- "Create independent memory for Agents" enabled for your project in the MemOS Dashboard.
“Create independent memory for Agents” is off by default. Once enabled, you can write and retrieve Agent Skills using agent_id as an independent entity.
Give the Agent the rules first
When a user says "my headphones have static noise", the support Agent can’t rely on user memory alone. It also needs to know which after-sales rules apply.
So during initialization, start by creating a policy knowledge base, uploading your after-sales policy documents, and waiting for file processing to finish:
def create_policy_knowledge_base():
headers = {
"Content-Type": "application/json",
"Authorization": f"Token {MEMOS_API_KEY}",
}
# 1. Create Knowledge Base:POST /create/knowledgebase
res = requests.post(
f"{MEMOS_BASE_URL}/create/knowledgebase",
headers=headers,
json={
"knowledgebase_name": "Consumer After-Sales Policy Knowledge Base",
"knowledgebase_description": "Consumer Return, Exchange, Warranty, Shipping, and Invoice Policies",
},
)
body = res.json()
if body.get("code") != 0:
sys.exit(f"Failed to create the knowledge base:{body.get('message')}")
policy_kb_id = body["data"]["id"]
# 2. Upload the policy document:POST /add/knowledgebase-file
# File content must be base64-encoded and passed as a data URL.
encoded = base64.b64encode(POLICY_DOC_MD.encode("utf-8")).decode("utf-8")
res = requests.post(
f"{MEMOS_BASE_URL}/add/knowledgebase-file",
headers=headers,
json={
"knowledgebase_id": policy_kb_id,
"file": [{
"type": "document",
"name": "consumer-after-sale-policy.md",
"content": f"data:text/markdown;base64,{encoded}",
}],
},
)
if res.json().get("code") != 0:
sys.exit(f"Failed to upload policy document:{res.json().get('message')}")
# 3. Poll file parsing status:POST /get/knowledgebase-file
print("Policy document uploaded. Wait for parsing...")
for _ in range(40):
time.sleep(3)
res = requests.post(
f"{MEMOS_BASE_URL}/get/knowledgebase-file",
headers=headers,
json={"knowledgebase_id": policy_kb_id, "page": 1, "page_size": 20},
)
files = res.json().get("data", {}).get("file_detail_list", [])
statuses = {str(item.get("status", "")).lower() for item in files}
if files and statuses <= {"completed", "available", "failed"}:
if "failed" in statuses:
sys.exit("Policy document parsing failed")
break
print(f"Policy knowledge base is ready":{policy_kb_id}")
return policy_kb_id
The knowledge base ID comes directly from the create endpoint's response, then be passed to the support assistant:
policy_kb_id = create_policy_knowledge_base()
assistant = CustomerServiceAssistant(policy_kb_id)
The policy_kb_id returned by the create endpoint is passed to the support Agent, and joins recall through knowledgebase_ids later.
Here's a easy-to-miss detail: a successful upload doesn’t mean the document is searchable yet.
The file may still be in a parsing state, so the demo waits for file processing to finish before moving on to the next steps.
One single task, distills two kinds of memory: User Memory and Agent Skill
A complete support task contains both the information the user provided and the process the Agent actually carried out.
So when writing back to MemOS, don't submit only the user's question and the final answer. Instead, assemble the user request, tool calls, tool results, and final reply into one complete task record.
user
→ assistant.tool_calls
→ tool
→ assistant
Use role_id to identify who actually spoke:
- User messages use
user_id - Support replies and tool calls use
agent_id
# Using the DAY 1 exchange request as an example: the full message list written back to MemOS for one conversation turn
memory_messages = [
# 1. User message: role_id is the current user
{
"role": "user",
"role_id": user_id, # e.g. customer_001
"content": "The left ear of my noise-canceling headphones has a crackling sound. I'd like to exchange them.",
},
# 2. Tool call: the support Agent looks up the order, role_id is agent_id
{
"role": "assistant",
"role_id": AGENT_ID,
"content": "",
"tool_calls": [{
"id": "call_1",
"type": "function",
"function": {
"name": "query_order",
"arguments": '{"order_id": "20260820-88"}',
},
}],
},
# 3. Tool result: matched to the call above via tool_call_id, no role_id needed
{
"role": "tool",
"tool_call_id": "call_1",
"content": '{"order_id": "20260820-88", "status": "delivered"}',
},
# 4. Final reply: also attributed to the support Agent
{
"role": "assistant",
"role_id": AGENT_ID,
"content": "I've created exchange ticket EX20260821-03 for you. Shipping is free both ways for exchanges...",
},
]
The full track is used for Skill distillation, while user facts and preferences are generated from the same record.
Separating Write Types
User writes and Agent writes use different views, while recall shares a single view configuration:
USER_WRITE_VIEWS = ["detail_factual", "preference"]
AGENT_SKILL_WRITE_VIEWS = ["skill"]
CONTEXT_VIEWS = ["detail_factual", "preference", "skill"]
This way, the same task record can be written to two memory spaces without producing duplicate memories of the same type.
Write user facts and preferences
The first /add/message call passes in both user_id and agent_id:
def add_user_memories(self, messages, user_id, conversation_id, channel):
"""First write: generates facts and preferences from the user's perspective only."""
user_data = {
"user_id": user_id,
"agent_id": AGENT_ID,
"conversation_id": conversation_id,
"info": {"channel": channel, "scene": "consumer_support"},
"allow_memory_view": USER_WRITE_VIEWS,
"messages": messages,
}
# POST /add/message: the write runs as an async task; once it returns a task_id,
# poll /get/status until completed (full implementation in the Demo at the end)
res = requests.post(
f"{MEMOS_BASE_URL}/add/message",
headers=self.headers,
json=user_data,
)
if res.json().get("code") != 0:
print(f" [MemOS] Failed to write user facts and preferences:{res.json().get('message')}")
Memory is owned by the current user_id and also attached to the support Agent‘s memory space, so it can later be precisely selected via related_id during the Agent-level combined recall.
MemOS decides from the conversation content whether a fact or preference has formed. If a given turn doesn’t express a stable preference, it can generate facts only.
Re-write Agent Skill
The second /add/message call passes agent_id and allows Skill generating:
def add_agent_skill(self, messages, conversation_id, channel):
"""Second write: request skill generation or update from the Agent's perspective only."""
skill_data = {
"agent_id": AGENT_ID,
"conversation_id": conversation_id,
"info": {"channel": channel, "scene": "consumer_support"},
"allow_memory_view": AGENT_SKILL_WRITE_VIEWS,
"custom_extract_prompt": {"skill": SKILL_EXTRACT_PROMPT},
"messages": messages,
}
# Still a POST to /add/message, but without user_id, and allow_memory_view contains only skill
res = requests.post(
f"{MEMOS_BASE_URL}/add/message",
headers=self.headers,
json=skill_data,
)
if res.json().get("code") != 0:
print(f" [MemOS] Failed to write Agent Skill: {res.json().get('message')}")
Skill writes are not bound to any user and exist only from the Agent's perspective, so they can be reused across users. This write runs after every turn.
The caller doesn't need to decide up front whether "this turn actually produced a Skill worth capturing." Every task can be submitted to MemOS, which weighs the task trajectory against existing Skills and decides on its own whether to generate, update, or skip.
Turning one exchange case into a general skill
A Skill is meant to capture a way of handling a task that can be reused next time.
You can further constrain what gets extracted as a Skill through custom_extract_prompt.skill:
- Decide whether tasks belong to the same Skill based on business goal, trigger conditions, core toolchain, and success criteria;
- For identical workflows, prefer merging into an existing Skill, and don’t generate a duplicate when there’s no new general information;
- Remove names, addresses, user IDs, order numbers, ticket numbers, and specific dates;
- Replace instance values with parameters such as
order_id,ticket_id, andshipping_address; - Don’t turn a user’s personal preferences or one-off timing requests into general rules;
- Keep only the execution steps that tool results can verify.
For example:
A user asks to “send exchange progress updates by SMS.”
That is a user preference.
Whereas:
“Send progress updates through the notification channel the user confirmed.”
is much closer to a general workflow.
This way, even if the next user would rather receive notifications by email, the Skill can still be reused.
Continuing the same case across sessions
By DAY 4, the user switches to email to follow up on the exchange progress.
This time, the customer service Agent needs three types of information at once:
- Facts and preferences the current user left behind earlier;
- General Skills the current customer service Agent has already captured;
- The after-sales policy currently in effect.
So before generating a reply, the customer service assistant calls /search/memory just once, retrieving all of these in a single request.
def search_memory(self, query, user_id):
"""One combined recall: user facts/preferences + Agent Skill + policy knowledge."""
context_data = {
"query": query,
"agent_id": AGENT_ID,
"knowledgebase_ids": self.knowledgebase_ids,
"include_memory_view": CONTEXT_VIEWS,
"memory_limit_number": 9,
"preference_limit_number": 6,
"filter": {
"user": {
"or": [
{"related_id": [user_id]},
{"memory_type": "SkillMemory"},
]
}
},
}
# POST /search/memory:User facts, preferences, Agent Skill and policy one-time retrieval
f"{MEMOS_BASE_URL}/search/memory",
headers=self.headers,
json=context_data,
)
result = res.json().get("data") or {}
return (
result.get("memory_detail_list", []), # User Facts
result.get("preference_detail_list", []), # User Preferences
result.get("skill_detail_list", []), # Agent Skill
)
The filter here joins two conditions with or: they are alternatives, and both don't need to be satisfied at the same time. related_id selects the current user's memories, keeping facts and preferences isolated between users. memory_type: "SkillMemory" lets Agent Skills be returned across users, unrestricted by who the user is. The policy knowledge base joins the same retrieval through knowledgebase_ids, so no separate call is needed.
Each of the 3 types of information solves its own problem in this retrieval:
- User memory answers "How far had this customer’s case progressed?"
- Agent Skills answer "How is this type of task usually handled?"
- The policy knowledge base answers "What are the current business rules?"
There's one more thing: memory is not real-time business state.
For example, if user memory records that “the exchange ticket has been created,” that only shows how far things had gotten earlier. If the user now asks “Has it shipped yet?”, the customer service Agent still needs to query the logistics or ticketing system for the latest status, then generate a reply based on that result.
New session, new user, and re-verifying the memory
Once the integration is in place, you can check the whole pipeline in 3 phases.
DAY 1: Check the first task
After the exchange task completes, confirm that user facts and preferences are written to the User Cube, and that the same task record is submitted to the Agent Cube for Skill evaluation. If a Skill is generated, also check whether it retains any user-specific information.
DAY 4: Switch to email
Use a new conversation_id and confirm that customer_001 can still recall the earlier order, defect, and notification requirements.
What this verifies: can user facts and preferences persist across sessions?
DAY 7: Switch to a different user
Switch to customer_002 and raise a similar headphone issue.
This verifies two things at once: whether the first user’s personal information stays isolated, and whether the current Agent can reuse the general Skill it has already formed.
These are the two most important checkpoints for the whole solution: remember who should be remembered, and reuse what should be reused.
Reference
MemOS Official Website: memos.openmem.net
GitHub: github.com/MemTensor/MemOS
MemOS Docs: memos-docs.openmem.net
Full Demo: memos-docs.openmem.net/cn/usecase/customer_service_assistant/
Top comments (0)