AI agents are good at responding to the conversation in front of them. The harder problem is remembering something after that conversation ends.
A user might tell a support agent:
- their preferred contact method
- an account identifier
- a product configuration
- the issue they previously reported
Without persistent memory, the next session begins from zero and the agent asks for the same information again.
In this tutorial, we’ll build a small Python CLI that sends a conversation to Telnyx Agent Memory, waits for the facts to be extracted, and then recalls those facts with a natural-language query.
The complete example is available here:
github.com/team-telnyx/telnyx-code-examples/tree/main/persistent-ai-agent-memory
What we’re building
The example follows a simple three-step workflow:
Conversation transcript
|
v
1. Ingest transcript
|
v
2. Poll async operation
|
v
3. Recall ranked facts
The CLI:
- Submits a support conversation for a profile.
- Receives an asynchronous operation ID.
- Polls until fact extraction finishes.
- Asks what the user’s preferred contact method is.
- Prints the matching memories in relevance order.
This pattern gives an agent continuity without placing every previous conversation into its prompt.
The Agent Memory model
Agent Memory organizes information using a few core concepts:
- Namespace: an isolation boundary for an application or environment.
- Profile: the person or entity the memories describe.
- Source: the original session or fact submitted to the API.
- Memory: an individual fact extracted from a source.
- Operation: an asynchronous write job that can be tracked to completion.
The example uses the default namespace and a profile named user_123.
A real application could use an existing customer, account, or caller identifier as the profile ID.
The three API calls
Every endpoint is relative to:
https://api.telnyx.com/v2/ai/memory
The example uses these calls:
POST /namespaces/{namespace}/profiles/{profile_id}/ingest
GET /namespaces/{namespace}/operations/{operation_id}
POST /namespaces/{namespace}/profiles/{profile_id}/recall
Let’s walk through each one.
1. Ingest a conversation
First, load the API key and prepare the authentication headers:
import os
import requests
from dotenv import load_dotenv
load_dotenv()
API_KEY = os.environ["TELNYX_API_KEY"]
BASE_URL = "https://api.telnyx.com/v2/ai/memory"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
"Accept": "application/json",
}
Here is a shortened version of the support transcript:
transcript = [
{
"role": "user",
"content": "Messages are failing for some international destinations.",
},
{
"role": "assistant",
"content": "Let me check your account configuration.",
},
{
"role": "user",
"content": "My preferred contact method is email at user@example.com.",
},
]
Submit the messages to a profile:
namespace = "default"
profile_id = "user_123"
session_id = "demo-session-001"
response = requests.post(
f"{BASE_URL}/namespaces/{namespace}/profiles/{profile_id}/ingest",
params={"session_id": session_id},
headers=headers,
json={"messages": transcript},
timeout=30,
)
response.raise_for_status()
operation_id = response.json()["data"]["operation_id"]
print(f"Ingest accepted: {operation_id}")
The API returns 202 Accepted, not a completed memory.
A typical response looks like this:
{
"data": {
"operation_id": "op_abc123",
"profile_id": "user_123",
"session_id": "demo-session-001",
"source_id": "src_xyz789"
}
}
Fact extraction happens asynchronously. The operation_id is the handle we use to track it.
2. Poll until the write finishes
A memory cannot be recalled until its write operation has completed.
The sample checks the operation every two seconds and stops after reaching a terminal state:
import time
terminal_statuses = {"completed", "failed", "cancelled"}
while True:
response = requests.get(
f"{BASE_URL}/namespaces/{namespace}/operations/{operation_id}",
headers=headers,
timeout=30,
)
response.raise_for_status()
status = response.json()["data"]["status"]
print(f"Operation status: {status}")
if status in terminal_statuses:
break
time.sleep(2)
The normal progression is:
pending -> processing -> completed
Production code should also enforce a timeout. The complete example stops polling after 60 seconds and raises an error rather than waiting forever.
This explicit operation lifecycle is useful because your application can distinguish between:
- a write that is still processing
- a write that completed
- a write that failed
- a write that was cancelled
If the client restarts while polling, it can continue checking the same operation instead of submitting the entire conversation again.
3. Recall relevant facts
Once ingestion completes, query the profile using natural language:
response = requests.post(
f"{BASE_URL}/namespaces/{namespace}/profiles/{profile_id}/recall",
headers=headers,
json={
"query": "What is the user's preferred contact method?",
"top_k": 10,
},
timeout=30,
)
response.raise_for_status()
facts = response.json()["data"]
The results are ranked by relevance:
for fact in facts:
print(f"[score={fact['score']}] {fact['text']}")
A result may look like this:
{
"id": "mem_abc123",
"text": "The user's preferred contact method is email at user@example.com.",
"recorded_at": "2026-10-05T12:00:05Z",
"score": 0.92
}
The response contains the extracted fact rather than requiring the application to search the original transcript itself.
That fact can now be added to an agent’s context when a new conversation begins.
Run the complete example
Clone the repository:
git clone https://github.com/team-telnyx/telnyx-code-examples.git
cd telnyx-code-examples/persistent-ai-agent-memory
Create a .env file:
echo "TELNYX_API_KEY=your_telnyx_api_key_here" > .env
Install the dependencies:
pip install -r requirements.txt
Run the CLI:
python app.py
The example requires only a Telnyx API key. It does not require a phone number, messaging profile, or separate database.
Details worth handling in production
The sample includes a few useful safeguards that are easy to overlook.
Encode path identifiers
Namespace, profile, and operation identifiers should be encoded before being inserted into a URL:
from urllib.parse import quote
profile_id = quote(profile_id, safe="")
This prevents reserved characters in an existing customer identifier from changing the request path.
Validate inputs before sending them
The example enforces several API constraints:
if len(session_id) > 128:
raise ValueError("session_id must be <= 128 characters")
if not query or len(query) > 4096:
raise ValueError("query must be between 1 and 4096 characters")
if not 1 <= top_k <= 100:
raise ValueError("top_k must be between 1 and 100")
Failing locally gives developers a clearer error than sending an invalid request and debugging it later.
Do not recall immediately after ingesting
An accepted write is not yet a recallable memory.
If recall returns an empty list immediately after ingestion, first confirm that the operation reached completed.
Choose isolation boundaries deliberately
Profiles are isolated from one another, and namespaces provide an additional boundary between applications or environments.
For example:
namespace: production-support
profile: customer_1024
You might use separate namespaces for development and production, while each customer or caller receives a unique profile inside that namespace.
Custom namespaces must be created before use. The default namespace is available without a separate provisioning step.
Plan for memory deletion
Persistent memory should come with a deletion strategy.
The Agent Memory API includes endpoints for deleting an individual source or an entire profile. That allows an application to remove one imported session or erase everything associated with a profile when required.
Where this pattern fits
The same ingest, poll, and recall workflow can support:
- Customer support agents that remember previous issues and preferences
- Sales assistants that maintain account context across conversations
- Scheduling agents that recall availability or communication preferences
- Voice agents that recognize returning callers
- Internal copilots that retain user-specific working context
The communication channel can change while the memory model stays the same. SMS, voice, email, and browser chat can all resolve to the same profile identifier.
Final thoughts
A longer prompt is not the same thing as long-term memory.
Prompt context is temporary. Persistent memory gives an agent a durable place to store what it has learned and retrieve only the facts relevant to the next interaction.
The implementation here stays intentionally small:
ingest -> poll -> recall
That is enough to move an agent from isolated sessions toward real continuity.
Top comments (0)