Building a chatbot that holds a coherent conversation for more than a few turns is harder than it looks. Language models are stateless — every API call starts fresh, with no memory of anything said before. The "memory" in polished chatbot products is maintained entirely by the application layer. Getting that right for production means making a set of architectural decisions up front that compound as the codebase grows.
The Core Problem: Stateless Models and Token Cost
A language model only knows what you put in the prompt. If you want it to remember that a user said "I prefer Go examples" four messages ago, you have to include that exchange in the current request.
The naive approach is full conversation replay: append every past turn to the messages array and send the whole thing on every call.
import os
import anthropic
client = anthropic.Anthropic(api_key=os.environ["LLM_API_KEY"])
def chat(history: list[dict], user_message: str) -> str:
history.append({"role": "user", "content": user_message})
response = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=1024,
system="You are a helpful technical assistant.",
messages=history,
)
reply = response.content[0].text
history.append({"role": "assistant", "content": reply})
return reply
This works until the conversation grows past the context window, or until your token bill does. A 50-turn conversation at 200 tokens per turn costs 10,000 input tokens on every single request — most of it repeated context the model already processed in the previous call.
Strategy 1: Sliding Window
The simplest fix is a sliding window: keep only the last N turns verbatim.
MAX_TURNS = 20 # each turn = 1 user + 1 assistant message
def trim_history(history: list[dict]) -> list[dict]:
max_entries = MAX_TURNS * 2
if len(history) > max_entries:
return history[-max_entries:]
return history
Call trim_history before every LLM request. Input tokens stay bounded regardless of session length.
The downside is abrupt amnesia: the model stops knowing anything said before turn 21. For task-oriented chatbots — support flows, form filling, one-session Q&A — this is usually fine. For anything requiring continuity across a longer conversation, it breaks the experience.
Strategy 2: Summarization-Based Memory
Instead of discarding old context, compress it. When history crosses a threshold, summarize everything before the last N turns and inject that summary into the system prompt.
def summarize_old_turns(old_history: list[dict], client) -> str:
prompt = (
"Summarize the key facts, preferences, and decisions from this conversation "
"in 3-5 bullet points. Be specific and concise.\n\n"
+ "\n".join(
f"{m['role'].upper()}: {m['content']}"
for m in old_history
)
)
response = client.messages.create(
model="claude-3-haiku-20240307",
max_tokens=512,
messages=[{"role": "user", "content": prompt}],
)
return response.content[0].text
def chat_with_memory(
history: list[dict],
warm_summary: str,
user_message: str,
client,
) -> tuple[str, str]:
KEEP_VERBATIM = 10 # turns to keep verbatim
if len(history) > KEEP_VERBATIM * 2:
old = history[: -KEEP_VERBATIM * 2]
new_summary = summarize_old_turns(old, client)
warm_summary = f"{warm_summary}\n{new_summary}".strip() if warm_summary else new_summary
history = history[-KEEP_VERBATIM * 2 :]
system = "You are a helpful technical assistant."
if warm_summary:
system += f"\n\nConversation context:\n{warm_summary}"
history.append({"role": "user", "content": user_message})
response = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=1024,
system=system,
messages=history,
)
reply = response.content[0].text
history.append({"role": "assistant", "content": reply})
return reply, warm_summary
The summarization call uses a faster, cheaper model — it is pure compression, not generation. In practice this approach preserves the semantics that matter while keeping prompt size predictable.
Strategy 3: Semantic Memory with Vector Retrieval
For persistent sessions — where users come back days later and expect continuity — a vector store gives you addressable long-term memory. Instead of replaying the whole session, you retrieve only the turns relevant to the current query.
from sentence_transformers import SentenceTransformer
import numpy as np
class SemanticMemory:
def __init__(self):
self.encoder = SentenceTransformer("all-MiniLM-L6-v2")
self._texts: list[str] = []
self._embeddings: list[np.ndarray] = []
def store(self, text: str) -> None:
vec = self.encoder.encode(text, normalize_embeddings=True)
self._texts.append(text)
self._embeddings.append(vec)
def retrieve(self, query: str, top_k: int = 3) -> list[str]:
if not self._embeddings:
return []
q_vec = self.encoder.encode(query, normalize_embeddings=True)
scores = [float(np.dot(q_vec, e)) for e in self._embeddings]
ranked = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)
return [self._texts[i] for i in ranked[:top_k]]
In production, swap the in-memory list for pgvector, Qdrant, or any other vector database — the interface stays the same, only the backend changes. Store one entry per significant exchange, not every single turn; store facts, not filler.
The Production Architecture
The pattern that holds up under load layers all three strategies:
- Hot memory: the last N turns, verbatim, always in the prompt
- Warm memory: a rolling summary of earlier turns, injected as a system note
- Cold memory: vector-retrieved facts from persistent storage, pulled per query
At each turn: retrieve relevant cold memories, prepend them to the system prompt, append the warm summary, append hot history, append the user message. Total input tokens stay bounded regardless of session length or how long the user has been a customer.
One operational detail teams miss: summarization and embedding writes should be asynchronous. If either happens synchronously in the request path, users feel the latency. Offload them to a background task queue and write through a cache.
For the storage layer, securing these memory backends matters — especially when they contain personal data or proprietary context. Our security hardening checklists cover encryption at rest and access control patterns for both relational and vector databases.
The Takeaway
A chatbot without memory management is a prototype. The model is stateless by design — your application layer owns the continuity. The three strategies (window trimming, summarization, semantic retrieval) are not alternatives; mature systems layer all three.
Start with the sliding window. Add summarization when context loss hurts output quality. Add semantic retrieval when sessions go persistent or cross-session. Do not add complexity before the problem is real — but do design the interface to support it from the start.
I run AYI NEDJIMI Consultants, a cybersecurity consulting firm. We publish free security hardening checklists — PDF and Excel.
Top comments (0)