A language model can make a game character say almost anything, which is the easy part. The hard part shows up the second time a player walks up to that character, and the NPC greets them like a stranger.
Memory is what separates a tech demo from a character players care about. Here is a practical way to layer it, from the conversation buffer up to structured facts, along with the parts of the surrounding architecture it plugs into.
Where Memory Fits in the NPC Pipeline
A typical LLM NPC system breaks into four modules. A context builder assembles the prompt, an LLM interface talks to a cloud API or a local inference engine, a response parser turns raw output into dialogue, emotion tags and game actions, and a memory manager decides what the character remembers.
The memory manager never talks to the model directly. Its job is to hand the context builder the right few hundred tokens of history, inside a strict token budget, so the prompt stays small and the character still feels continuous. This overview of how LLM powered NPCs are built walks through all four modules and how they connect.
Layer One: A Conversation Buffer That Summarizes
The simplest memory is a rolling buffer of recent messages, stored with role labels and timestamps. For most NPC conversations, 10 to 20 message pairs is enough for natural flow.
The mistake is what happens at the edge of the buffer. Dropping the oldest messages creates a sudden cutoff where the NPC forgets what you were just talking about. Compressing them into a two or three sentence summary at the top of the history keeps the topic alive for a fraction of the tokens, and it only costs an extra model call when the buffer overflows.
Layer Two: Vector Retrieval Across Sessions
Long term memory usually lives in a vector database such as ChromaDB, Pinecone, Qdrant or Weaviate. When a session ends, embed it and store it with a short text summary, a timestamp and the NPC's id. When the next session starts, embed the player's opening line and pull the closest matches back into the prompt.
Two settings matter more than the database choice. Store segments of 3 to 5 message pairs grouped by topic, since whole conversations retrieve too loosely and single messages bloat the index. And set a minimum similarity score with a cap of 2 to 4 memories, because irrelevant recollections waste tokens and confuse the character.
Layer Three: Structured Facts That Always Load
Semantic search can miss the facts that matter in every conversation: the player's name, a promise they made, their standing with the NPC. Extract those after each session into simple labeled records, such as the player promised to bring iron ore from the northern mines, and load them every time regardless of topic.
Some systems add memory decay on top, letting older, less important memories fade so retrieval stays focused. The step by step version, including how to format retrieved memories so the NPC references them naturally, is in this guide to giving NPCs memory and context.
None of these layers needs a bigger model. A well prompted small model with a good memory manager will feel more alive than a frontier model that forgets you every time you leave the room, and that is usually the cheaper problem to solve first.
Top comments (0)