You've hired a goldfish to guide you through the aquarium. He's enthusiastic, knowledgeable, and surprisingly articulate for someone with gills. There's just one small problem: he can only remember the last seven tanks you've visited together. Stand in front of the seahorses and he'll expertly recall everything about the stingrays, the coral reef, the tropical zone, and the four exhibits before that. But the jellyfish tank where you started? The one where you mentioned you're terrified of tentacles? Gone. Vanished from his mind completely. He will, with absolutely no sense of irony, suggest you circle back to see the jellyfish again.
This is how AI chatbots experience conversation with you.
The goldfish has a very specific memory problem
The technical term is a context window, which is the maximum amount of text an AI model can "see" at any given moment. That text gets measured in tokens (chunks of roughly three to four characters each, not full words, just fragments). A 4,000-token window holds about 3,000 words total, including everything you've written and everything the AI has responded with.
Here's the cruel part: once you exceed that limit, the oldest information doesn't get filed away for later or condensed into helpful notes. It gets shoved off the edge of a cliff. The model literally cannot see it anymore. Those early messages have ceased to exist in its world.
ChatGPT-3.5 has a 4,096-token limit, which sounds generous until you're 20 messages deep into brainstorming a project and realize the goals you outlined at the start have fallen into the void. GPT-4's 128,000-token window (roughly 96,000 words) gives you far more room before the amnesia kicks in, but even that has edges. Every window does.
By the time you reach the sharks, you're a stranger again
You established your name, your project constraints, your budget, and three absolute deal-breakers in your opening message. Twenty exchanges later, the AI cheerfully suggests something that violates deal-breaker number one. You point this out. It apologizes. Five messages after that, it asks for your name.
You are not talking to an AI that's being careless. You are talking to an AI that is meeting you for the first time, over and over, because the early part of your conversation has dropped out of the window.
Say you're debugging code with Claude across 25 messages. You defined a variable called userPreferences way back in message two. In message 23, you reference it. Claude tells you it doesn't see that variable anywhere in your code. This feels like gaslighting, but it's simpler than that: message two fell out of the window six exchanges ago. Claude is telling you the truth about what it can see, which is an incomplete picture.
The repetition gets eerie. You'll reject a suggestion, explain exactly why it won't work, and move on. Ten messages later, the AI will propose the same idea again with fresh enthusiasm. You'll say no a second time. It will apologize. Another dozen exchanges pass, and here comes that same suggestion, now dressed up with slightly different wording.
The AI has no memory of your earlier "no" because that rejection has slipped out of view. Every time it encounters your current question, it's working from scratch with only the visible context. Your previous refusal might as well have happened to someone else.
You tell a chatbot in message one that you're vegetarian. You're planning a dinner party together. By message 35, having discussed appetizers, drinks, and dessert, it enthusiastically suggests beef Wellington as the main course. Your dietary restriction dropped out of the window somewhere between the cheese plate and the cocktail napkins. The AI isn't messing with you. It genuinely has no idea.
So what can YOU do with this?
First, stop trying to have one infinitely long conversation about everything. Start fresh chats for new topics instead of threading 40 messages deep. Each conversation has a shelf life.
When you're stuck in a genuinely long session (debugging complex code, drafting a document through multiple revisions, planning something with lots of moving parts), repeat your critical constraints every 10 to 15 exchanges. Yes, this feels redundant. Do it anyway. A quick "Remember: this is for a teen audience, needs to stay under $500, and cannot include dairy" every so often keeps essential context alive.
Use custom instructions or system prompts where available. Many chatbots let you set persistent instructions that sit outside the regular window, so they don't get pushed out by the conversation itself.
Copy and paste key information back into the chat when you notice the amnesia. Those signs are easy to spot: contradictions of earlier statements, questions you've already answered, or suggestions you've already rejected.
Pick models with larger windows for multi-step projects. Claude 3 Opus and GPT-4 Turbo both offer significantly more room than their smaller cousins. If you're going to have a 50-message work session, you want the biggest tank available.
Structure long documents by working in smaller chunks rather than pasting 50 pages at once and asking for feedback. Break it into sections, work through each, and start fresh conversations as needed. When using ChatGPT for a long research project, create separate chats for literature review, outline, and drafting instead of one 100-message marathon.
TL;DR
- Context window is the maximum amount of text an AI can process at once, measured in tokens (roughly 3 to 4 characters each), typically between 3,000 and 96,000 words depending on the model.
- Once the conversation exceeds that limit, the oldest messages disappear completely from the AI's view, causing it to forget instructions, repeat rejected ideas, and ask previously answered questions.
- Work around it by starting fresh conversations for new topics, repeating critical context every 10 to 15 exchanges, using custom instructions, and choosing models with larger windows for complex projects.
At least he never gets bored of the seahorses.
Top comments (0)