<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: jj_uc</title>
    <description>The latest articles on DEV Community by jj_uc (@uchan135).</description>
    <link>https://dev.to/uchan135</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3839751%2F91b485cd-3011-4fb5-ac0e-a5a1cbafadc0.png</url>
      <title>DEV Community: jj_uc</title>
      <link>https://dev.to/uchan135</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/uchan135"/>
    <language>en</language>
    <item>
      <title>With Context Windows This Large, Why Do We Still Need Memory?</title>
      <dc:creator>jj_uc</dc:creator>
      <pubDate>Sun, 06 Sep 2026 11:43:19 +0000</pubDate>
      <link>https://dev.to/uchan135/with-context-windows-this-large-why-do-we-still-need-memory-5d8m</link>
      <guid>https://dev.to/uchan135/with-context-windows-this-large-why-do-we-still-need-memory-5d8m</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq1x0qmfo3lvn3qbvyeij.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq1x0qmfo3lvn3qbvyeij.png" alt=" " width="799" height="329"&gt;&lt;/a&gt;&lt;br&gt;
I've been reading through dev communities lately, and this exact topic keeps showing up in different forms: context windows hit a million tokens, so is memory even necessary anymore. Enough people are arguing both sides that I wanted to actually dig into it and put together my own take instead of just picking whichever post I read most recently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Just How Big Are They Now?
&lt;/h2&gt;

&lt;p&gt;A few years ago, a few thousand tokens felt generous. Now, 1 million is the baseline. Meta released the 10-million-token Llama 4 Scout last year, and a startup called Magic built a 100-million-token model (&lt;code&gt;LTM-2-mini&lt;/code&gt;). That means about 10 million lines of code or 750 novels can fit into a single prompt. At this point, you have to wonder what is left for a separate memory system to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Camp 1: It's Already Enough
&lt;/h2&gt;

&lt;p&gt;Fabio Akita pointed out that looking at the leaked Claude Code source code, Anthropic's own coding agent doesn't use a vector DB at all; it just uses the file system and grep. By his math, a 200,000-token query costs about $0.63 including caching, which is cheaper long-term than maintaining a vector DB pipeline. Meta made the same bet, marketing the Llama 4 Scout as capable of holding years of chat history "without a vector store."&lt;/p&gt;

&lt;p&gt;It's an attractive argument, and honestly, the cost aspect is the strongest part of this camp. But I think the question this camp is answering is narrower than reality. "Do we need a vector DB?" and "Do we need memory?" are different questions. Akita's point is really about search complexity, using grep instead of embeddings, not about whether state needs to persist between sessions. Meta's marketing also conveniently skips over the fact that a session with 10-million-token chat history eventually ends, and the next session starts at zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  Camp 2: Memory Solves a Completely Different Problem
&lt;/h2&gt;

&lt;p&gt;Mem0 compares the context window to RAM rather than storage: the moment a session ends, everything inside vanishes. Redis puts it more sharply: agents don't fail because a single invocation lacks space, but because they lack continuity, they can't carry what they learned in one session over to the next. No window size fixes this. Even a 100-million-token model completely forgets you the moment you open a fresh chat.&lt;/p&gt;

&lt;p&gt;The most reliable evidence here is Chroma's "context rot" research. Testing 18 models (including GPT-4.1, Claude 4, Gemini 2.5, Qwen3), they found that performance degrades as inputs get longer, well before hitting the model's actual limits. A single piece of irrelevant distractor info noticeably drops accuracy. In some tests, a scrambled mess of info actually performed better than an organized one. This is what genuinely convinced me: fitting into a window and a model actually utilizing it well are two different things, and sellers of larger windows have plenty of incentive to blur that boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  And Camp 3: The Window Is Simply Too Small
&lt;/h2&gt;

&lt;p&gt;Factory.ai points out that current 1-million to 2-million-token models are already smaller than their enterprise customers' codebases. CloudGeometry goes a step further, arguing that even if you gave them 100 million tokens tomorrow, it wouldn't be enough because codebases are graph structures while context windows are linear, no matter how large they get, that structural topology disappears.&lt;/p&gt;

&lt;p&gt;What's fascinating is that Factory and CloudGeometry draw the exact same fact, "therefore, we need better search," while Magic and Meta go with, "therefore, let's build bigger windows." Same observation, opposite prescriptions. It's also worth noting that the "bigger window" crowd usually sells models, while the "better search" crowd usually builds things on top of other people's models.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Conclusion
&lt;/h2&gt;

&lt;p&gt;Context windows need to keep growing, I'm not arguing against that. What I disagree with is treating "bigger windows will fix memory problems" as a given. These are two separate investments, and relying entirely on one to substitute for the other doesn't work. Why? Because a bigger window solves exactly one problem: &lt;strong&gt;the fitting problem&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It does nothing for session amnesia, accuracy loss from irrelevant tokens, or the cost of reprocessing the exact same context on every single invocation. Memory has to evolve on its own to solve these three, and scaling the context window won't do that work for it. So the real paradigm isn't "context versus memory." Rather, while context handles one problem, memory must solve the remaining three through its own evolution, not by waiting for someone else's context window to get bigger.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5cz26jr41qonl4e486jf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5cz26jr41qonl4e486jf.png" alt=" " width="799" height="495"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Strongest Counter-Argument to My Point
&lt;/h2&gt;

&lt;p&gt;I don't want to give memory a free pass, so let's briefly counter my own argument. Memory poisoning is a real attack where someone sneaks mundane content into an agent's long-term memory to plant false info, which the agent then retrieves and trusts in a completely unrelated situation later on. MINJA, presented at NeurIPS 2025, achieved this with just a few mundane queries.&lt;/p&gt;

&lt;p&gt;Add in silent failures (memory retrieves the wrong thing, yet the model speaks with absolute confidence and no errors) and stale information (studies measure that forcing answers yields stale info 15 to 40 percent of the time), and you have genuine reasons to doubt anyone saying, "just throw memory at it."&lt;/p&gt;

&lt;p&gt;Yet my answer doesn't change, and here is why. Silent failures and stale info are mistakes, not malicious acts. They are mostly fixed by displaying confidence scores and recency on retrieved memories, and automatically invalidating old info when new data arrives. This isn't a design flaw, it's engineering maturity, and it's precisely what improves as memory technology matures. Security, however, never fully goes away, just like defense and offense evolving together, much like anti-spam filtering never truly ends. But saying "this requires ongoing work" is different from saying "memory is a dead end." It's simply a trade-off you must accept the moment you ask an agent to persist anything, via memory or otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping Up
&lt;/h2&gt;

&lt;p&gt;So let's sum it up: if your agent needs to remember anything past the current session, waiting for context windows to get bigger is betting on the wrong lever.&lt;/p&gt;

&lt;p&gt;Change my mind. What is actually breaking in your production environment right now, massive contexts, RAG, structured memory, or something entirely different?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>discuss</category>
    </item>
    <item>
      <title>How Do ChatGPT and Claude Actually "Remember" Our Conversations?</title>
      <dc:creator>jj_uc</dc:creator>
      <pubDate>Thu, 03 Sep 2026 08:13:41 +0000</pubDate>
      <link>https://dev.to/uchan135/how-do-chatgpt-and-claude-actually-remember-our-conversations-6jm</link>
      <guid>https://dev.to/uchan135/how-do-chatgpt-and-claude-actually-remember-our-conversations-6jm</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;We’ve all experienced the frustration of yelling at an LLM, &lt;em&gt;"I literally just told you that!"&lt;/em&gt; But on the other hand, there are moments when it casually recalls a minor detail you mentioned weeks ago, making you wonder, &lt;em&gt;"Wait, how did it remember that?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Why is LLM memory so inconsistent? Is it actually storing our chat history somewhere, or is something else happening under the hood? Let's break down how modern AI services handle context and memory from a technical perspective.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. LLMs Don't Actually "Remember": The Reality of Context Windows
&lt;/h2&gt;

&lt;p&gt;The first thing to understand is that LLMs do not possess human-like long-term memory stored in an internal state. &lt;/p&gt;

&lt;p&gt;At their core, LLM-based chat applications are &lt;strong&gt;stateless&lt;/strong&gt;. Every single time you send a message, the system passes the &lt;strong&gt;entire conversation history up to that point&lt;/strong&gt; back into the model along with your new prompt. &lt;/p&gt;

&lt;p&gt;The bottleneck here is the &lt;strong&gt;Context Window&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Think of it like a desk with limited workspace. &lt;/li&gt;
&lt;li&gt;As the conversation grows and hits the token limit, older messages get pushed off the edge of the desk (pruning) and disappear from the model's field of view entirely.&lt;/li&gt;
&lt;li&gt;If the AI suddenly forgets a core architectural constraint you established earlier and starts hallucinating bad code, it's usually because those earlier tokens have fallen out of the active context window.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. Surviving Long Conversations: Summarization and Pruning
&lt;/h2&gt;

&lt;p&gt;So, do long conversations just completely wipe their early history? &lt;/p&gt;

&lt;p&gt;Not quite. When a conversation approaches the window size limit, services typically rely on &lt;strong&gt;background summarization&lt;/strong&gt;. The system compresses older messages into a brief summary to save tokens while attempting to preserve the main narrative flow.&lt;/p&gt;

&lt;p&gt;However, this introduces a severe engineering trade-off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Loss of Granularity:&lt;/strong&gt; During summarization, hyper-specific details—such as exact variable names, strict typing constraints, or edge-case handling logic—inevitably get squashed.&lt;/li&gt;
&lt;li&gt;If you find yourself asking, &lt;em&gt;"Why is it still recommending that function when I explicitly told it not to?"&lt;/em&gt; chances are that specific detail was lost in the summarization pipeline.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Persistent Memory Across Sessions: How Services Differ
&lt;/h2&gt;

&lt;p&gt;What about information that persists across weeks or in brand-new chat sessions? This is where the product design and underlying system architectures of different AI providers diverge.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI ChatGPT: Global User Profiling
&lt;/h3&gt;

&lt;p&gt;ChatGPT’s "Memory" feature continuously analyzes conversations in the background to build and update a &lt;strong&gt;global profile summary&lt;/strong&gt; of your preferences, tech stack, and coding style. When you open a new chat session, this summary is injected into the system prompt. This is why memory updates often take effect asynchronously or in subsequent sessions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Anthropic Claude: Structured Context &amp;amp; Projects
&lt;/h3&gt;

&lt;p&gt;Claude takes a slightly different approach through features like 'Projects', allowing developers to isolate files, documentation, and specific system instructions into a dedicated context container. When you correct Claude on a specific rule, it tends to apply updates more immediately through targeted prompt injection and RAG (Retrieval-Augmented Generation) pipelines.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Unified Limitation: Fragmented State
&lt;/h3&gt;

&lt;p&gt;The common denominator across all these systems is &lt;strong&gt;state fragmentation&lt;/strong&gt;. Information or preferences you teach ChatGPT remain completely unknown to Claude. It's the ultimate siloed developer experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;If you throw an infinite amount of context and code at a model, token costs and latency explode. If you summarize too aggressively, you lose all the engineering details. Current LLM services are constantly iterating to find the optimal trade-off between &lt;strong&gt;token efficiency, inference speed, and memory fidelity&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;How do you manage your AI interactions and prompt context in your daily workflow? Have you ever hit a bizarre debugging loop because of context window limitations? Let's discuss in the comments below!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>discuss</category>
    </item>
  </channel>
</rss>
