I've been reading through dev communities lately, and this exact topic keeps showing up in different forms: context windows hit a million tokens, ...
For further actions, you may consider blocking this person and/or reporting abuse
The context rot part is the one that changed how I work. Long threads get worse before they get full, and the failure looks like a confident wrong answer, so you only notice it much later.
The one I keep hitting is reprocessing. Same project, same files, paid for again every session. That is not a fitting problem and no bigger window fixes it.
Reprocessing is the one I underrated while writing this. Caching helps but it expires, and the moment anything early in the context changes you pay for the whole thing again, so the cost isn't really per session, it's per edit. And the confident wrong answer part is what makes context rot so hard to argue about, since there's no error to point at. If it threw an exception at 60 percent fill, everyone would have fixed it years ago.
The exception line is the useful part. You can build that error yourself, roughly. Keep a few questions with known answers, re-ask them as the thread grows, and watch where the answers start drifting.
Per edit rather than per session also explains why a long chat feels cheap right up until you go back and fix one thing near the top. That edit is the expensive keystroke and nothing tells you.
The canary idea is good and I hadn't thought to run it live. A few questions with known answers, re-asked as the thread grows, gives you the error signal you otherwise don't get, and it costs almost nothing. My guess is the drift point isn't fixed either, it moves depending on how much unrelated stuff is sitting in the thread, which would make it worth re-running rather than measuring once. The keystroke point is the part that stuck with me though. Editing near the top is the one action where the cost is immediate and invisible at the same time, and every tool I use presents it as if it's free.
Re-running rather than measuring once is right, and there is a catch inside it. The canary is part of the thread too, so every re-ask adds to the thing you are measuring. Keep it short, word for word identical each time, and read it as a trend rather than a number.
On the keystroke, the reason it looks free is that the counter you can see is tokens and what you actually spent is cache. Nothing in the interface shows the second one, so there is nothing on screen that could have warned you.
The canary being part of the thread is the bit I'd missed. In an offline setup you dodge it, since you can ask each question against a clean copy of the same context and nothing accumulates. Live you can't, so your third re-ask is measuring a slightly different thread than your first. Trend rather than number is the right way to read that. Word for word identical matters for a second reason too, since even a small rewording changes what the question is testing, and then you can't tell drift from a different question. On the cache thing, what gets me is that the visible number isn't just incomplete, it points the wrong way. Token count climbs slowly and smoothly, and the cost of an edit near the top is a spike that never shows up anywhere. A counter that moves calmly while the bill jumps is worse than no counter at all.
The RAM-vs-storage framing is the right one, and the "next session starts at zero" point is the one the 10M-token marketing skips. A context window is working memory: it's there for the duration of the invocation and then it's gone. Nothing about it persists.
Where I hit this in practice is the VRAM wall, not the token limit. On a 2x3090 the context size that actually fits fully in VRAM is a hard number ā push past it and the weights start offloading to system RAM and throughput collapses. So "how big can I make the context" has a concrete answer on my hardware, and I measure it rather than guess. The benchmark tab in homelab-monitor runs each model across a context ladder and reports the largest context that still fits in VRAM ā that's the cap I actually set, and it's the number that decides how much of "memory" I can keep in-context vs. have to persist.
The VRAM wall is a nicer problem to have than the one most people in this thread are describing, because it fails loudly. Throughput collapsing is impossible to miss. The ceiling everyone else is hitting is a quality one, and it sits well below the advertised limit with nothing announcing it. Your context ladder is the same shape as something another commenter suggested here, which was keeping a few questions with known answers and re-asking them as the thread grows to see where the answers start drifting. One measures where it stops fitting, the other measures where it stops being useful, and my guess is the second number is lower than the first on your setup. Have you ever run the ladder with an accuracy check on top, rather than just fit and throughput?
The uncached share spike is a diagnostic for what should have been in persistent memory instead of the context window. Facts that don't change between sessions don't need to be reprocessed every time something upstream shifts.
We built DataGrout's Logic tools around this. Facts stored outside context, retrieved per task. The reprocessing cost disappears because those tokens never entered the window in the first place.