Most engineering teams working on long-context agents hit the same billing wall around turn twenty. A coding agent runs twenty shell commands, read...
For further actions, you may consider blocking this person and/or reporting abuse
To me the architectural split between "environment output" and "agent intent" makes perfect sense from an infrastructure perspective. If we treat a coding agent like a system administrator, it needs an exact log of its own executed commands (the hard actions) to understand what broke.
However, the 10,000-line output of a crashed grep or a full core dump can be compressed.
It’s essentially separating the audit log from the payload.
That means keeping the decision layer byte-exact while distilling the "environment noise" is a much more robust design pattern than blind semantic truncation.
The audit log versus payload distinction is clean, but the edge case that always bites is where stdout turns into error diagnostics. If a harness collapses a 5,000-line build failure into a one-sentence summary, the agent loses the exact linker symbol or missing header on line 4,812 that tells it why the build broke.
What has worked well in my scripts is keeping the exact exit code plus a bounded head and tail of the stream in context, while dumping the full unedited output to a scratch file on disk. The agent gets the failure trace immediately without burning 20k tokens on compiler noise, and it can grep the scratch path directly if it actually needs the full traceback.
You map classic Linux log management directly to the LLM context window.
Effective troubleshooting never involves reading massive log files sequentially.
You isolate the exit code.
You read the tail.
You grep for specific linker symbols.
Dumping unedited stdout streams to disk treats the agent like a standard Unix process.
This shifts operational burden from expensive context memory to cheap local storage.
The agent uses standard diagnostic tools instead of brute-force reading comprehension.
Solid infrastructure architecture solves model limitations.
One thing I'd add from building in this space: past turn twenty, most of what's worth keeping isn't the raw history at all, it's a handful of facts. What was decided, what was tried and failed, what's still open. Pulling those into a small state block and starting fresh with it has worked better for me than compressing the transcript, both on cost and on the agent repeating old mistakes. Did you measure answer quality alongside token savings?
The ablation compared keeping K recent raw observations alongside compressed LOHA state blocks against naive head-truncation of plain text. At K=1, performance dropped because the model lost immediate execution feedback, while moving from K=3 to K=8 yielded diminishing returns on resolve rates while inflating token usage. The jump to 21.% came from retaining those three recent verbatim tool outputs so the agent had precise file paths and compiler errors for immediate next steps, while the compressed state block preserved long-range task intent.
The benchmark numbers in the post come from the Zou et al. paper on SWE-bench Verified, and they tracked resolve rate directly against token reduction. On unbounded windows with K=3 observations kept raw, resolve rate dipped from 14.5% to 12.1% on Qwen3-4B. But on long-horizon runs capped at 32K tokens, preserving latent observations resolved 21.1% compared to 11.1% for hard truncation.
That structured state block (decisions, failed attempts, open tasks) works well for keeping high-level intent intact. Where I see it hit friction in a coding harness is patch application. If an extraction pass summarizes an earlier traceback instead of keeping the exact line numbers and symbol names, the agent has to re-run the inspection tools just to get the raw strings back into scope.
That matches what I hit building Deiko. The state block is great for intent and bad for exact strings, so I keep two layers: a short note (decided, tried, open) goes into the next chat, and each line links back to the original brief, with the exact screen text, error output and file paths, which the agent fetches on demand through an MCP tool. In my tests, agents did pull the full originals when a question needed exact details, so the note could stay small. The 21.1% vs 11.1% at 32K is a big gap. Did the paper split out how much came from exact observations versus recency?
Keeping the small "decided / tried / open" state and fetching the exact details only when needed makes a lot of sense. A summary is great for intent, but it's a pretty bad place to store exact error text or file-level facts.
Exact error messages and stack traces degrade quickly once paraphrased. A model summarizing an error tends to strip out column offsets, exception classes, and exit codes, turning a concrete compiler failure into a generic complaint. Leaving the raw execution artifacts in a local SQLite table or structured append log while keeping only the high-level intent in the prompt window preserves both token budget and diagnostic accuracy.
I ran into this the opposite way. For a while I was compressing everything past a certain turn count because the bill was hurting, and I noticed that accuracy on any step requiring an exact identifier tanked while more open-ended reasoning seemed fine. It took a while to isolate because the failures looked like the model getting confused, not like it was missing a specific string it had seen earlier. What I actually needed was to keep the raw text for anything the model would need to quote back verbatim, and let everything else compress.
That exact identifier loss is where the failure chain usually starts. The model keeps the high-level intent, but file paths, UUIDs, or compiler flags turn into plausible hallucinations. I had the same issue with patch application, where a summarized diff dropped the exact leading whitespace and git rejected the hunk. Keeping raw stdout for tool calls and only compressing the reasoning turns saves the budget without breaking the downstream edits.
We ran into this with a document retrieval pipeline last year: I'd summarized older turns to cut costs and the agent started generating file paths that were close but not quite right, then wasting two or three extra tool calls to re-confirm what it had already seen. Didn't connect it to compression until I looked at which turns I'd trimmed. It's the agent's own reasoning steps that need to stay verbatim, not the grep outputs or test logs. Expanding the exact window back a bit fixed it almost immediately.
That path hallucination loop happens because lossy summaries replace exact identifier strings with generalized descriptions. The model remembers it inspected a router file, but loses the exact directory depth or filename extension, triggering redundant search tool calls to recover state it already discovered. Retaining structured reasoning traces while trimming raw tool output preserves the exact nouns without context bloat.
The 21.1% vs 11.1% result under the 32K limit is the part that really caught me.
It suggests context compression isn’t just about saving tokens so it’s about deciding what must remain exact. A small state summary plus on demand access to the original logs might be the better architecture.
The failure mode with pure summarization is that exact line numbers, compiler flags, and git hashes vanish first. When an agent needs to apply a patch three turns later, lossy prose in the context window forces it to guess the indentation or re-run the whole test suite just to recover stdout. Dumping raw command output to a local scratch directory and leaving only a two-line index in the active window gives you the token savings of a tight budget without turning file paths into approximations.
The recency scaling row is the underrated one: moving K from 3 to 8 takes Qwen3 from 12.1% back to 14.4% — essentially full recovery against 14.5% uncompressed — which means nearly all the recovered accuracy comes from simply widening the raw window, not from the latent compression itself. The honest headline of this paper is the 32K row: 21.1% vs 11.1% against truncation. LOHA doesn't preserve accuracy through compression; it preserves viability when the alternative is a truncated prompt.
The K scaling delta makes that mechanic very clear. Dropping K below three starves the model of the immediate tool outputs it needs to anchor file edits, which is why widening K to eight does the heavy lifting. The 32K row works as an insurance policy against hard context cliffs. An agent pays an overhead penalty across every intermediate turn to keep the task trajectory intact when truncation would otherwise drop the initial prompt.