Originally published on AI Tech Connect.
What the measurement actually showed Token reduction and cost reduction came apart. Across 2,908 paired executions, an aggressive compression arm removed 38.4 percent of estimated raw tool-output tokens and raised billed cost by 6.8 percent. You are optimising a rounding error. Prompt-cache traffic accounted for roughly 80 percent of actual cost. The compressed tool outputs were about 3.3 percent of total cost. The mechanism is extra turns. Aggressive compression triggered additional model turns that re-transmitted the entire cached prefix, so the local saving was offset and then some. Destroying verbatim anchors breaks the task. On 40 SWE-bench-derived Go rows, patch application fell from 27/40 raw to 15/40 compressed, because compression corrupted the exact strings the agent had to…
Top comments (0)