Coding agents resend old tool output on every turn, so replacing it with something smaller looks like an obvious saving. I built that as a Torana plugin, broke it once, rebuilt it, and measured it. On DeepSeek V4 Pro the saving was about 2% per session, smaller than ordinary run-to-run variation, because prompt caching had already made the repeated output nearly free.
This write-up covers the idea, what broke, the final design, the experiment setup and methodology, the results, and what's next.
How coding agents resend tool output
The model behind a coding agent can't read files or run commands itself. It asks for a tool call, the agent runs it locally, and the output is appended to the conversation. On every following turn, the agent sends the whole visible conversation back to the model, including every earlier tool result.
Each request re-sends the log as input. Context fills up and the output is billed again on every turn.
This applies to harnesses that resend visible history, typically over Chat Completions or Messages-style APIs. When a provider owns the history (for example OpenAI Responses with previous_response_id), Torana can't rewrite earlier items; native provider compaction is a separate mechanism.
The idea: a cheaper assistant for the hosted model
If the hosted model is the brain doing the real work, a cheaper model could summarise tool output before the brain reads it. A summary is only useful if it knows what matters, so Torana's intent plugin changes the tool definitions the model sees to request a short reason with every tool call, then strips that field before the harness sees the call. The compactor uses the captured intent to decide what to keep.
Reason field: added to tool schemas going out, removed from tool calls coming back. The agent never sees it.
Attempt 1: summarise every large output
The first compactor summarised any tool output over 2,000 characters. It broke coding agents immediately (issue #166). Coding agents edit files with exact search and replace; when a file read came back as a natural-language summary, edits stopped matching. The agent re-read the file, failed again, and looped, burning tokens instead of saving them.
Attempt 2: exact by default, compact only older output
The second design leaves everything exact by default. Policies are ordered, first-match and case-insensitive; unmatched tools are never changed:
- exact: never altered. Mutation tools, diffs, failed commands, errors, stack traces and non-zero exits are forced exact even if a broader rule matches them.
-
deterministic: keeps bounded head and tail evidence plus size, SHA-256, omitted-byte count and a rerun instruction. With
first_pass, it applies from the first time the model sees the output, so the prompt never needs a later rewrite. -
keyword (
keyword_compactor): keeps intent-matching lines, only after at least one exact exposure. -
model (
compactor): an intent-guided summary, only after at least one exact exposure, and only when the economic gate predicts a net saving. - source (file reads): accepted in configuration but behaves as exact (see what broke).
Keyword/model: exact on first exposure, eligible for compaction from the next request if the gate allows. Deterministic with first_pass: compact from the first exposure, avoiding a later prefix rewrite. Cached compact versions keep the same bytes on later requests.
Rewriting an older message changes the prompt prefix from that point on. The model-gated compactor therefore runs two economic gates: a preflight that checks whether even a best-case reduction could repay the cache rewrite (before paying for a summary), and a final check on the actual candidates. It applies a batch only when the estimated net saving is positive.
Experiment setup and methodology
Environment
- Date: 21 July 2026.
- Harness: OMP 16.2.9 (coding agent).
-
Target model: DeepSeek V4 Pro via
https://api.deepseek.com/beta(OpenAI-compatible format). - Summariser (model-gated arm): DeepSeek V4 Flash on the same endpoint.
- Torana: commit e09d225 from PR #179. That PR was the experiment branch; the plugins were later moved to torana-plugins, and current releases differ. This is the implementation used for the historical run; the unpublished task prompts and run artifacts below prevent exact reproduction from this article alone.
Prices used (July 2026, per million tokens)
Historical test inputs, not current prices. Cache-hit price is 0.83% of the cache-miss price for V4 Pro.
| Model | Input (cache miss) | Cache hit | Output |
|---|---|---|---|
| DeepSeek V4 Pro | $0.435 | $0.003625 | $0.87 |
| DeepSeek V4 Flash | $0.14 | $0.0028 | $0.28 |
Arms
- Control: Torana in the request path with no compaction plugin.
-
Deterministic:
schema_translator → intent → keyword_compactor, summariser disabled. This is the committedconfig.dogfood.json. The run itself used thehttps://api.deepseek.com/betaendpoint for the target model.
Deterministic arm configuration (JSON)
{
"providers": {
"deepseek": { "url": "https://api.deepseek.com", "format": "openai" },
"deepseek-anthropic": { "url": "https://api.deepseek.com/anthropic", "format": "anthropic" }
},
"plugins": {
"order": ["schema_translator", "intent", "keyword_compactor"],
"config": {
"keyword_compactor": {
"tool_policies": [
{ "match": "read*", "mode": "exact" },
{ "match": "view*", "mode": "exact" },
{ "match": "web_search", "mode": "deterministic", "first_pass": true,
"rerun": "Repeat the search to recover every result." },
{ "match": "grep*", "mode": "keyword" }
]
}
}
},
"offload": { "enabled": false, "provider": "deepseek", "model": "deepseek-v4-flash" },
"limits": { "concurrency": 10, "rpm": 100 }
}
-
Model-gated:
schema_translator → intent → compactor, with explicit prices,expected_applications: 6and V4 Flash as the summariser. Grep-like output was eligible for model summaries; the same exact and first-pass rules applied otherwise. Reconstructed from the run record in the shape of the committed example (the committed file covers only the deterministic arm).
DeepSeek has no separate cache-write charge, so the cache-write rate is set equal to the input rate: a rewritten span is billed as cache misses.Model-gated arm configuration (JSON)
{
"providers": {
"deepseek": {
"url": "https://api.deepseek.com/beta",
"format": "openai",
"pricing": {
"deepseek-v4-pro": { "input_usd_per_mtok": 0.435, "output_usd_per_mtok": 0.87,
"cache_read_usd_per_mtok": 0.003625, "cache_write_usd_per_mtok": 0.435 },
"deepseek-v4-flash": { "input_usd_per_mtok": 0.14, "output_usd_per_mtok": 0.28,
"cache_read_usd_per_mtok": 0.0028, "cache_write_usd_per_mtok": 0.14 }
}
}
},
"plugins": {
"order": ["schema_translator", "intent", "compactor"],
"config": {
"compactor": {
"expected_applications": 6,
"tool_policies": [
{ "match": "read*", "mode": "exact" },
{ "match": "view*", "mode": "exact" },
{ "match": "web_search", "mode": "deterministic", "first_pass": true,
"rerun": "Repeat the search to recover every result." },
{ "match": "grep*", "mode": "model" }
]
}
}
},
"offload": { "enabled": true, "provider": "deepseek", "model": "deepseek-v4-flash" }
}
Protocol
- Tasks: five repository tasks run with OMP: architecture, economics, regression review, safety (exercising mutations, patches, errors and non-zero exits), and a stress task.
- Design: five repetitions × three arms = 75 sessions. Within each paired block (task × repetition), the order of the three arms was randomised.
Guards, cost accounting, statistics and what isn't published
prompt_cache_hit_tokens.
Results
Medians per session over 25 runs per arm.
| Arm | Completed | Loop stops | Median cost | Median input tokens | Median requests | Cache-hit ratio |
|---|---|---|---|---|---|---|
| Control (no compaction) | 23/25 | 2 | $0.024820 | 450,810 | 16 | 92.23% |
| Deterministic | 24/25 | 1 | $0.026699 | 479,087 | 19 | 91.68% |
| Model-gated | 24/25 | 1 | $0.022465 | 444,625 | 17 | 92.75% |
Deterministic arm: 15 of 25 treatment runs transformed something; 399 applications (34 new transformations, 365 cache reuses) removed 4,809,164 bytes in total. Median paired saving $0.000524 (~2.1%), 95% interval −$0.003714 to +$0.003725. Mean saving −$0.0000559; cheaper in 13 of 25 pairs. Median paired input-token saving 29,117 (interval −57,989 to +60,161). Transformed-only runs had a median change of −$0.000262, a small increase.
Model-gated arm: only four runs transformed output, all through the deterministic rule (6 transformations, 79 reuses). Median paired saving $0.000478 (~1.9%), interval −$0.000998 to +$0.004973; mean $0.001569, cheaper in 15 of 25 pairs. The preflight made zero Flash calls: with six expected applications at DeepSeek's cache-hit price, no rewritten prefix could repay its cache misses. This measures the gate, not model-written summaries.
Loop stops (request-cap terminations) happened in the control arm and in treatment runs with no transformations, so they reflect agent variance rather than compaction.
Median cost per task. Differences in tool calls and answer length outweighed the value of the removed cached tokens.Median cost by task
Task
Control
Deterministic
Model-gated
Architecture
$0.026040
$0.022984
$0.025993
Economics
$0.024820
$0.026717
$0.021742
Regression review
$0.033672
$0.034278
$0.032411
Safety
$0.017640
$0.015142
$0.017297
Stress
$0.022256
$0.027260
$0.021505
Why it didn't pay: prompt-cache economics
About 92% of input tokens were cache hits in every arm, and a cache hit on V4 Pro cost 0.83% of a miss: about 120× cheaper. The repeated tool output being removed was already nearly free. A rewrite changes the prompt from that point on, so everything after it is billed as new input until the cache is rebuilt.
Removing bytes only pays if the saving on later turns exceeds the misses it causes. With 92% cache hits and a 120× price gap, the removed bytes were worth very little.
That is a property of DeepSeek's pricing. Providers that charge more for cached input, or have no prompt cache at all, could change the result. That is the next experiment (issue #193).
The break-even condition
Let N be future turns, R the rewritten cached span, W the provider's write/input rate, C its cache-read rate, D tokens removed per future turn, and S the summariser cost. Compaction needs:
N > (R × max(0, W − C) + S) / (D × C)
This simplified model assumes positive D × C and matching units: rates per token when S is in dollars. It doesn't model every provider's caching rules or changes in agent behaviour. When economic inputs are missing or the expected net is not positive, the gate does nothing, which is a correct outcome.
What broke along the way
- File-read markers caused re-read loops. An earlier, excluded pilot replaced an old file read with a "reread" marker. The agent fetched different line ranges of the same file until it hit the request cap; deduplicating identical arguments didn't help because the offsets kept changing. Source reads are now always exact (#178 tracks a bounded alternative).
- The gate missed a cost. An uncached summary candidate was mislabelled as a cache reuse, so the preflight left out the prefix-rewrite premium. Torana paid for Flash summaries that the final gate rejected. Uncached candidates are now treated as rewrites; the corrected run made zero uneconomic summariser calls.
-
Cache hits were invisible. DeepSeek reports cache usage as
prompt_cache_hit_tokens, not the OpenAI-styleprompt_tokens_details.cached_tokens. Torana now reads both in streaming, JSON and model-service responses, and/statsreports summariser token usage.
Limitations
- One provider, one harness, five tasks, five repetitions per task, and July 2026 prices.
- Agent trajectories varied between runs, and many treatment sessions were no-ops.
- Bytes are not tokenizer counts; token savings are provider-reported.
- The quality heuristic understated semantic quality and was only used for triage.
- The excluded pilot is retained only as negative safety evidence.
- Compaction can still matter where caching is expensive or absent, or when an agent runs out of context.
What's next
Providers with expensive cache reads. A fixed transcript replay to isolate pricing from agent behaviour, then live paired runs on providers whose cache pricing doesn't make history nearly free (#193).
Decide per tool call, not by size. Some output must stay exact: file reads, edits, errors. Some never needed to stay in the conversation: npm install logs, build output, script runs. Small local decision models such as Jev and Laya can classify a tool call into a fixed set of choices. I'm building a version where one decides, per tool call, whether its output stays exact or can be compacted. That is in progress, not a measured result.
Try it or build on it
compactor, keyword_compactor and intent are Torana plugins. Install them by name, review their permissions and bind a summariser following the compactor guide. Measure with paired runs on your own tool mix and prices, using provider-reported tokens and including summariser cost.
The same plugin SDK powers the local PII check. To build your own, start with the plugin-authoring guide. Results from other providers, especially on #193, are welcome.
How this experiment became Torana: the pivot story →




Top comments (2)
Nice that you published the 2% instead of quietly dropping the idea. Prompt caching making the repeated output nearly free is easy to forget when you're staring at the raw token count. Did compaction do anything for context length or answer quality on long sessions, even if the bill barely moved?
The % number highlights how prompt caching broke the old heuristic that token count equals billable cost. When cached input gets an 8% or 9% discount, an uncompressed append is essentially free. The moment middleware compacts an earlier turn in the trace, it invalidates the prefix cache and forces the provider to reprocess everything downstream at full input pricing. Unless a command dumped hundreds of kilobytes of throwaway build logs that the model never needs to reread, letting the raw output sit in cache is cheaper and keeps exact diff anchors intact.