Coding agents hit a cost wall when tool-call output bloats context windows. A Show HN project tackles this with a fine-tuned compression model that sits between agent output and model input, trimming tokens by 29.6% without breaking KV cache or multi-turn reasoning. The project exposes a pattern: developers are building custom middleware layers to manage the economic pressure points in agentic workflows.
The Problem: $700/Day API Bills
The team behind this project maxed out their Codex subscription and burned $700 per day per person on API calls. The culprit was not the agent's reasoning steps but the tool-call results that get fed back into the model. File retrieval, test output, and error traces accumulate fast. Each round trip inflates the input token count, and cache misses compound the cost.
Coding agents differ from conversational agents in context shape. A chat agent might reference a few messages. A coding agent drags in file trees, diff output, stack traces, and test results. The context window fills with structured data that the model needs for the next step but that also contains redundancy.
Architecture: Proxy Layer with Fine-Tuned Compression
The solution is a local proxy that wraps Codex and intercepts tool-call results before they return to the model. The proxy runs a fine-tuned Qwen model trained to preserve agent trajectory while removing redundant information from tool outputs.
Execution flow:
- Agent calls a tool (file read, test run, search).
- Tool returns structured output (file content, test logs, search results).
- Proxy intercepts the output and passes it to the compression model.
- Compression model trims redundant tokens while preserving semantic fidelity.
- Compressed output goes back to the agent's context window.
- KV cache remains untouched because the proxy operates before the model sees the input.
The fine-tuning objective is trajectory preservation. The model learns which parts of tool output the agent needs for subsequent reasoning steps and which parts are noise. File retrieval accuracy and context-heavy tasks see the biggest gains.
Implementation Details
The CLI is open source and installs as a shell wrapper around Codex. It runs on by default, and you can disable it with codex --uncompress when you need full output.
Key design choices:
- Local proxy: No data leaves your machine. The proxy wraps your local Codex instance and does not retain queries.
- Fine-tuned Qwen model: The compression model is trained on coding agent trajectories, not general text. This preserves the structure that agents need for multi-turn reasoning.
- KV cache preservation: The proxy compresses tool output before it enters the model's context window, so the cache does not invalidate.
-
Token counting: The proxy uses OpenAI's
response.usageto measure savings. You can runsavingsto see cumulative token reduction.
Installation:
curl -fsSL https://install.everestagi.com/install.sh | sh && \
source ~/.config/everest/shell.sh
The proxy runs as a local service and intercepts API calls. You point your Codex client at the proxy endpoint instead of the OpenAI endpoint.
Trade-Offs and Failure Modes
| Dimension | Benefit | Risk |
|---|---|---|
| Cost | 29.6% token reduction, lower API spend | Compression model adds latency and local compute overhead |
| Fidelity | Fine-tuned on agent trajectories to preserve reasoning steps | May remove information the agent needs for edge cases or complex tasks |
| Cache | Operates before model input, so KV cache stays valid | If compression changes output shape, downstream tools may break |
| Privacy | Local proxy, no data retention | Requires trust in the proxy binary and shell script installer |
| Portability | Works with any Codex-compatible client | Tied to Codex and Astra workflows, not general-purpose |
Failure modes:
- Over-compression: The model removes a file path or error message that the agent needs for the next step. The agent halts or makes incorrect assumptions.
- Cache invalidation: If the compression model changes output structure in a way that breaks tool-call contracts, the agent's cache becomes stale.
- Latency: Running a fine-tuned model locally adds milliseconds to each tool call. For high-frequency agents, this accumulates.
- Model drift: If Codex changes its tool-call format, the compression model may need retraining.
Observability and Debugging
The proxy exposes a savings command that shows cumulative token reduction. This is useful for tracking ROI, but it does not expose per-call compression ratios or failure cases.
What you cannot see:
- Which tool calls benefit most from compression.
- When the compression model removes information that causes downstream errors.
- Latency breakdown (tool call, compression, model inference).
For production use, you would want structured logs that capture input tokens, output tokens, compression ratio, and agent success rate per task type. You would also want a fallback mode that disables compression if the agent fails repeatedly on a specific task.
When to Use This
Good fit:
- You run coding agents daily and hit API cost ceilings.
- Your tasks are context-heavy (file retrieval, test output, large diffs).
- You can tolerate local compute overhead and occasional compression errors.
- You trust the proxy binary and are comfortable with shell script installers.
Poor fit:
- You run conversational agents with small context windows.
- Your tasks require exact tool-call output (legal, compliance, security).
- You need sub-100ms latency and cannot afford local model inference.
- You operate in environments where local proxies violate security policy.
Technical Verdict
This project demonstrates a practical response to the economic pressure points in agentic workflows. The architecture is sound: a local proxy with a fine-tuned compression model preserves KV cache and reduces input tokens without breaking multi-turn reasoning. The 29.6% token reduction is meaningful for teams burning hundreds of dollars per day on API calls.
The trade-off is fidelity. Fine-tuning helps, but compression always risks removing information the agent needs. For production use, you would want observability that tracks compression ratio per task type and a fallback mode that disables compression when the agent fails.
Use this if you are optimizing for cost and can tolerate occasional compression errors. Avoid it if you need exact tool-call output or operate in environments where local proxies are not allowed. The pattern is worth watching: as agents move from prototypes to daily workflows, cost-optimization middleware will become a standard layer in the stack.
Top comments (0)