DEV Community

Cover image for Headroom: How Context Compression Cuts Agent Token Costs by 60–95% Without Changing Answers
mech.app
mech.app

Posted on Originally published at mech.app

Headroom: How Context Compression Cuts Agent Token Costs by 60–95% Without Changing Answers

Production agents hit context limits fast. A coding agent that runs tests, reads logs, and pulls documentation can burn through 100k tokens in three turns. RAG pipelines dump entire chunks into the prompt. Tool outputs return verbose JSON. Every token costs money and adds latency.

Headroom is a compression layer that sits between your agent and the LLM. It shrinks tool outputs, logs, RAG chunks, and files before they reach the context window. The project claims 20% token reduction for coding agents and 60–95% for JSON-heavy workflows, with no change to the final answer. It ships as a Python library, a FastAPI proxy, and an MCP server, giving you three deployment patterns with different trade-offs in latency, observability, and integration complexity.

What Gets Compressed and What Stays Intact

Headroom uses a fine-tuned model (kompress-v2-base) to decide what to keep and what to discard. The compression is not blind truncation. The model learns to preserve semantic anchors: error messages, function signatures, stack traces, and critical log lines.

The hero image in the repo shows a 55,957-token agent prompt compressed to 24,340 tokens. The FATAL line at item 67 survives byte-for-byte. This is the key engineering question: how do you compress structured data without breaking parsers downstream?

Headroom handles:

  • JSON tool outputs: Strips redundant keys, collapses nested structures, preserves schema-critical fields.
  • Log files: Keeps ERROR and FATAL lines, drops repetitive INFO entries, maintains timestamps for correlation.
  • Code diffs: Preserves function signatures and changed lines, compresses unchanged context.
  • RAG chunks: Keeps sentences with high semantic density, drops boilerplate.

The model does not compress the user query or the assistant's response. It only touches the context that the agent reads before generating an answer.

Three Deployment Modes

Headroom ships in three forms. Each has different latency, observability, and integration costs.

Deployment Mode Latency Overhead Observability Integration Effort Best For
Library 50–200ms per call Full control, log everything Low (import and wrap) Agents you control end-to-end
Proxy 100–300ms per call Centralized metrics, request logs Medium (point agent at proxy URL) Multi-agent fleets, shared infra
MCP Server 150–400ms per call MCP protocol logs, tool-level tracing Low (add to MCP config) Claude Code, Cursor, agent harnesses

Library Mode

You import Headroom as a Python package and wrap your agent's context-building logic.

from headroom import compress_context

def build_agent_context(tool_outputs, logs, rag_chunks):
    raw_context = "\n".join([
        format_tool_outputs(tool_outputs),
        format_logs(logs),
        format_rag_chunks(rag_chunks)
    ])

    compressed = compress_context(
        raw_context,
        preserve_patterns=["ERROR", "FATAL", "def ", "class "],
        target_ratio=0.4
    )

    return compressed
Enter fullscreen mode Exit fullscreen mode

This gives you full control. You can log the before and after token counts, inspect what got dropped, and adjust preserve_patterns based on your domain. The latency overhead is 50–200ms per compression call, depending on context size.

Proxy Mode

Headroom runs as a FastAPI service. Your agent sends requests to the proxy instead of directly to OpenAI or Anthropic. The proxy compresses the context, forwards the request, and returns the response.

# Start the proxy
headroom serve --port 8080 --model kompress-v2-base

# Point your agent at the proxy
export OPENAI_API_BASE=http://localhost:8080/v1
Enter fullscreen mode Exit fullscreen mode

This centralizes compression logic. You get request logs, token savings metrics, and error rates in one place. The latency overhead is 100–300ms, which includes network round-trip and compression time. This mode works well for multi-agent fleets where you want shared observability.

MCP Server Mode

Headroom implements the Model Context Protocol. You add it to your MCP config, and it compresses tool outputs before they reach the agent.

{
  "mcpServers": {
    "headroom": {
      "command": "headroom",
      "args": ["mcp"],
      "env": {
        "HEADROOM_MODEL": "kompress-v2-base",
        "HEADROOM_TARGET_RATIO": "0.4"
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

This is the lowest-friction option for Claude Code, Cursor, and other MCP-compatible harnesses. The agent sees Headroom as just another tool. The latency overhead is 150–400ms because MCP adds protocol serialization on top of compression time.

How Compression Handles Structured Formats

The kompress-v2-base model is fine-tuned on JSON, logs, code, and Markdown. It does not use generic text compression (gzip, LZ4). It learns to preserve structure.

For JSON, the model keeps:

  • Schema-defining keys (type, id, status)
  • Non-null values
  • First and last items in large arrays
  • Error and warning fields

For logs, it keeps:

  • Lines with ERROR, FATAL, WARN
  • Timestamps for correlation
  • Stack traces
  • First occurrence of repeated patterns

For code, it keeps:

  • Function and class signatures
  • Changed lines in diffs
  • Import statements
  • Docstrings for public APIs

The model drops:

  • Repeated INFO log lines
  • Null or default JSON values
  • Unchanged code context in diffs
  • Boilerplate comments

This is not lossless compression. You lose detail. The bet is that the detail you lose does not change the agent's answer.

Observability and Failure Modes

Headroom adds a new failure surface. If the compression model drops a critical line, the agent might give the wrong answer. If the proxy goes down, your agent stops working.

Observability hooks you need:

  • Token delta logs: Before and after token counts for every compression call.
  • Preserved pattern matches: Log which patterns (ERROR, FATAL, def) triggered preservation.
  • Compression ratio distribution: Track min, max, and p95 compression ratios across requests.
  • Downstream accuracy: Sample agent outputs and compare compressed vs. uncompressed contexts.

Failure modes:

  • Over-compression: The model drops a critical error message. The agent misses the root cause and suggests the wrong fix.
  • Under-compression: The model preserves too much. You save 10% tokens instead of 60%.
  • Proxy downtime: If you run Headroom as a proxy, it becomes a single point of failure. Add health checks and failover.
  • Latency spikes: Compression time scales with context size. A 200k-token context might take 1–2 seconds to compress.

When Compression Breaks Semantic Integrity

Headroom works well for verbose, repetitive data. It struggles with dense, information-rich contexts where every token matters.

Good candidates for compression:

  • JSON API responses with many null fields
  • Log files with repeated INFO lines
  • RAG chunks with boilerplate introductions
  • Code diffs with large unchanged sections

Bad candidates:

  • Mathematical proofs where every step is load-bearing
  • Legal documents where exact wording matters
  • Short, dense error messages (already minimal)
  • Contexts where the agent needs to count occurrences

If your agent needs to count how many times a specific error appears, compression will break that. If your agent needs to verify exact JSON schema compliance, compression might drop the field that violates the schema.

Integration with Agent Frameworks

Headroom supports LangChain, LlamaIndex, and custom agent loops. For LangChain, you wrap the retriever:

from langchain.retrievers import BaseRetriever
from headroom import compress_context

class CompressedRetriever(BaseRetriever):
    def __init__(self, base_retriever):
        self.base_retriever = base_retriever

    def get_relevant_documents(self, query):
        docs = self.base_retriever.get_relevant_documents(query)
        compressed_docs = [
            compress_context(doc.page_content, target_ratio=0.5)
            for doc in docs
        ]
        return compressed_docs
Enter fullscreen mode Exit fullscreen mode

For LlamaIndex, you compress the retrieved nodes before they go into the prompt.

For custom agent loops, you compress the accumulated context at each turn:

context_history = []

for turn in agent_loop:
    tool_output = run_tool(turn.tool_call)
    context_history.append(tool_output)

    # Compress the full history before the next LLM call
    compressed_history = compress_context(
        "\n".join(context_history),
        target_ratio=0.4
    )

    response = llm.generate(
        prompt=turn.user_message,
        context=compressed_history
    )
Enter fullscreen mode Exit fullscreen mode

Cost and Latency Trade-Offs

Compression adds latency but saves token costs. The break-even depends on your LLM pricing and request volume.

Example calculation:

  • Agent sends 50k tokens per request
  • Headroom compresses to 20k tokens (60% reduction)
  • Compression adds 200ms latency
  • GPT-4 input tokens cost $0.03 per 1k tokens

Savings per request:

  • Token cost saved: (50k - 20k) * $0.03 / 1k = $0.90
  • Latency cost: 200ms added to request time

If your agent runs 1,000 requests per day, you save $900 per day. If your SLA requires sub-500ms response times, the 200ms overhead might break your budget.

Technical Verdict

Use Headroom when:

  • Your agent workflows routinely exceed 32k tokens per request.
  • You have verbose tool outputs (JSON APIs, log files, RAG chunks).
  • You can tolerate 100–300ms added latency.
  • You have observability to detect when compression drops critical data.

Avoid Headroom when:

  • Your contexts are already dense and minimal.
  • Your agent needs exact token counts or field-level accuracy.
  • You cannot afford the latency overhead.
  • Your agent operates in a domain where every token is load-bearing (legal, math, compliance).

Headroom is infrastructure, not magic. It trades latency for cost savings. If your agent is already optimized and you are not hitting context limits, you do not need it. If you are burning $10k per month on tokens because your RAG pipeline dumps 100k tokens into every prompt, Headroom will pay for itself in a week.

Source Links

Top comments (0)