DEV Community

Cover image for Context as a File: Baize's CLM-Inspired Long-Conversation Projection
rebornace
rebornace

Posted on

Context as a File: Baize's CLM-Inspired Long-Conversation Projection


Project: Baize — an AI assistant runtime for your team (Go 1.25+, MIT)
Repo:
https://github.com/rebornace/baize


1. The short version

Baize v0.3.x takes inspiration from the recent Context Language Models (CLM) papers and ships a constrained take: keep important facts in long chats. It treats the context that reaches the model as something that can be curated — pin the essentials, roll up rolling summaries, drop oversized tool results — but it never rewrites the saved chat messages, and on failure it automatically falls back to the existing compaction logic.

One line for the tradeoff: CLM hands editing rights over the context to the model; Baize puts the editing in a system-side projection layer, trading some freedom for reliability.

2. Why long conversations are a hard problem for agents

An agent's context is the classic append-only structure: every new message, tool result and reasoning trace gets appended to the history. Once a conversation gets long, the problems show up:

  • Attention dilution: with thousands of lines of history, one critical configuration item can get lost; the model may not reliably pick it up anymore.

  • Truncation is irreversible: cutting old messages silently drops early agreements and context the user never notices.

  • Compaction is irreversible: rolling summaries save space, but a summary is a lossy, one-way transformation; key details can simply disappear.

The traditional answer is to hand context management to an external framework: compress on a threshold, offload, retrieve. Those approaches work, but they are "the framework defines the actions, the model can only pick one".

3. What CLM brings

CLM is a paper from late September 2026 (arXiv:2609.37725, by the University of Washington, Meta Superintelligence Labs, MIT and others). It turns the problem around: context is not an append-only log — it is a file the model has write permission on. The model can append, but it can also edit, delete and rewrite that file, and every change is synced into the next turn's context.

The observations in the paper are direct:

  • Models can learn on their own what is most worth keeping in context, and behaviors show up that previous frameworks never had — maintaining a tracker table for multi-agent collaboration, or defining reusable context-management functions.

  • On the deep-research benchmark BrowseComp-Plus, out-of-the-box CLM beats the strongest baseline by 11.4% in accuracy while using 21.5% fewer inference FLOPs; on a 24-hour, multi-repo agent-swarm task it achieves 65% higher end-to-end speedup at the same compute.

  • Because context management becomes model behavior rather than an external harness policy, the strategy itself can keep improving through reinforcement learning.

The appeal is that context management stops being a hard-coded rule set and becomes a product the model can understand and learn. But handing the model full editing rights has real production gates: is the editing reliable, does the archive stay trustworthy, and what happens when it fails?

4. Baize's constrained take: keep important facts in long chats

Baize did not copy CLM wholesale. It extracted the core idea — context is curatable output, not an append-only log — and added three constraints:

1. The projection only lives on the model input side; saved chats are never rewritten.

When a thread gets long, the assistant keeps a model-facing projection of the facts it still needs: it pins what actually matters, summarizes the secondary stuff, and drops oversized tool results. The chat the user sees stays untouched, ready to review, fork or roll back at any time. This matters for team use: the archive is the basis of audit and trust. The model input can be curated aggressively; the archive must stay true.

2. Fail-open.

If the projection fails, the assistant falls back to the original compaction and the conversation keeps going — context curation must never become an availability bottleneck. This matches the philosophy Baize's decision layer already follows: optimization must not affect whether the tools themselves still work.

3. No promise of lower cloud token usage.

The projection itself costs tokens; the curated result is not necessarily cheaper than plain compaction. The docs say it plainly. What this feature solves is key-fact retention and context usability — not the bill.

The knob is optional (default off), hot-reloadable, and costs nothing when unused.

Why "constrained"? Baize talks to OpenAI-compatible APIs: it cannot read logits, and there is no native model-side context-editing capability. Rather than letting the model rewrite history in an uncontrolled place, the editing lives in the system-side projection layer, where it can be made reliable and reversible. CLM is the radical route of native model editing; Baize's projection is the pragmatic route of framework-style editing. Same goal, different cost and reliability.

5. A little design philosophy

Looking back at this upgrade, a few decisions are consistent:

  • Optional, rather than deciding for the user by default: the projection and the decision layer are both runtime knobs, off by default, taking effect via hot reload when enabled.

  • Fail-open first: projection falls back to compaction, routing falls back to full tools — availability is never harmed by an optimization.

If you are also working on long-conversation handling for agents, come discuss your tradeoffs at the repo: https://github.com/rebornace/baize. Everything mentioned here (the projection and workspaces) already exists in the public repository; docs live under docs/developers in both English and Chinese. I'll keep following up on issues and suggestions.


Reference: Context Language Models (arXiv:2609.37725): https://arxiv.org/abs/2609.37725

Top comments (0)