Two papers landed in the same 48-hour window this week, and together they tell a story that matters more than either one alone.
📖 Read the full version with charts and embedded sources on AgentConn →
On September 29, researchers from the University of Washington, MIT, Meta's Superintelligence Labs, and Trillium published Context Language Models — a paper arguing that the agent "harness" (the fixed summarize/truncate/RAG policies that manage context for long-running agents) is the binding constraint on performance, not the model itself. Their solution: train the model to edit its own context file directly. The results are striking — +11.4% accuracy on BrowseComp+ with 21.5% fewer FLOPs, and 5% higher scores on 12-hour EdgeBench runs while burning 59% less compute.
One day later, turbopuffer engineer Dan Harrison published RIP, vector database — announcing that turbopuffer, the company that built a $100M ARR business on serverless vector search, is reorganizing its entire storage architecture so that ANN (approximate nearest neighbor) search becomes "just another secondary index" rather than the foundation organizing all data.
These are independent developments from unrelated teams. But they converge on the same structural claim: the standalone context and retrieval middleware layer — the thing agent builders have been treating as infrastructure — is being squeezed from both sides. Models are learning to manage their own context from above. Storage engines are absorbing vector search as a feature from below. The middle layer, where an entire ecosystem of vector databases, RAG frameworks, and context management tools lives, is watching its TAM compress in real time.
The CLM numbers that matter: +11.4% accuracy on BrowseComp+ with 21.5% fewer FLOPs. +5% on 12-hour EdgeBench with 59% fewer FLOPs. Online RL improved Qwen3.5-9B by 47.6% while reducing compute by 12%. Suffix Cache Reuse cut server-side compute by 35% versus standard SGLang.
What CLM Actually Does
Context Language Models treat the model's context window as a mutable file. Instead of the harness deciding what to keep, compress, or discard — using fixed policies like "summarize everything older than N turns" or "retrieve the top-K chunks" — the model itself makes those decisions using ordinary bash commands.
The model can delete a stale search result. It can rewrite its own plan mid-run. It can compress verbose tool outputs into a few lines. It can keep a running scoreboard and update it in place. The context file is synchronized with the language model's actual context, so edits take effect immediately.
This is not an incremental improvement on RAG. It is an architectural argument that the entire harness-managed context paradigm — where an external system decides what the model sees — is the bottleneck for 12-hour-plus agent runs. The paper demonstrates this on BrowseComp+, which tests multi-step web research over extended sessions, and on EdgeBench, which evaluates agents running autonomously for half a day.
The research team built CLMs zero-shot with existing models, meaning you do not need a special model. You need a special relationship between the model and its context. The further result — that online reinforcement learning improved a relatively small Qwen3.5-9B model by 47.6% on these tasks — suggests that the technique scales down, not just up.
What turbopuffer's Pivot Tells You
turbopuffer's blog post is a different kind of evidence for the same thesis, coming from the storage side.
The company built a serverless vector database that, according to Sacra, hit $100M in annualized revenue by March 2026 with 2,400% year-over-year growth. Its customer list reads like a who's who of AI-native companies: Cursor, Anthropic, Notion, Linear, Atlassian, Ramp, Grammarly, Superhuman. If anyone has a reason to defend the "vector database as a standalone category" thesis, it is turbopuffer.
Instead, they published a post explaining why they are abandoning their vector-primary architecture. The three reasons they cite are structural, not cosmetic:
Storage amplification. Multi-vector documents require duplicating non-vector data for each vector representation. When your data model assumes everything is organized by vector, adding a second embedding means copying every attribute.
Write amplification. Document updates trigger cascading reorganization because content is keyed by vector location. Harrison writes that "updating just one vector can move hundreds of attributes and their indexes."
Limited vectorization. Modern query engines benefit from processing large data blocks (2,048+ rows), but ANN-based layout constrains operations to 100–200 document clusters, preventing optimal CPU utilization.
turbopuffer's V3 architecture reorganizes storage so that document retrieval no longer depends on ANN addresses. Vector search becomes one index type among many — the same way a B-tree index and a full-text index coexist in PostgreSQL. Their own full-text search redesign already proved the point: decoupling from ANN-cluster-based postings to fixed 256-document blocks achieved 10x smaller indexes and up to 20x faster queries.
Read the full turbopuffer post →
The Squeeze: Models Up, Storage Down
Here is what these two developments look like when you overlay them on the agent stack:
┌──────────────────────────────────────────┐
│ FRONTIER MODEL │
│ (CLM: model manages its own context) │
│ ↓ squeezing down ↓ │
├──────────────────────────────────────────┤
│ CONTEXT / RAG MIDDLEWARE LAYER │ ← under pressure
│ (vector DBs, memory frameworks, │
│ summarization pipelines, RAG chains) │
├──────────────────────────────────────────┤
│ ↑ absorbing up ↑ │
│ STORAGE / DATABASE LAYER │
│ (turbopuffer V3: ANN as secondary idx) │
└──────────────────────────────────────────┘
From above, CLM argues that the model should be doing the context management work currently delegated to harness middleware. The +11.4% on BrowseComp+ is the empirical case: let the model decide what to keep in its context, and long-horizon agent tasks get measurably better.
From below, turbopuffer argues that vector search should be a secondary index inside a general-purpose storage engine, not a separate database category. The 10x smaller / 20x faster numbers from their full-text search decoupling are the empirical case: stop organizing your entire data layout around ANN and everything gets faster.
The middleware layer — where LangChain's retrievers, LlamaIndex's pipelines, Mem0's memory abstractions, and a dozen standalone vector databases live — is caught in between.
This does not mean the middleware disappears overnight. But it does mean the category's durable value is migrating. The question is no longer "how do I build a retrieval pipeline" — it is "what parts of my agent's harness cannot be absorbed by either the model or the database?"
The Harness Infrastructure That Survives
Cloudflare's same-day launch of Clef (open-weight decision models) and K2 (Kafka-on-S3 event streaming) adds another data point. Cloudflare is systematically commoditizing backend infrastructure by building it on object storage. The pattern across all three announcements — CLM, turbopuffer V3, Cloudflare's Clef+K2 — is the same: push intelligence and storage capabilities down the stack, closer to the primitives.
So what survives in the middle? The answer maps cleanly to what CLM explicitly cannot do and what a database index cannot replace:
Tool orchestration and dispatch. CLM lets the model manage context. It does not let the model manage tool authentication, rate limiting, credential rotation, or multi-tenant tool scoping. The harness that routes a model's tool calls to the right API with the right credentials is not a context management problem.
Safety and guardrail enforcement. The CLM paper trains models to decide what context to keep. It does not train models to decide what actions are safe to execute. The harness that gates destructive operations — the human-in-the-loop checkpoint, the budget cap, the PII filter — is a control plane function, not a memory function.
Observability and audit. When an agent runs for 12 hours, you need to know what it did, why, and whether it drifted. Structured logging, distributed tracing, and cost tracking are harness responsibilities that no amount of model-managed context replaces.
Multi-agent coordination. The CLM paper's sub-agent spawning (writing a bash file) is elegant for simple cases. Production fleet orchestration — where agents have different privilege levels, shared state, and handoff protocols — requires harness infrastructure that is inherently external to any single model's context.
NVIDIA's OpenShell, trending at +2,503 GitHub stars today, is building exactly this kind of harness: sandboxed runtimes for autonomous agents with safety and privacy controls. openrig, at +640 stars today, is tackling multi-agent orchestration with persistent teams and shared context. These projects are betting that the control plane, not the memory layer, is where the durable value in agent infrastructure accrues.
⚠️ Contrarian corner: the harness does not die — it specializes. CLM's self-managed context requires models trained (or at least prompted) for it. Production agents still need deterministic tooling for auth, safety, auditing, and multi-tenancy. A context file cannot enforce a spending cap or rotate an OAuth token. The "kill the harness" reading of CLM oversimplifies. What dies is the memory management layer of the harness. What remains — and grows — is the control plane layer.
Community Reaction
The Hacker News thread on turbopuffer's post captured the practitioner consensus forming in real time. The top comment thread debated whether "vector database" was ever a category or always a feature — with multiple commenters pointing to LanceDB's embedded approach and PostgreSQL's pgvector extension as evidence that the standalone category was a temporary artifact of the 2023 RAG hype cycle.
On the CLM side, the Discover AI deep-dive noted the paper's most provocative implication: if models can manage their own context, the entire memory-framework and vector-DB tooling layer faces an existential question. The Dev.to analysis from Gaurav Dadhich was more measured, arguing that CLM changes the default for context management but does not eliminate the need for harness tooling in production settings.
Read the full Dev.to analysis →
View the Cloudflare Clef HN thread →
LangChain's own positioning is instructive. Their context engineering framework defines four strategies — Write, Select, Compress, Isolate — and their agent lifecycle content explicitly frames the shift from "demo agent to production agent" as a tooling-and-discipline story, not a model story. This is a company that sells harness infrastructure acknowledging that the value is in orchestration and lifecycle management, not in the retrieval layer.
Karpathy's observation — that LLMs sometimes need "more bits" to understand what you're trying to achieve — cuts to the heart of why CLM matters. The question is not whether the model is smart enough. The question is whether the harness is feeding it the right context. CLM's answer is to let the model feed itself.
What This Means for You
If you are building or evaluating agent infrastructure today, here is the practical read:
Stop treating your vector database as a moat. Whether you are running Pinecone, Weaviate, Qdrant, or a self-hosted Milvus, the directional bet from both the model side (CLM) and the storage side (turbopuffer, LanceDB, pgvector) is that vector search is converging toward a commodity feature. Plan your architecture so that your retrieval layer is swappable.
Invest in the control plane, not the memory plane. The harness primitives that CLM cannot absorb — tool orchestration, safety gates, observability, credential management, multi-agent coordination — are where durable competitive advantage lives. If your agent framework's primary value proposition is "we manage context for you," that value proposition is depreciating.
Watch for CLM reproductions. The paper's results are strong (+11.4% on BrowseComp+, +5% on 12-hour EdgeBench) but from a single research group. Independent reproductions will determine whether this is a technique that generalizes or one that requires specific training recipes. The fact that they got results with Qwen3.5-9B suggests breadth, but breadth needs confirmation.
Evaluate your stack layer by layer. For each component in your agent's harness, ask: "Can the model absorb this?" and "Can the database absorb this?" If the answer to either is yes, that component is a candidate for compression. The components where both answers are no — that is your durable infrastructure.
💡 The bottom line: The agent stack is repricing. Memory and retrieval are becoming commodities. The control plane — tool orchestration, safety, observability, human-in-the-loop — is where the moat is forming. Build there.
We have been tracking the harness-as-moat thesis on AgentConn since early 2026. This week's CLM paper and turbopuffer pivot are the strongest evidence yet that the thesis is correct — but the location of the moat is shifting from the memory layer to the control plane. Our earlier coverage of agent memory wars and the Jev decision-model layer both pointed in this direction. The convergence is now hard to ignore.
The model got smarter about context. The database got smarter about vectors. The middleware in between? It needs to find a new reason to exist — or get absorbed.
Originally published at AgentConn







Top comments (0)