Introduction
I asked an AI coding agent to read a paper about AI coding agents. The paper spent a good chunk of its length dissecting the exact tool I was typing into.
That isn't a riddle. The paper is Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents (arXiv:2609.00006), a July 2026 source-code study by Paul Barbaste, Tristan Darrigol, Germain Vu, and Tom Wiltberger of the Wavestone AI Lab. One of the eleven systems it takes apart is Pi, the harness I use every day. Another is OpenCode, the host for oh-my-opencode-slim, the plugin I sent a Korean README fix.
So I did the thing the paper invites. I checked its homework. I opened Pi's local documentation and compared three of its claims about Pi against the primary source. Two matched to the digit. The third was more interesting than either.
Here's what the study found, what it means whether you use these tools or build them, and what happened when I audited the auditors.
The Research Process
This is a descriptive study, not a competition. The authors don't benchmark or rank anything. They read source code.
- Eleven harnesses, pinned to July 2026 releases: Claude Code (Anthropic), Codex CLI (OpenAI), Gemini CLI (Google), Mistral Vibe (Mistral), plus OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode, and OpenClaw.
- A twelfth system, Omnigent (Databricks), is analyzed as a "meta-harness", an orchestration layer that drives eleven vendor harnesses behind one API.
- Roughly four million lines of Python, TypeScript, and Rust, inspected down to dependency manifests and import statements.
- Because eight systems were re-pinned rather than replaced from an April 2026 edition of the same study, the paper also contains a controlled 90-day source diff. The same codebases, three months apart.
Out of that came 13 cross-cutting observations, a catalog of 29 recurring design patterns, 18 design recommendations, and a ~90-line minimum-viable-harness scaffold in Python.
Two disclosures shape how you should read it. First, most performance numbers in the paper are self-reported by the systems' own maintainers. Second, the paper was written with substantial assistance from Claude Code, disclosed in an acknowledgments section, and worth knowing since Claude Code is one of the eleven systems under study.
Key Findings
The paper has thirteen observations. These are the four that changed how I think about the tools on my own machine.
1. An agent is a model plus a harness. The harness is the product.
The paper opens with a one-line equation: Agent = Model + Harness. The harness is everything except the model. The loop, the tools, the context management, the safety controls, the orchestration, the extension surfaces.
It then argues that every system, from Mini-SWE-Agent's hundred-line research baseline to Claude Code, has to take a position on the same seven subsystems:
- Agent loop
- LLM integration
- Memory and context
- Tool and action system
- Safety and permissions
- Extensibility (skills, hooks, plugins, MCP)
- Multi-agent orchestration
Even "no position" counts as a position. Pi isn't simply missing a sandbox. The paper documents that the absence is a stated design principle, argued for in its own docs. That's the pattern across the corpus: deliberate refusals are documented as carefully as features.
This is why the tools feel so different while running the same handful of frontier models. The model is shared. Everything you actually touch is the harness.
2. The twin absences
This is the finding I didn't expect.
Across four million lines and eleven production systems:
- Zero import a general-purpose agentic framework. Not LangChain, not LangGraph, not AutoGen, not CrewAI, not LlamaIndex, not Pydantic AI, not Semantic Kernel. Gemini CLI uses neither of Google's own frameworks, Genkit and ADK.
- Zero use vector embeddings to retrieve code. Instead they use ripgrep, tree-sitter, glob, and auto-discovered Markdown context files like
AGENTS.md.
The authors searched for counterexamples for weeks, re-ran the sweep after tripling the corpus, and the result held. Every loop is hand-rolled in the host language's native primitives. asyncio, Tokio, Promises. Every tool registry is custom. Every prompt template is plain Markdown or string concatenation.
There's a footnote that savors its own irony: the term harness engineering was named and defined from inside LangChain, the framework vendor whose libraries appear nowhere in the study's runtime code.
The nuance matters before you argue with this finding. Embeddings do appear in OpenClaw, on by default, but for chat recall, never for reading a source tree. Aider can install llama-index as an optional extra, but only to run RAG over its own documentation, not inside its agent loop.
3. Loop sophistication doesn't predict performance
The paper returns to this point repeatedly. Mini-SWE-Agent's linear loop is roughly fifty lines. Its self-reported SWE-Bench Verified score is 74%+.
Codex's workspace, meanwhile, nearly doubled in a single quarter — from 621,000 to about 1.12 million lines of Rust, across 89 to 126 crates. That growth isn't the loop changing. It's everything around the loop: sandboxing, approvals, memory pipelines, plugin marketplaces.
The paper's conclusion: architectural complexity doesn't predict how well an agent scores, but it does predict production readiness — safety, reliability, extension surfaces, and increasingly the client/server and transport layers, where the largest harnesses now carry most of their mass.
The authors back this up with a scaffold. Listing 3 is roughly 90 lines of Python implementing a linear loop, four tools (bash, read, write, search_replace), root-to-leaf AGENTS.md discovery, and threshold compaction. It deliberately omits sandboxing, multi-agent orchestration, MCP, and skills, because the corpus shows real disagreement there. Their advice is blunt: start here, measure, and add the minimum your observed failure modes demand.
4. The standards resolved, and then everyone copied each other
Two format fights settled during the study's window.
Skills beat MCP. SKILL.md skills are used by 9 of 11 systems; MCP by 8 of 11. The paper credits Pi with breaking the tie by implementing skills while rejecting MCP outright. The skills layer then grew a supply chain: registries, trust tiers, provenance verification, cross-vendor discovery (OpenCode reads Claude Code's skills directory), and the corpus's first agent-authored skills.
ACP found a second job. The Agent Client Protocol ships in 6 of 11 systems. It was designed for the editor-to-agent boundary, but the study documents a role outside its original brief: harness hosting. OpenHands can run Claude Code, Codex, or Gemini CLI as interchangeable backends behind its own interface.
Then there's the 90-day diff, which is the part I keep thinking about. In the April edition, the systems had converged by independent rediscovery. By July, the convergence was traceable. Codex adopted Claude Code's hook event vocabulary verbatim and shipped an importer for its sessions and settings. OpenHands adopted Claude Code's plugin manifest format. Patterns visibly diffused down the corpus.
The paper's assessment: the half-life of a competitive differentiator in this field is currently measurable in weeks.
What This Means for You
If you use coding agents
The study explains behaviors you've probably blamed on the model.
- The agent that gets vague after a long session is compacting. Claude Code fires compaction below a 13K-token buffer and post-restores files. Gemini CLI compacts at 50%, preserving the last 30% verbatim. Pi triggers at
contextWindow − 16,384and keeps the 20,000 most recent tokens. - The agent that keeps re-reading your repo instead of "remembering" it has no embeddings. And that's deliberate, not a missing feature.
- The agent that refuses to use your favorite framework isn't missing a dependency. It's following the entire field.
- Tool count matters more than you'd think. Past roughly 15 tools, the paper says, prompt bloat becomes prohibitive and systems switch to deferred tool loading. Claude Code's deferred flag reportedly cuts the initial prompt by about 40%.
If you build agents
Section 16 is the payoff: 18 recommendations, each anchored to an observed system and an explicit trade-off. The ones that stuck with me:
- Start with a linear loop and one
bashtool. Add more tools only in response to observed failure modes. - Graduate to a middleware pipeline when you have three or more independent turn-level policies, not before.
- Don't build RAG over code. Code has deterministic structure that semantic similarity can't replicate, and it changes minute-to-minute.
- Codify safety rules as data or policy files, not imperative code. And if you ship a YOLO mode, keep a floor beneath it.
- Stay single-agent until you can point at a specific phase where parallel context isolation clearly beats serial search.
What I Verified Myself
Pi's docs are on my machine, so I compared the paper's claims against the primary source.
Claim 1: Pi compacts at contextWindow − 16,384, keeping 20K recent, and merges summaries iteratively rather than re-summarizing from scratch. Verdict: exact match. docs/compaction.md lists a reserveTokens default of 16384, a keepRecentTokens default of 20k, and a summary step that passes "the previous summary as iterative context."
Claim 2: A session_before_compact hook lets extensions veto or replace a compaction result. Verdict: exact match. It's a documented event.
Claim 3: Pi is the only system in the corpus that documents the absence of safety infrastructure as a design argument. Verdict: the docs agree, in plainer language than the paper uses. From docs/security.md:
"Watching the transcript, using project trust, and reviewing changes do not create a security boundary."
That third check is the one I keep coming back to. The paper described a design philosophy I'd never seen stated so directly in the tool I use every day, and then the primary source said the same thing without any academic hedging. It changed how I read the rest of the paper and how I think about what "safe" means in a terminal that will run rm -rf on my behalf.
Limitations
The paper is careful about its own limits, and you should be too.
- It's descriptive. The authors explicitly don't benchmark or rank, so don't read "finding" as "winner."
- Most performance figures are self-reported by maintainers, not independently reproduced.
- Claude Code is analyzed from a circulated source snapshot, not a public repository. A different class of evidence than the other ten.
- The draft was AI-assisted, and one of the assistants is one of the studied systems' closest relatives.
- Some claims will age. The authors separate inventory claims (tool counts, feature cells, version pins), which decay in weeks, from structural claims (the loop taxonomy, the subsystem anatomy, the twin absences), which have proven durable and they label which is which.
Closing
The study's biggest claim is that coding agents stopped being tools and became platforms in the first half of 2026. The evidence is all in source: harnesses shipping as importable SDKs while framework vendors ship harnesses; cross-vendor session importers; marketplaces, registries, and enterprise governance layers; and a meta-harness making its own harnesses interchangeable.
Or as the paper puts it in Observation 12: "The competitive unit of the field is no longer the agent loop; it is the ecosystem surface around it."
If you use these tools daily, that's the finding worth sitting with. The loop isn't where the competition is anymore. And the platform around it is being built to hold you.
Call-to-Action
Read the paper yourself: Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents — the HTML version is fully readable.
Then do what I did. Pick your own harness, open its docs, and check the paper's claims about it. If you find a discrepancy, that's a better blog post than this one.
If you build with these tools, start with Section 16 and the 90-line scaffold in Listing 3.

Top comments (4)
Checking the paper against the primary docs is the part most readers skip, and it is the useful part. Harness details like context compaction and tool permission defaults change between releases, so a study pinned to July can drift from what you run today. Which of the three claims was the interesting mismatch, and did it change how you configure Pi?
iin1006h08
Hi, I've just been running default settings all the time. So no config changes for now
Pi's warning that watching the transcript isn't a security boundary puts a useful limit on the "start with a linear loop and one bash tool" advice. I'd happily start there for a prototype, but define what that tool can access before letting it touch customer data or production credentials. Your checks of the 16,384-token reserve and 20K recent-token window also make compaction a concrete engineering choice: I'd test whether unresolved constraints survive it, since a session can keep running smoothly while quietly losing the requirement that mattered most.
The finding on vector retrieval versus ripgrep and tree-sitter matches every practical coding harness in the wild.
Chunking source files into fixed token windows and calculating cosine similarity destroys the exact things an agent needs during edits, especially import graphs, symbol definitions, and line offsets. If an agent needs to know where a function is invoked, ripgrep returns the exact file and line in four milliseconds, while an embedding search returns three vaguely related comments and a README paragraph.
Frameworks hit the same wall. The moment a tool execution fails or a model emits malformed JSON, having multiple layers of middleware between the prompt string and the HTTP client makes debugging an agent loop unbearable. Plain async loops and string templates win simply because they keep the blast radius small.