Textbook: The Anatomy of an Agent and Five Waves of Evolution
Source version of CodeSmith:
v0.5.0(commit3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: readers new to Agent engineering who want a coordinate system of "loop–autonomy–multi-agent–evolution."
The prologue told of an invisible war, then tossed out a word: harness. Over the next twenty-odd installments we will plunge headlong into CodeSmith's organs — caches, handles, compaction, approval gates. But before the scalpel comes out, this installment wants to pause and dissect the word "Agent" itself: what the loop looks like, how many levels autonomy comes in, whether multi-agent can be trusted, and why in 2026 every engineering team is talking about harness engineering. These are the common knowledge of this craft, owned by no single project; and CodeSmith's source will keep surfacing to confirm them — including the passage of code pasted below, the closest thing to a textbook in the entire repository.
The Loop: Think–Act–Observe
The core abstraction of the modern Agent is the ReAct loop: the model first reasons about the current state and the actions available, executes one action (search, query a database, run code), the environment returns an observation, and the model reasons its next round from that observation. It sounds bland, but it carries a less-than-obvious corollary: an Agent's action trajectory cannot be reduced to one longer static answer.
The reason is that action changes what information is available. In pure-reasoning mode, if the context holds no information about "whether this code compiles," the model cannot "think it up" — it can only guess. In the loop, by contrast, the model runs the compile command, the compiler's error output becomes a new fact in the context, and subsequent reasoning is built on those new facts. An Agent that has run twenty steps owes its final judgment to the environmental feedback of every one of those steps — feedback that, before the actions took place, existed nowhere at all. So "replacing the loop with a longer prompt" holds only when the task involves no interaction with the environment.
Inside CodeSmith's engine sits a loop implementation that reads like a textbook. Not the 16,000-line host_executor on the production path, but the reference implementation in the framework crate, DefaultAgentExecutor::run_inner (crates/agent/src/executor/mod.rs:140), whose module-header comment bills itself as "The LangChain AgentExecutor analog". Strip away the peripheral details and its skeleton looks like this:
// crates/agent/src/executor/mod.rs:140 (excerpt)
let mut step: u32 = 0;
loop {
if step >= max_steps {
callback.on_complete(&StopReason::MaxSteps).await;
return Ok(StopReason::MaxSteps);
}
// ...assemble the MessageRequest (model, message history, system, tool catalog)...
let stream = client.create_message_stream(request).await?;
let (content, _stop_reason) = accumulate_stream(stream).await?;
// Persist the assistant turn.
history.push(Message { role: "assistant".to_string(), content: content.clone() });
// Collect tool calls (preserve order).
let tool_uses: Vec<(String, String, serde_json::Value)> = content
.into_iter()
.filter_map(|block| match block {
ContentBlock::ToolUse { id, name, input, .. } => Some((id, name, input)),
_ => None,
})
.collect();
if tool_uses.is_empty() {
return Ok(StopReason::NoToolCalls);
}
// Execute each tool sequentially and feed the result back as a
// `role:"user"` `ToolResult` block (Anthropic/OpenAI-compat shape).
for (id, name, input) in tool_uses {
let result = match tools.get(&name) {
Some(tool) => tool.run(input.clone()).await,
None => Err(ToolError::NotAvailable { /* no tool named '{name}' */ .. }),
};
history.push(Message {
role: "user".to_string(),
content: vec![ContentBlock::ToolResult { tool_use_id: id, /* ... */ }],
});
}
step += 1;
}
This code is doing exactly one thing — ask the model, collect tool calls from the reply, execute them and stuff the results back into the history under the user role, until the model stops asking for tools. "Thinking" is the assistant message, "acting" is the ToolUse block, "observing" is the backfilled ToolResult — ReAct's three beats, mapped word for word onto the message structure. Two details deserve a second look: tool results return to the history under the identity of role:"user" (this foreshadows Article 10's status-bar design), and calling a nonexistent tool is not a crash but a NotAvailable error result — fed back to the model, so that it corrects itself.
The stop conditions are gathered into a single enum (crates/agent/src/callback/mod.rs:23):
pub enum StopReason {
/// The model produced an assistant turn with no tool calls — the run is
/// finished.
NoToolCalls,
/// The step budget (`max_steps`) was exhausted mid-tool-loop.
MaxSteps,
/// The run aborted with an error.
Error(String),
/// The run was cancelled (user/external interruption). Distinct from
/// [`StopReason::Error`] so the host can surface "cancelled" rather than
/// "error" — mirrors production's `TurnOutcomeStatus::Interrupted`.
Interrupted,
}
Four values tell the whole story of the loop's exits: the model finished speaking, the steps ran out (max_steps defaults to 50), an error occurred, the user interrupted. "User interruption" gets its own variant rather than being folded into Error so that the interface can truthfully say "cancelled" instead of "error" — honest stop conditions and honest error reports are one and the same virtue.
The step-loop skeleton of the production HostAgentExecutor (crates/agent-runtime/src/engine/host_executor.rs:2642) matches the reference implementation, but roughly ten guardrails hang on before and after each step: cancellation checkpoints, system-prompt snapshot refreshes, compaction, capacity preflight, LSP diagnostics flush, the loop guard... Those guardrails are the protagonists of this series' second half; for now, they stay offstage.
Six "Fake Agents": Set Up the Targets Before Shooting
In industry usage, the word "Agent" suffers obvious marketing inflation. Lay the common misuses of the concept out on the table, and the boundaries of every discussion that follows become much clearer:
| Claim | What it actually is | One-sentence puncture |
|---|---|---|
| "Autonomous agent completed the search" | A single API call | A system that keeps no state, chooses no actions based on observations, and has no stop condition is not an Agent |
| "Autonomous workflow" | A fixed pipeline dressed up | All the decisions are hard-coded; the model only fills in the blanks |
| "The Agent gets better as it works" | Long-chain error propagation | One wrong file edit contaminates every subsequent operation; autonomy has nothing to do with it |
| "Experiment succeeded, tests passed" | Environmental hallucination | The "it's done" narrative had no real execution behind it — more dangerous than text hallucination |
| "The foundation model has this capability" | Capability misattribution | The gains from good retrieval, dedicated interfaces, and heavy retries get booked to the model's weights |
| "The Agent decided to…" | The autonomy myth | Choosing an action among given tools does not mean it "wants" or "decides" |
One touchstone for telling the real from the fake: when the environment returns an unexpected result, can the system change its sequence of actions? If it can, it has earned the right to talk about loops; if it cannot, it is nothing more than a fill-in-the-blank exam with a temperature setting. The touchstone cuts both ways — against the things we build ourselves too. CodeSmith's loop guard (Article 10) exists precisely because the distance between a system that "can change its action sequence" and one "destined to repeat the same action" is a single guardrail.
The Ablation Study: The Four Context Components Are Not Equally Important
Someone has run a systematic ablation study on an Agent's context: keep the full baseline, then remove one component at a time as the controls — tool definitions, tool execution results, the reasoning process, message history (the system prompt is the identity definition; ablate it and even running the test becomes meaningless, so it does not take part). The conclusions deserve to be memorized line by line:
- Tool definitions are the foundation of the capacity to act. Stripped of them, the model does not go silent — it still hands you a neatly formatted, confidently toned answer, except the data now comes from parametric memory, typeset exactly like a genuinely observation-based one;
- Tool results are the key to closed-loop control; without them the Agent executes "blind," retrying over and over until the iteration budget is exhausted;
- The reasoning process records "why this was done," while tool results record "what happened" — when the former can be rebuilt from the latter, dropping the thinking from the history costs almost nothing (this is the license under which Article 11's compaction dares to operate);
- Message history prevents redundant operations and spares you from making the same mistake twice.
The experiment's core insight compresses into one sentence: the context determines what the Agent can see, and the Agent can only make decisions from what it sees. And the components are not equivalent — the measure is whether the information a component carries can be rebuilt from elsewhere. One more finding matters even more to engineering practice: the typical failure under a mutilated context is not an error exit but a flawless-looking answer — "produced a reply" is not "completed the task." All the context engineering of Articles 12 through 15, and the checklist by which the improvement plan audits itself against the textbook, stand on this foundation.
ACI: Capability Is a Joint Product of Model and Interface
SWE-agent made a far-reaching discovery: the same foundation model performs worlds apart on codebase-modification tasks under a plain shell interface versus under a purpose-designed Agent-Computer Interface (ACI). How the interface presents file contents, how it formats edit commands, how it returns error messages — each of these bears on whether the model can operate code effectively. Put differently — an Agent's capability is not an attribute of the model weights alone, but a joint product of the model and the environment's interface.
This is the academic rendering of the prologue's "the difference is the harness." Take the harness apart and three groups of components fall out:
- The cognitive interface — the system prompt fixes the role and output format, tool schemas define what can be called, the ACI decides how results come back. Together they determine what the model can see and what it can do;
- The execution environment — filesystem, context organization, permission sandbox. Determines what happens when things go wrong;
- Audit and constraint — trajectory logging, budgets and stop mechanisms, evaluators, human takeover. Ensures that actions stay traceable and bounded.
Fit these three groups onto CodeSmith's 21 crates and you will find them lined up in tidy ranks: agent-runtime's Constitution and prompts and tool-impls' fifty weapons belong to the cognitive interface; agent/providers' message history and sandbox belong to the execution environment; execpolicy's command review, side-git snapshots, and loop guard belong to audit and constraint. Every installment this series has ahead of it is about one block of these three groups.
Five Levels of Autonomy and Five Forms of Loop
"Autonomy" is not a have-it-or-not property but a continuous variable, divisible into at least five levels. Level 1: the developer specifies each action step by step, and the model only fills in text. Level 2: the model chooses actions from a given tool set — the base mode of most Agent systems today. Level 3: the model can revise the plan, abandoning the original path when the environment returns a surprise. Level 4: the model can propose its own subgoals and decompose them. Level 5: the model can examine the task's goal and its evaluation criteria themselves — "is this task even worth doing?" The first four levels answer "how to get the task done"; the fifth pushes the question to "whether the task itself stands."
By this scale, CodeSmith lives between levels 2 and 3: the model freely chooses tools and arguments (level 2), and the loop guard and the capacity controller can force a VerifyAndReplan that resets the trajectory when it drifts (a passive level 3). Level 5 is, for now, a luxury for any production system.
The loop itself has evolved a family of variants. Along three dimensions — what is saved, what is read, and what triggers the next round — at least five distinct forms can be told apart:
| Form | What it saves | Echo in this series |
|---|---|---|
| ReAct (reactive) | No extra memory maintained across steps | The reference implementation above |
| Reflexion (reflective retry) | After failure, generates self-reflection, stores it in external memory, reads it back together next time | Article 23's experience distillation |
| LATS (search tree) | Each step is a node in a search tree; multiple candidate paths backtrack by value function | — |
| Voyager (skill accumulation) | Successful action sequences are encoded into reusable skill programs and filed into a skill library | Article 2's skills; Article 23 |
| MemGPT (layered memory) | A layered memory system with active read/write, managing history beyond the window | Article 8's virtual memory |
The differences among the five forms are, in essence, differences of memory strategy: ReAct is "no memory," Reflexion is "remember only on failure," Voyager is "remember only on success," MemGPT is "memory itself needs paging." In Article 8 you will see the bloodline connecting CodeSmith's VarHandle to MemGPT — that paper's metaphor of choice was the operating system's virtual memory.
Multi-Agent: A Bucket of Cold Water
When a task passes from one Agent to another, the delegation has to be contracted explicitly, or it will fail at some boundary. A complete delegation contract has at least eight clauses: objective, tool permissions, forbidden actions, resource ceilings, abort conditions, output format, lines of responsibility, and a renegotiation mechanism for when the environment changes. Leave out the resource ceiling and the delegate may burn compute without end; leave out the abort condition and a subtask that has already drifted from its goal will simply keep producing useless results.
To anyone too optimistic about multi-agent implementations, I would like to throw three buckets of cold water:
- First bucket: the relationship between team size and coordination gains is not monotonic — past some threshold, coordination costs climb steeply and marginal returns diminish or even go negative; on many tasks, adding more Agents does not beat one carefully designed single Agent.
- Second bucket: the unanimous agreement of ten same-origin Agents cannot be counted as ten independent pieces of evidence — ten instances of the same foundation model agreeing says more about the model's shared preferences than about any convergence of independent observations.
- Third bucket: multi-Agent debate may lose to simple majority voting — in debate, a weaker argument can ride rhetorical force past a stronger one, and debate amplifies initial anchoring.
So multi-Agent value demands specific conditions: the task must genuinely decompose, the subtasks must be independent enough, and the coordination mechanism must be able to handle dependencies. Article 6 shows how CodeSmith builds work crews under these constraints — every member with its own prompt, its own model, its own git worktree — which is, in essence, the responsibilities and resource boundaries from the "eight elements of the delegation contract" turned into configuration items.
Five Waves of Evolution: Competitive Advantage Moves Beyond the Model
Pull the camera back from any single system, and AI application engineering has traced a clean arc over the past few years:
Graph engineering — Agent loops, deterministic programs, and human approval organized into an explicit execution graph
└── Loop engineering — sustained autonomous operation across turns: who spots the next thing, when to verify, when it counts as done
└── Harness engineering — context and tool interfaces, constraints, verification, feedback loops, error recovery
└── Context engineering — systematically managing everything the model can see
└── Prompt engineering — optimizing the natural-language instructions fed to the model
Notice that these five waves do not replace one another — they nest: prompt engineering is a subset of context engineering, context engineering a subset of harness engineering — and the individual Agent loop is precisely one node in the execution graph. Each layer widens the engineer's field of attention beyond the one before it.
Why does this arc point beyond the model? LangChain's practice on Terminal Bench 2.0 (a benchmark that evaluates Agents completing complex tasks in a terminal environment) supplies a forceful footnote: the score climbed from 52.8% to 66.5%, vaulting from outside the top thirty on the leaderboard into the top five — what was swapped was not the model, it was the harness. The concrete measures included letting the Agent automatically check its own execution results, detecting whether it had fallen into a repetitive loop, and refining its thinking strategy. As the capabilities of the various models converge and cease to be the decisive differentiator, competitive advantage migrates to the engineering practices outside the model. That sentence is the reason this entire series exists.
On engineering principles, Anthropic gathers the lessons of successful Agents into three: keep it simple (direct API calls beat complex frameworks; every additional layer of abstraction is a fresh blind spot for future debugging); keep it transparent (planning, logs, and decision trajectories stay visible — an error inside a black box can be neither located nor corrected); and design the tool interface well — design it from the Agent's point of view, and where misuse comes easily, make the error impossible by design. Manufacturing has a term of art for the third: poka-yoke (mistake-proofing), out of the Toyota Production System — the notched corner of a SIM card makes inserting it backwards impossible, and a microwave with its door not properly shut will not heat. These three will keep echoing through every installment to come: the Constitution's layering is simplicity, the status bar is transparency, execpolicy's three-valued Decision is poka-yoke.
Coda: Anatomy Chart First, Surgery Second
Gather this installment's contents into six sentences:
- An Agent's essence is the think–act–observe loop, and the action trajectory cannot be reduced to one longer static answer;
- Context components are not equally important — the criterion is "can it be rebuilt from elsewhere" — and the typical symptom of a mutilated context is a perfect, useless answer;
- Capability is a joint product of model and interface, and the harness's three component groups (cognitive interface, execution environment, audit and constraint) are that interface's engineering checklist;
- Autonomy is a continuous variable, and the loop has five forms with distinct memory strategies;
- Multi-Agent works conditionally: delegation must be contracted, and same-origin consensus does not count as independent evidence;
- Engineering's center of gravity is moving along the five-wave arc to beyond the model — the harness is the current battlefield.
The anatomy chart is drawn. Next comes the first cut, beginning with the most valuable item on this checklist of distrust: the 100× price tag of a single byte.
Top comments (0)