Prologue: Like Jev — A Harness for Cheap Brains
Source version of CodeSmith:
v0.5.0(commit3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: anyone who has used a coding agent and wondered what parts a harness is actually assembled from.
Series index: README.
Start with a piece of code. This is filtering logic from the CodeSmith engine, the kind that runs while it streams and decodes model output — my favorite stretch of code in the entire repository:
// crates/agent-runtime/src/engine/streaming.rs:59
pub const TOOL_CALL_START_MARKERS: [&str; 5] = [
"[TOOL_CALL]",
"<codesmith:tool_call",
"<tool_call",
"<invoke ",
"<function_calls>",
];
Five "start markers," matched by five "end markers" (streaming.rs:67-73). What they filter out is, said out loud, a little funny: the model is pretending to call tools. To the person at the keyboard, this is information that never needed to be shown — Jev (the currently fashionable System One model from TypeSafe AI: one to two orders of magnitude faster, far cheaper, its engineering ability made up entirely by a harness bolted on from outside) works hard to screen exactly this kind of thing out of view.
An Invisible War
Here is what happens. Once you hook an open-source model up to a coding agent, you discover a problem that barely existed in the closed-source era: instead of using the API's tool-call channel like a good citizen, the model hand-writes a passage like this in its plain-text output —
[TOOL_CALL] {"name": "read_file", "arguments": {"path": "src/main.rs"}} [/TOOL_CALL]
It looks entirely earnest, and the syntax is all correct, but to the engine this is just ordinary text. At the API level, nothing has happened: no tool was called, no file was read — the model is just chanting an incantation. The bigger nuisance is that after finishing the incantation it keeps right on writing, as though the file's contents were already in hand — and so every step of reasoning that follows is built on hallucination.
Why does the model do this? Because its training data is soaked in conversations of this shape — other frameworks' formats, formats from the ChatGPT-plugin era, formats from tutorials that hand-simulated tool calls. The model learned "this is what a tool call looks like," not "a tool call goes through this protocol channel." The closed-source frontier models suppressed the habit with RLHF and oceans of well-formed samples; cheap open-source models flash it now and then.
CodeSmith's answer is not to beg in the prompt, "please don't do this" — that never works — but to accept reality: the moment any of these five wrappers appears in the output, it gets stripped on the spot. filter_tool_call_delta (streaming.rs:95-122) is a state machine that spans chunks: streaming output can cleave a marker in two, so an in_tool_call boolean remembers "I am currently inside a fake wrapper," and however a marker gets sliced by a chunk boundary, the halves still meet. After the strip, it also sends a notice up to the UI:
// crates/agent-runtime/src/engine/streaming.rs:79
pub const FAKE_WRAPPER_NOTICE: &str =
"Stripped non-API tool-call wrapper from model output (use the API tool channel)";
The very existence of that notice is a stance: the user deserves to know why their text got shorter. Tame the model's behavior — but never hide the taming from the user.
This small component is the whole project in miniature, and it is what this series is about: serving as the harness for cheap open-source models is engineering work full of exactly this kind of dirty work.
Harness: The Missing Layer Between a Model and an Agent
The CodeSmith README opens with a sentence that is the project's entire self-positioning (README.md:22):
A model answers a question; an agent finishes a task. CodeSmith is the harness in between.
A model answers questions; an agent finishes tasks. The difference between them is not a smarter model — it is that there is a harness: a system of rules, evidence, and feedback that keeps the model from drifting while it works through a multi-step task.
The word "harness" is a deliberate choice. However strong a horse, it cannot pull a plow without a harness on; however capable an open-source model, it cannot do long-chain engineering work without one. A harness generates no force; it transmits force and constrains it. You will find the word far more accurate than "framework": a framework is a stage the model stands on, while a harness is a restraint the model is strapped into.
What does this harness concretely look like? The one-sentence version: a written constitution (crates/agent-runtime/src/prompts/base.md, 297 lines); a nine-level hierarchy of authority (the user's current instruction > stale project rules; a tool's measured output > assumptions; verification > confidence); three modes (Plan read-only / Agent with approvals / YOLO auto-pass); one OS-level sandbox (macOS Seatbelt, Linux bubblewrap — Landlock has a type reserved for it but was never wired up, and Windows does not even have a helper; that bill is tallied in detail in Article 20); plus a side-git snapshot every turn and concurrent sub-agents on call.
But this series will not march down a feature checklist. What I want to tell is the ledger behind the design decisions — why each cut was made where it was, what it cost, and which ones I have regretted.
Why Me
A word about the lineage — it matters for understanding the project.
CodeSmith was not written from scratch. Its predecessor is CodeWhale (earlier known as deepseek-tui).
And the mindset behind this series is not "watch me reinvent a wheel," but rather: having read a great many open-source coding agents and agents with long-horizon task capability, here are the things I wanted to do — and did.
The engineering scale, laid out here up front — later installments will keep returning to this coordinate system:
-
21 crates in the Cargo workspace:
agent-runtime(engine core),tui(interactive runtime),agent/providers(LLM abstraction and adapters for 16 model services),execpolicy(command security),index(tree-sitter symbol index),mcp,hooks,extensions… - 548 Rust source files, 356,193 lines of code (find+wc counting, including comments and inline tests)
- 5,429 test functions — roughly one test per 66 lines of code, including end-to-end QA that runs a real TUI inside a sealed pseudo-terminal
Build something like this in Rust, and one frequently asked question is why not Python/TypeScript. The answer will surface naturally across the later installments; here is the one-line version: when you need a cross-chunk state machine inside streaming output, fat pointers crossing an extern "C" boundary, and a byte-level fingerprint of the system prompt before every request, you want a type system standing guard over those boundaries.
The Hidden Thread: A Catalog of Distrust
One hidden thread runs through the entire series — perceptive readers may already smell it: these beams are all the same thing — distrust. Distrust that the model will use the tool channel honestly (hence the streaming strip-out); distrust that the model can manage its own context (hence handles and RLM); distrust that the model finishes work without leaving garbage behind (hence the ledger and snapshots); even distrust of your future self evaluating model changes (hence a snapshot every turn). A good harness is, in essence, a clearly written catalog of distrust.
Top comments (0)