DEV Community

DogeKing
DogeKing

Posted on

CodeSmith:Like Jev A -Harness for Cheap Brains

Prologue: Like Jev — A Harness for Cheap Brains

Source version of CodeSmith: v0.5.0 (commit 3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: anyone who has used a coding agent and wondered what parts a harness is actually assembled from.
Series index: README.

Start with a piece of code. This is filtering logic from the CodeSmith engine, the kind that runs while it streams and decodes model output — my favorite stretch of code in the entire repository:

// crates/agent-runtime/src/engine/streaming.rs:59
pub const TOOL_CALL_START_MARKERS: [&str; 5] = [
    "[TOOL_CALL]",
    "<codesmith:tool_call",
    "<tool_call",
    "<invoke ",
    "<function_calls>",
];
Enter fullscreen mode Exit fullscreen mode

Five "start markers," matched by five "end markers" (streaming.rs:67-73). What they filter out is, said out loud, a little funny: the model is pretending to call tools. To the person at the keyboard, this is information that never needed to be shown — Jev (the currently fashionable System One model from TypeSafe AI: one to two orders of magnitude faster, far cheaper, its engineering ability made up entirely by a harness bolted on from outside) works hard to screen exactly this kind of thing out of view.

An Invisible War

Here is what happens. Once you hook an open-source model up to a coding agent, you discover a problem that barely existed in the closed-source era: instead of using the API's tool-call channel like a good citizen, the model hand-writes a passage like this in its plain-text output —

[TOOL_CALL] {"name": "read_file", "arguments": {"path": "src/main.rs"}} [/TOOL_CALL]
Enter fullscreen mode Exit fullscreen mode

It looks entirely earnest, and the syntax is all correct, but to the engine this is just ordinary text. At the API level, nothing has happened: no tool was called, no file was read — the model is just chanting an incantation. The bigger nuisance is that after finishing the incantation it keeps right on writing, as though the file's contents were already in hand — and so every step of reasoning that follows is built on hallucination.

Why does the model do this? Because its training data is soaked in conversations of this shape — other frameworks' formats, formats from the ChatGPT-plugin era, formats from tutorials that hand-simulated tool calls. The model learned "this is what a tool call looks like," not "a tool call goes through this protocol channel." The closed-source frontier models suppressed the habit with RLHF and oceans of well-formed samples; cheap open-source models flash it now and then.

CodeSmith's answer is not to beg in the prompt, "please don't do this" — that never works — but to accept reality: the moment any of these five wrappers appears in the output, it gets stripped on the spot. filter_tool_call_delta (streaming.rs:95-122) is a state machine that spans chunks: streaming output can cleave a marker in two, so an in_tool_call boolean remembers "I am currently inside a fake wrapper," and however a marker gets sliced by a chunk boundary, the halves still meet. After the strip, it also sends a notice up to the UI:

// crates/agent-runtime/src/engine/streaming.rs:79
pub const FAKE_WRAPPER_NOTICE: &str =
    "Stripped non-API tool-call wrapper from model output (use the API tool channel)";
Enter fullscreen mode Exit fullscreen mode

The very existence of that notice is a stance: the user deserves to know why their text got shorter. Tame the model's behavior — but never hide the taming from the user.

This small component is the whole project in miniature, and it is what this series is about: serving as the harness for cheap open-source models is engineering work full of exactly this kind of dirty work.

Harness: The Missing Layer Between a Model and an Agent

The CodeSmith README opens with a sentence that is the project's entire self-positioning (README.md:22):

A model answers a question; an agent finishes a task. CodeSmith is the harness in between.

A model answers questions; an agent finishes tasks. The difference between them is not a smarter model — it is that there is a harness: a system of rules, evidence, and feedback that keeps the model from drifting while it works through a multi-step task.

The word "harness" is a deliberate choice. However strong a horse, it cannot pull a plow without a harness on; however capable an open-source model, it cannot do long-chain engineering work without one. A harness generates no force; it transmits force and constrains it. You will find the word far more accurate than "framework": a framework is a stage the model stands on, while a harness is a restraint the model is strapped into.

What does this harness concretely look like? The one-sentence version: a written constitution (crates/agent-runtime/src/prompts/base.md, 297 lines); a nine-level hierarchy of authority (the user's current instruction > stale project rules; a tool's measured output > assumptions; verification > confidence); three modes (Plan read-only / Agent with approvals / YOLO auto-pass); one OS-level sandbox (macOS Seatbelt, Linux bubblewrap — Landlock has a type reserved for it but was never wired up, and Windows does not even have a helper; that bill is tallied in detail in Article 20); plus a side-git snapshot every turn and concurrent sub-agents on call.

But this series will not march down a feature checklist. What I want to tell is the ledger behind the design decisions — why each cut was made where it was, what it cost, and which ones I have regretted.

Why Me

A word about the lineage — it matters for understanding the project.

CodeSmith was not written from scratch. Its predecessor is CodeWhale (earlier known as deepseek-tui).

And the mindset behind this series is not "watch me reinvent a wheel," but rather: having read a great many open-source coding agents and agents with long-horizon task capability, here are the things I wanted to do — and did.

The engineering scale, laid out here up front — later installments will keep returning to this coordinate system:

  • 21 crates in the Cargo workspace: agent-runtime (engine core), tui (interactive runtime), agent / providers (LLM abstraction and adapters for 16 model services), execpolicy (command security), index (tree-sitter symbol index), mcp, hooks, extensions…
  • 548 Rust source files, 356,193 lines of code (find+wc counting, including comments and inline tests)
  • 5,429 test functions — roughly one test per 66 lines of code, including end-to-end QA that runs a real TUI inside a sealed pseudo-terminal

Build something like this in Rust, and one frequently asked question is why not Python/TypeScript. The answer will surface naturally across the later installments; here is the one-line version: when you need a cross-chunk state machine inside streaming output, fat pointers crossing an extern "C" boundary, and a byte-level fingerprint of the system prompt before every request, you want a type system standing guard over those boundaries.

The Hidden Thread: A Catalog of Distrust

One hidden thread runs through the entire series — perceptive readers may already smell it: these beams are all the same thing — distrust. Distrust that the model will use the tool channel honestly (hence the streaming strip-out); distrust that the model can manage its own context (hence handles and RLM); distrust that the model finishes work without leaving garbage behind (hence the ledger and snapshots); even distrust of your future self evaluating model changes (hence a snapshot every turn). A good harness is, in essence, a clearly written catalog of distrust.

Top comments (0)