DEV Community

DogeKing
DogeKing

Posted on AI-assisted

CodeSmith: Out of the Box, from Pi-Style Minimalism to a Full Coding Agent Suite

Out of the Box: One Key, from Pi-Style Minimalism to the Full Suite

Source version of CodeSmith: v0.5.0 (commit 3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: readers who have used coding agents such as pi or Claude Code, toyed with the idea of "reshaping one into your own Agent," but would rather not learn a whole new framework for the purpose.

CodeSmith v0.5.0 compresses "configuring an Agent" into a single key. One key — preset — and 13 resource switches, 11 feature flags, the tool surface, and the memory setting all snap into place at once: the simple tier runs 9 tools with zero background; the experiment tier throws everything open.

We put it at the head of this series because it produces no new "features" whatsoever, yet buys three things in one purchase: pi's minimalism, Claude Code's out-of-the-box readiness, and a DIY that belongs to you alone.

pi Is Hot, but a Bare Shell Is Uninhabitable

The hottest coding agent this year is, in all likelihood, pi by Mario Zechner (badlogic). Its philosophy is written on the storefront: other frameworks hand you a bigger box of parts; pi hands you a smaller one and leaves the rest for you to build. No forking, no source hacking — you grow it into what you want with extensions and skills.

I admire that philosophy — CodeSmith's extension system is itself a port of the pi-mono Extension model (docs/EXTENSIONS.md). But the very heat of pi's popularity lights up an industry truth: frameworks deliver a bare-shell unit, while what most people want is water, power, and broadband all live on move-in day.

DIY carries real costs. You must write TypeScript and read your way through its tool protocol; which tools to cut and which to keep, every one of those calls is yours to make alone; everyone on the team builds their own rig, then maintains their own copy of it through every version bump. None of these costs appear in the framework's README — they are all passed on to the user.

A saying makes the industry rounds: out-of-the-box readiness and customizability sit at the two ends of a spectrum — products pick one end, frameworks pick the other, and you cannot have your cake and eat it too. We disagree. The two are not born opposites; the configuration layer has simply never had a "baseline" semantics — "who has the final say" was never defined clearly. The answer v0.5.0 gives is called preset.

The Trouble with Configuration Overload

In a coding agent like CodeSmith that supports free-form DIY, the configuration keys proliferate without end. Imagine handing a new user a fully annotated, 1068-line config.example.toml: how do you get them "a working Agent in one keypress"?

The dumbest approach: ship no presets, expose everything, and let the README say, in all honesty, "please first read the one hundred configuration keys." Freedom in pi's eyes; an undecipherable scripture in a newcomer's.

First refinement: ship one recommended config, applied with a single key. But under what semantics? With override, the explicit config the user already wrote gets stomped on; without override, when the two sides fight, whose word wins? And this scheme promptly runs into a second bill left over from the old modes system.

Second refinement: the recommended config only fills blanks — keys the user never wrote get filled in; keys the user did write always count; and where the two sides disagree, it does not lie, but truthfully marks "deviated."

Four Tiers: A Table That May Not Be Reordered

The first impression a preset makes is four built-in tiers, each a step deeper than the last (crates/config/src/presets.rs:58), with middle as the factory default:

Tier Tools Memory Subagents Resource switches
simple whitelist of 9 goldfish (none) 0 index/LSP/snapshots/memory/update checks/audit all off
middle (default) inherited inherited 10 high-value, low-cost combinations on; experimental off
all full set notebook 20 middle + LSP warnings; preview flags still off
experiment full set elephant + KOD 20 everything on: vision, teams, the coordinator, the context manager, etc.

The true face of the simple tier is nothing more than a 49-line TOML (crates/config/src/presets/simple.toml), compiled into the binary (the include_str! at presets.rs:43-47):

name = "simple"
description = "Pi-style minimal: 9 core tools, thinking medium, no index/LSP/memory/snapshots/MCP/search."
app_mode = "agent"
reasoning_effort = "medium"
memory_level = "goldfish"
max_subagents = 0
index_enabled = false
lsp_enabled = false
#...... all remaining resource switches are false

[tools]
include = [
  "read_file", "write_file", "edit_file", "grep_files",
  "list_dir", "file_search", "exec_shell", "exec_shell_wait",
  "update_plan",
]
Enter fullscreen mode Exit fullscreen mode

Pi's minimalist posture — a small tool surface, medium thinking, no cross-session memory, zero background actions — is written out verbatim as data. The include list names only 9 tools, while the engine keeps 27 native tools on its books by default (the mirror list at presets.rs:710; the full catalog is 50+ — see Article 16).

More interesting still: this table may not fall out of order. A test at presets.rs:625 nails the entire four-tier switch matrix down:

#[test]
fn governed_switch_matrix_is_monotonic() {
    // Every governed key must be non-decreasing along
    // simple → middle → all → experiment: once a tier turns a switch
    // on, deeper tiers may not turn it off.
    for (key, row) in builtin_tier_matrix() {
        let mut seen_on = false;
        for value in row {
            let value = value.unwrap_or(true);
            if seen_on {
                assert!(value, "tier matrix regresses on {key}: {row:?}");
            }
            seen_on |= value;
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

In plain words: the 13 governed boolean switches (the bool_dials at presets.rs:255-271) may only be turned on, one by one, along simple → middle → all → experiment, and never off. A companion test checks the full spelling (presets.rs:671): every tier must write out every governed key explicitly — leave one unwritten, and the decision surface for diy quietly widens.

The predictability of the tiers is not a promise made in documentation; it is nailed down by tests. Moving up from simple to middle, your only mental burden is "there is more stuff now" — not a single familiar switch will suddenly vanish.

Why must the tiers be monotonic? Because a tier is, at bottom, a gradient of trust: simple is "I want none of it," experiment is "I want it all." If a middle tier turned one switch on here and another off there, every upgrade would force users to re-verify all 13 switches — and we would have degenerated right back into that 1068-line config file.

Baseline, Not Override: fill-if-unset Has Only Three Branches

The heart of the whole system is three helper functions cast from the same mold (crates/tui/src/presets.rs:252):

fn fill_bool(slot: &mut Option<bool>, value: Option<bool>, deviated: &mut bool) {
    match (&mut *slot, value) {
        (None, Some(value)) => *slot = Some(value),   //1. Slot empty: fill it in
        (Some(existing), Some(value)) if *existing != value => *deviated = true,  //2. Has a value and differs: book it
        _ => {}                                       //3. Has a value and matches: nothing happens
    }
}
Enter fullscreen mode Exit fullscreen mode

A preset value fills only empty slots; when a slot already has an owner and the preset disagrees, it does not overwrite — it only makes a bookkeeping entry (see the comment markers 1 and 2). fill_string and fill_usize are the string and integer editions of the same logic — that familiar recipe again.

A preset is a baseline, not an override: explicit configuration always wins.

apply_to_config (presets.rs:282-411) runs this three-branch pattern over the 13 switches, 11 feature flags, and 5 scalar dials, one by one, and its return value is a single boolean: whether anything deviated. Those 130 lines are really the same pattern copied and pasted over and over — they are doing exactly one thing: asking every configuration key, in turn, "has the user spoken?"

The startup path is therefore short (presets.rs:429-484); one resolution pass looks like this:

Preset resolution at startup (crates/tui/src/presets.rs)
|-- Tier selection: --preset CLI > preset key in config.toml > last choice in settings.toml > middle
|-- PresetCatalog::load
|   |-- Five built-in tiers (simple/middle/all/experiment/plan, compiled into the binary)
|   |-- ~/.codesmith/presets/*.toml        # your global presets
|   |-- <workspace>/.codesmith/presets/    # project presets; recommended to commit to the repo
|   `-- legacy modes/ directory still scanned, with a one-time migration warning
|-- apply_to_config: fill-if-unset fills blanks + deviation bookkeeping
`-- config.preset_deviated → effective label (tier name, or diy)
Enter fullscreen mode Exit fullscreen mode

The plan in that tree is the fifth built-in tier, dedicated to Plan mode — read-only reconnaissance plus planning tools, with write operations passing through approval (presets/plan.toml is all of 6 lines) — and it does not take part in the four-tier monotonic matrix.

A detail easy to miss: hot-switching and startup are two different semantics — on purpose. Startup folding uses fill-if-unset — you did not name a preset, so it only qualifies as a baseline. Mid-session /preset <name> uses override (the apply at presets.rs:85) — you just named one with your own mouth, so it takes full effect immediately, while config-class switches truthfully report "effective after restart." The same question, asked twice, answered differently — because it was asked differently.

diy: The Tier You Are Not Allowed to Choose

The outlet of all that deviation bookkeeping is a tier you cannot choose. First, watch it turn you down (crates/config/src/presets.rs:72):

if lowered == "diy" {
    bail!(
        "'diy' is a derived state, not a preset: it is shown when your \
         explicit config deviates from the selected tier. Pick simple | \
         middle | all | experiment (or a custom preset name)."
    );
}
Enter fullscreen mode Exit fullscreen mode

Then watch it get computed (crates/tui/src/config.rs:2975):

pub fn effective_preset(&self) -> &str {
    if self.preset_deviated {
        "diy"
    } else {
        self.preset_selection().unwrap_or("middle")
    }
}
Enter fullscreen mode Exit fullscreen mode

diy is not a configuration key; it is a computed state. The moment your explicit config differs from the selected tier in even one place, the status bar stops showing a tier name and truthfully writes diy. Want to pin it down by writing preset = "diy" into config.toml? The validator refuses that too (config.rs:1812).

This is my favorite stroke in the whole design. The value of a label is trust at zero cost: when you see middle, you know all 13 switches sit exactly where middle puts them — not one astray. Conversely, if diy were allowed to be saved as a tier, the status bar would be licensed to lie — persisting a derived state amounts to archiving a lie.

Wouldn't override be simpler? It would — but override hides "who has the final say" inside the runtime, and every read becomes a verdict rendered on the spot. fill-if-unset moves that reckoning up to load time and settles it once and for all, which earns a complimentary bonus: a preset can be taken off safely at any moment. After /preset off, your explicit config survives intact down to the last hair; a value that was stomped, by contrast, can only be recovered from memory.

DIY Without Writing Code: TOML Is Your Wrench

Back to the comparison with pi: pi's entrance to DIY is writing an extension; CodeSmith's is a TOML file dropped into one of two scanned directories — ~/.codesmith/presets/ (global) or <workspace>/.codesmith/presets/ (project). Later layers override the built-ins by name: drop a project file named simple.toml, and you have rewritten the definition of "minimal" for the whole team (presets.rs:401-427).

A bad file cannot crash startup: the directory loader skips files that fail to parse, collects them into warnings, and reports them faithfully in /preset list (presets.rs:449-505). The legacy modes/ directory is still scanned and gets only a migration hint — the rename does not cut off the old road.

The third entrance is export. Once you have dialed in a set of knobs, /preset export my-setup snapshots the current tool surface, thinking tier, approval posture, and memory level into TOML (crates/tui/src/presets.rs:583-640) and writes it into the project presets directory. Want to overwrite a built-in tier name? You must write /preset export! simple and force it explicitly — yet another gate against friendly fire.

And so "compatible with pi's usage" acquires three layers of meaning. Want the minimalist posture: --preset simple, 9 tools and zero background — the docs call it, verbatim, Pi-style minimal. Want programmability: the extension system is a port of the pi-mono model in the first place. Want the middle ground: TOML presets plus export, without writing a line of code.

The honest boundary must also be confessed: CodeSmith does not yet have pi's in-file tree browser, and docs/PRESETS_cn.md says outright that it is a larger UI/data-model project. What is compatible here is the usage — no stand-in is being impersonated.

One last easily misread design point: the built-in tiers never govern safety keys. yolo, approval_policy, sandbox_mode, telemetry, prompts and personality, provider credentials — the four tiers touch not a single one of them. The schema does keep approval_policy and sandbox_mode fields — those are for your own custom presets. A tier may decide how much capability you get; it never decides how much risk you take.

Coda: Turning pi's Multiple-Choice into Fill-in-the-Blank

String the mechanism into one chain:

  1. One key selects the tier, with precedence --preset > config.toml > the last session's choice > factory middle;
  2. The tier is a TOML baseline compiled into the binary; the four built-in tiers advance monotonically, nailed down by tests;
  3. At startup it is folded into the config with fill-if-unset, and explicit keys always win;
  4. The moment anything deviates, the effective label truthfully becomes diy — and diy may be neither chosen nor persisted;
  5. Customization and export go through TOML files in the two scanned directories — DIY with zero code.

Distill further, and you get three features, each with its price:

  1. Baseline, not override: a preset can be taken off safely at any time; the price is that a double-write — "the preset says one thing, the config says another" — must be read with fill-in semantics, not taken for granted.
  2. Monotonic tiers: upgrades hold zero surprises; the price is that inverted combinations (say, "every tool on, zero memory") can only be expressed through a custom preset.
  3. diy as a derived state: the label never lies; the price is that it cannot be pinned down, so when a team wants to share one deviation posture, it has to settle into a named preset — which is exactly the job /preset export does.

Series Roadmap

This series opens with two introductory lectures, then runs a six-act main line and closes on the evolution finale, with several Textbook lectures threaded in between. The whole series is aligned against v0.5.0 (commit 3a74c82f); line numbers and facts refer to that version. The incremental commits that landed between v0.3.0→v0.5.0 (layered presets, three-zone wiring, the write-to-disk syntax gate, the security settlement, streaming termination, claim verification, the evolution closed loop) are unpacked by theme in the respective installments.

Opening: getting started and customization (Articles 00-02). The first lecture of the opening (this very article) is about v0.5.0's layered presets — one key, from Pi-style minimalism to the full suite: the simple tier runs 9 tools with zero background, the experiment tier throws everything open, the baseline never overrides, and deviations are truthfully reported as diy.

Article 01 is the prologue: it starts from a stretch of streaming-filter code and a war over "open-source models faking tool calls," asks why the word harness fits better than "framework," and covers this project's pedigree and tonnage.

Article 02 covers the customizable surface: from a single line of personality = "playful" to swapping out the entire system prompt; from auto_allow whitelist exact matching (waving git status -s through, stopping git push) to replacing built-in tools with shell scripts. The Constitution's nine-level hierarchy guarantees that however far you go, "truthfulness" and "user sovereignty" remain untouchable — you can customize everything about the Agent, except making it lie to you.

Act One: An Architecture Born to Save Money (Articles 04-05). The greatest virtue of open-source models is that they are cheap — but cheapness comes with terms: a prefix cache hits only when the request's byte prefix matches exactly. Article 04 covers how the three-zone model bets the product's economics on that rule: the Constitution dares to run 297 lines precisely because, once the cache hits, "the per-turn cost drops roughly 100×."

Article 05 covers the v0.5.0 deed: the three-zone contract is wired into the engine's request path, and the conversation history's append-only-ness is promoted from discipline to a compile-time property of the type system — the methods that would rewrite history no longer compile.

Act Two: Signature Designs (Articles 06-07). Article 06 covers RLM: the Agent opens a resident private Python kitchen for itself, turning "can't read it all" into "can compute it all"; along the way come ten kinds of subagents and Team swarm orchestration.

Article 07 covers the two designs I am proudest of and quickest to mock myself over: the slop ledger, which turns the mess AI leaves behind into a persistent ledger; and side-git, which snapshots the workspace every turn, so that "the AI botched it" becomes an event you can roll back precisely.

Act Three: Context Engineering (Articles 08-11). Governing the model's field of view: fitting the model with "virtual memory" — big outputs become handles, paged on demand (08); a conversation history that never gets rewritten — cache friendliness is a typographic art (09); the dashboard hung at the end of the conversation — the truth about the user role hanging at the tail (10); six lines of defense after the window fills up (11).

Act Four: Meta-Ability (Articles 12-15). A probabilistic brain cannot compute deterministic problems, so let code think for the model — RLM's prompt writes "the LLM translates, the interpreter computes" into a contract (12). A rule written in natural language might as well not be written, so business rules get encoded: τ-bench's cancellation policy, execpolicy's three-valued verdicts, an 189-entry arity dictionary (13). The third ring is the Agent itself: doctor's deterministic boundary, example-based generation, genome replication with mutation for bootstrapping (14). The finale is v0.5.0's claim verifier: "tests all green" only counts after a replay, with four-tuple verdicts landing in a jsonl ledger (15).

Act Five: Tool Internals (Articles 16-18). Fifty weapons, one slot (16); house rules written for the model — the manual teaches, error messages correct, the engine intercepts (17); and 18 covers v0.5.0's write-to-disk syntax gate: judge the delta, not the state — a good file may not be written bad, a broken workspace still gets repaired, and the four disk-writing tools share a single gate.

Act Six: Injection Defense and Loop Robustness (Articles 19-22). Article 19 covers the six walls against prompt injection: the Constitution's hierarchy, Unicode disinfectant, the permission ratchet, static command analysis, the network boundary, and output-side anti-forgery. Article 20 covers v0.5.0's seven notes on settling the security debt: honest sandbox reporting, fail-closed, splitting compound commands, path-traversal guards, trust boundaries, poisoned-lock tolerance, atomic writes and release-chain pinning.

Article 21 covers loop robustness: a four-layer failure taxonomy, silent hangs and the adaptive watchdog, a graduated ladder of recovery with circuit breakers, and cross-vendor takeover. Article 22 covers v0.5.0's streaming cliff edge: when output is beheaded at the token limit, the tool call refuses to repair and refuses to execute — a severed stream no longer masquerades as a normal finish.

Finale: The Evolution Closed Loop (Article 23). First it sets the Textbook coordinates — three-layer verifiers, four kinds of update carriers, dual loops and three safety boundaries, the hardest of which is "safety mechanisms may not modify themselves"; then it looks at CodeSmith's implementation — the claim verifier reconciling the books online, doctor's LLM fallback, memory's sleep learning. Determinism as the foundation, proposals passing through verifiers — the loop closes here.

The Textbook lectures threaded in between fall into three groups: Article 03 sets up the coordinate system before any incisions — the anatomy of an Agent (the loop, autonomy, and the cold water thrown on multi-Agent) plus five waves of engineering evolution; Articles 21 and 22 teach robustness; Article 23 gathers the whole series in. They do not track any single module; they are the common knowledge on which every module stands.

The next installment (01) is the prologue — for whom this harness was built, and where to begin writing that inventory of distrust.

Top comments (0)