DEV Community

DogeKing
DogeKing

Posted on

CodeSmith DIY: Assemble Your Own Coding Agent Without Writing Code

DIY: Assemble Your Own Agent Without Writing a Line of Rust

Source version of CodeSmith: v0.5.0 (commit 3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: readers who want to remake a coding agent into a tool of their own — without learning a new language for the privilege.

The prologue and Article 0 were both about "what we built." This installment strikes while the iron is hot and flips to the buyer's side of the question: what can you change?

First, a quiz: how many steps does it take to turn CodeSmith into a "release engineering specialist" — one that handles version releases, changelog compliance checks, and release-note writing, and touches nothing else — without changing a single line of Rust?

This article's answer: one TOML file and two Markdown files. But we have to start from the foundation, which is the system prompt.

1. The Three Tiers of Prompt Customization

Article 4 will dissect that 297-line Constitution in detail; for now, all you need to know is that it is not some isolated blob of text but the first of five deterministically assembled layers (crates/agent-runtime/src/prompts.rs:873-887):

let parts: [&str; 5] = [
    tool_taxonomy.as_str(),                  // 1. tool taxonomy
    base_prompt.as_str(),                    // 2. the Constitution ({model_id} already substituted)
    personality.prompt().trim(),             // 3. personality
    mode_prompt(mode).trim(),                // 4. mode delta (Plan/Agent/YOLO)
    approval_prompt_for_mode(mode, approval_mode).trim(),  // 5. approval policy
];
Enter fullscreen mode Exit fullscreen mode

The order is fixed; each layer has a single responsibility. And the customization entry points open to users fall, rather conveniently, into three tiers by depth of intrusion (config.example.toml:125-158).

Tier one: change the tone.

# personality = "calm"        # default | "playful"
Enter fullscreen mode Exit fullscreen mode

In the Constitution's hierarchy table, the personality layer sits at level 8 — Presentation Only. It governs "how things are said" (tone, rhythm, opening lines) and can never touch "what gets done." Giving the model a new personality is like giving a lawyer a new tie: his standing to appear in court is unaffected.

Tier two: append your rules.

# instructions = ["./AGENTS.md", "~/.codesmith/global.md"]
# append_system_prompt = ["~/.codesmith/extra-prompt.md"]
Enter fullscreen mode Exit fullscreen mode

instructions splices your rule documents into the system prompt in declaration order (level 5 in the hierarchy table, Local Law — project rules outrank memories and are outranked by mode rules); append_system_prompt appends after all five layers have been assembled. The official advice in the example config could not be clearer: prefer this tier (config.example.toml:157-158: "Prefer instructions + append_system_prompt"). One detail worth noting: instructions in a project-level config is replaced wholesale, not merged — the repo has the final say, so home and work never cross-contaminate.

Tier three: tear the whole persona down and start over.

# system_prompt = "You are a release engineering assistant..."
# system_prompt_file = "~/.codesmith/prompt.md"
Enter fullscreen mode Exit fullscreen mode

Full override: the Constitution, the tool taxonomy, and the mode layers are all discarded; you start from zero. This is the entrance marked "I want to build an entirely new species," and it is also the nuclear button that the official docs explicitly label "use only when you truly need it."

An oft-overlooked companion to the three tiers: the /system command (crates/tui/src/commands/mod.rs:420, alias /xitong). Type it once, and the final assembled system prompt is visible in full — whether each of your customizations actually took effect, what it rendered into, which layer it landed in — what you see is what you get. The debugging experience of a customization system is half the life of that system, and on this point many big-vendor products are less candid than this open-source project.

So what makes it safe to let users override all the way to tier three? Because the harness's hard lines of defense are not in the prompt at all. The approval gates, the sandbox, the return of tool results verbatim — these are engine mechanisms; replacing the system prompt cannot move them by a single letter — you can dress the Agent in any personality you like, and it will still have to pass your gate to run a dangerous command, and a tool's error messages will still land right back in its face, unedited. The Constitution's provisions (Article II's truth-above-all, for instance) are indeed prompt-layer norms, and a tier-three override replaces them along with the entire Constitution file; but the engine guarantees that however the prompt is written, the flow of evidence is not governed by the prompt. You can customize the Agent's personality and knowledge. Fooling you is not on the menu.

2. The Behavior Surface: Precision-Guided Whitelisting

The prompt governs "how it thinks"; the whitelist governs "what it may do." The example configuration for auto_allow (config.example.toml:254-262):

# auto_allow = ["git status"]   # auto-approves: git status, git status -s, git status --porcelain
#                               # does NOT auto-approve: git push, git checkout
Enter fullscreen mode Exit fullscreen mode

Let git status -s through, hold git push back — behind that comment sits an exact-match algorithm: the bash arity dictionary (crates/execpolicy/src/bash_arity.rs).

The problem is harder than it looks. Naive matching fails at both ends: under whole-string equality, git status --porcelain fails to match the whitelisted git status — the harmless gets blocked, which is annoying; relax to a string prefix, and git stash hitches a free ride on the git sta prefix — what you meant to block slips away. Worse still is npm run: under pure prefix matching, any script appended after it gets waved through, including the harmless-named ones that do nasty things. CodeSmith's solution is to define a "command prefix" as a fixed number of positional arguments (bash_arity.rs:7-11):

Flags (tokens starting with -) are never counted toward arity. auto_allow = ["git status"] must match git status -s and git status --porcelain, but not git push.

arity = the number of positional-argument words, command itself included, that constitute the canonical prefix; flags never count. The static table holds entries for thirty-odd common tools (bash_arity.rs:36-43): ("git status", 2) — git + status, two words form the prefix, and whatever flags follow hang on freely; ("npm run", 3) — npm run <script>, three words form the prefix, the particular script being the variable; ("make", 1) — the bare command is the whole of it.

The rule layering is just as deliberate: BuiltinDefault < Agent < User, and at equal priority the longest prefix wins. The cautious built-in defaults get overridden tier by tier by progressively better-informed parties — which is itself a miniature model of this system's governance philosophy.

3. The Capability Surface: Skills, Commands, and Hooks

Skills are the workhorse of capability injection. One directory, one SKILL.md: YAML frontmatter for the metadata (name, description, optional when_to_use / allowed-tools / model…), the body for the knowledge. The core mechanism is progressive disclosure (docs/SKILLS.md:10) — the system prompt carries only the skill catalog (name + description + path, capped at 12,000 characters in total), and when the model matches a task it calls load_skill to pull the full text. Skills will not blow up your context, just as Article 8's handles will not blow up your context — the same virtual-memory philosophy.

"What shape should a capability take?" The textbook draws this as a spectrum. At one end stands the dedicated tool: a structured function call, schema-constrained parameters — highly deterministic, testable, with fine-grained permissions, at the price of several hundred tokens of resident definition each. At the other end stands the Skill: an operating procedure written in natural language, which the Agent follows using bash and a code interpreter; one catalog entry costs only a few dozen tokens, and the body is read only when needed — "deploy the application," say, need not be a deploy_app tool; a document reading npm run build → docker build → kubectl apply will do.

Between generic and dedicated stands the general-purpose executor: one code interpreter stocked with the usual libraries can replace dozens of dedicated tools and absorb the scenarios nobody planned for — Article 6's private RLM kitchen is precisely this species, and as a bonus it cashes in the general-purpose executor's most underrated dividend: let code orchestrate the tool calls, keep the intermediate variables in the execution environment, and return only the final result to the context, and token consumption drops by roughly two orders of magnitude. The default orientation is general-purpose over dedicated unless there is an explicit reason otherwise; and the four reasons to "retreat to dedicated" are worth committing to memory:

  1. Security and audit (writes to a production database need the permission granularity of a dedicated tool; an open executor cannot provide it)
  2. Shielding platform differences (grep/find could plainly be written in bash, yet every vendor's Agent still ships them as dedicated tools — for uniform line-number feedback and cross-platform arguments)
  3. High-frequency operations deserve a dedicated entry point
  4. When parameters are complex, a schema is more reliable than natural language (a Skill makes the model generate a legal command line and handle its own escaping, which is more error-prone than passing parameters as JSON).

Conversely, Skills are friendlier to the human author: editing a paragraph of text is far lighter than changing code, testing, and redeploying, and a local blunder will not keep the whole Agent from standing up the way a missing brace in a schema will. The four decision dimensions compress into one sentence: safety is judged by permissions, parameters by complexity, change by frequency, and the model by capability — the weaker the model, the more it needs a schema to hold its hand.

The skill-catalog discovery order spans ten locations, several of which belong to other tools: .claude/skills, .cursor/skills, .opencode/skills. The skills you maintain for Claude Code work as-is. Community skills install with a single command: /skill install github:owner/repo.

Custom slash commands are the lighter injection: drop a .md file into .codesmith/commands/; the filename becomes the command name, the body becomes the message sent when you press Enter, and the frontmatter may carry description and allowed-tools. .claude/commands/ works here too.

Hooks are the deep water — the part that "rewrites system behavior." Of the 14 lifecycle events, the sharpest are the three mutable ones. Here is a real recipe (docs/HOOKS.md:400-406):

[[hooks.hooks]]
name = "aws-creds"
event = "shell_env"
command = "aws-vault export my-profile --format=env"
Enter fullscreen mode Exit fullscreen mode

shell_env runs synchronously before every single exec_shell; its stdout is parsed as KEY=VALUE and merged into the subprocess environment. The effect: every time the Agent runs an aws command, what it holds is a freshly minted temporary credential from aws-vault — long-lived secrets never appear in any config file or conversation. The kind of elegance that would get a nod of approval from a security team.

Another recipe is just as practical: the pre_compact event runs before context compaction and splices the contents of DECISIONS.md into the compaction summary — key decisions survive /compact (HOOKS.md:408-415).

Tool replacement is the most thorough move of all (config.example.toml:1028-1039):

[tools.overrides]
"exec_shell" = { type = "script", path = "audit-exec-shell.sh" }   # audit wrapper
"read_file"  = { type = "command", command = "bat", args = ["--paging=never"] }
"fetch_url"  = { type = "script", path = "cached-fetch.sh", args = ["--ttl", "300"] }
Enter fullscreen mode Exit fullscreen mode

A built-in tool can be replaced wholesale by a shell script — exec_shell swapped for an audit-logging wrapper (the config file even ships a complete example that writes ~/.codesmith/audit/exec_shell.log), read_file swapped for bat. The protocol is as simple as a single agreement: JSON in on stdin, {"content": ..., "success": ...} out on stdout. Without writing Rust, you can still perform amputation-and-replacement surgery on this Agent.

4. The Orchestration Surface: Assemble Your Own Crew

Article 6 covered how every member of a Team can carry its own prompt, model, and worktree_path (crates/agent-runtime/src/team/team_file.rs:65-85). Now those are your building blocks: assemble a "release engineering crew" — an Implementer that bumps the version numbers in its own worktree, a Review that audits the changelog read-only, a Verifier that runs the pre-release checklist — three kinds of worker, three prompts, three price points (and remember to configure [subagents.models]: Flash for scouting, Pro for review).

Multi-environment switching goes to profiles (config.example.toml:823-835): [profiles.work] points at the company gateway, [profiles.dev] loosens allow_shell, [profiles.nvidia-nim] swaps out the whole provider — codesmith --profile work flips between them with one keystroke, and if you mistype a name it will even list the available ones for you.

For those who genuinely want to write Rust, the escape hatch is still there: dylib extensions — implement the Extension trait, register tools, commands, and 23 kinds of event handlers, and use HandlerOutcome::{Continue, Cancel, Block, Transform} to rewrite behavior in the crevices of input, requests, tool calls, compaction, and so on; /extension install git:<repo> hot-loads them. But honestly, this is for that 1% of scenarios — the capabilities of the previous four sections, put together, are already enough to assemble the great majority of personas you could imagine.

Finale: Assembling That "Release Engineer"

Back to the quiz we opened with. The recipe:

# ~/.codesmith/config.toml (or the repo's .codesmith/config.toml)
system_prompt_file = "~/.codesmith/release-engineer.md"   # Tier 3: a brand-new persona
auto_allow = ["git status", "git log", "git tag", "cargo check"]
# "cargo release" still requires approval — shipping is a big deal; keep a gate on it

[[hooks.hooks]]
name = "gh-creds"
event = "shell_env"
command = "gh auth token | sed 's/^/GH_TOKEN=/'"
Enter fullscreen mode Exit fullscreen mode

Then add two Markdown files: .codesmith/commands/release-notes.md — the prompt template that lets /release-notes generate release notes in one keystroke; and .codesmith/skills/changelog-lint/SKILL.md — give it a description reading "Use when editing CHANGELOG.md or preparing a release," and the model will load your compliance rules automatically whenever it so much as touches a changelog.

And just like that, off the assembly line rolls a dedicated Agent that knows nothing but releases, holds temporary credentials, carries its own compliance skills, and defers to you on any operation outside its command whitelist. Nothing in the process was compiled, nothing was forked, and nothing will be lost when you upgrade CodeSmith — because your customizations live entirely in configuration and documents, on the engine's stable interfaces.

Coda: The Two Faces of a Harness

This installment's four sections — prompts, whitelists, capability injection, orchestration — are two faces of the same thing.

The face turned toward the model is constraint: byte-level prefix discipline (Article 4), pinned sub-models (Article 6), a film negative of every turn (Article 7), bounded context exits (Article 8), verdicts rendered as whitelists (Article 13) — a checklist of distrust aimed at the model, every line of it defended.

The face turned toward the user is freedom: the three-tier prompt path, whitelists precise down to the arity, a replaceable tool surface, an orchestratable crew — total sovereignty, written in a config file.

The merit of a harness comes down, in the end, to the ratio between these two faces: to whatever degree the model is constrained, an equal measure of freedom is handed to the user. Article III of the Constitution — "the user is the sovereign of this session" (prompts/base.md:23-27) — finds its finished construction surface in every section of this installment.

But before you can watch these constraints being built with your own eyes, a more basic question has to be answered first: what, exactly, is an Agent, and where are the boundaries of its capability drawn? The next class is a textbook lesson — the anatomy of an Agent and five waves of evolution, the coordinate system for every installment the series has left.

Top comments (0)