Context Engineering: A Conversation History That Never Gets Rewritten
Source version of CodeSmith:
v0.5.0(commit3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: readers who know the KV cache exists, but have never stopped to think that "every byte of the conversation history is money."
Article 4 laid down an iron law: many large models grant a cache discount only when the request's byte prefix is exactly identical — drift the system prompt by a single byte, and every token after it pays full price. That installment was the legislation — the three-zone model, the SHA-256 fingerprint, the drift autopsy. The previous installment was about "don't stuff it in": large outputs become handles, paged in on demand. This one is the enforcement: how that law lands in every line of code — and in v0.5.0, it has grown from a discipline into a compile-time property of the type system (Article 5).
Start with the enforcer's motto. Tucked inside the engine is an unassuming method whose doc comment reads like a court verdict (crates/agent-runtime/src/engine/mod.rs:219-233):
/// Return the session transcript verbatim for a request snapshot.
///
/// The transcript is the session's `AppendLog` (#2264): this snapshot is
/// a plain `to_vec()` of its slice. `<turn_meta>` is stored on
/// user-text messages when the message is appended. Do not rewrite
/// historical messages at request time: doing so makes the API prefix
/// differ from the bytes sent in earlier turns and destroys DeepSeek's
/// KV prefix cache reuse.
// ... (a passage of module-relocation remarks omitted) ...
pub fn messages_with_turn_metadata(&self) -> Vec<Message> {
self.session.messages.to_vec()
}
In plain terms: rewriting historical messages at request time is forbidden — it would make this request's prefix disagree with the bytes actually sent in earlier turns and destroy KV prefix cache reuse. The method body is a single line, to_vec(): a verbatim snapshot, shipped as-is. And mind the comment's new wording: the transcript is now the session's AppendLog — the append-zone type from the three-zone contract — and the snapshot is nothing more than a to_vec() of its slice.
What this whole installment examines is how much "invisible discipline" the engine took on to defend the invariant behind that one line of to_vec().
The Anatomy of a Request
The conclusion first. The request sent to DeepSeek on every step ultimately looks like this (assembly order at prompts.rs:1027-1258, request snapshot at engine/host_executor.rs:2972-2994):
system (assembled in PromptBundle layers, ordered from most to least stable)
├─ locale_preamble [session] localization opener
├─ global_system_prefix [static] tool taxonomy → Constitution → persona → mode → approvals → …
├─ project_context [workspace] AGENTS.md four-tier merge
├─ project_context_pack [workspace] sorted, bounded JSON repo overview
├─ skills [workspace] skill catalog
├─ environment [session] language/platform/shell/pwd
└─ … (configured / memory / session_goal / authority_recap / locale_closer)
tools (sorted by name; built-in tools contiguous and first, MCP after)
messages (conversation history, append-only)
├─ user content[0] = <turn_meta> (date/routing/working-set summary), content[1] = user text
├─ assistant thinking + tool_use
├─ user tool_result
├─ …
└─ user the current turn's new message — the only new content in each request
Article 4 covered the four tiers of PromptSectionStability (Static/Workspace/Session/Dynamic) and dynamic_boundary_index. Here is one detail it never unpacked at the time: PromptBundle::push sets down the boundary at the first section that is neither Static/Workspace nor Cacheable (prompt_runtime.rs:145-179) — wherever the first volatile section appears is precisely where the "reusable prefix" ends. The whole layout logic of the system prompt is one sentence: stability in descending order — anything volatile gets herded to the back.
This art of layout is not CodeSmith's invention — Claude Code has pushed it even further, and three things there are worth setting side by side:
The system prompt is physically cut in two by a cache-breakpoint marker: everything before the marker is cached globally across users and across sessions, and only after it does user- and session-specific information appear. Every dynamic condition (operating system, mode, language) placed before the boundary doubles the variants of the cache key — N binary conditions make 2^N combinations — so the ordering of the prompt is decided first by cache economics and only second by semantic logic.
Sub-agents that inherit the parent context must stay byte-aligned with the parent — prompt, tool definitions, and message prefix identical byte for byte — to hit the same Prompt Cache. CodeSmith's sub-agents take the other road (each kind of worker gets its own prompt and tool surface, inheriting no parent prefix), and each choice keeps its own ledger: alignment saves money; independence saves worry.
A tool result's replacement string freezes on first appearance — once a large output is swapped for a summary preview, the replacement string is persisted and reused verbatim even across session restarts, so the restored message sequence stays consistent with the cached byte stream. This is isomorphic to CodeSmith's
<TOOL_RESULT_REF sha="...">: the original goes to disk keyed by its hash, and the second appearance of the same large result is reduced to a one-line reference.
Deterministic Assembly: Driving Randomness Out of the system prompt
A stable prefix presupposes an assembly process that is itself reproducible. In this codebase, "determinism" is not a slogan; it is a list of concrete military rules.
Rule 1: the builder does no I/O. The doc comment on the prompt-building function opens with the thesis (prompts.rs:32-37): no disk I/O happens inside the prompt builder — with no disk reads in the builder, the workspace-static part of the system prompt can be cache-friendly. Every file read happens outside the builder, at session start.
Rule 2: whatever enters the pack must be sorted, bounded, and timeless. The doc comment on the function that generates the repo-overview JSON pack (project_context.rs:155-205):
/// Generate a deterministic, cache-friendly project context pack.
///
/// The pack intentionally uses only stable workspace facts: relative paths,
/// sorted entries, bounded README text, and sorted JSON object fields. It does
/// not include timestamps, random ids, absolute temp paths, or live git state.
Even a seemingly harmless fact like the current git branch is excluded — the branch flips, the whole pack flips, and every byte after the pack pays full price. readdir order is not guaranteed to be consistent across filesystems, so subdirectories get an explicit sort (project_context.rs:224); JSON fields go through BTreeMap to guarantee ordering (project_context.rs:183).
Rule 3: rather auto-generate a file than rescan the disk every turn. This comment deserves to be quoted in full (project_context.rs:503-511):
Auto-generate
.codesmith/instructions.mdwhen NO tier resolved — no managed/user/project/local file anywhere. This avoids the per-turn filesystem scan fallback inprompts.rsthat breaks KV prefix cache stability.
After generating it, the merge runs once more so the new file gets stamped with the same <!-- tier: project --> tag as pre-existing files, "guaranteeing byte-identical repeated builds." For the sake of cache stability, the engine would rather write a file on your behalf.
Rule 4: recursion needs a depth cap. The @include reference chain is capped at 5 (claudemd.rs:77-82); the comment says the silent truncation beyond the cap exists to "bound recursion and keep prompt assembly cache-friendly," and it pointedly notes that this mirrors Claude Code's MAX_INCLUDE_DEPTH = 5.
Rule 5: localization uses bookends, not a full translation. The Constitution's base.md is forever English; every language only adds a pair of preamble/closer at the ends. The second reason is written right in the comment (prompts.rs:442-459): one English Constitution plus per-language bookends means the largest cacheable block is shared byte-for-byte, across sessions, among users of the same language; a full translation gives one cache per language, mutually unintelligible. Native-language reinforcement goes into the closer — the spot nearest the user's next message, where attention weights run highest (prompts.rs:484-495).
The military rules govern the determinism of assembly; how the content itself is written is a separate craft, and three lessons are worth writing down:
A two-layer structure of XML and Markdown: XML tag names carry their own semantics (
<turn_meta>,<section stability="static">— one glance tells you what the block is), while Markdown headings handle the human-readable hierarchy — the former is precise, the latter congenial.Process beats a pile of rules: handing the model a Step 1→2→3 standard operating procedure beats a hundred scattered rules, because at any moment the model knows which stage it is in, and when something goes wrong it acts by stage instead of walking the rule list looking for a match — the technique that lowers the load for a new human employee works just as well on a model.
Few-shot examples are one of the most expensive luxuries of the cache era: examples live in the prefix zone, so once fixed they must stay byte-stable — "dynamically selecting the most relevant examples per request" amounts to rewriting the prefix every time. Production systems therefore keep a fixed example set for each class of task; two or three examples that cover the edge cases beat ten near-duplicates.
Append, Don't Rewrite
The discipline at the message layer is harsher than at the system layer, because this is where "history" lives.
History always appends. The comment on the session-archival seam says it outright (engine/mod.rs:2545-2554): "Append the seam as an assistant message. This is an append-only operation — no messages are deleted. The prefix cache stays hot." Even "archival," an act that sounds like tidying up, is implemented as appending a new message rather than deleting or editing old ones.
The stealthiest entrance for rewriting history: conditional replay. This is the one comment I most want you to see in the entire installment (prompt_inspect.rs:945-962). Under DeepSeek's thinking mode, assistant messages carry the reasoning text; when the wire layer replays history, should it come along? It looks like nothing more than a question of presentation:
// Reasoning replay must be a function of the stored message ONLY,
// never of later history. DeepSeek's prefix cache hashes the raw
// bytes of every message; flipping `reasoning_content` on/off
// depending on whether a follow-up user turn exists rewrites a
// historical message between turns and busts the cache from that
// point onwards. Always emit `reasoning_content` when the model
// requires replay AND the stored message carries thinking text.
Without this line of defense, the bug would look roughly like this: turn one goes out with the reasoning and hits cache; turn two omits the reasoning "because newer messages follow" — history has changed, the cache is all gone, and the bill shows nothing amiss. This is the most insidious breed of cache bug: everything works perfectly, it just quietly gets more expensive.
Whether the system prompt gets replaced is decided by its hash. The system prompt is reassembled every turn, but before any replacement the hash is compared (engine/mod.rs:2846): the Session keeps a last_system_prompt_hash, and if the freshly assembled hash is identical, no replacement happens — sparing the downstream from pointless churn over a changed object identity. Even a "replacement that changes no content" is treated as contamination.
Dynamic Information: All of It Escorted to the Tail
Then what about when the model needs to know "what day is it today"? Or when the working directory changes? Information that may change from turn to turn lives entirely in the append zone.
The answer is the <turn_meta> block (engine/turn_meta.rs:1-15). The module docs: it carries the current date, the auto-routed model and thinking tier, a working-set summary, and any conditionally triggered skills, as the first ContentBlock of the user message — the model sees the state of its environment before it reads any injected text. Inside the content construction (turn_meta.rs:94-128) hides an exquisite choice of granularity:
let today = chrono::Local::now().format("%Y-%m-%d").to_string();
The date stops at the day. A minute-level timestamp changes every turn. Of course, <turn_meta> hangs off a new message in the append zone, so even a change would not wound the prefix — but day granularity pays a side dividend: within the same day, every turn's turn_meta is byte-identical, and the wire layer can deduplicate and fold them with a clear conscience. The granularity of a single format string feeds two mechanisms at once — caching and compression.
The working-set summary is another specimen of design-for-byte-stability (working_set.rs:683-691): at render time it deliberately discards turn-varying fields such as touches and "last seen N turns ago," and sorts with the turn-independent sorted_for_prompt — the comment stresses that the block is byte-stable across turns when "no new paths observed" (#280), because "one byte of drift here and every content cache after it misses." One honest observation in passing: that same comment still says "the block lives in the system prompt," while today it actually lives inside <turn_meta> (the parameter name at prompts.rs:933 carries an underscore prefix, marking it ignored) — the block moved out of system and into the append zone, and the comment never caught up. Documentation lagging behind code is itself fossil evidence of how the "stable sections up front" discipline evolved.
Even the Constitution's own recap obeys this art of layout. AUTHORITY_RECAP sits at the very tail of the system prompt, and a comment explains the choice of position (prompts.rs:840-844): it exploits recency bias so that the hierarchy table is the last thing the model reads before generating — while consuming none of the cache-stable prefix space. One reminder, collecting the best of both ends.
Even the Tool Catalog Stands in Formation
Article 4 mentioned tool-name sorting going into the fingerprint; here are the two follow-up repairs, both from issue #263.
Repair one: partitioned sorting. When the tool catalog is built, built-in tools and MCP tools are each sorted by name, and the built-ins stay contiguous and first (engine/tool_catalog.rs:98-113):
// Sort each partition by name for prefix-cache stability (#263). The
// upstream `to_api_tools()` already sorts the registry's HashMap output;
// this catalog is built from caller-supplied Vecs which the test harness
// and (future) caller refactors may not pre-sort. Built-ins stay as a
// contiguous prefix ahead of MCP tools so adding/removing an MCP tool
// never shifts a built-in's position.
Sorting guards against the engine's own randomness; partitioning guards against one tool's admission implicating the entire catalog.
Repair two: the two-pass approach to deferred activation. Some tools are not sent to the model by default; they are activated on demand. The naive approach re-inserts them at their original position once activated — one byte inserted mid-catalog, and everything after it shifts. The actual approach (tool_catalog.rs:284-304): resident tools keep their original order up front, and tools activated mid-flight are appended to the end of the catalog. The comment: "Otherwise activating a deferred tool shifts every later tool's byte offset and busts the cached prefix from that point onwards."
In the request, the tool catalog sits right beside the system prompt; it is part of the prefix. Adding or removing one tool = reordering the catalog = voluntarily forfeiting the cache, so index-class tools stay registered and unchanged for the entire session (docs/ARCHITECTURE.md:181, where the rationale field reads "catalog stability, KV prefix cache").
Teaching the Discipline to the Model Itself
Everything above is the engine guarding the cache on the model's behalf. This project goes one step further: it writes cache economics into the manual the model itself reads. The system prompt carries a section ### Prompt-cache awareness (prompts.rs:971-978) that explains to the model, straight out, that DeepSeek caches by the longest byte-stable prefix and that a hit costs roughly 1/100 of a miss, then lays out rules of conduct — four excerpts:
- Append, don't reorder. New context goes at the end. Reordering or rewriting earlier messages invalidates every cache entry past the point of change.
- Don't paraphrase quoted content. For files already read, cite the path and line numbers; do not re-cite them in a different format.
- Read once, refer back. Re-reading the same file produces a new tool-result envelope; paging back to look is cheaper than fetching it again.
-
Use
/compactas a hard reset, not a tweak. Compaction is by design the reset for when "the cache has already lost" — do not trigger it for small gains.
Even the cache-hit chip in the corner of the TUI spells out its color rules for the model (red <40%, yellow <80%) and adds the reminder that "a few red turns in a row means it is time to converge." The context rules written for the model and the interface drawn for the human share one semantics — Article 10 will show the full form of this design.
Going on Offense: Warm-up and Deduplication
Beyond keeping the discipline, there are two offensive moves.
Cache warm-up (prompt_inspect.rs:110-143): at session start, proactively send a request containing only the stable layers — the stable system plus the history minus its last user message — with one tail sentence CACHE_WARMUP_USER_TAIL = "请只回复 OK" and the parameters max_tokens: 8, temperature: 0.0. For the price of eight output tokens, the provider pours the prefix into the cache, and the first real request walks in pre-discounted. That Chinese tail sentence is a lovely touch: however talkative the model, eight tokens is a hard ceiling.
Wire-layer deduplication (prompt_inspect.rs:614-794): before sending, the message stream is folded in two ways. Adjacent <turn_meta> blocks with identical content fold into <turn_meta_unchanged />; duplicated large tool results (≥1,024 characters) are replaced by <TOOL_RESULT_REF sha="...">, with the original written to disk under ~/.codesmith/tool_outputs/ and fetched back by hash. Write-type tools are exempt — the comment (#1695) puts it plainly: "two identical large write_file calls must each keep their fulls confirmation inline" — not a single full receipt of a write operation may be spared.
The Honest Gaps
This installment has said a great deal about "discipline," but before closing, three facts about the implementation as it stands must be put on the table. The cardinal sin of source-code analysis is mistaking the scaffolding for the finished building, and these three are exactly the places most tempting to oversell.
First: the three-zone contract is genuinely wired in — no longer a diagram. In the v0.3.0 era it truly was a diagram: prompt_zones.rs drew the three types PinnedPrefix / AppendLog / TurnScratch, yet the module header described itself as "Phase 1 scaffolding — not yet wired into the engine request path," and even the /cache zones debug command's output admitted as much. v0.5.0 (#2264 Phase 2) wired the contract in, and the module header now reads "All six types are load-bearing." In code, this comes down to three things:
-
AppendLogis now theSession's official transcript storage, and the only day-to-day mutation left ispush— append. - To replace history wholesale, the sole entrance is
AppendLog::rebuild, and it must file aRebuildReason. - Each step's
MessageRequestis assembled byThreeZoneRequest, andmessagescan only ever be a slice of the log plus a scratch tail.
In other words, the methods for rewriting history no longer compile — append-only has been promoted from a discipline into a compile-time property of the type system (for the legislative history, see Article 5). To be clear, the exceptions were not exterminated: compaction, overflow recovery, /edit rollback, session resume, cycle seeds — every legitimate rewriting scenario is funneled into rebuild, each leaving a RebuildRecord for /cache zones to audit. The type system cannot govern "whether an exception ought to happen"; it governs "an exception cannot happen offhand."
Second: the fingerprint check has reported for duty, but only as a witness, not a police officer. The protagonist of Article 4, PrefixStabilityManager::check_and_update, still had no production call sites when the first draft of this installment was being written — the coroner was on the payroll but not on the schedule. Commit ab7f093b put it on the schedule: every step, while assembling the request, calls observe_prefix_stability once (engine/host_executor.rs:1739), and on fingerprint drift it emits an Event::PrefixCacheChange. The event fires only on drift; the comment explains why:
Routine heartbeats would flood the channel once per step; drift-only emission keeps the signal where the cost is.
But its authority must be drawn precisely: shadow mode — observe only, never enforce. On finding drift it merely reports the event; it intercepts no request, restores no scene, rewrites no content. Internally, the manager swaps the baseline to the new prefix, and the next drift is measured against that new baseline. Even the wiring is deliberate: the production path passes in a shared Arc attached to the Session via with_prefix_stability (host_executor.rs:1476), so fingerprint state survives across turns; embeds and tests receive None by default and behave exactly as before.
Third: there is no prompt_cache_key anywhere in the repo. Of the provider-side prefix-cache parameters, CodeSmith uses none — DeepSeek does implicit automatic caching, and there is exactly one thing the client can do: keep the bytes stable. This is a position, not an oversight: rely on no provider's special channel; cache friendliness is redeemed entirely by our own discipline of layout. The one explicit caching mechanism in the providers crate is the per-block cache_control passthrough on the Anthropic path (crates/providers/src/rig_adapter/shaper.rs:238-251) — a capability reserved for embedders; the default assembly produces no Blocks.
Conclusion: Cache Friendliness Is an Art of Layout
Strung together into one sentence, this installment's discipline reads: the stable is placed up front in order of stability; the volatile is driven to the tail in the form of appends; and every byte written into history is never rewritten for life.
Unpacked, the checklist: the builder does no I/O; the pack is sorted, bounded, and timeless; a missing four-tier instruction file gets auto-generated; include depth is capped; localization uses bookends; history is snapshotted verbatim; reasoning replay depends only on the stored message itself; whether the system prompt is replaced is decided by hash; dynamic information lives in <turn_meta>, with the date stopping at the day; the tool catalog is sorted and partitioned, with activation by the two-pass approach; even the model's own manual says "don't paraphrase, don't re-read, don't compact recklessly." Not one of these is a feature. Every one of them is money.
And what all this discipline buys is this: the append zone has become the only legitimate dynamic region in the entire system. Which raises the question — who lives there? And what does the "now" the model looks up to see on every turn actually look like?
Top comments (0)