Context Management: Virtual Memory for the Model
Source version of CodeSmith:
v0.5.0(commit3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: anyone whose context window has been stuffed to bursting by giant logs, long PDFs, or whole test-suite outputs.
The history of computer science contains a classic hard problem: the program is bigger than memory — what then?
The early answer was overlays — programmers carved the program into blocks by hand and ferried them into memory in turn, an exercise in misery. The later answer was virtual memory — the process believes it owns one vast, contiguous address space all to itself; the operating system quietly knows the truth, paging blocks in and out on demand. The programmer's burden vanished, because a healthy layer of abstraction sits between the program and the truth.
Today's large language models have crashed into the same problem: the material to read is bigger than the context window. An 80,000-character log, a 1,200-page PDF, the output of a full test-suite run — stuff it in, and it crowds out the budget; truncate it, and you lose the trail. Most tools handle it much the way the overlays era did: hand-carrying, by prompt trickery.
CodeSmith's answer is to bring the idea of virtual memory in wholesale: the model believes it can read everything; the engine guards the truth about the context.
The handle: a library card of a few dozen bytes
Start with the trigger. When an RLM session (the protagonist of Article 6) finishes a probe, stdout routinely runs to tens of thousands of characters. When do you give up on inlining and switch to a handle? The threshold is set surprisingly small (crates/tool-impls/src/tools/rlm.rs:31-34):
/// When `rlm_eval` stdout exceeds this many characters the full body is
/// stored as a `var_handle` instead of inlined into the parent transcript.
/// The model retrieves the body via `handle_read` using the returned handle.
const STDOUT_HANDLE_THRESHOLD_CHARS: usize = 1_000;
1,000 characters. Past that line, the full text never enters the parent context; what takes its place is a "card" — VarHandle (crates/agent-runtime/src/tools/handle.rs:36-45):
pub struct VarHandle {
pub kind: String,
pub session_id: String,
pub name: String,
#[serde(rename = "type")]
pub type_name: String,
pub length: usize,
pub repr_preview: String,
pub sha256: String,
}
Seven fields, a few dozen bytes: which session it belongs to, its name, its type, how long it is (length — how thick the book is), a short preview (repr_preview — one sentence quoted off the cover), and the SHA-256 of the contents. Whether what sits behind it is 1,000 characters or 1,000,000, what lies in the parent context is the same card.
Here is a small design that is easy to miss: the sha256 field. On read, the engine verifies that the hash carried by the handle matches the actual payload in storage (crates/tool-impls/src/tools/handle.rs:140-143); a mismatch fails outright. What this guards against is a subtle temporal skew: the model holds a handle from a previous version and goes to read a variable that has since been overwritten — reading the new page with an old page number should fail, not silently return the wrong content.
handle_read: paging on demand
How does a card become text? The handle_read tool. Its input schema (crates/tool-impls/src/tools/handle.rs:46-105) is itself a manual for a "projection language":
-
slice— a zero-based half-open slice, in units ofcharsorlines; -
range— 1-based, closed-interval line numbers (the more model-friendly "line X through line Y"); -
count— returns metadata counts only; -
jsonpath— a small JSONPath subset ($,.field,[index],[*],['field']), built precisely for large structured JSON; -
introspect— returns the projections this handle supports, size hints, and invocation examples you can copy-paste directly.
Each read is capped at 12,000 characters by default, with a hard ceiling of 50,000 (handle.rs:22-23). Faced with an 80,000-character log, the model's move is not "give me all of it" but something closer to a practiced engineer working less: count first to see how big it is, slice the head for a glance at the format, range to jump to the lines near the error, and, when necessary, jsonpath to pull fields straight out.
The introspect projection deserves a sentence of its own. It embeds "how this tool is used" into the tool itself — when the model is at a loss with an unfamiliar handle, one call to introspect brings back usage examples aimed at that specific handle. The manual is not written down in documentation, waiting on the model's luck to recall it; it is a first-class query. Tool discoverability is accessibility infrastructure for the model.
Step back, and every piece here has a straight-faced counterpart in computing: the handle is a page-table entry, length + repr_preview is a metadata page, slice/range is the paging window after a page fault, jsonpath is direct memory access, and the SHA-256 check is the page table's valid bit. The context window is RAM, the workbench is disk, and the model is the process that believes it owns all of memory.
In the language-model world, the most famous precursor of this metaphor is MemGPT — which goes so far as to define its theme as "putting an operating system on the LLM": the context window is main memory, everything outside the window is disk, and paging is done by the model's own hand — it offers a set of self-editing memory functions (append to and replace core memory, write to and retrieve from archival memory), and the model, like an operating system, decides for itself what stays in main memory and what gets swapped out to disk. It is also the representative of the "layered memory" lineage among the five shapes of the Agent loop (the family-tree table in Article 3): what the loop persists and reads is the layered memory itself.
VarHandle and MemGPT are two answers to the same problem, and the fork in the road is a single question: who plays the operating system. MemGPT's answer is the model itself — paging is a function call made by the model; CodeSmith's answer is the engine — the model merely pages on demand, and the page table is held in the engine's hands. For cheap brains, the latter assignment is the safer one: paging is deterministic mechanical work, and giving it to deterministic code removes a whole class of failure modes that a probabilistic brain would add — and it gives cache discipline (Article 4) a single unified gatekeeper.
The ruler: the CJK lesson in token counting
Before virtual memory can actually run, there has to be an honest ruler — you must know how much space each thing takes. CodeSmith's token counter has an origin worth recounting (crates/agent-runtime/src/tokenizer.rs:5-7):
Historically each site divided characters by 3 (or bytes by 4), which is a poor approximation for CJK-heavy text and drifts badly on JSON-heavy tool output.
Behind that sentence is a pit a Chinese-language team stepped into for real: one Chinese character counts as 1 character under Rust's chars().count(), but to a tokenizer it is frequently 1 to 2 tokens — chars/3 underestimates the bulk of Chinese text several times over. And every budget mechanism in the engine (compaction triggers, capacity pre-checks, large-output routing, truncation thresholds) feeds on this ruler's readings. If the ruler is crooked, every budget is crooked.
Today's answer is dual-mode (tokenizer.rs:32-40):
pub enum TokenCounter {
/// CJK-aware heuristic: one token per CJK codepoint, one token per
/// four non-CJK characters.
Heuristic,
/// Exact counts from a loaded HuggingFace tokenizer (feature
/// `hf-tokenizer`).
#[cfg(feature = "hf-tokenizer")]
Hf(Arc<tokenizers::Tokenizer>),
}
The default is still the heuristic — no extra dependency, works out of the box; and v0.5.0's heuristic has made up the CJK lesson: one token per CJK codepoint, one token per four remaining characters (heuristic_count, tokenizer.rs:101), so Chinese text is no longer systematically underestimated. For exact per-token values, simply configure [context].tokenizer_path to point at a HuggingFace tokenizer.json (with the hf-tokenizer feature enabled at compile time), and even the large-output routing estimates turn accurate along with it.
There is also a process-level detail: the counter is installed globally exactly once, first come first served (tokenizer.rs:130-140). Once the precise counter is installed at startup, anyone trying mid-session to switch back to the heuristic gets — a warning, and no switch. The comment gives the reason: changing rulers mid-session leaves the same session's budgets on inconsistent yardsticks. One crooked but stable ruler beats two accurate rulers used in rotation.
Two handy little bits of paging
The philosophy of virtual memory — "pay full price the first time, ride the cache afterward" — is everywhere in this codebase. Here are two of my favorites.
Lazy construction of the file index (crates/agent-runtime/src/working_set.rs). @ mentions and the file picker need fuzzy resolution: "the user typed handle — do they mean crates/xxx/src/handle.rs?" The naive implementation walks the whole tree on every query. CodeSmith's approach is written out plainly in the comments (working_set.rs:25-28):
...a lazy basename → paths index built once on first miss and reused for the rest of the session — without it, every mis-typed mention triggered a full
WalkBuildertraversal…
Loaded through a OnceLock (working_set.rs:34,101), capped at 50,000 entries (working_set.rs:299), with first-round latency kept under control. There is that familiar shape again: one cold start, and the hot path is cached forever.
Frecency decay for @-mentions (crates/tui/src/tui/file_frecency.rs). Between a file you dug into deep last week and a file you mentioned five minutes ago, which one ranks higher in the completion list? Rank by "total count," and last week's hotspots hold the top of the chart forever. The frecency (frequency + recency) answer is to let heat decay exponentially with time: a half-life of 7 days (file_frecency.rs:33), with a score of count × e^(−λ·age) (file_frecency.rs:72-76) — mention counts accumulate linearly, then get multiplied by a factor that decays exponentially with age. The comments carry a charming argument for the half-life's value: "long enough that frequently used files survive a working week, short enough that yesterday's deep dive does not haunt you forever." Persistence is append-only JSONL, compacted and merged in memory, capped at 1,000 entries with the lowest scores evicted first (file_frecency.rs:27). Incidentally, the comments candidly admit that the eviction heuristic borrows from OPENCODE's implementation — good borrowing never hides.
Conclusion: healthy abstractions all interpose a layer
The lesson operating systems taught us bears repeating: the greatness of virtual memory lies not in letting programs use more memory, but in letting programs use less memory without ever knowing it.
VarHandle is exactly that, for the context window. The model does not need to know the economics from Article 4 — "context is expensive, append-only cannot be rolled back, every token is billed" — it simply pages in when it needs to. The dirty work is left to the engine: threshold decisions, card issuance, hash verification, projection rate-limiting. And the foundation under all of it is a ruler with its CJK bias corrected.
Handles solve the "too much to read" problem: what needs reading is paged in on demand, and what needs computing goes to the private kitchen of Article 6. But virtual memory is only the first gate of context — it controls "what enters the window" and has not yet answered "what happens to those bytes once they are in". The answer is an iron law: every byte written into history stays unchanged for life. What that rule is worth, and what it is defended by, is the account Article 9 works out over an entire installment.
Top comments (0)