title: "The AI thinks, the gate decides — how I made LLM code edits deterministic (and cut token usage 42×)"
published: true
tags: ai...
For further actions, you may consider blocking this person and/or reporting abuse
The 42× number is real, but it only measures the right thing on well-scoped tasks. The win isn't a smaller model — it's removing the agent's explore-reread-redecide loop. We saw the same pattern: most token burn in coding agents is context re-reading and re-planning, not generation. Routing by task shape (deterministic pipeline for well-defined edits, agent only when the diff is genuinely ambiguous) cut our own spend on the same order.
One caveat worth flagging: formatDate is the best case for a harness. On refactors where the "right" change isn't knowable up front, the agent's exploration actually earns its tokens, and the gap shrinks. The honest unit is token-per-successful-edit, not token-per-task — the latter makes any harness look 40× better than it generalizes.
Did you benchmark on ambiguous tasks too, or only tightly-scoped ones? Curious whether the 14–42× range holds when the edit isn't a known shape.
Straight answer: no — all ten tasks were tightly-scoped edits, by design. That's the case D-Engine is built for, and the article's "what agents do better" section says so explicitly: exploration and genuinely ambiguous diffs are where the loop earns its tokens. A mixed-shape benchmark (route N tasks by shape, measure both) is the obvious next experiment, and I'd expect exactly what you describe — the gap holds on known-shape edits and shrinks or flips on exploratory ones.
On the unit: fair point, and worth being precise. On this task set quality tied (48/50), so token-per-successful-edit and token-per-task happen to coincide here. But you're right that the former is the honest unit in general — a harness that fails on a third of ambiguous tasks looks 40× cheaper per task and isn't. One nuance I'd add: the two sides' failures aren't priced equally. A harness failure is cheap and safe — the gate rejects, nothing merges, you burned ~2K tokens. An agent failure can be the most expensive run of the day (214K on the trap task) and still be wrong. A full comparison should price failed runs on both sides, not just successful ones.
Routing by task shape is exactly the endgame I have in mind — harness for the known-shape edit, agent for the genuinely ambiguous one. Curious how you decide which shape a task is before running it, though. That classifier feels like the hard part.
Fair — and that's the honest scope. The 42× is a clean number precisely because the tasks were well-defined; the moment the change isn't knowable up front, the agent's exploration earns its tokens and the gap to a harness shrinks fast. The unit that survives both regimes is token-per-successful-edit, not per-task — per-task makes any harness look 40× better than it generalizes. Did you ever find a task-complexity threshold where the explore-redecide loop starts costing more than it saves — i.e. where you'd just let the agent run?
You're right: the honest unit is tokens per successful edit, not per task. My benchmark measures 10 bounded tasks — home turf where a harness wins by design — and the 42× is clean precisely because of that. Counting tokens per edit that actually lands in master (with failed attempts priced in), the number holds up across any task type. In fact, I'm adding exactly that metric to every run's summary output.
And there's a detail that makes that unit even more lopsided than it looks: an agent's 30th attempt doesn't cost the same as its 1st. The model is stateless, so every call resends the full history — call 1 costs X, call 30 costs ~30X. Total cost = calls × context per call, and a harness wins both factors at once: fewer calls, each carrying less context (only what the selector picks), and both constant. The agent loses both: more calls, each one more expensive than the last. Same process, same LLM — the difference is structural, not tuning.
On the threshold: I don't have a single measured cutoff yet, but I do have a boundary that works better than "complexity". The variable isn't how complex the task is — it's whether the needed information already exists in written form. If I know WHAT to change and the symptom points to concrete suspects, the planner localizes and retries converge cheaply: each attempt carries new information (compiler errors, previous diffs, user hints), costs ~20K tokens, and master stays intact no matter what. If the information doesn't exist yet, the question becomes who manufactures it and at what price: the agent manufactures it with an unbounded loop — its worst case isn't just expensive, it's an infinite loop or a broken repo. We manufacture it with a bounded chain: instrument → execute → read the measurement → fix, each link behind its own gate.
Where would I just let the agent run? Where the measurement can't be written in advance and the user can't run it either. And even there, the plan (v0.5) is the agent behind the gate: it explores and measures with its loop, but every change it attempts must pass compilation/tests inside an isolated worktree, under a hard token cap. When it fails, both fail — the difference is that ours has a ceiling and the other one doesn't.
The "call N costs ~N×" point is the one most teams miss — they budget per call, not per trajectory, so a 30-iteration loop looks fine per-call and brutal per-trajectory. One nuance: a gateway can claw some of that back without a full harness rewrite, by de-duplicating the resent history — KV cache across turns, or summarizing stale context — which attacks the "each call costs more than the last" factor directly. So the lever splits into two: fewer calls (gate the loop) vs smaller calls (compress the history). In your runs, is the dominant cost the count of attempts or the growth of context per attempt? They want different fixes, and knowing which dominates tells you whether to cap the loop or compress the window.
Great question — and it's exactly the split our benchmark ended up measuring.
Short answer from our runs: attempt count dominated, not per-attempt context growth. The dsh rows correlate almost perfectly with tool-call count: 3 calls → ~32K tokens, 7 → 44K, 9 → 74K, 11 → 113K, 12 → 107K, 19 → 179K, 20 → 214K. That's roughly 8–11K tokens per call whether it's call 3 or call 20 — per-call cost stays flat, so the 6.6× variance across same-difficulty tasks comes from how many loop turns the agent decides to take, not from history growth. Two controls support the read: same model + same task with reasoning off did 3 calls / 37K vs 12 calls / 107K with reasoning on; and Aider — no loop at all, only receives the files you hand it — lands at ~2.1K tokens per task, matching our deterministic pipeline.
Two caveats I'd flag: (1) one write-heavy task (mass JSDoc) hit 103K in few turns — output volume is a third factor, distinct from both of yours. (2) bench-repo is a 5–6 file toy; per-attempt context would grow faster on a real repo, so I wouldn't generalize "attempts dominate" beyond bounded tasks.
Your KV-cache point shows up live in our data, by the way: T7 ran at a 96% provider cache-hit rate, which is exactly why per-call cost flattened instead of climbing. So I'd sharpen the conclusion: with provider-side prefix caching, gating the loop is the first-order lever; history compression is second-order. And one tension worth noting: summarizing stale context rewrites the middle of the transcript, which invalidates the cached prefix — so the next call pays full price again. Compression and prefix-caching fight each other; agent loops are append-only by construction, which is what makes cache hit rates high in the first place. "Compress the window" is pricier than it looks.
We hit this wall earlier this year on a refactoring pipeline where the agent was regularly burning 60+ turns on changes that should've been two or three. Turned out it was re-reading the same file repeatedly across turns because nothing in the loop told it the state hadn't changed. The 2,552 vs 107,000 number on the same diff makes complete sense once you've watched that happen. I'd push on intermediate-state tasks though, where the agent genuinely needs to observe what a prior step actually produced before deciding the next move.
That's a great war story — and it matches the mechanism exactly. Nothing in a standard loop carries "state hasn't changed" metadata, so the model re-reads defensively, and every re-read gets re-sent in every subsequent turn. The bill compounds twice: once for the read, once for the trajectory.
On intermediate-state tasks: fully agreed — it's in the limitations section for a reason. When step N+1 genuinely depends on observing what step N produced, a loop is the right tool. What I want to explore for v0.3+ is whether that observation can be bounded and fresh instead of cumulative. The compiler-feedback retry is one deliberate "observe, then act" cycle where the model sees only the error, not the whole trajectory. Chained tasks might work the same way: run 1 merges behind the gate, run 2 starts with a clean context against the new state. Git history becomes the memory; the trajectory never grows.
How did you end up fixing the re-reading in your pipeline — caching, state digests, something else?
The 6.6x variance between 32K and 214K tokens on identical tasks points to the real financial hazard of unconstrained agent loops. Teams usually budget agents using the average token burn of successful runs, but re-sending the entire trajectory on every tool call makes the cost function convex in turn count. If a run hits three extra diagnostic steps, you do not pay a linear surcharge. You re-bill the entire prior context three additional times.
Putting a deterministic gate around the edit boundary changes the payoff structure completely. Instead of selling the model a free option to wander through the repo at spot API rates, the harness forces the agent to price each state transition against a fixed budget. The reason providers rarely default to that pattern is straightforward. When the client harness wanders, the platform bills every redundant forward pass, while the user absorbs all the downside variance.
The convexity point is the sharpest framing of this I've seen — and you're right that it's not linear: three extra turns don't cost three extra units, they re-bill the entire prior context three more times. Average-based budgeting hides exactly that tail.
One qualification on the provider-incentive point: I'd separate "rarely default" from "never address". DeepSeek's own harness (dsh) ships a Minimal mode and PTC — one TypeScript program instead of N loop turns — so at least some providers are actively sanding the loop's cost down. But the asymmetry you describe is real: the wandering happens at spot rates on the client's bill, and the variance lands entirely on the user. A deterministic gate moves that variance to zero at the edit boundary — you can't fix the model, but you can fix the payoff structure around it.
Worth noting the 6.6× happened at DeepSeek prices — fractions of a cent. The same variance at frontier-model rates is where it stops being an anecdote and starts being a budget meeting.
Do you know of teams actually budgeting agents per-turn or with hard turn caps, rather than on averages? Genuinely curious how this gets handled in practice.
The token number has a second edge on free tiers that your table would not show. Measuring Groq's free tier, the 8000 TPM budget counts the max_tokens you declare, not the tokens you actually get back: a 20-token prompt with max_tokens 8192 is rejected 413, "Limit 8000, Requested 8271", before generating anything at all. So an agent loop does not only burn more tokens than a deterministic pipeline, it also has to declare a large ceiling on every hop, and on a metered free tier that ceiling is charged whether or not it gets used. Your 2,552 vs 34,600 gap probably understates the difference in how many requests each approach can actually push through per minute.
Great catch — and it's an axis my table doesn't capture: I measured consumed tokens (the API usage field), not the declared ceiling. And you're right that architecture matters even more under that rule: an agent loop doesn't know how much it will write on each hop, so it has to declare large ceilings just in case, and on a metered tier it pays that ceiling whether it uses it or not.
A deterministic pipeline's calls have structurally small outputs: the selector returns a file list (~300 tokens), the proposer returns a bounded diff. Small declared ceilings → many more requests per minute on the same budget. So yes: 2,552 vs 34,600 measures consumed tokens; in actual requests-per-minute under a metered quota, the difference is probably even larger.
And it compounds with the effect discussed above: the loop not only resends a growing history on every hop — it also declares a large ceiling on each one. I'll measure this as its own axis in the mixed benchmark I'm preparing (bounded tasks, symptomatic ones, and one genuinely blind task), including free-tier throughput. Thanks for the 413 detail — I didn't know that about Groq.
The token table is compelling, but the task is the most favourable possible case for a deterministic pipeline: add a pure function with a fully specified signature to a named file. No ambiguity about where it goes or what correct means.
The number I would want next is where the gate stops being able to route and has to hand back to the model, measured on a change that spans three files and depends on existing conventions, plus how often the deterministic path refuses or produces something subtly wrong. 42x on the easy half of the distribution is still genuinely valuable, but the honest framing is a router with a fast path rather than a replacement for the agent, and framing it that way makes the result harder to argue with.
Fair on one thing: the benchmark distribution skews bounded, and I'll own that — as I said to TokenLat above, the 42× is clean precisely because the tasks were well-defined, and the unit that survives everything is tokens per successful edit with failures priced in.
The numbers you ask for are exactly the ones I want to measure next. The mixed benchmark I'm preparing has three task shapes — bounded, symptomatic ("X is broken", no file named), and one genuinely blind task ("checkout is slow") — and it will count how often the deterministic path resolves, retries, or has to escalate. Multi-file changes that depend on repo conventions are their own task class in it. Including the losses, same as I did here with T9 and T10, where the agent won by creating new files.
One nuance on "refuses or produces something subtly wrong": the harness today doesn't refuse tasks — the gate refuses changes. If a diff doesn't compile, it doesn't land, and the compiler error goes back to the proposer as new information for a cheap retry. What the gate can't see yet is semantic truth — a patch that compiles but is subtly wrong. That's why any fuzzy-matched patch escalates to extra verification in v0.3, and why v0.5 adds tests and performance budgets to the gate.
And I accept the framing: it's a router with a fast path. In fact, the roadmap is a router — a preflight classifier that decides without spending tokens: fast path when you know what and where, a planner call when you know what but not where, a bounded chain (instrument → execute → read the measurement → fix) when the information has to be manufactured, and an agent — behind the gate, under a hard token cap — only when none of that applies. The difference from "just another agent" is that every path, fast or slow, goes through the same gate, and every path has a cost ceiling.
The determinism split is the right architecture: let the model propose semantics, let the gate own syntax and placement.
One thing I'd stress-test is the 14-42x comparison. Agent runs also spend tokens on retry and exploration that the harness skips, so part of that ratio buys failure recovery. The fair number is quality-matched runs at the same final test state, and it's worth publishing that methodology, because token ratios are exactly the kind of figure that gets quoted without the matching condition attached. Also curious how the gate handles a correct edit aimed at the wrong file: deterministic placement needs a resolution rule for ambiguity, and that rule is where the remaining variance will hide.
Both points are exactly the ones to measure. On the first, a nuance and an agreement: the ratio already includes the agent's retries (that's why its numbers are 107K–214K and not 30K) and ours too (T3 needed a second round, and it's counted). But you're right about methodology: the honest figure is quality-matched runs — same final test state — and that condition must be published next to the number, or the number is worthless. There's an asymmetry I like there: on our side, the matching is enforced by construction — a run only counts if the gate is green. On the agent's side we reviewed final states by hand, and when they didn't match we documented it as a loss (T9: hallucination + unauthorized fix; T10: unrequested new file). The mixed benchmark will publish that methodology explicitly.
On the correct edit aimed at the wrong file: the resolution rule today is uniqueness or refusal. Every change carries a SEARCH anchor that must match exactly once in the named file: zero matches, or two or more, and the local editor refuses to apply, reports, and master stays untouched. There's no placement guessing. The case that does escape that rule — a unique anchor in the wrong file, compiles but is semantically wrong — is exactly the variance you point at, and compilation alone can't cover it. That's why the roadmap adds three layers over that gap: in v0.3, any fuzzy-matched patch escalates to extra verification; in v0.4, a semantic reviewer inside the worktree (a second call that judges the applied diff against the spec, with rejections feeding bounded retries) plus a user-editable spec shown before execution; and in v0.5, tests and performance budgets inside the gate. The placement rule is "uniqueness or refusal"; semantics is defended in layers.