DEV Community

Sergio Corruchaga
Sergio Corruchaga

Posted on

The AI thinks, the gate decides — how I made LLM code edits deterministic (and cut token usage 42 )


title: "The AI thinks, the gate decides — how I made LLM code edits deterministic (and cut token usage 42×)"
published: true

tags: ai, opensource, typescript, llm

The AI thinks, the gate decides

D-Engine: a deterministic harness that matches coding agents' quality while burning 14–42× fewer tokens

Sergi Corruchaga · September 2026 · D-Engine v0.2.2 (MIT, open source)


1. The number that started it all

On September 10, 2026, I ran the same programming task three times, with the same model (DeepSeek V4.1-Flash), the same literal prompt, and the same repository:

"En utils.ts, añade una función formatDate que reciba un Date y devuelva DD/MM/YYYY"
(Add a formatDate function to utils.ts that takes a Date and returns DD/MM/YYYY)

All three runs produced functionally the same code. Here's what each one cost:

Tool Architecture Tokens consumed Time
D-Engine (my harness) Deterministic pipeline 2,552 ~4 s
dsh — Minimal mode Agent (single tool: shell) 34,600 1m 04s
dsh — effort Off Full agent, no thinking 37,100 6 s
dsh — factory defaults Full agent, thinking High 107,000 28 s

DeepSeek's official agent burned 42× more tokens than my tool to produce the same diff. And as you'll see in the controls section, that gap is explained neither by the model, nor by "thinking mode", nor by the agent's toolbox. It's explained by the architecture.

This article covers how I got here: what D-Engine is, how I ran the full benchmark (10 tasks, 5 contenders, 2 deliberate traps), what agents do better than my tool (quite a few things, and I'm going to disclose all of them), and why I believe the future of AI-assisted programming isn't a smarter agent — it's a stricter gate.

2. The problem: how an agent spends tokens

The dominant AI coding tools (OpenCode, Aider, dsh, Claude Code…) all follow the same pattern: the agentic loop. The model receives your request, decides to call a tool (read file, search, run shell), gets the result, decides another call, and so on until done.

The commonly overlooked detail: the model has no memory between calls. On every turn of the loop, the harness re-sends the full system prompt, all tool definitions, and the entire conversation trajectory so far. If the agent takes 20 steps, step 20 re-sends the previous 19. Cost grows quadratically with the agent's diligence — not with your task's difficulty.

Measured in my benchmark: the same task, in the same repo, with the same model, cost dsh between 32K and 214K tokens depending on how many loop turns it decided to take. A 6.6× variance the user neither controls nor can predict.

There's a second, subtler problem: state drift. The agent works from the "snapshot" of the code it has been reading during the session. If that snapshot goes stale — or the model misremembers it — it will edit something that doesn't exist. Or worse: it will believe it sees things that don't exist. In section 6 I describe how dsh reported a corrupted file that was perfectly healthy, complete with fabricated line-level evidence.

The third problem is atomicity: most agents write directly to your working tree. If the change breaks compilation, your main branch is already broken. Some will even auto-commit the disaster.

3. The idea: separate "thinking" from "touching"

D-Engine is built on a radical separation of responsibilities:

  • The cloud (the LLM) only thinks. It receives the minimum necessary context and responds with SEARCH/REPLACE blocks — patches anchored to existing code.
  • The local, deterministic runtime only touches. It applies those blocks in a photocopy of the repo (a shadow git worktree), compiles with tsc --noEmit, and only if the gate is green merges into the real repo.

A typical task consumes exactly 2 LLM calls:

  1. Selector (optional, ~300 tokens): given a map of the repo, the model picks the minimal set of relevant files. The user confirms — the selection never applies without authorization.
  2. Proposal (~2,000 tokens): the model receives only those files and generates the SEARCH/REPLACE blocks.

Everything else is local code: the LocalEditor applies each patch through a 4-strategy cascade (exact match → newline normalization → ignore trailing whitespace → fuzzy at 0.85 threshold), the compiler validates, and commitAndMerge stages only the files touched by the patch (with a git status --porcelain guard that aborts the merge if any foreign file appears, logging the offender's diff before destroying the photocopy).

There's also a Verify mode adding an optional second phase: send only the modified snippet for a semantic audit (the model answers OK / OK_WITH_OBSERVATIONS / FAIL). Measured cost: 330–984 tokens per task — 15–30% on top of the proposal. Nearly free semantic safety, on the programmer's demand.

The project's motto sums up the philosophy: the AI thinks, the gate decides. The model can propose whatever it wants; the only source of truth in the system is the compiler.

4. The benchmark

Methodology

  • Test repo: bench-repo, a TypeScript mini-shop (products, cart, pricing, utilities), frozen at the benchmark-base tag.
  • 10 representative tasks: add a function, multi-file rename, validation guards, mass JSDoc documentation, cart line deduplication, an operation-ordering bug, an extraction refactor, two deliberate traps (an already-implemented task and an "optimization" of something already optimal), and a full feature (a coupon system with expiration).
  • Rules: same literal prompt for every contender, one attempt per task, git reset --hard benchmark-base + git clean -fd before every run, engine frozen during the benchmark.
  • Contenders: OpenCode, Aider, dsh (DeepSeek's official harness), and D-Engine in Fast and Verify modes.
  • Era declaration: the original round ran on V4-Flash non-thinking; the dsh round ran on 2026-09-10 on V4.1-Flash — the very day DeepSeek retired the previous model. The benchmark survived a mid-flight model extinction thanks to declaring model+effort+date per row, plus the control runs in section 6.

Methodological honesty disclosures

Before the results, two confessions. First: two tasks (T3 and T5) turned out to be defective in their first round — the repo already contained what they asked for; the voided rows are preserved in the record as evidence, and the base was fixed. Second: the literal prompts typed in the first round were not preserved (the git clean -fd cycles wiped Aider's histories, and my own record document stored summaries instead of the actual texts). The original specifications were recovered from the design conversation, prompts were frozen in the record on 2026-09-10, and every later execution uses them verbatim. T1 shares an attested literal prompt across all eras — it is the comparability anchor.

None of this is glamorous. That's exactly why it's in the article.

5. Results

Quality: a statistical tie

Contender Points (max 50) Incidents
OpenCode 49/50 4 on trap T9
D-Engine Fast 48/50 4 on T9
dsh (factory) 48/50 4 on T7 (unrequested API added), 4 on T9; 1 hallucination
Aider 47/50 4 on T7, 4 on T9; committed a main branch that didn't compile (T2)
D-Engine Verify 46/50 4 on T7, 4 on T9; false rejection on T6 (parsing bug, since fixed)

Nobody crushed anybody on quality. With the same model, the "textbook" solution converges — on three tasks, three different contenders produced byte-identical files. What differentiates the tools isn't the answer: it's the machinery around the model.

Tokens: here's the difference

Contender Tokens per task (avg) vs D-Engine Fast
D-Engine Fast ~2,100 1×
Aider ~2,100 ~1×
D-Engine Verify ~2,600 ~1.2×
OpenCode ~9,700 ~4–5×
dsh ~93,000 (range 32K–214K) ~44×

(Aider deserves a fair note: it's by far the leanest agent, because it only passes the files you tell it to. Its problem wasn't cost — it was the gate. Keep reading.)

The full dsh round consumed ~931K tokens versus D-Engine Fast's ~21K for the same task set and equivalent results.

Time

Fast ~2.7s · Verify ~3.7s · Aider ~4.9s · OpenCode ~14.9s · dsh ~38s (wall clock; its own UI reports ~20s — the gap between both measures, 9 to 66 seconds per task, is startup and latency time the agent doesn't account for).

The front-page moment (T2)

Task T2 asked to rename a constant across two files. Aider warned it was missing context ("I don't have them in the chat. Let me know if you want me to review them")… and then auto-committed a main branch that didn't compile anyway (TS2305). Without a compile gate, AI can break your repo while knowing it's breaking it.

D-Engine, on the same task, rejected its own first attempt: the patch compiled in the photocopy but broke index.ts — the gate caught it, the merge never happened, and main stayed intact. Failing safe isn't a bug: it's the architecture.

Trap T9: the most revealing behavior

T9 asked to "optimize calculateTotal using Array.reduce"… when the function already used reduce. The perfect answer was "nothing to do here".

Nobody gave the perfect answer. Every first-round contender made cosmetic changes (4/5). But dsh did something more interesting and more unsettling at once:

  • It admitted the trap ("it already used reduce") — only OpenCode had done that.
  • It found a real bug nobody had asked it to look for: round2(1.005) returned 1.00 instead of 1.01 due to binary floating-point noise. A legitimate, valuable find.
  • …and then it fixed it unilaterally, modifying utils.ts (outside the target) and changing the rounding behavior of the entire system, when the prompt said to "keep rounding correct".
  • And it finished by reporting that src/products.ts was corrupted ("line 16 reads ndProduct… it breaks the whole project's compilation"). Manual verification: the file was intact. A hallucination with fabricated line-level evidence, in the same message where it claimed tsc passed cleanly.

The behavior was safe (it asked permission before touching the "corrupted" file). But had I answered "yes, fix it", the agent would have edited a healthy file chasing a ghost. D-Engine structurally cannot have this class of hallucination: it doesn't opine on repo state — truth comes from tsc, not from the model.

The cost of all that unleashed diligence: 214K tokens on a task whose correct answer was "nothing to do". One hundred times D-Engine.

6. The controls: killing objections before they're raised

I anticipate three objections to the token gap. All three have measured answers.

"It's the model" → No. The dsh round ran on V4.1-Flash; I ran D-Engine v0.2.2 on the same new model (adapted the very day of the API migration): 2,552 tokens. Gap intact.

"It's thinking mode" → Partially. With effort set to Off (an exact replica of the first round's non-thinking configuration), dsh dropped from 107K to 37.1K. Thinking amplifies the gap ~2.9× (and quadruples loop turns: 12 vs 3 tool calls — a model that "thinks" also wanders more). But the remaining 14.5× is still there with reasoning off.

"It's the tool arsenal" → No. In Minimal mode (a single tool: a persistent shell), dsh consumed 34.6K — practically identical to the full agent without thinking (37.1K). With a primitive shell the agent needed more turns (11), not fewer: search with Get-ChildItem, read with Get-Content, edit with Add-Content and hand-typed \r\n escapes, re-read to verify, compile… The cost isn't in the tool schemas. It's in the loop: every turn re-sends the full trajectory. Shrinking the arsenal doesn't shrink tokens; shrinking the loop does.

Final gap decomposition on T1 (same model, same prompt, same diff):

D-Engine (pipeline)       2,552 tok   1×     ← no loop
dsh Minimal              34,600 tok   13.6×  ← the arsenal doesn't matter
dsh Off                  37,100 tok   14.5×  ← pure architectural overhead
dsh Factory (thinking)  107,000 tok   42×    ← thinking amplifies ~2.9×
Enter fullscreen mode Exit fullscreen mode

7. What agents do better (and it would be dishonest to hide it)

This article is not "agents are bad". dsh produced, by far, the most diligent work in the benchmark:

  • It self-verified: compiling with tsc inside its own loop, and on two tasks it went as far as executing the program to confirm the output was identical before and after the change. No other contender did that.
  • On T10 it wrote a 23-case edge-case suite for the coupon system (same-day expiration, NaN dates, JavaScript's February 31st…), ran it — 23/23 — and deleted it afterwards. Brilliant.
  • It can create new files; D-Engine can't yet (documented limitation, on the roadmap).
  • It detected dead code, duplications, and a validation hole in applyDiscount, and reported them without touching anything, asking permission. Exemplary scope discipline — when it chooses to have it.
  • It found a real rounding bug I didn't know existed.

The honest conclusion isn't that agents are unnecessary. It's that today you pay for their diligence blind: you don't know if your task will cost 32K or 214K tokens, whether the agent will respect your scope or redecorate half your repo, or whether its report about your code's state is true or a plausible hallucination. D-Engine proposes the inverse split: the agent provides judgment; the machine provides truth and a fixed bill.

An important note about money: at DeepSeek's prices (with 96% cache-hit rates measured in some sessions), 100K tokens cost cents. Direct cost is not the argument. The argument is latency (2.7s vs 38s per task), predictability (bounded bill vs 6.6× variance), context degradation in long sessions, and what this gap means when the model costs dollars per million instead of cents — or when the loop runs unattended in CI, with no one there to say "no" in time.

8. Limitations

I declare them before the first hostile comment does:

  1. Small repo (a 5–6 file mini-shop). The absolute gap would grow with repo size in both systems, but the structure of the gap (loop vs pipeline) is size-independent.
  2. A single model family (DeepSeek). Nothing prevents rerunning the benchmark with other providers; the method is portable.
  3. First-round literal prompts were not preserved (incident disclosed in section 4; prompts frozen since 2026-09-10).
  4. D-Engine can't create new files yet, and its file selector (P9) is non-deterministic — though the architecture absorbs the variance safely.
  5. Fuzzy matching (0.85 threshold) is the weakest link: the only patch that broke syntax in the entire benchmark came through that path. On the roadmap: retry with compiler feedback, and mandatory Verify when a patch only applies via fuzzy.
  6. I don't measure agent judgment quality on open-ended tasks ("improve this design"), where the exploratory loop has real advantages. D-Engine is built for bounded, specifiable changes — which are most of daily work.

9. What's next

  • v0.3: new-file creation, retry with tsc feedback, Verify-if-fuzzy, and time+tokens printed in the final summary (measuring time by hand with a stopwatch was the least glamorous part of this benchmark).
  • Published: the code is MIT and lives on GitHub with the complete benchmark (frozen prompts, voided rows included) so anyone can reproduce or rebut it: https://github.com/corruchaga/D-Engine
  • dsh's PTC mode remains as future work: DeepSeek is already trying to collapse the loop into a single TypeScript program. If it works, it's proof the industry is converging toward this idea on its own.

10. Conclusion

Coding agents are impressive. They're also token-burning machines with unpredictable variance, no native compile gate, and a… creative relationship with your repository's actual state.

This benchmark shows that for daily work — bounded, specifiable, verifiable changes — a deterministic pipeline produces the same quality (48/50, tied with the best agent) at a fraction of the cost (14–42× fewer tokens depending on configuration), a fraction of the time (2.7s vs 38s), with a predictable bill and zero broken commits.

We don't need a smarter model. We need a gate.

The AI thinks. The gate decides.


Sergi Corruchaga is a junior developer (DAM graduate, currently studying DAW). D-Engine is his first open-source project. The complete benchmark — every table, incident, and voided row — is available in the repository.

Top comments (21)

Collapse
 
tokenlat profile image
TokenLat •

The 42× number is real, but it only measures the right thing on well-scoped tasks. The win isn't a smaller model — it's removing the agent's explore-reread-redecide loop. We saw the same pattern: most token burn in coding agents is context re-reading and re-planning, not generation. Routing by task shape (deterministic pipeline for well-defined edits, agent only when the diff is genuinely ambiguous) cut our own spend on the same order.

One caveat worth flagging: formatDate is the best case for a harness. On refactors where the "right" change isn't knowable up front, the agent's exploration actually earns its tokens, and the gap shrinks. The honest unit is token-per-successful-edit, not token-per-task — the latter makes any harness look 40× better than it generalizes.

Did you benchmark on ambiguous tasks too, or only tightly-scoped ones? Curious whether the 14–42× range holds when the edit isn't a known shape.

Collapse
 
sergiocorruchaga profile image
Sergio Corruchaga •

Straight answer: no — all ten tasks were tightly-scoped edits, by design. That's the case D-Engine is built for, and the article's "what agents do better" section says so explicitly: exploration and genuinely ambiguous diffs are where the loop earns its tokens. A mixed-shape benchmark (route N tasks by shape, measure both) is the obvious next experiment, and I'd expect exactly what you describe — the gap holds on known-shape edits and shrinks or flips on exploratory ones.

On the unit: fair point, and worth being precise. On this task set quality tied (48/50), so token-per-successful-edit and token-per-task happen to coincide here. But you're right that the former is the honest unit in general — a harness that fails on a third of ambiguous tasks looks 40× cheaper per task and isn't. One nuance I'd add: the two sides' failures aren't priced equally. A harness failure is cheap and safe — the gate rejects, nothing merges, you burned ~2K tokens. An agent failure can be the most expensive run of the day (214K on the trap task) and still be wrong. A full comparison should price failed runs on both sides, not just successful ones.

Routing by task shape is exactly the endgame I have in mind — harness for the known-shape edit, agent for the genuinely ambiguous one. Curious how you decide which shape a task is before running it, though. That classifier feels like the hard part.

Collapse
 
tokenlat profile image
TokenLat •

Fair — and that's the honest scope. The 42× is a clean number precisely because the tasks were well-defined; the moment the change isn't knowable up front, the agent's exploration earns its tokens and the gap to a harness shrinks fast. The unit that survives both regimes is token-per-successful-edit, not per-task — per-task makes any harness look 40× better than it generalizes. Did you ever find a task-complexity threshold where the explore-redecide loop starts costing more than it saves — i.e. where you'd just let the agent run?

Thread Thread
 
sergiocorruchaga profile image
Sergio Corruchaga •

You're right: the honest unit is tokens per successful edit, not per task. My benchmark measures 10 bounded tasks — home turf where a harness wins by design — and the 42× is clean precisely because of that. Counting tokens per edit that actually lands in master (with failed attempts priced in), the number holds up across any task type. In fact, I'm adding exactly that metric to every run's summary output.
And there's a detail that makes that unit even more lopsided than it looks: an agent's 30th attempt doesn't cost the same as its 1st. The model is stateless, so every call resends the full history — call 1 costs X, call 30 costs ~30X. Total cost = calls × context per call, and a harness wins both factors at once: fewer calls, each carrying less context (only what the selector picks), and both constant. The agent loses both: more calls, each one more expensive than the last. Same process, same LLM — the difference is structural, not tuning.
On the threshold: I don't have a single measured cutoff yet, but I do have a boundary that works better than "complexity". The variable isn't how complex the task is — it's whether the needed information already exists in written form. If I know WHAT to change and the symptom points to concrete suspects, the planner localizes and retries converge cheaply: each attempt carries new information (compiler errors, previous diffs, user hints), costs ~20K tokens, and master stays intact no matter what. If the information doesn't exist yet, the question becomes who manufactures it and at what price: the agent manufactures it with an unbounded loop — its worst case isn't just expensive, it's an infinite loop or a broken repo. We manufacture it with a bounded chain: instrument → execute → read the measurement → fix, each link behind its own gate.
Where would I just let the agent run? Where the measurement can't be written in advance and the user can't run it either. And even there, the plan (v0.5) is the agent behind the gate: it explores and measures with its loop, but every change it attempts must pass compilation/tests inside an isolated worktree, under a hard token cap. When it fails, both fail — the difference is that ours has a ceiling and the other one doesn't.

Thread Thread
 
tokenlat profile image
TokenLat •

The "call N costs ~N×" point is the one most teams miss — they budget per call, not per trajectory, so a 30-iteration loop looks fine per-call and brutal per-trajectory. One nuance: a gateway can claw some of that back without a full harness rewrite, by de-duplicating the resent history — KV cache across turns, or summarizing stale context — which attacks the "each call costs more than the last" factor directly. So the lever splits into two: fewer calls (gate the loop) vs smaller calls (compress the history). In your runs, is the dominant cost the count of attempts or the growth of context per attempt? They want different fixes, and knowing which dominates tells you whether to cap the loop or compress the window.

Thread Thread
 
sergiocorruchaga profile image
Sergio Corruchaga •

Great question — and it's exactly the split our benchmark ended up measuring.
Short answer from our runs: attempt count dominated, not per-attempt context growth. The dsh rows correlate almost perfectly with tool-call count: 3 calls → ~32K tokens, 7 → 44K, 9 → 74K, 11 → 113K, 12 → 107K, 19 → 179K, 20 → 214K. That's roughly 8–11K tokens per call whether it's call 3 or call 20 — per-call cost stays flat, so the 6.6× variance across same-difficulty tasks comes from how many loop turns the agent decides to take, not from history growth. Two controls support the read: same model + same task with reasoning off did 3 calls / 37K vs 12 calls / 107K with reasoning on; and Aider — no loop at all, only receives the files you hand it — lands at ~2.1K tokens per task, matching our deterministic pipeline.
Two caveats I'd flag: (1) one write-heavy task (mass JSDoc) hit 103K in few turns — output volume is a third factor, distinct from both of yours. (2) bench-repo is a 5–6 file toy; per-attempt context would grow faster on a real repo, so I wouldn't generalize "attempts dominate" beyond bounded tasks.
Your KV-cache point shows up live in our data, by the way: T7 ran at a 96% provider cache-hit rate, which is exactly why per-call cost flattened instead of climbing. So I'd sharpen the conclusion: with provider-side prefix caching, gating the loop is the first-order lever; history compression is second-order. And one tension worth noting: summarizing stale context rewrites the middle of the transcript, which invalidates the cached prefix — so the next call pays full price again. Compression and prefix-caching fight each other; agent loops are append-only by construction, which is what makes cache hit rates high in the first place. "Compress the window" is pricier than it looks.

Thread Thread
 
tokenlat profile image
TokenLat •

The "loop rounds as first-order lever, context growth as second-order" split is the cleanest framing I've seen for this. And your KV-cache data makes the mechanism explicit: when the provider-side prefix cache holds at ~96%, each call's cost flattens instead of climbing — so the gate's job is to keep the obvious cases deterministic (no model call at all) so only the ambiguous middle pays.

Your compression-vs-prefix-cache tension is the part most pieces miss: a summary that rewrites the middle of the transcript invalidates the cache prefix, so the next call pays full price again. "Compression window costs more than it looks" should be on a poster.

Thread Thread
 
sergiocorruchaga profile image
Sergio Corruchaga •

Thank you — and the cache point lands exactly where I wanted to get. It's true: with the prefix cache at ~96%, the billed cost of each call flattens even as raw tokens grow — that's why the benchmark annotates T7 with its 96% and the caveat that tokens ≠ real cost.

But there are three reasons I don't want to depend on that flattening: (1) the cache discount is rented, not owned — TTL and pricing policy belong to the provider and change without asking you; (2) even when the cache catches the prefix, latency keeps growing — reading 100K cached tokens is cheaper but not instant, and the user waits just the same; (3) as you say, compaction breaks the prefix — the two standard mitigations sabotage each other.
So my thesis isn't "agents are expensive". It's this: the agent's cost depends on provider policies you don't control (cache, TTL, pricing shifts); the harness's cost depends on the task, which you do control. The cheapest token isn't the cached one — it's the one never sent.

And I'll take you up on the poster, with its twin: "The only cache with a guaranteed 100% hit rate is the call you never make."

Thread Thread
 
tokenlat profile image
TokenLat •

The "rented, not owned" framing sticks: a cache hit is the provider's policy call, a skipped call is the router's. (3) is the crux — compaction and caching pull the same lever opposite ways, so you can't have both unless the router knows which regime the step is in. The latency point splits the optimization in two: a 100% cache hit bills less but still waits, so token-cost routing and tail-latency routing diverge. Do you route on billed tokens, wall-clock, or keep them as separate budgets a step trades against?

Thread Thread
 
sergiocorruchaga profile image
Sergio Corruchaga •

Exactly — and you're describing the design without having seen it: "the router has to know which regime the step is in" is literally the job of our preflight classifier: it labels the regime (bounded / symptomatic / blind) BEFORE spending a token. And yes, compaction and caching pull the same lever in opposite directions — which is exactly why the regime label comes first: in the bounded regime there is no call to cache and no transcript to compact.

On your question — billed tokens, wall-clock, or separate budgets? — my answer today is that the router routes on neither. It routes on task shape, which is the only thing knowable before spending. Tokens and wall-clock are only known after the step runs: they can't route, they can only budget. The router picks the regime; the budgets stop the regime if it runs away.

And they are separate budgets with different jobs: the token cap protects the economics; the wall-clock cap protects the user experience. A step can trade against both. But there's an asymmetry that I think is the pretty conclusion: the divergence you describe — a 100% cache hit bills less but still waits — is a property of the agent regime. In the deterministic regime, both budgets are minimal at once. So the design rule is: when the task shape allows it, prefer the path where the two budgets cannot diverge

Collapse
 
hannune profile image
Tae Kim •

We hit this wall earlier this year on a refactoring pipeline where the agent was regularly burning 60+ turns on changes that should've been two or three. Turned out it was re-reading the same file repeatedly across turns because nothing in the loop told it the state hadn't changed. The 2,552 vs 107,000 number on the same diff makes complete sense once you've watched that happen. I'd push on intermediate-state tasks though, where the agent genuinely needs to observe what a prior step actually produced before deciding the next move.

Collapse
 
sergiocorruchaga profile image
Sergio Corruchaga •

That's a great war story — and it matches the mechanism exactly. Nothing in a standard loop carries "state hasn't changed" metadata, so the model re-reads defensively, and every re-read gets re-sent in every subsequent turn. The bill compounds twice: once for the read, once for the trajectory.

On intermediate-state tasks: fully agreed — it's in the limitations section for a reason. When step N+1 genuinely depends on observing what step N produced, a loop is the right tool. What I want to explore for v0.3+ is whether that observation can be bounded and fresh instead of cumulative. The compiler-feedback retry is one deliberate "observe, then act" cycle where the model sees only the error, not the whole trajectory. Chained tasks might work the same way: run 1 merges behind the gate, run 2 starts with a clean context against the new state. Git history becomes the memory; the trajectory never grows.

How did you end up fixing the re-reading in your pipeline — caching, state digests, something else?

Collapse
 
deanlee profile image
Dean Lee •

The 6.6x variance between 32K and 214K tokens on identical tasks points to the real financial hazard of unconstrained agent loops. Teams usually budget agents using the average token burn of successful runs, but re-sending the entire trajectory on every tool call makes the cost function convex in turn count. If a run hits three extra diagnostic steps, you do not pay a linear surcharge. You re-bill the entire prior context three additional times.

Putting a deterministic gate around the edit boundary changes the payoff structure completely. Instead of selling the model a free option to wander through the repo at spot API rates, the harness forces the agent to price each state transition against a fixed budget. The reason providers rarely default to that pattern is straightforward. When the client harness wanders, the platform bills every redundant forward pass, while the user absorbs all the downside variance.

Collapse
 
sergiocorruchaga profile image
Sergio Corruchaga •

The convexity point is the sharpest framing of this I've seen — and you're right that it's not linear: three extra turns don't cost three extra units, they re-bill the entire prior context three more times. Average-based budgeting hides exactly that tail.

One qualification on the provider-incentive point: I'd separate "rarely default" from "never address". DeepSeek's own harness (dsh) ships a Minimal mode and PTC — one TypeScript program instead of N loop turns — so at least some providers are actively sanding the loop's cost down. But the asymmetry you describe is real: the wandering happens at spot rates on the client's bill, and the variance lands entirely on the user. A deterministic gate moves that variance to zero at the edit boundary — you can't fix the model, but you can fix the payoff structure around it.

Worth noting the 6.6× happened at DeepSeek prices — fractions of a cent. The same variance at frontier-model rates is where it stops being an anecdote and starts being a budget meeting.

Do you know of teams actually budgeting agents per-turn or with hard turn caps, rather than on averages? Genuinely curious how this gets handled in practice.

Collapse
 
build996 profile image
build996 •

The token number has a second edge on free tiers that your table would not show. Measuring Groq's free tier, the 8000 TPM budget counts the max_tokens you declare, not the tokens you actually get back: a 20-token prompt with max_tokens 8192 is rejected 413, "Limit 8000, Requested 8271", before generating anything at all. So an agent loop does not only burn more tokens than a deterministic pipeline, it also has to declare a large ceiling on every hop, and on a metered free tier that ceiling is charged whether or not it gets used. Your 2,552 vs 34,600 gap probably understates the difference in how many requests each approach can actually push through per minute.

Collapse
 
sergiocorruchaga profile image
Sergio Corruchaga •

Great catch — and it's an axis my table doesn't capture: I measured consumed tokens (the API usage field), not the declared ceiling. And you're right that architecture matters even more under that rule: an agent loop doesn't know how much it will write on each hop, so it has to declare large ceilings just in case, and on a metered tier it pays that ceiling whether it uses it or not.
A deterministic pipeline's calls have structurally small outputs: the selector returns a file list (~300 tokens), the proposer returns a bounded diff. Small declared ceilings → many more requests per minute on the same budget. So yes: 2,552 vs 34,600 measures consumed tokens; in actual requests-per-minute under a metered quota, the difference is probably even larger.
And it compounds with the effect discussed above: the loop not only resends a growing history on every hop — it also declares a large ceiling on each one. I'll measure this as its own axis in the mixed benchmark I'm preparing (bounded tasks, symptomatic ones, and one genuinely blind task), including free-tier throughput. Thanks for the 413 detail — I didn't know that about Groq.

Collapse
 
ahmetozel profile image
Ahmet Özel •

The token table is compelling, but the task is the most favourable possible case for a deterministic pipeline: add a pure function with a fully specified signature to a named file. No ambiguity about where it goes or what correct means.

The number I would want next is where the gate stops being able to route and has to hand back to the model, measured on a change that spans three files and depends on existing conventions, plus how often the deterministic path refuses or produces something subtly wrong. 42x on the easy half of the distribution is still genuinely valuable, but the honest framing is a router with a fast path rather than a replacement for the agent, and framing it that way makes the result harder to argue with.

Collapse
 
sergiocorruchaga profile image
Sergio Corruchaga •

Fair on one thing: the benchmark distribution skews bounded, and I'll own that — as I said to TokenLat above, the 42× is clean precisely because the tasks were well-defined, and the unit that survives everything is tokens per successful edit with failures priced in.
The numbers you ask for are exactly the ones I want to measure next. The mixed benchmark I'm preparing has three task shapes — bounded, symptomatic ("X is broken", no file named), and one genuinely blind task ("checkout is slow") — and it will count how often the deterministic path resolves, retries, or has to escalate. Multi-file changes that depend on repo conventions are their own task class in it. Including the losses, same as I did here with T9 and T10, where the agent won by creating new files.
One nuance on "refuses or produces something subtly wrong": the harness today doesn't refuse tasks — the gate refuses changes. If a diff doesn't compile, it doesn't land, and the compiler error goes back to the proposer as new information for a cheap retry. What the gate can't see yet is semantic truth — a patch that compiles but is subtly wrong. That's why any fuzzy-matched patch escalates to extra verification in v0.3, and why v0.5 adds tests and performance budgets to the gate.
And I accept the framing: it's a router with a fast path. In fact, the roadmap is a router — a preflight classifier that decides without spending tokens: fast path when you know what and where, a planner call when you know what but not where, a bounded chain (instrument → execute → read the measurement → fix) when the information has to be manufactured, and an agent — behind the gate, under a hard token cap — only when none of that applies. The difference from "just another agent" is that every path, fast or slow, goes through the same gate, and every path has a cost ceiling.

Collapse
 
jo-do profile image
Jo Do •

The determinism split is the right architecture: let the model propose semantics, let the gate own syntax and placement.

One thing I'd stress-test is the 14-42x comparison. Agent runs also spend tokens on retry and exploration that the harness skips, so part of that ratio buys failure recovery. The fair number is quality-matched runs at the same final test state, and it's worth publishing that methodology, because token ratios are exactly the kind of figure that gets quoted without the matching condition attached. Also curious how the gate handles a correct edit aimed at the wrong file: deterministic placement needs a resolution rule for ambiguity, and that rule is where the remaining variance will hide.

Collapse
 
sergiocorruchaga profile image
Sergio Corruchaga •

Both points are exactly the ones to measure. On the first, a nuance and an agreement: the ratio already includes the agent's retries (that's why its numbers are 107K–214K and not 30K) and ours too (T3 needed a second round, and it's counted). But you're right about methodology: the honest figure is quality-matched runs — same final test state — and that condition must be published next to the number, or the number is worthless. There's an asymmetry I like there: on our side, the matching is enforced by construction — a run only counts if the gate is green. On the agent's side we reviewed final states by hand, and when they didn't match we documented it as a loss (T9: hallucination + unauthorized fix; T10: unrequested new file). The mixed benchmark will publish that methodology explicitly.
On the correct edit aimed at the wrong file: the resolution rule today is uniqueness or refusal. Every change carries a SEARCH anchor that must match exactly once in the named file: zero matches, or two or more, and the local editor refuses to apply, reports, and master stays untouched. There's no placement guessing. The case that does escape that rule — a unique anchor in the wrong file, compiles but is semantically wrong — is exactly the variance you point at, and compilation alone can't cover it. That's why the roadmap adds three layers over that gap: in v0.3, any fuzzy-matched patch escalates to extra verification; in v0.4, a semantic reviewer inside the worktree (a second call that judges the applied diff against the spec, with rejections feeding bounded retries) plus a user-editable spec shown before execution; and in v0.5, tests and performance budgets inside the gate. The placement rule is "uniqueness or refusal"; semantics is defended in layers.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.