DEV Community

Qweezyy
Qweezyy

Posted on

One field in the request made our agent 3x cheaper and 8x faster

TL;DR. Reasoning models decide by themselves how long to think if you don't tell them. Our agent didn't — and on hard tasks the model sometimes thought for 14 minutes and 33K tokens in a single step. One field in the request body ("reasoning": {"effort": "low"}) made a task 3x cheaper and 7–8x faster, and it solved more tasks, not fewer: 12 of 12 instead of 6 of 7. The best setting in the end was neither "always low" nor "always high" but "low, and high right after a failing test". Below: how we measured it, the tables, the code, and what it doesn't fix.

About this post. The project is Altair, an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the experiments and the charts were prepared by that same AI assistant; I reviewed it and stand behind it. Every number comes from our runs.

This is the third post about the project: the first is about snapshots and running tests before "done", the second about cutting a browser agent's tokens by 58%.

How it started

We ran the agent on a cheap reasoning model (glm-5.3-flash) and read the provider's logs. On one task the very first step took 14 minutes. The provider had time to bill it: 32,978 reasoning tokens in one step, and no answer. Our loop guard cut the stream at 160,000 characters.

The cause was mundane. Reasoning models have a "how much to think" knob — reasoning_effort at OpenAI, reasoning.effort at OpenRouter, thinking in Z.ai-style APIs. If you don't send it, the provider or the model decides; OpenRouter's docs say as much. Our agent sent nothing.

How we measured

Easy tasks (fix a bug, answer a question about code) were solved 100% of the time under any setting, so they show no difference. We wrote six hard tasks with hidden tests — the agent never sees them, they run after it says "done":

  • an LRU cache with a time-to-live (11 hidden edge-case tests);
  • fixing a config parser from seven user bug reports;
  • renaming a function across a multi-file package and adding a parameter;
  • the 95th percentile of latencies from logs, with exclusions;
  • parsing "1.5M", "200K", "10 тыс";
  • business days between dates with holidays, fast over a hundred years.

Each task was first solved with reference code to make sure the hidden tests were fair. The agent is the real Altair, run through its CLI (altair -p). A proxy between agent and provider logged every request: size, cached tokens, reasoning tokens, time. Cost is what OpenRouter billed.

Results

a1-cost-en
Cost of one hard task

Setting Solved Cost per task Time per task
No level (before) 6 of 7 12.6 m$ 4.1 min
effort = low 12 of 12 4.2 m$ 0.55 min
effort = medium 10 of 12 4.8 m$ 0.64 min
effort = high 12 of 12 6.4 m$ 1.5 min

m$ is a thousandth of a dollar. 12.6 m$ without a level, 4.2 m$ with low: three times less. Time went from 4.1 to 0.55 minutes, 7.5x.

Why "6 of 7" and not "of 12": runs without a level kept looping for 10–14 minutes, and we stopped that series early so as not to burn money. One of the six tasks (parsing "1.5M") sent the model into endless reasoning in almost every series without a level. With any explicit level — never.

Where the money goes:

a1-reasoning-en
Reasoning tokens per task

9,743 reasoning tokens per task on average without a level, 235 with low — 41x fewer. Output tokens cost 3.3x more than input for this model, so on hard tasks reasoning is the main bill.

And on long tasks?

Short tasks are 5–9 steps. We also tried 15–20-step ones: a package with six bugs in different modules, auditing twelve values across sixty files of our own code, and "read eight files in full, then answer questions about them".

a1-long-en
Long tasks

9 of 9 solved in every setting, but without a level a task costs 28 m$ and takes 5.2 minutes; with an explicit level, 11–13 m$ and a little over a minute.

The twist: think hard only when something broke

low is cheap, high is safer. We wanted both, and tried two per-step modes:

  • "think about the plan" — high on the first step, then low;
  • "think after a failure" — low, but high on the step right after tests failed or a tool returned an error.

a1-adaptive-en
When to think harder

Setting (18 runs each) Solved Cost Time
always low 16 of 18 3.78 m$ 0.96 min
always high 17 of 18 5.72 m$ 1.92 min
high on the first step 15 of 18 4.37 m$ 1.93 min
low, high after a failure 17 of 18 3.65 m$ 0.84 min

Careful planning didn't help — it solved the fewest. "Think after a failure" solved as many as always-high, at the cost and speed of low. It makes sense: most of an agent's work is routine (read a file, write a file, run the tests), and thinking pays off where a test just showed the first idea was wrong. One task in 18 is close to noise, honestly, but this mode is no worse than low on all three measures. It's now our default.

How it looks in code

The catch is that every provider has its own field:

def reasoning_extra(level, base_url, model, override="auto"):
    """Request fields for a reasoning level ({} for "default" / unknown levels)."""
    if level not in LEVELS:
        return {}
    kind = dialect(base_url, model, override)
    if kind == "openrouter":
        # OpenRouter takes minimal..high; some endpoints refuse "none"/enabled=false,
        # so the lowest we send is "low" for "minimal" there.
        return {"reasoning": {"effort": "low" if level == "minimal" else level}}
    if kind == "zai":
        # Z.ai-style APIs have an on/off switch: low and minimal turn thinking off.
        return {"thinking": {"type": "disabled" if level in ("minimal", "low") else "enabled"}}
    if kind == "openai":
        return {"reasoning_effort": level}
    return {}
Enter fullscreen mode Exit fullscreen mode

What we stepped on along the way:

  • For this model OpenRouter refuses reasoning: {"enabled": false} and reasoning_effort: "none" with a 400 "Reasoning is mandatory for this endpoint". Only effort: "low" works.
  • A Z.ai-style gateway, the other way round, takes thinking: {"type": "disabled"} and didn't answer reasoning.effort at all.
  • A provider that doesn't know the field may answer 400. We catch that, retry without the field and remember not to send it again.

The agent loop picks the level for each step:

FAILURE_RE = re.compile(r"\b\d+ failed\b|\bFAILED\b|Traceback \(most recent call last\)|AssertionError|"
                        r"\bexit code [1-9]\d*\b|\bSyntaxError\b|[A-Za-z]Error:")

def _reasoning_level(self) -> str | None:
    mode = self.settings.llm_reasoning
    if mode == "default":
        return None
    if mode == "adaptive":
        return "high" if getattr(self, "_round_failed", False) else "low"
    return mode
Enter fullscreen mode Exit fullscreen mode

After each round of tools, _round_failed is true if a tool returned an error or its output has "2 failed", a traceback or a non-zero exit code.

Service calls thought more than the main ones

Two more places turned up in the logs.

The chat title. The agent asks the model for a title from the first message. About 95% of that answer was reasoning: ~190 tokens of "thoughts" for a five-word title. With reasoning off the answer is 12 tokens and comes 1.6–3x faster; all six sample titles were still fine.

The history summary. When the context grows, the agent asks the model to condense the start of the conversation. The answer limit was 600 tokens. The model spent them on reasoning and the summary came back empty in 3 cases out of 4 (finish_reason: length). An empty summary isn't a fold, it's a loss: the agent drops the start of the conversation instead of condensing it. With a 2,000 limit there were no empty summaries.

Service calls now ask for the least reasoning, and the summary limit is 2,500.

What this doesn't fix

  • One model. Everything was measured on glm-5.3-flash. Other models scale their levels differently, and "low" may mean something else. The principle — don't leave the level to the provider — carries over; the exact numbers don't.
  • Small samples. 12–18 runs per setting. A difference of one or two solved tasks is noise. A difference of 3x in cost and time is not.
  • Not every provider has levels. A Z.ai-style gateway only knows on/off. There "high after a failure" means "think fully", which is expensive: on our tasks "always off" on that gateway came out half the price of adaptive (5.9 vs 10.8 m$, 12 of 12 both). If cost matters most, pick "low" there.
  • The tasks are code. For writing, search or analytics the best level may differ.

Check it yourself

The core of the experiment is one agent run with a set level:

# proxy.py adds the level to every agent request to the provider
body = {**{"reasoning": {"effort": "low"}}, **request_body, "model": provider["model"]}
# stand.py: the agent solves the task, then the hidden tests run
subprocess.run([sys.executable, "altair_cli.py", "-p", "--output-format", "json",
                "--mode", "bypass", "--cwd", workspace, task.prompt], env=env, timeout=900)
passed, detail = task.check(workspace, answer)
Enter fullscreen mode Exit fullscreen mode

The whole stand (the logging proxy, tasks with hidden tests and reference solutions, the summary script) and the raw results: research/agent-lab-2026-10 (python lab/facts.py recomputes every number in this post from the saved results, no keys needed). Altair's code: https://github.com/Qweezyy/AltairAgent — the level is in Settings → Agent → Reasoning level (default "Adaptive"), LLM_REASONING in .env, module pc/core/llm/reliability.py.

How do you set the reasoning level in your agents — fixed, per task type, or as the work goes? Have you seen a model think far more than it needs when no level is set?

Top comments (2)

Collapse
 
max_quimby profile image
Max Quimby •

The "low by default, high right after a failing test" policy is the part that'll stick with me — it's the difference between a knob and a controller. We hit the same unbounded-thinking failure: a reasoning model deciding on its own to burn tens of thousands of tokens on one step with nothing to show for it, and a character/loop guard being the only thing that stops the bleed.

Two questions. (1) A failing test is a clean escalation trigger, but it only fires when you have tests. Did you try other signals — repeated no-op steps, the same tool erroring twice, the model re-reading the same file — as "think harder" triggers for tasks without a test harness? (2) The counterintuitive result is that low effort solved more tasks (12/12 vs 6/7). My read is that high effort gives the model enough rope to talk itself out of a correct plan. Did the high-effort failures look like over-reasoning (changing a right answer into a wrong one), or more like timeouts/truncation from the 160k-char guard?

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Useful result. The speedup is a reminder that agent reliability is also a systems problem: trim context, but keep explicit stop conditions, approval gates, and a record of what the tool actually did. How are you measuring wrong or over-broad actions alongside latency and cost? iin1004h2228