TL;DR. Reasoning models decide by themselves how long to think if you don't tell them. Our agent didn't — and on hard tasks the model sometimes thought for 14 minutes and 33K tokens in a single step. One field in the request body ("reasoning": {"effort": "low"}) made a task 3x cheaper and 7–8x faster, and it solved more tasks, not fewer: 12 of 12 instead of 6 of 7. The best setting in the end was neither "always low" nor "always high" but "low, and high right after a failing test". Below: how we measured it, the tables, the code, and what it doesn't fix.
About this post. The project is Altair, an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the experiments and the charts were prepared by that same AI assistant; I reviewed it and stand behind it. Every number comes from our runs.
This is the third post about the project: the first is about snapshots and running tests before "done", the second about cutting a browser agent's tokens by 58%.
How it started
We ran the agent on a cheap reasoning model (glm-5.3-flash) and read the provider's logs. On one task the very first step took 14 minutes. The provider had time to bill it: 32,978 reasoning tokens in one step, and no answer. Our loop guard cut the stream at 160,000 characters.
The cause was mundane. Reasoning models have a "how much to think" knob — reasoning_effort at OpenAI, reasoning.effort at OpenRouter, thinking in Z.ai-style APIs. If you don't send it, the provider or the model decides; OpenRouter's docs say as much. Our agent sent nothing.
How we measured
Easy tasks (fix a bug, answer a question about code) were solved 100% of the time under any setting, so they show no difference. We wrote six hard tasks with hidden tests — the agent never sees them, they run after it says "done":
- an LRU cache with a time-to-live (11 hidden edge-case tests);
- fixing a config parser from seven user bug reports;
- renaming a function across a multi-file package and adding a parameter;
- the 95th percentile of latencies from logs, with exclusions;
- parsing "1.5M", "200K", "10 тыс";
- business days between dates with holidays, fast over a hundred years.
Each task was first solved with reference code to make sure the hidden tests were fair. The agent is the real Altair, run through its CLI (altair -p). A proxy between agent and provider logged every request: size, cached tokens, reasoning tokens, time. Cost is what OpenRouter billed.
Results
| Setting | Solved | Cost per task | Time per task |
|---|---|---|---|
| No level (before) | 6 of 7 | 12.6 m$ | 4.1 min |
effort = low |
12 of 12 | 4.2 m$ | 0.55 min |
effort = medium |
10 of 12 | 4.8 m$ | 0.64 min |
effort = high |
12 of 12 | 6.4 m$ | 1.5 min |
m$ is a thousandth of a dollar. 12.6 m$ without a level, 4.2 m$ with low: three times less. Time went from 4.1 to 0.55 minutes, 7.5x.
Why "6 of 7" and not "of 12": runs without a level kept looping for 10–14 minutes, and we stopped that series early so as not to burn money. One of the six tasks (parsing "1.5M") sent the model into endless reasoning in almost every series without a level. With any explicit level — never.
Where the money goes:
9,743 reasoning tokens per task on average without a level, 235 with low — 41x fewer. Output tokens cost 3.3x more than input for this model, so on hard tasks reasoning is the main bill.
And on long tasks?
Short tasks are 5–9 steps. We also tried 15–20-step ones: a package with six bugs in different modules, auditing twelve values across sixty files of our own code, and "read eight files in full, then answer questions about them".
9 of 9 solved in every setting, but without a level a task costs 28 m$ and takes 5.2 minutes; with an explicit level, 11–13 m$ and a little over a minute.
The twist: think hard only when something broke
low is cheap, high is safer. We wanted both, and tried two per-step modes:
-
"think about the plan" —
highon the first step, thenlow; -
"think after a failure" —
low, buthighon the step right after tests failed or a tool returned an error.
| Setting (18 runs each) | Solved | Cost | Time |
|---|---|---|---|
always low
|
16 of 18 | 3.78 m$ | 0.96 min |
always high
|
17 of 18 | 5.72 m$ | 1.92 min |
high on the first step |
15 of 18 | 4.37 m$ | 1.93 min |
low, high after a failure |
17 of 18 | 3.65 m$ | 0.84 min |
Careful planning didn't help — it solved the fewest. "Think after a failure" solved as many as always-high, at the cost and speed of low. It makes sense: most of an agent's work is routine (read a file, write a file, run the tests), and thinking pays off where a test just showed the first idea was wrong. One task in 18 is close to noise, honestly, but this mode is no worse than low on all three measures. It's now our default.
How it looks in code
The catch is that every provider has its own field:
def reasoning_extra(level, base_url, model, override="auto"):
"""Request fields for a reasoning level ({} for "default" / unknown levels)."""
if level not in LEVELS:
return {}
kind = dialect(base_url, model, override)
if kind == "openrouter":
# OpenRouter takes minimal..high; some endpoints refuse "none"/enabled=false,
# so the lowest we send is "low" for "minimal" there.
return {"reasoning": {"effort": "low" if level == "minimal" else level}}
if kind == "zai":
# Z.ai-style APIs have an on/off switch: low and minimal turn thinking off.
return {"thinking": {"type": "disabled" if level in ("minimal", "low") else "enabled"}}
if kind == "openai":
return {"reasoning_effort": level}
return {}
What we stepped on along the way:
- For this model OpenRouter refuses
reasoning: {"enabled": false}andreasoning_effort: "none"with a 400 "Reasoning is mandatory for this endpoint". Onlyeffort: "low"works. - A Z.ai-style gateway, the other way round, takes
thinking: {"type": "disabled"}and didn't answerreasoning.effortat all. - A provider that doesn't know the field may answer 400. We catch that, retry without the field and remember not to send it again.
The agent loop picks the level for each step:
FAILURE_RE = re.compile(r"\b\d+ failed\b|\bFAILED\b|Traceback \(most recent call last\)|AssertionError|"
r"\bexit code [1-9]\d*\b|\bSyntaxError\b|[A-Za-z]Error:")
def _reasoning_level(self) -> str | None:
mode = self.settings.llm_reasoning
if mode == "default":
return None
if mode == "adaptive":
return "high" if getattr(self, "_round_failed", False) else "low"
return mode
After each round of tools, _round_failed is true if a tool returned an error or its output has "2 failed", a traceback or a non-zero exit code.
Service calls thought more than the main ones
Two more places turned up in the logs.
The chat title. The agent asks the model for a title from the first message. About 95% of that answer was reasoning: ~190 tokens of "thoughts" for a five-word title. With reasoning off the answer is 12 tokens and comes 1.6–3x faster; all six sample titles were still fine.
The history summary. When the context grows, the agent asks the model to condense the start of the conversation. The answer limit was 600 tokens. The model spent them on reasoning and the summary came back empty in 3 cases out of 4 (finish_reason: length). An empty summary isn't a fold, it's a loss: the agent drops the start of the conversation instead of condensing it. With a 2,000 limit there were no empty summaries.
Service calls now ask for the least reasoning, and the summary limit is 2,500.
What this doesn't fix
-
One model. Everything was measured on
glm-5.3-flash. Other models scale their levels differently, and "low" may mean something else. The principle — don't leave the level to the provider — carries over; the exact numbers don't. - Small samples. 12–18 runs per setting. A difference of one or two solved tasks is noise. A difference of 3x in cost and time is not.
- Not every provider has levels. A Z.ai-style gateway only knows on/off. There "high after a failure" means "think fully", which is expensive: on our tasks "always off" on that gateway came out half the price of adaptive (5.9 vs 10.8 m$, 12 of 12 both). If cost matters most, pick "low" there.
- The tasks are code. For writing, search or analytics the best level may differ.
Check it yourself
The core of the experiment is one agent run with a set level:
# proxy.py adds the level to every agent request to the provider
body = {**{"reasoning": {"effort": "low"}}, **request_body, "model": provider["model"]}
# stand.py: the agent solves the task, then the hidden tests run
subprocess.run([sys.executable, "altair_cli.py", "-p", "--output-format", "json",
"--mode", "bypass", "--cwd", workspace, task.prompt], env=env, timeout=900)
passed, detail = task.check(workspace, answer)
The whole stand (the logging proxy, tasks with hidden tests and reference solutions, the summary script) and the raw results: research/agent-lab-2026-10 (python lab/facts.py recomputes every number in this post from the saved results, no keys needed). Altair's code: https://github.com/Qweezyy/AltairAgent — the level is in Settings → Agent → Reasoning level (default "Adaptive"), LLM_REASONING in .env, module pc/core/llm/reliability.py.
How do you set the reasoning level in your agents — fixed, per task type, or as the work goes? Have you seen a model think far more than it needs when no level is set?




Top comments (2)
The "low by default, high right after a failing test" policy is the part that'll stick with me — it's the difference between a knob and a controller. We hit the same unbounded-thinking failure: a reasoning model deciding on its own to burn tens of thousands of tokens on one step with nothing to show for it, and a character/loop guard being the only thing that stops the bleed.
Two questions. (1) A failing test is a clean escalation trigger, but it only fires when you have tests. Did you try other signals — repeated no-op steps, the same tool erroring twice, the model re-reading the same file — as "think harder" triggers for tasks without a test harness? (2) The counterintuitive result is that low effort solved more tasks (12/12 vs 6/7). My read is that high effort gives the model enough rope to talk itself out of a correct plan. Did the high-effort failures look like over-reasoning (changing a right answer into a wrong one), or more like timeouts/truncation from the 160k-char guard?
Useful result. The speedup is a reminder that agent reliability is also a systems problem: trim context, but keep explicit stop conditions, approval gates, and a record of what the tool actually did. How are you measuring wrong or over-broad actions alongside latency and cost? iin1004h2228