DEV Community

Cover image for Your LLM bill is 80% hidden thinking tokens. One parameter fixes it.
Qubax AI
Qubax AI

Posted on Originally published at qubax.ai

Your LLM bill is 80% hidden thinking tokens. One parameter fixes it.

I asked three cheap, popular models the same simple question: "Write a 150-word explanation of how HTTP caching headers work, for a junior developer."

Each one answered in about 150 words. Each one billed me for 800 to 2,500 output tokens.

The difference is reasoning tokens: "thinking" the model does before it answers. You pay for every one of them at the output rate, they never appear in the response, and most dashboards don't break them out. For a task like this one, they are pure waste.

Here is how to measure them in your own app, and the one parameter that removed them for me.

The measurement

Every OpenAI-compatible API reports hidden tokens in usage.completion_tokens_details.reasoning_tokens. This script calls each model three times, first with default settings and then with reasoning_effort turned down, and prints the medians.

"""hidden_tokens.py: how much of your output bill is hidden thinking?"""
import os, statistics, time
from openai import OpenAI

client = OpenAI(base_url="https://api.qubax.ai/v1",
                api_key=os.environ["QUBAX_API_KEY"])
PROMPT = ("Write a 150-word explanation of how HTTP caching "
          "headers work, for a junior developer.")
# model -> (output $ per 1M tokens, lowest reasoning_effort it accepts)
MODELS = {
    "glm-5.3-flash": (0.0426, "low"),
    "deepseek-v4.1-flash": (0.058824, "none"),
    "claude-haiku-5.5": (0.340909, "low"),
}

def call(model, **extra):
    t0 = time.perf_counter()
    r = client.chat.completions.create(
        model=model, max_tokens=4000, temperature=0,
        messages=[{"role": "user", "content": PROMPT}], **extra)
    d = r.usage.completion_tokens_details
    hidden = (getattr(d, "reasoning_tokens", 0) or 0) if d else 0
    words = len((r.choices[0].message.content or "").split())
    return {"out": r.usage.completion_tokens, "hidden": hidden,
            "words": words, "s": time.perf_counter() - t0}

for model, (price, effort) in MODELS.items():
    for label, extra in [("default", {}),
                         (f"effort={effort}", {"reasoning_effort": effort})]:
        rows = [call(model, **extra) for _ in range(3)]
        med = lambda k: statistics.median(r[k] for r in rows)
        per_1k = med("out") * price / 1e6 * 1000
        print(f"{model:<22}{label:<14}{med('out'):>6.0f} out "
              f"{med('hidden'):>6.0f} hidden {med('words'):>4.0f} words "
              f"{med('s'):>5.1f}s  ${per_1k:.4f}/1k calls")
Enter fullscreen mode Exit fullscreen mode

Change base_url and the key to point at any OpenAI-compatible endpoint. The model names above are the ones I tested.

The results (Oct 11, 2026, median of 3 runs)

Model Setting Billed output tokens Hidden Words shown Time Cost per 1k calls
GLM 5.3 Flash default 2,483 2,256 152 39.6s $0.1058
GLM 5.3 Flash reasoning_effort="low" 232 0 159 8.4s $0.0099
DeepSeek V4.1 Flash default 1,525 1,293 150 12.8s $0.0897
DeepSeek V4.1 Flash reasoning_effort="none" 200 0 132 3.6s $0.0118
Claude Haiku 5.5 default 1,652 1,354 150 7.4s $0.5632
Claude Haiku 5.5 reasoning_effort="low" 289 0 129 2.7s $0.0985

The same answer, 83–91% cheaper and 3–5× faster, from one parameter.

Hidden-token counts vary a lot between runs, so I ran the whole script again. On the second run, defaults were 74–90% hidden, and the setting still cut cost by 69–91%. The exact numbers move; the direction never did.

At default settings, 82–91% of what I paid for on that run was text I never saw. Latency follows the bill: GLM 5.3 Flash went from 40 seconds to 8 because it stopped writing a hidden essay before the real one.

Two things that surprised me

1. Values aren't portable. DeepSeek accepts "none" and goes to zero hidden tokens. On GLM, "low" already gave zero. On the same models, "minimal" sometimes kept a few hundred hidden tokens. Measure each model; don't assume.

2. Some models ignore it. In a separate run, GPT-5.5 kept 400–600 hidden tokens whatever value I sent. If a model doesn't respond to the knob, the fix is choosing a different model for that job, not tuning.

When to keep thinking on

Hidden reasoning isn't always waste. It earns its cost on multi-step math, tricky refactors, planning agents, and anything where a wrong answer is expensive. It's waste on:

  • classification and routing ("which queue does this ticket go to?")
  • extraction to JSON
  • summaries, rewrites, translations
  • short chat replies and UI copy

A simple rule that works: set reasoning_effort per call site, not per app. Your "summarize this email" helper and your "plan this database migration" agent shouldn't share a setting.

FAST = {"reasoning_effort": "low"}   # extraction, labels, summaries
DEEP = {}                            # planning, math, hard code

def summarize(text):
    return client.chat.completions.create(
        model="claude-haiku-5.5", **FAST,
        messages=[{"role": "user", "content": f"Summarize in 2 lines:\n{text}"}],
    ).choices[0].message.content
Enter fullscreen mode Exit fullscreen mode

Check your own traffic

Log completion_tokens_details.reasoning_tokens next to completion_tokens for a day. If hidden tokens are more than half your output on endpoints that do simple work, you've found the cheapest optimization you'll make this year.

I ran these tests through Qubax, an OpenAI-compatible API with 400+ models (often well below OpenRouter prices) where you can switch models by changing one string, so the script runs unchanged on every model. Prices in the table are the per-token rates on Oct 11, 2026; live rates are on the price index.

FAQ

What are reasoning tokens?

They're tokens a model generates while "thinking" before it writes the visible answer. They're billed at the output-token price but never returned in the response text. You can see the count in usage.completion_tokens_details.reasoning_tokens.

Does lowering reasoning_effort make answers worse?

For simple tasks like summaries, extraction and labels, it made no visible difference in my tests. For multi-step math, planning or hard code, keep reasoning on and measure before switching it off.

Which reasoning_effort value should I use?

It depends on the model. In these tests, DeepSeek V4.1 Flash went to zero hidden tokens with "none", and GLM 5.3 Flash and Claude Haiku 5.5 with "low". Test each model you use.

References

What's the worst hidden-token ratio you've found in production? I'd love to see numbers from other stacks.


Originally published at qubax.ai. Qubax gives you every top AI model with one API key, cheaper than OpenRouter.

Top comments (0)