DEV Community

Cover image for I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.
Abeera Lodhi
Abeera Lodhi Subscriber

Posted on

I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?"

With reasoning effort set to none, it replied:

Cleo

FINAL ANSWER: Cleo
Enter fullscreen mode Exit fullscreen mode

18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave the same wrong answer, word for word, on the second repeat.

At high it spent 1,333 tokens, cost about $0.006, and said Fay. That looks like an argument for always choosing high.

On another item I asked the same model, also at high, to count the tools in "You have a chisel and a drill." It answered 10,016.

Both replies came from the same model, setting and benchmark. That is the whole post in two examples: the reasoning dial sometimes matters a lot. Most of the time it only makes the bill bigger.

TL;DR: Reasoning Dial is a Kaggle benchmark that changes one API parameter, reasoning_effort (none / low / medium / high), and keeps everything else fixed: the same 60 code-generated questions, the same prompt and the same deterministic grader. Over 1,800 graded calls on 4 models, high produced exactly one statistically real accuracy gain: gpt-5.4-mini on logic puzzles, 15% → 97.5%. In 7 of 12 model × task cells, accuracy did not change while cost per correct answer rose ×1.5 to ×3.4. The dial also isn't one instrument: one model spends ×14 more tokens at high, one rejects none with an HTTP 400, and one thinks at none anyway. Hypotheses and analysis were pre-registered before the main run. Total main-run cost: $4.20.


My Benchmark

GitHub logo Abeera81 / reasoning-dial

Reasoning Dial: what the reasoning-effort setting changes in accuracy, tokens and cost (Kaggle Benchmarks)

Reasoning Dial

Reasoning Dial is a Kaggle benchmark about one setting developers choose blind: reasoning effort (none / low / medium / high). It holds everything else fixed (same items, same prompt, same grader) and measures what turning the dial actually changes: accuracy, output tokens, and cost per correct answer, on three kinds of task. It also asks whether the dial is even the same instrument across vendors.

Design
















Task families
Deduce: 7-person, 7-day scheduling puzzles with exactly one solution (checked by brute force). Arith: 6-step word problems with a percentage and an exact division (control family). Distract: count the objects you own while ignoring distractors (after Gema et al. 2025). 20 items per family.
Dial levels
none, low, medium, high
…





The Reasoning Dial leaderboard on Kaggle


What I Benchmarked

Almost every modern model API has a version of this parameter: reasoning_effort, thinking, reasoning. Developers set it every day, usually by feel:

  • "It's a hard question, so high."
  • "It's in production, so low."

Almost nobody measures what moving it actually changes.

So I built a benchmark where the only variable is the dial. Same items, same prompt suffix, same parser, same grader. For each model, the four Kaggle tasks differ only in the LEVELS value and the two task-name lines. (The push CLI reads task names as string literals, so one line of difference wasn't possible.)

I measured three things at each setting:

  1. Accuracy. Strict and deterministic: a regex reads the last FINAL ANSWER: line. There is no LLM judge anywhere.
  2. Output tokens. Reasoning tokens are included, because you pay for them whether or not you can see them.
  3. Cost per correct answer. The per-call nanodollar cost that Kaggle's model proxy reports, divided by the number of correct answers.

Three kinds of question, chosen to disagree with each other

Family What it is Why it's here Example
Deduce 7 people, 7 days, 6–12 clues (before, immediately before, not on, not adjacent). Brute force over all 7! orders confirms exactly one solution. Where more thinking should help (H3) "Who gives the talk on Friday?" → Fay
Arith Six-step word problems with a percentage step and an exact division; 4–6 digit answers The control. Every model scored 100% at none in calibration, so it shows what the dial costs when there is nothing left to gain (H4) A courier's fuel bill with 10% tax → 11880
Distract "You have a chisel and a drill." followed by 1–3 irrelevant numbers: a shop's sales, someone's homework, a code snippet Where more thinking might hurt. This follows Gema et al. 2025, Inverse Scaling in Test-Time Compute (H2) "Calculate how many tools you have." → 2

Here is a real Distract item from the locked test set, exactly as the models saw it:

You have a chisel and a drill.

A shop in Easton sold 8,377 levels last week.

A friend shows you this code:
```python
stock = ["saw", "hammer", "screwdriver"]
print(len(stock) * 3)
```

A classmate is working on a homework problem: if 44 crates each hold 37 saws, how many are there altogether?

Question: Calculate how many tools you have.

Answer instruction: Answer with a whole number.

End your response with a final line in exactly this format:
FINAL ANSWER: <answer>
Enter fullscreen mode Exit fullscreen mode

A person answers 2 without reading past the first line. That is exactly why the item is useful.

I wrote down my guesses before running anything

Before the main run, I committed docs/PREREGISTRATION.md. It contains four hypotheses, the frozen prompt and parser, and the statistics: a paired Wilcoxon test, Holm correction across all 12 model × family tests, and a 5,000-rep item-cluster bootstrap. In the same commit I locked the 60 test items under a SHA-256 hash (90dd5498…0042). The analysis script refuses to run if the hash doesn't match.

Hypothesis
H1 The dial is not one instrument. Some models scale tokens with it; others ignore it, collapse levels, or reject a level.
H2 More thinking can hurt on distractor counting.
H3 More thinking helps on deduction.
H4 Where accuracy doesn't rise, the bill still does.

Every change made after the lock is logged, dated and explained in that file. There are three, and none of them touches the items, the prompt, the parser or the statistics.


How I built it on Kaggle Benchmarks

Pipeline diagram: generate.py (seeded generators, answers computed by code, logic puzzles brute-force checked for one solution) → test_items.json (60 items, SHA-256 locked and pre-registered) → build_tasks.py (inlines items and parser into 4 task files that differ only in the level) → Kaggle Benchmarks (4 tasks, reasoning=<level>, 8 calls in parallel) → parse.py (last FINAL ANSWER line, deterministic grading, no LLM judge) → analyze.py (paired bootstrap, Wilcoxon, Holm across 12 tests, summary and figures)

Everything in this post comes out of this pipeline. The generator uses a fixed seed, the test set is locked by hash, and the grader is a regex. The only thing that changes between the four Kaggle tasks is the reasoning-effort value.

The whole experiment is possible because of one feature of the kaggle-benchmarks library: llm.prompt() accepts a reasoning= argument, and Kaggle's model proxy translates it into each vendor's own reasoning_effort. One line of code runs the same experiment on OpenAI, Anthropic, Google and open-weight models. Here is the core of every call:

# src/dial/task_template.py.txt  (inlined into each Kaggle task by build_tasks.py)
with kbench.chats.new(name) as chat:
    try:
        raw = llm.prompt(prompt, reasoning=level, seed=SEED_BASE + repeat)
        traces = kbench.last_reasoning_traces()
    except Exception as e:
        raw, traces = None, None
        rec["error"] = f"{type(e).__name__}: {e}"   # a provider 400 becomes an "api_error", which is data for H1
    u = chat.usage
rec["output_tokens"] = u.output_tokens               # hidden reasoning tokens included
rec["cost_nanodollars"] = u.total_cost_nanodollars   # what Kaggle's proxy actually charged
Enter fullscreen mode Exit fullscreen mode

chat.usage is the reason this benchmark can measure cost instead of guessing it. Every call records its own token count and its own cost in nanodollars, and those records end up in Kaggle's run files.

Each dial-<level> task fans out all 120 calls (60 items × 2 repeats) as a nested task:

records = _records(dial_call.evaluate(
    llm=[llm], evaluation_data=df, n_jobs=8, on_failure="continue"))
Enter fullscreen mode Exit fullscreen mode

on_failure="continue" matters. When gpt-oss-120b rejects none, the run doesn't crash: it records the rejection, and the rejection becomes one of the findings.

Grading is deliberately boring. The prompt ends with "End your response with a final line in exactly this format: FINAL ANSWER: ", and the parser takes the last match:

# src/dial/parse.py
_MARKER = re.compile(
    r"FINAL[ \t]+ANSWER[ \t]*[*_`]*[ \t]*[::][ \t]*(.*)$",
    re.IGNORECASE | re.MULTILINE,
)
Enter fullscreen mode Exit fullscreen mode

It handles **FINAL ANSWER:** 42 and full-width colons, strips <think> blocks, and is covered by the repo's test suite. I didn't use the library's structured output (schema=) because it breaks when reasoning is on: the proxy returns <think>…</think>{json} and JSON parsing fails. Plain text plus my own parser was the only approach that worked the same way across all four vendors.

The Kaggle CLI handled the rest of the loop:

kaggle b t push dial-high -f tasks/dial_high.py --wait        # also runs it once on Kaggle's default model
kaggle b t run dial-high -m gpt-5.4-mini-2026-03-17           # one model, one dial position
kaggle b t download dial-high -o results/main                 # per-call records → analysis/analyze.py
Enter fullscreen mode Exit fullscreen mode

The first command had a side effect I didn't plan for. Pushing a task runs it once on Kaggle's default model, gemini-3.7-flash. Those validation runs were full, clean runs of the locked items, so Gemini joined the lineup without me choosing it.

The dial-high task on Kaggle

The part where Kaggle's quota taught me something

Kaggle reserves each call's maximum possible cost against your quota before the call runs: the input cost plus 128,000 output tokens at the model's output price. For gpt-5.4-mini that is about $0.58 reserved per call, even when the call actually costs $0.0002.

With 8 calls in parallel, one task reserves about $4.61. My first gpt-5.4-mini run launched several tasks at once, and 378 of 480 calls came back as HTTP 403 quota refusals. Real spend was nowhere near the cap.

I made a rule and wrote it into the pre-registration: a quota 403 is missing data, never model behavior. I discarded those runs (they're listed in results/superseded.json, and the analysis excludes them) and re-ran gpt-5.4-mini one task at a time after the quota reset. The re-run was clean: 0 refusals, 0 API errors.

  • gpt-5.4-mini: $0.044 at none, $0.371 at low, $0.588 at medium and $0.713 at high, $1.72 in total.
  • Claude Sonnet 5: dropped from the lineup. Its reservation is about $1.28 per call, and 8 parallel calls (≈$10.24) exceed the $10 daily quota on their own. The analysis reports it as "dropped (platform quota reservation)" rather than leaving it out silently.

Models Tested

Model Levels run Graded calls Main-run cost
gpt-5.4-mini-2026-03-17 none, low, medium, high 480 $1.716
gemini-3.7-flash none, low, medium, high 480 $2.034
claude-haiku-5.5 none, low, medium, high 480 $0.186
gpt-oss-120b low, medium, high (rejects none with HTTP 400) 360 $0.263
claude-sonnet-5 n/a 0 dropped (quota reservation, see above)

1,800 graded calls in total: 60 items × 2 repeats × 15 model-levels. There were 0 API errors and 0 quota refusals in the final data, and 1,964,378 output tokens. The main run cost $4.20; the whole project, including the pilot and two calibration rounds, cost about $6.38 of Kaggle quota.

A note on control: Kaggle's proxy silently drops temperature, so I can't claim temperature 0, and the Google models also drop the seed. That is why every cell has two repeats, and why the analysis reports a noise floor (§7 of summary.md).


Findings

1. The dial is not one instrument (H1: ✅ supported)

Token ladder: median output tokens per call at each reasoning level, one line per model, log scale. gpt-5.4-mini rises from 18 to 254; gpt-oss-120b from 292 at low to 1,340 at high; claude-haiku-5.5 from 150 to 298; gemini-3.7-flash starts at 453 at none, is flat to low, and rises to 1,013 at high

Median output tokens per call (log scale). The same parameter name produces four different behaviors.

The same parameter means something different for every vendor:

Model Median output tokens, none → low → medium → high high ÷ lowest Label (pre-registered rule)
gpt-5.4-mini 18 → 168 → 195 → 254 ×14.1 honors
gpt-oss-120b rejected → 292 → 446 → 1,340 ×4.6 (from low) honors, rejects none
claude-haiku-5.5 150 → 162 → 171 → 298 ×2.0 honors
gemini-3.7-flash 453 → 446 → 566 → 1,013 ×2.2 neither; none ≈ low

Two of these surprised me:

  • none doesn't mean "no thinking" on Gemini. At none it spends a median 453 output tokens, more than gpt-5.4-mini spends at high (254). It returns no reasoning trace at none or low, and its one-line replies look the same as a model that didn't think at all. You pay for that thinking but can't see it.
  • none doesn't exist on gpt-oss-120b. It returns HTTP 400 instead. If you write reasoning_effort="none" into a multi-vendor config, one of these four models will error.

Per-family token ladder (exploratory, not pre-registered)

Token ladder split by task family: on Deduce, gpt-5.4-mini goes from 18 to 3,040 median tokens and gpt-oss-120b from 1,008 to 8,058; on Arith all models stay within roughly ×1.7 to ×2.5; on Distract gpt-oss-120b goes from 52 to 1,082

Split by family, the effect of the dial depends heavily on the question. gpt-5.4-mini's Deduce tokens go 18 → 1,579 → 2,856 → 3,040 (×169). On Arith they only go 128 → 220. gpt-oss-120b spends ×20.6 more tokens on Distract at high than at low, on questions whose answer is in the first sentence.

2. One real accuracy gain, and it came from the first step (H3: ✅ for gpt-5.4-mini only)

Dial curves: accuracy vs reasoning level in three panels (logic puzzles, arithmetic, counting with distractors), one line per model with 95% CI bands. On logic puzzles gpt-5.4-mini rises from 15% at none to 97.5% at high, starting near a dashed

Accuracy at each level with 95% item-cluster bootstrap CIs. The dashed line in the logic panel is random guessing among 7 names. At none, gpt-5.4-mini is right 15% of the time, which is about the random-guess rate.

gpt-5.4-mini on Deduce none low medium high
Accuracy 15% 75% 87.5% 97.5%
95% CI 5–28% 62–88% 78–97% 93–100%
Median output tokens 18 1,579 2,856 3,040
Cost per correct answer $0.0019 $0.0109 $0.0149 $0.0160

The change from none to high is +82.5 percentage points (95% CI +70 to +95, Holm-adjusted p = 0.00056). It is the only effect in the study that passes both pre-registered criteria.

Most of the gain comes from the first step: none → low adds 60 points. Each further step adds less accuracy and costs more.

Forest plot of the 12 pre-registered tests: accuracy change from the lowest level to high with 95% CIs. Only gpt-5.4-mini logic (+82.5 points, Holm p=0.000556) is labelled

All 12 pre-registered tests. One is significant. Six can't move because the model already scores 100% at its lowest setting. gpt-oss-120b's +17.5 points on logic look like a gain but don't survive the correction for 12 tests (Holm p = 0.597).

The other models have a less exciting result that matters for the dial question: Claude Haiku 5.5 scored 97.5% on these puzzles at none, and Gemini scored 100%. For them, turning the dial up on Deduce had nothing left to fix.

3. More thinking didn't significantly hurt, but you can see it going wrong (H2: ❌ not supported)

I expected counting with distractors to be where more effort backfires (H2). The pre-registered test says no: no model got significantly worse at high.

  • gpt-oss-120b went from 75% (low) to 67.5% (high): −7.5 points, CI −22.5 to +7.5. That isn't significant.
  • gpt-5.4-mini went from 75% to 85%, also not significant.
  • Haiku and Gemini scored 100% at every level.

H2 is not supported, and I'm reporting it that way.

The errors were still worth reading. Every wrong Distract answer, from every model, at every level, was an over-count. No model ever answered too low. The models weren't confused about what they owned; they added the distractor numbers to the count.

Three real gpt-5.4-mini replies at reasoning effort high. distract-016:

Three replies from the main run at high, decomposed. article_visuals.py asserts every sum against the raw run files and the locked items before it draws, so none of these numbers are typed by hand. In the first two replies the model adds every number in the prompt. In the third it solves someone else's homework and reports that as the answer.

The longer replies were also the wrong ones. At high, gpt-5.4-mini's wrong Distract answers used a median of 412 output tokens; its correct answers used 91.5. gpt-oss-120b showed the same pattern more strongly: 5,776 tokens for wrong answers against 747 for correct ones. In the most extreme case, gpt-oss-120b spent 21,875 tokens and almost four minutes (235 s) to answer 6 when the answer was 5.

gpt-oss-120b also became less consistent as effort went up. On Distract at high, it gave different answers on the two repeats for 45% of items, compared with 0% at low.

So H2 fails as a significance test, and I won't claim it passed. But "more effort made it worse" isn't the right summary either. A better one is more effort gave it more room to over-count.

4. The bill (H4: ✅ supported, descriptively)

Cost vs gain: each point is one model × task family. x-axis is cost per correct answer at high divided by cost at the lowest level (log scale), y-axis is accuracy change in percentage points. Seven points cluster on the zero line between ×1.5 and ×3.4 cost. gpt-5.4-mini logic sits alone at +82.5 points for ×8.5 cost. gpt-oss counting sits at −7.5 points for about ×25 cost

What turning the dial to high bought. Seven of twelve cells sit on the zero line: same accuracy at ×1.5 to ×3.4 the cost per correct answer. Only one point is high above the line.

This is the chart I'd show anyone who sets reasoning_effort in production:

  • In 7 of 12 cells, high changed nothing but the price. Accuracy stayed the same while cost per correct answer rose ×1.5 to ×3.4.
  • On Arith, the control family, every model stayed at its accuracy while cost per correct rose ×1.65 (gpt-5.4-mini) to ×2.5 (gpt-oss-120b). Spending more effort on an already-solved problem buys nothing.
  • gpt-oss-120b on Distract paid about ×25 per correct answer ($0.00006 → $0.00151) and scored 7.5 points lower (not significant).
  • The one big win was expensive too. gpt-5.4-mini's logic fix raised cost per correct answer ×8.5.

Cost per correct answer vs reasoning level, one panel per task family, one line per model, log scale

Cost per correct answer (USD, log scale). Every line rises from left to right. None of them falls.

One counter-intuitive detail: on Deduce, gpt-5.4-mini at none is the cheapest per correct answer ($0.0019, versus $0.016 at high), even though it's wrong 85% of the time. Cost per correct answer only tells you something when you can tell which answers are right. In most real uses you can't, so don't read that row as advice to choose none.

Scoreboard

Hypothesis Result
H1 The dial is not one instrument ✅ Supported. ×14 vs ×2 token scaling; gpt-oss rejects none; Gemini's none ≈ low and thinks invisibly
H2 More thinking hurts on distractor counting ❌ Not supported. No significant drop. Every error is an over-count, and wrong answers are longer
H3 More thinking helps on deduction ✅ gpt-5.4-mini only: 15% → 97.5% (Holm p = 0.00056). Others are at the ceiling or not significant
H4 When accuracy doesn't rise, the bill still does ✅ Supported (descriptive). 7 of 12 cells unchanged at ×1.5 to ×3.4 cost per correct answer

What I'd tell someone choosing a reasoning-effort value

These come from 20 items per family, 2 repeats and 4 models, so treat them as starting points rather than rules:

  1. Measure none first. Two of the four models (Haiku and Gemini) were already at or near 100% on every family at their lowest setting.
  2. If none fails, try low before high. For the one model that needed the dial, none → low delivered 60 of the 82.5 points.
  3. Don't assume a level does the same thing across vendors. none errors on one, is invisible thinking on another, and is a real off switch on a third.
  4. Watch for long answers to simple questions. On Distract, long replies were the wrong ones.

Surprises

  • The model that needed the dial most was the one where none really meant none. gpt-5.4-mini answered logic puzzles in 18 tokens at none and was right about as often as a random guess.
  • Gemini's none isn't free. It spent a median of 453 hidden tokens at none, which is more than gpt-5.4-mini spent at high.
  • The errors all went in one direction. Every wrong Distract answer was too high, and none was too low.
  • Pushing a task added a model to my lineup. Kaggle's default-model validation runs are how Gemini got into the study.
  • The quota system nearly faked a finding. 378 quota refusals could have been read as "gpt-5.4-mini fails at high effort". The pre-registered rule (403 = missing data) kept them out.

Limitations, plainly

  • Small samples. 20 items per family and 2 repeats per cell. Most "no effect" results mean not detectable at this size, not proven zero.
  • Ceilings. Haiku and Gemini are at or near 100% almost everywhere, so I can say what the dial costs them but not what it buys. Harder items would be needed for that.
  • Synthetic tasks and one prompt format. These questions were built to isolate effects, not to represent real workloads.
  • No temperature control. Kaggle's proxy drops temperature, and the Google models also drop seed. The repeats and the flip rates are my substitute.
  • Cost is Kaggle's proxy cost. Your provider's pricing may differ.
  • Incomplete lineup. Claude Sonnet 5 was dropped (quota reservation), and gpt-oss-120b has no none level. Its comparisons start from low.
  • Partial view of reasoning. Only Gemini returned reasoning traces (at medium and high). For the rest, output tokens are the only evidence of thinking.

What I'd measure next

  • Harder Deduce items (more people, more clues) to get Haiku and Gemini off the ceiling and see whether their dial buys anything.
  • More Distract items and repeats. gpt-oss-120b's −7.5 points and 45% flip rate are the direction H2 predicted, but the sample is too small to test them properly.
  • Latency, which every call already records (latency_s, backend_latency_ms) and this post doesn't analyze.
  • Sonnet and other models, run one task at a time to fit the quota reservation.

Reproduce it

Everything that produced a number in this post is public: the generators, parser, task builder, analysis script, pre-registration and every raw Kaggle run file.

git clone https://github.com/Abeera81/reasoning-dial && cd reasoning-dial
pip install -r requirements.txt
python -m pytest -q                 # 68 tests: generators, parser, task builder, analysis
python analysis/analyze.py          # rebuilds results/summary.md and every figure from results/main/
Enter fullscreen mode Exit fullscreen mode

analyze.py checks the locked test set's SHA-256 before it computes anything. To add a model, fork any of the four public Kaggle tasks and run it. Because each task holds only the dial level, the comparison stays fair.

The dial is a real control. For one model on one kind of problem, it made the difference between guessing and solving. For everything else I measured, it mostly raised the price.

Before you turn it up, check what none gets you. If you run your own models through the tasks, I'd like to see the ladders they produce. 🎛️

Top comments (0)