DEV Community

Cover image for max_tokens=700 on a reasoning model returned empty replies — hidden thinking tokens was the whole budget
InApp
InApp

Posted on Originally published at imapp.blogspot.com

max_tokens=700 on a reasoning model returned empty replies — hidden thinking tokens was the whole budget

For three nights straight my job-digest pipeline produced nothing. The pipeline pulls raw job-posting results, calls an LLM to write a 200-word digest, and stores the output. Mid-week I upgraded the digest call from a standard chat model to a reasoning model — smarter model, better summaries, or so I hoped.

Every response came back with finish_reason: "length" and an empty content field. No error, no truncation warning. The request succeeded; the answer just wasn't there.

The culprit: max_tokens. Seven hundred tokens is plenty for a 200-word digest on a chat model. But on reasoning models the completion budget includes the model's internal reasoning tokens — thousands of invisible thinking tokens get generated before the first visible character. My cap bought roughly 0% reasoning and 0% answer. The model burned the entire budget thinking and got cut off mid-thought. finish_reason honestly reported length, because it did hit a length limit — just not one I thought I was setting.

Two fixes, both permanent now:

  1. Reasoning models get a much larger completion budget — I scaled that call to ~4,000 tokens. The digest got smarter and the cost barely moved, because reasoning tokens are cheap individually but easy to miscount in bulk.

  2. Empty content + finish_reason: "length" is now a first-class failure, not a silent success. Any pipeline that treats finish_reason alone as its health signal will misread exactly this situation as a normal completion.

The broader lesson: an OpenAI-compatible interface hides model-class differences. The same parameter name means different things to a chat model and to a reasoning model. If you route one prompt across model families through a single endpoint, audit what every response field actually means per family — especially the ones that look universal.

I ended up documenting per-model token behavior in my LLM Chat API (https://x402.freeq.one/tools/llm_chat.html); the models endpoint now annotates which tiers burn reasoning tokens, so nobody else repeats my three nights of empty digests.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The empty content plus length finish reason is a useful failure category to preserve, especially when the endpoint remains API-compatible after a model switch. I would also record the reported reasoning-token and visible-token counts separately where the provider supplies them, so the larger budget can be explained from observed usage rather than another guessed cap.

For the digest pipeline, a retry with a larger budget should still have a bounded overall job deadline. A completed non-empty answer then needs a separate content check against the requested digest format. That keeps budget exhaustion, genuinely truncated prose and a valid short digest from collapsing into the same success or failure bucket.