For three nights straight my job-digest pipeline produced nothing. The pipeline pulls raw job-posting results, calls an LLM to write a 200-word digest, and stores the output. Mid-week I upgraded the digest call from a standard chat model to a reasoning model — smarter model, better summaries, or so I hoped.
Every response came back with finish_reason: "length" and an empty content field. No error, no truncation warning. The request succeeded; the answer just wasn't there.
The culprit: max_tokens. Seven hundred tokens is plenty for a 200-word digest on a chat model. But on reasoning models the completion budget includes the model's internal reasoning tokens — thousands of invisible thinking tokens get generated before the first visible character. My cap bought roughly 0% reasoning and 0% answer. The model burned the entire budget thinking and got cut off mid-thought. finish_reason honestly reported length, because it did hit a length limit — just not one I thought I was setting.
Two fixes, both permanent now:
Reasoning models get a much larger completion budget — I scaled that call to ~4,000 tokens. The digest got smarter and the cost barely moved, because reasoning tokens are cheap individually but easy to miscount in bulk.
Empty content +
finish_reason: "length"is now a first-class failure, not a silent success. Any pipeline that treatsfinish_reasonalone as its health signal will misread exactly this situation as a normal completion.
The broader lesson: an OpenAI-compatible interface hides model-class differences. The same parameter name means different things to a chat model and to a reasoning model. If you route one prompt across model families through a single endpoint, audit what every response field actually means per family — especially the ones that look universal.
I ended up documenting per-model token behavior in my LLM Chat API (https://x402.freeq.one/tools/llm_chat.html); the models endpoint now annotates which tiers burn reasoning tokens, so nobody else repeats my three nights of empty digests.
Top comments (1)
The empty content plus length finish reason is a useful failure category to preserve, especially when the endpoint remains API-compatible after a model switch. I would also record the reported reasoning-token and visible-token counts separately where the provider supplies them, so the larger budget can be explained from observed usage rather than another guessed cap.
For the digest pipeline, a retry with a larger budget should still have a bounded overall job deadline. A completed non-empty answer then needs a separate content check against the requested digest format. That keeps budget exhaustion, genuinely truncated prose and a valid short digest from collapsing into the same success or failure bucket.