DEV Community

Cover image for The ML you need to operate LLMs, not train them
Raj Murugan
Raj Murugan

Posted on Originally published at rajmurugan.com

The ML you need to operate LLMs, not train them

You do not need to understand backpropagation to run a large language model well in production. You need a smaller, more practical thing: the operator's mental model. What a token actually costs you, what a context window actually bounds, what temperature and top_p actually do, and why evals catch what your own reading of the output will not. None of this is deep learning theory. All of it is the part that bites when you skip it.

Diagram of the operator's mental model: the same text tokenising to different counts on different model families, a context window shared between system prompt, tools, history and output, and a Bedrock Converse request rejected for sending both temperature and top_p.

Tokens are not words, and the count is not universal

A token is the unit the model actually reads and the unit you are actually billed on, and it is not a word. Common English words are usually one token. Longer or less common words split into pieces. Punctuation, whitespace, and anything outside the tokeniser's common vocabulary (a rare name, a non-English string, a stretch of code) can cost more tokens than it looks like it should.

The part that catches people operating across model families on Bedrock specifically: the same text does not tokenise to the same count on every model, because different model families ship different tokenisers. I measured this directly on a real production system prompt while working on prompt caching: the identical string came out to 8,156 tokens on the Nova tokeniser and 8,788 tokens on the Anthropic tokeniser, an 8 percent difference from the exact same characters. If you are estimating cost or checking against a context limit, "how many tokens is this" has a different answer depending on which model you ask, and a number measured against one family does not transfer to another.

A context window is a shared budget, not a per-turn allowance

"Context window" sounds like it means how much the model remembers. Operationally, it means something narrower and more consequential: the hard ceiling on everything in a single request, added together. Your system prompt, your tool definitions, the conversation history you send back each turn, and the model's own output all draw from the same number. None of them have their own separate budget.

This is why a long tool registry or a verbose system prompt is not free even before a user has said anything: it is the floor your actual conversation history has to fit above. On Bedrock today, that ceiling varies a lot by model: Claude Sonnet 4.6's model card lists a 1,000,000-token context window, though AWS's own beta-header table also lists context-1m-2025-08-07 against this model, so confirm empirically whether you need it before assuming the full million applies with no header set. Amazon Nova's own range spans from 128,000 tokens on Nova Micro up to 1,000,000 on Nova 2 Lite, with Nova Pro and Nova Lite sitting at 300,000. Picking a model by capability alone and finding out its context ceiling later is a real way to discover, mid-project, that your working budget just got a lot smaller.

temperature and top_p are not independent dials

Generic ML content treats temperature and top_p as two separate knobs you can tune together: temperature reshapes the probability distribution, top_p truncates it. On Bedrock, AWS documents this exactly for two specific models, Claude Sonnet 4.5 and Claude Haiku 4.5: "support specifying either the temperature or top_p parameter, but not both." That is a named, model-specific constraint, not a blanket rule for "4.x and up," and the same shape of failure keeps turning up on newer releases without AWS updating that page to match, so check per model rather than assuming generational coverage.

Send both temperature and top_p in the same inferenceConfig to one of these models and the Converse API rejects the call. The exact wording I have seen reported, consistently, across multiple independent client libraries and agent frameworks is "temperature and top_p cannot both be specified for this model. Please use only one." AWS's own documentation states the constraint in prose, not that literal string, so treat the error text itself as an observed runtime message, not a documented, versioned contract. Either way, this is not a rare edge case: the common trigger is not a developer deliberately setting both, it is a client SDK that sends both by default (one explicitly set, one left at its base-config default) without the caller ever intending to.

The operator fix is boring and worth stating plainly: pick one. Keep temperature unless you have a specific reason to reach for top_p, and when you touch a client library's defaults, check what it is silently sending alongside the parameter you meant to set.

"Pick one" stops being enough the moment extended thinking is on. AWS's own extended-thinking documentation is explicit that thinking is not compatible with temperature, top_p, or top_k modifications at all, so with thinking enabled the fix is to omit both, not choose between them.

# This fails on Claude Sonnet 4.5 / Haiku 4.5 and newer releases that share
# the same constraint, if a client library has already populated top_p from
# its own defaults. Check per model; AWS names only two here explicitly.
inference_config = {
    "temperature": 0.7,
    "topP": 0.9,      # remove this key entirely, don't set it to "neutral"
    "maxTokens": 1024,
}
Enter fullscreen mode Exit fullscreen mode

Why it hallucinates is not a bug report, it is the training objective

A model is trained to predict the most statistically plausible next token given everything before it. That objective has no separate channel for "is this true", because truth was never the thing being optimised. Fluent and plausible is the target; correct is a frequent side effect of fluent and plausible, not a guarantee of it.

This is why a confident, well-formatted, wrong answer is not a malfunction. The model did exactly what it was trained to do: continue the text in the most plausible way it could find. Instruction-tuning and RLHF push the model towards saying "I don't know" in more situations, but that is a learned behaviour layered on top of the same underlying objective, not a built-in truth detector. Treat "the model sounded confident" as evidence of nothing. It is optimising for fluency, every time, including the times it is wrong.

Why evals, not your own reading of the output

Once you have internalised all of the above, the honest next question is: how do you actually know the output is good? Not by reading a handful of examples and feeling satisfied. I have written at length about the ways that instinct fails: an LLM-as-judge score that is not a fixed property of the thing being graded, a golden dataset that is easier than production traffic, a quality threshold that needs a real sample size to mean anything. That is rung 7 of this roadmap, and it exists because the operator's mental model in this post is necessary and not sufficient. Knowing what a token costs and what a context window bounds tells you how the system behaves. It does not tell you whether the system is right.

Four rows: the same text tokenises differently per model family, measure on the model you call; a context window is one shared budget across prompt, tools, history and output; temperature and top_p cannot both be set on current Claude models on Bedrock; a confident answer is not evidence of a correct one.

What I now check

The token count for a string is not portable between model families; measure it on the model you are actually calling, not a general-purpose estimate.

A context window is one shared budget across system prompt, tools, history and output, not a per-turn allowance on top of them.

Check whether the model you are calling accepts both temperature and top_p before a client library finds out for you in production.

A confident answer is not a correct answer. The training objective never promised the second thing.

Which one of these cost you the most time to learn the hard way?

Find me on LinkedIn or via rajmurugan.com.

Top comments (0)