Passing a value to a parent function and that value actually reaching the server turned out to be two different statements
This is the English version of a post originally written in Korean for my algorithmic trading system devlog(new tab).
A few weeks ago, out of nowhere, I got asked: "Haven't we never actually discussed the temperature setting?"
The question was about the decoding temperature of the local LLM used in the AI recommendation pipeline(new tab) — the value that governs how much the answer wobbles from run to run on the same input — and whether I had really chosen it on purpose. When I went back and looked, the situation was half right and half wrong.
An experiment I thought was finished
A few weeks earlier I had already run an A/B test to check whether lowering the temperature makes the answers more consistent.
The result was "no difference." There was no statistically meaningful gap between the low-temperature arm and the default arm, so instead of touching the temperature I concluded I'd solve the consistency problem another way (re-drawing the same input several times and averaging). That conclusion then sat there, unquestioned, for weeks.
The real cause - the value was passed, but never reached the server
To answer the reopened question, I traced the code from the start.
The upper layer (the code that builds the graph) really was passing the temperature value as an argument. That part was right.
The problem was the client function that received that value and actually sent the request to the local LLM server. It accepted a few arguments like callbacks and passed them on, but temperature wasn't in its signature at all, so it was silently discarded. No error, no warning. It just vanished.
So the "low-temperature" arm of that A/B test had really been running on the server default. The two arms had effectively been under the same conditions, which means the "no difference" conclusion came from a measurement that hadn't verified anything to begin with.
One layer deeper - production wasn't on the intended value either
Since I'd confirmed the value was getting lost, I also checked what temperature production was actually running on.
The server was being launched without an explicit temperature flag. In that case the server adopts whatever default sampling profile is baked into the model file.
That default was the profile recommended for "thinking mode" (the reasoning-boosted mode), while the model in actual operation was configured not to use thinking mode. So the mode and the sampling profile were mismatched, and had been for weeks.
The belief that "we tuned it to this value" and the fact that "it's actually running on this value" had been separated, and never once verified.
The fix - and changing how I verify
I explicitly added temperature, top_p, and presence_penalty to the client function's signature so that values passed from the upper layer actually make it into the request. When no value is given, it behaves exactly as before, which prevented regressions.
I also changed how I verify this. "The argument was added in the code" is no longer enough for me. I launched a real server process and checked directly in the logs that the process was actually carrying those flags on its command line. I separately confirmed that the "applied" message shows up in the logs of a request actually being handled. The evidence was measurement of a running process, not code review.
The same incident, once more, in another system
Two days after deploying the fix, the same class of problem came up in a separate hardware performance comparison job.
That benchmark harness launched its own server process for testing, and its startup path was different from the path production uses to launch the server. The sampling-value wiring I had just fixed wasn't part of that path.
As a result, the performance-comparison trials were about to run on the pre-fix defaults (the thinking-mode profile). Fixing the production path did not automatically fix the other paths that build the same kind of request.
Prompted by this second recurrence, I added a general monitoring check that automatically catches any mismatch between the server settings production actually launches with and the intended values defined in code. The goal was to get rid of a setup where this only gets found when someone happens to ask a question.
The general lesson
"I passed the value to the parent function" doesn't prove "the value actually reached its destination." When a parameter travels through several layers, a function signature somewhere in the middle can silently drop it. This kind of loss leaves no error and no warning, so code review alone rarely catches it. You need a separate verification step that inspects the request actually going out, or the process actually running.
When a parameter silently disappears, an A/B test becomes an A-vs-A test. A "no difference" result can mean two things: there's really no difference, or the two things being compared never actually diverged. Before trusting an experiment's result, it's worth first checking that the two conditions really produce different values in the code.
Fixing a configuration bug doesn't fix every path that uses that configuration. If several code paths build the same kind of request, then after fixing one you must ask "is there another entry point that uses this value?" That's especially true when a verification harness launches the same server through a different path than production.
An incident found by accident should turn into a recurrence-prevention mechanism. The whole thing started with one offhand question. So that the next mismatch of this kind gets caught by a monitor rather than by a question, it's better to plant a generalized check right after the discovery.
Top comments (0)