Conclusion (the point of this article — about 8 min read)
On my home LLM (Large Language Model — an AI model that generates text) infrastructure, I changed the monitoring approach once — and reverted it the same day.
What I added was an active health check: make the model actually generate a single character, and judge from the response whether the node is usable. The goal was to catch the state where the process is alive but inference never returns. The problem: I applied a 10-second cutoff to an environment where normal inference takes about 31.8 seconds. Under that condition, 3 of 4 healthy nodes get removed from rotation.
- For inference that takes tens of seconds even when healthy, an active check shorter than the normal response time is unusable. It cannot tell "slow but normal" from "stalled and never returning," so it drops healthy nodes.
-
The original problem — status returns 200 while only inference stalls — is caught, in the end, by real user requests themselves. I stopped probing and instead watch actual traffic: when a real request fails to return within 120 seconds, or comes back 5xx, and that persists across the last 8 requests, the node is removed from rotation. A wedged
/api/chatsurfaces on its own — the real request simply times out, no probe needed. The liveness check went back to the light/api/version(0.25s), splitting "is it alive" and "is the work stuck" into two separate monitors. -
I did not abandon active checks as a technique. In fact, liveness is still checked actively (a light periodic hit to
/api/version). What I dropped was only the "probe inference actively" part. Actively hitting fast work to verify it remains valid as ever; the assumption only breaks when you use an active check to measure inference — work that is slow even when healthy.
Below, in order: what I was monitoring, where I misjudged, and how I went back.
What I was monitoring
I run a small setup for LLM inference. Four Macs carry a 14-billion-parameter (14B — a measure of model size) classification model, and in front of them sits the load-balancing feature of Caddy (a web server and reverse proxy — it takes incoming traffic and distributes it to the machines behind it). Incoming requests get routed to whichever Macs are alive.
The classification model is a small role: it sorts each incoming request by what kind of work it is. Heavy text generation goes to a separate lane; these four machines just do light triage, fast. Since the per-node work is light, the design is to line up several of them and handle requests in parallel. Why I run this myself and how it is assembled are covered in other articles; here I stick to monitoring.
Once you pool four machines, you must not keep sending requests to one that has gone bad. Remove the unhealthy machine from rotation automatically, and bring it back when it recovers. How to make that call is the subject of this article. Caddy, the component doing the routing, periodically checks the state of each of the four machines behind it and sends traffic only to the usable ones. That "how it checks" is what I rebuilt once — and promptly reverted.
The symptom — alive, but not answering
A node would occasionally hang so that only inference stopped responding. The wedge in the title refers to this state: stuck, never returning.
What makes it nasty is that the machine itself looks fine. /api/version, a light query that returns the version, answers in 0.25 seconds. But /api/chat, which carries the inference, never comes back. A monitor that only watches process liveness judges this node "healthy" and keeps sending it requests. From the user's side, requests just get swallowed and never return.
What tipped me off was the unevenness: overall the system was responding, yet some requests never came back at all. The load balancer rotates across the four live machines, so when one of them is wedged, only the requests that land on that machine vanish. Even in the first measurement I observed the mismatch: the liveness check returns in 0.25 seconds while inference on the same node times out. The process is alive. But it is not doing its job. A liveness-only monitor cannot catch that difference.
Wedges tended to happen when several heavy requests piled onto the same node, or when a request suddenly hit a node that had been idle for a while (a cold state). The model has to be loaded back into memory first, so requests in that window are kept waiting. The load balancer mechanically keeps routing to any node that is alive, so requests keep landing on the wedged one and stack up, and only the users who happened to hit that node keep waiting. The only way to stop this is to see the wedge and take the node out of rotation.
I have written before that getting a "200" response (in Hypertext Transfer Protocol, HTTP — the conventions of web traffic — the number that means "success") and the thing actually working are two different matters. This article is the sequel: what happens when you apply that lesson to inference, work that is slow even when perfectly normal.
The first fix — measure inference itself
So I decided to measure "ready or not" by actually making the node infer. An active probe (a small request sent to feel out the state) against /api/chat. The idea: flush out the wedge that liveness checks cannot catch, by running one real inference. Ask the machine not "are you alive?" but "can you work?" — as intentions go, a straightforward one.
Here is the payload. A short, 115-byte request that makes the classification model produce exactly one token (the smallest unit of generation — one character's worth here):
{"model":"qwen2.5-coder:14b","messages":[{"role":"user","content":"1"}],"stream":false,"options":{"num_predict":1}}
The decision logic was built on Caddy's own features. Send this request every 15 seconds; time out at 10 seconds; after 3 consecutive failures, remove the node from rotation; a single success brings it back. A response containing a message field counts as success. The 10-second timeout came from the going rate for watching web endpoints.
Before rolling it out, I ran 12 Bats tests (a test framework for configuration) and Caddy's syntax check (caddy validate). I also verified several assumptions against the real environment: the liveness check returns in 0.25 seconds; the probe traffic actually reaches the Macs from Caddy; success can be judged from the message field; tests and validation pass. Those four I confirmed. But the last assumption — the latency of real production inference — I left at the 10-second going rate without measuring it. I thought I was prepared.
It collapsed in production — three miscalculations
It fell apart immediately after rollout. There were three miscalculations.
First, normal inference exceeded 10 seconds. A 14B model, when busy or freshly started (cold — the model not yet warm in memory), takes about 31.8 seconds even to generate a single token. This is not a hang; it is normal operation. The model must be loaded into memory and the conversation so far recomputed, so requesting just one character via num_predict=1 does not make it shorter. The cost sits in preparing the model to run, not in how many characters come out. The very worst case is the first request landing on a node that has been resting: it starts over from loading the model into memory. Checking all four nodes, three exceeded 10 seconds. Under this criterion, most of the healthy pool gets removed.
This is the heart of the incident. An active probe treats "slow to return" the same as "stopped." For fast endpoints, slow = broken is roughly right. But for inference, slow = normal happens all the time, so the same yardstick does not transfer. The monitor itself knocks out healthy machines. The mechanism meant to protect the pool was shrinking it.
Second, extending the timeout does not make the problem go away. Raise 10 seconds to 60 and healthy machines stop being removed — but now a genuinely wedged node is considered alive for a full 60 seconds. The line between "slow but normal" and "never returning" cannot be drawn by duration alone. Wherever you place the threshold, one kind of error grows. Longer protects healthy nodes but dulls detection; shorter detects fast but catches healthy nodes in the blast. There was no way out of that tug-of-war with a single time-based yardstick.
Third, the request body embedded in the configuration was not sent as written. Caddy has a substitution mechanism that splices runtime values into strings in its config. The curly braces {...} of the JSON above (JavaScript Object Notation, a format for writing data) were misinterpreted as substitution placeholders and replaced with empty strings. The 115 bytes became 28, and Ollama (software for running LLMs locally), receiving a broken request, returned HTTP 400 (the number for "bad request").
What made this worse is that caddy validate cannot find it. The config is syntactically correct, so static validation passes. I only discovered it by watching the traffic on the wire and checking the transmitted bytes one by one. On the screen where the config was written, the request looked correct.
Tests passed; the real machine failed
The original change had 12 passing Bats tests and a passing syntax check. It was still wrong on the real machines.
What the tests guaranteed was the syntax of the configuration and the shape of its output. They never checked "healthy inference does not get removed from rotation by mistake." This is the kind of failure that design and unit tests cannot catch — it only appears when the thing actually runs. The idea that the pass line is the real machine is something I have written repeatedly in other articles. This time it simply showed up in the form of a monitoring threshold.
Put differently: tests verify that things work as written. What I got wrong was whether what I wrote matched reality. The 10-second threshold behaved exactly as written. The premise was simply out of line with the real world. Adding more tests does not close that gap. The only prevention is ordering: measure the real environment's numbers first, then make them the premise.
Back to passive, the same day
What I chose was to remove active inference monitoring entirely.
In its place I put passive monitoring that observes real traffic. Stop sending probes; instead, watch how actual user requests are being answered. A node whose responses are abnormally slow, or which returns errors, gets picked out of the records and removed from rotation. It adds not a single extra request for monitoring's sake, so it puts zero additional load on the pool — it only counts traffic that is already flowing.
The thresholds were set like this. Any request that takes more than 120 seconds marks its node as unhealthy. Healthy-but-long inference measured up into the 10–60 second range, so I raised the bar past where it could catch those. If I had kept the going rate of 8–10 seconds, I would have been removing healthy nodes all over again. Errors (5xx responses) also count as unhealthy, and when failures persist within the most recent window (8 requests), the node comes out of rotation. It stays out for 30 seconds; after that, real requests probe the waters again.
The key point is that the numbers are split in two. The yardstick for what counts as unhealthy (a response over 120 seconds, or a 5xx error) and the duration a node stays removed once judged unhealthy (30 seconds) are separate settings. Making one number do both jobs — a number premised on fast work — is exactly the hole I had just climbed out of.
The liveness side went back to the light /api/version: sent every 5 seconds, cut off at 3. That is plenty fast for noticing a dead process. Whether inference is wedged is the passive monitor's job. Having failed at making one yardstick measure both, I split the roles in two: one watcher for "is it alive," another for "is the work stuck."
Why active inference monitoring was rejected is written down in the config file's comments and in the commit message of the change. A separate design document has not been written so far. Still, for the next time the same temptation comes around, the reasons live where my future self will see that a past self already tried this once and pulled it out.
With active gone, how do wedges get caught?
An obvious question remains: if I stop probing, don't wedged nodes slip through?
The answer: they get caught by the latency and errors of real requests. A real request routed to a node fails to return within 120 seconds, or comes back 5xx. When that persists within the last 8 requests, the node is judged wedged and removed. A wedged /api/chat shows up as real requests timing out — no probe needed. The wedge surfaces in the traffic on its own; monitoring only has to count what surfaces.
This approach has a real weakness. An active probe could preempt: declare "this node is not ready" before any user request arrives. Passive monitoring only knows after a few real requests have already gotten stuck. It cannot preempt.
I chose passive anyway, after weighing the two. The active probe was removing healthy machines wholesale. Passive monitoring delays the handful of requests that land on a wedged node. Dropping three healthy nodes outright, versus a few requests running late on one wedged node — the latter is the lesser harm. And in practice, removing healthy nodes concentrates load on whatever remains, which then starts to wedge as well. A monitor that spreads the outage is the outcome I most wanted to avoid.
There are also guards against misreading. Connection failures cap at 2, and health is judged over a window of the last 8 requests. This keeps "one request happened to be slow" from being mistaken for "persistently wedged." If a single fluke pulls a node out and pushes it back in over and over, that churn itself becomes the instability.
How far does this generalize?
Let me draw one line.
Active probes are not bad. For fast work, hitting the endpoint for real and checking the response content is still effective. Don't accept a bare HTTP 200 as health; look inside the response. That principle has served me well.
It only broke when applied to inference — work that takes tens of seconds while perfectly healthy. The criterion is a single number: how many seconds the target takes to respond when normal. Hit it with a yardstick shorter than that, and the active probe switches sides — it becomes the thing that drops healthy machines. So: measure the normal time first, then set the threshold. It looks obvious, and it was the single most effective ordering in this incident.
This may sound specific to self-built AI infrastructure. But work that is slow while healthy exists elsewhere: converting large files, heavy aggregation, long calls to external services. Monitor those and the same trap is waiting. Do not carry a yardstick premised on fast responses into slow work. That one point is the whole lesson.
Conversely, for an endpoint that answers in one second, active monitoring is still the right fit. In fact, on my own system liveness is still checked actively by hitting /api/version — I did not drop active checks, I only stopped applying them to slow inference. For fast work you can honestly say slow = broken, so hitting it for real and checking the content bothers no one. The dividing line is not the speed of the work itself but whether it can be slow while normal. Passive monitoring for work that is slow even when healthy; active monitoring for work that is fast when healthy. Same discipline, chosen by the nature of the target — and that is all there is to it.
What this incident decided
- The time that counts as "unhealthy" gets decided after measuring the target's normal response time. Never place the threshold first and fit the measurements to it afterward. Treat going-rate values with suspicion — they were made for fast work.
- Whether inference is wedged is judged by observing real traffic. Active checks are limited to a light "is the process alive." One yardstick never gets to cover both liveness and wedging.
- Rejected options get their reasons recorded in both the config file and the change history. So the same choice never has to be made twice from scratch. Six months from now, I will not remember today's reasons — I plan for that.



Top comments (1)
We burned 3 out of 4 nodes doing exactly this on a 14B inference cluster. The culprit was the same: cold-start on a big model takes longer than any sane timeout, so the health check is just racing against normal behavior and losing. Dropped probing entirely, moved to watching real request outcomes. What stuck for us: treating "is the process alive" and "can it actually infer" as separate signals, since one's cheap to measure and the other isn't.