On my MacBook Air, Ollama 0.34.4 served qwen3:1.7b with a 4,096-token window. The model supports 40,960. I sent a 7,000-token prompt and got HTTP 200 back, but the model had only seen the last 2,050 tokens. There was no error and no warning.
That was the start of llm-doctor, a health check for local LLM setups.
The problems it looks for
Running models locally leaves a lot of quiet mess behind.
- Every runtime keeps its own copy of the same GGUF. Ollama hides it behind a sha256 blob name, so you cannot see that LM Studio holds the same 20 GB file.
- Quantizers fix chat templates after release, and your old download keeps the broken one. Tool calls then fail in ways that look like the model being bad.
- Ollama defaults to a small context window, and the OpenAI-compatible endpoint that coding agents use cannot raise it per request.
- Leaving Ollama means downloading everything again, because its blobs have no names.
Using it
uv tool install git+https://github.com/Arthur031221/llm-doctor
llm-doctor # scan every model store
llm-doctor fix # preview dedupe and cleanup
llm-doctor fix --yes # apply it
llm-doctor unbundle --to llama-server # expose Ollama models to llama.cpp
llm-doctor endpoint http://localhost:11434 # agent-readiness probe
The GIF above runs against a throwaway home directory where Ollama and LM Studio hold the same GGUF, plus one orphan blob. The scan reports both, and fix --yes replaces the LM Studio copy with a symlink to the Ollama blob and deletes the orphan. The demo fixture reclaims 140 MB, and every change is logged to ~/.llm-doctor/fix-log.jsonl.
The endpoint probe is the part I use most. It checks context length, tool calls, parallel tool calls, streaming tool calls, JSON schema output and think tags, then prints what to change for each failure. For the Ollama case above it reported 4,096 effective tokens of 40,960 advertised and marked the context check as failed.
Limits
It needs Python 3.10 or newer. It is not on PyPI yet, so install from GitHub as shown. I tested mostly on macOS with Ollama, so Linux and Windows reports are the most useful feedback I can get. If you run a store or runtime it does not detect, open an issue with the folder layout.
The code is at https://github.com/Arthur031221/llm-doctor under the MIT license.

Top comments (0)