Why We Compared Docker Model Runner and Ollama
We did not need another local chat command. We needed a repeatable way to give developers the same local model, configuration, and API surface without requiring a different setup on every laptop.
No Verified Speed Winner
Both tools commonly run llama.cpp underneath, so raw generation speed tends to be close and cannot be declared a winner here; the defensible decision turns on Docker-native workflow fit, model lifecycle control, and validated GPU-path support rather than an unmeasured cold-start or memory advantage.
Ollama already handles the fast-start path well: install it, run a named model, and connect an application to port 11434. Docker Model Runner targets a different source of friction. It treats models as artifacts that fit Docker-oriented workflows, exposes OpenAI- and Ollama-compatible APIs, and integrates model management into the Docker CLI and Docker Desktop.
That distinction matters more than a one-token-per-second performance difference.
Under its default configuration, Docker Model Runner uses llama.cpp for GGUF inference. It can also use vLLM for NVIDIA-backed, higher-throughput workloads and Diffusers for image generation. Ollama likewise uses llama.cpp as a core supported backend, so a default comparison often measures two orchestration layers around closely related inference machinery rather than fundamentally different inference architectures.
We therefore framed the decision around four questions:
- Can we install and operate the runtime consistently across developer machines?
- Can we measure cold starts without confusing downloads, model loading, prompt evaluation, and generation?
- Can we identify whether memory is resident, cached, or unloaded?
- Can applications and containers reach the API without accidental exposure or platform-specific networking hacks?
We also imposed a hard evidence rule: we would not publish invented time-to-first-token, RSS, or VRAM values. The available material did not contain controlled, same-hardware cold-start or memory results, and our isolated executable check did not run either inference server. We therefore focus on workflow comparisons and measurement limits rather than a benchmark ranking.
In our benchmark review, we examined April 2025 aggregate results of 11,982.18 ms mean duration and 23.65 mean tokens per second for Ollama, versus 12,872.06 ms and 24.53 mean tokens per second for Docker Model Runner. Median throughput was 24.31 versus 24.68 tokens per second. We did not use those figures to rank the tools: we could not establish the hardware, model artifact, quantization, context size, prompt set, runtime versions, or warm-up policy, and these were not measurements from our isolated SDK check.
The useful conclusion was narrower: raw generation speed can be close when both paths ultimately rely on llama.cpp. Workflow, lifecycle control, and hardware support should drive the purchase or standardization decision.
Teams evaluating adjacent infrastructure can also review our tools collection or work with us on an inference architecture review.
Hands-On Walkthrough: Setup, Execution & Output
We started with the officially supported installation paths rather than wrapper projects.
For Ollama on macOS or Linux, the published bootstrap command is:
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4
On Windows, the corresponding PowerShell installation path is:
irm https://ollama.com/install.ps1 | iex
Ollama exposes its native REST API on http://localhost:11434. A non-streaming request looks like this:
curl http://localhost:11434/api/chat \
-H 'Content-Type: application/json' \
-d '{
"model": "gemma4",
"messages": [
{
"role": "user",
"content": "Reply with exactly: ready"
}
],
"stream": false
}'
Docker Model Runner requires a supported Docker Desktop or Docker Engine installation with Model Runner enabled. We checked availability before pulling anything:
docker model status
docker model pull ai/smollm2
docker model run ai/smollm2
docker model ps
We could force the default llama.cpp backend when we wanted the execution choice to be explicit:
docker model run ai/smollm2 --backend llama.cpp
For an NVIDIA Linux host configured for vLLM, the setup path changes:
nvidia-smi
docker run --rm --gpus all nvidia/cuda:12.0-base nvidia-smi
docker model install-runner --backend vllm --gpu cuda
docker model status
We treated those GPU checks as gates, not proof of accelerated inference. A successful nvidia-smi confirms that the host sees the device. A successful CUDA container confirms Docker GPU access. Neither proves that a specific model loaded on the intended backend or that every layer was offloaded.
The following transcript is simulated to show the expected command sequence and output shape. It is not presented as measured benchmark evidence:
$ docker model status
Docker Model Runner is running
Status:
llama.cpp: running
vllm: not installed
diffusers: not installed
$ docker model pull ai/smollm2
Downloaded ai/smollm2
$ docker model run ai/smollm2
> Reply with exactly: ready
ready
$ docker model ps
MODEL BACKEND STATUS
ai/smollm2 llama.cpp running
$ ollama run gemma4 "Reply with exactly: ready"
ready
$ curl -s http://localhost:11434/api/chat \
-H 'Content-Type: application/json' \
-d '{"model":"gemma4","messages":[{"role":"user","content":"Reply with exactly: ready"}],"stream":false}'
{"model":"gemma4","message":{"role":"assistant","content":"ready"},"done":true}
We would not use those two model names for a performance comparison. A valid benchmark requires the same weights, quantization, context size, prompt, output-token ceiling, and sampling parameters. Successfully running ai/smollm2 and gemma4 would establish basic operation of the respective command paths, not comparative performance. The simulated transcript above does not establish that either command succeeded.
We also tested the response-validation boundary in ollama==0.4.7, using pydantic==2.10.6 and httpx==0.28.1. We supplied small synthetic response objects without contacting a server.
The SDK accepted a completed response with done: true even when all of these benchmark fields were absent:
total_durationload_durationprompt_eval_countprompt_eval_durationeval_counteval_duration
It also accepted a response with no completion marker, leaving done as null. When we supplied a malformed string for load_duration, validation correctly failed with an int_parsing error.
That result showed a limitation of SDK validation: successful SDK parsing does not guarantee that a response contains usable benchmark telemetry. Our harness must validate required fields independently and measure time-to-first-byte or time-to-first-token at the streaming transport layer.
Plugin Discovery, GPU Support, Memory, and API Limitations
Our first Docker-specific failure mode was plugin discovery. When docker model is not recognized on macOS, the CLI may not see the Model Runner plugin in its expected directory. The workaround is a symlink:
mkdir -p "$HOME/.docker/cli-plugins"
ln -s \
/Applications/Docker.app/Contents/Resources/cli-plugins/docker-model \
"$HOME/.docker/cli-plugins/docker-model"
docker model status
We would inspect the target first and avoid replacing an existing plugin blindly. This is a Docker Desktop integration issue, not a model or GPU problem.
The second gotcha was the phrase “GPU passthrough.” On Apple silicon, Docker Model Runner’s llama.cpp engine uses Metal acceleration, but the engine does not run as an ordinary container with a passed-through GPU device. On macOS, the engine runs in a host-side sandbox. On Linux, Model Runner and its inference engines run inside a container, making the isolation and device-access path materially different.
That means a successful Mac test does not validate Linux CUDA deployment.
Our platform matrix was also narrower than the marketing-level phrase “runs locally” suggests:
- The llama.cpp backend covers macOS, Windows, and Linux, including CPU execution.
- Apple silicon uses Metal automatically.
- Linux can use NVIDIA CUDA, AMD ROCm, Vulkan, or CPU paths, subject to hardware and runtime setup.
- vLLM requires NVIDIA CUDA and does not provide a CPU fallback.
- vLLM is supported on Linux x86-64 and Windows through WSL2, with Docker Desktop 4.54 or later required for the documented Windows path.
- Diffusers requires an NVIDIA GPU on supported Linux architectures.
- Docker Model Runner’s stated NVIDIA driver floor is 575.57.08 or later on Linux.
We also had to control context size before discussing memory. Docker Model Runner’s defaults vary by model and commonly fall between 2,048 and 8,192 tokens. That range can materially change key-value cache allocation, so an RSS or VRAM comparison without a pinned context is not meaningful.
We used the configuration shape:
docker model configure --context-size 2048 ai/qwen2.5-coder
We did not obtain trustworthy same-host RSS or VRAM measurements from the supplied execution environment. We therefore cannot state that either runtime used fewer gigabytes. Docker Model Runner is designed to load models when requested and unload them when idle, but we still need to measure that behavior against Ollama with an explicit keep-alive policy, fixed observation window, and process-tree accounting.
Another issue was API exposure. Docker Model Runner’s API has no built-in authentication. Any client that can reach it can submit inference requests and may be able to pull, load, or run models. We would not expose its port on an untrusted LAN, shared CI network, or publicly routed developer host. Our workaround is network scoping plus an authenticated reverse proxy where multi-user access is unavoidable.
Finally, we found that “OpenAI-compatible” does not mean operationally identical. Docker Model Runner uses engine-aware paths such as:
/engines/llama.cpp/v1/chat/completions
/engines/vllm/v1/chat/completions
/engines/v1/chat/completions
Ollama’s native chat API uses:
/api/chat
Both can fit existing clients, but we still test request fields, streaming chunks, error schemas, usage data, and completion markers before switching providers.
Scale, Latency & Cost vs. Alternatives
Our comparison favors operational fit over unsupported precision:
| Decision area | Docker Model Runner | Ollama | Our assessment |
|---|---|---|---|
| Fastest standalone setup | Requires Docker integration and Model Runner enablement | One-line installer and direct CLI | Ollama wins |
| Model distribution | OCI artifacts and registry-oriented workflows | Ollama model library and Modelfiles | Docker wins for registry governance |
| Default local engine | llama.cpp | llama.cpp-supported architecture | Likely similar when artifacts and settings match |
| OpenAI-compatible API | Yes | Available alongside native API support | Both are usable; test client behavior |
| Native API compatibility | Ollama-compatible API available | Native Ollama API | Ollama is the reference path |
| macOS acceleration | Apple silicon with Metal | Apple silicon acceleration | Both are practical |
| Linux NVIDIA path | llama.cpp CUDA and optional vLLM | NVIDIA acceleration | Docker offers clearer multi-engine expansion |
| AMD Linux path | ROCm or Vulkan options for supported backends | Hardware support varies by release | Validate the exact GPU |
| CPU fallback | Yes with llama.cpp | Yes | Both |
| vLLM integration | Built into the Model Runner engine model | Not the primary Ollama workflow | Docker wins for this transition |
| Compose and Testcontainers alignment | Compose and Java/Go Testcontainers support | Comparative integration behavior not established here | We favor Docker for ecosystem fit, not verified workflow parity |
| API authentication | None by default | Must also be network-scoped carefully | Neither should be exposed casually |
| Verified cold-start winner | Not established | Not established | No defensible winner |
| Verified memory winner | Not established | Not established | No defensible winner |
For local development, software licensing cost is usually zero. The real cost is engineering time and workstation capacity.
We use this break-even model:
monthly local cost =
workstation amortization
+ electricity
+ setup and support hours
+ CI maintenance
+ developer waiting time
monthly hosted cost =
input token charges
+ output token charges
+ provisioned GPU hours
+ network and storage
+ privacy or compliance overhead
Assume a team spends eight engineering hours standardizing a runtime at a fully loaded engineering cost of $150 per hour. The initial integration cost is $1,200. If the chosen workflow saves ten developers six minutes per working day, the monthly recovery is approximately:
10 developers × 0.1 hours × 20 days × $150 = $3,000 per month
Under those illustrative assumptions, standardization pays back within the first month. The result does not depend on Docker Model Runner producing one more token per second. It depends on avoiding setup drift, broken GPU paths, duplicate model downloads, and client-specific configuration.
For sustained production traffic, neither default local-developer workflow is automatically the right answer. Docker Model Runner’s vLLM backend simplifies the transition to concurrent NVIDIA inference, but we would still benchmark dedicated vLLM, SGLang, managed endpoints, or a Kubernetes-serving stack before declaring a production standard.
A laptop benchmark answers developer-experience questions. It does not establish production throughput, tail latency, admission control, or multi-tenant isolation.
Our Final Verdict: When to Deploy, When to Skip
We would deploy Docker Model Runner when:
- We already standardize development around Docker Desktop, Docker Engine, Compose, registries, and Testcontainers.
- We want model artifacts managed through familiar OCI distribution controls.
- We need one Docker-oriented interface spanning llama.cpp today and potentially vLLM later.
- We can enforce supported Docker, driver, and GPU versions.
- We are willing to secure the unauthenticated API at the network boundary.
- We value reproducible model packaging more than the shortest possible install path.
We would choose Ollama when:
- We want the fastest route from a clean laptop to a local model.
- We are building a standalone local application rather than a Docker-centered development platform.
- Our tools already target Ollama’s native API and model library.
- We want fewer Docker Desktop dependencies on developer machines.
- We do not need OCI-packaged model governance.
We would hold off on either as a production standard when:
- We need proven high-concurrency serving or strict latency objectives.
- We require built-in authentication, quotas, tenancy, or audit controls.
- We have not pinned the model artifact, quantization, context, runtime version, sampling settings, and unload policy.
- We are comparing Mac Metal results with Linux CUDA deployment plans.
- We cannot observe process RSS, GPU memory, load duration, first-token latency, and steady-state throughput independently.
Our final call is straightforward: Ollama remains the better default for individual developers who want local inference with minimal ceremony. Docker Model Runner is the stronger organizational choice when models must fit the same artifact, registry, Compose, and testing workflows as the rest of the stack.
What We Could Not Verify
We did not establish a cold-start or memory winner through controlled, same-hardware measurements. Our isolated executable check tested SDK response validation without running either inference server.
Before committing across a team, we would run the same GGUF on the same machine, pin context to 2,048 tokens, record streamed first-token time externally, poll the complete process tree and GPU allocator, force unload between cold runs, and preserve raw JSON responses. If that level of validation is important to your rollout, contact Effloow before workstation convenience turns into infrastructure policy.
Top comments (0)