DEV Community

Nerav Doshi
Nerav Doshi

Posted on Edited on Originally published at pipelineandprompts.com

Containerized Ollama and Found the Real Memory Overhead

On Linux, containers run natively — straight on the host kernel. On macOS, they can't, because containers need a Linux kernel underneath and your Mac isn't running one. So Podman and Docker Desktop quietly spin up a small Linux VM in the background and put every container inside that instead. Which means every container on a Mac is sharing a fixed slice of memory carved out for that VM — not your Mac's actual RAM — and that distinction is about to matter a lot.

Started Ollama as a container, exposed on a different port so it wouldn't collide with the native app already running:

podman run -d --name ollama-container -p 11435:11434 ollama/ollama
podman exec -it ollama-container ollama pull llama3.2:1b
podman exec -it ollama-container ollama run llama3.2:1b
Enter fullscreen mode Exit fullscreen mode

Pull went fine, after one transient network hiccup. Loading the model didn't:

Error: 500 Internal Server Error: model requires more system memory (1.3 GiB) than is available (620.8 MiB)
Enter fullscreen mode Exit fullscreen mode

podman machine list explained why immediately — the VM backing every container on this machine was set to 2GiB total, for the OS, the runtime, and every container combined. After overhead, there was only ~620MB actually free. Nowhere close to enough.

Fixed by resizing the VM (has to be stopped first):

podman machine stop
podman machine set --memory 4096
podman machine start
Enter fullscreen mode Exit fullscreen mode

That restart also killed the container, which then refused to exec into ("container state improper") until I explicitly restarted it too:

podman start ollama-container
podman exec -it ollama-container ollama run llama3.2:1b
Enter fullscreen mode Exit fullscreen mode

Worked that time. Asked it "what is kubernetes?" to force a real response, then pulled the memory number with Podman's version of docker stats:

podman stats ollama-container --no-stream
Enter fullscreen mode Exit fullscreen mode
Entry 01 (bare metal, native app) Containerized
Process memory ~1.24 GB RSS 1.663 GB
Processor 100% GPU (Metal) CPU only — no Metal passthrough in a container

About 34% more memory for the exact same model. Some of that's Podman/Ollama server overhead, some is probably the lack of GPU acceleration pushing more work onto the CPU side — that second part's an inference from the numbers, not something I directly measured, so I'm not going to state it more confidently than that. Either way, "same model, same memory" turned out to be wrong. Containerizing this isn't memory-neutral.

Which means the resource request I'd actually set for this model in a container should be closer to 1.663GB than the 1.24GB bare-metal number from Entry 01 — sizing off the bare-metal figure alone would under-provision it. And separately: Podman's VM has its own memory ceiling, completely independent of how much RAM your Mac actually has. That's the first thing worth checking before assuming a model is just too big to run.

Top comments (0)