DEV Community

Cover image for Run Ollama with GPU on Ubuntu Using Docker and Open WebUI
Ethan Vance
Ethan Vance

Posted on Originally published at migservers.com AI-assisted

Run Ollama with GPU on Ubuntu Using Docker and Open WebUI

Running large language models locally gives developers more control over their AI environment. With Ollama, Docker, and Open WebUI, you can set up a self-hosted AI chatbot on an Ubuntu server with NVIDIA GPU acceleration.

In this guide, we'll cover the essential components required to get started with GPU-powered local AI.

What You'll Need

Before deploying Ollama, prepare the following:

  • Ubuntu 22.04 or 24.04 LTS
  • A compatible NVIDIA GPU
  • NVIDIA drivers
  • Docker Engine
  • NVIDIA Container Toolkit
  • Sufficient GPU VRAM, RAM, and storage for your chosen model

Step 1: Update the server and check your GPU

Log in over SSH and update the system:

sudo apt update && sudo apt upgrade -y
Enter fullscreen mode Exit fullscreen mode

Confirm Ubuntu can see your GPU:

lspci | grep -i nvidia
Enter fullscreen mode Exit fullscreen mode

You should see a line with your GPU model. If nothing appears, the GPU is not detected: check with your hosting provider before continuing.

Step 2: Install the NVIDIA driver

List the recommended driver for your GPU:

sudo ubuntu-drivers devices
Enter fullscreen mode Exit fullscreen mode

Install it automatically (this picks the recommended version):

sudo ubuntu-drivers autoinstall
Enter fullscreen mode Exit fullscreen mode

Reboot:

sudo reboot
Enter fullscreen mode Exit fullscreen mode

After reconnecting, verify:

nvidia-smi
Enter fullscreen mode Exit fullscreen mode

You should see a table showing your GPU name, driver version and VRAM. If you see it, the driver works.

⚠️ Secure Boot warning If Secure Boot is enabled in BIOS, the NVIDIA kernel module may refuse to load and nvidia-smi will fail. Either disable Secure Boot in BIOS/IPMI, or enroll a MOK key when the installer asks. See Troubleshooting, Problem 1.

Step 3: Install Docker

curl -fsSL https://get.docker.com | sudo sh
Enter fullscreen mode Exit fullscreen mode

Allow your user to run Docker without sudo:

sudo usermod -aG docker $USER
Enter fullscreen mode Exit fullscreen mode

Log out and log back in (or run newgrp docker), then test:

docker run --rm hello-world
Enter fullscreen mode Exit fullscreen mode

Step 4: Install the NVIDIA Container Toolkit

This is the piece that lets Docker containers use your GPU.
Add NVIDIA's repository:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
  sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
Enter fullscreen mode Exit fullscreen mode

Install it:

sudo apt update
sudo apt install -y nvidia-container-toolkit
Enter fullscreen mode Exit fullscreen mode

Tell Docker to use the NVIDIA runtime, then restart Docker:

sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Enter fullscreen mode Exit fullscreen mode

Test GPU access inside a container

docker run --rm --gpus all ubuntu nvidia-smi
Enter fullscreen mode Exit fullscreen mode

If you see the same GPU table as in Step 2, GPU passthrough to Docker works. Do not continue until this test passes

Step 5: Create the project folder

mkdir -p ~/ai-server && cd ~/ai-server
Enter fullscreen mode Exit fullscreen mode

5.1 Create the docker-compose.yml

nano docker-compose.yml
Enter fullscreen mode Exit fullscreen mode
Paste this:
services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    volumes:
      - ollama_data:/root/.ollama
    environment:
      - OLLAMA_KEEP_ALIVE=30m        # keep model loaded in VRAM for 30 min
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    # No "ports:" section on purpose. Ollama has no login,
    # so it must NOT be reachable from the internet.

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    restart: unless-stopped
    depends_on:
      - ollama
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY}
      - ENABLE_SIGNUP=${ENABLE_SIGNUP:-true}
    volumes:
      - open_webui_data:/app/backend/data

  caddy:
    image: caddy:2
    container_name: caddy
    restart: unless-stopped
    depends_on:
      - open-webui
    ports:
      - "80:80"
      - "443:443"
    volumes:
      - ./Caddyfile:/etc/caddy/Caddyfile:ro
      - caddy_data:/data
      - caddy_config:/config

volumes:
  ollama_data:
  open_webui_data:
  caddy_data:
  caddy_config:
Enter fullscreen mode Exit fullscreen mode

Save with Ctrl+O, Enter, then Ctrl+X.

5.2 Create the .env file

Generate a random secret key:

echo "WEBUI_SECRET_KEY=$(openssl rand -hex 32)" > .env
echo "ENABLE_SIGNUP=true" >> .env
Enter fullscreen mode Exit fullscreen mode

5.3 Create the Caddyfile

Option A: with a domain (recommended, free automatic HTTPS). First point your domain's DNS A record to your server IP, then:

cat > Caddyfile << 'EOF'
ai.yourdomain.com {
    reverse_proxy open-webui:8080
}
EOF
Enter fullscreen mode Exit fullscreen mode

Replace ai.yourdomain.com with your real domain.

Option B: no domain (testing only, plain HTTP).

cat > Caddyfile << 'EOF'
:80 {
    reverse_proxy open-webui:8080
}
EOF
Enter fullscreen mode Exit fullscreen mode

Step 6: Open the firewall

Only SSH, HTTP and HTTPS are needed:

sudo ufw allow 22/tcp
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
Enter fullscreen mode Exit fullscreen mode

⭐ Important: Docker can bypass UFW for any port you publish with ports:. That is why Ollama has no published port in the compose file above. Never publish 11434 to the internet.

Step 7: Start everything

docker compose up -d
Enter fullscreen mode Exit fullscreen mode

Check that all three containers are running:

docker compose ps
Enter fullscreen mode Exit fullscreen mode

Step 8: Download your first model

docker exec -it ollama ollama pull llama3.1:8b
Enter fullscreen mode Exit fullscreen mode

Test it from the command line:

docker exec -it ollama ollama run llama3.1:8b "Explain what a dedicated server is in two sentences."
Enter fullscreen mode Exit fullscreen mode

Verify the model is using the GPU (the most important check)

While a model is loaded, run:

docker exec -it ollama ollama ps
Enter fullscreen mode Exit fullscreen mode

Look at the PROCESSOR :

  • 100% GPU: Perfect: the whole model is in VRAM
  • 40%/60%: Model is too big for VRAM, partly on CPU (slow). Use a smaller model or lower quantization
  • CPU/GPU: Model is too big for VRAM, partly on CPU (slow). Use a smaller model or lower quantization
  • 100% CPU: GPU is not being used. See Troubleshooting, Problem 3

You can also watch GPU usage live in a second terminal:

watch -n 1 nvidia-smi
Enter fullscreen mode Exit fullscreen mode

Step 9: Open the chat interface and create your admin account

Visit https://ai.yourdomain.com (or http://YOUR_SERVER_IP for Option B).

  • Click Get Started.
  • Create your account. The first account becomes the administrator.
  • Pick your model from the dropdown at the top and start chatting.

Lock down registration (do this right after creating your admin)
Otherwise anyone who finds your URL can sign up.

sed -i 's/ENABLE_SIGNUP=true/ENABLE_SIGNUP=false/' .env
docker compose up -d
Enter fullscreen mode Exit fullscreen mode

⚠️ Note In some Open WebUI versions, the signup setting is also stored in the admin panel after first start. If signups are still open, go to Admin Panel → Settings → General and turn off Enable New Sign Ups.

Which model should I run? (VRAM guide)

Rule of thumb for 4-bit quantized models (the Ollama default): model size in GB is roughly VRAM needed, plus 1 to 2 GB for context.

Your GPU VRAM: 6 to 8 GB
Comfortable model size: 3B to 8B
Example use: Chat, summaries, simple coding help

Your GPU VRAM: 12 to 16 GB
Comfortable model size: 8B to 14B
Example use: Better reasoning, coding

Your GPU VRAM: 24 GB
Comfortable model size: up to ~30B
Example use: Strong general assistant

Your GPU VRAM: 48 GB
Comfortable model size: up to ~70B
Example use: Near top-tier open models

Your GPU VRAM: 80 GB+
Comfortable model size: 70B+ with long context
Example use: Heavy or multi-user workloads

Browse available models at ollama.com/library and pull any with:

docker exec -it ollama ollama pull MODEL_NAME:TAG
Enter fullscreen mode Exit fullscreen mode

List installed models, and delete ones you no longer need:

docker exec -it ollama ollama list
docker exec -it ollama ollama rm MODEL_NAME:TAG
Enter fullscreen mode Exit fullscreen mode

Using the API from your own apps

Ollama is OpenAI-compatible. Because we did not publish the port, call it from inside the Docker network, or temporarily test from the server itself:

docker exec -it ollama curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1:8b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Enter fullscreen mode Exit fullscreen mode

For remote API access, put it behind an authenticated gateway or use a VPN such as WireGuard. Never expose raw Ollama to the internet.

Troubleshooting

Problem 1 - nvidia-smi says "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver"

Likely causes and fixes:

  • You have not rebooted after installing the driver. Run sudo reboot.
  • Secure Boot is on. Check with mokutil --sb-state. If enabled, disable Secure Boot in BIOS/IPMI or enroll the MOK key.
  • Kernel and driver mismatch after a kernel update. Reinstall:
sudo apt install --reinstall linux-headers-$(uname -r)
sudo ubuntu-drivers autoinstall
sudo reboot
Enter fullscreen mode Exit fullscreen mode
  • The open-source nouveau driver is loaded. Check with lsmod | grep nouveau. If it shows, blacklist it:
echo -e "blacklist nouveau\noptions nouveau modeset=0" | sudo tee /etc/modprobe.d/blacklist-nouveau.conf
sudo update-initramfs -u && sudo reboot
Enter fullscreen mode Exit fullscreen mode

Problem 2 -  docker: Error response from daemon: could not select device driver "" with capabilities: [[gpu]]

The NVIDIA Container Toolkit is missing or Docker was not configured. Re-run:

sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Enter fullscreen mode Exit fullscreen mode

Then retest: docker run - rm - gpus all ubuntu nvidia-smi

Problem 3 -  Ollama is running on CPU, not GPU ("Ollama not using GPU")

Work through this checklist in order:

  • Does nvidia-smi work on the host? If not, go to Problem 1.
  • Does docker run --rm --gpus all ubuntu nvidia-smi work? If not, go to Problem 2.
  • Does your compose file contain the deploy.resources.reservations.devices block for the ollama service? Without it, the container gets no GPU.
  • Check the Ollama logs for GPU detection messages:
docker logs ollama 2>&1 | grep -i -E "gpu|cuda|vram"
Enter fullscreen mode Exit fullscreen mode
  • Is the model simply too large for your VRAM? Run ollama ps. A split such as 30%/70% CPU/GPU means it is overflowing. Use a smaller model.

Problem 4 -  GPU works at first, then Ollama falls back to CPU after some time

This is a known issue on some Linux setups where the GPU driver state is lost. Try reloading the UVM module and restarting the container:

sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm
docker restart ollama
Enter fullscreen mode Exit fullscreen mode

If it keeps happening, make sure you are on a recent driver version and keep the system updated.

Problem 5 - Open WebUI shows no models, or "Ollama connection error"

  • Confirm OLLAMA_BASE_URL=http://ollama:11434 is set and both containers are in the same compose project.
  • Check Ollama is up: docker compose ps and docker logs ollama.
  • Confirm you actually pulled a model: docker exec -it ollama ollama list.
  • Restart: docker compose restart open-webui.

Problem 6 - HTTPS does not work / Caddy shows certificate errors

  • DNS A record must point to your server IP before starting Caddy. Check with dig +short ai.yourdomain.com.
  • Ports 80 and 443 must be open in UFW and any provider-level firewall.
  • Read the logs: docker logs caddy.
  • If you hit certificate rate limits from repeated failed attempts, wait and retry, or test with Option B first.

Problem 7 - "CUDA out of memory" or the model fails to load

  • Use a smaller model or a smaller quantization.
  • Reduce context length in Open WebUI (Settings → Advanced Params → Context Length).
  • Make sure no other process is holding VRAM: nvidia-smi shows processes using the GPU.

Problem 8 - Responses are very slow

Run ollama ps: if it is not 100% GPU, that is your cause.

  • Large context windows slow things down. Lower the context length.
  • First request after idle is slower because the model loads into VRAM. OLLAMA_KEEP_ALIVE=30m (already set) reduces this.

Problem 9 - Port 3000 or 11434 is reachable from the internet

You should not publish these ports. Remove any ports: entries for ollama and open-webui, then run docker compose up -d. Remember Docker bypasses UFW for published ports.

Maintenance

Update to the latest versions:

cd ~/ai-server
docker compose pull
docker compose up -d
docker image prune -f
Enter fullscreen mode Exit fullscreen mode

Back up your data (chats, users, settings):

docker run --rm -v ai-server_open_webui_data:/data -v $(pwd):/backup ubuntu \
  tar czf /backup/open-webui-backup.tar.gz /data
Enter fullscreen mode Exit fullscreen mode

(The volume name is your folder name plus _open_webui_data. Check yours with docker volume ls.)

Useful daily commands:

  • See running containers:docker compose ps
  • View logs:docker compose logs -f
  • Loaded models:docker exec -it ollama ollama ps
  • Stop everything:docker compose down
  • Start everything:docker compose up -d 
  • GPU usage:nvidia-smi

Security checklist

  • SSH key login only (disable password login)
  • Ollama port is not published
  • Sign-ups disabled after creating the admin
  • HTTPS enabled via Caddy
  • WEBUI_SECRET_KEY set to a random value
  • Firewall allows only ports 22, 80, 443
  • Regular updates and backups scheduled

Top comments (0)