Our ML Services need LLMs to process the large documents and data. We started using Ollama since 2024 and the learnings below:
Ollama, an open-source tool that packages LLMs into a simple CLI and REST server, has emerged as a game-changer for developers and organizations seeking privacy, cost-efficiency, and offline capabilities
*What is Ollama and Why Use It? *
Ollama is often described as "Docker for LLMs" – you pull a model by name, and Ollama handles quantization, memory mapping, and inference runtime behind the scenes. The core reasons developers choose Ollama over cloud APIs include:
- Privacy: Your prompts and data never leave your machine
- Cost: Zero per-token charges after the one-time model download
- Latency: No network round-trip, especially fast on modern GPUs
- Offline use: Works without any internet connection once models are downloaded
Our ML Services Environment
Our setup runs a variety of models optimized for different tasks, including:
- Qwen3 series (8B, 4B-instruct, 0.6B) for general-purpose inference
- Llama3.2 and custom variants for specialized tasks
- Gemma3:4b and Gemma2:2b for lightweight processing
- BGE-M3 for embedding generation
- llm – a custom model created from GGUF format
We also maintain two dedicated services:
- ner: For Named Entity Recognition in forms
- summarizer-llm: For summarizing signals and communications
Complete Setup Walkthrough
- Install Dependencies
Start by updating your system and installing essential packages:
apt update && apt upgrade -y
apt install git docker-compose curl
- Install NVIDIA Drivers and Container Toolkit
For GPU acceleration (which provides 10-100x speedup over CPU inference) , install NVIDIA drivers and the NVIDIA Container Toolkit:
distribution=$(. /etc/os-release;echo $ID$VERSION_ID)
curl -s -L https://nvidia.github.io/nvidia-docker/gpgkey | sudo apt-key add -
curl -s -L https://nvidia.github.io/nvidia-docker/$distribution/nvidia-docker.list | sudo tee /etc/apt/sources.list.d/nvidia-docker.list
sudo apt-get install -y nvidia-docker2
sudo systemctl restart docker
sudo ubuntu-drivers autoinstall
Reboot the system after driver installation.
- Clone and Configure the Ollama Repository
git clone https://github.com/sujithrpillai/ollama.git
cd ollama
Create a docker-compose.yml file with GPU support:
version: '3.8'
services:
ollama:
image: ollama/ollama:latest
hostname: ollama
ports:
- "11434:11434"
volumes:
- ./models:/root/.ollama/models
networks:
- genai-network
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
runtime: nvidia
restart: always
open-webui:
image: ghcr.io/open-webui/open-webui:main
ports:
- "3000:8080"
environment:
- OLLAMA_BASE_URL=http://ollama:11434
volumes:
- ./backend/data:/app/backend/data
networks:
- genai-network
restart: always
networks:
genai-network:
driver: bridge
name: genai-network
The GPU configuration in the deploy section enables hardware acceleration for supported NVIDIA GPUs .
- Start the Services
docker-compose up -d
- Pull Models
Access the Ollama container and download models:
docker exec -it ollama ollama pull qwen3:8b
docker exec -it ollama ollama pull gemma2:2b
docker exec -it ollama ollama pull llama3.2:latest
- Create Custom Models from GGUF Files
Ollama allows creating custom models from GGUF format files :
sudo apt-get install git-lfs
git clone https://huggingface.co/org/llm-v3-270k-GGUF
curl -fsSL https://ollama.com/install.sh | sh
Create a Modelfile:
FROM /home/RatnaDveLinga-v3-270k-GGUF/RatnaDveLinga-v3-270k.gguf
Then create the model:
ollama create llmname -f Modelfile
- Verify Setup
Check models available:
root@ollama:/ ollama list
NAME ID SIZE MODIFIED
qwen3:8b 500a1f067a9f 5.2 GB 8 months ago
qwen3:4b-instruct 0edcdef34593 2.5 GB 8 months ago
qwen2.5:1.5b 65ec06548149 986 MB 9 months ago
gemma3:4b a2af6cc3eb7f 3.3 GB 13 months ago
llama3.2:latest a80c4f17acd5 2.0 GB 13 months ago
bge-m3:latest 790764642607 1.2 GB 13 months ago
RatnaDveLinga:latest decc819a53b1 1.7 GB 21 months ago
Model Management Commands
Essential CLI commands for managing your Ollama setup :
| Command | Description |
|---|---|
ollama list |
List all downloaded models |
ollama pull <model> |
Download a model without running it |
ollama run <model> |
Run a model interactively |
ollama stop <model> |
Stop a running model |
ollama rm <model> |
Remove a model from disk |
ollama ps |
Show currently running models |
ollama show <model> |
Display model metadata |
Performance Considerations
GPU acceleration can provide 10-100x faster inference compared to CPU-only operation . To verify GPU usage:
docker exec -it ollama nvidia-smi
The Open WebUI interface available at http://localhost:3000 provides a user-friendly way to interact with models, pull new ones, and manage your setup without using the command line .
Conclusion
This setup provides a robust, self-hosted LLM environment with GPU acceleration, a comprehensive model library, and a web interface for easy interaction. Whether you're running specialized NER services, summarization tasks, or general-purpose inference, Ollama combined with Open WebUI delivers a powerful and private AI platform.
Happy prompting!
Top comments (0)