As developers, we’ve all been there: you build an AI-powered feature, deploy it to production, and suddenly get slapped with a massive OpenAI or Anthropic API bill because a user ran a recursive loop, or because your team shared raw API keys across staging environments. 💸
Sharing master API keys is a major security risk, and tracking individual user costs across multiple models (GPT-4o, Claude 3.5 Sonnet, Llama-3) is a management nightmare.
There is a better way. By self-hosting LiteLLM, you can spin up a unified, private LLM Gateway that acts as a proxy between your applications and your AI providers.
With LiteLLM, you get:
- One Unified API: Call 100+ LLMs using the exact same OpenAI SDK structure.
- Granular Cost Tracking: Create virtual keys for teams, developers, or users with strict spending limits (e.g., max $5/day).
- Load Balancing & Failover: Automatically route requests to backup API keys or models if rate limits are hit.
- Semantic Caching: Cache identical queries in Redis to slash API costs by up to 40%.
In this guide, we'll deploy a production-ready LiteLLM Instance with a beautiful UI Admin Dashboard on a high-speed Vultr High-Performance Cloud Instance in under 10 minutes.
Why Host on Vultr?
When proxying API requests, latency is everything. Adding a proxy layer shouldn't introduce lag. Vultr's High-Performance Cloud offers high-frequency CPU cores and NVMe storage in over 32 global locations. This means your proxy can be co-located right next to your target audience (or your main application servers) to ensure sub-millisecond routing overhead.
Step 1: Provision Your Vultr Instance
- Sign up or log into your account via Vultr.
- Click Deploy Server and select Cloud Compute.
- Choose High Performance (AMD or Intel) to ensure ultra-fast request routing.
- Select a server location closest to your main applications.
- Choose Ubuntu 24.04 LTS as your Operating System.
- Select the 2 GB RAM / 1 vCPU plan (which is more than enough to handle thousands of concurrent proxy requests).
- Add your SSH key and click Deploy Now.
Once your instance is ready, copy the IP address and connect to it via your terminal:
ssh root@YOUR_VULTR_IP
Step 2: Install Docker and Docker Compose
Update your system packages and install Docker to manage our containers easily:
sudo apt update && sudo apt upgrade -y
sudo apt install -y docker.io docker-compose
Verify the installation:
docker --version && docker-compose --version
Step 3: Create the LiteLLM Configuration
Create a dedicated directory for LiteLLM:
mkdir litellm-gateway && cd litellm-gateway
Now, create a litellm_config.yaml file. This is where you configure the models you want to proxy. For this demo, we'll configure OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet:
model_list:
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: "os.environ/OPENAI_API_KEY"
- model_name: claude-3-5-sonnet
litellm_params:
model: anthropic/claude-3-5-sonnet-20240620
api_key: "os.environ/ANTHROPIC_API_KEY"
# Configure global database for key management and logs
router_settings:
routing_strategy: latency-based-routing
general_settings:
master_key: "os.environ/LITELLM_MASTER_KEY"
Step 4: Define the Docker Compose Stack
We will set up three components:
- LiteLLM Core: The proxy service.
- PostgreSQL: Database to store your custom API keys, user budgets, and audit logs.
- LiteLLM Admin UI: A clean dashboard to create keys, set limits, and view live metrics.
Create a docker-compose.yml file in the same directory:
version: '3.8'
services:
db:
image: postgres:16-alpine
container_name: litellm-db
environment:
POSTGRES_DB: litellm
POSTGRES_USER: litellm_user
POSTGRES_PASSWORD: super_secret_db_password
volumes:
- pgdata:/var/lib/postgresql/data
ports:
- "5432:5432"
restart: always
litellm:
image: ghcr.io/berriai/litellm:main-latest
container_name: litellm-proxy
ports:
- "4000:4000"
volumes:
- ./litellm_config.yaml:/app/config.yaml
environment:
- DATABASE_URL=postgresql://litellm_user:super_secret_db_password@db:5432/litellm
- LITELLM_MASTER_KEY=sk-your-super-secure-master-admin-key-12345
- OPENAI_API_KEY=your_actual_openai_api_key_here
- ANTHROPIC_API_KEY=your_actual_anthropic_api_key_here
depends_on:
- db
command: ["--config", "/app/config.yaml", "--detailed_debug"]
restart: always
volumes:
pgdata:
💡 Note: Replace
your_actual_openai_api_key_hereandyour_actual_anthropic_api_key_herewith your real API keys, and change theLITELLM_MASTER_KEYto a secure, random string.
Step 5: Start Your Gateway
Run the Docker stack in detached mode:
docker-compose up -d
Check if everything is running correctly:
docker-compose ps
Step 6: Accessing the Admin Dashboard
Your gateway is now active!
- Open your browser and go to
http://YOUR_VULTR_IP:4000/ui. - Log in using the
LITELLM_MASTER_KEYyou configured in yourdocker-compose.yml(e.g.,sk-your-super-secure-master-admin-key-12345).
From this dashboard, you can:
- Create Virtual Keys for external developers or internal microservices.
- Set Budgets: Limit a key to a specific dollar amount (e.g., $10 total limit, or resetting daily).
- View Real-time Analytics: Track exactly which developer, key, or model is consuming the most tokens.
How to Use Your Private Gateway in Code
To consume your self-hosted gateway, you simply point your existing OpenAI SDK to your Vultr server. No code changes required!
Python Example:
from openai import OpenAI
client = OpenAI(
api_key="sk-the-virtual-key-you-created-in-dashboard",
base_url="http://YOUR_VULTR_IP:4000"
)
# Call any model you defined in your LiteLLM config!
response = client.chat.completions.create(
model="claude-3-5-sonnet",
messages=[{"role": "user", "content": "Explain quantum computing in 2 sentences."}]
)
print(response.choices[0].message.content)
Secure Your Setup for Production
Before pointing production applications to your gateway, ensure you secure it with SSL. You can easily install Nginx and Certbot on your Vultr instance to get a free Let's Encrypt SSL certificate:
sudo apt install -y nginx certbot python3-certbot-nginx
Configure Nginx to reverse proxy traffic from port 80/443 to http://localhost:4000 and run sudo certbot --nginx to enable HTTPS.
Ready to get your team's AI costs under lock and key? Build your high-performance, private LLM gateway today on Vultr High-Performance Cloud!
Liked this resource? Join our daily Telegram channel for more developer tools and cloud insights: @Libretech2026
Top comments (0)