Engineering teams worldwide face a common dilemma when adopting generative AI:
- Cloud AI privacy & security risks: Sending proprietary source code to third-party APIs (OpenAI, Anthropic) triggers compliance alarms.
- Uncontrolled cloud billing: A handful of developers running autonomous agents (Cline, Roo Code, Cursor) can easily run up thousands of dollars in surprise monthly token bills.
- Hardware fragmentation: On-premise clusters often consist of heterogeneous hardwareβa few servers with NVIDIA A100/RTX 4090s, some mid-range workstation GPUs (RTX 3060/4060), and fallback CPU clusters. Without smart routing, high-end GPUs saturate while other machines sit idle.
To solve this, we built and open-sourced Sarrera ("Entryway / Portal" in Basque) β an enterprise local AI inference gateway, access control (RBAC), and observability platform packaged into a single Docker Compose deployment.
In this article, I will walk you through the architecture, multi-tier subscription quotas, dynamic node management, and how you can deploy your own private AI hub in under 5 minutes.
ποΈ High-Level Architecture
Sarrera decouples client IDEs from physical compute hardware, shielding your internal network behind a single perimeter reverse proxy.
[ Internet / Corporate LAN ]
β
βΌ (Ports 80 / 443 Only)
βββββββββββββββββββββββββββββββββββββ
β ai-caddy (Caddy v2) β
β - Central TLS Termination β
β - Sarrera Service Hub & Portal β
β - Security Headers (HSTS) β
βββββββββββββββββββ¬ββββββββββββββββββ
β
ββββββββββββββββββββββββββ¬βββββββββββββ΄βββββββββββββ¬βββββββββββββββββββββββββ
βΌ βΌ βΌ βΌ
ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
β Open WebUI β β LiteLLM Proxy β β Langfuse v2 β β MinIO Console β
β (Chat Portal) β β (Gateway/Admin) β β (Observability) β β (Trace Storage) β
β Internal :8080 β β Internal :4000 β β Internal :3000 β β Internal :9001 β
ββββββββββββββββββββ βββββββββββ¬βββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
β
ββββββββββββββββββΌβββββββββββββββββ
βΌ (least-busy) βΌ βΌ
[GPU Premium] [GPU Standard] [CPU Cluster]
(A100 / RTX 4090) (RTX 3060/4060) (AVX-512 CPU)
:11434 :11434 :11434
The 7 Core Building Blocks
-
Edge Reverse Proxy (
Caddy v2): Serves as the single exposed internet entrypoint on ports 80/443. Manages automated TLS certificate lifecycles (Let's Encrypt / ZeroSSL / Internal CA), and serves the integrated Sarrera Service Hub. -
AI Gateway & Router (
LiteLLM Proxy): Standard/v1/chat/completionsOpenAI-compatible API gateway. Enforces subscription Tiers, virtual API keys, and weightedleast-busyload balancing. -
Observability & Auditing (
Langfuse v2): Asynchronous, non-blocking telemetry engine recording every prompt, completion, Time-To-First-Token (TTFT), token sum, and department cost attribution. -
Chat & Prompt Portal (
Open WebUI): Interactive chat interface for non-developer staff, document RAG, and Active Directory / LDAP authentication. -
Relational State (
PostgreSQL 16): Dedicated persistent databases for LiteLLM keys/budgets and Langfuse trace metadata. -
Object Storage (
MinIO S3): High-performance S3 storage bucket retaining large trace payloads and prompt attachments. -
Local Stand-in Engine (
Ollama): Local container with network aliases (gpu-entry-node,cpu-cluster-node,gpu-premium-node) for instant testing without external GPU dependencies.
π‘οΈ 3-Tier Subscription & Token Quota Governance
To avoid budget blowouts, Sarrera maps developer virtual API keys to LiteLLM Teams representing corporate subscription tiers:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β SARRERA SUBSCRIPTION TIERS β
βββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββ€
β tier-basic β tier-standard β tier-premium β
β (Junior / Entry) β (Pro / Standard) β (Lead / Expert) β
βββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββ€
β β’ basic-coder (7B) β β’ basic-coder (7B) β β’ basic-coder (7B) β
β β β’ premium-coder (32B) β β’ premium-coder (32B) β
β β β β’ premium-reasoning β
β β β (DeepSeek-R1) β
βββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββ€
β Budget: 15 EUR/month β Budget: 50 EUR/month β Budget: 100 EUR/month β
β 60 RPM Β· 30k TPM β 120 RPM Β· 60k TPM β 180 RPM Β· 120k TPM β
βββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββ
Sub-Millisecond Enforcement
When a developer in VS Code triggers autocomplete:
- LiteLLM checks if the requested model is whitelisted for their tier.
- If a junior developer with
tier-basicrequestspremium-reasoning(DeepSeek-R1), the gateway immediately returnsHTTP 403 Forbiddenin 10 milliseconds without touching the GPU cluster. - If accumulated monthly spend exceeds the team's cap, it returns
HTTP 400 Budget Exceeded. - If rate limits are exceeded, it returns
HTTP 429 Too Many Requests.
β‘ Heterogeneous Load Balancing: Do You Need HAProxy?
A frequent question from infrastructure engineers is: "Do we need HAProxy or Nginx to balance traffic across our inference servers?"
The short answer is NO. Traditional Layer 4 / Layer 7 proxies like HAProxy only see HTTP byte streams and status codes. They do not understand:
- VRAM memory footprints per request.
- The difference between a 20-token autocomplete versus a 4,000-token multi-file refactoring.
- Token streaming (
text/event-stream), TTFT latencies, or Out-Of-Memory (OOM) GPU states.
LiteLLM is an LLM-aware Application Router. In config/litellm-config.yaml:
model_list:
# Route 1: Mid-range dedicated GPU (RTX 4060)
- model_name: basic-coder
litellm_params:
model: ollama/qwen2.5-coder:7b
api_base: http://192.168.1.50:11434
weight: 8
rpm: 60
# Route 2: Fallback CPU cluster
- model_name: basic-coder
litellm_params:
model: ollama/qwen2.5-coder:7b
api_base: http://192.168.1.51:11434
weight: 2
rpm: 20
router_settings:
routing_strategy: "least-busy"
timeout: 45
num_retries: 2
-
least-busyrouting: Dynamically tracks active requests in flight and routes the next query to the least saturated host. -
Failover & Retries (
num_retries: 2): If a GPU server crashes or encounters a kernel stall, LiteLLM transparently retries the query against the alternate node before returning an error to the developer.
π₯οΈ Live Dynamic Node Administration (Zero-Downtime)
Editing configuration files and restarting containers during office hours is impractical.
We built an Authenticated Admin Control Center directly into the Caddy web interface (https://localhost/):
- Click
π Admin Control Panelin the top navigation. - Sign in with the credentials defined in
.env(ADMIN_USERNAMEandADMIN_PASSWORD). - You can:
-
Register new GPU/CPU nodes on the fly: Provide the node IP (
api_base), backend model, engine (Ollama, vLLM, TGI), weight, and RPM limits. - Ping & test latency: Verify network connectivity before saving.
- Persist in PostgreSQL: Saves the node directly into LiteLLM without touching YAML files or restarting Docker.
-
Issue scoped developer keys: Create keys for
tier-basic,tier-standard, ortier-premiumwith one click.
-
Register new GPU/CPU nodes on the fly: Provide the node IP (
π» Developer Client Integration (VS Code Continue)
Developers consume the platform just like OpenAI:
In ~/.continue/config.json:
{
"models": [
{
"title": "Sarrera (tier-standard)",
"provider": "openai",
"model": "premium-coder",
"apiKey": "sk-your-virtual-key-here",
"apiBase": "https://ai.company.local/v1"
}
],
"tabAutocompleteModel": {
"title": "Sarrera Autocomplete",
"provider": "openai",
"model": "basic-coder",
"apiKey": "sk-your-virtual-key-here",
"apiBase": "https://ai.company.local/v1"
}
}
Every autocomplete and chat interaction appears in Langfuse in real time:
- Exact prompt and generated code snippet.
- Exact tokens consumed (
prompt_tokens,completion_tokens). - Time-To-First-Token (TTFT) and total latency.
- Tagged with the developer's username and department for monthly chargeback.
π Quickstart: Deploy in 5 Minutes
1. Clone & Configure
git clone https://github.com/Sarrera/sarrera.git
cd sarrera
cp .env.example .env
2. Launch the Stack
docker compose up -d
3. Bootstrap the 3 Subscription Tiers
./scripts/bootstrap-tiers.sh
4. Issue a Virtual API Key
./scripts/issue-key.sh alex tier-standard 90d "Core-Engineering"
5. Run the End-to-End Validation Suite
./scripts/smoke-test.sh
Open https://localhost/ in your browser to access the Sarrera Central Portal!
π¦ What's Next & Resources
- GitHub Repository: https://github.com/Sarrera/sarrera
-
Documentation Portal: Hosted on GitHub Pages inside the repository (
/docs), featuring complete runbooks for Jekyll, Docsify, and MkDocs.
If your organization is exploring private, compliant local AI inference, check out the repository, give it a star β, and let me know your thoughts in the comments!
Top comments (0)