DEV Community

Cover image for Building Sarrera: Self-Hosted Enterprise AI Inference Gateway with RBAC, Token Quotas & Telemetry
Mario Ezquerro for Google Developer Experts

Posted on Originally published at github.com

Building Sarrera: Self-Hosted Enterprise AI Inference Gateway with RBAC, Token Quotas & Telemetry

Engineering teams worldwide face a common dilemma when adopting generative AI:

  1. Cloud AI privacy & security risks: Sending proprietary source code to third-party APIs (OpenAI, Anthropic) triggers compliance alarms.
  2. Uncontrolled cloud billing: A handful of developers running autonomous agents (Cline, Roo Code, Cursor) can easily run up thousands of dollars in surprise monthly token bills.
  3. Hardware fragmentation: On-premise clusters often consist of heterogeneous hardwareβ€”a few servers with NVIDIA A100/RTX 4090s, some mid-range workstation GPUs (RTX 3060/4060), and fallback CPU clusters. Without smart routing, high-end GPUs saturate while other machines sit idle.

To solve this, we built and open-sourced Sarrera ("Entryway / Portal" in Basque) β€” an enterprise local AI inference gateway, access control (RBAC), and observability platform packaged into a single Docker Compose deployment.

In this article, I will walk you through the architecture, multi-tier subscription quotas, dynamic node management, and how you can deploy your own private AI hub in under 5 minutes.


πŸ›οΈ High-Level Architecture

Sarrera decouples client IDEs from physical compute hardware, shielding your internal network behind a single perimeter reverse proxy.

                                [ Internet / Corporate LAN ]
                                              β”‚
                                              β–Ό (Ports 80 / 443 Only)
                             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                             β”‚       ai-caddy (Caddy v2)         β”‚
                             β”‚  - Central TLS Termination        β”‚
                             β”‚  - Sarrera Service Hub & Portal   β”‚
                             β”‚  - Security Headers (HSTS)        β”‚
                             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                               β”‚
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β–Ό                        β–Ό                         β–Ό                        β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Open WebUI      β”‚    β”‚  LiteLLM Proxy   β”‚      β”‚  Langfuse v2     β”‚     β”‚  MinIO Console   β”‚
β”‚  (Chat Portal)   β”‚    β”‚  (Gateway/Admin) β”‚      β”‚  (Observability) β”‚     β”‚  (Trace Storage) β”‚
β”‚  Internal :8080  β”‚    β”‚  Internal :4000  β”‚      β”‚  Internal :3000  β”‚     β”‚  Internal :9001  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β–Ό (least-busy)   β–Ό                β–Ό
          [GPU Premium]    [GPU Standard]    [CPU Cluster]
         (A100 / RTX 4090) (RTX 3060/4060)  (AVX-512 CPU)
           :11434           :11434            :11434
Enter fullscreen mode Exit fullscreen mode

The 7 Core Building Blocks

  1. Edge Reverse Proxy (Caddy v2): Serves as the single exposed internet entrypoint on ports 80/443. Manages automated TLS certificate lifecycles (Let's Encrypt / ZeroSSL / Internal CA), and serves the integrated Sarrera Service Hub.
  2. AI Gateway & Router (LiteLLM Proxy): Standard /v1/chat/completions OpenAI-compatible API gateway. Enforces subscription Tiers, virtual API keys, and weighted least-busy load balancing.
  3. Observability & Auditing (Langfuse v2): Asynchronous, non-blocking telemetry engine recording every prompt, completion, Time-To-First-Token (TTFT), token sum, and department cost attribution.
  4. Chat & Prompt Portal (Open WebUI): Interactive chat interface for non-developer staff, document RAG, and Active Directory / LDAP authentication.
  5. Relational State (PostgreSQL 16): Dedicated persistent databases for LiteLLM keys/budgets and Langfuse trace metadata.
  6. Object Storage (MinIO S3): High-performance S3 storage bucket retaining large trace payloads and prompt attachments.
  7. Local Stand-in Engine (Ollama): Local container with network aliases (gpu-entry-node, cpu-cluster-node, gpu-premium-node) for instant testing without external GPU dependencies.

πŸ›‘οΈ 3-Tier Subscription & Token Quota Governance

To avoid budget blowouts, Sarrera maps developer virtual API keys to LiteLLM Teams representing corporate subscription tiers:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           SARRERA SUBSCRIPTION TIERS                          β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚      tier-basic       β”‚         tier-standard         β”‚     tier-premium      β”‚
β”‚   (Junior / Entry)    β”‚        (Pro / Standard)       β”‚     (Lead / Expert)   β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ β€’ basic-coder (7B)    β”‚ β€’ basic-coder (7B)            β”‚ β€’ basic-coder (7B)    β”‚
β”‚                       β”‚ β€’ premium-coder (32B)         β”‚ β€’ premium-coder (32B) β”‚
β”‚                       β”‚                               β”‚ β€’ premium-reasoning   β”‚
β”‚                       β”‚                               β”‚   (DeepSeek-R1)       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Budget: 15 EUR/month  β”‚ Budget: 50 EUR/month          β”‚ Budget: 100 EUR/month β”‚
β”‚ 60 RPM Β· 30k TPM      β”‚ 120 RPM Β· 60k TPM             β”‚ 180 RPM Β· 120k TPM    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Enter fullscreen mode Exit fullscreen mode

Sub-Millisecond Enforcement

When a developer in VS Code triggers autocomplete:

  1. LiteLLM checks if the requested model is whitelisted for their tier.
  2. If a junior developer with tier-basic requests premium-reasoning (DeepSeek-R1), the gateway immediately returns HTTP 403 Forbidden in 10 milliseconds without touching the GPU cluster.
  3. If accumulated monthly spend exceeds the team's cap, it returns HTTP 400 Budget Exceeded.
  4. If rate limits are exceeded, it returns HTTP 429 Too Many Requests.

⚑ Heterogeneous Load Balancing: Do You Need HAProxy?

A frequent question from infrastructure engineers is: "Do we need HAProxy or Nginx to balance traffic across our inference servers?"

The short answer is NO. Traditional Layer 4 / Layer 7 proxies like HAProxy only see HTTP byte streams and status codes. They do not understand:

  • VRAM memory footprints per request.
  • The difference between a 20-token autocomplete versus a 4,000-token multi-file refactoring.
  • Token streaming (text/event-stream), TTFT latencies, or Out-Of-Memory (OOM) GPU states.

LiteLLM is an LLM-aware Application Router. In config/litellm-config.yaml:

model_list:
  # Route 1: Mid-range dedicated GPU (RTX 4060)
  - model_name: basic-coder
    litellm_params:
      model: ollama/qwen2.5-coder:7b
      api_base: http://192.168.1.50:11434
      weight: 8
      rpm: 60

  # Route 2: Fallback CPU cluster
  - model_name: basic-coder
    litellm_params:
      model: ollama/qwen2.5-coder:7b
      api_base: http://192.168.1.51:11434
      weight: 2
      rpm: 20

router_settings:
  routing_strategy: "least-busy"
  timeout: 45
  num_retries: 2
Enter fullscreen mode Exit fullscreen mode
  • least-busy routing: Dynamically tracks active requests in flight and routes the next query to the least saturated host.
  • Failover & Retries (num_retries: 2): If a GPU server crashes or encounters a kernel stall, LiteLLM transparently retries the query against the alternate node before returning an error to the developer.

πŸ–₯️ Live Dynamic Node Administration (Zero-Downtime)

Editing configuration files and restarting containers during office hours is impractical.

We built an Authenticated Admin Control Center directly into the Caddy web interface (https://localhost/):

  1. Click πŸ” Admin Control Panel in the top navigation.
  2. Sign in with the credentials defined in .env (ADMIN_USERNAME and ADMIN_PASSWORD).
  3. You can:
    • Register new GPU/CPU nodes on the fly: Provide the node IP (api_base), backend model, engine (Ollama, vLLM, TGI), weight, and RPM limits.
    • Ping & test latency: Verify network connectivity before saving.
    • Persist in PostgreSQL: Saves the node directly into LiteLLM without touching YAML files or restarting Docker.
    • Issue scoped developer keys: Create keys for tier-basic, tier-standard, or tier-premium with one click.

πŸ’» Developer Client Integration (VS Code Continue)

Developers consume the platform just like OpenAI:

In ~/.continue/config.json:

{
  "models": [
    {
      "title": "Sarrera (tier-standard)",
      "provider": "openai",
      "model": "premium-coder",
      "apiKey": "sk-your-virtual-key-here",
      "apiBase": "https://ai.company.local/v1"
    }
  ],
  "tabAutocompleteModel": {
    "title": "Sarrera Autocomplete",
    "provider": "openai",
    "model": "basic-coder",
    "apiKey": "sk-your-virtual-key-here",
    "apiBase": "https://ai.company.local/v1"
  }
}
Enter fullscreen mode Exit fullscreen mode

Every autocomplete and chat interaction appears in Langfuse in real time:

  • Exact prompt and generated code snippet.
  • Exact tokens consumed (prompt_tokens, completion_tokens).
  • Time-To-First-Token (TTFT) and total latency.
  • Tagged with the developer's username and department for monthly chargeback.

πŸš€ Quickstart: Deploy in 5 Minutes

1. Clone & Configure

git clone https://github.com/Sarrera/sarrera.git
cd sarrera
cp .env.example .env
Enter fullscreen mode Exit fullscreen mode

2. Launch the Stack

docker compose up -d
Enter fullscreen mode Exit fullscreen mode

3. Bootstrap the 3 Subscription Tiers

./scripts/bootstrap-tiers.sh
Enter fullscreen mode Exit fullscreen mode

4. Issue a Virtual API Key

./scripts/issue-key.sh alex tier-standard 90d "Core-Engineering"
Enter fullscreen mode Exit fullscreen mode

5. Run the End-to-End Validation Suite

./scripts/smoke-test.sh
Enter fullscreen mode Exit fullscreen mode

Open https://localhost/ in your browser to access the Sarrera Central Portal!


πŸ“¦ What's Next & Resources

  • GitHub Repository: https://github.com/Sarrera/sarrera
  • Documentation Portal: Hosted on GitHub Pages inside the repository (/docs), featuring complete runbooks for Jekyll, Docsify, and MkDocs.

If your organization is exploring private, compliant local AI inference, check out the repository, give it a star ⭐, and let me know your thoughts in the comments!

Top comments (0)