Quick Answer
Deploying Agentic Workloads: Learn how senior engineers can reliably deploy agentic AI workloads on Azure AI Foundry using Kubernetes, with concrete architecture, code, cost, and failure‑mode guidance.
Deploying Agentic Workloads at Scale: A Practical Blueprint
Deploying Agentic Workloads is not a one‑liner. In production the orchestration layer, state store, and inference engine have to be decoupled, observable, and cost‑aware. Below is a hardened reference architecture that has survived dozens of 200‑RPS workloads in a regulated fintech environment. It shows why a hand‑rolled container on a VM quickly breaks and how Azure AI Foundry + AKS can keep you in business.
Monolith Bottlenecks in Agent Workloads
When you prototype a LLM‑driven agent, you often start with a single container that calls the model once. That works for a few users, but once you add:
- Multi‑step planning and tool‑calling
- Per‑tenant state that must survive restarts
- Dynamic token budgets that can blow past the 4 k request limit
- Fine‑grained RBAC that spans the entire stack
the naive “run‑the‑model‑in‑a‑container” approach explodes in latency, cost, and security. The core problem is that the inference engine becomes a monolith that couples the request lifecycle to GPU usage, making it impossible to scale the dispatcher independently.
Real‑World Example
Our team built a financial‑advisor bot that pulls real‑time market data, runs a 15‑step decision chain, and returns a recommendation. The request pipeline looked like this:
Client → Front Door → APIM → Dispatcher (AKS) → Service Bus → Foundry (GPU) → Cosmos DB → Event Grid → APIM → Client
At peak load (≈250 RPS) we hit the following production pain points:
- APIM throttled 30 % of requests due to the 4 k payload limit.
- Service Bus dead‑letter queue grew to 1 k messages in 2 hours after an OpenAI 429 burst.
- Cosmos DB RU/s consumption spiked to 15 k RU/s, pushing the account into the next pricing tier.
- Each GPU node ran at 95 % CPU for 18 h/day, but the dispatcher pods were idle for 2 h/day.
Trade‑Offs
| Aspect | Foundry + AKS | DIY Inference + K8s | Serverless (Functions Premium) |
|---|---|---|---|
| Operational Overhead | Low – managed GPU lifecycle, auto‑patching | High – manual driver updates, pod restarts | Very Low – fully managed |
| Latency SLA | 200 ms–2 s (GPU warm) | 200 ms–2 s (depends on pod load) | ~100 ms (warm) |
| Token‑budget Control | Built‑in, per‑request cost attribution | Custom instrumentation needed | Manual, risk of runaway loops |
| Scaling Granularity | Pod‑level + GPU pool | Pod‑level only | Instance‑level |
| Security Surface | Private endpoint, VNet integration, IAM‑based secrets | Expose model port, rely on network policies | Private, but requires function‑level RBAC |
| Cost Predictability | Per‑hour GPU + per‑request metrics | Hard to attribute costs per user | Serverless metering, but hidden cold‑start cost |
Stack Fit for SLA, Cost, Compliance
Use the table below to decide which stack fits your SLA, cost model, and regulatory constraints:
| Requirement | Foundry + AKS | DIY Inference | Functions Premium |
|---|---|---|---|
| Latency < 500 ms | ✓ (warm GPU) | ✗ (pod spin‑up) | ✓ (warm function) |
| Token‑budget per request < 8 k | ✓ (built‑in guard) | ✗ (need custom) | ✗ (risk runaway) |
| Tenant isolation with per‑tenant secrets | ✓ (Key Vault + Managed Identity) | ✗ (harder to enforce) | ✗ (function secrets per tenant) |
| Budget < $5k/month | ✗ (GPU cost dominates) | ✓ (run on spot VMs) | ✓ (serverless) |
Token Budgets, Batching, Caching, Observability
- Token Budget: Enforce a hard ceiling (e.g., 12 steps or 8 k tokens) before calling OpenAI. This keeps GPU usage predictable.
-
Batching: Group short requests into a single GPU batch when possible; Foundry supports
batch_sizevia the Compute Profile. - Cache: Cache frequently used embeddings or prompt templates in Redis; this reduces OpenAI calls by ~30 % for 80 % of traffic.
-
Observability: Use OpenTelemetry with a
tenant.idtag; set astep.durationmetric to spot slow tool calls.
Scaling Notes
- Service Bus Premium: Use
1 k messages/sthroughput tier; scalemaxDeliveryCountto 5 to avoid endless retries. - Cosmos DB: Pre‑allocate 12 k RU/s for writes; enable
sessionconsistency for read‑your‑writes guarantee. - Foundry GPU pool: Min 1, max 10 nodes; target 70 % CPU to keep headroom for unexpected bursts.
- AKS dispatcher: Auto‑scale based on
request.countmetric; keepmaxPodsPerNodeat 20 to avoid node saturation.
When This Fails in Production
-
Data Leakage Across Tenants: A tenant can inject a
systemrole message that overrides the guardrail, exposing other tenants’ data. - Runaway Token Loops: An agent that calls itself recursively without a step limit can generate > 50 k tokens, blowing the budget and the GPU queue.
- Dead‑Letter Queue Buildup: If the orchestrator completes a Service Bus message before an OpenAI error, the state is lost and the dead‑letter queue fills.
- Cold‑Start Latency Spike: When the GPU pool scales up, the first few requests suffer 1–2 s latency until the GPU warms.
- Cost Surprises: Uncontrolled token usage, high Cosmos RU/s, or exceeding the Service Bus premium tier can push costs beyond the forecast.
Common Mistakes Engineers Make
- Using a single API key for all tenants – this breaks isolation and makes key rotation painful.
- Relying on the default APIM request size – 4 k is insufficient for 15‑step chains.
- Ignoring the Service Bus
visibilityTimeout– leading to duplicate processing. - Storing per‑tenant secrets in environment variables – they get leaked in container logs.
- Not enabling
sessionconsistency – read‑your‑write guarantees fail, causing stale data to be served.
Better Approach Based on Experience
From the failures above, the following hardened pattern emerged:
- Use Managed Identities everywhere – dispatcher, orchestrator, and tool functions. Pull secrets from Key Vault on demand.
- Wrap every external call (OpenAI, SQL, REST) in a Circuit Breaker with exponential back‑off.
- Persist a
stepIdcounter in Cosmos DB and use it as the primary key for idempotent writes. - Introduce a Durable Functions orchestrator for short, deterministic flows (<5 steps). Keep Foundry for the heavy, multi‑step chains.
- Implement a global rate‑limit per tenant that throttles to 10 RPS; this protects the GPU pool and keeps costs predictable.
- Deploy a dead‑letter monitor that alerts when the DLQ depth > 50 messages; automatically trigger a re‑queue with a back‑off.
- Use Azure Policy to enforce VNet isolation for Foundry and restrict public IPs for the dispatcher.
- Run a synthetic load test every week that simulates a 200 RPS burst; capture the GPU utilization curve and adjust auto‑scale thresholds accordingly.
Production Checklist
- Enable Managed Identity on AKS and grant
KeyVault Secrets Userto the dispatcher. - Configure Cosmos DB with
sessionconsistency and amaxRUPerSecondlimit. - Set Service Bus
maxDeliveryCountto 5 and enable DLQ monitoring. - Add OpenTelemetry exporters for traces and metrics; create dashboards for
step.durationandtoken_usage. - Implement a
maxStepsguard in the orchestrator and surface the limit in the API response. - Run a load test (e.g., k6) targeting 200 RPS with a mix of 1‑step and 10‑step workloads; verify latency < 2 s and CPU < 70 % on dispatcher pods.
- Document the incident‑response run‑book: how to flush the dead‑letter queue, how to rotate the OpenAI API key, and how to scale the GPU pool.
Conclusion
Deploying agentic workloads at scale is a multi‑layer problem. The naive “one‑container” approach fails because it entangles state, policy, and inference. By decoupling the dispatcher, orchestration, and GPU compute, and by leveraging Azure AI Foundry’s managed GPU lifecycle, we can achieve predictable latency, fine‑grained RBAC, and cost control. The key trade‑offs—operational overhead vs. latency, token budget control vs. cost predictability—are clear when you map them onto real‑world metrics. Follow the checklist, avoid the common pitfalls, and iterate on the scaling knobs; you’ll end up with a production‑grade agent that can handle hundreds of concurrent users without breaking the bank.
Top comments (0)