DEV Community

Cover image for Deploying Agentic Workloads: Debugging 200‑RPS Outages in AKS
Amitesh0512
Amitesh0512

Posted on Originally published at amiteshsurwar.com

Deploying Agentic Workloads: Debugging 200‑RPS Outages in AKS

Quick Answer

Deploying Agentic Workloads: Learn how senior engineers can reliably deploy agentic AI workloads on Azure AI Foundry using Kubernetes, with concrete architecture, code, cost, and failure‑mode guidance.

Deploying Agentic Workloads at Scale: A Practical Blueprint

Deploying Agentic Workloads is not a one‑liner. In production the orchestration layer, state store, and inference engine have to be decoupled, observable, and cost‑aware. Below is a hardened reference architecture that has survived dozens of 200‑RPS workloads in a regulated fintech environment. It shows why a hand‑rolled container on a VM quickly breaks and how Azure AI Foundry + AKS can keep you in business.

Monolith Bottlenecks in Agent Workloads

When you prototype a LLM‑driven agent, you often start with a single container that calls the model once. That works for a few users, but once you add:

  • Multi‑step planning and tool‑calling
  • Per‑tenant state that must survive restarts
  • Dynamic token budgets that can blow past the 4 k request limit
  • Fine‑grained RBAC that spans the entire stack

the naive “run‑the‑model‑in‑a‑container” approach explodes in latency, cost, and security. The core problem is that the inference engine becomes a monolith that couples the request lifecycle to GPU usage, making it impossible to scale the dispatcher independently.

Real‑World Example

Our team built a financial‑advisor bot that pulls real‑time market data, runs a 15‑step decision chain, and returns a recommendation. The request pipeline looked like this:

Client → Front Door → APIM → Dispatcher (AKS) → Service Bus → Foundry (GPU) → Cosmos DB → Event Grid → APIM → Client
Enter fullscreen mode Exit fullscreen mode

At peak load (≈250 RPS) we hit the following production pain points:

  • APIM throttled 30 % of requests due to the 4 k payload limit.
  • Service Bus dead‑letter queue grew to 1 k messages in 2 hours after an OpenAI 429 burst.
  • Cosmos DB RU/s consumption spiked to 15 k RU/s, pushing the account into the next pricing tier.
  • Each GPU node ran at 95 % CPU for 18 h/day, but the dispatcher pods were idle for 2 h/day.

Trade‑Offs

Aspect Foundry + AKS DIY Inference + K8s Serverless (Functions Premium)
Operational Overhead Low – managed GPU lifecycle, auto‑patching High – manual driver updates, pod restarts Very Low – fully managed
Latency SLA 200 ms–2 s (GPU warm) 200 ms–2 s (depends on pod load) ~100 ms (warm)
Token‑budget Control Built‑in, per‑request cost attribution Custom instrumentation needed Manual, risk of runaway loops
Scaling Granularity Pod‑level + GPU pool Pod‑level only Instance‑level
Security Surface Private endpoint, VNet integration, IAM‑based secrets Expose model port, rely on network policies Private, but requires function‑level RBAC
Cost Predictability Per‑hour GPU + per‑request metrics Hard to attribute costs per user Serverless metering, but hidden cold‑start cost

Stack Fit for SLA, Cost, Compliance

Use the table below to decide which stack fits your SLA, cost model, and regulatory constraints:

Requirement Foundry + AKS DIY Inference Functions Premium
Latency < 500 ms ✓ (warm GPU) ✗ (pod spin‑up) ✓ (warm function)
Token‑budget per request < 8 k ✓ (built‑in guard) ✗ (need custom) ✗ (risk runaway)
Tenant isolation with per‑tenant secrets ✓ (Key Vault + Managed Identity) ✗ (harder to enforce) ✗ (function secrets per tenant)
Budget < $5k/month ✗ (GPU cost dominates) ✓ (run on spot VMs) ✓ (serverless)

Token Budgets, Batching, Caching, Observability

  • Token Budget: Enforce a hard ceiling (e.g., 12 steps or 8 k tokens) before calling OpenAI. This keeps GPU usage predictable.
  • Batching: Group short requests into a single GPU batch when possible; Foundry supports batch_size via the Compute Profile.
  • Cache: Cache frequently used embeddings or prompt templates in Redis; this reduces OpenAI calls by ~30 % for 80 % of traffic.
  • Observability: Use OpenTelemetry with a tenant.id tag; set a step.duration metric to spot slow tool calls.

Scaling Notes

  • Service Bus Premium: Use 1 k messages/s throughput tier; scale maxDeliveryCount to 5 to avoid endless retries.
  • Cosmos DB: Pre‑allocate 12 k RU/s for writes; enable session consistency for read‑your‑writes guarantee.
  • Foundry GPU pool: Min 1, max 10 nodes; target 70 % CPU to keep headroom for unexpected bursts.
  • AKS dispatcher: Auto‑scale based on request.count metric; keep maxPodsPerNode at 20 to avoid node saturation.

When This Fails in Production

  1. Data Leakage Across Tenants: A tenant can inject a system role message that overrides the guardrail, exposing other tenants’ data.
  2. Runaway Token Loops: An agent that calls itself recursively without a step limit can generate > 50 k tokens, blowing the budget and the GPU queue.
  3. Dead‑Letter Queue Buildup: If the orchestrator completes a Service Bus message before an OpenAI error, the state is lost and the dead‑letter queue fills.
  4. Cold‑Start Latency Spike: When the GPU pool scales up, the first few requests suffer 1–2 s latency until the GPU warms.
  5. Cost Surprises: Uncontrolled token usage, high Cosmos RU/s, or exceeding the Service Bus premium tier can push costs beyond the forecast.

Common Mistakes Engineers Make

  • Using a single API key for all tenants – this breaks isolation and makes key rotation painful.
  • Relying on the default APIM request size – 4 k is insufficient for 15‑step chains.
  • Ignoring the Service Bus visibilityTimeout – leading to duplicate processing.
  • Storing per‑tenant secrets in environment variables – they get leaked in container logs.
  • Not enabling session consistency – read‑your‑write guarantees fail, causing stale data to be served.

Better Approach Based on Experience

From the failures above, the following hardened pattern emerged:

  • Use Managed Identities everywhere – dispatcher, orchestrator, and tool functions. Pull secrets from Key Vault on demand.
  • Wrap every external call (OpenAI, SQL, REST) in a Circuit Breaker with exponential back‑off.
  • Persist a stepId counter in Cosmos DB and use it as the primary key for idempotent writes.
  • Introduce a Durable Functions orchestrator for short, deterministic flows (<5 steps). Keep Foundry for the heavy, multi‑step chains.
  • Implement a global rate‑limit per tenant that throttles to 10 RPS; this protects the GPU pool and keeps costs predictable.
  • Deploy a dead‑letter monitor that alerts when the DLQ depth > 50 messages; automatically trigger a re‑queue with a back‑off.
  • Use Azure Policy to enforce VNet isolation for Foundry and restrict public IPs for the dispatcher.
  • Run a synthetic load test every week that simulates a 200 RPS burst; capture the GPU utilization curve and adjust auto‑scale thresholds accordingly.

Production Checklist

  1. Enable Managed Identity on AKS and grant KeyVault Secrets User to the dispatcher.
  2. Configure Cosmos DB with session consistency and a maxRUPerSecond limit.
  3. Set Service Bus maxDeliveryCount to 5 and enable DLQ monitoring.
  4. Add OpenTelemetry exporters for traces and metrics; create dashboards for step.duration and token_usage.
  5. Implement a maxSteps guard in the orchestrator and surface the limit in the API response.
  6. Run a load test (e.g., k6) targeting 200 RPS with a mix of 1‑step and 10‑step workloads; verify latency < 2 s and CPU < 70 % on dispatcher pods.
  7. Document the incident‑response run‑book: how to flush the dead‑letter queue, how to rotate the OpenAI API key, and how to scale the GPU pool.

Conclusion

Deploying agentic workloads at scale is a multi‑layer problem. The naive “one‑container” approach fails because it entangles state, policy, and inference. By decoupling the dispatcher, orchestration, and GPU compute, and by leveraging Azure AI Foundry’s managed GPU lifecycle, we can achieve predictable latency, fine‑grained RBAC, and cost control. The key trade‑offs—operational overhead vs. latency, token budget control vs. cost predictability—are clear when you map them onto real‑world metrics. Follow the checklist, avoid the common pitfalls, and iterate on the scaling knobs; you’ll end up with a production‑grade agent that can handle hundreds of concurrent users without breaking the bank.

Top comments (0)