Quick Answer
Learn how to architect, secure, and scale Deploying Multi‑Agent RAG workflows with Semantic Kernel on Azure AI Foundry—complete with code, observability, and production‑ready checklists.
Token Budget Limits in Multi‑Agent RAG
Multi‑agent RAG is attractive because it lets you split a complex request into micro‑tasks—retrieval, summarisation, domain‑specific reasoning, and final answer generation—each handled by a specialised LLM call. In a lab this works fine, but when you start routing hundreds of requests per minute across a shared vector store and a fleet of autonomous agents, you hit three hard limits:
- Token budget exhaustion – every agent adds its own prompt, system messages, and function‑call overhead to the same quota.
- State leakage – a tenant’s vector search results can bleed into another tenant’s workflow if the filter is mis‑wired.
- Latency tail – a single slow agent can drag the whole orchestration past the SLA, and there’s no deterministic way to recover.
These issues surface only under load; a handful of test requests will never trigger the 429 or 500 errors that kill your service during a traffic spike.
Real‑World Example
At a mid‑size fintech, we built a “Regulatory Compliance Bot” that answers questions from legal teams. The bot used five agents:
- Document‑retrieval agent – pulls relevant clauses from a 200‑GB Azure Cognitive Search index.
- Summarisation agent – condenses 5‑page excerpts into 200‑token briefs.
- Context‑fusion agent – stitches summaries with the user query.
- LLM reasoning agent – runs a 32‑k token prompt through GPT‑4o.
- Response‑formatter agent – turns the raw LLM output into a compliance‑grade JSON payload.
Under peak load (≈10 k requests/min) the bot hit 4‑second tail latency in 18% of requests, and the Azure AI Foundry API returned 400‑Bad‑Request errors 3% of the time due to token budget overruns. The root cause was an un‑guarded DAG where each agent added its own prompt without a shared budget and the vector store was shared across all tenants without a strict tenantId filter.
Trade‑offs
| Approach | Pros | Cons |
|---|---|---|
| Raw Azure OpenAI SDK + hand‑rolled orchestrator | Full control over prompt shape, minimal abstraction overhead, can push experimental features straight to production. | Duplication of prompt logic across agents, risk of inconsistent token budgeting, no built‑in retry or circuit‑breaker plumbing, higher maintenance cost. |
| LangChain.NET + custom middleware | Rich ecosystem of tools, easy plug‑in of third‑party memory stores, community support. | Auth layers are extra wrappers, less seamless Azure AI Foundry integration, extra latency from wrapper layers. |
| Semantic Kernel + Azure AI Foundry | Unified prompt templating, built‑in semantic memory, first‑class Azure auth, auto‑retries via Kernel, easy model switching via MCP. | Learning curve, fewer community plugins, some advanced features still in preview. |
Stack Choice by Operational Constraints
Pick the stack that aligns with your operational constraints:
- Control & experimentation – if you need to push the newest Azure OpenAI SDK features or fine‑tune prompt engineering at a low level, go raw SDK + orchestrator.
- Rapid prototyping & third‑party tooling – if you want to mix in vector stores from Pinecone or use a pre‑built chain library, LangChain.NET is the way.
- Enterprise‑grade, Azure‑centric – for multi‑tenant services that need tight auth, token budgeting, and observability out of the box, Semantic Kernel + Azure AI Foundry wins.
In our compliance bot, we switched to Semantic Kernel after the first month of production. The built‑in SemanticMemory reduced redundant vector searches by 70%, cutting retrieval latency from 250 ms to 75 ms on average.
Latency, Token Budgeting, and Warm‑Starts
-
Latency per agent – aim for <200 ms on retrieval, <400 ms on LLM calls. Use
KernelBuilder.AddOpenTelemetry()to surface per‑step spans and identify bottlenecks. -
Token budgeting – enforce a global budget across the DAG with
Kernel.SetTokenBudget(); a 2 500‑token cap keeps you under the 400 MB per‑minute quota for GPT‑4o. -
Batching & pipelining – batch vector queries (up to 10 per request) and reuse HTTP connections via
HttpClientFactoryto shave 30 ms per call. -
Cold starts – keep the
Kernelinstance warm in a durable function host; a warm start is ~50 ms vs 300 ms cold.
Scaling Notes
When you scale to 10 k RPS:
- Deploy the orchestrator as an Azure Function with
functionAppScaleLimitset to 200 instances. - Use Azure Service Bus Premium tier for request queuing; it guarantees 10 k messages per second with
MaxDeliveryCount=10to avoid starvation. - Leverage Azure AI Foundry’s model‑specific scaling controls: set
maxConcurrencyper model to avoid throttling. - Persist intermediate results in Azure Table Storage every 5 steps to keep Durable Function state below 10 MB.
When This Fails in Production
-
Token budget overrun on hot paths – a sudden surge in user query length pushes the DAG over 2 500 tokens. Result: 400‑Bad‑Request. Fix: Validate input size before queuing; insert a
SummariseFirstagent that truncates the prompt. -
Cross‑tenant vector leakage – a mis‑configured
tenantIdfilter causes tenant B’s clauses to appear in tenant A’s results. Fix: Add a mandatorytenantIdfilter to every search and audit ingestion pipelines. - Durable Functions state explosion – orchestration history exceeds 10 MB after 30 min of idle processing, causing the function to terminate. Fix: Periodically checkpoint state to Azure Blob Storage and resume from the checkpoint.
Common Mistakes Engineers Make
- Assuming each agent’s
MaxTokenssetting is isolated – it isn’t; all prompts share the same budget. - Ignoring the cost of function calls – Azure OpenAI counts function‑call tokens separately, leading to hidden spend.
- Using a shared vector index without a strict tenant filter – data leakage is the most common security breach in multi‑tenant RAG services.
- Not instrumenting the orchestrator – without OpenTelemetry you can’t correlate a 2‑second tail to a specific agent.
Better Approach Based on Experience
In production, I would:
- Wrap the entire DAG in a
Kernelinstance that carries aTokenBudgetand aTenantContext. - Use
SemanticMemorywith a per‑tenant Azure Cognitive Search index; this eliminates duplicate retrievals and guarantees isolation. - Implement a
RetryPolicy(Polly) at the agent level and aCircuitBreakerat the orchestrator level. - Expose a
/healthzendpoint that aggregates per‑agent latency and token usage, so ops can see SLA drift in real time. - Run a k6 load test with 5 k concurrent users and a 10 k RPS burst to validate the 99th percentile latency stays below 2 s.
- Set Azure Cost Management alerts at 80% of the monthly token budget to catch runaway spend early.
Actionable Checklist for Your Next Rollout
- Define a
tenantIdfield in every document and enforce strict filtering in Azure Cognitive Search. - Instantiate a
Kernelper orchestration withKernel.SetTokenBudget(2500)andKernel.SetTenantId(tenantId). - Add
PromptGuardrailsthat whitelist allowed characters and log any injection attempts. - Configure Polly
RetryPolicywith exponential back‑off for all LLM calls. - Enable OpenTelemetry and export to Azure Monitor; build a dashboard that shows per‑agent latency, token usage, and error rates.
- Set up Azure Key Vault for tenant‑specific OpenAI keys, accessed via Managed Identity.
- Run a load test with k6 k RPS, 5 k concurrent, 5 min duration; validate 99th percentile latency < 2 s.
- Document fallback flows: when the circuit breaker opens, return a cached “service busy” message.
Conclusion
Deploying multi‑agent RAG at scale isn’t a matter of throwing more GPUs at the problem. It’s about disciplined orchestration, shared token budgeting, tenant‑aware vector storage, and observability that ties every step of the DAG to a single source of truth. By embracing Semantic Kernel’s built‑in memory, prompt templating, and Azure‑native auth, you can turn a prototype that once hiccupped on 50 RPS into a production service that consistently serves 10 k RPS with <2 s tail latency and predictable cost. The trade‑off is a steeper learning curve, but the payoff is a maintenance‑free, audit‑ready, and secure RAG platform that scales with your business, not your code.
Top comments (0)