Provisioned Throughput in Microsoft Foundry: The Capacity Math Nobody Does Until the Bill Arrives
Most teams discover Provisioned Throughput Units (PTUs) the hard way: either a 429 storm during a product launch on a standard deployment, or a five-figure invoice after someone "just provisioned a bit extra to be safe." Both outcomes trace back to the same root cause — PTUs are not a pricing tier you toggle on, they're a capacity-engineering decision, and Microsoft Foundry expects you to do the arithmetic yourself before you click deploy.
This article is that arithmetic. We'll go under the hood of what a PTU actually represents, how Foundry's sizing model converts your traffic shape into a PTU count, why quota and capacity are two entirely different constraints that both have to clear before a deployment succeeds, how spillover changes your reliability posture, and where the reservation math flips in your favor — or doesn't.
Table of Contents
- Why This Matters
- Deployment Types
- What a PTU Actually Is
- Quota vs. Capacity
- The Sizing Formula, Derived
- Worked Example With Code
- Spillover
- Hourly Billing vs. Reservations
- Production Scenario
- Common Mistakes
- When Not to Use Provisioned Throughput
- Practical Recommendations
- Conclusion
- References
Why This Matters
Every LLM-backed production system eventually hits the same wall: standard (pay-as-you-go) deployments share inference capacity across every tenant hitting that regional pool. That's fine for prototyping. It is not fine for a checkout assistant, a fraud-review copilot, or any workload where a customer-facing SLA depends on a model responding inside a tight latency budget at 2 p.m. on Black Friday. Standard deployments give you no throughput guarantee — only best-effort service with token-based rate limits (TPM/RPM) that can tighten under regional load.
Provisioned Throughput exists to solve exactly this problem: dedicated, isolated model capacity that is yours whether you use it or not, with a defined latency SLA per model. The trade-off is that dedicated capacity is billed by the hour regardless of utilization, which means an undersized deployment throttles your users and an oversized one burns budget for idle silicon. Getting the PTU count right is the whole game, and Foundry gives you the formulas and a calculator to do it — but almost nobody reads past the pricing page to the sizing methodology, and that's where the expensive mistakes happen.
Deployment Types
Foundry Models expose four deployment shapes, and each is a genuinely different point in the latency/cost/predictability space, not a marketing tier:
| Deployment type | Billing | Latency SLA | Best for |
|---|---|---|---|
| Standard | Pay per token | None | Dev/test, unpredictable or low-volume production traffic |
| Priority processing | Pay per token (priority rate) | Defined latency target, no commitment | Latency-sensitive production without long-term commitment |
| Provisioned | Per PTU per hour (or reservation) | Defined latency target per model | Mission-critical, high-scale, guaranteed-throughput workloads |
| Batch | Discounted per-token, async | None | Bulk, non-interactive processing (embeddings backfills, offline eval runs) |
The decision isn't "which is cheapest per token" — standard deployments will usually win that comparison at low volume. The decision is "what does an unpredictable multi-second latency spike cost my product," and provisioned throughput is the answer when that cost is unacceptable.
What a PTU Actually Is
A Provisioned Throughput Unit is a fixed slice of model-serving compute that Foundry reserves exclusively for your deployment the moment you create it — it sits there, staffed and warm, whether or not a single request arrives. Four properties define how PTUs behave, and all four matter when you're capacity planning:
- Model-independent purchase, model-dependent yield. You don't buy "GPT-4.1 PTUs" — you buy PTUs, and then point them at any supported model. But the tokens per minute a given PTU count delivers is entirely model-specific. A heavier model (larger context, more parameters engaged per token) needs more PTUs to hit the same TPM as a lighter one. This is the single most common source of bad sizing: reusing last quarter's PTU count for a newly upgraded model version without re-running the math.
- Region- and pool-scoped quota. PTU quota is granted per subscription, per Azure region, and per deployment type (Global / Data Zone / Regional are separate pools). Quota approved in East US is worthless in West Europe. Nothing carries over.
- Model-specific minimums and scale increments. Every model has a minimum deployable PTU count and a scale increment (e.g., "round up to the nearest 5 or 50 PTUs"), so your calculated number almost never matches your purchased number — you always round up.
- Fixed cost regardless of traffic. The billing meter runs from the moment the deployment exists until you delete it. A provisioned deployment sitting at 5% utilization overnight costs exactly the same per hour as one at 95% utilization. This is the property that makes over-provisioning expensive and under-provisioning throttling, with almost no comfortable middle ground unless your sizing is accurate.
[IMAGE: Diagram showing the Microsoft Foundry provisioned throughput request flow — a client application sending requests to a Foundry resource endpoint, which routes to a fixed-capacity PTU pool (shown as a gauge), with three parallel deployment scope options (Global Provisioned, Data Zone Provisioned, Regional Provisioned) feeding into it, and a dashed overflow arrow labeled "Spillover (429)" routing excess traffic to a standard pay-as-you-go deployment.]
Quota vs. Capacity
This is the distinction that trips up almost every team provisioning PTUs for the first time, because the two failure modes look identical from the outside (deployment creation fails) but require completely different remediation.
Quota is a policy ceiling. It's the maximum number of PTUs Azure's control plane will let your subscription request in a given region and deployment-type pool. It costs nothing to hold, and a default allotment is granted automatically to eligible subscriptions. If you need more, you file the quota request form and wait — sometimes days — for approval.
Capacity is a physical reality. It's the actual silicon available in that region, for that model, at that moment, across every customer competing for it. Capacity is allocated at deployment creation and held for the deployment's entire lifetime — but it is not reserved for you in advance just because you have quota. Two engineering consequences follow directly from this:
- Having quota does not guarantee you can deploy. You can hold 500 PTUs of approved Data Zone quota and still have a deployment fail because the region's capacity pool for that specific model is currently exhausted by other tenants' demand.
- Deleting a deployment releases capacity back to the shared pool, permanently, from your perspective. If you scale a provisioned deployment down (or delete it to save cost overnight, which is exactly the anti-pattern the reservation model discourages), there is no guarantee the same capacity is available when you try to scale back up an hour later. Capacity availability fluctuates through the day based on aggregate demand you have no visibility into.
Practically: before you commit to an architecture that scales provisioned deployments up and down with traffic (the instinct every Kubernetes-minded engineer has), check capacity via the Foundry portal's deployment experience or the model capacities REST API — and understand that the "elastic PTU" pattern that works beautifully for compute (VMs, containers) does not translate cleanly to a resource pool this contested.
The Sizing Formula, Derived
Foundry's sizing methodology reduces three independent variables — your request rate, your prompt/response shape, and your cache hit rate — into one number: normalized TPM, which you then divide by the model's throughput-per-PTU constant.
Three inputs come from your traffic profile:
- Peak RPM — requests per minute at your busiest sustained period, not your average. Sizing to average RPM guarantees throttling at peak.
- Average prompt size and average response size, in tokens.
- Cache rate — the fraction of input tokens served from Foundry's prompt cache. Cached tokens are excluded entirely from PTU consumption, so a well-designed prompt-caching strategy (stable system prompts, repeated few-shot blocks, consistent tool schemas at the front of the context window) directly reduces your PTU bill.
One model-specific constant matters here: the output-to-input ratio. Generating an output token costs meaningfully more compute than processing an input token (the model has to run a full forward pass per generated token versus batched parallel processing for prompt tokens), so Foundry weights output tokens heavier — for GPT-4.1-class models and later, this ratio matches the model's standard pricing ratio between output and input token cost, which is a convenient mental shortcut: if output tokens are priced 4x input tokens on standard billing, they'll consume roughly 4x the PTU capacity too.
The formula:
Input TPM = Peak RPM × avg prompt tokens
Output TPM = Peak RPM × avg response tokens
Normalized TPM = (Input TPM × (1 − cache rate)) + (output_to_input_ratio × Output TPM)
PTUs required = Normalized TPM ÷ Input_TPM_per_PTU [then round up to the model's scale increment]
Input_TPM_per_PTU and output_to_input_ratio are both published per model in Foundry's deployment parameters tables — they are not universal constants, and they change (usually downward, i.e. better) as Microsoft optimizes serving infrastructure for a given model version, which is why re-checking sizing after a model upgrade isn't optional.
Worked Example With Code
Here's a small, honest Python helper that encodes the formula so you can run "what-if" sizing before you ever open the Foundry portal calculator. This is a simplified planning tool, not a production billing calculator — treat its output as a starting estimate to validate against real benchmark traffic, per Microsoft's own guidance.
import math
def estimate_ptus(
peak_rpm: int,
avg_prompt_tokens: int,
avg_response_tokens: int,
input_tpm_per_ptu: int,
output_to_input_ratio: float,
cache_rate: float = 0.0,
scale_increment: int = 5,
) -> dict:
"""
Estimate required PTUs for a Foundry provisioned deployment.
Args:
peak_rpm: expected peak requests per minute (size to PEAK, not average)
avg_prompt_tokens: average input tokens per request
avg_response_tokens: average output tokens per request
input_tpm_per_ptu: model-specific constant from Foundry docs
output_to_input_ratio: model-specific output cost weighting
cache_rate: fraction (0.0-1.0) of input tokens served from prompt cache
scale_increment: model's minimum deployment scale step (e.g. 5, 50)
Returns:
dict with intermediate and final sizing values
"""
if not (0.0 <= cache_rate <= 1.0):
raise ValueError("cache_rate must be between 0.0 and 1.0")
input_tpm = peak_rpm * avg_prompt_tokens
output_tpm = peak_rpm * avg_response_tokens
normalized_tpm = (input_tpm * (1 - cache_rate)) + (output_to_input_ratio * output_tpm)
raw_ptus = normalized_tpm / input_tpm_per_ptu
# Always round UP to the nearest scale increment — Foundry will not let
# you deploy a partial increment, and under-rounding silently throttles you.
rounded_ptus = math.ceil(raw_ptus / scale_increment) * scale_increment
return {
"input_tpm": input_tpm,
"output_tpm": output_tpm,
"normalized_tpm": normalized_tpm,
"raw_ptus": round(raw_ptus, 2),
"rounded_ptus": rounded_ptus,
}
# Example: gpt-5.2 on Data Zone Provisioned, 1,000 peak RPM,
# 200-token prompts, 20-token responses, no caching yet.
result_no_cache = estimate_ptus(
peak_rpm=1000,
avg_prompt_tokens=200,
avg_response_tokens=20,
input_tpm_per_ptu=3400, # published per-model constant
output_to_input_ratio=8, # published per-model constant
cache_rate=0.0,
)
print("Without caching:", result_no_cache)
# -> {'input_tpm': 200000, 'output_tpm': 20000, 'normalized_tpm': 360000,
# 'raw_ptus': 105.88, 'rounded_ptus': 110}
# Same workload, but 50% of input tokens now hit the prompt cache
# (e.g., a stable system prompt + tool schema shared across requests)
result_with_cache = estimate_ptus(
peak_rpm=1000,
avg_prompt_tokens=200,
avg_response_tokens=20,
input_tpm_per_ptu=3400,
output_to_input_ratio=8,
cache_rate=0.5,
)
print("With 50% cache hit rate:", result_with_cache)
# -> {'input_tpm': 200000, 'output_tpm': 20000, 'normalized_tpm': 260000,
# 'raw_ptus': 76.47, 'rounded_ptus': 80}
That 30-PTU delta between the cached and uncached run (110 vs. 80) is not a rounding artifact — at typical Data Zone Provisioned rates, that's real monthly cost, entirely earned back by restructuring your prompts so the static portions sit at the front of the context window where the cache can hit them. This is the single highest-leverage, lowest-effort optimization available to teams running provisioned deployments, and it's almost always left on the table because prompt-caching discipline is treated as a "nice to have" rather than a capacity-planning lever.
[IMAGE: Infographic-style diagram showing the PTU sizing formula as a left-to-right pipeline — peak RPM multiplied by average prompt tokens equals input TPM, peak RPM multiplied by average response tokens equals output TPM, both feeding into a normalized TPM calculation that accounts for cache rate and output-to-input ratio, then divided by input TPM per PTU to yield required PTUs, then rounded up to the deployment's scale increment.]
Spillover
Spillover is Foundry's answer to the obvious question: what happens when your sized-for-peak deployment meets a traffic spike above peak? Without spillover, Foundry returns 429 once the PTU pool is saturated, and your application has to handle that itself (retry-with-backoff, queue, degrade gracefully — whatever your resilience layer does).
With spillover configured, those same overflow requests are automatically redirected to a standard (pay-as-you-go) deployment in the same Foundry resource, instead of failing. You can enable it resource-wide, or control it per request using the x-ms-spillover-deployment header — which is the more interesting pattern operationally, because it lets you make the spillover decision per request-class rather than globally. A synchronous, user-facing chat completion might be worth spilling over to standard (accept variable latency rather than a hard failure); a background batch-style call inside the same application might be better served by queuing and retrying against the PTU deployment once capacity frees up, rather than silently degrading to shared capacity.
One constraint worth flagging clearly: spillover currently only works for Azure OpenAI models in Foundry. Third-party Foundry Models (Meta Llama, Azure DeepSeek, and others) don't support it as of this writing — so if your provisioned deployment is running a non-Azure-OpenAI model, you need your own overflow/backpressure strategy at the application layer, because Foundry won't do it for you.
import httpx
def call_with_explicit_spillover(client: httpx.Client, endpoint: str, payload: dict, standard_deployment: str):
"""
Force this specific call to spill over to a named standard deployment
if the primary provisioned deployment is saturated, rather than relying
on resource-wide spillover configuration.
"""
headers = {
"Content-Type": "application/json",
"x-ms-spillover-deployment": standard_deployment,
}
response = client.post(endpoint, json=payload, headers=headers)
response.raise_for_status()
return response.json()
Hourly Billing vs. Reservations
Provisioned deployments support two billing modes, and conflating them is how teams end up with a bill that doesn't match their mental model.
Hourly billing meters PTUs at a flat $/PTU/hour rate from deployment creation to deletion, independent of tokens consumed. It's the right mode for short, bounded experiments: benchmarking a candidate model before committing, or scaling up temporarily for a known event. It is explicitly not meant to be a scale-to-zero mechanism you toggle daily to save money, for the two reasons already covered under capacity: your capacity isn't guaranteed to still be there when you scale back up, and sustained high-utilization hourly billing typically costs more than the equivalent reservation.
Azure Reservations are a financial commitment layered on top of the same PTU meter — a 1-month or 1-year commitment in exchange for a discounted effective hourly rate. The subtlety worth internalizing: reservations and deployments are loosely coupled. Buying a reservation doesn't create or reserve a deployment, and creating a deployment doesn't require a reservation. The correct order of operations is:
- Create the provisioned deployment first and confirm the capacity you need is actually available in your target region.
- Only then purchase the reservation to lock in the discounted rate against that ongoing usage.
Buying a reservation before confirming capacity availability is a common and entirely avoidable mistake — you can be sitting on a paid-for financial commitment with no deployment able to consume it if the region runs out of capacity for your model in the interim.
Production Scenario
Consider a mid-size SaaS company shipping an internal support-ticket triage copilot to 400 support agents. Requirements: sub-second p95 latency (agents are watching a live queue), predictable monthly cost for finance, and strict EU data residency because tickets contain customer PII from European accounts.
Walking through the decision tree:
- Data residency rules out Global Provisioned immediately — routing across regions globally is incompatible with the EU-only constraint. Data Zone Provisioned (EU) is the correct scope: it stays within the geographic zone while still getting better availability than pinning to one specific region.
- Sizing: agents peak at roughly 300 triage requests per minute during EU morning hours, each with a ~600-token prompt (ticket text + retrieved knowledge-base snippets) and a ~150-token structured response. The knowledge-base retrieval instructions and output schema are static and sit at the front of the prompt, giving a realistic ~35% cache rate once prompt caching is tuned.
- Reliability: rare backlog spikes (a P1 incident generating a flood of tickets) are handled via per-request spillover to a standard deployment scoped explicitly for the triage service's own summarization calls — not silently applied resource-wide, so a spike in one workload doesn't quietly degrade the latency guarantee for a different, more latency-sensitive workload sharing the resource.
- Cost: because this is a permanent, 24/7 production workload rather than a bursty experiment, the team validates capacity with a deployment first, confirms sustained utilization over two weeks, and then buys a 1-year reservation against that Data Zone Provisioned usage to lock in the discounted rate — using Cost Management's reservation utilization view to confirm they're not paying for meaningfully idle capacity.
This is the pattern worth generalizing: residency constraints pick the deployment scope, traffic shape and cache design pick the PTU count, and workload criticality/duration picks the billing mode — three separate decisions, frequently collapsed into one guess.
Common Mistakes
- Sizing to average RPM instead of peak RPM. Averages hide the exact moments your SLA matters most.
-
Re-using a PTU count across model version upgrades.
Input TPM per PTUand the output-to-input ratio are model-specific and change between versions — a "like for like" model swap can silently under-provision you. - Treating provisioned deployments as elastic infrastructure. Scaling down nightly to save cost, expecting to scale back up seamlessly, ignores that capacity is a shared, fluctuating pool with no guarantee of return.
- Buying a reservation before confirming capacity. Reservations don't create capacity — they discount a meter that only runs if a deployment can actually be created.
- Ignoring prompt-cache design as a cost lever. Cache rate is a first-class term in the sizing formula, not an afterthought; restructuring prompts to front-load static content is often cheaper than buying more PTUs.
- Applying spillover resource-wide by default. This can leak latency-sensitive traffic's overflow policy onto unrelated workloads sharing the same Foundry resource; use the per-request header when workloads have different risk tolerances.
- Forgetting third-party model spillover limitations. If you're running Llama or DeepSeek models on provisioned throughput, don't assume the same 429-to-standard safety net exists — verify it explicitly for your model provider.
When Not to Use Provisioned Throughput
Skip PTUs when your traffic is genuinely unpredictable or low-volume — development, testing, early-stage products still finding product-market fit, or internal tools used sporadically. The hourly billing floor means an idle or lightly-used provisioned deployment is nearly always more expensive than standard pay-as-you-go for the same workload. Also reconsider if you can't yet produce a reasonably confident peak-RPM estimate; sizing formulas amplify bad inputs, and an under-informed PTU purchase just moves your risk from "unpredictable latency" to "unpredictable cost with the same unpredictable latency once you exceed your guess." Priority processing is frequently the better middle ground here — a defined latency target without a capacity commitment.
Practical Recommendations
- Run the sizing formula against peak traffic windows, not steady-state averages, and re-run it after every model version change.
- Treat prompt-cache rate as a design target during development, not a bonus you discover in production — structure prompts with static content first.
- Validate capacity availability with a real deployment before purchasing any reservation.
- Use per-request spillover headers to scope overflow behavior to the workloads that should actually tolerate variable latency.
- Monitor reservation utilization continuously via Cost Management; a reservation bought for peak sizing but running well under utilization most of the time is a signal to right-size, not a sunk cost to ignore.
- For multi-model or multi-workload Foundry resources, size and reserve independently per workload rather than pooling estimates — the sizing constants are model-specific enough that averaging across models produces meaningless numbers.
Conclusion
Provisioned Throughput is Foundry's mechanism for buying certainty — guaranteed latency and guaranteed capacity — in exchange for giving up the elasticity that makes pay-as-you-go deployments forgiving of bad estimates. That trade only pays off when the estimate behind it is good, and a good estimate requires taking peak traffic shape, model-specific throughput constants, and prompt-cache design seriously as inputs, not decorations, to the sizing formula. Do the arithmetic before you deploy, validate capacity before you reserve, and treat spillover as a scoped safety net rather than a blanket assumption — and PTUs become a predictable cost center instead of the place production incidents and surprise invoices both originate.
References
- Provisioned throughput for Foundry Models — Microsoft Learn
- Determine PTU sizing for a workload — Microsoft Learn
- Manage traffic with spillover for provisioned deployments — Microsoft Learn
- Deployment types for Microsoft Foundry Models — Microsoft Learn
- Microsoft Cost Management documentation for Azure Reservations utilization and chargeback (verify latest links before publishing)
This is Day 17 of the Microsoft Foundry 100 Days / 100 Blogs series — a daily deep dive into the architecture, APIs, and production realities of building on Microsoft Foundry.
Top comments (0)