DEV Community

Cover image for GPT-5.6 Luna on Foundry: PTU Sizing, PayGo vs. PTU + Spillover Pricing
Burak Unuvar
Burak Unuvar

Posted on AI-assisted

GPT-5.6 Luna on Foundry: PTU Sizing, PayGo vs. PTU + Spillover Pricing

A quick note before we start: While this article focuses on GPT-5.6 Luna to make the pricing and PTU calculations concrete, the same methodology applies to other models when their model-specific throughput and pricing values are substituted.

Provisioned Throughput provides a dedicated, fixed amount of processing capacity exclusively for your model deployment. Unlike Standard/PayGo, it provides a model-specific latency SLA, and its capacity is not shared across tenants.

  • PTU is a good fit for predictable, sustained traffic with consistent latency and high-throughput requirements.
  • PTU quota is model-independent, so the same quota pool can be allocated across supported models. Throughput per PTU remains model- and version-specific.
  • PTU quota is granted per subscription, region, and deployment type. Quota in East US does not carry over to West Europe, and Global Provisioned quota does not carry over to Data Zone Provisioned.

PTU Sizing and Estimation

Input Description
Model and Version The model determines which Input TPM per PTU and output-to-input ratio values to use. Each model has a minimum PTU count and specific PTU throughput.
Deployment type The provisioned deployment type: Global Provisioned, Data Zone Provisioned, or Regional Provisioned.
Peak RPM The expected peak number of calls per minute sent to the model.
Average prompt size The average number of input tokens per request.
Average response size The average number of output tokens per request.
Cache rate The percentage of input tokens served from the prompt cache. Cached tokens don't consume any PTU capacity.
FORMULAS
Input TPM = Peak RPM × Average input tokens per request
Output TPM = Peak RPM × Average output tokens per request
Effective Input TPM = Input TPM × (1 - Cache rate)
Normalized TPM = Effective Input TPM + (Output-to-input ratio × Output TPM)
Estimated PTUs = Normalized TPM / Input TPM per PTU

Note: Input TPM is the workload-specific calculated volume, whereas Input TPM per PTU is a model-specific sizing constant. For example, some listed Input TPM per PTU values are 30,000 for GPT-5.6 Luna, 3,000 for GPT-5.6 Terra, and 1,200 for GPT-5.6 Sol.


Sample Pricing Calculations for GPT-5.6 Luna on Microsoft Foundry

Representative sample:

Let's suppose your application sends requests at a peak rate of 1,000 RPM, with an average prompt size of 1,200 tokens and an average response size of 200 tokens, using the gpt-5.6-luna model with a Global Provisioned deployment.

Based on the Microsoft Foundry PTU sizing table, gpt-5.6-luna has these constants:

GPT-5.6 LUNA SIZING CONSTANT VALUE
Input TPM per PTU 30,000
Output-to-input ratio 6
Minimum Global Provisioned deployment 15 PTUs
Global Provisioned scale increment 5 PTUs

1. Without Prompt Caching

CALCULATION
Input TPM = 1,000 × 1,200 = 1,200,000
Output TPM = 1,000 × 200 = 200,000
Normalized TPM = 1,200,000 + (6 × 200,000) = 2,400,000
Estimated PTUs = 2,400,000 / 30,000 = 80 PTUs
PTUs deployed = 80 PTUs

2. With a 50% Prompt-Cache Hit Rate

CALCULATION
Input TPM = 1,000 × 1,200 = 1,200,000
Effective Input TPM = 1,200,000 × (1 - 0.50) = 600,000
Output TPM = 1,000 × 200 = 200,000
Normalized TPM = 600,000 + (6 × 200,000) = 1,800,000
Estimated PTUs = 1,800,000 / 30,000 = 60 PTUs
PTUs deployed = 60 PTUs
Peak RPM Prompt size Response size Cache rate Effective Input TPM Output TPM Normalized TPM Estimated PTUs PTUs deployed
1,000 1,200 200 0% 1,200,000 200,000 2,400,000 80.00 80
1,000 1,200 200 50% 600,000 200,000 1,800,000 60.00 60

Notes and Remarks

  • For this simplified example, each representative prompt is assumed to contain 1,200 tokens. This exceeds the 1,024-token minimum for prompt caching. A cache hit also requires at least the first 1,024 tokens to be identical across requests.
  • Prompt caching is also available for Provisioned Throughput deployments. Cached input tokens don't consume PTU capacity. In this example, caching reduces the calculated requirement from 80 PTUs to 60 PTUs—a reduction of 20 PTUs (25%). Both values are above the 15-PTU minimum and divisible by the 5-PTU scale increment.
  • No additional rounding is required in this example. Rounding would be required if an estimate fell below the 15-PTU minimum or between supported 5-PTU increments; for example, 62.4 PTUs would be rounded up to 65 PTUs.

Handling Spiky Traffic

Illustrative 24-Hour RPM Profile and 30-Day Estimate

Representative sample: Let's suppose traffic fluctuates between 0 and 2,500 RPM over a typical 24-hour period, with an average input size of 1,200 tokens and an average response size of 200 tokens, using the gpt-5.6-luna model with a Global Standard (pay-as-you-go) deployment. For the 30-day estimate, this daily traffic profile is assumed to repeat every day.

We'll also assume that 50% of input tokens are cache reads and the remaining 50% are cache misses. Of all input tokens, 10 percentage points are cache writes, leaving 40 percentage points as regular input that is neither read from nor written to the cache. Prompt caching applies only to input tokens; output tokens are always charged at the regular output-token rate.

Illustrative RPM distribution over 24 hours
(RPM changes every 4 hours)

2500 ┤                         ██████
2250 ┤                         ██████
2000 ┤                         ██████  ██████
1750 ┤                         ██████  ██████
1500 ┤                         ██████  ██████
1250 ┤                         ██████  ██████
1000 ┤                 ██████  ██████  ██████  ██████
 750 ┤                 ██████  ██████  ██████  ██████
 500 ┤         ██████  ██████  ██████  ██████  ██████
 250 ┤         ██████  ██████  ██████  ██████  ██████
   0 ┼─────────────────────────────────────────────────
HOUR │ 00–04 │ 04–08 │ 08–12 │ 12–16 │ 16–20 │ 20–24
 RPM │     0 │   500 │ 1,000 │ 2,500 │ 2,000 │ 1,000
Enter fullscreen mode Exit fullscreen mode

1. PayGo-Only Cost Calculation

For Standard/PayGo rates, as of August 27, 2026, the Azure OpenAI pricing page lists these USD rates for gpt-5.6-luna Global Standard:

Meter Symbol Price per 1M tokens
Regular input P_regular $0.20
Cached input P_cached $0.02
Cache writes P_write $0.25
Output P_output $1.20

50% cache reads + 10% cache writes + 40% regular input = 100%

Symbol Description
RPM Requests per minute
a.i.T, a.o.T Average input and output tokens per request
Hours Duration of the batch in hours
r_cached, r_write Cache-read and cache-write rates
C_batch Total input-and-output token cost for one batch
Formula Calculation
Total input tokens T_input = RPM × a.i.T × 60 × Hours
Cached input tokens T_cached = T_input × r_cached
Cache-write input tokens T_write = T_input × r_write
Regular input tokens T_regular = T_input - T_cached - T_write
Total output tokens T_output = RPM × a.o.T × 60 × Hours
Batch cost C_batch = (T_regular·P_regular + T_cached·P_cached + T_write·P_write + T_output·P_output) / 1,000,000

Each column represents one 4-hour batch on a typical day. The daily traffic profile is assumed to repeat itself for 30 days. Token values are shown in billions (B).

Metric 00–04 04–08 08–12 12–16 16–20 20–24 Total per day
Batch duration 4 hours 4 hours 4 hours 4 hours 4 hours 4 hours 24 hours
RPM 0 500 1,000 2,500 2,000 1,000
T_regular 0 0.05760B 0.11520B 0.28800B 0.23040B 0.11520B 0.80640B
T_cached 0 0.07200B 0.14400B 0.36000B 0.28800B 0.14400B 1.00800B
T_write 0 0.01440B 0.02880B 0.07200B 0.05760B 0.02880B 0.20160B
T_output 0 0.02400B 0.04800B 0.12000B 0.09600B 0.04800B 0.33600B
C_4-hour batch $0.00 $45.36 $90.72 $226.80 $181.44 $90.72 $635.04

30-day token totals:

  • Regular input (T_regular): 24.19200B
  • Cached input (T_cached): 30.24000B
  • Cache writes (T_write): 6.04800B
  • Output (T_output): 10.08000B

Estimated pay-as-you-go cost per day: $635.04.

Estimated cost per 30-day month: $19,051.20.

Estimated cost per 365-day year: $231,789.60.


2. PTU + Spillover to PayGo Cost Calculation

The following Sweden Central PTU rates were retrieved on August 27, 2026, using the Azure Retail Prices API. Sample queries and commands are included in the Appendix as a reference.

Tier Retail API rate Monthly equivalent per PTU
Hourly PTU $1.00/PTU/hour $720.00
Monthly reservation $260.00/PTU/month $260.00
One-year reservation $2,652.00/PTU/year $221.00

The hourly PTU monthly equivalent assumes 720 hours (30 days).

Warning: Hourly, non-reserved PTU is best suited to temporary or uncertain workloads, such as testing, benchmarking, capacity validation, short pilots, or migration exercises.

2.A Provisioned Baseline of 250 RPM

Let's take 250 RPM as the provisioned baseline and use Standard/PayGo spillover for bursts above it.

CALCULATION
Input TPM = 250 × 1,200 = 300,000
Effective Input TPM = 300,000 × (1 - 0.50) = 150,000
Output TPM = 250 × 200 = 50,000
Normalized TPM = 150,000 + (6 × 50,000) = 450,000
Estimated PTUs = 450,000 / 30,000 = 15 PTUs
Provisioned baseline = 15 PTUs
Time Incoming RPM Normalized TPM Demand 15-PTU Capacity (Normalized TPM) Potential Spillover Demand
00–04 0 0 450,000 0
04–08 500 900,000 450,000 450,000
08–12 1,000 1,800,000 450,000 1,350,000
12–16 2,500 4,500,000 450,000 4,050,000
16–20 2,000 3,600,000 450,000 3,150,000
20–24 1,000 1,800,000 450,000 1,350,000
Incoming traffic ──▶ 15-PTU deployment (450,000 normalized TPM capacity)
                  X
                  X if throttled (HTTP 429)
                  │
                  └──▶ Automated Spillover to Standard/PayGo deployment if configured 

Enter fullscreen mode Exit fullscreen mode

Important: Spillover is optional and must be configured either for the provisioned deployment or per request. Once configured, Microsoft Foundry automatically routes eligible requests that the provisioned deployment cannot serve—such as requests receiving HTTP 429, 500, or 503—to the associated Standard deployment. Without this configuration, the application must implement its own fallback logic.

Provisioned-Capacity Pricing for 15 PTUs

Pricing option Retail API rate 1 PTU/month 15 PTUs/month 15 PTUs/day (approx.)
Hourly PTU $1.00/PTU/hour $720.00 $10,800.00 $360.00
Monthly reservation $260.00/PTU/month $260.00 $3,900.00 $130.00
One-year reservation $2,652.00/PTU/year $221.00 $3,315.00 $108.99

Monthly cost of a yearly PTU reservation is an approximate equivalent calculated by dividing the annual price by 12.

Daily cost of a yearly PTU reservation is an approximate equivalent calculated by dividing the annual price by 365.

Daily cost of monthly PTU reservation is a rough estimation calculated by dividing the monthly price by 30.

Monthly Reserved PTU + PayGo Spillover Pricing

Metric 00–04 04–08 08–12 12–16 16–20 20–24 Daily total
Incoming RPM 0 500 1,000 2,500 2,000 1,000
Estimated spillover RPM 0 250 750 2,250 1,750 750
T_regular 0 0.02880B 0.08640B 0.25920B 0.20160B 0.08640B 0.66240B
T_cached 0 0.03600B 0.10800B 0.32400B 0.25200B 0.10800B 0.82800B
T_write 0 0.00720B 0.02160B 0.06480B 0.05040B 0.02160B 0.16560B
T_output 0 0.01200B 0.03600B 0.10800B 0.08400B 0.03600B 0.27600B
C_PTU reservation $21.67 $21.67 $21.67 $21.67 $21.67 $21.67 $130.00
C_PayGo spillover $0.00 $22.68 $68.04 $204.12 $158.76 $68.04 $521.64
C_PTU + spillover $21.67 $44.35 $89.71 $225.79 $180.43 $89.71 $651.64

Daily estimate using a monthly reservation: The $3,900 monthly PTU reservation amortizes to $130.00 per day over a 30-day month. Estimated PayGo spillover is $521.64 per day, for a combined daily estimate of $651.64.

30-day estimate using a monthly reservation: PTU reservation $3,900.00; PayGo spillover $15,649.20; combined cost $19,549.20.

365-day estimate using a one-year reservation: PTU reservation $39,780.00; PayGo spillover $190,398.60; combined cost $230,178.60.

Interval amounts are rounded independently. Daily and longer-term totals are calculated using unrounded values. Actual throughput and costs can vary with request concurrency, token-length distribution, caching behavior, model version, regional pricing, and throttling characteristics.

2.B Provisioned Baseline of 500 RPM

Let's take 500 RPM as the provisioned baseline and use Standard/PayGo spillover for bursts above it.

CALCULATION
Input TPM = 500 × 1,200 = 600,000
Effective Input TPM = 600,000 × (1 - 0.50) = 300,000
Output TPM = 500 × 200 = 100,000
Normalized TPM = 300,000 + (6 × 100,000) = 900,000
Estimated PTUs = 900,000 / 30,000 = 30 PTUs
Provisioned baseline = 30 PTUs
Time Incoming RPM Normalized TPM Demand 30-PTU Capacity (Normalized TPM) Potential Spillover Demand
00–04 0 0 900,000 0
04–08 500 900,000 900,000 0
08–12 1,000 1,800,000 900,000 900,000
12–16 2,500 4,500,000 900,000 3,600,000
16–20 2,000 3,600,000 900,000 2,700,000
20–24 1,000 1,800,000 900,000 900,000
Incoming traffic ──▶ 30-PTU deployment (900,000 normalized TPM capacity)
                  X
                  X if throttled (HTTP 429)
                  │
                  └──▶ Automated Spillover to Standard/PayGo deployment if configured 
Enter fullscreen mode Exit fullscreen mode

Important: Spillover is optional and must be configured either for the provisioned deployment or per request. Once configured, Microsoft Foundry automatically routes eligible requests that the provisioned deployment cannot serve—such as requests receiving HTTP 429, 500, or 503—to the associated Standard deployment. Without this configuration, the application must implement its own fallback logic.

Provisioned-Capacity Pricing for 30 PTUs

Pricing option Retail API rate 1 PTU/month 30 PTUs/month 30 PTUs/day (approx.)
Hourly PTU $1.00/PTU/hour $720.00 $21,600.00 $720.00
Monthly reservation $260.00/PTU/month $260.00 $7,800.00 $260.00
One-year reservation $2,652.00/PTU/year $221.00 $6,630.00 $217.97

Monthly cost of a yearly PTU reservation is an approximate equivalent calculated by dividing the annual price by 12.

Daily cost of a yearly PTU reservation is an approximate equivalent calculated by dividing the annual price by 365.

Daily cost of monthly PTU reservation is a rough estimation calculated by dividing the monthly price by 30.

Monthly Reserved PTU + PayGo Spillover Pricing

Metric 00–04 04–08 08–12 12–16 16–20 20–24 Daily total
Incoming RPM 0 500 1,000 2,500 2,000 1,000
Estimated spillover RPM 0 0 500 2,000 1,500 500
T_regular 0 0 0.05760B 0.23040B 0.17280B 0.05760B 0.51840B
T_cached 0 0 0.07200B 0.28800B 0.21600B 0.07200B 0.64800B
T_write 0 0 0.01440B 0.05760B 0.04320B 0.01440B 0.12960B
T_output 0 0 0.02400B 0.09600B 0.07200B 0.02400B 0.21600B
C_PTU reservation $43.33 $43.33 $43.33 $43.33 $43.33 $43.33 $260.00
C_PayGo spillover $0.00 $0.00 $45.36 $181.44 $136.08 $45.36 $408.24
C_PTU + spillover $43.33 $43.33 $88.69 $224.77 $179.41 $88.69 $668.24

Daily estimate using a monthly reservation: The $7,800 monthly PTU reservation amortizes to $260.00 per day over a 30-day month. Estimated PayGo spillover is $408.24 per day, producing a combined daily estimate of $668.24.

30-day estimate using a monthly reservation: PTU reservation $7,800.00; PayGo spillover $12,247.20; combined cost $20,047.20.

365-day estimate using a one-year reservation: PTU reservation $79,560.00; PayGo spillover $149,007.60; combined cost $228,567.60.

Interval amounts are rounded independently. Daily and longer-term totals are calculated using unrounded values. Actual throughput and costs can vary with request concurrency, token-length distribution, caching behavior, model version, regional pricing, and throttling characteristics.

PayGo Only vs. 15 PTUs + PayGo Spillover vs. 30 PTUs + PayGo Spillover

Period and pricing basis PayGo only 250-RPM baseline (15 PTUs) + spillover Difference vs. PayGo 500-RPM baseline (30 PTUs) + spillover Difference vs. PayGo
Daily — monthly reservation $635.04 $651.64 +$16.60 (+2.61%) $668.24 +$33.20 (+5.23%)
30-day month — monthly reservation $19,051.20 $19,549.20 +$498.00 (+2.61%) $20,047.20 +$996.00 (+5.23%)
365-day year — one-year reservation $231,789.60 $230,178.60 −$1,611.00 (−0.70%) $228,567.60 −$3,222.00 (−1.39%)

With monthly reservation pricing, PayGo-only is the least expensive option. The 250-RPM baseline costs $498.00 more per 30-day month, while the 500-RPM baseline costs $996.00 more.

With one-year reservation pricing, the result reverses. The 250-RPM baseline saves $1,611.00 per year relative to PayGo-only, while the 500-RPM baseline saves $3,222.00 per year. Under this representative traffic profile, the 500-RPM baseline therefore provides the lowest annual cost of the three options.

Key Takeaway

If cost is the primary objective, use PTUs to cover the workload's stable, sustained baseline and route variable or burst traffic to PayGo. Avoid reserving PTUs for capacity that may remain idle; each additional PTU block should save more in PayGo charges than it costs to reserve.

That said, cost is not the only objective of PTUs. A correctly sized provisioned deployment also provides dedicated throughput, more predictable latency, a defined latency SLA, and more consistent benchmark results than shared PayGo capacity. The best choice therefore depends on both economics and performance requirements.


Billing and Reservation: Some Lessons Learnt by Blood

You are billed on PTUs deployed, not tokens processed: an idle deployment costs exactly the same as a saturated one. Deployments cannot be paused, so under hourly billing, charges stop only when the deployment is deleted.

  • Hourly billing charges per PTU per hour, prorated for partial hours. Use it for benchmarking, evaluation, and short-lived capacity.
  • Azure Reservations discount the effective PTU rate for a one-month or one-year commitment. This is the intended mode for sustained production workloads.

Reservations and deployments are created independently, with two consequences:

  • A reservation is a billing discount, not a capacity guarantee. Create the deployment first to confirm capacity exists, then reserve the PTUs you actually deployed.
  • If that deployment is later scaled down or deleted, the reservation keeps billing its original quantity. Deployed PTUs below it become unused coverage; PTUs above it bill at the hourly rate.

Scaling down also releases capacity back to the regional pool with no guarantee of reclaiming it, so cycling a production deployment up and down is a poor cost-control strategy. A reservation on a steady deployment is usually cheaper and safer.


References and Further Reading


Appendix

How to Query the Pricing API for This Example

{
    curl -s "https://prices.azure.com/api/retail/prices?api-version=2023-01-01-preview&\$filter=productName%20eq%20%27Azure%20OpenAI%27%20and%20armRegionName%20eq%20%27swedencentral%27"

    curl -s "https://prices.azure.com/api/retail/prices?api-version=2023-01-01-preview&\$filter=productName%20eq%20%27Azure%20AI%20Foundry%20Provisioned%20Throughput%20Reservation%27%20and%20armRegionName%20eq%20%27swedencentral%27"
} | jq -rs '
    ["Pricing option", "USD / PTU", "Reservation term"],
    (
        .[].Items[]
        | select(.skuName == "Provisioned Managed Global")
        | [
                (if .reservationTerm == null
                 then "Hourly"
                 else .reservationTerm
                 end),
                .retailPrice,
                (.reservationTerm // "None")
            ]
    )
    | @tsv'

Enter fullscreen mode Exit fullscreen mode

Top comments (0)