DEV Community

Cover image for Introduction to Azure API Management
Martin Oehlert
Martin Oehlert

Posted on AI-assisted

Introduction to Azure API Management

Which API Management tier goes in front of an order API that gets 5 million calls a month and none at night? Consumption bills about $14 for that month and nothing while the API sits idle, Basic v2 bills about $150, and you can't move from one to the other later without building a new instance. That last part turns a pricing choice into a day-one architecture decision.

Azure API Management (APIM) is a managed gateway between the callers of your API and the backends that serve it: Azure Functions, Container Apps, App Service, or anything else that answers HTTP. Callers see one hostname and one contract. The gateway takes over the jobs that otherwise end up scattered: function keys mailed to partners and never rotated, rate limiting copied into every function, an API reference in a wiki that drifted two releases ago, a reverse proxy on a VM you patch yourself.

The running example for this series is an order API: an Azure Functions app on the isolated worker model and .NET 10, in the ApimOrdersDemo folder of the companion repo. It starts small and public on purpose. If you already need the enterprise version (Front Door at the edge, APIM Standard v2 on a private endpoint, Container Apps behind it), What Each Hop Trusts walks that chain one hop at a time.

Pick the tier before you deploy anything

The sentence that decides most of this section sits in the scaling docs: "Currently, you cannot upgrade from or downgrade to the Consumption tier" (upgrade and scale). The other families have walls too. Classic tiers change among themselves in place, Basic v2 and Standard v2 change between each other, and nothing crosses from classic or Consumption into v2, because v2 is available for newly created service instances only.

Crossing a wall means a new instance, with a new *.azure-api.net hostname and new subscription keys. The cheap insurance is a custom domain on the gateway from the first deploy (api.contoso.com instead of apim-orders.azure-api.net), which turns a later move into a DNS change on your side. The keys still change unless you copy their values across; subscriptions accept explicit primary and secondary key values, which is the escape hatch for that day.

Three families, eight tiers

APIM comes in three families (key concepts):

  • Consumption: a "serverless gateway ... that scales based on demand and bills per execution". No units, no idle cost.
  • Classic: Developer, Basic, Standard, Premium. You pay per unit per hour and calls are unlimited.
  • v2: Basic v2, Standard v2, Premium v2. Also billed per unit; Basic v2 and Standard v2 include a call allowance and meter the calls past it.

v2 does not replace classic. The v2 overview says "there are no changes to the classic Developer, Basic, Standard, or Premium tiers", and there is no retirement date. If you start fresh, "classic is legacy" is not a reason to pick v2; the features further down are.

Price, included calls, SLA and maximum units for the eight API Management tiers, West Europe list prices, October 2026

Prices are pay-as-you-go list prices in USD for West Europe, an hourly unit price times 730 hours, from the pricing page and the Azure Retail Prices API on 2026-10-02. SLAs come from the same pricing page, unit limits from the feature comparison.

Developer is the tier without an SLA. It has almost every feature, VNet injection included, but Microsoft scopes it to "non-production use cases and evaluations", it can't scale past one unit, and moving to or from it causes downtime. Labs only.

Consumption has an SLA: 99.95%, the same number as Basic, Standard, Basic v2 and Standard v2. What it doesn't have is a unit that stays warm. Capacity is assigned when traffic arrives, so the first call after a quiet spell waits for the gateway. Microsoft publishes no number for that delay; the closest thing to an official statement is an accepted answer on Microsoft Q&A that says it "spins up resources on-demand. This can lead to cold start delays, especially if there are infrequent or intermittent requests." An availability SLA promises that the gateway answers, not how long the first answer takes.

So I measured it on the sample from this article: a Consumption gateway in front of a Flex Consumption function, West Europe, 2 October 2026, curl from outside Azure with a new TLS connection per call, and Application Insights on both layers so every call splits into gateway time and backend time. Warm, a call through the gateway took a median of 0.29 seconds (29 calls), against 0.20 seconds straight to the function. After 31 minutes without traffic, five times: with both layers cold, the first call took 46.0, 9.4 and 17.0 seconds. With the function woken first, so only the gateway was cold, it took 50.5 seconds once; the other time the call hung in the TLS handshake for 120 seconds and failed, and the retry took 51.9 seconds.

Application Insights shows where the time went. The function added about 2 seconds whenever it was cold, and the gateway's own processing never took more than a third of a second (333 ms at worst, 0 to 2 ms warm). The rest, 15 to 50 seconds in four of the five windows, passed after the TLS handshake and before the gateway logged the request at all, which points at the Consumption gateway getting capacity assigned. The 9.4-second call was the exception, mostly my own DNS lookup. An earlier run the same day without Application Insights saw 59.8 seconds once, which fits that pattern, but also gateway-only cold calls of 1.3 and 1.5 seconds, so the long wait doesn't happen every time. Nine cold windows on two deployments are not a distribution, but they put the tail on the gateway side, at tens of seconds, and nothing in the SLA bounds it.

The 99.99% on Premium and Premium v2 isn't automatic either: it needs the instance spread across two or more availability zones (or, on classic Premium, regions), which means at least two units on the bill.

The capabilities that decide it

Inbound private endpoint, VNet, availability zones, multi-region, self-hosted gateway, cache, developer portal and per-key limits for each API Management tier

Sources: feature comparison, gateways overview.

The network rows decide more often than price does. If the backend lives in a VNet, Consumption, Basic, Basic v2 and classic Standard are out. Standard v2 is the cheapest tier that takes an inbound private endpoint and also reaches into a VNet outbound, which is why the enterprise chain in What Each Hop Trusts runs on it. Zones leave the two Premium tiers, and multi-region or a self-hosted gateway in your own datacenter leaves classic Premium alone.

The Consumption column has gaps that come back later in this series. The per-key policies rate-limit-by-key and quota-by-key don't run there, so "100 calls per minute per partner" needs another tier (plain rate-limit and quota scoped to a subscription do work; Part 2 covers the difference). There's no built-in cache, no developer portal, which is all of Part 4, and request logs go to Application Insights only.

Where the per-call meter stops winning

The usual advice is "Consumption for low volume, a unit-based tier once traffic grows". The crossover sits much higher than that suggests, because Basic v2 and Standard v2 meter calls too once the allowance runs out:

monthly cost in USD, n = millions of calls per month

Consumption  = 3.50 * (n - 1)
Basic v2     = 150 + 3.00 * max(0, n - 10)
Standard v2  = 700 + 2.50 * max(0, n - 50)

Consumption = Basic v2     when  n = ~247
Consumption = Standard v2  when  n = ~579
Enter fullscreen mode Exit fullscreen mode

At 5M calls that's $14 against $150 against $700. The Consumption and Basic v2 lines only meet at about 247M calls a month, roughly 95 requests per second around the clock, and by then one Basic v2 unit may need a second. On list price alone, very few APIs leave Consumption to save money. You leave it for the first call, for the features above, or because the network forces you to.

Basic v2, and the order the constraints bite

Microsoft positions Basic v2 as "designed for development and testing scenarios, and it comes with an SLA" (v2 overview). The label is narrower than the product: Basic v2 is the cheapest tier with an SLA and a unit that stays warm, and it runs the full policy set and the developer portal. For a low-volume production API where a person waits on the first call of the morning, it's what I'd run. Its ceiling is no private endpoint, no VNet, no zones and 10 units, and the way past it is the one in-place move v2 allows: Basic v2 to Standard v2, same instance, same hostname, same keys. Consumption has no such path.

Put together, the tier falls out of three questions:

  1. Does the backend live in a VNet, or must the gateway be private? Standard v2. Add zones or very high scale and it's Premium v2. Multi-region or a self-hosted gateway means classic Premium, the only tier with either.
  2. Does a person wait on the first call after a quiet period, or do you need per-key limits, the developer portal or Log Analytics request logs? Basic v2.
  3. Neither? Consumption. Machine-to-machine traffic, spiky or idle most of the day, in front of a serverless backend, is what it was built for.

Classic Basic and Standard still make sense when you need something v2 lacks today, such as backup and restore or a static IP. The order API in this series starts on Consumption because it costs nothing while you read, and the sample's apimSku parameter switches it to Basic v2.

Four nouns: API, operation, product, subscription

Everything you configure hangs off one service resource, and the nesting already tells you most of the model:

Microsoft.ApiManagement/service     the instance: gateway, management plane, developer portal
├── apis                            API: a facade over one backend
│   └── operations                  operation: GET /orders/{orderId}
├── products                        product: what a consumer signs up for
│   └── apis                        which APIs the product contains
└── subscriptions                   subscription: two keys, scoped to a product, an API or all APIs
Enter fullscreen mode Exit fullscreen mode

A partner request passes the API Management gateway, which maps its key to a subscription, the subscription to a product, the product to the orders API and the URL to an operation before forwarding to the function app

APIs and operations

An API in APIM is a facade in front of the service. Its two most important settings point in opposite directions: path is the URL suffix callers see after the gateway host (apim-orders.azure-api.net/sales), and serviceUrl is where the gateway forwards (func-orders.azurewebsites.net/api). Because they're separate, you can move the backend from Functions to Container Apps without a caller noticing; the migration article uses APIM backends for exactly that. Every API also carries subscriptionRequired, which defaults to true: a freshly imported API rejects calls without a key.

An operation is one method plus one URL template, such as GET /orders/{orderId}. You rarely create them by hand, because an OpenAPI import creates one per spec operation and names it after the operationId. Operations are also the narrowest scope a policy attaches to: a cache on GET /orders/{orderId} without touching POST /orders.

Products and subscriptions

These two get confused because they always appear together. The split is access versus credential.

A product is what a consumer gets access to: one or more APIs, plus the rules for joining. "Requires subscription" decides whether a key is needed at all, and "requires approval" whether a person signs off on new subscriptions. Partners, internal teams and a free trial can be three products over the same API, each with its own limits. Some tiers create two sample products, Starter and Unlimited, on a new instance; lock them down or delete them. Consumption doesn't create them.

A subscription is the credential: a named pair of keys, primary and secondary, scoped to a product, a single API or all APIs. Despite the name, it has nothing to do with your Azure subscription. Every instance, Consumption included, also has a built-in all-access subscription called master for the portal's test console. It opens every API, so never hand its keys to anyone.

One trap inverts the model: a product that doesn't require a subscription makes its APIs callable without a key, and APIM then ignores an invalid key instead of rejecting it. What Each Hop Trusts has the details and the fix.

The gateway and where policies live

An APIM instance has three parts (key concepts). The gateway is the data plane: the *.azure-api.net endpoint that receives every call, checks the key, runs policies and forwards. The management plane is Azure Resource Manager, so the portal, the CLI and Bicep all configure the same resources from the tree above. The developer portal is the generated site where consumers read docs and request keys.

What the gateway does to a request is defined in policies: XML documents attached at four main scopes (all APIs, product, API, operation). Every one has the same skeleton:

<policies>
    <inbound>
        <base />
    </inbound>
    <backend>
        <base />
    </backend>
    <outbound>
        <base />
    </outbound>
    <on-error>
        <base />
    </on-error>
</policies>
Enter fullscreen mode Exit fullscreen mode

inbound runs before the call reaches the backend, backend around the forward itself, outbound on the response, and on-error when anything above fails. <base /> runs the policies from the next scope up, so an operation policy without it silently drops everything defined for its API and product. That one element is where Part 2 starts.

Import the order API from its OpenAPI spec

The contract comes first. The sample keeps a hand-written OpenAPI 3.0 file next to the code, and everything APIM knows about the order API comes from it:

# openapi/orders.yaml (excerpt)
openapi: 3.0.3
info:
  title: Orders API
  version: 1.0.0
servers:
  # APIM replaces this with the API's serviceUrl on import.
  - url: https://func-orders.azurewebsites.net/api
paths:
  /orders:
    post:
      operationId: create-order
      summary: Place an order
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CreateOrderRequest'
      responses:
        '201':
          description: Order accepted.
  /orders/{orderId}:
    get:
      operationId: get-order
      summary: Get one order
      parameters:
        - name: orderId
          in: path
          required: true
          schema:
            type: string
      responses:
        '200':
          description: The order.
        '404':
          description: No order with that id.
Enter fullscreen mode Exit fullscreen mode

The full file also has GET /orders and the schemas. Keep the operationIds stable: renaming one later looks like a delete plus an add. The backend is an isolated-worker Functions app on .NET 10 whose function names and routes match the spec one to one (create-order on POST orders, and so on); the code is in ApimOrdersDemo.

With Bicep

This is the path the sample deploys, and the one to keep once the API is real:

resource apim 'Microsoft.ApiManagement/service@2024-05-01' = {
  name: 'apim-orders-${resourceToken}'
  location: location
  tags: tags
  sku: {
    name: apimSku
    capacity: apimSku == 'Consumption' ? 0 : 1
  }
  properties: {
    publisherEmail: publisherEmail
    publisherName: 'Orders demo'
  }
}

resource ordersApi 'Microsoft.ApiManagement/service/apis@2024-05-01' = {
  parent: apim
  name: 'orders'
  properties: {
    displayName: 'Orders API'
    path: 'sales'
    protocols: ['https']
    format: 'openapi'
    value: loadTextContent('../openapi/orders.yaml')
    serviceUrl: functionBaseUrl
    subscriptionRequired: true
  }
}
Enter fullscreen mode Exit fullscreen mode

format: 'openapi' means "OpenAPI 3, inline YAML", and loadTextContent reads the file at compile time, so the spec travels inside the deployment and APIM never has to fetch it. serviceUrl overrides the spec's servers entry, which is why the placeholder hostname in the YAML never matters. path: 'sales' puts the operations at apim-orders-<token>.azure-api.net/sales/orders/{orderId}.

capacity must be 0 on Consumption and 1 or more on every other tier; the conditional keeps the switch to BasicV2 a one-parameter change. On Consumption the service resource took 2 minutes 48 seconds to provision in West Europe, and the whole template, function app included, under five minutes.

With the Azure CLI, or the portal

For an instance that already exists, or to try an import before you commit it to a template:

az apim api import -g rg-apim-orders -n apim-orders-<token> \
  --api-id orders --path sales \
  --specification-format OpenApi --specification-path openapi/orders.yaml \
  --service-url https://func-orders-<token>.azurewebsites.net/api
Enter fullscreen mode Exit fullscreen mode

An API imported this way belongs to no product yet, so the partner key from the next section gets a 401 on it ("invalid subscription key") even though the key is valid: a key only opens the APIs in its scope. The sample README covers the other spec formats and the display-name clash you hit when you import the same file twice.

The portal does the same import under APIs > Add API > OpenAPI, but nothing records what you clicked; use it to explore and to check the result. Under APIs > Orders API you should see three operations named after the spec's operationIds (list-orders, get-order, create-order), and the Settings tab should show the serviceUrl from Bicep with Subscription required checked. The Test tab sends a call with a key from the built-in master subscription, so a 200 there proves the import and the backend, not your product setup. That gets its own test below.

The backend credential

The function keeps AuthorizationLevel.Function, so a caller who finds the *.azurewebsites.net hostname still needs a function key. APIM holds that key in a secret named value and sends it on every forwarded call through a backend entity:

resource functionKeyValue 'Microsoft.ApiManagement/service/namedValues@2024-05-01' = {
  parent: apim
  name: 'orders-function-key'
  properties: {
    displayName: 'orders-function-key'
    secret: true
    value: functionKey
  }
}

resource ordersBackend 'Microsoft.ApiManagement/service/backends@2024-05-01' = {
  parent: apim
  name: 'orders-func'
  properties: {
    url: functionBaseUrl
    protocol: 'http'
    credentials: {
      header: {
        'x-functions-key': ['{{${functionKeyValue.name}}}']
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

A backend entity does nothing until a policy routes to it, so the API policy has one line for that, <set-backend-service backend-id="orders-func" />. Function keys are still shared secrets; Part 3 replaces them with the gateway's managed identity.

Before you re-import a changed spec into an existing API, know that the import replaces the operation set. An operation that disappeared from the spec disappears from APIM, policies attached to it included. That's why Parts 5 and 6 put the spec and the policies in Git and deploy them together.

Subscription keys: the first thing a caller hits

The sample's access model is three resources: a product that requires a subscription, the order API inside it, and one subscription for one partner.

resource partnerProduct 'Microsoft.ApiManagement/service/products@2024-05-01' = {
  parent: apim
  name: 'orders-partners'
  properties: {
    displayName: 'Orders for partners'
    description: 'Order intake for partner shops. One subscription per partner.'
    subscriptionRequired: true
    approvalRequired: false
    state: 'published'
  }
}

resource partnerProductOrders 'Microsoft.ApiManagement/service/products/apis@2024-05-01' = {
  parent: partnerProduct
  name: ordersApi.name
}

resource contosoShop 'Microsoft.ApiManagement/service/subscriptions@2024-05-01' = {
  parent: apim
  name: 'contoso-shop'
  properties: {
    displayName: 'Contoso Shop'
    scope: '/products/${partnerProduct.name}'
    state: 'active'
  }
}
Enter fullscreen mode Exit fullscreen mode

APIM generates the keys and never returns them as deployment outputs. Read them through the management API:

az rest --method post \
  --url "https://management.azure.com/subscriptions/<azure-subscription-id>/resourceGroups/rg-apim-orders/providers/Microsoft.ApiManagement/service/apim-orders-<token>/subscriptions/contoso-shop/listSecrets?api-version=2024-05-01" \
  --query primaryKey -o tsv
Enter fullscreen mode Exit fullscreen mode

Then call the gateway the way a partner would. These come from the sample's requests.http:

### No key: 401 from the gateway, the function never sees the call
GET {{gateway}}/orders

### Key in the header: 200
GET {{gateway}}/orders
Ocp-Apim-Subscription-Key: {{key}}

### Key in the query string: also 200, but now the key sits in every access log on the way
GET {{gateway}}/orders?subscription-key={{key}}
Enter fullscreen mode Exit fullscreen mode

Against the deployed sample, the first call came back as:

HTTP 401
{ "statusCode": 401, "message": "Access denied due to missing subscription key. Make sure to include subscription key when making requests to an API." }
Enter fullscreen mode Exit fullscreen mode

The second and third returned 200 with the order list. A made-up key gets a 401 with "Access denied due to invalid subscription key", and the function's own hostname without x-functions-key gets a 401 from the Functions host, so nobody gets around the gateway without a secret you never gave them.

APIM reads the Ocp-Apim-Subscription-Key header and only falls back to the subscription-key query parameter when the header is missing. Use the header: a key in a URL ends up in proxy logs, browser history and your own request logs. Both names can be changed per API if a partner's client already sends something else; the sample README shows how.

Give partners product-scoped subscriptions, as the sample does. API-scoped and all-APIs keys skip product-scope policies, where per-partner rate-limit and quota usually live, and an all-APIs key also opens every API you add later.

Rotating and revoking keys

Two keys exist so you can rotate without an outage: partners run on the primary, you regenerate the secondary and hand it over, and once they've switched you regenerate the primary, which becomes the spare. Nothing schedules that for you: APIM "doesn't provide built-in features to manage the lifecycle of subscription keys, such as setting expiration dates or automatically rotating keys", so the schedule lives in your runbook or a pipeline. To cut a partner off, set the subscription's state to cancelled; in the test their next call got a 401 within five seconds, as long as no open product contained the API.

One default catches people: APIM forwards the subscription key header to your backend, where it lands in whatever the backend logs. The sample's API policy removes it before the call leaves the gateway:

<set-header name="Ocp-Apim-Subscription-Key" exists-action="delete" />
Enter fullscreen mode Exit fullscreen mode

Conclusion

Almost everything in this article is configuration you'll redeploy many times: the import, the product, the keys, the backend credential, and from Part 5 on all of it lives in a Git repository. The tier is the exception. Consumption can't become Basic v2 or Standard v2 in place, so the one decision to make before the first az deployment is how the gateway may behave after a quiet night. Part 2 opens the policy file this sample already has and starts putting real rules in it.

Would you put a Consumption-tier gateway in front of an API a person calls first thing in the morning: yes or no?

Top comments (1)

Collapse
 
manojkagitha profile image
Manoj Kumar Kagitha •

Those cold-start numbers are the most honest measurement I have seen on Consumption tier. A 46-second tail on a cold gateway window is why a low-volume production API should skip Consumption even when the crossover math looks great.