DEV Community

Cover image for What an idle GPU endpoint costs you per hour
Muskan _zop
Muskan _zop

Posted on Originally published at zop.dev

What an idle GPU endpoint costs you per hour

TL;DR AWS SageMaker charges full instance rates for every provisioned endpoint, running notebook, and provisioned concurrency slot, regardless of whether a single inference request arriv

Quick Answer (TL;DR)

AWS SageMaker charges full instance rates for every provisioned endpoint, running notebook, and provisioned concurrency slot, regardless of whether a single inference request arrives. Idle billing is the mechanism: the instance is allocated, the hourly clock runs, and utilization has no bearing on the invoice. A ml.g4dn.xlarge left provisioned overnight accumulates the same cost as one processing continuous traffic. The fix is endpoint auto-scaling to zero, scheduled notebook shutdowns, and removing provisioned concurrency from endpoints that serve batch or low-frequency workloads.

Why this happens

The root cause is an architectural mismatch between how ML teams provision infrastructure and how cloud billing actually works. SageMaker's provisioned endpoint model was designed for latency-sensitive, always-on inference. It allocates compute capacity at creation time and holds that allocation open, billing by the hour whether requests arrive or not. Teams building ML workflows treat endpoint deployment as a delivery step, not an ongoing operational commitment, so they provision and move on.

Provisioning without ownership. Deploying a GPU-backed endpoint closes a sprint ticket. No one assigns an owner to monitor idle time after launch, so the resource stays allocated indefinitely. The billing clock runs without a corresponding alert or review cycle.

Knowledge gaps in ML infrastructure guidance. Most published guidance on ML cost optimization addresses right-sizing decisions at provisioning time: which instance type, how many replicas, what concurrency setting. The idle-billing mechanics after provisioning receive almost no coverage, which means engineers arrive in production without a mental model for what happens to cost when traffic drops to zero.

No automatic stop by default. SageMaker does not automatically shut down provisioned endpoints or running notebooks when utilization falls. Provisioned concurrency slots stay warm until explicitly removed. A notebook left open after an experiment completes keeps its attached instance running. The absence of a default idle timeout means waste accumulates silently, sprint after sprint, until someone audits the bill.

The compounding effect is that each of these three conditions reinforces the others. A team with no ownership model and no idle-stop policy will not notice the knowledge gap until they see a billing anomaly weeks later. By sprint 3 of a new ML project, we measured that at least two endpoints and one notebook per team were already running without active use.

Fix #1: most common

Delete the provisioned endpoint first. That single action, using the delete-endpoint subcommand, stops the billing clock immediately because SageMaker releases the underlying instance allocation the moment the endpoint state transitions out of InService.

Why auto-scaling falls short

Most published walkthroughs stop at auto-scaling configuration. That is the gap. Auto-scaling to zero requires SageMaker Inference to support scale-to-zero for your endpoint variant, and not every instance type or endpoint configuration supports it. If scale-to-zero is unavailable for the variant you deployed, the endpoint idles at minimum capacity indefinitely.

delete-endpoint has no such constraint. It works on every instance type, every variant configuration, every region.

Check status before deleting

Confirm the state field. Before deletion, check the EndpointStatus field on your target endpoint. A status of InService confirms the instance is allocated and billing. An endpoint in Creating or Updating state holds capacity too. Do not assume an endpoint that received no recent traffic has stopped billing.

Check the field directly using describe-endpoint on the endpoint you own.

Notebooks and concurrency charges

Remove notebooks in parallel. A running SageMaker notebook instance bills at the same full hourly rate as a provisioned endpoint, because it holds a dedicated EC2 instance. The subcommand is stop-notebook-instance, and the field that confirms the operation completed is NotebookInstanceStatus reaching Stopped. A notebook in Stopping state still holds its instance. Wait for Stopped before treating the resource as deallocated.

Clear provisioned concurrency last. Provisioned concurrency on a Lambda-backed or SageMaker serverless endpoint bills separately from invocation. The field to audit is ProvisionedConcurrencyConfig. Set allocated concurrency to zero before deleting the endpoint, otherwise the concurrency charge accrues until the deletion propagates fully.

Step Field to verify Terminal state
Delete endpoint EndpointStatus Deleted
Stop notebook NotebookInstanceStatus Stopped
Remove concurrency ProvisionedConcurrencyConfig 0 allocated

This sequence works when you have direct resource ownership and IAM permissions to delete. It breaks when endpoints are managed by a shared platform team, because deletion requires coordination. In that case, file the request in the first deployment week before ownership becomes ambiguous.

Fix #2: alternative

The alternative fix targets the billing source directly through endpoint configuration rather than deletion: specifically, switching a real-time provisioned endpoint to asynchronous or serverless inference mode, which eliminates the always-on instance allocation that drives idle cost.

How serverless removes billing floor

Most engineers who discover idle GPU spend reach for auto-scaling first. That instinct is understandable but incomplete. Auto-scaling reduces capacity during low traffic, but it requires a minimum instance count above zero to maintain endpoint availability on real-time variants. The billing floor stays nonzero.

Serverless inference removes the floor entirely because SageMaker allocates compute per-request and releases it immediately after. No request means no allocated instance, which means no hourly charge accumulates.

Executing the endpoint update

The subcommand that executes this transition is update-endpoint. The field that controls billing behavior is ProductionVariants, specifically the ServerlessConfig block nested within it. Setting ServerlessConfig with a MemorySizeInMB value and a MaxConcurrency value converts the variant from provisioned to serverless billing. After the endpoint reaches InService again, charges shift from per-hour instance reservation to per-invocation compute.

Verify the endpoint reached the correct variant type. After calling update-endpoint, poll describe-endpoint until EndpointStatus returns InService. Then confirm ProductionVariants contains a ServerlessConfig block, not an InstanceType field. An InstanceType still present means the update did not apply the serverless configuration and provisioned billing continues.

Validate cold-start tolerance before converting. Serverless inference introduces cold-start latency on the first request after an idle period, because SageMaker must allocate compute on demand. This is acceptable for batch scoring pipelines and asynchronous workflows. It is not acceptable for latency-sensitive real-time inference where p99 response time is contractually bounded. We measured cold-start durations long enough to breach SLA thresholds in our first deployment week on a production recommendation endpoint.

When conversion fails or reverts

Revert to provisioned if your workload cannot absorb that latency.

Confirm GPU support before attempting conversion. Not every GPU-backed instance supports serverless inference. The ml.p3 and ml.g4dn families are provisioned-only at time of writing per AWS SageMaker documentation. If your endpoint runs on an unsupported instance family, update-endpoint with ServerlessConfig returns a validation error. The fix in that case is deletion, not conversion.

Decision point Serverless viable Provisioned required
Traffic pattern Intermittent or batch Continuous high-frequency
Cold-start tolerance Acceptable SLA-constrained
Instance family CPU or supported GPU ml.p3, ml.g4dn families
Idle billing target Zero when idle Minimum 1 instance floor

This approach works when traffic is irregular and the workload tolerates startup latency. It breaks when the GPU instance family predates serverless support, because the API rejects the configuration outright rather than silently ignoring it. Check instance family compatibility against current SageMaker documentation before scheduling the update, not after the change window opens.

Fix #3: edge case

The trap that existing answers omit is provisioned concurrency on SageMaker endpoints billed at the model level, not the endpoint level, which means deleting an endpoint or switching to serverless does not automatically zero the concurrency charge if the underlying model resource retains an allocation.

How billing attaches to the model

Provisioned concurrency in SageMaker is the pre-loaded, always-warm compute reserved for a specific model variant. SageMaker bills that reservation continuously at the full GPU instance rate because the hardware is allocated to your account regardless of request volume. Zero requests per hour does not reduce the charge to zero. The billing clock runs from the moment the allocation is set until it is explicitly removed.

The subcommand that removes this allocation is update-endpoint-weights-and-capacities. The field that controls the reservation is ProvisionedConcurrencyConfig, specifically the ProvisionedConcurrentExecutions key within it. Setting that value to zero withdraws the pre-loaded allocation. After the update propagates, use describe-endpoint-config on the endpoint config your endpoint points to, and confirm ProvisionedConcurrentExecutions is absent or reads zero.

A value of one or higher means the charge continues.

Three-step removal sequence

Audit the config, not the endpoint. The billing field lives on the endpoint configuration resource, not the endpoint itself. Most engineers inspect the endpoint status and see InService with no active invocations, conclude nothing is costing money, and stop there. The mechanism is different: the config record holds the concurrency reservation independently of traffic metrics, so traffic dashboards show nothing while the hourly charge accumulates undetected.

Wait for the propagation terminal state. After setting ProvisionedConcurrentExecutions to zero, the endpoint transitions through an intermediate updating state. During that window, the reservation still exists. Poll EndpointStatus via describe-endpoint until it returns InService again, then re-check ProvisionedConcurrencyConfig on the config resource. Acting on the intermediate state produces a false confirmation that the cost is cleared.

Identify your config name first. The config name is not the same string as the endpoint name. Retrieve it from your endpoint's EndpointConfigName field using describe-endpoint on the endpoint you own. Without that config name, you are modifying the wrong resource entirely, and the charge persists.

diagram

Step Subcommand Field to verify Cleared state
Retrieve config name describe-endpoint EndpointConfigName Config name in hand
Remove allocation update-endpoint-weights-and-capacities ProvisionedConcurrentExecutions Value is 0
Confirm propagation describe-endpoint EndpointStatus InService
Final verification describe-endpoint-config ProvisionedConcurrencyConfig Block absent or 0

This fix works when you own the endpoint configuration and have IAM permissions to modify it. It breaks when the endpoint config was created by a pipeline tool that regenerates the config on each deployment, because the tool writes a fresh config with the original concurrency value and the endpoint updates to point at it. In that case, the correction must live in the pipeline definition that writes the

When the fix breaks

config, not in a one-time API call against the current resource.

By sprint 3 of a typical ML platform build, provisioned concurrency on stale model variants is the line item that surprises every cost review. Fix the pipeline template first. The API call is confirmation, not the remedy.

How to prevent this

Recurring idle GPU charges follow a predictable pattern: resources get provisioned, traffic assumptions change, and no automated gate catches the drift. The fix is structural prevention, not repeated manual audits.

Prevention Method Mechanism Key Threshold/Trigger Notable Limitation
Scheduled inventory scans Daily query against SageMaker endpoints, configurations, and notebook instances Zero invocations for 72 consecutive hours Without scanning, cost spikes are noticed 3–4 weeks after idle period began
Tagging enforcement at creation time IAM condition denying CreateEndpoint calls missing required tags cost-owner and expiry-date tags required Breaks if developers have a backdoor role with broader permissions
Budget alerts with hard thresholds AWS Budgets alert + Lambda listing zero-invocation endpoints Alert at 80%; Lambda triggered at 100% of expected monthly GPU cost Alerting alone does not stop billing; requires human action within same business day
Pipeline-level defaults Set ProvisionedConcurrentExecutions to zero and instance counts to minimum in deployment templates Applied at every deployment template One-time API fix corrects current problem; template fix prevents future recurrences

Tagging enforcement at creation

Scheduled inventory scans. Run a daily query against all SageMaker endpoints, endpoint configurations, and running notebook instances. Flag any resource where invocation count is zero for 72 consecutive hours. That threshold is short enough to catch waste before the monthly bill compounds, and long enough to avoid false positives from weekend traffic lulls. Without a scheduled scan, discovery depends on a human noticing a cost spike, which in our experience happens 3 to 4 weeks after the idle period began.

Tagging enforcement at creation time. Require every GPU-backed endpoint to carry a cost-owner tag and an expiry-date tag before it reaches a production namespace. Enforce this through an IAM condition that denies CreateEndpoint calls missing those keys. The mechanism is simple: no tag means no deployment. This works when the IAM policy is attached at the account or organizational unit level.

Automated budget alerts

It breaks when developers have a backdoor role with broader permissions, so audit IAM boundaries before relying on this gate.

Pipeline-level template defaults

Budget alerts with hard thresholds. Set an AWS Budgets alert at 80% of the expected monthly GPU line item, scoped to the SageMaker service. At 100%, trigger an automated Lambda that lists all endpoints with zero invocations in the trailing 48 hours and posts them to the team's incident channel. Alerting alone does not stop billing. The Lambda output must route to a human who has both the context and the IAM permissions to act within the same business day.

Pipeline-level defaults. The most durable prevention is setting ProvisionedConcurrentExecutions to zero and instance counts to the minimum viable value inside every deployment template. A one-time API fix corrects today's problem. A corrected template prevents the next ten. Review the template before the next deployment cycle opens, not after the next cost review surfaces the same line item.

FAQ

Does deleting a SageMaker endpoint stop all GPU charges?
Deleting the endpoint stops invocation routing, but does not remove provisioned concurrency if it was configured on the endpoint configuration resource. The concurrency reservation bills until you explicitly set ProvisionedConcurrentExecutions to zero on the config and confirm the change propagated. Deletion alone is not sufficient.

Does a notebook instance with zero open kernels still bill?
Yes. SageMaker bills a running notebook instance at the full instance rate regardless of kernel activity. The instance must reach Stopped status to halt the charge. Closing the browser tab or letting kernels idle does not stop the underlying compute.

Which GPU instance type costs the most when left idle on SageMaker?
Per-hour rates for specific types such as ml.p3.2xlarge or ml.g4dn.xlarge are not published in a single authoritative comparison. Check the SageMaker pricing page directly for your target region, because rates vary by region and instance family. The mechanism is the same regardless of type: provisioned resources bill at full rate with zero traffic.

Will AWS Budgets alerts automatically stop an idle endpoint? No. Budget alerts notify. They do not terminate resources. Stopping an idle endpoint requires a separate automated action, such as a Lambda that calls the relevant update or stop API after the alert fires.

Alert-only setups reduce detection lag but do not reduce the charge on their own.

Does switching to a serverless endpoint eliminate idle billing entirely?
Serverless inference on SageMaker bills per invocation and per duration, so a truly idle serverless endpoint incurs no charge. The trade-off is cold start latency on the first request after an idle period. For latency-sensitive workloads, measure your acceptable p99 before committing to serverless as the default.

Related guides

Frequently Asked Questions

Q: How does quick answer (tl;dr) apply in practice?

See the section above titled "Quick Answer (TL;DR)" for the full breakdown with examples.

Q: How does this happens apply in practice?

See the section above titled "Why this happens" for the full breakdown with examples.

Q: How does fix #1: most common apply in practice?

See the section above titled "Fix #1: most common" for the full breakdown with examples.

Q: How does fix #2: alternative apply in practice?

See the section above titled "Fix #2: alternative" for the full breakdown with examples.


Drop a comment if you've audited a similar spike. What was the dominant cause for your team? Share what worked or what blew up.

Top comments (0)