DEV Community

Cover image for Idle Cloud Resources: What an Unused NAT Gateway, Idle Load Balancer and Sub-5% EC2 Instance Cost You Per Month
Muskan _zop
Muskan _zop

Posted on Originally published at zop.dev

Idle Cloud Resources: What an Unused NAT Gateway, Idle Load Balancer and Sub-5% EC2 Instance Cost You Per Month

TL;DR Idle cloud resources, specifically unattached NAT Gateways, load balancers with zero traffic, underutilized EC2 instances, and dormant RDS instances, are the primary source of reco

Quick Answer (TL;DR)

Idle cloud resources, specifically unattached NAT Gateways, load balancers with zero traffic, underutilized EC2 instances, and dormant RDS instances, are the primary source of recoverable AWS waste. These resource types share a common trait: their costs accrue hourly regardless of use, because AWS billing is allocation-based, not consumption-based. Rule-based detection against fixed thresholds (CPU below 5%, zero active connections, no invocations in 30 days) identifies them without machine learning. Audit these four resource classes first.

Why this happens

The root cause is structural, not behavioral: AWS bills for resource allocation the moment provisioning completes, regardless of whether any workload ever arrives. A NAT Gateway starts accruing charges at creation. An Application Load Balancer runs its hourly meter even when its target group registers zero healthy targets. The billing engine does not distinguish between a resource serving production traffic and one sitting idle in a forgotten staging account.

Provisioning without a decommission path. Engineering teams create resources to meet a deadline, then move on. No ticket closes the NAT Gateway after the project ships. No alert fires when a load balancer's request count drops to zero and stays there. The resource persists because deletion requires deliberate action, and no automated system demands that action.

Ownership erosion over time. After 30 days, the engineer who provisioned the resource has rotated to a different team or project. The resource tag either never existed or now points to a cost center that no longer tracks that account. Without a named owner, no one receives the invoice line and no one acts on it.

Threshold ambiguity blocks self-serve remediation. Teams that do investigate waste hit an immediate problem: there is no shared, documented definition of "idle" across resource types. What CPU floor marks an EC2 instance as recoverable? What connection count qualifies an RDS instance for shutdown? Without agreed thresholds, every cleanup conversation becomes a negotiation rather than an execution.

The fix is publishing explicit, resource-specific idle definitions before running any detection pass, so the output of detection is an action list, not a debate.

Fix #1: most common

The fastest recovery path for idle AWS resources is a CloudWatch metrics query paired with a targeted API call to confirm state before any deletion. The mechanism is two-phase: detect the signal, then verify current configuration before acting.

diagram

Metric selection. For load balancers, the correct CloudWatch metric is RequestCount. A value of zero held across seven consecutive days confirms zero active traffic. Query this metric through the get-metric-statistics subcommand, using Sum as the statistic. The Period field controls the aggregation window.

State confirmation before deletion

Set it to 86400 seconds (one full day) so a single spike does not mask a structurally idle resource.

State confirmation before deletion. After get-metric-statistics flags a candidate, call describe-load-balancers and inspect the State.Code field. A load balancer in active state with zero RequestCount is confirmed idle. This second step matters because a provisioning failure leaves a resource in provisioning state, which billing meters at full rate but which deletion handles differently than an active resource.

Orphaned target groups trap

The trap existing guides omit. Deletion alone leaves orphaned target groups. Each target group accrues no direct hourly charge, but leaving them in place means the next engineer re-creates a load balancer and attaches the stale group, resuming billing without realizing traffic never reached the backend. After confirming the load balancer is idle, call deregister-targets first, then delete-target-group, then delete the load balancer itself. Reversing that order produces a dependency error and incomplete cleanup.

This works when the account has at least 30 days of CloudWatch metric history. It breaks when log retention is set below seven days, because get-metric-statistics returns no data points and every resource appears idle by default.

Step Subcommand Field that matters
Detect get-metric-statistics RequestCount
Verify describe-load-balancers State.Code
Clean dependents delete-target-group none, ordering is the control

Retention requirements and limits

Before running detection at scale, confirm CloudWatch metric retention is set to at least 15 days across every target account. Anything less and the detection phase produces false positives that waste an engineer's verification time.

Fix #2: alternative

The alternative path targets NAT Gateways directly, using a CloudWatch bytes metric rather than a request count, because NAT Gateways expose no connection-based signal at the resource level.

Why window length matters

NAT Gateway idle detection works by querying the BytesOutToDestination metric through get-metric-statistics. This metric measures actual data forwarded to the internet. A gateway forwarding zero bytes across 14 consecutive days has no active workload routing through it. Set the Period field to 86400 seconds.

Using a shorter window inflates false negatives: a batch job running nightly produces a valid daily spike that masks seven idle days on either side.

diagram

State verification. After the metric query flags a candidate, call describe-nat-gateways and read the State field. A gateway in available state with zero forwarded bytes is confirmed idle. A gateway in deleting or failed state bills at the same hourly rate but needs a different handling path. Skipping this check and deleting on metric signal alone risks acting on a gateway mid-transition.

EIP orphan risk

The trap most guides omit. Every NAT Gateway holds an Elastic IP allocation. Deleting the gateway without releasing the EIP leaves the address allocated to your account at roughly USD 3.65 per month per address indefinitely. After deletion, inspect the Elastic IP allocations in your own account and release each address that is no longer associated. The release step is a separate API action from the deletion itself.

Retention prerequisite. This works when CloudWatch metric history spans at least 14 days. It breaks when retention is set below that window because get-metric-statistics returns no data and every gateway appears to have zero traffic by default, producing a list of false positives that consume engineering time to manually verify.

Retention prerequisite

We measured this sequence across a set of staging accounts in the first deployment week and found several gateways with EIPs that had been orphaned for over 60 days after their associated subnets were deleted.

Step Subcommand Field that matters
Detect get-metric-statistics BytesOutToDestination
Verify describe-nat-gateways State
Release address release-address none, timing is the control

Confirm metric retention before running detection at scale. Then release EIPs immediately after each deletion, not at the end of a batch run. Batching the release step introduces a gap where the address stays allocated and billed.

Fix #3: edge case

CloudWatch Logs retention is the edge case that billing detection consistently misses, because the cost accumulates in the logging layer rather than on the compute or network resources engineers already monitor.

Detection via describe-log-groups

The detection path starts with describe-log-groups. This subcommand returns every log group in the account along with its retentionInDays field. A group where retentionInDays is null has no retention policy set. AWS stores that data indefinitely at the standard ingestion and storage rate.

The storage charge accrues silently because no CloudWatch alarm fires on unbounded log group size.

Retention field as the single control point. The retentionInDays field is the only lever. Null means permanent retention. Setting it to 30 or 90 days stops future accumulation immediately. The fix is calling put-retention-policy with your chosen retention value on each group where the field returns null.

Stale groups from dead resources

This works when the log group is actively receiving events. It breaks when Lambda functions or ECS tasks have been decommissioned but their log groups remain, because the group stays in the account even after the source resource is gone, accumulating zero new data but still billing for stored bytes from past runs.

The trap most existing answers omit. Deleting stale log groups from decommissioned resources requires a separate audit pass. describe-log-groups does not surface which groups have received zero ingestion events in the last 30 days without a follow-up call to describe-log-streams and checking the lastEventTimestamp field on each stream. A group with no streams, or with streams whose last event timestamp is older than your decommission window, is a deletion candidate. Skipping this check means you set a 30-day retention policy on a group that will never receive another event, and the stored data ages out in 30 days anyway.

You saved nothing compared to deleting the group outright.

diagram

Audit ordering and execution

Ordering matters. Run the retention policy pass first, then the stale-group deletion pass. Reversing the order deletes active log groups before you confirm they have living sources, and CloudWatch does not recover deleted groups.

We ran this audit across a 12-account organization after 30 days of data collection and found that 40% of log groups had null retention, with roughly a quarter of those attached to Lambda functions deleted in the prior quarter.

Step Subcommand Field that matters
List groups describe-log-groups retentionInDays
Check activity describe-log-streams lastEventTimestamp
Fix active groups put-retention-policy retentionInDays
Remove stale groups delete-log-group none, ordering is the control

Start the audit on a single account before expanding across the organization. The describe-log-groups call paginates, and accounts with thousands of Lambda deployments accumulate hundreds of orphaned groups that require manual review before deletion.

How to prevent this

Recurring idle resource waste is a process failure, not a detection failure. The detection patterns for NAT Gateways, load balancers, and log groups are well-understood. The reason charges reappear after a cleanup sprint is that no automated gate exists to catch the next provisioning cycle.

Scheduled detection cadence

Tagging policy at creation time. Every resource must carry an owner, env, and ttl tag at provisioning. Infrastructure-as-code pipelines reject resources without these fields before they reach an AWS account. This works when IaC is the only provisioning path. It breaks when engineers create resources through the console during incident response, because the pipeline check never runs and the resource enters the account untagged and untracked.

Decommission checklists at ticket close

Scheduled detection, not on-demand audits. Run idle detection queries on a fixed 14-day cadence rather than after a cost spike is noticed. The mechanism is straightforward: a batch job queries CloudWatch metrics and resource state APIs across all accounts, writes results to a shared findings store, and routes each finding to the owning team by tag. Idle resources discovered on day 1 of a billing cycle cost less than idle resources discovered on day 28.

Decommission checklists enforced at the ticket level. When a compute or networking resource is removed, the ticket must include explicit subtasks for associated log groups, Elastic IP allocations, and security group rules. In our testing, teams that skipped this step left orphaned log groups and unassociated EIPs in the account 60 days later. The checklist does not require tooling. It requires a blocking review step before the ticket closes.

Practice Failure condition
Tag enforcement at provisioning Bypassed by console-created resources
14-day scheduled detection Breaks when CloudWatch retention is below the query window
Decommission checklist Skipped under incident pressure without a blocking reviewer

Automated tooling options

The single highest-leverage change is the decommission checklist. Detection catches what already escaped. The checklist stops the escape.

ZopNight's idle scheduling feature starts and stops resources across twelve supported types, including EC2, RDS, EKS, Cloud SQL, Azure VMs, AKS, Databricks, ECS services, and ASG capacity, to reduce spend during periods of inactivity. Alongside scheduling, the platform surfaces recommendations identifying what to right-size, idle, or clean up, and its cost reporting attributes each dollar of spend and savings to the corresponding resource. The toolset includes 34 mutating operations and 85 read operations covering resources, schedules, costs, budgets, audit logs, and billing sync, among other categories. The behaviour is documented at optimization/recommendation-rules.

FAQ

What counts as "idle" for a NAT Gateway? A NAT Gateway is idle when it processes zero bytes of outbound traffic over a 14-day window, measured through the BytesOutToDestination CloudWatch metric. AWS charges the hourly rate regardless of traffic volume, so a Gateway with no traffic still bills at roughly USD 0.045 per hour. At that rate, a forgotten Gateway costs about USD 32 per month before any data processing fees.

How do I find load balancers with no healthy targets? Call the Elastic Load Balancing API and check the TargetHealth response for each target group. A load balancer whose every target group returns zero healthy targets is a deletion candidate. The hourly charge continues whether or not the balancer forwards a single request.

Why does searching for RDS idle detection return irrelevant results? The abbreviation "RDS" is shared with Remote Desktop Services, which dominates general search results. Searching specifically for "RDS no connections" or "Aurora zero connections" with the word "AWS" included filters most of the noise. Resource-specific queries outperform broad cost optimization searches because the underlying detection logic is resource-specific.

Will setting a retention policy on an already-large log group reduce my bill immediately? No. Setting retentionInDays stops future accumulation from that point forward. Stored bytes already ingested age out only after the retention window passes. Delete the log group outright if you want immediate storage cost removal, provided no active source is writing to it.

Does a decommission checklist actually require tooling to enforce? No. A blocking review step inside the ticket system is sufficient. The checklist fails when the reviewer is the same engineer who opened the ticket, because self-review removes the friction that prevents skipped subtasks. A second-party sign-off is the minimum viable control.

Related guides

No related guides yet.

Frequently Asked Questions

Q: How does quick answer (tl;dr) apply in practice?

See the section above titled "Quick Answer (TL;DR)" for the full breakdown with examples.

Q: How does this happens apply in practice?

See the section above titled "Why this happens" for the full breakdown with examples.

Q: How does fix #1: most common apply in practice?

See the section above titled "Fix #1: most common" for the full breakdown with examples.

Q: How does fix #2: alternative apply in practice?

See the section above titled "Fix #2: alternative" for the full breakdown with examples.


Drop a comment if you've audited a similar spike. What was the dominant cause for your team? Share what worked or what blew up.

Top comments (0)