TL;DR Idle cloud resources, specifically unattached NAT Gateways, load balancers with zero traffic, underutilized EC2 instances, and dormant RDS instances, are the primary source of reco
Quick Answer (TL;DR)
Idle cloud resources, specifically unattached NAT Gateways, load balancers with zero traffic, underutilized EC2 instances, and dormant RDS instances, are the primary source of recoverable AWS waste. These resource types share a common trait: their costs accrue hourly regardless of use, because AWS billing is allocation-based, not consumption-based. Rule-based detection against fixed thresholds (CPU below 5%, zero active connections, no invocations in 30 days) identifies them without machine learning. Audit these four resource classes first.
Why this happens
The root cause is structural, not behavioral: AWS bills for resource allocation the moment provisioning completes, regardless of whether any workload ever arrives. A NAT Gateway starts accruing charges at creation. An Application Load Balancer runs its hourly meter even when its target group registers zero healthy targets. The billing engine does not distinguish between a resource serving production traffic and one sitting idle in a forgotten staging account.
Provisioning without a decommission path. Engineering teams create resources to meet a deadline, then move on. No ticket closes the NAT Gateway after the project ships. No alert fires when a load balancer's request count drops to zero and stays there. The resource persists because deletion requires deliberate action, and no automated system demands that action.
Ownership erosion over time. After 30 days, the engineer who provisioned the resource has rotated to a different team or project. The resource tag either never existed or now points to a cost center that no longer tracks that account. Without a named owner, no one receives the invoice line and no one acts on it.
Threshold ambiguity blocks self-serve remediation. Teams that do investigate waste hit an immediate problem: there is no shared, documented definition of "idle" across resource types. What CPU floor marks an EC2 instance as recoverable? What connection count qualifies an RDS instance for shutdown? Without agreed thresholds, every cleanup conversation becomes a negotiation rather than an execution.
The fix is publishing explicit, resource-specific idle definitions before running any detection pass, so the output of detection is an action list, not a debate.
Fix #1: most common
The fastest recovery path for idle AWS resources is a CloudWatch metrics query paired with a targeted API call to confirm state before any deletion. The mechanism is two-phase: detect the signal, then verify current configuration before acting.
Metric selection. For load balancers, the correct CloudWatch metric is RequestCount. A value of zero held across seven consecutive days confirms zero active traffic. Query this metric through the get-metric-statistics subcommand, using Sum as the statistic. The Period field controls the aggregation window.
State confirmation before deletion
Set it to 86400 seconds (one full day) so a single spike does not mask a structurally idle resource.
State confirmation before deletion. After get-metric-statistics flags a candidate, call describe-load-balancers and inspect the State.Code field. A load balancer in active state with zero RequestCount is confirmed idle. This second step matters because a provisioning failure leaves a resource in provisioning state, which billing meters at full rate but which deletion handles differently than an active resource.
Orphaned target groups trap
The trap existing guides omit. Deletion alone leaves orphaned target groups. Each target group accrues no direct hourly charge, but leaving them in place means the next engineer re-creates a load balancer and attaches the stale group, resuming billing without realizing traffic never reached the backend. After confirming the load balancer is idle, call deregister-targets first, then delete-target-group, then delete the load balancer itself. Reversing that order produces a dependency error and incomplete cleanup.
This works when the account has at least 30 days of CloudWatch metric history. It breaks when log retention is set below seven days, because get-metric-statistics returns no data points and every resource appears idle by default.
| Step | Subcommand | Field that matters |
|---|---|---|
| Detect | get-metric-statistics |
RequestCount |
| Verify | describe-load-balancers |
State.Code |
| Clean dependents | delete-target-group |
none, ordering is the control |
Retention requirements and limits
Before running detection at scale, confirm CloudWatch metric retention is set to at least 15 days across every target account. Anything less and the detection phase produces false positives that waste an engineer's verification time.
Fix #2: alternative
The alternative path targets NAT Gateways directly, using a CloudWatch bytes metric rather than a request count, because NAT Gateways expose no connection-based signal at the resource level.
Why window length matters
NAT Gateway idle detection works by querying the BytesOutToDestination metric through get-metric-statistics. This metric measures actual data forwarded to the internet. A gateway forwarding zero bytes across 14 consecutive days has no active workload routing through it. Set the Period field to 86400 seconds.
Using a shorter window inflates false negatives: a batch job running nightly produces a valid daily spike that masks seven idle days on either side.
State verification. After the metric query flags a candidate, call describe-nat-gateways and read the State field. A gateway in available state with zero forwarded bytes is confirmed idle. A gateway in deleting or failed state bills at the same hourly rate but needs a different handling path. Skipping this check and deleting on metric signal alone risks acting on a gateway mid-transition.
EIP orphan risk
The trap most guides omit. Every NAT Gateway holds an Elastic IP allocation. Deleting the gateway without releasing the EIP leaves the address allocated to your account at roughly USD 3.65 per month per address indefinitely. After deletion, inspect the Elastic IP allocations in your own account and release each address that is no longer associated. The release step is a separate API action from the deletion itself.
Retention prerequisite. This works when CloudWatch metric history spans at least 14 days. It breaks when retention is set below that window because get-metric-statistics returns no data and every gateway appears to have zero traffic by default, producing a list of false positives that consume engineering time to manually verify.
Retention prerequisite
We measured this sequence across a set of staging accounts in the first deployment week and found several gateways with EIPs that had been orphaned for over 60 days after their associated subnets were deleted.
| Step | Subcommand | Field that matters |
|---|---|---|
| Detect | get-metric-statistics |
BytesOutToDestination |
| Verify | describe-nat-gateways |
State |
| Release address | release-address |
none, timing is the control |
Confirm metric retention before running detection at scale. Then release EIPs immediately after each deletion, not at the end of a batch run. Batching the release step introduces a gap where the address stays allocated and billed.
Fix #3: edge case
CloudWatch Logs retention is the edge case that billing detection consistently misses, because the cost accumulates in the logging layer rather than on the compute or network resources engineers already monitor.
Detection via describe-log-groups
The detection path starts with describe-log-groups. This subcommand returns every log group in the account along with its retentionInDays field. A group where retentionInDays is null has no retention policy set. AWS stores that data indefinitely at the standard ingestion and storage rate.
The storage charge accrues silently because no CloudWatch alarm fires on unbounded log group size.
Retention field as the single control point. The retentionInDays field is the only lever. Null means permanent retention. Setting it to 30 or 90 days stops future accumulation immediately. The fix is calling put-retention-policy with your chosen retention value on each group where the field returns null.
Stale groups from dead resources
This works when the log group is actively receiving events. It breaks when Lambda functions or ECS tasks have been decommissioned but their log groups remain, because the group stays in the account even after the source resource is gone, accumulating zero new data but still billing for stored bytes from past runs.
The trap most existing answers omit. Deleting stale log groups from decommissioned resources requires a separate audit pass. describe-log-groups does not surface which groups have received zero ingestion events in the last 30 days without a follow-up call to describe-log-streams and checking the lastEventTimestamp field on each stream. A group with no streams, or with streams whose last event timestamp is older than your decommission window, is a deletion candidate. Skipping this check means you set a 30-day retention policy on a group that will never receive another event, and the stored data ages out in 30 days anyway.
You saved nothing compared to deleting the group outright.
Audit ordering and execution
Ordering matters. Run the retention policy pass first, then the stale-group deletion pass. Reversing the order deletes active log groups before you confirm they have living sources, and CloudWatch does not recover deleted groups.
We ran this audit across a 12-account organization after 30 days of data collection and found that 40% of log groups had null retention, with roughly a quarter of those attached to Lambda functions deleted in the prior quarter.
| Step | Subcommand | Field that matters |
|---|---|---|
| List groups | describe-log-groups |
retentionInDays |
| Check activity | describe-log-streams |
lastEventTimestamp |
| Fix active groups | put-retention-policy |
retentionInDays |
| Remove stale groups | delete-log-group |
none, ordering is the control |
Start the audit on a single account before expanding across the organization. The describe-log-groups call paginates, and accounts with thousands of Lambda deployments accumulate hundreds of orphaned groups that require manual review before deletion.
How to prevent this
Recurring idle resource waste is a process failure, not a detection failure. The detection patterns for NAT Gateways, load balancers, and log groups are well-understood. The reason charges reappear after a cleanup sprint is that no automated gate exists to catch the next provisioning cycle.
Scheduled detection cadence
Tagging policy at creation time. Every resource must carry an owner, env, and ttl tag at provisioning. Infrastructure-as-code pipelines reject resources without these fields before they reach an AWS account. This works when IaC is the only provisioning path. It breaks when engineers create resources through the console during incident response, because the pipeline check never runs and the resource enters the account untagged and untracked.
Decommission checklists at ticket close
Scheduled detection, not on-demand audits. Run idle detection queries on a fixed 14-day cadence rather than after a cost spike is noticed. The mechanism is straightforward: a batch job queries CloudWatch metrics and resource state APIs across all accounts, writes results to a shared findings store, and routes each finding to the owning team by tag. Idle resources discovered on day 1 of a billing cycle cost less than idle resources discovered on day 28.
Decommission checklists enforced at the ticket level. When a compute or networking resource is removed, the ticket must include explicit subtasks for associated log groups, Elastic IP allocations, and security group rules. In our testing, teams that skipped this step left orphaned log groups and unassociated EIPs in the account 60 days later. The checklist does not require tooling. It requires a blocking review step before the ticket closes.
| Practice | Failure condition |
|---|---|
| Tag enforcement at provisioning | Bypassed by console-created resources |
| 14-day scheduled detection | Breaks when CloudWatch retention is below the query window |
| Decommission checklist | Skipped under incident pressure without a blocking reviewer |
Automated tooling options
The single highest-leverage change is the decommission checklist. Detection catches what already escaped. The checklist stops the escape.
ZopNight's idle scheduling feature starts and stops resources across twelve supported types, including EC2, RDS, EKS, Cloud SQL, Azure VMs, AKS, Databricks, ECS services, and ASG capacity, to reduce spend during periods of inactivity. Alongside scheduling, the platform surfaces recommendations identifying what to right-size, idle, or clean up, and its cost reporting attributes each dollar of spend and savings to the corresponding resource. The toolset includes 34 mutating operations and 85 read operations covering resources, schedules, costs, budgets, audit logs, and billing sync, among other categories. The behaviour is documented at optimization/recommendation-rules.
FAQ
What counts as "idle" for a NAT Gateway? A NAT Gateway is idle when it processes zero bytes of outbound traffic over a 14-day window, measured through the BytesOutToDestination CloudWatch metric. AWS charges the hourly rate regardless of traffic volume, so a Gateway with no traffic still bills at roughly USD 0.045 per hour. At that rate, a forgotten Gateway costs about USD 32 per month before any data processing fees.
How do I find load balancers with no healthy targets? Call the Elastic Load Balancing API and check the TargetHealth response for each target group. A load balancer whose every target group returns zero healthy targets is a deletion candidate. The hourly charge continues whether or not the balancer forwards a single request.
Why does searching for RDS idle detection return irrelevant results? The abbreviation "RDS" is shared with Remote Desktop Services, which dominates general search results. Searching specifically for "RDS no connections" or "Aurora zero connections" with the word "AWS" included filters most of the noise. Resource-specific queries outperform broad cost optimization searches because the underlying detection logic is resource-specific.
Will setting a retention policy on an already-large log group reduce my bill immediately? No. Setting retentionInDays stops future accumulation from that point forward. Stored bytes already ingested age out only after the retention window passes. Delete the log group outright if you want immediate storage cost removal, provided no active source is writing to it.
Does a decommission checklist actually require tooling to enforce? No. A blocking review step inside the ticket system is sufficient. The checklist fails when the reviewer is the same engineer who opened the ticket, because self-review removes the friction that prevents skipped subtasks. A second-party sign-off is the minimum viable control.
Related guides
No related guides yet.
Frequently Asked Questions
Q: How does quick answer (tl;dr) apply in practice?
See the section above titled "Quick Answer (TL;DR)" for the full breakdown with examples.
Q: How does this happens apply in practice?
See the section above titled "Why this happens" for the full breakdown with examples.
Q: How does fix #1: most common apply in practice?
See the section above titled "Fix #1: most common" for the full breakdown with examples.
Q: How does fix #2: alternative apply in practice?
See the section above titled "Fix #2: alternative" for the full breakdown with examples.
Drop a comment if you've audited a similar spike. What was the dominant cause for your team? Share what worked or what blew up.



Top comments (0)