DEV Community

Cover image for Why a deleted backup Lambda kept billing 9,400 EBS snapshots
Muhammad Hassaan Javed for Infraforge

Posted on Edited on Originally published at infraforge.agency

Why a deleted backup Lambda kept billing 9,400 EBS snapshots

The EBS Snapshot line on the monthly bill was $1,830. The only snapshot policy on the account was an AWS Backup plan that accounted for under $20 of it. The backup Lambda that had produced the rest had been deleted thirteen months earlier, replaced by AWS Backup, and forgotten. Nobody had deleted what it created. Two volumes, snapshotted every two hours for 392 days, came to 9,408 orphans sitting on 36 TB of storage, billed at the us-east-1 EBS Snapshot rate of $0.05 per GB-month every month since.

Problem signals:

  • EBS Snapshot line is several hundred dollars a month more than the active backup policy can account for
  • describe-snapshots --owner-ids self returns thousands of entries when you expect dozens
  • Sampling a few snapshots shows source volumes that describe-volumes reports as InvalidVolume.NotFound, and descriptions written by a pipeline nobody runs any more
  • A backup Lambda or custom snapshot script was deprecated in the last 12 to 24 months
  • AWS Backup is the active tool and its dashboard shows normal counts, but the cost line tells a different story

Over $1,800 a month on a backup pipeline the account no longer used

The line item that should have been small

The EBS Snapshot line had sat at the same level for thirteen months, after climbing for the year before that. Nobody had flagged it through eight quarterly cost reviews, four while it climbed and four since. The ninth surfaced it only because a new finance lead asked what the sixth-largest line item on the account was for, and the team's mental model said it should have ranked nowhere. The only snapshot policy running was AWS Backup's, which had taken over RDS and EBS backups a year earlier, with the old Lambda plus EventBridge pipeline retired the same week.

The first instinct in the room was to pull AWS Backup's plan and see if a retention window had widened. The plan was clean. Snapshot counts there were in the low dozens, exactly what the new policy specified. So the snapshots driving the bill were coming from somewhere else.

$ aws ec2 describe-snapshots --owner-ids self \
    --query 'length(Snapshots)' --output json
9446
Enter fullscreen mode Exit fullscreen mode

The number that turned a routine cost review into an incident.

That number was the moment the room got quiet. AWS Backup writes maybe forty snapshots a month on this account. Nine thousand was a different category of problem.

AWS Backup was clean, so who made these 9,408 snapshots

Ruling out the obvious suspect

AWS Backup's vault listed 38 of the 9,446 as its own recovery points, all of volumes that still existed. With those set aside and no other named pipeline running, the question became: who created the other 9,408, and is anything still creating more. We pulled the StartTime field on the most recent hundred of them. The newest one was thirteen months old. Whatever pipeline made them had stopped, which meant we were looking at a stable population, not a leak that was still growing. That mattered because it meant the cleanup had a known size.

The next question was whether the source volumes were still around. We sampled twenty random snapshots and ran describe-volumes against their VolumeId. All twenty came back InvalidVolume.NotFound, and all twenty carried the description the Lambda had written on every snapshot it took: ebs-backup-lambda: followed by the volume ID. The pattern was clear: the snapshots were referencing two specific volume IDs (the Lambda snapshotted two production EBS volumes every two hours), both of which had been deleted along with the EC2 instances they served when the application moved to a managed service.

rm -f orphan-source-volumes.txt
aws ec2 describe-snapshots --owner-ids self \
    --query 'Snapshots[*].[SnapshotId,VolumeId,StartTime,Description]' \
    --output text > all-snapshots.tsv || exit 1

# Copies carry a placeholder volume ID such as vol-ffffffff; never treat it
# as a source. Only InvalidVolume.NotFound means gone: any other error
# (access denied, throttling, an expired session) stops the loop, and the
# list exists only if every volume was checked in this run.
awk -F'\t' '$2 !~ /^vol-f+$/ {print $2}' all-snapshots.tsv | sort -u \
  | ( while read vid; do
        if err=$(aws ec2 describe-volumes --volume-ids "$vid" 2>&1 >/dev/null); then
          continue
        fi
        case "$err" in
          *InvalidVolume.NotFound*) echo "$vid orphan" ;;
          *) echo "stopping at $vid: $err" >&2; exit 1 ;;
        esac
      done ) > orphan-source-volumes.tmp \
  && mv orphan-source-volumes.tmp orphan-source-volumes.txt
Enter fullscreen mode Exit fullscreen mode

Group snapshots by their source volume and check which source volumes still exist. A missing volume is a lead, not a verdict: copies, kept final snapshots and failed API calls all look like one.

Only two volume IDs appeared in the orphan list. Two volumes, one snapshot every two hours, 392 days of runtime before the Lambda was deleted: 2 x 12 x 392 is 9,408. The arithmetic closed exactly, which told us we were looking at the whole population and not a subset of it. The Lambda that created them was gone, but AWS does not garbage-collect snapshots when their creator disappears. Snapshots are first-class objects with their own lifecycle, and nothing in CreateSnapshot sets an expiry: a snapshot lives until something deletes it. The Lambda had written a description on each one and set no tags.

Why we sampled twenty before touching the other 9,388

What we did before running delete-snapshot in a loop

The temptation at this point is to write a one-line loop and delete everything. delete-snapshot is irreversible unless a Recycle Bin retention rule catches the snapshot, and this account had none. The cost was real, over $1,800 a month for storage of data that referenced infrastructure that no longer existed. Two reasons we did not run the loop immediately.

First, orphan is sometimes a transient state. A volume gets deleted on Tuesday during a planned migration. On Wednesday the orphan-finder runs. A snapshot taken two hours before the volume's deletion looks orphaned but is actually the most recent backup of a service that was just migrated. Deleting it would destroy the only remaining copy of that data. We checked when the two source volumes had been deleted. The migration runbook dated both to thirteen months earlier, the same week the Lambda went, so even the newest snapshot of either volume was more than a year old. The cohort was uniformly historical. No active workflow could be depending on any of them.

Second, we needed to be sure these snapshots were not being referenced as the base for any AMI or any live AWS Backup recovery point. We ran describe-images with a block-device-mapping.snapshot-id filter across the full list, a batch of IDs per call, expecting nothing, and got nothing. We checked the AWS Backup recovery point inventory. None of the orphan snapshot IDs appeared there. The deletion was safe.

The deletes needed little API capacity, and we spread them across three calendar days on purpose. EC2 throttles each account per Region, and DeleteSnapshot gets a token bucket of 100 requests that refills at 5 per second, so 9,400 deletes plus retries on the occasional RequestLimitExceeded fit in well under an hour of API throughput. The loop's 250 ms sleep alone came to about forty minutes across 9,408 iterations, and each CLI call added its own time on top. Nobody wanted nine thousand irreversible deletes running in one unsupervised burst, so we split the list into three batches of about 3,100, one per day, each batch gated on someone reading the previous day's failed.log and signing off before the next one started. The loop used its append-only deleted.log as the checkpoint, so we could resume after any interruption without re-trying the ones it had recorded; an interruption between a delete and its log line only sends that one ID to failed.log on the rerun. If you run a cleanup this size, a temporary Recycle Bin retention rule for EBS snapshots is the safety net: a wrong delete becomes a restore. Untagged snapshots like these need a Region-level rule, or tags added first for a tag-level one, and a Region-level rule also holds anything else deleted in the Region meanwhile, AWS Backup's expired snapshots included. Every day of retention keeps the deleted snapshots billed, about $60 a day for this cleanup. Confirm the rule holds one test delete before the real run: AWS warns that the first rule for a resource type in a Region might not start retaining immediately.

awk '{print $1}' orphan-source-volumes.txt > orphan-vols.txt
# A snapshot makes the list only if its volume is gone AND it carries the
# Lambda's own description: copies, AWS Backup's snapshots and anything
# made by hand are left alone.
awk -F'\t' 'NR==FNR{o[$1];next}
  ($2 in o) && index($4, "ebs-backup-lambda: ") == 1 {print $1}' \
    orphan-vols.txt all-snapshots.tsv > orphan-snapshot-ids.txt

touch deleted.log
while read sid; do
  if grep -qx "$sid" deleted.log; then continue; fi
  aws ec2 delete-snapshot --snapshot-id "$sid" \
    && echo "$sid" >> deleted.log \
    || echo "$sid" >> failed.log
  sleep 0.25
done < orphan-snapshot-ids.txt
Enter fullscreen mode Exit fullscreen mode

The list comes from two independent facts, a vanished volume and the creator's own description. The loop is resumable and rate-limited, and the checkpoint file is the load-bearing part.

After three days the EBS Snapshot line on the next monthly forecast dropped to under $20, which is what AWS Backup's own EBS snapshots cost. The thirty-six terabytes of orphan storage was gone.

Tag at creation, schedule the cleanup, watch the lines that should be zero

The rule that meant the next deprecated pipeline could not do this

The deletion fixed the symptom. The interesting part of this engagement was the cause. AWS does not couple a snapshot's lifecycle to the lifecycle of whatever process created it. A Lambda gets deleted, an EventBridge rule gets removed, the IAM role goes with them, and the snapshots they made keep existing and keep being billed, forever, until something explicitly deletes them. There is no warning email. There is no dashboard widget. The only signal is the monthly bill, and on this account the bill had been loud for more than a year before anyone asked about it.

Two changes went in after the cleanup. The first was a tag-at-creation rule. Every snapshot the account creates now carries Owner (a team or service name) and CreatedBy (the pipeline that made it). The handful of custom Lambdas that survived the migration were rewritten to apply those two plus a third, Retention (an ISO date past which the snapshot is safe to delete), computed at creation time. AWS Backup cannot supply that third tag: a backup plan rule's RecoveryPointTags is a static key/value map with no variable or date substitution, and tag-copy only propagates the source resource's existing tags, so there is no field in a plan that produces a per-snapshot expiry date. AWS Backup gets static Owner and CreatedBy recovery point tags instead, and expresses expiry the way it is designed to, through the rule's lifecycle (DeleteAfterDays). A weekly cleanup Lambda walks the account, deletes anything past its Retention date, and flags anything older than 90 days with no Retention tag. It skips any snapshot that one of the account's backup vaults lists as a recovery point, because the plan lifecycle governs those and EC2 refuses to delete them anyway; when one of those genuinely needs removing it goes through AWS Backup, not the EC2 API. For the first 60 days the Lambda posted a Slack message and waited for a thumbs-up before deleting. After that it ran unattended.

The second change was to the quarterly cost review process. It now starts with the line items that should be zero or near zero, not the ones that are already big. The big lines get watched constantly by capacity planners. The lines that should be zero or near it are where deleted infrastructure leaves footprints, and they are the ones least likely to be on anybody's dashboard. An EBS Snapshot line far above what the backup plan writes. Lambda invocations on a service that was migrated to ECS six months ago. NAT Gateway hours on a VPC with nothing left in its private subnets. These are the lines where deprecated pipelines keep paying rent.

The lifecycle every snapshot now goes through. Untagged snapshots cannot live past 90 days without an explicit decision.

The lifecycle every snapshot now goes through. Untagged snapshots cannot live past 90 days without an explicit decision.

Cost archaeology on accounts where a deprecated pipeline is still paying rent

When the bill is the only thing telling you what you forgot

The shape of this incident is common. A pipeline gets shipped, the engineer who wrote it leaves, the policy gets replaced but the outputs survive, and the bill never comes back down. EBS snapshots are the most common shape we see. Detached EIPs are close behind. Idle NAT gateways and orphaned ElastiCache clusters round out the top four. None of these line items alarm on a CloudWatch dashboard because nothing is actively misbehaving. The deprecated pipeline is the misbehavior, and the pipeline no longer exists.

We run these cost-archaeology engagements regularly. In the last quarter we walked through three accounts where a single deprecated backup pipeline accounted for more than half of the account's EBS Snapshot line. We have an inventory script that finds orphan snapshots, detached volumes, unused EIPs, and idle NAT gateways across an account in about 20 minutes, plus a sample-then-delete workflow we walk the team through live so nothing irreversible happens on autopilot. The deletion is always the easy half. The work is figuring out which orphans are safe and writing the tag-at-creation policy that stops the next one.

If your bill has a line that does not match anything that should be running, the orphan audit is usually the fastest way to find out where it is going. Request an infrastructure review and we will run the audit with your team on a 30-minute diagnostic call this week. You can also see the broader pattern in our services overview for cloud cost spike work.


Originally published at https://infraforge.agency/insights/orphan-ebs-snapshots-deleted-backup-pipeline-cost-spike/.

If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.

Top comments (0)