As cloud infrastructure scales, idle resources like unattached EBS volumes, obsolete EC2 instances, and abandoned Elastic IPs quietly drain your budget. Managing this manually in an enterprise environment with hundreds of resources is impossible.
With over 12 years in IT and 6+ years managing production workloads on AWS, I’ve found that the best way to control cloud spend is to shift from reactive monitoring to automated remediation.
In this post, I will share a production-ready blueprint to automatically identify and clean up idle AWS resources using AWS Lambda, Amazon EventBridge, and AWS Systems Manager (SSM).
🏗️ The Automated Architecture
Instead of just looking at cost dashboards, we can build a serverless automation engine that cleans up resources during off-peak hours:
- Amazon EventBridge triggers a cron job every Friday at 6:00 PM.
- An AWS Lambda function scans the account for unattached EBS volumes and idle Elastic IPs.
- If an idle resource is found, Lambda triggers an AWS Systems Manager (SSM) Automation Document to snapshot the volume for safety and then delete it.
🛠️ Step 1: The Lambda Automation Script (Python & Boto3)
Here is a snippet of the Lambda function used to detect and remove unattached EBS volumes. It checks for volumes that have been in an available (unattached) state for more than 7 days.
import boto3
from datetime import datetime, timedelta, timezone
def lambda_handler(event, context):
ec2 = boto3.resource('ec2')
client = boto3.client('ec2')
# Threshold for idle resources (7 days)
retention_days = 7
cutoff_date = datetime.now(timezone.utc) - timedelta(days=retention_days)
print("Scanning for unattached EBS volumes...")
for volume in ec2.volumes.all():
if volume.state == 'available':
# Check how long the volume has been unattached/available
if volume.create_time < cutoff_date:
print(f"Found idle volume: {volume.id}. Creating final snapshot and deleting...")
# Create snapshot for data safety
client.create_snapshot(
VolumeId=volume.id,
Description=f"Automated pre-deletion snapshot for {volume.id}"
)
# Delete the volume
volume.delete()
print(f"Successfully deleted volume: {volume.id}")
💡 Lessons Learned from Production Deployments
Before rolling this out in your own environments, keep these two enterprise guardrails in mind:
-
The Opt-Out Tag: Always give your development teams a way to bypass automation. Implement a tag check (e.g.,
SkipCostAutomation = True) in your Lambda code so critical resources aren't deleted. -
Dry-Run Mode First: When deploying this to production for the first time, configure your script to log the idle resources to Amazon CloudWatch or send a Slack alert via AWS Chatbot before actually executing the
volume.delete()command.
🚀 What's Next?
Automating resource cleanup not only saves your organization money, but it also reduces your cloud attack surface by eliminating unmanaged infrastructure.
In my next post, I'll share the Terraform templates to deploy this entire infrastructure as code (IaC) across a multi-account AWS organization.
How does your team handle cloud cost optimization? Do you rely on manual checks, or have you automated your cleanup processes? Let's discuss in the comments below!
Top comments (0)