October 2026 | ~16 min read
The $2,000/Month Creep
Three months after our 71% EKS cost reduction, we watched our spend climb from $4,180/month back up to $6,200/month. Same teams, same cluster, same Karpenter configs. We'd right-size pods, and two months later the same engineers were back to requesting 4x what they needed. The problem wasn't tools — it was that nobody saw their own costs.
We needed a feedback loop.
So we built a four-layer monitoring stack: Kubecost for cost allocation, Prometheus and Grafana for dashboards, CloudWatch alarms for automated alerts, and a weekly Slack report that named names. Within 30 days, spend stabilized at $3,800/month. Teams self-corrected when they could see their own numbers. No new policies. No enforcement automation. Just visibility.
This post is the complete build guide — every Helm value, every Grafana dashboard JSON, every Terraform alarm, and the Python report that changed how our teams think about resources.
The Numbers That Matter
┌──────────────────────────────────────────────────────────────┐
│ COST MONITORING IMPACT │
├──────────────────────────┬──────────┬──────────┬─────────────┤
│ Metric │ Before │ After │ Change │
├──────────────────────────┼──────────┼──────────┼─────────────┤
│ Monthly EKS spend │ $6,200 │ $3,800 │ -39% │
│ Pod CPU waste │ 52% │ 14% │ -73% │
│ Pod memory waste │ 48% │ 16% │ -67% │
│ Cluster efficiency score │ 41/100 │ 78/100 │ +90% │
│ Spot instance ratio │ 42% │ 68% │ +62% │
│ Time to detect waste │ Months │ Hours │ — │
│ Monitoring stack cost │ — │ $52/mo │ — │
│ Production incidents │ 0 │ 0 │ No change │
└──────────────────────────┴──────────┴──────────┴─────────────┘
The monitoring stack itself costs $52/month and saves us $2,400/month. That's a 46:1 return.
Table of Contents
- Architecture Overview
- Layer 1: Kubecost for Cost Allocation
- Layer 2: Prometheus + Grafana Dashboards
- Layer 3: CloudWatch Alarms
- Layer 4: Automated Weekly Cost Reports
- The Behavioral Shift
- Results and ROI
- Action Items Checklist
Architecture Overview
The stack has four layers, each serving a different audience:
┌─────────────────────────────────────────────────────────┐
│ COST MONITORING STACK │
│ │
│ ┌──────────┐ ┌────────────┐ ┌──────────────────┐ │
│ │ Kubecost │───▶│ Prometheus │───▶│ Grafana │ │
│ │ │ │ │ │ (3 dashboards) │ │
│ └──────────┘ └─────┬──────┘ └──────────────────┘ │
│ │ │
│ ▼ │
│ ┌────────────────┐ ┌────────────────┐ │
│ │ CloudWatch │───▶│ SNS → Slack │ │
│ │ (3 alarms) │ │ #cloud-costs │ │
│ └────────────────┘ └────────────────┘ │
│ │
│ ┌────────────────┐ ┌────────────────┐ │
│ │ Weekly Report │───▶│ Slack Bot │ │
│ │ (CronJob) │ │ #cloud-costs │ │
│ └────────────────┘ └────────────────┘ │
└─────────────────────────────────────────────────────────┘
- Kubecost scrapes pod-level resource usage and maps it to actual AWS costs via the Cost and Usage Report (CUR). Engineers self-serve their own team's spend.
- Prometheus stores cost metrics and powers custom recording rules for efficiency scoring.
- Grafana gives us three dashboards: cluster overview, namespace detail, and efficiency trends over time.
- CloudWatch alarms fire when efficiency drops below 60%, a namespace exceeds its budget, or spot ratio falls under 50%.
- Weekly Slack report posts a cost breakdown by team every Monday morning with week-over-week trends.
Layer 1: Kubecost for Cost Allocation
Kubecost is the foundation. Without per-team cost allocation, everything else is just graphs nobody looks at.
Helm Deployment
We run Kubecost open-source (not enterprise) with Athena integration for accurate AWS pricing:
# kubecost-values.yaml
kubecostProductConfigs:
clusterName: "prod-eks-cluster"
currencyCode: "USD"
# CUR integration via Athena — this is what makes costs accurate
athenaProjectID: "123456789012"
athenaBucketName: "s3://our-cur-reports/athena-results"
athenaRegion: "us-east-1"
athenaDatabase: "cur_database"
athenaTable: "cost_and_usage_report"
athenaWorkgroup: "kubecost"
# Custom pricing for negotiated rates (we have an EDP)
customPricesEnabled: true
defaultModelPricing:
CPU: 0.031611
RAM: 0.004237
storage: 0.00013889
spotCPU: 0.010000
spotRAM: 0.001340
# Shared cost allocation
sharedNamespaces: "kube-system,monitoring,istio-system"
shareNamespaces: true
sharedOverhead: 280 # $280/month for control plane + NAT gateways
kubecostModel:
# Enable Prometheus integration
promClusterIDLabel: "cluster_id"
networkCosts:
enabled: true
config:
services:
amazon-web-services: true
prometheus:
server:
# Use our existing Prometheus — don't deploy a second one
enabled: false
nodeExporter:
enabled: false
global:
prometheus:
enabled: true
fqdn: "http://prometheus-server.monitoring.svc.cluster.local:80"
serviceAccount:
create: true
annotations:
eks.amazonaws.com/role-arn: "arn:aws:iam::123456789012:role/kubecost-s3-reader"
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
Deploy it:
helm repo add kubecost https://kubecost.github.io/cost-analyzer/
helm repo update
helm upgrade --install kubecost kubecost/cost-analyzer \
--namespace monitoring \
--create-namespace \
-f kubecost-values.yaml \
--version 2.2.3
Cost Allocation Labels
Labels are what make cost allocation work. Without them, Kubecost just shows you namespace totals. We enforced a standard label schema across every deployment:
# Standard labels — every pod must have these
metadata:
labels:
app.kubernetes.io/name: orders-api
team: platform # Maps to cost center
cost-center: engineering # Rolls up to department
tier: critical # critical | important | flexible
environment: production
We already had an OPA Gatekeeper constraint enforcing these labels (from post #7), so adoption was automatic.
Kubecost API for Programmatic Access
The Kubecost API is how we pull data into our weekly reports and custom tooling:
# Per-namespace costs for the last 7 days
curl -s "http://kubecost.monitoring:9090/model/allocation?window=7d&aggregate=namespace" \
| jq '.data[0] | to_entries[] | {namespace: .key, totalCost: .value.totalCost}'
# Per-team costs (using label aggregation)
curl -s "http://kubecost.monitoring:9090/model/allocation?window=7d&aggregate=label:team" \
| jq '.data[0] | to_entries[] | {team: .key, totalCost: .value.totalCost}'
# Efficiency by namespace
curl -s "http://kubecost.monitoring:9090/model/allocation?window=7d&aggregate=namespace" \
| jq '.data[0] | to_entries[] | {
namespace: .key,
cpuEfficiency: .value.cpuEfficiency,
ramEfficiency: .value.ramEfficiency,
totalEfficiency: .value.totalEfficiency
}'
Sample output that changed everything — when the payments team saw their row, they right-sized within 48 hours:
┌───────────────┬───────────┬────────────┬────────────┬─────────────┐
│ Team │ Weekly $ │ CPU Eff. │ RAM Eff. │ Total Eff. │
├───────────────┼───────────┼────────────┼────────────┼─────────────┤
│ platform │ $312.40 │ 71% │ 68% │ 69% │
│ payments │ $487.20 │ 22% │ 19% │ 20% │ ← 😬
│ data-pipeline │ $198.60 │ 65% │ 58% │ 61% │
│ frontend │ $89.30 │ 74% │ 72% │ 73% │
│ ml-inference │ $156.80 │ 48% │ 42% │ 45% │
│ shared │ $64.70 │ — │ — │ — │
├───────────────┼───────────┼────────────┼────────────┼─────────────┤
│ Total │ $1,309.00 │ 52% │ 48% │ 50% │
└───────────────┴───────────┴────────────┴────────────┴─────────────┘
Layer 2: Prometheus + Grafana Dashboards
Kubecost gives us the raw numbers. Prometheus and Grafana turn those numbers into trends that drive action.
Prometheus Recording Rules
Recording rules pre-compute expensive queries so dashboards load fast and we don't hammer Prometheus:
# prometheus-cost-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: cost-recording-rules
namespace: monitoring
labels:
release: prometheus
spec:
groups:
- name: cost.efficiency
interval: 5m
rules:
# CPU efficiency per namespace
- record: namespace:cpu_efficiency:ratio
expr: |
sum by (namespace) (
rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
)
/
sum by (namespace) (
kube_pod_container_resource_requests{resource="cpu", container!=""}
)
# Memory efficiency per namespace
- record: namespace:memory_efficiency:ratio
expr: |
sum by (namespace) (
container_memory_working_set_bytes{container!="", container!="POD"}
)
/
sum by (namespace) (
kube_pod_container_resource_requests{resource="memory", container!=""}
)
# Overall cluster efficiency score (0-100)
- record: cluster:efficiency_score:gauge
expr: |
(
(
sum(rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m]))
/
sum(kube_pod_container_resource_requests{resource="cpu", container!=""})
)
+
(
sum(container_memory_working_set_bytes{container!="", container!="POD"})
/
sum(kube_pod_container_resource_requests{resource="memory", container!=""})
)
) / 2 * 100
# Cost per namespace (estimated from resource usage + node costs)
- record: namespace:estimated_hourly_cost:gauge
expr: |
(
sum by (namespace) (
kube_pod_container_resource_requests{resource="cpu", container!=""}
) * 0.031611
+
sum by (namespace) (
kube_pod_container_resource_requests{resource="memory", container!=""}
) / 1073741824 * 0.004237
)
# Idle resources per namespace (requested but unused)
- record: namespace:idle_cpu_cores:gauge
expr: |
sum by (namespace) (
kube_pod_container_resource_requests{resource="cpu", container!=""}
)
-
sum by (namespace) (
rate(container_cpu_usage_seconds_total{container!="", container!="POD"}[5m])
)
# Spot instance ratio
- record: cluster:spot_ratio:gauge
expr: |
count(kube_node_labels{label_karpenter_sh_capacity_type="spot"})
/
count(kube_node_labels)
# Node count by capacity type
- record: cluster:node_count_by_type:gauge
expr: |
count by (label_karpenter_sh_capacity_type) (kube_node_labels)
Apply it:
kubectl apply -f prometheus-cost-rules.yaml
Grafana Dashboards
We built three dashboards. Here's the JSON for each — import them via Grafana's dashboard import or provision them with a ConfigMap.
Dashboard 1: Cluster Cost Overview
This is the executive view — one screen that tells you if the cluster is healthy or bleeding money.
{
"dashboard": {
"title": "EKS Cost Overview",
"uid": "eks-cost-overview",
"tags": ["cost", "eks", "overview"],
"timezone": "browser",
"refresh": "5m",
"panels": [
{
"title": "Cluster Efficiency Score",
"type": "gauge",
"gridPos": { "h": 8, "w": 6, "x": 0, "y": 0 },
"targets": [
{
"expr": "cluster:efficiency_score:gauge",
"legendFormat": "Efficiency"
}
],
"fieldConfig": {
"defaults": {
"min": 0,
"max": 100,
"unit": "percent",
"thresholds": {
"steps": [
{ "color": "red", "value": 0 },
{ "color": "orange", "value": 40 },
{ "color": "yellow", "value": 60 },
{ "color": "green", "value": 75 }
]
}
}
}
},
{
"title": "Estimated Monthly Cost",
"type": "stat",
"gridPos": { "h": 8, "w": 6, "x": 6, "y": 0 },
"targets": [
{
"expr": "sum(namespace:estimated_hourly_cost:gauge) * 730",
"legendFormat": "Monthly Cost"
}
],
"fieldConfig": {
"defaults": {
"unit": "currencyUSD",
"thresholds": {
"steps": [
{ "color": "green", "value": 0 },
{ "color": "yellow", "value": 4000 },
{ "color": "red", "value": 5000 }
]
}
}
}
},
{
"title": "Spot Instance Ratio",
"type": "gauge",
"gridPos": { "h": 8, "w": 6, "x": 12, "y": 0 },
"targets": [
{
"expr": "cluster:spot_ratio:gauge * 100",
"legendFormat": "Spot %"
}
],
"fieldConfig": {
"defaults": {
"min": 0,
"max": 100,
"unit": "percent",
"thresholds": {
"steps": [
{ "color": "red", "value": 0 },
{ "color": "yellow", "value": 40 },
{ "color": "green", "value": 50 }
]
}
}
}
},
{
"title": "Node Count by Type",
"type": "piechart",
"gridPos": { "h": 8, "w": 6, "x": 18, "y": 0 },
"targets": [
{
"expr": "cluster:node_count_by_type:gauge",
"legendFormat": "{{label_karpenter_sh_capacity_type}}"
}
]
},
{
"title": "Cost by Namespace (Monthly Est.)",
"type": "barchart",
"gridPos": { "h": 10, "w": 12, "x": 0, "y": 8 },
"targets": [
{
"expr": "sort_desc(namespace:estimated_hourly_cost:gauge * 730)",
"legendFormat": "{{namespace}}"
}
],
"fieldConfig": {
"defaults": { "unit": "currencyUSD" }
}
},
{
"title": "Cluster Efficiency Trend (30d)",
"type": "timeseries",
"gridPos": { "h": 10, "w": 12, "x": 12, "y": 8 },
"targets": [
{
"expr": "cluster:efficiency_score:gauge",
"legendFormat": "Efficiency Score"
}
],
"fieldConfig": {
"defaults": {
"unit": "percent",
"custom": {
"fillOpacity": 15,
"lineWidth": 2
}
}
}
}
]
}
}
Dashboard 2: Namespace Cost Detail
The team-level view — every engineer can drill into their own namespace.
{
"dashboard": {
"title": "EKS Namespace Cost Detail",
"uid": "eks-namespace-detail",
"tags": ["cost", "eks", "namespace"],
"timezone": "browser",
"refresh": "5m",
"templating": {
"list": [
{
"name": "namespace",
"type": "query",
"query": "label_values(kube_namespace_labels, namespace)",
"multi": false,
"includeAll": true,
"current": { "text": "All", "value": "$__all" }
}
]
},
"panels": [
{
"title": "CPU: Requests vs Actual Usage",
"type": "timeseries",
"gridPos": { "h": 9, "w": 12, "x": 0, "y": 0 },
"targets": [
{
"expr": "sum(kube_pod_container_resource_requests{resource='cpu', namespace=~'$namespace', container!=''}) ",
"legendFormat": "CPU Requested"
},
{
"expr": "sum(rate(container_cpu_usage_seconds_total{namespace=~'$namespace', container!='', container!='POD'}[5m]))",
"legendFormat": "CPU Actual"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"custom": { "fillOpacity": 10 }
}
}
},
{
"title": "Memory: Requests vs Actual Usage",
"type": "timeseries",
"gridPos": { "h": 9, "w": 12, "x": 12, "y": 0 },
"targets": [
{
"expr": "sum(kube_pod_container_resource_requests{resource='memory', namespace=~'$namespace', container!=''}) / 1073741824",
"legendFormat": "Memory Requested (GiB)"
},
{
"expr": "sum(container_memory_working_set_bytes{namespace=~'$namespace', container!='', container!='POD'}) / 1073741824",
"legendFormat": "Memory Actual (GiB)"
}
],
"fieldConfig": {
"defaults": {
"unit": "decgbytes",
"custom": { "fillOpacity": 10 }
}
}
},
{
"title": "CPU Efficiency",
"type": "gauge",
"gridPos": { "h": 6, "w": 6, "x": 0, "y": 9 },
"targets": [
{
"expr": "namespace:cpu_efficiency:ratio{namespace=~'$namespace'} * 100",
"legendFormat": "{{namespace}}"
}
],
"fieldConfig": {
"defaults": {
"min": 0, "max": 100, "unit": "percent",
"thresholds": {
"steps": [
{ "color": "red", "value": 0 },
{ "color": "orange", "value": 30 },
{ "color": "yellow", "value": 50 },
{ "color": "green", "value": 65 }
]
}
}
}
},
{
"title": "Memory Efficiency",
"type": "gauge",
"gridPos": { "h": 6, "w": 6, "x": 6, "y": 9 },
"targets": [
{
"expr": "namespace:memory_efficiency:ratio{namespace=~'$namespace'} * 100",
"legendFormat": "{{namespace}}"
}
],
"fieldConfig": {
"defaults": {
"min": 0, "max": 100, "unit": "percent",
"thresholds": {
"steps": [
{ "color": "red", "value": 0 },
{ "color": "orange", "value": 30 },
{ "color": "yellow", "value": 50 },
{ "color": "green", "value": 65 }
]
}
}
}
},
{
"title": "Estimated Monthly Cost Trend",
"type": "timeseries",
"gridPos": { "h": 6, "w": 12, "x": 12, "y": 9 },
"targets": [
{
"expr": "namespace:estimated_hourly_cost:gauge{namespace=~'$namespace'} * 730",
"legendFormat": "{{namespace}}"
}
],
"fieldConfig": {
"defaults": {
"unit": "currencyUSD",
"custom": { "fillOpacity": 15 }
}
}
},
{
"title": "Idle CPU (Requested but Unused)",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 15 },
"targets": [
{
"expr": "namespace:idle_cpu_cores:gauge{namespace=~'$namespace'}",
"legendFormat": "{{namespace}} idle cores"
}
],
"fieldConfig": {
"defaults": {
"unit": "short",
"custom": {
"fillOpacity": 20,
"lineWidth": 2
}
}
}
},
{
"title": "Top Resource Consumers (Pods)",
"type": "table",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 15 },
"targets": [
{
"expr": "topk(10, sum by (pod, namespace) (rate(container_cpu_usage_seconds_total{namespace=~'$namespace', container!='', container!='POD'}[5m])))",
"legendFormat": "{{namespace}}/{{pod}}",
"format": "table",
"instant": true
}
]
}
]
}
}
Dashboard 3: Efficiency Trends
The long-term view — are we getting better or worse over time?
{
"dashboard": {
"title": "EKS Efficiency Trends",
"uid": "eks-efficiency-trends",
"tags": ["cost", "eks", "trends"],
"timezone": "browser",
"refresh": "1h",
"time": { "from": "now-90d", "to": "now" },
"panels": [
{
"title": "Cluster Efficiency Score (90d)",
"type": "timeseries",
"gridPos": { "h": 8, "w": 24, "x": 0, "y": 0 },
"targets": [
{
"expr": "avg_over_time(cluster:efficiency_score:gauge[1d])",
"legendFormat": "Daily Avg Efficiency",
"interval": "1d"
}
],
"fieldConfig": {
"defaults": {
"unit": "percent",
"min": 0, "max": 100,
"custom": {
"fillOpacity": 20,
"lineWidth": 2,
"thresholdsStyle": { "mode": "line" }
},
"thresholds": {
"steps": [
{ "color": "red", "value": 40 },
{ "color": "green", "value": 60 }
]
}
}
}
},
{
"title": "Monthly Cost Trend",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 0, "y": 8 },
"targets": [
{
"expr": "sum(namespace:estimated_hourly_cost:gauge) * 730",
"legendFormat": "Est. Monthly Cost",
"interval": "1d"
}
],
"fieldConfig": {
"defaults": {
"unit": "currencyUSD",
"custom": {
"fillOpacity": 15,
"lineWidth": 2,
"thresholdsStyle": { "mode": "area" }
},
"thresholds": {
"steps": [
{ "color": "green", "value": 0 },
{ "color": "yellow", "value": 4000 },
{ "color": "red", "value": 5000 }
]
}
}
}
},
{
"title": "Spot vs On-Demand Ratio (90d)",
"type": "timeseries",
"gridPos": { "h": 8, "w": 12, "x": 12, "y": 8 },
"targets": [
{
"expr": "cluster:spot_ratio:gauge * 100",
"legendFormat": "Spot %"
},
{
"expr": "(1 - cluster:spot_ratio:gauge) * 100",
"legendFormat": "On-Demand %"
}
],
"fieldConfig": {
"defaults": {
"unit": "percent",
"custom": {
"fillOpacity": 30,
"stacking": { "mode": "normal" }
}
}
}
},
{
"title": "Efficiency by Namespace (30d)",
"type": "timeseries",
"gridPos": { "h": 8, "w": 24, "x": 0, "y": 16 },
"targets": [
{
"expr": "(namespace:cpu_efficiency:ratio + namespace:memory_efficiency:ratio) / 2 * 100",
"legendFormat": "{{namespace}}"
}
],
"fieldConfig": {
"defaults": {
"unit": "percent",
"custom": { "lineWidth": 2 }
}
}
}
]
}
}
Provision all three as ConfigMaps so they survive Grafana restarts:
# grafana-dashboards-configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: grafana-cost-dashboards
namespace: monitoring
labels:
grafana_dashboard: "true"
data:
eks-cost-overview.json: |
# (paste Dashboard 1 JSON above)
eks-namespace-detail.json: |
# (paste Dashboard 2 JSON above)
eks-efficiency-trends.json: |
# (paste Dashboard 3 JSON above)
Make sure your Grafana Helm values include the sidecar that picks up dashboard ConfigMaps:
# grafana-values.yaml (relevant section)
sidecar:
dashboards:
enabled: true
label: grafana_dashboard
searchNamespace: monitoring
Layer 3: CloudWatch Alarms
Dashboards are passive — people have to look at them. Alarms are active — they come to you. We set up three alarms that catch the most common failure modes.
Alarm 1: Cluster Efficiency Below 60%
If the cluster drops below 60% efficiency for two consecutive hours, something is wrong — either a team deployed over-provisioned pods or Karpenter isn't consolidating.
Alarm 2: Namespace Budget Exceeded
Each namespace has a monthly budget. If projected spend exceeds it, the owning team gets alerted.
Alarm 3: Spot Ratio Below 50%
We target 65% spot. If it drops below 50%, we're paying too much for on-demand capacity — usually means a Karpenter NodePool misconfiguration or an availability event.
Terraform Module
Here's the complete Terraform module that deploys all three alarms with SNS and Slack:
# modules/eks-cost-alarms/variables.tf
variable "cluster_name" {
type = string
description = "EKS cluster name"
}
variable "efficiency_threshold" {
type = number
default = 60
description = "Minimum cluster efficiency percentage"
}
variable "spot_ratio_threshold" {
type = number
default = 50
description = "Minimum spot instance percentage"
}
variable "namespace_budgets" {
type = map(number)
default = {
"payments" = 1600
"platform" = 1400
"data-pipeline" = 900
"ml-inference" = 700
"frontend" = 400
}
description = "Monthly budget per namespace in USD"
}
variable "slack_webhook_url" {
type = string
sensitive = true
description = "Slack incoming webhook URL for #cloud-costs channel"
}
# modules/eks-cost-alarms/main.tf
resource "aws_sns_topic" "cost_alerts" {
name = "${var.cluster_name}-cost-alerts"
tags = {
Environment = "production"
ManagedBy = "terraform"
Purpose = "eks-cost-monitoring"
}
}
resource "aws_sns_topic_subscription" "slack" {
topic_arn = aws_sns_topic.cost_alerts.arn
protocol = "lambda"
endpoint = aws_lambda_function.slack_notifier.arn
}
# Lambda function to format SNS → Slack messages
resource "aws_lambda_function" "slack_notifier" {
filename = data.archive_file.slack_notifier.output_path
function_name = "${var.cluster_name}-cost-slack-notifier"
role = aws_iam_role.slack_notifier.arn
handler = "index.handler"
runtime = "python3.12"
source_code_hash = data.archive_file.slack_notifier.output_base64sha256
timeout = 10
environment {
variables = {
SLACK_WEBHOOK_URL = var.slack_webhook_url
CLUSTER_NAME = var.cluster_name
}
}
}
data "archive_file" "slack_notifier" {
type = "zip"
source_file = "${path.module}/lambda/index.py"
output_path = "${path.module}/lambda/slack_notifier.zip"
}
resource "aws_iam_role" "slack_notifier" {
name = "${var.cluster_name}-cost-slack-notifier"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Action = "sts:AssumeRole"
Effect = "Allow"
Principal = {
Service = "lambda.amazonaws.com"
}
}]
})
}
resource "aws_iam_role_policy_attachment" "slack_notifier_basic" {
role = aws_iam_role.slack_notifier.name
policy_arn = "arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole"
}
resource "aws_lambda_permission" "sns_invoke" {
statement_id = "AllowSNSInvoke"
action = "lambda:InvokeFunction"
function_name = aws_lambda_function.slack_notifier.function_name
principal = "sns.amazonaws.com"
source_arn = aws_sns_topic.cost_alerts.arn
}
# Alarm 1: Cluster efficiency below threshold
resource "aws_cloudwatch_metric_alarm" "cluster_efficiency" {
alarm_name = "${var.cluster_name}-low-efficiency"
alarm_description = "Cluster efficiency dropped below ${var.efficiency_threshold}% for 2 consecutive hours"
comparison_operator = "LessThanThreshold"
evaluation_periods = 2
period = 3600
threshold = var.efficiency_threshold
statistic = "Average"
namespace = "EKS/CostOptimization"
metric_name = "ClusterEfficiencyScore"
treat_missing_data = "breaching"
dimensions = {
ClusterName = var.cluster_name
}
alarm_actions = [aws_sns_topic.cost_alerts.arn]
ok_actions = [aws_sns_topic.cost_alerts.arn]
tags = {
Environment = "production"
AlertType = "cost-efficiency"
}
}
# Alarm 2: Namespace budget exceeded (one per namespace)
resource "aws_cloudwatch_metric_alarm" "namespace_budget" {
for_each = var.namespace_budgets
alarm_name = "${var.cluster_name}-${each.key}-budget-exceeded"
alarm_description = "Namespace ${each.key} projected spend exceeds monthly budget of $${each.value}"
comparison_operator = "GreaterThanThreshold"
evaluation_periods = 1
period = 86400
threshold = each.value
statistic = "Maximum"
namespace = "EKS/CostOptimization"
metric_name = "NamespaceProjectedMonthlyCost"
treat_missing_data = "notBreaching"
dimensions = {
ClusterName = var.cluster_name
Namespace = each.key
}
alarm_actions = [aws_sns_topic.cost_alerts.arn]
tags = {
Environment = "production"
AlertType = "cost-budget"
Team = each.key
}
}
# Alarm 3: Spot ratio below threshold
resource "aws_cloudwatch_metric_alarm" "spot_ratio" {
alarm_name = "${var.cluster_name}-low-spot-ratio"
alarm_description = "Spot instance ratio dropped below ${var.spot_ratio_threshold}%"
comparison_operator = "LessThanThreshold"
evaluation_periods = 3
period = 1800
threshold = var.spot_ratio_threshold
statistic = "Average"
namespace = "EKS/CostOptimization"
metric_name = "SpotInstancePercentage"
treat_missing_data = "breaching"
dimensions = {
ClusterName = var.cluster_name
}
alarm_actions = [aws_sns_topic.cost_alerts.arn]
ok_actions = [aws_sns_topic.cost_alerts.arn]
tags = {
Environment = "production"
AlertType = "cost-spot-ratio"
}
}
# modules/eks-cost-alarms/outputs.tf
output "sns_topic_arn" {
value = aws_sns_topic.cost_alerts.arn
description = "ARN of the cost alerts SNS topic"
}
output "alarm_arns" {
value = {
efficiency = aws_cloudwatch_metric_alarm.cluster_efficiency.arn
spot_ratio = aws_cloudwatch_metric_alarm.spot_ratio.arn
budgets = { for k, v in aws_cloudwatch_metric_alarm.namespace_budget : k => v.arn }
}
description = "ARNs of all cost alarms"
}
Lambda: SNS to Slack Formatter
The Lambda formats CloudWatch alarm messages into readable Slack blocks:
# modules/eks-cost-alarms/lambda/index.py
import json
import os
from urllib import request, error
SLACK_WEBHOOK_URL = os.environ["SLACK_WEBHOOK_URL"]
CLUSTER_NAME = os.environ["CLUSTER_NAME"]
ALARM_EMOJIS = {
"ALARM": ":rotating_light:",
"OK": ":white_check_mark:",
"INSUFFICIENT_DATA": ":warning:",
}
def handler(event, context):
for record in event["Records"]:
message = json.loads(record["Sns"]["Message"])
alarm_name = message.get("AlarmName", "Unknown")
state = message.get("NewStateValue", "UNKNOWN")
reason = message.get("NewStateReason", "")
emoji = ALARM_EMOJIS.get(state, ":question:")
color = "#cc0000" if state == "ALARM" else "#36a64f"
slack_payload = {
"channel": "#cloud-costs",
"username": f"EKS Cost Monitor ({CLUSTER_NAME})",
"icon_emoji": ":chart_with_downwards_trend:",
"attachments": [
{
"color": color,
"blocks": [
{
"type": "header",
"text": {
"type": "plain_text",
"text": f"{emoji} {alarm_name}",
},
},
{
"type": "section",
"fields": [
{
"type": "mrkdwn",
"text": f"*State:*\n{state}",
},
{
"type": "mrkdwn",
"text": f"*Cluster:*\n{CLUSTER_NAME}",
},
],
},
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": f"*Reason:*\n{reason}",
},
},
],
}
],
}
req = request.Request(
SLACK_WEBHOOK_URL,
data=json.dumps(slack_payload).encode("utf-8"),
headers={"Content-Type": "application/json"},
method="POST",
)
try:
request.urlopen(req)
except error.HTTPError as e:
print(f"Slack webhook failed: {e.code} {e.read().decode()}")
raise
return {"statusCode": 200}
Use the module:
# environments/production/cost-monitoring.tf
module "eks_cost_alarms" {
source = "../../modules/eks-cost-alarms"
cluster_name = "prod-eks-cluster"
slack_webhook_url = var.slack_webhook_url
efficiency_threshold = 60
spot_ratio_threshold = 50
namespace_budgets = {
"payments" = 1600
"platform" = 1400
"data-pipeline" = 900
"ml-inference" = 700
"frontend" = 400
}
}
Layer 4: Automated Weekly Cost Reports
This is the piece that changed team behavior the most. Every Monday at 9 AM, a CronJob posts a cost report to #cloud-costs in Slack. It shows each team's spend, week-over-week change with trend arrows, and calls out the top 3 wasters and top 3 improvers by name.
This is the expanded version of the cost_report.py script from post #7, now with week-over-week comparison and Slack formatting:
#!/usr/bin/env python3
"""
Weekly EKS cost report — posts to Slack every Monday.
Pulls data from Kubecost API, compares week-over-week, and highlights
top wasters and improvers.
"""
import json
import os
from datetime import datetime
from urllib import request, error, parse
KUBECOST_ENDPOINT = os.environ.get(
"KUBECOST_ENDPOINT", "http://kubecost-cost-analyzer.monitoring:9090"
)
SLACK_WEBHOOK_URL = os.environ["SLACK_WEBHOOK_URL"]
CLUSTER_NAME = os.environ.get("CLUSTER_NAME", "prod-eks-cluster")
EXCLUDED_NAMESPACES = {"kube-system", "monitoring", "istio-system", "kube-node-lease"}
def fetch_kubecost_data(window):
"""Fetch cost allocation data from Kubecost API."""
url = (
f"{KUBECOST_ENDPOINT}/model/allocation"
f"?window={window}&aggregate=label:team&accumulate=true"
)
req = request.Request(url)
with request.urlopen(req, timeout=30) as resp:
return json.loads(resp.read().decode())
def get_team_costs(window):
"""Extract per-team costs from Kubecost response."""
data = fetch_kubecost_data(window)
teams = {}
for entry in data.get("data", []):
for team_name, metrics in entry.items():
if team_name in ("__idle__", "__unallocated__"):
continue
teams[team_name] = {
"total": round(metrics.get("totalCost", 0), 2),
"cpu_cost": round(metrics.get("cpuCost", 0), 2),
"ram_cost": round(metrics.get("ramCost", 0), 2),
"cpu_eff": round(metrics.get("cpuEfficiency", 0) * 100, 1),
"ram_eff": round(metrics.get("ramEfficiency", 0) * 100, 1),
}
return teams
def build_report():
"""Build the weekly cost comparison report."""
current = get_team_costs("7d")
previous = get_team_costs("14d,7d") # Previous 7-day window
rows = []
for team in sorted(current.keys()):
curr_cost = current[team]["total"]
prev_cost = previous.get(team, {}).get("total", curr_cost)
cpu_eff = current[team]["cpu_eff"]
ram_eff = current[team]["ram_eff"]
if prev_cost > 0:
change_pct = ((curr_cost - prev_cost) / prev_cost) * 100
else:
change_pct = 0
if change_pct > 5:
trend = ":arrow_up_small: :red_circle:"
elif change_pct < -5:
trend = ":arrow_down_small: :large_green_circle:"
else:
trend = ":left_right_arrow:"
rows.append({
"team": team,
"current": curr_cost,
"previous": prev_cost,
"change_pct": change_pct,
"trend": trend,
"cpu_eff": cpu_eff,
"ram_eff": ram_eff,
})
# Sort by absolute change to find biggest movers
rows.sort(key=lambda r: r["change_pct"], reverse=True)
wasters = [r for r in rows if r["change_pct"] > 5][:3]
improvers = [r for r in rows if r["change_pct"] < -5][:3]
# Sort final table by cost descending
rows.sort(key=lambda r: r["current"], reverse=True)
total_current = sum(r["current"] for r in rows)
total_previous = sum(r.get("previous", 0) for r in rows)
total_change = ((total_current - total_previous) / total_previous * 100) if total_previous else 0
return rows, wasters, improvers, total_current, total_change
def format_slack_message(rows, wasters, improvers, total, total_change):
"""Format the report as Slack blocks."""
today = datetime.now().strftime("%B %d, %Y")
trend_emoji = ":large_green_circle:" if total_change <= 0 else ":red_circle:"
# Build the cost table
table_lines = ["```
"]
table_lines.append(f"{'Team':<16} {'This Week':>10} {'Last Week':>10} {'Change':>8} {'CPU%':>5} {'RAM%':>5}")
table_lines.append("─" * 68)
for r in rows:
change_str = f"{r['change_pct']:+.1f}%"
table_lines.append(
f"{r['team']:<16} ${r['current']:>8,.2f} ${r['previous']:>8,.2f} "
f"{change_str:>8} {r['cpu_eff']:>4.0f}% {r['ram_eff']:>4.0f}%"
)
table_lines.append("─" * 68)
table_lines.append(
f"{'TOTAL':<16} ${total:>8,.2f} {total_change:+.1f}%"
)
table_lines.append("
```")
blocks = [
{
"type": "header",
"text": {
"type": "plain_text",
"text": f":bar_chart: Weekly EKS Cost Report — {today}",
},
},
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": (
f"*Cluster:* `{CLUSTER_NAME}` | "
f"*Weekly Spend:* ${total:,.2f} ({total_change:+.1f}% WoW) {trend_emoji}"
),
},
},
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": "\n".join(table_lines),
},
},
]
# Top wasters callout
if wasters:
waster_lines = [f"• *{w['team']}*: +{w['change_pct']:.1f}% (${w['current']:,.2f})" for w in wasters]
blocks.append({
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":rotating_light: *Top Wasters (cost increase)*\n" + "\n".join(waster_lines),
},
})
# Top improvers callout
if improvers:
improver_lines = [f"• *{i['team']}*: {i['change_pct']:.1f}% (${i['current']:,.2f})" for i in improvers]
blocks.append({
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":trophy: *Top Improvers (cost reduction)*\n" + "\n".join(improver_lines),
},
})
blocks.append({
"type": "context",
"elements": [
{
"type": "mrkdwn",
"text": (
f"Data source: Kubecost | Cluster: {CLUSTER_NAME} | "
"<https://grafana.internal/d/eks-cost-overview|View Dashboard>"
),
}
],
})
return blocks
def post_to_slack(blocks):
"""Send the formatted report to Slack."""
payload = {
"channel": "#cloud-costs",
"username": f"EKS Cost Report ({CLUSTER_NAME})",
"icon_emoji": ":bar_chart:",
"blocks": blocks,
}
req = request.Request(
SLACK_WEBHOOK_URL,
data=json.dumps(payload).encode("utf-8"),
headers={"Content-Type": "application/json"},
method="POST",
)
try:
request.urlopen(req)
print("Report posted to #cloud-costs")
except error.HTTPError as e:
print(f"Slack post failed: {e.code} {e.read().decode()}")
raise
def main():
rows, wasters, improvers, total, total_change = build_report()
blocks = format_slack_message(rows, wasters, improvers, total, total_change)
post_to_slack(blocks)
if __name__ == "__main__":
main()
Deploy it as a Kubernetes CronJob:
# weekly-cost-report-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: weekly-cost-report
namespace: monitoring
spec:
schedule: "0 9 * * 1" # Every Monday at 9 AM UTC
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 4
failedJobsHistoryLimit: 2
jobTemplate:
spec:
backoffLimit: 2
activeDeadlineSeconds: 300
template:
metadata:
labels:
app: weekly-cost-report
team: platform
cost-center: engineering
spec:
restartPolicy: OnFailure
serviceAccountName: cost-reporter
containers:
- name: report
image: python:3.12-slim
command: ["python3", "/scripts/weekly_cost_report.py"]
env:
- name: KUBECOST_ENDPOINT
value: "http://kubecost-cost-analyzer.monitoring:9090"
- name: CLUSTER_NAME
value: "prod-eks-cluster"
- name: SLACK_WEBHOOK_URL
valueFrom:
secretKeyRef:
name: slack-webhook
key: url
volumeMounts:
- name: scripts
mountPath: /scripts
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
cpu: 200m
memory: 128Mi
volumes:
- name: scripts
configMap:
name: weekly-cost-report-script
---
apiVersion: v1
kind: ConfigMap
metadata:
name: weekly-cost-report-script
namespace: monitoring
data:
weekly_cost_report.py: |
# (paste the Python script above)
---
apiVersion: v1
kind: Secret
metadata:
name: slack-webhook
namespace: monitoring
type: Opaque
stringData:
url: "https://hooks.slack.com/services/YOUR/WEBHOOK/URL"
Here's what the Slack report looks like every Monday:
┌──────────────────────────────────────────────────────────────────┐
│ 📊 Weekly EKS Cost Report — September 30, 2026 │
│ │
│ Cluster: prod-eks-cluster │
│ Weekly Spend: $878.50 (-3.2% WoW) 🟢 │
│ │
│ Team This Week Last Week Change CPU% RAM% │
│ ────────────────────────────────────────────────────────────── │
│ payments $298.40 $387.20 -22.9% 64% 61% │
│ platform $274.80 $268.10 +2.5% 73% 70% │
│ data-pipeline $148.20 $152.40 -2.8% 67% 62% │
│ ml-inference $102.30 $96.80 +5.7% 52% 48% │
│ frontend $54.80 $53.10 +3.2% 76% 74% │
│ ────────────────────────────────────────────────────────────── │
│ TOTAL $878.50 $907.60 -3.2% │
│ │
│ 🏆 Top Improvers │
│ • payments: -22.9% ($298.40) — right-sized after last report │
│ │
│ 🚨 Top Wasters │
│ • ml-inference: +5.7% ($102.30) — new model deployment │
│ │
│ Data source: Kubecost | View Dashboard │
└──────────────────────────────────────────────────────────────────┘
The Behavioral Shift
The technical stack was the easy part. The hard part was getting 40 engineers to actually care about resource costs. Here's what moved the needle:
Making It Personal
Before the monitoring stack, cost was an ops problem. Platform team owned the bill, nobody else thought about it. The weekly report changed that by putting a number next to every team name. When the payments team saw they were spending $487/week at 20% efficiency — in front of every other team — they right-sized within 48 hours without anyone asking them to.
The Cost Champion Program
We assigned one "cost champion" per team — a senior engineer who got view access to the Grafana dashboards and was CC'd on budget alarms. Their job wasn't enforcement. It was awareness. They'd mention cost in sprint planning: "hey, our spend jumped 15% this week, probably that new worker deployment — can we review the resource requests?"
The champions met monthly for 30 minutes to share what worked. The data-pipeline team's champion shared a trick for right-sizing batch jobs that the ML team adopted and saved $40/week. This kind of cross-pollination never happened when costs were invisible.
Monthly Cost Review
We added a 15-minute cost review to the existing monthly architecture review. One slide: the efficiency trends dashboard screenshotted from Grafana. No finger-pointing — just "here's where we are, here's the trend." When the trend line was going up, teams self-corrected before the next meeting. Nobody wanted to be the team that broke the streak.
What Didn't Work
- Automated enforcement (rejecting deployments above a cost threshold): Engineers gamed it by splitting workloads across namespaces. We scrapped this after two weeks.
- Daily reports: Too frequent. People stopped reading them by day 3. Weekly was the sweet spot — frequent enough to catch problems, infrequent enough to feel meaningful.
- Shame-based messaging: The first version of the report highlighted "worst teams." We changed it to "top improvers" within a week. Positive reinforcement drove 3x more action than negative.
Results and ROI
Six months after deploying the monitoring stack, here's where we landed:
┌──────────────────────────────────────────────────────────────────┐
│ 6-MONTH RESULTS │
├──────────────────────────────┬──────────┬──────────┬─────────────┤
│ Metric │ Month 0 │ Month 6 │ Change │
├──────────────────────────────┼──────────┼──────────┼─────────────┤
│ Monthly EKS spend │ $6,200 │ $3,800 │ -39% │
│ Waste (requested - used) │ 52% │ 14% │ -73% │
│ Teams self-correcting │ 0 / 5 │ 5 / 5 │ 100% │
│ Avg time to detect waste │ ~90 days │ < 24 hrs │ -99% │
│ VPA recommendations adopted │ 12% │ 78% │ +550% │
│ Spot instance ratio │ 42% │ 68% │ +62% │
│ Budget alarm triggers │ — │ 7 │ All resolved│
│ Production incidents │ 0 │ 0 │ No change │
└──────────────────────────────┴──────────┴──────────┴─────────────┘
Savings Breakdown
How the $2,400/month savings broke down:
Team self-corrections (post-report) $1,100/mo (46%)
├── Payments team right-sizing $480
├── ML team batch job optimization $340
└── Data pipeline schedule tuning $280
Alarm-triggered fixes $720/mo (30%)
├── Spot ratio recovery (2 incidents) $380
├── Namespace budget catches $220
└── Efficiency drops caught early $120
Proactive optimization (dashboards) $580/mo (24%)
├── Identified zombie deployments $310
└── Right-sized based on trend data $270
─────────────────────────────────────────────────────
Total monthly savings $2,400/mo
Monitoring stack cost $52/mo
Net savings $2,348/mo
ROI 46:1
Monitoring Stack Cost Breakdown
Kubecost (open-source) $0/mo
Prometheus (additional storage for metrics) $18/mo (50GB retention)
Grafana (t3.small instance) $15/mo
CloudWatch alarms (3 alarms) $3/mo
Lambda invocations (~100/mo) $0/mo (free tier)
CronJob resources (50m CPU, 64Mi) $16/mo
─────────────────────────────────────────────────
Total $52/mo
Combined Program Savings
Across both optimization posts (post #7 and this one):
Original EKS spend (pre-optimization) $14,200/mo
After post #7 optimizations $4,180/mo (-71%)
Cost creep after 3 months $6,200/mo (waste returning)
After monitoring stack (stabilized) $3,800/mo (-73% from original)
─────────────────────────────────────────────────────────────
Total annual savings $124,800/yr
Total investment (tooling + labor) $21,100
Payback period ~2 months
First-year ROI 492%
The monitoring stack didn't just recover the savings from post #7 — it actually pushed costs below our original optimization target because teams found efficiencies we'd missed.
Action Items Checklist
Week 1: Foundation
- [ ] Deploy Kubecost with Athena/CUR integration
- [ ] Verify per-namespace cost data is accurate (compare against AWS bill)
- [ ] Enforce cost allocation labels via OPA Gatekeeper (see post #7)
- [ ] Test Kubecost API endpoints manually
Week 2: Visibility
- [ ] Deploy Prometheus recording rules for cost metrics
- [ ] Import Grafana dashboards (cluster overview, namespace detail, efficiency trends)
- [ ] Verify dashboards populate correctly with live data
- [ ] Share dashboard URLs with engineering leads
Week 3: Alerts
- [ ] Deploy Terraform alarm module (efficiency, budget, spot ratio)
- [ ] Deploy SNS → Lambda → Slack integration
- [ ] Test each alarm by temporarily lowering thresholds
- [ ] Set per-namespace budgets based on current baseline + 10% buffer
Week 4: Reports + Culture
- [ ] Deploy weekly cost report CronJob
- [ ] Verify Slack report posts correctly on Monday
- [ ] Assign cost champions per team (one senior engineer each)
- [ ] Schedule first monthly cost review (15 min in existing architecture review)
Ongoing
- [ ] Review and adjust namespace budgets quarterly
- [ ] Update Kubecost custom pricing when AWS rates change
- [ ] Review alarm thresholds monthly (tighten as teams improve)
- [ ] Archive weekly reports for trend analysis
Conclusion
We spent $52/month on monitoring tools and got $2,400/month in savings — not from automation, but from making costs visible to the people creating them. The payments team didn't need a policy to right-size their pods. They needed a Slack message showing they were spending 2.5x more than every other team at 20% efficiency.
The four layers work together: Kubecost for accurate cost data, Prometheus and Grafana for trend visibility, CloudWatch alarms for automated detection, and weekly Slack reports for accountability. Skip any layer and the system doesn't drive behavior change.
If you've already optimized your EKS cluster (right-sizing, Karpenter, spot instances), the monitoring stack is what prevents regression. Without it, you'll be re-optimizing the same cluster every quarter. With it, teams self-correct — and the savings compound.
This is Part 9 of the AWS Cost Optimization Series. Part 7 covers the six optimization strategies this monitoring stack protects. Part 8 covers Lambda vs. Fargate cost analysis.
Top comments (0)