Been staring at six browser tabs, three Grafana instances, and a Slack channel full of alerts trying to figure out if your EKS cluster is healthy? We've all been there. Someone pages you at 2 AM, you open your monitoring setup, and you spend the first ten minutes just figuring out where to look before you even start diagnosing what went wrong.
Let's be real β most EKS monitoring setups evolve organically. Someone adds a CPU widget here, a pod restart counter there, maybe a log query panel that made sense six months ago. Before you know it, you've got a Frankenstein dashboard that tells you everything and nothing at the same time. It's like trying to read a novel where every chapter is from a different book.
The actual problem isn't a lack of data. Kubernetes and AWS services generate mountains of metrics. The problem is operational opinion β deciding what matters at 2 AM versus what's just noise. That's exactly what a proper NOC dashboard encodes. AWS recently published a detailed walkthrough on their Containers Blog showing how to build one for EKS using Amazon CloudWatch, and the approach is honestly worth stealing.
The Monitoring Mess We Keep Building π€
Here's what the typical EKS monitoring story looks like:
- π Too many dashboards, no single pane: Node metrics in one tool, application golden signals in another, AWS backing-service health in a third
- π Blank widgets during healthy state: Your PromQL query returns nothing when there are zero errors, so the panel justβ¦ sits there empty. Is it broken? Is it healthy? Nobody knows.
- π PromQL and CloudWatch living in different worlds: Kubernetes state metrics speak PromQL, but your DynamoDB and ELB metrics live in CloudWatch namespaces β most setups don't combine them
- π¨ Alert fatigue: Every metric gets an alarm, every alarm fires independently, and your on-call engineer's phone buzzes so much they start ignoring it
- π No incident-specific view: When something actually breaks, you need a tight, focused dashboard β not a sprawling 3-hour lookback
Sound familiar? Yeah, thought so.
Enter the CloudWatch NOC Dashboard
Say hello to a single-pane NOC dashboard approach that AWS laid out on their Containers Blog this week. The idea is honestly elegant: instead of one mega-dashboard, you get 4 purpose-built dashboards that serve different audiences and moments in your operational lifecycle.
Here's what the solution actually produces:
| Dashboard | Widgets | Lookback | Purpose |
|---|---|---|---|
retail-store-noc |
35 | 3 hours | Executive rollup plus condensed sections |
retail-store-eks-monitoring |
9 | 3 hours | Infrastructure: nodes, namespaces, control plane |
retail-store-application-health |
11 | 3 hours | Golden signals, dependencies, JVM heap, ingress, backing services |
retail-store-incident-response |
12 | 1 hour | Triage: failure counters, alarms, saturation, log search |
35 widgets on the main NOC. 12 on the incident response board with a tight 1-hour window. That's an operational opinion baked right into the architecture β when you're triaging, you don't need three hours of history cluttering your view.
How the Magic Actually Works π§
The secret sauce is that no single widget type can reach all the signals you need. The NOC combines PromQL chart widgets for Kubernetes state, node/pod resources, and control-plane metrics with CloudWatch metric widgets for Application Signals golden signals and AWS service namespaces like ELB, DynamoDB, RDS, ElastiCache, and MQ.
Think of it like a bilingual translator sitting between two conversations. Your Kubernetes metrics speak Prometheus, your managed AWS services speak CloudWatch, and the dashboard stitches them into one coherent story.
The Zero-Value Trick
This is one of those details that separates a good dashboard from a great one. Every counter on the incident strip uses or vector(0) for PromQL queries and FILL(m,0) for CloudWatch metric widgets. Why? Because when everything is healthy, you want to see an explicit 0 β not a blank panel that makes your NOC operator wonder if the data pipeline is broken.
Simple but powerful. A blank widget is ambiguous. A zero is a statement.
Prerequisites Before You Start
Before we get our hands dirty, here's what you need:
For a standard EKS cluster (non-Auto Mode), you need Kubernetes 1.28 or later and version 6.2.0 or later of the Amazon CloudWatch Observability EKS add-on with OTel Container Insights enabled. You'll also need to configure EKS Pod Identity or IAM roles for service accounts (IRSA), attach CloudWatchAgentServerPolicy to the node IAM role, and allow outbound access to CloudWatch endpoints.
If you're running EKS Auto Mode, the blog post validated all PromQL queries with add-on v6.4.0-eksbuild.1 on EKS Auto Mode running Kubernetes 1.34. Lucky you.
Let's Get Our Hands Dirty π οΈ
Alright, enough talking. Let's actually build this thing.
Step 1: Enable the CloudWatch Observability Add-on
If you haven't already installed it, add the observability add-on to your cluster. Replace the placeholders with your own cluster name and region:
aws eks create-addon \
--cluster-name your-cluster-name \
--addon-name amazon-cloudwatch-observability \
--addon-version v6.4.0-eksbuild.1 \
--configuration-values '{"containerInsights":{"enabled":true}}' \
--region us-east-1
Verify it's running:
aws eks describe-addon \
--cluster-name your-cluster-name \
--addon-name amazon-cloudwatch-observability \
--query "addon.status" \
--output text
You should see ACTIVE. Go grab a coffee if it takes a moment βοΈ.
Step 2: Build the NOC Dashboard with Combined Widget Types
Here's where the PromQL and CloudWatch metric widgets come together. The main NOC dashboard combines both types. Here's a snippet showing the zero-value pattern for a CloudWatch metric widget β notice the FILL(m,0) in the metric math that keeps healthy state reading an explicit zero:
{
"type": "metric",
"properties": {
"metrics": [
[{ "expression": "FILL(m,0)", "id": "e1", "label": "Pod Restarts" }],
["ContainerInsights", "pod_restart_count",
"ClusterName", "your-cluster-name",
{ "id": "m", "visible": false }]
],
"view": "singleValue",
"region": "us-east-1",
"title": "Pod Restarts",
"period": 300
},
"width": 3,
"height": 3
}
For PromQL widgets, you'd append or vector(0) to your query string β same idea, different syntax. That explicit zero is what keeps your NOC operator sane at 3 AM.
Step 3: Set Up Composite Alarms
Individual alarms are noisy. It's about damn time we stopped treating every threshold breach as a crisis. Composite alarms let you encode logic like "only fire if both node CPU is saturated AND pods are failing health checks":
aws cloudwatch put-composite-alarm \
--alarm-name "eks-node-and-pod-health" \
--alarm-rule 'ALARM("eks-node-cpu-high") AND ALARM("eks-pod-restarts-high")' \
--alarm-description "Fires only when node saturation coincides with pod health failures" \
--actions-enabled
This is an operational opinion in code. You're telling CloudWatch: "I don't care about high CPU alone. I care about high CPU plus application impact." That's the difference between a dashboard that alerts and one that actually helps.
Step 4: Watch the Magic Happen β¨
With Container Insights flowing and your dashboards deployed, open the CloudWatch console. Navigate to Dashboards and you should see your four dashboards listed. The main NOC board β all 35 widgets β gives you the executive rollup. Click through to the incident response dashboard and you'll see those tight 1-hour windows with explicit zeros on every healthy counter.
So what just happened behind the scenes? With OTel Container Insights enabled through the add-on, the NOC combines PromQL chart widgets β covering Kubernetes state, node/pod resources, and control-plane metrics β with CloudWatch metric widgets for Application Signals golden signals and AWS service namespaces like ELB, DynamoDB, RDS, ElastiCache, and MQ. Two widget types, two metric languages, one coherent view. Magical, honestly.
When to Use This (and When Not To) π
β Great fit:
- You're running EKS and already paying for CloudWatch β no extra tooling cost
- Your team wants a single-pane NOC without bolting on Grafana or Datadog
- You need to combine Kubernetes metrics with AWS managed service health in one view
- You want composite alarms that encode operational opinions, not just threshold breaches
β Probably not for you:
- You're deep in the Grafana/Prometheus ecosystem and don't want to migrate dashboards
- Your clusters span multiple cloud providers β CloudWatch is AWS-only
- You need custom long-term metric retention beyond what CloudWatch offers
The Reality Check (Because Nothing's Perfect)
Look, there's no free lunch. Here are the things you should actually think about:
You need specific add-on versions. Standard clusters require Kubernetes 1.28 or later and the CloudWatch Observability add-on v6.2.0 or later with OTel Container Insights enabled. If you're running older versions, you'll need to upgrade before any of this works. That's not a trivial change in production.
PromQL query compatibility varies by version. The PromQL queries were validated with add-on v6.4.0-eksbuild.1 on Kubernetes 1.34 running EKS Auto Mode. If you're on a different version combination, your mileage may vary β and debugging PromQL query compatibility issues isn't exactly a fun Friday afternoon.
Widget count adds up fast. Between 4 dashboards you're looking at 67 total widgets (35 + 9 + 11 + 12). That's a lot of configuration to maintain. Dashboard-as-code helps, but it's still real surface area that someone owns.
It's AWS-only. If your organization runs multi-cloud Kubernetes, this NOC approach only covers the EKS leg. You'll still need something else for your GKE or AKS clusters.
But honestly? If EKS is your primary platform and CloudWatch is already in your stack, the trade-off is overwhelmingly worth it. You get a production-ready NOC with encoded operational opinions built into the structure β not just a pile of metrics hoping someone will interpret them correctly.
Wrapping Everything Up
Remember that 2 AM page where you spent ten minutes figuring out where to look? That's exactly what this dashboard eliminates. Four purpose-built views β executive rollup, infrastructure drilldown, application health, and incident response β each scoped to the right time window with the right metrics.
The real takeaway here isn't the widgets or the PromQL syntax. It's the idea that a great NOC dashboard is an opinion. It tells your team what matters, what to suppress, and what "healthy" actually looks like. Hint: it's an explicit zero, not a blank panel.
Go build this for your cluster. Customize the operational opinions to match your team's runbooks. And next time that 2 AM page hits, you'll open one dashboard instead of six tabs. And honestly? That's what we all really want, isn't it? β¨
Resources
Getting Started
Documentation
- Amazon CloudWatch Container Insights
- Amazon CloudWatch Observability EKS Add-on
- CloudWatch Composite Alarms
Community Resources
I used an AI assistant to research and draft this post, then reviewed, tested and edited it before publishing.
Top comments (0)