Observability, Incident Response & Root-Cause Analysis
Scenario: You are the DevOps/SRE on-call team for an e-commerce company.
Architecture:
CUSTOMER
│
▼
order-service
│
▼
Service
│
┌─────────┼─────────┐
▼ ▼ ▼
Pod-1 Pod-2 Pod-3
│
▼
Dependencies
Today we will deliberately create these incidents:
INCIDENT 1 → New deployment → CrashLoopBackOff → Rollback
INCIDENT 2 → Production slow → High traffic / CPU → HPA investigation
INCIDENT 3 → Running Pods but application unavailable → Service selector
INCIDENT 4 → Running 0/1 → Readiness failure
INCIDENT 5 → OOMKilled → Memory investigation
INCIDENT 6 → Application dependency problem → Logs
LAB 0 — Check the Cluster
Everybody starts here.
kubectl cluster-info
Then:
kubectl get nodes
Expected:
NAME STATUS ROLES AGE
ip-192-168-10-10.ec2... Ready <none> ...
ip-192-168-20-20.ec2... Ready <none> ...
Check:
kubectl get nodes -o wide
Now create our production namespace:
kubectl create namespace production
Check:
kubectl get namespace production
LAB 1 — Build Healthy Production First
Before troubleshooting production, we need healthy production.
Create:
nano production.yaml
Paste:
apiVersion: apps/v1
kind: Deployment
metadata:
name: order-service
namespace: production
spec:
replicas: 3
selector:
matchLabels:
app: order-service
template:
metadata:
labels:
app: order-service
spec:
containers:
- name: order-service
image: nginx:1.27
ports:
- containerPort: 80
resources:
requests:
cpu: "100m"
memory: "64Mi"
limits:
cpu: "500m"
memory: "256Mi"
readinessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 10
periodSeconds: 10
---
apiVersion: v1
kind: Service
metadata:
name: order-service
namespace: production
spec:
selector:
app: order-service
ports:
- port: 80
targetPort: 80
type: ClusterIP
Apply:
kubectl apply -f production.yaml
Watch:
kubectl get pods -n production -w
Eventually:
NAME READY STATUS
order-service-xxxxxxxxxx-abcde 1/1 Running
order-service-xxxxxxxxxx-fghij 1/1 Running
order-service-xxxxxxxxxx-klmno 1/1 Running
Press:
Ctrl+C
Check the Entire Production Environment
kubectl get all -n production
Check labels:
kubectl get pods -n production --show-labels
You should see:
app=order-service
Check Service:
kubectl get svc -n production
Now:
kubectl get endpoints order-service -n production
Depending on Kubernetes version, you may also use:
kubectl get endpointslice -n production
You should have backend addresses.
Concept:
Service
│
│ selector:
│ app=order-service
│
▼
Pod Pod Pod
app= app= app=
order order order
-service -service -service
Everything matches.
Production is healthy.
Verify the Application
Create a temporary troubleshooting Pod:
kubectl run test-client \
--image=curlimages/curl \
--restart=Never \
-n production \
-- sleep 3600
Check:
kubectl get pod test-client -n production
Then:
kubectl exec -n production test-client -- \
curl -s http://order-service
You should receive the Nginx HTML page.
Now tell students:
This is our known-good production baseline.
That's important.
INCIDENT 1 — New Deployment Breaks Production
STUDENTS ONLY RECEIVE THIS MESSAGE
🚨 INCIDENT: Customers cannot place orders. The Order Service started failing immediately after today's deployment.
Ask students:
What changed?
Answer:
New deployment.
First check:
kubectl get pods -n production
But currently everything works.
INSTRUCTOR — Break Production
Don't show students the command.
Run:
kubectl set image deployment/order-service \
order-service=nginx:this-version-does-not-exist \
-n production
Now students run:
kubectl get pods -n production
They should eventually see something similar to:
order-service-old 1/1 Running
order-service-new 0/1 ImagePullBackOff
order-service-new 0/1 ImagePullBackOff
Important teaching point:
Because this is a rolling update, old healthy Pods may remain available while new Pods fail. So production may be degraded rather than completely unavailable.
Student Investigation
First:
kubectl rollout status deployment/order-service -n production
It will not complete normally.
Then:
kubectl get pods -n production
Pick one failing Pod:
kubectl describe pod <NEW-POD-NAME> -n production
Look at Events.
Students should find something similar to:
Failed to pull image
and:
ErrImagePull
then:
ImagePullBackOff
Now ask:
Is
ImagePullBackOffthe root cause?
More precisely, it's the Kubernetes symptom/state. The useful cause in the Events is that the requested image/tag cannot be pulled.
Check Deployment:
kubectl get deployment order-service \
-n production \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
It should show:
nginx:this-version-does-not-exist
Check Rollout History
kubectl rollout history deployment/order-service -n production
You'll have revisions.
Now ask:
Production worked before the release. The new rollout cannot start. Should we sit in production and experiment with the image?
No.
Mitigate first.
kubectl rollout undo deployment/order-service -n production
Then:
kubectl rollout status deployment/order-service -n production
Check:
kubectl get pods -n production
Eventually:
READY STATUS
1/1 Running
1/1 Running
1/1 Running
Check image:
kubectl get deployment order-service \
-n production \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Should be back to:
nginx:1.27
Verify from customer path:
kubectl exec -n production test-client -- \
curl -s http://order-service
Incident #1 conclusion
New Deployment
↓
New Pods fail
↓
ImagePullBackOff
↓
Events
↓
Invalid/unavailable image tag
↓
ROLLBACK
↓
Previous known-good release
↓
VERIFY
INCIDENT 2 — Production Is Slow
Now we make the lab more professional.
Tell students only:
🚨 INCIDENT: Production is slow. Customers report that Order Service responses are taking much longer than normal. There has been no new deployment.
Ask:
What's the cause?
Students should NOT answer:
“Traffic.”
They should say:
We need evidence.
First Establish Baseline
Check:
kubectl get pods -n production
All should be:
1/1 Running
Now:
kubectl top pods -n production
If this command works, save/observe the CPU usage.
If it returns something like:
error: Metrics API not available
that itself is an observability problem: Metrics Server isn't available/configured. Don't confuse that with the application incident.
Check HPA:
kubectl get hpa -n production
There probably isn't one yet.
Create HPA
kubectl autoscale deployment order-service \
--cpu-percent=50 \
--min=3 \
--max=6 \
-n production
Check:
kubectl get hpa -n production
Then:
kubectl describe hpa order-service -n production
If Metrics Server is functioning, after metrics become available you'll see current/target CPU utilization.
Generate Traffic
Run a load-generator Pod:
kubectl run load-generator \
--image=busybox:1.36 \
--restart=Never \
-n production \
-- /bin/sh -c \
'while true; do wget -q -O- http://order-service >/dev/null; done'
One loop may not produce enough CPU to demonstrate scaling. Create several:
for i in 1 2 3 4 5; do
kubectl run load-$i \
--image=busybox:1.36 \
--restart=Never \
-n production \
-- /bin/sh -c \
'while true; do wget -q -O- http://order-service >/dev/null; done'
done
Now watch:
kubectl top pods -n production
Then:
kubectl get hpa -n production -w
And in another terminal:
kubectl get pods -n production -w
What should students investigate?
Request/load increased
↓
CPU may increase
↓
HPA observes CPU relative to requests
↓
Desired replicas may increase
↓
Deployment creates additional Pods
Important: don't promise a particular CPU percentage or replica count. It depends on your cluster, Nginx workload, Metrics Server, timing, and how much load the generators produce.
That's actually good for the class: students must observe evidence rather than copy an expected number.
Stop watchers with:
Ctrl+C
Make Scaling Intentionally Too Restrictive
Now create a more interesting scenario.
kubectl patch hpa order-service \
-n production \
-p '{"spec":{"maxReplicas":3}}'
Check:
kubectl get hpa order-service -n production
Now you have:
MINPODS = 3
MAXPODS = 3
The HPA cannot scale beyond three.
Keep load running.
Check:
kubectl top pods -n production
Then:
kubectl describe hpa order-service -n production
Now ask:
What evidence would you need before saying traffic/capacity is the reason for slowness?
Students should correlate:
Load/request rate ↑
CPU/resource pressure ↑
Latency ↑
HPA target exceeded
Replicas already at configured maximum
Not:
Production slow
→ probably traffic.
Stop Traffic
kubectl delete pod load-generator -n production --ignore-not-found
And:
kubectl delete pod \
load-1 load-2 load-3 load-4 load-5 \
-n production \
--ignore-not-found
Restore HPA:
kubectl patch hpa order-service \
-n production \
-p '{"spec":{"maxReplicas":6}}'
INCIDENT 3 — All Pods Running, Application Doesn't Work
This is one of the best incidents.
Tell students:
🚨 INCIDENT: All Order Service Pods show
Running 1/1, but requests to the Order Service fail.
Ask:
If Pods are Running, does that prove production is working?
No.
INSTRUCTOR — Break Service
Run:
kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"wrong-order-service"}}}'
Students don't see this command.
Student Investigation
They start:
kubectl get pods -n production
Result:
Running
Running
Running
Then try:
kubectl exec -n production test-client -- \
curl --max-time 5 http://order-service
It fails/times out depending on the exact networking behavior.
Now:
kubectl get svc -n production
Service exists.
Next:
kubectl get endpoints order-service -n production
Likely:
ENDPOINTS <none>
On newer clusters also inspect:
kubectl get endpointslice -n production
Now compare:
kubectl describe service order-service -n production
with:
kubectl get pods -n production --show-labels
Service selector:
app=wrong-order-service
Pod label:
app=order-service
There is your evidence.
Pods Running ✓
↓
Service exists ✓
↓
Endpoints ✗
↓
Selector does not match labels
↓
ROOT CAUSE
Fix It
kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"order-service"}}}'
Check:
kubectl get endpoints order-service -n production
Addresses should return.
Verify:
kubectl exec -n production test-client -- \
curl -s http://order-service
Production works again.
INCIDENT 4 — Running 0/1
Tell students:
🚨 INCIDENT: Pods are running, but they are not becoming Ready.
Instructor breaks readiness:
kubectl patch deployment order-service \
-n production \
--type='json' \
-p='[
{
"op":"replace",
"path":"/spec/template/spec/containers/0/readinessProbe/httpGet/path",
"value":"/health-does-not-exist"
}
]'
Watch:
kubectl rollout status deployment/order-service -n production
Students check:
kubectl get pods -n production
New Pods may show:
0/1 Running
Ask:
What's the difference between STATUS and READY?
Now:
kubectl describe pod <POD-NAME> -n production
They should find readiness probe failures such as an HTTP failure status.
Now check:
kubectl get endpoints order-service -n production
Ready backends can disappear as the bad rollout progresses.
Concept:
Container
↓
Running
↓
Readiness check
↓
FAIL
↓
Pod not Ready
↓
Not a healthy Service backend
What Do We Do in Production?
Again ask:
This happened immediately after a new configuration/deployment. Previous revision worked. What is our fastest safe mitigation?
Rollback.
kubectl rollout history deployment/order-service -n production
Then:
kubectl rollout undo deployment/order-service -n production
Verify:
kubectl rollout status deployment/order-service -n production
Then:
kubectl get pods -n production
And:
kubectl get endpoints order-service -n production
Finally:
kubectl exec -n production test-client -- \
curl -s http://order-service
INCIDENT 5 — OOMKilled
For this one use a separate workload so we don't destroy the main application.
Create:
nano memory-app.yaml
Paste:
apiVersion: v1
kind: Pod
metadata:
name: memory-app
namespace: production
spec:
containers:
- name: memory-app
image: python:3.12-alpine
resources:
requests:
memory: "20Mi"
limits:
memory: "40Mi"
command:
- python
- -c
- |
import time
data = []
while True:
data.append(bytearray(10 * 1024 * 1024))
print("Allocated more memory", flush=True)
time.sleep(1)
Apply:
kubectl apply -f memory-app.yaml
Watch:
kubectl get pod memory-app -n production -w
The container should exceed its memory limit and be terminated/restarted. Depending on timing, you may see OOMKilled most clearly in describe rather than as the STATUS column.
Run:
kubectl describe pod memory-app -n production
Look for:
Last State:
Terminated
Reason:
OOMKilled
Now:
kubectl logs memory-app -n production
And critically:
kubectl logs memory-app -n production --previous
This is a very good place to teach --previous.
Ask:
Was there a deployment?
No.
Should we automatically rollback?
No.
Evidence points toward memory behavior/resource limit.
Remove:
kubectl delete pod memory-app -n production
INCIDENT 6 — Dependency Failure
Now we'll simulate the idea of:
Order Service
↓
Kafka / Database / API
without requiring Kafka installation.
Create:
nano dependency-app.yaml
Paste:
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-service
namespace: production
spec:
replicas: 1
selector:
matchLabels:
app: payment-service
template:
metadata:
labels:
app: payment-service
spec:
containers:
- name: payment-service
image: busybox:1.36
command:
- /bin/sh
- -c
- |
echo "Starting payment-service"
echo "Connecting to Kafka..."
sleep 2
echo "ERROR: Cannot connect to kafka.production.svc:9092"
echo "Application exiting"
exit 1
Apply:
kubectl apply -f dependency-app.yaml
Watch:
kubectl get pods -n production -w
Eventually:
CrashLoopBackOff
Now tell students only:
🚨 INCIDENT: Payment Service is unavailable.
They should investigate:
kubectl get pods -n production
Then:
kubectl logs deployment/payment-service -n production
Eventually use the actual Pod name:
kubectl get pods -n production \
-l app=payment-service
Then:
kubectl logs <PAYMENT-POD> \
-n production \
--previous
They'll see:
Starting payment-service
Connecting to Kafka...
ERROR: Cannot connect to kafka.production.svc:9092
Application exiting
Now ask:
Have we proven Kafka itself is down?
No.
This is very important.
We have proven:
Payment application reports that it cannot connect to the Kafka endpoint.
We haven't yet proven why.
Potential investigation branches:
Kafka connection failure
│
┌─────────────────┼─────────────────┐
↓ ↓ ↓
DNS/name Network Kafka
wrong? blocked? unhealthy?
│ │ │
↓ ↓ ↓
Service? NetworkPolicy? Brokers?
│
↓
Port?
That's production thinking.
Delete:
kubectl delete deployment payment-service -n production
FINAL INCIDENT — Students Get No Hints
Now combine everything.
Tell them:
🚨 SEV-1 INCIDENT
Customers report that the Order Service is unavailable.
You are the on-call DevOps team. Investigate and restore service.
Don't tell them anything else.
Instructor secretly breaks:
kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"orders"}}}'
Then sit back.
Students should develop this process themselves:
What's affected?
↓
Order Service
↓
Any deployment?
↓
No
↓
Pods?
↓
Running / Ready
↓
Resources?
↓
Healthy
↓
Service?
↓
Exists
↓
Endpoints?
↓
NONE
↓
Why?
↓
Selector vs labels
↓
Mismatch
↓
ROOT CAUSE
They fix:
kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"order-service"}}}'
But don't accept:
“Fixed.”
Ask:
PROVE IT.
They need to verify:
kubectl get endpoints order-service -n production
Then:
kubectl exec -n production test-client -- \
curl -s http://order-service
Then:
kubectl get pods -n production
Only now:
Production restored.
The Troubleshooting Command Sheet
This is the part I would have every student save.
# =====================================================
# 1. WORKLOAD
# =====================================================
kubectl get pods -n production
kubectl get pods -n production -o wide
kubectl describe pod <pod> -n production
# =====================================================
# 2. LOGS
# =====================================================
kubectl logs <pod> -n production
kubectl logs <pod> -n production --previous
# =====================================================
# 3. EVENTS
# =====================================================
kubectl get events \
-n production \
--sort-by=.metadata.creationTimestamp
# =====================================================
# 4. METRICS
# =====================================================
kubectl top pods -n production
kubectl top nodes
# =====================================================
# 5. DEPLOYMENT
# =====================================================
kubectl get deployment -n production
kubectl describe deployment <deployment> -n production
kubectl rollout status deployment/<deployment> -n production
kubectl rollout history deployment/<deployment> -n production
kubectl rollout undo deployment/<deployment> -n production
# =====================================================
# 6. NETWORK
# =====================================================
kubectl get svc -n production
kubectl describe svc <service> -n production
kubectl get endpoints -n production
kubectl get endpointslice -n production
kubectl get ingress -n production
kubectl get networkpolicy -n production
# =====================================================
# 7. STORAGE
# =====================================================
kubectl get pvc -n production
kubectl get pv
kubectl get storageclass
kubectl describe pvc <pvc> -n production
# =====================================================
# 8. SCALING
# =====================================================
kubectl get hpa -n production
kubectl describe hpa <hpa> -n production
# =====================================================
# 9. SECURITY
# =====================================================
kubectl auth can-i get pods -n production
kubectl auth can-i delete pods -n production
# =====================================================
# 10. CONFIGURATION
# =====================================================
kubectl get configmap -n production
kubectl get secret -n production
The Production Rule Students Need to Remember
Don't teach them this:
Production Down
↓
kubectl get pods
↓
restart everything
Teach them this:
INCIDENT
│
▼
WHAT IS THE IMPACT?
│
▼
WHAT CHANGED?
│
▼
IDENTIFY AFFECTED COMPONENT
│
▼
┌─────────────┼─────────────┐
▼ ▼ ▼
LOGS METRICS EVENTS
│ │ │
└─────────────┼─────────────┘
▼
HYPOTHESIS
│
▼
EVIDENCE
│
▼
ROOT CAUSE
│
┌─────────┴─────────┐
▼ ▼
MITIGATE FIX
e.g. rollback
│
▼
VERIFY
│
▼
PRODUCTION RESTORED
│
▼
ROOT CAUSE ANALYSIS
And throughout the lab, whenever a student says:
“I think it's traffic.”
Your response should be:
“Show me the evidence.”
If they say:
“It's Kubernetes.”
Ask:
“Which Kubernetes component? Show me the evidence.”
If they say:
“Kafka is down.”
Ask:
“Have you proven Kafka is down, or have you only proven that your application cannot connect to Kafka?”
If they say:
“I fixed it.”
Top comments (0)