DEV Community

Aisalkyn Aidarova
Aisalkyn Aidarova

Posted on

Production Kubernetes Troubleshooting Lab

Observability, Incident Response & Root-Cause Analysis

Scenario: You are the DevOps/SRE on-call team for an e-commerce company.

Architecture:

                    CUSTOMER
                        │
                        ▼
                   order-service
                        │
                        ▼
                    Service
                        │
              ┌─────────┼─────────┐
              ▼         ▼         ▼
            Pod-1     Pod-2     Pod-3
                        │
                        ▼
                  Dependencies
Enter fullscreen mode Exit fullscreen mode

Today we will deliberately create these incidents:

INCIDENT 1 → New deployment → CrashLoopBackOff → Rollback

INCIDENT 2 → Production slow → High traffic / CPU → HPA investigation

INCIDENT 3 → Running Pods but application unavailable → Service selector

INCIDENT 4 → Running 0/1 → Readiness failure

INCIDENT 5 → OOMKilled → Memory investigation

INCIDENT 6 → Application dependency problem → Logs
Enter fullscreen mode Exit fullscreen mode

LAB 0 — Check the Cluster

Everybody starts here.

kubectl cluster-info
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get nodes
Enter fullscreen mode Exit fullscreen mode

Expected:

NAME                         STATUS   ROLES    AGE
ip-192-168-10-10.ec2...      Ready    <none>   ...
ip-192-168-20-20.ec2...      Ready    <none>   ...
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get nodes -o wide
Enter fullscreen mode Exit fullscreen mode

Now create our production namespace:

kubectl create namespace production
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get namespace production
Enter fullscreen mode Exit fullscreen mode

LAB 1 — Build Healthy Production First

Before troubleshooting production, we need healthy production.

Create:

nano production.yaml
Enter fullscreen mode Exit fullscreen mode

Paste:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: order-service
  namespace: production
spec:
  replicas: 3

  selector:
    matchLabels:
      app: order-service

  template:
    metadata:
      labels:
        app: order-service

    spec:
      containers:
      - name: order-service
        image: nginx:1.27

        ports:
        - containerPort: 80

        resources:
          requests:
            cpu: "100m"
            memory: "64Mi"
          limits:
            cpu: "500m"
            memory: "256Mi"

        readinessProbe:
          httpGet:
            path: /
            port: 80
          initialDelaySeconds: 5
          periodSeconds: 5

        livenessProbe:
          httpGet:
            path: /
            port: 80
          initialDelaySeconds: 10
          periodSeconds: 10

---
apiVersion: v1
kind: Service
metadata:
  name: order-service
  namespace: production
spec:
  selector:
    app: order-service

  ports:
  - port: 80
    targetPort: 80

  type: ClusterIP
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f production.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Eventually:

NAME                             READY   STATUS
order-service-xxxxxxxxxx-abcde   1/1     Running
order-service-xxxxxxxxxx-fghij   1/1     Running
order-service-xxxxxxxxxx-klmno   1/1     Running
Enter fullscreen mode Exit fullscreen mode

Press:

Ctrl+C
Enter fullscreen mode Exit fullscreen mode

Check the Entire Production Environment

kubectl get all -n production
Enter fullscreen mode Exit fullscreen mode

Check labels:

kubectl get pods -n production --show-labels
Enter fullscreen mode Exit fullscreen mode

You should see:

app=order-service
Enter fullscreen mode Exit fullscreen mode

Check Service:

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

Now:

kubectl get endpoints order-service -n production
Enter fullscreen mode Exit fullscreen mode

Depending on Kubernetes version, you may also use:

kubectl get endpointslice -n production
Enter fullscreen mode Exit fullscreen mode

You should have backend addresses.

Concept:

Service
   │
   │ selector:
   │ app=order-service
   │
   ▼
Pod        Pod        Pod
app=       app=       app=
order      order      order
-service   -service   -service
Enter fullscreen mode Exit fullscreen mode

Everything matches.

Production is healthy.


Verify the Application

Create a temporary troubleshooting Pod:

kubectl run test-client \
  --image=curlimages/curl \
  --restart=Never \
  -n production \
  -- sleep 3600
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get pod test-client -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl exec -n production test-client -- \
curl -s http://order-service
Enter fullscreen mode Exit fullscreen mode

You should receive the Nginx HTML page.

Now tell students:

This is our known-good production baseline.

That's important.


INCIDENT 1 — New Deployment Breaks Production

STUDENTS ONLY RECEIVE THIS MESSAGE

🚨 INCIDENT: Customers cannot place orders. The Order Service started failing immediately after today's deployment.

Ask students:

What changed?

Answer:

New deployment.
Enter fullscreen mode Exit fullscreen mode

First check:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

But currently everything works.


INSTRUCTOR — Break Production

Don't show students the command.

Run:

kubectl set image deployment/order-service \
order-service=nginx:this-version-does-not-exist \
-n production
Enter fullscreen mode Exit fullscreen mode

Now students run:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

They should eventually see something similar to:

order-service-old        1/1   Running
order-service-new        0/1   ImagePullBackOff
order-service-new        0/1   ImagePullBackOff
Enter fullscreen mode Exit fullscreen mode

Important teaching point:

Because this is a rolling update, old healthy Pods may remain available while new Pods fail. So production may be degraded rather than completely unavailable.


Student Investigation

First:

kubectl rollout status deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

It will not complete normally.

Then:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Pick one failing Pod:

kubectl describe pod <NEW-POD-NAME> -n production
Enter fullscreen mode Exit fullscreen mode

Look at Events.

Students should find something similar to:

Failed to pull image
Enter fullscreen mode Exit fullscreen mode

and:

ErrImagePull
Enter fullscreen mode Exit fullscreen mode

then:

ImagePullBackOff
Enter fullscreen mode Exit fullscreen mode

Now ask:

Is ImagePullBackOff the root cause?

More precisely, it's the Kubernetes symptom/state. The useful cause in the Events is that the requested image/tag cannot be pulled.

Check Deployment:

kubectl get deployment order-service \
-n production \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Enter fullscreen mode Exit fullscreen mode

It should show:

nginx:this-version-does-not-exist
Enter fullscreen mode Exit fullscreen mode

Check Rollout History

kubectl rollout history deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

You'll have revisions.

Now ask:

Production worked before the release. The new rollout cannot start. Should we sit in production and experiment with the image?

No.

Mitigate first.

kubectl rollout undo deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl rollout status deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Eventually:

READY   STATUS
1/1     Running
1/1     Running
1/1     Running
Enter fullscreen mode Exit fullscreen mode

Check image:

kubectl get deployment order-service \
-n production \
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
Enter fullscreen mode Exit fullscreen mode

Should be back to:

nginx:1.27
Enter fullscreen mode Exit fullscreen mode

Verify from customer path:

kubectl exec -n production test-client -- \
curl -s http://order-service
Enter fullscreen mode Exit fullscreen mode

Incident #1 conclusion

New Deployment
      ↓
New Pods fail
      ↓
ImagePullBackOff
      ↓
Events
      ↓
Invalid/unavailable image tag
      ↓
ROLLBACK
      ↓
Previous known-good release
      ↓
VERIFY
Enter fullscreen mode Exit fullscreen mode

INCIDENT 2 — Production Is Slow

Now we make the lab more professional.

Tell students only:

🚨 INCIDENT: Production is slow. Customers report that Order Service responses are taking much longer than normal. There has been no new deployment.

Ask:

What's the cause?

Students should NOT answer:

“Traffic.”

They should say:

We need evidence.


First Establish Baseline

Check:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

All should be:

1/1 Running
Enter fullscreen mode Exit fullscreen mode

Now:

kubectl top pods -n production
Enter fullscreen mode Exit fullscreen mode

If this command works, save/observe the CPU usage.

If it returns something like:

error: Metrics API not available
Enter fullscreen mode Exit fullscreen mode

that itself is an observability problem: Metrics Server isn't available/configured. Don't confuse that with the application incident.

Check HPA:

kubectl get hpa -n production
Enter fullscreen mode Exit fullscreen mode

There probably isn't one yet.


Create HPA

kubectl autoscale deployment order-service \
--cpu-percent=50 \
--min=3 \
--max=6 \
-n production
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get hpa -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe hpa order-service -n production
Enter fullscreen mode Exit fullscreen mode

If Metrics Server is functioning, after metrics become available you'll see current/target CPU utilization.


Generate Traffic

Run a load-generator Pod:

kubectl run load-generator \
--image=busybox:1.36 \
--restart=Never \
-n production \
-- /bin/sh -c \
'while true; do wget -q -O- http://order-service >/dev/null; done'
Enter fullscreen mode Exit fullscreen mode

One loop may not produce enough CPU to demonstrate scaling. Create several:

for i in 1 2 3 4 5; do
  kubectl run load-$i \
  --image=busybox:1.36 \
  --restart=Never \
  -n production \
  -- /bin/sh -c \
  'while true; do wget -q -O- http://order-service >/dev/null; done'
done
Enter fullscreen mode Exit fullscreen mode

Now watch:

kubectl top pods -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get hpa -n production -w
Enter fullscreen mode Exit fullscreen mode

And in another terminal:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

What should students investigate?

Request/load increased
        ↓
CPU may increase
        ↓
HPA observes CPU relative to requests
        ↓
Desired replicas may increase
        ↓
Deployment creates additional Pods
Enter fullscreen mode Exit fullscreen mode

Important: don't promise a particular CPU percentage or replica count. It depends on your cluster, Nginx workload, Metrics Server, timing, and how much load the generators produce.

That's actually good for the class: students must observe evidence rather than copy an expected number.

Stop watchers with:

Ctrl+C
Enter fullscreen mode Exit fullscreen mode

Make Scaling Intentionally Too Restrictive

Now create a more interesting scenario.

kubectl patch hpa order-service \
-n production \
-p '{"spec":{"maxReplicas":3}}'
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get hpa order-service -n production
Enter fullscreen mode Exit fullscreen mode

Now you have:

MINPODS = 3
MAXPODS = 3
Enter fullscreen mode Exit fullscreen mode

The HPA cannot scale beyond three.

Keep load running.

Check:

kubectl top pods -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe hpa order-service -n production
Enter fullscreen mode Exit fullscreen mode

Now ask:

What evidence would you need before saying traffic/capacity is the reason for slowness?

Students should correlate:

Load/request rate ↑
CPU/resource pressure ↑
Latency ↑
HPA target exceeded
Replicas already at configured maximum
Enter fullscreen mode Exit fullscreen mode

Not:

Production slow
→ probably traffic.
Enter fullscreen mode Exit fullscreen mode

Stop Traffic

kubectl delete pod load-generator -n production --ignore-not-found
Enter fullscreen mode Exit fullscreen mode

And:

kubectl delete pod \
load-1 load-2 load-3 load-4 load-5 \
-n production \
--ignore-not-found
Enter fullscreen mode Exit fullscreen mode

Restore HPA:

kubectl patch hpa order-service \
-n production \
-p '{"spec":{"maxReplicas":6}}'
Enter fullscreen mode Exit fullscreen mode

INCIDENT 3 — All Pods Running, Application Doesn't Work

This is one of the best incidents.

Tell students:

🚨 INCIDENT: All Order Service Pods show Running 1/1, but requests to the Order Service fail.

Ask:

If Pods are Running, does that prove production is working?

No.


INSTRUCTOR — Break Service

Run:

kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"wrong-order-service"}}}'
Enter fullscreen mode Exit fullscreen mode

Students don't see this command.


Student Investigation

They start:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Result:

Running
Running
Running
Enter fullscreen mode Exit fullscreen mode

Then try:

kubectl exec -n production test-client -- \
curl --max-time 5 http://order-service
Enter fullscreen mode Exit fullscreen mode

It fails/times out depending on the exact networking behavior.

Now:

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

Service exists.

Next:

kubectl get endpoints order-service -n production
Enter fullscreen mode Exit fullscreen mode

Likely:

ENDPOINTS   <none>
Enter fullscreen mode Exit fullscreen mode

On newer clusters also inspect:

kubectl get endpointslice -n production
Enter fullscreen mode Exit fullscreen mode

Now compare:

kubectl describe service order-service -n production
Enter fullscreen mode Exit fullscreen mode

with:

kubectl get pods -n production --show-labels
Enter fullscreen mode Exit fullscreen mode

Service selector:

app=wrong-order-service
Enter fullscreen mode Exit fullscreen mode

Pod label:

app=order-service
Enter fullscreen mode Exit fullscreen mode

There is your evidence.

Pods Running ✓
      ↓
Service exists ✓
      ↓
Endpoints ✗
      ↓
Selector does not match labels
      ↓
ROOT CAUSE
Enter fullscreen mode Exit fullscreen mode

Fix It

kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"order-service"}}}'
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get endpoints order-service -n production
Enter fullscreen mode Exit fullscreen mode

Addresses should return.

Verify:

kubectl exec -n production test-client -- \
curl -s http://order-service
Enter fullscreen mode Exit fullscreen mode

Production works again.


INCIDENT 4 — Running 0/1

Tell students:

🚨 INCIDENT: Pods are running, but they are not becoming Ready.

Instructor breaks readiness:

kubectl patch deployment order-service \
-n production \
--type='json' \
-p='[
  {
    "op":"replace",
    "path":"/spec/template/spec/containers/0/readinessProbe/httpGet/path",
    "value":"/health-does-not-exist"
  }
]'
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl rollout status deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

Students check:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

New Pods may show:

0/1 Running
Enter fullscreen mode Exit fullscreen mode

Ask:

What's the difference between STATUS and READY?

Now:

kubectl describe pod <POD-NAME> -n production
Enter fullscreen mode Exit fullscreen mode

They should find readiness probe failures such as an HTTP failure status.

Now check:

kubectl get endpoints order-service -n production
Enter fullscreen mode Exit fullscreen mode

Ready backends can disappear as the bad rollout progresses.

Concept:

Container
   ↓
Running
   ↓
Readiness check
   ↓
FAIL
   ↓
Pod not Ready
   ↓
Not a healthy Service backend
Enter fullscreen mode Exit fullscreen mode

What Do We Do in Production?

Again ask:

This happened immediately after a new configuration/deployment. Previous revision worked. What is our fastest safe mitigation?

Rollback.

kubectl rollout history deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl rollout undo deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

Verify:

kubectl rollout status deployment/order-service -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

And:

kubectl get endpoints order-service -n production
Enter fullscreen mode Exit fullscreen mode

Finally:

kubectl exec -n production test-client -- \
curl -s http://order-service
Enter fullscreen mode Exit fullscreen mode

INCIDENT 5 — OOMKilled

For this one use a separate workload so we don't destroy the main application.

Create:

nano memory-app.yaml
Enter fullscreen mode Exit fullscreen mode

Paste:

apiVersion: v1
kind: Pod
metadata:
  name: memory-app
  namespace: production
spec:
  containers:
  - name: memory-app
    image: python:3.12-alpine

    resources:
      requests:
        memory: "20Mi"
      limits:
        memory: "40Mi"

    command:
      - python
      - -c
      - |
        import time
        data = []
        while True:
            data.append(bytearray(10 * 1024 * 1024))
            print("Allocated more memory", flush=True)
            time.sleep(1)
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f memory-app.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pod memory-app -n production -w
Enter fullscreen mode Exit fullscreen mode

The container should exceed its memory limit and be terminated/restarted. Depending on timing, you may see OOMKilled most clearly in describe rather than as the STATUS column.

Run:

kubectl describe pod memory-app -n production
Enter fullscreen mode Exit fullscreen mode

Look for:

Last State:
  Terminated
Reason:
  OOMKilled
Enter fullscreen mode Exit fullscreen mode

Now:

kubectl logs memory-app -n production
Enter fullscreen mode Exit fullscreen mode

And critically:

kubectl logs memory-app -n production --previous
Enter fullscreen mode Exit fullscreen mode

This is a very good place to teach --previous.

Ask:

Was there a deployment?

No.

Should we automatically rollback?

No.

Evidence points toward memory behavior/resource limit.

Remove:

kubectl delete pod memory-app -n production
Enter fullscreen mode Exit fullscreen mode

INCIDENT 6 — Dependency Failure

Now we'll simulate the idea of:

Order Service
     ↓
Kafka / Database / API
Enter fullscreen mode Exit fullscreen mode

without requiring Kafka installation.

Create:

nano dependency-app.yaml
Enter fullscreen mode Exit fullscreen mode

Paste:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: payment-service
  namespace: production
spec:
  replicas: 1

  selector:
    matchLabels:
      app: payment-service

  template:
    metadata:
      labels:
        app: payment-service

    spec:
      containers:
      - name: payment-service
        image: busybox:1.36

        command:
        - /bin/sh
        - -c
        - |
          echo "Starting payment-service"
          echo "Connecting to Kafka..."
          sleep 2
          echo "ERROR: Cannot connect to kafka.production.svc:9092"
          echo "Application exiting"
          exit 1
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f dependency-app.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Eventually:

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Now tell students only:

🚨 INCIDENT: Payment Service is unavailable.

They should investigate:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl logs deployment/payment-service -n production
Enter fullscreen mode Exit fullscreen mode

Eventually use the actual Pod name:

kubectl get pods -n production \
-l app=payment-service
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl logs <PAYMENT-POD> \
-n production \
--previous
Enter fullscreen mode Exit fullscreen mode

They'll see:

Starting payment-service
Connecting to Kafka...
ERROR: Cannot connect to kafka.production.svc:9092
Application exiting
Enter fullscreen mode Exit fullscreen mode

Now ask:

Have we proven Kafka itself is down?

No.

This is very important.

We have proven:

Payment application reports that it cannot connect to the Kafka endpoint.

We haven't yet proven why.

Potential investigation branches:

                 Kafka connection failure
                           │
         ┌─────────────────┼─────────────────┐
         ↓                 ↓                 ↓
      DNS/name           Network           Kafka
      wrong?             blocked?          unhealthy?
         │                 │                 │
         ↓                 ↓                 ↓
      Service?        NetworkPolicy?      Brokers?
         │
         ↓
       Port?
Enter fullscreen mode Exit fullscreen mode

That's production thinking.

Delete:

kubectl delete deployment payment-service -n production
Enter fullscreen mode Exit fullscreen mode

FINAL INCIDENT — Students Get No Hints

Now combine everything.

Tell them:

🚨 SEV-1 INCIDENT

Customers report that the Order Service is unavailable.

You are the on-call DevOps team. Investigate and restore service.

Don't tell them anything else.

Instructor secretly breaks:

kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"orders"}}}'
Enter fullscreen mode Exit fullscreen mode

Then sit back.

Students should develop this process themselves:

What's affected?
       ↓
Order Service
       ↓
Any deployment?
       ↓
No
       ↓
Pods?
       ↓
Running / Ready
       ↓
Resources?
       ↓
Healthy
       ↓
Service?
       ↓
Exists
       ↓
Endpoints?
       ↓
NONE
       ↓
Why?
       ↓
Selector vs labels
       ↓
Mismatch
       ↓
ROOT CAUSE
Enter fullscreen mode Exit fullscreen mode

They fix:

kubectl patch service order-service \
-n production \
-p '{"spec":{"selector":{"app":"order-service"}}}'
Enter fullscreen mode Exit fullscreen mode

But don't accept:

“Fixed.”

Ask:

PROVE IT.

They need to verify:

kubectl get endpoints order-service -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl exec -n production test-client -- \
curl -s http://order-service
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Only now:

Production restored.


The Troubleshooting Command Sheet

This is the part I would have every student save.

# =====================================================
# 1. WORKLOAD
# =====================================================

kubectl get pods -n production

kubectl get pods -n production -o wide

kubectl describe pod <pod> -n production


# =====================================================
# 2. LOGS
# =====================================================

kubectl logs <pod> -n production

kubectl logs <pod> -n production --previous


# =====================================================
# 3. EVENTS
# =====================================================

kubectl get events \
-n production \
--sort-by=.metadata.creationTimestamp


# =====================================================
# 4. METRICS
# =====================================================

kubectl top pods -n production

kubectl top nodes


# =====================================================
# 5. DEPLOYMENT
# =====================================================

kubectl get deployment -n production

kubectl describe deployment <deployment> -n production

kubectl rollout status deployment/<deployment> -n production

kubectl rollout history deployment/<deployment> -n production

kubectl rollout undo deployment/<deployment> -n production


# =====================================================
# 6. NETWORK
# =====================================================

kubectl get svc -n production

kubectl describe svc <service> -n production

kubectl get endpoints -n production

kubectl get endpointslice -n production

kubectl get ingress -n production

kubectl get networkpolicy -n production


# =====================================================
# 7. STORAGE
# =====================================================

kubectl get pvc -n production

kubectl get pv

kubectl get storageclass

kubectl describe pvc <pvc> -n production


# =====================================================
# 8. SCALING
# =====================================================

kubectl get hpa -n production

kubectl describe hpa <hpa> -n production


# =====================================================
# 9. SECURITY
# =====================================================

kubectl auth can-i get pods -n production

kubectl auth can-i delete pods -n production


# =====================================================
# 10. CONFIGURATION
# =====================================================

kubectl get configmap -n production

kubectl get secret -n production
Enter fullscreen mode Exit fullscreen mode

The Production Rule Students Need to Remember

Don't teach them this:

Production Down
     ↓
kubectl get pods
     ↓
restart everything
Enter fullscreen mode Exit fullscreen mode

Teach them this:

                    INCIDENT
                       │
                       ▼
              WHAT IS THE IMPACT?
                       │
                       ▼
               WHAT CHANGED?
                       │
                       ▼
          IDENTIFY AFFECTED COMPONENT
                       │
                       ▼
         ┌─────────────┼─────────────┐
         ▼             ▼             ▼
       LOGS          METRICS       EVENTS
         │             │             │
         └─────────────┼─────────────┘
                       ▼
                  HYPOTHESIS
                       │
                       ▼
                    EVIDENCE
                       │
                       ▼
                  ROOT CAUSE
                       │
             ┌─────────┴─────────┐
             ▼                   ▼
         MITIGATE               FIX
       e.g. rollback
             │
             ▼
           VERIFY
             │
             ▼
      PRODUCTION RESTORED
             │
             ▼
       ROOT CAUSE ANALYSIS
Enter fullscreen mode Exit fullscreen mode

And throughout the lab, whenever a student says:

“I think it's traffic.”

Your response should be:

“Show me the evidence.”

If they say:

“It's Kubernetes.”

Ask:

“Which Kubernetes component? Show me the evidence.”

If they say:

“Kafka is down.”

Ask:

“Have you proven Kafka is down, or have you only proven that your application cannot connect to Kafka?”

If they say:

“I fixed it.”

Top comments (0)