DEV Community

Aisalkyn Aidarova
Aisalkyn Aidarova

Posted on

JumpToTech — LIVE Production Troubleshooting Lab

JumpToTech — LIVE Production Troubleshooting Lab

StatefulSet + PVC + Service + DNS + Failures + Argo CD

Duration: 3 hours

Rule for today:

Do NOT immediately delete the Pod.

For every incident:

OBSERVE → DESCRIBE → LOGS → EVENTS → DEPENDENCIES → FIX → VERIFY
Enter fullscreen mode Exit fullscreen mode

PART 1 — CREATE THE LAB

Step 1 — Create the folder

mkdir stateful-production-lab
cd stateful-production-lab

mkdir k8s
cd k8s
Enter fullscreen mode Exit fullscreen mode

Check:

pwd
Enter fullscreen mode Exit fullscreen mode

Step 2 — Namespace

Create:

touch 00-namespace.yaml
Enter fullscreen mode Exit fullscreen mode

Open it:

code 00-namespace.yaml
Enter fullscreen mode Exit fullscreen mode

Add:

apiVersion: v1
kind: Namespace
metadata:
  name: production
Enter fullscreen mode Exit fullscreen mode

Step 3 — Secret

Create:

touch 01-secret.yaml
Enter fullscreen mode Exit fullscreen mode

Add:

apiVersion: v1
kind: Secret
metadata:
  name: mysql-secret
  namespace: production
type: Opaque
stringData:
  MYSQL_ROOT_PASSWORD: jumptotech123
  MYSQL_DATABASE: restaurant
  MYSQL_USER: appuser
  MYSQL_PASSWORD: apppassword
Enter fullscreen mode Exit fullscreen mode

This is a classroom Secret.

Do not use plaintext production passwords in Git like this in a real environment.


Step 4 — Headless Service

Create:

touch 02-mysql-headless.yaml
Enter fullscreen mode Exit fullscreen mode

Add:

apiVersion: v1
kind: Service
metadata:
  name: mysql-headless
  namespace: production
spec:
  clusterIP: None

  selector:
    app: mysql

  ports:
    - name: mysql
      port: 3306
      targetPort: 3306
Enter fullscreen mode Exit fullscreen mode

Explain:

mysql-0.mysql-headless.production.svc.cluster.local
Enter fullscreen mode Exit fullscreen mode

StatefulSet uses stable Pod identity.


Step 5 — StatefulSet

Create:

touch 03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Add:

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: mysql
  namespace: production

spec:
  serviceName: mysql-headless

  replicas: 1

  selector:
    matchLabels:
      app: mysql

  template:
    metadata:
      labels:
        app: mysql

    spec:
      terminationGracePeriodSeconds: 30

      containers:
        - name: mysql
          image: mysql:8.0

          ports:
            - name: mysql
              containerPort: 3306

          envFrom:
            - secretRef:
                name: mysql-secret

          resources:
            requests:
              cpu: 100m
              memory: 256Mi

            limits:
              cpu: 500m
              memory: 512Mi

          readinessProbe:
            exec:
              command:
                - sh
                - -c
                - mysqladmin ping -h 127.0.0.1 -uroot -p"$MYSQL_ROOT_PASSWORD"

            initialDelaySeconds: 15
            periodSeconds: 10
            timeoutSeconds: 5

          livenessProbe:
            exec:
              command:
                - sh
                - -c
                - mysqladmin ping -h 127.0.0.1 -uroot -p"$MYSQL_ROOT_PASSWORD"

            initialDelaySeconds: 30
            periodSeconds: 20
            timeoutSeconds: 5

          volumeMounts:
            - name: mysql-data
              mountPath: /var/lib/mysql

  volumeClaimTemplates:
    - metadata:
        name: mysql-data

      spec:
        accessModes:
          - ReadWriteOnce

        resources:
          requests:
            storage: 2Gi
Enter fullscreen mode Exit fullscreen mode

Step 6 — Apply

Return to the project root:

cd ..
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Wait until:

NAME      READY   STATUS
mysql-0   1/1     Running
Enter fullscreen mode Exit fullscreen mode

Check StatefulSet:

kubectl get sts -n production
Enter fullscreen mode Exit fullscreen mode

Check PVC:

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

Check PV:

kubectl get pv
Enter fullscreen mode Exit fullscreen mode

Check Service:

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

STOP HERE AND EXPLAIN

Ask students:

What did Kubernetes create?

Draw:

StatefulSet
     |
     v
 mysql-0
     |
     v
mysql-data-mysql-0
     |
     v
    PV
     |
     v
Persistent Storage
Enter fullscreen mode Exit fullscreen mode

Now:

kubectl get pod mysql-0 \
  -n production \
  -o jsonpath='{.metadata.name}'
Enter fullscreen mode Exit fullscreen mode

Output:

mysql-0
Enter fullscreen mode Exit fullscreen mode

Stable identity.


PART 2 — PROVE THAT DATA SURVIVES

Step 7 — Create actual data

Enter MySQL:

kubectl exec -it mysql-0 \
  -n production \
  -- mysql -u root -pjumptotech123
Enter fullscreen mode Exit fullscreen mode

Run:

USE restaurant;
Enter fullscreen mode Exit fullscreen mode

Create table:

CREATE TABLE customers (
    id INT AUTO_INCREMENT PRIMARY KEY,
    name VARCHAR(100),
    email VARCHAR(100)
);
Enter fullscreen mode Exit fullscreen mode

Insert data:

INSERT INTO customers (name, email)
VALUES
('Emma', 'emma@jumptotech.com');
Enter fullscreen mode Exit fullscreen mode

Check:

SELECT * FROM customers;
Enter fullscreen mode Exit fullscreen mode

Expected:

+----+-------+----------------------+
| id | name  | email                |
+----+-------+----------------------+
|  1 | Emma  | emma@jumptotech.com |
+----+-------+----------------------+
Enter fullscreen mode Exit fullscreen mode

Exit:

exit;
Enter fullscreen mode Exit fullscreen mode

Step 8 — Delete the Pod

Now deliberately delete the database Pod:

kubectl delete pod mysql-0 -n production
Enter fullscreen mode Exit fullscreen mode

Immediately:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

StatefulSet creates:

mysql-0
Enter fullscreen mode Exit fullscreen mode

again.

Wait until:

mysql-0   1/1   Running
Enter fullscreen mode Exit fullscreen mode

Step 9 — Did we lose the database?

Connect again:

kubectl exec -it mysql-0 \
  -n production \
  -- mysql -u root -pjumptotech123
Enter fullscreen mode Exit fullscreen mode

Run:

USE restaurant;

SELECT * FROM customers;
Enter fullscreen mode Exit fullscreen mode

The data should still exist.

Explain:

POD DELETED
    X

PVC NOT DELETED
    ✓

PV NOT DELETED
    ✓

DATA NOT DELETED
    ✓


New mysql-0
     |
     v
same PVC
     |
     v
persistent data
Enter fullscreen mode Exit fullscreen mode

This is the foundation of today's incidents.


INCIDENT 1 — CLIENT CANNOT CONNECT TO DATABASE

Time: approximately 35 minutes into class

Tell students only:

Production alert:

Application cannot connect to database.

Find the problem.

Do NOT tell them what is broken.

First:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Output:

mysql-0   1/1   Running
Enter fullscreen mode Exit fullscreen mode

Ask:

Is production healthy just because the Pod is Running?

No.

Check Service:

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

Check endpoints:

kubectl get endpoints -n production
Enter fullscreen mode Exit fullscreen mode

Check EndpointSlices:

kubectl get endpointslices -n production
Enter fullscreen mode Exit fullscreen mode

Now create a troubleshooting Pod:

kubectl run troubleshoot \
  --image=busybox:1.36 \
  --restart=Never \
  -n production \
  -- sleep 3600
Enter fullscreen mode Exit fullscreen mode

Wait:

kubectl get pod troubleshoot -n production
Enter fullscreen mode Exit fullscreen mode

Test DNS

kubectl exec troubleshoot \
  -n production \
  -- nslookup mysql-headless
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl exec troubleshoot \
  -n production \
  -- nslookup mysql-headless.production.svc.cluster.local
Enter fullscreen mode Exit fullscreen mode

If DNS works, say:

Stop.

We just proved Kubernetes DNS can resolve the Service.

Therefore, don't continue saying “maybe DNS.”

Move to the next layer.


Test the Port

BusyBox commonly provides nc.

Run:

kubectl exec troubleshoot \
  -n production \
  -- nc -zv mysql-headless 3306
Enter fullscreen mode Exit fullscreen mode

We want connectivity to:

mysql-headless:3306
Enter fullscreen mode Exit fullscreen mode

Architecture:

Troubleshooting Pod
        |
        | DNS
        v
mysql-headless
        |
        | TCP 3306
        v
     mysql-0
Enter fullscreen mode Exit fullscreen mode

NOW BREAK IT

Open:

code k8s/02-mysql-headless.yaml
Enter fullscreen mode Exit fullscreen mode

Change:

selector:
  app: mysql
Enter fullscreen mode Exit fullscreen mode

to:

selector:
  app: mysql-broken
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/02-mysql-headless.yaml
Enter fullscreen mode Exit fullscreen mode

Check Pods:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Still:

mysql-0   1/1   Running
Enter fullscreen mode Exit fullscreen mode

Now:

kubectl get endpoints mysql-headless -n production
Enter fullscreen mode Exit fullscreen mode

Students should notice the Service no longer has the expected backend.

Check labels:

kubectl get pods \
  -n production \
  --show-labels
Enter fullscreen mode Exit fullscreen mode

Compare:

SERVICE SELECTOR

app=mysql-broken


POD LABEL

app=mysql
Enter fullscreen mode Exit fullscreen mode

Root cause:

Service
   |
   X
Endpoint
   |
   X
mysql-0
Enter fullscreen mode Exit fullscreen mode

FIX INCIDENT 1

Change back:

selector:
  app: mysql
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/02-mysql-headless.yaml
Enter fullscreen mode Exit fullscreen mode

Verify:

kubectl get endpoints mysql-headless -n production
Enter fullscreen mode Exit fullscreen mode

Test again:

kubectl exec troubleshoot \
  -n production \
  -- nc -zv mysql-headless 3306
Enter fullscreen mode Exit fullscreen mode

Lesson:

Pod Running ≠ Application Reachable
Enter fullscreen mode Exit fullscreen mode

INCIDENT 2 — WRONG PORT

Tell students:

New incident.

Database Pod is healthy.

DNS resolves.

Application still cannot connect.

Break Service.

Change:

targetPort: 3306
Enter fullscreen mode Exit fullscreen mode

to:

targetPort: 3307
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/02-mysql-headless.yaml
Enter fullscreen mode Exit fullscreen mode

Now:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Healthy.

DNS:

kubectl exec troubleshoot \
  -n production \
  -- nslookup mysql-headless
Enter fullscreen mode Exit fullscreen mode

DNS works.

Now connectivity:

kubectl exec troubleshoot \
  -n production \
  -- nc -zv mysql-headless 3306
Enter fullscreen mode Exit fullscreen mode

It should fail to reach the actual MySQL listener through the misconfigured Service path.

Ask:

DNS works. Pod works. What is between them?

Answer:

SERVICE
Enter fullscreen mode Exit fullscreen mode

Inspect:

kubectl describe svc mysql-headless -n production
Enter fullscreen mode Exit fullscreen mode

Look at:

Port:        3306
TargetPort:  3307
Enter fullscreen mode Exit fullscreen mode

But MySQL listens on:

3306
Enter fullscreen mode Exit fullscreen mode

Therefore:

Application
     |
     | 3306
     v
Service
     |
     | WRONG → 3307
     X
MySQL
     |
     | listening 3306
Enter fullscreen mode Exit fullscreen mode

FIX INCIDENT 2

Restore:

targetPort: 3306
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/02-mysql-headless.yaml
Enter fullscreen mode Exit fullscreen mode

Verify:

kubectl exec troubleshoot \
  -n production \
  -- nc -zv mysql-headless 3306
Enter fullscreen mode Exit fullscreen mode

INCIDENT 3 — DATABASE POD RUNNING BUT NOT READY

Tell students:

Monitoring says database Pod exists, but it is not Ready.

Break readiness probe.

Open:

code k8s/03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Find:

readinessProbe:
Enter fullscreen mode Exit fullscreen mode

Change its command to intentionally use the wrong port:

readinessProbe:
  exec:
    command:
      - sh
      - -c
      - mysqladmin ping -h 127.0.0.1 -P 3307 -uroot -p"$MYSQL_ROOT_PASSWORD"
  initialDelaySeconds: 15
  periodSeconds: 10
  timeoutSeconds: 5
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

You may see:

mysql-0   0/1   Running
Enter fullscreen mode Exit fullscreen mode

Ask:

Is the container Running?

Yes.

Is Kubernetes considering it Ready?

No.

Investigate:

kubectl describe pod mysql-0 -n production
Enter fullscreen mode Exit fullscreen mode

Go to:

Events:
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get events \
  -n production \
  --sort-by=.lastTimestamp
Enter fullscreen mode Exit fullscreen mode

Check logs too:

kubectl logs mysql-0 -n production
Enter fullscreen mode Exit fullscreen mode

Explain:

STATUS
Running

READY
0/1
Enter fullscreen mode Exit fullscreen mode

means:

PROCESS EXISTS

but

POD IS NOT READY FOR NORMAL SERVICE TRAFFIC
Enter fullscreen mode Exit fullscreen mode

FIX INCIDENT 3

Restore:

readinessProbe:
  exec:
    command:
      - sh
      - -c
      - mysqladmin ping -h 127.0.0.1 -uroot -p"$MYSQL_ROOT_PASSWORD"
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Wait for:

mysql-0   1/1   Running
Enter fullscreen mode Exit fullscreen mode

INCIDENT 4 — IMAGEPULLBACKOFF

Tell students:

Developer released a new database version.

Production deployment started failing.

Break:

image: mysql:8.0
Enter fullscreen mode Exit fullscreen mode

Change to:

image: mysql:jumptotech-does-not-exist
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Possible:

ErrImagePull

ImagePullBackOff
Enter fullscreen mode Exit fullscreen mode

Ask students:

What is your first investigation command?

Use:

kubectl describe pod mysql-0 -n production
Enter fullscreen mode Exit fullscreen mode

Look at Events.

Then:

kubectl get events \
  -n production \
  --sort-by=.lastTimestamp
Enter fullscreen mode Exit fullscreen mode

Root cause:

Requested image/tag doesn't exist.
Enter fullscreen mode Exit fullscreen mode

Possible real causes:

wrong repository
wrong tag
private registry authentication
imagePullSecret
registry/network issue
Enter fullscreen mode Exit fullscreen mode

FIX INCIDENT 4

Restore:

image: mysql:8.0
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Watch recovery:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Then verify the data again:

kubectl exec -it mysql-0 \
  -n production \
  -- mysql -u root -pjumptotech123
Enter fullscreen mode Exit fullscreen mode
USE restaurant;

SELECT * FROM customers;
Enter fullscreen mode Exit fullscreen mode

Your data should still exist.


10-MINUTE BREAK

At this point students have diagnosed:

✓ DNS

✓ Service selector

✓ Endpoints

✓ Wrong port

✓ Readiness

✓ ImagePullBackOff

✓ Stateful persistence
Enter fullscreen mode Exit fullscreen mode

Now move specifically into storage.


INCIDENT 5 — STORAGE TROUBLESHOOTING

Start healthy:

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

Expected:

mysql-data-mysql-0   Bound
Enter fullscreen mode Exit fullscreen mode

Ask:

If the Pod becomes Pending or ContainerCreating because of storage, where do we look?

First:

kubectl describe pod mysql-0 -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe pvc mysql-data-mysql-0 \
  -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get pv
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get sc
Enter fullscreen mode Exit fullscreen mode

Then Events:

kubectl get events \
  -n production \
  --sort-by=.lastTimestamp
Enter fullscreen mode Exit fullscreen mode

Teach the dependency chain:

POD
 |
 v
PVC
 |
 v
STORAGE CLASS
 |
 v
CSI / PROVISIONER
 |
 v
PV
 |
 v
ACTUAL STORAGE
Enter fullscreen mode Exit fullscreen mode

Do not delete the working PVC just to manufacture a failure, because it contains the database created earlier.

Instead create a safe broken claim.

Create:

cat > k8s/04-broken-pvc.yaml <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: broken-storage
  namespace: production
spec:
  accessModes:
    - ReadWriteOnce
  storageClassName: storage-class-does-not-exist
  resources:
    requests:
      storage: 1Gi
EOF
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/04-broken-pvc.yaml
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

Students should see:

broken-storage   Pending
Enter fullscreen mode Exit fullscreen mode

Now ask:

Don't tell me the answer. Show me how you find it.

Run:

kubectl describe pvc broken-storage -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get sc
Enter fullscreen mode Exit fullscreen mode

Students discover:

storage-class-does-not-exist
Enter fullscreen mode Exit fullscreen mode

is invalid.

Delete only the training PVC:

kubectl delete pvc broken-storage -n production
Enter fullscreen mode Exit fullscreen mode

Delete the lab file:

rm k8s/04-broken-pvc.yaml
Enter fullscreen mode Exit fullscreen mode

Important:

WE DID NOT TOUCH:

mysql-data-mysql-0
Enter fullscreen mode Exit fullscreen mode

INCIDENT 6 — CRASHLOOPBACKOFF

Now create a separate failure so the database stays safe.

Create:

cat > k8s/05-crash.yaml <<'EOF'
apiVersion: v1
kind: Pod
metadata:
  name: broken-application
  namespace: production
spec:
  containers:
    - name: application
      image: busybox:1.36
      command:
        - sh
        - -c
        - |
          echo "Starting JumpToTech application..."
          sleep 3
          echo "ERROR: Database configuration invalid"
          exit 1
EOF
Enter fullscreen mode Exit fullscreen mode

Apply:

kubectl apply -f k8s/05-crash.yaml
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pod broken-application \
  -n production \
  -w
Enter fullscreen mode Exit fullscreen mode

Eventually:

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Now ask:

What do you do?

Logs:

kubectl logs broken-application -n production
Enter fullscreen mode Exit fullscreen mode

Then the very important command:

kubectl logs broken-application \
  -n production \
  --previous
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe pod broken-application \
  -n production
Enter fullscreen mode Exit fullscreen mode

Explain:

APPLICATION STARTS
       |
       v
APPLICATION FAILS
       |
       v
CONTAINER EXITS
       |
       v
KUBERNETES RESTARTS
       |
       v
FAILS AGAIN
       |
       v
CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Clean up:

kubectl delete -f k8s/05-crash.yaml
rm k8s/05-crash.yaml
Enter fullscreen mode Exit fullscreen mode

PART 3 — ARGO CD

Now tell students:

Up to now we intentionally changed the cluster ourselves.

In a GitOps production environment, we want Git to define desired state.

Argo CD's automated sync can reconcile Git changes, and selfHeal: true can reconcile supported live-state drift back toward the desired state in Git. :chatgpt-content-reference{index="1"}

First make sure the healthy manifests are restored.

Check:

kubectl get pods -n production
kubectl get svc -n production
kubectl get endpoints -n production
kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

We want healthy state.


Create Git Repository

From project root:

git init
git branch -M main

git add .
git commit -m "add production stateful application"
Enter fullscreen mode Exit fullscreen mode

Create an empty GitHub repository for the lab.

Then:

git remote add origin <YOUR-REPOSITORY-URL>

git push -u origin main
Enter fullscreen mode Exit fullscreen mode

Create Argo CD Application

Create:

mkdir -p argocd
Enter fullscreen mode Exit fullscreen mode

Then:

cat > argocd/application.yaml <<'EOF'
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: stateful-production
  namespace: argocd
spec:
  project: default

  source:
    repoURL: YOUR_REPOSITORY_URL
    targetRevision: main
    path: k8s

  destination:
    server: https://kubernetes.default.svc
    namespace: production

  syncPolicy:
    automated:
      prune: true
      selfHeal: true

    syncOptions:
      - CreateNamespace=true
EOF
Enter fullscreen mode Exit fullscreen mode

Replace:

YOUR_REPOSITORY_URL
Enter fullscreen mode Exit fullscreen mode

with your repository.

Commit:

git add .
git commit -m "add Argo CD application"
git push
Enter fullscreen mode Exit fullscreen mode

Bootstrap:

kubectl apply -f argocd/application.yaml
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get applications -n argocd
Enter fullscreen mode Exit fullscreen mode

Argo CD's Application spec supports automated sync, pruning, and self-healing in this configuration. :chatgpt-content-reference{index="2"}


INCIDENT 7 — ARGO CD DRIFT

Now the class becomes very interesting.

Git says:

replicas: 1
Enter fullscreen mode Exit fullscreen mode

Check:

grep -n "replicas" \
  k8s/03-mysql-statefulset.yaml
Enter fullscreen mode Exit fullscreen mode

Now somebody manually changes production:

kubectl scale sts mysql \
  --replicas=2 \
  -n production
Enter fullscreen mode Exit fullscreen mode

Immediately:

kubectl get sts mysql -n production
Enter fullscreen mode Exit fullscreen mode

You may briefly see:

DESIRED: 2
Enter fullscreen mode Exit fullscreen mode

But Git says:

1
Enter fullscreen mode Exit fullscreen mode

Ask:

Who owns desired state?

GIT
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get sts mysql \
  -n production \
  -w
Enter fullscreen mode Exit fullscreen mode

With self-healing enabled, Argo CD should reconcile the managed StatefulSet back toward the Git-defined state. :chatgpt-content-reference{index="3"}

Explain:

              GIT
          replicas: 1
               |
               v
            Argo CD
               |
          compare state
               |
               v
          Kubernetes
          replicas: 2
               |
               v

             DRIFT

               |
               v
         RECONCILIATION

               |
               v
          replicas: 1
Enter fullscreen mode Exit fullscreen mode

INCIDENT 8 — BAD PRODUCTION DEPLOYMENT THROUGH GIT

This is the FINAL incident.

Tell students:

It is 2:00 PM.

Production is working.

Developer pushes a change.

Argo CD deploys it.

Five minutes later users report an outage.

You are on call.

Now deliberately introduce the bad image through Git.

Change:

image: mysql:8.0
Enter fullscreen mode Exit fullscreen mode

to:

image: mysql:broken-production-release
Enter fullscreen mode Exit fullscreen mode

Do NOT run:

kubectl apply
Enter fullscreen mode Exit fullscreen mode

Instead:

git add .
git commit -m "release database v2"
git push
Enter fullscreen mode Exit fullscreen mode

Now Argo CD sees Git change and syncs according to its automated-sync configuration. :chatgpt-content-reference{index="4"}

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Production begins failing.


STUDENTS TROUBLESHOOT WITHOUT HELP

Give them 5–10 minutes.

Their process should be:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe pod mysql-0 -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get events \
  -n production \
  --sort-by=.lastTimestamp
Enter fullscreen mode Exit fullscreen mode

Then ask:

What changed?

Check:

git log --oneline -5
Enter fullscreen mode Exit fullscreen mode

Now they should connect:

HEALTHY PRODUCTION
       |
       v
NEW GIT COMMIT
       |
       v
ARGO CD SYNC
       |
       v
BAD IMAGE
       |
       v
ImagePullBackOff
       |
       v
DATABASE UNAVAILABLE
       |
       v
PRODUCTION INCIDENT
Enter fullscreen mode Exit fullscreen mode

PRODUCTION RECOVERY

Don't fix production manually with:

kubectl set image...
Enter fullscreen mode Exit fullscreen mode

because Git would still contain the bad desired state.

For today's simple lab, restore the known-good image in the manifest:

image: mysql:8.0
Enter fullscreen mode Exit fullscreen mode

Then:

git add .
git commit -m "fix: restore known-good mysql image"
git push
Enter fullscreen mode Exit fullscreen mode

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

Wait:

mysql-0   1/1   Running
Enter fullscreen mode Exit fullscreen mode

Now verify DATA, not just Pod status:

kubectl exec -it mysql-0 \
  -n production \
  -- mysql -u root -pjumptotech123
Enter fullscreen mode Exit fullscreen mode

Run:

USE restaurant;

SELECT * FROM customers;
Enter fullscreen mode Exit fullscreen mode

We should still have:

Emma
Enter fullscreen mode Exit fullscreen mode

That is your production verification.


FINAL INCIDENT REVIEW

Ask students:

Incident 1

Pod = Running
Client cannot connect
Enter fullscreen mode Exit fullscreen mode

Root cause:

Service selector mismatch
Enter fullscreen mode Exit fullscreen mode

How did we find it?

Service → Endpoints → Labels
Enter fullscreen mode Exit fullscreen mode

Incident 2

DNS works
Connection doesn't work
Enter fullscreen mode Exit fullscreen mode

Root cause:

Wrong targetPort
Enter fullscreen mode Exit fullscreen mode

How did we find it?

DNS test
     ↓
Port test
     ↓
Service inspection
Enter fullscreen mode Exit fullscreen mode

Incident 3

mysql-0

STATUS = Running
READY = 0/1
Enter fullscreen mode Exit fullscreen mode

Root cause:

Readiness probe
Enter fullscreen mode Exit fullscreen mode

Incident 4

ImagePullBackOff
Enter fullscreen mode Exit fullscreen mode

Root cause:

Invalid image
Enter fullscreen mode Exit fullscreen mode

Primary evidence:

kubectl describe pod
Events
Enter fullscreen mode Exit fullscreen mode

Incident 5

PVC = Pending
Enter fullscreen mode Exit fullscreen mode

Root cause in our exercise:

Invalid StorageClass
Enter fullscreen mode Exit fullscreen mode

Investigation:

PVC
 ↓
describe PVC
 ↓
StorageClass
 ↓
Events
Enter fullscreen mode Exit fullscreen mode

Incident 6

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Investigation:

logs
logs --previous
describe
events
Enter fullscreen mode Exit fullscreen mode

Incident 7

Git = 1 replica

Cluster = 2 replicas
Enter fullscreen mode Exit fullscreen mode

Problem:

Configuration Drift
Enter fullscreen mode Exit fullscreen mode

Argo CD:

reconciles desired state
Enter fullscreen mode Exit fullscreen mode

Incident 8

Production was working

Git change

Argo CD sync

Production down
Enter fullscreen mode Exit fullscreen mode

Investigation:

What changed?
    ↓
Pods
    ↓
Describe
    ↓
Events
    ↓
Git history
    ↓
Root cause
    ↓
Restore known-good desired state
    ↓
Verify application AND data
Enter fullscreen mode Exit fullscreen mode

FINAL INTERVIEW QUESTION

Your interviewer says:

“Our stateful production application suddenly became unavailable. How would you troubleshoot it?”

Answer:

“First I would clarify the scope of the incident and determine what changed around the time of failure. I would check the workload state and then use describe, logs and events to collect evidence. I would verify the application's dependencies, including Services, endpoints, DNS, ports, StatefulSet replicas, PVC/PV status and storage. If the environment uses GitOps, I would also compare the Git desired state with the live cluster and review the latest Argo CD synchronization. After identifying the root cause, I would apply or revert the appropriate change through the normal deployment process and verify both service availability and data integrity.”


Commands Students Should Know After This Lab

kubectl get pods -n production -o wide

kubectl describe pod <pod> -n production

kubectl logs <pod> -n production

kubectl logs <pod> -n production --previous

kubectl get events \
  -n production \
  --sort-by=.lastTimestamp

kubectl get sts -n production

kubectl get svc -n production

kubectl describe svc <service> -n production

kubectl get endpoints -n production

kubectl get endpointslices -n production

kubectl get pods -n production --show-labels

kubectl get pvc -n production

kubectl describe pvc <pvc> -n production

kubectl get pv

kubectl get sc
Enter fullscreen mode Exit fullscreen mode

The Rule to Remember

                 PRODUCTION DOWN
                        |
                        v
                  WHAT CHANGED?
                        |
                        v
                       GET
                        |
                        v
                    DESCRIBE
                        |
                        v
                      LOGS
                        |
                        v
                     EVENTS
                        |
                        v
                  DEPENDENCIES
                        |
        +---------------+---------------+
        |               |               |
        v               v               v
     NETWORK          STORAGE          CONFIG
        |               |               |
     Service           PVC           Secret
     Endpoint           PV            ConfigMap
     DNS                SC
     Port               CSI
        |               |
        +-------+-------+
                |
                v
            ROOT CAUSE
                |
                v
               FIX
                |
                v
             VERIFY
Enter fullscreen mode Exit fullscreen mode

Do not restart first. Investigate first.

How I would run the three hours

First 30–40 minutes: build the MySQL StatefulSet, create real database data, delete mysql-0, and prove persistence.

Next 60 minutes: Service selector → endpoints → DNS → wrong targetPort → readiness → ImagePullBackOff. Make the students tell you which command to run before you run it.

10-minute break.

Next 35–40 minutes: PVC Pending + CrashLoopBackOff + storage troubleshooting.

Final 30–40 minutes: turn on Argo CD ownership, demonstrate drift, then finish with the bad Git release causing a production outage. Make the students diagnose and recover it themselves.

One important production distinction to say during class: a StatefulSet gives each Pod stable identity and can associate it with stable storage, but that alone does not give MySQL database replication or high availability. This lab deliberately uses one MySQL replica so students don't leave believing that replicas: 3 automatically creates a valid MySQL HA cluster. Kubernetes documents StatefulSet's stable network/storage identity behavior explicitly. :chatgpt-content-reference{index="5"}

Top comments (0)