DEV Community

Aisalkyn Aidarova
Aisalkyn Aidarova

Posted on

JumpToTech DevOps — Production Troubleshooting Lab

StatefulSet + PostgreSQL + Helm + Argo CD + Production Troubleshooting

Duration: 3 hours

Level: Intermediate → Advanced

Role: DevOps / SRE Engineer


🎯 LAB OBJECTIVE

You joined the production support rotation.

The Kubernetes platform already exists.

The company already uses:

  • Kubernetes
  • StatefulSet
  • Deployment
  • PostgreSQL
  • PV/PVC
  • Helm
  • Argo CD
  • GitOps

You are NOT building everything from scratch.

A developer released a new application version.

After the release, production started experiencing problems.

Your responsibility is to:

OBSERVE
   ↓
IDENTIFY THE SYMPTOM
   ↓
COLLECT EVIDENCE
   ↓
ISOLATE THE LAYER
   ↓
FIND ROOT CAUSE
   ↓
FIX IN GIT / HELM
   ↓
ARGO CD SYNC
   ↓
VERIFY
Enter fullscreen mode Exit fullscreen mode

⚠️ PRODUCTION RULE

Do NOT immediately run:

kubectl delete pod
Enter fullscreen mode Exit fullscreen mode

Do NOT immediately restart everything.

Do NOT immediately change YAML.

First ask:

WHAT IS THE SYMPTOM?

WHAT CHANGED?

WHICH LAYER IS FAILING?

WHAT EVIDENCE DO I HAVE?
Enter fullscreen mode Exit fullscreen mode

🏗️ PRODUCTION ARCHITECTURE

                        USER
                         |
                         v
                   Load Balancer
                         |
                         v
                    Web Service
                         |
                         v
                 Web Deployment
                    3 replicas
                         |
                         v
                  Backend / API
                         |
                         v
                PostgreSQL Service
                         |
                         v
               PostgreSQL StatefulSet
                         |
                         v
                     postgres-0
                         |
                         v
                        PVC
                         |
                         v
                         PV
                         |
                         v
                Persistent Storage


Git Repository
      |
      v
    Helm
      |
      v
   Argo CD
      |
      v
 Kubernetes
Enter fullscreen mode Exit fullscreen mode

🧠 FIRST INTERVIEW QUESTION

Why is the frontend a Deployment but PostgreSQL a StatefulSet?

Deployment

Frontend Pods are replaceable.

web-abc123 dies

        ↓

Kubernetes creates

web-xyz456
Enter fullscreen mode Exit fullscreen mode

We normally do not care about the individual Pod identity.


StatefulSet

PostgreSQL has state.

postgres-0
     |
     v
PVC
     |
     v
PV
     |
     v
DATA
Enter fullscreen mode Exit fullscreen mode

The Pod may disappear.

The data must NOT disappear.

StatefulSet also provides stable Pod identity and ordered behavior.


🧠 SECOND INTERVIEW QUESTION

Does every application need a PVC?

NO.

Example:

Frontend
Enter fullscreen mode Exit fullscreen mode

User:

opens page
clicks button
browses products
changes page
Enter fullscreen mode Exit fullscreen mode

The frontend itself may not need persistent local storage.

If the Pod disappears:

Pod gone
   ↓
new Pod
   ↓
application continues
Enter fullscreen mode Exit fullscreen mode

But persistent business data such as:

orders
customers
inventory
transactions
Enter fullscreen mode Exit fullscreen mode

must be stored somewhere durable.

That could be:

PostgreSQL StatefulSet + PVC
Enter fullscreen mode Exit fullscreen mode

or an external managed database such as:

Amazon RDS
Enter fullscreen mode Exit fullscreen mode

🛠️ TROUBLESHOOTING TOOLBOX

Before starting the incidents, remember these commands.

Cluster

kubectl get nodes
Enter fullscreen mode Exit fullscreen mode

Pods

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get pods -n production -o wide
Enter fullscreen mode Exit fullscreen mode

Describe

kubectl describe pod <POD> -n production
Enter fullscreen mode Exit fullscreen mode

Logs

kubectl logs <POD> -n production
Enter fullscreen mode Exit fullscreen mode

Previous crashed container:

kubectl logs <POD> -n production --previous
Enter fullscreen mode Exit fullscreen mode

Events

kubectl get events \
  -n production \
  --sort-by=.metadata.creationTimestamp
Enter fullscreen mode Exit fullscreen mode

Services

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

Endpoints

kubectl get endpoints -n production
Enter fullscreen mode Exit fullscreen mode

StatefulSet

kubectl get statefulset -n production
Enter fullscreen mode Exit fullscreen mode
kubectl describe statefulset postgres -n production
Enter fullscreen mode Exit fullscreen mode

Storage

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get pv
Enter fullscreen mode Exit fullscreen mode
kubectl get sc
Enter fullscreen mode Exit fullscreen mode

Deployment

kubectl get deployment -n production
Enter fullscreen mode Exit fullscreen mode
kubectl rollout status deployment/web -n production
Enter fullscreen mode Exit fullscreen mode
kubectl rollout history deployment/web -n production
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 1 — Developer Released a Bad Image

At 2:00 PM, the developer releases a new version.

Five minutes later:

PRODUCTION ALERT

Application unavailable.
Enter fullscreen mode Exit fullscreen mode

Run:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

You discover:

web-7f8d9c      0/1    ImagePullBackOff
Enter fullscreen mode Exit fullscreen mode

Your job

Do NOT immediately fix it.

Investigate.

kubectl describe pod <BROKEN-POD> -n production
Enter fullscreen mode Exit fullscreen mode

Look at:

Events
Enter fullscreen mode Exit fullscreen mode

Also check:

kubectl get events \
  -n production \
  --sort-by=.metadata.creationTimestamp
Enter fullscreen mode Exit fullscreen mode

Interview Question

What does ImagePullBackOff mean?

Think about:

wrong image
wrong tag
private registry
authentication
imagePullSecret
ECR permissions
network access
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 2 — Pod Is Running, Application Is Down

Developer says:

My Pod is Running. Kubernetes is fine.

You run:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Result:

NAME                READY     STATUS
web-xxxxx            0/1      Running
Enter fullscreen mode Exit fullscreen mode

Question:

Is this Pod healthy?

NO.

Remember:

RUNNING != READY
Enter fullscreen mode Exit fullscreen mode

Investigate:

kubectl describe pod <POD> -n production
Enter fullscreen mode Exit fullscreen mode

Look for:

Readiness probe failed
Enter fullscreen mode Exit fullscreen mode

Possible developer mistake:

readinessProbe:
  httpGet:
    path: /health-does-not-exist
    port: 80
Enter fullscreen mode Exit fullscreen mode

But the real application endpoint is:

/
Enter fullscreen mode Exit fullscreen mode

🧠 INTERVIEW QUESTION

What is the difference between readiness and liveness?

Readiness

Can this Pod receive traffic?
Enter fullscreen mode Exit fullscreen mode

Failure:

Pod remains Running
BUT
Service stops sending traffic to it
Enter fullscreen mode Exit fullscreen mode

Liveness

Is the application still alive?
Enter fullscreen mode Exit fullscreen mode

Repeated failure can cause Kubernetes to restart the container.


🚨 INCIDENT 3 — All Pods Are Running But Website Is Down

This is one of the most important production scenarios.

You run:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Everything shows:

1/1 Running
Enter fullscreen mode Exit fullscreen mode

But customers say:

WEBSITE DOWN
Enter fullscreen mode Exit fullscreen mode

What do you check next?

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get endpoints -n production
Enter fullscreen mode Exit fullscreen mode

You discover:

web
ENDPOINTS: <none>
Enter fullscreen mode Exit fullscreen mode

Now compare:

kubectl get svc web -n production -o yaml
Enter fullscreen mode Exit fullscreen mode

with:

kubectl get pods -n production --show-labels
Enter fullscreen mode Exit fullscreen mode

Developer accidentally changed:

labels:
  app: frontend
Enter fullscreen mode Exit fullscreen mode

while Service still expects:

selector:
  app: web
Enter fullscreen mode Exit fullscreen mode

Traffic flow is broken:

User
 |
 v
Service
 |
 X
 |
Pods
Enter fullscreen mode Exit fullscreen mode

🧠 INTERVIEW QUESTION

Pods are Running but Service has no endpoints. What do you check?

Answer:

Service selector
      ↓
Pod labels
Enter fullscreen mode Exit fullscreen mode

They must match.


🚨 INCIDENT 4 — Application Cannot Connect to Database

Frontend works.

Application starts.

But when users try to save something:

ERROR

Database connection failed
Enter fullscreen mode Exit fullscreen mode

First:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

PostgreSQL:

postgres-0    1/1    Running
Enter fullscreen mode Exit fullscreen mode

Do NOT conclude:

Database is fine.
Enter fullscreen mode Exit fullscreen mode

Check Service:

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode

Check endpoints:

kubectl get endpoints postgres -n production
Enter fullscreen mode Exit fullscreen mode

Check DNS from another Pod:

kubectl run debug \
  --rm -it \
  --image=busybox:1.36 \
  -n production \
  -- sh
Enter fullscreen mode Exit fullscreen mode

Inside:

nslookup postgres
Enter fullscreen mode Exit fullscreen mode

Test port:

nc -zv postgres 5432
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 5 — Developer Changed Database Configuration

Developer changes:

DB_HOST: postgres
Enter fullscreen mode Exit fullscreen mode

to:

DB_HOST: postgres-production
Enter fullscreen mode Exit fullscreen mode

But Kubernetes Service is actually:

postgres
Enter fullscreen mode Exit fullscreen mode

Application logs show database connection failures.

Investigate:

kubectl logs <APPLICATION-POD> -n production
Enter fullscreen mode Exit fullscreen mode

Then inspect configuration:

kubectl get configmap -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get secret -n production
Enter fullscreen mode Exit fullscreen mode
kubectl describe pod <APPLICATION-POD> -n production
Enter fullscreen mode Exit fullscreen mode

🧠 INTERVIEW QUESTION

Pod can start but cannot connect to database. What do you check?

Think layer by layer:

Application logs
       ↓
Environment variables
       ↓
Secret / ConfigMap
       ↓
DNS
       ↓
Service
       ↓
Endpoints
       ↓
Port
       ↓
NetworkPolicy / Firewall
       ↓
Database authentication
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 6 — PostgreSQL Pod Is Pending

Run:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Result:

postgres-0     0/1     Pending
Enter fullscreen mode Exit fullscreen mode

Do NOT delete it.

Run:

kubectl describe pod postgres-0 -n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

You discover:

postgres-data-postgres-0     Pending
Enter fullscreen mode Exit fullscreen mode

Now investigate:

kubectl describe pvc postgres-data-postgres-0 -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get sc
Enter fullscreen mode Exit fullscreen mode

Possible developer mistake:

storageClassName: gp3-wrong
Enter fullscreen mode Exit fullscreen mode

but the cluster does not have:

gp3-wrong
Enter fullscreen mode Exit fullscreen mode

🧠 INTERVIEW QUESTION

Pod is Pending. What can cause it?

Do NOT immediately say storage.

Possible causes include:

Insufficient CPU
Insufficient memory
Node selector
Affinity
Taints / tolerations
PVC Pending
StorageClass problem
Volume topology
Scheduling constraints
Enter fullscreen mode Exit fullscreen mode

Use:

kubectl describe pod <POD>
Enter fullscreen mode Exit fullscreen mode

The Events section usually provides important evidence.


🚨 INCIDENT 7 — StatefulSet Pod Was Deleted

Someone runs:

kubectl delete pod postgres-0 -n production
Enter fullscreen mode Exit fullscreen mode

Students may panic.

Watch:

kubectl get pods -n production -w
Enter fullscreen mode Exit fullscreen mode

StatefulSet recreates:

postgres-0
Enter fullscreen mode Exit fullscreen mode

Now check:

kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

PVC remains.

Connect:

kubectl exec -it postgres-0 \
  -n production \
  -- psql \
  -U appuser \
  -d appdb
Enter fullscreen mode Exit fullscreen mode

Run:

SELECT * FROM students;
Enter fullscreen mode Exit fullscreen mode

The previous data should still exist.


🧠 INTERVIEW QUESTION

Why didn't deleting the Pod delete our database?

Because:

Pod lifecycle
Enter fullscreen mode Exit fullscreen mode

and:

Persistent Volume lifecycle
Enter fullscreen mode Exit fullscreen mode

are different.

Conceptually:

postgres-0 deleted
       ↓
StatefulSet recreates postgres-0
       ↓
PVC still exists
       ↓
Persistent storage still exists
       ↓
Database sees existing data
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 8 — CrashLoopBackOff

Developer releases a configuration change.

Now:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

shows:

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

First commands:

kubectl describe pod <POD> -n production
Enter fullscreen mode Exit fullscreen mode
kubectl logs <POD> -n production
Enter fullscreen mode Exit fullscreen mode

And extremely important:

kubectl logs <POD> \
  -n production \
  --previous
Enter fullscreen mode Exit fullscreen mode

Possible causes:

missing environment variable
wrong command
application exception
database unavailable
bad configuration
permissions
bad Secret
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 9 — OOMKilled

Developer changed memory limits:

resources:
  limits:
    memory: "50Mi"
Enter fullscreen mode Exit fullscreen mode

Application requires more memory.

Pod starts restarting.

Investigate:

kubectl describe pod <POD> -n production
Enter fullscreen mode Exit fullscreen mode

Look for:

Reason: OOMKilled
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl top pod -n production
Enter fullscreen mode Exit fullscreen mode

if Metrics Server is installed.


🧠 INTERVIEW QUESTION

Requests vs Limits

Requests

Used primarily by the scheduler when deciding where the Pod can run.

Limits

Maximum resource usage allowed for the container.

For memory, exceeding the limit can result in:

OOMKilled
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 10 — HELM PRODUCTION FAILURE

Now we move from raw Kubernetes YAML to the way many companies manage applications.

Structure:

restaurant-company/
│
├── Chart.yaml
├── values.yaml
│
└── templates/
    ├── deployment.yaml
    ├── service.yaml
    ├── statefulset.yaml
    └── configmap.yaml
Enter fullscreen mode Exit fullscreen mode

Developer normally does NOT rewrite every YAML file.

They may change:

image:
  repository: nginx
  tag: "1.27"

replicaCount: 3

service:
  port: 80

database:
  host: postgres
  port: 5432

persistence:
  enabled: true
  size: 2Gi
Enter fullscreen mode Exit fullscreen mode

Developer accidentally commits:

service:
  port: 8080
Enter fullscreen mode Exit fullscreen mode

But the container listens on:

80
Enter fullscreen mode Exit fullscreen mode

Before deploying, inspect rendered manifests:

helm template production-app .
Enter fullscreen mode Exit fullscreen mode

Validate:

helm lint .
Enter fullscreen mode Exit fullscreen mode

Compare values:

helm get values production-app -n production
Enter fullscreen mode Exit fullscreen mode

Check releases:

helm list -n production
Enter fullscreen mode Exit fullscreen mode

History:

helm history production-app -n production
Enter fullscreen mode Exit fullscreen mode

🧠 INTERVIEW QUESTION

Why use Helm if Kubernetes YAML already works?

Because Helm allows teams to reuse templates and change environment-specific configuration through values.

Example:

Same Helm Chart
      |
      +------ values-dev.yaml
      |
      +------ values-stage.yaml
      |
      +------ values-prod.yaml
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 11 — ARGO CD SHOWS OUTOFSYNC

Now imagine someone executes:

kubectl edit deployment web -n production
Enter fullscreen mode Exit fullscreen mode

and changes:

replicas: 5
Enter fullscreen mode Exit fullscreen mode

But Git says:

replicaCount: 3
Enter fullscreen mode Exit fullscreen mode

Now:

GIT
3 replicas

KUBERNETES
5 replicas
Enter fullscreen mode Exit fullscreen mode

Argo CD detects:

OutOfSync
Enter fullscreen mode Exit fullscreen mode

🧠 INTERVIEW QUESTION

What is configuration drift?

Configuration drift occurs when the live environment differs from the desired configuration stored in Git.

With GitOps:

Git = desired state
Enter fullscreen mode Exit fullscreen mode

The correct workflow is generally:

Change code/config
       ↓
git add
       ↓
git commit
       ↓
git push
       ↓
Argo CD detects change
       ↓
Sync
       ↓
Kubernetes updated
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 12 — ARGO CD SYNCED BUT APPLICATION IS BROKEN

This is a critical interview concept.

Argo CD shows:

Synced
Enter fullscreen mode Exit fullscreen mode

Students say:

Everything is working!

Wrong.

Synced means the live Kubernetes configuration matches the desired Git configuration.

Git itself may contain a bad configuration.

For example:

database:
  host: wrong-postgres
Enter fullscreen mode Exit fullscreen mode

Argo CD can successfully deploy it.

Therefore:

Synced ≠ Application Healthy
Enter fullscreen mode Exit fullscreen mode

Investigate:

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode
kubectl logs <POD> -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get endpoints -n production
Enter fullscreen mode Exit fullscreen mode

🚨 INCIDENT 13 — MANUAL FIX KEEPS DISAPPEARING

Engineer discovers a problem and runs:

kubectl edit deployment web -n production
Enter fullscreen mode Exit fullscreen mode

They fix the problem.

Five minutes later:

THE PROBLEM RETURNS
Enter fullscreen mode Exit fullscreen mode

Why?

Because Git still contains the incorrect desired state.

Argo CD reconciles the cluster back to Git.


🧠 GOLDEN GITOPS RULE

In an Argo CD managed production environment:

DO NOT ONLY FIX THE CLUSTER.

FIX THE SOURCE OF TRUTH.
Enter fullscreen mode Exit fullscreen mode

Usually:

values.yaml
       ↓
Git commit
       ↓
Git push
       ↓
Argo CD
       ↓
Kubernetes
Enter fullscreen mode Exit fullscreen mode

Emergency procedures vary by company, but the durable fix must be reconciled with the source of truth.


🚨 FINAL INCIDENT — 3:00 AM PRODUCTION OUTAGE

This is the final interview simulation.

You receive:

SEV-1 PRODUCTION INCIDENT

Customers cannot place orders.

Website opens successfully.

Browsing works.

Checkout fails.

Deployment happened 10 minutes ago.
Enter fullscreen mode Exit fullscreen mode

You are the on-call DevOps/SRE engineer.

Nobody tells you what is broken.


STEP 1 — CHECK APPLICATION

kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode

Ask:

Are Pods Running?

Are they Ready?

Any restarts?
Enter fullscreen mode Exit fullscreen mode

STEP 2 — CHECK EVENTS

kubectl get events \
  -n production \
  --sort-by=.metadata.creationTimestamp
Enter fullscreen mode Exit fullscreen mode

STEP 3 — CHECK APPLICATION LOGS

kubectl logs <APPLICATION-POD> -n production
Enter fullscreen mode Exit fullscreen mode

Previous container:

kubectl logs <APPLICATION-POD> \
  -n production \
  --previous
Enter fullscreen mode Exit fullscreen mode

STEP 4 — CHECK DATABASE

kubectl get statefulset -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get pods -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get pvc -n production
Enter fullscreen mode Exit fullscreen mode

STEP 5 — CHECK NETWORKING

kubectl get svc -n production
Enter fullscreen mode Exit fullscreen mode
kubectl get endpoints -n production
Enter fullscreen mode Exit fullscreen mode

Test DNS:

nslookup postgres
Enter fullscreen mode Exit fullscreen mode

Test port:

nc -zv postgres 5432
Enter fullscreen mode Exit fullscreen mode

STEP 6 — CHECK CONFIGURATION

kubectl get configmap -n production
Enter fullscreen mode Exit fullscreen mode
kubectl describe pod <APPLICATION-POD> -n production
Enter fullscreen mode Exit fullscreen mode

Ask:

Did DB_HOST change?

Did DB_PORT change?

Did Secret reference change?

Did Service name change?
Enter fullscreen mode Exit fullscreen mode

STEP 7 — CHECK HELM

helm list -n production
Enter fullscreen mode Exit fullscreen mode
helm get values production-app -n production
Enter fullscreen mode Exit fullscreen mode
helm history production-app -n production
Enter fullscreen mode Exit fullscreen mode

Ask:

What changed between releases?
Enter fullscreen mode Exit fullscreen mode

STEP 8 — CHECK ARGO CD

Ask:

Synced?

OutOfSync?

Healthy?

Degraded?

What commit was deployed?
Enter fullscreen mode Exit fullscreen mode

Then inspect Git history:

git log --oneline -10
Enter fullscreen mode Exit fullscreen mode
git show HEAD
Enter fullscreen mode Exit fullscreen mode

or compare commits:

git diff HEAD~1 HEAD
Enter fullscreen mode Exit fullscreen mode

🎯 STUDENT FINAL REPORT

Every student must report the incident using this structure:

1. SYMPTOM

2. IMPACT

3. WHAT CHANGED?

4. EVIDENCE

5. FAILED LAYER

6. ROOT CAUSE

7. FIX

8. VERIFICATION

9. PREVENTION
Enter fullscreen mode Exit fullscreen mode

Example:

SYMPTOM:
Checkout failed after deployment.

IMPACT:
Customers could browse products but could not place orders.

EVIDENCE:
Application logs showed database hostname resolution failure.

ROOT CAUSE:
Developer changed database.host in values.yaml
from postgres to postgres-production.

FIX:
Corrected values.yaml and committed the change to Git.

DEPLOYMENT:
Argo CD synchronized the corrected desired state.

VERIFICATION:
Application became Ready.
Database connection succeeded.
Checkout test succeeded.

PREVENTION:
Add Helm validation and automated integration tests
before production deployment.
Enter fullscreen mode Exit fullscreen mode

🎤 RAPID-FIRE INTERVIEW ROUND

Teacher asks students without giving them commands.

Question 1

Pod is Pending.

What do you run?

kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

Question 2

Pod is CrashLoopBackOff.

What do you check?

kubectl logs <pod>
kubectl logs <pod> --previous
kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

Question 3

Pod is Running but 0/1.

Think:

Readiness
Enter fullscreen mode Exit fullscreen mode

Question 4

All Pods are Running but application is inaccessible.

Think:

Service
Endpoints
Selectors
Labels
Ingress / Load Balancer
Network
Enter fullscreen mode Exit fullscreen mode

Question 5

Service has no endpoints.

Think:

selector != Pod labels
Enter fullscreen mode Exit fullscreen mode

Question 6

StatefulSet Pod is Pending.

Think:

Scheduling
PVC
PV
StorageClass
Events
Enter fullscreen mode Exit fullscreen mode

Question 7

Database Pod was deleted. Will data necessarily disappear?

No.

Check persistent storage and PVC/PV lifecycle.


Question 8

Argo CD says Synced. Does that prove the application works?

No.

Synced != Healthy application
Enter fullscreen mode Exit fullscreen mode

Question 9

You manually fixed production, but Argo CD changed it back.

Why?

Git still contains the desired state Argo CD is reconciling.
Enter fullscreen mode Exit fullscreen mode

Question 10

Where should the permanent fix go?

SOURCE OF TRUTH

Git / Helm values
Enter fullscreen mode Exit fullscreen mode

🏆 FINAL PRODUCTION TROUBLESHOOTING MODEL

Memorize this:

USER REPORTS PROBLEM
        ↓
WHAT CHANGED?
        ↓
kubectl get
        ↓
kubectl describe
        ↓
kubectl logs
        ↓
events
        ↓
Service / Endpoints / DNS
        ↓
ConfigMap / Secret
        ↓
StatefulSet / PVC / PV
        ↓
Helm values
        ↓
Git diff
        ↓
Argo CD
        ↓
ROOT CAUSE
        ↓
FIX SOURCE OF TRUTH
        ↓
SYNC / DEPLOY
        ↓
VERIFY APPLICATION
        ↓
MONITOR
        ↓
POSTMORTEM / PREVENTION
Enter fullscreen mode Exit fullscreen mode

🎓 WHAT YOU SHOULD BE ABLE TO EXPLAIN AFTER THIS LAB

By the end of this lab you should be comfortable explaining:

  • Deployment vs StatefulSet
  • Stateless vs stateful applications
  • StatefulSet Pod identity
  • PV vs PVC
  • StorageClass
  • Data persistence after Pod deletion
  • Headless Service
  • Service selectors
  • Pod labels
  • Kubernetes endpoints
  • Kubernetes DNS
  • Readiness vs liveness
  • Pending Pods
  • CrashLoopBackOff
  • ImagePullBackOff
  • OOMKilled
  • ConfigMap vs Secret
  • Database connectivity troubleshooting
  • Helm charts
  • values.yaml
  • helm lint
  • helm template
  • Helm release history
  • Argo CD
  • GitOps
  • Desired state
  • Configuration drift
  • Synced vs application health
  • Why manual production changes can be reverted
  • How to investigate a production incident
  • Root cause analysis
  • How to verify recovery

⭐ MOST IMPORTANT LESSON

A DevOps Engineer is not valuable because they can memorize:

kubectl apply -f deployment.yaml
Enter fullscreen mode Exit fullscreen mode

The important skill is being able to answer:

The application worked yesterday.

A deployment happened.

Now production is broken.

WHAT CHANGED?

WHERE IS THE FAILURE?

WHAT EVIDENCE PROVES IT?

HOW DO I FIX IT SAFELY?

HOW DO I KNOW IT IS REALLY FIXED?
Enter fullscreen mode Exit fullscreen mode

That is production troubleshooting.

This complements the lab you already posted particularly well because your existing article already establishes the healthy PostgreSQL StatefulSet, persistent-storage test, and initial incidents. :chatgpt-content-reference{index="1"}

One teaching change I'd make: you should introduce the mistakes secretly. Don't show students “change this line to the wrong value.” Give them only the incident ticket. That turns it from a Kubernetes tutorial into a realistic DevOps/SRE troubleshooting interview simulation.

Top comments (0)