StatefulSet + PostgreSQL + Helm + Argo CD + Production Troubleshooting
Duration: 3 hours
Level: Intermediate → Advanced
Role: DevOps / SRE Engineer
🎯 LAB OBJECTIVE
You joined the production support rotation.
The Kubernetes platform already exists.
The company already uses:
- Kubernetes
- StatefulSet
- Deployment
- PostgreSQL
- PV/PVC
- Helm
- Argo CD
- GitOps
You are NOT building everything from scratch.
A developer released a new application version.
After the release, production started experiencing problems.
Your responsibility is to:
OBSERVE
↓
IDENTIFY THE SYMPTOM
↓
COLLECT EVIDENCE
↓
ISOLATE THE LAYER
↓
FIND ROOT CAUSE
↓
FIX IN GIT / HELM
↓
ARGO CD SYNC
↓
VERIFY
⚠️ PRODUCTION RULE
Do NOT immediately run:
kubectl delete pod
Do NOT immediately restart everything.
Do NOT immediately change YAML.
First ask:
WHAT IS THE SYMPTOM?
WHAT CHANGED?
WHICH LAYER IS FAILING?
WHAT EVIDENCE DO I HAVE?
🏗️ PRODUCTION ARCHITECTURE
USER
|
v
Load Balancer
|
v
Web Service
|
v
Web Deployment
3 replicas
|
v
Backend / API
|
v
PostgreSQL Service
|
v
PostgreSQL StatefulSet
|
v
postgres-0
|
v
PVC
|
v
PV
|
v
Persistent Storage
Git Repository
|
v
Helm
|
v
Argo CD
|
v
Kubernetes
🧠 FIRST INTERVIEW QUESTION
Why is the frontend a Deployment but PostgreSQL a StatefulSet?
Deployment
Frontend Pods are replaceable.
web-abc123 dies
↓
Kubernetes creates
web-xyz456
We normally do not care about the individual Pod identity.
StatefulSet
PostgreSQL has state.
postgres-0
|
v
PVC
|
v
PV
|
v
DATA
The Pod may disappear.
The data must NOT disappear.
StatefulSet also provides stable Pod identity and ordered behavior.
🧠 SECOND INTERVIEW QUESTION
Does every application need a PVC?
NO.
Example:
Frontend
User:
opens page
clicks button
browses products
changes page
The frontend itself may not need persistent local storage.
If the Pod disappears:
Pod gone
↓
new Pod
↓
application continues
But persistent business data such as:
orders
customers
inventory
transactions
must be stored somewhere durable.
That could be:
PostgreSQL StatefulSet + PVC
or an external managed database such as:
Amazon RDS
🛠️ TROUBLESHOOTING TOOLBOX
Before starting the incidents, remember these commands.
Cluster
kubectl get nodes
Pods
kubectl get pods -n production
kubectl get pods -n production -o wide
Describe
kubectl describe pod <POD> -n production
Logs
kubectl logs <POD> -n production
Previous crashed container:
kubectl logs <POD> -n production --previous
Events
kubectl get events \
-n production \
--sort-by=.metadata.creationTimestamp
Services
kubectl get svc -n production
Endpoints
kubectl get endpoints -n production
StatefulSet
kubectl get statefulset -n production
kubectl describe statefulset postgres -n production
Storage
kubectl get pvc -n production
kubectl get pv
kubectl get sc
Deployment
kubectl get deployment -n production
kubectl rollout status deployment/web -n production
kubectl rollout history deployment/web -n production
🚨 INCIDENT 1 — Developer Released a Bad Image
At 2:00 PM, the developer releases a new version.
Five minutes later:
PRODUCTION ALERT
Application unavailable.
Run:
kubectl get pods -n production
You discover:
web-7f8d9c 0/1 ImagePullBackOff
Your job
Do NOT immediately fix it.
Investigate.
kubectl describe pod <BROKEN-POD> -n production
Look at:
Events
Also check:
kubectl get events \
-n production \
--sort-by=.metadata.creationTimestamp
Interview Question
What does ImagePullBackOff mean?
Think about:
wrong image
wrong tag
private registry
authentication
imagePullSecret
ECR permissions
network access
🚨 INCIDENT 2 — Pod Is Running, Application Is Down
Developer says:
My Pod is Running. Kubernetes is fine.
You run:
kubectl get pods -n production
Result:
NAME READY STATUS
web-xxxxx 0/1 Running
Question:
Is this Pod healthy?
NO.
Remember:
RUNNING != READY
Investigate:
kubectl describe pod <POD> -n production
Look for:
Readiness probe failed
Possible developer mistake:
readinessProbe:
httpGet:
path: /health-does-not-exist
port: 80
But the real application endpoint is:
/
🧠 INTERVIEW QUESTION
What is the difference between readiness and liveness?
Readiness
Can this Pod receive traffic?
Failure:
Pod remains Running
BUT
Service stops sending traffic to it
Liveness
Is the application still alive?
Repeated failure can cause Kubernetes to restart the container.
🚨 INCIDENT 3 — All Pods Are Running But Website Is Down
This is one of the most important production scenarios.
You run:
kubectl get pods -n production
Everything shows:
1/1 Running
But customers say:
WEBSITE DOWN
What do you check next?
kubectl get svc -n production
Then:
kubectl get endpoints -n production
You discover:
web
ENDPOINTS: <none>
Now compare:
kubectl get svc web -n production -o yaml
with:
kubectl get pods -n production --show-labels
Developer accidentally changed:
labels:
app: frontend
while Service still expects:
selector:
app: web
Traffic flow is broken:
User
|
v
Service
|
X
|
Pods
🧠 INTERVIEW QUESTION
Pods are Running but Service has no endpoints. What do you check?
Answer:
Service selector
↓
Pod labels
They must match.
🚨 INCIDENT 4 — Application Cannot Connect to Database
Frontend works.
Application starts.
But when users try to save something:
ERROR
Database connection failed
First:
kubectl get pods -n production
PostgreSQL:
postgres-0 1/1 Running
Do NOT conclude:
Database is fine.
Check Service:
kubectl get svc -n production
Check endpoints:
kubectl get endpoints postgres -n production
Check DNS from another Pod:
kubectl run debug \
--rm -it \
--image=busybox:1.36 \
-n production \
-- sh
Inside:
nslookup postgres
Test port:
nc -zv postgres 5432
🚨 INCIDENT 5 — Developer Changed Database Configuration
Developer changes:
DB_HOST: postgres
to:
DB_HOST: postgres-production
But Kubernetes Service is actually:
postgres
Application logs show database connection failures.
Investigate:
kubectl logs <APPLICATION-POD> -n production
Then inspect configuration:
kubectl get configmap -n production
kubectl get secret -n production
kubectl describe pod <APPLICATION-POD> -n production
🧠 INTERVIEW QUESTION
Pod can start but cannot connect to database. What do you check?
Think layer by layer:
Application logs
↓
Environment variables
↓
Secret / ConfigMap
↓
DNS
↓
Service
↓
Endpoints
↓
Port
↓
NetworkPolicy / Firewall
↓
Database authentication
🚨 INCIDENT 6 — PostgreSQL Pod Is Pending
Run:
kubectl get pods -n production
Result:
postgres-0 0/1 Pending
Do NOT delete it.
Run:
kubectl describe pod postgres-0 -n production
Then:
kubectl get pvc -n production
You discover:
postgres-data-postgres-0 Pending
Now investigate:
kubectl describe pvc postgres-data-postgres-0 -n production
kubectl get sc
Possible developer mistake:
storageClassName: gp3-wrong
but the cluster does not have:
gp3-wrong
🧠 INTERVIEW QUESTION
Pod is Pending. What can cause it?
Do NOT immediately say storage.
Possible causes include:
Insufficient CPU
Insufficient memory
Node selector
Affinity
Taints / tolerations
PVC Pending
StorageClass problem
Volume topology
Scheduling constraints
Use:
kubectl describe pod <POD>
The Events section usually provides important evidence.
🚨 INCIDENT 7 — StatefulSet Pod Was Deleted
Someone runs:
kubectl delete pod postgres-0 -n production
Students may panic.
Watch:
kubectl get pods -n production -w
StatefulSet recreates:
postgres-0
Now check:
kubectl get pvc -n production
PVC remains.
Connect:
kubectl exec -it postgres-0 \
-n production \
-- psql \
-U appuser \
-d appdb
Run:
SELECT * FROM students;
The previous data should still exist.
🧠 INTERVIEW QUESTION
Why didn't deleting the Pod delete our database?
Because:
Pod lifecycle
and:
Persistent Volume lifecycle
are different.
Conceptually:
postgres-0 deleted
↓
StatefulSet recreates postgres-0
↓
PVC still exists
↓
Persistent storage still exists
↓
Database sees existing data
🚨 INCIDENT 8 — CrashLoopBackOff
Developer releases a configuration change.
Now:
kubectl get pods -n production
shows:
CrashLoopBackOff
First commands:
kubectl describe pod <POD> -n production
kubectl logs <POD> -n production
And extremely important:
kubectl logs <POD> \
-n production \
--previous
Possible causes:
missing environment variable
wrong command
application exception
database unavailable
bad configuration
permissions
bad Secret
🚨 INCIDENT 9 — OOMKilled
Developer changed memory limits:
resources:
limits:
memory: "50Mi"
Application requires more memory.
Pod starts restarting.
Investigate:
kubectl describe pod <POD> -n production
Look for:
Reason: OOMKilled
Check:
kubectl top pod -n production
if Metrics Server is installed.
🧠 INTERVIEW QUESTION
Requests vs Limits
Requests
Used primarily by the scheduler when deciding where the Pod can run.
Limits
Maximum resource usage allowed for the container.
For memory, exceeding the limit can result in:
OOMKilled
🚨 INCIDENT 10 — HELM PRODUCTION FAILURE
Now we move from raw Kubernetes YAML to the way many companies manage applications.
Structure:
restaurant-company/
│
├── Chart.yaml
├── values.yaml
│
└── templates/
├── deployment.yaml
├── service.yaml
├── statefulset.yaml
└── configmap.yaml
Developer normally does NOT rewrite every YAML file.
They may change:
image:
repository: nginx
tag: "1.27"
replicaCount: 3
service:
port: 80
database:
host: postgres
port: 5432
persistence:
enabled: true
size: 2Gi
Developer accidentally commits:
service:
port: 8080
But the container listens on:
80
Before deploying, inspect rendered manifests:
helm template production-app .
Validate:
helm lint .
Compare values:
helm get values production-app -n production
Check releases:
helm list -n production
History:
helm history production-app -n production
🧠 INTERVIEW QUESTION
Why use Helm if Kubernetes YAML already works?
Because Helm allows teams to reuse templates and change environment-specific configuration through values.
Example:
Same Helm Chart
|
+------ values-dev.yaml
|
+------ values-stage.yaml
|
+------ values-prod.yaml
🚨 INCIDENT 11 — ARGO CD SHOWS OUTOFSYNC
Now imagine someone executes:
kubectl edit deployment web -n production
and changes:
replicas: 5
But Git says:
replicaCount: 3
Now:
GIT
3 replicas
KUBERNETES
5 replicas
Argo CD detects:
OutOfSync
🧠 INTERVIEW QUESTION
What is configuration drift?
Configuration drift occurs when the live environment differs from the desired configuration stored in Git.
With GitOps:
Git = desired state
The correct workflow is generally:
Change code/config
↓
git add
↓
git commit
↓
git push
↓
Argo CD detects change
↓
Sync
↓
Kubernetes updated
🚨 INCIDENT 12 — ARGO CD SYNCED BUT APPLICATION IS BROKEN
This is a critical interview concept.
Argo CD shows:
Synced
Students say:
Everything is working!
Wrong.
Synced means the live Kubernetes configuration matches the desired Git configuration.
Git itself may contain a bad configuration.
For example:
database:
host: wrong-postgres
Argo CD can successfully deploy it.
Therefore:
Synced ≠ Application Healthy
Investigate:
kubectl get pods -n production
kubectl logs <POD> -n production
kubectl get svc -n production
kubectl get endpoints -n production
🚨 INCIDENT 13 — MANUAL FIX KEEPS DISAPPEARING
Engineer discovers a problem and runs:
kubectl edit deployment web -n production
They fix the problem.
Five minutes later:
THE PROBLEM RETURNS
Why?
Because Git still contains the incorrect desired state.
Argo CD reconciles the cluster back to Git.
🧠 GOLDEN GITOPS RULE
In an Argo CD managed production environment:
DO NOT ONLY FIX THE CLUSTER.
FIX THE SOURCE OF TRUTH.
Usually:
values.yaml
↓
Git commit
↓
Git push
↓
Argo CD
↓
Kubernetes
Emergency procedures vary by company, but the durable fix must be reconciled with the source of truth.
🚨 FINAL INCIDENT — 3:00 AM PRODUCTION OUTAGE
This is the final interview simulation.
You receive:
SEV-1 PRODUCTION INCIDENT
Customers cannot place orders.
Website opens successfully.
Browsing works.
Checkout fails.
Deployment happened 10 minutes ago.
You are the on-call DevOps/SRE engineer.
Nobody tells you what is broken.
STEP 1 — CHECK APPLICATION
kubectl get pods -n production
Ask:
Are Pods Running?
Are they Ready?
Any restarts?
STEP 2 — CHECK EVENTS
kubectl get events \
-n production \
--sort-by=.metadata.creationTimestamp
STEP 3 — CHECK APPLICATION LOGS
kubectl logs <APPLICATION-POD> -n production
Previous container:
kubectl logs <APPLICATION-POD> \
-n production \
--previous
STEP 4 — CHECK DATABASE
kubectl get statefulset -n production
kubectl get pods -n production
kubectl get pvc -n production
STEP 5 — CHECK NETWORKING
kubectl get svc -n production
kubectl get endpoints -n production
Test DNS:
nslookup postgres
Test port:
nc -zv postgres 5432
STEP 6 — CHECK CONFIGURATION
kubectl get configmap -n production
kubectl describe pod <APPLICATION-POD> -n production
Ask:
Did DB_HOST change?
Did DB_PORT change?
Did Secret reference change?
Did Service name change?
STEP 7 — CHECK HELM
helm list -n production
helm get values production-app -n production
helm history production-app -n production
Ask:
What changed between releases?
STEP 8 — CHECK ARGO CD
Ask:
Synced?
OutOfSync?
Healthy?
Degraded?
What commit was deployed?
Then inspect Git history:
git log --oneline -10
git show HEAD
or compare commits:
git diff HEAD~1 HEAD
🎯 STUDENT FINAL REPORT
Every student must report the incident using this structure:
1. SYMPTOM
2. IMPACT
3. WHAT CHANGED?
4. EVIDENCE
5. FAILED LAYER
6. ROOT CAUSE
7. FIX
8. VERIFICATION
9. PREVENTION
Example:
SYMPTOM:
Checkout failed after deployment.
IMPACT:
Customers could browse products but could not place orders.
EVIDENCE:
Application logs showed database hostname resolution failure.
ROOT CAUSE:
Developer changed database.host in values.yaml
from postgres to postgres-production.
FIX:
Corrected values.yaml and committed the change to Git.
DEPLOYMENT:
Argo CD synchronized the corrected desired state.
VERIFICATION:
Application became Ready.
Database connection succeeded.
Checkout test succeeded.
PREVENTION:
Add Helm validation and automated integration tests
before production deployment.
🎤 RAPID-FIRE INTERVIEW ROUND
Teacher asks students without giving them commands.
Question 1
Pod is Pending.
What do you run?
kubectl describe pod <pod>
Question 2
Pod is CrashLoopBackOff.
What do you check?
kubectl logs <pod>
kubectl logs <pod> --previous
kubectl describe pod <pod>
Question 3
Pod is Running but 0/1.
Think:
Readiness
Question 4
All Pods are Running but application is inaccessible.
Think:
Service
Endpoints
Selectors
Labels
Ingress / Load Balancer
Network
Question 5
Service has no endpoints.
Think:
selector != Pod labels
Question 6
StatefulSet Pod is Pending.
Think:
Scheduling
PVC
PV
StorageClass
Events
Question 7
Database Pod was deleted. Will data necessarily disappear?
No.
Check persistent storage and PVC/PV lifecycle.
Question 8
Argo CD says Synced. Does that prove the application works?
No.
Synced != Healthy application
Question 9
You manually fixed production, but Argo CD changed it back.
Why?
Git still contains the desired state Argo CD is reconciling.
Question 10
Where should the permanent fix go?
SOURCE OF TRUTH
Git / Helm values
🏆 FINAL PRODUCTION TROUBLESHOOTING MODEL
Memorize this:
USER REPORTS PROBLEM
↓
WHAT CHANGED?
↓
kubectl get
↓
kubectl describe
↓
kubectl logs
↓
events
↓
Service / Endpoints / DNS
↓
ConfigMap / Secret
↓
StatefulSet / PVC / PV
↓
Helm values
↓
Git diff
↓
Argo CD
↓
ROOT CAUSE
↓
FIX SOURCE OF TRUTH
↓
SYNC / DEPLOY
↓
VERIFY APPLICATION
↓
MONITOR
↓
POSTMORTEM / PREVENTION
🎓 WHAT YOU SHOULD BE ABLE TO EXPLAIN AFTER THIS LAB
By the end of this lab you should be comfortable explaining:
- Deployment vs StatefulSet
- Stateless vs stateful applications
- StatefulSet Pod identity
- PV vs PVC
- StorageClass
- Data persistence after Pod deletion
- Headless Service
- Service selectors
- Pod labels
- Kubernetes endpoints
- Kubernetes DNS
- Readiness vs liveness
- Pending Pods
- CrashLoopBackOff
- ImagePullBackOff
- OOMKilled
- ConfigMap vs Secret
- Database connectivity troubleshooting
- Helm charts
values.yamlhelm linthelm template- Helm release history
- Argo CD
- GitOps
- Desired state
- Configuration drift
-
Syncedvs application health - Why manual production changes can be reverted
- How to investigate a production incident
- Root cause analysis
- How to verify recovery
⭐ MOST IMPORTANT LESSON
A DevOps Engineer is not valuable because they can memorize:
kubectl apply -f deployment.yaml
The important skill is being able to answer:
The application worked yesterday.
A deployment happened.
Now production is broken.
WHAT CHANGED?
WHERE IS THE FAILURE?
WHAT EVIDENCE PROVES IT?
HOW DO I FIX IT SAFELY?
HOW DO I KNOW IT IS REALLY FIXED?
That is production troubleshooting.
This complements the lab you already posted particularly well because your existing article already establishes the healthy PostgreSQL StatefulSet, persistent-storage test, and initial incidents. :chatgpt-content-reference{index="1"}
One teaching change I'd make: you should introduce the mistakes secretly. Don't show students “change this line to the wrong value.” Give them only the incident ticket. That turns it from a Kubernetes tutorial into a realistic DevOps/SRE troubleshooting interview simulation.
Top comments (0)