Kubernetes is easy when everything is working.
You run:
kubectl get pods
and see:
NAME READY STATUS
frontend-7d8f9c6d7-x2k4 1/1 Running
backend-5f7c8d9b8-m4p2 1/1 Running
Perfect.
But real Kubernetes work begins when you see:
CrashLoopBackOff
ImagePullBackOff
Pending
0/1
Error
Terminating
That's when many beginners start running random commands from Google.
Don't.
Kubernetes troubleshooting is a process, not a collection of magic commands.
In this guide, we'll build a systematic troubleshooting workflow you can use for real Kubernetes clusters and CKA-style problems.
The Golden Rule of Kubernetes Troubleshooting
When something breaks, don't immediately ask:
"What command should I run?"
Ask:
"What is the system telling me?"
Start with evidence.
A useful troubleshooting flow is:
Problem
↓
Observe
↓
Identify the failing layer
↓
Collect evidence
↓
Find the root cause
↓
Fix
↓
Verify
This mindset is far more valuable than memorizing hundreds of commands.
Step 1 — Check the Pod
Start simple:
kubectl get pods
You might see:
NAME READY STATUS
web-7d9f8c6d4-x7k2 1/1 Running
api-5f7d8c9b4-m2p8 0/1 CrashLoopBackOff
Immediately, you know:
web → probably healthy
api → something is wrong
Now investigate the problematic Pod.
Step 2 — Describe the Pod
Run:
kubectl describe pod api-5f7d8c9b4-m2p8
This gives you important information about:
- Node placement
- Containers
- Images
- Environment variables
- Volumes
- Events
- Probes
- Resource configuration
The Events section is particularly useful.
You might see:
Failed to pull image
or:
Back-off restarting failed container
or:
FailedMount
Now you have evidence.
Step 3 — Check the Logs
If the container starts and then crashes, check its logs:
kubectl logs api-5f7d8c9b4-m2p8
You might discover:
Error: DATABASE_URL is missing
Now the problem isn't mysterious anymore.
The application is telling you exactly what it needs.
For a Pod with multiple containers:
kubectl logs <pod-name> -c <container-name>
And if the container has restarted:
kubectl logs <pod-name> --previous
That last command is extremely useful for crash-looping containers.
Understanding CrashLoopBackOff
One of the most common statuses beginners encounter is:
CrashLoopBackOff
It does not mean "Kubernetes is broken."
It generally means the container is repeatedly starting and failing, and Kubernetes is backing off before restarting it again.
Think:
Container starts
↓
Application crashes
↓
Container exits
↓
Kubernetes restarts it
↓
Application crashes again
↓
Backoff
↓
Retry
Your first questions should be:
Why did the process exit?
Check:
kubectl logs <pod>
and:
kubectl logs <pod> --previous
Then inspect:
kubectl describe pod <pod>
Common Cause #1 — Wrong Image
Suppose your Deployment contains:
image: myapp:v99
but that image doesn't exist in your container registry.
You may see:
ImagePullBackOff
or:
ErrImagePull
Check:
kubectl describe pod <pod>
Look at the Events.
You might find:
Failed to pull image
Now investigate:
- Is the image name correct?
- Is the tag correct?
- Does the registry contain the image?
- Does the cluster have permission to pull it?
- Is authentication required?
Don't randomly restart the Pod.
Fix the image problem.
Common Cause #2 — Application Configuration
Your container may start perfectly but immediately crash because configuration is missing.
For example:
DATABASE_HOST
DATABASE_USER
DATABASE_PASSWORD
Your application expects these values.
But your Pod doesn't receive them.
The result might be:
Application starts
↓
Reads configuration
↓
Configuration missing
↓
Application exits
↓
CrashLoopBackOff
Check your:
ConfigMap
Secret
Deployment
Environment variables
You can inspect the Deployment:
kubectl get deployment <name> -o yaml
And inspect configuration resources:
kubectl get configmap
kubectl get secret
Remember that sensitive values should be handled carefully and not casually exposed in terminal output or source repositories.
Common Cause #3 — Readiness Probe Failure
Sometimes your application is actually running.
But Kubernetes doesn't consider it ready.
For example:
Pod:
Running
Ready:
0/1
This can happen because a readiness probe is failing.
Example:
readinessProbe:
httpGet:
path: /health
port: 8080
Kubernetes may repeatedly check:
http://container:8080/health
If the endpoint fails, the Pod may remain unavailable to Service traffic.
This gives us an important distinction:
Running does not always mean Ready.
Liveness vs Readiness
These two concepts are easy to confuse.
Liveness
Asks:
"Is this container still healthy enough to keep running?"
If the liveness probe repeatedly fails, Kubernetes may restart the container.
Readiness
Asks:
"Is this application ready to receive traffic?"
If readiness fails, Kubernetes can stop sending traffic to that Pod while the container continues running.
Think:
Liveness
↓
Should the container keep running?
Readiness
↓
Should the container receive traffic?
Understanding this distinction is extremely important when troubleshooting production workloads.
Common Cause #4 — Pod Is Pending
Now suppose you run:
kubectl get pods
and see:
api-7d9f8c6d4-x8m2 0/1 Pending
The application isn't necessarily broken.
It might not have been scheduled onto a node.
Start with:
kubectl describe pod <pod-name>
Look at Events.
You might discover:
Insufficient cpu
or:
Insufficient memory
or a scheduling constraint is preventing placement.
This leads to an important troubleshooting question:
Is the problem inside the application, or is the Pod unable to be scheduled?
Don't jump into application logs if the container never started.
Common Cause #5 — Service Isn't Working
Your Pods might be perfectly healthy:
Pod 1 → Running
Pod 2 → Running
Pod 3 → Running
But users still can't reach the application.
Now inspect the Service:
kubectl get svc
Then:
kubectl describe svc <service-name>
One important thing to investigate is whether the Service has the expected endpoints.
Check:
kubectl get endpoints
or, depending on your Kubernetes setup:
kubectl get endpointslices
If the Service isn't selecting the correct Pods, you can have:
Healthy Pods
↓
Service
↓
No matching endpoints
The application is running.
The networking configuration is the problem.
Labels Are Extremely Important
Services commonly select Pods using labels.
For example:
selector:
app: backend
Your Pod needs a matching label:
labels:
app: backend
If you accidentally have:
labels:
app: frontend
the Service won't select that Pod.
You can inspect labels with:
kubectl get pods --show-labels
This simple command can save a lot of debugging time.
Common Cause #6 — DNS Problems
Imagine your frontend needs to communicate with:
backend-service
Kubernetes provides service discovery through DNS.
If DNS isn't working or you're using the wrong Service name, communication can fail.
Start by checking:
kubectl get svc
Then verify the Service name and namespace.
For example:
backend-service.default.svc.cluster.local
The exact DNS name depends on the Service and namespace.
A useful debugging technique is launching a temporary Pod with networking tools and testing name resolution or connectivity from inside the cluster.
Common Cause #7 — Wrong Port
This is another classic problem.
Suppose your application listens on:
8080
But your Service sends traffic to:
5000
You could have:
User
↓
Service :80
↓
targetPort :5000
↓
Application listens on :8080
The Pod may be healthy.
The Service may exist.
But traffic still fails.
Always compare:
Service port
targetPort
containerPort
application listening port
These are related concepts, but they are not automatically the same thing.
Common Cause #8 — Volume Mount Problems
Suppose your Pod refuses to start and Events show:
FailedMount
Now investigate:
PersistentVolume
PersistentVolumeClaim
StorageClass
Volume
Mount path
Check:
kubectl get pv
kubectl get pvc
Then:
kubectl describe pvc <name>
The key question is:
What resource is Kubernetes waiting for?
Again, follow the evidence.
Common Cause #9 — RBAC Problems
Sometimes your application or Kubernetes user doesn't have the required permissions.
You might see errors like:
Forbidden
For example:
User cannot get pods
This points toward Kubernetes authorization.
Investigate:
Role
ClusterRole
RoleBinding
ClusterRoleBinding
ServiceAccount
You can test permissions using:
kubectl auth can-i get pods
For a particular ServiceAccount, you can also check its permissions explicitly.
This is much better than blindly changing permissions.
A Real Troubleshooting Scenario
Imagine you deploy an application.
You run:
kubectl get pods
and get:
NAME READY STATUS
backend-6f8d9c7d-x8p2 0/1 CrashLoopBackOff
Don't panic.
Follow the workflow.
1. Check logs
kubectl logs backend-6f8d9c7d-x8p2
Output:
DATABASE_URL not found
Now inspect the Deployment.
kubectl get deployment backend -o yaml
You discover the environment variable is missing.
You check the ConfigMap/Secret.
kubectl get configmap
kubectl get secret
You find the expected configuration.
You fix the Deployment.
Then:
kubectl rollout status deployment/backend
Finally:
kubectl get pods
And:
backend-6f8d9c7d-x8p2 1/1 Running
That's troubleshooting.
Not guessing.
The Kubernetes Debugging Pyramid
A useful way to think about troubleshooting is from broad to specific.
Application
▲
│
Container
▲
│
Pod
▲
│
Service
▲
│
Network
▲
│
Node/Cluster
Start at the layer where the failure appears.
For example:
Pod Pending
Start with scheduling.
Not application logs.
If:
Pod Running
Service unreachable
Investigate:
Service
Endpoints
Ports
Network
Ingress
If:
Pod CrashLoopBackOff
Investigate:
Container
Logs
Configuration
Probes
Dependencies
This simple mental model makes troubleshooting much faster.
Your Essential Kubernetes Debugging Toolkit
You don't need hundreds of commands.
Start with these:
kubectl get pods
kubectl describe pod <pod>
kubectl logs <pod>
kubectl logs <pod> --previous
kubectl get svc
kubectl describe svc <service>
kubectl get endpointslices
kubectl get events
kubectl get nodes
kubectl describe node <node>
kubectl get deployment
kubectl rollout status deployment/<name>
kubectl rollout history deployment/<name>
kubectl auth can-i <verb> <resource>
Learn what these commands tell you, not just how to type them.
A Better Production Troubleshooting Workflow
When an incident happens, use this sequence:
1. What is broken?
↓
2. Which resource is affected?
↓
3. What does Kubernetes report?
↓
4. What do Events say?
↓
5. What do application logs say?
↓
6. Which layer is failing?
↓
7. What changed recently?
↓
8. Apply the smallest safe fix
↓
9. Verify
↓
10. Document the root cause
This approach scales much better than memorizing random commands.
Don't Just Fix It — Understand Why
Suppose you fix:
CrashLoopBackOff
by changing an environment variable.
Don't stop there.
Ask:
Why was the environment variable missing?
Maybe:
Developer
↓
Changed application
↓
Didn't update ConfigMap
↓
CI/CD deployed new image
↓
Application crashed
Now you've discovered a process problem.
Maybe the real solution is:
Application Change
↓
Automated Validation
↓
Deployment
↓
Health Check
This is where Kubernetes troubleshooting connects directly to DevOps.
Build a Kubernetes Troubleshooting Lab
If you want to become good at Kubernetes, deliberately break things.
Create an application and then introduce failures.
Try:
Experiment 1
Use an invalid image:
image: nginx:this-does-not-exist
Find the cause.
Experiment 2
Break an environment variable.
Experiment 3
Use the wrong Service selector.
Experiment 4
Use the wrong target port.
Experiment 5
Create an intentionally failing readiness probe.
Experiment 6
Request more CPU than your cluster can provide.
Experiment 7
Create an RBAC permission problem.
Experiment 8
Break a volume configuration.
Then solve every problem without immediately searching for the answer.
This is how you develop real troubleshooting skills.
If You're Preparing for the CKA
Troubleshooting isn't just another Kubernetes topic.
It's a core skill for anyone working with Kubernetes.
You should be comfortable with:
Pods
Deployments
Services
Networking
Scheduling
Storage
RBAC
Logs
Events
Resource Management
Cluster Components
More importantly, you should be able to reason:
Symptom
↓
Evidence
↓
Root Cause
↓
Fix
↓
Verification
That mindset is useful far beyond an exam.
My Kubernetes Book
If you're learning Kubernetes from scratch or preparing for the Certified Kubernetes Administrator (CKA), I've created:
CKA Complete Study Guide — Certified Kubernetes Administrator
The goal is to give you a structured path through Kubernetes concepts, administration, troubleshooting, and practical learning.
📘 Get the CKA Complete Study Guide
Don't use a Kubernetes book only for reading.
Use it alongside a cluster or hands-on lab:
Read
↓
Understand
↓
Practice
↓
Break
↓
Troubleshoot
↓
Repeat
That's how Kubernetes becomes a skill instead of just another certification on your resume.
More Resources for Your DevOps Journey
If you're building a complete DevOps skill set, you can also explore:
### 🏗️ Terraform Associate Crash Course
Final Thoughts
Kubernetes troubleshooting isn't about knowing the most commands.
It's about knowing how to think when something breaks.
When you see:
CrashLoopBackOff
don't panic.
When you see:
Pending
don't randomly delete the Pod.
When a Service doesn't work, don't immediately restart everything.
Instead:
Observe
↓
Collect evidence
↓
Identify the failing layer
↓
Understand the root cause
↓
Fix it
↓
Verify
That is the real Kubernetes skill.
Because in production, your value isn't measured by how quickly you can type:
kubectl delete pod
It's measured by how well you can answer:
What broke, why did it break, and how can we prevent it from happening again?
Learn Kubernetes.
Build clusters.
Break things intentionally.
Troubleshoot them.
And eventually, Kubernetes won't feel complicated anymore.
It will feel like a system you know how to reason about.
📘 Ready to go deeper?
CKA Complete Study Guide — Certified Kubernetes Administrator
Top comments (0)