DEV Community

Yash Sonawane
Yash Sonawane

Posted on

Kubernetes Troubleshooting: A Practical Guide to Debugging Pods That Don't Work

Kubernetes is easy when everything is working.

You run:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

and see:

NAME                       READY   STATUS
frontend-7d8f9c6d7-x2k4    1/1     Running
backend-5f7c8d9b8-m4p2     1/1     Running
Enter fullscreen mode Exit fullscreen mode

Perfect.

But real Kubernetes work begins when you see:

CrashLoopBackOff
ImagePullBackOff
Pending
0/1
Error
Terminating
Enter fullscreen mode Exit fullscreen mode

That's when many beginners start running random commands from Google.

Don't.

Kubernetes troubleshooting is a process, not a collection of magic commands.

In this guide, we'll build a systematic troubleshooting workflow you can use for real Kubernetes clusters and CKA-style problems.


The Golden Rule of Kubernetes Troubleshooting

When something breaks, don't immediately ask:

"What command should I run?"

Ask:

"What is the system telling me?"

Start with evidence.

A useful troubleshooting flow is:

Problem
   ↓
Observe
   ↓
Identify the failing layer
   ↓
Collect evidence
   ↓
Find the root cause
   ↓
Fix
   ↓
Verify
Enter fullscreen mode Exit fullscreen mode

This mindset is far more valuable than memorizing hundreds of commands.


Step 1 — Check the Pod

Start simple:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

You might see:

NAME                    READY   STATUS
web-7d9f8c6d4-x7k2      1/1     Running
api-5f7d8c9b4-m2p8      0/1     CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Immediately, you know:

web → probably healthy
api → something is wrong
Enter fullscreen mode Exit fullscreen mode

Now investigate the problematic Pod.


Step 2 — Describe the Pod

Run:

kubectl describe pod api-5f7d8c9b4-m2p8
Enter fullscreen mode Exit fullscreen mode

This gives you important information about:

  • Node placement
  • Containers
  • Images
  • Environment variables
  • Volumes
  • Events
  • Probes
  • Resource configuration

The Events section is particularly useful.

You might see:

Failed to pull image
Enter fullscreen mode Exit fullscreen mode

or:

Back-off restarting failed container
Enter fullscreen mode Exit fullscreen mode

or:

FailedMount
Enter fullscreen mode Exit fullscreen mode

Now you have evidence.


Step 3 — Check the Logs

If the container starts and then crashes, check its logs:

kubectl logs api-5f7d8c9b4-m2p8
Enter fullscreen mode Exit fullscreen mode

You might discover:

Error: DATABASE_URL is missing
Enter fullscreen mode Exit fullscreen mode

Now the problem isn't mysterious anymore.

The application is telling you exactly what it needs.

For a Pod with multiple containers:

kubectl logs <pod-name> -c <container-name>
Enter fullscreen mode Exit fullscreen mode

And if the container has restarted:

kubectl logs <pod-name> --previous
Enter fullscreen mode Exit fullscreen mode

That last command is extremely useful for crash-looping containers.


Understanding CrashLoopBackOff

One of the most common statuses beginners encounter is:

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

It does not mean "Kubernetes is broken."

It generally means the container is repeatedly starting and failing, and Kubernetes is backing off before restarting it again.

Think:

Container starts
      ↓
Application crashes
      ↓
Container exits
      ↓
Kubernetes restarts it
      ↓
Application crashes again
      ↓
Backoff
      ↓
Retry
Enter fullscreen mode Exit fullscreen mode

Your first questions should be:

Why did the process exit?
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl logs <pod>
Enter fullscreen mode Exit fullscreen mode

and:

kubectl logs <pod> --previous
Enter fullscreen mode Exit fullscreen mode

Then inspect:

kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

Common Cause #1 — Wrong Image

Suppose your Deployment contains:

image: myapp:v99
Enter fullscreen mode Exit fullscreen mode

but that image doesn't exist in your container registry.

You may see:

ImagePullBackOff
Enter fullscreen mode Exit fullscreen mode

or:

ErrImagePull
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

Look at the Events.

You might find:

Failed to pull image
Enter fullscreen mode Exit fullscreen mode

Now investigate:

  • Is the image name correct?
  • Is the tag correct?
  • Does the registry contain the image?
  • Does the cluster have permission to pull it?
  • Is authentication required?

Don't randomly restart the Pod.

Fix the image problem.


Common Cause #2 — Application Configuration

Your container may start perfectly but immediately crash because configuration is missing.

For example:

DATABASE_HOST
DATABASE_USER
DATABASE_PASSWORD
Enter fullscreen mode Exit fullscreen mode

Your application expects these values.

But your Pod doesn't receive them.

The result might be:

Application starts
      ↓
Reads configuration
      ↓
Configuration missing
      ↓
Application exits
      ↓
CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Check your:

ConfigMap
Secret
Deployment
Environment variables
Enter fullscreen mode Exit fullscreen mode

You can inspect the Deployment:

kubectl get deployment <name> -o yaml
Enter fullscreen mode Exit fullscreen mode

And inspect configuration resources:

kubectl get configmap
kubectl get secret
Enter fullscreen mode Exit fullscreen mode

Remember that sensitive values should be handled carefully and not casually exposed in terminal output or source repositories.


Common Cause #3 — Readiness Probe Failure

Sometimes your application is actually running.

But Kubernetes doesn't consider it ready.

For example:

Pod:
Running

Ready:
0/1
Enter fullscreen mode Exit fullscreen mode

This can happen because a readiness probe is failing.

Example:

readinessProbe:
  httpGet:
    path: /health
    port: 8080
Enter fullscreen mode Exit fullscreen mode

Kubernetes may repeatedly check:

http://container:8080/health
Enter fullscreen mode Exit fullscreen mode

If the endpoint fails, the Pod may remain unavailable to Service traffic.

This gives us an important distinction:

Running does not always mean Ready.


Liveness vs Readiness

These two concepts are easy to confuse.

Liveness

Asks:

"Is this container still healthy enough to keep running?"

If the liveness probe repeatedly fails, Kubernetes may restart the container.

Readiness

Asks:

"Is this application ready to receive traffic?"

If readiness fails, Kubernetes can stop sending traffic to that Pod while the container continues running.

Think:

Liveness
   ↓
Should the container keep running?

Readiness
   ↓
Should the container receive traffic?
Enter fullscreen mode Exit fullscreen mode

Understanding this distinction is extremely important when troubleshooting production workloads.


Common Cause #4 — Pod Is Pending

Now suppose you run:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

and see:

api-7d9f8c6d4-x8m2    0/1    Pending
Enter fullscreen mode Exit fullscreen mode

The application isn't necessarily broken.

It might not have been scheduled onto a node.

Start with:

kubectl describe pod <pod-name>
Enter fullscreen mode Exit fullscreen mode

Look at Events.

You might discover:

Insufficient cpu
Enter fullscreen mode Exit fullscreen mode

or:

Insufficient memory
Enter fullscreen mode Exit fullscreen mode

or a scheduling constraint is preventing placement.

This leads to an important troubleshooting question:

Is the problem inside the application, or is the Pod unable to be scheduled?

Don't jump into application logs if the container never started.


Common Cause #5 — Service Isn't Working

Your Pods might be perfectly healthy:

Pod 1 → Running
Pod 2 → Running
Pod 3 → Running
Enter fullscreen mode Exit fullscreen mode

But users still can't reach the application.

Now inspect the Service:

kubectl get svc
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe svc <service-name>
Enter fullscreen mode Exit fullscreen mode

One important thing to investigate is whether the Service has the expected endpoints.

Check:

kubectl get endpoints
Enter fullscreen mode Exit fullscreen mode

or, depending on your Kubernetes setup:

kubectl get endpointslices
Enter fullscreen mode Exit fullscreen mode

If the Service isn't selecting the correct Pods, you can have:

Healthy Pods
      ↓
Service
      ↓
No matching endpoints
Enter fullscreen mode Exit fullscreen mode

The application is running.

The networking configuration is the problem.


Labels Are Extremely Important

Services commonly select Pods using labels.

For example:

selector:
  app: backend
Enter fullscreen mode Exit fullscreen mode

Your Pod needs a matching label:

labels:
  app: backend
Enter fullscreen mode Exit fullscreen mode

If you accidentally have:

labels:
  app: frontend
Enter fullscreen mode Exit fullscreen mode

the Service won't select that Pod.

You can inspect labels with:

kubectl get pods --show-labels
Enter fullscreen mode Exit fullscreen mode

This simple command can save a lot of debugging time.


Common Cause #6 — DNS Problems

Imagine your frontend needs to communicate with:

backend-service
Enter fullscreen mode Exit fullscreen mode

Kubernetes provides service discovery through DNS.

If DNS isn't working or you're using the wrong Service name, communication can fail.

Start by checking:

kubectl get svc
Enter fullscreen mode Exit fullscreen mode

Then verify the Service name and namespace.

For example:

backend-service.default.svc.cluster.local
Enter fullscreen mode Exit fullscreen mode

The exact DNS name depends on the Service and namespace.

A useful debugging technique is launching a temporary Pod with networking tools and testing name resolution or connectivity from inside the cluster.


Common Cause #7 — Wrong Port

This is another classic problem.

Suppose your application listens on:

8080
Enter fullscreen mode Exit fullscreen mode

But your Service sends traffic to:

5000
Enter fullscreen mode Exit fullscreen mode

You could have:

User
 ↓
Service :80
 ↓
targetPort :5000
 ↓
Application listens on :8080
Enter fullscreen mode Exit fullscreen mode

The Pod may be healthy.

The Service may exist.

But traffic still fails.

Always compare:

Service port
targetPort
containerPort
application listening port
Enter fullscreen mode Exit fullscreen mode

These are related concepts, but they are not automatically the same thing.


Common Cause #8 — Volume Mount Problems

Suppose your Pod refuses to start and Events show:

FailedMount
Enter fullscreen mode Exit fullscreen mode

Now investigate:

PersistentVolume
PersistentVolumeClaim
StorageClass
Volume
Mount path
Enter fullscreen mode Exit fullscreen mode

Check:

kubectl get pv
kubectl get pvc
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl describe pvc <name>
Enter fullscreen mode Exit fullscreen mode

The key question is:

What resource is Kubernetes waiting for?

Again, follow the evidence.


Common Cause #9 — RBAC Problems

Sometimes your application or Kubernetes user doesn't have the required permissions.

You might see errors like:

Forbidden
Enter fullscreen mode Exit fullscreen mode

For example:

User cannot get pods
Enter fullscreen mode Exit fullscreen mode

This points toward Kubernetes authorization.

Investigate:

Role
ClusterRole
RoleBinding
ClusterRoleBinding
ServiceAccount
Enter fullscreen mode Exit fullscreen mode

You can test permissions using:

kubectl auth can-i get pods
Enter fullscreen mode Exit fullscreen mode

For a particular ServiceAccount, you can also check its permissions explicitly.

This is much better than blindly changing permissions.


A Real Troubleshooting Scenario

Imagine you deploy an application.

You run:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

and get:

NAME                     READY   STATUS
backend-6f8d9c7d-x8p2    0/1     CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Don't panic.

Follow the workflow.

1. Check logs

kubectl logs backend-6f8d9c7d-x8p2
Enter fullscreen mode Exit fullscreen mode

Output:

DATABASE_URL not found
Enter fullscreen mode Exit fullscreen mode

Now inspect the Deployment.

kubectl get deployment backend -o yaml
Enter fullscreen mode Exit fullscreen mode

You discover the environment variable is missing.

You check the ConfigMap/Secret.

kubectl get configmap
kubectl get secret
Enter fullscreen mode Exit fullscreen mode

You find the expected configuration.

You fix the Deployment.

Then:

kubectl rollout status deployment/backend
Enter fullscreen mode Exit fullscreen mode

Finally:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

And:

backend-6f8d9c7d-x8p2    1/1    Running
Enter fullscreen mode Exit fullscreen mode

That's troubleshooting.

Not guessing.


The Kubernetes Debugging Pyramid

A useful way to think about troubleshooting is from broad to specific.

                 Application
                     ▲
                     │
                  Container
                     ▲
                     │
                     Pod
                     ▲
                     │
                  Service
                     ▲
                     │
                  Network
                     ▲
                     │
                 Node/Cluster
Enter fullscreen mode Exit fullscreen mode

Start at the layer where the failure appears.

For example:

Pod Pending
Enter fullscreen mode Exit fullscreen mode

Start with scheduling.

Not application logs.

If:

Pod Running
Service unreachable
Enter fullscreen mode Exit fullscreen mode

Investigate:

Service
Endpoints
Ports
Network
Ingress
Enter fullscreen mode Exit fullscreen mode

If:

Pod CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

Investigate:

Container
Logs
Configuration
Probes
Dependencies
Enter fullscreen mode Exit fullscreen mode

This simple mental model makes troubleshooting much faster.


Your Essential Kubernetes Debugging Toolkit

You don't need hundreds of commands.

Start with these:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode
kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode
kubectl logs <pod>
Enter fullscreen mode Exit fullscreen mode
kubectl logs <pod> --previous
Enter fullscreen mode Exit fullscreen mode
kubectl get svc
Enter fullscreen mode Exit fullscreen mode
kubectl describe svc <service>
Enter fullscreen mode Exit fullscreen mode
kubectl get endpointslices
Enter fullscreen mode Exit fullscreen mode
kubectl get events
Enter fullscreen mode Exit fullscreen mode
kubectl get nodes
Enter fullscreen mode Exit fullscreen mode
kubectl describe node <node>
Enter fullscreen mode Exit fullscreen mode
kubectl get deployment
Enter fullscreen mode Exit fullscreen mode
kubectl rollout status deployment/<name>
Enter fullscreen mode Exit fullscreen mode
kubectl rollout history deployment/<name>
Enter fullscreen mode Exit fullscreen mode
kubectl auth can-i <verb> <resource>
Enter fullscreen mode Exit fullscreen mode

Learn what these commands tell you, not just how to type them.


A Better Production Troubleshooting Workflow

When an incident happens, use this sequence:

1. What is broken?
        ↓
2. Which resource is affected?
        ↓
3. What does Kubernetes report?
        ↓
4. What do Events say?
        ↓
5. What do application logs say?
        ↓
6. Which layer is failing?
        ↓
7. What changed recently?
        ↓
8. Apply the smallest safe fix
        ↓
9. Verify
        ↓
10. Document the root cause
Enter fullscreen mode Exit fullscreen mode

This approach scales much better than memorizing random commands.


Don't Just Fix It — Understand Why

Suppose you fix:

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

by changing an environment variable.

Don't stop there.

Ask:

Why was the environment variable missing?

Maybe:

Developer
   ↓
Changed application
   ↓
Didn't update ConfigMap
   ↓
CI/CD deployed new image
   ↓
Application crashed
Enter fullscreen mode Exit fullscreen mode

Now you've discovered a process problem.

Maybe the real solution is:

Application Change
       ↓
Automated Validation
       ↓
Deployment
       ↓
Health Check
Enter fullscreen mode Exit fullscreen mode

This is where Kubernetes troubleshooting connects directly to DevOps.


Build a Kubernetes Troubleshooting Lab

If you want to become good at Kubernetes, deliberately break things.

Create an application and then introduce failures.

Try:

Experiment 1

Use an invalid image:

image: nginx:this-does-not-exist
Enter fullscreen mode Exit fullscreen mode

Find the cause.

Experiment 2

Break an environment variable.

Experiment 3

Use the wrong Service selector.

Experiment 4

Use the wrong target port.

Experiment 5

Create an intentionally failing readiness probe.

Experiment 6

Request more CPU than your cluster can provide.

Experiment 7

Create an RBAC permission problem.

Experiment 8

Break a volume configuration.

Then solve every problem without immediately searching for the answer.

This is how you develop real troubleshooting skills.


If You're Preparing for the CKA

Troubleshooting isn't just another Kubernetes topic.

It's a core skill for anyone working with Kubernetes.

You should be comfortable with:

Pods
Deployments
Services
Networking
Scheduling
Storage
RBAC
Logs
Events
Resource Management
Cluster Components
Enter fullscreen mode Exit fullscreen mode

More importantly, you should be able to reason:

Symptom
  ↓
Evidence
  ↓
Root Cause
  ↓
Fix
  ↓
Verification
Enter fullscreen mode Exit fullscreen mode

That mindset is useful far beyond an exam.


My Kubernetes Book

If you're learning Kubernetes from scratch or preparing for the Certified Kubernetes Administrator (CKA), I've created:

CKA Complete Study Guide — Certified Kubernetes Administrator

The goal is to give you a structured path through Kubernetes concepts, administration, troubleshooting, and practical learning.

📘 Get the CKA Complete Study Guide

Don't use a Kubernetes book only for reading.

Use it alongside a cluster or hands-on lab:

Read
  ↓
Understand
  ↓
Practice
  ↓
Break
  ↓
Troubleshoot
  ↓
Repeat
Enter fullscreen mode Exit fullscreen mode

That's how Kubernetes becomes a skill instead of just another certification on your resume.


More Resources for Your DevOps Journey

If you're building a complete DevOps skill set, you can also explore:

### 🐳 Docker Mastery

### 🏗️ Terraform Associate Crash Course

### 🔀 Git Mastery

### ⚙️ DevOps Complete Pack

### 🐹 Mastering Go


Final Thoughts

Kubernetes troubleshooting isn't about knowing the most commands.

It's about knowing how to think when something breaks.

When you see:

CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

don't panic.

When you see:

Pending
Enter fullscreen mode Exit fullscreen mode

don't randomly delete the Pod.

When a Service doesn't work, don't immediately restart everything.

Instead:

Observe
   ↓
Collect evidence
   ↓
Identify the failing layer
   ↓
Understand the root cause
   ↓
Fix it
   ↓
Verify
Enter fullscreen mode Exit fullscreen mode

That is the real Kubernetes skill.

Because in production, your value isn't measured by how quickly you can type:

kubectl delete pod
Enter fullscreen mode Exit fullscreen mode

It's measured by how well you can answer:

What broke, why did it break, and how can we prevent it from happening again?

Learn Kubernetes.

Build clusters.

Break things intentionally.

Troubleshoot them.

And eventually, Kubernetes won't feel complicated anymore.

It will feel like a system you know how to reason about.

📘 Ready to go deeper?

CKA Complete Study Guide — Certified Kubernetes Administrator

Top comments (0)