DEV Community

Cover image for Why a Kubernetes Pod Being 'Running' Doesn't Mean It's Healthy
Sheryar Ahmed
Sheryar Ahmed

Posted on

Why a Kubernetes Pod Being 'Running' Doesn't Mean It's Healthy

Why a Kubernetes Pod Being Running Doesn't Mean It's Healthy

A Pod can be Running while your application is broken, unable to receive traffic, or quietly failing every job.

One of the easiest Kubernetes mistakes to make is also one of the most dangerous:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

You see:

NAME                          READY   STATUS    RESTARTS
worker-7f9    1/1     Running   0
Enter fullscreen mode Exit fullscreen mode

And you think:

"The Pod is healthy."

Not necessarily.

Running is not a health check.

It's a statement about Pod lifecycle.

And understanding that difference completely changes how you debug Kubernetes.


The first mental model

A Kubernetes Pod has several different questions that can be answered independently:

Is the Pod running?
        ↓
Is the container ready?
        ↓
Can it receive traffic?
        ↓
Is the application actually functioning?
        ↓
Can it successfully perform its real work?
Enter fullscreen mode Exit fullscreen mode

These are not the same question.

Think about a normal server.

A process can be alive:

PID 1234
Enter fullscreen mode Exit fullscreen mode

while the application is:

  • stuck
  • unable to connect to a dependency
  • returning errors
  • unable to process requests
  • consuming jobs and failing them

Kubernetes is no different.


Running means something very specific

Kubernetes Pod phases include:

Pending
Running
Succeeded
Failed
Unknown
Enter fullscreen mode Exit fullscreen mode

When a Pod is Running, Kubernetes knows that the Pod has been scheduled and that its containers have started.

It does not automatically mean:

✓ application is healthy
✓ application can serve traffic
✓ dependencies are working
✓ requests are succeeding
✓ background jobs are succeeding
Enter fullscreen mode Exit fullscreen mode

That's the first distinction to remember:

Running describes lifecycle. Healthy describes behavior.


So how does Kubernetes know whether a container is healthy?

This is where probes enter the picture.

Kubernetes gives us three primary probe types:

Startup Probe
Readiness Probe
Liveness Probe
Enter fullscreen mode Exit fullscreen mode

They answer three different questions.


1. Startup Probe

"Has the application finished starting?"

Imagine your application takes 40 seconds to initialize.

During those 40 seconds:

container = alive
application = still starting
Enter fullscreen mode Exit fullscreen mode

If you immediately run a liveness check, Kubernetes might conclude:

"This application is broken."

And restart it.

Then it starts again.

Then the probe fails again.

Then Kubernetes restarts it again.

You can accidentally create:

start
 ↓
probe fails
 ↓
restart
 ↓
start
 ↓
probe fails
 ↓
restart
Enter fullscreen mode Exit fullscreen mode

A startup probe protects slow-starting applications from that situation.

Example:

startupProbe:
  httpGet:
    path: /health/live
    port: 8080
  periodSeconds: 2
  failureThreshold: 15
Enter fullscreen mode Exit fullscreen mode

With these settings, Kubernetes gives the application roughly:

2 × 15 = 30 seconds
Enter fullscreen mode Exit fullscreen mode

to successfully start before treating startup as failed.


2. Readiness Probe

This is one of the most important concepts in production Kubernetes.

Readiness asks:

"Should this Pod receive work right now?"

Suppose you have:

Service
   ↓
Pod A
Pod B
Pod C
Enter fullscreen mode Exit fullscreen mode

If Pod B fails its readiness probe:

Pod A → Ready
Pod B → NotReady
Pod C → Ready
Enter fullscreen mode Exit fullscreen mode

Kubernetes can stop sending Service traffic to Pod B.

But here's the important part:

A failed readiness probe does not restart the container.

That's intentional.

Maybe the application is temporarily unable to serve traffic but is expected to recover.

For example:

database temporarily unavailable
        ↓
readiness fails
        ↓
Pod stays alive
        ↓
Pod removed from traffic
        ↓
database recovers
        ↓
readiness succeeds
        ↓
Pod receives traffic again
Enter fullscreen mode Exit fullscreen mode

That's very different from liveness.


3. Liveness Probe

Liveness asks:

"Is this application so broken that restarting the container is appropriate?"

For example, suppose your application enters a deadlock:

process = alive
HTTP server = technically running
application = permanently stuck
Enter fullscreen mode Exit fullscreen mode

A liveness probe can detect that condition.

If it repeatedly fails:

liveness fails
      ↓
Kubernetes restarts container
      ↓
application starts fresh
Enter fullscreen mode Exit fullscreen mode

That's useful.

But it can also be dangerous.


The dangerous liveness probe

Imagine your liveness endpoint checks:

database
Redis
external API
third-party service
Enter fullscreen mode Exit fullscreen mode

Now Redis goes down.

Your application itself is perfectly capable of running, but your liveness probe says:

Redis unavailable
      ↓
liveness fails
      ↓
restart application
      ↓
application starts
      ↓
Redis unavailable
      ↓
liveness fails
      ↓
restart
Enter fullscreen mode Exit fullscreen mode

You've just turned a dependency outage into a restart storm.

That's why liveness should generally answer:

"Is this process/application locally stuck or irrecoverably broken?"

rather than:

"Are all of my dependencies healthy?"


The three probes in one picture

Remember this:

                 APPLICATION
                     │
          ┌──────────┼──────────┐
          ↓          ↓          ↓
       STARTUP    READINESS   LIVENESS
          │          │          │
          ↓          ↓          ↓
     "Can you      "Should     "Should I
      start?"       receive      restart
                    work?"      you?"
Enter fullscreen mode Exit fullscreen mode

Or even shorter:

Startup   → Can you start?
Readiness → Can you work?
Liveness  → Should I restart you?
Enter fullscreen mode Exit fullscreen mode

That's the mental model.


But there's another trap

Here's where real production systems become interesting.

Imagine a worker that consumes jobs from a queue.

Its architecture is:

Queue
  ↓
Worker
  ↓
Browser
  ↓
Result
Enter fullscreen mode Exit fullscreen mode

Now imagine the browser inside the worker becomes unhealthy.

You might create a readiness check:

browser unhealthy
       ↓
readiness = false
Enter fullscreen mode Exit fullscreen mode

Sounds correct.

But what happens to the worker?

If the worker is still consuming jobs:

Queue
 ↓
Worker
 ↓
Browser ❌
Enter fullscreen mode Exit fullscreen mode

the worker may continue receiving jobs even though it cannot successfully process them.

Now you've created a worse situation:

job arrives
   ↓
worker receives it
   ↓
browser unhealthy
   ↓
job fails
   ↓
next job
   ↓
fails
   ↓
next job
Enter fullscreen mode Exit fullscreen mode

The Pod can be:

STATUS = Running
Enter fullscreen mode Exit fullscreen mode

while the actual business operation is failing.

This is exactly why readiness needs to be designed around how the workload actually receives work, not simply around whether a health endpoint returns 200. Your infrastructure notes flag this as an area requiring careful review.


The subtle lesson: Readiness doesn't stop everything

This is especially important for background workers.

For a normal HTTP application:

Load Balancer
     ↓
Service
     ↓
Readiness
     ↓
Pod
Enter fullscreen mode Exit fullscreen mode

Readiness has a very direct effect.

But for a worker:

Pub/Sub / Queue
      ↓
Worker
Enter fullscreen mode Exit fullscreen mode

the application may pull work directly.

There may be no Kubernetes Service sitting between the queue and the worker.

Therefore:

readiness = false
Enter fullscreen mode Exit fullscreen mode

doesn't automatically mean:

stop consuming jobs
Enter fullscreen mode Exit fullscreen mode

That's an application-level behavior.

This is one of the biggest differences between:

"Not receiving HTTP traffic"

and

"Not receiving background work."


Kubernetes health is layered

A better mental model is:

Layer 1
Pod lifecycle
      ↓
Running?

Layer 2
Container health
      ↓
Startup / Liveness

Layer 3
Traffic readiness
      ↓
Ready?

Layer 4
Application health
      ↓
Can requests actually succeed?

Layer 5
Business health
      ↓
Are jobs actually completing successfully?
Enter fullscreen mode Exit fullscreen mode

And that last layer is often forgotten.


A Pod can look perfectly healthy and still be failing

Consider:

NAME                    READY   STATUS
worker-abc              1/1     Running
Enter fullscreen mode Exit fullscreen mode

Everything looks good.

But your metrics say:

jobs_received = 10,000
jobs_completed = 0
jobs_failed = 10,000
Enter fullscreen mode Exit fullscreen mode

From Kubernetes' perspective:

Pod = Running
Container = alive
Enter fullscreen mode Exit fullscreen mode

From your business perspective:

System = broken
Enter fullscreen mode Exit fullscreen mode

Both can be true at the same time.

That's why production observability cannot rely on:

kubectl get pods
Enter fullscreen mode Exit fullscreen mode

alone.


What should you actually monitor?

For a production workload, I'd think in layers.

Kubernetes signals

Pod phase
Container restarts
OOMKilled
CPU throttling
Memory usage
Pending Pods
Scheduling failures
Enter fullscreen mode Exit fullscreen mode

Probe signals

Startup failures
Readiness failures
Liveness failures
Time spent NotReady
Enter fullscreen mode Exit fullscreen mode

Application signals

Request success rate
Error rate
Latency
Dependency failures
Queue consumption
Enter fullscreen mode Exit fullscreen mode

Business signals

Jobs completed
Jobs failed
Jobs expired
Queue wait time
Success rate
Customer-visible failures
Enter fullscreen mode Exit fullscreen mode

The further down you go, the closer you get to the question:

"Is the system actually doing what users need?"


A practical debugging flow

When someone says:

"The Pod is unhealthy."

Don't immediately restart it.

Ask:

Step 1: Is it running?

kubectl get pod <pod>
Enter fullscreen mode Exit fullscreen mode

Step 2: What are the conditions?

kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

Look at:

Conditions
Events
Restart Count
Last State
Reason
Enter fullscreen mode Exit fullscreen mode

Step 3: Which probe is failing?

Look for:

Startup probe
Readiness probe
Liveness probe
Enter fullscreen mode Exit fullscreen mode

Step 4: Test the endpoint yourself

For example:

curl localhost:8080/health/live
curl localhost:8080/health/ready
Enter fullscreen mode Exit fullscreen mode

Step 5: Check application logs

kubectl logs <pod>
Enter fullscreen mode Exit fullscreen mode

Step 6: Check actual business behavior

For a worker:

Are jobs being consumed?
Are jobs succeeding?
Are they failing?
Are they expiring?
Enter fullscreen mode Exit fullscreen mode

That's where you stop debugging Kubernetes and start debugging the actual system.


A useful production rule

Here's the rule I now keep in my head:

Don't design health checks around what is easy to check. Design them around what Kubernetes needs to know.

For example:

Readiness

Good:

Can this Pod safely receive work?
Enter fullscreen mode Exit fullscreen mode

Potentially dangerous:

Is every dependency in the entire system healthy?
Enter fullscreen mode Exit fullscreen mode

Liveness

Good:

Is this process/application fundamentally stuck?
Enter fullscreen mode Exit fullscreen mode

Potentially dangerous:

Is every external dependency healthy?
Enter fullscreen mode Exit fullscreen mode

Business monitoring

Good:

Are jobs actually succeeding?
Enter fullscreen mode Exit fullscreen mode

This belongs in metrics and alerting, not necessarily in liveness.


The final mental model

Next time you see:

STATUS: Running
Enter fullscreen mode Exit fullscreen mode

don't translate it to:

"Everything is healthy."

Translate it to:

"The Pod is alive according to Kubernetes' lifecycle state. Now I need to determine whether it is ready, responsive, and actually doing useful work."

That's a much safer production mindset.


The 10-second revision

If you remember nothing else:

Running
  ≠
Healthy
Enter fullscreen mode Exit fullscreen mode
Startup   → Can you start?
Readiness → Can you receive work?
Liveness  → Should you be restarted?
Business  → Are you actually succeeding?
Enter fullscreen mode Exit fullscreen mode

And the most important production lesson:

Kubernetes can tell you that your process is alive. Your application metrics need to tell you whether your system is actually working.


What I learned from this

The interesting part of Kubernetes isn't memorizing:

Pod
Deployment
Service
Probe
Enter fullscreen mode Exit fullscreen mode

It's understanding the failure semantics between them.

A readiness failure doesn't restart a container.

A liveness failure does.

A Pod being Running doesn't make it Ready.

A Pod being Ready doesn't guarantee that a business operation succeeds.

And for background workers, being NotReady doesn't necessarily mean the worker has stopped consuming work.

That's where Kubernetes knowledge starts becoming production engineering.


Top comments (0)