DEV Community

Cover image for From Provisioned to Ready: Tracking Application Health in My Kubernetes Platform
shubham goel
shubham goel

Posted on

From Provisioned to Ready: Tracking Application Health in My Kubernetes Platform

In the previous posts, I built the first real platform flow:

Application
     |
     v
Platform Operator
     |
     +------ Deployment
     |
     +------ Service
Enter fullscreen mode Exit fullscreen mode

The platform could also handle updates.

If I changed:

spec:
  image: greeting-service:2.0.0
  replicas: 3

  port:
    containerPort: 9090
Enter fullscreen mode Exit fullscreen mode

the existing Deployment and Service were reconciled to the new desired state.

That worked well.

But there was still one important problem.

The platform could tell me:

I created your Deployment.

It could not tell me:

Your application is actually running.

Those are not the same thing.

That is what I wanted to solve next.


Provisioned Does Not Mean Ready

Suppose the Platform Operator successfully creates:

Deployment/greeting-service
Enter fullscreen mode Exit fullscreen mode

and:

Service/greeting-service
Enter fullscreen mode Exit fullscreen mode

From the controller's perspective, provisioning succeeded.

But the Pods might still be:

Pending
ImagePullBackOff
CrashLoopBackOff
Enter fullscreen mode Exit fullscreen mode

So this:

Deployment exists
Enter fullscreen mode Exit fullscreen mode

does not mean:

Application Ready=True
Enter fullscreen mode Exit fullscreen mode

That was the main difference compared with my earlier Greeting operator.

For Greeting, the managed resource was a ConfigMap.

Greeting
    |
    v
ConfigMap
Enter fullscreen mode Exit fullscreen mode

If the ConfigMap was successfully reconciled, that was basically enough to consider the Greeting ready.

For an Application:

Application
    |
    v
Deployment
    |
    v
Pods
Enter fullscreen mode Exit fullscreen mode

there is another runtime layer.


Where Should Readiness Come From?

The Platform Operator does not need to manage Pods directly.

Kubernetes already has the Deployment controller for that.

The responsibilities are:

Platform Operator
=
Application → Deployment + Service
Enter fullscreen mode Exit fullscreen mode

while:

Kubernetes Deployment Controller
=
Deployment → ReplicaSets + Pods
Enter fullscreen mode Exit fullscreen mode

So instead of watching Pods directly, I decided to use:

Deployment.status
Enter fullscreen mode Exit fullscreen mode

as the first source of Application readiness.

The flow becomes:

Application
      |
      v
Platform Operator
      |
      v
Deployment
      |
      v
Kubernetes Deployment Controller
      |
      v
Pods
      |
      v
Deployment.status
      |
      v
Platform Operator
      |
      v
Application.status
Enter fullscreen mode Exit fullscreen mode

What Deployment Status Gives Me

For the first readiness implementation, I use values such as:

spec.replicas
status.readyReplicas
status.availableReplicas
status.updatedReplicas
status.observedGeneration
Enter fullscreen mode Exit fullscreen mode

Suppose the Application wants:

3 replicas
Enter fullscreen mode Exit fullscreen mode

and the Deployment reports:

readyReplicas     = 3
availableReplicas = 3
updatedReplicas   = 3
Enter fullscreen mode Exit fullscreen mode

and the latest Deployment generation has been observed.

Then the Application can be considered ready.


Adding Readiness to Application Status

The Application status now includes:

status:
  observedGeneration: 2

  deploymentName: greeting-service
  serviceName: greeting-service

  readyReplicas: 3

  conditions:
    - type: Ready
      status: "True"
      observedGeneration: 2
      reason: DeploymentReady
      message: Deployment has 3/3 ready replicas
      lastTransitionTime: "2026-10-03T14:20:00Z"
Enter fullscreen mode Exit fullscreen mode

The new parts are:

readyReplicas
Ready condition
Enter fullscreen mode Exit fullscreen mode

So the Application API is starting to expose workload state instead of only provisioning information.


Ready=False

Suppose the Deployment wants:

3 replicas
Enter fullscreen mode Exit fullscreen mode

but only one is currently ready.

Then Application status may look like:

status:
  readyReplicas: 1

  conditions:
    - type: Ready
      status: "False"
      reason: DeploymentProgressing
      message: Deployment has 1/3 ready replicas
Enter fullscreen mode Exit fullscreen mode

The platform is saying:

The Kubernetes resources exist, but the workload is not fully available yet.


Ready=True

Once all replicas become available:

desired replicas   = 3
updated replicas   = 3
ready replicas     = 3
available replicas = 3
Enter fullscreen mode Exit fullscreen mode

the Application can report:

conditions:
  - type: Ready
    status: "True"
    reason: DeploymentReady
    message: Deployment has 3/3 ready replicas
Enter fullscreen mode Exit fullscreen mode

So the lifecycle becomes:

Application created
        |
        v
Deployment created
        |
        v
Pods starting
        |
        v
Ready=False
        |
        v
Pods become available
        |
        v
Ready=True
Enter fullscreen mode Exit fullscreen mode

Separating Readiness Logic

I did not want to put all the readiness calculations directly inside:

ApplicationReconciler
Enter fullscreen mode Exit fullscreen mode

So I created:

ApplicationReadinessEvaluator.java
Enter fullscreen mode Exit fullscreen mode

and a small result model:

ApplicationReadiness
Enter fullscreen mode Exit fullscreen mode

The evaluator returns:

readyReplicas
ready
reason
message
Enter fullscreen mode Exit fullscreen mode

Conceptually:

Deployment
     |
     v
ApplicationReadinessEvaluator
     |
     v
ApplicationReadiness
     |
     v
Application.status
Enter fullscreen mode Exit fullscreen mode

This keeps the controller focused on reconciliation while the evaluator contains the actual platform rule.


The Basic Readiness Rule

For now, the rule is intentionally simple.

The Application is ready when:

latest Deployment generation observed

AND

updatedReplicas == desiredReplicas

AND

readyReplicas == desiredReplicas

AND

availableReplicas == desiredReplicas
Enter fullscreen mode Exit fullscreen mode

Then:

Ready=True
Enter fullscreen mode Exit fullscreen mode

Otherwise:

Ready=False
Enter fullscreen mode Exit fullscreen mode

I expect this rule to evolve later, but it is enough for the first version.


Why Deployment Generation Matters

There are actually two different generations involved now.

The Application has:

Application.metadata.generation
Enter fullscreen mode Exit fullscreen mode

and the Deployment has:

Deployment.metadata.generation
Enter fullscreen mode Exit fullscreen mode

The Platform Operator tracks the Application generation using:

Application.status.observedGeneration
Enter fullscreen mode Exit fullscreen mode

But for Deployment readiness, I also need to check:

Deployment.status.observedGeneration
Enter fullscreen mode Exit fullscreen mode

For example:

Deployment generation         = 3
Deployment observedGeneration = 2
Enter fullscreen mode Exit fullscreen mode

means the Kubernetes Deployment controller has not processed the latest Deployment specification yet.

Even if some replica values look okay, I should not assume that the latest desired Deployment state is ready.

So the Application stays:

Ready=False
Enter fullscreen mode Exit fullscreen mode

until the Deployment controller catches up.


Continuous Readiness

One thing I wanted to understand was whether readiness should only be calculated during deployment.

It should not.

Application health can change later.

Suppose everything is healthy:

Ready=True
3/3 replicas ready
Enter fullscreen mode Exit fullscreen mode

Then one Pod becomes unhealthy.

The flow is:

Pod becomes unhealthy
        |
        v
Deployment Controller notices
        |
        v
Deployment.status changes
        |
        v
Platform Operator receives update
        |
        v
Application reconciled again
        |
        v
readyReplicas decreases
        |
        v
Ready=False
Enter fullscreen mode Exit fullscreen mode

Then Kubernetes tries to recover the Deployment.

Once the replacement Pod becomes ready:

new Pod ready
      |
      v
Deployment.status changes
      |
      v
Application reconciled
      |
      v
Ready=True
Enter fullscreen mode Exit fullscreen mode

So Application status can move like:

Ready=True
3/3
    |
    v
Ready=False
2/3
    |
    v
Ready=True
3/3
Enter fullscreen mode Exit fullscreen mode

That is much more useful than setting Ready=True once and never checking again.


The Platform Operator Does Not Heal Pods

This was another important distinction.

If a Pod goes down, the Platform Operator should not do this:

Pod failed
    |
    v
Platform Operator creates another Pod
Enter fullscreen mode Exit fullscreen mode

That is already Kubernetes responsibility.

Instead:

Pod failed
      |
      v
Deployment Controller
      |
      v
replacement Pod
Enter fullscreen mode Exit fullscreen mode

The Platform Operator only observes:

Deployment.status
Enter fullscreen mode Exit fullscreen mode

and reflects that into:

Application.status
Enter fullscreen mode Exit fullscreen mode

So each controller has a clear responsibility.


Why I Am Not Watching Pods Yet

I could watch individual Pods and inspect things like:

CrashLoopBackOff
ImagePullBackOff
Pending
Enter fullscreen mode Exit fullscreen mode

But I decided not to add that yet.

For basic readiness, Deployment status already tells me enough:

how many replicas are desired
how many are updated
how many are ready
how many are available
Enter fullscreen mode Exit fullscreen mode

So the current flow stays:

Application
     |
     v
Deployment
     |
     v
Deployment.status
     |
     v
Application.status
Enter fullscreen mode Exit fullscreen mode

Later, Pod-level observation can be added if the platform needs better failure explanations.


Ready Condition

The Application currently uses one condition:

Ready
Enter fullscreen mode Exit fullscreen mode

Some reasons are:

DeploymentPending
DeploymentProgressing
DeploymentReady
Enter fullscreen mode Exit fullscreen mode

For example:

conditions:
  - type: Ready
    status: "False"
    reason: DeploymentProgressing
    message: Deployment has 1/3 ready replicas
Enter fullscreen mode Exit fullscreen mode

and later:

conditions:
  - type: Ready
    status: "True"
    reason: DeploymentReady
    message: Deployment has 3/3 ready replicas
Enter fullscreen mode Exit fullscreen mode

lastTransitionTime

The condition also contains:

lastTransitionTime
Enter fullscreen mode Exit fullscreen mode

I wanted this to represent an actual condition transition.

For example:

Ready=False
0/3
Enter fullscreen mode Exit fullscreen mode

then:

Ready=False
1/3
Enter fullscreen mode Exit fullscreen mode

then:

Ready=False
2/3
Enter fullscreen mode Exit fullscreen mode

does not mean the Ready condition changed.

It is still:

False
Enter fullscreen mode Exit fullscreen mode

So I preserve the existing lastTransitionTime.

But:

Ready=False
      |
      v
Ready=True
Enter fullscreen mode Exit fullscreen mode

is a real transition.

That gets a new timestamp.


kubectl Output

I also added a ReadyReplicas printer column to the Application CRD.

Now:

kubectl get papp
Enter fullscreen mode Exit fullscreen mode

can show something like:

NAME               READY   READYREPLICAS   IMAGE                    REPLICAS
greeting-service   True    3               greeting-service:2.0.0   3
Enter fullscreen mode Exit fullscreen mode

If one replica becomes unavailable:

NAME               READY   READYREPLICAS   IMAGE                    REPLICAS
greeting-service   False   2               greeting-service:2.0.0   3
Enter fullscreen mode Exit fullscreen mode

That is closer to the developer experience I want.

A developer should not always need to run:

kubectl get deployment
kubectl get pods
Enter fullscreen mode Exit fullscreen mode

just to understand the basic state of their Application.


Adding Java Unit Tests

This feature also gave me a good reason to introduce Java unit tests into the real Platform Operator.

Until now, most tests were using the Kubernetes API directly.

But readiness evaluation is platform logic.

So I added tests for:

ApplicationReadinessEvaluator
Enter fullscreen mode Exit fullscreen mode

For example:

desired = 3
ready = 3
available = 3
updated = 3
latest generation observed

→ Ready=True
Enter fullscreen mode Exit fullscreen mode

Another test covers:

desired = 3
ready = 1

→ Ready=False
Enter fullscreen mode Exit fullscreen mode

And another checks:

Deployment generation = 3
observedGeneration = 2

→ Ready=False
Enter fullscreen mode Exit fullscreen mode

Now the test layers look like:

API tests
      |
      v
CRD validation


Java unit tests
      |
      v
platform logic


controller integration tests
      |
      v
real Kubernetes reconciliation
Enter fullscreen mode Exit fullscreen mode

This feels much more natural than introducing all the testing infrastructure before there was real platform logic to test.


Running the Tests

The API tests are still:

task application:test:api
Enter fullscreen mode Exit fullscreen mode

The Java tests are:

task platform-operator:test
Enter fullscreen mode Exit fullscreen mode

And the controller integration tests are:

task application:test:controller
Enter fullscreen mode Exit fullscreen mode

So the testing strategy is growing together with the platform.


Current Platform Architecture

At this point the architecture looks like:

                           Application
                                |
                                v
                      ApplicationReconciler
                                |
                 +--------------+--------------+
                 |                             |
                 v                             v
            Deployment                      Service
                 |
                 v
       Kubernetes Deployment
             Controller
                 |
                 v
                Pods
                 |
                 v
         Deployment.status
                 |
                 v
         Platform Operator
                 |
                 v
         Application.status
                 |
        +--------+--------+
        |                 |
        v                 v
 readyReplicas       Ready Condition
Enter fullscreen mode Exit fullscreen mode

This is the first time the platform is not only creating infrastructure but also reporting workload state.


What This Still Does Not Tell Me

The current Application status can say:

Ready=False
0/3 replicas ready
Enter fullscreen mode Exit fullscreen mode

but it cannot yet tell me exactly why.

For example, it does not currently distinguish between:

ImagePullBackOff

CrashLoopBackOff

failed readiness probe

Pod unschedulable
Enter fullscreen mode Exit fullscreen mode

That would require deeper workload inspection.

For now, I want to keep the first readiness model simple.


Source Code

The complete project is available in my Platform Lab repository.

Repository: Platform Lab

The changes covered in this post are available in:

Pull Request: Add Application Readiness

The main changes include:

readyReplicas in Application status

Ready condition

ApplicationReadinessEvaluator

Deployment status observation

continuous readiness updates

Java unit tests

ReadyReplicas kubectl printer column
Enter fullscreen mode Exit fullscreen mode

What I Learned

The biggest lesson from this step was the difference between:

resource provisioning
Enter fullscreen mode Exit fullscreen mode

and:

workload readiness
Enter fullscreen mode Exit fullscreen mode

Creating a Deployment successfully only tells me that Kubernetes accepted the desired workload configuration.

It does not tell me that the workload is actually running.

The Deployment controller already knows how to manage Pods and recover failed replicas.

So instead of duplicating that responsibility, the Platform Operator can observe Deployment status and expose a simpler platform-level view.

That gives me:

Application
=
developer desired state
+
platform workload status
Enter fullscreen mode Exit fullscreen mode

which is much closer to what I want the Platform API to eventually become.


What's Next?

Right now my sample Application still uses placeholder images.

So the next useful step is to build a small real application image and run it through Platform Lab.

That will let me observe a real lifecycle:

Application created
        |
        v
Ready=False
        |
        v
Pod starts
        |
        v
Deployment becomes available
        |
        v
Application Ready=True
Enter fullscreen mode Exit fullscreen mode

Then I also want to deliberately make the workload unhealthy and observe:

Ready=True
      |
      v
Ready=False
      |
      v
Kubernetes recovery
      |
      v
Ready=True
Enter fullscreen mode Exit fullscreen mode

That should give me a real end-to-end test of the readiness model before moving into configuration and Secrets.

Top comments (0)