DEV Community

Cover image for Making My Platform API Reconcile Application Updates
shubham goel
shubham goel

Posted on

Making My Platform API Reconcile Application Updates

In the previous post, I built the first real behavior for Platform Lab.

A developer could create:

apiVersion: platform.shubforge.dev/v1alpha1
kind: Application

metadata:
  name: greeting-service

spec:
  image: greeting-service:1.0.0
  replicas: 1

  port:
    containerPort: 8080
Enter fullscreen mode Exit fullscreen mode

and the Platform Operator would create:

Application
     |
     v
Platform Operator
     |
     +------ Deployment
     |
     +------ Service
Enter fullscreen mode Exit fullscreen mode

That was the first point where the Application API actually resulted in real Kubernetes resources.

But there was another important question.

What happens when the Application changes?

Creating resources once is useful, but a Kubernetes API should represent desired state.

If I change:

image: greeting-service:1.0.0
Enter fullscreen mode Exit fullscreen mode

to:

image: greeting-service:2.0.0
Enter fullscreen mode Exit fullscreen mode

I expect the platform to update the existing Deployment.

The same should happen for replicas and ports.

That is what I explored next.


Application Is Desired State

Initially, my Application looked like:

spec:
  image: greeting-service:1.0.0
  replicas: 1

  port:
    containerPort: 8080
Enter fullscreen mode Exit fullscreen mode

The Platform Operator created:

Deployment
  image = 1.0.0
  replicas = 1
  port = 8080

Service
  port = 8080
  targetPort = 8080
Enter fullscreen mode Exit fullscreen mode

Now I changed the Application to:

spec:
  image: greeting-service:2.0.0
  replicas: 3

  port:
    containerPort: 9090
Enter fullscreen mode Exit fullscreen mode

What I wanted was:

same Deployment
    |
    +---- image = 2.0.0
    +---- replicas = 3
    +---- port = 9090

same Service
    |
    +---- port = 9090
    +---- targetPort = 9090
Enter fullscreen mode Exit fullscreen mode

The important word here is:

same
Enter fullscreen mode Exit fullscreen mode

I did not want the platform to delete everything and recreate it.


Updating the Application

For the experiment, I used:

kubectl patch application greeting-service \
  --type=merge \
  -p '{
    "spec": {
      "image": "greeting-service:2.0.0",
      "replicas": 3,
      "port": {
        "containerPort": 9090
      }
    }
  }'
Enter fullscreen mode Exit fullscreen mode

I also added a Taskfile helper:

task application:update
Enter fullscreen mode Exit fullscreen mode

so I can repeat the same experiment more easily.

After the update, Kubernetes increments:

metadata.generation
Enter fullscreen mode Exit fullscreen mode

For example:

generation = 1
Enter fullscreen mode Exit fullscreen mode

becomes:

generation = 2
Enter fullscreen mode Exit fullscreen mode

That tells the controller that the desired Application specification changed.


The Interesting Part: Almost No New Controller Logic

I initially expected that I might need code like:

if (imageChanged) {
    updateDeployment();
}

if (replicasChanged) {
    updateDeployment();
}

if (portChanged) {
    updateDeployment();
    updateService();
}
Enter fullscreen mode Exit fullscreen mode

But that is not how I had structured the controller.

The Deployment dependent resource already builds its desired state from:

application.getSpec()
Enter fullscreen mode Exit fullscreen mode

The Service dependent resource does the same.

So when the Application changes, the flow is simply:

Application spec changes
        |
        v
controller receives update
        |
        v
desired Deployment recalculated
        |
        v
desired Service recalculated
        |
        v
actual resources reconciled
Enter fullscreen mode Exit fullscreen mode

That means creation and updates use the same model.


Creation and Update Are the Same Reconciliation Problem

For creation:

desired Deployment exists
actual Deployment does not exist
        |
        v
create Deployment
Enter fullscreen mode Exit fullscreen mode

For updates:

desired Deployment changed
actual Deployment exists
but differs
        |
        v
update Deployment
Enter fullscreen mode Exit fullscreen mode

The same applies to the Service.

This is one of the things I like about the managed dependent resource approach.

Instead of thinking in terms of:

CREATE
UPDATE
DELETE
Enter fullscreen mode Exit fullscreen mode

I can think in terms of:

What should the resource look like now?
Enter fullscreen mode Exit fullscreen mode

and let reconciliation handle the difference.


Verifying the Deployment Update

After changing the Application, I checked:

kubectl get deployment greeting-service \
  -o jsonpath='{.spec.replicas}'
Enter fullscreen mode Exit fullscreen mode

and got:

3
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get deployment greeting-service \
  -o jsonpath='{.spec.template.spec.containers[0].image}'
Enter fullscreen mode Exit fullscreen mode

returned:

greeting-service:2.0.0
Enter fullscreen mode Exit fullscreen mode

And:

kubectl get deployment greeting-service \
  -o jsonpath='{.spec.template.spec.containers[0].ports[0].containerPort}'
Enter fullscreen mode Exit fullscreen mode

returned:

9090
Enter fullscreen mode Exit fullscreen mode

So the Deployment had moved to the new desired state.


Verifying the Service Update

The port change also affected the Service.

I checked:

kubectl get service greeting-service \
  -o jsonpath='{.spec.ports[0].port}'
Enter fullscreen mode Exit fullscreen mode

and got:

9090
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl get service greeting-service \
  -o jsonpath='{.spec.ports[0].targetPort}'
Enter fullscreen mode Exit fullscreen mode

also returned:

9090
Enter fullscreen mode Exit fullscreen mode

So one higher-level platform change:

port:
  containerPort: 9090
Enter fullscreen mode Exit fullscreen mode

updated both Kubernetes resources that depend on that value.

Application.port
      |
      +------ Deployment containerPort
      |
      +------ Service port
      |
      +------ Service targetPort
Enter fullscreen mode Exit fullscreen mode

That is exactly the kind of translation I want the platform to handle.


Were the Resources Recreated?

I also wanted to make sure the operator was updating the existing resources rather than deleting and recreating them.

Before the Application update, I captured the Deployment UID:

kubectl get deployment greeting-service \
  -o jsonpath='{.metadata.uid}'
Enter fullscreen mode Exit fullscreen mode

and the Service UID:

kubectl get service greeting-service \
  -o jsonpath='{.metadata.uid}'
Enter fullscreen mode Exit fullscreen mode

Then I updated the Application and checked them again.

The UIDs stayed the same.

Conceptually:

Before

Deployment UID = A
Service UID    = B


After

Deployment UID = A
Service UID    = B
Enter fullscreen mode Exit fullscreen mode

So the resources were updated in place.


Checking the Service Identity Too

I also checked:

kubectl get service greeting-service \
  -o jsonpath='{.spec.clusterIP}'
Enter fullscreen mode Exit fullscreen mode

before and after the update.

The ClusterIP remained unchanged.

That gave me another useful signal that the Service was being reconciled rather than replaced.

The behavior I wanted was:

Application update
        |
        v
existing Service updated
Enter fullscreen mode Exit fullscreen mode

not:

delete Service
      |
      v
create another Service
Enter fullscreen mode Exit fullscreen mode

observedGeneration

The Application status already contains:

status:
  observedGeneration:
Enter fullscreen mode Exit fullscreen mode

This becomes more useful once updates exist.

After changing the Application:

metadata.generation = 2
Enter fullscreen mode Exit fullscreen mode

the controller eventually reports:

status.observedGeneration = 2
Enter fullscreen mode Exit fullscreen mode

So:

metadata.generation
=
latest desired state
Enter fullscreen mode Exit fullscreen mode

while:

status.observedGeneration
=
latest state processed by the controller
Enter fullscreen mode Exit fullscreen mode

When they match:

generation          = 2
observedGeneration  = 2
Enter fullscreen mode Exit fullscreen mode

I know the controller has processed the current Application specification.


Extending the Integration Test

I already had a black-box controller test that checked the initial provisioning flow:

create Application
      |
      v
Deployment created
      |
      v
Service created
      |
      v
status updated
Enter fullscreen mode Exit fullscreen mode

For this feature, I extended the same test.

The new flow is:

create Application
        |
        v
verify Deployment + Service
        |
        v
capture resource UIDs
        |
        v
update Application
        |
        v
wait for reconciliation
        |
        v
verify Deployment changes
        |
        v
verify Service changes
        |
        v
verify UIDs unchanged
        |
        v
verify observedGeneration
Enter fullscreen mode Exit fullscreen mode

This made the test much more representative of an actual controller lifecycle.


What the Test Verifies

The test starts with:

spec:
  image: example/test:1.0.0
  replicas: 2

  port:
    containerPort: 8080
Enter fullscreen mode Exit fullscreen mode

and later updates it to:

spec:
  image: example/test:2.0.0
  replicas: 3

  port:
    containerPort: 9090
Enter fullscreen mode Exit fullscreen mode

Then it verifies:

Deployment replicas = 3
Deployment image = example/test:2.0.0
Deployment port = 9090

Service port = 9090
Service targetPort = 9090
Enter fullscreen mode Exit fullscreen mode

It also verifies that:

Deployment UID did not change
Service UID did not change
Service ClusterIP did not change
Enter fullscreen mode Exit fullscreen mode

Finally, it waits until:

metadata.generation
=
status.observedGeneration
Enter fullscreen mode Exit fullscreen mode

Why the Test Waits

One important thing with controller tests is that reconciliation is asynchronous.

After running:

kubectl patch ...
Enter fullscreen mode Exit fullscreen mode

I cannot immediately assume the Deployment has already changed.

The real flow is:

Application updated
        |
        v
Kubernetes stores update
        |
        v
watch event
        |
        v
controller receives event
        |
        v
reconciliation happens
        |
        v
Deployment / Service updated
Enter fullscreen mode Exit fullscreen mode

So the test waits until the actual resources match the expected state.

This is much better than assuming everything happens immediately.


The Current Platform Lifecycle

The platform now supports:

Application created
        |
        v
Deployment + Service created

Application updated
        |
        v
same Deployment + Service updated
Enter fullscreen mode Exit fullscreen mode

So Application is starting to behave like a proper desired-state API.

The controller is not performing one-time provisioning.

It keeps asking:

Does the current Kubernetes state match the Application?


Desired State vs Provisioning Request

This distinction became clearer during this step.

A provisioning API might behave like:

POST /applications
        |
        v
create resources
        |
        v
done
Enter fullscreen mode Exit fullscreen mode

But the Kubernetes controller model is different:

Application exists
        |
        v
continuously describe desired state
        |
        v
reconcile actual state toward it
Enter fullscreen mode Exit fullscreen mode

That means updates naturally become part of the same model.

There is no separate:

update application workflow
Enter fullscreen mode Exit fullscreen mode

inside the controller.

There is only:

reconciliation
Enter fullscreen mode Exit fullscreen mode

Current Architecture

The current platform now looks like:

                        Developer
                            |
                            v
                       Application
                            |
                            v
                    Platform Operator
                            |
              +-------------+-------------+
              |                           |
              v                           v
         Deployment                    Service
              |                           |
              v                           |
             Pods <-----------------------+
Enter fullscreen mode Exit fullscreen mode

When the Application changes:

Application v1
      |
      v
Deployment + Service

      |
      | spec changes
      v

Application v2
      |
      v
same Deployment + Service
reconciled to new state
Enter fullscreen mode Exit fullscreen mode

What This Still Does Not Mean

Even though the Deployment has been updated successfully, I am still not setting:

Ready=True
Enter fullscreen mode Exit fullscreen mode

on the Application.

That is because:

Deployment spec updated
Enter fullscreen mode Exit fullscreen mode

does not mean:

application rollout completed
Enter fullscreen mode Exit fullscreen mode

For example:

Application updated
      |
      v
Deployment updated
      |
      v
new ReplicaSet created
      |
      v
Pods starting
Enter fullscreen mode Exit fullscreen mode

Those Pods could still end up in:

ImagePullBackOff
CrashLoopBackOff
Pending
Enter fullscreen mode Exit fullscreen mode

So there is another layer we still need to understand:

provisioned
Enter fullscreen mode Exit fullscreen mode

vs:

ready
Enter fullscreen mode Exit fullscreen mode

That will be the next feature.


Source Code

The full implementation is available in my Platform Lab repository.

Repository: Platform Lab

The changes covered in this post are available in:

Pull Request: Reconcile Application Updates

The main changes are:

Application update Taskfile command
extended controller integration test
in-place resource update verification
observedGeneration verification
Enter fullscreen mode Exit fullscreen mode

Interestingly, there was very little controller implementation change.

Most of this feature was about proving that the existing reconciliation model worked the way I expected.


What I Learned

The main thing I learned from this step is that a controller should not think only in terms of operations.

Instead of:

create
update
delete
Enter fullscreen mode Exit fullscreen mode

the more useful question is:

What should the resource look like right now?
Enter fullscreen mode Exit fullscreen mode

The Application spec defines that desired state.

The dependent resources calculate their desired Kubernetes representation.

The operator then reconciles the actual resources toward that state.

Because of that, adding Application updates did not require building a completely separate update path.

The same reconciliation logic handled both creation and modification.

That made the platform design feel much more aligned with Kubernetes.


What's Next?

The Platform Operator can now create and update resources.

But it still cannot answer:

Is my application actually ready?

That is the next problem I want to work on.

The Deployment exposes information such as:

desired replicas
ready replicas
available replicas
updated replicas
Enter fullscreen mode Exit fullscreen mode

The platform can use that information to start building meaningful Application status.

Eventually I want something like:

status:
  observedGeneration: 2
  readyReplicas: 3

  conditions:
    - type: Ready
      status: "True"
      reason: DeploymentReady
Enter fullscreen mode Exit fullscreen mode

But first I want to understand what Ready should actually mean.

So the next step is:

Application
      |
      v
Deployment
      |
      v
Deployment status
      |
      v
Application status
Enter fullscreen mode Exit fullscreen mode

That will move the platform from:

I created your Kubernetes resources
Enter fullscreen mode Exit fullscreen mode

toward:

I know whether your application is actually running.
Enter fullscreen mode Exit fullscreen mode

Top comments (0)