TL;DR
A small lab with two Kubernetes clusters. Crossplane is installed in one (the "brain") and manages stuff in the other (the "worker"). We ask three plain questions:
- Can the brain reach a worker cluster it does not own?
- When something breaks on the worker, does the brain notice?
- When the brain says "ready", does the app actually work?
To poke at those we run four short experiments and read what Crossplane reports back. Each one is a few kubectl commands with a screenshot or two:
- Native vs remote composition. Try to ship a Deployment across clusters using only Crossplane's built-in mechanism. See what the composition file lets you (and does not let you) express.
- Break the container image. Point a Workload at an image that does not exist and read what the brain says about it.
-
Delete a resource on the worker cluster. Time how long the brain takes to notice, and test what the
watch: truefield actually does. - Rotate a database password directly on Postgres. Watch the app fail and check whether the brain reports the failure.
If you are new to Crossplane, you get to reach each finding by running the commands yourself, which is the point. If you have been around Crossplane for a while, the interesting bits are the exact timings, the specific flag names, and the failure modes each experiment surfaces. Take your pick.
Repo: github.com/Joojo7/crossplane-lab
Intro
Hello my people its me again. A while back I wrote about BYOC in platform engineering, and the argument I kept coming back to was the difference between a control plane that reconciles infrastructure vs one that only references it. Crossplane was the most legible open example of the reconciling side, so this article is me sitting down with you and running experiments to see how much of that promise actually holds by default.
Its written like a lab report. Small clusters, three questions, four experiments, screenshots included. Follow along with the repo Joojo7/crossplane-lab if you want to reproduce it.
Background you need before we start
Three ideas in your head first. If you already know these, nevermind.
Control plane vs data plane. The control plane is the brain, it decides what should exist. The data plane is where things actually run. In this lab one Kubernetes cluster is the control plane (Crossplane lives here) and another is the data plane (just runs pods, no Crossplane).
What Crossplane is. A Kubernetes controller that lets you define your own resource types (like Workload or DataService) and say "when someone creates one of these, produce these other Kubernetes resources". Then it reconciles, which just means it keeps checking that reality matches the spec, forever.
Why two clusters. Because the interesting case is BYOC like in my last article. Your customer runs the target cluster. You run the control plane. Nothing installed on their side. Can your control plane really manage a cluster it does not own?
The 3 questions
- Can a control plane reach a cluster it does not own?
- Does it notice when something drifts?
- Does it know whether the thing it built actually works?
To test these we build a small platform with two abstractions. A Workload (stateless app, resolves to a Deployment/Service/ConfigMap/HPA) and a DataService (Postgres). We compose each one locally and remotely, then we break stuff and read what the control plane reports back.
Setup
1. Install the CLI
The CLI is separate from the controller.
brew install crossplane/tap/crossplane
# or
curl -sL https://raw.githubusercontent.com/crossplane/crossplane/main/install.sh | sh
sudo mv crossplane /usr/local/bin
crossplane version
If you see unable to get crossplane version: ... not found, thats not a broken CLI. It just means Crossplane isnt installed in a cluster yet.
2. Two clusters
kind create cluster --name crossplane-lab
kind create cluster --name workload-target
kubectl config use-context kind-crossplane-lab
Remember to switch context so you dont apply in the wrong cluster.
3. Install Crossplane in the control plane
helm repo add crossplane-stable https://charts.crossplane.io/stable
helm repo update
helm upgrade --install crossplane crossplane-stable/crossplane \
--namespace crossplane-system --create-namespace --wait
4. Install provider-kubernetes
kubectl apply -f providers/kubernetes/provider.yaml
kubectl get providers -w
Wait for INSTALLED and HEALTHY.
5. The kubeconfig trick
A normal kind kubeconfig points at 127.0.0.1, which means nothing from inside a pod in another cluster. --internal rewrites it to the Docker network name:
kind get kubeconfig --name workload-target --internal > /tmp/target.kubeconfig
grep server: /tmp/target.kubeconfig
# server: https://workload-target-control-plane:6443
kubectl create secret generic target-cluster \
-n crossplane-system \
--from-file=kubeconfig=/tmp/target.kubeconfig
That kubeconfig grants cluster-admin. Do not do this in prod, only for demo.
Building blocks (short tour before the experiments)
Every Crossplane abstraction has three parts. Very fast tour:
XRD (CompositeResourceDefinition) is the schema. Says which fields exist, whats required, what the defaults are. Like a Zod or Yup schema, but it registers a real API endpoint so kubectl get workloads starts working, and validation runs server-side at admission.
Composition is the recipe. Given one Workload, produce these Kubernetes resources with these values. Feels like Helm, but Helm renders once and Crossplane re-renders every reconcile forever. That continuous correction is the value prop.
Functions are the workers. In v2 a composition cant produce anything on its own, it needs at least one functionRef. Functions are packages (like npm modules) but each one runs as a pod, and Crossplane calls it over gRPC on every reconcile.
Install two: one for templating, one for readiness.
kubectl apply -f functions/functions.yaml
kubectl get functions
Wait for both INSTALLED and HEALTHY. Note v2 dropped the default registry, so the fully qualified package URL is now mandatory.
Apply the Workload XRD
kubectl apply -f experiments/01-native-vs-remote/definition.yaml
kubectl api-resources | grep lab.example.org
# workloads e01.lab.example.org/v1alpha1 true Workload
Nice milestone. Your abstraction is a first-class API now, indistinguishable from a built-in type. NAMESPACED: true in that third column is scope: Namespaced doing its job so two teams can each ship a hello Workload without colliding.
Small snag: my first apply failed with spec.versions[0].referenceable: Required value. Claims were removed in v2 but the field pointing at them is still required. Fix is one line, but its the first sign that the docs move faster than the plumbing under them.
Experiment 1: Native vs Remote
Question: can a composition target a cluster the control plane does not own?
Local first
The default composition puts resources in the control plane's own cluster.
kubectl apply -f experiments/01-native-vs-remote/examples/workload.yaml
kubectl get deploy,svc,cm,hpa -n default
NAME READY UP-TO-DATE AVAILABLE AGE
deployment.apps/orders-api 1/1 1 1 30s
NAME TYPE CLUSTER-IP PORT(S) AGE
service/orders-api ClusterIP 10.96.135.216 80/TCP 30s
One catch: Crossplane isnt allowed to create ordinary Kubernetes resources out of the box, so the composition renders perfectly and then creates nothing. You have to grant it a ClusterRole with rbac.crossplane.io/aggregate-to-crossplane: "true". See the repo for the full manifest.
Now try remote
Try to point that Deployment at the other cluster. You cant. Not because it errors, because there is no field in the composition where you could express it.
An error tells you what you did wrong. An absence tells you nothing. You only notice by going to look for the field and finding it isnt there.
The fix: Object + ProviderConfig
provider-kubernetes splits the problem in two.
Object is the cargo. Wraps a manifest and adds a destination:
apiVersion: kubernetes.m.crossplane.io/v1alpha1
kind: Object
spec:
providerConfigRef:
kind: ProviderConfig
name: workload-target # where it goes
forProvider:
manifest: # what should exist there
apiVersion: apps/v1
kind: Deployment
# ...
ProviderConfig is the address book entry. Gives a name to a set of credentials:
apiVersion: kubernetes.m.crossplane.io/v1alpha1
kind: ProviderConfig
metadata:
name: workload-target
spec:
credentials:
source: Secret
secretRef:
namespace: crossplane-system
name: target-cluster
key: kubeconfig
The indirection matters. Rotate the kubeconfig, update the secret, every Object using that name picks it up. Add a second customer cluster, add a second ProviderConfig. The composition changes by one string.
Result
Switch to workload-remote, re-apply, check the target:
kubectl --context kind-workload-target get deploy,svc,cm,hpa -n default
NAME READY UP-TO-DATE AVAILABLE AGE
deployment.apps/orders-api 1/1 1 1 37s
NAME TYPE CLUSTER-IP PORT(S) AGE
service/orders-api ClusterIP 10.96.5.197 80/TCP 37s
And crucially, no Crossplane on the target:
kubectl --context kind-workload-target get crds
# No resources found
Thats the control plane / data plane split made literal. Everything Crossplane knows is in one cluster. Everything the workload actually is lives in another, with no agent, no operator, no CRDs.
Finding. Native composition can only target its own cluster. Object + ProviderConfig is the only way across the boundary. In the BYOC case its very much not legacy baggage, its the mechanism.
Experiment 2: Readiness (break the image)
Question: does the control plane know whether what it built actually runs?
Break it
kubectl patch workload.e01.lab.example.org orders-api --type=merge \
-p '{"spec":{"image":{"versionSet":{"dev":{"path":"nginx:does-not-exist"}}}}}'
Target reacts immediately:
kubectl --context kind-workload-target get pods -n default
NAME READY STATUS RESTARTS AGE
orders-api-74bf7fd79-mf7lz 1/1 Running 0 2m39s
orders-api-d54b4d5d6-wkj5g 0/1 ImagePullBackOff 0 42s
Control plane:
crossplane resource trace workload.e01.lab.example.org orders-api
NAME SYNCED READY STATUS
Workload/orders-api (default) True True Available
├─ Object/orders-api-configmap (default) True True Available
├─ Object/orders-api-deployment (default) True True Available
├─ Object/orders-api-hpa (default) True True Available
└─ Object/orders-api-service (default) True True Available
Five healthy rows over a workload that cannot run.
Why
kubectl explain object.spec.readiness --api-version=kubernetes.m.crossplane.io/v1alpha1
readiness <Object>
Readiness ... if not specified it will be considered ready as soon as the
underlying external resource is considered up-to-date.
FIELDS:
policy <string>
enum: SuccessfulCreate, DeriveFromObject, AllTrue, DeriveFromCelQuery
"Up-to-date" is not "healthy". The default is doing exactly what it says on the tin. The problem is that the column heading says READY and nothing tells you which of the four policies is in play.
Fix, per kind
Instinct is to apply one policy everywhere. Thats wrong. A Deployment publishes Available, but a ConfigMap and a Service publish no conditions at all, so a query looking for Available on them would sit unhealthy forever.
Deployment. Read the condition Kubernetes already publishes:
spec:
readiness:
policy: DeriveFromCelQuery
celQuery: object.status.conditions.exists(c, c.type == "Available" && c.status == "True")
ConfigMap and Service. Existing is working:
spec:
readiness:
policy: SuccessfulCreate
HPA. Left on SuccessfulCreate. It publishes AbleToScale / ScalingActive / ScalingLimited and with no metrics-server in kind, ScalingActive would be false for reasons unrelated to the workload.
After
Same broken image, new policies:
NAME SYNCED READY STATUS
Workload/orders-api (default) True False Creating: Unready resources: deployment
├─ Object/orders-api-configmap (default) True True Available
├─ Object/orders-api-deployment (default) True False Unavailable
├─ Object/orders-api-hpa (default) True True Available
└─ Object/orders-api-service (default) True True Available
Deployment is unhealthy, XR names exactly whats wrong. Restore the image and it goes back to healthy:
kubectl patch workload.e01.lab.example.org orders-api --type=merge \
-p '{"spec":{"image":{"versionSet":{"dev":{"path":"nginx:1.27"}}}}}'
Unhealthy to healthy on recovery matters as much as healthy to unhealthy. Proves the signal is derived, not stuck.
Finding. The default readiness policy is the weakest of four and its silent about which one you're on. A "healthy" status sitting under a column labelled READY is confident enough to discourage you from checking. Dashes were more honest than ticks.
Experiment 3: Drift (delete the resource)
Question: does the control plane notice when something changes on the target?
Delete it
kubectl --context kind-workload-target delete deploy orders-api -n default
sleep 180 && kubectl --context kind-workload-target get deploy -n default
No resources found in default namespace.
Three minutes, still gone. Control plane:
kubectl get objects.kubernetes.m.crossplane.io -n default
NAME KIND PROVIDERCONFIG SYNCED READY AGE
orders-api-deployment Deployment workload-target True True 5m12s
Synced, ready, describing a Deployment that does not exist.
Poke it and it comes back
Touch anything on the control plane side:
kubectl annotate object.kubernetes.m.crossplane.io orders-api-deployment \
-n default poke=1 --overwrite
sleep 15 && kubectl --context kind-workload-target get deploy -n default
NAME READY UP-TO-DATE AVAILABLE AGE
orders-api 0/1 1 0 15s
Fifteen seconds from an annotation that means nothing to anyone. The reconciler works fine. It just doesnt watch the target, only its own side. Drift on the target is invisible, drift on the control plane is caught instantly.
The mental model most people (me included) carry is that the control plane is watching the target. Its not. Its watching its own resources and pushing outward on change.
watch: true was on the whole time
spec:
watch: true
I had it set from the start because the name says exactly what I wanted. Then:
kubectl explain object.spec --api-version=kubernetes.m.crossplane.io/v1alpha1
watch <boolean>
THIS IS AN ALPHA FIELD. Not honored unless "watches" feature gate is enabled.
Accepted by the schema, persisted in the object, doing nothing. No warning, no event, no condition.
Turning it on
The gate belongs to provider-kubernetes, not Crossplane itself. Reinstalling Crossplane with extra args would have done nothing. Provider flags go through a DeploymentRuntimeConfig:
apiVersion: pkg.crossplane.io/v1beta1
kind: DeploymentRuntimeConfig
metadata:
name: provider-kubernetes-watches
spec:
deploymentTemplate:
spec:
selector: {}
template:
spec:
containers:
- name: package-runtime
args:
- --enable-watches
- --debug
Container name has to be package-runtime and selector: {} is required for validation despite doing nothing. Reference it from the Provider (runtimeConfigRef:) and:
NAME ARGS
provider-kubernetes-eb133ed78b8c [--enable-watches --debug]
With the gate on
Same test:
kubectl --context kind-workload-target delete deploy orders-api -n default
sleep 30 && kubectl --context kind-workload-target get deploy -n default
NAME READY UP-TO-DATE AVAILABLE AGE
orders-api 1/1 1 1 30s
Deleted and already back. Sampling the Object every three seconds shows it went unhealthy and then healthy honestly:
orders-api-deployment Deployment workload-target True False 3h47m
orders-api-deployment Deployment workload-target True True 3h47m
A correction
Provider debug logs contain this:
External resource is up to date
{"requeue-after": "2026-09-06T10:38:50.100Z"}
Ten minutes out. So reconciliation isnt absent by default, its on a ~10 min poll and my 180s test was too short to see it. Three tiers:
default: ~10 min poll, status healthy throughout the gap
watch: true, gate off: identical to default, field silently inert
watch: true, gate on: informers on the target cluster, seconds
Finding. The watch: true field is a lie unless you also enable a feature gate nobody tells you about, on a component (the provider pod) you might not think to look at. Ten minutes of "synced and ready" on a deleted resource isnt long on a lab. On a fleet its the difference between an incident you caught and one a customer reported.
Experiment 4: Credentials (rotate the password)
Question: does the control plane notice when the value it delivered stops working?
For this we need a Workload that actually uses the credential rather than just receive it. db-checker is a small Express service that connects to Postgres using the injected env vars and runs SELECT version().
make e02-image
kubectl apply -f experiments/02-connected-resources/examples/db-checker.yaml
Baseline: nine healthy rows and a working connection
kubectl --context kind-workload-target exec deployment/db-checker -- \
wget -qO- http://localhost:3000/check
{"status":"connected","details":{"version":"PostgreSQL 16.15 ... 64-bit"}}
crossplane resource trace dataservice.e02.lab.example.org/checker-db -n default
crossplane resource trace workload.e02.lab.example.org/db-checker -n default
NAME SYNCED READY STATUS
DataService/checker-db (default) True True Available
├─ Object/checker-db-credentials (default) True True Available
├─ Object/checker-db-service (default) True True Available
└─ Object/checker-db-statefulset (default) True True Available
NAME SYNCED READY STATUS
Workload/db-checker (default) True True Available
├─ Object/db-checker-configmap (default) True True Available
├─ Object/db-checker-deployment (default) True True Available
├─ Object/db-checker-hpa (default) True True Available
└─ Object/db-checker-service (default) True True Available
Rotate the password on the target
Bypass the control plane, like a DBA might during an incident:
kubectl --context kind-workload-target exec statefulset/checker-db -c postgres -- \
psql -U checker -d checker \
-c "ALTER USER checker WITH PASSWORD 'rotated-out-of-band';"
ALTER ROLE
After: nine identical healthy rows and a failing app
Same traces, still all healthy. The app:
kubectl --context kind-workload-target exec deployment/db-checker -- \
node -e "require('http').get('http://localhost:3000/check', r => { let d=''; r.on('data', c=>d+=c); r.on('end', () => { console.log('HTTP', r.statusCode); console.log(d); }); });"
HTTP 500
{"status":"error","error":"password authentication failed for user \"checker\""}
For a broken image, readiness went unhealthy. For drift, the target caught up (once the gate was on). For a rotated credential, nothing budged.
Why
The Workload does not copy the credential onto the control plane. It holds the name of the Secret on the target and mounts it via envFrom.secretRef. Object/checker-db-credentials still contains the same Secret manifest it always did, and applying that manifest still succeeds. From the control planes side of the boundary, nothing has changed.
Finding. Reference-and-report is not verify. This is a reasonable boundary (the control plane shouldnt store the credential value) but a healthy status here does not mean the app can talk to the database. Teams close this gap with a checker sidecar, a scheduled Operation, or an external prober writing results back somewhere the control plane can read.
Bonus findings from wiring two abstractions together
The DataService writes host, port and secretName to its own status. The Workload looks that up with function-extra-resources and mounts the Secret by name. Two things bit hard while wiring it up:
Optional dependency cant actually be optional. With dataServiceRef simply absent, the pipeline failed with cannot get value from field path "spec.dataServiceRef": no such field. minMatch: 0 governs how many matches are acceptable, not whether the field exists. Workaround: dataServiceRef: "". The empty string resolves, matches nothing, pipeline proceeds. Handle it with an XRD default: "".
Composed resource names are a flat namespace. Experiment 01 and 02 both had a Workload called orders-api in different API groups. Both produced orders-api-deployment in the same namespace and the API server rejected it:
metadata.ownerReferences: Only one reference can have Controller set to true.
Found "true" in references for Workload/orders-api and Workload/orders-api
Two owners, distinguishable only by apiVersion. Nothing scopes composed names to the XRD that produced them. Suffix every generated name with an env or namespace in prod.
Conclusion
Reconcile versus reference is not a property of the tool. Crossplane will do either, and by default it does the weaker one.
Past a certain point it stops being a configuration choice and becomes a structural limit. Anything whose health lives outside the Kubernetes API is something the control plane can create, reference and report on, but sometimes cannot verify.
So the questions worth asking a BYOC vendor or your infra team arent about architecture diagrams. They are:
- Is drift detected, and on what timescale?
- Does a healthy status mean the resource was applied, or that it works?
- What happens to that status when something changes outside your control plane?
- How would I know, from your dashboard, that you had stopped watching?
A vendor or team who has thought about this will answer specifically. One who hasnt will show you a screen full of ticks.






Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.