I've been running EKS with Karpenter in production for a while. I can debug a node that won't provision, I know what a NodeClaim is, and I've written the Terraform that installs the whole thing.
What I couldn't do was explain what Karpenter's process is actually doing between "pod is unschedulable" and "EC2 instance appears."
So I wrote a controller from scratch. No controller-runtime, no Kubebuilder — just client-go, so there was nowhere to hide. About 250 lines.
It works. But the three things that went wrong taught me more than the working version did, and none of them are in the tutorials.
What I built:
A controller that watches nodes for a disruption taint and reports which pods on them won't come back if the node dies.
Not all pods are equal when a node goes away:
Owned by a ReplicaSet, StatefulSet or Job → rescheduled elsewhere, fine
Owned by a DaemonSet → goes with the node, expected
No controller owner → gone, permanently
That last case is what kubectl drain refuses to evict without --force. I wanted to surface it before the drain, which is when it's actually useful.
The shape is the same as every controller:
API server
│ WATCH (opened once, held open)
▼
Node informer ────► local cache
│
│ AddFunc / UpdateFunc / DeleteFunc
▼
enqueue(key) "ip-10-0-1-42.ec2.internal"
│
▼
┌─────────────┐
│ workqueue │ dedupes · rate limits · tracks in-flight
└─────────────┘
│
▼
worker goroutine
│
▼
reconcile(key)
Everything left of reconcile is plumbing. All the decisions live in about 40 lines inside it.
Right — the bugs.
1. The control plane looked like it was about to die
First working version. I tainted the node and got this:
`node ip-10-0-1-42 tainted: spot-interruption=true:NoSchedule
kube-system/etcd-... — bare pod, will NOT be recreated
kube-system/kube-apiserver-... — bare pod, will NOT be recreated
kube-system/kube-scheduler-... — bare pod, will NOT be recreated
kube-system/kube-controller-manager-... — bare pod, will NOT be recreated`
My controller had just told me the entire Kubernetes control plane was at risk.
The check itself was fine:
`go
owner := metav1.GetControllerOf(pod)
if owner == nil {
// nothing will recreate this
}`
And it was correct. Those pods genuinely have no controller owner in the API.
They're static pods. The kubelet runs them from manifest files on disk — which it has to, because the API server can't be created by the thing it is. The kubelet then creates read-only mirror pods in the API so they show up in kubectl get pods, and sets their owner to the Node.
So nothing in the API would recreate them, and my conclusion was still wrong, because the kubelet restarts them from disk regardless.
The fix is one annotation the kubelet sets on every mirror pod:
go
if _, isMirror := pod.Annotations["kubernetes.io/config.mirror"]; isMirror {
continue
}
What I actually learned: a correct check can produce a wrong answer. The ownership model in the API isn't the whole lifecycle story — there's a second lifecycle running on each node that the API only mirrors. I'd read about static pods. I didn't know about them until one showed up in my own output claiming it was about to be lost.
2. My controller could make a claim but not retract it
Second version annotated at-risk pods so the finding outlived my terminal:
node-drain-controller/at-risk: ip-10-0-1-42.ec2.internal
Worked nicely. Then I removed the taint, and the annotation stayed. Forever.
The pod was now carrying a note saying it was at risk on a node that was completely healthy. Nothing was ever going to remove it.
Here's the line that caused it:
go
if taint == nil {
return nil // "no taint, nothing to do"
}
I'd treated "this node has no taint" as nothing to do. It isn't. It's information: this node is healthy. And if it's healthy, anything on it claiming otherwise is wrong and needs correcting.
go
if taint == nil {
return c.clearAnnotations(name)
}
**
What I actually learned:** I'd written half a reconcile loop. A reconcile function isn't "do the thing when the trigger fires" — that's an event handler. It's make reality match what it should be, whatever reality currently is. And that has to work in both directions, or the state you write drifts away from the truth and nobody can tell which annotations still mean anything.
There's a subtlety in the cleanup too. It keys off the annotation I left behind, not the conditions I used when leaving it:
go
if pod.Annotations[riskAnnotation] == "" {
continue // I never annotated this one
}
If I later change which pods get annotated, cleanup still finds everything I've ever written. Re-deriving the original condition would orphan state I could no longer locate — which is how controllers leave debris nobody can trace back to them.
3. Making it configurable broke state written under the old config
Last addition: a CRD so the taint key wasn't hardcoded.
yaml
apiVersion: jk.io/v1alpha1
kind: DisruptionPolicy
metadata:
name: default
spec:
taintKey: spot-interruption
Changed it to maintenance with kubectl patch, restarted, and got:
`loaded policy: taintKey=maintenance annotate=true
caches synced — starting workers
cleared annotation on default/orphan — node no longer tainted`
The node still had spot-interruption on it. The pod was still annotated under that key. But the controller was now looking for maintenance, found none, concluded the node was fine, and cleaned up.
Which is correct, and I only got away with it because the bidirectional cleanup from bug 2 existed and the controller restarted and re-evaluated everything. If I'd implemented a live policy watch instead of reading at startup, the same change would have left stale annotations with no path back.
What I actually learned: changing a controller's configuration invalidates state it wrote under the old configuration, and that's a design question, not an implementation detail. "Make the taint key configurable" sounds like a one-line change. It isn't. It means deciding what happens to everything you've already written.
I left it reading at startup and wrote the limitation down, which I think is the right call — a half-built live watch would have been worse than an honest restart requirement.
The one I didn't hit
Worth mentioning because it's in a lot of tutorials.
In the worker loop, this placement is wrong:
go
func (c *Controller) runWorker() {
for {
key, _ := c.queue.Get()
defer c.queue.Done(key) // ← never fires
c.reconcile(key)
}
}
defer is function-scoped, not block-scoped. runWorker loops forever, so the defer doesn't run until shutdown. Every key stays marked in-flight permanently, the queue stops redelivering them, and the controller silently stops reconciling. No error. No log.
Splitting the loop so each item gets its own function is what makes it fire:
go
func (c *Controller) runWorker() {
for c.processNextItem() {
}
}
func (c *Controller) processNextItem() bool {
key, shutdown := c.queue.Get()
if shutdown {
return false
}
defer c.queue.Done(key) // ← fires per item
...
}
I avoided this one because a commenter on the tutorial I was following had already caught it. I'd have written the broken version otherwise, and I don't think I'd have found it quickly — a controller that processes each object exactly once and then goes quiet looks a lot like a controller with nothing to do.
What changed about how I debug Karpenter
This is the part I actually wanted.
Memory sizing stopped being arbitrary. Karpenter's controller requests 2Gi. That's the informer caches — every Pod, Node, NodeClaim and NodePool held in RAM. When I size a controller now I'm sizing its caches, not guessing.
The restart delay makes sense. After a Karpenter restart there's a window where nothing provisions, plus a spike of API server load. That's every informer doing its initial LIST, and WaitForCacheSync holding the workers back until it's done — because a controller acting on a half-populated cache is worse than one that's briefly idle.
The "why is it doing nothing" question has a better answer. Controllers are level-triggered: they react to what's true now, not to which event fired. A controller sitting silent is usually correct, not stuck. Mine runs reconcile several times a minute and almost always decides there's nothing to do.
Our Terraform's ugliest code has a reason. We have a destroy-time provisioner with a bash loop — sleep 10, a counter, give up after 30 — waiting for NodeClaims to drain, and eventually force-stripping finalizers. I'd always read that as a hack. It is, but it's a necessary one: Karpenter puts a finalizer on every Node so it can drain pods and terminate the EC2 instance before the object disappears, and Terraform has no way to wait on a watch. The workqueue does that properly — exponential backoff, per-key, non-blocking. The bash loop is the same idea with none of the machinery.
Would I use raw client-go again?
No. For anything real I'd use controller-runtime — SetupWithManager does five things in one line that took me two evenings, and it does them better.
But I can now say what those five things are, and that's worth more to me than the two evenings. When something in Karpenter behaves oddly I'm no longer guessing at a black box; I'm reasoning about an informer, a queue and a reconcile loop I've written myself.
If you operate Kubernetes and haven't written a controller, I'd recommend it on exactly those grounds. Not because you'll ship it. Because the next time you debug one, you'll know what's inside.
Code is on GitHub. The README goes into the design decisions and the limitations in more detail.
Conclusion:
I went in expecting something that remembers and reacts. What's actually there is something that forgets everything, looks at the world fresh, and fixes the difference. That's the whole trick.
Top comments (1)
Great article