DEV Community

Yu Ting Chen
Yu Ting Chen

Posted on AI-assisted

A Team Asked for One More Service, and My Kubernetes Platform Quietly Refused

I'm building a small internal developer platform on Kubernetes. The idea is simple. A team writes a short request that says "this is my team, this is my service, run it". The platform does the rest: it gives the team its own space in the cluster, the right access, a resource budget, and it deploys the service.

It worked for the first service. Then I added a second service for the same team. Kubernetes accepted the request. No error came back. The new service just never became ready.

This post is about what was going on, how I thought it through, and what I changed. If you build things on Kubernetes that create other things, you may run into the same wall.

What the platform does

Kubernetes lets you add your own kinds of objects, called custom resources. You then write a controller: a program that watches those objects and makes the cluster match what they describe.

My platform had one custom resource, a ServiceClaim. A team's request looked like this:

kind: ServiceClaim
metadata:
  name: payments          # the service name
spec:
  team: payments
  image: nginx:1.27-alpine   # a stand-in for the real service; any image works
  replicas: 2
  resources:              # the team's resource budget
    cpu: "2"
    memory: 4Gi
    pods: 10
Enter fullscreen mode Exit fullscreen mode

For each claim, the controller created four things:

  1. a namespace for the team, team-payments (a namespace is a separate space inside the cluster)
  2. a RoleBinding that gives the team edit access in that namespace
  3. a ResourceQuota that caps what the team can use
  4. an ArgoCD Application that deploys the service (ArgoCD is a tool that keeps the cluster in sync with manifests in Git)

Why the second service got stuck

In Kubernetes, an object can record who owns it. One of those owners can be marked as the controller: the one object responsible for it. Kubernetes allows only one controller owner per object. The library my controller uses, controller-runtime, checks this rule before it even sends an update, and returns an AlreadyOwnedError if the object already has one.

My controller made the ServiceClaim the controller owner of all four objects. That's fine for one service. But look at the list again. Only the last one belongs to a service. The namespace, the access and the budget belong to the team.

So the second claim for the same team asked to own a namespace that the first claim already owned. Kubernetes' rule said no. The controller stopped at step one and retried, over and over. The claim's status showed it:

NamespaceReady=False (NamespaceError) Object /team-payments is already owned by another ServiceClaim controller payments
Ready=False (ResourcesNotReady) one of the claim's resources is not ready
Enter fullscreen mode Exit fullscreen mode

Nothing crashed. The request was accepted, the first service kept running, and the second one waited forever. That kind of failure is easy to miss, because nothing looks broken until someone asks where their service is.

Before: one ServiceClaim owns the namespace, RoleBinding, quota and its Application, so a second claim for the same team hits AlreadyOwnedError. After: a Tenant owns the three team-level objects and each ServiceClaim owns only its own Application.

How I thought about it

The error message was clear. The harder question was why my design asked for this at all. So before fixing anything, I listed the obvious fixes and asked what each one would lock in.

Block the second claim. A validation check could reject any second claim for the same team. It's the simplest fix, and it's the wrong one. It turns a bug into a rule. Teams still can't run two services, they just get told so earlier.

Let claims share the team's objects. Each claim could add itself to the team's objects and keep a count. The namespace would go away only when the last claim did. But the controller would then have to add up every claim's budget into one quota, and track who is still around. That's a lot of code to work around the shape of one object.

Make one object per team, with a list of services inside. No conflict, since there is only one owner. But changing one service would mean editing the whole team's list. That's not how teams think. They want to "deploy a service", not "re-declare my team".

There was also a second problem hiding behind the first. Even if the ownership conflict were solved another way, each claim carried its own resources budget. Two claims would keep writing two different budgets into the same quota. Nobody had hit that yet, because the second claim never got past the namespace. But on paper, the design was already wrong.

Looking at the three options side by side made the real issue easier to see. One object was carrying two different sizes of thing: team things and service things. The fix was to split it along that line.

The fix: one object for the team, one for each service

I added a second custom resource, the Tenant, for the team:

kind: Tenant
metadata:
  name: payments          # the team name
spec:
  resources:
    cpu: "2"
    memory: 4Gi
    pods: 10
Enter fullscreen mode Exit fullscreen mode

The Tenant owns the namespace, the access and the budget. Its name is the team name, so "one Tenant per team" is guaranteed by Kubernetes itself. A Tenant is cluster-wide, and two cluster-wide objects of the same kind can't share a name. I didn't have to write any code for it.

The ServiceClaim became smaller. It keeps only what a service needs, team, image and replicas, and it owns only its own ArgoCD Application. A team can now have as many claims as it wants, all pointing at the same Tenant.

The split raised new questions

Two objects that depend on each other bring questions that one object never had. I answered two of them before calling the fix done.

What if the service arrives before the team? Teams often apply a folder with a Tenant and several ServiceClaims in one go, and Kubernetes doesn't promise any order. If a claim were rejected because its Tenant wasn't there yet, a valid folder would fail or succeed by luck. So a claim without a ready Tenant doesn't fail. It reports TenantReady=False and waits, and it continues on its own when the Tenant is ready. "Not yet" and "wrong" are different states, and only "wrong" deserves an error.

What happens on delete? Kubernetes cleans up owned objects automatically when their owner is deleted, but it doesn't do it in any order. Deleting a Tenant could remove the namespace while services were still running in it. To control the order, I used finalizers. A finalizer is a marker on an object that makes Kubernetes wait, before deleting it, until a controller has done its cleanup. Three of them work together:

  • The ServiceClaim finalizer deletes the claim's ArgoCD Application, and waits until it's really gone.
  • The Application carries ArgoCD's own finalizer. That's what makes ArgoCD remove the running service before the Application itself disappears. I found this out while building the delete path: without it, deleting the Application only makes ArgoCD forget the service, and the service keeps running with nothing managing it.
  • The Tenant finalizer refuses to delete the team while any claim still points at it. So the namespace stays alive long enough for every service to be removed first.

One gap is known and written down rather than fixed: deleting a Tenant with kubectl delete --cascade=foreground removes the namespace before the finalizer runs. The default delete is the supported path.

What it cost

The split isn't free. One simple object became two objects with a relationship: a reference, a waiting state, and finalizers that have to run in the right order. That's more to reason about and more to test. I think it's the right trade here, but it's a trade, not a pure win.

If you run into something similar

This is the approach I'd take again:

  1. Treat the error as a question about the design. AlreadyOwnedError wasn't the bug. It was Kubernetes pointing at an object that owned things at two different levels.
  2. For each quick fix, ask what it would lock in. The simplest fix here would have made the limitation permanent.
  3. Split objects by what they really belong to. Team things and service things have different lifetimes. Putting them in one object is where the conflict came from.
  4. Answer the questions the fix creates. Order of arrival and order of deletion were new problems that only existed after the split.
  5. Write down the questions you're not solving yet. A month before this bug, I had noted that "who owns what" between my custom resources would need real design once there was more than one. The note didn't predict this bug, but when it came, the question was already on paper.

Try it yourself

The original code that failed isn't in the repository's history, because the history was squashed. So I rebuilt that earlier version to show the failure. It runs in its own small local cluster, and it prints the same status you saw above.

  • The failing version: scenarios/ownership-conflict. The README there has the steps. You need Docker, k3d, kubectl, the task runner and Go.
  • The current design: idp-platform-lab. The README's quick start brings up the whole platform, and docs/verification.md walks from a request to a running service.

I used AI to help draft and edit the English. I checked the technical claims against the code and a live reproduction.

Top comments (0)