Namespaces are a good isolation boundary right up until two teams need different cluster-scoped things. The first time one team wants a newer operator CRD, or a webhook configuration that does not match the rest of the cluster, the argument stops being about YAML and starts being about tenancy. A full cluster per team fixes it, but that is an expensive answer to a problem which is usually about control-plane isolation rather than worker-node count.
vCluster fills that gap. Each team gets its own Kubernetes API server, its own CRDs and its own RBAC, while their workloads keep running on a shared host cluster. The tenant team gets a real cluster experience; the platform team gets a much smaller footprint. One caveat belongs up here rather than buried at the bottom: this raises the isolation bar, but on shared nodes it is not a hard security boundary. Teams still share the host kernel and the same underlying worker estate.
Why namespaces stop being enough
Namespace isolation covers a lot. Where it stops is the set of things Kubernetes makes cluster-scoped by design:
- CRDs and their versions are global to the cluster.
- Admission webhooks are global to the cluster.
- ClusterRoles and bindings are global to the cluster.
- API-server flags and behaviour are global to the cluster.
The failure mode is familiar. Team A wants one version of an operator, Team B wants another, and both assume the cluster is theirs to shape. In a shared cluster one of them loses, and the workaround is political rather than technical: a platform review, a freeze window, or a quiet agreement that somebody will wait a quarter.
Once you are arbitrating CRD versions in a meeting, the abstraction is wrong. The question is no longer "how do we organise namespaces?" but "how do we let teams own cluster-shaped things without buying a cluster each?"
What a vCluster actually is
A vCluster is a Kubernetes control plane running inside another cluster. The tenant sees a normal Kubernetes API surface. The host sees a StatefulSet and some pods in one namespace.
The shared-nodes model is the one most platform teams mean when they say vCluster. The control plane runs as a StatefulSet in the host cluster, backed by its own datastore, with a syncer that maps selected resources between the tenant and the host. The default distribution is vanilla Kubernetes, and the default datastore is an embedded SQLite database.
The architecture explains both the appeal and the limits:
- the tenant gets its own API server, CRDs and RBAC model
- the host keeps the operational density of a shared cluster
- the syncer decides which resources reach the host and which stay virtual
- worker isolation is only as strong as the deployment model underneath
In one line: vCluster virtualises the control plane, not the physics of the host nodes.
The default image is not the OSS build
Before you install anything, check which image you are about to run, because the default is not the one most people assume.
The Helm chart defaults controlPlane.statefulSet.image.repository to loft-sh/vcluster-pro. The chart's own comment is explicit about it: that image carries the optional pro modules, turned off by default, and you set the repository to loft-sh/vcluster-oss if you want the pure Apache-2.0 build.
Neither choice is wrong. But an OSS-first install should name the image it wants, rather than inherit a default that brings the pro modules along with it.
Create one with the OSS image pinned
You need two things to follow along: a Kubernetes cluster with kubectl pointed at it, and a default StorageClass, because the tenant control plane runs as a StatefulSet and wants a volume. Any cluster will do — the host does not have to be large to hold a few tenant control planes.
The CLI comes from a tap, or from the get-started docs if you would rather not use Homebrew:
brew install loft-sh/tap/vcluster
A minimal values file for an OSS-first, shared-nodes cluster:
controlPlane:
statefulSet:
image:
repository: loft-sh/vcluster-oss
policies:
podSecurityStandard: baseline
resourceQuota:
enabled: true
limitRange:
enabled: true
Then create it:
vcluster create demo \
--namespace team-a \
--values vcluster.yaml \
--connect=false
--connect defaults to true, so without that flag the CLI rewrites your kubeconfig context the moment creation finishes. Fine interactively, surprising in a script that had other plans. When you do want to move into the tenant context:
vcluster connect demo --namespace team-a
vcluster list --namespace team-a
vcluster disconnect
disconnect only puts your kubeconfig context back; the tenant cluster keeps running. Removing it is a separate verb, and it takes the namespace with it:
vcluster delete demo --namespace team-a
If you would rather drive Helm yourself, the chart lives at https://charts.loft.sh and installs through the usual helm upgrade --install path.
Prove it: a CRD the host never sees
This is the demo that explains why a team puts up with a second control plane. Connect to the tenant cluster and install something that is normally cluster-scoped:
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: widgets.demo.coresolutions.ltd
spec:
group: demo.coresolutions.ltd
scope: Namespaced
names:
plural: widgets
singular: widget
kind: Widget
versions:
- name: v1alpha1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
size:
type: string
Apply it inside the tenant cluster and it behaves exactly as it would in a dedicated one:
vcluster connect demo --namespace team-a
kubectl apply -f widget-crd.yaml
kubectl get crd widgets.demo.coresolutions.ltd
Now step back out to the host and ask the same question:
vcluster disconnect
kubectl get crd widgets.demo.coresolutions.ltd
Error from server (NotFound): customresourcedefinitions.apiextensions.k8s.io
"widgets.demo.coresolutions.ltd" not found
The CRD exists, it is served, and the host control plane has never heard of it. Swap the CRD for a competing operator version and you have the whole argument in two commands.
What the host does have is one namespace:
kubectl get pods -n team-a
You will see the demo-0 control-plane pod, plus any synced workload pods under names like web-x-default-x-demo — the syncer flattens <name>-x-<tenant-namespace>-x-<vcluster> into the single host namespace. The platform team keeps one namespace to reason about. The tenant keeps its own API surface, CRDs, ClusterRoles and admission chain. Nobody has to let a tenant mutate the real host control plane to get there.
What the host actually sees
Sync defaults are where vCluster stops being magical and becomes something you can predict. The shipped defaults enable host sync for a small, workload-facing set:
pods · services · endpoints · endpointSlices · persistentVolumeClaims · configMaps · secrets
Plenty of other types stay off in the host direction, including ingresses, networkPolicies, serviceAccounts, storageClasses, persistentVolumes and namespaces.
Host-to-tenant sync is narrower still. events are on. Nodes and secrets are off. A few storage-related types — csiDrivers, csiNodes, csiStorageCapacities, storageClasses — sit on auto, meaning they switch on only when something in your configuration needs them.
So the host only ever sees the subset required to actually run tenant workloads, which is why the shared-nodes model stays lighter than a real cluster while still feeling like one from the inside.
Two layers of policy, not one
Once the host sees a single tenant namespace plus synced workloads, platform controls get much easier to place. You wrap the host namespace as the outer boundary, then tighten things inside the tenant cluster where it matters.
The built-in policies: block is opinionated enough to be a starting point rather than an empty frame:
-
resourceQuota— ships onauto, with counts and CPU/memory request and limit ceilings already filled in -
limitRange— alsoauto; defaults ofcpu: 1andmemory: 512Mi, with requests of100mand128Mi -
networkPolicy— scaffolding present, off by default -
podSecurityStandard— unset by default;baseline,restrictedorprivilegedwhen you want enforcement
That gives you two real layers: host-cluster controls around the tenant namespace as a whole, and tenant-cluster controls for the workloads the team actually runs. The team gets a real-enough cluster; the platform still defines the blast radius.
Ephemeral environments per pull request
Per-PR environments are one of the cleanest fits, because the lifecycle is obvious — create, test, destroy:
vcluster create pr-123 --namespace pr-123 --connect=false --values vcluster.yaml
vcluster connect pr-123 --namespace pr-123 -- kubectl get ns
# install and test here
vcluster delete pr-123 --namespace pr-123
Creating a control plane in a namespace is operationally lighter than provisioning a cluster, which is the entire reason this pattern is attractive for short-lived environments. How much lighter depends on your host cluster, your CI, and what you install into each environment, so measure it on your own estate rather than trusting anyone's benchmark. If these environments are driven from Git, the same argument applies as in GitOps for Kubernetes with Argo CD: make the teardown as automatic as the creation, or you will be paying for PR environments from last March.
Where the shared-nodes model stops
Two limits are worth naming before you build on this, because neither is obvious from a working demo.
The first is the one from the top of this post: on shared nodes this is not a hard security boundary. Tenants get their own control plane, and they still share the host kernel and the same worker nodes. It is the right tool for teams that need autonomy from each other, and the wrong answer to a hostile-tenant or compliance requirement. vCluster does offer stronger models — dedicated, private and standalone nodes — and the step between them is larger than it first appears, which is a post of its own.
The second is the datastore. That embedded SQLite default is fine for a PR environment that lives for forty minutes, and it is one writer in one pod against one volume. Nothing replicates it, and nothing warns you when a two-week experiment becomes the thing three teams deploy to. A tenant cluster you would be upset to lose needs a real backing store and a backup you have actually restored from.
Where to start
Create one vCluster for the team currently waiting on a CRD argument, pin loft-sh/vcluster-oss, and let them use it properly for a fortnight with their own operators in it. That gets you the demo above with real workloads instead of a toy CRD, and it surfaces the questions worth answering before anything depends on the answers.
The first one that will bite: whether the thing you just built is a scratch environment or a cluster someone is relying on. Those want different datastores, and it is cheaper to decide early than to migrate a tenant later.
Top comments (0)