DEV Community

Cover image for vCluster tutorial: virtual Kubernetes clusters per team
Billy Walker for Core Solutions

Posted on Originally published at coresolutions.ltd

vCluster tutorial: virtual Kubernetes clusters per team

Namespaces are a good isolation boundary right up until two teams need different cluster-scoped things. The first time one team wants a newer operator CRD, or a webhook configuration that does not match the rest of the cluster, the argument stops being about YAML and starts being about tenancy. A full cluster per team fixes it, but that is an expensive answer to a problem which is usually about control-plane isolation rather than worker-node count.

vCluster fills that gap. Each team gets its own Kubernetes API server, its own CRDs and its own RBAC, while their workloads keep running on a shared host cluster. The tenant team gets a real cluster experience; the platform team gets a much smaller footprint. One caveat belongs up here rather than buried at the bottom: this raises the isolation bar, but on shared nodes it is not a hard security boundary. Teams still share the host kernel and the same underlying worker estate.

Why namespaces stop being enough

Namespace isolation covers a lot. Where it stops is the set of things Kubernetes makes cluster-scoped by design:

  • CRDs and their versions are global to the cluster.
  • Admission webhooks are global to the cluster.
  • ClusterRoles and bindings are global to the cluster.
  • API-server flags and behaviour are global to the cluster.

The failure mode is familiar. Team A wants one version of an operator, Team B wants another, and both assume the cluster is theirs to shape. In a shared cluster one of them loses, and the workaround is political rather than technical: a platform review, a freeze window, or a quiet agreement that somebody will wait a quarter.

Once you are arbitrating CRD versions in a meeting, the abstraction is wrong. The question is no longer "how do we organise namespaces?" but "how do we let teams own cluster-shaped things without buying a cluster each?"

What a vCluster actually is

A vCluster is a Kubernetes control plane running inside another cluster. The tenant sees a normal Kubernetes API surface. The host sees a StatefulSet and some pods in one namespace.

The shared-nodes model is the one most platform teams mean when they say vCluster. The control plane runs as a StatefulSet in the host cluster, backed by its own datastore, with a syncer that maps selected resources between the tenant and the host. The default distribution is vanilla Kubernetes, and the default datastore is an embedded SQLite database.

The architecture explains both the appeal and the limits:

  • the tenant gets its own API server, CRDs and RBAC model
  • the host keeps the operational density of a shared cluster
  • the syncer decides which resources reach the host and which stay virtual
  • worker isolation is only as strong as the deployment model underneath

In one line: vCluster virtualises the control plane, not the physics of the host nodes.

The default image is not the OSS build

Before you install anything, check which image you are about to run, because the default is not the one most people assume.

The Helm chart defaults controlPlane.statefulSet.image.repository to loft-sh/vcluster-pro. The chart's own comment is explicit about it: that image carries the optional pro modules, turned off by default, and you set the repository to loft-sh/vcluster-oss if you want the pure Apache-2.0 build.

Neither choice is wrong. But an OSS-first install should name the image it wants, rather than inherit a default that brings the pro modules along with it.

Create one with the OSS image pinned

You need two things to follow along: a Kubernetes cluster with kubectl pointed at it, and a default StorageClass, because the tenant control plane runs as a StatefulSet and wants a volume. Any cluster will do — the host does not have to be large to hold a few tenant control planes.

The CLI comes from a tap, or from the get-started docs if you would rather not use Homebrew:

brew install loft-sh/tap/vcluster
Enter fullscreen mode Exit fullscreen mode

A minimal values file for an OSS-first, shared-nodes cluster:

controlPlane:
  statefulSet:
    image:
      repository: loft-sh/vcluster-oss
policies:
  podSecurityStandard: baseline
  resourceQuota:
    enabled: true
  limitRange:
    enabled: true
Enter fullscreen mode Exit fullscreen mode

Then create it:

vcluster create demo \
  --namespace team-a \
  --values vcluster.yaml \
  --connect=false
Enter fullscreen mode Exit fullscreen mode

--connect defaults to true, so without that flag the CLI rewrites your kubeconfig context the moment creation finishes. Fine interactively, surprising in a script that had other plans. When you do want to move into the tenant context:

vcluster connect demo --namespace team-a
vcluster list --namespace team-a
vcluster disconnect
Enter fullscreen mode Exit fullscreen mode

disconnect only puts your kubeconfig context back; the tenant cluster keeps running. Removing it is a separate verb, and it takes the namespace with it:

vcluster delete demo --namespace team-a
Enter fullscreen mode Exit fullscreen mode

If you would rather drive Helm yourself, the chart lives at https://charts.loft.sh and installs through the usual helm upgrade --install path.

Prove it: a CRD the host never sees

This is the demo that explains why a team puts up with a second control plane. Connect to the tenant cluster and install something that is normally cluster-scoped:

apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: widgets.demo.coresolutions.ltd
spec:
  group: demo.coresolutions.ltd
  scope: Namespaced
  names:
    plural: widgets
    singular: widget
    kind: Widget
  versions:
    - name: v1alpha1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              properties:
                size:
                  type: string
Enter fullscreen mode Exit fullscreen mode

Apply it inside the tenant cluster and it behaves exactly as it would in a dedicated one:

vcluster connect demo --namespace team-a
kubectl apply -f widget-crd.yaml
kubectl get crd widgets.demo.coresolutions.ltd
Enter fullscreen mode Exit fullscreen mode

Now step back out to the host and ask the same question:

vcluster disconnect
kubectl get crd widgets.demo.coresolutions.ltd
Enter fullscreen mode Exit fullscreen mode
Error from server (NotFound): customresourcedefinitions.apiextensions.k8s.io
"widgets.demo.coresolutions.ltd" not found
Enter fullscreen mode Exit fullscreen mode

The CRD exists, it is served, and the host control plane has never heard of it. Swap the CRD for a competing operator version and you have the whole argument in two commands.

What the host does have is one namespace:

kubectl get pods -n team-a
Enter fullscreen mode Exit fullscreen mode

You will see the demo-0 control-plane pod, plus any synced workload pods under names like web-x-default-x-demo — the syncer flattens <name>-x-<tenant-namespace>-x-<vcluster> into the single host namespace. The platform team keeps one namespace to reason about. The tenant keeps its own API surface, CRDs, ClusterRoles and admission chain. Nobody has to let a tenant mutate the real host control plane to get there.

What the host actually sees

Sync defaults are where vCluster stops being magical and becomes something you can predict. The shipped defaults enable host sync for a small, workload-facing set:

pods · services · endpoints · endpointSlices · persistentVolumeClaims · configMaps · secrets

Plenty of other types stay off in the host direction, including ingresses, networkPolicies, serviceAccounts, storageClasses, persistentVolumes and namespaces.

Host-to-tenant sync is narrower still. events are on. Nodes and secrets are off. A few storage-related types — csiDrivers, csiNodes, csiStorageCapacities, storageClasses — sit on auto, meaning they switch on only when something in your configuration needs them.

So the host only ever sees the subset required to actually run tenant workloads, which is why the shared-nodes model stays lighter than a real cluster while still feeling like one from the inside.

Two layers of policy, not one

Once the host sees a single tenant namespace plus synced workloads, platform controls get much easier to place. You wrap the host namespace as the outer boundary, then tighten things inside the tenant cluster where it matters.

The built-in policies: block is opinionated enough to be a starting point rather than an empty frame:

  • resourceQuota — ships on auto, with counts and CPU/memory request and limit ceilings already filled in
  • limitRange — also auto; defaults of cpu: 1 and memory: 512Mi, with requests of 100m and 128Mi
  • networkPolicy — scaffolding present, off by default
  • podSecurityStandard — unset by default; baseline, restricted or privileged when you want enforcement

That gives you two real layers: host-cluster controls around the tenant namespace as a whole, and tenant-cluster controls for the workloads the team actually runs. The team gets a real-enough cluster; the platform still defines the blast radius.

Ephemeral environments per pull request

Per-PR environments are one of the cleanest fits, because the lifecycle is obvious — create, test, destroy:

vcluster create pr-123 --namespace pr-123 --connect=false --values vcluster.yaml
vcluster connect pr-123 --namespace pr-123 -- kubectl get ns
# install and test here
vcluster delete pr-123 --namespace pr-123
Enter fullscreen mode Exit fullscreen mode

Creating a control plane in a namespace is operationally lighter than provisioning a cluster, which is the entire reason this pattern is attractive for short-lived environments. How much lighter depends on your host cluster, your CI, and what you install into each environment, so measure it on your own estate rather than trusting anyone's benchmark. If these environments are driven from Git, the same argument applies as in GitOps for Kubernetes with Argo CD: make the teardown as automatic as the creation, or you will be paying for PR environments from last March.

Where the shared-nodes model stops

Two limits are worth naming before you build on this, because neither is obvious from a working demo.

The first is the one from the top of this post: on shared nodes this is not a hard security boundary. Tenants get their own control plane, and they still share the host kernel and the same worker nodes. It is the right tool for teams that need autonomy from each other, and the wrong answer to a hostile-tenant or compliance requirement. vCluster does offer stronger models — dedicated, private and standalone nodes — and the step between them is larger than it first appears, which is a post of its own.

The second is the datastore. That embedded SQLite default is fine for a PR environment that lives for forty minutes, and it is one writer in one pod against one volume. Nothing replicates it, and nothing warns you when a two-week experiment becomes the thing three teams deploy to. A tenant cluster you would be upset to lose needs a real backing store and a backup you have actually restored from.

Where to start

Create one vCluster for the team currently waiting on a CRD argument, pin loft-sh/vcluster-oss, and let them use it properly for a fortnight with their own operators in it. That gets you the demo above with real workloads instead of a toy CRD, and it surfaces the questions worth answering before anything depends on the answers.

The first one that will bite: whether the thing you just built is a scratch environment or a cluster someone is relying on. Those want different datastores, and it is cheaper to decide early than to migrate a tenant later.

Top comments (0)