DEV Community

Cover image for Kubernetes Operator for BitCloudPhone Device Pools: From Zero to Production
Claude DEL
Claude DEL

Posted on

Kubernetes Operator for BitCloudPhone Device Pools: From Zero to Production

Managing 200 cloud phone instances through a REST client is fine until the day one region goes offline, half your devices die, and you find yourself typing curl commands at 2 a.m. to rebuild the pool. The Kubernetes Operator pattern was invented for exactly this problem: declare what you want, let a controller reconcile reality to match.

This article walks through building a production Operator for BitCloudPhone device pools using Kubebuilder and Go. CRD schema, reconcile loop, finalizers, metrics, the three gotchas that will bite you in week two. Everything is code you can actually paste into a fresh cluster.

What a Kubernetes Operator is

An Operator is a custom Kubernetes controller paired with one or more Custom Resource Definitions (CRDs). The CRD extends the Kubernetes API with a new object type (like DevicePool or MobileFleet). The controller watches those objects and takes action to make the real world match the spec.

For cloud phones, this means you write YAML like this:

apiVersion: mobile.example.com/v1alpha1
kind: DevicePool
metadata:
  name: tiktok-us-pool
spec:
  provider: bitcloudphone
  platform: android
  region: us-west
  osVersion: "14"
  replicas: 50
  proxy:
    poolRef: proxyseller-us
    stickyTTL: 900
Enter fullscreen mode Exit fullscreen mode

You kubectl apply -f that file, and 50 US Android devices get provisioned, health-checked, wired to the right proxy pool, and kept alive. Delete the file, they get torn down. Change replicas: 50 to replicas: 80, thirty new devices show up within two minutes.

Why cloud phone pools fit this pattern

Three properties make cloud phones a natural fit for the Operator model:

  1. The provider exposes a REST API. BitCloudPhone Android has endpoints for device creation, status queries, remote reboot, and teardown. Same shape as any cloud VM API, which is the pattern Operators were designed around.
  2. Devices have lifecycle state. provisioning, ready, busy, offline, terminated. Kubernetes already excels at reconciling state machines, and the controller-runtime library gives you the wiring for free.
  3. Pools grow and shrink predictably. TikTok Shop campaigns spike traffic on weekends. A declarative replicas field beats a Slack message that says "spin up 40 more phones by Friday."

Project scaffolding with Kubebuilder

Kubebuilder is the reference SDK for building Operators in Go. Install it, then:

mkdir bitcloudphone-operator && cd bitcloudphone-operator
kubebuilder init --domain example.com --repo github.com/you/bitcloudphone-operator
kubebuilder create api --group mobile --version v1alpha1 --kind DevicePool
Enter fullscreen mode Exit fullscreen mode

Two files matter after scaffolding:

  • api/v1alpha1/devicepool_types.go holds the CRD schema (Go structs with +kubebuilder markers)
  • internal/controller/devicepool_controller.go holds the reconcile logic

The CRD schema

type DevicePoolSpec struct {
    // +kubebuilder:validation:Enum=bitcloudphone
    Provider string `json:"provider"`

    // +kubebuilder:validation:Enum=android;ios
    Platform string `json:"platform"`

    // +kubebuilder:validation:MinLength=1
    Region string `json:"region"`

    OSVersion string `json:"osVersion,omitempty"`

    // +kubebuilder:validation:Minimum=0
    // +kubebuilder:validation:Maximum=500
    Replicas int32 `json:"replicas"`

    Proxy ProxyRef `json:"proxy,omitempty"`
}

type DevicePoolStatus struct {
    ReadyDevices     int32              `json:"readyDevices"`
    ProvisioningDevices int32           `json:"provisioningDevices"`
    OfflineDevices   int32              `json:"offlineDevices"`
    LastReconciled   metav1.Time        `json:"lastReconciled,omitempty"`
    DeviceIDs        []string           `json:"deviceIds,omitempty"`
    Conditions       []metav1.Condition `json:"conditions,omitempty"`
}
Enter fullscreen mode Exit fullscreen mode

Cap Replicas at 500 in the CRD itself. This is the single line that will save you from a bad kubectl apply accidentally requesting 50,000 devices at $0.10 an hour each.

The reconcile loop

The reconcile function is called every time something changes about a DevicePool object, or every 30 seconds if nothing changes (the requeue interval). It has one job: read spec, compare to reality, close the gap.

func (r *DevicePoolReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
    var pool mobilev1alpha1.DevicePool
    if err := r.Get(ctx, req.NamespacedName, &pool); err != nil {
        return ctrl.Result{}, client.IgnoreNotFound(err)
    }

    // Query BitCloudPhone for the current device count under this pool tag
    currentDevices, err := r.BCPClient.ListDevices(ctx, pool.Name)
    if err != nil {
        return ctrl.Result{RequeueAfter: 30 * time.Second}, err
    }

    delta := int(pool.Spec.Replicas) - len(currentDevices)

    switch {
    case delta > 0:
        // Provision (delta) new devices
        for i := 0; i < delta; i++ {
            _, err := r.BCPClient.CreateDevice(ctx, pool.Spec.ToCreateRequest())
            if err != nil {
                return ctrl.Result{RequeueAfter: 15 * time.Second}, err
            }
        }
    case delta < 0:
        // Terminate the (delta) oldest idle devices
        idle := filterIdle(currentDevices)
        toTerminate := idle[:min(len(idle), -delta)]
        for _, d := range toTerminate {
            if err := r.BCPClient.TerminateDevice(ctx, d.ID); err != nil {
                return ctrl.Result{RequeueAfter: 15 * time.Second}, err
            }
        }
    }

    // Update status
    pool.Status = buildStatus(currentDevices)
    if err := r.Status().Update(ctx, &pool); err != nil {
        return ctrl.Result{}, err
    }

    return ctrl.Result{RequeueAfter: 30 * time.Second}, nil
}
Enter fullscreen mode Exit fullscreen mode

Two things worth flagging in this loop:

  • Never provision or terminate in bulk without rate-limiting. BitCloudPhone caps concurrent device creation at 20 per minute per account. Fire 100 at once and 80 fail with 429. The loop above ignores that; a production version wraps each call in a token bucket.
  • Terminate only idle devices. Killing a busy device mid-session burns proxy binding, loses cookies, and orphans whatever automation was running. filterIdle() checks the device's last-activity timestamp and the current session flag.

Finalizers for clean teardown

When a DevicePool gets deleted, Kubernetes removes it from etcd immediately unless a finalizer blocks that. Without a finalizer, the controller never gets a chance to tear down the actual cloud phone instances.

const finalizerName = "mobile.example.com/finalizer"

if pool.DeletionTimestamp != nil {
    if controllerutil.ContainsFinalizer(&pool, finalizerName) {
        if err := r.BCPClient.TerminatePool(ctx, pool.Name); err != nil {
            return ctrl.Result{RequeueAfter: 30 * time.Second}, err
        }
        controllerutil.RemoveFinalizer(&pool, finalizerName)
        r.Update(ctx, &pool)
    }
    return ctrl.Result{}, nil
}
if !controllerutil.ContainsFinalizer(&pool, finalizerName) {
    controllerutil.AddFinalizer(&pool, finalizerName)
    r.Update(ctx, &pool)
}
Enter fullscreen mode Exit fullscreen mode

Without this, kubectl delete devicepool tiktok-us-pool returns success while 50 phones keep running and billing until you notice on the invoice.

Health checks and self-healing

The reconcile loop already handles the count. Health goes in a second controller loop that runs every 60 seconds:

  • Query each device's status via GET /api/v1/devices/{id}/status
  • If status is offline for more than 3 minutes, mark the device as failed and remove it from the pool
  • The main reconcile loop then sees readyDevices < replicas and provisions a replacement

The important design choice: never try to "restart" a broken device. Terminate it and provision a new one. Cloud phone provisioning is under 60 seconds; a restart of a wedged device can take 15 minutes and often fails anyway. Replace, do not repair.

Extending to iOS and the browser side


The same Operator handles iOS if you set platform: ios in the CRD spec. The controller routes the API call to BitCloudPhone iOS instead of the Android endpoint. Fingerprint behavior differs between the two platforms in ways that matter for account isolation; the difference is covered in this fingerprint isolation on iOS cloud writeup.

For desktop browser profiles, BitBrowser runs outside Kubernetes (GUI dependency, see the Docker Swarm article in this series). But you can add a second CRD, BrowserPool, whose controller talks to the BitBrowser LocalAPI on tagged worker nodes. Two CRDs, one Operator binary, one control plane for both surfaces.

Deployment and RBAC

The Operator itself runs as a Deployment in a namespace like mobile-system. RBAC needs:

  • Read/write access to devicepools and devicepools/status in the custom API group
  • Read access to Secrets in the same namespace (for the BitCloudPhone API key)
  • No cluster-wide permissions

Package it with Helm or Kustomize. Ship the CRD, the Deployment, the ServiceAccount, the Role and RoleBinding, and a ConfigMap for the reconcile interval and rate-limit values. Total footprint is under 200 lines of YAML.

Metrics that matter

Expose four Prometheus metrics from the Operator:

  • devicepool_replicas_desired (gauge, labeled by pool name and platform)
  • devicepool_replicas_ready (gauge)
  • devicepool_reconcile_duration_seconds (histogram)
  • devicepool_api_errors_total (counter, labeled by error type)

The first two catch drift. The third catches slow reconciles before they become timeouts. The fourth catches provider-side outages within one scrape interval.

Production gotchas

Three things break in the first month, none of them documented:

  • CRD schema changes require a controlled rollout. Renaming a spec field breaks every existing DevicePool object. Use conversion webhooks between v1alpha1 and v1beta1 when you evolve the API, or write a migration Job that reads all objects, transforms them, and re-applies. Do not just delete and recreate; the finalizers will hang.
  • BitCloudPhone rate-limits by account, not by API key. Multiple API keys under the same account share the same 20-per-minute cap. Splitting keys does nothing. To scale past that, run separate accounts and route pools to different accounts via a providerCredentialRef in the spec.
  • etcd fills up with high-frequency status updates. If your reconcile loop updates status every 30 seconds and you run 40 pools, that is 115,200 status writes per day. Cluster etcd was not designed for that. Update status only when values change, not on every reconcile.

FAQ

Do I need Kubebuilder specifically, or can I use Operator SDK or Kopf?
All three work. Kubebuilder and Operator SDK both produce Go code and share the same controller-runtime under the hood. Kopf lets you write Operators in Python if the team is Python-native. Pick the language your on-call rotation can debug at 2 a.m.

Can this run on a managed Kubernetes service like GKE or EKS?
Yes. The Operator has no special requirements; it needs outbound HTTPS to the BitCloudPhone API and standard RBAC. A t3.small node runs the Operator itself; the cloud phones live on the provider's infrastructure, not yours.

What happens if the Operator pod crashes?
Kubernetes restarts it. On restart, the reconcile loop reads every DevicePool object and compares to reality. Any drift accumulated during the outage gets fixed on the next reconcile. This is the whole point of the pattern: state lives in etcd and the provider API, not in the controller's memory.

How do I test the reconcile logic locally?
Use envtest from controller-runtime. It runs a real etcd and kube-apiserver in-process, mocks the BitCloudPhone client, and lets you assert reconcile behavior end to end. A full test suite for a two-CRD Operator runs in about 40 seconds.

Is this overkill for 20 devices?
Yes. Under 50 devices, a cron job calling curl is fine. The Operator earns its complexity above 100 devices, multiple regions, or when device pools are managed by more than one person.

Where to take it from here

Start with a single-namespace deployment, one DevicePool object, five replicas. Prove the reconcile loop, the finalizer, and the metrics work end to end against a real BitCloudPhone account. Only then add the second CRD, the second controller, and the migration webhooks.

The Operator pattern rewards patience. Every production Operator I have shipped started as a 500-line controller that did one thing well, then grew. The teams that started with 2,000 lines never got them working.


Affiliate disclosure: This post contains affiliate links. If you sign up for a paid BitCloudPhone or BitBrowser plan through the links above, I may earn a small commission at no additional cost to you. All numbers above come from Operators I run against my own paid accounts.

Top comments (0)