Managing 200 cloud phone instances through a REST client is fine until the day one region goes offline, half your devices die, and you find yourself typing curl commands at 2 a.m. to rebuild the pool. The Kubernetes Operator pattern was invented for exactly this problem: declare what you want, let a controller reconcile reality to match.
This article walks through building a production Operator for BitCloudPhone device pools using Kubebuilder and Go. CRD schema, reconcile loop, finalizers, metrics, the three gotchas that will bite you in week two. Everything is code you can actually paste into a fresh cluster.
What a Kubernetes Operator is
An Operator is a custom Kubernetes controller paired with one or more Custom Resource Definitions (CRDs). The CRD extends the Kubernetes API with a new object type (like DevicePool or MobileFleet). The controller watches those objects and takes action to make the real world match the spec.
For cloud phones, this means you write YAML like this:
apiVersion: mobile.example.com/v1alpha1
kind: DevicePool
metadata:
name: tiktok-us-pool
spec:
provider: bitcloudphone
platform: android
region: us-west
osVersion: "14"
replicas: 50
proxy:
poolRef: proxyseller-us
stickyTTL: 900
You kubectl apply -f that file, and 50 US Android devices get provisioned, health-checked, wired to the right proxy pool, and kept alive. Delete the file, they get torn down. Change replicas: 50 to replicas: 80, thirty new devices show up within two minutes.
Why cloud phone pools fit this pattern
Three properties make cloud phones a natural fit for the Operator model:
- The provider exposes a REST API. BitCloudPhone Android has endpoints for device creation, status queries, remote reboot, and teardown. Same shape as any cloud VM API, which is the pattern Operators were designed around.
-
Devices have lifecycle state.
provisioning,ready,busy,offline,terminated. Kubernetes already excels at reconciling state machines, and the controller-runtime library gives you the wiring for free. -
Pools grow and shrink predictably. TikTok Shop campaigns spike traffic on weekends. A declarative
replicasfield beats a Slack message that says "spin up 40 more phones by Friday."
Project scaffolding with Kubebuilder
Kubebuilder is the reference SDK for building Operators in Go. Install it, then:
mkdir bitcloudphone-operator && cd bitcloudphone-operator
kubebuilder init --domain example.com --repo github.com/you/bitcloudphone-operator
kubebuilder create api --group mobile --version v1alpha1 --kind DevicePool
Two files matter after scaffolding:
-
api/v1alpha1/devicepool_types.goholds the CRD schema (Go structs with+kubebuildermarkers) -
internal/controller/devicepool_controller.goholds the reconcile logic
The CRD schema
type DevicePoolSpec struct {
// +kubebuilder:validation:Enum=bitcloudphone
Provider string `json:"provider"`
// +kubebuilder:validation:Enum=android;ios
Platform string `json:"platform"`
// +kubebuilder:validation:MinLength=1
Region string `json:"region"`
OSVersion string `json:"osVersion,omitempty"`
// +kubebuilder:validation:Minimum=0
// +kubebuilder:validation:Maximum=500
Replicas int32 `json:"replicas"`
Proxy ProxyRef `json:"proxy,omitempty"`
}
type DevicePoolStatus struct {
ReadyDevices int32 `json:"readyDevices"`
ProvisioningDevices int32 `json:"provisioningDevices"`
OfflineDevices int32 `json:"offlineDevices"`
LastReconciled metav1.Time `json:"lastReconciled,omitempty"`
DeviceIDs []string `json:"deviceIds,omitempty"`
Conditions []metav1.Condition `json:"conditions,omitempty"`
}
Cap Replicas at 500 in the CRD itself. This is the single line that will save you from a bad kubectl apply accidentally requesting 50,000 devices at $0.10 an hour each.
The reconcile loop
The reconcile function is called every time something changes about a DevicePool object, or every 30 seconds if nothing changes (the requeue interval). It has one job: read spec, compare to reality, close the gap.
func (r *DevicePoolReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
var pool mobilev1alpha1.DevicePool
if err := r.Get(ctx, req.NamespacedName, &pool); err != nil {
return ctrl.Result{}, client.IgnoreNotFound(err)
}
// Query BitCloudPhone for the current device count under this pool tag
currentDevices, err := r.BCPClient.ListDevices(ctx, pool.Name)
if err != nil {
return ctrl.Result{RequeueAfter: 30 * time.Second}, err
}
delta := int(pool.Spec.Replicas) - len(currentDevices)
switch {
case delta > 0:
// Provision (delta) new devices
for i := 0; i < delta; i++ {
_, err := r.BCPClient.CreateDevice(ctx, pool.Spec.ToCreateRequest())
if err != nil {
return ctrl.Result{RequeueAfter: 15 * time.Second}, err
}
}
case delta < 0:
// Terminate the (delta) oldest idle devices
idle := filterIdle(currentDevices)
toTerminate := idle[:min(len(idle), -delta)]
for _, d := range toTerminate {
if err := r.BCPClient.TerminateDevice(ctx, d.ID); err != nil {
return ctrl.Result{RequeueAfter: 15 * time.Second}, err
}
}
}
// Update status
pool.Status = buildStatus(currentDevices)
if err := r.Status().Update(ctx, &pool); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{RequeueAfter: 30 * time.Second}, nil
}
Two things worth flagging in this loop:
- Never provision or terminate in bulk without rate-limiting. BitCloudPhone caps concurrent device creation at 20 per minute per account. Fire 100 at once and 80 fail with 429. The loop above ignores that; a production version wraps each call in a token bucket.
-
Terminate only idle devices. Killing a busy device mid-session burns proxy binding, loses cookies, and orphans whatever automation was running.
filterIdle()checks the device's last-activity timestamp and the current session flag.
Finalizers for clean teardown
When a DevicePool gets deleted, Kubernetes removes it from etcd immediately unless a finalizer blocks that. Without a finalizer, the controller never gets a chance to tear down the actual cloud phone instances.
const finalizerName = "mobile.example.com/finalizer"
if pool.DeletionTimestamp != nil {
if controllerutil.ContainsFinalizer(&pool, finalizerName) {
if err := r.BCPClient.TerminatePool(ctx, pool.Name); err != nil {
return ctrl.Result{RequeueAfter: 30 * time.Second}, err
}
controllerutil.RemoveFinalizer(&pool, finalizerName)
r.Update(ctx, &pool)
}
return ctrl.Result{}, nil
}
if !controllerutil.ContainsFinalizer(&pool, finalizerName) {
controllerutil.AddFinalizer(&pool, finalizerName)
r.Update(ctx, &pool)
}
Without this, kubectl delete devicepool tiktok-us-pool returns success while 50 phones keep running and billing until you notice on the invoice.
Health checks and self-healing
The reconcile loop already handles the count. Health goes in a second controller loop that runs every 60 seconds:
- Query each device's status via
GET /api/v1/devices/{id}/status - If status is
offlinefor more than 3 minutes, mark the device as failed and remove it from the pool - The main reconcile loop then sees
readyDevices < replicasand provisions a replacement
The important design choice: never try to "restart" a broken device. Terminate it and provision a new one. Cloud phone provisioning is under 60 seconds; a restart of a wedged device can take 15 minutes and often fails anyway. Replace, do not repair.
Extending to iOS and the browser side

The same Operator handles iOS if you set platform: ios in the CRD spec. The controller routes the API call to BitCloudPhone iOS instead of the Android endpoint. Fingerprint behavior differs between the two platforms in ways that matter for account isolation; the difference is covered in this fingerprint isolation on iOS cloud writeup.
For desktop browser profiles, BitBrowser runs outside Kubernetes (GUI dependency, see the Docker Swarm article in this series). But you can add a second CRD, BrowserPool, whose controller talks to the BitBrowser LocalAPI on tagged worker nodes. Two CRDs, one Operator binary, one control plane for both surfaces.
Deployment and RBAC
The Operator itself runs as a Deployment in a namespace like mobile-system. RBAC needs:
- Read/write access to
devicepoolsanddevicepools/statusin the custom API group - Read access to Secrets in the same namespace (for the BitCloudPhone API key)
- No cluster-wide permissions
Package it with Helm or Kustomize. Ship the CRD, the Deployment, the ServiceAccount, the Role and RoleBinding, and a ConfigMap for the reconcile interval and rate-limit values. Total footprint is under 200 lines of YAML.
Metrics that matter
Expose four Prometheus metrics from the Operator:
-
devicepool_replicas_desired(gauge, labeled by pool name and platform) -
devicepool_replicas_ready(gauge) -
devicepool_reconcile_duration_seconds(histogram) -
devicepool_api_errors_total(counter, labeled by error type)
The first two catch drift. The third catches slow reconciles before they become timeouts. The fourth catches provider-side outages within one scrape interval.
Production gotchas
Three things break in the first month, none of them documented:
-
CRD schema changes require a controlled rollout. Renaming a spec field breaks every existing DevicePool object. Use conversion webhooks between
v1alpha1andv1beta1when you evolve the API, or write a migration Job that reads all objects, transforms them, and re-applies. Do not just delete and recreate; the finalizers will hang. -
BitCloudPhone rate-limits by account, not by API key. Multiple API keys under the same account share the same 20-per-minute cap. Splitting keys does nothing. To scale past that, run separate accounts and route pools to different accounts via a
providerCredentialRefin the spec. - etcd fills up with high-frequency status updates. If your reconcile loop updates status every 30 seconds and you run 40 pools, that is 115,200 status writes per day. Cluster etcd was not designed for that. Update status only when values change, not on every reconcile.
FAQ
Do I need Kubebuilder specifically, or can I use Operator SDK or Kopf?
All three work. Kubebuilder and Operator SDK both produce Go code and share the same controller-runtime under the hood. Kopf lets you write Operators in Python if the team is Python-native. Pick the language your on-call rotation can debug at 2 a.m.
Can this run on a managed Kubernetes service like GKE or EKS?
Yes. The Operator has no special requirements; it needs outbound HTTPS to the BitCloudPhone API and standard RBAC. A t3.small node runs the Operator itself; the cloud phones live on the provider's infrastructure, not yours.
What happens if the Operator pod crashes?
Kubernetes restarts it. On restart, the reconcile loop reads every DevicePool object and compares to reality. Any drift accumulated during the outage gets fixed on the next reconcile. This is the whole point of the pattern: state lives in etcd and the provider API, not in the controller's memory.
How do I test the reconcile logic locally?
Use envtest from controller-runtime. It runs a real etcd and kube-apiserver in-process, mocks the BitCloudPhone client, and lets you assert reconcile behavior end to end. A full test suite for a two-CRD Operator runs in about 40 seconds.
Is this overkill for 20 devices?
Yes. Under 50 devices, a cron job calling curl is fine. The Operator earns its complexity above 100 devices, multiple regions, or when device pools are managed by more than one person.
Where to take it from here
Start with a single-namespace deployment, one DevicePool object, five replicas. Prove the reconcile loop, the finalizer, and the metrics work end to end against a real BitCloudPhone account. Only then add the second CRD, the second controller, and the migration webhooks.
The Operator pattern rewards patience. Every production Operator I have shipped started as a 500-line controller that did one thing well, then grew. The teams that started with 2,000 lines never got them working.
Affiliate disclosure: This post contains affiliate links. If you sign up for a paid BitCloudPhone or BitBrowser plan through the links above, I may earn a small commission at no additional cost to you. All numbers above come from Operators I run against my own paid accounts.
Top comments (0)