DEV Community

Claude DEL
Claude DEL

Posted on

Running Playwright Across 100 Cloud Phone Devices with BitBrowser API


TL;DR — A single Playwright process on a laptop caps at 6–8 concurrent Chromium instances before memory pressure kills throughput. To run 100 real-device sessions in parallel, you offload the browser and the device fingerprint to a cloud phone fleet, keep Playwright as the orchestrator, and let the BitBrowser API allocate profiles per session. This guide walks the full stack: architecture, provisioning, session pool, proxy assignment, retry logic, monitoring, and the actual monthly cost.


Contents

  • Why a laptop can't do this
  • The stack in one diagram
  • Three pieces you actually need
  • Provisioning 100 cloud phones from a script
  • Connecting Playwright to Android and iOS
  • Session pool and queue pattern
  • Proxy assignment: one exit IP per device, always
  • Concurrency ceilings that actually hold
  • Retry logic that doesn't burn accounts
  • Monitoring 100 devices without a NOC
  • Real 2026 cost breakdown
  • Pitfalls that took me weeks to debug
  • FAQ

Why a laptop can't do this

I tried the naive setup first. A Ryzen 9 with 64 GB of RAM, headless Chromium, 30 Playwright contexts, all pointed at the same TikTok flow. Around context 8 the OOM killer hit. At context 12 the whole machine froze.

The problem isn't Playwright. Every context runs a full Chromium process with its own JS engine, GPU sandbox, and cache. Add a residential proxy per context and you pay network latency on every request. Add a real device fingerprint and you can't do it at all. A laptop only has one WebGL vendor string, one AudioContext hash, one canvas.

Cloud phones fix both problems in one move. Each phone is a real ARM device with its own hardware fingerprint. The browser runs on the phone. Your laptop only sends commands and receives DOM events. RAM stays free.

The stack in one diagram

┌──────────────┐    WebSocket / CDP    ┌──────────────────────┐
│  Playwright  │◄──────────────────────►│  Cloud Phone (x100)  │
│  Orchestrator│                        │  Android or iOS      │
│  (Node.js)   │                        │  Chrome / Safari     │
└──────┬───────┘                        └──────────┬───────────┘
       │                                           │
       │ REST                                      │ Egress
       ▼                                           ▼
┌──────────────┐                        ┌──────────────────────┐
│  BitBrowser  │                        │  Mobile/Residential  │
│  Local API   │                        │  Proxy (per device)  │
└──────────────┘                        └──────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The orchestrator holds the queue and the session pool. It talks to each cloud phone through a WebSocket that bridges to the on-device browser via the Chrome DevTools Protocol. It calls the BitBrowser local API to spin up isolated profiles for the browser-only tasks that don't need a physical device (email signup, admin panels, affiliate dashboards). Every device gets its own outbound proxy assigned once and reused for the lifetime of the account.

Three pieces you actually need

  1. BitBrowser runs the browser-side fingerprints: canvas, WebGL, WebRTC, fonts, ClientRects, AudioContext, timezone, screen. Its local REST API lets you create, open, close, and rotate profiles from Node.js in about 200 ms per profile.
  2. BitCloudPhone handles the Android side. Each instance is a real ARM device sitting in a data-center rack, not an emulator, so Play Integrity API returns MEETS_DEVICE_INTEGRITY and TikTok, Instagram, and Facebook don't flag the session as an emulator.
  3. BitCloudPhone iOS covers the Apple side. Physical iPhones in a rack, exposed over a WebDriverAgent endpoint. This is the only reliable way to run iMessage-linked, iCloud-locked, or Snapchat US accounts from outside the US without device-attestation failures.

Playwright is the glue. It uses chromium.connectOverCDP() for Android Chrome and an Appium bridge for iOS Safari.

Provisioning 100 cloud phones from a script

You don't click "Create device" 100 times. The provider exposes a REST API that returns device IDs and connection endpoints. A minimal batch provisioner in Node.js:

import fetch from 'node-fetch';

const CLOUD_API = 'https://api.your-provider.tld/v1';
const TOKEN = process.env.CLOUD_PHONE_TOKEN;

async function provisionBatch(count, region = 'us-east-1') {
  const devices = [];
  for (let i = 0; i < count; i++) {
    const res = await fetch(`${CLOUD_API}/devices`, {
      method: 'POST',
      headers: {
        'Authorization': `Bearer ${TOKEN}`,
        'Content-Type': 'application/json'
      },
      body: JSON.stringify({
        os: 'android',
        model: 'pixel_7',
        region,
        proxy_id: null // assigned separately
      })
    });
    devices.push(await res.json());
    await new Promise(r => setTimeout(r, 500)); // rate-limit courtesy
  }
  return devices;
}

const fleet = await provisionBatch(100);
console.log(`Provisioned ${fleet.length} devices`);
Enter fullscreen mode Exit fullscreen mode

Two things this snippet gets right that the naive version won't:

  • 500 ms delay between calls. Cloud APIs rate-limit at around 5–10 req/s. Batching without a delay gets you a 429 and a partial fleet.
  • proxy_id: null at creation. Attaching the proxy after the device is running lets you rotate proxies without device reboots.

Store the returned device IDs in Redis or a Postgres table. You'll refer to them every time you dispatch a job.

Connecting Playwright to Android and iOS

Playwright doesn't natively understand a remote Android Chrome. What it does understand is the Chrome DevTools Protocol. Every cloud phone provider exposes a WebSocket URL that proxies CDP messages to the on-device Chrome instance.

For Android:

import { chromium } from 'playwright';

async function attachToAndroid(deviceId) {
  const wsEndpoint = await getCdpEndpoint(deviceId); // provider-specific
  const browser = await chromium.connectOverCDP(wsEndpoint);
  const context = browser.contexts()[0];
  const page = context.pages()[0] ?? await context.newPage();
  return { browser, page };
}
Enter fullscreen mode Exit fullscreen mode

For iOS, Playwright can't drive Safari over CDP because Safari doesn't speak it. You bridge through Appium instead:

import { remote } from 'webdriverio';

async function attachToIos(deviceId) {
  const client = await remote({
    hostname: `ios-${deviceId}.your-provider.tld`,
    port: 4723,
    path: '/wd/hub',
    capabilities: {
      platformName: 'iOS',
      'appium:automationName': 'XCUITest',
      'appium:bundleId': 'com.apple.mobilesafari'
    }
  });
  return client;
}
Enter fullscreen mode Exit fullscreen mode

You lose some Playwright niceties on iOS (video recording, first-class trace viewer). You gain the ability to run Safari-only flows against real iPhones.

Session pool and queue pattern

100 devices doesn't mean 100 always-open sessions. Most jobs take 30–90 seconds. A pool with checkout/return semantics keeps utilization high without leaking browsers.

class DevicePool {
  constructor(devices) {
    this.available = [...devices];
    this.inUse = new Map();
    this.waiters = [];
  }

  async acquire(timeoutMs = 30000) {
    if (this.available.length > 0) {
      const device = this.available.shift();
      this.inUse.set(device.id, device);
      return device;
    }
    return new Promise((resolve, reject) => {
      const timer = setTimeout(() => {
        this.waiters = this.waiters.filter(w => w.resolve !== resolve);
        reject(new Error('pool acquire timeout'));
      }, timeoutMs);
      this.waiters.push({ resolve, timer });
    });
  }

  release(device) {
    this.inUse.delete(device.id);
    if (this.waiters.length > 0) {
      const w = this.waiters.shift();
      clearTimeout(w.timer);
      this.inUse.set(device.id, device);
      w.resolve(device);
    } else {
      this.available.push(device);
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Wrap every job:

async function runJob(pool, task) {
  const device = await pool.acquire();
  try {
    const { page } = await attachToAndroid(device.id);
    await task(page, device);
  } finally {
    pool.release(device);
  }
}
Enter fullscreen mode Exit fullscreen mode

The finally block is not optional. Any thrown error inside the task leaks the device otherwise, and after 5–10 leaks your pool is dead.

Proxy assignment: one exit IP per device, always

The single biggest mistake in multi-account automation is rotating the proxy per session on the same account. TikTok, Instagram, and Facebook read the IP + country + ASN triple and log it as part of the account's history. A US mobile carrier ASN on Monday and a Brazilian datacenter ASN on Tuesday is a ban signal, not fingerprint noise.

Pin the proxy to the device at provisioning time. Store the mapping. Never rotate unless the account is being reset.

async function assignProxy(deviceId, proxyEndpoint) {
  await fetch(`${CLOUD_API}/devices/${deviceId}/proxy`, {
    method: 'PUT',
    headers: { 'Authorization': `Bearer ${TOKEN}` },
    body: JSON.stringify({
      type: 'socks5',
      host: proxyEndpoint.host,
      port: proxyEndpoint.port,
      username: proxyEndpoint.username,
      password: proxyEndpoint.password
    })
  });
}
Enter fullscreen mode Exit fullscreen mode

Mobile 4G/5G proxies cost more (roughly $3 to $8 per GB in 2026), but they match what TikTok Shop, Snapchat, and Reddit's newer detection expects to see on a phone-shaped User-Agent. Residential works too at $2 to $4 per GB, but bans arrive faster on the strictest platforms.

Concurrency ceilings that actually hold

Just because you have 100 devices doesn't mean you can run 100 parallel actions on the same target platform. Real ceilings I've hit on TikTok Shop from a 100-device fleet:

Action type Safe concurrency Notes
Product page views 100 Read-only, no signal
Cart additions 40 Bursts above 60 trigger challenge
Account creation 8 to 12 Higher = SMS verification wall
Live viewer join 60 Live rooms count IPs closely
Comment posting 15 Comment velocity is heavily rate-limited

Enforce these at the queue level, not the pool level. A semaphore per target platform works:

import pLimit from 'p-limit';

const tiktokLimit = pLimit(12); // account creation
const igLimit = pLimit(20);

await tiktokLimit(() => runJob(pool, createTikTokAccount));
Enter fullscreen mode Exit fullscreen mode

Retry logic that doesn't burn accounts

Naive retry logic is the second-fastest way to lose accounts. If a Playwright action fails, the wrong response is to hammer the same account with 3 retries in 30 seconds. That pattern looks like a bot on any behavioral-scoring platform.

The rules I follow:

  • Transient network errors (ECONNRESET, WebSocket close): retry up to 2 times, with 5 s and 20 s backoff, on the same device.
  • HTTP 429 or platform "slow down" screen: park the device for 15 minutes. Don't move the job to a fresh device. The platform's rate limit is per-account, not per-IP.
  • HTTP 403 or account-flag screen: quarantine the device. Log it. Don't retry. Investigate before dispatching to that account again.
  • Element-not-found: retry once after a 3 s wait, then fail. UI changes are the number-one cause and retrying harder won't fix a moved selector.
async function withRetry(fn, { maxAttempts = 2, delayMs = 5000 } = {}) {
  let lastErr;
  for (let attempt = 1; attempt <= maxAttempts; attempt++) {
    try {
      return await fn();
    } catch (err) {
      lastErr = err;
      if (err.status === 429 || err.status === 403) throw err;
      if (attempt < maxAttempts) {
        await new Promise(r => setTimeout(r, delayMs * attempt));
      }
    }
  }
  throw lastErr;
}
Enter fullscreen mode Exit fullscreen mode

Monitoring 100 devices without a NOC

A dashboard doesn't need Grafana on day one. A single Postgres table plus a Slack webhook gets you 80% of what you need:

CREATE TABLE device_health (
  device_id TEXT PRIMARY KEY,
  last_heartbeat TIMESTAMPTZ,
  last_job_status TEXT,
  quarantined BOOLEAN DEFAULT false,
  proxy_endpoint TEXT,
  account_id TEXT
);
Enter fullscreen mode Exit fullscreen mode

Every job writes its status. A cron task every 5 minutes checks for devices with no heartbeat in 15 minutes and posts to Slack. When a device gets quarantined 3 times in 24 hours, page yourself.

My last fleet hit around 97% uptime across 100 devices. The 3% loss was nearly always proxy provider issues, not the cloud phones. A Playwright with BitBrowser reference covers the browser-side monitoring in more depth.

Real 2026 cost breakdown

Numbers from a fleet I ran through Q2 2026, 100 devices, 24/7:

Line item Monthly cost (USD)
100× Android cloud phones (mid-tier) ~$1,200
20× iOS cloud phones (US-restricted apps) ~$800
Mobile proxies (avg 4 GB/device/month) ~$1,600
BitBrowser subscription (team tier) ~$200
Managed Postgres + Redis (small) ~$60
Node.js orchestrator VPS (8 vCPU) ~$80
Total ~$3,940

Cost per active session-hour: about $0.055. That's the number to beat if you're evaluating build-vs-buy on a managed farm.

Pitfalls that took me weeks to debug

CDP WebSocket timeouts silently corrupt state. If the connection drops mid-action, Playwright keeps the context object valid client-side but the browser is already gone. Heartbeat the CDP connection every 30 s and rebuild on failure.

Time-zone drift. Cloud phones default to the datacenter's time zone. If your proxy exit is in Los Angeles and the device reports Asia/Shanghai, every Facebook login triggers a review. Set device time zone to match the proxy country at provisioning.

Chromium version mismatch. Playwright pins a specific Chromium build. The on-device Chrome updates on its own schedule. Some CDP calls disappear or change signature between versions. Pin the on-device Chrome, or handle version detection in your CDP layer.

iOS keychain persistence. iOS cloud phones sometimes retain Apple ID sessions across "reset" operations. Verify keychain clearance manually the first few times you rotate accounts.

Silent throttling from the cloud provider. After ~50 concurrent CDP connections to the same billing account, some providers throttle without a status code. Split fleets across two accounts if you exceed 50 concurrent sessions.

FAQ

Can I use free BrowserStack or Sauce Labs sessions instead of paid cloud phones?
No. Both platforms mark their sessions as testing devices in the User-Agent and network headers. TikTok, Instagram, and Facebook fingerprint them within one page load. Paid cloud phone providers give you real consumer devices with clean signatures.

How is this different from Playwright's built-in device emulation?
Device emulation only changes the User-Agent, viewport, and touch flag. Canvas, WebGL, AudioContext, and hardware sensors still report desktop values. Any platform that checks two or more of those signals flags the session as spoofed.

Do I need BitBrowser if I run everything on cloud phones?
For pure mobile flows, not always. Most real automation mixes browser-based tasks (email signup, admin panels, affiliate networks) with mobile-only tasks. Keeping profile isolation consistent between the two sides is easier when the same tool manages both fingerprint layers.

What's the minimum viable version of this stack?
10 Android cloud phones, one shared BitBrowser account, a single-process Node.js orchestrator, no Redis. That gets you 5–7 parallel jobs and costs under $500/month. Scale from there.

Will TikTok or Instagram detect Playwright itself?
Playwright leaves navigator.webdriver = true by default. Launch with --disable-blink-features=AutomationControlled and remove the flag with an init script. Also patch navigator.plugins.length and chrome.runtime to match a fresh Chrome install.

Android or iOS cloud phones — which one first?
Android for volume ops (TikTok, Instagram, Facebook, WhatsApp, Telegram). iOS only when the target app requires it (Snapchat US, iMessage-linked dating apps, BeReal, Cash App US). iOS is roughly 3× the per-device cost.


Disclosure: this post contains affiliate links. If you sign up through them I may receive a commission at no extra cost to you. I've used every tool referenced here on production workloads.

Top comments (0)