DEV Community

Cover image for What does it take to run a headless browser farm for ChatGPT at scale? The engineering build
Ricardo Batista
Ricardo Batista

Posted on Originally published at cloro.dev AI-assisted

What does it take to run a headless browser farm for ChatGPT at scale? The engineering build

A headless browser farm for ChatGPT at scale takes five components in every browser and three systems around the browsers. Each browser needs a stealth-patched Chromium, a sticky residential proxy session, a network interceptor on the conversation endpoint, a completion detector and a citation expander. Around them, the farm needs a pool of warm sessions, routing per region and monitoring per field.

The build is the smaller cost. One published estimate from cloro puts a self-run ChatGPT scraper at roughly $4,800 to $14,100 a month at 500,000 to 1 million requests, with engineer time as the largest line. The alternatives are rented browser infrastructure such as Bright Data or Browserbase, where the parser stays yours, and a managed endpoint such as cloro, which bills 5 credits for each ChatGPT web-search answer.

  • The response is a Server-Sent Events stream, so the scraper reads the network and not the DOM.
  • The failures that cost the most return a success status with part of the data missing.
  • Playwright wins when answer data is your product. A paid API wins when the data is an input.

What are the components of a headless browser farm for ChatGPT?

Each browser in the farm runs the same five parts. Each part answers one specific failure.

  1. Stealth-patched Playwright Chromium. A WebDriver-controlled browser exposes navigator.webdriver as true, and detection scripts read that flag on the first navigation. Apply playwright-stealth or equivalent canvas, WebGL and audioContext patches.
  2. A residential proxy with a sticky session. Datacenter IPs get a CAPTCHA on the first prompt. Cloudflare sets the cf_clearance cookie after a visitor solves a challenge and skips further challenges while the cookie is valid. The session must keep that cookie across the round trip.
  3. A response interceptor scoped to backend-api/f/conversation. ChatGPT makes other API calls during the page lifecycle. Only this URL carries the stream.
  4. A completion detector that looks for [DONE]. DOM checks for "is the response complete" return false positives during streaming. The [DONE] sentinel in the stream body is the reliable signal.
  5. A citation expander. Citations are not in the stream. The scraper clicks button.group/footnote after the response completes and parses the flyout at [data-testid="screen-threadFlyOut"].

The interceptor is a Playwright response handler:

async def setup_page_interceptor(self, page: Page):
    """Set up network request interception."""

    async def handle_response(response):
        # Capture conversation API responses
        if 'backend-api/f/conversation' in response.url:
            response_body = await response.text()
            self.captured_responses.append(response_body)

    page.on('response', handle_response)
Enter fullscreen mode Exit fullscreen mode

The captured body is raw SSE text. The parser splits it on data: lines, drops the sentinel and keeps the valid JSON objects:

import json
from typing import List, Dict, Any

def extract_raw_response(input_string: str) -> List[Dict[str, Any]]:
    """Parse ChatGPT's Server-Sent Events stream."""
    json_objects = []

    # Split by lines that start with "data: "
    lines = input_string.split("\n")

    for line in lines:
        # Skip empty lines and non-data lines
        if not line.strip() or not line.startswith("data: "):
            continue

        # Remove "data: " prefix
        json_str = line[6:].strip()

        # Skip special markers like [DONE]
        if json_str == "[DONE]":
            continue

        # Try to parse as JSON
        try:
            json_obj = json.loads(json_str)

            # Only include if it's a dictionary (object), not string or other types
            if isinstance(json_obj, dict):
                json_objects.append(json_obj)
        except json.JSONDecodeError:
            # Skip invalid JSON
            continue

    return json_objects
Enter fullscreen mode Exit fullscreen mode

The full scraper class that these two excerpts come from launches Chromium with headless=False. That choice is deliberate. Headless defaults are the first signal that bot detection reads, so a production farm runs full browser environments with consistent canvas, font and TLS fingerprints.

What actually breaks when you build an LLM scraper in-house?

Five failure modes hit a ChatGPT scraper in production. We run about 1,000 engine answers a day across six engines for cloro's own tracking corpus, and these are the failures that operation meets.

Failure What you see Why a retry does not fix it
Response format shift The answer arrives as HTML fragments. The model name and search queries are absent, and the source list is shorter. OpenAI serves a share of logged-out traffic in a mobile-web format and selects the format per request.
Model rotation Fields that depend on the model, such as fan-out queries under gpt-5-3, appear on some answers only. ChatGPT selects the serving model per request. No parameter controls it.
Regional login wall Anonymous sessions fail for one geography. The provider changed its access rules for that region.
Anti-bot escalation Rate limits, session churn and fingerprint checks start at volume. A browser that works at 10 answers a day meets them at 10,000.
Silent quality drift The scrape succeeds and the source list is missing half its entries. A DOM class changed. No request failed.

The first and last rows cost the most. A parser built against one format reports the other as a partial success, and a partial success sends no alert.

The targets also move. The backend-api/f/conversation URL changed twice in 2025. The shopping-card event moved from a flat products array to a nested offers block. OpenAI's build pipeline generates CSS class names that change between deploys, so a scraper that relies on class selectors breaks roughly weekly. ARIA labels, role attributes and text content last longer than class names, and both still break.

How do you keep the farm from getting blocked at scale?

The farm must look and behave like the logged-out users that ChatGPT already serves. Do these four things.

  1. Keep a pool of established sessions. Engines challenge new sessions harder than warm ones, and one cleared session serves many requests before rotation.
  2. Match the exit region to the request. An answer requested for Germany must exit in Germany. Traffic concentrated in one region meets volume ceilings.
  3. Space the retries and add jitter. Immediate identical retries are a bot signature.
  4. Treat a degraded format as a block. It costs the engine less to serve a reduced response to suspected automation than to refuse it.

The IP type sets the ceiling. In testing published by proxies.sx, real mobile IPs lasted 50 to 100 or more queries, and datacenter IPs were blocked within the first few requests. That figure is the vendor's and was not reproduced for this article.

How do you know the farm still returns good data?

Status codes do not show the two failures that cost the most, so monitor the fields. Track these rates per day:

  • answers with an empty model name
  • answers with a short source list
  • answers with no fan-out queries where you expect them

A change in any of the three is a format change or a model rotation in your sample.

One invariant helps with ChatGPT. We measured grounding on roughly 2,500 prompts per engine, and ChatGPT grounded 98.4% of its answers in live sources at 14.1 sources each. A re-measurement on 2026-08-19 gave 78.9% and 12.7 sources. An empty citation list from a ChatGPT scrape is more often a parse failure than an ungrounded answer, so assert a source list that is not empty in your tests.

Playwright or a paid API for scraping AI chat answers?

Playwright is the correct choice when answer data is your product, your volume pays for a dedicated team and you need control that a vendor cannot give. It is free and open source, every layer is yours to tune and there is no vendor lock-in. A team with no budget and strong engineers can also run it at low volume. The roundup linked below budgets 8 to 15 engineer hours a month for a Playwright stack at 1,000 queries a day.

A paid API is the correct choice when answer data is an input to another product. Under roughly 1,000 answers a day, a provider costs less as soon as you price your own time. Two kinds of paid API exist, and they are not the same purchase:

  • Browser and proxy infrastructure (Bright Data, Oxylabs, Browserbase, Browserless). The access layer is solved. The SSE parser, the citation expander and the selectors stay with you. Bright Data prices its hosted Browser API from $5 per GB of traffic.
  • A managed answer endpoint (cloro is one). The provider owns all five components and returns parsed JSON.

The managed call looks like this:

import requests
import json

# Your prompt
prompt = "Compare the top 3 programming languages for web development in 2026"

# API request to cloro
response = requests.post(
    'https://api.cloro.dev/v1/monitor/chatgpt',
    headers={'Authorization': 'Bearer YOUR_API_KEY'},
    json={
        'prompt': prompt,
        'country': 'US',
        'include': {
            'markdown': True,
            'rawResponse': True,
            'searchQueries': True
        }
    }
)

result = response.json()
print(json.dumps(result, indent=2))
Enter fullscreen mode Exit fullscreen mode

A ChatGPT web-search answer costs 5 credits there as an async request. A synchronous request, as in the sample above, adds 2 credits. 5,000 prompts a day is about 750,000 credits a month, which fits in the Growth plan at $500 a month for 1,350,000 credits. That is $1.85 per 1,000 answers. Failed requests are not billed. The trade is control: you cannot tune a session pool that you do not run.

Neither option fits if you want the model's output and not the answer that ChatGPT shows its users. That is the OpenAI API's job, priced in tokens.

How these figures were collected

The failure modes come from cloro's own tracking operation, about 1,000 engine answers a day across six engines. The components and the code come from a working Playwright scraper for the logged-out ChatGPT surface. The $4,800 to $14,100 figure is an estimate and not a measured bill. It adds list prices ($999 for 332 GB of residential proxy bandwidth on Bright Data, $496 a month for each c6i.4xlarge browser instance) to a fifth to a half of one engineer at the US median software developer salary of $131,450, loaded.

This article is published by cloro, which is one of the managed options in it. Figures for other vendors are those vendors' published figures. The cloro prices are self-reported and were checked against its price list on 2026-10-07.

The source pages

Top comments (0)