DEV Community

Cover image for 2Captcha-compatible API: migrate off 2Captcha by changing one base URL (Python diff)
Bassem Shahin
Bassem Shahin

Posted on Edited on Originally published at dev.to

2Captcha-compatible API: migrate off 2Captcha by changing one base URL (Python diff)

If your scraper or test suite already calls 2Captcha's in.php / res.php API, trying another provider does not need a new integration. Several services accept the same requests and return the same responses (a "2Captcha-compatible API"), so the migration is two values: the base URL and the API key. method=userrecaptcha, googlekey, pageurl, the OK|<id> submit response, CAPCHA_NOT_READY polling and the token you put into g-recaptcha-response all stay the same.

This post covers the exact diff, a small client that makes the provider a config value (Python and Node.js), which reCAPTCHA parameters carry over, the five things that do not, how to point the official 2captcha-python SDK at another host, and an A/B harness so you decide on your own traffic.

Scope: this is about the classic in.php / res.php interface. If your code uses 2Captcha's newer JSON createTask / getTaskResult endpoints, check that the new provider accepts that format before you plan a one-line change.

The diff

 import os

-SOLVER_BASE = "https://2captcha.com"
-SOLVER_KEY = os.environ["TWOCAPTCHA_KEY"]
+SOLVER_BASE = "https://ocr.captchaai.com"
+SOLVER_KEY = os.environ["SOLVER_API_KEY"]
Enter fullscreen mode Exit fullscreen mode

Both hosts serve /in.php and /res.php, so every request built from SOLVER_BASE keeps working. If the base URL is hard-coded in several files, the first real step is to make it one setting. The client below does that, so switching (or switching back) becomes an environment variable instead of a deploy.

A client where the provider is config (Python)

Prerequisites: Python 3.9+, pip install requests, and three environment variables: SOLVER_API_KEY, SITE_KEY (the data-sitekey value, or the k= parameter of the reCAPTCHA anchor URL) and PAGE_URL (the full URL of the page that shows the widget).

# solver_client.py  (Python 3.9+, pip install requests)
import os
import time

import requests

# Worth retrying: the server is busy or briefly failing.
RETRYABLE = {
    "ERROR_SERVER_ERROR", "ERROR_INTERNAL_SERVER_ERROR",
    "ERROR_ZERO_BALANCE",       # on thread-based billing this can mean "all threads busy"
    "ERROR_NO_SLOT_AVAILABLE",  # 2Captcha's queue-full code
}


class SolverError(Exception):
    pass


class CompatSolver:
    """Talks the 2Captcha in.php / res.php protocol to any compatible host."""

    def __init__(self, base, api_key, initial_wait=15, poll=5, timeout=120):
        if not api_key or api_key == "YOUR_API_KEY":
            raise ValueError("Set SOLVER_API_KEY to a real key")
        self.base = base.rstrip("/")
        self.key = api_key
        self.initial_wait, self.poll, self.timeout = initial_wait, poll, timeout
        self.http = requests.Session()

    def _call(self, endpoint, params, method="GET"):
        params = {**params, "key": self.key, "json": 1}
        url = f"{self.base}/{endpoint}"
        delay, code = 5, None
        for attempt in range(5):
            if attempt:
                time.sleep(delay)
                delay = min(delay * 2, 30)  # exponential backoff, capped at 30 s
            try:
                if method == "POST":
                    r = self.http.post(url, data=params, timeout=30)
                else:
                    r = self.http.get(url, params=params, timeout=30)
                r.raise_for_status()
                data = r.json()
            except (requests.RequestException, ValueError) as exc:
                code = f"NETWORK: {exc.__class__.__name__}"
                continue
            code = data.get("request")
            if data.get("status") == 1 or code == "CAPCHA_NOT_READY":
                return data
            if code not in RETRYABLE:
                # bad key, bad params, bad proxy, ERROR_CAPTCHA_UNSOLVABLE:
                # sending the same request again will not fix these
                raise SolverError(code)
        raise SolverError(f"gave up after 5 attempts: {code}")

    def solve(self, **task):
        task_id = self._call("in.php", task, method="POST")["request"]
        time.sleep(self.initial_wait)
        deadline = time.monotonic() + self.timeout
        attempt = 0
        while time.monotonic() < deadline:
            attempt += 1
            data = self._call("res.php", {"action": "get", "id": task_id})
            if data.get("status") == 1:
                data.setdefault("request", data.get("result"))  # Enterprise tokens can come back as "result"
                return data  # data["request"] is the token
            print(f"poll {attempt}: not ready")
            time.sleep(self.poll)
        raise SolverError(f"no result after {self.timeout}s (task {task_id})")


if __name__ == "__main__":
    solver = CompatSolver(
        base=os.environ.get("SOLVER_BASE", "https://ocr.captchaai.com"),
        api_key=os.environ.get("SOLVER_API_KEY", "YOUR_API_KEY"),
    )
    result = solver.solve(
        method="userrecaptcha",
        googlekey=os.environ["SITE_KEY"],
        pageurl=os.environ["PAGE_URL"],
        # invisible=1,                   # v2 invisible
        # version="v3", action="login",  # v3: action comes from grecaptcha.execute
    )
    print("token:", result["request"][:40] + "...")
Enter fullscreen mode Exit fullscreen mode

Expected output:

poll 1: not ready
poll 2: not ready
token: 03AHJ_Vuve5Asa4koK3KSMyUkCq0vUFCR5Im4CwB...
Enter fullscreen mode Exit fullscreen mode

What it handles:

  • json=1 on every call, so every answer is {"status": ..., "request": ...}. Both providers use that shape, with one exception: on the endpoint used here, a solved reCAPTCHA Enterprise task returns the token as result, next to a user_agent field, so solve() copies result into request.
  • A 15-second wait after submitting, then a poll every 5 seconds, with a 120-second cap.
  • Network errors, HTTP 5xx and ERROR_SERVER_ERROR / ERROR_INTERNAL_SERVER_ERROR: retried with exponential backoff capped at 30 seconds.
  • A bad key, bad parameters, a bad proxy or ERROR_CAPTCHA_UNSOLVABLE: raised at once, because repeating the same request will not change the answer. If you want, resubmit an unsolvable task once from the caller.

The same client in Node.js

Node 18 or later has fetch built in, so there are no dependencies. Save it as solver-client.mjs and run node solver-client.mjs with the same environment variables.

// solver-client.mjs  (Node 18+, no dependencies)
const RETRYABLE = new Set([
  "ERROR_SERVER_ERROR", "ERROR_INTERNAL_SERVER_ERROR",
  "ERROR_ZERO_BALANCE", "ERROR_NO_SLOT_AVAILABLE",
]);
const sleep = (s) => new Promise((resolve) => setTimeout(resolve, s * 1000));

class SolverError extends Error {}

class CompatSolver {
  constructor(base, apiKey, { initialWait = 15, poll = 5, timeout = 120 } = {}) {
    if (!apiKey || apiKey === "YOUR_API_KEY") throw new Error("Set SOLVER_API_KEY to a real key");
    Object.assign(this, { base: base.replace(/\/+$/, ""), apiKey, initialWait, poll, timeout });
  }

  async call(endpoint, params, method = "GET") {
    const query = new URLSearchParams({ ...params, key: this.apiKey, json: "1" });
    const url = `${this.base}/${endpoint}` + (method === "GET" ? `?${query}` : "");
    let delay = 5, code;
    for (let attempt = 0; attempt < 5; attempt++) {
      if (attempt) { await sleep(delay); delay = Math.min(delay * 2, 30); }
      let data;
      try {
        const res = await fetch(url, {
          method,
          body: method === "POST" ? query : undefined,
          signal: AbortSignal.timeout(30_000),
        });
        if (!res.ok) throw new Error(`HTTP ${res.status}`);
        data = await res.json();
      } catch (err) {
        code = `NETWORK: ${err.message}`;
        continue;
      }
      code = data.request;
      if (data.status === 1 || code === "CAPCHA_NOT_READY") return data;
      if (!RETRYABLE.has(code)) throw new SolverError(code);
    }
    throw new SolverError(`gave up after 5 attempts: ${code}`);
  }

  async solve(task) {
    const taskId = (await this.call("in.php", task, "POST")).request;
    await sleep(this.initialWait);
    const deadline = Date.now() + this.timeout * 1000;
    for (let attempt = 1; Date.now() < deadline; attempt++) {
      const data = await this.call("res.php", { action: "get", id: taskId });
      if (data.status === 1) return { request: data.result, ...data }; // Enterprise tokens can come back as "result"
      console.log(`poll ${attempt}: not ready`);
      await sleep(this.poll);
    }
    throw new SolverError(`no result after ${this.timeout}s (task ${taskId})`);
  }
}

const { SOLVER_BASE, SOLVER_API_KEY, SITE_KEY, PAGE_URL } = process.env;
if (!SITE_KEY || !PAGE_URL) throw new Error("Set SITE_KEY and PAGE_URL");

const solver = new CompatSolver(SOLVER_BASE ?? "https://ocr.captchaai.com", SOLVER_API_KEY);
const result = await solver.solve({ method: "userrecaptcha", googlekey: SITE_KEY, pageurl: PAGE_URL });
console.log(`token: ${result.request.slice(0, 40)}...`);
Enter fullscreen mode Exit fullscreen mode

The output is the same three lines as the Python version.

Which reCAPTCHA parameters carry over

Every variant uses method=userrecaptcha, googlekey and pageurl. The extra parameters are the same names on both sides:

Variant Extra in.php parameters What to do with the token
v2 checkbox none Put it in g-recaptcha-response, then submit the form
v2 invisible invisible=1 (required; leaving it out can make solves fail) Same field, then trigger the protected action right away
v2 with a callback none Set the field, then call the page's callback function with the token
v3 version=v3 and action (from the page's grecaptcha.execute call; verify if the page sets none) Send it in the backend request with the same action
v2 Enterprise enterprise=1, plus action if the anchor URL has sa= Poll with json=1; the endpoint used here returns the token as result plus a user_agent, and you submit using that user agent
v3 Enterprise version=v3, enterprise=1, action Same user-agent rule; a mismatched user agent fails validation

Not sure which one a site runs? I wrote up how to tell reCAPTCHA v2, v3 and Enterprise apart from the page source: api.js vs enterprise.js, render=, and the execute call.

Five things that do not carry over

  1. min_score on v3. 2Captcha lets you ask for a minimum score. On the endpoint used here it is not a parameter, so remove it. Judge v3 by whether the target accepts the token (the A/B harness below records that).
  2. One error name. 2Captcha's docs call a malformed sitekey ERROR_WRONG_GOOGLEKEY. The error docs for the endpoint used here list it as ERROR_WRONG_SITEKEY. If your code compares error strings, match both, and treat any string you don't recognise as a fatal input error, as the client above does: when I tested, a sitekey of the wrong length came back as the plain message Sitekey length is invalid.
  3. What ERROR_ZERO_BALANCE means. With per-solve billing it means the wallet is empty, so you stop. With thread-based billing, the error docs for the endpoint used here describe it as "insufficient balance/threads" and suggest reducing concurrency and retrying. So it can simply mean you are at your concurrency limit. That is why the client above backs off and retries on it instead of halting.
  4. Proxies. On the endpoint used here, proxy support is off until support enables it for your account, and reCAPTCHA v3 tasks do not take a proxy at all. Strip proxy and proxytype from v3 tasks before you switch.
  5. 2Captcha-only extras. soft_id is 2Captcha's developer-attribution ID and means nothing to another provider. If you rely on pingback callbacks, confirm the new provider supports them, or keep polling as the client above does.

Pointing the 2captcha-python SDK at another host

If you use the official 2captcha-python package, you don't need to replace it. I checked the source (twocaptcha/solver.py and twocaptcha/api.py): the constructor takes a server argument, default 2captcha.com, and builds https://{server}/in.php and https://{server}/res.php from it.

import os

from twocaptcha import TwoCaptcha  # pip install 2captcha-python

solver = TwoCaptcha(
    os.environ["SOLVER_API_KEY"],
    server="ocr.captchaai.com",  # hostname only: the SDK adds https:// and /in.php
    softId=None,                 # the default, 4580, is 2Captcha's software-catalog ID
    pollingInterval=5,
)

result = solver.recaptcha(sitekey=os.environ["SITE_KEY"], url=os.environ["PAGE_URL"])
print(result["captchaId"], result["code"][:40])

# v3: pass version and the page's action. Do not pass score=, the SDK sends it as min_score.
# solver.recaptcha(sitekey=..., url=..., version="v3", action="login")
Enter fullscreen mode Exit fullscreen mode

Two details from the source: the SDK always sends version (default v2) and enterprise (default 0), and it raises ApiException with the raw error string, so point 2 above applies to SDK users too. Run one real solve for each variant you use before sending production traffic.

A/B the switch on your own traffic

Don't move everything at once. Keep the old provider, send a fixed share of jobs to the new one, and record what matters: did you get a token, did the target accept it, and how long did it take. A token the target rejects costs you the same as no token, so "accepted" is the number to decide on. The harness reads TWOCAPTCHA_KEY, SOLVER_BASE and SOLVER_API_KEY from the environment and reuses the client above.

# ab_router.py
import hashlib
import os
import statistics
import time

from solver_client import CompatSolver, SolverError

providers = {
    "incumbent": CompatSolver("https://2captcha.com", os.environ["TWOCAPTCHA_KEY"]),
    "candidate": CompatSolver(os.environ["SOLVER_BASE"], os.environ["SOLVER_API_KEY"]),
}
CANDIDATE_PERCENT = 10
stats = {name: {"solved": 0, "failed": 0, "accepted": 0, "secs": []} for name in providers}


def pick(job_id):
    # Same job -> same provider every time, so a retry never mixes providers.
    bucket = int(hashlib.sha256(job_id.encode()).hexdigest(), 16) % 100
    return "candidate" if bucket < CANDIDATE_PERCENT else "incumbent"


def solve_and_submit(job_id, task, submit):
    """submit(token) is your code: post the form, return True if the target accepted it."""
    name = pick(job_id)
    s = stats[name]
    started = time.monotonic()
    try:
        token = providers[name].solve(**task)["request"]
    except SolverError:
        s["failed"] += 1
        raise
    s["solved"] += 1
    s["secs"].append(time.monotonic() - started)
    if submit(token):
        s["accepted"] += 1


def report():
    for name, s in stats.items():
        n = s["solved"] + s["failed"]
        if not n:
            continue
        secs = sorted(s["secs"]) or [float("nan")]  # nan: no successful solve yet
        p95 = secs[int(0.95 * (len(secs) - 1))]
        print(f"{name:9} n={n:5}  token={s['solved'] / n:6.1%}  accepted={s['accepted'] / n:6.1%}"
              f"  median={statistics.median(secs):5.1f}s  p95={p95:5.1f}s")
Enter fullscreen mode Exit fullscreen mode

report() prints one line per provider. The format looks like this (placeholder numbers, not a benchmark):

incumbent n= 1804  token= 96.8%  accepted= 91.2%  median= 21.4s  p95= 47.9s
candidate n=  196  token= 97.4%  accepted= 90.8%  median= 19.6s  p95= 38.2s
Enter fullscreen mode Exit fullscreen mode

Run it for at least a day so peak hours are included, and compare on the sites you actually target. If the candidate holds up, raise CANDIDATE_PERCENT. If it doesn't, 90% of your traffic never left the old provider.

Sizing: per-solve vs per-thread billing

The main reason teams move is cost at volume. Per-solve billing charges for every token. Thread-based billing charges for how many solves can run at the same time, with no limit on the number of solves. So the number you plan around changes from "solves per month" to "solves in progress at peak". Little's law gives it directly:

threads needed ≈ peak solves per second × average seconds per solve
Enter fullscreen mode Exit fullscreen mode

Use your own measured solve time, not a vendor's figure: statistics.mean(stats["candidate"]["secs"]) from the harness gives it. At 2 solves per second and 20 seconds per solve you have about 40 solves in progress; add 20 to 30% headroom for bursts and slow tails and plan for about 50. I worked through the per-1,000 vs thread-based numbers from 10k to 10M solves a month in The real cost of solving reCAPTCHA at scale.

FAQ

Is a 2Captcha-compatible API really a drop-in replacement?

At the protocol level, yes: same endpoints, same parameter names for reCAPTCHA, same OK|<id> and CAPCHA_NOT_READY responses. Check the five points above before you cut over, especially min_score and exact error-string matching.

Do I have to change my token injection code?

No. The token goes into the same place as before: g-recaptcha-response for v2, your backend request for v3, and the page's callback if it uses one. The one addition is for Enterprise: send the request with the user_agent returned alongside the token.

Can I keep 2Captcha as a fallback?

Yes. Because the client takes the base URL and key as arguments, a fallback is a second CompatSolver instance. Catch SolverError from the first and retry the job on the second.

Does this work for Cloudflare Turnstile too?

The same client works; only the task changes (method=turnstile, sitekey, pageurl), and the token goes into cf-turnstile-response. Check the new provider's parameter list for any type you use besides reCAPTCHA.

Trying it

A compatible provider is a drop-in at the API level. Whether it is the right choice depends on acceptance rate, latency and cost on your own traffic, and compatibility is what makes that test cheap.

Disclosure: I work on CaptchaAI, the service behind ocr.captchaai.com in the examples above. If you want to run the A/B without paying for anything, the free thread gives you one thread for 30 days with no card, which is enough for a small test slice. The parameter reference for each reCAPTCHA variant is in the reCAPTCHA v2 guide and the guides linked from it.

Top comments (0)