When a scraper or a test run hits Cloudflare, "the Cloudflare captcha" can mean four different things: a Managed Challenge page, a JS Challenge page, an Interactive Challenge page, or a Turnstile widget in the site's own form. The first three are full-page challenges, and passing one sets a cf_clearance cookie. The fourth sits inside a normal page, and passing it gives you a one-time token. One HTTP response is enough to tell which one you have, and that answer decides the whole fix.
The short version:
- The response has the header
cf-mitigated: challenge: it is a challenge page (managed, JS or interactive). - The response is a normal
200page with an element of classcf-turnstile: it is a Turnstile widget. - The response is a
403or429witherror code: 1020(or 1006, 1010, 1015 and so on) and nocf-mitigated: it is a firewall decision, and there is nothing to solve.
Below are the signals, a classifier in Python and Node.js, what it found on 19 Cloudflare-fronted sites, and the fix for each case. I also correct a check that an earlier version of this post got wrong.
Four things called "the Cloudflare captcha"
| Type | What a visitor sees | Where it appears | Passing it gives you |
|---|---|---|---|
| Managed Challenge | "Just a moment...", usually with no click; a checkbox only when Cloudflare is unsure | Full page, instead of the URL you asked for |
cf_clearance cookie |
| JS Challenge | "Just a moment..." while the browser runs JavaScript for a few seconds; never a click | Full page |
cf_clearance cookie |
| Interactive Challenge | A full page with a checkbox the visitor must click | Full page |
cf_clearance cookie |
| Turnstile widget | A small box inside a login, signup or checkout form | Inside a normal 200 page |
A token in the cf-turnstile-response field |
The first three are actions a site owner picks in a WAF or bot rule: managed_challenge, js_challenge and challenge in Cloudflare's rules API. Cloudflare recommends Managed Challenge over the other two, because it lets Cloudflare decide per request how much interaction to ask for.
Two naming collisions cause most of the confusion:
- The checkbox on a challenge page is drawn by Turnstile. It is the same widget technology sites embed in their own forms, so seeing a Turnstile box does not tell you which of the two you are in. In DevTools it is an iframe titled "Widget containing a Cloudflare security challenge" either way.
- Turnstile widgets have their own "managed" mode. The widget modes are managed, non-interactive and invisible. A "managed Turnstile widget" on a signup form is not a Managed Challenge.
This is why a screenshot is useless for diagnosis. The response is not.
How to tell them apart from the response
| Signal in the response | What it is | Fix |
|---|---|---|
Header cf-mitigated: challenge, usually status 403, <title>Just a moment...</title>, "Enable JavaScript and cookies to continue", body loads /cdn-cgi/challenge-platform/h/.../orchestrate/chl_page/v1
|
A challenge page (managed, JS or interactive) | Clearance flow: get cf_clearance and carry it |
Inside that page, window._cf_chl_opt = {... cType: 'managed' ...}
|
Which variant (managed, non-interactive for JS, interactive) |
Same fix; log it for diagnosis |
Status 200 with real content, an element with class="cf-turnstile" and a data-sitekey, or a turnstile.render(...) call |
A Turnstile widget | Token flow: get a token and submit it with the form |
403/429, server: cloudflare, body error code: 1020 (or 1003, 1006-1008, 1010, 1015), no cf-mitigated
|
A Cloudflare firewall rule, IP ban or rate limit | Nothing to solve: change egress IP, slow down, or stop matching the rule |
403, server: cloudflare, no cf-mitigated, but an X-DataDome header and a datadome cookie |
A different vendor's block served through Cloudflare's network | Not a Cloudflare challenge at all |
Three details matter here:
-
Build on the header. Cloudflare documents
cf-mitigated: challengeas present on every challenge page response, whatever the challenge type, and the content type is alwaystext/html, even when you asked for JSON.cTypeis an internal field in the page script. I log it, but I never branch on it, because Cloudflare can rename it at any time. -
Do not key on the string
challenge-platform. An earlier version of this post did, and it is wrong. Many ordinary pages load/cdn-cgi/challenge-platform/scripts/jsd/main.js, Cloudflare's JavaScript detections script, on a normal200response. In the spot check below, that string test flagged 14 of 19 responses as challenges. Only 7 were. -
Do not key on the status code. All 7 challenge pages in the spot check returned
403, and so did the DataDome block. Older write-ups mention503, which is what the old "I'm Under Attack" page used. The header is the stable signal.
A classifier you can run (Python)
# classify.py - pip install requests ; python classify.py URL [URL ...]
import re
import sys
import requests
UA = ("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36")
CF_ERRORS = {
"1003": "direct IP access not allowed",
"1006": "your IP is banned by the site",
"1007": "your IP is banned by the site",
"1008": "your IP is banned by the site",
"1010": "your browser signature is banned",
"1015": "you are being rate limited",
"1020": "a firewall rule blocked you",
}
def classify(resp):
"""Return (kind, detail) for a response that might be a Cloudflare block."""
body, h = resp.text, resp.headers # requests headers are case-insensitive
# 1. The documented signal: every Cloudflare challenge page sets this header.
if h.get("cf-mitigated", "").lower() == "challenge":
m = re.search(r"""cType:\s*['"]([\w-]+)['"]""", body) # internal field: log it only
return "challenge_page", m.group(1) if m else "unknown"
# 2. A Cloudflare 1xxx error: a firewall/IP/rate decision, nothing to solve.
if resp.status_code >= 400 and h.get("server", "").lower() == "cloudflare":
m = re.search(r'(?:error code:\s*|cf-error-code">|Error\s+)(1\d{3})\b', body)
if m:
return "cloudflare_error", f"{m.group(1)} ({CF_ERRORS.get(m.group(1), 'see Cloudflare docs')})"
if "x-datadome" in h:
return "other_vendor", "DataDome, served through Cloudflare's network"
# 3. A normal page with a Turnstile widget embedded in a form.
tag = re.search(r"""<[^>]+class=["'][^"']*\bcf-turnstile\b[^"']*["'][^>]*>""", body)
if tag:
key = re.search(r"""data-sitekey=["']([^"']+)["']""", tag.group(0))
return "turnstile_widget", key.group(1) if key else "sitekey is passed to turnstile.render()"
if "challenges.cloudflare.com/turnstile" in body:
return "turnstile_widget", "sitekey is passed to turnstile.render()"
return "no_cloudflare_block", str(resp.status_code)
if __name__ == "__main__":
for url in sys.argv[1:]:
try:
r = requests.get(url, headers={"User-Agent": UA}, timeout=20)
except requests.RequestException as exc:
print(f"{url}\n network error: {exc}")
continue
kind, detail = classify(r)
print(f"{url}\n {r.status_code} server={r.headers.get('server')} -> {kind}: {detail}")
Here is what it printed on 2026-09-29 for Cloudflare's community forum, a public bot-detection test page, a bare Cloudflare IP and Cloudflare's docs:
$ python classify.py https://community.cloudflare.com/ https://nowsecure.nl/ http://104.16.132.229/ https://developers.cloudflare.com/
https://community.cloudflare.com/
403 server=cloudflare -> challenge_page: managed
https://nowsecure.nl/
200 server=cloudflare -> turnstile_widget: 3x00000000000000000000FF
http://104.16.132.229/
403 server=cloudflare -> cloudflare_error: 1003 (direct IP access not allowed)
https://developers.cloudflare.com/
200 server=cloudflare -> no_cloudflare_block: 200
The test page's widget uses Cloudflare's test sitekey 3x00000000000000000000FF, which always forces the interactive checkbox. The forum's answer depends on your IP and client, which is the point: classify every response, not just the first one.
The same classifier in Node.js
Node 18 or later, no dependencies:
// classify.mjs - node classify.mjs URL [URL ...]
const UA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 " +
"(KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36";
export function classify(status, headers, body) {
if ((headers.get("cf-mitigated") ?? "").toLowerCase() === "challenge") {
const m = body.match(/cType:\s*['"]([\w-]+)['"]/);
return ["challenge_page", m ? m[1] : "unknown"];
}
if (status >= 400 && (headers.get("server") ?? "").toLowerCase() === "cloudflare") {
const m = body.match(/(?:error code:\s*|cf-error-code">|Error\s+)(1\d{3})\b/);
if (m) return ["cloudflare_error", m[1]];
if (headers.has("x-datadome")) return ["other_vendor", "DataDome"];
}
const tag = body.match(/<[^>]+class=["'][^"']*\bcf-turnstile\b[^"']*["'][^>]*>/);
if (tag) {
const key = tag[0].match(/data-sitekey=["']([^"']+)["']/);
return ["turnstile_widget", key ? key[1] : "sitekey is passed to turnstile.render()"];
}
if (body.includes("challenges.cloudflare.com/turnstile")) {
return ["turnstile_widget", "sitekey is passed to turnstile.render()"];
}
return ["no_cloudflare_block", String(status)];
}
for (const url of process.argv.slice(2)) {
try {
const res = await fetch(url, {
headers: { "User-Agent": UA },
signal: AbortSignal.timeout(20_000),
});
const [kind, detail] = classify(res.status, res.headers, await res.text());
console.log(`${url}\n ${res.status} server=${res.headers.get("server")} -> ${kind}: ${detail}`);
} catch (err) {
console.log(`${url}\n network error: ${err.message}`);
}
}
On the same four URLs it gave the same four answers under Node 22; it just leaves out the error-code descriptions.
What a spot check of 19 sites showed
On 2026-09-29 I ran the Python classifier from one machine against 25 well-known sites, with a Chrome User-Agent and no browser. Two did not answer from my network and four are not behind Cloudflare, which left 19 Cloudflare-fronted responses:
-
7 were challenge pages. All 7 returned
403, and all 7 hadcType: 'managed'. No JS or Interactive Challenge pages showed up, which fits Cloudflare recommending Managed Challenge to site owners. -
1 had a Turnstile widget on an otherwise normal
200page. -
1 was a DataDome block. It came back
403withserver: cloudflare, but the block was DataDome's, not Cloudflare's. A Cloudflare fix does nothing there. - 10 returned the normal page.
The old "challenge-platform" in r.text test called 14 of the 19 a challenge: the 7 real ones, 5 normal pages, the DataDome block and the Turnstile page. Half of its "challenges" would have sent you down the wrong fix.
How the type changes your fix
| Type | From a plain HTTP client | From a real browser (Playwright, Puppeteer) |
|---|---|---|
| JS Challenge | Cannot run it. You need a cf_clearance cookie. |
Often clears on its own if you wait for "Just a moment..." to go away; automated browsers can still fail it |
| Managed Challenge | Same: you need cf_clearance
|
Passes without a click when the signals look clean; otherwise shows a checkbox |
| Interactive Challenge | Same: you need cf_clearance
|
Always needs the checkbox click |
| Turnstile widget | Needs only a valid token at submit time | The widget fills cf-turnstile-response itself, sometimes after a click |
The two expensive mistakes are the mirror image of each other:
-
Fetching a Turnstile token for a challenge page. The checkbox on a challenge page is part of Cloudflare's own flow. A token you post somewhere does not set
cf_clearance, so the next request is challenged again. - Starting a full browser session for a form widget. The page already loaded. All you are missing is a token at submit time.
Both fixes below use one small helper for an in.php/res.php solver API (the 2Captcha-style protocol). It submits a task, waits, polls every 5 seconds, and raises on any error code instead of polling forever:
# solve.py - pip install requests
import time
import requests
API = "https://ocr.captchaai.com"
API_KEY = "YOUR_API_KEY"
class SolveError(RuntimeError):
pass
def solve(params, first_wait=15, poll_every=5, timeout=180):
"""Submit a task to an in.php/res.php solver API and wait for the answer."""
sub = requests.post(f"{API}/in.php", data={"key": API_KEY, "json": 1, **params},
timeout=30).json()
if sub.get("status") != 1:
raise SolveError(sub.get("request")) # e.g. ERROR_ZERO_BALANCE, ERROR_BAD_PROXY
task_id = sub["request"]
time.sleep(first_wait)
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
res = requests.get(f"{API}/res.php", params={
"key": API_KEY, "action": "get", "id": task_id, "json": 1}, timeout=30).json()
if res.get("status") == 1:
return res
if res.get("request") != "CAPCHA_NOT_READY":
raise SolveError(res.get("request")) # e.g. ERROR_CAPTCHA_UNSOLVABLE
time.sleep(poll_every)
raise SolveError(f"no answer after {timeout}s")
Fix 1 (Turnstile widget): the token flow
The widget's script adds a hidden cf-turnstile-response field to the form when it renders, which is why you often will not find that field in the raw HTML. The site checks the token server-side with Cloudflare's siteverify API. A token is single-use and valid for 300 seconds, so get it right before you submit, not at the start of a long job.
Read the sitekey (and data-action, if the widget has one) from the widget, ask for a token for that sitekey and the exact page URL, then post it with the rest of the form fields. This version runs against Cloudflare's public Turnstile demo, whose form posts to /handler:
# token_flow.py - a Turnstile widget inside a normal page
import re
import requests
from solve import solve
PAGE = "https://demo.turnstile.workers.dev/" # Cloudflare's public Turnstile demo
html = requests.get(PAGE, timeout=30).text
tag = re.search(r'<[^>]+class="[^"]*\bcf-turnstile\b[^"]*"[^>]*>', html)
key = tag and re.search(r'data-sitekey="([^"]+)"', tag.group(0))
if not key:
raise SystemExit("No data-sitekey in the HTML: look for turnstile.render() in the page scripts.")
action = re.search(r'data-action="([^"]+)"', tag.group(0))
params = {"method": "turnstile", "sitekey": key.group(1), "pageurl": PAGE}
if action:
params["action"] = action.group(1)
token = solve(params, first_wait=10)["request"]
print("token:", token)
# Submit straight away: a Turnstile token is single-use and expires after 300 s.
r = requests.post(PAGE + "handler", data={"cf-turnstile-response": token}, timeout=30)
print(r.status_code, r.text.splitlines()[0])
What it prints:
token: XXXX.DUMMY.TOKEN.XXXX
200 Turnstile token successfuly validated.
The demo uses Cloudflare's always-pass test sitekey, so the token is Cloudflare's dummy value, and its backend uses a test secret key, which accepts only that dummy token. (The typo is on Cloudflare's demo.) So this run proves the plumbing: sitekey extraction, the solve, the field name and the post. On a real site the token is a long opaque string, and production secret keys reject the dummy one. If you post a real token twice, or more than 300 seconds after you got it, the site's siteverify check fails with timeout-or-duplicate. If a fresh token still gets a 403 from your HTTP client, I went through the usual causes in Turnstile tokens that work in the browser but 403 from requests.
Fix 2 (challenge page): the clearance flow
Here a token does not help. What lets you in is the cf_clearance cookie that Cloudflare sets once the challenge passes. Three things follow from how that cookie works:
- It is checked against the IP address and User-Agent that passed the challenge. So the challenge has to be passed through the same proxy you will scrape from, and every later request must send that exact User-Agent.
- It expires. Its lifetime is the zone's Challenge Passage setting, which defaults to 30 minutes.
-
Your client's fingerprint still counts. A plain
requestsTLS handshake next to a Chrome User-Agent can raise the bot score enough to get challenged again.curl_cffiwithimpersonate="chrome"sends a browser-like handshake if that happens.
The cloudflare_challenge method takes the page URL plus a proxy, which is mandatory, and returns the cookie value and the User-Agent to use with it:
# clearance_flow.py - a challenge page (managed, JS or interactive)
from urllib.parse import urlsplit
import requests
from classify import classify
from solve import solve
PAGE = "https://target.example/products"
PROXY = "user:pass@203.0.113.10:8080" # the same exit IP you will scrape from
session = requests.Session()
session.proxies = {"http": f"http://{PROXY}", "https": f"http://{PROXY}"}
def mint_clearance():
ans = solve({"method": "cloudflare_challenge", "pageurl": PAGE,
"proxy": PROXY, "proxytype": "HTTP"}, first_wait=20)
session.headers["User-Agent"] = ans["user_agent"] # the cookie is tied to this UA
cookie = ans.get("result") or ans.get("request") # the docs return it in "result"
session.cookies.set("cf_clearance", cookie, domain=urlsplit(PAGE).hostname)
def fetch(url):
r = session.get(url, timeout=30)
if classify(r)[0] == "challenge_page": # no cookie yet, expired, or IP/UA changed
mint_clearance()
r = session.get(url, timeout=30)
return r
r = fetch(PAGE)
print(r.status_code, *classify(r))
When it works, the last line prints 200 no_cloudflare_block 200. The fetch() wrapper matters more than the first solve. When cf-mitigated: challenge comes back later, the cookie has expired or your IP or User-Agent has changed. Mint a new cookie; do not keep retrying the old one. If you are stuck in that loop, see why cf_clearance expires and how to stop the re-challenge loop.
FAQ
What is a Cloudflare challenge page?
It is the interstitial ("Just a moment...") that Cloudflare serves instead of the page you asked for, when a security rule, a bot setting or "I'm Under Attack" mode decides to challenge the request. It carries cf-mitigated: challenge, and passing it sets a cf_clearance cookie for that site.
What is the difference between a managed challenge and a JS challenge?
A JS Challenge always runs a non-interactive JavaScript check. A Managed Challenge lets Cloudflare choose for each request, from no interaction at all up to a checkbox. From an HTTP client they look the same: the same header, the same cf_clearance at the end. Only the cType value differs.
Is Cloudflare Turnstile the same as a Cloudflare challenge?
No. Turnstile is a widget that site owners put in their own forms, and it produces a token that their backend verifies. Challenge pages use Turnstile to draw their checkbox, but they are a separate, full-page layer that produces a cookie.
Why do I get a 403 from Cloudflare with no challenge at all?
Look for error code: 10xx in the body. 1020 is a firewall rule, 1006 to 1008 are IP bans, 1010 is a banned browser signature and 1015 is rate limiting (usually 429). None of these can be solved; change the IP, slow down, or stop matching the rule. If there is no error code and no cf-mitigated, check for another vendor's headers, such as X-DataDome.
TL;DR
-
cf-mitigated: challengemeans a challenge page (managed, JS or interactive). Use the clearance flow:cf_clearanceplus the same IP and User-Agent, and re-mint it when the challenge comes back. - A
200page withclass="cf-turnstile"means a widget. Use the token flow: sitekey plus page URL, then submit within 300 seconds, once. -
error code: 1020,1006-1008,1010or1015means a firewall decision. There is nothing to solve. - Never classify on the string
challenge-platformor on the status code alone.
The solver endpoint in these examples is CaptchaAI's. It speaks the 2Captcha-style in.php/res.php protocol and has a method for each flow: turnstile for the widget and cloudflare_challenge for challenge pages, whose parameters are in the Cloudflare Challenge guide. To run both scripts against your own target, one free thread for 30 days, no card needed is enough.
Top comments (0)