The captcha widget is not in your page (and that changes everything)
Here is a thing I got wrong for longer than I want to admit. I had a scraper that needed to tick a captcha checkbox, and I spent an evening trying to querySelector my way to it. The checkbox exists. I could see it on screen. I could even click it with a mouse. I just could not reach it from JavaScript, and the error was not helpful because there was no error.
The reason is boring once you know it: the checkbox is not in your document.
Nearly every captcha is a third-party iframe
When a site embeds reCAPTCHA, hCaptcha, Cloudflare Turnstile, or FunCaptcha, what actually lands in the page is usually a small <iframe> pointed at the vendor's domain. The widget renders inside that iframe, in a document that belongs to someone else.
That is not an implementation detail you can route around. It is the same-origin policy working exactly as designed, and it has consequences that only show up when you try to automate:
document.querySelector will never find the checkbox. It is not in your DOM tree. You can find the <iframe> element itself, and that is where your access stops. This is the single most common reason a script that "works in devtools" does nothing when you run it.
CSS cannot reach inside either. Styling that iframe's contents, or hiding the badge, needs the vendor to cooperate. You are looking at a rendered box, not a subtree.
Anything you compute from getBoundingClientRect() is a screen rectangle, not page structure. Which is still useful — it is how you find where to click — but it tells you nothing about what you are clicking.
There is a softer variant of the same thing: some sites keep the widget same-origin but render it into a shadow root. Same conclusion, different mechanism. If querySelector returns null and devtools shows the element, suspect a boundary of this kind before you suspect a race condition.
What this means for the two ways people automate
Once you accept the boundary, the design space narrows to two shapes:
Drive a real browser. You are not reaching into the iframe's DOM here either, but you do not have to: you click at coordinates. A CDP-level mouse event goes through iframes, shadow roots and cross-origin boundaries, because it is dispatched at the compositor rather than through the page's DOM. This is why coordinate clicking is the default approach for browser automation, and why it keeps working when every DOM-based selector fails.
Or hand the image to something that specialises in images. If you can obtain the pixels, you do not need the DOM at all. That is the whole idea behind the recognition half of a captcha solving API: you send a picture, you get back what to do with it — which tiles to click, which text to type, or where to drag.
Most of my own automation ends up combining the two: browser for the clicking, an API for the reading.
The part that trips people up: getting the pixels
If you go the image route, the hard part is not the model. It is extracting the picture you are actually looking at. I have hit three distinct versions of this and they all look like "the OCR is bad":
The iframe's default height. A cross-origin iframe with no explicit size renders at a default height (150px is the classic value) while the content inside keeps its own layout. Screenshot the iframe element and you get a correctly-sized image of a clipped puzzle. Mine came out 280x150 and I had assumed a 3x3 grid of ~100px tiles, because that is what a vendor demo page looks like.
The tainted canvas. Reading pixels back out of a canvas is only permitted if nothing drawn into it came from another origin. Draw a cross-origin image into a canvas without CORS approval and the canvas is tainted: getImageData(), toDataURL() and toBlob() all throw a SecurityError instead of returning pixels rather than silently returning nothing (the rule is documented on MDN's CORS-enabled image page). The nasty version of this failure is not the exception. It is the blank capture: right dimensions, plausible file size, no content, and nothing in the logs to say so. Compare a headed run against a headless run — mine only failed headless.
Someone else's challenge. This is the one that stung. A test page of mine still had a leftover anti-bot script, so the solver ran happily at 2am and solved a challenge that a third-party script had injected to verify me. The call succeeded, the answer was correct, and the result went nowhere, because it belonged to a different form. The fix is one line: log the frame URL alongside the challenge type. If the two disagree, you are solving a puzzle that was not yours.
A shape that survives all of this
What I do now, and what I would tell anyone starting:
-
Assume every challenge is cross-origin until you check. Find the iframe, log its
src, and confirm which vendor you are actually dealing with before writing any solving logic. - Compute click targets from rectangles, not from selectors. Take the iframe's bounding box, get your offset from the challenge response, and convert to real pixels. Coordinates that come back from recognition services are typically percentages (0-100) for exactly this reason: the service does not know your image's size, so it tells you the position in a form you can scale.
- Dump the first frame to disk and look at it with your own eyes. Not a benchmark — the actual image. This one habit caught two of the three bugs above.
- Distrust a successful call whose result you did not observe being consumed. A green checkmark means the API answered, not that the answer was used.
None of this is about model accuracy. The model was fine the entire time. I was feeding it the wrong picture, or the right picture for the wrong form, and both of those look identical from the outside.
Note on the extension route
There is a third option worth knowing about if you do not want to own this pipeline: ship it as a browser extension. The extension runs inside the browser, so it can take the screenshot, do the cropping and inject the result without you writing any of that glue, and the vendor's iframe boundary stops mattering because the extension is not a page script. I maintain a captcha solver extension that handles this for the common challenge types; if you would rather build your own, the four rules above are the parts that cost me the most time.
Top comments (0)