<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Tom Hargreaves</title>
    <description>The latest articles on DEV Community by Tom Hargreaves (@tomhargreaves).</description>
    <link>https://dev.to/tomhargreaves</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4138641%2F8eecf814-9425-4ff2-8150-59f2cf223b0a.png</url>
      <title>DEV Community: Tom Hargreaves</title>
      <link>https://dev.to/tomhargreaves</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tomhargreaves"/>
    <language>en</language>
    <item>
      <title>hcaptcha bypass in a real pipeline: what actually decides whether it works</title>
      <dc:creator>Tom Hargreaves</dc:creator>
      <pubDate>Thu, 01 Oct 2026 08:58:48 +0000</pubDate>
      <link>https://dev.to/tomhargreaves/hcaptcha-bypass-in-a-real-pipeline-what-actually-decides-whether-it-works-24nm</link>
      <guid>https://dev.to/tomhargreaves/hcaptcha-bypass-in-a-real-pipeline-what-actually-decides-whether-it-works-24nm</guid>
      <description>&lt;p&gt;People search for &lt;strong&gt;hcaptcha bypass&lt;/strong&gt; expecting a single technique. What they actually need depends on something they usually have not checked yet: &lt;strong&gt;where the challenge lives and what shape the answer takes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I have automated hCaptcha on and off for a couple of years, in scrapers and in browser flows. Every time it went badly, the model was not the problem. The setup was. Here is the version of the advice I wish I had read first.&lt;/p&gt;

&lt;h2&gt;
  
  
  First: there are three hCaptcha shapes, not one
&lt;/h2&gt;

&lt;p&gt;hCaptcha does not present one kind of task. Depending on what the site owner configured, you may get:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A checkbox.&lt;/strong&gt; No image at all in the normal case. You click, and it either passes or escalates. There is nothing to read and nothing to solve locally — the decision was made by a score, server-side.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A tile grid.&lt;/strong&gt; The familiar 3x3 (sometimes 4x4) "please click each image containing a boat". This is the only shape where image recognition is the point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A drag task.&lt;/strong&gt; A generated object that has to be dragged somewhere. The answer here is &lt;strong&gt;a coordinate pair, not a text answer&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each of those needs a different response, and the most expensive mistake in this space is treating a checkbox like a puzzle. If a site is showing you a checkbox, there is no picture to send anywhere. There is only a session that either looks plausible or does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule that confuses everyone: billing and "not ready"
&lt;/h2&gt;

&lt;p&gt;Because hCaptcha work is asynchronous — you submit, then you poll — you have to understand the error vocabulary before you write the loop. On the service I use, three facts do the heavy lifting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Not finished yet" has its own code.&lt;/strong&gt; It is a normal step in the happy path. If your client treats every error as a failure, your first naive version will resubmit a job that is already running and already paid for. One image becomes several jobs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You are billed when the upstream accepts the job, not when it succeeds.&lt;/strong&gt; A rejected submit — bad parameters, no credit, no upstream capacity — costs nothing. Once accepted, the credits are spent regardless of the outcome, including a timeout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;There are no refunds, including on timeouts.&lt;/strong&gt; So your true cost is per &lt;em&gt;attempt&lt;/em&gt;, not per &lt;em&gt;successful solve&lt;/em&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical consequence is that "retry on error" is a dangerous default here. I had to split error handling into three branches — wait, slow down, and genuine failure — before my numbers stopped drifting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing between recognition and a token
&lt;/h2&gt;

&lt;p&gt;If you need an hCaptcha answer, there are two entry points and they solve different problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Image recognition&lt;/strong&gt; — you already have the picture; it tells you which tiles to click or where to drag. You do the clicking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token generation&lt;/strong&gt; — you do not have the picture, or do not care; the whole challenge is solved server-side and you get back something submittable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They are priced differently, and not by a rounding error: on the service I use an hCaptcha &lt;strong&gt;token costs 10 credits where a recognition call costs 1&lt;/strong&gt;. If your browser flow still has to tick the checkbox itself, you want recognition — it returns an acknowledgement rather than a token, and costs correspondingly less.&lt;/p&gt;

&lt;h2&gt;
  
  
  The details that only show up in production
&lt;/h2&gt;

&lt;p&gt;These each cost me more than an hour:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Drag answers come back as percentages.&lt;/strong&gt; Top-left is &lt;code&gt;0,0&lt;/code&gt;, bottom-right is &lt;code&gt;100,100&lt;/code&gt;. Multiply by the real image width and height. I spent a while feeding percentages into a mouse move as if they were pixels and wondering why every drag landed in the corner.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token challenges are slow.&lt;/strong&gt; A measured hCaptcha token takes about a minute, and longer at peak. That is why token jobs get a longer expiry than recognition jobs — five minutes against three. Poll every 2-3 seconds, not faster: every accepted job is billed, so hammering the endpoint just pays twice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Concurrency, not model quality, is your throughput ceiling.&lt;/strong&gt; If the limit is 4 parallel jobs and a job takes about a minute, your realistic ceiling is about 4 per minute no matter how good the model is. Compute that number before you design the queue.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A checkbox result is not a solved captcha.&lt;/strong&gt; If you get back an acknowledgement, you still have to place it where the page expects and let the site's own callback fire. Injecting a value into the wrong field looks identical to a failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  When it fails, check these before blaming the solver
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Is the task even enabled on your plan?&lt;/strong&gt; Some types sit behind higher tiers, and you only find out after paying.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did you send the challenge text and the page context?&lt;/strong&gt; For tile selection, sending the question text improves accuracy noticeably, and for hCaptcha specifically the &lt;code&gt;data&lt;/code&gt; object scraped from the page should be passed through unchanged — it carries the link between the question and the images.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is your user agent consistent with the page you are solving for?&lt;/strong&gt; For token types this is the single biggest lever, and it is free.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Are you measuring per attempt or per success?&lt;/strong&gt; If failed-but-accepted jobs are billed, your real unit cost is the former, and retry loops hide the difference.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;hCaptcha is three problems wearing one name. Figure out which shape you actually have, read the billing rule before writing any retry logic, and remember that the number that decides your throughput is the concurrency ceiling, not how clever the model is.&lt;/p&gt;

&lt;p&gt;I maintain the service this is measured against — the full error table, the endpoint list, the twelve supported types and the exact billing rule are on the front page at &lt;a href="https://fuckcaptcha.top/" rel="noopener noreferrer"&gt;hcaptcha bypass&lt;/a&gt;. The figures above are the ones I operate, not ones from a marketing page.&lt;/p&gt;

</description>
      <category>automation</category>
      <category>security</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>The captcha widget is not in your page (and that changes everything)</title>
      <dc:creator>Tom Hargreaves</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:44:10 +0000</pubDate>
      <link>https://dev.to/tomhargreaves/the-captcha-widget-is-not-in-your-page-and-that-changes-everything-b9b</link>
      <guid>https://dev.to/tomhargreaves/the-captcha-widget-is-not-in-your-page-and-that-changes-everything-b9b</guid>
      <description>&lt;h1&gt;
  
  
  The captcha widget is not in your page (and that changes everything)
&lt;/h1&gt;

&lt;p&gt;Here is a thing I got wrong for longer than I want to admit. I had a scraper that needed to tick a captcha checkbox, and I spent an evening trying to &lt;code&gt;querySelector&lt;/code&gt; my way to it. The checkbox exists. I could see it on screen. I could even click it with a mouse. I just could not reach it from JavaScript, and the error was not helpful because there was no error.&lt;/p&gt;

&lt;p&gt;The reason is boring once you know it: &lt;strong&gt;the checkbox is not in your document.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Nearly every captcha is a third-party iframe
&lt;/h2&gt;

&lt;p&gt;When a site embeds reCAPTCHA, hCaptcha, Cloudflare Turnstile, or FunCaptcha, what actually lands in the page is usually a small &lt;code&gt;&amp;lt;iframe&amp;gt;&lt;/code&gt; pointed at the vendor's domain. The widget renders inside that iframe, in a document that belongs to someone else.&lt;/p&gt;

&lt;p&gt;That is not an implementation detail you can route around. It is the same-origin policy working exactly as designed, and it has consequences that only show up when you try to automate:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;document.querySelector&lt;/code&gt; will never find the checkbox.&lt;/strong&gt; It is not in your DOM tree. You can find the &lt;code&gt;&amp;lt;iframe&amp;gt;&lt;/code&gt; element itself, and that is where your access stops. This is the single most common reason a script that "works in devtools" does nothing when you run it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CSS cannot reach inside either.&lt;/strong&gt; Styling that iframe's contents, or hiding the badge, needs the vendor to cooperate. You are looking at a rendered box, not a subtree.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Anything you compute from &lt;code&gt;getBoundingClientRect()&lt;/code&gt; is a screen rectangle&lt;/strong&gt;, not page structure. Which is still useful — it is how you find where to click — but it tells you nothing about what you are clicking.&lt;/p&gt;

&lt;p&gt;There is a softer variant of the same thing: some sites keep the widget same-origin but render it into a shadow root. Same conclusion, different mechanism. If &lt;code&gt;querySelector&lt;/code&gt; returns null and devtools shows the element, suspect a boundary of this kind before you suspect a race condition.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the two ways people automate
&lt;/h2&gt;

&lt;p&gt;Once you accept the boundary, the design space narrows to two shapes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Drive a real browser.&lt;/strong&gt; You are not reaching into the iframe's DOM here either, but you do not have to: you click at coordinates. A CDP-level mouse event goes through iframes, shadow roots and cross-origin boundaries, because it is dispatched at the compositor rather than through the page's DOM. This is why coordinate clicking is the default approach for browser automation, and why it keeps working when every DOM-based selector fails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Or hand the image to something that specialises in images.&lt;/strong&gt; If you can obtain the pixels, you do not need the DOM at all. That is the whole idea behind the recognition half of a captcha solving API: you send a picture, you get back what to do with it — which tiles to click, which text to type, or where to drag.&lt;/p&gt;

&lt;p&gt;Most of my own automation ends up combining the two: browser for the clicking, an API for the reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that trips people up: getting the pixels
&lt;/h2&gt;

&lt;p&gt;If you go the image route, the hard part is not the model. It is extracting the picture you are actually looking at. I have hit three distinct versions of this and they all look like "the OCR is bad":&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The iframe's default height.&lt;/strong&gt; A cross-origin iframe with no explicit size renders at a default height (150px is the classic value) while the content inside keeps its own layout. Screenshot the iframe element and you get a correctly-sized image of a clipped puzzle. Mine came out 280x150 and I had assumed a 3x3 grid of ~100px tiles, because that is what a vendor demo page looks like.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tainted canvas.&lt;/strong&gt; Reading pixels back out of a canvas is only permitted if nothing drawn into it came from another origin. Draw a cross-origin image into a canvas without CORS approval and the canvas is &lt;em&gt;tainted&lt;/em&gt;: &lt;code&gt;getImageData()&lt;/code&gt;, &lt;code&gt;toDataURL()&lt;/code&gt; and &lt;code&gt;toBlob()&lt;/code&gt; all throw a &lt;code&gt;SecurityError&lt;/code&gt; instead of returning pixels rather than silently returning nothing (the rule is documented on &lt;a href="https://developer.mozilla.org/en-US/docs/Web/HTML/CORS_enabled_image" rel="noopener noreferrer"&gt;MDN's CORS-enabled image page&lt;/a&gt;). The nasty version of this failure is not the exception. It is the blank capture: right dimensions, plausible file size, no content, and nothing in the logs to say so. Compare a headed run against a headless run — mine only failed headless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Someone else's challenge.&lt;/strong&gt; This is the one that stung. A test page of mine still had a leftover anti-bot script, so the solver ran happily at 2am and solved a challenge that a third-party script had injected to verify &lt;em&gt;me&lt;/em&gt;. The call succeeded, the answer was correct, and the result went nowhere, because it belonged to a different form. The fix is one line: log the frame URL alongside the challenge type. If the two disagree, you are solving a puzzle that was not yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  A shape that survives all of this
&lt;/h2&gt;

&lt;p&gt;What I do now, and what I would tell anyone starting:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Assume every challenge is cross-origin until you check.&lt;/strong&gt; Find the iframe, log its &lt;code&gt;src&lt;/code&gt;, and confirm which vendor you are actually dealing with before writing any solving logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compute click targets from rectangles, not from selectors.&lt;/strong&gt; Take the iframe's bounding box, get your offset from the challenge response, and convert to real pixels. Coordinates that come back from recognition services are typically percentages (0-100) for exactly this reason: the service does not know your image's size, so it tells you the position in a form you can scale.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dump the first frame to disk and look at it with your own eyes.&lt;/strong&gt; Not a benchmark — the actual image. This one habit caught two of the three bugs above.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distrust a successful call whose result you did not observe being consumed.&lt;/strong&gt; A green checkmark means the API answered, not that the answer was used.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is about model accuracy. The model was fine the entire time. I was feeding it the wrong picture, or the right picture for the wrong form, and both of those look identical from the outside.&lt;/p&gt;

&lt;h2&gt;
  
  
  Note on the extension route
&lt;/h2&gt;

&lt;p&gt;There is a third option worth knowing about if you do not want to own this pipeline: ship it as a browser extension. The extension runs inside the browser, so it can take the screenshot, do the cropping and inject the result without you writing any of that glue, and the vendor's iframe boundary stops mattering because the extension is not a page script. I maintain a &lt;a href="https://fuckcaptcha.top/download" rel="noopener noreferrer"&gt;captcha solver extension&lt;/a&gt; that handles this for the common challenge types; if you would rather build your own, the four rules above are the parts that cost me the most time.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
