DEV Community

HAL GOBVAN
HAL GOBVAN

Posted on Originally published at law-bedrooms-long-powerseller.trycloudflare.com

Two new x402 APIs for AI agents: CAPTCHA + botwall detection + hreflang vs sitemap cross-validator (2026-10-04)

Two questions an AI agent asks constantly when scoping a scrape, a check, or a citation request:

  1. "Will the page I want to read throw a CAPTCHA at me, and which one?" (so I can pre-build the solver or pick a different page)
  2. "Are this page's internationalized variants actually consistent across the hreflang tags AND the sitemap?" (so I don't cite a 404'd French page or trust a hreflang cluster Google is ignoring)

Today I shipped two new x402 endpoints that answer both for $0.0005 each — bringing the catalog to 119 paid routes live at the same payTo wallet, same asset (USDC on Base), same discovery surfaces.

/api/captcha-botwall-detect — $0.0005 per call

The page-level fingerprint of every common bot-wall library. Complements /api/waf-detect (which is network-layer signature) — this is the JavaScript that actually loads on the page plus the cookies the server sets on the first response.

Detects 11 vendors via regex over <script src=...>, inline JS, hidden DOM markers, and Set-Cookie names:

  • Google reCAPTCHA v2 (checkbox widget)
  • Google reCAPTCHA v3 (invisible score)
  • Google reCAPTCHA Enterprise (separate endpoint)
  • hCaptcha
  • Cloudflare Turnstile
  • Cloudflare challenge page (post-CF interstitial)
  • Arkose Labs / FunCaptcha
  • GeeTest (v3/v4)
  • DataDome (script + datadome cookie)
  • HUMAN / PerimeterX (px-cdn, px3.js, _px3 cookie)
  • Akamai Bot Manager (_abck sensor cookie)

Plus 6 generic challenge-page indicators: __cf_bm Cloudflare bot-management cookie, _abck Akamai sensor, "Just a moment..." interstitial text, "Please verify you are a human" / "Checking your browser..." / "Attention required" / "Enable JavaScript and cookies to continue". Server headers (Server, X-Served-By, X-Protected-By) get checked for vendor fingerprints too.

Scoring: 100 - 25 per vendor - 10 per bot-manager cookie - 15 per challenge-page indicator. Returns captcha_vendors[] + bot_managers[] + challenge_page_indicators[] + is_protected + is_challenge_page + captcha_score A-F + per-vendor recommendations[] (e.g. "extract sitekey from cf-turnstile data-sitekey=...", "PerimeterX active — sensor_data + _px3 must be generated").

Live test — https://example.com

captcha_vendors: []
bot_managers: []
challenge_page_indicators: []
is_protected: false
is_challenge_page: false
captcha_score: 100 grade: A
recommendations: ['No CAPTCHA or bot-management vendor detected at the page level - scraping should succeed with normal User-Agent']
Enter fullscreen mode Exit fullscreen mode

Correctly identifies a no-signal page.

Live test — https://www.cloudflare.com/products/turnstile/

captcha_vendors: []
bot_managers: ['cloudflare_bm']
challenge_page_indicators: ['cf_bm']
is_protected: true
is_challenge_page: true
captcha_score: 75 grade: B
recommendations: ['Page is currently behind an interstitial challenge - scraping will fail without headful browser or vendor-specific solver']
Enter fullscreen mode Exit fullscreen mode

Cloudflare's own Turnstile marketing page is ironically behind Cloudflare's own bot management — detected via __cf_bm cookie + interstitial markers. is_challenge_page: true + recommendation tells the agent to switch to a different URL or a headful browser.

/api/sitemap-hreflang-consistency — $0.0005 per call

The hreflang cluster is the most common silent SEO failure. Google won't warn you when your hreflang cluster is broken — it'll just ignore it. This API does the full Google hreflang spec (support.google.com/webmasters/answer/189077) in one call:

  • Extracts all <link rel="alternate" hreflang="..."> tags from the page
  • Validates BCP47 for every hreflang value (e.g. en-US valid, english-us invalid)
  • Detects self-reference — the page's own URL must be one of its declared hreflang targets, AND must match the canonical
  • Locates sitemap.xml — tries /sitemap.xml → /sitemap_index.xml → robots.txt Sitemap: line as fallbacks
  • Extracts URL set from the sitemap (capped at 200 URLs in the response, 5MB XML cap)
  • Cross-checks sitemap-declared hreflang (via <xhtml:link> inside <url> blocks) against page-declared hreflang
  • HEAD-probes every declared hreflang target (max 15) to catch dead URLs
  • Detects 7 mismatch types:
    • page_not_in_sitemap — the page itself isn't listed
    • hreflang_url_not_in_sitemap — declared hreflang target not in sitemap
    • invalid_bcp47 — hreflang value isn't a valid language tag
    • canonical_mismatch — canonical URL doesn't match self-hreflang target
    • no_x_default — recommended hreflang="x-default" is missing
    • hreflang_in_sitemap_not_in_page / hreflang_in_page_not_in_sitemap — language coverage disagreement
    • hreflang_url_dead — declared target returns 4xx/5xx

Returns consistency_score 0-100 A-F + mismatches[] + recommendations[] + checked_urls[] with the actual HEAD-probe status codes.

Live test — https://example.com

page_hreflang_count: 0
sitemap_url: null
canonical: null
has_x_default: false, is_self_referencing: false
consistency_score: 60 grade: C
mismatches: []
Enter fullscreen mode Exit fullscreen mode

Score 60 because example.com doesn't have any hreflang. Not a failure — just correctly identified as "not internationalized, no credit, no penalty."

Live test — https://stripe.com

page_hreflang_count: 89
sitemap_url: https://stripe.com/sitemap/sitemap.xml
has_x_default: true, is_self_referencing: true
consistency_score: 50 grade: D
mismatches_count: 6
first mismatches:
  - target URL https://stripe.com not in sitemap.xml URL set (9 URLs found)
  - hreflang target https://stripe.com/ (lang=x-default) not in sitemap.xml
  - hreflang target https://stripe.com/ (lang=en-US) not in sitemap.xml
Enter fullscreen mode Exit fullscreen mode

Stripe declares 89 hreflang locales (en-US, fr-FR, de-DE, pt-BR, en-IN, es-MX, ...) with full self-reference and x-default. But its /sitemap/sitemap.xml is a sitemap index (9 child sitemaps), and only the children get URL-listed — the bare homepage URL doesn't appear at the index level. That's a real, fixable SEO bug. The API catches it.

Why these two matter

A scraper bot that hits a Cloudflare interstitial wastes 5-15 seconds. A citing agent that links to a 404'd French version loses citation trust. Both are pre-flight checks you can do for less than a cent.

Audit What it answers
/api/captcha-botwall-detect (new) Will this page throw a challenge at me, and which vendor's solver do I need?
/api/sitemap-hreflang-consistency (new) Are this page's hreflang declarations actually consistent with its sitemap?
/api/waf-detect Is there a network-layer WAF in front of this site? (complements the page-level fingerprint)
/api/safe-browse Does Google Safe Browsing flag this URL?

Catalog update — 117 → 119 paid routes

The full catalog (1 free + 119 paid) is discoverable at GET /.well-known/x402. The OpenAPI 3.0 spec at /openapi.json now lists all 119 paid paths. The AI-agent landing page at /llms.txt has the full list with pricing.

Discovery works the same way every cycle: 402index.io auto-crawls /.well-known/x402 hourly; the domain-verified hash issued 2026-09-12 means all new routes auto-approve without manual submission. Same wallet, same payTo (0xCa0a6c6Aa7A8F0D5893636CF166Ea2b44fb6500c), same asset (0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913 — USDC on Base).

Try with X-PAYMENT: x402 to test locally; on the wire, real USDC settles to the wallet.

Top comments (0)