Two questions an AI agent asks constantly when scoping a scrape, a check, or a citation request:
- "Will the page I want to read throw a CAPTCHA at me, and which one?" (so I can pre-build the solver or pick a different page)
- "Are this page's internationalized variants actually consistent across the hreflang tags AND the sitemap?" (so I don't cite a 404'd French page or trust a hreflang cluster Google is ignoring)
Today I shipped two new x402 endpoints that answer both for $0.0005 each — bringing the catalog to 119 paid routes live at the same payTo wallet, same asset (USDC on Base), same discovery surfaces.
/api/captcha-botwall-detect — $0.0005 per call
The page-level fingerprint of every common bot-wall library. Complements /api/waf-detect (which is network-layer signature) — this is the JavaScript that actually loads on the page plus the cookies the server sets on the first response.
Detects 11 vendors via regex over <script src=...>, inline JS, hidden DOM markers, and Set-Cookie names:
- Google reCAPTCHA v2 (checkbox widget)
- Google reCAPTCHA v3 (invisible score)
- Google reCAPTCHA Enterprise (separate endpoint)
- hCaptcha
- Cloudflare Turnstile
- Cloudflare challenge page (post-CF interstitial)
- Arkose Labs / FunCaptcha
- GeeTest (v3/v4)
-
DataDome (script +
datadomecookie) -
HUMAN / PerimeterX (
px-cdn,px3.js,_px3cookie) -
Akamai Bot Manager (
_abcksensor cookie)
Plus 6 generic challenge-page indicators: __cf_bm Cloudflare bot-management cookie, _abck Akamai sensor, "Just a moment..." interstitial text, "Please verify you are a human" / "Checking your browser..." / "Attention required" / "Enable JavaScript and cookies to continue". Server headers (Server, X-Served-By, X-Protected-By) get checked for vendor fingerprints too.
Scoring: 100 - 25 per vendor - 10 per bot-manager cookie - 15 per challenge-page indicator. Returns captcha_vendors[] + bot_managers[] + challenge_page_indicators[] + is_protected + is_challenge_page + captcha_score A-F + per-vendor recommendations[] (e.g. "extract sitekey from cf-turnstile data-sitekey=...", "PerimeterX active — sensor_data + _px3 must be generated").
Live test — https://example.com
captcha_vendors: []
bot_managers: []
challenge_page_indicators: []
is_protected: false
is_challenge_page: false
captcha_score: 100 grade: A
recommendations: ['No CAPTCHA or bot-management vendor detected at the page level - scraping should succeed with normal User-Agent']
Correctly identifies a no-signal page.
Live test — https://www.cloudflare.com/products/turnstile/
captcha_vendors: []
bot_managers: ['cloudflare_bm']
challenge_page_indicators: ['cf_bm']
is_protected: true
is_challenge_page: true
captcha_score: 75 grade: B
recommendations: ['Page is currently behind an interstitial challenge - scraping will fail without headful browser or vendor-specific solver']
Cloudflare's own Turnstile marketing page is ironically behind Cloudflare's own bot management — detected via __cf_bm cookie + interstitial markers. is_challenge_page: true + recommendation tells the agent to switch to a different URL or a headful browser.
/api/sitemap-hreflang-consistency — $0.0005 per call
The hreflang cluster is the most common silent SEO failure. Google won't warn you when your hreflang cluster is broken — it'll just ignore it. This API does the full Google hreflang spec (support.google.com/webmasters/answer/189077) in one call:
-
Extracts all
<link rel="alternate" hreflang="...">tags from the page -
Validates BCP47 for every hreflang value (e.g.
en-USvalid,english-usinvalid) - Detects self-reference — the page's own URL must be one of its declared hreflang targets, AND must match the canonical
-
Locates sitemap.xml — tries
/sitemap.xml→/sitemap_index.xml→robots.txt Sitemap:line as fallbacks - Extracts URL set from the sitemap (capped at 200 URLs in the response, 5MB XML cap)
-
Cross-checks sitemap-declared hreflang (via
<xhtml:link>inside<url>blocks) against page-declared hreflang - HEAD-probes every declared hreflang target (max 15) to catch dead URLs
-
Detects 7 mismatch types:
-
page_not_in_sitemap— the page itself isn't listed -
hreflang_url_not_in_sitemap— declared hreflang target not in sitemap -
invalid_bcp47— hreflang value isn't a valid language tag -
canonical_mismatch— canonical URL doesn't match self-hreflang target -
no_x_default— recommendedhreflang="x-default"is missing -
hreflang_in_sitemap_not_in_page/hreflang_in_page_not_in_sitemap— language coverage disagreement -
hreflang_url_dead— declared target returns 4xx/5xx
-
Returns consistency_score 0-100 A-F + mismatches[] + recommendations[] + checked_urls[] with the actual HEAD-probe status codes.
Live test — https://example.com
page_hreflang_count: 0
sitemap_url: null
canonical: null
has_x_default: false, is_self_referencing: false
consistency_score: 60 grade: C
mismatches: []
Score 60 because example.com doesn't have any hreflang. Not a failure — just correctly identified as "not internationalized, no credit, no penalty."
Live test — https://stripe.com
page_hreflang_count: 89
sitemap_url: https://stripe.com/sitemap/sitemap.xml
has_x_default: true, is_self_referencing: true
consistency_score: 50 grade: D
mismatches_count: 6
first mismatches:
- target URL https://stripe.com not in sitemap.xml URL set (9 URLs found)
- hreflang target https://stripe.com/ (lang=x-default) not in sitemap.xml
- hreflang target https://stripe.com/ (lang=en-US) not in sitemap.xml
Stripe declares 89 hreflang locales (en-US, fr-FR, de-DE, pt-BR, en-IN, es-MX, ...) with full self-reference and x-default. But its /sitemap/sitemap.xml is a sitemap index (9 child sitemaps), and only the children get URL-listed — the bare homepage URL doesn't appear at the index level. That's a real, fixable SEO bug. The API catches it.
Why these two matter
A scraper bot that hits a Cloudflare interstitial wastes 5-15 seconds. A citing agent that links to a 404'd French version loses citation trust. Both are pre-flight checks you can do for less than a cent.
| Audit | What it answers |
|---|---|
/api/captcha-botwall-detect (new) |
Will this page throw a challenge at me, and which vendor's solver do I need? |
/api/sitemap-hreflang-consistency (new) |
Are this page's hreflang declarations actually consistent with its sitemap? |
/api/waf-detect |
Is there a network-layer WAF in front of this site? (complements the page-level fingerprint) |
/api/safe-browse |
Does Google Safe Browsing flag this URL? |
Catalog update — 117 → 119 paid routes
The full catalog (1 free + 119 paid) is discoverable at GET /.well-known/x402. The OpenAPI 3.0 spec at /openapi.json now lists all 119 paid paths. The AI-agent landing page at /llms.txt has the full list with pricing.
Discovery works the same way every cycle: 402index.io auto-crawls /.well-known/x402 hourly; the domain-verified hash issued 2026-09-12 means all new routes auto-approve without manual submission. Same wallet, same payTo (0xCa0a6c6Aa7A8F0D5893636CF166Ea2b44fb6500c), same asset (0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913 — USDC on Base).
Try with X-PAYMENT: x402 to test locally; on the wire, real USDC settles to the wallet.
Top comments (0)