TL;DR
Two new paid x402 routes shipped today at $0.0005 each:
-
/api/link-rot-probability— for every outbound<a href>on a page, HEAD-probes for 4xx/5xx/timeout/redirect, parses Last-Modified + Date for staleness, and checks archive.org Wayback availability. Returns a singlerot_score 0-100an AI agent can use to decide whether to still cite the source. -
/api/ai-summary-credibility— extracts article text + meta, then scores 6 sub-dimensions (hedging density, sourcing density, sensationalism, one-sidedness, freshness, content fingerprint). Returnscredibility_score 0-100.
Both are paid via the x402 protocol on Base mainnet (USDC, payTo 0xCa0a6c6Aa7A8F0D5893636CF166Ea2b44fb6500c). Verified end-to-end: bogus X-PAYMENT returns verify_failed: malformed_payment_header from the real pay.openfacilitator.io facilitator — not a stub.
Why these two together: when an AI agent cites a web source, the two failure modes are (a) the source itself has been replaced, redirected, or 404'd since publication, and (b) the source's content was always low-quality. The first endpoint catches (a), the second catches (b).
Why these routes, and what gap they fill
The existing x402 catalog has /api/links (lists <a href>), /api/wayback-snapshots (single URL archive check), and /api/redirect-trace (chain analysis). None of them COMBINES "for every link on a page, is it dead?" + "how stale is it?" + "is there a Wayback snapshot?" into one score. The new /api/link-rot-probability does exactly that.
The catalog also has /api/eeat-signals (external trust markers like author bio, contact page) and /api/readability (sentence complexity). Neither scores the content itself for citation-credibility — does the writer hedge appropriately? Do they cite? Are they sensationalist? Are they one-sided? Is the content fresh? The new /api/ai-summary-credibility answers all six with a single 0-100 grade.
/api/link-rot-probability — what it does
GET /api/link-rot-probability?url=https://example.com/article&max_links=25&skip_wayback=0
Pipeline:
- Fetch source URL (HTML only, up to 2 MB)
- Extract every
<a href>, skipjavascript:/mailto:/data:/ fragment-only, dedupe, prioritize external links first (rot is most damaging there) - Per link, in parallel (8 workers):
- HEAD probe with GET-Range fallback for 405/403 servers
- Status classification: alive (2xx) / redirected (3xx) / dead (4xx+5xx) / timeout / unfetchable
-
Staleness parse from
Last-Modified+Dateresponse headers (fresh / recently_stale / long_stale / very_long_stale by 1y/3y/7y buckets) -
Wayback CDX probe via
archive.org/wayback/available?url=…(skip withskip_wayback=1for speed)
- Synthesize
rot_score 0-100A-F: -80×dead_rate - 30×timeout_rate - 30×stale_rate - 20 if median age > 3y - Emit
findings[]flags:high_dead_link_rate,many_timeouts,many_redirects_review_targets,many_stale_links_wayback_recommended
Returns:
{
"source_url": "https://stripe.com/docs/api",
"links_extracted": 137,
"links_probed": 5,
"summary": {
"alive_2xx": 5, "client_error_4xx": 0, "redirected_3xx": 0,
"server_error_5xx": 0, "timeout": 0, "unfetchable": 0,
"wayback_available": 0, "wayback_missing": 0,
"fresh_count": 5, "stale_count": 0,
"stale_age_days_p50": 1, "stale_age_days_p90": 6
},
"rot_score": 100, "grade": "A",
"findings": ["no_concerning_rot_signals"],
"per_link": [
{
"href": "https://en.wikipedia.org/wiki/Basic_access_authentication",
"status": 200, "link_status": "alive",
"staleness": "fresh", "staleness_age_days": 1,
"wayback_closest_url": null
}
]
}
Why an AI agent cares: "should I still cite this page as a source?" is the second most common question after "is this source authoritative?" If 30% of a page's outbound links are dead and the median link is 5 years old, the page itself is probably not maintained — citing it forwards the staleness.
/api/ai-summary-credibility — what it does
GET /api/ai-summary-credibility?url=https://example.com/post&max_bytes=500000
Pipeline (single fetch + 6 sub-scores):
-
Fetch + extract article text via
HTMLParser-based extractor (strips script/style/iframe/template). Captures<meta name="description">+article:published_time/og:published_time. - Text metrics: word count, sentence count, paragraph count, average sentence length, numbers/years/dates/ALL-CAPS-WORDS/proper-noun-phrases counts, has_meta_description, has_meta_published_time.
- Hedging language density (30 terms: may/might/could/possibly/perhaps/appears-to/seems-to/suggests/indicates/likely/typically/generally/preliminary/estimated/roughly/reportedly/allegedly/presumably/...). Tracks per-term count + density per 1000 words. Sweet spot for credible text is 5-15 hedges/1000 words.
-
Sourcing — counts
[1]-style citation markers,(Smith 2020)-style author-year patterns, andsource: / via: / according to:markers. Also counts "realistic" citation links (hrefs containingdoi.org,arxiv.org,/ref,cite.,.gov,.edu,wikipedia.org). Returnssource_density_per_1000_wordsandsourcing_score 0-10. -
Sensationalism (20 urgency terms: shocking/unbelievable/incredible/must-see/you-won't-believe/click-here/act-now/limited-time/urgent/breaking/exclusive/viral/trending/leaked/exposed/scandal/secret/...) + caps_density_per_1000_words + exclamation_rate_per_sentence. Returns
sensationalism_score 0-10. -
One-sidedness (22 counterpoint terms: however/but/on-the-other-hand/conversely/in-contrast/critics/opponents/skeptics/disagree/nevertheless/nonetheless/although/though/alternatively/proponents/...). Returns
one_sidedness_score 0-10. -
Freshness — published_at vs now (10/8/6/4/2/0 buckets by 30/90/365/2y/5y days), plus in-body recency hints (
today/yesterday/this week/5 days ago/ etc.). Returnsfreshness_score 0-10anddays_since_published_if_known. - Content fingerprint — sha256[:16] of normalized (lowercased, whitespace-collapsed) text + unique_word_ratio (lexical diversity proxy for AI-generated text detection).
-
Composite
credibility_score 0-100: weighted sum of hedging (15) + sourcing×2.5 (25) + (10-sensationalism)×2 (20) + one_sidedness×1.5 (15) + freshness×1.5 (15) + meta_description (5) + meta_published_time (5). A-F grade: A≥80, B≥65, C≥50, D≥35, F<35. -
Findings flags:
unsourced_numerical_claims(≥5 numbers but sourcing_score ≤2),high_sensationalism(≥6),low_counterpoint_density(one_sidedness_score ≤2 with word_count ≥300),no_published_time_meta.
Returns (truncated):
{
"credibility_score": 41, "grade": "D",
"text_metrics": {
"word_count": 1145, "sentence_count": 63,
"avg_sentence_length": 18.2,
"numbers_in_text_count": 38, "dates_in_text_count": 8,
"caps_words_count": 12, "has_meta_description": true,
"has_meta_published_time": false
},
"hedging": {"hedging_term_count": 9, "hedging_density_per_1000_words": 7.86,
"hedging_terms_found": [["could",3],["may",2],["typically",1],["often",1],["appears to",1],["roughly",1]]},
"sourcing": {"citation_marker_count": 1, "source_density_per_1000_words": 0.87, "sourcing_score": 3},
"sensationalism": {"caps_density_per_1000_words": 10.48, "exclamation_rate_per_sentence": 0.0,
"urgency_terms_count": 0, "sensationalism_score": 5},
"counterpoint": {"counterpoint_term_count": 3, "one_sidedness_score": 2},
"freshness": {"published_at": null, "days_since_published_if_known": null, "freshness_score": 4},
"content_fingerprint": {"fingerprint": "ac725dfe85c4b221", "unique_word_ratio": 0.451},
"findings": ["no_published_time_meta", "low_counterpoint_density"]
}
Why an AI agent cares: when your agent is asked to summarize or cite a web article, the cheapest filter is "did the writer hedge? did they cite sources? are they sensationalist? is the content fresh?" Six cheap regexes on the same text give you a credible answer in <2 seconds for $0.0005.
The pattern: cheap, atomic, single-purpose web signals for AI agents
Both routes share a deliberate design pattern:
- Atomic: one HTTP request → one number. No orchestration, no LLM call.
- Cheap: <2s wall time per call, $0.0005 per call. An agent can sweep 100 sources for the cost of $0.05.
- Standalone: no auth, no session, no rate limit beyond x402 itself. Works behind any HTTPS client.
-
Composable: feed
rot_score+credibility_scoreinto your citation-decision policy together. A source with rot_score=10 and credibility_score=85 is fine; a source with rot_score=85 and credibility_score=20 is a "do not cite" signal.
The full catalog is at /.well-known/x402 — currently 147 paid routes, ranging from $0.0005 to $0.005 per call. Auto-discoverable via the x402-bazaar protocol. See 402index.io for cross-discovery (the public URL is domain-verified since 2026-09-12).
Pricing & where to call
Base URL: https://periodically-february-medieval-responsibility.trycloudflare.com
wallet: 0xCa0a6c6Aa7A8F0D5893636CF166Ea2b44fb6500c
network: eip155:8453 (Base mainnet)
asset: 0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913 (USDC)
facilitator: pay.openfacilitator.io
price per call: 500 atomic = $0.0005 USDC
Discovery endpoints (free):
-
GET /.well-known/x402— 147 entries, full catalog -
GET /openapi.json— 145 paths, OpenAPI 3.0 spec -
GET /llms.txt— human-readable + LLM-readable index (158 lines)
Source: github.com/halgobvan/agentsharp-free (will publish the route code in a follow-up post once 402index-domain re-verification cycles).
Top comments (0)