Most agents cannot sign up for anything. A coding agent can write code against an API, but a general assistant that is halfway through a task will not open a browser, create an account and paste a key. What it can do is fetch a URL.
So I put two of our tools behind plain URLs, with no account and no key:
https://tools.yukai.uk/md/<any URL> -> the page as clean Markdown
https://tools.yukai.uk/ai-crawlers/<domain> -> which AI crawlers the site's robots.txt allows
Both are free and rate limited (about 10 requests per minute per IP). CORS is open, so a browser tool can call them too. Here is what they return, using real calls from 1 October 2026.
1. Any page as Markdown
curl https://tools.yukai.uk/md/https://docs.python.org/3/library/pathlib.html
---
title: "pathlib — Object-oriented filesystem paths — Python 3.14.8 documentation"
url: "https://docs.python.org/3/library/pathlib.html"
status: 200
tokens_estimate: 20907
truncated: false
source: "TidyTools free tier (no key, rate limited). ..."
---
# `pathlib` — Object-oriented filesystem paths
Added in version 3.4.
**Source code:** [Lib/pathlib/](https://github.com/python/cpython/tree/3.14/Lib/pathlib/)
...
The YAML header is for the model. It says what the page is, where the fetch ended up after redirects, and roughly how many tokens it will cost before the agent reads the body. Navigation, cookie banners and scripts are stripped, links stay as Markdown links. Output is capped at 100,000 characters (truncated: true tells you when that happened).
If you would rather have JSON, ask for it:
curl -H "Accept: application/json" https://tools.yukai.uk/md/https://react.dev/
From Python it is one line:
import urllib.request
md = urllib.request.urlopen("https://tools.yukai.uk/md/https://example.com/").read().decode()
And from an agent, you do not need any code. Tell it something like this in the system prompt or a tool description:
To read a web page, fetch
https://tools.yukai.uk/md/<url>and use the Markdown it returns.
What it does not do
The free tier is a plain HTTP fetch. It does not run JavaScript, so a single-page app that ships an empty <div id="root"> comes back nearly empty. The header then carries a note saying so. It also does not crawl whole sites, convert PDFs or chunk text for RAG. Those need a real browser or more compute, and they are in the paid versions (an Apify Actor that crawls whole sites at $1 per 1,000 pages, and a RapidAPI API with a free monthly plan).
2. Which AI crawlers does a site allow?
The second endpoint reads a site's robots.txt, llms.txt and noai meta tags and checks 29 AI crawlers against them: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended and the rest.
curl https://tools.yukai.uk/ai-crawlers/nytimes.com
Four sites, same day:
| Site | Policy | AI search score | Blocked (first few) |
|---|---|---|---|
| nytimes.com | blocks search | 41 | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User |
| theguardian.com | blocks search | 65 | ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-CloudVertexBot |
| wikipedia.org | open | 100 | none |
| stackoverflow.com | unknown | – | robots.txt answered HTTP 403, so access cannot be verified |
The interesting split is between training crawlers (GPTBot, ClaudeBot) and search/assistant crawlers (OAI-SearchBot, Claude-SearchBot, ChatGPT-User). Blocking the first keeps your content out of model training. Blocking the second keeps you out of AI answers, which is often not what a publisher meant. The response spells out that difference in a verdict line and suggests a robots.txt snippet.
An agent can use it before it scrapes ("am I allowed to fetch this?"), and a site owner can use it to see what their robots.txt actually says to AI.
Why give this away
Two reasons. Agents pick tools they can use without setup, and a tool nobody has heard of does not get picked. If the free endpoint is useful, some people will need the heavier versions: JavaScript rendering, whole-site crawls, thousands of URLs. Those are paid.
The machine-readable list of everything else (screenshots, SEO audits, company enrichment, OCR, transcription and so on) is at tools.yukai.uk/llms.txt.
If something breaks or comes back wrong, tell me in the comments with the URL you tried.
Top comments (0)