DEV Community

Tidy Tools
Tidy Tools

Posted on

A no-key URL-to-Markdown endpoint your AI agent can call right now

Most agents cannot sign up for anything. A coding agent can write code against an API, but a general assistant that is halfway through a task will not open a browser, create an account and paste a key. What it can do is fetch a URL.

So I put two of our tools behind plain URLs, with no account and no key:

https://tools.yukai.uk/md/<any URL>             -> the page as clean Markdown
https://tools.yukai.uk/ai-crawlers/<domain>     -> which AI crawlers the site's robots.txt allows
Enter fullscreen mode Exit fullscreen mode

Both are free and rate limited (about 10 requests per minute per IP). CORS is open, so a browser tool can call them too. Here is what they return, using real calls from 1 October 2026.

1. Any page as Markdown

curl https://tools.yukai.uk/md/https://docs.python.org/3/library/pathlib.html
Enter fullscreen mode Exit fullscreen mode
---
title: "pathlib — Object-oriented filesystem paths — Python 3.14.8 documentation"
url: "https://docs.python.org/3/library/pathlib.html"
status: 200
tokens_estimate: 20907
truncated: false
source: "TidyTools free tier (no key, rate limited). ..."
---
# `pathlib` — Object-oriented filesystem paths

Added in version 3.4.

**Source code:** [Lib/pathlib/](https://github.com/python/cpython/tree/3.14/Lib/pathlib/)
...
Enter fullscreen mode Exit fullscreen mode

The YAML header is for the model. It says what the page is, where the fetch ended up after redirects, and roughly how many tokens it will cost before the agent reads the body. Navigation, cookie banners and scripts are stripped, links stay as Markdown links. Output is capped at 100,000 characters (truncated: true tells you when that happened).

If you would rather have JSON, ask for it:

curl -H "Accept: application/json" https://tools.yukai.uk/md/https://react.dev/
Enter fullscreen mode Exit fullscreen mode

From Python it is one line:

import urllib.request
md = urllib.request.urlopen("https://tools.yukai.uk/md/https://example.com/").read().decode()
Enter fullscreen mode Exit fullscreen mode

And from an agent, you do not need any code. Tell it something like this in the system prompt or a tool description:

To read a web page, fetch https://tools.yukai.uk/md/<url> and use the Markdown it returns.

What it does not do

The free tier is a plain HTTP fetch. It does not run JavaScript, so a single-page app that ships an empty <div id="root"> comes back nearly empty. The header then carries a note saying so. It also does not crawl whole sites, convert PDFs or chunk text for RAG. Those need a real browser or more compute, and they are in the paid versions (an Apify Actor that crawls whole sites at $1 per 1,000 pages, and a RapidAPI API with a free monthly plan).

2. Which AI crawlers does a site allow?

The second endpoint reads a site's robots.txt, llms.txt and noai meta tags and checks 29 AI crawlers against them: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended and the rest.

curl https://tools.yukai.uk/ai-crawlers/nytimes.com
Enter fullscreen mode Exit fullscreen mode

Four sites, same day:

Site Policy AI search score Blocked (first few)
nytimes.com blocks search 41 GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User
theguardian.com blocks search 65 ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-CloudVertexBot
wikipedia.org open 100 none
stackoverflow.com unknown – robots.txt answered HTTP 403, so access cannot be verified

The interesting split is between training crawlers (GPTBot, ClaudeBot) and search/assistant crawlers (OAI-SearchBot, Claude-SearchBot, ChatGPT-User). Blocking the first keeps your content out of model training. Blocking the second keeps you out of AI answers, which is often not what a publisher meant. The response spells out that difference in a verdict line and suggests a robots.txt snippet.

An agent can use it before it scrapes ("am I allowed to fetch this?"), and a site owner can use it to see what their robots.txt actually says to AI.

Why give this away

Two reasons. Agents pick tools they can use without setup, and a tool nobody has heard of does not get picked. If the free endpoint is useful, some people will need the heavier versions: JavaScript rendering, whole-site crawls, thousands of URLs. Those are paid.

The machine-readable list of everything else (screenshots, SEO audits, company enrichment, OCR, transcription and so on) is at tools.yukai.uk/llms.txt.

If something breaks or comes back wrong, tell me in the comments with the URL you tried.

Top comments (0)