Most AI agents need one boring tool: "read this URL". Raw HTML wastes tokens. A typical article page is 80–95% menus, scripts, cookie banners and share buttons. What the model actually needs is the main text as Markdown, with headings, lists, tables and links.
Here's the pipeline I ended up with, and the one trick that keeps it fast.
1. Download first, browser only if needed
A headless browser handles everything, but it's slow (seconds per page) and heavy. Most pages don't need one: blogs, docs, news and company sites ship their text in the HTML.
So: plain fetch first. Only if the result is nearly empty, open the page in Chrome:
const APP_SHELL = /<div[^>]+id=["'](root|app|__next|__nuxt)["'][^>]*>\s*<\/div>/i;
const html = await (await fetch(url)).text();
let c = convert(html, url);
const thin = c.text.length < 400 && (APP_SHELL.test(html) || c.text.length < 150);
if (thin) {
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForLoadState('networkidle', { timeout: 10000 }).catch(() => {});
c = convert(await page.content(), page.url());
}
Two details that matter:
-
Block images, fonts and video in the browser (
context.route). You only want text, and it halves load time. - Check the browser's HTTP status. A 403 error page renders beautifully and converts into confident-looking Markdown about "Request blocked". Treat status 400+ as a failure, not content.
2. Find the main content with Readability
Mozilla Readability is the engine behind Firefox Reader View. It works on any DOM, so on a server you can pair it with linkedom (much lighter than jsdom):
import { parseHTML } from 'linkedom';
import { Readability } from '@mozilla/readability';
const { document } = parseHTML(html);
const article = new Readability(document, { charThreshold: 200 }).parse();
// article.content = cleaned HTML, article.title, article.byline
If Readability returns too little (home pages, product listings), fall back to <main> or <body> minus nav, header, footer, aside and cookie banners.
3. HTML to Markdown with Turndown
import TurndownService from 'turndown';
import { gfm } from 'turndown-plugin-gfm';
const td = new TurndownService({ headingStyle: 'atx', codeBlockStyle: 'fenced' });
td.use(gfm); // tables and strikethrough
td.remove(['script', 'style', 'iframe', 'form', 'svg']);
// Icon-only links (share buttons) carry no text: drop them
td.addRule('emptyLinks', {
filter: (n) => n.nodeName === 'A' && !n.textContent.trim() && !n.querySelector('img'),
replacement: () => '',
});
const markdown = td.turndown(article.content);
Make links and image URLs absolute before converting (new URL(href, pageUrl)), or the Markdown breaks the moment it leaves the site.
4. Be a polite bot
- Respect
robots.txt(longest matching rule wins, like Google). - Send an honest User-Agent.
- If the page title is "Just a moment…" or "Access denied", stop. That's a bot check, and getting around it isn't your call to make.
The packaged version
I run this as Website to Markdown for AI on Apify. You can crawl whole sites into RAG chunks, or use it as an instant API: one HTTP request with a URL, Markdown back in about 2 seconds, including JavaScript apps:
curl "https://swiftkit--web-to-markdown.apify.actor/?url=https://docs.apify.com&format=markdown" \
-H "Authorization: Bearer YOUR_APIFY_TOKEN"
It's a fixed price per converted page, and failed or blocked pages are free.
What does your agent use for reading pages today: raw HTML, a reader API, or something home-made?
Written with AI assistance; the code and numbers were tested before publishing.
Top comments (0)