DEV Community

SwiftKit
SwiftKit

Posted on

Give your AI agent a 'read this URL' tool: clean Markdown in ~2 seconds, even for JavaScript apps

Most AI agents need one boring tool: "read this URL". Raw HTML wastes tokens. A typical article page is 80–95% menus, scripts, cookie banners and share buttons. What the model actually needs is the main text as Markdown, with headings, lists, tables and links.

Here's the pipeline I ended up with, and the one trick that keeps it fast.

1. Download first, browser only if needed

A headless browser handles everything, but it's slow (seconds per page) and heavy. Most pages don't need one: blogs, docs, news and company sites ship their text in the HTML.

So: plain fetch first. Only if the result is nearly empty, open the page in Chrome:

const APP_SHELL = /<div[^>]+id=["'](root|app|__next|__nuxt)["'][^>]*>\s*<\/div>/i;

const html = await (await fetch(url)).text();
let c = convert(html, url);
const thin = c.text.length < 400 && (APP_SHELL.test(html) || c.text.length < 150);
if (thin) {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'domcontentloaded' });
  await page.waitForLoadState('networkidle', { timeout: 10000 }).catch(() => {});
  c = convert(await page.content(), page.url());
}
Enter fullscreen mode Exit fullscreen mode

Two details that matter:

  • Block images, fonts and video in the browser (context.route). You only want text, and it halves load time.
  • Check the browser's HTTP status. A 403 error page renders beautifully and converts into confident-looking Markdown about "Request blocked". Treat status 400+ as a failure, not content.

2. Find the main content with Readability

Mozilla Readability is the engine behind Firefox Reader View. It works on any DOM, so on a server you can pair it with linkedom (much lighter than jsdom):

import { parseHTML } from 'linkedom';
import { Readability } from '@mozilla/readability';

const { document } = parseHTML(html);
const article = new Readability(document, { charThreshold: 200 }).parse();
// article.content = cleaned HTML, article.title, article.byline
Enter fullscreen mode Exit fullscreen mode

If Readability returns too little (home pages, product listings), fall back to <main> or <body> minus nav, header, footer, aside and cookie banners.

3. HTML to Markdown with Turndown

import TurndownService from 'turndown';
import { gfm } from 'turndown-plugin-gfm';

const td = new TurndownService({ headingStyle: 'atx', codeBlockStyle: 'fenced' });
td.use(gfm); // tables and strikethrough
td.remove(['script', 'style', 'iframe', 'form', 'svg']);
// Icon-only links (share buttons) carry no text: drop them
td.addRule('emptyLinks', {
  filter: (n) => n.nodeName === 'A' && !n.textContent.trim() && !n.querySelector('img'),
  replacement: () => '',
});
const markdown = td.turndown(article.content);
Enter fullscreen mode Exit fullscreen mode

Make links and image URLs absolute before converting (new URL(href, pageUrl)), or the Markdown breaks the moment it leaves the site.

4. Be a polite bot

  • Respect robots.txt (longest matching rule wins, like Google).
  • Send an honest User-Agent.
  • If the page title is "Just a moment…" or "Access denied", stop. That's a bot check, and getting around it isn't your call to make.

The packaged version

I run this as Website to Markdown for AI on Apify. You can crawl whole sites into RAG chunks, or use it as an instant API: one HTTP request with a URL, Markdown back in about 2 seconds, including JavaScript apps:

curl "https://swiftkit--web-to-markdown.apify.actor/?url=https://docs.apify.com&format=markdown" \
  -H "Authorization: Bearer YOUR_APIFY_TOKEN"
Enter fullscreen mode Exit fullscreen mode

It's a fixed price per converted page, and failed or blocked pages are free.

What does your agent use for reading pages today: raw HTML, a reader API, or something home-made?


Written with AI assistance; the code and numbers were tested before publishing.

Top comments (0)