I keep a few single-purpose web tools running on the Cloudflare Workers free tier. Each one does one job, calls no LLM API at request time, and loads nothing from external CDNs. Here are three of them, with requests you can copy, and two related things I offer at the end.
1. AI Crawler Checker: what does your robots.txt say to GPTBot?
URL: https://ai-crawler-checker.hdg-os.workers.dev
Enter a domain and the tool fetches its robots.txt, then gives a per-crawler verdict for 15 AI user agents: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-User, Claude-SearchBot, Google-Extended, PerplexityBot, Perplexity-User, Bytespider, CCBot, Amazonbot, Applebot-Extended, Diffbot and meta-externalagent. The UI is in Korean, but the API returns plain JSON:
curl 'https://ai-crawler-checker.hdg-os.workers.dev/api/check?domain=example.com'
curl 'https://ai-crawler-checker.hdg-os.workers.dev/api/bots'
Each bot gets one of four verdicts:
-
allowed: a matching group exists, but no rule blocks it -
partial: the root path is open, some paths are disallowed -
blocked: the root path is disallowed, so the whole site is off limits -
no-rules: no group applies at all (effectively allowed)
Matching follows RFC 9309: user agent tokens are case-insensitive, a bot without its own group falls back to the * group, the longer pattern wins when allow and disallow collide, and allow wins on equal length. An empty Disallow: means "no restriction", not a rule. The response also includes a ready-to-paste blockSnippet that disallows every bot not yet blocked.
A second page at /llms (English UI) checks and generates llms.txt. The checks come from the llmstxt.org proposal text: the only required section is the H1, so a missing H1 is the only error; a missing summary blockquote, non-H2 sections, list items without a markdown link, and H3-or-deeper headings are warning.
curl 'https://ai-crawler-checker.hdg-os.workers.dev/api/llms-check?domain=example.com'
2. Character Limit Checker: where exactly does the text get cut?
URL: https://char-limit-checker.hdg-os.workers.dev
Paste a caption, bio, post or meta tag once, and the page compares it against the limits of X (Twitter), Instagram, LinkedIn, Threads, Bluesky, Mastodon, TikTok, YouTube, Discord, Reddit, Pinterest, Google Ads and search result title/description, showing pass and fail side by side. When something fails, it marks the exact cut point.
Three counts are shown separately because they disagree on real text: grapheme clusters, UTF-8 bytes and words. A family emoji is one grapheme but several UTF-16 code units, and platforms count differently. For SMS the tool decides between GSM-7 and UCS-2 encoding and reports the segment count, since one character outside GSM-7 shrinks your per-segment budget.
Everything runs in the browser; the worker only serves static files, so the text you paste never leaves your machine.
Four sub-pages each add one operation the main page lacks:
-
/meta-description-length-checkersimulates pixel-width truncation withcanvas.measureText(920 px desktop, 680 px mobile, both editable), because Google states there is no fixed character limit -
/title-tag-length-checkerruns the same pixel engine plus six rewrite-risk checks mapped to Google's documented title rewriting cases -
/youtube-title-character-limitseparates the 100 character hard cap from the visible cut (70 by default) -
/discord-character-limit-checkercompares the 2,000 / 4,000 (Nitro) / embed field limits and splits overflow at blank lines, then line breaks, then spaces, keeping grapheme clusters intact and re-opening code fences across chunks
3. Korean Terms Summary: ToS;DR-style cards for Korean services
URL: https://terms-summary.hdg-os.workers.dev
This one is in Korean and is for anyone who has to read Korean terms of service. Search a service name and you get plain-language summaries of the clauses worth noticing (automatic renewal, refund limits, liability caps, third-party data sharing) next to the article number and the quoted source text. ToS;DR does not cover Korean services, which is the gap this fills.
The database currently holds 10 services: Naver, Kakao T, KakaoPay, Toss, Gmarket, Netflix, Google, Airbnb, Albamon and Apple Media Services. Summaries are written by a person from the original text, not generated on request, so a page view costs nothing.
curl 'https://terms-summary.hdg-os.workers.dev/api/search?q=토스'
curl 'https://terms-summary.hdg-os.workers.dev/api/stats'
The collector honors robots.txt with RFC 9309 matching, waits at least 2 seconds between requests, respects Crawl-delay, and never publishes full original text, only summaries and source URLs.
How they are built and tested
Each tool is one static HTML file plus a worker.js. The rules live in exactly one place, and a Python script re-implements the same rules independently, then cross-checks fixtures, the Python result and the JS result. If one side drifts, the check fails. The crawler checker's regression set has 240 cases; the character checker compared 9,544 values with 0 mismatches on its 2026-09-04 run.
Two related things I offer
Auto-check badge. I sell an AI Crawler Policy Auto-Check Badge on Gumroad for $19 (one domain, 365 days). Running the check itself is free at https://verify-badge.hdg-os.workers.dev/check?domain=example.com; buying unlocks a permanent result page and a small SVG badge that shows only your domain and the date the automatic check passed. It is an automatic check, not a certification, and the result page lists what it does not cover.
Mobile app AI agent. As freelance work I build agents that run on an always-on PC and control an Android phone over wireless debugging, with nothing installed on the phone: they scan an app on a schedule, let an AI judge the detail screen, draft the first message, and wait for your approval via Telegram before sending. Approval gates and rate limits are on by default, and I do not bypass app policies or identity checks.
max hwang, full-stack developer in Korea
Top comments (1)
The independent Python/JS fixture check is useful, and the badge's explicit "not a certification" boundary is worth keeping.
One classification case I'd add:
Disallow: /together withAllow: /public/. Under the longest-pattern rule you describe, the root is blocked but/public/is allowed. Would that get apartialverdict rather than "the whole site is off limits"? Showing the matched rule and an example path beside the verdict would make that distinction easier to inspect, especially before someone pastes the generated block snippet.