DEV Community

max
max

Posted on

3 free web tools I built: AI crawler check, text limits, Korean ToS

I keep a few single-purpose web tools running on the Cloudflare Workers free tier. Each one does one job, calls no LLM API at request time, and loads nothing from external CDNs. Here are three of them, with requests you can copy, and two related things I offer at the end.

1. AI Crawler Checker: what does your robots.txt say to GPTBot?

URL: https://ai-crawler-checker.hdg-os.workers.dev

Enter a domain and the tool fetches its robots.txt, then gives a per-crawler verdict for 15 AI user agents: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-User, Claude-SearchBot, Google-Extended, PerplexityBot, Perplexity-User, Bytespider, CCBot, Amazonbot, Applebot-Extended, Diffbot and meta-externalagent. The UI is in Korean, but the API returns plain JSON:

curl 'https://ai-crawler-checker.hdg-os.workers.dev/api/check?domain=example.com'
curl 'https://ai-crawler-checker.hdg-os.workers.dev/api/bots'
Enter fullscreen mode Exit fullscreen mode

Each bot gets one of four verdicts:

  • allowed: a matching group exists, but no rule blocks it
  • partial: the root path is open, some paths are disallowed
  • blocked: the root path is disallowed, so the whole site is off limits
  • no-rules: no group applies at all (effectively allowed)

Matching follows RFC 9309: user agent tokens are case-insensitive, a bot without its own group falls back to the * group, the longer pattern wins when allow and disallow collide, and allow wins on equal length. An empty Disallow: means "no restriction", not a rule. The response also includes a ready-to-paste blockSnippet that disallows every bot not yet blocked.

A second page at /llms (English UI) checks and generates llms.txt. The checks come from the llmstxt.org proposal text: the only required section is the H1, so a missing H1 is the only error; a missing summary blockquote, non-H2 sections, list items without a markdown link, and H3-or-deeper headings are warning.

curl 'https://ai-crawler-checker.hdg-os.workers.dev/api/llms-check?domain=example.com'
Enter fullscreen mode Exit fullscreen mode

2. Character Limit Checker: where exactly does the text get cut?

URL: https://char-limit-checker.hdg-os.workers.dev

Paste a caption, bio, post or meta tag once, and the page compares it against the limits of X (Twitter), Instagram, LinkedIn, Threads, Bluesky, Mastodon, TikTok, YouTube, Discord, Reddit, Pinterest, Google Ads and search result title/description, showing pass and fail side by side. When something fails, it marks the exact cut point.

Three counts are shown separately because they disagree on real text: grapheme clusters, UTF-8 bytes and words. A family emoji is one grapheme but several UTF-16 code units, and platforms count differently. For SMS the tool decides between GSM-7 and UCS-2 encoding and reports the segment count, since one character outside GSM-7 shrinks your per-segment budget.

Everything runs in the browser; the worker only serves static files, so the text you paste never leaves your machine.

Four sub-pages each add one operation the main page lacks:

  • /meta-description-length-checker simulates pixel-width truncation with canvas.measureText (920 px desktop, 680 px mobile, both editable), because Google states there is no fixed character limit
  • /title-tag-length-checker runs the same pixel engine plus six rewrite-risk checks mapped to Google's documented title rewriting cases
  • /youtube-title-character-limit separates the 100 character hard cap from the visible cut (70 by default)
  • /discord-character-limit-checker compares the 2,000 / 4,000 (Nitro) / embed field limits and splits overflow at blank lines, then line breaks, then spaces, keeping grapheme clusters intact and re-opening code fences across chunks

3. Korean Terms Summary: ToS;DR-style cards for Korean services

URL: https://terms-summary.hdg-os.workers.dev

This one is in Korean and is for anyone who has to read Korean terms of service. Search a service name and you get plain-language summaries of the clauses worth noticing (automatic renewal, refund limits, liability caps, third-party data sharing) next to the article number and the quoted source text. ToS;DR does not cover Korean services, which is the gap this fills.

The database currently holds 10 services: Naver, Kakao T, KakaoPay, Toss, Gmarket, Netflix, Google, Airbnb, Albamon and Apple Media Services. Summaries are written by a person from the original text, not generated on request, so a page view costs nothing.

curl 'https://terms-summary.hdg-os.workers.dev/api/search?q=토스'
curl 'https://terms-summary.hdg-os.workers.dev/api/stats'
Enter fullscreen mode Exit fullscreen mode

The collector honors robots.txt with RFC 9309 matching, waits at least 2 seconds between requests, respects Crawl-delay, and never publishes full original text, only summaries and source URLs.

How they are built and tested

Each tool is one static HTML file plus a worker.js. The rules live in exactly one place, and a Python script re-implements the same rules independently, then cross-checks fixtures, the Python result and the JS result. If one side drifts, the check fails. The crawler checker's regression set has 240 cases; the character checker compared 9,544 values with 0 mismatches on its 2026-09-04 run.

Two related things I offer

Auto-check badge. I sell an AI Crawler Policy Auto-Check Badge on Gumroad for $19 (one domain, 365 days). Running the check itself is free at https://verify-badge.hdg-os.workers.dev/check?domain=example.com; buying unlocks a permanent result page and a small SVG badge that shows only your domain and the date the automatic check passed. It is an automatic check, not a certification, and the result page lists what it does not cover.

Mobile app AI agent. As freelance work I build agents that run on an always-on PC and control an Android phone over wireless debugging, with nothing installed on the phone: they scan an app on a schedule, let an AI judge the detail screen, draft the first message, and wait for your approval via Telegram before sending. Approval gates and rate limits are on by default, and I do not bypass app policies or identity checks.

max hwang, full-stack developer in Korea

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

The independent Python/JS fixture check is useful, and the badge's explicit "not a certification" boundary is worth keeping.

One classification case I'd add: Disallow: / together with Allow: /public/. Under the longest-pattern rule you describe, the root is blocked but /public/ is allowed. Would that get a partial verdict rather than "the whole site is off limits"? Showing the matched rule and an example path beside the verdict would make that distinction easier to inspect, especially before someone pastes the generated block snippet.