DEV Community

Shahid Saleem
Shahid Saleem

Posted on

Why modern sites are invisible to AI search engines (and how to test raw HTML)

Most modern web applications look great in the browser. But if you rely on client-side rendering or heavy hydration, AI search engines—like OpenAI's OAI-SearchBot, PerplexityBot, and Claude-SearchBot—might be seeing a completely blank page.

The Headless Bot Bottleneck

Unlike Google's full-rendering web crawlers, AI search engines crawl on strict compute and latency budgets. When an LLM crawler fetches your URL:

  • It frequently captures raw HTML without waiting for client-side JavaScript to execute.
  • It scans for structured metadata (schema.jsonld) and clean text blocks directly in the initial response.
  • It inspects your robots.txt for explicit bot permissions.

If your content only lives after hydration, answer engines will silently skip citing your domain.

How to Inspect What Bots See

You can run a quick terminal test using curl with a bot user-agent to see your raw server response:

curl -A "OAI-SearchBot" -L "[https://usecrawlable.com](https://usecrawlable.com)"
Enter fullscreen mode Exit fullscreen mode

If that command returns empty <div> containers or boilerplate loading states, bots cannot cite you.

We built an open diagnostic engine called Crawlable to automate this. It crawls up to 40 pages without JavaScript, evaluates token density and bot policies, and generates production Fix Kits (robots.txt, llms.txt, and JSON-LD).

How is your engineering team preparing for Generative Engine Optimization (GEO)?

Top comments (0)