Most of the time when a client tells me "we are not showing up in ChatGPT," the site is a React app that renders everything on the client. The HTML that comes over the wire is a <div id="root"></div> and a script tag. A human with a browser sees a full page. A crawler that does not execute JavaScript sees nothing.
Google runs a rendering queue and will eventually execute your JavaScript. Most AI crawlers do not, or do it inconsistently. If you want to be cited by an AI assistant, the safest assumption is that the bot reads exactly what curl reads. Here is how to check, and what to change.
Step 1: Look at what the bot actually receives
Skip DevTools. DevTools shows you the DOM after hydration, which is not what a crawler gets. Use curl and pretend to be a bot.
curl -s -A "GPTBot" https://example.com/services/plumbing | head -c 3000
Now do the same with a few other user agents that matter right now:
for ua in "GPTBot" "ClaudeBot" "PerplexityBot" "Google-Extended" "Googlebot"; do
echo "== $ua"
curl -s -o /dev/null -w "%{http_code} %{size_download} bytes\n" -A "$ua" https://example.com/services/plumbing
done
Two things to read from the output. First, the status code: a 403 for one bot and a 200 for another means a WAF or bot rule is blocking you and you probably did not choose that. Second, the byte count: if the page is 1.8 KB for a bot and 90 KB in the browser, your content is not in the HTML.
Then grep the raw response for something that only exists in your real content, like a phone number or a service name:
curl -s -A "GPTBot" https://example.com/services/plumbing | grep -c "water heater"
Zero means the text is being injected by JavaScript after load. That is the whole problem in one number.
Step 2: Check robots.txt for accidental blocks
A lot of teams copied a robots.txt from a template in 2023 that blocked AI crawlers wholesale, and nobody revisited it. Pull it and look:
curl -s https://example.com/robots.txt
If you see User-agent: GPTBot followed by Disallow: /, you have opted out of being cited. That is a legitimate choice for some businesses. It is a bad choice for a local service company that wants referrals from AI assistants. Decide it on purpose, do not inherit it.
Also check the CDN layer. Cloudflare, Vercel, and others have toggles for AI bot blocking that override anything in your file. The curl status codes in step 1 will reveal that even when robots.txt looks clean.
Step 3: Render the important pages on the server
This is the fix, and in most modern frameworks it is a configuration change, not a rewrite.
In Next.js with the App Router, a page is a server component by default. If your service pages are client components because someone added "use client" at the top for a single interactive widget, move the widget into its own client component and keep the page itself on the server. The HTML will contain the full text on the first response.
// app/services/plumbing/page.tsx (server component, no "use client")
import QuoteForm from "@/components/QuoteForm"; // this one is "use client"
export default async function PlumbingPage() {
const data = await getServiceContent("plumbing");
return (
<main>
<h1>{data.title}</h1>
<p>{data.intro}</p>
<QuoteForm />
</main>
);
}
If you are on Vite plus React with no server, you have three realistic options: prerender the static routes at build time (vite-plugin-ssr, Astro, or a simple prerender script with Playwright), move the marketing pages to a framework that renders on the server and leave the app where it is, or accept that those pages are invisible to non-rendering crawlers. The third one is fine for an authenticated dashboard. It is not fine for the pages you want customers to find.
When we take on a client site, the marketing pages are always server rendered from the start. We build them as static or server-rendered Next.js sites on Vercel, and it costs nothing extra at that scale. The dashboard behind login can be as client-heavy as it wants.
Step 4: Make sure the content is real HTML, not a JSON blob
Some sites technically ship the content in the initial response, but only inside a __NEXT_DATA__ or similar JSON script tag, with the visible markup still rendered client side. Crawlers vary in whether they read that. Do not depend on it. The test is simple: view the raw response and confirm the <h1>, the paragraphs, and the internal links exist as markup.
While you are there, confirm the links are real anchors with href attributes, not onClick handlers on divs. A crawler cannot follow a click handler.
Step 5: Recheck after every deploy
Put the curl checks in CI. A small script that fetches five key URLs as GPTBot, asserts a 200, and asserts that a known phrase exists in the raw body takes ten minutes to write and catches the regression where a developer wraps a layout in "use client" and silently blanks the site for every crawler.
#!/usr/bin/env bash
set -e
for url in "$@"; do
body=$(curl -s -A "GPTBot" "$url")
echo "$body" | grep -q "<h1" || { echo "FAIL no h1 in $url"; exit 1; }
echo "$body" | grep -q "levelupdigitalmarketinggroup\|Call\|Contact" || { echo "FAIL no contact text in $url"; exit 1; }
echo "ok $url"
done
The short version
Crawlers read source, not screens. Curl your pages as a bot, check the status and the byte count, grep for real content, fix the robots and CDN rules you forgot about, and render the pages that matter on the server. Everything else in AI search optimization sits on top of that, and none of it works if the bot receives an empty div.
Carlynn Espinoza runs Level Up Digital Marketing Group, a San Diego agency that builds and markets websites for local service businesses.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.