AI assistants answer questions by citing a small set of sources. Before you can be one of them, a crawler has to be able to read you at all.
Many sites return almost nothing without JavaScript. I wrote a small script to check, ran it on my own site, and the results were uncomfortable.
The script
No dependencies. Saves as readcheck.py.
import re
import sys
import urllib.request
UA = "Mozilla/5.0 (compatible; Crawler/1.0)"
def meta(html, name):
"""Return a meta tag's content, whatever the attribute order is."""
for tag in re.findall(r"<meta[^>]*>", html, re.I):
if name not in tag:
continue
found = re.search(r'content="([^"]*)"', tag, re.I)
if found:
return found.group(1)
return "(none)"
def main():
url = sys.argv[1]
req = urllib.request.Request(url, headers={"User-Agent": UA})
with urllib.request.urlopen(req, timeout=30) as response:
html = response.read().decode("utf-8", "replace")
# 1. The one-line definition a model is most likely to quote
print("description:", meta(html, "og:description")[:160] or meta(html, "description")[:160])
# 2. Machine-checkable facts
blocks = re.findall(
r'<script[^>]*type="application/ld\+json"[^>]*>(.*?)</script>', html, re.S | re.I
)
print("structured-data blocks:", len(blocks))
# 3. Text a crawler can actually read, without executing JavaScript
text = re.sub(r"<(script|style)[^>]*>.*?</\1>", " ", html, flags=re.S | re.I)
text = re.sub(r"<[^>]+>", " ", text)
text = " ".join(text.split())
print("readable characters:", len(text))
print("first 200:", text[:200])
if __name__ == "__main__":
main()
Run it:
python readcheck.py https://yoursite.com
How to read the output
Three numbers, in order of importance:
-
description— the sentence an assistant is most likely to echo back. If it is stale or vague, that is the sentence you are being summarized as. -
structured-data blocks— the machine-checkable facts. Zero means a model has to guess. -
readable characters— how much text survives without JavaScript. Under about 1,000 and a crawler has almost nothing to work with.
What I got on my own site
| Page | Readable characters | Structured-data blocks |
|---|---|---|
| A profile page | 24,167 | 3 |
| The homepage | 14,529 | 3 |
| The directory index | 3,286 | 1 |
The deepest page won. The page I had actually been optimizing lost.
Two things I took from it:
Depth beats decoration when a machine is doing the reading. One page with concrete facts (what it is, who it is for, evidence, dates) out-produced my hero section by 7x in readable text.
My metadata was lying. The description on my index said "100+" while the site lists 759. That number sits in the exact place machines read first — and it was wrong. Fixed the same day.
The bigger gap: read versus cited
Then I checked the other side of the ledger. I logged every automated reader that hit one of our platforms for 28 days:
- 97,955 reads from AI crawlers
- 96,468 of them from one crawler alone
- the rest spread across a dozen others
And how many of those turned into an actual citation inside an answer?
Zero.
That gap is the point. Being read is not being cited. A crawler stocking a warehouse is not a model recommending you.
Three changes I made after seeing it:
- I stopped counting crawler traffic as a win. It is an input metric, not an outcome.
- I started measuring citations the only way I found that works: asking the assistants the questions we should win, on a schedule, and logging whether we were named.
- I made our facts machine-checkable — one canonical name, one URL, structured data, evidence with dates. A model cites what it can verify, not what it can read.
Takeaway
Run the script on your own site before you spend money on anything called "AI SEO".
If it returns under 1,000 readable characters, you do not have a marketing problem. You have a rendering problem.
I build CitableHub, a free directory that publishes machine-readable profiles for software projects — that is the site I measured above. The script works on any site.
Top comments (0)