DEV Community

citablehub
citablehub

Posted on

What AI crawlers actually read on your site (and a 40-line script to check)

AI assistants answer questions by citing a small set of sources. Before you can be one of them, a crawler has to be able to read you at all.

Many sites return almost nothing without JavaScript. I wrote a small script to check, ran it on my own site, and the results were uncomfortable.

The script

No dependencies. Saves as readcheck.py.

import re
import sys
import urllib.request

UA = "Mozilla/5.0 (compatible; Crawler/1.0)"


def meta(html, name):
    """Return a meta tag's content, whatever the attribute order is."""
    for tag in re.findall(r"<meta[^>]*>", html, re.I):
        if name not in tag:
            continue
        found = re.search(r'content="([^"]*)"', tag, re.I)
        if found:
            return found.group(1)
    return "(none)"


def main():
    url = sys.argv[1]
    req = urllib.request.Request(url, headers={"User-Agent": UA})
    with urllib.request.urlopen(req, timeout=30) as response:
        html = response.read().decode("utf-8", "replace")

    # 1. The one-line definition a model is most likely to quote
    print("description:", meta(html, "og:description")[:160] or meta(html, "description")[:160])

    # 2. Machine-checkable facts
    blocks = re.findall(
        r'<script[^>]*type="application/ld\+json"[^>]*>(.*?)</script>', html, re.S | re.I
    )
    print("structured-data blocks:", len(blocks))

    # 3. Text a crawler can actually read, without executing JavaScript
    text = re.sub(r"<(script|style)[^>]*>.*?</\1>", " ", html, flags=re.S | re.I)
    text = re.sub(r"<[^>]+>", " ", text)
    text = " ".join(text.split())
    print("readable characters:", len(text))
    print("first 200:", text[:200])


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Run it:

python readcheck.py https://yoursite.com
Enter fullscreen mode Exit fullscreen mode

How to read the output

Three numbers, in order of importance:

  1. description — the sentence an assistant is most likely to echo back. If it is stale or vague, that is the sentence you are being summarized as.
  2. structured-data blocks — the machine-checkable facts. Zero means a model has to guess.
  3. readable characters — how much text survives without JavaScript. Under about 1,000 and a crawler has almost nothing to work with.

What I got on my own site

Page Readable characters Structured-data blocks
A profile page 24,167 3
The homepage 14,529 3
The directory index 3,286 1

The deepest page won. The page I had actually been optimizing lost.

Two things I took from it:

Depth beats decoration when a machine is doing the reading. One page with concrete facts (what it is, who it is for, evidence, dates) out-produced my hero section by 7x in readable text.

My metadata was lying. The description on my index said "100+" while the site lists 759. That number sits in the exact place machines read first — and it was wrong. Fixed the same day.

The bigger gap: read versus cited

Then I checked the other side of the ledger. I logged every automated reader that hit one of our platforms for 28 days:

  • 97,955 reads from AI crawlers
  • 96,468 of them from one crawler alone
  • the rest spread across a dozen others

And how many of those turned into an actual citation inside an answer?

Zero.

That gap is the point. Being read is not being cited. A crawler stocking a warehouse is not a model recommending you.

Three changes I made after seeing it:

  1. I stopped counting crawler traffic as a win. It is an input metric, not an outcome.
  2. I started measuring citations the only way I found that works: asking the assistants the questions we should win, on a schedule, and logging whether we were named.
  3. I made our facts machine-checkable — one canonical name, one URL, structured data, evidence with dates. A model cites what it can verify, not what it can read.

Takeaway

Run the script on your own site before you spend money on anything called "AI SEO".

If it returns under 1,000 readable characters, you do not have a marketing problem. You have a rendering problem.


I build CitableHub, a free directory that publishes machine-readable profiles for software projects — that is the site I measured above. The script works on any site.

Top comments (0)