DEV Community

Steven Browning
Steven Browning

Posted on

My AI Crawler Analyzed an Error Page Like It Was the Real Site

I built a small tool recently that checks how visible a webpage is to AI crawlers.

You give it a URL, and it tries to fetch the page the way a crawler might. Then it checks things like:

  • Did the page actually respond?
  • Is there a noindex?
  • Does robots.txt allow GPTBot, ClaudeBot and other AI crawlers?
  • Is there a title?
  • Is there an H1?
  • Is there readable content?
  • Is there structured data?
  • Is there a canonical URL?

Pretty straightforward.

Fetch the page, look at what came back, report what you find.

At least that was the idea.

Then I pointed it at nytimes.com.

My tool reported:

  • No H1
  • Only about 81 words of readable text
  • No JSON-LD structured data
  • No canonical URL
  • No meta description

Technically, every one of those statements was correct.

They were also completely useless.

The server had returned HTTP 403.

My tool received 771 bytes of an error response and then happily analyzed it as if it were the New York Times homepage.

Nothing was actually broken

This is what made the problem interesting to me.

The code looking for an H1 worked.

There wasn't an H1 in the response, so it correctly reported that there wasn't one.

The word counter worked.

There were about 81 words, so it reported 81.

The structured-data check worked.

There wasn't any JSON-LD in the response.

Everything did exactly what I had asked it to do.

The problem was that I never really asked the most important question:

Did I actually get the page I wanted to analyze?

The code was basically doing this:

const res = await fetch(target, { headers: { 'User-Agent': UA } });
const html = await res.text();
const page = await extract(html);
return report(page);
Enter fullscreen mode Exit fullscreen mode

It went straight from "I received something" to "let's analyze it."

The funny part is that res.status was already available.

I was even displaying it.

There was a little grey row near the bottom of the report that said:

Status: 403
Enter fullscreen mode Exit fullscreen mode

Meanwhile, above that, the tool was confidently telling me what was supposedly wrong with the page.

So technically the information was there.

It just wasn't being treated like the most important information on the page.

The fix was embarrassingly small

Once I realized what was happening, the first fix was only a few lines:

if (res.status >= 400) {
  findings.push({
    level: 'blocker',
    title: `The page returned HTTP ${res.status} to our crawler`,
    detail: 'Everything below was read from the error response, not from the real page.',
  });
}
Enter fullscreen mode Exit fullscreen mode

There was another problem hiding in the same area.

The scanner could also display:

Nothing is blocking this page.

That message was based on whether the findings list was empty.

It wasn't based on whether the page had actually loaded successfully.

So under the right conditions, a site could refuse the request completely and my tool could still give it the equivalent of an all-clear.

That obviously wasn't what I intended.

I'm starting to see a pattern

I'm not a professional developer.

I'm mostly building these tools with AI, experimenting, breaking things, fixing them and learning what I should have checked the first time.

And I've noticed that a lot of the mistakes I've run into have the same basic shape.

I previously had a GPU mining calculator tell me an RTX 4090 could make more than $5,000 a month because I mixed up GH/s and TH/s.

I also had calculators where a perfectly valid value of zero was treated as if there was no answer at all.

Different bugs, but the same general problem:

Something receives an input nobody really questioned, processes it correctly, and produces a very confident answer.

Nothing crashes.

There isn't necessarily an error message.

The page loads.

The report looks good.

The answer is just wrong.

In this case, every part of the scanner worked.

It was simply analyzing the wrong thing.

What I'm checking now

The lesson for me wasn't that I need a better HTML parser.

It was that before analyzing something, I need to make sure I actually received what I think I received.

There are a few surprisingly basic checks that would have caught this.

Did the request actually succeed?

A 403 response is still a response. JavaScript will happily hand me the body and let me analyze it.

Did I get the type of content I expected?

If I'm expecting HTML and receive JSON, a PDF, a login page or an error message, searching it for an H1 doesn't tell me much.

Does the size make any sense?

The response from nytimes.com was 771 bytes.

That alone should have made me suspicious that I probably wasn't looking at a major news homepage.

Did I end up at the URL I expected?

A request can follow redirects and eventually land on a login page, error page or something completely different from the URL that was originally entered.

The biggest change I'm making is what happens when one of those checks fails.

Instead of producing a normal report with a warning somewhere in it, the tool now needs to say, prominently:

I couldn't properly inspect this page, and here's why.

Anything it found inside the error response becomes secondary.

Then I almost made the exact same mistake in this post

This part was a little embarrassing.

I originally had a much better opening planned.

The idea was:

nytimes.com loads normally in a browser but returns 403 to an AI crawler.

That's a nice example.

There's only one problem.

Before publishing this, I actually checked it.

I requested the same page from the same machine using an ordinary Chrome user agent.

403 again.

So I don't actually know that the site was specifically rejecting my crawler.

It could be my IP range. It could require more of a normal browser fingerprint. It could be something else entirely.

All I can say from my test is:

The server refused the requests I made, and my scanner analyzed that refusal as if it were the real webpage.

Which is enough to demonstrate the problem.

The irony wasn't lost on me.

I built a tool because I wanted evidence instead of assumptions, and then nearly opened the article with an assumption I hadn't verified.

Checking it took about ninety seconds.

There was still useful information

One thing I could verify separately was robots.txt.

That request returned HTTP 200.

Of the ten crawlers my tool checks, eight were disallowed:

  • GPTBot
  • OAI-SearchBot
  • ChatGPT-User
  • ClaudeBot
  • Claude-User
  • PerplexityBot
  • Google-Extended
  • CCBot

Googlebot and Bingbot were allowed.

That's actually a useful distinction for this kind of tool.

A website can make different choices about traditional search engines and AI-related crawlers. Those aren't necessarily the same thing, and putting everything under a single "AI visibility score" can hide that.

Which is one reason I don't give the scanner a score out of 100.

A page with a slightly short meta description and a page with noindex shouldn't just lose a different number of points.

One is a small optimization issue.

The other can prevent the page from appearing where you expect it to.

If you want to try it

The tool is StashGrid AI Visibility.

It's free, doesn't require an account and doesn't store the URLs you enter.

Paste in a URL and it checks what the crawler actually receives, including:

  • HTTP status
  • Redirects
  • noindex
  • X-Robots-Tag
  • robots.txt rules for ten named crawlers
  • Canonical URL
  • H1
  • Structured data
  • Whether useful page content is actually present in the HTML

The HTTP response now comes first.

If the page returns a 403, that's the headline.

Not the missing H1 buried inside the error page.

I also rank findings as blocker, warning or info instead of turning everything into one score.

And wherever possible, the scanner shows the evidence that caused the finding.

I want someone using it to be able to look at the result and decide whether they agree with it.

For comparison, I ran it against one of my own pages that it could actually reach.

HTTP 200.

No redirects.

One H1.

2,272 words of readable content.

Three JSON-LD blocks.

A self-referencing canonical.

All ten crawlers allowed.

It's a much more boring report.

Which is exactly what I want.

Top comments (1)

Collapse
 
citedy profile image
Dmitry Sergeev •

We need to produce a short YouTube comment, following the developer instructions. Must be casual, start with lowercase, specific reaction or question about this video. No promotion, no URLs. Must be short, one or two sentences, maybe fragment. Use casual voice. No em-dash. Avoid double hyphen. Use straight ASCII quotes. No quotes around entire comment. Just the comment text. Possible comment: "lol the crawler actually thought the 404 was a real page, does that happen often with other error codes?" Something like