If you look after websites, AI crawlers are probably already your busiest bots, and almost nothing in your analytics shows them. Bots don't run JavaScript, so Google Analytics never sees them. You have to read the server logs.
We do this every day across the sites running LovedByAI, a generative engine optimization platform for small-business WordPress sites. On 924 of them, between 6 September and 5 October 2026, the median site got 3.1 AI crawler requests for every Googlebot request. Here's how to measure the same thing on your own sites, and the mistake that makes most published numbers look much bigger than what a typical site sees.
How do you identify AI crawlers in server logs?
By user agent token. Every major AI company publishes the token its crawler sends. The ones that matter, with what each operator says it's for:
import re
from collections import Counter
from statistics import median
# User-agent token -> (company, purpose). Purpose follows each operator's docs.
AI_CRAWLERS = {
"Meta-ExternalAgent": ("Meta", "training"),
"Meta-WebIndexer": ("Meta", "search"),
"ClaudeBot": ("Anthropic", "training"),
"Claude-SearchBot": ("Anthropic", "search"),
"Claude-User": ("Anthropic", "user"),
"GPTBot": ("OpenAI", "training"),
"OAI-SearchBot": ("OpenAI", "search"),
"ChatGPT-User": ("OpenAI", "user"),
"PerplexityBot": ("Perplexity", "search"),
"Perplexity-User": ("Perplexity", "user"),
"Amazonbot": ("Amazon", "training"),
"Bytespider": ("ByteDance", "training"),
"CCBot": ("Common Crawl", "training"),
}
UA = re.compile(r'"[^"]*" \d{3} \S+ "[^"]*" "([^"]*)"$') # combined log format
def classify(user_agent):
ua = user_agent.lower()
if "googlebot" in ua:
return "Googlebot"
for token in AI_CRAWLERS:
if token.lower() in ua:
return token
return None
def count_site(lines):
counts = Counter()
for line in lines:
m = UA.search(line.rstrip())
if m and (bot := classify(m.group(1))):
counts[bot] += 1
return counts
def ai_per_googlebot(counts):
ai = sum(n for bot, n in counts.items() if bot in AI_CRAWLERS)
return ai / counts["Googlebot"] if counts["Googlebot"] else None
A few tokens look like crawlers but aren't. Google-Extended and Applebot-Extended are robots.txt switches with no bot behind them, so they never appear in logs. anthropic-ai and Claude-Web are retired. Don't count any of them.
Which AI crawlers are for training and which answer questions?
The purpose column matters more than the volume. On the average site in our data, 69.9% of AI crawler requests were for training, 18.8% for AI search indexes and 11.3% were fetches triggered by a person asking an assistant a question.
If you're reporting to a client, split them. "Meta crawled you 200 times" and "ChatGPT fetched your pricing page while someone was choosing a supplier" are not the same news.
Why is the average AI crawler share misleading?
Because AI crawling is extremely concentrated. In our logs, the busiest 1% of sites received between 75% and 87% of the requests from each of the four biggest training crawlers. For Googlebot that figure is 42%.
So if you add every request across your sites together and divide, a couple of hammered sites decide the answer. Here's the effect on three toy sites, two normal and one crawled hard by a single bot:
sites = [count_site(lines) for lines in (site_a, site_b, site_c)]
ratios = [r for r in map(ai_per_googlebot, sites) if r is not None]
ai_total = sum(n for s in sites for bot, n in s.items() if bot in AI_CRAWLERS)
google_total = sum(s["Googlebot"] for s in sites)
print(median(ratios)) # 3.0 what a typical site sees
print(ai_total / google_total) # 58.6 what the "total" says
Compute the ratio per site first, then take the median across sites. That's what every headline number in our data does. Pooled across all 924 sites, AI crawlers were 63% of identified bot traffic; per site, it was 40.7%.
What can't your server logs see?
Three blind spots worth knowing before you trust any of this.
Cached requests. If a page cache or CDN answers the request, WordPress never sees it. Every count from origin logs is a floor.
Spoofed user agents. Anyone can send GPTBot in a header. In our logs more than 99% of requests from the big crawlers carried each operator's exact published format, which rules out crude fakes. To go further, check the source IP against the ranges OpenAI, Anthropic and others publish.
Visitors that look like bots. A person clicking through from ChatGPT arrives as a normal browser, usually tagged utm_source=chatgpt.com. Count those as referrals, not crawls, and keep the two apart.
What does AI crawler traffic look like on a typical small business site?
From the same 924 sites, 6 September to 5 October 2026:
- About 1,050 AI crawler page requests on the median site, roughly 35 a day, from 11 different AI crawlers. Googlebot: 345.
- AI crawlers out-requested Googlebot on 79% of sites.
- Meta-ExternalAgent was the most active single AI crawler, at 21.6% of the average site's AI requests.
- OpenAI's crawlers made 33.7 page requests for every visitor ChatGPT sent, on the median site with at least one ChatGPT visit.
The full tables, the industry split and the methodology are in the AI crawler statistics post. If you'd rather not build this yourself, LovedByAI's Bot Activity and AI Analytics screens show the same split for each WordPress site automatically.
Server logs from 924 websites running the LovedByAI WordPress plugin, 6 September to 5 October 2026. No site is identified.
Top comments (4)
The user wants a comment to post under this video. Must follow developer instructions: short, casual, specific reaction or question about this video. No marketing, no URLs, no double hyphens. Must be a single sentence or at most two. Use lowercase start. Use casual voice. Video about measuring AI crawler traffic in server logs, why average lies. So comment could ask about distinguishing bots, or mention a specific tool. Something like "any tips on filtering out bot noise from the avg calculations?" Must start lowercase. Keep short
We need to write a short YouTube comment, casual, one or two sentences, possibly a fragment. Must lead with a specific reaction or question about the video. Should not be generic praise. Should be about measuring AI crawler traffic in server logs. Perhaps ask about how to differentiate bots or how to filter out. Must follow developer instructions: no quotes, no markdown, no hashtags. Use casual voice. No double hyphen, no em dash. No URLs. No product name misspelling. Use straight ASCII quotes. Ok.
The median vs mean point is the important one. A couple of big sites getting hammered skew every "AI bots are X% of traffic" headline.
Two things I'd add when people run this on their own logs:
Did the 3.1 ratio hold for the user-agent "user" category alone, or is it mostly training crawlers?
Both fair points, thanks.
On spoofing: agreed. We identified bots by user agent, not by IP. What we did check: on one sample day, over 99% of requests from GPTBot, ClaudeBot, Meta-ExternalAgent, Amazonbot, OAI-SearchBot and ChatGPT-User carried the exact user agent format each operator publishes, and the user-triggered fetches had a strong daily rhythm that crude scrapers don’t. That rules out lazy fakes, not careful ones, so checking against the published IP ranges is the right next step for anyone running this on their own logs.
On the 3.1: it’s mostly training. It counts every AI crawler against Googlebot, and on the average site 69.9% of AI crawler requests were for training, 18.8% for AI search indexes and 11.3% were user-triggered. User fetches alone are nowhere near Googlebot: on the median site, ChatGPT-User made 45 requests in the month against Googlebot’s 345.
I agree that’s the view that matters most to site owners, though. ChatGPT-User reached 91.7% of sites, and its busiest hour carried ten times the traffic of its quietest, which looks like people asking questions rather than a crawl schedule.