Most teams trying to understand AI search start by querying ChatGPT. There's a cheaper signal sitting on your servers already: your access logs.
When ChatGPT, Claude or Perplexity browse the web to answer a user's question, they fetch your page with an identifiable user agent. Each of those hits means a live conversation pulled in that URL, which makes it about the closest thing you have to a server-side "citation candidate" log.
This post shows how to pull those hits out of nginx logs, tell the different kinds of AI bot apart, check they're genuine, and turn them into a list of pages AI assistants actually read.
Three kinds of AI bot, and why the difference matters
AI companies run several crawlers with different jobs. Lumping them together hides the useful signal.
| Purpose | Examples | What a hit means |
|---|---|---|
| Training |
GPTBot, ClaudeBot, Google-Extended*, CCBot
|
Content may go into a future model. No user is waiting. |
| Search indexing |
OAI-SearchBot, Claude-SearchBot, PerplexityBot
|
Your page is being indexed for the engine's search feature. |
| User-triggered fetch |
ChatGPT-User, Claude-User, Perplexity-User
|
A person's conversation fetched this page right now. |
* Google-Extended is a robots.txt control token, not a separate crawler. You won't see it as a user agent in your logs. Google fetches with its normal crawlers.
The third row is the interesting one. A ChatGPT-User hit on /pricing means someone asked ChatGPT a question and it opened your pricing page to answer it.
User agent strings change, so check the vendors' own bot documentation pages (OpenAI, Anthropic and Perplexity each publish one) for the current list.
Step 1: Pull AI bot hits out of the log
Assuming the default nginx combined log format:
# Count hits per AI bot over the current log
grep -oE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|CCBot|Bytespider|Amazonbot|Applebot-Extended' \
/var/log/nginx/access.log | sort | uniq -c | sort -rn
4120 GPTBot
1873 ClaudeBot
912 PerplexityBot
388 OAI-SearchBot
141 ChatGPT-User
57 Perplexity-User
12 Claude-User
(Illustrative numbers.) That shape is typical: training crawlers dominate the volume, and user-triggered fetches are a small but far more meaningful slice.
Step 2: Which pages do live conversations fetch?
Filter to user-triggered agents and count by path:
grep -E 'ChatGPT-User|Claude-User|Perplexity-User' /var/log/nginx/access.log* \
| awk '{print $7}' \
| sed -E 's/\?.*$//' \
| sort | uniq -c | sort -rn | head -20
Over a few weeks of logs, this gives you a ranked list of the pages AI assistants open on behalf of users. Compare it with your Google top pages. The two lists are rarely the same. Docs pages, pricing pages and comparison pages are often overrepresented in AI fetches relative to their search traffic.
Also check the status codes:
grep -E 'ChatGPT-User|Claude-User|Perplexity-User' /var/log/nginx/access.log \
| awk '{print $9}' | sort | uniq -c
Any 403 or 429 in that output means a live conversation tried to read your page and you turned it away. That's usually a WAF or rate-limit rule that nobody set up with AI agents in mind.
Step 3: Make sure the bots are real
User agents are trivially spoofed, and scrapers love pretending to be GPTBot. Before you trust the numbers (or allowlist anything), verify the source IPs.
OpenAI and Perplexity publish JSON files of the IP ranges their bots use, linked from their bot documentation. Here's a verifier that loads those ranges and checks log IPs against them:
# verify_bots.py
import ipaddress, json, re, sys, urllib.request
# Published by each vendor. Confirm on their bot docs pages before relying on them.
RANGES_URLS = {
"ChatGPT-User": "https://openai.com/chatgpt-user.json",
"OAI-SearchBot": "https://openai.com/searchbot.json",
"GPTBot": "https://openai.com/gptbot.json",
"Perplexity-User": "https://www.perplexity.ai/perplexity-user.json",
"PerplexityBot": "https://www.perplexity.ai/perplexitybot.json",
}
def load(url):
data = json.load(urllib.request.urlopen(url, timeout=15))
nets = []
for p in data.get("prefixes", []):
cidr = p.get("ipv4Prefix") or p.get("ipv6Prefix")
if cidr:
nets.append(ipaddress.ip_network(cidr))
return nets
networks = {bot: load(u) for bot, u in RANGES_URLS.items() if u.startswith("http")}
line_re = re.compile(r'^(\S+) .*"([^"]*)"$')
stats = {}
for line in sys.stdin:
m = line_re.match(line.rstrip())
if not m:
continue
ip, ua = m.groups()
for bot, nets in networks.items():
if bot in ua:
addr = ipaddress.ip_address(ip)
ok = any(addr in n for n in nets)
s = stats.setdefault(bot, [0, 0])
s[0 if ok else 1] += 1
for bot, (good, bad) in stats.items():
print(f"{bot:15} verified {good:6} spoofed {bad:6}")
cat /var/log/nginx/access.log | python verify_bots.py
The ranges files use a prefixes list with ipv4Prefix/ipv6Prefix keys, the same format Google uses for Googlebot. If a vendor doesn't publish ranges, fall back to reverse DNS plus forward-confirm, where the vendor supports it.
Step 4: Log it properly going forward
grep works for a quick look. For ongoing tracking, tag AI traffic at the edge so it's queryable without regex archaeology:
map $http_user_agent $ai_bot {
default "";
~*ChatGPT-User "chatgpt-user";
~*Claude-User "claude-user";
~*Perplexity-User "perplexity-user";
~*OAI-SearchBot "openai-search";
~*Claude-SearchBot "claude-search";
~*PerplexityBot "perplexity-search";
~*GPTBot "openai-train";
~*ClaudeBot "claude-train";
}
log_format ai '$time_iso8601\t$remote_addr\t$ai_bot\t$status\t$request_uri';
server {
# ...
access_log /var/log/nginx/ai.log ai if=$ai_bot;
}
The if= parameter writes only requests where $ai_bot is non-empty, so ai.log stays small and tab-separated. Ship it to whatever you already use (Loki, BigQuery, a cron job that loads it into SQLite).
On Cloudflare or another CDN, your origin logs may miss requests the edge served from cache or blocked outright. Use the CDN's own logs or AI-bot analytics for the complete picture.
What logs can't tell you
Server logs show that an assistant read a page. They don't show:
- what the user asked,
- whether the answer cited you or used a competitor instead,
- what the assistant said about you,
- anything from answers built from training data or a search index without a live fetch.
That last one is a big gap. Many answers never trigger a live fetch at all. So logs are a strong supply-side signal ("these pages are being read") and need pairing with a demand-side one ("this is what the answers say").
That pairing is what citation tracking covers. Vista AI's AI Citation Tracking records the prompt, the full response and the cited source for each brand mention across the major AI engines, scores sentiment and flags inaccurate claims. Put the two side by side: a page that's fetched often in your logs but rarely cited in answers is a page the engines find and then pass over, which makes it a good first candidate for a rewrite.
Quick checklist
- [ ] Count AI bot hits by type (training, search, user-triggered)
- [ ] List the top 20 paths fetched by
*-Useragents - [ ] Check those fetches for 403/429 responses
- [ ] Verify bot IPs against published ranges
- [ ] Add an nginx
map+ dedicated AI log for ongoing tracking


Top comments (0)