DEV Community

Cover image for Your Access Logs Already Show Which Pages AI Assistants Are Reading
Furqan Khalid
Furqan Khalid

Posted on

Your Access Logs Already Show Which Pages AI Assistants Are Reading

Most teams trying to understand AI search start by querying ChatGPT. There's a cheaper signal sitting on your servers already: your access logs.

When ChatGPT, Claude or Perplexity browse the web to answer a user's question, they fetch your page with an identifiable user agent. Each of those hits means a live conversation pulled in that URL, which makes it about the closest thing you have to a server-side "citation candidate" log.

This post shows how to pull those hits out of nginx logs, tell the different kinds of AI bot apart, check they're genuine, and turn them into a list of pages AI assistants actually read.


Three kinds of AI bot, and why the difference matters

AI companies run several crawlers with different jobs. Lumping them together hides the useful signal.

Purpose Examples What a hit means
Training GPTBot, ClaudeBot, Google-Extended*, CCBot Content may go into a future model. No user is waiting.
Search indexing OAI-SearchBot, Claude-SearchBot, PerplexityBot Your page is being indexed for the engine's search feature.
User-triggered fetch ChatGPT-User, Claude-User, Perplexity-User A person's conversation fetched this page right now.

* Google-Extended is a robots.txt control token, not a separate crawler. You won't see it as a user agent in your logs. Google fetches with its normal crawlers.

The third row is the interesting one. A ChatGPT-User hit on /pricing means someone asked ChatGPT a question and it opened your pricing page to answer it.

User agent strings change, so check the vendors' own bot documentation pages (OpenAI, Anthropic and Perplexity each publish one) for the current list.

Donut chart: training crawlers 80%, search indexing 17%, user-triggered fetches 3% of AI bot hits

Step 1: Pull AI bot hits out of the log

Assuming the default nginx combined log format:

# Count hits per AI bot over the current log
grep -oE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|CCBot|Bytespider|Amazonbot|Applebot-Extended' \
  /var/log/nginx/access.log | sort | uniq -c | sort -rn
Enter fullscreen mode Exit fullscreen mode
   4120 GPTBot
   1873 ClaudeBot
    912 PerplexityBot
    388 OAI-SearchBot
    141 ChatGPT-User
     57 Perplexity-User
     12 Claude-User
Enter fullscreen mode Exit fullscreen mode

(Illustrative numbers.) That shape is typical: training crawlers dominate the volume, and user-triggered fetches are a small but far more meaningful slice.

Step 2: Which pages do live conversations fetch?

Filter to user-triggered agents and count by path:

grep -E 'ChatGPT-User|Claude-User|Perplexity-User' /var/log/nginx/access.log* \
  | awk '{print $7}' \
  | sed -E 's/\?.*$//' \
  | sort | uniq -c | sort -rn | head -20
Enter fullscreen mode Exit fullscreen mode

Over a few weeks of logs, this gives you a ranked list of the pages AI assistants open on behalf of users. Compare it with your Google top pages. The two lists are rarely the same. Docs pages, pricing pages and comparison pages are often overrepresented in AI fetches relative to their search traffic.

Also check the status codes:

grep -E 'ChatGPT-User|Claude-User|Perplexity-User' /var/log/nginx/access.log \
  | awk '{print $9}' | sort | uniq -c
Enter fullscreen mode Exit fullscreen mode

Any 403 or 429 in that output means a live conversation tried to read your page and you turned it away. That's usually a WAF or rate-limit rule that nobody set up with AI agents in mind.

Bar chart of pages fetched by ChatGPT-User, Claude-User and Perplexity-User, with two pages returning 403 and 429

Step 3: Make sure the bots are real

User agents are trivially spoofed, and scrapers love pretending to be GPTBot. Before you trust the numbers (or allowlist anything), verify the source IPs.

OpenAI and Perplexity publish JSON files of the IP ranges their bots use, linked from their bot documentation. Here's a verifier that loads those ranges and checks log IPs against them:

# verify_bots.py
import ipaddress, json, re, sys, urllib.request

# Published by each vendor. Confirm on their bot docs pages before relying on them.
RANGES_URLS = {
    "ChatGPT-User":    "https://openai.com/chatgpt-user.json",
    "OAI-SearchBot":   "https://openai.com/searchbot.json",
    "GPTBot":          "https://openai.com/gptbot.json",
    "Perplexity-User": "https://www.perplexity.ai/perplexity-user.json",
    "PerplexityBot":   "https://www.perplexity.ai/perplexitybot.json",
}

def load(url):
    data = json.load(urllib.request.urlopen(url, timeout=15))
    nets = []
    for p in data.get("prefixes", []):
        cidr = p.get("ipv4Prefix") or p.get("ipv6Prefix")
        if cidr:
            nets.append(ipaddress.ip_network(cidr))
    return nets

networks = {bot: load(u) for bot, u in RANGES_URLS.items() if u.startswith("http")}
line_re = re.compile(r'^(\S+) .*"([^"]*)"$')

stats = {}
for line in sys.stdin:
    m = line_re.match(line.rstrip())
    if not m:
        continue
    ip, ua = m.groups()
    for bot, nets in networks.items():
        if bot in ua:
            addr = ipaddress.ip_address(ip)
            ok = any(addr in n for n in nets)
            s = stats.setdefault(bot, [0, 0])
            s[0 if ok else 1] += 1

for bot, (good, bad) in stats.items():
    print(f"{bot:15} verified {good:6}  spoofed {bad:6}")
Enter fullscreen mode Exit fullscreen mode
cat /var/log/nginx/access.log | python verify_bots.py
Enter fullscreen mode Exit fullscreen mode

The ranges files use a prefixes list with ipv4Prefix/ipv6Prefix keys, the same format Google uses for Googlebot. If a vendor doesn't publish ranges, fall back to reverse DNS plus forward-confirm, where the vendor supports it.

Step 4: Log it properly going forward

grep works for a quick look. For ongoing tracking, tag AI traffic at the edge so it's queryable without regex archaeology:

map $http_user_agent $ai_bot {
    default                 "";
    ~*ChatGPT-User          "chatgpt-user";
    ~*Claude-User           "claude-user";
    ~*Perplexity-User       "perplexity-user";
    ~*OAI-SearchBot         "openai-search";
    ~*Claude-SearchBot      "claude-search";
    ~*PerplexityBot         "perplexity-search";
    ~*GPTBot                "openai-train";
    ~*ClaudeBot             "claude-train";
}

log_format ai '$time_iso8601\t$remote_addr\t$ai_bot\t$status\t$request_uri';

server {
    # ...
    access_log /var/log/nginx/ai.log ai if=$ai_bot;
}
Enter fullscreen mode Exit fullscreen mode

The if= parameter writes only requests where $ai_bot is non-empty, so ai.log stays small and tab-separated. Ship it to whatever you already use (Loki, BigQuery, a cron job that loads it into SQLite).

On Cloudflare or another CDN, your origin logs may miss requests the edge served from cache or blocked outright. Use the CDN's own logs or AI-bot analytics for the complete picture.

What logs can't tell you

Server logs show that an assistant read a page. They don't show:

  • what the user asked,
  • whether the answer cited you or used a competitor instead,
  • what the assistant said about you,
  • anything from answers built from training data or a search index without a live fetch.

That last one is a big gap. Many answers never trigger a live fetch at all. So logs are a strong supply-side signal ("these pages are being read") and need pairing with a demand-side one ("this is what the answers say").

That pairing is what citation tracking covers. Vista AI's AI Citation Tracking records the prompt, the full response and the cited source for each brand mention across the major AI engines, scores sentiment and flags inaccurate claims. Put the two side by side: a page that's fetched often in your logs but rarely cited in answers is a page the engines find and then pass over, which makes it a good first candidate for a rewrite.

Quick checklist

  • [ ] Count AI bot hits by type (training, search, user-triggered)
  • [ ] List the top 20 paths fetched by *-User agents
  • [ ] Check those fetches for 403/429 responses
  • [ ] Verify bot IPs against published ranges
  • [ ] Add an nginx map + dedicated AI log for ongoing tracking

Top comments (0)