If you grep your access logs for ChatGPT-User to see what ChatGPT can't find on your site, most of what comes back won't be ChatGPT.
We looked at one day of bot traffic, Wednesday 7 October 2026, across the small business websites running LovedByAI, a generative engine optimization platform for small-business WordPress sites. Requests carrying the ChatGPT-User user agent got a 404 1,022 times across 111 sites, and 703 of those (69%) were exploit probes, not ChatGPT. A vulnerability scanner is wearing the name.
Meanwhile the 404s that do matter, real people ChatGPT sent to pages that don't exist, never show up under that user agent at all. This post covers how to tell the two apart, with a script you can run on your own logs. The full study behind the numbers is ChatGPT broken links: why visitors land on 404s.
What is the ChatGPT-User user agent?
ChatGPT-User is the user agent OpenAI uses when ChatGPT fetches a page because a person asked it something mid-conversation. It's one of three OpenAI crawlers, and they do different jobs:
| User agent | What it does | OpenAI's published IP ranges |
|---|---|---|
GPTBot |
Collects training data | https://openai.com/gptbot.json |
OAI-SearchBot |
Builds ChatGPT's search index | https://openai.com/searchbot.json |
ChatGPT-User |
Fetches a page for a user, on demand | https://openai.com/chatgpt-user.json |
A ChatGPT-User request looks like this in a combined-format log (a demo line, using an address from OpenAI's list):
104.208.184.197 - - [07/Oct/2026:10:14:02 +0000] "GET /services/ HTTP/1.1" 200 512 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"
The user agent string is trivially easy to copy. The IP ranges are not, which is why OpenAI publishes them.
How many ChatGPT-User 404s are fake?
On 7 October 2026, 5.8% of the 823,928 bot requests to small business websites in our logs got a 404. For user agents claiming to be an AI assistant fetching a page for a user (ChatGPT-User, Claude-User, Perplexity-User and the like), the 404 rate was 43.7% (4,597 of 10,528). For AI search crawlers it was 10.0%, search engines 7.2% and AI training crawlers 4.4%.
The gap is the scanner. Of the 1,022 ChatGPT-User 404s that day, 703 matched a strict exploit pattern:
-
/cgi-bin/php-cgiwith anallow_url_includeargument, the PHP-CGI argument injection tracked as CVE-2024-4577 -
.envfiles -
/proc/selfreads and path traversal - Vite dev server
/@fs/file reads - Kubernetes service account tokens
Most of the remainder asked for config and debug files (runtime-config.js, /debug/pprof) that no small business website serves. The same probes also arrive dressed as GrokBot, MistralAI-User and DuckAssistBot. Some examples, as they appear in a log (demo lines, addresses from the documentation ranges):
203.0.113.24 - - [07/Oct/2026:10:14:02 +0000] "GET /cgi-bin/php-cgi.exe?%ADd+allow_url_include%3d1+%ADd+auto_prepend_file%3dphp://input HTTP/1.1" 404 512 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"
203.0.113.24 - - [07/Oct/2026:10:14:02 +0000] "GET /.env HTTP/1.1" 404 512 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"
198.51.100.7 - - [07/Oct/2026:10:14:02 +0000] "GET /@fs/etc/passwd?raw HTTP/1.1" 404 512 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"
198.51.100.7 - - [07/Oct/2026:10:14:02 +0000] "GET /var/run/secrets/kubernetes.io/serviceaccount/token HTTP/1.1" 404 512 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot"
On a WordPress site these should all 404, which is the correct answer. The problem is only in the reporting: count them as ChatGPT and your "ChatGPT can't find our pages" number is mostly noise.
How do you verify a ChatGPT-User request?
Check the client IP against OpenAI's published ranges. This script does that, separates real ChatGPT-User fetches from fakes, and then does the more useful job: it finds the human visits ChatGPT sent that landed on a 404.
#!/usr/bin/env python3
"""Split an access log's ChatGPT traffic into what's real and what isn't.
python3 chatgpt_404s.py access.log
Reads the combined log format (nginx and Apache default) and downloads
OpenAI's published ChatGPT-User IP ranges. Standard library only.
"""
import ipaddress
import json
import re
import sys
import urllib.request
from collections import Counter
from urllib.parse import parse_qs, urlsplit
RANGES_URL = "https://openai.com/chatgpt-user.json"
LINE = re.compile(
r'(?P<ip>\S+) \S+ \S+ \[[^\]]+\] "(?P<method>\S+) (?P<path>\S+)[^"]*" '
r'(?P<status>\d{3}) \S+ "(?P<referer>[^"]*)" "(?P<ua>[^"]*)"'
)
# Paths no small business website serves. These are what the scanner wearing
# the ChatGPT-User name asked for in our logs.
PROBE = re.compile(
r"php-cgi|allow_url_include|/\.env|/proc/self|/@fs/|serviceaccount/token"
r"|\.\./|%2e%2e|/debug/pprof|runtime-config\.js",
re.IGNORECASE,
)
def load_ranges(url=RANGES_URL):
req = urllib.request.Request(url, headers={"User-Agent": "chatgpt-404s/1.0"})
with urllib.request.urlopen(req, timeout=20) as resp:
prefixes = json.load(resp)["prefixes"]
return [ipaddress.ip_network(p.get("ipv4Prefix") or p["ipv6Prefix"]) for p in prefixes]
def from_openai(ip, nets):
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return False
return any(addr in net for net in nets if net.version == addr.version)
def sent_by_chatgpt(path, referer):
"""A person clicking a link in a ChatGPT answer, not a bot."""
sources = parse_qs(urlsplit(path).query).get("utm_source", [])
return "chatgpt.com" in sources or urlsplit(referer).hostname in ("chatgpt.com", "chat.openai.com")
def main(log_path):
nets = load_ranges()
bots = Counter()
bot_404s = Counter()
visitor_404s = Counter()
visitors = 0
with open(log_path, encoding="utf-8", errors="replace") as log:
for raw in log:
m = LINE.match(raw)
if not m:
continue
page = urlsplit(m["path"]).path
if "ChatGPT-User" in m["ua"]:
if not from_openai(m["ip"], nets):
bots["fake (not from OpenAI's IP ranges)"] += 1
if PROBE.search(m["path"]):
bots[" of which exploit probes"] += 1
else:
bots["real ChatGPT-User"] += 1
if m["status"] == "404":
bot_404s[page] += 1
elif sent_by_chatgpt(m["path"], m["referer"]):
visitors += 1
if m["status"] == "404":
visitor_404s[page] += 1
print("ChatGPT-User requests")
for label, n in bots.items():
print(f" {n:>5} {label}")
print("\nPages real ChatGPT-User fetched that returned 404")
for page, n in bot_404s.most_common(20):
print(f" {n:>5} {page}")
lost = sum(visitor_404s.values())
print(f"\nVisitors ChatGPT sent you: {visitors}, landed on a 404: {lost}")
for page, n in visitor_404s.most_common(20):
print(f" {n:>5} {page}")
if __name__ == "__main__":
main(sys.argv[1])
Standard library only, Python 3.8 or later. Run it against a log:
python3 chatgpt_404s.py /var/log/nginx/access.log
On a small demo log (invented lines: one real ChatGPT-User address from OpenAI's list, two fake ones from the documentation ranges, and a handful of visitors), it prints:
ChatGPT-User requests
2 real ChatGPT-User
6 fake (not from OpenAI's IP ranges)
5 of which exploit probes
Pages real ChatGPT-User fetched that returned 404
1 /experiences/
Visitors ChatGPT sent you: 5, landed on a 404: 3
2 /contact-us/
1 /product/discontinued-perfume/
Three things to know before you trust it on production logs:
-
Behind Cloudflare or another proxy, log the real client IP. Otherwise every request comes from the proxy's addresses and everything looks fake. On nginx that means
real_ip_header CF-Connecting-IPwith Cloudflare's ranges inset_real_ip_from. -
Fetch the ranges fresh. OpenAI updates
chatgpt-user.json; the copy we fetched on 10 October 2026 was dated 7 October. -
Gzipped, rotated logs need decompressing first (
zcat access.log.2.gz > access.log.2).
How do you find the 404s real ChatGPT visitors hit?
Look for the click, not the bot. When a person clicks a link in a ChatGPT answer, their browser makes the request, with a normal browser user agent, and ChatGPT adds utm_source=chatgpt.com to the URL (the referrer is often https://chatgpt.com/ too). That's what sent_by_chatgpt() matches.
Those are the ones that cost you something. Across 323 small business websites that received a ChatGPT visit between 30 September and 9 October 2026, 36 of 4,053 ChatGPT visits (about 1 in 113) landed on a 404. When we rechecked each address on 10 October against the live site and the Wayback Machine's CDX index, they split like this:
The four "garbled" ones are worth a developer's attention: the page existed, but the link ChatGPT wrote had its percent-encoding doubled (%25 where a % belonged), one byte changed, or bytes missing from the middle so it was no longer valid UTF-8. Long, encoded slugs are the ones a model breaks, because it writes a URL one token at a time instead of copying it.
What should each kind of ChatGPT 404 return?
| What the 404 is | What to serve |
|---|---|
| Page moved (old URL structure, renamed slug) | 301 to the new address |
| Product, listing or job removed | 301 to the replacement or the category, or keep a "no longer available" page with alternatives. Not the homepage: Google treats a mass redirect to the homepage as a soft 404 |
| Address ChatGPT invented that keeps getting visits | Build a page there, or 301 to the closest real one |
| Garbled version of a real URL | 301 the broken variant to the real page |
| Spam or test pages that should stay gone | 410 Gone |
Exploit probes (php-cgi, .env, /@fs/) |
Leave the 404, or block at the WAF |
A threshold that keeps this manageable: redirect a dead address once it has had more than two or three human visits and you have a live page that answers the same need, or immediately if other sites link to it.
The same fixes, walked through in under five minutes:
What we couldn't prove
- Our 69% is by path, not by IP. We classified the 7 October probes by what they requested; we didn't check each one's address against OpenAI's ranges. The script above is what we recommend, not how the study counted.
- Ten days is a first read. 36 dead-page visits shows the pattern, not a precise rate.
- CDN redirects are invisible to us. Redirects done by Cloudflare or the web server never reach WordPress, so a site that fixed things there looks clean.
For the owner's side of this, without the code, our Medium piece walks through it: Dead links from ChatGPT: it's still sending customers to pages you deleted.
LovedByAI's plugin records the status code of every AI referral visit on WordPress sites, which is where the 4,053 visits above come from. If you're not on WordPress, the script gets you the same list from raw logs.
Original study and full dataset: ChatGPT broken links: why visitors land on 404s

Top comments (1)
tr.ee/dev-to