In July I wrote that my open-source LLM visibility checker tested Claude only and that multi-model support was "planned but not yet implemented." That's the kind of line that's easy to write and easy to forget.
I didn't forget. I just needed a reason to finish it.
The reason showed up when SearchApi launched endpoints for ChatGPT and Gemini, on top of the Perplexity and Bing Copilot endpoints they already had, and their growth engineer Sam Gale offered API credits to test them. Around the same time, Sam posted his own tool in SearchApi's Discord: ai-visibility-tracker, a dashboard that scores brand mentions across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Mode. Everything except Claude.
That's not a coincidence. His tool was missing the one engine mine already had. Mine was missing the four his already had. So instead of building a second comparison tool from scratch, I extended the one I'd already shipped.
What changed
llm_visibility.py used to do one thing: send a query to Claude Haiku, regex-match your domain against the response, print a score. Simple, and honest about its limits. Claude only knows what was in its training data, so anything published recently was invisible to it by design.
Now it does five things. Claude still runs the same way: direct API call, training-data knowledge, same regex match. The other four go through a new client, searchapi_client.py, that hits SearchApi's /api/v1/search endpoint with engine=chatgpt, engine=gemini, engine=perplexity, or engine=bing_copilot. All four share one endpoint shape and return a reference_links array (title, link, source) that I check against your domain the same way I check Claude's response text.
ChatGPT only returns cited sources if you pass web_search=true. Without it, you get an answer with no citations at all. Caught it in SearchApi's docs before I ran anything, so it's been in the client from the first version, but easy to miss if you're skimming past the optional parameters.
One more gotcha, unrelated to SearchApi: my existing serp_features.py module already talked to an API called SerpApi, for classic Google SERP feature detection. SearchApi and SerpApi are two different companies with confusingly similar names. I kept the two clients in separate files with separate env vars (SERPAPI_KEY vs SEARCHAPI_KEY) on purpose, and I'd recommend anyone doing this kind of work do the same before they mix up a bill.
The mix-up isn't hypothetical, either. When I asked Perplexity about open-source SEO agent tools during testing, it cited github.com/serpapi/seo-research-agent, SerpApi's own official project, built the same way as mine (LLM plus search API). Two similarly-named companies, two similarly-shaped tools, and an AI engine happily citing both without distinguishing them.
What I found
I ran it against two of my own domains. The queries came from real Google Search Console exports, not ones I picked to make a point.
dannwaneri.com, 10 queries (mostly "hire freelance [library] developer" searches plus my own name):
| Engine | Score |
|---|---|
| Claude | 0/10 |
| ChatGPT | 0/10 |
| Gemini | 2/10 (20%) |
| Perplexity | 0/10 |
| Copilot | 0/9 — 1 error |
Gemini was the only engine that cited me at all. It got "daniel nwaneri" right, correctly linking my homepage. It also cited me for "hire freelance scipy developer," except the page it pulled was /hire-python-developer/, not a scipy-specific page. Close, not exact. That gap is the difference between an engine understanding your content and one pattern-matching on adjacency.
naija-vpn.com, 10 queries (real buyer-intent searches: Twitch payments, Fiverr payouts, dollar accounts for Nigerian freelancers):
| Engine | Score |
|---|---|
| Claude | 0/10 |
| ChatGPT | 1/10 (10%) |
| Gemini | 0/10 |
| Perplexity | 0/10 |
| Copilot | 0/9 — 1 error |
Here it flipped. ChatGPT was the only one that cited me, correctly pulling /twitch-payments-nigeria for "how to receive money from twitch in nigeria." Gemini, the engine that carried dannwaneri.com, found nothing on this domain at all.
Two domains. Two different engines doing the only citing. Zero overlap between them. Claude, in both cases, found nothing. That tracks: neither domain existed in a form Claude's training data would have caught.
The finding
If I'd only tested Claude, like the July version of this tool did, I'd have told you both domains were invisible to AI. If I'd only tested Gemini, I'd have said dannwaneri.com was fine and naija-vpn.com wasn't. Backwards, if you'd only checked ChatGPT.
Wrong story either way. Each one only saw a fifth of the picture. The only way to know your actual AI visibility is to check all of them, because you can't predict which engine will happen to cite you this month.
That's the whole argument for a tool like this existing as multi-engine from the start, and it's the same argument for using one API across five engines instead of scraping each separately.
Two failure modes, for anyone building on this
Copilot errored on both test runs: a 503 ("unable to generate an answer for this query") on one, a request timeout on the other. Different failures, same engine, two separate runs. If you're building on top of SearchApi's Copilot endpoint, plan for it to occasionally just not answer, and don't let one failed query kill the whole batch. Mine logs the error against that query and keeps going.
Where to look
- The code: github.com/dannwaneri/seo-agent, specifically
modules/searchapi_client.pyandmodules/llm_visibility.py - Sam's tool, which covers the engine mine doesn't (Google AI Mode): SamJale/ai-visibility-tracker
- The engines this runs against: ChatGPT, Gemini, Perplexity, Bing Copilot
- SearchApi itself, who provided the credits this was tested with: searchapi.io
*This piece was produced as part of SearchApi's Developer Ambassador program. They provided API credits; I built and tested the integration myself.
Top comments (3)
The zero overlap is useful, but with 10 queries the bigger design risk is conflating engine coverage with run stability. In a 284-brand Korean DTC scan we ran across 50 AI shopping questions per brand, 65.5% of brands had zero appearances and the mean was only 0.648 out of 50. That sparse distribution made the denominator and repeatability as important as the engine split. For Copilot, I’d report 0/9 answered plus 1/10 unavailable rather than only 0/9, then rerun the identical prompts two or three times. Have you checked whether the Gemini/ChatGPT ownership flip persists across repeated runs?
Bulti, no. Single run per query, no repeats. Didn't check whether the split holds.
10 queries once each tells you almost nothing about whether Gemini citing dannwaneri.com is stable or a coin flip that happened to land there that day. Your 0.648-of-50 mean is a good reminder of how sparse this signal actually is at scale. At n=10 with one pass, I'm probably closer to noise than to a measurement.
The 0/9-plus-1-unavailable framing is right too. Folding the Copilot error into the denominator hides that it's a different failure mode than "answered, didn't cite you."
At 284 brands times 50 questions times multiple runs, what's your actual query volume before rate limits become the bottleneck?
that is why worlds need something which can prove the ai's (ai agents, ai models) responses,