In August I published a number on my studio's site: named unprompted in 14 of 18 blind answers across ChatGPT, Perplexity and Gemini.
It was wrong, and the way it was wrong is worth more than the number.
What I did
I wrote a set of questions and pasted them into one session per engine. Most never mention my studio by name. Two do, because I also wanted to check the engines had the facts right.
You can see it already. Those two questions name the company. Everything in that paste shares one context window. By the time an engine answered question one, the answer was already in front of it.
I was not measuring whether an engine finds my studio unprompted. I was measuring whether it can read.
How I caught it
I did not, at first. I re-ran it, got a near-perfect sweep, and my first reaction was that things had improved.
That reaction is the tell. The output also contained a giveaway: text from one question spliced into the answer to another, mid-sentence. An engine mangling the prompt into the response is not producing a clean read on anything.
The corrected method
Two sessions per engine, not one. Session A holds only the blind questions and never contains the company name. Session B holds the named fact-checks, in a separate fresh incognito session. Three engines, six sessions. Twelve blind questions, so 36 blind answers.
What it actually scored
Nine of 36. Twenty-five percent, against the 78 I had published.
The shape underneath matters more than the total. ChatGPT 9 of 12. Perplexity 0 of 12. Gemini 0 of 12.
It is not a 25% problem. It is a one-engine-in-three problem, and I would never have seen that from the contaminated run, because contamination had lifted all three engines to roughly the same place.
The part I did not expect
Both engines scoring zero describe the studio accurately the moment you ask them directly. They hold the facts. They never volunteer them.
So the problem is distribution, not knowledge and not the website. I checked: every source one of them cited when answering a vendor question was somebody else's roundup. Not one was my own site. It was reading published lists and repeating the names on them, rather than evaluating websites and ranking them.
No amount of schema on my site reaches that.
If you are tracking your own AI visibility
Split the sessions. Never let a question that names you share a context window with a question testing whether you are found.
And if your number goes up after a change, check the method before you celebrate. Mine went up 53 points because I broke the test.
I have published the corrected figure with every miss next to it, along with the twelve questions and the full protocol, at reidify.design/research/measuring-ai-visibility. A flattering number produced by a broken method is worth less than an unflattering one produced by a sound one.
Top comments (5)
The discrepancy in visibility scores highlights the importance of consistent measurement methodologies. A 78% to 25% drop indicates that the metrics used can significantly influence perceived performance. It's crucial to standardize how we assess visibility, especially in AI contexts, to ensure actionable insights.
it's wild how much the measurement method changes the actual data. makes me wonder how many "stats" we see online are just skewed by the metric used lol
ngl it's crazy how misleading these automated scores are. makes you wonder what else we're measuring wrong lol
This is also a good case for CI checks: route returns 200 only for eligible combinations, canonical matches the route, the sitemap contains no previews, and structured data matches the visible page.
man, the way these metrics are calculated is so wild. really makes you wonder if any of these ai visibility scores are actually legit tbh