When I ran the first AI visibility test for an established software brand, the result looked almost perfect.
The brand appeared in every unbranded buyer prompt.
At first glance, the conclusion seemed obvious:
The brand had 100% AI visibility.
But the result looked too clean.
So I repeated the same prompts.
The numbers changed.
Across repeated runs, the average mention rate fell to 65%, the recommendation rate fell to 59%, and the brand was consistently recommended in only 11 of 18 unbranded buyer prompts.
The more honest result was not 100%.
It was:
61% stable recommendation coverage.
That experiment changed how I think AI visibility should be measured.
A single AI answer can be useful.
It is not a ranking.
Why AI answers are unstable
Traditional search gives marketers a relatively familiar object to measure.
A page has a position. That position may change, but there is still a search-results page that can be checked repeatedly.
AI answers are different.
The response may depend on:
- the wording of the question;
- the model or assistant being used;
- whether web search is available;
- the sources retrieved during that request;
- the date and location;
- previous conversation context;
- normal model variation.
Ask the same question twice and you may get:
- a different order of products;
- a different top recommendation;
- new competitors;
- fewer products;
- different sources;
- or no citations at all.
This does not make AI visibility impossible to measure.
It means the unit of measurement cannot be a single screenshot.
Buyer prompts are not keywords
Another mistake is treating every question about a category as equivalent.
A buyer looking for code-review software might ask:
What are the best AI code review tools?
But they might also ask:
How can we catch bugs before merging a pull request?
Or:
What is an affordable PR review tool for a small engineering team?
Or:
What are the alternatives to CodeRabbit?
These prompts represent different stages of awareness.
Category prompts
The buyer already understands the product category.
Problem prompts
The buyer has a problem but may not know which type of product solves it.
Segment prompts
The buyer adds constraints such as company size, budget, workflow, or technical stack.
Comparison prompts
The buyer is already evaluating known products.
A brand can perform well in category prompts while remaining invisible in problem-oriented discovery.
That distinction matters more than a single overall score.
Branded and unbranded visibility are different
Consider these two prompts:
Best AI code review tools
and:
Is CodeRabbit worth it?
The second question already contains the brand.
A mention is expected.
That result tells us how an AI system describes a company when the buyer already knows it.
It does not tell us whether the company will be discovered by a buyer who has never heard of it.
This is why branded and unbranded prompts should be reported separately.
Branded visibility helps reveal:
- reputation;
- objections;
- comparisons;
- perceived strengths;
- perceived weaknesses.
Unbranded visibility helps reveal:
- category discovery;
- problem discovery;
- new customer acquisition;
- whether a brand enters the consideration set.
For most growth teams, unbranded visibility is the more valuable test.
Mentioned does not mean recommended
AI visibility tools often combine several outcomes into one number.
But there is a major difference between these statements:
CodeRabbit is one of several available tools.
and:
CodeRabbit is the best option for a small team that wants fast pull-request feedback.
The first is a mention.
The second is a recommendation.
There is another distinction:
Greptile may be a better option for teams that need deeper repository context.
In this answer, CodeRabbit may still be discussed, but another brand wins the recommendation.
A useful AI visibility analysis should therefore separate:
- missing;
- mentioned;
- recommended;
- top recommendation;
- branded discovery;
- unbranded discovery;
- stable and unstable results.
Without these distinctions, a brand may appear highly visible while consistently losing the actual buyer recommendation.
What repeated runs revealed
For the pilot, I tested 24 buyer prompts related to AI code-review products.
The prompts covered four groups:
- category discovery;
- segment and constraint;
- problem-oriented discovery;
- comparison and alternatives.
Each prompt was repeated three times under the same documented conditions.
The first single-run result suggested near-perfect visibility.
The repeated result was more nuanced:
- average unbranded mention share: 65%;
- average unbranded recommendation share: 59%;
- consistently recommended in at least two of three runs: 11 of 18 unbranded prompts;
- stable recommendation coverage: 61%.
The brand was clearly visible.
But its visibility was not universal or perfectly stable.
That is a much more useful conclusion than “100% visibility.”
Stable coverage is more useful than a screenshot
Imagine two brands.
Brand A
It appears in 15 of 20 prompts during one run.
When the prompts are repeated, only six recommendations remain consistent.
Brand B
It appears in 11 of 20 prompts during one run.
Ten of those recommendations remain consistent across repeated runs.
A single-run report would favour Brand A.
A consistency-first report would favour Brand B.
For a founder or marketer, Brand B probably has the stronger position.
This is why the metric I now care about most is:
Stable recommendation coverage: the percentage of buyer prompts where a brand is recommended consistently across repeated runs.
It does not remove uncertainty.
It makes the uncertainty visible.
Different AI engines represent different discovery environments
There is no single universal AI search result.
Different AI assistants can return different answers to the same buyer prompt.
They may rely on different:
- models;
- search systems;
- sources;
- ranking signals;
- answer formats;
- citation behaviour.
A company might be strongly recommended by one engine and rarely mentioned by another.
This gives AI visibility at least three dimensions.
Prompt coverage
In how many buyer situations does the brand appear?
Repeatability
Does the same engine produce a similar result across repeated runs?
Cross-engine agreement
Do different AI systems recommend the same company?
These dimensions should not be hidden inside one unexplained score.
A better report shows where the systems agree, where they disagree, and where a result is volatile.
Citations reveal why competitors may be winning
Brand mentions show what appeared in an answer.
Citations can reveal which sources influenced the answer.
A competitor may be repeatedly supported by:
- documentation;
- independent comparison pages;
- review platforms;
- GitHub repositories;
- industry publications;
- Reddit discussions;
- customer stories;
- category roundups.
This can expose a source gap.
The problem may not be that an AI model dislikes a product.
The product may simply be missing from the sources repeatedly used to explain the category.
That leads to a more practical question:
Where does AI find convincing information about competitors, and where is our brand absent?
The answer can suggest concrete work:
- create a clearer use-case page;
- publish a comparison page;
- improve documentation;
- add customer evidence;
- answer problem-oriented buyer questions;
- earn inclusion in relevant third-party sources;
- repeat the test after the changes.
AI visibility should be treated as an experiment
The wrong workflow is:
Run one prompt → get a score → celebrate or panic.
A better workflow is:
- Define a representative set of buyer prompts.
- Separate branded and unbranded discovery.
- Record a baseline.
- Repeat prompts to identify unstable results.
- Compare multiple AI engines.
- Find consistently lost buyer intents.
- Review which sources support competitor recommendations.
- Make one specific change.
- Repeat the same test.
- Measure directional movement.
This is closer to experimentation than traditional rank tracking.
There is no guarantee that a brand will hold a fixed position inside an AI-generated answer.
But it is still possible to observe whether visibility becomes:
- broader;
- more consistent;
- stronger across engines;
- better supported by relevant sources.
The most valuable result is often a lost prompt
A company may already appear for:
Best AI code review tools
But remain absent for:
How can a growing engineering team reduce manual review workload?
The second prompt may represent a larger opportunity.
The buyer has a problem but has not yet chosen a category or vendor.
When a brand repeatedly disappears from these prompts, the result can lead to a useful action:
- publish content around the problem;
- explain the workflow before introducing the product;
- create material for a specific team size;
- clarify how the product differs from alternatives;
- improve external evidence.
This is more actionable than being told that an “AI visibility score” increased from 42 to 47.
What teams should measure
A useful AI visibility report should answer:
- Where are we mentioned without the buyer naming us?
- Where are we genuinely recommended?
- Where are we the first recommendation?
- Which results remain stable across repeated runs?
- Which engines favour our competitors?
- Which buyer intents do we consistently lose?
- Which sources are repeatedly cited?
- What change should we test next?
The goal is not to manufacture a permanent AI ranking.
The goal is to replace random screenshots with repeatable evidence.
What this approach still cannot prove
Repeated testing makes the result more defensible, but it does not eliminate uncertainty.
A visibility snapshot cannot prove that:
- every user will receive the same answer;
- an AI assistant will keep recommending the same brand;
- citations directly caused a recommendation;
- higher visibility will automatically produce revenue;
- a content change will improve every engine equally.
It is still a sampled observation.
The purpose of repetition is not to create false certainty.
It is to identify which results appear stable enough to investigate and which ones may have been one-off variations.
Final thought
For years, marketers asked:
Where do we rank in Google?
The equivalent AI-search question should not be:
What is our ChatGPT rank?
A more useful question is:
Across the buyer questions that matter, how consistently do AI systems recommend us—and where do competitors win instead?
One AI answer cannot answer that.
A representative prompt set, repeated observations, multiple engines, and prompt-level evidence can get much closer.
The public research snapshot and future updates are available at BuyerPrompt.
Disclosure: I used an AI writing assistant to help structure and edit this article. The experiment, results, interpretations, and final review are my own.

Top comments (0)