We thought we were measuring AI recommendations.
We weren't. At least, not completely.
For most of our previous research, we analyzed brands that AI had already recommended.
That created a statistical problem.
We were studying the population that had already passed the gate.
So we changed the question: What separates the stores AI ever recommends from the stores it never recommends?
We tested the full population, and the result changed how we think about AI recommendation systems.
We analyzed 60,924 e-commerce stores.
- 599 were recommended at least once
- 60,325 were never recommended
- Recommended stores averaged 55.4 on our AI Commerce Score
- Never-recommended stores averaged 56.0
- Intent was the one score component that clearly separated the two groups
- Intent was roughly 39% higher among recommended stores
- But once a store was already in the recommendation set, Intent explained only 1.2% of recommendation-frequency variation
The key finding:
A factor can help a store become a candidate without helping it win once it becomes a candidate.
We call this distinction Candidacy vs Selection.
The statistical problem: range restriction
Our previous recommendation research focused on brands that the model already recommended.
That lets you ask: Among brands already being recommended, does store quality correlate with recommendation frequency?
But it does not answer: What separates stores that are ever recommended from stores that are never recommended?
Those are different populations.
If a variable only matters at the recommendation gate, restricting the sample to already-recommended brands can hide the effect.
So for this study, we removed that restriction.
The dataset
We tested:
- 60,924 total stores
- 599 recommended at least once
- 60,325 never recommended
We compared the total AI Commerce Score and all seven score components across the full population.
We stopped looking only at the winners and tested the full population.
Result #1: Recommended stores were not better overall
If AI recommendation were simply a reflection of store quality, we would expect the recommended population to have a higher average score.
It didn't.
| Group | Average AI Commerce Score |
|---|---|
| Recommended | 55.4 |
| Never recommended | 56.0 |
The recommended group actually scored slightly lower.
So overall store quality did not separate the two populations.
That was the first signal that we were dealing with something more complicated than: Better store → more AI recommendations
Result #2: One score component stood out
We then compared the seven score components.
Intent was the one score component that consistently separated stores that were ever recommended from stores that were never recommended.
The values were:
| Component | Recommended | Never recommended |
|---|---|---|
| Intent | 5.13 | 3.70 |
| Visual | 7.43 | 7.22 |
| Schema | 3.83 | 3.53 |
| Technical | 11.27 | 12.13 |
| Trust | 10.42 | 11.06 |
| Price | 8.25 | 8.94 |
| Brand | 3.07 | 3.14 |
Intent showed roughly a 39% gap.
That made it the obvious candidate for further testing.
Does the Intent signal hold up?
A single correlation can be misleading.
So we checked several possible explanations.
1. Category mix
The Intent gap appeared across 9 of 9 niches.
The gap ranged from +0.78 to +1.59 points.
That makes a simple category-composition explanation less likely.
2. Circularity
We checked Intent against the other score components.
Its correlations with the other factors ranged from approximately:
0.20 to 0.29
So Intent was not simply duplicating another score component.
3. Brand ownership
The gap was:
1.40 vs 1.53
for the relevant brand-owned comparison.
The difference was nearly identical.
4. Selection after candidacy
This was the critical test.
Once we restricted the sample to the 599 stores that had already been recommended, Intent had almost no relationship with recommendation frequency.
r = 0.112
R² = 1.2%
That's the key result.
Candidacy ≠ Selection
The same factor can matter at one stage and become almost irrelevant at another.
We can represent the process like this:
60,924 stores
|
v
CANDIDACY
|
v
599 ever recommended
|
v
SELECTION
|
v
Who actually wins?
Getting into the recommendation set and winning inside that set are different problems.
Intent appears to help separate: stores that ever enter the set from
stores that never enter the set.
But once the store is already in the set, Intent predicts almost nothing about how often it wins.
Two anonymized brands showed the same split
We then reused controlled experiments from earlier studies in the series.
These weren't observational correlations.
They involved controlled manipulations of context.
Two anonymized brands showed almost opposite bottlenecks.
Brand J: a candidacy problem
Brand J had an overall recommendation rate of: 1.2%
across 400 real buyer-question observations.
But under supportive exposure conditions:
100% winner rate
Then we injected a single verified fact.
Its mention rate moved from: 1.25% → 97.5%
Nothing else about the brand changed.
This strongly suggests that Brand J's bottleneck was getting into the room.
Brand H: a selection problem
Brand H could also reach full candidacy under supportive conditions.
But in a multi-turn conversation, it survived to the final pick only:
5%
in one tested condition.
Across Hidden Context tests, its winner rate ranged from: 0% → 53%
So getting Brand H into the recommendation set didn't solve the problem.
Its bottleneck appeared later.
Something about the specific matchup still determined the outcome.
Why this matters for AI optimization
This creates an important distinction.
A brand can have a problem with:
Discovery
AI cannot find the brand.
Understanding
AI cannot correctly interpret what the brand sells.
Candidacy
AI knows the brand exists but rarely puts it into the consideration set.
Selection
AI considers the brand but repeatedly chooses a competitor.
These problems can look identical from the outside:
"AI doesn't recommend me."
But they may require completely different interventions.
Correlation still cannot tell us causality
There is an important limitation here.
We can show that Intent separates recommended and never-recommended stores.
We cannot use this observational result alone to say:
Increasing Intent will cause AI to recommend a store.
There are alternative explanations.
For example, sharply positioned brands might be easier for AI to identify.
Or successful brands might simply have invested more heavily in positioning.
Both could produce the same observed correlation.
That's why controlled experimentation matters.
If we want to understand causality, we need to:
Change one variable
Hold other variables constant
Run the same test again
Measure the change in AI behavior
That's a very different research problem from collecting correlations across thousands of stores.
A platform confound worth mentioning
One result initially looked dramatic.
Magento stores appeared in the recommended group at: 3.70%
Shopify stores: 0.79%
That's approximately a 4.7x difference.
But Magento stores in this dataset skew older and larger.
So we don't interpret Magento itself as the cause.
The platform is better treated as a marker for other underlying characteristics.
This is another reason to be careful when turning observational correlations into optimization advice.
What we are not reporting
An earlier pass also compared raw review counts and average price between the two groups.
We removed both from the final study.
A later data-quality check found:
Review count was populated for only 9.4% of the 66,085 scanned stores
Average price was populated for 66.7%
That meant those comparisons were largely comparing missing fields rather than real values.
The score components used in this study were computed for every scanned store, so those are the comparisons we report.
The bigger idea
This study changed one assumption for us.
We used to think about AI recommendation as one problem.
Now we think it may be at least two.
Candidacy asks:
Can the brand get into the recommendation set?
Selection asks:
Once it's there, why does it beat the alternatives?
And the data suggests that the signals involved may be different.
That's important because most AI visibility measurement still collapses these stages together.
But:
Being visible is not the same as being considered.
And:
Being considered is not the same as being chosen.
What we're testing next
This study gives us a stronger reason to move toward controlled experiments.
We don't just want to know which signals correlate with recommendations.
We want to know which changes actually move the decision.
Change a rating.
Change a specification.
Change a price.
Add or remove a verified fact.
Change the available context.
Then observe what happens.
That's where we think the next generation of AI Commerce Intelligence will come from.
Not more signals for the sake of more signals.
But better evidence about which signals actually matter.
The takeaway
The most important finding isn't: "Intent is the ranking factor."
We don't have evidence for that.
The evidence supports something narrower: Intent separates stores that ever enter the recommendation set from stores that never do.
And:
Once a store is already in the recommendation set, Intent explains almost nothing about which store wins more often.
That's the distinction.
Candidacy gets you into the game.
Selection wins the game.
And we're trying to understand what happens between those two stages.


Top comments (0)