At the baseline, Google AI Mode did most of the citing. Perplexity barely did any.
Thirty days later, when I replayed the same prompt panel, they had almost swapped roles.
- Google AI Mode: 29 citation observations fell to 11.
- Perplexity: 5 citation observations rose to 35.
The aggregate moved from 35 to 50 citation-positive observations. Read alone, that looks like a straightforward improvement.
It was not. It was a redistribution.
That distinction is the most useful result of this experiment: an AI visibility total can be accurate and still hide the change that matters.
What I repeated
I froze 25 non-branded French prompts about measuring visibility in AI-generated answers. Each prompt was tested three times across four visible product environments:
- ChatGPT;
- Claude;
- Perplexity;
- Google AI Mode.
That produced 300 planned runs per wave. Five planned runs in each wave were non-valid for documented reasons, leaving 295 valid observations at J0 and 295 at J+30. The J+30 archive also retains one supplemental calibration row outside the 300-run panel.
The corpus, wording, coding rules, planned run count, and comparison method remained fixed. A run recorded brand mention, linked citation, cited URL, total visible sources, technical validity, and the conditions needed to interpret the answer.
This is a panel for one brand, one French corpus, and two dates. It is not market share, a platform ranking, or evidence that an editorial change caused the movement.
The total hid a platform reversal
Here is the comparison that changed my interpretation:
| Platform | J0 citations | J+30 citations | J0 mentions | J+30 mentions |
|---|---|---|---|---|
| ChatGPT | 0/75 · 0% | 2/71 · 2.82% | 0 | 2 |
| Claude | 1/74 · 1.35% | 2/75 · 2.67% | 1 | 2 |
| Perplexity | 5/74 · 6.76% | 35/74 · 47.30% | 6 | 3 |
| Google AI Mode | 29/72 · 40.28% | 11/75 · 14.67% | 25 | 11 |
At J0, Google AI Mode accounted for 29 of the 35 citation-positive observations. At J+30, Perplexity accounted for 35 of 50.
A dashboard containing only the global citation rate would show this:
11.86% → 16.95% (+5.08 percentage points)
True, but incomplete. The observed citation counts did not rise uniformly; they shifted sharply across products.
That matters operationally. A team reacting only to the aggregate might congratulate itself, continue the same work, and miss a large loss in the environment that previously carried most of its citations.
Mentions and citations moved in opposite directions
The second surprise was the split between being named and being used as a source.
| Signal | J0 | J+30 | Change · 95% interval |
|---|---|---|---|
| Mention rate | 32/295 · 10.85% | 18/295 · 6.10% | −4.75 pp [−9.57; +0.34] |
| Citation rate | 35/295 · 11.86% | 50/295 · 16.95% | +5.08 pp [−1.45; +11.82] |
| Edikka share of visible sources | 36/1,752 · 2.05% | 51/1,974 · 2.58% | +0.53 pp [−0.43; +1.56] |
| Prompts with a mention | 14/25 · 56% | 7/25 · 28% | −28 pp [−52; −4] |
| Prompts with a citation | 14/25 · 56% | 15/25 · 60% | +4 pp [−20; +28] |
Perplexity makes the distinction concrete. It cited an Edikka URL in 35 observations at J+30 but named the brand in only 3.
If I had treated a citation as a mention, I would have erased that behavior. If I had treated either as a visit, I would have invented another one.
I now keep four layers separate:
- Mention: the answer names the brand.
- Citation: the interface exposes a brand URL as a source.
- Visit: an attributable referral reaches the site.
- Outcome: that visit contributes to a useful action.
The first two are answer-layer observations. The last two require analytics and business evidence. No arithmetic can recover the missing link between them.
Only one headline change had a 95% interval excluding zero
It is tempting to write “citations increased by 43%” or “AI visibility improved.” I do not think the study supports either headline.
The prompt-paired 95% interval for the citation-rate change runs from −1.45 to +11.82 percentage points. It includes zero. The same is true for the change in source share and citation coverage.
The clearest global movement was narrower: mention coverage fell from 14 of 25 prompts to 7 of 25, with an interval of −52 to −4 percentage points.
That does not make the other changes meaningless. It changes the language I use for them:
- “rose in this observed wave” is descriptive;
- “improved reliably” would be stronger than the evidence;
- “the platform distribution reversed” is visible in the platform-level counts;
- “the work caused the change” is not established at all.
Uncertainty is not a footnote added after the story. It determines which story can be told.
Five rules I would use in an AI visibility dashboard
After this replay, my minimum reporting model is:
1. Never publish a total without the platform table
The aggregate is useful for orientation. The platform breakdown shows where the movement occurred.
2. Keep mentions, citations, source share, and prompt coverage separate
They answer different questions. Combining them into a proprietary “authority score” makes the result easier to sell and harder to inspect.
3. Show numerators, denominators, and invalid runs
50/295 carries more information than 16.95%. Five non-valid planned runs remained in the audit trail and stayed out of the visibility denominator; they were not silently retried into success or converted into absence. The supplemental calibration row remained outside the panel.
4. Compare at prompt level
Three repetitions of one prompt are not three independent strategic topics. The comparison used 10,000 bootstrap replications clustered by prompt, rather than pretending every run was independent.
5. Preserve the dated evidence
Interfaces, retrieval systems, indexes, and source panels change. A result without its date, conditions, coding rules, and evidence is a screenshot, not a time series.
The quality check also changed how I report confidence
The J+30 documentary review covered 30 rows in a separate AI-assisted pass. Twenty-nine could be fully recoded; one archived response was empty. Citation presence agreed on all 30 rows. Mention presence agreed on 27 of the 29 comparable rows, but Cohen's κ was 0: positive mentions were rare, and there was no positive-positive agreement in that sample. Raw agreement alone would overstate inter-coder reliability. The ordinal quality and factual-fidelity scores varied more substantially.
This was not an independent second human review. I am publishing that limitation because “double-checked” would imply more than happened. The two mention disagreements remain in the quality-control table instead of being forced into consensus.
The raw third-party answer text and screenshots are not redistributed in the public package. The structured observations, methodology, limitations, manifest, and checksums are public; the evidence archive remains controlled.
What I would do next — and what I would not do
I would not rewrite pages to chase a one-month spike in Perplexity or a one-month decline in Google AI Mode.
I would:
- inspect the prompt-platform pairs that changed repeatedly;
- check whether the cited URL is the expected page;
- compare source fidelity, not only source presence;
- retain a small weekly sentinel panel for breakage;
- replay the complete frozen corpus monthly;
- connect answer-layer citations to referrals and outcomes only when attribution exists.
The next useful finding is not another blended score. It is whether the reversal persists.
The complete J+30 results, accessible tables, and limitations are public. The MIA-FR protocol and the frozen J0 dataset have their own versioned records.
If your dashboard had shown 35 → 50 citation-positive observations (and 36 → 51 Edikka sources), would you have called that improvement before looking at the platform table?
AI-assistance disclosure: I used AI assistance to challenge the structure and edit the English of this DEV edition. The study design, collection protocol, interpretation, limitations, and publication decision remain my responsibility. The J+30 documentary quality check was also AI-assisted and is not presented as an independent human review.
Top comments (0)