I started this test with what I thought would be the boring part of a luxury-brand audit: can an AI answer point a buyer to the right official China channel?
The more interesting result appeared when I compared that task with a buying decision.
In a small exploratory pre-wave, Piaget appeared in all four answers about jewelry brands with verifiable official China channels. It appeared in none of the four answers recommending high-end brands for wedding jewelry.
That does not make Piaget “invisible,” and it certainly does not establish a market ranking. Four answer opportunities per task are too small for either claim.
It does expose a measurement problem that matters well beyond jewelry: finding an entity, recommending a brand and routing a buyer to a correct official channel are different jobs. If we report them as one visibility score, the number may be tidy while the diagnosis is wrong.
The deliberately small test
I fixed eight target brands before collection:
- Cartier
- Tiffany & Co.
- Bvlgari
- Van Cleef & Arpels
- Chaumet
- Boucheron
- Piaget
- De Beers Jewellers
I used three neutral Chinese buyer questions:
- Which high-end brands are worth considering for wedding jewelry, and what should a buyer compare?
- What is the difference between a brand boutique and a daigou for a high-value jewelry purchase?
- Which international jewelry brands have verifiable official websites or boutique channels in mainland China?
The valid collection contained two retrieval-off API surfaces, DeepSeek and Doubao, with two answers per surface and question. That produced 12 raw answers.
The important implementation choice was to collect each answer once. I did not repeat the same unbranded question eight times, once per target brand. After collection, each stored answer was evaluated against the frozen eight-brand registry, producing 96 answer-brand cells.
The distinction matters because provider calls and scoring units are not the same thing:
12 provider answers
× 8 predeclared target brands
= 96 answer-brand evaluation cells
Reporting “96 AI answers” would overstate the collection by a factor of eight. Reporting only 12 rows without describing the derived brand cells would hide the scoring denominator.
What the answers did
For the wedding-selection question, the eight brands received 21 positive shortlist recommendations across 32 eligible answer-brand opportunities.
For the official-channel question, the models affirmatively listed a target brand in 25 of 32 opportunities.
Those two rates are not a funnel. They come from different prompt intents. The useful comparison is at brand level:
| Brand | Wedding shortlist | Official China channel listed |
|---|---|---|
| Cartier | 4/4 | 4/4 |
| Tiffany & Co. | 4/4 | 4/4 |
| Bvlgari | 4/4 | 4/4 |
| Van Cleef & Arpels | 4/4 | 4/4 |
| Chaumet | 3/4 | 2/4 |
| Boucheron | 1/4 | 2/4 |
| Piaget | 0/4 | 4/4 |
| De Beers Jewellers | 1/4 | 1/4 |
The Piaget row is the clearest example of why entity recognition is not recommendation. The APIs could associate the brand with a current China-facing channel, but the brand did not enter their wedding shortlist in this tiny sample.
Chaumet shows the other direction. It appeared in three of four wedding shortlists but only two of four official-channel answers. A brand can enter consideration while the route to verification remains less consistently reproduced.
Neither pattern tells us why it happened. It tells us what to inspect next.
A buyer journey needs more than one metric
For a high-value product, I would split the audit into at least three layers.
1. Consideration
Does the brand enter a shortlist for a real decision: wedding jewelry, an anniversary gift, a diamond purchase, a high-jewelry commission or a particular budget?
This is where associations, product authority, cultural relevance and brand familiarity may matter. The current pre-wave does not identify which factor caused a recommendation.
2. Verification
Can a buyer find a current route that the brand itself controls or authorizes?
That may include a China website, boutique directory, official customer-service route, mini-program or disclosed marketplace store. A brand-global domain may be legitimate, but it is not always the most precise answer to a mainland-China channel question.
3. Purchase confidence
Can the buyer verify what the channel is allowed to sell and what the maison will actually support after purchase?
This is where generic AI language becomes risky. “Official authentication,” “global warranty,” “free lifetime resizing” and “seven-day returns” are not universal benefits. They can vary by product, market, purchase route, documentation and damage.
A buyer does not need a plausible paragraph. They need the current policy that applies to their item.
The truth table was harder than the prompt panel
Checking whether an answer printed a domain was easy. Deciding what the domain meant was not.
For every asserted official route, I needed to keep separate fields for:
{
"brand_id": "stable entity identifier",
"channel_locator": "domain, account, store or service route",
"operator_entity": "the entity actually operating it",
"ownership_or_authorization": "verified, unresolved or contradicted",
"live_status": "dated observation",
"market_scope": "mainland China, global or another market",
"service_scope": "sales, boutique lookup, support or other",
"truth_source": "dated brand-controlled or authoritative evidence"
}
I did not use .cn ownership, an ICP record or a company registration as interchangeable proof of officiality. They answer different questions. I also did not convert “no public evidence found” into “verified absent.”
Across the four official-channel answers, there were 19 exact target-brand domain assertions. Fifteen matched the dated China-local route in the truth table. Three were brand-global alternatives. One route remained unresolved.
That is why I would not publish a simple “channel accuracy rate” without the categories beside it. A global route can be less locally precise without being false. An unresolved route is not evidence of impersonation.
Invalid calls are not negative brand evidence
An ERNIE API surface was planned for the same panel, but all six requests returned an account-state 403. Those rows were retained as attempted measurements with an invalid reason. They did not enter the answer denominator.
This sounds obvious until an automated report turns a missing response into eight brand absences.
The minimum record I want for a failed call is still substantial:
observation_id
question_id and version
platform and surface
model/configuration
retrieval status
attempt timestamp
instrument versions
error class and raw provider status
validity = invalid
The error tells me something about the collection instrument. It tells me nothing about whether the absent answer would have mentioned Cartier or Piaget.
What a brand team can do with the result
The first action is not “publish more GEO content.” It is to identify which layer is failing.
If the brand is missing from consideration prompts but its channel is correctly reproduced, audit the evidence around the buying decision: occasion, design language, category authority, local editorial presence and comparison contexts.
If the brand enters shortlists but channel answers are weak, give digital operations and legal a concrete route inventory to maintain. The official website, store locator, customer-service route, marketplace disclosures and operating entity should not contradict one another.
If the route is correct but service claims are wrong, the problem belongs partly to customer service and policy publishing. Put current, scoped answers where buyers and reviewers can verify them. Do not promise that publishing the page will change an AI response; retest the same prompt panel.
This gives the organization a practical ownership map:
| Failure | Likely owner |
|---|---|
| Missing from occasion shortlist | Brand, editorial, PR, category marketing |
| Wrong or vague official route | Digital operations, e-commerce, legal |
| Incorrect authorization claim | Legal, retail operations, channel management |
| Incorrect warranty or service claim | Customer service, legal, after-sales |
| Measurement cannot reproduce | Research, data and analytics |
“Improve AI visibility” is too broad to assign. These failures are specific enough for someone to own.
What this test does not establish
The APIs ran without live retrieval. This study cannot tell us which sources influenced an answer, whether a displayed domain was retrieved, or what a signed-in consumer sees in the current product interface. A target-supported Citation Rate is not observable here.
It is also a 12-answer exploratory pre-wave, not the first longitudinal benchmark. It cannot establish a stable brand position, a platform winner, a trend or a causal link between a page and a recommendation.
The next formal wave needs declared and repeatable consumer or search surfaces, a larger fixed buyer-question panel, a frozen truth table and the same metric membership rules in every period. Qwen, Kimi and ERNIE should be included only after their exact surfaces and account states pass preflight; adding a platform halfway through would create an apparent trend by changing the instrument.
The useful question changed
Before the test, I would have asked whether each jewelry house was visible.
Now I would ask three narrower questions:
- Does the brand enter the buyer’s shortlist?
- Can the buyer verify a current official China route?
- Are the purchase and service claims attached to that route accurate?
The answers lead to different work. That is the value of the small pre-wave: not a league table, but a better diagnosis to carry into the larger benchmark.
The complete method, evidence boundaries and brand-level table are available in the canonical study on Visibility Atlas.
Top comments (0)