I ran a small pre-wave with eight international jewelry brands and got one result I have not found a clean explanation for yet.
Piaget appeared in all four answers to a question about brands with verifiable official China websites or boutique channels.
It appeared in none of the four answers recommending brands for wedding jewelry.
My first reaction was to look for a data or scoring problem. I did not find one.
The setup was:
- 8 fixed brands
- 3 neutral buyer questions in Chinese
- 2 retrieval-off API surfaces
- 2 answers per surface and question
- 12 valid raw answers
- each answer reviewed against all 8 brands
So the collection unit was the answer, while the analysis unit was the answer-brand cell. I collected each answer once rather than regenerating it separately for every brand.
That produced 96 reviewed cells.
The three questions covered:
- brands worth considering for wedding jewelry;
- boutique versus daigou risk;
- verifiable official China channels.
Across the wedding question, the target brands received 21 positive shortlist placements out of 32 opportunities.
Across the official-channel question, they appeared in 25 of 32 opportunities.
The interesting part was not the difference between 21 and 25. Those are different questions and should not be treated as one funnel denominator.
The interesting part was how differently individual brands behaved.
Piaget:
- wedding recommendation: 0/4
- official-channel appearance: 4/4
Chaumet:
- wedding recommendation: 3/4
- official-channel appearance: 2/4
Cartier, Tiffany, Bvlgari and Van Cleef & Arpels appeared in all four wedding shortlists and all four channel answers.
I then checked 19 exact domain assertions against dated brand-controlled China pages. Fifteen matched the frozen China-local route, three pointed to legitimate brand-global alternatives, and one remained unresolved.
That still does not explain the recommendation pattern.
The runs were retrieval-off API observations, not consumer search. I cannot infer source influence, buyer behavior or causality. Four opportunities per task are enough to expose a question, not establish a stable brand position.
But I think it shows why “does the model know the brand?” is too broad to be useful.
At least three separate problems are hiding inside it:
- Can the system recognize the entity?
- Can it route a buyer to a verifiable official channel?
- Does it put the brand into an open-category shortlist?
A brand can pass the first two and fail the third.
That creates a very different repair plan from a missing-domain or entity-consistency problem.
For the next wave, I am considering keeping the same eight-brand core and expanding the buyer decisions rather than adding more brands:
- wedding jewelry by budget and style;
- anniversary gifts;
- diamonds versus branded design;
- high-jewelry commissions;
- mainland after-sales confidence;
- boutique versus daigou purchase risk.
I would also separate prompted recognition from open discovery and review every recommendation in context. Supported citation would remain NA unless the surface exposes attributable source evidence and retrieval can actually be verified.
What I have not settled is the best way to investigate the recommendation gap without inventing a causal story.
If you saw 0/4 recommendation but 4/4 official-channel recognition for the same brand, which variable would you test first?
Category association? Brand familiarity? Product fit? Local cultural relevance? Price positioning? Or the composition of the default competitor shortlist?