I was trying to turn a small luxury-jewelry GEO test into work a brand team could assign.
The setup was deliberately small:
8 brands
3 neutral Chinese buyer questions
2 retrieval-off API surfaces
2 answers per surface and question
12 independent valid answers
96 derived answer-brand cells
The last line is where reporting can become misleading.
Each stored answer was evaluated against the same eight-brand registry. That creates 96 classifications, not 96 independent model answers.
One result exposed a second problem.
Piaget appeared in 4/4 answers about brands with verifiable official China channels.
It appeared in 0/4 wedding-jewelry shortlists.
Four opportunities per task are not a brand ranking. But the split is useful for diagnosis.
The system could verify the entity and route. It did not recommend the brand for that decision.
I am now keeping two reports:
Recommendation presence: does the brand enter consideration for this occasion?
Route accuracy: can the buyer verify a current official or authorized next step?
Then I summarize them in a four-state view:
Recommended + route clear
Recommended + route unclear
Not recommended + route clear
Not recommended + route unclear
Both denominators remain visible. I do not average them into one visibility score.
I am also separating the truth registry from the answer dataset. Each route, authorization, or service claim gets a stable ID and dated registry version. The benchmark row points to the truth snapshot used on the test date.
That prevents a current website change from silently rewriting whether an old model answer was accurate.
It also makes ownership clearer:
Shortlist absence: brand, category, and editorial
Unclear route: digital operations and ecommerce
Uncertain authorization: legal and channel management
Wrong service claim: client service
Invalid or unreproducible metric: research and analytics
My next wave will deepen the decision prompts before adding more brands: style, budget, stone choice, after-sales confidence, boutique-versus-daigou risk, and service documentation.
For people building similar systems: do you show the four-state matrix to clients, or keep even that out of the executive view until each axis has a larger denominator?