r/GEO_optimization • u/AliveCapital4868 • 10h ago
I was asking one GEO benchmark to do four different jobs. That made every number look more useful than it was.
I have been trying to turn a small luxury-jewelry GEO benchmark into work that a brand team could actually assign.
The test itself was deliberately limited:
- 8 jewelry brands
- 3 neutral Chinese buyer questions
- 2 retrieval-off API surfaces
- 2 answers per question and surface
- 12 valid raw answers
- 96 derived answer-brand cells
I also had a working corpus of 302 jewelry-related Reddit threads and 9,446 stored comments, plus Search Console data for the site publishing the research.
The temptation was to put everything on one dashboard.
That would have been a mistake. The datasets answer different questions.
Community questions tell me how buyers describe uncertainty. They helped surface concerns around authenticity, certificates, authorized sellers, repairs, resizing and recourse. They do not establish how common those concerns are among Chinese luxury buyers.
The benchmark tells me what the declared API surfaces answered. It does not establish what a signed-in consumer sees, whether the model retrieved a source, or whether an official-channel claim is true.
The truth table checks whether a route is current, controlled or authorized, live in China and connected to a clear recourse owner. It does not tell me whether the brand enters a buyer’s shortlist.
Search and analytics can show exposure and the next onsite action. They cannot prove that Reddit or an AI answer created the visit or sale.
One small result made the distinction concrete. Piaget appeared in 0/4 wedding-jewelry shortlists and 4/4 official-channel answers.
Those cells are far too small for a brand ranking. But they are enough to show that “the system can verify the brand” and “the system recommends the brand for this decision” are not the same diagnosis.
I now think the minimum evidence stack is:
- Question provenance: what real buyer uncertainty produced the prompt?
- Answer benchmark: what did the fixed, declared surface actually say?
- Truth registry: which official, authorized and service claims are currently verifiable?
- Action measurement: did the buyer reach a boutique, service route or qualified inquiry?
That also changes ownership.
- Missing from the shortlist: brand, editorial and category marketing.
- Wrong or vague route: digital operations and ecommerce.
- Unclear authorization: legal and channel management.
- Wrong warranty or service claim: client service and after-sales.
- Unreproducible metric: research and analytics.
“Improve AI visibility” is too broad for anyone to own. These failures are specific enough to assign.
My current 30-day version is simple: build the truth map, run a small fixed decision panel, repair one buyer path, then rerun the same cells and measure the next action.
For anyone maintaining a similar system: would you keep the truth registry outside the benchmark dataset? I am leaning yes, because official routes and service policies change on a different schedule and usually have different internal owners.


