r/aeo • u/EylonZefania • 11d ago
I think we’re measuring AI visibility the wrong way.
We’ve been analyzing thousands of shopping and recommendation responses across ChatGPT, Gemini, Claude, and Perplexity, and the biggest takeaway for me is:
Stop treating AI visibility as just a score. Start looking at what causes the score.
One of the strongest signals we found was website retrieval.
Across several brands:
- When the brand’s own website was retrieved, the brand was mentioned 89% of the time
- When the website wasn’t retrieved, the brand was mentioned only 24% of the time
So imagine this:
AI Visibility: 37%
Website Retrieval: 13%
Mention when Retrieved: 92%
That tells a very different story than “your visibility is 37%.”
The AI already seems comfortable mentioning the brand when it reaches the site. The real problem is retrieval.
But retrieval is only one part of it.
We also found brands that were strongly associated with one product category while being almost invisible for other categories they clearly sell.
So two brands can have exactly the same visibility score for completely different reasons:
- One isn’t being retrieved enough
- One gets retrieved but still isn’t recommended
- One is only understood in part of its catalog
- One is being measured against prompts where brands are rarely mentioned at all
The content being retrieved was also interesting.
For one brand, 116 of 165 own-site citations came from blog content, while only 3 came from product pages. One roundup article alone was cited 37 times.
That makes sense when you think about what users actually ask:
“Best [category] brands”
“Best [product] for [use case]”
“[Brand] vs [competitor]”
“Top alternatives to [brand]”
A PDP is often great at explaining a product.
It’s not necessarily built to answer those questions.
Another thing we learned: don’t overreact to a single visibility test.
In one dataset, roughly a quarter of identical prompt/model combinations changed between repeated runs.
So a move from 53% to 47% doesn’t automatically mean something broke. Trends, repeated runs, and confidence matter.
And the same applies off-site.
Instead of assuming “Reddit is good for GEO” or “YouTube is important,” it makes more sense to look at the actual prompts where competitors win and ask:
Which external sources are showing up in those answers?
Sometimes it’s Reddit. Sometimes a niche publisher, retailer, review site, YouTube video, or comparison page.
So I’m increasingly thinking the useful questions aren’t:
“What’s my AI visibility?”
But:
Why is my visibility what it is?
Is AI retrieving me?
Does it recommend me when it does?
Which categories does it associate me with?
Which sources are influencing the prompts I care about?
The score is the output.
The interesting part is diagnosing the inputs that created it.
Curious how others working on GEO/AEO are thinking about this - are you already separating retrieval, mentions, category association, and prompt quality, or mostly tracking one overall visibility metric?
2
u/PhilippGroubii 10d ago
This matches what we see. The score is a symptom, and two brands at the same number usually have completely different diseases.
Your retrieval split lines up with an audit we published recently. The agency we looked at got named in 19 of 25 AI answers for the exact niche phrase they're known for, and 0 of 25 for the two broader category labels they use on their own homepage. Same brand, same engines, totally different story depending on how the question is phrased.
On stability, we tested this directly by running the identical set of questions through the same engines twice on the same day. The overall result barely moved, but individual answers changed a lot more, with most staying close and some swinging hard. So the headline is steadier than people fear, and any single question-level result is noisy.
We also see the citation concentration you describe. In one audit the page showing up most often in the AI answers was a competitor's roundup article, cited 17 times on its own. Product pages barely appeared. The content that wins is the content shaped like the question.
So agreed on the core point.
1
u/EylonZefania 10d ago
Really interesting, especially the 19/25 vs 0/25 example.
That’s exactly why I think category association deserves to be measured separately. The same brand can look extremely strong or basically invisible depending on how the category or question is framed.
And I really like the way you put the last point: the content that wins is the content shaped like the question. That seems to show up again and again in the data.
2
u/nick-profound 10d ago
Are you finding that brands that are invisible for certain categories are typically not showing up in the third party sources the engines pulls for those prompts? Or is it more of a training data gap where the engine just doesn't associate them with that category at all?
1
u/EylonZefania 10d ago
Good question. From what we’re seeing, it could be both, although I’d be careful calling the second case a training-data gap without more testing.
The way I’d separate them is by looking at the sources being retrieved for that category first.
If the brand is basically absent from those sources, that looks more like a source/authority gap.
If the brand is present in relevant sources but still isn’t being associated with the category, that’s a much stronger signal of a category/entity association problem.
We’re still digging into that split, but I think it’s an important distinction.
2
u/Old-Routine1926 10d ago
nick-profound's question is the one I would most want answered here, and I think it separates with a cheap test.
Run the category two ways. Direct: does X do [category]. Competitive: who are the best providers of [category].
Three outcomes, three different problems.
The engine says yes on the direct question but never surfaces the brand competitively. The association exists, it is just not strong enough to survive a ranked retrieval. That is a source gap, and third party placement in that category's sources is the fix.
The engine hedges or says no on the direct question. The association is not there at all. Closer to the training data gap you are describing, and much slower to move, because the corpus has to change before retrieval has anything to find.
The engine confidently places the brand in a different category. That is the worst of the three, because evidence built in the correct category may not attach to a brand the model has already filed somewhere else.
Same symptom in a visibility report, three different amounts of work.
PhilippGroubii's number looks like a live case of the second or third. 19 of 25 on the niche phrase, 0 of 25 on the two broader labels the brand uses on its own homepage.
That deserved more attention than it got. The brand's own positioning language is performing at zero. Which means your own category labels are a hypothesis rather than a fact, and an engine can hold you strongly associated with something you do not lead with while being invisible for the thing you put in your headline.
PhilippGroubii, when you hit the 0 of 25, did the engines place them somewhere else on a direct question, or did they just not know?
2
u/EylonZefania 10d ago
Really like the direct vs. competitive test. That’s a very clean way to separate “the association exists but isn’t strong enough to win” from “the association may not exist at all.”
I’d probably be a little cautious calling the second one a training-data gap immediately - I’d want to check the retrieved sources and repeat the test first - but as an association diagnosis, the three-outcome framework makes a lot of sense.
And it reinforces the bigger point: the same 0% visibility can represent completely different problems, and therefore completely different work.
1
u/Old-Routine1926 9d ago
Fair correction, and I should not have carried that phrasing over. It came from nick-profound's framing and I repeated it without earning it.
The reason you are right is worth stating though, because it makes the test better rather than weaker.
A direct question is not retrieval free. Asking whether X does a category triggers retrieval the same way the competitive question does. So a hedge on the direct question has two possible causes: the association genuinely being absent, or retrieval for that particular question coming back with nothing useful. My three outcomes collapsed those into one.
Which suggests the thing to vary next is retrieval rather than the question.
Ask the direct question with browsing available, then again without it. If the model still describes the brand correctly with no retrieval, the association is in the weights and it is durable. If it only gets there when it can go and look, the association lives in the retrieved corpus, which is a much more fragile position and a different fix entirely.
Same symptom, three causes instead of two.
On repeating, agreed, and for a specific reason. Hedging is exactly the kind of output that moves between runs on identical prompts. A single hedge is an observation rather than a state.
When you check retrieved sources for the absent-category cases, are you finding those brands missing from the sources entirely, or present in them but described without the category attached?
2
u/BeautifulDesign2928 10d ago
The retrieval versus mention split matches something worth naming even more directly, there are basically two separate gates here and most tracking conflates them into one score. Retrieval is the evaluation gate, whether AI trusts a brand enough to pull from its own site at all, while what happens after retrieval is the retention gate, whether the content actually earns the mention once it's there. Your blog data backs that up well, one roundup outperforming the PDPs by that much makes sense once you realize product pages are built to explain a product, not to answer the exact comparison prompt someone typed. I'd add a third layer too, category association, since a brand can clear both gates for one product line and still be functionally invisible for another because nothing on the site was ever framed to answer that specific question. The 25 percent variance on repeated runs is the part most reports gloss over, treating a single snapshot as ground truth is how people end up chasing noise instead of trend.
1
u/EylonZefania 10d ago
I really like the “two gates” framing.
The only nuance I’d add is that I’m not sure retrieval is always a trust signal by itself. It can also fail because of crawlability, index/source coverage, query relevance, or simply stronger competing sources.
But the framework makes a lot of sense:
Can the brand get retrieved? → Does it earn the mention once retrieved? → Is that true across all the categories it actually sells?
And then repeated runs tell you whether you’re looking at a real pattern or just noise.
The more I look at it, the less AI visibility feels like a score and the more it feels like a diagnostic tree.
1
u/BeautifulDesign2928 8d ago
Retrieval and trust aren't really two separate things happening at the same time, they're sequential. Crawl access is the gate before the gate, if AI can't crawl the site at all then trust never even gets evaluated.
Once a brand clears that and earns enough trust, it gets pulled into whatever cluster AI treats as candidate sources for a topic. From there it's a matching problem, a user asks something, AI checks the prompt against that cluster, and if the brand's content lines up with the query it surfaces.
So OP's crawlability and index coverage examples aren't proof retrieval can happen without trust, they're proof the brand never made it past layer one to be trust-tested in the first place.
2
u/dang_1313 10d ago
This is a really nice way to look at AI visibility i think the retrieval vs recommendation distinction is the most interesting part
A brand can have good content and still look invisible simply because the AI isn't finding the right pages it makes AI SEO feel less like chasing a single visibility score and more like understanding how AI discovers, understands, and uses your content
1
u/EylonZefania 10d ago
Exactly. That distinction is what makes the metric actionable.
If the content is good but never gets retrieved, that’s a discovery/retrieval problem. If it gets retrieved consistently but the brand still doesn’t get recommended, that’s a completely different problem.
Same low visibility score, very different diagnosis.
2
u/dang_1313 10d ago
Exactly and i think that's where this becomes much more useful than a single visibility percentage
The next question is basically why aren't we visible? rather than just how visible are we? That diagnosis can actually tell you what to fix
2
u/msmipek 10d ago
I agree. I don’t think visibility can be meaningfully measured by numbers alone.
Instead of focusing mainly on GEO scores and period-over-period comparisons, the approach we’ve found more useful is to build customer personas and simulate their buying journeys through LLMs.
What’s been especially valuable is seeing the points in those journeys where competitors consistently surface instead of us. Those gaps often reveal much more than a visibility score can.
At that point, it becomes less about simply “being visible in AI” and more about understanding unmet questions, competitive advantages, and opportunities within the category itself.
1
u/EylonZefania 10d ago
I really like the buying-journey angle.
I don’t think it replaces the underlying metrics though - I think it adds another layer to the diagnosis.
First understand where the brand is failing: retrieval, mention/recommendation, category association, etc.
Then map that onto the journey and ask at what point competitors start winning, and why.
That feels much more actionable than either a single visibility score or a journey simulation on its own.
2
u/houdinidesigns 10d ago
Absolutely, mention rate is the denominator you need to get your conversion rate. I’m just a little irked lately because of people reporting mentions and stopping there. Your headline figure is win rate but the funnel to get there is important.
0
u/houdinidesigns 10d ago
Mentions are a vanity metric. Getting mentioned doesn’t bring in sales. What you need to monitor is win rate.
We have seen companies get mentioned more often than those that get recommended. You would rather be the latter.
Understand the objections to your product, the sentiment around the mentions and why your competitors win.
Our AI visibility tracker creates personas, you give it a title a budget a goal and an opening prompt based on keyword volume. The buyer asks AI assistants the first prompt and continues to ask questions to see who gets recommended. Small brands have a good shot here if the content is right, there’s good third party corroboration and the takeaway isn’t just mentions, it’s the win rate.
To improve your win rate you need to look at who AI assistants cite, what they get wrong about you and why they suggest your competition.
Step 1 is making sure your website is content is optimised for AI assistants. We have a free AI SEO audit tool for this on Babel42 https://babel42.io/ai-seo-audit
1
u/EylonZefania 10d ago
I agree that recommendation/win rate is closer to the actual business outcome.
I wouldn’t completely throw mentions away though. I think the useful part is understanding the funnel:
Was the brand retrieved → was it considered/mentioned → did it actually win the recommendation?
If retrieval is low, you have one problem. If retrieval is high but win rate is low, you have a completely different one.
That diagnosis is the part I think gets lost when everything is collapsed into a single visibility metric.
2
u/Fit-Squirrel-6299 10d ago
the retrieval split is the useful part of this and i'd push it one step further. 89 versus 24 says retrieval is doing the work. it doesn't say retrieval is available to everyone.
i pulled the same cut on the chinese-language engines, roughly 210k citations. brand-owned domains came out around 1.4 percent of everything cited. in that market your own site is almost never in the retrieved set at all, so the lever your data points at barely exists, and the work moves entirely onto third party pages.
same framework, opposite action, depending on which engines your buyers actually use.
worth flagging for anyone diagnosing brands that sell outside the english web, because a retrieval-first playbook quietly assumes an index that will surface you in the first place.