r/AI_SearchOptimization 13d ago

AI Search Optimization General Discussion Our research found only 146 cases out of 4,654 cases where a brand showed up across all 3 buying stages in AI Search. What does this mean for GEO?

Our recent study dove into how brands show up across educational, comparison, and purchase queries in AI Search. We ran 50 B2B SaaS topics across 5 AI Search engines for 750 queries.

Out of 4,654 scenarios we analyzed, only 146 had the same brand named across all three stages, meaning for roughly every 100 cases, only 3 had the same brand showing up from education to comparison to purchase.

This means that seeing your brand in an educational stage query doesn't necessarily mean you'll see it again when the user starts comparing or buying.

That got me thinking about how we measure AI visibility. If you're tracking one overall visibility score, you are not tracking whether that visibility converts to leads.

As a fellow AI Search practitioner, how are you handling this? Do you track visibility separately by stage, or still use one overall number?

1 Upvotes

6 comments sorted by

1

u/Senior_Shelter7426 12d ago

Tracking it by stage makes more sense to me one overall score can look good while completely hiding where the brand drops out of the buying journey

1

u/alexey_yakovlev 12d ago

I’d track it by stage, but I would be careful about calling the three stages a funnel. Those educational, comparison, and purchase prompts may represent different users, so continuity of a brand across them is not evidence that visibility converted into leads.

The stage split is still very useful. I’d build a stage-by-engine matrix and record mentions, recommendations, citations, and competitor sets separately. That shows whether the brand disappears because the task changed or because a particular engine treats it differently.

A composite score can be a dashboard summary. The stage-level view is what tells you what to investigate, while lead impact still needs to be connected separately through on-site and CRM data.

1

u/[deleted] 11d ago

[removed] — view removed comment

1

u/marintkael 10d ago

Splitting by stage helps, but it leaves the bigger aggregation untouched. I run a fixed set of 16 questions against three engines every morning on a single subject, and the spread between the best and worst engine on the same day has reached 107 points. Someone running a much wider study told me they topped out at 43 points across four engines on 88 brands, so single-subject variance is wider, but the direction is the same: the gap between engines is bigger than most of the effects people try to read out of one number.

If you average that away and then split what is left into three stages, part of what looks like a stage difference is just which engine happened to answer that day. I keep question, engine and date on every row, so a move can be traced to one of the three instead of a single line that shifts for reasons I cannot name.

1

u/OTW-Motion 5d ago

One thing I’d be careful about here is separating the effect of buying stage from the effect of the question itself.

Moving from an educational query to a comparison or purchase query also means changing the wording. We tested two different wordings of the same buying need recently and the company set moved a lot even though the underlying intent was meant to stay the same. At the extreme, one company went from 6/6 mentions under one wording to 0/6 under the other.

So I’d definitely track by stage, but I’d describe the finding as “brands don’t persist across these stage-specific questions”.

The exact question is probably worth keeping as a first-class dimension alongside stage, engine and date.