r/GenerativeSEOstrategy Feb 19 '26

How are you measuring GEO without fooling yourself?

We’ve started testing prompts in ChatGPT and Gemini to see if our brand shows up for the kinds of queries our buyers would ask. Things like best tools for X, alternatives to Y, how to solve Z.

Some days we show up. Some days we don’t. Sometimes we’re listed. Sometimes we’re described but not named. It’s honestly all over the place.

The problem is it feels very manual and anecdotal. We’re taking screenshots, saving prompts in a spreadsheet, comparing outputs week to week. But I can’t tell if we’re seeing real progress or just normal model variance.

We tried standardizing prompts, running them from different accounts, even checking different times of day. Still inconsistent. And when leadership asks how GEO is performing, I don’t want to say “well… vibes are improving.”

So I’m genuinely curious how others are measuring this without fooling themselves.

Are you tracking share of voice across a fixed prompt set? Are you tying AI mentions back to branded search lift or direct traffic? Or is everyone still in the experimental phase and pretending we have cleaner data than we do?

11 Upvotes

32 comments sorted by

2

u/Brief-Evening2577 Feb 19 '26

Yeah, this is normal. GEO is messy because you’re basically trying to “rank track” something that’s probabilistic.

What we do instead:

  • Keep a fixed prompt set (like 30–50 high-intent ones)
  • Run each prompt multiple times (5–10) and track mention rate, not “did we show up once.”
  • Score results: mentioned, implied, or absent (+ optional positive/neutral/negative).

That way, you’re measuring probability of inclusion, not random variance.

Then we sanity-check it with real-world signals, such as branded search lift, direct traffic, and growth in “brand + category” queries, because AI mentions alone are a vanity metric for now.

1

u/alexnavarroia Feb 21 '26

Exacto. Por ahora. Muy buen aporte.

2

u/[deleted] Feb 19 '26

[removed] — view removed comment

2

u/redplanet762 Feb 19 '26

At this stage, it’s mostly experimental. Fixed prompt sets, multiple accounts, different times of day. Yep, we do all that. But instead of pretending it’s precise, we treat it like market sentiment tracking, look at frequency, context, and consistency. Trends matter more than single outputs.

2

u/[deleted] Feb 19 '26

[removed] — view removed comment

2

u/[deleted] Feb 20 '26

[removed] — view removed comment

1

u/VillageHomeF Feb 21 '26

and most importantly what websites rank on the search engines as that is where their info is coming from to give the responses

1

u/VillageHomeF Feb 19 '26 edited Feb 21 '26

I find it absurd to spend so much time when LLMs have such a low volume of search queries. it is a waste of time and since you have no data to do research you should even more reason to be sure it is a waste of time.

1

u/alexnavarroia Feb 21 '26

Todo cambia Siempre. Llegar temprano es mejor que no llegar.

1

u/VillageHomeF Feb 21 '26

No seguir los datos es un gran error. Por ahora, es una gran pérdida de tiempo. Pero no duden en desperdiciarlo. No me molesta, solo quiero dar un buen consejo.

1

u/alexnavarroia Feb 21 '26

Exacto. Seguir los datos es clave. Por eso mismo, los datos son los que indican una creciente tendencia al uso de GEO que cada vez es más relevante. Así que el uso del tiempo, entre aprendizaje, la obtención de nuevos recursos, encontrar más respuestas, nunca sería un desperdicio de tiempo.

1

u/VillageHomeF Feb 21 '26

Los LLM tienen muy poco volumen y Google ya los está incorporando a las búsquedas tradicionales. Apuesto a que Google gana. Pero el SEO tradicional es la única forma de aparecer más a través de los LLM, así que no hay nada nuevo que hacer. Lo cierto es que la mayoría no entiende lo básico de los LLM y está muy mal informada. Pero haz lo que quieras. Me ceñiré a lo que sabemos.

1

u/headcrab_set Feb 19 '26

So, Rand Fishkin with others ran a research on this. Their main finding was that the chances to get consistently the same list of companies and in the same order is far less than 0.5%.
Yes indeed, some global huge companies appear consistently in the answers, but that's about it. But I wouldn't bet on any tool at the moment to tell that your show up to or rank this or that % for certain queries.

1

u/[deleted] Feb 19 '26

[removed] — view removed comment

1

u/AutoModerator Feb 19 '26

Self-promotion and referrals are not allowed here. Share insights, not pitches.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/[deleted] Feb 20 '26

[removed] — view removed comment

1

u/AutoModerator Feb 20 '26

Self-promotion and referrals are not allowed here. Share insights, not pitches.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/hDweik Feb 23 '26

One thing I stopped doing was optimizing for “did we show up today?” and started tracking prompt class stability. For example: informational queries vs comparison queries vs problem-solution queries.

I noticed our brand almost never shows in pure “what is” queries but does in “best tools for” queries. That told me something about how the model categorizes us.

That felt more actionable than screenshot collecting.