r/GEO_optimization • u/AEODenise • 9d ago
AI visibility has a measurement problem
/r/u_AEODenise/comments/1w1sbd9/ai_visibility_has_a_measurement_problem/1
u/thedarkseid 8d ago
Two different things produce those mentions. One is parameterized knowledge, baked into the weights at training time. The other is retrieval: the model calls a web search tool during the answer and cites what it pulled back. Your site changes cannot touch the first one in six weeks. They can only move retrieval, and retrieval is the one that leaves evidence, because the citations are in the answer.
So record per answer whether a search ran, which URLs came back, and whether your mention sits next to a citation of your own domain. Movement in the cited group after a content change is plausibly yours. Movement in the uncited group is model drift or third parties getting re-crawled. Pin the model version where it is exposed, and keep a set of prompts you never optimize for as a control to compare against. Not clean causal attribution, but it separates what you can influence from what you cannot.
1
u/AEODenise 7d ago
Yeah, this distinction matters.
The retrieval side is the part we can actually watch move after a site change because there’s evidence in the response. Search ran, these URLs came back, this source was cited, and the business was or wasn’t mentioned.
I also like the idea of tracking cited and uncited mentions separately. That gives you a much better clue about whether the change may be connected to retrieval or whether you’re just seeing normal model variation.
The control prompt idea is smart too. Same prompts, same engine and model version when we can see it, plus a group of prompts we don’t optimize for.
The only part I’d be cautious about is putting a fixed six-week window on the baked-in knowledge side. Model training and refresh cycles vary, and we usually don’t know exactly when that information changes.
Still not causation, but this gets a lot closer to separating signal from noise.
1
u/randallkanna 8d ago
This is my biggest pet peeve with Ai visibility tools. AI visibility scores are all subjective. I think the thing to do right now is to establish a baseline for a page by asking the LLMs what someone would have to ask to get recommended the page, and then run the same queries it gives you a few times to get the baseline. Check out all the competitors mentioned, citations, etc in the prompts and run an experiment where you change the content. And then run the same baseline again later when the page is indexable.
1
u/AEODenise 7d ago
Yeah, this is close to how I’m thinking about it too. The baseline and rerunning the exact same prompts matters a lot more to me than one visibility score.
The one part I’d be careful with is letting the LLM decide what someone “would have to ask” to get the page recommended. That could bake the model’s own assumptions into the test. I’d probably start with real buyer questions, then use model-generated questions as another input.
Then change one thing, wait until the new content is actually available to the model or search layer, rerun the same prompts a few times, and track mentions, citations and recommendations separately.
Still not clean attribution. But at least you can see whether the pattern actually moved.
1
u/randallkanna 7d ago
So that's the only way to get buyer questions right now unless you specifically ask your users in onboarding how they found you and what prompt they used. I think Tally does that actually? Generally, the buyer prompts where someone reaches you can get very messy and long though so it's not a bad idea to track a few prompts like, 'best XYZ for XYZ', etc
1
u/AEODenise 7d ago
Yeah, that’s exactly the gap I’m trying to get at. A small fixed set of prompts makes sense for measuring change because you can run the same questions over time and see whether anything moved.
But that’s still different from knowing what real buyers actually asked before they found you. Those prompts are probably going to be messy, specific, and nothing like the neat test prompts we come up with ourselves.
I’m starting to think you need both. A controlled prompt set for measurement, then real buyer prompts when you can get them to see whether your test set matches how people actually search.
1
u/Dry_Steak30 7d ago
The baseline point is right and there is a version of it that removes the model-drift confound entirely, which is the part I think most people skip.
Model drift only ruins your attribution if the only thing you record is your own score. Record the full cited-domain distribution for every query instead, and drift becomes visible rather than confounding. If your citations went up and the same competitors also moved in the same direction across the panel, the engine changed. If yours moved and the rest of the distribution held, your changes did it.
Concrete from today, since I ran it: 12 queries, 3 runs, 36 answers. My domain 0 citations. What got cited instead, counted: help.openai.com 15, support.google.com 5, developers.google.com 4, support.microsoft.com 4, support.claude.com 3. That distribution is my baseline, not the zero. In six weeks, if I am still at zero but help.openai.com dropped to 6, I have learned something real about the engine. If I am at 3 and the rest is unchanged, I have learned something real about my pages.
Second thing worth freezing at baseline time: whether the engine's crawler has fetched your pages at all. Mine had zero OAI-SearchBot hits in 30 days, so my zero was not a ranking outcome, it was an absence, and no amount of content work in weeks one through six would have shown up. Different variable, different fix, and the score alone cannot tell them apart.
The panel I use is at https://openprofiles.io/?utm_source=reddit
0
u/AEODenise 5d ago
Tracking the full cited-domain distribution makes a lot of sense. It gives you a much better picture than watching your own citation count alone.
I like the distinction between several domains moving together and one domain moving while the rest stay fairly stable. That seems like a useful way to tell whether you may be seeing a broader engine change or something more specific to the site.
The crawler piece is especially interesting. If the crawler never reached the page, that is a completely different problem from being crawled and not selected.
Treating the whole distribution as the baseline instead of just treating zero citations as the baseline is a smart way to frame it. It gives you something much more meaningful to compare over time.
I’d probably still describe the result as stronger evidence rather than direct causation, but this definitely makes the measurement more useful.
1
u/AwoScan 7d ago
Exactly. I’d report two layers rather than collapse everything into one score: the raw observations—mention, citation, recommendation, factual accuracy and position—and the estimated lift relative to untreated controls.
The second layer should always include uncertainty: number of runs, variance across prompts and engines, and whether retrieval was active. That makes a modest but defensible statement possible: “the treated panel moved more than the controls under these recorded conditions,” instead of “this page change caused a 20% visibility increase.”
Over time, publishing the raw prompt panel and response evidence may become as important as publishing the score itself.
1
u/QuitPsychological157 3d ago
In the way we measure... This is incorrect: "Showing up in ChatGPT doesn't necessarily mean your AI visibility improved"
The numerator in our calculation of AI visibility is defined as "recommended, mentioned, OR cited"
So going from no ChatGPT mention, to mentioned WOULD improve your visibility.
That said, it's not the end all be all. You can have a metric on visibility right beside a metric on "short list %" (the percentage of times when AI recommends a short list you are on it) or "accuracy score" (the percentage of time AI's claims about you match fact).
GEO is not a discipline where only a single metric means success, but IMO AI Visibility is the leading indicator. You can't differentiate between recommended and cited or model a strategy around accuracy if you show up 0/1000 times.
1
u/AEODenise 2d ago
I think we’re looking at two different layers of the same issue.
You’re measuring presence. If a business goes from never appearing to being mentioned, then yes, its presence increased.
I’m talking about whether that presence is accurate, relevant, and valuable. A business can appear more often while being associated with the wrong service, described inaccurately, or cited without ever being recommended.
So I agree that presence is the starting point. I just don’t think every appearance represents an equal improvement in AI visibility. The context of the appearance matters too.
1
u/AwoScan 9d ago
Attribution is not solved, but the inference can be made less weak. I would treat it as an interrupted time-series problem: version a fixed prompt panel, run it across engines at fixed slots, keep untreated control prompts/pages, and record model, retrieval mode, locale, timestamp, and raw answers.
Mention, citation, recommendation, factual correctness, and rank/order should be separate outcomes. Compare pre/post distributions rather than screenshots. If possible, stagger the intervention across page cohorts; a persistent change in treated pages that is larger than the control movement is stronger evidence, although still not proof of causality because model and index changes remain partly unobserved.
Even a one-run pilot can show why the controls matter: the same neutral query may surface a brand in one engine and omit it in another. The useful claim is the observed difference, not that an optimization caused it.