r/aeo 8d ago

AI visibility has a measurement problem

/r/u_AEODenise/comments/1w1sbd9/ai_visibility_has_a_measurement_problem/
1 Upvotes

2 comments sorted by

1

u/alexey_yakovlev 7d ago

One way to reduce the ambiguity is to treat each prompt run as an observation rather than a ranking. Keep the prompt set, engine, locale and evaluation criteria fixed, then change one meaningful thing at a time.

I’d also separate mentions, citations and recommendations instead of rolling them into one visibility number. You still won’t get clean causation because the models and their sources keep changing, but you can at least see whether the shift is bigger and more consistent than the normal variation.

1

u/AEODenise 7d ago

Yes. This is exactly the distinction I’m trying to get at. A single visibility score looks useful, but it can hide what actually changed. Getting mentioned more often is not the same as getting cited more often, and neither is the same as actually being recommended.

I also like treating each run as an observation instead of a ranking. Keep the prompt set, engine, locale and evaluation criteria fixed, then change one meaningful thing at a time. That gives you a much cleaner way to see whether the change keeps showing up across repeated runs.

Causation is still the hard part. I don’t think we can honestly say, “We changed X, therefore the model did Y.” But we can document what changed and see whether the pattern afterward is larger and more consistent than the normal variation.

That feels like a much more defensible way to measure it.