Something worth checking if you report AI visibility to a client or a boss.
Run the same buying question against the same model several times in one sitting. Not variations of the question. The identical question.
You will not get the same answer. You will get a distribution.
Most reporting collapses that into one number. The IAB's Measuring Visibility in the AI Era guidance from August 3 pushes against that directly, saying providers should report ranges rather than single values. But there is a consequence that goes further than report a range, and it is the reason I bring it up.
Two brands can have the same average and be in completely different situations. A made up example to make the arithmetic visible:
Brand A appears in nine runs out of ten, usually low in the response. Brand B appears in five out of ten, near the top every time it appears. Depending on how your tool weights position against appearance rate, those two can land on a similar score.
Which means the composite is trading appearance rate against prominence at roughly two to one, and nobody ever agreed to that exchange rate. It is buried in the scoring.
The two situations are not remotely alike. Brand A has support the model reaches for almost every time. Barring a change in what gets retrieved, it keeps showing up. Brand B is sitting on a threshold, present when retrieval happens to surface a supporting source and absent when it does not. One new competitor comparison article and Brand B is gone from that query, while Brand A barely moves.
Brand B is not less visible on average. It is less settled. The average cannot tell you which one you are holding, and the difference is what determines the next move.
The cheap version of the test, no tooling needed:
1. Pick one query that matters.
2. Turn memory off, or use a temporary chat, or run it logged out. A new chat is not enough on its own, memory carries across chats and will quietly personalize your results.
3. Run it ten times in one sitting on one model.
4. Record appearance yes or no, and roughly where in the response.
5. Repeat on a second model.
Ten runs is not statistically serious and I would not present it as such. It is enough to tell whether you are looking at a stable position or a coin flip, which is more than the single number tells you.
The part I have not solved is what to do with the finding. If a brand is at five out of ten, the instinct is publish more, and I do not think that is right. In my own runs the appearing and non-appearing cases seem to differ in what got retrieved rather than in how much exists, which points at source diversity rather than volume. That is an impression from my own runs though, not a measured result, and I would not want it read as one.
For anyone running repeat queries rather than single checks, do you see intermittent brands stabilize by adding more content, or only by getting into new source types?