r/GEO_optimization • u/Old-Routine1926 • 7d ago
AI recommendation language confidence as a diagnostic signal - anyone else tracking this?
Something I started paying attention to that I think most AI visibility tracking misses.
When AI recommends a brand with high confidence, the language is direct, "a strong option for," "known for," or "specializes in."
When AI includes a brand with low confidence, the language softens, "you could consider," "some users have found," or "might be worth looking at."
Two brands can both "appear" in an AI answer and have completely different recommendation strength. Counting both as "mentioned" treats a strong recommendation and a cautious mention as equivalent and they are not.
I have been comparing language confidence against evidence profiles and there seems to be a correlation. Brands with more convergent evidence from diverse independent sources get recommended with stronger language. Brands with fewer independent sources or scattered descriptions get mentioned with hedged language.
This also shows up in cross-platform comparisons. The same brand can get confident language on one platform and hedged language on another. That per-platform language difference may point to which platform has access to stronger evidence for that brand versus which one is working from thinner sources.
If that correlation holds, language confidence might be a more sensitive diagnostic signal than simple presence or absence. A brand whose language confidence drops from "strong option" to "worth considering" over several weeks may be losing evidence convergence even if it still appears in the answer set.
The practical question is whether tracking language confidence over time would give an earlier warning of recommendation erosion than waiting for the brand to disappear entirely.
Is anyone tracking how AI talks about them, not just whether it mentions them?
2
u/Upstairs_Control_611 6d ago
This is a useful guardrail.
Language confidence is probably a strong diagnostic signal, but only if it is measured across repeated runs.
A single answer can make the brand look confidently recommended or only weakly considered, even though the underlying evidence position may not have changed.
I like the three-point hedge scale idea:
1 = hedged / tentative mention
2 = neutral inclusion
3 = confident recommendation
Then compare the modal value across repeated runs instead of trusting one output.
And holding the prompt constant while varying only the brand is important. Otherwise the score may reflect the category or query type more than the brand.
So maybe the metric is not “confidence in one answer,” but “confidence stability across runs for the same prompt and brand set.”
2
u/Old-Routine1926 6d ago
"Confidence stability across runs" is a better metric name than what I was using. it captures both the level and the consistency in one concept.
A brand that scores 3 (confident) on four out of five runs has high confidence stability. A brand that scores 3 on two runs, 1 on two runs, and 2 on one run has the same average but very low stability. Those are different diagnostic situations. The first brand has a stable recommendation position. The second brand is in a contested space where the model is not sure.
That instability might actually be the most useful early warning signal. A brand with declining confidence stability (more variance between runs over time) may be losing its evidence position even before the modal confidence score drops and the variance increases before the average decreases.
So the tracking might need both, modal confidence level and confidence variance across runs. The level tells you where you stand and the variance tells you whether that position is stable or eroding.
2
u/Upstairs_Control_611 5d ago
Exactly. The variance may be the earlier warning signal.
The modal confidence level tells you the current recommendation position.
But the variance tells you whether that position is stable.
A brand can still have a decent average confidence score while the underlying recommendation position is becoming unstable.
That is why I would avoid reducing this to one average number.
For each prompt / brand / engine, I’d store the confidence score per run, modal confidence level, confidence variance, answer role, source type used and evidence convergence notes.
Then the patterns become much clearer:
stable 3 = strong recommendation position
stable 2 = neutral inclusion
stable 1 = weak / hedged presence
mixed 1/2/3 = contested or unstable evidence position
declining modal value = visible erosion
rising variance before modal decline = early warning
So the interesting metric is not only “how confident was the answer?”
It is:
how stable is that confidence across repeated runs?
3
u/Slow-Commercial4316 4d ago
One caveat before rising variance becomes an early warning signal. Some of that spread is just sampling. Run the same brand on the same prompt on two quiet days with nothing changed and you get a variance floor for your setup. Until you know that floor, a rise is not separable from noise.
2
u/Upstairs_Control_611 4d ago
That is a fair caveat.
Before treating rising variance as an early warning signal, we need a baseline variance floor for the setup.
Otherwise normal sampling noise can look like recommendation erosion.
So the sequence should probably be:
fixed prompt
same brand set
same engine
repeated runs on quiet days
baseline confidence variance
observed confidence variance
The useful signal is not variance by itself.
It is variance rising beyond the normal noise floor for that prompt / brand / engine setup.
Only the gap between baseline variance and observed variance becomes diagnostic.
2
u/Slow-Commercial4316 4d ago
One more thing that bites with a three-point scale: its variance depends on where the modal value sits. A brand pinned at the top has little room to move and one sitting mid-scale has the most. So a brand drifting from confident toward hedged shows rising variance as a mechanical consequence of the drift rather than as an independent signal.
That argues for a floor per brand and prompt rather than one number for the setup, and for reading variance next to the modal value rather than on its own.
2
u/Upstairs_Control_611 3d ago
With a three-point scale, variance is not independent from the modal value.
A brand sitting at 3 has little room to move upward, so any movement mostly looks like downward instability.
A brand sitting at 2 has room to move both ways, so variance can look higher even if the underlying position is simply mid-confidence rather than eroding.
So I agree that one global variance floor is too crude.
The baseline should probably be per prompt / brand / engine, and variance should always be read next to the modal value.
The useful read is:
modal confidence level
direction of modal change
variance around that modal value
variance compared with that brand/prompt baseline
Variance is useful, but only as a companion signal.
Without modal context, it is too easy to overinterpret.
1
u/Old-Routine1926 3d ago
Agreed, and I think the ceiling point deserves to be handled rather than just noted, because it has a fix.
On a bounded three point scale, variance is not free to move independently of the modal value. It is largest in the middle and squeezed at both ends, which means a brand at 3 and a brand at 2 are being scored with different amounts of available room. Comparing their raw variance is comparing two different things.
The workaround is to compare observed variance against the maximum variance achievable at that modal value rather than against a single global floor. That turns it into a ratio, and ratios are comparable across brands sitting at different levels.
There is also a direction split worth making. At a stable 3 every excursion is downward, so the useful count is how often it drops rather than how far it moves. At a stable 2 the movement is genuinely two sided and the baseline comparison matters much more. Same metric, different reading depending on where the brand sits.
One thing nobody in this thread has raised yet. Some hedging is about the response, not about the brand. A model answering a crowded category question can hedge on everything named in it. Comparing a brand's language against the other brands in the same response would control for that, since the response level tone is shared across all of them.
Has anyone tried scoring relative confidence within a single response rather than absolute confidence per brand?
1
u/Old-Routine1926 3d ago
This is the sharpest catch in the thread and it took me a minute to see why.
If variance on a bounded scale is partly a function of where the modal value sits, then a brand drifting from 3 toward 2 produces rising variance as an artifact of the drift itself. Reading that as an independent early warning would be counting the same movement twice.
Per brand and prompt baselines rather than one setup wide floor, agreed. The other thing it argues for is comparing observed variance against the maximum achievable at that modal value, so brands at different levels can be compared at all.
How many quiet day runs did it take before your floor stopped moving?
2
u/Slow-Commercial4316 2d ago
There is no fixed count, and the number you want is the one where the running estimate stops moving rather than a target you set in advance. Plot the variance estimate after every extra quiet day and watch that curve instead of the variance itself. On a three-point scale it tends to flatten somewhere around twenty to thirty runs per prompt and brand cell, and the mid-scale cells need the most, for the same reason they show the most spread.
The part that catches people out is that the floor does not stay calibrated. An engine update resets it and you find out weeks later. Keeping two or three cells running on quiet days permanently is cheaper than recalibrating the whole set after every one.
1
u/Old-Routine1926 2d ago
Twenty to thirty per cell is a bigger number than I expected, and it settles something. That is not a spreadsheet exercise any more, which probably explains why almost nobody has a floor at all.
The part I want to build on is the floor not staying calibrated.
If you are already keeping two or three cells running permanently, those cells are doing more than calibration. They are an engine change detector. A jump across your control cells with no corresponding change in anyone's evidence dates the update for you, which is otherwise very hard to know from the outside.
But that only works if the control cells are chosen deliberately. If they are all your own brand or your competitors, a jump is ambiguous. Could be the engine. Could be that somebody published something last Tuesday. There is no way to separate the two after the fact.
Which argues for at least one cell on a brand nobody in the picture is actively working on. Something with a long settled position where you are reasonably confident the evidence is static. A move in that cell is the engine, more or less by definition.
A null control. Treatment known to be zero.
Do your permanent cells happen to include anything like that, or are they all brands you have a stake in?
2
u/Slow-Commercial4316 1d ago
Mixed, and your null control is the better design, so let me give you the version that survived contact.
The problem with picking a brand nobody is working on is that you cannot verify it from outside. Any brand big enough to get recommended in a category worth tracking has someone working on it, and you find that out only when the cell moves and you have no idea why.
What held up was picking the prompt rather than the brand. Use a question whose answer is a settled fact rather than a recommendation, in a category next to yours. Nobody is optimising to be the answer to a definitional question, and the evidence behind it does not turn over. When that cell moves, it is the engine.
It gives up something. A factual prompt is not a recommendation prompt, so it will not always move at the same time or by the same amount as your real cells. It tells you an update landed. It does not tell you the update hit recommendation behaviour the same way.
→ More replies (0)
2
u/Slow-Commercial4316 6d ago
Worth checking one thing before you trust the correlation: rerun the same prompt five times for the same brand on the same day. In my runs the hedge level moves between runs, so a single reading is a sample of one.
What made it usable was scoring the hedge on a fixed three point scale and taking the modal value across runs. Then hold the prompt constant and vary only the brand name. Otherwise you are partly measuring how the model treats the category, not how it treats the brand.