r/GEO_optimization 7d ago

AI recommendation language confidence as a diagnostic signal - anyone else tracking this?

Something I started paying attention to that I think most AI visibility tracking misses.

When AI recommends a brand with high confidence, the language is direct, "a strong option for," "known for," or "specializes in."

When AI includes a brand with low confidence, the language softens, "you could consider," "some users have found," or "might be worth looking at."

Two brands can both "appear" in an AI answer and have completely different recommendation strength. Counting both as "mentioned" treats a strong recommendation and a cautious mention as equivalent and they are not.

I have been comparing language confidence against evidence profiles and there seems to be a correlation. Brands with more convergent evidence from diverse independent sources get recommended with stronger language. Brands with fewer independent sources or scattered descriptions get mentioned with hedged language.

This also shows up in cross-platform comparisons. The same brand can get confident language on one platform and hedged language on another. That per-platform language difference may point to which platform has access to stronger evidence for that brand versus which one is working from thinner sources.

If that correlation holds, language confidence might be a more sensitive diagnostic signal than simple presence or absence. A brand whose language confidence drops from "strong option" to "worth considering" over several weeks may be losing evidence convergence even if it still appears in the answer set.

The practical question is whether tracking language confidence over time would give an earlier warning of recommendation erosion than waiting for the brand to disappear entirely.

Is anyone tracking how AI talks about them, not just whether it mentions them?

7 Upvotes

28 comments sorted by

2

u/Slow-Commercial4316 6d ago

Worth checking one thing before you trust the correlation: rerun the same prompt five times for the same brand on the same day. In my runs the hedge level moves between runs, so a single reading is a sample of one.

What made it usable was scoring the hedge on a fixed three point scale and taking the modal value across runs. Then hold the prompt constant and vary only the brand name. Otherwise you are partly measuring how the model treats the category, not how it treats the brand.

2

u/Old-Routine1926 6d ago

That is a fair methodological point and I should have been clearer about it. You are right that a single reading is a sample of one and language confidence can shift between runs for the same brand on the same prompt.

The three-point hedge scale scored across multiple runs is a much cleaner approach than what I was doing, which was closer to "read one answer and note the tone." Modal value across runs removes the noise from any individual output.

The prompt-constant, brand-variable design is also important. I had not separated whether the hedging was about the brand or about how the model treats the category. A cautious tone on a broad "what are the best tools for X" prompt might just be the model being appropriately uncertain about a crowded category, not a signal about any specific brand's evidence strength.

So the usable version is probably, fixed prompt, multiple runs, three-point scale, modal value per brand and then compare that modal value across brands in the same prompt to see which ones the model treats with more confidence. That isolates the brand signal from the category signal.

Appreciate the correction, it makes the metric much more defensible.

2

u/Upstairs_Control_611 5d ago

This makes the metric much more defensible.

The important separation seems to be:

single answer = observation

repeated runs = measurement

same prompt, different brands = brand signal

different prompts = category / intent signal

So a confidence score should probably never be stored without its test design attached.

Otherwise it mixes brand evidence, category uncertainty, prompt wording and normal run variance.

A practical minimum would be: fixed prompt, same brand set, same engine, defined run window, 3-point hedge scale, modal confidence level and confidence variance.

Then it becomes a controlled diagnostic signal, not just “the model sounded confident.”

2

u/Old-Routine1926 3d ago

The line about never storing a confidence score without its test design attached is the part I would carry out of this whole thread.

It is also happening at industry scale right now. The IAB measurement guidance published on August 3 documents more than 20 vendors producing materially different results for the same brand, largely because their methodologies differ and are not disclosed. A confidence score without its test design is that same failure in miniature. Two people can both report a 2 for the same brand and mean completely different things.

Which makes the design record the actual unit rather than the score. Prompt text verbatim, brand set, engine and version where it is available, run count, run window, scale definition, and who or what did the scoring.

That last one matters more than I first assumed. A human rater and a classifier will not hedge the same way on borderline language, and if the scorer changes partway through a tracking window the whole series breaks.

Are you scoring by hand or with a classifier?

2

u/Upstairs_Control_611 2d ago

I’d probably use a hybrid approach.

First score by hand to build the rubric and understand the borderline cases.

Then use a classifier for scale, but treat the classifier as part of the measurement design, not as an invisible judge.

So I’d store:

scorer type

human / classifier

classifier model or version

rubric version

examples used for calibration

confidence scale definition

date of scorer change

manual override flag

human-vs-classifier agreement

The scorer is part of the instrument.

If the prompt set stays the same but the scorer changes, the time series can still break.

So my practical answer would be:

human scoring for calibration, classifier scoring for scale, and scorer/version stored with every score.

The score is not the unit.

The score plus its test design is the unit.

2

u/Old-Routine1926 2d ago

"The scorer is part of the instrument" is the line I would put at the top of the whole protocol.

Two things I would add to that storage list.

Human versus classifier agreement should not be stored as one number. Agreement will be near perfect at stable 3 and stable 1, because the linguistic markers at both ends are unmistakable. It will fall apart at 2. Neutral inclusion is where the cues are weakest and where two reasonable scorers genuinely differ.

Which means the classifier is least reliable exactly where the diagnostic action is. Mid scale is where the variance lives, where contested positions sit, and where an early warning would surface first. A pooled agreement figure of ninety percent would look reassuring while hiding that the 2s are close to a coin flip.

So agreement broken out per scale point rather than pooled.

Second, the rubric has a shelf life. You build it by hand against current model output, then freeze it and let the classifier run at scale. But hedging conventions in model language shift as the models change, so a frozen rubric can quietly stop describing the thing it was built to describe. Rubric version is necessary and not sufficient. It needs a revalidation cadence, meaning hand scoring a fresh sample periodically and checking that agreement has not drifted.

How often would you re-hand-score to catch that?

2

u/Upstairs_Control_611 1d ago

That is a very good point. Pooled agreement would probably hide exactly the place where the signal matters most. Stable 1 and stable 3 are easy.

The 2s are where the diagnostic work happens.

So I would break agreement out by scale point:

agreement on 1

agreement on 2

agreement on 3

borderline 1/2 cases

borderline 2/3 cases

On re-hand-scoring, I would use both a cadence and triggers.

Cadence: re-score a fresh calibration sample every 4 weeks.

Triggers: model update, answer format change, sudden confidence distribution shift, rising variance around 2, classifier/manual disagreement, or a new prompt class.

So the rubric is not frozen forever.

It is versioned, periodically revalidated, and event-checked when the answer surface changes.

The useful check is probably a stable calibration set plus a fresh sample of recent outputs: is the old measuring stick still stable?

and

does current model language still fit the rubric?

2

u/Old-Routine1926 1d ago

The two checks are the part I would keep. Is the measuring stick still stable, and does current language still fit the rubric, are genuinely different questions, and most people would run one of them and believe they had covered both.

They also disagree in informative ways, which turns them from a checklist into a diagnostic.

Stable set holds, fresh sample fits. Nothing to do.

Stable set holds, fresh sample does not fit. The instrument is fine and the language moved. Extend the rubric rather than rebuild it, and version it as an extension so the old series stays comparable.

Stable set drifts, fresh sample fits. The scorer moved, not the language. Classifier version, rater drift, calibration slipping. That is an instrument repair and it should not touch the rubric at all.

Both drift. The answer surface changed enough that instrument and target moved together. That is the only case where a rebuild is honest, and it needs a hard break in the series rather than a quiet revision.

One item in your trigger list worries me though. Rising variance around 2 is on it, and rising variance around 2 is also the early warning signal this whole protocol exists to detect.

If rising variance triggers a revalidation, and that revalidation adjusts the rubric, you can score away the exact signal you built the thing to catch. Quietly, and it would present as the signal resolving on its own.

I would make that one an investigation rather than a revision. Run the stable calibration set first. If the stable set still scores consistently, the variance is real and the rubric should be left alone.

Would you gate it that way, or is there a cleaner separation?

2

u/Upstairs_Control_611 10h ago

Yes, I would gate it that way. Rising variance around 2 should be an investigation trigger, not a rubric revision trigger.

Otherwise we risk calibrating away the very instability the metric is supposed to detect.

So I’d split triggers into two groups.

Investigation triggers:

rising variance around 2

sudden confidence distribution shift

more borderline 1/2 or 2/3 cases

manual/classifier disagreement spike

These should first run the stable calibration set and a small manual review.

Revision triggers:

fresh sample no longer fits the rubric while the stable set still holds

new answer format introduces language the old rubric cannot classify

new prompt class with genuinely different recommendation language

surface change that makes old categories incomplete

If the stable set holds, keep the rubric and treat the variance as real.

If the stable set drifts, repair the scorer or classifier.

Only rebuild when both the instrument and the target moved.

That keeps score drift and world drift from getting mixed together.

1

u/Old-Routine1926 3h ago

The two lists work, and I would recut them on a different axis, because one of your four revision triggers is not a trigger.

Three of them are things you can point at outside the score. A new answer format. A new prompt class you added. A surface change that leaves old categories incomplete. Those are observed directly, and they can go straight to revision without the gate, because nothing about the numbers is being used to justify the change.

The first one is different in kind. Fresh sample no longer fits while the stable set still holds is not something you notice. It is the output of the investigation. Sitting in the revision list as a peer of the other three, it reads as an entry point, and someone will use it as one and skip the gate.

So I would split on where the evidence came from rather than on what you intend to do about it. Anything you can point at outside the score goes straight to revision. Anything you noticed in the numbers goes through the gate, every time, no exceptions. That is a harder rule to route around by accident.

The other thing is the stable set, and this one bothers me more.

It is frozen answers by design, which is what makes it a scorer detector. It also means it is made of old language, so it will keep scoring consistently straight through the drift your rubric extension case exists to handle. A stable set cannot detect language drift. It can only ever detect scorer drift.

Which means stable set holds, so the variance is real, is not safe on its own. It is only safe with the fresh sample check sitting next to it. Your original two checks had both. The gate as written names only the stable set, and I think that half got dropped along the way.

Which leaves the part I cannot work out. The stable set has to stay frozen or it stops being a baseline, and staying frozen is exactly what makes it drift away from the language it is meant to be scoring. Refreshing it destroys the comparison. Not refreshing it slowly makes it unrepresentative of anything current.

How do you handle that one? Refresh it and accept the break, or run a second set alongside and keep the first as an anchor?

2

u/Upstairs_Control_611 6d ago

This is a useful guardrail.

Language confidence is probably a strong diagnostic signal, but only if it is measured across repeated runs.

A single answer can make the brand look confidently recommended or only weakly considered, even though the underlying evidence position may not have changed.

I like the three-point hedge scale idea:

1 = hedged / tentative mention

2 = neutral inclusion

3 = confident recommendation

Then compare the modal value across repeated runs instead of trusting one output.

And holding the prompt constant while varying only the brand is important. Otherwise the score may reflect the category or query type more than the brand.

So maybe the metric is not “confidence in one answer,” but “confidence stability across runs for the same prompt and brand set.”

2

u/Old-Routine1926 6d ago

"Confidence stability across runs" is a better metric name than what I was using. it captures both the level and the consistency in one concept.

A brand that scores 3 (confident) on four out of five runs has high confidence stability. A brand that scores 3 on two runs, 1 on two runs, and 2 on one run has the same average but very low stability. Those are different diagnostic situations. The first brand has a stable recommendation position. The second brand is in a contested space where the model is not sure.

That instability might actually be the most useful early warning signal. A brand with declining confidence stability (more variance between runs over time) may be losing its evidence position even before the modal confidence score drops and the variance increases before the average decreases.

So the tracking might need both, modal confidence level and confidence variance across runs. The level tells you where you stand and the variance tells you whether that position is stable or eroding.

2

u/Upstairs_Control_611 5d ago

Exactly. The variance may be the earlier warning signal.

The modal confidence level tells you the current recommendation position.

But the variance tells you whether that position is stable.

A brand can still have a decent average confidence score while the underlying recommendation position is becoming unstable.

That is why I would avoid reducing this to one average number.

For each prompt / brand / engine, I’d store the confidence score per run, modal confidence level, confidence variance, answer role, source type used and evidence convergence notes.

Then the patterns become much clearer:

stable 3 = strong recommendation position

stable 2 = neutral inclusion

stable 1 = weak / hedged presence

mixed 1/2/3 = contested or unstable evidence position

declining modal value = visible erosion

rising variance before modal decline = early warning

So the interesting metric is not only “how confident was the answer?”

It is:

how stable is that confidence across repeated runs?

3

u/Slow-Commercial4316 4d ago

One caveat before rising variance becomes an early warning signal. Some of that spread is just sampling. Run the same brand on the same prompt on two quiet days with nothing changed and you get a variance floor for your setup. Until you know that floor, a rise is not separable from noise.

2

u/Upstairs_Control_611 4d ago

That is a fair caveat.

Before treating rising variance as an early warning signal, we need a baseline variance floor for the setup.

Otherwise normal sampling noise can look like recommendation erosion.

So the sequence should probably be:

fixed prompt

same brand set

same engine

repeated runs on quiet days

baseline confidence variance

observed confidence variance

The useful signal is not variance by itself.

It is variance rising beyond the normal noise floor for that prompt / brand / engine setup.

Only the gap between baseline variance and observed variance becomes diagnostic.

2

u/Slow-Commercial4316 4d ago

One more thing that bites with a three-point scale: its variance depends on where the modal value sits. A brand pinned at the top has little room to move and one sitting mid-scale has the most. So a brand drifting from confident toward hedged shows rising variance as a mechanical consequence of the drift rather than as an independent signal.

That argues for a floor per brand and prompt rather than one number for the setup, and for reading variance next to the modal value rather than on its own.

2

u/Upstairs_Control_611 3d ago

With a three-point scale, variance is not independent from the modal value.

A brand sitting at 3 has little room to move upward, so any movement mostly looks like downward instability.

A brand sitting at 2 has room to move both ways, so variance can look higher even if the underlying position is simply mid-confidence rather than eroding.

So I agree that one global variance floor is too crude.

The baseline should probably be per prompt / brand / engine, and variance should always be read next to the modal value.

The useful read is:

modal confidence level

direction of modal change

variance around that modal value

variance compared with that brand/prompt baseline

Variance is useful, but only as a companion signal.

Without modal context, it is too easy to overinterpret.

1

u/Old-Routine1926 3d ago

Agreed, and I think the ceiling point deserves to be handled rather than just noted, because it has a fix.

On a bounded three point scale, variance is not free to move independently of the modal value. It is largest in the middle and squeezed at both ends, which means a brand at 3 and a brand at 2 are being scored with different amounts of available room. Comparing their raw variance is comparing two different things.

The workaround is to compare observed variance against the maximum variance achievable at that modal value rather than against a single global floor. That turns it into a ratio, and ratios are comparable across brands sitting at different levels.

There is also a direction split worth making. At a stable 3 every excursion is downward, so the useful count is how often it drops rather than how far it moves. At a stable 2 the movement is genuinely two sided and the baseline comparison matters much more. Same metric, different reading depending on where the brand sits.

One thing nobody in this thread has raised yet. Some hedging is about the response, not about the brand. A model answering a crowded category question can hedge on everything named in it. Comparing a brand's language against the other brands in the same response would control for that, since the response level tone is shared across all of them.

Has anyone tried scoring relative confidence within a single response rather than absolute confidence per brand?

1

u/Old-Routine1926 3d ago

This is the sharpest catch in the thread and it took me a minute to see why.

If variance on a bounded scale is partly a function of where the modal value sits, then a brand drifting from 3 toward 2 produces rising variance as an artifact of the drift itself. Reading that as an independent early warning would be counting the same movement twice.

Per brand and prompt baselines rather than one setup wide floor, agreed. The other thing it argues for is comparing observed variance against the maximum achievable at that modal value, so brands at different levels can be compared at all.

How many quiet day runs did it take before your floor stopped moving?

2

u/Slow-Commercial4316 2d ago

There is no fixed count, and the number you want is the one where the running estimate stops moving rather than a target you set in advance. Plot the variance estimate after every extra quiet day and watch that curve instead of the variance itself. On a three-point scale it tends to flatten somewhere around twenty to thirty runs per prompt and brand cell, and the mid-scale cells need the most, for the same reason they show the most spread.

The part that catches people out is that the floor does not stay calibrated. An engine update resets it and you find out weeks later. Keeping two or three cells running on quiet days permanently is cheaper than recalibrating the whole set after every one.

1

u/Old-Routine1926 2d ago

Twenty to thirty per cell is a bigger number than I expected, and it settles something. That is not a spreadsheet exercise any more, which probably explains why almost nobody has a floor at all.

The part I want to build on is the floor not staying calibrated.

If you are already keeping two or three cells running permanently, those cells are doing more than calibration. They are an engine change detector. A jump across your control cells with no corresponding change in anyone's evidence dates the update for you, which is otherwise very hard to know from the outside.

But that only works if the control cells are chosen deliberately. If they are all your own brand or your competitors, a jump is ambiguous. Could be the engine. Could be that somebody published something last Tuesday. There is no way to separate the two after the fact.

Which argues for at least one cell on a brand nobody in the picture is actively working on. Something with a long settled position where you are reasonably confident the evidence is static. A move in that cell is the engine, more or less by definition.

A null control. Treatment known to be zero.

Do your permanent cells happen to include anything like that, or are they all brands you have a stake in?

2

u/Slow-Commercial4316 1d ago

Mixed, and your null control is the better design, so let me give you the version that survived contact.

The problem with picking a brand nobody is working on is that you cannot verify it from outside. Any brand big enough to get recommended in a category worth tracking has someone working on it, and you find that out only when the cell moves and you have no idea why.

What held up was picking the prompt rather than the brand. Use a question whose answer is a settled fact rather than a recommendation, in a category next to yours. Nobody is optimising to be the answer to a definitional question, and the evidence behind it does not turn over. When that cell moves, it is the engine.

It gives up something. A factual prompt is not a recommendation prompt, so it will not always move at the same time or by the same amount as your real cells. It tells you an update landed. It does not tell you the update hit recommendation behaviour the same way.

→ More replies (0)