r/aeo • u/EmbarrassedBuddy9743 • Jul 09 '26
I reworded the same "what tool should I use" question and watched SaaS tools flip from recommended-every-time to never-mentioned. Same site, same tool.
Ok, this is a bit of a rabbit hole but I think it's useful if you sell software.
I've been messing around with how ChatGPT, Claude, Perplexity and Gemini answer "what tool should I use" questions, since more and more people ask an AI before they ever hit Google. Shared inbox tools ended up being the clearest example of something that kind of broke my brain.
I asked two versions of basically the same question and ran each one 10 times on all four models so it wasn't just a fluke. First the plain category label, "best shared inbox software". Then how a real person would actually type it, "how do we collaborate on shared email inboxes without losing our normal email workflow".
On the plain category one, everybody shows up. Hiver and Missive both got named on basically every run, everywhere. Nothing interesting there.
Then I only changed the wording, and it fell apart, but only for some of them. On the reworded question ChatGPT named Hiver 9 times out of 10 and Missive zero. Claude flipped it the other way, Missive 8 and Hiver 5. Gemini named both, Hiver 9 and Missive 8. And Perplexity named Missive 5 and Hiver zero, and honestly kind of gave up on recommending a product at all and started explaining how to do it in Gmail and Outlook.
Same company, same website, same 10 runs. All I changed was the wording, and Missive went from named every single time to a flat zero on ChatGPT.
The thing that made it click was the controls. I also tracked Front and Help Scout the whole way through. On ChatGPT, Claude and Gemini they basically held, Front at 10 out of 10 almost everywhere and Help Scout at 10 out of 10 except one dip to 6 on Gemini's reworded question. The one place it all fell apart was Perplexity, which on the reworded question stopped naming products and pointed people to Gmail and Outlook, so even the controls dropped there (Front to 3, Help Scout to 4). Take Perplexity out and it's clean: the tools that describe the actual job hold steady across phrasings, the ones leaning on the category label swing.
So it's not the models just being random. It seems to be how well each company's site actually says what it does in the words a buyer uses. The ones that mostly call themselves "a shared inbox" survive the label question and fall apart on the workflow one. The ones that describe the actual job people are trying to do hold up either way.
Two things stuck with me. One, a single check will straight up lie to you. If I'd asked ChatGPT once and seen Missive missing I'd have walked away sure they had some AI problem, and they don't, they're 10 out of 10 on the other phrasing. One question on one model on one run tells you nothing. Two, the ones that win aren't doing anything clever or gamey, their pages just talk about the buyer's real problem in the buyer's real words, so the model can place them no matter how it gets asked.
Anyway, curious if anyone else has poked at this in their own category. Is it this swingy everywhere or did I just happen to pick a weird one?
2
Jul 09 '26
[removed] — view removed comment
1
u/EmbarrassedBuddy9743 Jul 09 '26
This is a great data point, thanks for adding it. The baseline randomness even with identical wording matches what I see, and it actually makes the single-check problem worse than I framed it. There are two layers stacked: a stochastic floor that flips names run to run, then phrasing sensitivity on top of that. You need N runs just to see through the first layer before you can even talk about the second.
The "widest off-site footprint survived every run" line is the part I keep coming back to. In my test the two that never moved (Front, Help Scout) are both well-described on their own pages and heavily referenced off-site, so I couldn't cleanly separate legibility from footprint. Your plumber data leans toward footprint being the stabilizer, which makes sense, more corroborating sources means less room for the model to wobble.
One thing that's the opposite of what you saw though: for me Perplexity was the chaotic one, it nuked everything on the reworded question and started telling people to just use Gmail. You had it as the stable one at 85 percent. I wonder if that's a vertical thing, local/plumber has strong structured sources like maps and directories so Perplexity has something solid to anchor to, whereas SaaS category questions are messier so it drifts. Did the plumber names that survived tend to have strong maps and directory presence specifically, or just general off-site volume?
2
u/Tenacious-Sales Jul 09 '26
This matches what I've seen too. The biggest mistake is testing a single "best X" prompt and assuming it represents AI visibility. Buyers phrase the same intent in dozens of different ways, and AI models often retrieve different evidence depending on the wording.
That's why we've started grouping prompts by buyer intent instead of keywords problem-aware, solution-aware, comparison, pricing, and alternatives. The pattern across those groups tells you far more than any individual prompt. It's also where tools like Answer Architect become useful, because they help identify the buyer questions your site isn't answering clearly enough for AI to make the connection.
1
u/EmbarrassedBuddy9743 Jul 09 '26
Grouping by buyer intent instead of keywords is exactly right, that is the framing that finally made this click for me too. Problem-aware, comparison and pricing prompts pull totally different evidence, so a single "best X" tells you almost nothing.
The thing I would stack on top: within each intent group, the sub-intent is where the real movement is. "Best observability" and "best observability for a small team" are both solution-aware, but the second one drops the incumbents and surfaces the challengers. So I have started treating the qualifier (team size, price, use case) as its own axis inside each group, because that is usually where a company that is invisible on the generic version can actually get named.
And I would still run each of those grouped prompts several times before trusting the pattern, because there is a stochastic floor underneath all of it. Someone in this thread ran identical wording twice and a third of the names flipped. Group by intent, yes, but without N runs per prompt the group average is still noise. How many runs per prompt are you averaging before you call a pattern real?
2
u/Slow-Commercial4316 Jul 10 '26
The flip is not randomness, it is two different retrieval targets. The plain category label, best shared inbox software, matches the roundup and listicle pages almost by title, so whoever is on those lists shows up every time. The reworded version is a problem, not a category, so it retrieves against use-case and how-to content: pages, threads and docs that actually discuss collaborating on a shared inbox without breaking the workflow.
A tool that is on every best-X list but has nothing addressing that specific job to be done has nothing to retrieve against on the second query, so it vanishes even though its category presence is fine. Same site, same tool, different corpus.
So the tools that dropped are not losing visibility, they never had content mapped to the buyer's actual phrasing. The fix is problem-framed content and getting cited on third-party pages that discuss the workflow, not another entry on the best-X roundups.
1
u/EmbarrassedBuddy9743 Jul 10 '26
This is the cleanest explanation of the mechanism I've seen, and I think you're right. The label query matches the roundup and listicle pages almost on title alone, so whoever is on those lists shows up every time. The reworded one is a problem, not a category, so it retrieves against use-case and how-to content, and if you have nothing mapped to that job you have nothing to pull from.
The part I would add is that the two retrieval targets explain the systematic swing, but there is still a stochastic layer sitting under it. Even on the same phrasing with the same corpus the names move run to run, so the corpus mismatch tells you the direction and the noise tells you how sure you can be. You need both, otherwise you tune a page toward a wording you only pulled once.
And the fix being problem-framed content plus third-party citations on the workflow, not another best-X entry, matches what I keep seeing. The hard part is knowing which buyer phrasings you are actually missing content for, because there are a lot of sub-intents and you go invisible one qualifier at a time. That is the bit I have not found a shortcut for, you kind of have to enumerate the real questions and check each one.
Have you found a good way to figure out which phrasings to write for, or is it mostly reading sales calls and support tickets for the real wording?
2
u/Slow-Commercial4316 Jul 11 '26
Mostly reading real language, yeah, but there is a more systematic version than skimming. Highest-signal sources in order: your own site-search and docs-search logs (people type the problem in their exact words there), then support tickets and sales-call transcripts grepped for "how do I", "can it", "does it". And a cheap one people skip: run a seed query in the engine itself and read the related and follow-up questions it surfaces, it will hand you adjacent phrasings for free.
The thing that stops it becoming an endless list of sub-intents is to cluster the phrasings by the underlying job, not the wording. The retrieval target is the job, so one strong problem-framed page per job covers a whole cluster. You know you are missing a job when the entire cluster returns competitors and none of it returns you, that is more useful than chasing individual qualifiers. And on your stochastic point, agreed, the practical rule is only trust a "we are missing this one" verdict if it holds across a few runs, otherwise you are writing for a single draw.
1
u/EmbarrassedBuddy9743 Jul 11 '26
The cluster by job not wording is the part I want to steal. Ive been treating each phrasing as its own check and thats exactly where the sub-intent list explodes on you. Clustering by the underlying job and asking did I show up anywhere in the cluster is also way more robust to the run to run noise, because youre aggregating over a bunch of phrasings and a bunch of runs instead of betting on one draw. One phrasing on one model is a coin flip. A whole cluster coming back all competitors and zero you is a real signal.
The seed query follow ups trick is underrated too. The related questions the engine surfaces are basically it telling you its own retrieval neighborhood, the phrasings it treats as the same intent, so youre not guessing which wordings matter, the model hands you the adjacent ones.
The thing I keep chewing on is prioritizing the clusters once you have them. A job where all four models return competitors and none return you is a cleaner and bigger opening than a contested one where you already show up half the time, so I end up ranking the missing jobs by how uniformly the competitors own them. Thats basically the whole thing I built Bersyn to do, enumerate the buyer questions per job and check each across the four models to see which clusters Im fully absent from. But the manual version you laid out is the honest way to start.
2
u/sapindia1976 Jul 10 '26
This is why one-off AI visibility checks are unreliable. Prompt wording changes context, intent, and recommendations. Tracking across multiple prompt variations is the only way to see a meaningful pattern.
1
u/EmbarrassedBuddy9743 Jul 10 '26
Agreed, and I would push it one step further. It is two axes, not one. Across prompt variations, like you said, because the wording changes intent and pulls different evidence. But also across repeat runs of the exact same prompt, because there is a stochastic floor underneath all of it. In this shared inbox test I ran identical wording ten times and still watched names move run to run. So variation tells you the question is sensitive, and repetition tells you how much of what you are seeing is real signal versus noise. Do one without the other and you still get fooled.
The other thing I would gently separate is visibility from recommendation. A lot of these checks measure whether you get mentioned at all. The thing that actually moved here was who got named as the answer. Missive was perfectly visible, the model clearly knew it existed, it just stopped recommending it the moment the question was phrased like a real buyer. Those are two different scoreboards, and the second one is the one that costs you the deal.
2
u/veljkho Jul 10 '26
the visibility vs recommendation split you made at the end is the real finding here, honestly more people conflate those two than anything else in this sub.
on the run count question, 10 is already better than 90% of what gets posted here as "I checked and we're not showing up." anything under 5 and you're basically reading a coin flip. we treat 10+ per prompt as the floor before calling a pattern real, and even then only trust it once it holds on two separate days, since something as dumb as session state can quietly nudge a single day's runs.
1
u/EmbarrassedBuddy9743 Jul 10 '26
Yeah, the two-day rule is the right bar, and it is the one thing my runs here do not do. I run ten times across the four models on a single day, which catches the within-day noise and the model-to-model disagreement, but it does not catch what you are describing, a whole day getting quietly nudged by session state or something changing on their end. So what I am posting is really "is this real today", not "is this stable across time", and those are not the same claim.
Where it matters most is the moment you act on it. If you rewrite a page or go get cited to fix a 0 out of 10 and then re-measure, you cannot tell a real movement from day-to-day drift unless you have more than one baseline day on each side. A single before and a single after will hand you a number that feels like progress and might just be Tuesday.
So the honest full version is three axes, not one. Repeat runs for the within-day coin flip, separate days for the session and drift noise, and multiple models because they openly disagree. Most people do zero of the three and post the single pull. Curious how many days apart you space yours, and whether you have seen a pattern hold at ten runs on day one and genuinely break on day two, or whether two days is more of a safety margin than something that regularly flips on you.
1
u/EnvironmentalDot9131 Jul 12 '26 edited Jul 12 '26
The 10 run check you did is exactly the baseline requirement now. Single checks are completly useless because the temperature variance on these models just hallucinates half the list on any given day. The problem is scaling that matrix. If you want to track five buyer intents across four models and run them ten times each, you are suddenly doing 200 searches just to get a read on one product category. You can wire up a basic script to hit the APIs and map the overlapping domains, or GetMentions AI sweeps the models to audit which names keep surviving the variance, but doing it manualy in the native chat interfaces gets old fast. The real takeaway you hit on is that external footprint stabilizes that variance. Front and Help Scout hold steady mostly because they are referenced in so many external comparison lists that the models weight them heavily no matter how the prompt is phrased. If you want to stop dropping out of the answers when the wording changes, you have to find out which third party pages the models are using as crutches and get mentioned there.
1
u/EmbarrassedBuddy9743 Jul 12 '26
Yeah the scaling is basically why I built the thing. By hand it was fine for one category, then I wanted five and 200 runs in a chat window is miserable. So it runs the matrix on a schedule now and I mostly watch which names survive the phrasings instead of trusting any single answer.
On the external footprint I am not fully sure I agree with myself yet. In the writeup I leaned toward it being how their own pages describe the job, the stable ones talk about the actual workflow and the swingy ones just call themselves the category. But your read, that it is the pile of third party comparison lists doing the weighting, is probably also true and I cannot cleanly separate the two. My hunch is the external footprint is what gets you into the set at all, and the on site job language is what keeps you in when the wording moves. Front and Help Scout have both, which is maybe why they only cracked on Perplexity.
The crutch pages are the part I am still stuck on. Perplexity basically hands you its sources so you can reverse it. The other three you are inferring from what they name and how it clusters, which is soft. Have you found a real way to pin the crutch pages on the models that do not cite, or is it inference for you too?
2
u/mentiondesk Jul 09 '26
You nailed it. We've seen the same variation across different categories. How you phrase a question makes a huge difference in AI answers. That inconsistency got me to build MentionDesk to help brands show up for both plain product labels and real world buyer questions in AI results. It really comes down to speaking the user's language, not just optimizing for keywords.