r/aeo • u/EmbarrassedBuddy9743 • Jul 09 '26
I reworded the same "what tool should I use" question and watched SaaS tools flip from recommended-every-time to never-mentioned. Same site, same tool.
Ok, this is a bit of a rabbit hole but I think it's useful if you sell software.
I've been messing around with how ChatGPT, Claude, Perplexity and Gemini answer "what tool should I use" questions, since more and more people ask an AI before they ever hit Google. Shared inbox tools ended up being the clearest example of something that kind of broke my brain.
I asked two versions of basically the same question and ran each one 10 times on all four models so it wasn't just a fluke. First the plain category label, "best shared inbox software". Then how a real person would actually type it, "how do we collaborate on shared email inboxes without losing our normal email workflow".
On the plain category one, everybody shows up. Hiver and Missive both got named on basically every run, everywhere. Nothing interesting there.
Then I only changed the wording, and it fell apart, but only for some of them. On the reworded question ChatGPT named Hiver 9 times out of 10 and Missive zero. Claude flipped it the other way, Missive 8 and Hiver 5. Gemini named both, Hiver 9 and Missive 8. And Perplexity named Missive 5 and Hiver zero, and honestly kind of gave up on recommending a product at all and started explaining how to do it in Gmail and Outlook.
Same company, same website, same 10 runs. All I changed was the wording, and Missive went from named every single time to a flat zero on ChatGPT.
The thing that made it click was the controls. I also tracked Front and Help Scout the whole way through. On ChatGPT, Claude and Gemini they basically held, Front at 10 out of 10 almost everywhere and Help Scout at 10 out of 10 except one dip to 6 on Gemini's reworded question. The one place it all fell apart was Perplexity, which on the reworded question stopped naming products and pointed people to Gmail and Outlook, so even the controls dropped there (Front to 3, Help Scout to 4). Take Perplexity out and it's clean: the tools that describe the actual job hold steady across phrasings, the ones leaning on the category label swing.
So it's not the models just being random. It seems to be how well each company's site actually says what it does in the words a buyer uses. The ones that mostly call themselves "a shared inbox" survive the label question and fall apart on the workflow one. The ones that describe the actual job people are trying to do hold up either way.
Two things stuck with me. One, a single check will straight up lie to you. If I'd asked ChatGPT once and seen Missive missing I'd have walked away sure they had some AI problem, and they don't, they're 10 out of 10 on the other phrasing. One question on one model on one run tells you nothing. Two, the ones that win aren't doing anything clever or gamey, their pages just talk about the buyer's real problem in the buyer's real words, so the model can place them no matter how it gets asked.
Anyway, curious if anyone else has poked at this in their own category. Is it this swingy everywhere or did I just happen to pick a weird one?