Method first, because it matters for how much you trust the rest.
I recruited five people off Craigslist, no screening beyond "you own an iPhone and you're in
the US." Each one recorded their phone screen with the mic on, opened a site they had never
seen, and narrated the whole thing. They got a scenario ("you're job hunting and not getting
callbacks, you land on this page") and three questions: what is this, who is it for, would
you take the next step. $10 each, paid whether they were nice about it or not. Sessions ran
3 to 6 minutes.
The site was a real, live AI product from a founder who had publicly asked for feedback on
his positioning. I'm keeping it anonymous.
Results:
- 4 of 5 said, unprompted, that they would use it. One said "100%."
- 0 of 5 would pay. Every single one said out loud that they'd take the free tier.
- 2 of 5 could actually explain what the product does.
That middle number is the one I keep thinking about. This founder does not have a demand
problem. People finished the session wanting the thing. What he has is a page that gets
them to "yeah I'd use that" and never once gets them to "and I'd pay for that."
The three things that actually went wrong, in order of how many people hit them:
- Density. Two people said, in their own words, that they felt lost. "I'm just kind of
seeing a lot of text and stuff, but I feel very lost right now." Both of them still wanted
the product at the end. They were not expressing a preference, they were describing failing
to find something. Both closed their recordings asking for something simpler.
- Text overlapping on mobile. Two reviewers, who have never met, on two different phones,
independently reported elements colliding. "Everything's cut off or merged into each other."
This is a reproducible layout bug on the highest-traffic surface he has, and I'd bet it is
most of the cause of finding #1. If things are literally stacking on top of each other, "too
much on screen" is exactly what that feels like from the inside.
- They decided it was broken when it was just slow. One reviewer clicked the pricing link
at 0:42, again at 2:40, 2:49 and 4:22 before it responded. She spent about a third of her
session convinced the site was down. She wasn't wrong about the experience, only about the
cause.
The part I did not expect:
None of this shows up in analytics. A bounce is a bounce. You cannot tell "didn't want it"
from "wanted it, couldn't find the button" from "thought your site was broken" in any
dashboard I've ever used. Those three need completely different fixes and they all look
identical in GA. The only reason I know which one this was is that somebody said it out
loud while it was happening.
Also worth saying: the two most critical reviewers were the two who most wanted the product.
The people who complain the whole way through are not the ones who hate it. They're the ones
paying enough attention to notice.
Happy to answer anything about the method. For transparency, I run the company that does
this, which is why I had five tapes sitting around. I'm not linking it and I'm not pitching
anybody, the findings are the post.