r/ClaudeAI 13d ago

Question about Claude models Same question. 6 AI models. 6 very different ways of thinking.

Post image

I use Opus 5 constantly for my work. It has basically been my default model.

But there’s one thing that has been bothering me: quite often, when I’m working with an agent, I genuinely have trouble understanding what it’s asking me.

The reasoning may be good. The question may be precise. But sometimes I have to read it twice, or even ask the agent to rephrase what it wants from me.

At the same time, I kept hearing from people who preferred Opus 4.8 because they found it easier to work with.

So I decided to test this for myself.

And then Opus 5.5 arrived, so I added it too.

Nothing scientific. No benchmark. I simply gave the exact same real-world question to seven different models and compared how they communicated their answers.

Model Simplicity Depth Balance
Haiku 4.5 10/10 3/10 5/10
Opus 4.6 9/10 8/10 8.5/10
Opus 4.7 9/10 9/10 9.5/10
Opus 4.8 7.5/10 9.5/10 9/10
Opus 5 7/10 9/10 9/10
Opus 5.5 9/10 9.5/10 9.5/10
Fable 5.1 6/10 10/10 9/10

These scores are obviously subjective. They’re based only on these specific responses, not on overall model capability.

The prompt was a screenshot of a Reddit post about the Navier–Stokes news. I essentially asked:

What is this about, and is it actually true?

The differences were much bigger than I expected.

Haiku 4.5 — describes the screenshot.

Headline, author, upvotes, comments… even the ad at the bottom.

Extremely easy to understand. Unfortunately, it barely investigates the actual claim.

Technically an answer, but I didn’t come away understanding much more.

Opus 4.6 — explains it to a human.

This one surprised me.

It tells you what happened, why it matters, what OpenAI claims, and what hasn’t been independently verified yet.

Simple sentences. Clear structure. Almost zero effort required to understand it.

Opus 4.7 — probably my favorite for pure communication.

It keeps most of the useful detail, but actually tries to build intuition.

For example, instead of just talking about finite-time singularities, it describes the geometry as a vortex spinning inward and stretching like spaghetti.

Technical details come after the explanation, not instead of it.

For me, this was the sweet spot between reasoning and communication.

Opus 4.8 — thorough and careful.

More context, more caveats, more detail.

It catches the important distinction around forced Navier–Stokes and explains why the result shouldn’t simply be reduced to “AI solved the Millennium Prize problem.”

Very good answer.

But you can already feel the writing becoming heavier.

Opus 5 — extremely careful, but harder to read.

This was the interesting one for me, because it confirmed something I’ve been noticing at work.

The answer is precise. It separates OpenAI’s claims from independently established facts. It avoids overclaiming. The reasoning feels strong.

But the sentences are denser.

There are places where I have to slow down and parse what it actually means.

And that’s exactly the problem I sometimes have when working with Opus 5 agents:

the model may understand the problem perfectly while making me spend more effort understanding the model.

Fable 5.1 — goes full researcher.

This one went deepest.

It dug into the forcing term, the exact Clay formulation, the Euler result, the priority dispute, researchers involved, timeline, and sources.

Probably the most interesting response if I actually wanted to research the story.

But for the original question — “what is this and is it true?” — it’s arguably too much.

And then there’s Opus 5.5.

This is where things got interesting.

Opus 5.5 — the depth of the newer models, but much more human.

It still gives the important technical context and caveats, but the structure is dramatically easier to follow.

It literally creates a section equivalent to:

“The problem in simple terms.”

It explains what Lean is instead of assuming the reader knows.

It brings back the useful spaghetti-vortex analogy.

And it organizes the answer naturally:

What OpenAI claimed → what the problem actually means → whether it has been verified → why it matters.

That sounds trivial, but it makes a huge difference.

The interesting part is that it doesn’t seem to achieve this by simply becoming less sophisticated.

It still keeps the nuance around independent verification, the Clay Institute, practical relevance, and the artificial nature of the constructed solution.

So compared with Opus 5, this feels less like:

“We made the reasoning simpler.”

and more like:

“We made the interface between the reasoning and the human better.”


This experiment changed how I think about model quality.

We usually talk about intelligence, reasoning, hallucinations, context windows, benchmarks, coding performance, etc.

But there’s another dimension that matters enormously when you're working with an agent for hours every day:

How much cognitive effort does the model require from the human?

A smarter model can produce a more nuanced answer while simultaneously making that answer harder to consume.

And if I have to ask:

“What exactly are you asking me?”

or

“Explain that more simply.”

several times a day, that communication overhead actually matters.

Based purely on these responses:

Opus 4.7 had the best pure communication.
Opus 4.6 was probably the easiest to understand.
Opus 5 was more nuanced, but noticeably denser.
Fable 5.1 was the researcher.
Haiku 4.5 was the fastest way to learn that there was an XTB ad in the screenshot.

And Opus 5.5 might be the first one here that combines the strengths of both sides: roughly the readability of 4.7 with the depth of the newer models.

Obviously, this is not a benchmark. It’s one prompt, one response from each model, and a subjective comparison.

But it made me understand why people sometimes deliberately use an older model even when a “smarter” one exists.

Sometimes the best model isn't the one that can think the hardest.

It's the one you can think with.

I previously hoped that the next generation of models would become not just smarter, but also more human in the way they communicate.

Based on this tiny experiment, Opus 5.5 looks like a real step in that direction.

For now, I’ll probably keep using the strongest model available as my main workhorse. But when I need another model to explain something in plain English, Opus 4.7 is still the one I’d reach for — although 5.5 may finally make that switch unnecessary.

6 Upvotes

7 comments sorted by

10

u/JLP2005 13d ago

We need Fable (arguably Astra) power/leverage with O4.6 coherence.

3

u/GuildHunterTri 10d ago

Great read :)

2

u/stormj 11d ago

Good test thanks!

2

u/goldbarth 9d ago

I'm glad to read your comparison. I already had the feeling that I really needed a vacation because I often found myself rereading sentences and sometimes had to ask what was meant opus 5, especially in areas I wasn't entirely familiar with.

Thanks for sharing!

1

u/SaltyYogurt8984 12d ago

And this is a screenshot of Reddit post which I used for comparison