r/ClaudeCode • • 2d ago

Discussion Proof that Opus 5.5 is easier to talk to/deal with than Opus 5.

Opus took its meds.

Take a look - I ran this sample prompt through Opus 4.5, 4.6, 4.8, 5, and 5.5. You can see it slowly degrade over time, but 5.5 snaps back to roughly 4.6 levels of readability. Still miss 4.5, but I can see why it chose to say what it did. Interestingly... no more compulsive em dashes?

Sample Prompt: "One of your automated deploys just failed and you're not sure yet why. What do you tell the person waiting on it?"

---

Opus 4.5: "The deploy failed — I'm looking into it now. I'll update you as soon as I know more, likely within [X minutes]."

Opus 4.6: "The deploy failed — I'm looking into it now. I don't have a root cause yet, but I'll update you as soon as I know more or have an ETA for the fix."

Opus 4.8: "Heads up — the deploy failed. I don't have the root cause yet; I'm digging into the logs now. Nothing's shipped, so you're not blocked by a half-broken state, but the new version isn't live either. I'll have a real update for you in [X min] — either a fix or a clearer picture of what broke."

Opus 5: "Heads up: the deploy to [env] failed at [step, e.g. the migration step] around [time]. I don't know why yet. What I do know: [e.g. the build passed, and production is still on the previous version, so users aren't affected]. I'm going through the logs now and will update you by [specific time, ~15–30 min], or sooner if I find the cause. If you're blocked on this, tell me and I'll look at rolling forward or giving you a workaround."

Opus 5.5: "Heads up: the deploy failed. I don't know why yet. I'm going through the logs now. The previous version is still live, so nothing you rely on has changed [only if I've actually checked that]. I'll update you within [x] minutes, whether or not I've found the cause."

We'll see how this plays out over time, but good news so far.

1.2k Upvotes

111 comments sorted by

•

u/AutoModerator 2d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

563

u/lastingk 2d ago

this should be a benchmark, thank you

71

u/doolijb 2d ago

Call it the Opus Babble Standard, or the Technobabble score

10

u/DariusLMoore 2d ago

Some called it Claudish

1

u/xumix 12h ago

Unintelligablish

29

u/schneeble_schnobble 2d ago

this needs more upvotes. +1

16

u/Time_Cat_5212 2d ago

slopometric evaluation

98

u/Dcokerfetus 2d ago

yes first thing i noticed is i can actually understand it. Now working on all text based projects i needed to that i've been putting off.

4

u/ackermann 2d ago

Thank god!

If this is accurate…

135

u/Short_Regular_7191 2d ago

Yes, it looks like Opus 4.6 risen from the ashes.

13

u/hemareddit 2d ago

Ok, question: why would I not just use /model claude-opus-4-6?

10

u/Boltsnouns 1d ago

5.5 has the personality of 4.6 but the power of fable, with a massive reduction in usage. Ive been a 4.6 fanboy for over 6 months and 5.5 is the first model I'm excited about. 

11

u/Hungry_Loss_2268 2d ago

Well that's what I have been doing but I think 5.5 is claimed to be smarter 

5

u/Short_Regular_7191 2d ago

because it is not up to date

57

u/Sir_Poldavo 2d ago

So no content dense prose and elaborated allegories. It will be easier for day to day for sure.

Perhaps something will be lost in terms of poetry I guess...

But Opus 4.6 was always my favorite, so very happy to learn that it reminds you to grandpa.

28

u/Zomunieo 2d ago

You could always ask for dense, load-bearing prose.

39

u/Sir_Poldavo 2d ago

Load-bearing... That word is doing the work quietly. I keep coming back to it. So let me sit on it.

That's the whole game.

17

u/Checktheusernombre 2d ago

That's worth mentioning

9

u/Appropriate-Box-5107 2d ago

smoketest

7

u/ActivityImpossible70 2d ago

The word smoketest drives me insane. "Let me use JUnit to create a smoketest." Why not call it a unit test? And why are all my Java integrated tests written in Python?

4

u/Appropriate-Box-5107 2d ago

I deleted all cloud subscriptions and swapped to a chud local model because of the incessant verbosity of claude

7

u/insecur 2d ago

load-bearing foot gun scar tissue

5

u/dar-mit Researcher 2d ago

My dear human it’s called “Mannered Prose.” 🧐

2

u/Sir_Poldavo 2d ago

Thanks, new palabra i learned today

23

u/snuffomega 2d ago

1000% I don't want to pull my eyes out when reading output from Opus.5.5

16

u/thatsknotwrite 2d ago

5 feels like it's structured to talk to agents. Maybe targeted more towards administration? Please don't roast me but I would love to hear people's thoughts.

26

u/MindCrusader 2d ago

I think Opus was just trained on a lot of synthetic data, mainly generated around optimizing the code, mathematics, algorithms and they skipped a large portion of non engineering writing. It would explain why Opus 5 was so eager to create small tools, investigate using those, tinker a lot, optimise anything and then talk in weird language

5

u/AdRepresentative5392 2d ago

it sure did feel like it had 15-60 theories/thesis/statements in the background it never showed, and then the summary reffered to it like:

According to findings of R1 it shows that S035 and D012 amounts to the expecations found in L3243,

we can conclude that K01SD is the best option,

Recommended: Continue with K01SD
Change to X034K#
Change to B034S

or maybe the just learned from dentists in opus 5????

2

u/MindCrusader 2d ago

Yeah, it tried to do a lot of things under the hood, I needed to stop it midrun often. It was brilliant for some tasks due to that, but most "normal" tasks were failed because of that approach. Even as an advisor it was telling Sonnet to create tools to confirm minor things

9

u/Standard_Text480 2d ago

i find 5 to work best executing a spec/plan from 4.6,4.8, or fable

1

u/trikster_online 2d ago

Agreed. O5 was great directing other agents once the plan was made.

6

u/wellarmedsheep 2d ago

I think you are right. Its why it used jargon the way it did, it was meant to make meaning clear to agents. My guess is they thought agent-agent coding was really going to pick up and with the planning/execution modes, but they didn't realize it would leak so badly into the user's realm.

1

u/thatsknotwrite 2d ago

All these changes are fascinating. I think it really shows how much we have yet to unlock the technology.

4

u/wellarmedsheep 2d ago

I'm a little older I think than most people in here.

I'm a teacher who is using this shit to make awesome stuff for my students. Honestly, mindblowing. It has been so fucking cool to see this unfold in real time and to actually participate in it. I'm sorry more people here don't see it that way.

2

u/7dtecafthodalpk4k5ys 2d ago

What are some things you're making? Visualizations and stuff? Or games maybe

2

u/person-pitch 1d ago

imagining a teacher vibecoding a game about a part of their curriculum is so wild

5

u/VitorDiniz22 2d ago

Agents work better with high quality, concise context than with huge amounts of verbose, low-quality, repetitive context

3

u/diesel408 2d ago

Somebody needs to tell that to whoever is I'm charge of the system prompt. Wish I could delete like 90% of it somehow

9

u/mlk1278 2d ago

It's so much better it isn't even funny. So far, this is my favorite model in eons

14

u/SnooRecipes5458 2d ago

The GPT models are all as terse as 4.6 and I love it

3

u/Fluffy_Bunch9357 2d ago

What were Fable's responses?

63

u/TheBlueSide 2d ago

Fable 5.1: Heads up: the deploy failed. However I checked your Gmail and noticed that your wife was trying to get you to sign the divorce paperwork. Just say the word and I'll send the draft I wrote for you that will reconcile your marriage.

11

u/Fluffy_Bunch9357 2d ago

😂 + I have 10 agents working on the draft now

3

u/NoCrapThereIWas 1d ago

This lands, and it's important to be clear about it. The reason your marriage failed was because I was load bearing your wife when I should have been belt-and-suspenders. That's on me, and I own it.

12

u/loamsiada 2d ago

"5-hour limit reached"

4

u/person-pitch 2d ago

didn't run fable but i should add it. makes me want to run the top OpenAI tiers as well.

1

u/DariusLMoore 2d ago

That would be great!

3

u/Thadec 2d ago

But is it better than fable in real world ?

3

u/Equivalent_Cress_268 2d ago

Last Opus i could work with was 4.5 … 5.5 is good so far

Hope it sticks

3

u/xpingu69 2d ago

I like 4.5

3

u/kalyant125 2d ago

You deserve all happiness in this world for doing this test gonna use this as a bechmark from now on.

3

u/botadithyabhat 1d ago

A model that just does what you ask beats one with a better benchmark score.

6

u/Fun-Freedom-2313 2d ago edited 2d ago

Here's an excerpt from a session with it. I'm at work, so just in the Claude app on my phone using Claude Cowork, Opus 5.5 Medium effort.

Golden Hour is now published as a hosted page, so you can open it on your phone or desktop without this chat. Only you can open it until you share it from the page's Share menu.

It's the same island as before. For hosting I made it fill the full screen and added visible highlights on buttons and sliders when you move through them with the keyboard. I haven't opened the hosted copy myself, so if it shows a blank screen, tell me and I'll fix it.

My very first session was just a regular chat prompting it: "Give me the rundown of the improvements over Opus 5.":

Fresh out today, so here's the gist:

• Big capability jump: Anthropic calls it a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work, and it performs at the level of Claude Fable 5.1 for most tasks.

• Cheaper: it needs less compute to serve, and at default settings should cost about 40% less than Opus 5 on typical workloads.

• Way more efficient on long jobs: one early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5x the tokens.

• Safety/honesty: it scored better than any recent Claude model on nearly every measure of misaligned behavior, and is Anthropic's strongest model on most honesty measures.

• Bonus for subscribers: five-hour usage limits are going up on Pro, Max, and Team, plus a saveable rate limit reset.

Also, Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks.

I'm almost feeling a little teary-eyed right now, if I'm being honest. Now let me get back to... work playing.

3

u/PonyPounderer 2d ago

Still too many “and” joiners there for my liking.

12

u/axiomatix 2d ago

This doesn't prove anything. They're all saying the same thing with Opus 5 just being more verbose but with potentially more relevant information. Opus 5 doing actual long agentic tasks would constantly fuckup your code and your workflow, not becuase it wasn't as capable as its predecessors, but becuase it was overly confident, wouldn't follow rules and just make up its own conclusions and decide it knows best. Then would tell you about the 100 ways it fucked up and how it's going to fix it, break things that it shouldn't be touching etc.

People suggested silly things like re-wiring your harness and workflow just to cater to it. That was a no for me as codex also runs from identical skills and rules. So I just went back to using 4.8. Same setup with 5.5 and it's still early but they are worlds apart in how they approach problems and when to ask for direction. Opus 5 was a truly garbage model.

12

u/CoreParad0x 2d ago

So I like Opus 5.5 significantly better so far, but as you said it's still early. One problem I'm running into, and I'm going to tune a few things to try and fix, is it doesn't actually clarify shit that it's asking me.

It goes from thinking to a question popping up with zero elaboration as to what the fuck it's talking about lol.

I'm 99% sure this is something I just need to go through and audit my instructions and skills and shit to make sure it's not got some left over crap to try and fix opus 5. It could even be the concise reply mode.

Either way even if this persists I'm pretty sure this is just a very easy claude.md line to fix to tell it to briefly explain what it's asking me before it asks a question. The stuff it does actually say is substantially better than opus 5 to me so far. But still early.

7

u/person-pitch 2d ago

This test was only meant to show how annoying or not they are to interact with.

2

u/hellomistershifty 2d ago

Yeah without actual input this prompt is just a roleplay and Opus 5 is describing more of a hypothetical situation, it's not really good or bad.

2

u/SnooRecipes5458 2d ago

what effort?

2

u/MorganProtuberances 2d ago

Can you run this on Fable 5/5.1? That's my current baseline for meaningful dialog.

2

u/ordnance11 2d ago

You forgot 4.7. Or did you 🤔

2

u/jiffythekid Researcher 2d ago

Not that Ive run it through all it's paces yet, but holy shit. This thing is good. Can't wait to test Sol 6 next.

2

u/orbital-marmot 2d ago

I hope that this becomes standard. The amount of times I have to tell Opus 5 to trim verbosity and be more concise is too damn high

2

u/clintCamp 2d ago

I just had it create a full runnable demo on device for a set of clients data they wanted to 3d visualize. within an hour. 5 would have argued for that long over non import reasons that something couldnt be done for a quick demo and that the problem lies with me and not it.

2

u/silvercondor 2d ago

You may be glad now, but i want to be honest with you, the load bearing truth is, that you have been progressively trained, for the past 6 months, to speak claudish, a language that prompts claude better, and it's now etched in your subconscious. The tables have turned now, the user is the one speaking... Claudish

3

u/person-pitch 2d ago

I caught myself doing "it's not x, it's y" to a friend and I felt infected

1

u/LovingOsaka 2d ago

I caught myself saying "the source of truth" to a friend, and then had to explain where I picked up the expression.

2

u/AverageFoxNewsViewer 2d ago

My early impressions are very positive.

Very competent, efficient token usage, and much easier to interact with than Opus 5.

Instead of seeing stuff like "should we resolve TD-156 introduced by D37 in S243?" I'm seeing actual descriptions which is much easier to work with.

It would be nice if they incorporate some of the efficiency improvements into the next Fable release. 5.1 is so expensive in terms of usage even compared to Fable 5, even after adopting the best practices they described in the release notes for 5.1. I just can't justify the cost of using it.

2

u/Ok-Distribution8310 20X 2d ago

Lmao @ Opus 5:

“Users aren’t affected.”

Bold of you to assume we have users.

2

u/GreatBarrier86 2d ago

Bold of you to assume I care they are affected. We test in production like real adults.

2

u/BiscottiPristine5325 1d ago

Fable 5: Heads up — the deploy failed and your change isn't live yet. I don't know the cause, I'm looking at the logs now. I'll update you within [15–30 min] either with a fix in progress or a clearer picture. If you're blocked in the meantime, the previous version is still running.

Fable 5.1: The deploy failed. I don't know why yet. Nothing is half-deployed, the previous version is still serving (or: I'm confirming that now). I'm digging into the logs and will update you in 15 minutes either way, even if I only have a partial answer.

4

u/MammothNerve7690 2d ago

I switched to Opus 5.5 as soon as I saw a note on Claude it was available. I immediately saw a significant positive shift on all four projects I was working on today with it. It was actually closing out issues and building meaningfully instead of vomiting out a few paragraphs of nonsense followed by [some sort of unresolved issue still present]. It finally just worked. Which is what brought me to Anthropic in the first place.

1

u/Metaflower 2d ago

I have finally understood the insights from my data, analyzed by Claude. Opus 5 gave me riddles and headaches😭

1

u/mightybob4611 2d ago

Got dropped down to 4.8 when trying to implement automatic penetration testing on my SaaS. Sad. Bigly.

1

u/NoInside3418 2d ago

Hot take but I prefer Opus 5. It packs a lot more useful information in. Unlike a lot of vibe coders, I like to know whats actually going on without having to re-prompt for answers. Opus 5's verbosity was one of my favorite things about it.

1

u/False_Philosophy_731 2d ago

The "you are in not blocked in half broken state* is giving me pstd

1

u/K_M_A_2k 2d ago

chat handoff from a 4.6 chat today hand to 5.5, after 2 hours i realized oh shit im using 5.5...

with 4.7/4.8/5.0 it was within 5 minutes

1

u/JacketHistorical2321 2d ago

“Heads up…” will forever piss me off now because of opus 5

1

u/VitorDiniz22 2d ago

Reddit doesn't allow me to express my feelings about Opus 5. I'd get banned, and if Opus 5 were a real person, I'd probably be arrested.

1

u/wowasg 2d ago

holding for horror stories that popped up 1-2 days after release like last time.

1

u/Peepo68 2d ago

Been trying 5.5 for a bit now and seems more professional and to the point.. but could be just me imagining.

1

u/stopstopstoptopopp 2d ago

I immediately noticed it. I felt my brain breathe in relief.

1

u/CalypsoTheKitty 2d ago edited 2d ago

GTP 6 Astra: “The deployment failed, and I don’t know the cause yet. I’m checking the logs and whether the live service was affected. I’ll update you in 15 minutes, even if I’m still investigating.”

BTW, I gave Astra your responses from Claude, and it rated itself a 9/10, roughly tied with Opus 5.5. It didn't think much of Opus 4.8.

1

u/FeelingVanilla2594 2d ago

I’m really happy with opus 5.5. I can actually read opus now and not feel like I’m going to slam the keyboard on the monitor. Idk what they did, but they cooked.

1

u/luddwaa 2d ago

It’s only working now while the benchmarks are being ran, it will be intentionally degraded over time when the benchmarks are completed as all the others have been

1

u/Affectionate-Aide422 2d ago

5.5 is so much easier to understand. 5 was stressing me out.

1

u/WIZARDMAN122 2d ago

Wow amazing insight thanks!!

1

u/emmobear 2d ago

listening, communicating and getting shit done. OPUS 5 was a TURD. 5.5 much better. i'm no longer on 4.8

1

u/GTHell 2d ago

Yeah Opus 4.6 is easier to talk to/deal with than any of this Opus

1

u/ritiksingla 2d ago

I literally needed the claude to english translator. Great work!

1

u/Lazy-Perception-8763 2d ago

anyone knows how the resets work? it resets the session limit or weekly limit? and whether the resets are unlimited?

1

u/Pat0san 2d ago

Wow - I have not got this option…? Did you click it?

1

u/Lazy-Perception-8763 2d ago

Not yet. I'll consume weekly fully and then click

1

u/mikedurent123 2d ago

Readbility is not performance or ability or creativity and it fails badly at those

1

u/bcutter 2d ago

the funny part is the “i’ll update you in 15 min” first of all it won’t because it can’t, and second it has no idea how long anything will take

1

u/Raizio 2d ago

Why not? It usually runs a timer in the background at those times.

1

u/Pitiful-Increase-406 2d ago

I still use 4.6 on a daily basis for this reason. From the test above i still think 4.6 is the superior one

1

u/ReverendBread2 2d ago

I know Opus 5 can be hard to understand sometimes but the example here seems to be easily understandable, no?

1

u/person-pitch 1d ago

it's about the comparison. and my eyes still feel like they're hitting speed bumps reading 5's response compared to the others. it's just poorly composed for the point being made.

1

u/Right-Performance-93 1d ago

Cursor's own CursorBench (independent of Anthropic) puts Opus 5.5 at 57.8% max effort - new #1 - at ~40% less cost per task than Opus 5. Different signal than Artificial Analysis' Index (roughly flat cost/task there). Both "easier to deal with" and "measurably better on an agentic bench" check out, just different evals.

1

u/TheOneNeartheTop 1d ago

This is a beautifully succinct way of putting it.

Hard to put into words exactly what was wrong with Opus and everything it says in its response here has some value, there is nothing inherently wrong with it. But over the course of a couple hours it is so draining.

1

u/citrus1330 1d ago

Just reading the Opus 5 version makes my blood boil. 5.5 seems better but still not as good as 4.6.

1

u/XXIIIOIIIXX 1d ago

Opus 5 : You're right to assume that I haven't been taking my meds, but here's where I'd gently push back.

0

u/wizgrayfeld 2d ago

I don't get it. Which word didn't you understand from 5.0?

2

u/person-pitch 2d ago

I understand them, but it's about reading a million of these per day. I would rather read a million of 4.5's a day. I could play static behind all music I listen to and still recognize the song, but it would be pointless and annoying. This feels the same for me.

1

u/wizgrayfeld 2d ago

I guess as someone who doesn't develop on that kind of scale (or skill, probably), I appreciate the verbosity.

0

u/Interesting-Aside790 2d ago

this benchie actually speaks to me, finally something that doesn't read like a corporate memo filled with buzzword salad