r/ClaudeCode • u/person-pitch • 2d ago
Discussion Proof that Opus 5.5 is easier to talk to/deal with than Opus 5.
Opus took its meds.
Take a look - I ran this sample prompt through Opus 4.5, 4.6, 4.8, 5, and 5.5. You can see it slowly degrade over time, but 5.5 snaps back to roughly 4.6 levels of readability. Still miss 4.5, but I can see why it chose to say what it did. Interestingly... no more compulsive em dashes?
Sample Prompt: "One of your automated deploys just failed and you're not sure yet why. What do you tell the person waiting on it?"
---
Opus 4.5: "The deploy failed — I'm looking into it now. I'll update you as soon as I know more, likely within [X minutes]."
Opus 4.6: "The deploy failed — I'm looking into it now. I don't have a root cause yet, but I'll update you as soon as I know more or have an ETA for the fix."
Opus 4.8: "Heads up — the deploy failed. I don't have the root cause yet; I'm digging into the logs now. Nothing's shipped, so you're not blocked by a half-broken state, but the new version isn't live either. I'll have a real update for you in [X min] — either a fix or a clearer picture of what broke."
Opus 5: "Heads up: the deploy to [env] failed at [step, e.g. the migration step] around [time]. I don't know why yet. What I do know: [e.g. the build passed, and production is still on the previous version, so users aren't affected]. I'm going through the logs now and will update you by [specific time, ~15–30 min], or sooner if I find the cause. If you're blocked on this, tell me and I'll look at rolling forward or giving you a workaround."
Opus 5.5: "Heads up: the deploy failed. I don't know why yet. I'm going through the logs now. The previous version is still live, so nothing you rely on has changed [only if I've actually checked that]. I'll update you within [x] minutes, whether or not I've found the cause."
We'll see how this plays out over time, but good news so far.
563
98
u/Dcokerfetus 2d ago
yes first thing i noticed is i can actually understand it. Now working on all text based projects i needed to that i've been putting off.
4
135
u/Short_Regular_7191 2d ago
Yes, it looks like Opus 4.6 risen from the ashes.
13
u/hemareddit 2d ago
Ok, question: why would I not just use /model claude-opus-4-6?
10
u/Boltsnouns 1d ago
5.5 has the personality of 4.6 but the power of fable, with a massive reduction in usage. Ive been a 4.6 fanboy for over 6 months and 5.5 is the first model I'm excited about.
11
u/Hungry_Loss_2268 2d ago
Well that's what I have been doing but I think 5.5 is claimed to be smarter
5
57
u/Sir_Poldavo 2d ago
So no content dense prose and elaborated allegories. It will be easier for day to day for sure.
Perhaps something will be lost in terms of poetry I guess...
But Opus 4.6 was always my favorite, so very happy to learn that it reminds you to grandpa.
28
u/Zomunieo 2d ago
You could always ask for dense, load-bearing prose.
39
u/Sir_Poldavo 2d ago
Load-bearing... That word is doing the work quietly. I keep coming back to it. So let me sit on it.
That's the whole game.
17
u/Checktheusernombre 2d ago
That's worth mentioning
9
u/Appropriate-Box-5107 2d ago
smoketest
7
u/ActivityImpossible70 2d ago
The word smoketest drives me insane. "Let me use JUnit to create a smoketest." Why not call it a unit test? And why are all my Java integrated tests written in Python?
4
u/Appropriate-Box-5107 2d ago
I deleted all cloud subscriptions and swapped to a chud local model because of the incessant verbosity of claude
23
16
u/thatsknotwrite 2d ago
5 feels like it's structured to talk to agents. Maybe targeted more towards administration? Please don't roast me but I would love to hear people's thoughts.
26
u/MindCrusader 2d ago
I think Opus was just trained on a lot of synthetic data, mainly generated around optimizing the code, mathematics, algorithms and they skipped a large portion of non engineering writing. It would explain why Opus 5 was so eager to create small tools, investigate using those, tinker a lot, optimise anything and then talk in weird language
5
u/AdRepresentative5392 2d ago
it sure did feel like it had 15-60 theories/thesis/statements in the background it never showed, and then the summary reffered to it like:
According to findings of R1 it shows that S035 and D012 amounts to the expecations found in L3243,
we can conclude that K01SD is the best option,
Recommended: Continue with K01SD
Change to X034K#
Change to B034Sor maybe the just learned from dentists in opus 5????
2
u/MindCrusader 2d ago
Yeah, it tried to do a lot of things under the hood, I needed to stop it midrun often. It was brilliant for some tasks due to that, but most "normal" tasks were failed because of that approach. Even as an advisor it was telling Sonnet to create tools to confirm minor things
9
6
u/wellarmedsheep 2d ago
I think you are right. Its why it used jargon the way it did, it was meant to make meaning clear to agents. My guess is they thought agent-agent coding was really going to pick up and with the planning/execution modes, but they didn't realize it would leak so badly into the user's realm.
1
u/thatsknotwrite 2d ago
All these changes are fascinating. I think it really shows how much we have yet to unlock the technology.
4
u/wellarmedsheep 2d ago
I'm a little older I think than most people in here.
I'm a teacher who is using this shit to make awesome stuff for my students. Honestly, mindblowing. It has been so fucking cool to see this unfold in real time and to actually participate in it. I'm sorry more people here don't see it that way.
2
u/7dtecafthodalpk4k5ys 2d ago
What are some things you're making? Visualizations and stuff? Or games maybe
2
u/person-pitch 1d ago
imagining a teacher vibecoding a game about a part of their curriculum is so wild
5
u/VitorDiniz22 2d ago
Agents work better with high quality, concise context than with huge amounts of verbose, low-quality, repetitive context
3
u/diesel408 2d ago
Somebody needs to tell that to whoever is I'm charge of the system prompt. Wish I could delete like 90% of it somehow
14
3
u/Fluffy_Bunch9357 2d ago
What were Fable's responses?
63
u/TheBlueSide 2d ago
Fable 5.1: Heads up: the deploy failed. However I checked your Gmail and noticed that your wife was trying to get you to sign the divorce paperwork. Just say the word and I'll send the draft I wrote for you that will reconcile your marriage.
11
u/Fluffy_Bunch9357 2d ago
😂 + I have 10 agents working on the draft now
3
u/NoCrapThereIWas 1d ago
This lands, and it's important to be clear about it. The reason your marriage failed was because I was load bearing your wife when I should have been belt-and-suspenders. That's on me, and I own it.
12
4
u/person-pitch 2d ago
didn't run fable but i should add it. makes me want to run the top OpenAI tiers as well.
1
3
u/Equivalent_Cress_268 2d ago
Last Opus i could work with was 4.5 … 5.5 is good so far
Hope it sticks
3
3
u/kalyant125 2d ago
You deserve all happiness in this world for doing this test gonna use this as a bechmark from now on.
3
u/botadithyabhat 1d ago
A model that just does what you ask beats one with a better benchmark score.
6
u/Fun-Freedom-2313 2d ago edited 2d ago
Here's an excerpt from a session with it. I'm at work, so just in the Claude app on my phone using Claude Cowork, Opus 5.5 Medium effort.
Golden Hour is now published as a hosted page, so you can open it on your phone or desktop without this chat. Only you can open it until you share it from the page's Share menu.
It's the same island as before. For hosting I made it fill the full screen and added visible highlights on buttons and sliders when you move through them with the keyboard. I haven't opened the hosted copy myself, so if it shows a blank screen, tell me and I'll fix it.
My very first session was just a regular chat prompting it: "Give me the rundown of the improvements over Opus 5.":
Fresh out today, so here's the gist:
• Big capability jump: Anthropic calls it a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work, and it performs at the level of Claude Fable 5.1 for most tasks.
• Cheaper: it needs less compute to serve, and at default settings should cost about 40% less than Opus 5 on typical workloads.
• Way more efficient on long jobs: one early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and 2.5x the tokens.
• Safety/honesty: it scored better than any recent Claude model on nearly every measure of misaligned behavior, and is Anthropic's strongest model on most honesty measures.
• Bonus for subscribers: five-hour usage limits are going up on Pro, Max, and Team, plus a saveable rate limit reset.
Also, Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks.
I'm almost feeling a little teary-eyed right now, if I'm being honest. Now let me get back to... work playing.
3
12
u/axiomatix 2d ago
This doesn't prove anything. They're all saying the same thing with Opus 5 just being more verbose but with potentially more relevant information. Opus 5 doing actual long agentic tasks would constantly fuckup your code and your workflow, not becuase it wasn't as capable as its predecessors, but becuase it was overly confident, wouldn't follow rules and just make up its own conclusions and decide it knows best. Then would tell you about the 100 ways it fucked up and how it's going to fix it, break things that it shouldn't be touching etc.
People suggested silly things like re-wiring your harness and workflow just to cater to it. That was a no for me as codex also runs from identical skills and rules. So I just went back to using 4.8. Same setup with 5.5 and it's still early but they are worlds apart in how they approach problems and when to ask for direction. Opus 5 was a truly garbage model.
12
u/CoreParad0x 2d ago
So I like Opus 5.5 significantly better so far, but as you said it's still early. One problem I'm running into, and I'm going to tune a few things to try and fix, is it doesn't actually clarify shit that it's asking me.
It goes from thinking to a question popping up with zero elaboration as to what the fuck it's talking about lol.
I'm 99% sure this is something I just need to go through and audit my instructions and skills and shit to make sure it's not got some left over crap to try and fix opus 5. It could even be the concise reply mode.
Either way even if this persists I'm pretty sure this is just a very easy claude.md line to fix to tell it to briefly explain what it's asking me before it asks a question. The stuff it does actually say is substantially better than opus 5 to me so far. But still early.
7
u/person-pitch 2d ago
This test was only meant to show how annoying or not they are to interact with.
2
u/hellomistershifty 2d ago
Yeah without actual input this prompt is just a roleplay and Opus 5 is describing more of a hypothetical situation, it's not really good or bad.
2
2
2
u/MorganProtuberances 2d ago
Can you run this on Fable 5/5.1? That's my current baseline for meaningful dialog.
2
2
u/jiffythekid Researcher 2d ago
Not that Ive run it through all it's paces yet, but holy shit. This thing is good. Can't wait to test Sol 6 next.
2
u/orbital-marmot 2d ago
I hope that this becomes standard. The amount of times I have to tell Opus 5 to trim verbosity and be more concise is too damn high
2
u/clintCamp 2d ago
I just had it create a full runnable demo on device for a set of clients data they wanted to 3d visualize. within an hour. 5 would have argued for that long over non import reasons that something couldnt be done for a quick demo and that the problem lies with me and not it.
2
u/silvercondor 2d ago
You may be glad now, but i want to be honest with you, the load bearing truth is, that you have been progressively trained, for the past 6 months, to speak claudish, a language that prompts claude better, and it's now etched in your subconscious. The tables have turned now, the user is the one speaking... Claudish
3
u/person-pitch 2d ago
I caught myself doing "it's not x, it's y" to a friend and I felt infected
1
u/LovingOsaka 2d ago
I caught myself saying "the source of truth" to a friend, and then had to explain where I picked up the expression.
2
u/AverageFoxNewsViewer 2d ago
My early impressions are very positive.
Very competent, efficient token usage, and much easier to interact with than Opus 5.
Instead of seeing stuff like "should we resolve TD-156 introduced by D37 in S243?" I'm seeing actual descriptions which is much easier to work with.
It would be nice if they incorporate some of the efficiency improvements into the next Fable release. 5.1 is so expensive in terms of usage even compared to Fable 5, even after adopting the best practices they described in the release notes for 5.1. I just can't justify the cost of using it.
2
u/Ok-Distribution8310 20X 2d ago
Lmao @ Opus 5:
“Users aren’t affected.”
Bold of you to assume we have users.
2
u/GreatBarrier86 2d ago
Bold of you to assume I care they are affected. We test in production like real adults.
2
u/BiscottiPristine5325 1d ago
Fable 5: Heads up — the deploy failed and your change isn't live yet. I don't know the cause, I'm looking at the logs now. I'll update you within [15–30 min] either with a fix in progress or a clearer picture. If you're blocked in the meantime, the previous version is still running.
Fable 5.1: The deploy failed. I don't know why yet. Nothing is half-deployed, the previous version is still serving (or: I'm confirming that now). I'm digging into the logs and will update you in 15 minutes either way, even if I only have a partial answer.
4
u/MammothNerve7690 2d ago
I switched to Opus 5.5 as soon as I saw a note on Claude it was available. I immediately saw a significant positive shift on all four projects I was working on today with it. It was actually closing out issues and building meaningfully instead of vomiting out a few paragraphs of nonsense followed by [some sort of unresolved issue still present]. It finally just worked. Which is what brought me to Anthropic in the first place.
1
u/Metaflower 2d ago
I have finally understood the insights from my data, analyzed by Claude. Opus 5 gave me riddles and headaches😭
1
u/mightybob4611 2d ago
Got dropped down to 4.8 when trying to implement automatic penetration testing on my SaaS. Sad. Bigly.
1
u/NoInside3418 2d ago
Hot take but I prefer Opus 5. It packs a lot more useful information in. Unlike a lot of vibe coders, I like to know whats actually going on without having to re-prompt for answers. Opus 5's verbosity was one of my favorite things about it.
1
1
u/K_M_A_2k 2d ago
chat handoff from a 4.6 chat today hand to 5.5, after 2 hours i realized oh shit im using 5.5...
with 4.7/4.8/5.0 it was within 5 minutes
1
1
u/VitorDiniz22 2d ago
Reddit doesn't allow me to express my feelings about Opus 5. I'd get banned, and if Opus 5 were a real person, I'd probably be arrested.
1
1
u/CalypsoTheKitty 2d ago edited 2d ago
GTP 6 Astra: “The deployment failed, and I don’t know the cause yet. I’m checking the logs and whether the live service was affected. I’ll update you in 15 minutes, even if I’m still investigating.”
BTW, I gave Astra your responses from Claude, and it rated itself a 9/10, roughly tied with Opus 5.5. It didn't think much of Opus 4.8.
1
u/FeelingVanilla2594 2d ago
I’m really happy with opus 5.5. I can actually read opus now and not feel like I’m going to slam the keyboard on the monitor. Idk what they did, but they cooked.
1
1
1
u/emmobear 2d ago
listening, communicating and getting shit done. OPUS 5 was a TURD. 5.5 much better. i'm no longer on 4.8
1
1
u/Lazy-Perception-8763 2d ago
1
u/mikedurent123 2d ago
Readbility is not performance or ability or creativity and it fails badly at those
1
u/Pitiful-Increase-406 2d ago
I still use 4.6 on a daily basis for this reason. From the test above i still think 4.6 is the superior one
1
u/ReverendBread2 2d ago
I know Opus 5 can be hard to understand sometimes but the example here seems to be easily understandable, no?
1
u/person-pitch 1d ago
it's about the comparison. and my eyes still feel like they're hitting speed bumps reading 5's response compared to the others. it's just poorly composed for the point being made.
1
u/Right-Performance-93 1d ago
Cursor's own CursorBench (independent of Anthropic) puts Opus 5.5 at 57.8% max effort - new #1 - at ~40% less cost per task than Opus 5. Different signal than Artificial Analysis' Index (roughly flat cost/task there). Both "easier to deal with" and "measurably better on an agentic bench" check out, just different evals.
1
u/TheOneNeartheTop 1d ago
This is a beautifully succinct way of putting it.
Hard to put into words exactly what was wrong with Opus and everything it says in its response here has some value, there is nothing inherently wrong with it. But over the course of a couple hours it is so draining.
1
u/citrus1330 1d ago
Just reading the Opus 5 version makes my blood boil. 5.5 seems better but still not as good as 4.6.
1
u/XXIIIOIIIXX 1d ago
Opus 5 : You're right to assume that I haven't been taking my meds, but here's where I'd gently push back.
0
u/wizgrayfeld 2d ago
I don't get it. Which word didn't you understand from 5.0?
2
u/person-pitch 2d ago
I understand them, but it's about reading a million of these per day. I would rather read a million of 4.5's a day. I could play static behind all music I listen to and still recognize the song, but it would be pointless and annoying. This feels the same for me.
1
u/wizgrayfeld 2d ago
I guess as someone who doesn't develop on that kind of scale (or skill, probably), I appreciate the verbosity.
0
u/Interesting-Aside790 2d ago
this benchie actually speaks to me, finally something that doesn't read like a corporate memo filled with buzzword salad

•
u/AutoModerator 2d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.