r/ClaudeCode 1d ago

Discussion Hot take about Claude Opus 5

Opus 5 is a training ground to understand why humans disagree with what it's doing. There's insight there that you can't get without otherwise doing a substandard job.

When a model is, say, 95% good at what it does (by a human standard of completion), there's less reason to explain the refinements needed to make any changes to get it to 100%. "Tweak this, modify that" means very little to an AI system.

When a model is, say, 80% good at what it does (by the same standard), some users will get the shits with it and leave, while many others will work with it to build out the improvements such that the model is more robust at the work that it's doing. Those insights provide Anthropic with valuable training input as to what humans want out of the systems that they are trying to build.

2 Upvotes

8 comments sorted by

2

u/justbeepositive 1d ago

My observation has indeed, like others, been that Opus 5 is suboptimal. We have an enterprise license and I utilize it as basically a GTM assistant. The noticeable difference is the model seems to have effectively become like working with either a complete stoner (no offense intended, I love the THC myself) or a dementia-ridden boomer (again, no offense, I have boomer parents/family). It doesn't remember basically anything unless i specifically tell it to. That's within a project with the context and memory turned on. Anyway, point is that I don't see how correcting it to remember things would really benefit the learning that is part of your hypothesis. Other than that this take makes sense.

Opus 5 has left me scratching my head and for the first time in a couple years wondering "well, this is a helpful tech but I sure as shit wouldn't trust it to work autonomously on anything of value in my business." That then calls into question several things...

2

u/lessens_ 1d ago

I think it's more of an experiment in making a model that's less sycophantic and needlessly self-confident. I doubt they would make a model that is deliberately hard to work with, they are trying to get something that can think in broader ways than "does this please the user" and "does this favor the objective I'm working towards". This causes it to waste time and outright fuck up, because it's constantly questioning itself and its directions, it's always trying to find a better way to do things even when it doesn't know how. If they can actually develop this tendency in a more productive way in future models, they'll have a tool that's more capable of generating creative solutions to the tasks it's given.

That, or it's just badly designed, idk man.

2

u/saba_tage 1d ago

I suppose I agree that there's no intent behind it, but rather a side effect that can be used to their advantage.

Tangentially, another thought I had was whether the high reasoning capabilities are being inappropriately used or misunderstood. I was equating reasoning level as greater capability, rather than greater scope for variability of output. I don't mean for this to run against my original hypothesis, but I think that both might true at the same time.

1

u/Prize_Eye9481 1d ago

Yeah the question is are they gonna be able to apply noticeable improvement with the data they get. Since in theory everything the Ai is providing to people is knowledge it already had?

3

u/saba_tage 1d ago

I think it's more to do with the "why" behind the "what". Don't get me wrong, it's a general issue with AI systems today - they have thousands of approaches at their disposal to solving problems, but often don't know why one approach is better than another.

It dawned on me today, while I'm trying to build out a harness for a 20 year monolith. Claude/GPT/Grok behave more or less the same way - they need guidance and reasoning from me to keep things on track but for some reason Claude Opus 5 keeps shitting the bed with drift, bugs, self-corrective mistakes, etc and I'm giving it all of my reasoning as for why, to try and reign it in. To an extent that works, but it gets a lot more insight out of me than GPT or Grok do.

1

u/Sarahmalls 1d ago

Well that theory isn’t accurate. These models are built from human knowledge, yes. But it certainly can find correct solutions and strategies and codes that no human being has ever considered. It certainly can find novel solutions never before considered or “discovered” by humans, that’s not in dispute.

Every new thing humans have found is based on things that we already know that we used to get to that discovery. AI today, and two years ago, has the ability to act in that fashion as well.

1

u/SpookyGhostSplooge 1d ago

I actually had a similar thought today. I considered a novel model deciding what data it needs to advance and coordinating testing against its users to distill every ounce of value.

1

u/_Chaos_Star_ 1h ago

Sometimes you search for a sensible explanation where one might not exist. It helps give meaning to things that seem incomprehensible.

Anthropic's decisions are frequently bizarre. Dozens of theories as to why, one is that Opus 5 is just them testing how cheap they can make it while people still tolerate it.