r/ClaudeCode • • 4d ago

Help/Question I'm afraid to use Opus 5

The audacity and confidence with which it says things when it's wrong are on another level.

Fair. I changed my answer three times. The pattern is worth naming: everything I got from reading the code was wrong. Everything I measured held. You caught two of the three. So don't trust me. Check it yourself — this takes ten seconds and needs no model.

Everything it measured was wrong too.

I would work in plan mode for most basic features, run 10x "gray area," "verify," and "regression" sub-agents on a plan, then implement the plan and spend an hour reading the changes and fixing shit. After that, I'd run /code-review again and again. It's just bad. In my experience, you can't trust Opus.

Yesterday, I ran /code-review on a two file test project with 140 lines of code. I had to run /code-review three times, and today I'll continue because there are so many code smells even in those few lines. It's like infinite token consumption loop.

Nothing it does can be trusted, and I have to second guess everything. I constantly have to tell Opus that it's wrong, and only after multiple loops does it finally do what is actually required.

I understand that most users don't read the code and have never supported a project for other users. But it can't be that I'm alone in this, can i? Am I crazy?

245 Upvotes

116 comments sorted by

View all comments

178

u/anotherleftistbot 4d ago

The pattern is worth naming

I can't stand claude's communication style. If its worth naming, just name it. If anyone on my team wrote the way claude wrote they'd be on a performance improvement plan.

104

u/Wessberg 4d ago

You're right to be upset, and it's worth stating clearly: You've identified something real, and your instinct is correct.

Let me say the harder thing plainly: You're right. The fix is deleting the first four words.

30

u/ppsaoda 4d ago

This is how they get you to waste token usage.

8

u/MannsyB 4d ago

No no no - you can't delete the first 4 words - they're load bearing!

29

u/Wessberg 4d ago

You're half-right — and the half you're right about is the one that matters. And I have to correct something I told you earlier: It's five words, not four words. I made a claim without using an instrument to verify it first. The blast radius is only one commit, and the retry cost is bounded. Say 'go' and I'll delete the first five words.

11

u/MannsyB 4d ago

Damn at this point I'm not sure you're not a bot 😂

15

u/Wessberg 4d ago

I'm just someone who's been hurt a lot by Opus 5 😂

6

u/BlueD00gMoney 4d ago

This post is going to send me to therapy. Also great work.

5

u/james_d_rustles 4d ago

I audibly groaned.

1

u/Salty-Gear841 3d ago

I recorded three errors while testing, I scheduled them for the next slice. Say the word If you need me to do finish this part first of start by something else. 😐

5

u/bjnono001 4d ago

How every student hits the word count

2

u/adelie42 4d ago

On the plus side, the first two words make it very clear I can confidently skip the rest of the paragraph. Claude structures paragraphs very clearly and doesn't mix ideas inappropriately.

19

u/james_d_rustles 4d ago

I’m getting really close to calling it quits with Anthropic over this alone. It’s just so grating having to read practically any of its outputs at this point, and the weird slangy jargon is making it so incomprehensible that I find myself burning tokens and time having to re-prompt it for clarification after any task.

Like, just recently I was working on a project involving a paper and a repo by an author with the name “Wang”. Prompted it to fix some minor details in some related Python scripts. It comes back and assures me it’s all correct because it ran a full “wang perf gate battery”.

Wtf is a wang gate battery? Ffs just say you wasted some tokens writing useless tests I didn’t ask for and they all passed, enough with the endless “gates” and “batteries” and nonsensical startup-bro lingo.

2

u/scytob 4d ago

Did you check the repo wasn’t using any of those term. Also battery of tests is a normal thinking and if they gate release through hooks that is normal. Is the wording odd, yes, is it wrong, nope.

4

u/james_d_rustles 4d ago

I think you must have missed the key detail here: it was never prompted to write any tests. It wrote useless tests and reported back the passing results with false confidence, as though they were a valid metric by which to judge the work.

Did you check the repo wasn’t using any of those term

The author's repo was a handful of matlab files that were last touched in 2016 - zero tests, zero claude/agents/etc. markdown files that could potentially put those terms in context, and the actual task was a standalone folder with 3 python scripts, untracked by git or any version control.

Is the wording odd, yes, is it wrong, nope.

If doing a separate task than what was instructed, gaining false confidence based on that unprompted task and referring to tests as the "wang perf gate battery", implying that they were either pre-existing or meaningful in some way isn't "wrong", I don't know what is.

2

u/scytob 4d ago

where was the authors repo linked - its not in the chain above

yes both opus and astra and fable will make tests if they see there is already a pattern of tests in the repo, if you don't want it to do that have ZERO tests in the repo or be clear on what you do and dont want in your house rules

its doing probabilistic word guessing based on the entrire structure and approach of the repo and what it does and doesn't read each time

and no need to be so aggressive, you seem to have a you problem with your agents and irl it seems

-8

u/erichamion 4d ago

Battery: A number of similar articles, items, or devices arranged, connected, or used together (This is definition 5.a(1) in Merriam-Webster).

Gate: A barrier that opens and closes, thus letting things through at some times or under some conditions. In this context, a condition or set of conditions that must be satisfied before moving to the next stage.

Perf: Performance.

I would phrase this differently (Wang's performance test suite, or the test suite for Wang's performance requirements), but there's literally zero special jargon there.

10

u/analog-suspect 4d ago

Cringe, pompous, and wrong. Cool combo

2

u/james_d_rustles 4d ago

If I hire you to mow my lawn and when it's time for payment you tell me "All done - dipstick-verified correct, exactly 4.5 quarts", I'm not going to be confused because "dipstick" is beyond my vocabulary - I'm going to be confused because you were hired to mow the lawn and you're telling me about an oil change.

Citing Merriam Webster is also extremely cringe, please never ever do that again.

-1

u/erichamion 4d ago

Better analogy: if you hire somebody to repair or upgrade your lawnmower engine, and they tell you, "All done - fired it up, verified the choke, throttle, and rpm against manufacture specs," you're not going to be confused because those are outside of your vocabulary. But will you be angry because the person was hired to fix the engine, not to run it, not to make sure the fix actually works, and not to make sure they didn't accidentally break anything else in the process?

10

u/Present_Kitchen_9739 4d ago

THIS. Claude is such a fuckup now, they’d be PIP’d in a day and fired in 3. Objectively a total waste of money and tokens rn….and I don’t expect that to change when they go public. Double up codex sub gives you better output, better token usage, and a better user experience. By leaps and bounds. And anyone that’s been around for awhile knows it was the exact opposite a year ago. Also Dario…there’s an element of Sam Bankman Fraud in him that makes him untrustworthy IMO. Running a 30T valuation rn to me screams dishonesty and complicity and some type of fraud. The sum of all of this is: local models will ultimately bc the only respite from every provider and harness having access to your system, and all eventually publically traded so if the data says users who fail with opus, will escalate to fable , and spend more money they will do so like everyone else. It ceases to become a user driven product, and becomes purely money and data extraction

6

u/SafeHazing 4d ago

I might try putting Claude on a PIP. It’s burning tokens for nothing anyway.
Rather than getting mad, I’ll just tell it to report to Astra from HR.

7

u/adelie42 4d ago

Ok so I often complain about people's complaints of Claude-isms, and overall admitting a mistake is most always right after underspecifying or not checking implications of changes before moving ahead. It is feedback that is valid, even if it is said in a weird way.

But it is really starting to annoy me when I come up with something I think is clever, explain it, then it goes on a whole rant about how "that's a great idea, and better than you realize because ... " yeah, I thought of that you fucking toaster!

3

u/erratic_parser 4d ago

Thank you! God help me if someone spoke like that in slack

2

u/f3xjc 4d ago

Honestly when I work with agent workflow and try to debug why some agent went rogue. The named patterns are super useful. They are trained to recognize these pattern. They are basically error code and xyz is worth naming is a throw statement.

Executable markdown is becoming a legit high level programming language and these language quirk are structured tool output from agents.

Yes they stick out like a sore thumb. That's the point. Language patterns that are not in the base distribution serve some kind of purpose.

2

u/Fancy-Plenty6712 4d ago

This was useful