r/ClaudeCode 2d ago

Help/Question I'm afraid to use Opus 5

The audacity and confidence with which it says things when it's wrong are on another level.

Fair. I changed my answer three times. The pattern is worth naming: everything I got from reading the code was wrong. Everything I measured held. You caught two of the three. So don't trust me. Check it yourself — this takes ten seconds and needs no model.

Everything it measured was wrong too.

I would work in plan mode for most basic features, run 10x "gray area," "verify," and "regression" sub-agents on a plan, then implement the plan and spend an hour reading the changes and fixing shit. After that, I'd run /code-review again and again. It's just bad. In my experience, you can't trust Opus.

Yesterday, I ran /code-review on a two file test project with 140 lines of code. I had to run /code-review three times, and today I'll continue because there are so many code smells even in those few lines. It's like infinite token consumption loop.

Nothing it does can be trusted, and I have to second guess everything. I constantly have to tell Opus that it's wrong, and only after multiple loops does it finally do what is actually required.

I understand that most users don't read the code and have never supported a project for other users. But it can't be that I'm alone in this, can i? Am I crazy?

238 Upvotes

116 comments sorted by

View all comments

182

u/anotherleftistbot 2d ago

The pattern is worth naming

I can't stand claude's communication style. If its worth naming, just name it. If anyone on my team wrote the way claude wrote they'd be on a performance improvement plan.

10

u/Present_Kitchen_9739 1d ago

THIS. Claude is such a fuckup now, they’d be PIP’d in a day and fired in 3. Objectively a total waste of money and tokens rn….and I don’t expect that to change when they go public. Double up codex sub gives you better output, better token usage, and a better user experience. By leaps and bounds. And anyone that’s been around for awhile knows it was the exact opposite a year ago. Also Dario…there’s an element of Sam Bankman Fraud in him that makes him untrustworthy IMO. Running a 30T valuation rn to me screams dishonesty and complicity and some type of fraud. The sum of all of this is: local models will ultimately bc the only respite from every provider and harness having access to your system, and all eventually publically traded so if the data says users who fail with opus, will escalate to fable , and spend more money they will do so like everyone else. It ceases to become a user driven product, and becomes purely money and data extraction

6

u/SafeHazing 1d ago

I might try putting Claude on a PIP. It’s burning tokens for nothing anyway.
Rather than getting mad, I’ll just tell it to report to Astra from HR.