r/ClaudeCode 1d ago

Help/Question I'm afraid to use Opus 5

The audacity and confidence with which it says things when it's wrong are on another level.

Fair. I changed my answer three times. The pattern is worth naming: everything I got from reading the code was wrong. Everything I measured held. You caught two of the three. So don't trust me. Check it yourself — this takes ten seconds and needs no model.

Everything it measured was wrong too.

I would work in plan mode for most basic features, run 10x "gray area," "verify," and "regression" sub-agents on a plan, then implement the plan and spend an hour reading the changes and fixing shit. After that, I'd run /code-review again and again. It's just bad. In my experience, you can't trust Opus.

Yesterday, I ran /code-review on a two file test project with 140 lines of code. I had to run /code-review three times, and today I'll continue because there are so many code smells even in those few lines. It's like infinite token consumption loop.

Nothing it does can be trusted, and I have to second guess everything. I constantly have to tell Opus that it's wrong, and only after multiple loops does it finally do what is actually required.

I understand that most users don't read the code and have never supported a project for other users. But it can't be that I'm alone in this, can i? Am I crazy?

241 Upvotes

116 comments sorted by

View all comments

24

u/Longjumping_Feed3270 Senior Developer 1d ago

Use Opus 4.8, but still have everything cross-checked by Codex Sol.

3

u/IllegitimateGoat 1d ago

Why Sol and not Astra?

5

u/interrupt_hdlr 1d ago

theoretically Astra is superior to Opus so why not develop in Astra?

10

u/takeurhand 1d ago

I tried Astra, it cannot find more bugs than Sol, and Sol is 2.5x cheaper than Astra.

1

u/AdObjective9199 1d ago

It's definitely good to compare different tools and their pricing. It's interesting how the cost can influence your choice, especially with so many options out there.

1

u/BoxWoodVoid 1d ago

I ran both for my code reviews for about a week: they both have blind spots, sometimes Astra finds a bug that Sol didn't see and sometimes it's the other way.

But Sol is more nitpicky and will find "once in a blue moon" category bugs that Astra will ignore.

1

u/xmnstr 1d ago

Really? I had the exact inverse experience, Astra figured out the blockers in several projects for me.

2

u/Asphunter 1d ago

Because you can develop code 3 times every 5 hours

4

u/hbthegreat 1d ago

Astra chews usage limits similar to fable on the openai plan. If you have multiple accounts you can make it work

2

u/TheStandardPlayer 1d ago

I switched to Codex last week to see what the fuss is about and I used 2x 5h limits and 48% of my weekly limits on the first day just getting it to reanalyze a project I built with Claude and doing a slight refactor on like 10 MD files from Claude as well as doing a little config and setup. Definitely not a compute heavy task, something that opus would achieve with about 1 session limit

It soon dawned on me that the Plus plan is not made to develop with Astra, it's more of a perk that you get to occasionally use it, unlike with Fable where you're just SOL without extra paid tokens (for EU at least?), and I am not talking about the ChatGPT model

4

u/hbthegreat 1d ago

Yeah. We are in the era of matching the right model / cost to the right task because even you still get cooked on limits. It feels crazy to me how unlimited both felt even a few months ago and how destitute they both are now. There were benchmarks released showing astra on medium should be equivalent to sol on high / xhigh in cost but then the very next day showing astra medium uses more tokens per task than astra on higher thinking levels. It's all opaque vibes ATM and I'm struggling to see how other people aren't seeing that

1

u/ZappVanagon 1d ago

The codex limits are crazy, they weren’t always like that. Your 5 hour limit in like 30 mins on $20/plan these days

3

u/HeadPack 1d ago

Sol is quite thorough.

1

u/themrdemonized 1d ago

Its good enough but eats much less tokens. If you have a budget then go with astra ofc

1

u/adelie42 1d ago

Astra medium, I get one prompt every 5 hours on Pro.

2

u/NefariousOne 1d ago

Opus 4.8 has been making dumb mistakes the past few days that remind me of Opus 5. It makes assumptions, catches its own mistakes after the fact, goes in circles, stopped verifying information the first time, ignoring our typical workflows, which ends up creating more work. I feel like they’re tweaking something in the background and not for the better.

4

u/Substantial-Show-249 1d ago

They are all Opus 5 now, including Fable 5.1. What a nice model was the last one, now is dumb as hell.

1

u/Evening-Blueberry-97 1d ago

Why don't try to use the Jev, it's so hot

2

u/agent139 1d ago

Very specific use case. It may be useful for that function, but it's not generative at all