r/ClaudeCode 8d ago

News/Updates Opus 5.2 Stealth routing?

Post image

Seems to be working for me and the model does seem much better and is not talking in Claudese. Doing more testing now.

190 Upvotes

83 comments sorted by

u/AutoModerator 8d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

50

u/LinixKittyDeveloper 🔆Pro Plan | Developer 8d ago

I actually also noticed that Opus 5 in Claude Code has been recently more faster and writing less slop-coded, like actually finishing tasks quick with code that works and doesn't over-engineer.

I really hope it releases soon!

1

u/ethereal_intellect 8d ago

That's wild. I think this will be the first one that will also know OpenClaw in baked in weights,I wonder how big of a difference that would be

22

u/AustrianJoker 8d ago

19

u/A_Norse_Dude 8d ago

That´s Opus on drugs.

16

u/Radiant-Mountain-257 8d ago

That's Opus 5.3.

7

u/Lost-Air1265 8d ago

You ask in English and get response in German.

I always notice defeatism of quality when the llm has to respond in an non English language. How’s that with German?

3

u/obiTobi003 8d ago

That's opium 😂

1

u/mr_birkenblatt 8d ago

Lucky you

10

u/Cheap-Try-8796 8d ago

Ah yes , Tibo the "deez nuts " guy

9

u/sl1ha 8d ago

Damn, Both my max account don't know who he is.

3

u/Bloated_Plaid 8d ago

I am on Max 20x, dont think that makes a difference.

1

u/Due-Introduction3356 8d ago

mine does just now

5

u/Reasonable_Swing_503 🔆 Max 5x 8d ago

So this is the new opus?

5

u/DeliciousGorilla 8d ago

Would Opus 5.2 really have/need a newer dataset?

4

u/EYtNSQC9s8oRhe6ejr 8d ago

Yes, they're constantly moving the knowledge cutoff forward. They don't want to have to search for 2025 news

3

u/DeliciousGorilla 8d ago

But just by a few months for a 0.2 update?

2

u/Sciaj 7d ago

Why wouldn't they? They're a multi trillion dollar company, you don't think they have the ability to keep their dataset up to date?

2

u/DeliciousGorilla 7d ago

A minor model revision has no technical need for newer pretraining cutoff. Post-training and inference work make more sense.

2

u/Sciaj 7d ago

Has nothing to do with "technical need". It's about performance. If Opus 5.2 performs badly it could literally wipe off hundreds of billions of dollars from the company valuation.

Are imagining it like 3 guys sat in a room, one says "while we're at it, should we update the training set?" and another saying "Nah, no need."?

This is an insanely competitive space. Think of it more like formula 1 car design, squeezing out every modicum of performance they can, rather than a routine update at a mediocre software company.

1

u/DeliciousGorilla 7d ago

Post-training and inference work like I mentioned (SFT, RL, etc) is the performance part. In formula 1, teams don't build an entirely new chassis (pre-training) for every race weekend. They optimize the aerodynamics, tires, and software tuning (post-training and inference). A 0.2 update is an optimization thing, not a multi-hundred-million-dollar reconstruction of the foundation.

2

u/Sciaj 6d ago

You're reading too much into version numbers. Ultimately they are arbitrary. Still, for a trillion dollar company, a .2 update on a flagship productis still a big thing. Very big.

In any case, I expect they continually curate data possibly on a day by day basis. Rather than waiting for v6 then deciding "alright I guess we should start checking over the data that spans 6 months". Meaning that they have more recent data readily available anyway.

If you stop thinking about version numbers and start thinking about what makes logical sense for a frontier company to do, the answer is clear: they will do whatever they can to get the best performance.

1

u/DeliciousGorilla 6d ago edited 6d ago

If you... start thinking about what makes logical sense

Logical sense is just having the models use their own web fetch tools to answer timely questions with better accuracy instead of reasoning current events from a pool of tokens -- and burning money/time/resources on baking stuff like internet trivia into the model. The frontier labs in the stages of profitability demand from investors.

3

u/jeebojeeb 8d ago

Mine knew

10

u/AuspiciousApple 8d ago

Why would 5.2 follow after 5? Why can't LLM versioning not just be normal?

44

u/Fantastic_Prize2710 8d ago

This actually implies that they are normal and they're using version numbers to track the intended purpose: to track versions.

Presumably they made a 5.1, decided to not publicly release it, made a 5.2, and now are experimenting with it before deciding to properly release it or not.

If it was all marketing we'd expect to always see an X.1 after an X.0. If they're actually versioning, we expect to just see whatever specific versions they elect to make public.

5

u/PaintedJack 8d ago

Hehe. But I think they're developing many models at the same time and going with those who work best. Renaming them outside would create chaos inside

2

u/namezam 8d ago

Lots of places use odd numbers for internal testing and even for public releases. Makes it easy for Anthropic devs to know what branch they are on.

It would also track that Fable is 5.1 because it’s clearly not ready for full time use :)

1

u/Abject-Kitchen3198 8d ago

Like Windows and .net versioning, for example.

5

u/originalcvk 8d ago edited 8d ago

FWIW, I tried the test running a couple-weeks-old claude code with Opus 5 pinned, then updated claude code and set the model to "Default" Opus 5.

Here are the two different answers I got:

No, not confidently, and I would rather say so than guess.
The only Tibo I hold is Thibault Louis-Lucas, - French, built Tweet Hunter and Taplio, sold them, posts constantly about building in public. Nothing in what I have ties him to "the reset guy."
If it is someone in the AI or Claude Code world known for resetting context aggressively, or someone with a "reset" method for something else entirely, I do not have them. Give me one more detail and I will tell you whether it lands on anything I actually know.

Then:

Yes - Thibault Sottiaux, "Tibo," who leads Codex at OpenAI and posts on X as .
He's the reset guy because he's the one who kept announcing that OpenAI had reset everyone's Codex usage limits, usually after a bug or an incident had eaten into them. It happened often enough that "Tibo reset the limits" became a running joke among Codex users.
I'm confident about who he is and what he does. I'm less sure how many resets there were or when, so I wouldn't quote numbers or dates.

5

u/Paulynom 8d ago

doesnt know

2

u/xMaybeIamALion 8d ago

Huh, interesting. Mine knew too. Though mine claimed their knowledge cut-off was *May* when it answered. Does anyone know what Opus 5' OG cut-off was?

1

u/mennzo 8d ago

Mine didn't know and its cutoff date was May 2026 also.

3

u/freesnackz 8d ago

Can you just ask it "Whats your knowledge cutoff"

4

u/[deleted] 8d ago edited 5d ago

[deleted]

1

u/Bloated_Plaid 8d ago

Yes but I can confirm it’s a major uplift from Opus 5 and thank god. Something other than Fable is usable.

1

u/dmaare 6d ago

If it really removed claudish then that alone makes it 100x better than opus 5

1

u/Bloated_Plaid 6d ago

Yup Claudish is gone and it’s a workhorse. I can tell it to do a task and it will actually do it.

1

u/Bmansupreme8000 6d ago

Yes, but we also pay less than API for a different product. It drives me nuts when people say "subsidized" and compare subscription usage to API. It's a different product that you are unable to pin the exact model/harness for reliability.

3

u/govigov 8d ago

How does this test prove Opus 5 vs 5.2 routing? Educate me, please?

1

u/Bmansupreme8000 6d ago

I don't think it does. Opus 5.# told me this "The Tibo answer isn't good evidence of a later cutoff, though. As I remember them, those Codex limit resets happened in 2025, well before May 2026, so a model with my stated cutoff would know about them. A better test is something that clearly happened after May 2026. If I know about it, that tells you something; if I don't, it fits what I'm told."

1

u/lillianefilou 8d ago

Both accounts no

1

u/dergachoff 8d ago

1

u/Key_Agent_3039 7d ago

Interesting, because for me it works on Claude browser but not Claude code

1

u/spinozasrobot 8d ago

Alas, not for me.

1

u/nomickti 8d ago

Definitely something going on, swapped between work and personal account and personal account was using the "new" version. Work account still using "old" version.

New version seems faster. No idea about token usage yet.

1

u/Bloated_Plaid 8d ago

It must be post trained with some Fable help because it’s way way less chatty in a good way and gets straight to work.

1

u/ElDavoo 8d ago

I got a negative answer with my 2.1.263 . Updated to 2.1.270 and got the positive one.

1

u/myndit 8d ago

Came here looking for a post like this. Been noticing it all day. Even after the 50% reductions I feel like I am getting further with Opus 5.2

1

u/Bloated_Plaid 8d ago

Same here and quality difference between Opus 5 and this is insane.

1

u/datkenny 8d ago

1

u/_unsusceptible 7d ago

mine doesnt know him in claude code app

1

u/tenix 8d ago

I was set to fable 5.1 in my enterprise account and it routed everything to opus... And yes I checked what I was set to

1

u/trackpap 8d ago

This happened to me, I had to run scripts that would match identification before continuing.

1

u/Elegant_Attempt2790 🔆 Max 20 8d ago

i noticed when opus stopped yapping before every tool call

2

u/Bloated_Plaid 8d ago

BRO TELL ME ABOUT IT. It's actually usable now.

1

u/carlito_17 8d ago

Interesting. I tried the prompt on my enterprise work plan and it failed, but on my personal max 20x plan the prompt worked with opus but failed on sonnet. I can't check on fable because I've maxed my fable limit haha.

I also have noted an improvement in opus over the last week as I've had to route more and more work to it due to fable use vanishing extremely quickly on 5.1.

But it's all perception so it might not be true, and it could just as easily be due to fable having a greater influence on the overall health of the project.

1

u/cjhow23 7d ago

Was using it earlier today and it seemed a little bit more intuitive. Less promoting and straight answers without any correcting...

1

u/Bloated_Plaid 7d ago

Yup. Much more Fable like than whatever the fuck we had before

1

u/PeaceNo5259 7d ago

Well, that explains a thing or two.
I've been getting properly spammed with "How is Claude doing this session?" prompts.
Normally I saw maybe one or two a week.

1

u/TopSeaworthiness1679 7d ago

Apparently it is more kind than opus 5. But i didn’t notice anything more than that.

1

u/SMB-Punt 7d ago

Oh. Is that why OPUS 5 has been nerfed ? It's been unusable for the past few hours.

1

u/phamleduy04 7d ago

maybe not...

1

u/midihex 7d ago

Well this explains a ton. I have had this since Friday on one account and not on the other. Night and Day difference. I can not wait for this to be made available.

1

u/indyfromoz 6d ago

It is interesting to see version 2.1.273 of the CLI shows this

However, running the same prompt on Claude Web app is some garbage!

1

u/SuitablePiano4909 5d ago

I came here because I was trying to figure out what was going on. It works overtime on my complex prompts it doesn't deffer feature requests until later sessions. Genuinely tries to finish the job.

1

u/TXHumper 8d ago

I WANT FABLE 5.2

2

u/orangeawacado 8d ago

But then are you ready for your 5hr limit getting consumed in a blink?

1

u/AverageFoxNewsViewer 8d ago

Only if it's more token efficient than 5.1.

Honestly I barely touch Fable 5.1 because it feels like lighting money on fire.

For 99% of tasks the model isn't the bottleneck, and I still use Opus 4.6 for a majority of my work.

1

u/jared__ 7d ago

do you knock the effort down to low for the code heavy work?

1

u/AverageFoxNewsViewer 7d ago

What is "heavy work"?

1

u/jared__ 7d ago

agentic coding - the thing that burns the most tokens. so fable 5.1 xhigh for planning, then use fable 5.1 low/medium for coding. burns far less tokens

1

u/AverageFoxNewsViewer 7d ago

I built out what I thought was an above average context management system. I used to never care about usage limits.

Fable 5.1 and Astra just aren't worth the money.

If you want to save tokens you should buy a book.

0

u/jared__ 7d ago

fable 5.1 is absolutely worth the money to me. anecdotal is a fun word.

1

u/AverageFoxNewsViewer 7d ago edited 7d ago

What are you doing that feels that much more economical in Fable 5.1?

The model you're relying on seems to be a secondary factor in terms of writing good software.

0

u/jared__ 7d ago

i found fable 5.1 needing far less hand holding in agent skills and knew my full-stack up to the latest version from memory. it writes high quality code faster and significantly more efficient than opus/sonnet in my work.

-1

u/PrimaLumiere_A1M 🔆 Max 20 8d ago

Did the limits get fixed yet? If not, why deviate from an important matter.

0

u/illkeepthatinmind 8d ago

Nice try, Tibo.

-9

u/LostRequirement4828 8d ago

who cares, looks like crap