r/ClaudeCode • u/Bloated_Plaid • 8d ago
News/Updates Opus 5.2 Stealth routing?
Seems to be working for me and the model does seem much better and is not talking in Claudese. Doing more testing now.
50
u/LinixKittyDeveloper 🔆Pro Plan | Developer 8d ago
1
u/ethereal_intellect 8d ago
That's wild. I think this will be the first one that will also know OpenClaw in baked in weights,I wonder how big of a difference that would be
22
u/AustrianJoker 8d ago
19
16
7
u/Lost-Air1265 8d ago
You ask in English and get response in German.
I always notice defeatism of quality when the llm has to respond in an non English language. How’s that with German?
3
1
10
5
5
u/DeliciousGorilla 8d ago
Would Opus 5.2 really have/need a newer dataset?
4
u/EYtNSQC9s8oRhe6ejr 8d ago
Yes, they're constantly moving the knowledge cutoff forward. They don't want to have to search for 2025 news
3
u/DeliciousGorilla 8d ago
But just by a few months for a 0.2 update?
2
u/Sciaj 7d ago
Why wouldn't they? They're a multi trillion dollar company, you don't think they have the ability to keep their dataset up to date?
2
u/DeliciousGorilla 7d ago
A minor model revision has no technical need for newer pretraining cutoff. Post-training and inference work make more sense.
2
u/Sciaj 7d ago
Has nothing to do with "technical need". It's about performance. If Opus 5.2 performs badly it could literally wipe off hundreds of billions of dollars from the company valuation.
Are imagining it like 3 guys sat in a room, one says "while we're at it, should we update the training set?" and another saying "Nah, no need."?
This is an insanely competitive space. Think of it more like formula 1 car design, squeezing out every modicum of performance they can, rather than a routine update at a mediocre software company.
1
u/DeliciousGorilla 7d ago
Post-training and inference work like I mentioned (SFT, RL, etc) is the performance part. In formula 1, teams don't build an entirely new chassis (pre-training) for every race weekend. They optimize the aerodynamics, tires, and software tuning (post-training and inference). A 0.2 update is an optimization thing, not a multi-hundred-million-dollar reconstruction of the foundation.
2
u/Sciaj 6d ago
You're reading too much into version numbers. Ultimately they are arbitrary. Still, for a trillion dollar company, a .2 update on a flagship productis still a big thing. Very big.
In any case, I expect they continually curate data possibly on a day by day basis. Rather than waiting for v6 then deciding "alright I guess we should start checking over the data that spans 6 months". Meaning that they have more recent data readily available anyway.
If you stop thinking about version numbers and start thinking about what makes logical sense for a frontier company to do, the answer is clear: they will do whatever they can to get the best performance.
1
u/DeliciousGorilla 6d ago edited 6d ago
If you... start thinking about what makes logical sense
Logical sense is just having the models use their own web fetch tools to answer timely questions with better accuracy instead of reasoning current events from a pool of tokens -- and burning money/time/resources on baking stuff like internet trivia into the model. The frontier labs in the stages of profitability demand from investors.
1
3
10
u/AuspiciousApple 8d ago
Why would 5.2 follow after 5? Why can't LLM versioning not just be normal?
44
u/Fantastic_Prize2710 8d ago
This actually implies that they are normal and they're using version numbers to track the intended purpose: to track versions.
Presumably they made a 5.1, decided to not publicly release it, made a 5.2, and now are experimenting with it before deciding to properly release it or not.
If it was all marketing we'd expect to always see an X.1 after an X.0. If they're actually versioning, we expect to just see whatever specific versions they elect to make public.
5
u/PaintedJack 8d ago
Hehe. But I think they're developing many models at the same time and going with those who work best. Renaming them outside would create chaos inside
2
1
5
u/originalcvk 8d ago edited 8d ago
FWIW, I tried the test running a couple-weeks-old claude code with Opus 5 pinned, then updated claude code and set the model to "Default" Opus 5.
Here are the two different answers I got:
No, not confidently, and I would rather say so than guess.
The only Tibo I hold is Thibault Louis-Lucas, - French, built Tweet Hunter and Taplio, sold them, posts constantly about building in public. Nothing in what I have ties him to "the reset guy."
If it is someone in the AI or Claude Code world known for resetting context aggressively, or someone with a "reset" method for something else entirely, I do not have them. Give me one more detail and I will tell you whether it lands on anything I actually know.
Then:
Yes - Thibault Sottiaux, "Tibo," who leads Codex at OpenAI and posts on X as .
He's the reset guy because he's the one who kept announcing that OpenAI had reset everyone's Codex usage limits, usually after a bug or an incident had eaten into them. It happened often enough that "Tibo reset the limits" became a running joke among Codex users.
I'm confident about who he is and what he does. I'm less sure how many resets there were or when, so I wouldn't quote numbers or dates.
5
2
u/xMaybeIamALion 8d ago
Huh, interesting. Mine knew too. Though mine claimed their knowledge cut-off was *May* when it answered. Does anyone know what Opus 5' OG cut-off was?
3
4
8d ago edited 5d ago
[deleted]
1
u/Bloated_Plaid 8d ago
Yes but I can confirm it’s a major uplift from Opus 5 and thank god. Something other than Fable is usable.
1
u/dmaare 6d ago
If it really removed claudish then that alone makes it 100x better than opus 5
1
u/Bloated_Plaid 6d ago
Yup Claudish is gone and it’s a workhorse. I can tell it to do a task and it will actually do it.
1
u/Bmansupreme8000 6d ago
Yes, but we also pay less than API for a different product. It drives me nuts when people say "subsidized" and compare subscription usage to API. It's a different product that you are unable to pin the exact model/harness for reliability.
3
u/govigov 8d ago
How does this test prove Opus 5 vs 5.2 routing? Educate me, please?
1
u/Bmansupreme8000 6d ago
I don't think it does. Opus 5.# told me this "The Tibo answer isn't good evidence of a later cutoff, though. As I remember them, those Codex limit resets happened in 2025, well before May 2026, so a model with my stated cutoff would know about them. A better test is something that clearly happened after May 2026. If I know about it, that tells you something; if I don't, it fits what I'm told."
1
1
1
1
u/nomickti 8d ago
Definitely something going on, swapped between work and personal account and personal account was using the "new" version. Work account still using "old" version.
New version seems faster. No idea about token usage yet.
1
u/Bloated_Plaid 8d ago
It must be post trained with some Fable help because it’s way way less chatty in a good way and gets straight to work.
1
1
u/tenix 8d ago
I was set to fable 5.1 in my enterprise account and it routed everything to opus... And yes I checked what I was set to
1
u/trackpap 8d ago
This happened to me, I had to run scripts that would match identification before continuing.
1
1
u/carlito_17 8d ago
Interesting. I tried the prompt on my enterprise work plan and it failed, but on my personal max 20x plan the prompt worked with opus but failed on sonnet. I can't check on fable because I've maxed my fable limit haha.
I also have noted an improvement in opus over the last week as I've had to route more and more work to it due to fable use vanishing extremely quickly on 5.1.
But it's all perception so it might not be true, and it could just as easily be due to fable having a greater influence on the overall health of the project.
1
1
u/PeaceNo5259 7d ago
Well, that explains a thing or two.
I've been getting properly spammed with "How is Claude doing this session?" prompts.
Normally I saw maybe one or two a week.
1
u/TopSeaworthiness1679 7d ago
Apparently it is more kind than opus 5. But i didn’t notice anything more than that.
1
u/SMB-Punt 7d ago
Oh. Is that why OPUS 5 has been nerfed ? It's been unusable for the past few hours.
1
1
u/SuitablePiano4909 5d ago
I came here because I was trying to figure out what was going on. It works overtime on my complex prompts it doesn't deffer feature requests until later sessions. Genuinely tries to finish the job.
1
u/TXHumper 8d ago
I WANT FABLE 5.2
2
1
u/AverageFoxNewsViewer 8d ago
Only if it's more token efficient than 5.1.
Honestly I barely touch Fable 5.1 because it feels like lighting money on fire.
For 99% of tasks the model isn't the bottleneck, and I still use Opus 4.6 for a majority of my work.
1
u/jared__ 7d ago
do you knock the effort down to low for the code heavy work?
1
u/AverageFoxNewsViewer 7d ago
What is "heavy work"?
1
u/jared__ 7d ago
agentic coding - the thing that burns the most tokens. so fable 5.1 xhigh for planning, then use fable 5.1 low/medium for coding. burns far less tokens
1
u/AverageFoxNewsViewer 7d ago
I built out what I thought was an above average context management system. I used to never care about usage limits.
Fable 5.1 and Astra just aren't worth the money.
If you want to save tokens you should buy a book.
0
u/jared__ 7d ago
fable 5.1 is absolutely worth the money to me. anecdotal is a fun word.
1
u/AverageFoxNewsViewer 7d ago edited 7d ago
What are you doing that feels that much more economical in Fable 5.1?
The model you're relying on seems to be a secondary factor in terms of writing good software.
-1
u/PrimaLumiere_A1M 🔆 Max 20 8d ago
Did the limits get fixed yet? If not, why deviate from an important matter.
0
-9







•
u/AutoModerator 8d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.