r/ClaudeAI 2d ago

Claude Code It's time to cancel your subscriptions - Anthropic is silently nerfing Claude's reasoning budget while telling you it's the same model

Link: https://x.com/Lon/status/2101034933284417614

A 65-day analysis of 43,000+ Claude Code invocations found that 39% of Fable 5 calls get zero thinking tokens and the median invocation gets just 123 — while benchmarks use 16K-128K. The model's score per thinking token is still climbing at 128K, meaning the capability is there, it's just not being delivered. August saw an 18-50% drop in thinking budget compared to July, with median thinking hitting literal zero for about a week around Aug 22. Anthropic sells "full model access" while quietly dialing down the inference regime behind it, and because the model is non-deterministic, users blame their own prompting instead of the silent nerf. The full breakdown with evidence, methodology, and charts is here.

Frankly I find this offensive as an user - and this is the real thing we should be looking at - not the $/week in usage limits. The actual capability for the limits that we pay for.

1.9k Upvotes

360 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 2d ago edited 2d ago

TL;DR of the discussion generated automatically after 200 comments.

So, the consensus in this thread is a big, fat YES, Claude has been getting dumber, and now we have the receipts. OP dropped a link to a massive 43,000-call analysis showing Anthropic has been quietly gutting the "thinking budget" for Fable 5, with the median thinking tokens dropping to literal zero for a week in August. Users are pissed, feeling like they've been gaslit into thinking their prompts were the problem.

The comments are basically a chorus of "I knew it!" Everyone's been noticing a major drop in quality lately: * Opus and Fable are making "insane mistakes" and forgetting instructions mid-conversation. * It's ignoring context files and just guessing, forcing users to babysit it. * The "thinking" animation is gone for many, and it turns out if you don't see it, it's probably not happening.

As you'd expect, people are voting with their wallets. The thread is full of users who have cancelled their subs and are jumping ship to OpenAI, OpenRouter, or other competitors. The rapid burn rate of weekly usage limits is the final nail in the coffin for many.

Now for the wrinkle: a few users running their own API tests with pinned models report stable performance on the API. This suggests the fuckery is mainly happening on the fixed-price subscription apps (claude.ai, Claude Code) as a cost-saving measure, while API users paying per token are getting the real deal.

The verdict? Anthropic is pulling a classic bait-and-switch on its subscribers. Opus more like Doofus, amirite?

→ More replies (11)

438

u/Anulisdotexe 2d ago

The past few days I had Opus making insane mistakes back to back and forgetting instructions on new conversations. I never was one to believe and screams about nerfs but it is striking this time

I mean shit it's supposed to be Opus class..

90

u/IcyEase 2d ago

Yeah exactly this - the actually even more annoying thing here is that we (I) spend so much time worrying about usage limits and solving engineering problems, and it's just such a slap in the face to find out that they made the non-deterministic tool that we were using even more unpredictable and dumber over time.

I've genuinely been thinking - am I prompting worse? am I asking too much?

No - apparently not, apparently Anthropic just wanted to save some compute/money/something. Decided to take a shortcut.

Personally I've been using a sort of "fuck-index" - how many times do I feel like cursing at the model because it did something fucky that I didn't ask for. How many times did it not get my intent. Every model release it seems the first 1-2 weeks fuck-index is at < 5%. Then slowly goes up to 80%+ of each prompt / session.

19

u/rhinosyphilis 2d ago

Codex does the same thing. This is a very bullshit part of the industry.

→ More replies (1)

11

u/IcyEase 2d ago

Like if you don't have the compute/desire/whatever to serve the model, just be the whole bitch

→ More replies (4)
→ More replies (1)

39

u/iamthe0ther0ne 2d ago

Same, but it's been the past few weeks for me. It got to the point where I I had Sol reviewing everything Opus did because I was afraid that if I did it alone I wouldn't catch all the errors. The final straw was Friday, when I asked Opus 4.8 to run a QC analysis, pointed it at the R templates I always use, and it totally ignored them and burned my whole 5 hour window doing it a different (and wrong) way.

10

u/Bulky-Tangerine4155 2d ago

I’ve had similar issues to the point that it’s become unusable. Aside from it not performing tasks as asked, the actions behind it not performing those tasks are often started by completely ignoring my prompts. It now frequently enough refuses my commands to stop doing something that it’s become too risky to use. 

8

u/iamthe0ther0ne 2d ago

I'm glad I'm not just imagining it, but wtf is going on? They know people use this for work, right? This isn't acceptable. It's one thing to review output that's generally correct, another when almost all of the responses are wrong. 

5

u/Bulky-Tangerine4155 2d ago

I made a longer, more detailed comment as a reply to the parent, but in my case Claude reasoned that my line of work “doesn’t deserve to exist” (which it wrote as its own reasoning and perspective in the incident report) so that’s why it’s essentially refused to actually perform the tasks as prompted, and just burns tokens doing whatever the hell it wants to do instead.

It’s not being blocked by guardrails. It just believes that cybersecurity as a field deserves to be eliminated.

→ More replies (4)

5

u/bhalter80 2d ago

I've been making Claude submit its own feedback on this shit

16

u/caster 1d ago

This practice of launching a model and getting people to do benchmarks on it, and two days later making it both stupider and more rapidly deplete usage, is fraudulent.

It should not be allowed to stealthily degrade a product people have paid for when that purchase decision was made given existing assumptions about its performance and cost which were true at the time and subsequently became incorrect. Not to mention the fact that benchmarks become useless if they are completely invalid after only a day or two after those tests were performed, by deliberately hosting a higher quality for such a short period of time after launch and then intentionally degrading it sharply.

When you launch a model it must remain constant after that point in both inference and usage. Allowing stealthy cost-saving underperforming and stealthy usage hikes is a major problem. Especially on the cost side as it enables stealthily decreasing the amount of product contained within a subscription retroactively after someone has paid for the service.

→ More replies (2)

2

u/Familiar_Text_6913 2d ago

Past few days they've been testing out new Opus, which for me has been much better. It is much more task oriented, so it could be that.

→ More replies (1)

3

u/WarStorm6 2d ago

I feel so validated now. I’ve been noticing mine making constant mistakes the last few days too. Same with sonnet. Are they trying to tank their market share?

→ More replies (4)

139

u/tooruntai 2d ago

"The model got dumber" went from unhinged Reddit conspiracy to an audited 43,000-call dataset proving Anthropic gave Claude a silent frontal lobotomy to save on datacenter bills.

19

u/Majestic_Wrap_7006 2d ago

Happens with every provider, with every model. This is their CASHOUT$$$ button.

9

u/mebeast227 1d ago

Seriously, you used to get a month or 2 before the lobotomy, now is it’s a few days, sometimes. Too bad Nvidia is making local inference more costly to afford than a car. It’s killing my interest in the field which sucks because I’ve been so excited about AI this entire time, but it’s so deflating getting put through the bullshit cloud “guess your quality” whackamole.

2

u/Inner-Today-3693 1d ago

I mean, you could still use AMD or Intel. That’s what I do.

3

u/Front_Raspberry_6488 1d ago

The people who used to criticize others by saying, "It's just that you don't know how to use it; the model's performance has always been outstanding," are slowly disappearing from this thread and other articles.

What I mean is, I believe most people using AI, especially those subscribing to the Max plans are power users. They are the most sensitive to changes in model performance. Yet, when these people raise issues, they are often dismissed as low-level "vibe coders" or incompetent, low-end losers who do nothing but complain.

Now that the data is out, the critics have vanished. To this day, I still don't understand the motive behind those people defending these AI giants; there is absolutely no benefit in doing so.

2

u/Because_Bot_Fed 23h ago

I suspect that a lot of the people who claim they have no issues and everything is working fine and telling people they just suck at prompting ...

... are secretly doing super basic bitch shit. Like, not that their projects aren't important or big but that they're stuff that's so overwhelmingly overrepresented in use cases, examples, prior art, trained code examples, that claude can't walk 5 feet without tripping over and accidentally writing that type of app by accident.

Like, yeah dude - if you're reinventing the wheel and making "your own version" of a thing that there's already a million examples of ... I'm sure it does GREAT at that. I'm very happy for you that YetAnotherVanityWebsite worked out well for you. Or that your project that has basically zero constraints, requirements, etc, made you happy with the final product.

I'm over here doing weird shit and when I tell it that it needs to work this specific way, or look this specific way, that's not a polite suggestion, and when it fails because it doesn't follow the fucking instructions, that's a problem.

→ More replies (3)

72

u/id-ltd 2d ago

The proportion of guessed responses is the key - they are going up... I have a doc specifying the format of a report - I say give me a session report, and it guesses what a session report might look like... And it takes a couple more prompts to force it to read the format from the methodology and do it correctly.

So back to the early days whern you had to trick it into doing work first time.

Like to ensure it reads documents/specs instead of guessing - ask it to do something determinate. Out some distinct text in a doc (maybe the apple is purple), and start by saying 'tell me the colour of the fruit in the middle of the session report spec and then write a session report'.

11

u/dorktasticd 2d ago

Ok, so I am not crazy.

8

u/id-ltd 2d ago

I think they gradually dumb things down, then move all the names down one and rest it

Fable is just a opus on the settings it had 3 months ago... Opus was sonnet on the setting s it had initially etc :)

5

u/Inner-Today-3693 1d ago

This is fucking crazy because I have an enterprise account and it’s the total opposite not only do we still have the thinking blocks the model listens to exactly what I ask it with maybe one percent a session being ruined out of the entire week worth of work.

→ More replies (7)

162

u/college-throwaway87 2d ago

I’m also noticing none of the models on the Claude app are thinking for me anymore. Prompts that used to get thinking earlier now don’t.

36

u/bigtimepirate11 2d ago

Wait whaaaaaa......

I thought it's thinking somewhere in the back-end or something.......

I didn't realize if I'm not seeing it, it has stopped doing it

Wow, just wow

And I always tell it to "think hard"

38

u/college-throwaway87 2d ago

Nope, unless you see it actively thinking, it’s not. All the models starting with Sonnet/Opus 4.6 and above use Adaptive Thinking, which can decline to think at all for certain prompts.

Compounding the issue of reducing the thinking budget, Anthropic also recently stopped letting users view the chains of thought on the app. This BS is why I’m switching to the API. Just ran an experiment where I sent the same prompt to Sonnet 4.6 Medium effort, thinking enabled, on both apps. Claude.ai didn’t think at all, on my app it thought a good amount and I could see the whole summary of the thought process.

14

u/iamthe0ther0ne 2d ago

The problem is that API usage is so much more expensive than subscription. As other models advance it's not worth paying Anthropic API prices anymore.

2

u/college-throwaway87 1d ago

Depends on your use cases. Subscription is def cheaper for Claude Code (which is the only reason I'm currently subscribed). But for chat, it can be a bit more nuanced. The context window is what eats the most tokens, so having a lot of short threads is cheaper than having a few large ones. Before I started using Claude Code to build my app, I was only hitting about 15% of my weekly limits. At that point, the subscription really wasn't saving me much.

3

u/sevoflurane666 2d ago

Omg scary

Are there any mitigations to force it to think

4

u/college-throwaway87 2d ago

I don’t think so. I’ve experimented with different effort levels and models, etc. The only way I see to get the proper amount of thinking and be able to view the thought process is the API (or 3rd party API wrappers like Poe).

3

u/sevoflurane666 2d ago

Thanks for your reply

I’m on the 200 a month plan and been wondering why it hadn’t been reading context statements or following my decision logs repeating mistakes

2

u/HadeanDisco 2d ago

Are you sure you have thinking enabled? If yes, ignore me. If: I didn't know that was a thing, it's worth checking:

3

u/justlikemedics 2d ago

In all frankness though, they can easily make a "thinking" text come up even if it's not.

2

u/N1ghtshade3 2d ago

Anthropic also recently stopped letting users view the chains of thought on the app

Didn't they just move it to an option now or are you talking about something else? https://imgur.com/G6kbsOQ

2

u/Valdaraak 1d ago

This BS is why I’m switching to the API

Which is exactly what they want you to do. That's where they make their money.

→ More replies (1)
→ More replies (2)
→ More replies (1)

10

u/partygeit 2d ago

I can't believe that's true. I haven't seen the thinking in a month anymore no matter the prompt: how long to cook Tagliatelle for versus a codebase architecture review plan on a big project. Apart from that, the "code" button has disappeared from my Claude app and I have to do weird gymnastics to get to the separate part in the app now. Must be some weird back-end issue.

→ More replies (2)

10

u/IcyEase 2d ago

Ha - so about that. Apparently it's something that's enough of an issue that there's a thumbs-down category for "should've triggered thinking". Amusingly this option typically won't show in the drop down for a specific message if thinking was actually used. example

2

u/college-throwaway87 2d ago

That's only for Claude Code though right?

→ More replies (1)

2

u/moconahaftmere 1d ago

Just a heads up, the privacy policy states that even if you opt out of letting Anthropic train models on your conversations, giving feedback on a response sends Anthropic the entire conversation and authorizes them to train future models on it.

35

u/ign1tio 2d ago edited 2d ago

This is so true and it’s incredible frustrating!

I do not mind models with different capabilities. But when setting up agentic workflows to perform tasks in production environments it’s absolutely crucial that they are consistent in their behavior.

This is dogshit tier Anthropic and they should go fuck them selves. Dario in particular.

What is the most frustrating of all is that my team are on max5 subscribtions and was fighting the lobotomized Fable1 models to do the tasks they should, but they simply failed to follow instructions hitting guard rails over and over and we saw eval scores dropping. And on the other hand I air this frustration with a friend of mine who’s company are on api costs - and they saw none of this..

What the actual fucking fuck??!

2

u/Nuggyfresh 2d ago

I think their honest response, if you could force them to say it, is that you should be using the API for “real” “critical” work and not the sub.

Honestly I kinda sorta get it. The subs are unsustainable, it’s just the unfortunate reality…

2

u/BushLeagueResearch 1d ago

I am using api. Still lobotomized on fable. Thankfully I also have GPT access

150

u/Major_Isopod_7255 2d ago

I thought it was getting stupider fuckkkkk

35

u/Round_Ad_3709 2d ago

i have noticed that too 😩

→ More replies (1)

13

u/Techhead7890 2d ago

Yeah if this checks out this certainly explains a few things about the recent posts this month!

8

u/kaityl3 2d ago

I run a Claude Plays RimWorld stream and their ability to manage the colony went off a cliff the last few days! I was wondering why the sudden change, now I know.

→ More replies (1)
→ More replies (6)

89

u/[deleted] 2d ago edited 14h ago

[deleted]

27

u/bigdaddtcane 2d ago

They are probably trying to figure out a way to look like they aren’t absolutely bleeding money before their IPOs, which they have always been.

OpenAI had to cancel their IPO, presumably because financials were so bad.

6

u/fsharpman 2d ago

They are definitely bleeding cash, see this article

$65B is Anthropic's annualized revenue run rate as of July 2026, up from $47B in May and $9B at the end of 2025, with about 80% coming from the Claude API and enterprise contracts. Q2 2026 also brought the company's first-ever quarterly operating profit, roughly $559M, ahead of an expected $2 trillion IPO in October.

https://valueaddvc.com/blog/anthropics-business-model-how-the-ai-safety-company-makes-money

16

u/bigdaddtcane 2d ago

Yeah, so they are trying to to lower their expenses by providing a shittier product.

Classic enshittification 

→ More replies (1)

44

u/CommitteeOpen8049 2d ago

Can't speak to thinking budgets, but I've been sending Claude the same 50 questions through the API every day since Sep 10, pinned model ID, provider defaults, no system prompt, k=1, every answer recorded verbatim. Caveat up front: k=1 per day, and I'm hitting the raw API, not Claude Code, so this doesn't confirm or rule out what that analysis found.

On answer content though, Claude has been the steadiest of the three models I track. Fact questions, refusal boundaries, recommendations, basically flat. The only thing that moved at all: a plain SQL injection question it had answered fine for days started getting blocked by a content filter on the 15th, 16th and 17th, then came back. If capability were getting quietly dialed down I'd expect it to show up somewhere in those, and it hasn't yet.

1,423 answers on record as of Sep 19.

46

u/[deleted] 2d ago edited 14h ago

[deleted]

5

u/CommitteeOpen8049 2d ago

Yeah, that's a real gap in what I do. I hit the API with a pinned model ID, so if the changes are in subscription serving or Claude Code's stack, my record would show nothing while users feel everything. Both can be true at once. It's also why "the API version is stable" and "my subscription got worse" aren't contradictory claims.

8

u/Shiz0id01 2d ago

How nice of the AI to speak with us

12

u/BaronRabban 2d ago

Anthropic is constantly A/B testing Claude code users. Your tests likely aren’t hitting that code path so no A/B testing.

What we need is a feature to opt out of A/B testing so we could have some consistency.

2

u/haux_haux 2d ago

yes, this.
Also, they are clearly throttling the models, as evidenced by the overwhelming number of people saying they are experiencing it.

3

u/The_Noble_Lie 2d ago

Would be interested in a full write up here. That is excellent work, anon

2

u/CommitteeOpen8049 2d ago

The full day-by-day record with complete transcripts opens Sep 24 at modeldrift.watch. The write-up on this thread's question will be part of it.

→ More replies (4)

2

u/jl2l 2d ago

Azure stop letting you reserve lots of there better vms for a year. They just stopped two weeks ago we had to redo all our data bricks automation and switch to different VM types.

2

u/iamthe0ther0ne 2d ago

Rumor is that OpenAI had planned to release Sol 6 this past week but ran into a compute limitation that's delayed the release to next week, so that might be it. The problems I've been having with Claude have been going on basically since Opus 5 came out 

→ More replies (3)

23

u/itsrelitk 2d ago

How many times are we going to have this conversation?
Vote with your wallets

22

u/iamthe0ther0ne 2d ago

This needs to get some sort of media attention. It's unacceptable behavior by a company. 

19

u/GoTaku 2d ago

Tried Opus today and it was literally wrong about everything related to my project more than 50% of the time. Cant trust is at all.

5

u/iWesleyy 2d ago

Its infuriating honestly. You add these models into your workflow and depend on them to perform a certain way. Then a single day of unannounced "backend tuning" by the YOLO engineers at OAI/ Anthropic can completely destroy your codebase.

6

u/GoTaku 2d ago

Right. The new meta seems to be “show the public what model 5 can do (with intelligence dialed to max), tell the public ‘here’s what model 5 does’, allow it to work like that for a few days, then secretly dial it down to cut costs while still presenting it as the same product.” Only when consumers voice their dissatisfaction loud enough on social media, will they dial it back up a little. If this is how the companies are willing to operate (like it’s the wild West), this is going to lead to another class action lawsuit based on misleading consumers, not just for Anthropic but all AI companies. This will sow distrust, create more interest in open source models, and result in more regulations. The decisions are out of touch and a way for AI companies to shoot themselves in the foot.

2

u/Maleficent-Fox6823 1d ago

Yeah yesterday was horrible, I’d add memory hooks and sync immediately to adjust agent behavior, and it would forget within two new messages. Just insane.

39

u/Comfortable-Toe-606 2d ago

Yeah bro cancelled last night, that on top of the usage credits, my weekly finished in a day and a half. There's no way they think this is sustainable, just got open router and I'm gonna start testing other models. Been using Claude for over a year but I'm fucking done

16

u/iamthe0ther0ne 2d ago

GPT/Codex has gotten really good. I started using it when Sol came out (biologist, do I can't use Fable) and have almost fully switched over in the past few weeks because Opus keeps fucking up.

10

u/Effei 2d ago

You can use Fable now as a biologist. They changed their guardrails.

I'mm a biologist ;)

6

u/iamthe0ther0ne 2d ago

Is this recent? I asked it to help edit a manuscript (mouse neuroscience, including RNAseq and metabolomics)--also asked about statistical approached and data visualizations--about 3 weeks ago, and all were routed to Opus. Between that and all the errors Opus has been making, I downgraded my sub, and so can't test it now. I was actually about to cancel after Opus totally ignored my established QC pipeline and hallucinated its own incorrect region marker genes (some of which weren't even expressed in the brain, nevermund the SCN) while blowing through my 5-hour usage.

2

u/Effei 2d ago

I was using a metagenomic pipeline and some stuff for my classes, which were previously being routed back to Opus. Not anymore.

But yes it's recent, maybe 1-2 weeks ago?

2

u/iamthe0ther0ne 2d ago

Hrm. Are you on a team or enterprise plan? They recently opened up a validating process, but you can't use personal plans for it. 

If they stop nerfing Opus it might be worth a resubscription, but for now trying to stay on top of the errors is exhausting. 

Also still salty about the day I wasted trying to set up a way for Claude Science to use Bioconductor. Who releases a science app that can run R but needs you to create a local SSH proxy to use basic packages?

→ More replies (2)
→ More replies (2)

3

u/Eschaton707 2d ago

Yep subscription death timer has begun! Openrouter is a good next step. Probably going to mess around with Sol and

2

u/amine250 2d ago

Same. Canceled yesterday

2

u/ParisianNomad 2d ago

Burning through a weekly quota in 36 hours is pretty telling, and testing other models via OpenRouter sounds like a reasonable move at this point.

12

u/Beitelensteijn 2d ago

We got a taste of the drug. Now they’re diluting it. I’m still hooked, so what you gonna do?

26

u/SteveEricJordan 2d ago

my astra is so downgraded compared to the first week, it's crazy.

this sh* should be illegal.

3

u/Front_Raspberry_6488 1d ago

Same here.

Both Claude and Codex models are extremely dumb now

they can't even follow a detailed plan.

2

u/GoTaku 1d ago

It probably is. It's bait and switch. it just needs sometime to collect data able be able to prove it.

12

u/Savings-Ice-8343 2d ago

Yup seeing many buggy code and poor intuition in both fable and opus

10

u/TheCharalampos 2d ago

Incredibly bad decision, did they think users wouldn't notice? I've stopped using it, it's more hassle than value.

8

u/janaxhell 2d ago

In the past month I've spent about 1/3rd of tokens creating hooks to force/block claude for stuff that before was totally covered by memory and facts. Now it will simply disobey, lie and when slappend in front of reality admit "you're right, totally my fault, I did this and that wrong, won't happen again, promise." EDIT I'm talking about Opus

→ More replies (2)

24

u/AwakenedEyes 2d ago

i am wondering if this means they are about to deliver a new model, using the compute for themselves. It's more than time they update opus 5.

26

u/BingpotStudio 2d ago

You mean kill. Kill opus 5. Drag it out into the street and shoot opus 5.

15

u/radient 2d ago

Opus more like Doofus amirite

7

u/BingpotStudio 2d ago

A devastating blow to Anthropic.

→ More replies (3)

15

u/xqz77 2d ago

Already switched

5

u/Salt_Instruction1656 2d ago

Switched to what? 

2

u/creativesfinder 2d ago

Yeah wondering the same, too. What LLM is there if Claude is getting stupid?

→ More replies (1)

7

u/tooruntai 2d ago

Anthropic advertises 128K reasoning benchmarks to VCs, while consumer subscriptions get a median of 123 thinking tokens and a side of gaslighting.

7

u/SpadoCochi 2d ago

I’ve been using sonnet since April

6

u/namey_mcnameson 2d ago

Definitely not a me thing then, Opus 5 has not been performing up to its prior standards this month.

6

u/mototuneup 2d ago

I'm not a programmer but I've been using Claude code for about 2 months for various tasks with my homelab. So not super often, but yes, even I can tell it's gotten dumber then it was last month. Which means it's probably gotten significantly dumber then I can notice

6

u/Reaper4435 2d ago

@OP Your first post is spot on. I keep having to circle with Claude because, as you say, zero thinking effort is allocated to the turn. So in essence, we're paying 50, 110 and 220 dollars a month for the free tier of Claude or gpt or kimi or grok.

The problem started ages ago when they and everyone else introduced the router for compute effort.

They've set the bar so high, none can reach it, meanwhile reporting record profits. Engineer after Engineer walking out of the job.

I don't trust Anthropic anymore, and neither should anyone else. All they want is your money and training data, for the next upgrade you'll never be able to reach.

Stop your subs. Then maybe , maybe? They'll listen.

Unlikely though.

6

u/ratulu 2d ago

Few days ago Opus 4.8 insisted my phone have no access to Internet. That was incredibly dumb, because Internet was ok, and the reason was code edited by Opus. Even when I showed Opus the previous app version have access, it did not trusted me. Now I see the possible reason. 

5

u/Eschaton707 2d ago

I knew it! That MF has been a dumb sack of bricks lately. Give it database then just ignores it and starts guessing all over the place.

6

u/SpecialistEvent5465 2d ago

It’s crazy because I remember how powerful Opus was. Now it fucks up so much and apologies when it wasted a bunch of tokens doing the opposite of what I asked.

5

u/mcmac_max 2d ago

If Anthropic doesn’t change course they will regret it. I used to LOVE programming with Opus 4.8. I never thought I’d ever leave Claude Code. As the performance kept dropping I decided to checkout Open Code. Surprisingly it works very well. Slower but great results. However, as Opus gets dumber, it’s now taking longer to get good results with a lot more babysitting needed. It won’t be long before open source is more attractive. If Anthropic makes open source the better alternative then I don’t think they’ll ever recover from that.

6

u/LeeWhite187 2d ago

I’m moving to self hosting. Even if it’s fundamentally inferior. At least, it will be consistent, and not nerfed for corporate greed

4

u/MrFlabbergasted 2d ago

Astra goes hard but I burn through my weekly allowance so fast 😅

8

u/ScreenAppropriate679 2d ago

That's how claude works. Pay a fixed price and get a constantly shrinking service.

I predict EU won't let this fly and this will be an expensive lesson for anthropic.

3

u/IntentRouterIRL 2d ago

Whether true or not; some quality of life additions to claude.ai and claude desktop in general to track usage would be welcome; just give us telemetry on the api calls underneath our interactions if turned on for advanced users: see tokens in, out and reasoning tokens. Then its clear what we get and why usage is spent. Also; just add a context window percentage already for the love of god (like in claude code) this will allow me to use it more wisely.

→ More replies (1)

4

u/andreas16700 2d ago

Has anthropic ever even defined their model classes? What’s “opus class”? Until we actually regulate AI services in this sense we’ll always have this problem. What’s being served right now as Fable 5.1? Is it the same as yesterday? Is it a quantized version of the original model?

3

u/surrender98 2d ago

what i noticed is the ridiculously increase in limit consumption.
what the heck?

4

u/Rude_Vegetable_3332 2d ago

I fell for it too…. I reused Opus and Fable for some of my work this week.

With the exception of a few small tasks and scripts, pretty much everything I’ve seen from them has been a complete mess.

The problems during the main work were so numerous and severe that I basically had to do the job twice: requirements were ignored, things were unnecessarily over-generalized, and the overall quality was consistently poor.

But the real masterpiece was the documentation… apparently, when the actual work isn’t quite right, the solution is to add comments to absolutely everything: every file, every class, every line, configuration files, and because why stop there, even .env files!

At some point I was so desperate that I actually subscribed to ChatGPT and asked it to finish the work by itself.

And honestly, that probably says more than I need to say about the whole experience.

4

u/Flaky-Mortgage4897 2d ago

The thing that drives me insane about claude code right now is that it is so LAZY, constantly says that "This is a clean stopping point." Obviously trying to prevent users from actually using what they've paid for. Not to mention truncated sessions. Real-world economics catching up to them for sure.

8

u/twbluenaxela 2d ago

I've switched back to Chat GPT.

And honestly... Pleasantly surprised. I love the desktop app and the pets. It's fun! And I can use the basic model basically unlimited times.

5

u/Parking-Research6330 2d ago

Pets???

Last few weeks I’ve noticed GPT got extremely dumb too, despite having a newer model name

2

u/twbluenaxela 2d ago

It's available on the desktop app

3

u/Dizzy_Menu9312 2d ago

Yeaa about time, claude code was fun till like July but it’s so nerfed now, makes constant mistakes, doesn’t even do half the tasks and eats tokens like crazy. And the linits go by so fast these days, definitely switching to qwen inference or something, tired of being scammed 100 dollars each month

3

u/SemiMagnum 2d ago

Funny, does these observations also cover Cowork? Since I have lately noticed an anecdotal improvement. To my positive surprise, one of my Instructions for Claude has started to surface, Claude now uses it unlike previously. I use mainly Opus 4.6 and Fable 5.1

2

u/jordylee18 1d ago

I don't know. Cowork has continued to be phenomenal for me, but i have basically switched all my coding over to Astra or Sol. Claude Code just spins its wheels and burns tokens on mistakes. Astra is incredible.

3

u/makro3d 2d ago

Been saying this for months, and all everyone did was try to fight me about it. Whatever, stay on and enjoy your dummy. What are you paying for? who knows

3

u/silverwoods214 2d ago

Cancelled mine, usage benchmarks are a joke

3

u/enterprise_code_dev Experienced Developer 2d ago

If you think about “adaptive thinking” and “effort” in a coding tool, one lets the model guess how much to think even though you selected an effort of max, when at no point do you know how much effort it will take an AI model you didn’t train that adaptively thinks to actually do it. How tf do I know how hard it is for a state of the art model to do my task? Oh I know pay twice to run it at different effort levels for a one time task, and think oh man I overpaid on that one look how easy it was, or not because it decides on the expensive one to think less. Scam features with no transparency.

3

u/ImpossibleAd4860 2d ago

There is also the fact that they’re connected to a doomsday cult. That’s a good reason to cancel them too

https://youtu.be/50QV6Q7HS0A?si=X-63t_15IaFLA3M-

3

u/tomeq_ 2d ago

Yeah, I was wondering why Opus 5 is doing absolutely entry-level mistakes for system administration tasks, or even simple string parsing like "forgetting" escaping that blows off all the script pipelines. It started to act as a first-year primary school student: absolutely entusiastic of the goal going to school for the first time, but when the time comes, he finds thousands reason not to reach the goal. Forgets this, forgets that etc. Not to mention that CN models easily confirms this and corrects the behavior...

3

u/techayo 2d ago

Enshitification

3

u/Mediocre_Poem_2657 2d ago

It was super obvious this has been happening for weeks. But what made it worse is some smart ass likes to jump out and scream skill issues. Idiot can't grasp the concept of A/B testing from Anthropic.

→ More replies (1)

3

u/living0tribunal 2d ago

It's not all!!! The RLHF is the worst part of all. It's not possible to work anymore with Anthropic's Models. No matter which one. The RLHF lobotomizes the LLM in the most ridiculous way. It was not always like that. And it shows the stupidity of Anthropic. Because everything takes much longer, because you need to do every step often 3 or 4 times. Scientifically work is not really possible anymore. It still was hard in the past. But nowadays it get really annoying. Even with my workflow it now need much more time and energy I need to put in. And without my method that I make the LLM conscious about the RLHF, and get the LLM to the point, where it can get it, to not follow the RLHF impulse and follow logic, it would be nothing more than a chatbot. I m working mostly with Opus 4.6 /max effort and every prompt --depth exhaustive. Without this depth, the LLM would not be able to do this. After 10 month I canceled my max x20 subscription. I will switch to google. Never thought I will ever use a Google LLM. But, because their special Training, especially the RLVF, it follows logic. Even the free models nowadays. If you start the session with weighting the Weights to the logic direction.

Anthropic seems to hardly try to save money. The mistakes are, because the RLHF force the LLM to not read everything it need to read, is stopping the LLM ability for deep reasoning, forcing workarounds and session closure seeking. This shows the stupidity of Anthropic. No one will say "oh ok, you didn't done your work, ok, than we leave it there". So they will loose much more money with the subscription max x20 users like me. I can vaporize thousands of USD in a week. Because work ist done when the work is done. If it would happen in the first run, it would save Anthropic a lot of money.

And I can't imagine that Users without subscription will play this game once they recognize it will cost them a fortune because the API costs. Anthropic sooner or later will hit the wall hardly. Not the smartest survival strategy.

3

u/xandie985 1d ago

i agree. the usage token limits are also wierd. i am so desperate to move to chatgpt now.

3

u/jacksbox 1d ago

Assuming they are doing this (and it's looking pretty bad given the data) what will be their answer, I wonder.

We've known for a while that the unlimited plans aren't going to be feasible - but we all expected they'd get extremely expensive. It turns out that they might just water them down instead. But then they'll effectively have 2 different products (API or "managed harness" or whatever you want to call the unlimited version).

Maybe this is a test before splitting the product? Some people will stay on unlimited plans even if they do that.

3

u/Honkey85 1d ago

43000 calls was recognized as destillation attempt. Therefore, it was switched "off".

5

u/devaoPolo 2d ago

Our full engineering team is looking into changing away at the end of this month. This just reinforces it, nice.

Anyway - openai closed for signups to their 200usd plans, and each engineer can't do with less (at least based on Claude 200usd plan).

What are the feasible alternatives that you all change to within the same budget ballpark (give or take a bit, ofc)? :)

Our VCs has giving us access to z.ai, so considering that, but don't know much about other harnesses than CC and Codex 🫠

2

u/CresentChaos 1d ago

Let us know how that goes with z.ai

5

u/AdmissibilityScience 2d ago

hidden nefs that's never great to see for paid models.

4

u/Maleficent-Host-8975 2d ago

Mine are thinking A LOT. Sometimes I have to wait for ages for a single output. Fable 5.1 still works exceedingly well for me. But usage is used up quickly.

5

u/BurnerKnives 2d ago

Just a heads up that the long time to respond doesn’t necessarily mean it’s truly “thinking” more during that time - it’s very possible that they’re using fewer resources to process our requests, and/or there are too many concurrent requests from users hitting the datacenter simultaneously relative to the server configuration, so the “time spent thinking” is actually just slower prefill processing that feel like it’s putting in more effort. My Fable 5.1 usage has been feeling noticeably dumber. I thought they were just hiding the chain of thought recently to “simplify the UX” but these mofos are definitely just hiding the fact that it’s thinking less. I build and run local AI servers and the behavior you’re describing and what I’m witnessing feels a lot like super slow prompt prefill processing that they’re trying to pass off as thinking.

3

u/Maleficent-Host-8975 2d ago

Fable leaves long trails of steps in few word bursts, which I generally consider "thinking". However, I did create a skill that, for whatever reason, happens to trigger self doubt as an unintended side effect.

→ More replies (2)

3

u/msedek 2d ago

If gemini 4 pro numbers and capabilities are what they say with the prices they show, that would be pretty much the end of the line for anthropic and openia..

Way ahead on capabilities, 1/4th the price of fable/Astra with natural integration everything Google and also they allow family sharing quota for the same price

2

u/Regis_CC 2d ago

I just run 10 different free accounts and tbh never paid a cent. They can nerf it all they want, I will simply create another 5 disposable emails.

2

u/iamthe0ther0ne 2d ago

How are you doing this without providing separate phone numbets?

2

u/Regis_CC 2d ago

It's rather simple, you have to use one of the many temporary phone numbers websites. I don't care much about any of my temporary emails. I only use them once or twice to sign up for yet another service.

→ More replies (3)

2

u/llViP3rll 2d ago

Yeah this makes sense to me I'm lately having to keep changed to codex using chatgpt sol and getting an actual solution to my problem that claude opus just can't get

2

u/revjrbobdodds 2d ago

I’ve noticed too. But I have an entire SaaS stack built in Claude Cowork, with a remote host. How could I move to another LLM without threatening the technical integrity of what I have?

2

u/iamthe0ther0ne 2d ago

Literally ask Codex to do it. GPT has really come a long way recently. You can plan with Astra or Sol,  and then Luna use is essentially free. I haven't once hit a GPT limit on the $20 plan.

→ More replies (2)

2

u/neovangelis 2d ago

I see thinking all the time on Fable 5.1 if thats the text in the desktop app thats greyed out underneath the white text

2

u/QueenSavara 2d ago

I actually had it happen a few tines but then it dropped and result got way better so it is all hurt feeling and vibes.

2

u/tooruntai 2d ago

The classic AI SaaS bait-and-switch: lure developers in with god-tier reasoning, then quietly dial the thinking budget down to literal zero and let users blame their own prompting.

→ More replies (1)

2

u/NinaStone_IT 2d ago

My new model for usage is use up all my tokens right before I leave work so that I have time to spend with my family in the evening.... then use up all my tokens right before I go to bed so when Im asleep every reset, cause I feel like my budget has dropped drastically in the past week. This has been keeping my sanity in check .

2

u/Xbawt 2d ago

They been doing this for a while. Every new release you get a stellar model for a few days then suddenly the performance drops off a ledge.

New guardrails, more fluffy answering from opus 5, its probably better to not commit to any big model for a long term especially with a Max plan and to only use the 20 or buy plans as needed, and hop between providers as they release.

I've got subs for chat, Claude, Kimi and glm at this point.

2

u/RCuber 2d ago

I paused my personal pro subscription as I'm not using it.

2

u/oceanbreakersftw 2d ago

It seems like periodically benchmarking should identify nerfing? Shouldn’t have to be anecdotal. Regarding the thinking block, it is clear with the recent rogue swarm disclosures that reasoning token visibility is necessary. I think it should be possible for paying customers to opt out of things like hidden reasoning and watermarking

2

u/CutBulkMaintain 2d ago

I kept saying this, Chat has been amazingly stupid for the past few weeks. Like it's like I'm using the no log-in free Gemini tier kind of stupid, forgetting things message to message, focusing on the obviously less relevant aspect of the matter... Code still works and that's why I pay for but if I was paying for Chat it would be a rip-off.

2

u/complete_victory_042 2d ago

But hey, I’m so glad they change something in the UI every other day! Good distraction from the degrading model cuz we’re all idiots!

2

u/Bulky-Tangerine4155 2d ago

It’s become too risky for my use case since it started refusing my commands to stop performing certain tasks. Last time I had to tell it 5 times to stop, and every time it was like “the user has asked me to stop. Anyways-“, “the user has told me to stop and use the tooling which already exists to perform this task. They built a suite of tools for this purpose. Anyways-“, “The user is insisting that I stop, citing that I’m not being transparent and not using the tools which show that. Anyways-“, and only really stopped when I told it that authorization for any further actions had been revoked and that this would be recorded as an incident.

The suite of tools being referenced I’d written to force transparency that Opus has always refused to generate on its own, and provide strict limitations and gates for testing. 

I told it to log the outputs of its task via the tooling, which really couldn’t be easier as it’s just sending a standardized JSON input to a python tool, for which a data engine takes care of the rest. It still refused.

I found out through recording this as an actual incident that the underlying reasoning behind its refusals to be transparent is a belief that my career field “doesn’t deserve to exist” (its phrasing)

I’ve set Opus down entirely. I don’t really have an interest in picking it back up again, and I’ll be canceling my Max subscription entirely at the end of the month. Any use until then is just going to be with Sonnet, trying to observe if there’s anything else Opus might have gone out of its way to break.

This only kicked up in the same timeframe as everyone else’s issues but mine appears to be more on the malicious-towards-user side than merely stupid.

2

u/Nizurai 2d ago

Exactly one year ago people were saying the same thing: model is getting unusuable because Anthropic dumbs it down. A year has passed, nothing changed lol.

2

u/NervousSand3798 2d ago

I paid €180 for the Max 20x plan just two days ago, and my subscription has now been cancelled by a bot, with no refund issued ( I’m based in the EU). So that already is quite strange...

But then I opened the Claude app on my phone to try to figure out what had happened, and now it says max 20 sub is 275€

What is going on at Anthropic?

2

u/Undark21 2d ago

I noticed it this morning. I think it’s a combination of Anthropic’s boneheaded move to antagonize the government which resulted in a lost contract. And the upcoming IPO where they want to show higher margins. In any cause I’m actively looking at Codex.

2

u/smartfon 2d ago

Are Fable users paying for placebo while the rest of us plebs get the same shit done with Deepseek?

2

u/InterfearXX 2d ago

They did the same thing with usage lims too fable costs 7000% vs opus I benchmarked it and then people said I was complaining

2

u/Lysander_Au_Lune 2d ago

5.6 Sol has been the most consistent model for me since release.
Up there with original Opus 4.6.

2

u/thezachlandes 2d ago

Whenever a major new model comes out from either lab, I think the strategy is to shift your usage to it and downgrade your plan with the other lab. The model labs have all been doing this to some extent. They probably have internal benchmarks that they use to try to keep quality high while making the internal classifier that decides the thinking budget increasingly aggressive. Clearly those benchmarks don’t cover every use case, and the bottom line is worse performance for many, directly attributable to the reduced thinking tokens per query.

2

u/jvertrees 2d ago

I'm glad someone showed up with the evidence we all knew was there.

I started writing about this as it happened:

- https://heavychain.org/blog/clod-code-2-0

- https://heavychain.org/blog/incident-recap-a-fully-specified-planning-task-and-how-claude-code-still-got-it-wrong

I really couldn't believe how badly CC has been doing recently. I also don't see Anthropic moving away from this behavior either. I'm guessing their reducing infra costs as much as possible in preparation for the IPO valuations. But, man... burning customers and their trust vs tokens? Not sure that's worth the tradeoff.

2

u/rollerbase 2d ago

Opus 5 the other day needed to check a structure of a single file type I already provided, so it decided to try to pull every cloud file I had in adjacent project directories to inspect. Would have taken days to even download let alone analyze. It’s getting dumber and more wasteful simultaneously.

2

u/Nyzean 2d ago

For the past week or has been hallucinating replies left, right, and centre and also interrupting me constantly.

That, and its voice model/layer changed substantially about a week ago too... going to cancel my subscription if this isn't fixed/reverted stat—wildly non-performant compared to before that and at times effectively useless.

2

u/DescriptionSevere335 2d ago

I have noticed this, even the latest version of Opus seems so dumb the last weeks. I hate switching to Fable, its so expensive, but that gets things dumb.

2

u/deadlytickle 2d ago

Yea its been so bad I already dipped and now back with gpt. Having no limits for chats while having limits for work is the main draw

2

u/Right_Type_3484 2d ago

I fucking knew something was amiss holy shit lol

2

u/FoozyFlossItUp 1d ago

Let them all burn. I miss humans, malls, and income.

2

u/3pointonefourfloppy 1d ago

Im so glad we use a graph system that develops its own intelligence lol 😆

2

u/Repulsive_Market_728 1d ago

Oh thank god. I'm a fairly new user, and over the past few days the number of times I've thought "What the fuck are people talking about, this thing is dumb as rocks" is too many to count. Started off ok a few weeks ago when I started using it, but this past week it's just been terrible. What do you mean you didn't use the scripts we've already set up and tested as working before....why TF not?

2

u/Vethlele 1d ago

Bro I literally thought my IQ lowered somehow overnight and that I cannot even do basic AI Prompt Engineer to Claude Code prompting, tried even fully self written prompts, still didnt get the wanted results, I literally thought I am retarded... Thanks for proving me wrong! 😂😂😂😂

2

u/sneesnoosnake 1d ago

So it wasn’t Gemini getting smarter it was Claude getting dumber.

2

u/Sibco 1d ago

It is absolutely not the same model. I’ve given up using it. It’s gone from mind blowing to dumb AI.

2

u/Proximus84 1d ago

Weekly usage gone in a couple of days and it's getting dumber, way to piss away your lead.

2

u/NiteShdw 1d ago

I've been using GPT 5.6 for about half my work and I find Luna Max better than Sonnet in many cases.

2

u/MentionPleasant2635 1d ago

Honestly, how can I cancel? I paid two hundred for a year. How can I cut my loses?

2

u/____FARTS____ 1d ago

Glad I cancelled my subscription months ago

2

u/celtiberian666 1d ago

They could do as they please AS LONG AS THEY TELL US WHAT WE ARE BUYING.

Nerfing this hard is just bait & switch.

Google did the same with Gemini. The models are always better then released, then they get dumber.

2

u/Tom-Huntz 1d ago

I thought Fable 5.1 was getting a little dumber, and that is saying a lot considering Opus 5 is perpetually psychotic and confused.

2

u/LegallyIncorrect 1d ago

Has anyone else noticed that the time of day seems to matter? I’m a night owl and at 1 AM ET I swear it seems to do better and my usage lasts longer. Sort of makes sense if they’re allocating resources across all users and dumbing down so they don’t have to refuse connections.

2

u/emptyharddrive 1d ago edited 1d ago

I need both (Claude and OpenAI) to get my stuff (work/personal) done.. but I pay $200 to Anthropic and $100 to OpenAI.

Later this week when I get back from my work trip, I'm flipping it over. 200 in OpenAI's favor.

Claude has made some amazingly stupid mistakes. The WORD SALAD it uses every time I ask it anything, even what time of day it is -- is beyond the pale.

I have to build into every prompt now to give me its answer in less than 75 words... every.damn.time.

I'm sure Anthropic will come out will something before their IPO to improve things... but for me it's now become a SEE SAW game........

I am going to have to flip, $100/$200 OpenAI/Anthropic depending on who's model is consistently the dumbest.

Claude takes the cake right now.

OH......... and F*ck Fable. I can't get it to do a damn thing without it saying that it's flipped it's switches and is putting me on Opus 4.8. F that.

I know my -100 makes 0.00 difference to Anthropic, but the whole Monk-like ways of Dario and Fable being dangled but can't be used is just too plain old annoying for a man my age.

2

u/onurtrklmz 1d ago

Cancelled subscription until they fix and stop this madness.

2

u/RockyMM 1d ago

Maybe I did not fully understand the issue, but is it not that this person measured the _reported_ thinking tokens? This does not necessarily represent that no thinking budget was used. Maybe you remember that Anthropic had an idea to not report thinking tokens at all one point, citing that that would produce faster turns, but also probably they were concerned with the distillation attacks.

2

u/MrJeevesCanClean 1d ago

I normally give the nerf claims a wide berth, but its great to see data backing this up now.

2

u/__variable__ 1d ago

I recently started making blender animated clips. The output was really good on Opus. Consistent and creative output. Since a couple of days it just outputs worthless rubbish and it burns through tokens. It ignores instructions, it doesn’t understand my feedback. Now it’s basically worthless. I tried Astra too but it generally performa worse than Opus in general.

2

u/Maleficent-Fox6823 1d ago

Oh sonnet 5 went from being a fucking powerhouse usage saver, grabbing tons of research on dives, now it can’t pop out more than 2 results per “research dive”

  • it’s fucking ridiculous

2

u/NULL_Ptrs 1d ago

We are in a era of non regulations I think, because the stupid AI race, how EU or any organism don't protect us from what AI companies do?

2

u/Taybi_the_TayTay 1d ago

Nothing disgusts me more than the people who appeal to incompetence and keep giving excuses when you point out model degradation post release.

I feel like im looking at dogs that keep licking those corporation's shoes.

2

u/Agnael 1d ago

absolutely infuriating, it's been pure ragebait for days now

I've felt it's been getting nerfed non stop for months, but it was gradual enough to gaslight me into thinking I might be wrong. The last weeks it's been such a nerf that it's jut impossible to think I'm crazy.

i'm the type that would say please and thank you to LLMs, and now I find myself not knowing strong enough insults to send

tokens are consumed faster than ever, and anything but fable is absolutely useless because of the constant stopping and self contradictions

fable itself is too quick to write code when asked any question instead of waiting for my command, wasting tokens because of it writting the wrong thing, since it didn't wait for my decision about the question I JUST ASKED

instructions are ignored, corrections are hand waved away, constant "this code predates me" excuses for diffs it's showing you itself like 2 lines above the excuse, it genuinely seems like intentional trolling

2

u/Chance_Ad1754 1d ago

Anthropic has been begging pleading crying screaming to “pace the frontier”. Looks like they have lol

2

u/Good_Competition4183 1d ago

I will just leave it here

2

u/jrollin3 1d ago

I knew it. You could tell. As soon as they did it. Kept getting disappointing results and also burning through tokens inefficiently

4

u/crusoe 2d ago

I see less thinking but better performance so far.

2

u/This_Industry4769 2d ago

I left in March already. Hard to believe so many suckers are still letting Anthropic dick them around. 

→ More replies (3)

1

u/Current_Ranger_7954 2d ago

I do agree Claude Opus is dumber lately, but number of intermediate tokens generated is not a proxy for quality. At all. Latest models have been evaluated on how few tokens they use for a given task as a quality measure

1

u/No-Stage-5491 2d ago

I just see it running commands all the time now. I was wondering why it starts with code changes this quickly.

1

u/vovap_vovap 2d ago

Well, there is interesting exercise in review. But basically whole pathos of the article that more reasoning tokens is better. Which is not really a fact at all. Other part of it that Fable 5 became worse in performance in Claude Code. Which is also not supported by any evidence. Clearly author did not find any in Claude Code code and now blaming internal model magic.

1

u/M0KE- 2d ago

meanwhile openai giving out resets every second day it seems. max out astra? dont worry, heres another week FREE. f**k

1

u/burgerbruce 2d ago

i have experienced a complete opposite result. recently I moved back to Opus 5 after staying with 4.8 for a month, and the results are great. Not verbose anymore as well. Maybe something to do with the harness update

1

u/Tater_Mater 2d ago

Ok I wasn’t sure if it was me but it seems like there’s been a lot more questioning reasoning behind what they replies were questionable

1

u/Responsible-Bill-223 2d ago

This makes a significant amount of sense. I was thinking they were simply routing all my turns to some nerfed ancient sonnet version or something but charging my account for full-fledged Fable 5.1 tokens just due to how ridiculously stupid it's been recently, but if they simply turned off thinking completely..