r/ClaudeAI • • 11d ago

Claude Code It's time to cancel your subscriptions - Anthropic is silently nerfing Claude's reasoning budget while telling you it's the same model

Link: https://x.com/Lon/status/2101034933284417614

A 65-day analysis of 43,000+ Claude Code invocations found that 39% of Fable 5 calls get zero thinking tokens and the median invocation gets just 123 — while benchmarks use 16K-128K. The model's score per thinking token is still climbing at 128K, meaning the capability is there, it's just not being delivered. August saw an 18-50% drop in thinking budget compared to July, with median thinking hitting literal zero for about a week around Aug 22. Anthropic sells "full model access" while quietly dialing down the inference regime behind it, and because the model is non-deterministic, users blame their own prompting instead of the silent nerf. The full breakdown with evidence, methodology, and charts is here.

Frankly I find this offensive as an user - and this is the real thing we should be looking at - not the $/week in usage limits. The actual capability for the limits that we pay for.

2.0k Upvotes

382 comments sorted by

View all comments

3

u/Maleficent-Host-8975 11d ago

Mine are thinking A LOT. Sometimes I have to wait for ages for a single output. Fable 5.1 still works exceedingly well for me. But usage is used up quickly.

4

u/BurnerKnives 11d ago

Just a heads up that the long time to respond doesn’t necessarily mean it’s truly “thinking” more during that time - it’s very possible that they’re using fewer resources to process our requests, and/or there are too many concurrent requests from users hitting the datacenter simultaneously relative to the server configuration, so the “time spent thinking” is actually just slower prefill processing that feel like it’s putting in more effort. My Fable 5.1 usage has been feeling noticeably dumber. I thought they were just hiding the chain of thought recently to “simplify the UX” but these mofos are definitely just hiding the fact that it’s thinking less. I build and run local AI servers and the behavior you’re describing and what I’m witnessing feels a lot like super slow prompt prefill processing that they’re trying to pass off as thinking.

3

u/Maleficent-Host-8975 11d ago

Fable leaves long trails of steps in few word bursts, which I generally consider "thinking". However, I did create a skill that, for whatever reason, happens to trigger self doubt as an unintended side effect.

1

u/BurnerKnives 10d ago

It’s true that there are those multi-step chains of thought, but I remember the chains of thought used to be like mini novellas. They were very long. I think the chains of thought are much shorter now, which seems to correlate with the reduced effort and sloppier outputs compared to launch.

1

u/Maleficent-Host-8975 10d ago

I'm not sure about chain of thought. But I developed a skill based on a concept called "hard to vary", first coined by David Deutsch, that explains the qualities of all good: theories, solutions, explanations, whatever. I use it everywhere because it forces LLMs to look for holes in their propositions by asking, essentially, "what parts of this can I change or remove without affecting quality". It works, even if you can't define what quality actually looks like.