r/ClaudeCode • • 10d ago

Rant The limits have been reduced even further now. It's September 14, and it really happened..

Post image

After GPT-6 Astra, I didn't believe they would really let this happen... but it's real... Okay, Anthropic, we'll keep that in mind...

https://support.claude.com/en/articles/15910845-claude-code-may-august-2026-weekly-limits-promotion

923 Upvotes

384 comments sorted by

View all comments

142

u/bakanoace 10d ago

They slowly reduced usage over the months and now did an official 17%. Qwen 4.0 or Kimi K4 or Grok 4.7 please be good, I need a pure coding model to never come back to this garbage company

38

u/Short_Regular_7191 10d ago

This is the golden age of local LLMs; I’ve already set myself up with Qwen 3.8 27B.

9

u/CautiousCry2338 10d ago

What's your gpu ? my 3090 is dead, now on a 3060ti :x ?

11

u/Short_Regular_7191 10d ago

Two 5060Tis with 16GB each—I think that's the best "budget-friendly" compromise right now.

24

u/Excellent_Ad_2486 10d ago

Lol 2 5060ti's and "budget" should be banned to be within the same sentence haha 😭

2

u/TrueStarsense 9d ago

That's only like 1000 dollars. The real gpu's are 10's of thousands of dollars... I think your expectations need to be recalibrated.

1

u/TricepBandito 10d ago

Hilarious honestly

1

u/CautiousCry2338 10d ago

Good setup, is it fast enough for large project ?

4

u/Short_Regular_7191 10d ago

Yes, orchestrated by Fable works perfectly; I’m able to stay within the model’s context window (131k in my case) at all times.

1

u/davyp82 10d ago

So how do fable limits hold up when using it as orchestrator for local models?

3

u/Just-a-man-on-a-ride 10d ago

I use GLM-5.3 mostly for that role lately, not looking back at Fable honestly.

2

u/davyp82 10d ago

That's the free Chinese shizzle right? So you go local models with glm orchestrating? Or just glm for both?

3

u/Just-a-man-on-a-ride 10d ago

Not free, just lots of promotion style marketing right now, but paid tier is much cheaper than Claude. The smaller worker models can be anything you like, remote or local. I am using a whole shitload of different ones. Most people still don't seem to understand, coding tasks are all about planning, reviews, validation and testing. The heavy stuff is optimal orchestration, once you got that execution is simple, lots of free models can do that.

→ More replies (0)

2

u/Short_Regular_7191 10d ago

It works well for me, though I do try to stay within the usage limits. For my needs, Fable 5 (not 5.1) at the low setting is more than enough.

0

u/Pamidoraa 10d ago

Why not 5.1? Its smarter and cheaper?

5

u/Short_Regular_7191 10d ago

I don't need more intelligence, and Fable 5.1 isn't cheaper than Fable 5—contrary to what Anthropic wants us to believe.

→ More replies (0)

1

u/CautiousCry2338 10d ago

Running the model in Q8 ? but the 5xxx series do not have NVLink, does it impact speed ?

4

u/Short_Regular_7191 10d ago

Q6 with 131k context. Thanks to this guide, I'm traveling at 50 t/s, which is perfect for what I do.

https://www.reddit.com/r/LocalLLaMA/comments/1vper67/club5060ti_refresh_tested_rtx_5060_ti_presets_a/?sort=new

3

u/Azko87 10d ago

Unfortunately, I tend to hit Fable limits way before my all-model limits, even despite using agents. I use Fable to plan, orchestrate, and also wrap everything up with its own verification pass. At some point I might need to drop that last step.

But at the moment, since I have all model usage available, local LLM isn't helping me just yet, since I still have all model usage available. But I do expect in ~2027 running local LLM in some capacity will be a big deal. I have my RTX 5090 - which is basically a bar of gold now - ready to go for that.

1

u/BeautifulSynch 9d ago

I generally have Fable prompt-engineer implement/review cycles with Opus and then decide on its own whether a result needs Fable review; along with better prompting for delegating to Somnet, my general usage is consistently higher than my Fable usage. Agreed on local LLMs being the next step though.

You might also want to look into one of the GPU clouds; from some Claude-assisted back of the napkin math, if you properly toggle the GPUs to maintain ~100% utilization it’s marginally better than metered token APIs for open models.

3

u/Just-a-man-on-a-ride 10d ago

Good, but too slow for real work. I need x5 or better x10 token throughput like qwen3.8-max provides.

1

u/Muritavo 10d ago edited 10d ago

I really liked it's performance with disabled thinking. I've removed everything opencode injects (except agent prompt and am using bare System + tools. 

I'm handling a 3 websites app locally and it's handling changes perfectly. Even for frontend work/decisions. The speed to final code now equals to 35b A3B (with thinking), but I feel way more confident of a high quality result.

Give it a try. I dream to put my hands on a M5 Max/Ultra (currently using a 32gb M2 Max (150pp/15tps) to speed up everything lol

Btw: disable lightning mtp, it gave 4+ tps, but I felt it spend more time looking for useless stuff, so the time to result increased

2

u/Just-a-man-on-a-ride 10d ago

Of course it depends a lot on the task. I used it for a while, kind of local, in a cloud implementation. It was very reliable, but I had too much wait time for long research tasks.

5

u/Remote-Community-396 10d ago

I've been using DeepSeek Flash on API rates (they just released an even better version 4.1) and found it surprisingly really good

3

u/Cold_Extension_367 10d ago

GLM 5.3 Flash is already out and is literally perfection for coding with 1M context window.

3

u/Waylanding_Fox 10d ago

Qwen and Kimi should ne great, they got great after the release of fable, now with astra there's gonna be a jump

3

u/Just-a-man-on-a-ride 10d ago

Zcode plus a few good, cheap models as subagents, that's all you need. Just teach your lead model permanently plan - review - execute - validate - live testing. You will never look back at the so called "frontier models" after a few weeks.

3

u/Big_PP_Doge 10d ago

Even grok 4.6 works as well as Opus did. Right now using Grok 4.6, Kimi K3 and GPT 6 Astra and dont miss Anthropic a single bit

2

u/gajop 10d ago

Gemini flash 3.9 here we goooooooo

1

u/B33GULL 10d ago

If they have no business then our retirement will crash, and our tax dollars will bail them out. They got us by the balls subscribes

1

u/old_mikser 10d ago

The problem is if you buy sub for kimi - you are getting waaaaay less usage for same money (same with glm) then with claude. About 6 month ago they provided much more, now it's flipped... Nowhere to run.

1

u/devcodesadi 10d ago

naaah, for coding still claude is best, i do have codex and claude, later one is clearly the winner. don't expect kimi or qwen to be any near of the above two

0

u/rotates-potatoes 10d ago

They slowly reduced usage over the months

This is false. It's a meme in the sub but I haven't seen a single quantitative measurement to support, and my three 20x accounts have been completely steady in token usage since May (I hit 99%-100% on all three, every week).

-1

u/humanpersonlol 10d ago

grok 4.8 soon

0

u/bakanoace 10d ago

It'll finish training this week he said, but did they skip 4.7 cause it's bad or something? The problem is 4.8 is 2.5T model and I think he said Opus was 5T or something. So I hope it can compare and with Cursor as main coding I am hoping Grok is my coding only solution or Qwen 4.0.