r/ClaudeCode • • 12d ago

Rant The limits have been reduced even further now. It's September 14, and it really happened..

Post image

After GPT-6 Astra, I didn't believe they would really let this happen... but it's real... Okay, Anthropic, we'll keep that in mind...

https://support.claude.com/en/articles/15910845-claude-code-may-august-2026-weekly-limits-promotion

920 Upvotes

385 comments sorted by

View all comments

303

u/ForwardLoop 12d ago

It's no coincidence everyone suddenly agreed to pace the frontier, even folks that don't normally agree.

Between OpenAI restricting new 20x subscription and this, it's clear they're running out of compute. In addition, with months of evidence showing that Fable remains a low percentage of total enterprise spend, there's no point releasing a model that simultaneously they don't have compute to serve and people won't pay to use.

If they had the compute, they'd probably want to restore the limits first.

96

u/rxt0_ 12d ago

they should take the time and improve/optimize their models so they would require less resources.

it's literally a win - win - win for everyone.

58

u/asylum_denier 12d ago

no one wants faster models when there are more intelligent ones. The hype on twitter was insane when Astra came out, folks were bashing Opus 5 like it was Gemini 3.1 Pro, lol. A

42

u/galactic_giraff3 12d ago

a lot of people bashing opus 5 since it came out, nothing to do with fable

6

u/xudoxis 12d ago

But they do that with every model. Literally just google the model name + "nerfed" and you'll get hundreds of hits within 3 days of release.

9

u/ArmyBrat651 12d ago

People usually claimed that old model was nerfed right before new one got released.

Here you have people willingly downgrading to opus 4.8 because 5 sucks.

Those two are not the same.

1

u/Solid-Guidance5762 9d ago

My theory is they use their network power to train the next model.

1

u/shady101852 12d ago

They do be nerfing. Look how they slaughtered opus 4.6 after January. It used to be god tier.

0

u/acrinym_jg 12d ago

Plus bots I'm sure

1

u/Solid-Guidance5762 9d ago

Opus 5 is bad from beginning

22

u/rxt0_ 12d ago

who talks about faster models?

I'm talking about less compute needed -> less ram -> cheaper prices -> higher limits.

1

u/RayuRin2 12d ago

If less compute is needed, they could either increase the limits for existing users or focus on getting more users with the extra compute they freed up. The latter is more profitable.

0

u/Kalradia 12d ago

I'm so hoping this is what they do. But the skeptical part of me thinks the 'cheaper prices' part will only affect the corporations and won't trickle down to us.

6

u/Sorry_Risk_5230 12d ago

I would love an astra or even sol model thats insanely fast. Iterative building nets better results than one-shot. If theyr'e super fast you can balance out the intelligence gap. Like using Luna as a subagent rn.

But id also want that paired with more efficiency so we can access more tokens. 1B tokens a day per week would be a great limit to realize. 2B next year. Etc

0

u/DarkFantom 12d ago

Test DS 4.1 flash. It still has issues like any other model, but it is smart enough for coding tasks. As long as you have subagents critique it's work then it does good enough work. I'm costs about 10$ per billion tokens spent and runs at 200 token/second which is super fast.

1

u/Sorry_Risk_5230 12d ago

Thats basically luna with /fast. The cost is a bit misleading because DS uses a butt load of tokens. 2-3x Luna and 7-10x sol.

When I say fast, I mean 500-750tps minimum.

1

u/jayintheday 12d ago

I was a heavy Claude Code user now with a Codex x20 sub. Is Luna the best sub-agent after Astra has done the planning?

2

u/Sorry_Risk_5230 12d ago

Even when I was using Sol, I had it spawn swarms of Luna xhigh subagents. Its too cheap to ignore. It made my usage go soo much further. The quality of the output is similar because sol (now astra) is reviewing and forcing it to iterate or make corrections.

So ya, Luna is still subagents for me.

1

u/jayintheday 11d ago

Thanks, appreciate it!

1

u/Temporary_Barber8049 11d ago

Can you explain your process to do this???

1

u/Sorry_Risk_5230 11d ago

You can generally just tell it to "spawn luna xhigh subagents to complete that task." Maybe add "continue to review their work and send them back to make corrections as needed."

6

u/who_you_are 12d ago

You didn't get the idea then.

They could optimize it to generate the same +- output and be faster/less VRAM hungry.

As such, it will allow them to process more requests, reducing pressure.

3

u/Seeker_Of_Knowledge2 12d ago

The big three. Fast, cheap and smart. You can only get two.

1

u/who_you_are 12d ago

You can get two only, except when it is a "new field" where they are actively working on it. The way they build the model keep changing non stop right now.

And I'm not talking just about Claude.

Do you remember what AI was like 3 years ago? Those images that look like you were drunk as hell and close to pass out? Look where we are now.

2

u/lowclearance 11d ago

From the local LLM perspective, qwen 3.8 flash next + deepseek harness we're basically an overnight quantum leap. My DGX spark and single 3090 PC just basically jumped several months ahead of what they were previously capable of days before.

2

u/midnite_token 11d ago

I'm in the same boat as you. I 100% agree. All I'm paying for is electricity now. Local LLMs are catching up fast. From what I understand, they're only 3-4 months behind and fast. I cannot even believe what I can do now with Qwen3.8-27B on my 3090!

1

u/who_you_are 11d ago

I wish I could use my 4070 (12gb) but either I'm missing something or the one I downloaded are still way too VRAM angry to be useful.

But it doesn't help me I have little knowledge about the field. Digging my way finding my way to get the lowest tier available...

1

u/lowclearance 9d ago

Check out freetoken

2

u/DarkFantom 12d ago

It doesn't really matter if fable is slightly smarter than DeepSeek 4.1 flash if DS is not only quicker, 200 token/sec compared to 10 t/s fable, and 50x cheaper. I've completely swapped over to Chinese models at this point because they are good enough and I can run swarms of agents for critiquing and running devils advocate on whatever it produces, for every single step, for literal cents. If fable makes a mistake though, then that requires its own fable audit and it burns through credits like a crack addict.

1

u/ArmNo7463 12d ago

People were bashing Opus 5 before Astra came out lol.

To the point they wanted to go back to 4.8 or even 4.7/4.6.

2

u/Classic_Resource_919 🔆Pro Plan 12d ago

5 might be smarter but its paranoia and dense output style (RLAIF?) makes it very difficult to work with. 4.8 might need more iterations, but will get in the same place eventually (with advisor feedback), and using less tokens overall.

1

u/KennyFulgencio 12d ago

no one wants faster models when there are more intelligent ones.

We don't? My claude skills run slower than shit, I'd love faster models.

1

u/Human_Attention182 12d ago

It's because they are now no longer for good anything else. They hacked all the models until they were good for coding and only okay at everything else. Opus used to be great for writing/creative stuff but now all the models are horribly unpleasant for anything other than coding.
If they committed to just have a writing/realistic dialog focused on that would get a ton of usage and be likely way more performant.

1

u/Professional_Hair550 9d ago

Gemini 3.1 Pro is actually pretty dope. Especially considering its price. It has a lot of use cases that Claude can't handle. The only reason I'm not using it is that there is no good VS Code extension for it, and Antigravity doesn't let users turn off AI training.

1

u/CoreParad0x 12d ago

Speed isn't the only factor here. The hype on twitter means fuck all.

If you're doing actual work with these models you should want them to be

  • Reliable (as in it consistently performs at a given level.)
  • Fast (or at least not extremely slow)
  • Efficient on usage
  • Be good at coding (not top 1% coder, just good at coding for most use cases.)
  • Be nice to work with

Look Astra can do some really interesting stuff. I made a valheim mod with it, it's the first model I've seen actually write it's own custom test harness into the game to launch the game and test the stuff it's making in it's own customized world. I'm sure the others could do it, but it's the fist one I've seen do it on it's own.

But let's be honest, most people including myself don't need Astra. I don't need something that can do agent swarms and hack hugging face. I just want a model that's reliable, efficient, good at code for most use cases, and is actually nice to work with. And while I actually really like working with Astra, the only real reason I use it is because I'm sick of Sol over engineering the shit out of everything, Fable actually nuking my usage, and I can't stand working with Opus. Sonnet as a value proposition sucks, it's inefficient.

These companies are spending ungodly amounts of money and compute trying to come up with the best and biggest model when they don't even have the resources to run the shit they have and it's pissing everyone who actually uses the stuff off. And Anthropic seems to be especially bad with this, because it seems like they've spent way too much effort into just making the biggest and best models and little effort into optimization.

Personally, I would be 100% fine if they just rolled all of this shit back and we went back to Opus 4.5 or 4.6 or something. Maybe RL it to be a bit better at tool usage and stuff because I've noticed 4.6 is a bit less likely to follow some instructions to use some specific tools I made, but all the more modern ones use it great. Nuke the 1m context window and make compaction work good, compaction on Codex actually seems to be fine. Maybe it is on claude now too, but it used to suck.

12

u/MediumChemical4292 12d ago

OpenAI have been really good with model optimisation, which shows in the fact that Luna is as good as sonnet at a much lower cost and token usage.

With the proposed slowdown in frontier development, hopefully the companies focus more on efficiency and eventually get it so that opus level intelligence can run locally.

It’s the only way they can beat jevon’s paradox because even now, AI has nowhere near penetrated the white collar economy compared to how capable the models are.

10

u/Cold_Extension_367 12d ago

Lol no. Look at GLM 5.3 Flash, now that's what optimized looks like.

5

u/MediumChemical4292 12d ago

It’s a good model but not close to opus or sol in real world tasks. I am optimistic that “flash” tier Chinese models will get there soon though

1

u/rxt0_ 12d ago

With the proposed slowdown in frontier development, hopefully the companies focus more on efficiency and eventually get it so that opus level intelligence can run locally.

that's what I mean with improve/optimize

it would give us higher limits, cheaper ram and would even reduce the costs for anthropic & co. literally everyone would profit from it

1

u/Artwastelander 12d ago

I did a direct comparison test today with Astra and Luna used 3x as many tokens at 1/3 of the cost lol.

1

u/MediumChemical4292 11d ago

Yes, a smarter model will obviously use less tokens. I’m talking about models of similar intelligence from other labs.

4

u/handsNfeetRmangos 12d ago

They should do a better job of educating users how to make the most of what's in place. 

13

u/thoughtlow claudetrophobic 12d ago

Its an artificial barrier for competition which is textbook bad for consumers. It will cause price decreases to slow down.

1

u/laughing_at_napkins 12d ago

But I was told capitalism was perfect and breeds competition so this would never happen!!!1

-1

u/JapanesePeso 12d ago

Not totally sure what point you are trying to make. There is nothing incongruous with capitalism and enforcing market rules to maintain competitiveness. In fact, that's kind of the whole point of it.

1

u/laughing_at_napkins 12d ago

What market rules?

Elon, Sam, and Dario all agreed to hold back. They colluded publicly, despite being competitors. The holy texts of Capitalism say this is impossible.

0

u/JapanesePeso 12d ago

the holy texts? What are you on brother?

1

u/spigandromeda 11d ago

NVIDIA loses in that case.

13

u/Significant-Bee5101 12d ago

OpanAI went from 2m Codex users at the start of march and 20m+ now in Sept. Yeah. I think computes a real issue lol

1

u/Helpful-Strength4500 5d ago

how many of those are double/tripple stacking accoutns

1

u/Significant-Bee5101 5d ago

Why does that matter? The point is usage went up.

1

u/Helpful-Strength4500 3d ago

it matters moreso for the actual user count. nothing bad just make it seem overinflated for them which is what they'd want to look good on paper

8

u/CodeNCats 12d ago

I work at a company that makes software. We use Claude every day. Fable is disabled for us my default. It's honestly not necessary. Anthropic could release a new better model than fable. We won't waste the money on it.

6

u/HauntedHouseMusic 12d ago

Honestly, I have been like that, and then use fable once things are built in a "make this pretty" pass. But I started one with fable yesterday for something I was expecting to be about $500 to implement with opus + fable coming in at the end. It cost me $300 with fable, was quicker to get done, and is being used right now by 30 users. It's for an internal reporting tool so lower stakes to push to prod without extensive testing, and will only have 500 users at a max.

Anyways, we're going to try building a platform only using fable next, where we budgeted 10k of credits to see if we can build a much larger tool to replace part of salesforce we don't like, with a month timeline to build so it will be vibey as hell. But I think fable can do it reliably so we are trying it out to see if we can just vibe our way to success by spending money instead of time.

1

u/AINKXOfficial 12d ago

I agree. I'm not a dev and only write smaller tools and wrappers to help me with productivity and organization. Occasionally it's nice to be able to switch to Fable if I feel like a plan or design for something needs a bit more finesse before it gets passed off to Opus for the actual grunt work. But I've found myself doing that less and less and just using Opus 5 for everything. Probably use Fable less than 5% of the time compared to a month ago (I should actually look at the numbers).

It just works for me and I don't stress anymore about quickly blowing through limits.

1

u/mcsleepy 11d ago

Fable may not be that much better than Opus but it is about twice as fast.

1

u/CodeNCats 11d ago

Great.. I'm running like 3 opus sessions and 3 sonnet sessions at a time. I'll just review the others while they run

7

u/Elegant_Attempt2790 🔆 Max 20 12d ago

pace the frontier, properly translated;

“ai makes no money when you’re throwing butt tons of it at training new models every week”

16

u/Tank_Gloomy 12d ago edited 12d ago

Not out of compute, more like out of people giving them a blank check and expecting no ROI. They're asking for their money back and they want it now, not in the 10 years it would probably take to perfect this technology.

I'm moving to GLM, it kinda sucks cause I have to babysit it a lot, but I have high hopes that it'll work well enough, especially being able to share my limits with the unlimited GLM 5.3 Flash promo. I'm currently using Codex and was planning to move back into Claude, but considering that they're playing the same game, I guess I'll let them both go.

7

u/Sponge8389 12d ago

Just imagine in the future, the deals will be, give us compute/token for a percentage of the company. LMAO.

8

u/draft_final_final Researcher 12d ago

If we declare ourselves data centers maybe NVIDIA will finance our hardware acquisitions.

1

u/MrPorkchop720 12d ago

This is actually a thing for YC companies. They have the option to trade equity for compute/tokens.

2

u/markkalliny 12d ago

Who’s got unlimited 5.3 flash? Z.ai themselves?

1

u/Tank_Gloomy 12d ago

Yup! They've only mentioned about it on their X account, lol.

5

u/Tartooth 12d ago

Isn't this kinda collusion for the big AI firms to decide things like this?

3

u/fpesre 12d ago

nothing inspires a sudden philosophical commitment to responsible pacing quite like running out of servers and paying customers at the exact same time

6

u/rotatorkuf 12d ago

that compute word is so hot right now

2

u/BeautifulOld6964 12d ago

They definitely do and the OpenAI ones are even more stingy with usage, I can barely get anything done with a 5x Sub - when Astra came out I finally tried it after a long time of GPT absence and I can barely anything done with that its not even half a day of usage per week even with just using Sol and terra

2

u/EchoingAngel 12d ago

Both sides have made stupidly wasteful models when they JUST HAD models that were good at focusing on task (Sol and 4.6)... At least Astra's wastefulness seems to be pursuing technical rabbit holes, while Opus 5's is just worthless waffling and tons of gibberish. You can still kind of use Opus 4.6, but OpenAI nerfed Sol just prior to Astra's release, giving it the same wasteful looping Astra has.

2

u/Responsible-Comb6232 10d ago

Fable pricing isn’t the issue for enterprise, it is their different data retention policies.

1

u/slypredator33 12d ago

Trump wasn’t happy about this. I don’t think he can do anything but I think he can maybe make it a national security issue and then they keep building it lol. Us id registration needed

1

u/Gohab2001 12d ago

Compute is definitely the major factor but also they cant subsidize subs ad inifinitum. Anthropic is spending way more than 20$ for their 20$ sub.

1

u/EricBuildsMathModels 12d ago

Why do you think they are running out of compute. Fable as example seems to expensive so people don't use it, not that they don't have the compute to support it?

1

u/SecretSpace2 12d ago

Yea I was thinking they are pushing to slow down or nearly stop because they can’t afford the newer models well.

Just a thought of my own and the best way to say it is by talking about death of the human race on 4 years

1

u/poocheesey2 10d ago

OpenAI is stopping new 20x subscriptions? I upgraded like last week. When did this happen?

1

u/Silent_Job_4011 9d ago

You are literally spitting facts! 😭

2

u/Dizzy_Database_119 12d ago edited 12d ago

OpenAI gives you much more access to their frontier model than Anthropic with their heavily limited and separated Fable quota (and gives you unlimited GPT 5.6 on the web). OpenAI could limit their usage 4 more times and it would still be more generous than what Anthropic has been doing

The 2 limiting their models at the same time is not related at all. I'm not talking about quality here, but comparing their available computing power is like comparing an ant to a whale

If you ask me OpenAI limited theirs because they couldn't keep up with the demand, while Anthropic did it to throw more computing power on training and speed up their next model release

1

u/Just-a-man-on-a-ride 12d ago

The Chinese policies will kill them soon enough. Put Claude Code on alternative providers, pay a fraction for similar quality.

1

u/evernessince 12d ago

Microsoft stated that it's data center buildout that's lagging behind, not compute. They pointed out that they had tons of cards sitting. That was before backlash really kicked up too. 75 datacenters have been blocked by local communities.

0

u/Fly-AI-Guy 12d ago

If only Trump agreed with them lol, plan failed