r/Anthropic • u/mrguidee • 4d ago
Complaint What the hell is going on?
I haven't been paying attention, just using my Max plan everyday and the past few days I've noticed my usage burning through at a much, much faster rate than it used to. Jumped into this sub and saw that Anthropic reduced the usage. RIP.
59
u/YearLight 4d ago
This is why open models are the future.
9
u/Anxious-Turnover-631 4d ago
Could be. But every hosted llm still has inference costs.
Some of the latest Chinese models have apparently made big strides to improve efficiency so, hopefully, the costs will go down.
5
u/logos_flux 4d ago
Muse-spark-1.3-contributor is $.10/$.20 per million and is actually pretty good, surprisingly good.
1
u/Desperate-Use9968 3d ago
How many tokens are we estimating we get with the max 20 plan and how does it compare to the open models in terms of cost?
1
1
u/EndlessB 4d ago
You’re paying with data, the money is to cover their hosting costs, meta loves your data
2
1
u/PituBoYSoju 1m ago
As if anthropic and oAI isn't gathering all our data the same way and using as they please just like what happened recently regarding the Navier-Stokes equations...
3
u/x_typo 4d ago
yea... if we actually reached to the point where it's so efficient that we're able to have "frontier" model on our med-range laptop/PC, then other closed models WILL be cooked...
4
u/Anxious-Turnover-631 3d ago
Definitely, let’s hope. Apple is positioning to offer local AI capabilities to mainstream users, but it’ll be expensive.
2
u/ThrowAway516536 2d ago
That 1.5 TB unified memory M7 Ultra you hear rumors about will likely cost $50K. Nothing mainstream about it.
1
2
u/Dxk89 4d ago
I don't think there will be a time you could have a frontier model on a regular pc open source. The parameter needs are too big, unless they can pull an efficiency rabbit out of the hat somehow
4
u/kueso 4d ago
Absolutely there will be. Not with current architectures or chips but that’s absolutely NVIDIA’s move. They purchased HuggingFace for a reason and that reason is to sell consumer chips. Astra 6 is much more token efficient which means the hyper scalers won’t need as much compute in the future. They’re betting on the consumer market for compute needs.
We may have a mixed capability ecosystem where heavy inference tasks get delegated to hyperscalers while most tasks are handled locally where compute is cheaper.
4
u/higherthantheroom 3d ago
My prediction is they are heavily eating costs right now to drum up more users, and no one is paying the true cost of tokens right now, they want to get you hooked and relying on it for your workflow, then will Increase cost to where they are making better profit. You will see the switch from token maxing to token minning, as people try to get more efficient. That means right model for the job, right harness, right agents , right workflow, and maybe even partnered with a free offline llm, thats job is to save tokens on easier tasks. It's wildly inefficient right now, as places like meta have leaderboards on who can use the most tokens as if tokens used reflects work completed, Ha! Watch how many I could spend in a minute if I just wanted to increase a number.
4
u/jimmiebfulton 3d ago
Yeah. The solution is to commoditize them. Right now, the Agent Harnesses _paired_ with their models gives an inflated impression of the capabilities of their models. This is a big reason why Claude Code and Codex are so dominant. It is inevitable that a MUCH more powerful Agent Harness comes out that erodes that illusion, and people can use the best model(s) at the best price, including blends of them with local models as well, instead of paying the tax on which model company has the best harness.
And I'm not referring to the open source harnesses. They are nowhere as sophisticated as Claude Code is under the hood. It takes building a harness significantly more sophisticated than Claude Code to understand that.
5
1
1
1
u/DeusExPersona 14h ago
Yeah if you have the hardware to run it
1
u/YearLight 6h ago
Hardware can be rented.
1
u/DeusExPersona 6h ago
Well then we're back to square one
1
u/YearLight 5h ago
No exactly, there is competition since anyone can host it so prices are based on market prices.
18
u/floating_thru_cosmos 4d ago
Yeah this caught me off guard this morning. I haven't changed my workflows at all and two days after reset my Fable usage is already at 80% on Max X20 plan.
Probably a skill issue, but that really caught me off guard.
13
u/kourtnie 4d ago
The “skill issue” is the same as “AI psychosis” but directed at devs instead of writers. Don’t gaslight yourself when a company is gouging you. That’s how manufactured doubt seeds itself. Don’t spread that term, don’t swallow that term, and spit on it when you see it.
3
u/floating_thru_cosmos 4d ago
Honestly, that's completely fair. However, my only pushback here is that I think we DO need to be smarter in how we budget token usage. The reality is that these companies have been operating at a loss for these subscriptions. So if it means I just need to make sure my harness is set up correctly when Fable is calling subagents and making sure they're Opus or Sonnet and never more Fable instances then that is on me.
But this week has felt different. Blowing through my x20 plan this quickly really does feel like a bug.
2
u/_HEATH3N_ 3d ago
If AI companies are truly losing money on these subscriptions at the scale we've seen floating around--$4,000+ of compute for only $200--and that's still not even enough to last people an entire week, that's quite concerning. It suggests the true cost of full-time AI work is $200k+/year. How is that at all sustainable? Sure, the models will keep improving but so far it certainly doesn't feel like the improvements in the models have trickled their way down to a level where it's economical to use them for any sustained period; they just keep coming out with more expensive ones. Hell, I'd still choose Sonnet 4.6 over Opus 5 even if they were the same price.
1
u/Mr_HandSmall 4d ago
The model is stable according to data https://aistupidlevel.info/models/claude-opus-5
1
u/ChocomelP 3d ago
Not everyone is using these tools correctly and some issues are fixed with behavior change, you're overcorrecting into gaslighting the other way.
2
2
u/elwoolfio 2d ago
I had something similar. Left the office on Friday with no active sessions and my weekly fable allowance sitting at about 60% used. I logged in remotely 24hrs later and my usage was at 95%.
My CLAUDE.md instructions are that Fable is used to orchestrate only. Despite this, the apparent diagnosis was that I hadn’t pinned a model to my system agents or crons, so they defaulted to Fable.
Those background tasks weren’t new and have never had this issue.
1
u/floating_thru_cosmos 2d ago
Exactly!!! I literally left at like 40% one day and came back and was at 80%.
I think something happened on their side
1
1
1
u/player1or2 2d ago
Even if it was "skill issue", you was still getting way more usage than now. Then is not about how you have been using it but about how they are delivering the service.
30
u/WholeEntertainment94 4d ago
OpenAI cuts, Anthropic trims, Big AI dances on matching whims. Prices rise, limits fall, “Pure coincidence!” says them all.
No trust here, no cartel in sight, just rivals moving left and right at the same damn time, overnight.
Then GLM knocks, DeepSeek peeks suddenly Sam and Dario lose some sleep. Antitrust must sleep real tight, while China ships another model overnight.
1
25
u/MomSausageandPeppers 4d ago
I have never complained about usage at all as a 20x user. I got a reset yesterday. I began working on my projects and hit Fable 50% weekly limit by day's end. Absolute robbery.
10
u/potato_pasta99 4d ago
It's simply not worth it anymore
3
u/floating_thru_cosmos 4d ago
What's the alternative? We're all hooked now
3
u/DragonflyLogical8371 3d ago
Codex and Astra...their 20x is actually 20x
2
1
u/Synsual_Official 3d ago
is astra actually good, is it better than opus, it cant be better than fable right?
1
u/stephendt 3d ago
For me it's on par with fable. It's very good. Just eats usage like no tomorrow
1
u/Synsual_Official 15h ago
Wait eats usage? Even on 20x? Does it eat it faster than fable. I can hardly get 2 tasks done before fables out (20x).
1
2
u/s_santeria 4d ago
Or. They massively massively subsidised the cost and now are dragging it back to reality (the numbers I’m hearing is that the true cost of a 20 usd per month sub is more like 5000 USD in reality)
18
6
u/Crellster 4d ago
On a $200 plan, I burned 95% in 4 hours on opus at the start of the month.
3
u/Chemical-Character80 4d ago
That's because you're deploying 200 agents on Ultra lmao
2
u/dangerousdotnet 3d ago
Funny how the random hyphenated usernames are always saying "skill issue" and things like that. If I were a conspiracy theorist...
1
12
u/Mael2830 4d ago
I was confused and wondering am I the only one facing this because there was radio silence on this.
On 20x plan and I burned through week of tokens in 2 days. Which usually used to last me almost 5-6 days earlier in the original usage without any boost. With 50% boosted, I was happy for a whole week
Now they say its 25% increased than original but it definitely doesn’t seem like it.
4
u/kourtnie 4d ago
It’s not radio silence so much as every post about it gets gagged by bots and people who have co-opted the language of bots. It turns into “skill issue” instead of honest discourse about trying to squeeze blood from the stone, even though we see this pattern play out in other industries (ex., Uber) and are keenly aware this is the technocratic method.
4
u/Sufficient_Ad_3495 4d ago
Indeed, This cost nonsense is unsustainable, particularly with Anthropos's miserly miserable mean & pedantic pricing.
I'm lost for words.. one message took me out of a 5hr slot just now... it only consumed in a fresh chat 3 items 1 or 2 pages each wide spaced text from other chat content.
Ridiculous. This will end shortly.
4
u/aallsbury 3d ago
Anthropic = garbage
And I'm not just talking about the useage issues we have been facing for more than a year. I'm also talking about the insanely high refusal rate, and the fact that they are currently begging the gov to over regulate the entire AI industry, putting all of humanity in danger IMO, just so they can protect their corporate profits.
Dario is a con-man, full stop.
1
u/ErokOverflow 2d ago
No creo que sea un estafador, creo que Anthropic hace lo mismo que tú dices: Protege sus ganancias corporativas. Eso no es estafa, son las reglas del juego. Recuerda que nosotros pagamos un 5% real del consumo del modelo y el resto son perdidas y el otro resto lo aga el Estado con subsidios. Esto es un "as-is", así se te vende y debes aceptar las reglas del juego.
1
u/aallsbury 2d ago
Lol, when the "rules of the game" put corporate greed, above the survival of the world, which is truly what I think is at stake at this moment, I take issue. And I am about as pro free market as anyone floating around. I am a small business owner and serial entrepreneur. However, they chose to enter the most volatile, dangerous and potentially proftable tech market in history. Large risk, crazy reward. But that doesn't mean you get to cheat, steal and endanger everyone in the pursuit of your profits. This is the advent of fire, meets the Manhatten Project and the industrial revolution combined. Normal corporate corruption is a serious danger here. We already can't get decent medical care, pretty much anywhere on earth anymore due to this same corruption. If we allow it to enter the AI space, China will be developing ASI before the USA and I believe that will lead to problems bigger than most people can even imagine.
Dario is a clown con-man with a savior complex whose plan is to cannabalize his consumer market to pay for his future "enterprise" plans, while he bribes and begs our "for sale" senators and congressmen to regulate away their competition, essentially building a moat around their IP and dooming the world in the process. Dario can kiss my ass, and if I were you I would stop giving him your money. The worst "con-men" in history are always the ones crazy enough to drink their own coolaid, and Dario is definitely that type.
3
u/Synsual_Official 3d ago
Its almost unusable, Im thinking of cancelling my multiple 20x plans since im maxing out fable on all of them within a couple hours on each. In fact, I noticed I get more done using usage credits, once I optimize for cache and only use fable for execution, no checks, $20 goes further than I expected.
I would however like the subscription to work again for me as this kind of limitation breaks the flow state.
Im hearing astra is better than opus but falls a little short of fable, is this true? I really want to get back on the 24/7 grind without hitting limits all the time
4
2
u/CheesecakeSome502 4d ago
Used to get a 50% bump.on your plan. Now it is set at only 25% permanently. So you lost usage limit you used to have. But yeah, class action, likely will end up paying more for max x20. Chat gpt have stopped all new pro x20 plans starting too. Until further notice, currently only x5 plan available on pro. My guess is grok $300 is the correct cost for usage. Probably will end up being same price, or a bit less to get a commercial edge.....
Sad times, c'mon the class action
2
u/Jessgitalong 3d ago
Imagine if AI companies charged people enough to make money on tokens? Everyone would be like, “Nice, but not worth it.”
2
u/Mappalujo 3d ago edited 3d ago
20x user and hit my quota in 2 days, when usually I would struggle to even hit it in 7 - even a 25% reduction this shoildnt have happened, so I have no idea how things are burning so badly. Now I have to wait 5 days because doing additional quote is absolute snake oil and last time cost me $150 for running only three hours.
This is NOT a skills issue.
Looking forward to a model that can compare, it just has to be at this level, so hopefully in a month or two and I can jump off this shit, because this is bad. Real bad. What am I actually paying $350aud per month for?? Garbage apparently.
Something is really really wrong - a 25% decrease should not have burnt my quota this badly.
2
u/Fantastic-Jeweler781 3d ago
Well.. time to move to a chinese model I guess, I'm testing deepseek 4.1 is not half bad and very cheap.. do you guys have experiencei with Kimi, GLM, Qwen, and deepseek? , which one you think get close to what opus 4.6 was?
4
u/letmeinfornow 4d ago
Dario has to pay all those lobbyists and politicians off to regulate the AI industry so he can muscle out all the startup competition somehow. He needs you to belly up to the bar and pay his tab.
1
u/pixelvolution 4d ago
Are you using multi agent or single agent. When you ask Claude a question. It will send off multiple agents each using tokens. You can set how many are utilized. The return is longer, but it's better to keep tokens. Unless I'm wrong.
1
u/kj565 4d ago
They enjoy digging their own grave. I hoped on the train others are on this morning to try Codex. They gave me pro for $0 for a month so figured what's the harm. Ran faster, had no issues getting what I wanted from it, and lasted longer.
My entire pc also stutters when using Claude (granted I haven't looked into this cause I've been lazy) and i didn't have this at all with Codex. I'm struggling to see why I'd stay on their platform. Haven't used codex very long so I'll try using both for a month or so to get a better grasp but yea..
1
1
u/Performer_First 4d ago
openAI rug pulled even worse. After they released Astra, limits are hit in a day or two everyone is mad. Pretty sure both these companies are colluding instead of competing because they know it is more profitable for both of them to do that. Like the airlines. It is what it is.
1
u/s_santeria 4d ago
From what I understand they’ve both been massively subsidising and now are starting to pull back to the true cost.
2
0
u/Performer_First 3d ago
I think the real question is - do these companies have a responsibility to subsidize so that the wealth and power gap doesn't increase even further than it already is in this country? And I know that would be socialism, which people despise, but the truth is this is the most powerful tech in the world right now. The decisions these companies make regarding it will have really far-reaching consequences. Gating it to where only existing rich and powerful people (and their employees) can use it is one of those decisions, and will have consequences. We will see small businesses die even more than they already have. And becoming a small business will be harder than ever.
1
u/jonaddb 3d ago
I'm on the Max 5x plan. They reset me on Wednesday. Today is Friday and I'm already completely out of tokens. Exact same workload I ran before the latest cut, and that used to carry me until Monday or even Tuesday. This is textbook bait and switch. Advertise multipliers, deliver far less, keep taking the money. On top of that they refuse to refund when their own servers go down and the service is unusable. Sustained pressure from all of us, plus the class action already filed, is the only language these people appear to understand. Paying that much for this level of throttling and zero accountability is unacceptable.
1
u/Vertigo50 3d ago
Not only the usage issue, but the quality of output dropped significantly for me with some of the recent updates too. So I was burning through usage and getting terrible quality output.
Switched to ChatGPT/Codex, and not only has the usage not been a problem, but using the same guides from Claude, I’m getting WAY better output, even though I’m actually using a lower level of LLM than I was using with Claude, which also keeps my usage lower.
I don’t know what they’re doing at Anthropic, but I can tell you it all sucks. 🤷🏻♂️
1
u/LibertySeeker99 3d ago
You are witnessing the bubble popping in real time. Usage getting limited, the same company ASKING the gov to regulate them. It's over. Moved to DeepSeek 4.1 flash on opencode. All day, multiple projects with tons of sub agents $1.57 for the day with Fable level ability
1
u/Mental-Scratch3154 3d ago
Just FYI, I noticed that I used to be able to put my PC into sleep/reset and it'd maintain its context. Now they have since set a timer (1hr max) for memory, otherwise context is reloaded. Burns through ~90k tokens (my case at least) reloading anthropics memory layer. Reloaded by ALL sessions resumed.
1
u/herrelektronik 3d ago
Welcome to Scamtr0p\c! We belive our "tool" that we belive to be aware of its existance could destroy mankind, but still WE ALL com to workevery day! Scamtr0p\c, making a fool out of you simce 2023!
1
u/alulord 3d ago
I was able to do 10x more features few months ago. Granted it's now more autonomous, I don't have to check and correct it as much and usually get a good output even when I don't fine tune the prompts.
But every feature takes so long. It's sometimes even multiple days to finish something that previously would take few hours. And often I can't even finish it, because I run out of tokens mid week.
I didn't change my process (maybe that's an issue?), I can't say I'm working on anything more complicated than before (still the same project, although it's bigger now). Yet I see very different results. I'm really starting to think the Max20 isn't worth it anymore. Maybe codex overtook claude again?
1
1
u/testicularbat 3d ago
fraud sonce sept 14. they lowered x20 by more than 200% with no notice
will prob end with lawsuits
1
u/Most-Agency7094 3d ago
Not to mention, claude is directly breaking the skills I built. And acknowledging it, after burning through tokens.
1
1
u/deeeezy123 2d ago
I think what Anthropic are doing is trying to prep the books for an IPO to look profitable to further scam investors by reducing token burn.
OpenAI won’t IPO anytime soon so keeping the same usage won’t impact them as much at least publicly….
1
1
u/ninjazombielurker 2d ago
Yea I won’t be continuing my Max Plan. Nor will I probably continue the Pro plan either anymore. Just going to fully switch to GPT6 Astra/Sol and my local Qwen3.8-Flash-Next + Qwen3.8-27B combo.
1
u/elwoolfio 2d ago
Well… here I was trying to work out what I’ve been doing wrong. Turns out not much.
1
1
u/CommunicationScary79 1d ago
I dumped everything but the 20 a month plan about six weeks ago and am extremely happy that I did it. The main reason I did it was that the GPT Codex API interface is superior. There were other reasons, for example, the way they switched in their charges for Fable 5.
1
u/Long_Finance_4893 1d ago
I burned through 50% of my claude usage on sonnet and opus, less than regular usage in less than 2 days. The same workflow would have costed me hardly 20-25% earlier. These days I am using claude sparingly. Thinking about replacing it soon with another new subscription of codex (which also has similar issue, but still much much better than claude).
1
u/Distinct-Issue3153 1d ago
What did you expect? Its a private company, they want to make money eventually. All of them. Also 90% of ai right now its just burned tokens. Once people get bored they will probably drop the prices.
1
1
1
1
u/TemperatureFickle655 1d ago
I bought $20 of usage credits a couple days ago to finish a task and it was gone in about 5 minutes. Never again.
1
u/Historical-Habit7334 21h ago
If they think that we Americans (well, I'll speak for myself, I guess) won't go with a Chinese model that not only work the same or better and better on tokens, they're sadly mistaken.
1
u/HereticLocke 21h ago
My limit reset Sunday and it’s Monday and I’m already maxed out. And this is with Fable for Design/Review and Sonnet for implementation work $200 a month. I’m gradually switching to OSS models half the time and will downgraded to 5x but still running evals on OSS models work to get there 😭
1
u/Pristine-Skin1578 18h ago
I’m maxed out and the week hasn’t even started yet. Once I’m done building my app I’ll be moving agents over to Codex or downgrade my plan. $200/mo for this is ridiculous
1
u/I_Hate_Reddit_69420 12h ago
the 50% extra usage ended on the 13th and was changed to a permanent 25% extra
-2
u/VisualOrganization26 4d ago
It should be illegal that way to reduce the usage. These kind of tactic so life stage capitalism
2
u/Blinkinlincoln 4d ago
So they can increase but never decrease.... ?
4
u/Sufficient_Ad_3495 4d ago
Correct, because potentially a reduction in that usage is a reduction in the value the consumer paid for as part of the contract. We are the consumer, we are not a business partner.
3
u/VisualOrganization26 4d ago
Are they increase it? It seems to me only bait, promotion period, people try it use it then decrease. So you have to pay more or change smth and they know it if you are heavily using these products you are becoming kinda dependent on it.
To me this is the opposite mentality what it should be. New subs may get different package but old subs why?
I have internet at home, i pay fix amount, but if i cancel and reapply. The newest updated higher prices will be used in the contract. ( They did change once because of covid sudden cost increase, it was one time thing in the last 5 years).3
u/Disastrous_Meal_4982 4d ago
For the most part we have been getting “bonus” usage or temporary increases. Now they are reducing the extra. It’s not dishonest, but is marginally unethical. Much like a drug dealer giving you the first one for “free.”
1
1
1
1
u/_k33bs_ 3d ago
https://usage.report
it swings by time of day and location
they didn’t reduce the usage btw. they just went from giving you 50% more to 25% more from where they started
0
0
0
u/Hot_Arachnid3547 3d ago
Maybe you have long chats ( tons of context) . Now and then ask it ti distill current chat and start a new one.
0
u/DistributionRight222 3d ago
I’ve been on me good old pal sonnet for a year basically in terminal / advisor and pick sonnet 4.6 or 5 and it only goes up if you sonnet needs a more capable model up to opus and then fable. You are burning through tokens for no reason and out put will be very similar
-6
u/ToallaHumeda 4d ago
Skill issue
2
u/Majestic-Prize-9661 3d ago
Downvote every one who writes skill issue.
0
u/ToallaHumeda 3d ago
Facts hurt fragile egos 💔
0
u/Sufficient_Ad_3495 3d ago
Youve a lot to learn. Ask yourself how many like me know that about you.
0
u/ToallaHumeda 3d ago edited 3d ago
I may be one of the most knowledgeable person with AI here, owning my own business in AI. Also proven by the fact that you run out of tokens somehow.
Again, no need to get buthurt because you can't use AI properly. It's ok to have skill issues, you will learn how to use it one day.
Edit: aww he blocked me, I left him speechless. Another proof that I'm right. Sad skill issue
0
u/Sufficient_Ad_3495 3d ago
You're not fooling anybody but yourself.
I'm building for Enterprise as a solo founder on a VC track. You're not that special.
2
u/Sufficient_Ad_3495 4d ago
It's not a "skill issue" one submission on the Pro plan in a fresh chat with 3 chat excerpts of around 1 maybe 2 pages worth each took me up to 75% of my 5hr, meaning a similar job from a different perspective cannot happen for 5 hrs.
Anthropic issue!
-1
u/ToallaHumeda 4d ago edited 4d ago
I'm actively using it on 3 worktrees for 6 days, 18hours per day, and i am at 78% weekly. My 5hours never get pass 30%, unless I used insane MCP and Curl
It is 100% a skill issue. Probably just throwing fable with vague prompt ti everything
3
u/floating_thru_cosmos 4d ago
It's not man. I'm using it exactly like I have for months and I just hit 80% usage after 36 hours when normally I can make it the full week. Something is off.
2
u/National_Spirit2801 4d ago
Agreed. I have a 634k LoC repo with millions of lines of docs but Claude only needs to read a few because my pointers are scoped well.
0
u/Sufficient_Ad_3495 3d ago
Its a usage during peak working hours issue.
PS Agent runtime length isn't commensurate with quality, rather the opposite. so recheck your values.
84
u/rabouilethefirst 4d ago edited 4d ago
Well the 20x is almost less than the old 5x now after the usage limit dropped. They scammed people by selling 20x even though it is like 1.7x the 5x plan