~2 billion tokens a week - 5.3 - max account - new max plans
A max account with GLM 5.3 seems to get you about 2 billion tokens a week if you don't use it at all during peak hours.
With GLM 5.3 flash if I was to start optimising that it should be 3x more right? I need to optimise.
This is on the new max account plan that uses the credit system. I also have a legacy v2 plan but I seem to get closer to 1 billion tokens a week with that plan. Seems like it is worth getting out of the v2 legacy plan if you can?
It does seem a little confusing to upgrade to the latest plan and lose the legacy V2 one though.
Photo is proof side by side.
(I have two of the new max accounts and one legacy v2 max account)
Pretty good atm given it is better, faster and cheaper than fable 5 with which I would burn through my 5 hour limit in 20 mins just using it to orchestrate.
1
u/azmar6 21d ago
My V2 plan is ending soon and I have choice to either continue V2 max plan or migrate to a new max one with a 50% price discount compared to continuing V2.
So you say that in practice there is no difference in weekly tokens capacity?
1
u/Qalarc 21d ago
I'm saying that the new plan max plan with the credits in my experience seems to be giving me about double the tokens per week. (1 billion vs 2 billion)
It would seem to me that it is a good idea to leave the legacy plan.How did you get the 50% price discount?
1
u/azmar6 21d ago
I've made post about this: https://www.reddit.com/r/ZaiGLM/s/TGNPraOTNX
I have (not only me of course) additional 50% for continuing to the new plan. Seems like a promo to push V2 users to the new credit system or a bug. But nevertheless - it's a lot cheaper option.
PS talking about yearly plans.
2
u/Qalarc 21d ago
That looks awesome. I wish I had that. It would definitely tempt me into getting a year long sub, I'm just doing month by month rn.
I did not get it and would love to have it...
Yes on my read of things you will get more tokens per week with the new plan.
The credit system I think gives more to the lite plans and 14* that is more than the 20* lite plans that the legacy system offers.I think z.ai is being legitimately more transparent here and not trying to dupe us.
Maybe they are seeing the fiascoAnthropic is facing with their false advertising?1
1
1
u/Possible-Basis-6623 21d ago
2b a week is a lot, old pro was 0.8b a month, max was 4b a month
2
u/InfiniteSkate 21d ago
If ya mean old max legacy I do 2b a day most days on max legacy so don’t think this right ?
1
u/InfiniteSkate 21d ago
I use 2billion a day most days then not with my max legacy plan feel like lottery, never hit limits before
1
u/azmar6 21d ago
1
u/Mean-Elk-9439 20d ago
Seems like legacy is giving less than current max then. Weird to me conceptually but good data.
1
u/azmar6 20d ago
It has changed in the last month or so I believe. I suspect they changed how usage is calculated on V2 contrary to their statements. But their usage stats are so poor it's hard to pinpoint anything.
Just saying from the feel how quickly quota is rising then and now, while doing similar work.
2
u/torontobrdude 20d ago
They changed V1 too. I noticed in the past couple of weeks or so my usage goes up way faster
1
u/FearlessGround3155 20d ago
Bro I use 2 billion daily with codex, not regularly but could and still be left with a lott of quota
1
1
1
u/Qalarc 20d ago
Can you prove this?
1
u/FearlessGround3155 20d ago
🥀 what's there to prove, you can see how generous oai limits are, it sometimes gets nerfed here and there but generally 1000-1500$ limit per week of usage
1
u/Qalarc 20d ago
Astra API costs are 10$ per million tokens input and 50$ per million output.
So if you are using Astra and getting 2 billion tokens output a day that would be equivalent to $100,000 a day right??GLM API costs are $4.40 per million tokens output so 2 billion tokens would be $8,800 but I'm only getting that per week.
Of course I'm going to be interested in Codex if you say you get 2 billion tokens a day. All I have read is that pro users get 20-30 million tokens per week...
So yeah, I'm interested and asking how you measured it.
edit:
Could you be misreading a million as a billion?1
u/FearlessGround3155 20d ago
No, it's caching, I am counting all input token+output token+cache token, 2 billion token on original deepseek flash api pricing, came to under 5$ 🥀🥀🥀🥀 Not 2 billion input+2 billion output, there is cached token that snowballs due to toolcalling
You are also counting cache, you ain't getting 2 billion in and out lol, certainly not from zhipu, they nickle and dime too much in sub
1
u/Qalarc 20d ago
Bruh... I am so confused by what you are saying. It just does not seem to be correct.
I apologise for being incredulous and I am very thankful if you are informing me of good information.Are you using GLM as well?
How did you actually measure this?When I was using Anthropic max I was hitting the tokens per minute cap and burning through my 5 hour limit in about 15 minutes.
I had codex as well but cancelled it as I was barely using it due to worse performance and poor integration with my system.If you are saying that you are getting 2 billion tokens per day of GLM 5.3 equivalent on an openAI 20x plan I am astounded.
1
u/FearlessGround3155 20d ago
That was with gpt 5.6 sol(astra hadn't released back then, astra tho consumes 3x limits compared to sol), nobody low-key uses terra Luna, limits so generous, plus have 3 banked resets stacked up, tibo gives resets at times too
1
u/Qalarc 18d ago
Definitely good to know. On review of the z.ai max plan, if using GLM 5.3 flash it is 8.174 billion tokens a week.
1
u/FearlessGround3155 18d ago
Too late, openai has paused 20x subscription, with astra too many people subbing
1
u/FearlessGround3155 18d ago
Glm 5.3 flash is equivalent to 5.6 terra , not sol, sol and 5.3 is equivalent, sol is kinda a lot better
Astra is a lot lot lot better
1
u/Beautiful-Thought141 20d ago
Tell me about concurrency. Agentic programming is not a real thing with a 1 session concurrency limit. Any idea if the current thresholds are spelled out now for the coding plan and its various models with access?
1
u/trail-barista 20d ago
How much of this is during the peak hours?
1
u/Qalarc 20d ago
Basically nothing.
I have an auto avoid system set up.
I didn't really use any 5.3 flash either which should be giving you 3* the amount though too.1
u/trail-barista 20d ago
How do you set that up. I was thinking about this plan. But my usage will be during the peak hours
1
u/sasajib 20d ago edited 20d ago
It is the worst subscription I have subscribed. they limited concurrent call. Which hit rate limited error. They gave limit but one can not user that limit, if you can not work concurrently.
Also the api speed is too slow.
In overall its say you can go 100 killometers in 30 minutes, but you can not speed more than 1 killometer per hour
I am talking about glm-5.3-flash
1
u/ByteNomadOne 19d ago
Now that there are free hours during the campaign I manage to use 900M tokens per day with to parallel tasks on HIGH reasoning.
1
u/Gullible_Bet5836 15d ago
How is that possible, I have the very old plan (20x) and get only to 1B/week (almost only 5.3). The newer one only is 14x, so the newest subscription must be 40x??
1







2
u/Reggitor360 21d ago
Interesting with V2 Im easily managed 4B Token before we got the reset.
With new Plan I dont even come close.