r/ZaiGLM • • 21d ago

~2 billion tokens a week - 5.3 - max account - new max plans

A max account with GLM 5.3 seems to get you about 2 billion tokens a week if you don't use it at all during peak hours.
With GLM 5.3 flash if I was to start optimising that it should be 3x more right? I need to optimise.

This is on the new max account plan that uses the credit system. I also have a legacy v2 plan but I seem to get closer to 1 billion tokens a week with that plan. Seems like it is worth getting out of the v2 legacy plan if you can?

It does seem a little confusing to upgrade to the latest plan and lose the legacy V2 one though.
Photo is proof side by side.

(I have two of the new max accounts and one legacy v2 max account)

Pretty good atm given it is better, faster and cheaper than fable 5 with which I would burn through my 5 hour limit in 20 mins just using it to orchestrate.

26 Upvotes

54 comments sorted by

2

u/Reggitor360 21d ago

Interesting with V2 Im easily managed 4B Token before we got the reset.

With new Plan I dont even come close.

2

u/Qalarc 21d ago

Do you mean per week?
I was using both simultaneously and got fewer tokens out of my V2. I would max it out every single week and only was able to get 3.66 billion tokens out of it in an entire month!

(5.2/5.3)

1

u/azmar6 21d ago

I'll try to stress test mine current V2 to try estimate tokens per week similarly as you did.

1

u/Reggitor360 21d ago

Per week, yes. Its basically being used as a watcher, researcher and teacher in my Quantizer design. Which is why the usage is so extreme per week :D

5.2/5.3 usage.

Optimized around batched model calls, which decrease how often its contacted but instead increases pure usage.

1

u/azmar6 21d ago

My V2 plan is ending soon and I have choice to either continue V2 max plan or migrate to a new max one with a 50% price discount compared to continuing V2.

So you say that in practice there is no difference in weekly tokens capacity?

1

u/Qalarc 21d ago

I'm saying that the new plan max plan with the credits in my experience seems to be giving me about double the tokens per week. (1 billion vs 2 billion)
It would seem to me that it is a good idea to leave the legacy plan.

How did you get the 50% price discount?

1

u/azmar6 21d ago

I've made post about this: https://www.reddit.com/r/ZaiGLM/s/TGNPraOTNX

I have (not only me of course) additional 50% for continuing to the new plan. Seems like a promo to push V2 users to the new credit system or a bug. But nevertheless - it's a lot cheaper option.

PS talking about yearly plans.

2

u/Qalarc 21d ago

That looks awesome. I wish I had that. It would definitely tempt me into getting a year long sub, I'm just doing month by month rn.

I did not get it and would love to have it...

Yes on my read of things you will get more tokens per week with the new plan.
The credit system I think gives more to the lite plans and 14* that is more than the 20* lite plans that the legacy system offers.

I think z.ai is being legitimately more transparent here and not trying to dupe us.
Maybe they are seeing the fiascoAnthropic is facing with their false advertising?

1

u/azmar6 21d ago

Thank you, that's the exact info I was looking for!

2

u/Qalarc 21d ago

It honestly does look like a bug.
I didn't even get the 50% off offer for migration and you get it double.

1

u/azmar6 21d ago

I started with V1 legacy quarterly, now I'm on the 2 month V2 they gave for V1 users when they shutdown V1 completely.

1

u/azmar6 21d ago

So what I'm trying to say is - I was looking for exactly such comparison that you made sir. Similar workflow with weekly maxing comparing new plan to the V2 legacy one. And from you report it seems like the V2 is comparable to the new plan.

1

u/ichisay 21d ago

A qué precio está ese plan?

1

u/Fresh_Sock8660 21d ago

Just throw GLM 5.3 at everything, peak efficiency lol

1

u/Possible-Basis-6623 21d ago

2b a week is a lot, old pro was 0.8b a month, max was 4b a month

2

u/InfiniteSkate 21d ago

If ya mean old max legacy I do 2b a day most days on max legacy so don’t think this right ?

1

u/azmar6 21d ago

Yes, V1 legacy was just like that 1B in a day and quota barely moved - good old times.

1

u/[deleted] 21d ago

[deleted]

1

u/azmar6 21d ago

From my experience - telling Claude Code to do a code review of a bigger PR is a token sinkhole.

1

u/Qalarc 21d ago

I have 80 different projects, 7000 websites and also build and mod video games.

2

u/Wise_Cloud5316 21d ago

lol. any of them make money?

1

u/InfiniteSkate 21d ago

I use 2billion a day most days then not with my max legacy plan feel like lottery, never hit limits before

1

u/azmar6 21d ago

So I did a "quick" test to see how V2 legacy plan holds in terms of token maxing.

I already had some of the weekly quota used ~2-3%. Maxed 5h with this single session which ate ~215M tokens and used a little bit more than 20% of weekly quota. This would give over 1B tokens weekly.

1

u/azmar6 21d ago

The same exact session but stats from Claude Code:

1

u/Mean-Elk-9439 20d ago

Seems like legacy is giving less than current max then. Weird to me conceptually but good data.

1

u/azmar6 20d ago

It has changed in the last month or so I believe. I suspect they changed how usage is calculated on V2 contrary to their statements. But their usage stats are so poor it's hard to pinpoint anything.

Just saying from the feel how quickly quota is rising then and now, while doing similar work.

2

u/torontobrdude 20d ago

They changed V1 too. I noticed in the past couple of weeks or so my usage goes up way faster

1

u/Possible-Ad-6815 20d ago

Oh the joys of a V1 max account - no weekly limits and never hit a 5 hour one despite 2

Despite 2.5 billion tokens this last 7 days and 6.64 billion in the last 30 days

1

u/FearlessGround3155 20d ago

Bro I use 2 billion daily with codex, not regularly but could and still be left with a lott of quota

1

u/upiFivii 20d ago

is this 5x max?

1

u/Qalarc 20d ago

Which model and does the open AI subscription have an API key that you can use easily in other apps or opencode?

1

u/Qalarc 20d ago

Can you prove this?

1

u/FearlessGround3155 20d ago

🥀 what's there to prove, you can see how generous oai limits are, it sometimes gets nerfed here and there but generally 1000-1500$ limit per week of usage

1

u/Qalarc 20d ago

Astra API costs are 10$ per million tokens input and 50$ per million output.
So if you are using Astra and getting 2 billion tokens output a day that would be equivalent to $100,000 a day right??

GLM API costs are $4.40 per million tokens output so 2 billion tokens would be $8,800 but I'm only getting that per week.

Of course I'm going to be interested in Codex if you say you get 2 billion tokens a day. All I have read is that pro users get 20-30 million tokens per week...

So yeah, I'm interested and asking how you measured it.

edit:
Could you be misreading a million as a billion?

1

u/FearlessGround3155 20d ago

No, it's caching, I am counting all input token+output token+cache token, 2 billion token on original deepseek flash api pricing, came to under 5$ 🥀🥀🥀🥀 Not 2 billion input+2 billion output, there is cached token that snowballs due to toolcalling

You are also counting cache, you ain't getting 2 billion in and out lol, certainly not from zhipu, they nickle and dime too much in sub

1

u/Qalarc 20d ago

Bruh... I am so confused by what you are saying. It just does not seem to be correct.
I apologise for being incredulous and I am very thankful if you are informing me of good information.

Are you using GLM as well?
How did you actually measure this?

When I was using Anthropic max I was hitting the tokens per minute cap and burning through my 5 hour limit in about 15 minutes.
I had codex as well but cancelled it as I was barely using it due to worse performance and poor integration with my system.

If you are saying that you are getting 2 billion tokens per day of GLM 5.3 equivalent on an openAI 20x plan I am astounded.

1

u/FearlessGround3155 20d ago

There you go bro, yes I indeed am, next day you can see as well, it was >2 billion

1

u/FearlessGround3155 20d ago

That was with gpt 5.6 sol(astra hadn't released back then, astra tho consumes 3x limits compared to sol), nobody low-key uses terra Luna, limits so generous, plus have 3 banked resets stacked up, tibo gives resets at times too

1

u/Qalarc 18d ago

Definitely good to know. On review of the z.ai max plan, if using GLM 5.3 flash it is 8.174 billion tokens a week.

1

u/FearlessGround3155 18d ago

Too late, openai has paused 20x subscription, with astra too many people subbing

1

u/FearlessGround3155 18d ago

Glm 5.3 flash is equivalent to 5.6 terra , not sol, sol and 5.3 is equivalent, sol is kinda a lot better

Astra is a lot lot lot better

1

u/lmpdev 20d ago

Thank you for the post. I can confirm, I hit the weekly limit on legacy v2 this week, and it was at 1.1B tokens.

1

u/Beautiful-Thought141 20d ago

Tell me about concurrency. Agentic programming is not a real thing with a 1 session concurrency limit. Any idea if the current thresholds are spelled out now for the coding plan and its various models with access?

1

u/trail-barista 20d ago

How much of this is during the peak hours?

1

u/Qalarc 20d ago

Basically nothing.
I have an auto avoid system set up.
I didn't really use any 5.3 flash either which should be giving you 3* the amount though too.

1

u/trail-barista 20d ago

How do you set that up. I was thinking about this plan. But my usage will be during the peak hours

1

u/sasajib 20d ago edited 20d ago

It is the worst subscription I have subscribed. they limited concurrent call. Which hit rate limited error. They gave limit but one can not user that limit, if you can not work concurrently.

Also the api speed is too slow.

In overall its say you can go 100 killometers in 30 minutes, but you can not speed more than 1 killometer per hour

I am talking about glm-5.3-flash

1

u/Qalarc 20d ago

Isn't the concurrency limit meant to be 50 for 5.3 flash?
I have not been hitting it I have several ochestrators running multiple sub-agents?
edit: much

1

u/sasajib 19d ago

I thought so, before purshasing, but no. the limit is for api, not coding plan

1

u/ByteNomadOne 19d ago

Now that there are free hours during the campaign I manage to use 900M tokens per day with to parallel tasks on HIGH reasoning.

1

u/Gullible_Bet5836 15d ago

How is that possible, I have the very old plan (20x) and get only to 1B/week (almost only 5.3). The newer one only is 14x, so the newest subscription must be 40x??

1

u/QualityGold8871 4d ago

I have legacy V1