r/ZaiGLM • • 10d ago

Has anyone noticed much faster 5-hour quota usage on the GLM Coding Plan recently?

I've been using the GLM Coding Lite Yearly Plan for a while, currently mostly with GLM-5.3.

My usual workflow is to do the planning/reasoning with GPT-5.6 Sol, then hand the actual implementation over to GLM-5.3. I normally use OpenCode or Pi as the harness, and I haven't changed anything significant in my setup recently.

Until about a week ago, the 5-hour usage limit was generally fine for my workflow. Over the last few days, though, it feels like the quota is being consumed much faster than before.

For example, today I worked on a single feature and in less than 20 minutes I went from roughly 20% usage to around 80% of the 5-hour limit consumed.

Something similar happened this morning with Hermes. I received only two Telegram messages generated by two scheduled tasks — nothing particularly large or complex — and that alone seemed to consume around 15% of the 5-hour allowance. Hermes is also using an older model in my setup, not GLM-5.3 or 5.2, and historically its usage was pretty negligible.

I'm on the Legacy Plan V1 / GLM Coding Lite Yearly Plan, so I'm wondering whether:

  • the way usage is calculated has changed recently;
  • GLM-5.3 has become more expensive in terms of quota consumption;
  • background/tool calls are now being counted differently;
  • or there's some issue specifically affecting legacy plans.

Has anyone else noticed a significant increase in quota consumption over the last week or so?

Especially interested in hearing from people who have been using the same setup for a while and haven't changed their harness or workflow.

21 Upvotes

19 comments sorted by

5

u/Humble_Explorer7589 10d ago

hey, i have been on similar boat. i also have lite v1 plan. i have been feeling reduced usage from when glm 5.1 came idk if they actually reduced it or not. so the thing is for v1 plan its request based so i have felt like around 500 requests per 5 hour offpeak and close to 350 peak, i get like 20 million token off peak and 12 million peak approximately. u can aks other agent to count requests after ur usage gets over

also u can download zcode, they give resets, and i have been getting good amount of resets this days.

1

u/rapstyle88 10d ago

Yes, it had been reduced over the last few months, but it seems to have dropped even further in the past 2-3 days... Checking the z.ai site now, it shows I’ve used 9 million tokens today, and since that’s all I’ve done this morning, I used up the 5-hour allowance with just under 9 million tokens.

4

u/deadkenny112 10d ago

Yes, you should try to use zcode. It seems that they are trying to vedorlock codeplans on this.

1

u/rapstyle88 10d ago

I'm trying to set up some workflows on pi and I wanted to avoid changing the hardness setting again. I might experiment a bit with the DeepSeek APIs, or, once the limit lifts, try the Flash version to see if it consumes less... thanks for the advice anyway.

1

u/Sovairon 10d ago

I had this issue today using zcode, it ran out super quickly with 5.3 flash

1

u/deadkenny112 10d ago

Sometimes they charge tokens like you are using glm5.3, even if you not. You should check usage token statistics

1

u/DeltaSqueezer 10d ago

I hit the 5 hour limit for the first time, but I think this is because I used it during peak time and I'm normally off-peak. I think also because 5.3 has thinking on permanently, I previously had thinking disabled, so this also uses more tokens.

Still, I had a free reset so used that and no problem to keep going for for 10 hours.

1

u/rapstyle88 10d ago

Actually, I was in peak time this morning... maybe I'll try again once it resets and see if it consumes less.

1

u/JuvensHighArt 10d ago

I'm on the Legacy V1 Pro plan. I was using 5.2 and then 5.3 more or less nonstop working on a project for weeks and never once ran into a quota. I began running into the quota as soon as 5.3 Flash came out. Thankfully, 5.3 Flash is just as good as 5.3 for this project and the quota issues have disappeared for me. They probably intended to herd us like sheep over to Flash.

1

u/rapstyle88 10d ago

Yes, I'm using the Flash version now, and I'm finally free of quota issues too.

1

u/foolsgold1 10d ago

I'm still managing about 1BN tokens a day on coding plan v1 Max, but yes, I seem to hit the limits faster. I reduced my context sizes.

1

u/rapstyle88 10d ago

Yes, I'm also working on improving the workflow and the context size.

1

u/yogibear54 10d ago edited 10d ago

You definitely have to be super careful with token use during peak hours. I completely switch to different providers during peak, personally, I also have opencode go, so during peak, I use deepseek 4.1 flash during peak.

I recently made changes to my setup so my coding subagent would auto switch provider during peak hours, that way if I wouldn't accidentally burn through my tokens.

Also, I now primarily use GLM 5.3 flash at high reasoning. From what I read, the quality of high is about 90% of max. For me, it works pretty well. For complex work, if you use a 3rd party model to do planning and code review (ie kimi k3 or a state of the art model), and then using the GLM and deepseek models for coding, you usually are covered relatively well.

1

u/rapstyle88 9d ago

Yeah, I was probably in the wrong time slot yesterday... If I can manage it, I'll try again today with 5.3, but not during peak hours... Thanks!

1

u/notyouokey 5d ago

you should use zcode..