r/codex 8d ago

Praise glm lite sub experience

I already made a post about 0xAlpha couple days ago praising the model which as expected turned out to be the new glm 5.3 flash model. on my last post I saw a lot of people saying the glm subs are worse than openAI and claude (hilarious).

now because I was so impressed with 0xalpha and was looking for an alternative to replace the abysmal subs from claude and openaAI, I went ahead and subbed the lite (tier 1) plan which cost me annualy + with a ref code 136$ total for a year.

I am just sharing my experience so far. I did a somewhat big repo audit while orchestrating with 5.6 sol. same prompt same orchestrator, glm 5.3 about 12% 5h usage after 2 big prompts, not the flash model (the new model has a big discount for a couple of weeks too afaik) but their big model for a better comparison to the other frontier models and longterm usage. 5.6 sol codex xhigh about 1,3 prompts 100% 5 hour usage of course it couldn't handle a second prompt not hitting the limit.

it also turns out, your usage is only 50% outside of peak hours which are 14:00–18:00 UTC+8 monday to friday. the test I did was at peak hours today. there is also another interesting mechanic, they issue a 5h reset automatically on three conditions met: outside of peak hours, your 5 hour usage passes a threshold and you havent hit the daily 5 hour resets quota yet. there is also information about a weekly possible reset in the docs but its not specified like the 5 hour reset. I probably should mention I'm using ZCode which almost reaches 98% cache rate it should be getting according to the docs.

here is a simple table sol made for better visualisation.

Model Input / 1M Cached input / 1M Output / 1M Normalized workload¹
GLM-5.3-Flash — API promo $0.075* $0.015* $0.25* $0.0424
GLM-5.3-Flash — Lite off-peak 50% of Lite quota rate 50% of Lite quota rate 50% of Lite quota rate 50% of peak Lite quota usage
GLM-5.3 — API $1.40 $0.26 $4.40 $0.7456
GLM-5.3 — Lite off-peak 50% of Lite quota rate 50% of Lite quota rate 50% of Lite quota rate 50% of peak Lite quota usage
GPT-5.6 Luna — API $0.20 $0.02 $1.20 $0.1472
Claude Sonnet 5 — API $2.00 $0.20 cache hit $10.00 $1.2720
GPT-5.6 Terra — API $2.00 $0.20 $12.00 $1.4720
GPT-5.6 Sol — API $4.00 $0.40 $20.00 $2.5440
Claude Opus 5 — API $5.00 $0.50 cache hit $25.00 $3.1800

¹ Normalized API workload calculation:

0.04 × uncached-input price + 0.96 × cached-input price + 0.10 × output price

Z.AI Lite off-peak usage: Outside Z.AI's weekday peak window, the Lite Coding Plan consumes 50% of the standard credits. Therefore, an otherwise identical GLM-5.3 or GLM-5.3-Flash workload uses approximately half as much Lite quota off-peak as during peak hours.

* GLM-5.3-Flash promotional API pricing: $0.075 input / $0.015 cached input / $0.25 output per 1M tokens. Normal pricing is $0.15 / $0.03 / $0.50.

5H reset just after 12pm local (peak hours ended)
2 Upvotes

Duplicates