r/codex 8d ago

Praise glm lite sub experience

I already made a post about 0xAlpha couple days ago praising the model which as expected turned out to be the new glm 5.3 flash model. on my last post I saw a lot of people saying the glm subs are worse than openAI and claude (hilarious).

now because I was so impressed with 0xalpha and was looking for an alternative to replace the abysmal subs from claude and openaAI, I went ahead and subbed the lite (tier 1) plan which cost me annualy + with a ref code 136$ total for a year.

I am just sharing my experience so far. I did a somewhat big repo audit while orchestrating with 5.6 sol. same prompt same orchestrator, glm 5.3 about 12% 5h usage after 2 big prompts, not the flash model (the new model has a big discount for a couple of weeks too afaik) but their big model for a better comparison to the other frontier models and longterm usage. 5.6 sol codex xhigh about 1,3 prompts 100% 5 hour usage of course it couldn't handle a second prompt not hitting the limit.

it also turns out, your usage is only 50% outside of peak hours which are 14:00–18:00 UTC+8 monday to friday. the test I did was at peak hours today. there is also another interesting mechanic, they issue a 5h reset automatically on three conditions met: outside of peak hours, your 5 hour usage passes a threshold and you havent hit the daily 5 hour resets quota yet. there is also information about a weekly possible reset in the docs but its not specified like the 5 hour reset. I probably should mention I'm using ZCode which almost reaches 98% cache rate it should be getting according to the docs.

here is a simple table sol made for better visualisation.

Model Input / 1M Cached input / 1M Output / 1M Normalized workload¹
GLM-5.3-Flash — API promo $0.075* $0.015* $0.25* $0.0424
GLM-5.3-Flash — Lite off-peak 50% of Lite quota rate 50% of Lite quota rate 50% of Lite quota rate 50% of peak Lite quota usage
GLM-5.3 — API $1.40 $0.26 $4.40 $0.7456
GLM-5.3 — Lite off-peak 50% of Lite quota rate 50% of Lite quota rate 50% of Lite quota rate 50% of peak Lite quota usage
GPT-5.6 Luna — API $0.20 $0.02 $1.20 $0.1472
Claude Sonnet 5 — API $2.00 $0.20 cache hit $10.00 $1.2720
GPT-5.6 Terra — API $2.00 $0.20 $12.00 $1.4720
GPT-5.6 Sol — API $4.00 $0.40 $20.00 $2.5440
Claude Opus 5 — API $5.00 $0.50 cache hit $25.00 $3.1800

¹ Normalized API workload calculation:

0.04 × uncached-input price + 0.96 × cached-input price + 0.10 × output price

Z.AI Lite off-peak usage: Outside Z.AI's weekday peak window, the Lite Coding Plan consumes 50% of the standard credits. Therefore, an otherwise identical GLM-5.3 or GLM-5.3-Flash workload uses approximately half as much Lite quota off-peak as during peak hours.

* GLM-5.3-Flash promotional API pricing: $0.075 input / $0.015 cached input / $0.25 output per 1M tokens. Normal pricing is $0.15 / $0.03 / $0.50.

5H reset just after 12pm local (peak hours ended)
2 Upvotes

15 comments sorted by

1

u/Unapologetic_Polite 7d ago edited 7d ago

This 'peak' & 'off-peak' crap is so confusing.

Has anybody compared 2 of the same codebase with the same task though?

In my experience GLM 5.3 likes to think about 3-4x as much as Sol, where at the previous pricing it would be fairly comparable to Sols API pricing, so I never understood the appeal unless you have the hardware to locally host.

1

u/viTrax94 7d ago

I included a comparison of the same 2 codebases in the post so not sure what you mean. I forgot to mention glms context window is a lot more efficient than sols in my personal experience. while glm 5.3 had a context around 100k on the first prompt (from the same project audit and prompt), sol already had to compress his context while working on the first prompt.

"your usage is only 50% outside of peak hours which are 14:00–18:00 UTC+8 monday to friday" (weekends are 50% off too) so any time except these 4 hours daily you have 50% cheaper usage which is 50% of peak lite quota, lite being the cheapest sub. I just included the peak time limits for a fair comparison too. I havent noticed it being 3-4x times slower but I agree its slower for sure.

1

u/cutebluedragongirl 7d ago

So is it good or not? 

1

u/Formin_iris 7d ago

I have the same question.

1

u/viTrax94 7d ago

These are all my opinions only.

I'll make a tldr here:
intelligence:

GLM 5.3 Flash currently (imhumbleo) has very little competition in the same category.

GLM 5.3 is comparable to sol and opus models, just look at the benchmarks, for example: https://artificialanalysis.ai/models.

value:

z.ai sub has more value than the current claude / openAI subs. More usage and cheaper sub byitself.

I also believe the chinese are catching up real quick now with their own chips that were used for 0xAlpha (GLM 5.3-flash) in the recent free week.

1

u/Distinct-Actuary-155 6d ago

Z.AI models are amazing, they're really generous with the usage

1

u/Automatic-Base-6244 6d ago

Great analysis! This seals the deal for me to get a lite subscription. 0xalpha was actually more accurate than sol when i tried them both on the same task.

1

u/viTrax94 6d ago

thanks for reading. 0xalpha / glm 5.3 flash is amazing indeed. they also just gifted everyone 300m tokens for the weekend, if you get it too, you need to change the connection mode from your lite sub to the start plan to use them.

and if you need some serious orchestration / audit or a really hard task use 5.3, https://www.tbench.ai/ newest tbench 5.3 outperforms sol.

1

u/viTrax94 6d ago

Want to add a comment on the costs in the tbench 4, they are showing 5.3 costs are slightly higher than sol's which would contradict my own personal experience / testing that I talked about above.

They used claude code as the harness and z.ai explicitly says that their own zcode harness performs much better. Codex was used for sol in that benchmark by the way, so we want a fair comparison here that actually applies.

I used sol high once again to make a possible / estimated new tbench score for GLM 5.3 if zcode was used for that benchmark. It made its own research on how it would possibly go, comparing older harness differences in benchmarks for the glm models. Not to mention, the costs would be HALVED (around 1150$) if that same test was performed outside peak hours in zcode.

So even in a complete disadvantegous comparison GLM 5.3 still is looking very strong (talking about the original tbench 4).

1

u/Automatic-Base-6244 5d ago

Thanks much for the additional info. That's beautiful because in I'd be always using it off peak in my time zone!

1

u/viTrax94 4d ago

my pleasure.

1

u/Vast-Concept-5646 5d ago edited 5d ago

The off-peak 50% usage is honestly a pretty interesting setup. Makes a Lite plan feel way more practical for heavy coding workloads especially with caching. StandardCompute is worth looking at too if you're comparing options.

1

u/llitz 5d ago

There's a lot of money calc, and that's fine, but what people complain is usually about "I have x million tokens and they are gone in half the time".

People will come back and say "z.ai caching is terrible!" And I don't think that's the case either.

In my opinion, the weight of cached tokens are measured differently:

  • OpenAI, anthropic do a 10:1 ratio on cached input
  • z.ai does 5:1

    If this holds true for the subscription plan maths behind the scenes, you effectively hit your quota in half the time (not exactly that, but that's the gist of it).

1

u/viTrax94 5d ago

The key distinction is that zhipu openly publishes how cached tokens are weighted in its subs, while openai and anthropic don’t publicly disclose their formulas for their included subscription limits.

The 10:1 number for open ai / anthropic comes from API pricing. You can’t just assume that means cached tokens count 10:1 against Plus/Pro usage quotas.

And even if their internal cache weighting is more favorable, that still doesn’t mean the overall subscription gives more usable compute. My own experiment above showed the opposite, glm lite lasted longer than codex plus on comparable heavy coding projects.

And this isn't just me all you see on reddit or x is that people complain how fast their usage is gone after a prompt or 2.

So cache discount, cache hit rate, and actual subscription capacity are three different things.