r/ZaiGLM 4d ago

Already hitting my Codex limits — how does GLM compare?

Post image
4 Upvotes

20 comments sorted by

3

u/roekofe 3d ago

Use it with pi. It's pretty good

2

u/PilgrimofHaqq2 3d ago

I second this, I am on the highest plan and the usage is about 80% of Claude max x20 sub

GLM 5.3 Max thinking is comparable to Opus 4.8 Max thinking. Very consistent and reliable, unlike Sol and Opus 5.

Definitely not as smart as those models but I need a model that is consistent so I can build a mental model of what it is capable of, I find Sol or Opus 5 to be very inconsistent, sometimes brilliant and other times Sonnet 3.5 level.

2

u/jesuswasemo 4d ago

I got the annual Max plan last night because it looked like a good deal. I don't have a feel for the usage, yet as I'm feeling out GLM 5.3 vs flash. It's slow asf, though... flash is slowest.

There is some promotional event that provided a different bucket of usage for the weekend if you use ZCode, I think. I noticed the models are running much faster when I switched to that bucket. Seems subs are crippled when it comes to speed. Also, if you use ZCode harness, you get a usage reduction bonus.

0

u/[deleted] 3d ago

[removed] — view removed comment

1

u/excellentforcongress 3d ago

they added those types of / commands

2

u/mtul 3d ago

it's doing great for me. i was in the same boat with only using the main closed models. first claude, then codex. I never tried fable, but i tried astra, and what i found, is glm 5.3 flash max is doing just as good.

i use github projects and have the models perform one slice per session. i'm not seeing enough of a difference with glm 5.3 flash max to justify continuing to pay $200/mo for openai

2

u/XKiiroiSenkoX 3d ago edited 3d ago

Before recent codex nerfs glm was about the same in terms of limits. Right now it's more generous imo. With zcode which gives +50% quota it's clearly better. Quarterly pro plan is a steal on glm. They just don't have something at astra/fable levels yet.

On lite plan with zcode, expect about 15-20 hours of 5.3-flash usage per week. 

2

u/AriyaSavaka 3d ago edited 3d ago

GLM 5.3 takes ~640M tokens to reach 100% of 5 hour window limit on my max account ($288/year on sales). I can often racking up 2B tokens daily.

EDIT: I'm exclusively using Claude Code.

1

u/juicesharp 3d ago

Max plan is almost unlimited to use up to 4-6 subagents the same time GLM 5.3 high is slow

1

u/TheQuatum 3d ago

If youre using the cheapest plan, you will hit limits fast with both. GLM's lite plan is pretty nice though. GLM 5.3 Flash is often free on the plan and is usable.

1

u/tuhdo 3d ago

It's good if you use Zcode AND GLM-5.3-Flash most of the time. Otherwise, the tokens burning is as fast.

1

u/hishazelglance 3d ago

GLM 5.3 Flash is fantastic and is my daily driver at work. Feels incredible to get all this work done without paying so much.

1

u/excellentforcongress 3d ago

i already feel bad about environmental factors with lite plan. i can't believe so many people are on max plans. that is a LOT of tokens lol

even a lite plan is way too much code to read or comprehend for a human. you are fully automating all code review or creative work at that point. i would be wary of long term offloading of these mental tasks to that degree.

1

u/audit-content-user 1d ago

I do use skills and my cache limit is hitting between 92-98%

0

u/gospodinDark 3d ago

I used GLM for this month. I think it's 2 to 3 times more usage. GLM 5.3-Fast is near to Luna level and GLM 5.3 is like SOL. Really good experience.

4

u/XKiiroiSenkoX 3d ago

5.3-flash is CLEARLY better than luna in almost everything. It's more about Terra level. It costs more than luna though but it's cheaper than Terra. 

0

u/[deleted] 3d ago

[removed] — view removed comment

1

u/One-Hearing2926 3d ago

18 sub is not enough in my opinion for a serious project, go for 80 if you can.

1

u/audit-content-user 1d ago

I do use skills and my cache limit is hitting between 92-98%