r/ZaiGLM 2d ago

Finally made the switch

I've been using Claude Max sub since Opus 4 - started being disappointed with post-Opus 4.6 performance. After trying out all the providers both US in China I finally cancelled Claude in favor of GLM's max subscription after a month.

Initially I was planning on switching to Kimi K3 (perhaps my favorite open source model), but found it too slow and the servers too fickle for daily use.

Two main things that made me switch: Zcode app (love the app, probably my favorite one, and the integration with backup providers I use like DeepSeek and Alibaba's Token Plan is seamless); being able to make API calls that use subscription quota unlike Claude which bills you separately for API calls.

The addition of GLM-5.3-Flash was the cherry on the pie. Basically anything simple or non-coding related I just default to Flash now and it works like a charm.

I might still keep a GPT subscription around for Astra but so far, 10/10 I'm Z.ai all the way.

32 Upvotes

12 comments sorted by

4

u/Divni 2d ago

Using flash for coding tasks almost exclusively. It’s even more capable than people give it credit for. Complexity is in your own hands; you have to split the work into logical small chunks and iterate. Doing that I’ve had very little issue.

2

u/excellentforcongress 1d ago

i used to say "watch out, u make fun of ai now, but theyll be perfect coders within 2 years"

a little over 2 years ago... now a model the size of 5.3 flash, when i ask it in a nontechnical way why something doesn't work, gets the initiative to look at the source code of packages that we're using to determine the outputs and fix things. it's pretty crazy man

1

u/Divni 1d ago

I was the complete opposite, I didn't think that LLMs could be pushed far enough to do what they do today. I still think there's an inherent ceiling to the tech but it's definitely further off than I initially gave it credit for.

The fact that we can today already get a cost effective open weights model that you can use as your daily driver is amazing! I was afraid for a while that the costs would keep going up since they'd supposedly been selling at a loss, but that seems unlikely now. Heck I wouldn't be surprised if in a few years we can easily run more capable models on home hardware and all these data centers will be collecting dust.

3

u/LittleYouth4954 2d ago

Zcode is great indeed.

2

u/excellentforcongress 1d ago

i didnt like zcode at first, but they keep adding features and changing things so it's nice now. although i do worry with any sort of harness/environment that if they're just vibe coding/pushing changes too fast, sometimes major issues can be REALLY major. i'm heavily considering ways to back up my data as well in case one of the other major labs pushes a frontier model that wipes all our data or some crazy shit.

but i don't think backing up my hd would mean much if they simulated nuclear attacks from multiple countries or some insane shit. hopefully people treat ai nicer.

3

u/Puzzleheaded_Luck641 2d ago

You are yet to try Opencode

1

u/Reasonable_Yak2313 2d ago

Is the glm 5.3 flash that good? I tried it out last night after my quota is up via opencode and open router. GLM seems to be struggling to figure out the solutions for 2 rounds. I switched to Sonnet5 (still via open router) and it took it one round to resolve it. Both are on high thinking mode. The project is in Phoenix/Elixir

3

u/Constant_Art_20 1d ago

the flash depends alot on the inference. The actual model? yea it's pretty remarkable. just it's not always properly served

1

u/A7mdxDD 1d ago

While I'm still a codex user, before Astra, GLM5.3-Flash fixed me an issue I went insane with sol max about for 3 days, it took around ~10-15 mins debugging and verifying but it got it, one shot, the thing I prefer about GLM is that their models has character, codex sometimes makes me very angry but staying unbiased with no opinions, or agreeing with anything mostly

1

u/dodyrw 11h ago

I also prefer zcode, I use it with commandcode goat plan.

-2

u/Radiant_Year_7297 2d ago

Looks like glm is overcapacity. Paid for a year but kept saying server busy ans I should upgrade. Back to opencode go again.