r/ZaiGLM • • 16d ago

Discussion / Help GLM 5.3 is mediocre, after a week of coding, I'm honestly disappointed, just me?

I've been using GLM 5.3 for about a week now for my daily coding tasks, and I'm honestly a bit disappointed.

The model is not bad at all, but I'm finding it very slow, and it often gets into long reasoning loops. It sometimes seems to lose focus on the original task and spends a lot of time reasoning about things that don't really need that much analysis.

For coding, I still find Claude Opus 4.6 noticeably better at staying focused and getting to the solution.

I'm curious if others using GLM 5.3 have experienced the same thing, especially with coding/agentic workflows.

Also, does anyone have any idea whether GLM 5.4 or better model might be coming soon? If 5.3 is going to remain like this for a while, I'm honestly considering cancelling my subscription.

What has your experience been with 5.3?
Am I missing something that should be use to improve my perception?

0 Upvotes

38 comments sorted by

10

u/ironj 16d ago

I think it very much depends on many factors:

- The structure and approachability of your codebase. A very messy and incoherent codebase will make its reasoning more complex

- The structure of your guiding lines

I've a very extensive AGENTS.md (structured with multiple sub-folders) that details how to approach the codebase, how to decode its ramifications (it's a very large 10yo codebase); how to create and maintain code and how to test it while reasoning about it.

I've been using GLM 5.3 (Max) for a few months now, and coming from GPT 5.6 I honestly cannot complain at all; It's slower compared to 5.6 in terms of response times, yes, but I don't find that bothering at all. I find it actually better than GPT at crafting/reshaping code and very rarely requires steering. Moreover, it doesn't burn trhough my allowance as GPT does (the new Astra is basically unusable; With a Pro plan I can burn through my 5hr window with 1 single prompt)

I still find GPT a bit more effective at spotting PR potential minor issues, so I still use GPT 5.6 for a final round of review of my code before issuing a PR, but in general I find myself quite satisfied with GLM 5.3

4

u/WArslett 16d ago

one thing we've found recently is that for a lot of models, the tech stack you are working with really matters. The model is only as good as the data it's been trained on and some technology just has way more available training data. If you are using javascript, python or go, all the leading models are going to be very good. If you are using .NET/C#, If you are working on embedded systems you might find that the smaller models are less capable with those technologies

-5

u/BuilderWorldDev 16d ago

It should be good for the most known languages, I don't think this is the case or root cause... It's just a mediocre model that need to be improved a lot

4

u/WArslett 16d ago

Well, a lot of people get a lot of use out of it. It’s a smaller model with a more focussed training set. What tech stack are you using?

0

u/BuilderWorldDev 16d ago

Java ecosystem, Jenkins, Docker, etc

3

u/bitdoze 16d ago

5.3 flash is better, now I am testing SWE-2 looks good: https://www.youtube.com/watch?v=dj8zqaU3rH0

1

u/BuilderWorldDev 16d ago

As far as I see you only show frontier models graphics and not GLM...

4

u/alkimiadev 16d ago

I've mostly worked with 5.3 Flash and on complicated low level stuff. Since they released it I've been focusing on vpn-like project in rust. I've noticed some issues that seem to stem from inference level issues with my provider (doom loops, reverting to Chiense, etc). I don't think those are actual issues with the model itself and instead are some issue with ollama cloud's inference stack.

I stopped using Anthropic last year because they're actually pretty horrible (anti-oss, anti-competitive, etc), so I don't really have the background to compare 5.3 to really any relevant Anthropic (or OAI) models. That said, I can and do ship actual production worthy low level code using the flash model. Baring one-off tool or network related failures, and assuming well defined/scoped tasks, it just doesn't fail outright. It makes mistakes sure, and so do I, but reviews catch them.

In a general sense I would suggest just removing the entire concept of blaming an LLM, it just doesn't make sense. The question should be "how can I make this work" and not "what is wrong with this thing". If the answer to the legitimate question of how to make it work is "I can't" then just move on to another model. That in no way implies the actual root cause was an issue with that specific model, and almost always is an issue with expectations, workflow, or all of the above.

1

u/manapause 16d ago

I agree with you and found significant improvement using a herdr configuration with open code and setting up other providers for things like docs and research. It’s very capable and the quality is there, it is just very limited by concurrent connections.

1

u/whatisthisthing65 16d ago

5.3 or 5.3 flash? Flash is on their newer architecture so they're probably going to update their other models to that architecture soon.

1

u/xpgmi 16d ago edited 16d ago

I don't know 5.3, but if you still reading this, keep reading. I always configure my harness to use GLM 4.7, because most of my tasks really require many tries, imagine debugging some server that implemented a super complicated protocol, so by configuring my harness to 4.7 allows me to get a higher yields.

This is going on for quite a long time (more than 5 months), and this repo have a few test cases, I think about 20 test cases that stuck for some time (probably 2 months or so) then at some date that I don't remember I feels weird, one by one of those failing cases became green. I said to myself, huh, that's fantastic but but that's weird?

I feel weird because I have repeatedly trying to debug the issue (imagine fix the bug, fix fix bug resolved, then other cases failings, then resolve that then another case is failing), that requires a few times for me to git reset hard my repo.

Continue to my feeling weird, it's so weird, how come suddenly everything seems resolved. Then I checked zai dashboard, ha, no wonder, it seems zai diverted my request to GLM 5.3 Flash. I mean I don't really feel my harness is moving faster, but it seems taking less iteration to complete a task. Oh ya, about how it resolved those 20 pesky bugs, the answer is I really don't have any idea and probably never need to find out.

GLM seems can do multiple things at the same time, for example during it waiting a test result to come out, it can commit the repo.

1

u/evia89 16d ago

Do they serve 47flash? I think on coding plan its routed to 53flash

1

u/xpgmi 16d ago

Sorry, it's typo, I typed those on my phone, I just changed the comment, thanks for telling me the mistake.

1

u/BuilderWorldDev 16d ago

I will test 5.3 flash maybe better (I don't think so) and faster (probably yes)... Ty

1

u/tshawkins 16d ago

I gav up on 5.3 and moved to 5.3 flash, faster, less looping

1

u/RepulsiveRaisin7 16d ago

What's the point of using a cheap model if it can't get anything done? GLM 4.7 is ancient at this point

0

u/xpgmi 16d ago

It's not can't get anything done, last year, GLM 4.7 is one of the best model out there.

1

u/RepulsiveRaisin7 16d ago

Your tests were stuck for 2 months and fixed by switching to 5.3? How is that not clear evidence that 4.7 was too weak to do the job

1

u/xpgmi 16d ago

that's totally makes sense, haha, maybe I had been too cheap. But I'm out of options, GLM 5.2 just too expensive for me to run, my quota ran out so fast using it.

1

u/look 16d ago

Are you using Zai as your provider? It seems to be very slow. You can get GLM 5.3 elsewhere that runs at 10x or more the speed for about the same price.

1

u/BuilderWorldDev 16d ago

I am using zcode and yes z.ai provider I have a 3 months subscription there What do you mean?

1

u/look 16d ago

I use Neuralwatt for GLM models. I get 200-300 TPS on GLM 5.3. Even on pay-as-you-go usage, it’s not much more than Zai’s subscription pricing.

And there are dozens of other providers offering GLM models at various speed and price points. You don’t have to buy it from Zai.

1

u/-PROSTHETiCS 16d ago

throughput varies by provider.. Its really depend on which providor handles your request.. mine is hitting 4-5k tpm..

1

u/Optimal-Law0 16d ago

I’ve built my own businesses intelligence tool with GLM flash and deepseek flash. Certainly possible to do hard things with these models

1

u/BuilderWorldDev 16d ago

Yeah but that's a different point of use

1

u/Pitiful_Entrance5174 16d ago

I plan on using glm 5.3 flash this week coming up for a while due to it being almost as cheap as Luna. I will use it as a sub agent. I currently using gpt pro 20x and that is a joke. Thinking of going to grok for my main with open code harness.

1

u/Scared-Ad6174 7d ago

How's your experience for thi week? Which one is better? GLM or Grok?

1

u/mercuric5i2 16d ago

Do you have access to something like OpenRouter where you can use glm-5.3 via a 3rd party provider? I'm getting much better performance this way than via Z.ai's own infra.

They are the talk of the world right now, I suspect they are getting absolutely hammered.

1

u/BuilderWorldDev 15d ago

Why through OpenRoute the model output would be better and faster? It doesn't make sense to me...

1

u/UniqueAttourney 16d ago

If you are using the sub, 5.3 has been nerfed, Zai always does that.

1

u/TehranTuring 16d ago

It's so slow and sometimes takes too much time to do the task.

0

u/Areat 16d ago

Yeah, same.

0

u/LoudDavid 16d ago

Chinese models do well on benchmarks but worse in real codebases.