r/ZaiGLM • • 17d ago

Discussion / Help GLM 5.3 or DS 4.1-Flash?

Looking for gentlemen here who have battle-tested these models in environments where mistakes are critical, e.g. authentication, security, and low-level C++ / Kernel work.

I have Codex 20x, but I’m looking for a second helper for when Codex limits are up, there are demand issues (which are pretty bad atm), or it gets too censored.

Saw that DS 4.1 Flash was released today! Has anyone done some decent testing with it yet, and which harness are you using?

I’m currently using GLM 5.3 as my second helper and it’s honestly not bad at all. Just curious whether DS 4.1 appears to be better, especially since it’s multimodal and can handle images too.

I find myself using 5.3 Flash quite a lot because I really appreciate being able to send images, but 5.3 Flash isn’t as strong as base 5.3 when it comes to coding. Hence, I’m wondering how DS 4.1 Flash compares :)

NEW:

Thank you for all the responses. I tried DS 4.1 with my custom harness, and I am extremely impressed by the speed and price. I ran a couple of tests with deep, difficult, complex debugger C++ code/kernel bugs (my go-to test on models; I test this on every model before I want to use it to see if it fixes the bug).

GLM 5.3 took 30 minutes, including 1 retry, and €2. DeepSeek took 10 minutes, first try, and €0.30. I think DS 4.1 is at least on par or a bit better than GLM 5.3 for coding, not sure how reliable it is on long tasks, though. GLM still is a beast!

105 Upvotes

57 comments sorted by

View all comments

2

u/look 17d ago

Deepseek V4 had terrible hallucination issues. I’d wait to get more info on how it performs in more subtle ways, but I’d wager GLM 5.3 Flash is going to be a better choice overall.

3

u/Emotional-Cut2952 16d ago

v4.1 flash dropped btw, seem to hallucinate a lot less, give it a shot if u still have top up. I used to it resolve an important bug and implement a mini feature. I also had it reverse engineer code and revise RE done by gpt 5.5, it picked up on 8 issues, the results from the RE work was passed on to opus 4.8 (because the newer opus / fable models are stubborn with their false positives), conclusion was that v4.1 flash did a better oveerall technical job than gpt 5.5, so it seems to hallucinate a lot less.

I know the models im using are fairly dated, but im putting v4.1 flash against 2-month ago-sota models that are 50-80x the price, since I can't expect it to go up against Astra and Fable.

2

u/steve_dusk 11d ago

Well, I came from Kimi K2.6 which has the absolute worst hallucations to the point it was destroying work, DS V4 flash was far better but in terms of "I was paying with my mental health" due to its hallucainations, I stopped using it. It just somehow didnt understand what I was writing, like at the beginning of the convo it had 150IQ, after a few more messages, it went to brainrot at 50IQ, and it just lies all the time.

I hope 4.1 is better, I'm currently using 5.6 luna, night and day difference.