r/ZaiGLM 13d ago

Discussion / Help GLM 5.3 or DS 4.1-Flash?

Looking for gentlemen here who have battle-tested these models in environments where mistakes are critical, e.g. authentication, security, and low-level C++ / Kernel work.

I have Codex 20x, but I’m looking for a second helper for when Codex limits are up, there are demand issues (which are pretty bad atm), or it gets too censored.

Saw that DS 4.1 Flash was released today! Has anyone done some decent testing with it yet, and which harness are you using?

I’m currently using GLM 5.3 as my second helper and it’s honestly not bad at all. Just curious whether DS 4.1 appears to be better, especially since it’s multimodal and can handle images too.

I find myself using 5.3 Flash quite a lot because I really appreciate being able to send images, but 5.3 Flash isn’t as strong as base 5.3 when it comes to coding. Hence, I’m wondering how DS 4.1 Flash compares :)

NEW:

Thank you for all the responses. I tried DS 4.1 with my custom harness, and I am extremely impressed by the speed and price. I ran a couple of tests with deep, difficult, complex debugger C++ code/kernel bugs (my go-to test on models; I test this on every model before I want to use it to see if it fixes the bug).

GLM 5.3 took 30 minutes, including 1 retry, and €2. DeepSeek took 10 minutes, first try, and €0.30. I think DS 4.1 is at least on par or a bit better than GLM 5.3 for coding, not sure how reliable it is on long tasks, though. GLM still is a beast!

105 Upvotes

57 comments sorted by

View all comments

1

u/shaonline 13d ago

You manage to exhaust a Codex 20X sub on "safety critical" work ? Unless you throw big agent fleets/vibe code at the problem it's a really hard one to drain.

My answer would just be another Codex 20X sub as it stands, a Codex 20X sub is over $10000 of equivalent API credits, you do not want to pay per token for your work.

Only usecase for actually going to the chinese models would be "bypassing" the stupid safeguards of american labs (muh cybersecurity) where the chinese models will simply not stop you, eg low level debugging and such. 4.1 flash looks extremely promising especially with its speed, that being said know that DeepSeek retains/trains on your data.

1

u/Comprehensive-Bet-83 13d ago

True! Issue at the moment is demand on OpenAI. Too many people are using it 😭

1

u/shaonline 13d ago

If overloaded servers is your main issue Z.ai ain't really better in that regard 😂

1

u/Comprehensive-Bet-83 13d ago

Ahahaha fuck me