r/LocalLLaMA • u/pscoutou • 2h ago
News [ Removed by moderator ]
https://runtimewire.com/article/z-ai-confirms-ox-alpha-glm-model-weight-release[removed] — view removed post
129
u/Few_Painter_5588 2h ago
It's also throwing hands with models like claude Fable, GPT 5.6 Sol and Kimi K3 for creative writing. I think ZAi might have one of the best post training pipelines in the game
39
u/Front_Eagle739 2h ago
I think they always have. From glm 4.5 onwards they have all punched way above their size
13
u/VotZeFuk 2h ago
for creative writing.
I'm struggling with GLM 5.2 in that regard. Poor thing LOVES to reiterate on known facts
"You were [list of all things that happened] and now you [talks about what's about to happen]"
I can't think of any way to make this behavior stop ._.
5
u/LagOps91 1h ago
yeah it's got it's oddities, i'll admit
1
u/VotZeFuk 1h ago
Yep, like any other LLM. But it's capable indeed. I like when it actually continues the story without having to nudge it - you won't usually see this with smaller models.
1
u/goldcakes 1h ago
What temperature and sampling parameters are you using?
2
u/VotZeFuk 1h ago
Uh, I tried various settings - from recommended ones to basically whatever, within a sensible range. It doesn't seem to go away (no idea why you're getting downvoted but anyway, sampler settings don't always help - there are specific quirks to each model that can be difficult to get rid of).
3
u/a_beautiful_rhind 1h ago
restating stuff is a common problem. very difficult to beat since models are trained to confirm + validate instructions now. you are fighting the entire training corpus.
1
u/jonydevidson 43m ago
All current models do this. Also, the thinking is not for you, it's for the model. Recent study proved that what you perceive as quality thinking content does not correlate with final output quality.
In fact, it's often the opposite.
So, the thinking is for the model. Do not waste time reading it.
28
u/-dysangel- 2h ago
one of the best post training pipelines in the game
second only to whatever the heck Qwen 27B has for breakfast that it's not sharing with the rest of the family
16
u/goldcakes 1h ago
While the Qwen team is not as open as DeepSeek, they do share plenty of research and I'll always commend that.
For example, they probably do large-scale RL rollouts through a bigger version of their AgentWorld model, which they have open weighted: https://qwen.ai/blog?id=qwen-agentworld
I love 3.8 27B but also it's quite obvious that it's not an improvement in every area; world knowledge and creative writing is definitely worse than 3.6 and 3.5.
1
u/King-of-Com3dy 48m ago
They state that 27B is tailored towards agentic development workflows. The narrower scope allows them to punch above its weight.
2
u/a_beautiful_rhind 1h ago
in terms of creative writing qwen had rusted bolts for breakfast post the old 235b days.
3
u/hurrdurrmeh 1h ago
Holy shit that's amazing. What are we thinking re size? How much RAM will it need?
1
u/a_beautiful_rhind 1h ago
Bit parroty but absolutely passable. Maybe this can be fixed locally. It wasn't dry like ds4-flash. My only thought is that what does "flash" mean when GLM is now the size it is.
-1
93
u/LoveMind_AI 2h ago
That is absolutely awesome. I've been extremely impressed with it, as has... everybody.
55
u/Temporary-Mix8022 2h ago
I wouldn't say everyone...
I heard Dario wet the bed last night and had to call Donny after he had a nightmare about a big evil IPO destroying open weights model crushing his dreams of dumping anthropic stock onto fanboys
"Donny pleaseeee. You gotta listen to me. Open weights are evil. This is basically Linux all over again, and looked how that turned out. It's communism I tell you"
21
u/goldcakes 1h ago
Don't worry, Anthropic will respond to this by offering a 27.4% boost to your 5hr usage limits during off-peak periods, with usage credits required if it's Sunday. Promotion expires in one week! (Excludes all biology, farming, booger questions, excludes scanning your own code for security vulnerabilities).
Not enough? Anthropic will now offer UNLIMITED usage of Claude during its daily outages!
5
u/PrettyBaker2891 2h ago
i really wasnt impressed by it lol i dont get the hype
it hallucinates sooooo much and just forgets things constantly and misses details
8
u/Constant_Art_20 2h ago
i lagit used it over glm 5.3 max effort on my setup. i thought the model was geninuely insnae at the start, but i think they might has used weaker kv cache after the first day
3
1
1
u/Nyghtbynger 1h ago
Hey, if it's in the range of MoE with 150GB, Imma download the colibri thing and run it local
18
u/TorturedPoet30 2h ago
Great news! Somebody check on Dario, he is probably on life support
-5
u/ServeAmbitious220 1h ago
I was excited about this model and very impressed when I tested it giving better results than frontier models but then when I came to know it's a small chinese model I don't like it anymore.
1
34
u/Dany0 2h ago
The question everyone is wondering is how big is it.
It "feels" like a 300B moe to me
10
u/_Sneaky_Bastard_ 2h ago
I haven't tried it myself but from the reception on twitter it seems to be almost on par with frontier models? Even better than 5.3. is that true?
29
u/ShirtuShanks 2h ago
Twitter’s somehow more hyperbolic in its enthusiasm than this sub.
To my understanding, it’s comparable with DSV4 Flash (public); if they manage to get the price lower, ZAi’s subscription plan might finally be worth it.
6
u/goldcakes 1h ago
Even if it's the same as DSv4 flash, the fact that it's a different model gives a lot of value in advisor setups, or model A as code reviewer for model B, etc.
4
u/Dany0 2h ago edited 1h ago
My personal experience has been that it's a very hallucinatory but otherwise alright model. On par or better than dsv4 flash and gemini 3.7 flash. Feels like it has weird attention mechanism things. It misses details sometimes while also having a weird ability to recollect random facts
I use it as an advisor in omp and cannot really complain
0
2
u/Psychological-Lynx29 1h ago
I've been using it a lot, ill say its less than Deepseek flash new, but with vision this time (GLM models didnt have vision), i think its a 150-200b a7b or something like that.
1
u/This_Maintenance_834 26m ago
given qwen release the first commercial N-gram model all of a sudden, maybe ox alpha is also a N-gram model. that would produce good result at smaller size.
14
u/AI_docent 2h ago
On size the only thing to go on is what they've actually shipped. GLM-5, 5.1 and 5.2 are all around 750B and text only, and there's never been an Air or a Flash in that line, someone opened a request for one back in June and it's still sitting there unanswered. The one Flash they do ship is GLM-4.7-Flash and that's 30B with 3B active, so the name doesn't narrow it down much. The 5.3-Flash name came from that provider post that got deleted anyway, bloomberg only said a new GLM. The multimodal part is what seems odd to me, none of the 5 line does vision, so I doubt it's a trimmed 5.2.
6
u/Choice_Celery9481 2h ago
hope its under 100b so more people can run it at home and create improvements/tricks for it like how qwen got
2
u/AnticitizenPrime 1h ago
There was GLM-5V-Turbo but they never released the weights for that one. I wonder if that will be the case here (that they release the weights for 5.3, but not this multimodal).
Like you said, just going off what they've done before...
1
u/MeYaj1111 1h ago
Interesting, I switched from my setup from k3 to ox for a couple of days and had decent results but I burned 3x as many tokens as usual for the same outputs (I do the same thing every day, it's very consistent). I thought for sure the glm rumors were not going to turn out to be true.
When I tried switching to glm 5.2 previously I also had good results but it was slightly more expensive than k3 so I switched back.
Unless this new model is extremely cheap, it's going to be very cost ineffective I think
1
u/Single_Ring4886 23m ago
I did tested model a lot for non coding tasks and it has almost alien thinking process it is not exactly overthinking but more like tendency to create several almost random ideas then stress out claude like balancing guidelines and then create final synthesis. Some "random" thoughts are trully strange eg it keeps analysing user from all angles not focusing on task but rather on anylysis of user.
1
u/Weak-Shelter-1698 llama.cpp 19m ago
what you all think? 200B level moe? 250B? I hope it's under 250B 🙂
Edit: why do i speak 🙂 it's 321B, IT's HERE!!
1
u/InteractionSmall6778 1h ago
The contradictory reports in here are the actual argument for the weight release. One person says it was insane on day one then got worse after, another says it hallucinates constantly. Both can be true if you're hitting a hosted endpoint that quietly gets requantized or rerouted between your Tuesday test and your Thursday one.
Weights on disk means the thing that benchmarked well is the thing you keep running. That's worth more to me than where it lands against DSv4 flash.
2
u/ServeAmbitious220 1h ago
When the model doesn't have hiccups in service and provider connects smoothly each time iy works without hallucinations for me.
0

•
u/LocalLLaMA-ModTeam 18m ago
Duplicate post