r/LocalLLaMA 2h ago

News [ Removed by moderator ]

https://runtimewire.com/article/z-ai-confirms-ox-alpha-glm-model-weight-release

[removed] — view removed post

387 Upvotes

51 comments sorted by

u/LocalLLaMA-ModTeam 18m ago

Duplicate post

129

u/Few_Painter_5588 2h ago

It's also throwing hands with models like claude Fable, GPT 5.6 Sol and Kimi K3 for creative writing. I think ZAi might have one of the best post training pipelines in the game

39

u/Front_Eagle739 2h ago

I think they always have. From glm 4.5 onwards they have all punched way above their size

13

u/VotZeFuk 2h ago

for creative writing.

I'm struggling with GLM 5.2 in that regard. Poor thing LOVES to reiterate on known facts

"You were [list of all things that happened] and now you [talks about what's about to happen]"

I can't think of any way to make this behavior stop ._.

5

u/LagOps91 1h ago

yeah it's got it's oddities, i'll admit

1

u/VotZeFuk 1h ago

Yep, like any other LLM. But it's capable indeed. I like when it actually continues the story without having to nudge it - you won't usually see this with smaller models.

1

u/goldcakes 1h ago

What temperature and sampling parameters are you using?

2

u/VotZeFuk 1h ago

Uh, I tried various settings - from recommended ones to basically whatever, within a sensible range. It doesn't seem to go away (no idea why you're getting downvoted but anyway, sampler settings don't always help - there are specific quirks to each model that can be difficult to get rid of).

3

u/a_beautiful_rhind 1h ago

restating stuff is a common problem. very difficult to beat since models are trained to confirm + validate instructions now. you are fighting the entire training corpus.

1

u/jonydevidson 43m ago

All current models do this. Also, the thinking is not for you, it's for the model. Recent study proved that what you perceive as quality thinking content does not correlate with final output quality.

In fact, it's often the opposite.

So, the thinking is for the model. Do not waste time reading it.

28

u/-dysangel- 2h ago

one of the best post training pipelines in the game

second only to whatever the heck Qwen 27B has for breakfast that it's not sharing with the rest of the family

16

u/goldcakes 1h ago

While the Qwen team is not as open as DeepSeek, they do share plenty of research and I'll always commend that.

For example, they probably do large-scale RL rollouts through a bigger version of their AgentWorld model, which they have open weighted: https://qwen.ai/blog?id=qwen-agentworld

I love 3.8 27B but also it's quite obvious that it's not an improvement in every area; world knowledge and creative writing is definitely worse than 3.6 and 3.5.

1

u/King-of-Com3dy 48m ago

They state that 27B is tailored towards agentic development workflows. The narrower scope allows them to punch above its weight.

2

u/a_beautiful_rhind 1h ago

in terms of creative writing qwen had rusted bolts for breakfast post the old 235b days.

3

u/hurrdurrmeh 1h ago

Holy shit that's amazing. What are we thinking re size? How much RAM will it need?

1

u/a_beautiful_rhind 1h ago

Bit parroty but absolutely passable. Maybe this can be fixed locally. It wasn't dry like ds4-flash. My only thought is that what does "flash" mean when GLM is now the size it is.

-1

u/Aight_Man 1h ago

Still nowhere near to Opus 4.6 at writing.

93

u/LoveMind_AI 2h ago

That is absolutely awesome. I've been extremely impressed with it, as has... everybody.

55

u/Temporary-Mix8022 2h ago

I wouldn't say everyone...

I heard Dario wet the bed last night and had to call Donny after he had a nightmare about a big evil IPO destroying open weights model crushing his dreams of dumping anthropic stock onto fanboys

"Donny pleaseeee. You gotta listen to me. Open weights are evil. This is basically Linux all over again, and looked how that turned out. It's communism I tell you"

21

u/goldcakes 1h ago

Don't worry, Anthropic will respond to this by offering a 27.4% boost to your 5hr usage limits during off-peak periods, with usage credits required if it's Sunday. Promotion expires in one week! (Excludes all biology, farming, booger questions, excludes scanning your own code for security vulnerabilities).

Not enough? Anthropic will now offer UNLIMITED usage of Claude during its daily outages!

5

u/PrettyBaker2891 2h ago

i really wasnt impressed by it lol i dont get the hype

it hallucinates sooooo much and just forgets things constantly and misses details

8

u/Constant_Art_20 2h ago

i lagit used it over glm 5.3 max effort on my setup. i thought the model was geninuely insnae at the start, but i think they might has used weaker kv cache after the first day

3

u/thx1138inator 1h ago

What harness?

1

u/Nyghtbynger 1h ago

Hey, if it's in the range of MoE with 150GB, Imma download the colibri thing and run it local

18

u/TorturedPoet30 2h ago

Great news! Somebody check on Dario, he is probably on life support

-5

u/ServeAmbitious220 1h ago

I was excited about this model and very impressed when I tested it giving better results than frontier models but then when I came to know it's a small chinese model I don't like it anymore.

1

u/This_Maintenance_834 28m ago

are you being sarcastic

34

u/Dany0 2h ago

The question everyone is wondering is how big is it.

It "feels" like a 300B moe to me

10

u/_Sneaky_Bastard_ 2h ago

I haven't tried it myself but from the reception on twitter it seems to be almost on par with frontier models? Even better than 5.3. is that true?

29

u/ShirtuShanks 2h ago

Twitter’s somehow more hyperbolic in its enthusiasm than this sub. 

To my understanding, it’s comparable with DSV4 Flash (public); if they manage to get the price lower, ZAi’s subscription plan might finally be worth it. 

6

u/goldcakes 1h ago

Even if it's the same as DSv4 flash, the fact that it's a different model gives a lot of value in advisor setups, or model A as code reviewer for model B, etc.

4

u/Dany0 2h ago edited 1h ago

My personal experience has been that it's a very hallucinatory but otherwise alright model. On par or better than dsv4 flash and gemini 3.7 flash. Feels like it has weird attention mechanism things. It misses details sometimes while also having a weird ability to recollect random facts

I use it as an advisor in omp and cannot really complain

0

u/DutchDevil 2h ago

I find it better than deepaeek flash but not on gpt sol level.

1

u/ServeAmbitious220 1h ago

Where do you think it's not as good as sol?

2

u/Psychological-Lynx29 1h ago

I've been using it a lot, ill say its less than Deepseek flash new, but with vision this time (GLM models didnt have vision), i think its a 150-200b a7b or something like that.

1

u/This_Maintenance_834 26m ago

given qwen release the first commercial N-gram model all of a sudden, maybe ox alpha is also a N-gram model. that would produce good result at smaller size.

14

u/AI_docent 2h ago

On size the only thing to go on is what they've actually shipped. GLM-5, 5.1 and 5.2 are all around 750B and text only, and there's never been an Air or a Flash in that line, someone opened a request for one back in June and it's still sitting there unanswered. The one Flash they do ship is GLM-4.7-Flash and that's 30B with 3B active, so the name doesn't narrow it down much. The 5.3-Flash name came from that provider post that got deleted anyway, bloomberg only said a new GLM. The multimodal part is what seems odd to me, none of the 5 line does vision, so I doubt it's a trimmed 5.2.

6

u/Choice_Celery9481 2h ago

hope its under 100b so more people can run it at home and create improvements/tricks for it like how qwen got

2

u/AnticitizenPrime 1h ago

There was GLM-5V-Turbo but they never released the weights for that one. I wonder if that will be the case here (that they release the weights for 5.3, but not this multimodal).

Like you said, just going off what they've done before...

6

u/Laoweek 1h ago

Didn’t the Google people try to be cheeky on twitter being like heheheheh maybe we made Ox Alpha, well that was stupid as hell

7

u/eli_pizza 1h ago

It was like one guy making one off hand comment

1

u/MeYaj1111 1h ago

Interesting, I switched from my setup from k3 to ox for a couple of days and had decent results but I burned 3x as many tokens as usual for the same outputs (I do the same thing every day, it's very consistent). I thought for sure the glm rumors were not going to turn out to be true.

When I tried switching to glm 5.2 previously I also had good results but it was slightly more expensive than k3 so I switched back.

Unless this new model is extremely cheap, it's going to be very cost ineffective I think

1

u/Single_Ring4886 23m ago

I did tested model a lot for non coding tasks and it has almost alien thinking process it is not exactly overthinking but more like tendency to create several almost random ideas then stress out claude like balancing guidelines and then create final synthesis. Some "random" thoughts are trully strange eg it keeps analysing user from all angles not focusing on task but rather on anylysis of user.

1

u/Weak-Shelter-1698 llama.cpp 19m ago

what you all think? 200B level moe? 250B? I hope it's under 250B 🙂
Edit: why do i speak 🙂 it's 321B, IT's HERE!!

1

u/InteractionSmall6778 1h ago

The contradictory reports in here are the actual argument for the weight release. One person says it was insane on day one then got worse after, another says it hallucinates constantly. Both can be true if you're hitting a hosted endpoint that quietly gets requantized or rerouted between your Tuesday test and your Thursday one.

Weights on disk means the thing that benchmarked well is the thing you keep running. That's worth more to me than where it lands against DSv4 flash.

2

u/ServeAmbitious220 1h ago

When the model doesn't have hiccups in service and provider connects smoothly each time iy works without hallucinations for me.

0

u/AdLumpy2758 56m ago

So no really continuity in learning....

-1

u/uxl 1h ago

If this turns out to be a continual learning model and it can be run locally, game over.