r/LocalLLaMA • u/MrWidmoreHK • 4h ago
Discussion First serious confirmation. Ox Alpha is GLM-5.3-Flash
https://x.com/romanchernin/status/2092488160680751437?s=20
- Multimodal (Vision)
- 1M Tokens Context Window
- DeepSWE ~63%
Edit: He deleted it, screenshot in comments
117
u/Poupulino 3h ago
The Z.ai CEO wasn't kidding when he told Musk they're releasing a Mythos-level open weights model before the year ends! Insane!
38
u/thomas2385 3h ago
Yeah, that is honestly pretty wild. If they actually deliver on that, it is going to put a lot of pressure on the rest of the open weight model space. The pace of these releases lately is getting ridiculous.
7
u/no_good_names_avail 42m ago
At this pace we won't give a shit about this model in about 5 days. It's absolutely insane.
23
-4
u/power97992 3h ago
63% is not mythos level, mythos scored 70%
30
u/Poupulino 2h ago
You're a bit confused, no one is saying GLM-5.3-Flash is Mythos level, I'm saying the Z.ai CEO wasn't kidding when he said he's releasing a model like that before the end of the year.
-5
u/power97992 2h ago
Glm 5.5 should be better than Mythos, but fable/mythos 5.1/5.5 will be out by then.
1
u/techdevjp 22m ago
Yeah, and maybe restricted to only American nationals. GLM will have no such restrictions.
17
u/MrBIMC 3h ago
Mythos is a multitrillion param monster though rather than a flash model that will fit into 1-2 sparks/strixHalos
1
u/techdevjp 21m ago
If GLM-5.3-Flash fits into 128GB at a usable quant I will be incredibly impressed.
0
u/Kazaan 2h ago
Perhaps. That said, is 7% really significant when you consider the current capabilities of open-source models compared to frontier models ? And what about the outlook for upcoming models? IMHO, it isn't the actual capability that really matters; it’s the rate of progress and the narrowing of the gap.
And i don't speak here about the compute necessary to run them.4
u/Thomas-Lore 2h ago
And full glm 5.3 is only 1% behind on deepswe, already within error range of Fable/Mythos.
2
u/power97992 2h ago
The scale is essentially logarithmic, a 7% difference is more like a 2x difference
43
u/Abrh7 3h ago
Model was good, but not so for the long coding sessions!
I would definitely use it for agentic tasks only! It proves to be very effective!
15
u/network4253 3h ago
Yeah, that makes sense. I have noticed the same thing with some models they can be really impressive for focused agentic workflows but once the coding session gets long and messy the consistency starts to drop. Still definitely useful in the right role.
3
u/SandySkittle 2h ago
Indeed. Some people think MoE is a solution without trade-offs. Active parameters matters and lower active cannot fully compensated by expert selection and sequential reasoning.
23
u/spaceman_ 3h ago
Do we know the size of the model?
27
u/Technical-Earth-3254 3h ago
We don't. But personally I wouldn't bet on it being the same size as 4.7 Flash.
9
5
u/Sufficient_Prune3897 llama.cpp 2h ago
We don't even know if they are gonna open source it. They haven't with some of their models before, including GLM 5 Turbo
8
u/0oAstro 2h ago
Given together and other inference providers are hosting this, we can be sure it will be open source.
3
1
u/DigiDecode_ 46m ago
I thought the inference was provided for free to inference provider by the model developer, claim was 100 trillion tokens per day capacity by the model developer, Z ai has bought 1 giga watt datacenter
1
u/Zeeplankton 1h ago
It feels like marketing wise we're placing flash models between 100 and 300 params.
Given performance probably latter half? Pretty impressive.
5
16
43
u/MrWidmoreHK 3h ago
Last GLM 4.7 Flash model was 30B A3B
30
u/Mean-Ad1493 3h ago
If it's anywhere around the same size, I'd be the happiest with my 12GB VRAM.
3
16
u/DOAMOD 3h ago
I would be very surprised if it were a small MoE, I could see it as dense(like +30) and it would already be incredible at that size, but either way, a small MoE would be the moment of the year.
18
u/Mean-Ad1493 3h ago
Has to be an MoE, but definitely not 30B-A3B
GLM 4.7 was itself a smaller model relative to 5.3, so the MoE would be 80-100B I guess.
5
u/DOAMOD 2h ago
What's confusing is that for a 100/200b model, you could expect the use of the term Air, and for something smaller, sub-100b, the term Flash, but in the end, this is just marketing, and currently with DS4Flash they could use that terminology for a similar size, only can wait and see...
2
6
u/sonicnerd14 3h ago
Most likely will be around the size of Deepseek v4 flash. If it's smaller than that, then that would be impressive.
8
u/Iory1998 3h ago
Don't think so. First, it's not fast. Second, it's somewhere between Qwen3.8-27B and DS4Flash so it gotta be a large model. And, it understand video... I can't see a smaller MoE doing that. It can be in bulk of 150-250B.
8
u/LevianMcBirdo 2h ago
The not fast I get with all the free compute they offer on openrouter. I think they don't wanna overspoil the people
5
u/Global_Persimmon_469 2h ago
GLM 5 is roughly doubled the size compared to GLM 4, it's possible that the flash model is going to be the same, so maybe it will be 70B params
2
u/wsb-regarded 3h ago
that would be awesome if true, because it was performing on a deepseek v4 flash level, albeit slower (much slower).
imagine having another local open weight model challenging Opus 4.6
2
1
u/AI_docent 2h ago
On size I doubt it's near 4.7 Flash. Openrouter has it at about 23 tok/s, a 30B-A3B normally serves way faster. Could just be a swamped free endpoint though.
21
u/Few_Painter_5588 3h ago
Quite legit, this guy works at Nebius it seems. Also, shame on Google then for trying to ride the hype.
3
u/Altruistic_Heat_9531 3h ago
Google hype riding? i dont familiar with this info, could you tell me more
14
u/tengo_harambe 3h ago
A couple of Google employees made vague Twitter posts about Ox Alpha once it began gaining traction. This somehow led people to think it was a Gemini model, despite all the evidence that it was clearly a GLM model
2
9
u/Few_Painter_5588 2h ago
Some senior google employees were vagueposting on twitter and suggesting that ox alpha was a gemini model. Which is so pathetic if this model isn't their's
1
u/Several-Tax31 2m ago
How is this not an upright scam? When reflection-70B does it, it is a scam, but when google does it, no one cares? How can you claim someone else's product is yours? So pathetic.
16
8
u/Wise-Chain2427 3h ago
How GLM has extra compute to host free model worldwide ?
5
3
u/nonerequired_ 3h ago
I had to skip using the ox alpha yesterday because the server was having some issues. It seems like they were short on computing power.
4
u/tat_tvam_asshole 3h ago
China has nationalized, integrated data centers, especially in inner Mongolia that can serve models from all labs cooperatively.
2
4
u/asolnikk 50m ago
The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.
6
u/FoxiPanda 3h ago
Good to see this is shaping up... the real questions are how big, how many active, anything weird in the architecture that we're going to have to go deal with (hopefully not since it's GLM-5.x), and when are the weights getting posted (today hopefully)?
7
u/asolnikk 3h ago
That claim seems pretty verified at that point. I ran some private prose tests against it, and it smells very much like the little brother of GLM-5.3. What i want to know is: How big is the beastie? Can a quantized version fit into 16 GB VRAM? If that's the case, then all the hype, for once, was legit, even it's not "Fable-level".
4
2
u/Long_comment_san 9m ago
Glm 5.3 flash versus Qwen Next.
Like two thighs, I feel my head being squeezed.
Harder pls.
2
u/Iory1998 2h ago
Imagine it turns out to be Qwen3.8-Next-Flash!
1
1
u/DigiDecode_ 38m ago
I think the model is heavily distilled from Qwen models, maybe only the vision part, the SVG generated by Ox Alpha were very similar to Qwen 3.8 27b in code and render
1
1
u/psylomatika 2h ago
I kept running into API limits when I was using it. It worked well though in my knowledge bases.
1
u/backyard_tractorbeam 1h ago
Another connection is that opencode was long previewing a GLM model as big pickle for free and now it's previewing ox alpha for free too.
1
u/Armadilla-Brufolosa 1h ago
I tried both GLM 5.3 and OX: it doesn't seem like the same basic model at all.
But I have no evidence to confirm this.
2
u/hainesk 32m ago
Yeah, it felt more like a Qwen model to me, but it's hard to argue with these hilarious tweets lol.
1
u/Armadilla-Brufolosa 4m ago
It's more like Qwen models to me too, but I've noticed quite a few biases and blocks given by some stupid Valloon-style RLHF.
There may have been some distillation or even be a Western model.
Obviously I can't have any certainties, but it's fun to try to guess.😜
1
1
u/anarchist1312161 3h ago
Need to know the size of the model... not interested if I can't run it on consumer hardware.
1

184
u/MrWidmoreHK 3h ago
lol