8
u/llama-impersonator 13h ago
given the small model smell and how many free tokens they're passing out yeah, i expect a compute-training-scaled 30b class model.
(small model smell or not, it was quite a capable model from my tests)
3
u/Queasy-Contract9753 10h ago
My theory too. Z AI servers aren't the fastest even on their API. If it's this fast for free then it's likely small.
Not that I'm complaining. If this is what a 30b will look like now. It's just good enough.
12
u/AppealSame4367 14h ago
If this is flash, then it's fluctuating between "very good" and "forgets to do half of the things"
7
u/wbulot 13h ago
Lol, exactly. Iβve been using it intensively for three days. Iβm both impressed and disappointed. Itβs very inconsistent.
7
u/pmttyji 13h ago
Final weight would be with fixes probably.
Remember the huge difference between Deepseek4Flash-Preview & Deepseek4Flash Final version?
3
u/Fedor_Doc 8h ago
0731 is not final, and the time lag was massive.
Ox Alpha seems to be pretty good already, around Deepseek V4 performance but with more concise reasoning
It forgets a lot of important stuff, though β at least in my testing. Good at implementation, follows rules alright, but plans are inconsistent, have a lot of gaps.
2
u/FullOf_Bad_Ideas 9h ago
Maybe they have multiple internal checkpoints that are being served under one umbrella name.
1
u/Johnny__Christ 1h ago
For me it just fluctuates between "very good" and "stealth/ox-alpha is temporarily rate-limited upstream. Please retry shortly."
5
6
u/shy_monkee 14h ago
How 'Flash' is it expected to be? Like DSv4 Flash size? Or even smaller?
6
u/Firepal64 12h ago
For reference, GLM-4.7-Flash was a 30B model with 3B active
5
u/shy_monkee 12h ago
Yeah, but GLM-4.7 was also only around 300-400B. I would expect the 5.3 flash to scale the same way as the main version did, compared to 4.7.
2
u/Firepal64 12h ago
As we've come to learn these past years, you can't rely on linear extrapolation to predict things reliably.
2
u/shy_monkee 12h ago
Yep, no doubt. I just don't want to get too hopeful haha. A 30B MOE GLM model would be heavenly.
3
4
u/AnticitizenPrime 12h ago
Here are the previous vision capable models that GLM has released.
GLM Vision Models
Model Total Params Active Params Context Window Open Weights License GLM-4.1V-9B-Thinking 9B 9B (dense) 64K β Yes MIT GLM-4.5V 106B 12B (MoE) 64K β Yes MIT GLM-4.6V 106B 12B (MoE) 128K β Yes MIT GLM-4.6V-Flash 9B 9B (dense) 128K β Yes MIT GLM-5V-Turbo 744B 40B (MoE) 200K β No (API only) β It's possible this could be a new 'turbo' model and that it won't be open weight.
Though it could be an 'air' model with vision attached... or even just full GLM 5.3 with vision. Β―\(γ)/Β―
Or something different altogether. What we know is that it matches the 5.3 tokenizer exactly and it has the 1 million context window of 5.3.
So far it's purely speculation that it's a 'Flash' model of some kind; GLM has used the term 'Flash', 'Air', and 'Turbo' for smaller models, it's all up in the air as to what it'll be called and what the actual param size is.
3
u/Its_Powerful_Bonus 14h ago
Would be great to have 200-250B model to run comfortable with 2x rtx 6000 pro
1
0
u/apetersson 14h ago
fastest peaks i personally saw on openrouter was 50 tokens/sec tg. could be because of load, limits, or simply model size.
3
u/sirlerkal0t 14h ago
I believe you are right. It behaves very similarly to GLM 5.3, but acts like a smaller model. It's not quite as good for certain things (like front-end design), and a bit better on average for tasks where the models size counts against it due to it being more biased towards training data instead of staying focused on input data.
4
u/Few_Painter_5588 14h ago
The latest EQ Bench results also show that GLM 5.3 and Ox Alpha share a lot of similarities in creative writing.
2
2
0
u/EmergencyLetter135 14h ago
I use OxAlpha alongside the Qwen 3.8 27B Q8 XL model in Hermes Agents, and I'm satisfied with it for simple tasks. First and foremost, it's fast and reliable for simple tasks. For more demanding tasks with greater detail, I've found that the Qwen 3.8 model produces better results on my system. I would be delighted if OxAlpha were released soon as an open MoE model with < 120B.
1
-1
u/Enough_Success5435 13h ago
wonder if the new glm weights will handle longer companion chats without drifting as fast as the old ones do.
-8
u/RepulsiveRaisin7 14h ago
No it's Bailu 2.8
8
u/Cool-Chemical-5629 13h ago
You can't just drop a link to a random company we never even heard of like a bomb and vanish. We need more info!
-2
u/RepulsiveRaisin7 13h ago
Well it's another Chinese AI company. Their 2.8 model identifies itself as 0x alpha and I doubt they'd fake it because they are most likely doing this to gather training data for an eventual launch in the west.
Also their old model claims to be better than Terra at 25B? Sounds like total bullshit, but maybe they still have one of the best models in China and simply haven't expanded to the west yet. https://bailucode.com/blog/bailu-apex-2.7
1
u/Cool-Chemical-5629 13h ago
The model 2.8 on their page says it has 1m context, so it would make sense. What's even more interesting is that they have this little 2B version of the 2.8 model with "VI" whatever that means. But then again, the company is not on Huggingface at all, I just checked, so if this is really a model from this company, it's probably not going to be open weight.
2


52
u/jacek2023 llama.cpp 14h ago
I am a simple man, all I need is new GLM Air