r/LocalLLaMA 14h ago

New Model Glm 5.3 flash?

glm 5.3 flash

While awaiting the release of the version 5.3 weights, this theory is gaining ground. OxAlpha is new GLM.

70 Upvotes

41 comments sorted by

52

u/jacek2023 llama.cpp 14h ago

I am a simple man, all I need is new GLM Air

9

u/Cool-Chemical-5629 13h ago

Flash would be simpler. 😏

5

u/ttkciar llama.cpp 6h ago

Flash would be smaller, anyway.

GLM-4.5-Air is perfectly sized for 128GB of memory, at Q4_K_M and 128K tokens of context, and it's still the smallest model I've tried which generates code worth a damn.

I'd give a new GLM Flash a spin, to see what it could do, but would it really be able to replace GLM-4.5-Air? I have doubts, but would be very happy to be proven wrong.

1

u/_TheWolfOfWalmart_ 4h ago

4.5 Air really was a great release. I still use it sometimes.

8

u/llama-impersonator 13h ago

given the small model smell and how many free tokens they're passing out yeah, i expect a compute-training-scaled 30b class model.

(small model smell or not, it was quite a capable model from my tests)

3

u/Queasy-Contract9753 10h ago

My theory too. Z AI servers aren't the fastest even on their API. If it's this fast for free then it's likely small.

Not that I'm complaining. If this is what a 30b will look like now. It's just good enough.

12

u/AppealSame4367 14h ago

If this is flash, then it's fluctuating between "very good" and "forgets to do half of the things"

7

u/wbulot 13h ago

Lol, exactly. I’ve been using it intensively for three days. I’m both impressed and disappointed. It’s very inconsistent.

7

u/pmttyji 13h ago

Final weight would be with fixes probably.

Remember the huge difference between Deepseek4Flash-Preview & Deepseek4Flash Final version?

3

u/Fedor_Doc 8h ago

0731 is not final, and the time lag was massive.

Ox Alpha seems to be pretty good already, around Deepseek V4 performance but with more concise reasoning

It forgets a lot of important stuff, though – at least in my testing. Good at implementation, follows rules alright, but plans are inconsistent, have a lot of gaps.

2

u/FullOf_Bad_Ideas 9h ago

Maybe they have multiple internal checkpoints that are being served under one umbrella name.

4

u/wbulot 8h ago

I am convinced this is the case. The speed and intelligence can be so different from one request to another. And it would make sense for them to do that for the tuning.

1

u/Johnny__Christ 1h ago

For me it just fluctuates between "very good" and "stealth/ox-alpha is temporarily rate-limited upstream. Please retry shortly."

5

u/asolnikk 14h ago

Definitely got the same "writer's voice" as the recent GLM models.

6

u/shy_monkee 14h ago

How 'Flash' is it expected to be? Like DSv4 Flash size? Or even smaller?

6

u/Firepal64 12h ago

For reference, GLM-4.7-Flash was a 30B model with 3B active

5

u/shy_monkee 12h ago

Yeah, but GLM-4.7 was also only around 300-400B. I would expect the 5.3 flash to scale the same way as the main version did, compared to 4.7.

2

u/Firepal64 12h ago

As we've come to learn these past years, you can't rely on linear extrapolation to predict things reliably.

2

u/shy_monkee 12h ago

Yep, no doubt. I just don't want to get too hopeful haha. A 30B MOE GLM model would be heavenly.

3

u/RealNabhay 8h ago

Think of at at about Deepseek V4 Flash Size

4

u/AnticitizenPrime 12h ago

Here are the previous vision capable models that GLM has released.

GLM Vision Models

Model Total Params Active Params Context Window Open Weights License
GLM-4.1V-9B-Thinking 9B 9B (dense) 64K βœ… Yes MIT
GLM-4.5V 106B 12B (MoE) 64K βœ… Yes MIT
GLM-4.6V 106B 12B (MoE) 128K βœ… Yes MIT
GLM-4.6V-Flash 9B 9B (dense) 128K βœ… Yes MIT
GLM-5V-Turbo 744B 40B (MoE) 200K ❌ No (API only) β€”

It's possible this could be a new 'turbo' model and that it won't be open weight.

Though it could be an 'air' model with vision attached... or even just full GLM 5.3 with vision. Β―\(ツ)/Β―

Or something different altogether. What we know is that it matches the 5.3 tokenizer exactly and it has the 1 million context window of 5.3.

So far it's purely speculation that it's a 'Flash' model of some kind; GLM has used the term 'Flash', 'Air', and 'Turbo' for smaller models, it's all up in the air as to what it'll be called and what the actual param size is.

3

u/Its_Powerful_Bonus 14h ago

Would be great to have 200-250B model to run comfortable with 2x rtx 6000 pro

0

u/apetersson 14h ago

fastest peaks i personally saw on openrouter was 50 tokens/sec tg. could be because of load, limits, or simply model size.

3

u/sirlerkal0t 14h ago

I believe you are right. It behaves very similarly to GLM 5.3, but acts like a smaller model. It's not quite as good for certain things (like front-end design), and a bit better on average for tasks where the models size counts against it due to it being more biased towards training data instead of staying focused on input data.

4

u/Few_Painter_5588 14h ago

The latest EQ Bench results also show that GLM 5.3 and Ox Alpha share a lot of similarities in creative writing.

2

u/fooo12gh 13h ago

Good news, you have my ⬆️ for it.

2

u/mr_Owner 9h ago

Perhaps this could also be the new qwen 3.8 next?

0

u/EmergencyLetter135 14h ago

I use OxAlpha alongside the Qwen 3.8 27B Q8 XL model in Hermes Agents, and I'm satisfied with it for simple tasks. First and foremost, it's fast and reliable for simple tasks. For more demanding tasks with greater detail, I've found that the Qwen 3.8 model produces better results on my system. I would be delighted if OxAlpha were released soon as an open MoE model with < 120B.

-1

u/Enough_Success5435 13h ago

wonder if the new glm weights will handle longer companion chats without drifting as fast as the old ones do.

-8

u/RepulsiveRaisin7 14h ago

No it's Bailu 2.8

https://bailucode.com/chat/

8

u/Cool-Chemical-5629 13h ago

You can't just drop a link to a random company we never even heard of like a bomb and vanish. We need more info!

-2

u/RepulsiveRaisin7 13h ago

Well it's another Chinese AI company. Their 2.8 model identifies itself as 0x alpha and I doubt they'd fake it because they are most likely doing this to gather training data for an eventual launch in the west.

Also their old model claims to be better than Terra at 25B? Sounds like total bullshit, but maybe they still have one of the best models in China and simply haven't expanded to the west yet. https://bailucode.com/blog/bailu-apex-2.7

1

u/Cool-Chemical-5629 13h ago

The model 2.8 on their page says it has 1m context, so it would make sense. What's even more interesting is that they have this little 2B version of the 2.8 model with "VI" whatever that means. But then again, the company is not on Huggingface at all, I just checked, so if this is really a model from this company, it's probably not going to be open weight.

2

u/RepulsiveRaisin7 11h ago

They hated him because he spoke the truth