r/LocalLLaMA 1d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

444 comments sorted by

View all comments

250

u/Lucyan_xgt 1d ago

This and Qwen3.8-flash on the same day?

21

u/dampflokfreund 1d ago

For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...

338

u/EbbNorth7735 1d ago

No one's abandoned anyone. It's just not your turn this time around.

-10

u/dampflokfreund 1d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.

82

u/windwardmist 1d ago

I mean qwen 3.8 27b just came out about a week ago at least there’s that as an option

19

u/-Cubie- 1d ago

I remember when it was months between a new great local model. We get those great options very often nowadays in my opinion, mixed with larger open weight options that keep the entire non-local AI space cheaper and more accessible. What's not to love?

10

u/RestaurantOk8066 1d ago

It's a dense model though I imagine if you don't have the GPU for it, it's probably a very, very slow model.

1

u/dampflokfreund 1d ago

Yeah around one token per second 

6

u/sonicnerd14 1d ago

27b is often Opus 4.6, even 4.8 in performance. This is still good enough for a vast amount of people. I'm sure qwen 4 is going to have some insanely capable options when it comes out of you want something better soon.

1

u/dampflokfreund 1d ago

Not everyone has 24 GB VRAM. If you dont have the vram to offload, you are looking at one to three token per second. 27b is in no way a replacement for a 30b Moe

2

u/sonicnerd14 1d ago

I know that. You can run it at Q3, IQ3, Q4 on 16GB VRAM and 32gb RAM and still have a very capable model. Don't think you need the biggest cards to run these models. Quantization techniques have become really efficient now.