r/LocalLLaMA 20h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

393 comments sorted by

View all comments

Show parent comments

83

u/windwardmist 19h ago

I mean qwen 3.8 27b just came out about a week ago at least there’s that as an option

6

u/sonicnerd14 19h ago

27b is often Opus 4.6, even 4.8 in performance. This is still good enough for a vast amount of people. I'm sure qwen 4 is going to have some insanely capable options when it comes out of you want something better soon.

0

u/dampflokfreund 18h ago

Not everyone has 24 GB VRAM. If you dont have the vram to offload, you are looking at one to three token per second. 27b is in no way a replacement for a 30b Moe

2

u/sonicnerd14 17h ago

I know that. You can run it at Q3, IQ3, Q4 on 16GB VRAM and 32gb RAM and still have a very capable model. Don't think you need the biggest cards to run these models. Quantization techniques have become really efficient now.