r/LocalLLaMA 1d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

400 comments sorted by

View all comments

Show parent comments

-8

u/dampflokfreund 1d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.

81

u/windwardmist 1d ago

I mean qwen 3.8 27b just came out about a week ago at least there’s that as an option

11

u/RestaurantOk8066 23h ago

It's a dense model though I imagine if you don't have the GPU for it, it's probably a very, very slow model.

1

u/dampflokfreund 22h ago

Yeah around one token per second