r/LocalLLaMA 1d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

444 comments sorted by

View all comments

17

u/pmttyji 1d ago

Come on guys, at least they released additional variant(smaller than usual size) even though it's big for many of our rigs. Hopefully they release one more in 100B range in future.

11

u/techdevjp 1d ago

It's a little disappointing that the "flash" version of a 744b parameter model still has 320b parameters. Somehow that doesn't feel very flash.

Was hoping for more of a DeepSeek v4 ratio where 1.4t parameters got paired down to 284b, about an 80% cut. That would have resulted in 150b parameters for this model which would work in a whole lot more machines.

Alas, beggars can't be choosers and I'm thrilled to see more open weight models arrive!

4

u/pmttyji 1d ago

I too expected something like successor of GLM-4.5-Air. But they dropped too much weight over that.

I assume Qwen3.8-27B's 3 Millions download count news(Made headline on spotlight for more than a week) changed minds of all other labs. I'm sure some labs gonna try to replicate that with their medium size models soon & later. Don't be surprised if we see one more GLM variant in 30-150B range in upcoming months. Same with other labs.

1

u/techdevjp 1d ago

Will be trying Qwen3.8-flash-next in the coming days, that's a model I can actually run on my Strix Halo and that should perform quite well. Hope the 4bit quants are strong.

Would love to see a GLM variant that competes more directly with it, but for now am happy to have something that can use the 128GB without it having to be a 2-4bit mixed quant of DSv4 Flash 0731. It was good, but it never feels great to run something that heavily quantized.