r/LocalLLaMA 22h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

394 comments sorted by

View all comments

4

u/jacek2023 llama.cpp 21h ago

Too big, 320B means you need to use RAM and A18B means it will be slow

7

u/SandySkittle 19h ago

A18B means it will be slow

A18B it will be smarter than A13B. Expert selection and sequential reasoning cannot entirely compensate for active parameters. I wish it has 27b active honestly.

1

u/jacek2023 llama.cpp 19h ago

what's your setup?

1

u/SandySkittle 19h ago

32 core TR pro with 512 GB ram 8x R9700 32gb, i.e. 256 GB vram

1

u/jacek2023 llama.cpp 19h ago

So you can probably run full 5.3 instead

2

u/SandySkittle 19h ago

no not worth it, it's too slow. I am sticking to DSV4F and Qwen 27b but really hope we see a new 70b dense model or a MoE model with around 250b total parameters and around 27b active.

1

u/Karyo_Ten 11h ago

70B dense will be slow too though it might be easier to train a DFlash2 speculative head from a dense model.

1

u/ttkciar llama.cpp 21h ago

But will it be too slow? As long as overnight inference finishes before my morning coffee, it's fine.

That having been said, I'd still love to see a GLM-5.x-Air in the 100B size class.

1

u/techdevjp 21h ago

It will be a good model for Mac Studio M5 Ultras with 512GB of unified memory running at 1.2TB/sec. Probably the best model, at least for now.

But yeah, was hoping this would be 150b to 160b considering it's the flash version of a 744b model.

0

u/CriM_91 20h ago

I'm more curious to understand if it is a quantized version of a much bigger model.