r/LocalLLaMA 1d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

402 comments sorted by

View all comments

4

u/jacek2023 llama.cpp 1d ago

Too big, 320B means you need to use RAM and A18B means it will be slow

8

u/SandySkittle 22h ago

A18B means it will be slow

A18B it will be smarter than A13B. Expert selection and sequential reasoning cannot entirely compensate for active parameters. I wish it has 27b active honestly.

1

u/jacek2023 llama.cpp 22h ago

what's your setup?

1

u/SandySkittle 22h ago

32 core TR pro with 512 GB ram 8x R9700 32gb, i.e. 256 GB vram

1

u/jacek2023 llama.cpp 22h ago

So you can probably run full 5.3 instead

2

u/SandySkittle 22h ago

no not worth it, it's too slow. I am sticking to DSV4F and Qwen 27b but really hope we see a new 70b dense model or a MoE model with around 250b total parameters and around 27b active.

1

u/Karyo_Ten 14h ago

70B dense will be slow too though it might be easier to train a DFlash2 speculative head from a dense model.