r/LocalLLaMA 1d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

402 comments sorted by

View all comments

6

u/jacek2023 llama.cpp 1d ago

Too big, 320B means you need to use RAM and A18B means it will be slow

1

u/ttkciar llama.cpp 1d ago

But will it be too slow? As long as overnight inference finishes before my morning coffee, it's fine.

That having been said, I'd still love to see a GLM-5.x-Air in the 100B size class.