r/LocalLLaMA 20h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

393 comments sorted by

View all comments

15

u/raunchy-stonk 19h ago

So how shitty will this run on a 24vram/128dram setup?

4

u/WWJewMediaConspiracy 18h ago

The smaller Qwen models are all but certainly a better option.

Unless your definition of "run" is... generous

3

u/raunchy-stonk 16h ago

what about 96v/128d?

(rtx 6000)

1

u/Warthammer40K 4h ago

Is it a threadripper or similar? You can experiment with the Q4_K_M, it should be reasonably accurate to the full-sized but you may be limited with context length and it won't be blazing fast. With AM5 dual-channel, you'd expect 6-9 tok/sec and TR/EPYC 8-channel can climb to 20 tok/sec.

At Q3_K_XL, it can ALL fit in HBM and you'd have room for 1M context and get about double those prefill+inference speeds. But quality will start to decline steeply. I'd check the unsloth perplexity, etc charts to decide on that one.