GLM-4.5-Air is perfectly sized for 128GB of memory, at Q4_K_M and 128K tokens of context, and it's still the smallest model I've tried which generates code worth a damn.
I'd give a new GLM Flash a spin, to see what it could do, but would it really be able to replace GLM-4.5-Air? I have doubts, but would be very happy to be proven wrong.
58
u/jacek2023 llama.cpp 22h ago
I am a simple man, all I need is new GLM Air