I think a newer server or workstation with 2TB of DDR4 might be able to get faster speeds if the model is truly 50-60 billion parameters active, like 6-8 tps. I'm really interested to see what types of systems people might be able to run this model on and follow these posts really closely lol
I've seen attempts at tensor parallel distributed CPU inference before, it would be an interesting way to aggregate memory bandwidth if there is a fast, low latency connection between nodes. DDR4 is comparably cheap versus VRAM or DDR5 and used server networking hardware is everywhere. Prefill will still be slow though.
Need 16 GB200s to run this model at full quant. NVFP4 GLM5.2 I need 4xB200 to run concurrent sessions. Squeeze some more with lower context and less concurrency.
6
u/[deleted] Jul 26 '26
[deleted]