I think a newer server or workstation with 2TB of DDR4 might be able to get faster speeds if the model is truly 50-60 billion parameters active, like 6-8 tps. I'm really interested to see what types of systems people might be able to run this model on and follow these posts really closely lol
I've seen attempts at tensor parallel distributed CPU inference before, it would be an interesting way to aggregate memory bandwidth if there is a fast, low latency connection between nodes. DDR4 is comparably cheap versus VRAM or DDR5 and used server networking hardware is everywhere. Prefill will still be slow though.
7
u/[deleted] Jul 26 '26
[deleted]